Posts about natural language generation


Talk transcript: Fine-tuning GPT-2 on World of Warcraft quests


Tags: natural language generation, video games, projects, research

This is a textual version of the talk I gave at Foundations of Digital Games 2021, for the research paper titled Fine-tuning GPT-2 on annotated RPG quests for NPC dialogue generation. The talk was transcribed using Whisper.cpp and co-edited with the open beta of ChatGPT to improve the flow of the text.

Good morning everyone. My name is Judith van Stegeren and I'm here to present a fun project that I worked on together with my co-author Jakub.

Read more...

A comparison of GPT-2 and BERT


Tags: natural language generation, research

GPT-2 and BERT are two methods for creating language models, based on neural networks and deep learning. GPT-2 and BERT are fairly young, but they are 'state-of-the-art', which means they beat almost every other method in the natural language processing field.

GPT-2 and BERT are extra useable because they come with a set of pre-trained language models, which anyone can download and use. Pre-trained models have as main advantage that user don't have to train a language model from scratch, which is computationally expensive and requires a huge dataset. Instead, you can take a smaller dataset and "fine-tune" the large, pre-trained model to your specific dataset with a bit of additional training, which is much cheaper.

Read more...

Interview: NaNoGenMo and coherence


Tags: natural language generation, research

A few months ago I was interviewed by Wired about NaNoGenMo 2019, an online programming challenge where participants try to build a novel-generator in 30 days. The interview was part of the background research for this Wired article.

I figured some people might be interested in this interview too, so I decided to put the questions and their answers up on my blog. I've rewritten some answers of my answers to clarify them -- the reporter who interviewed me had read my NaNoGenMo paper, but perhaps you haven't. :)

Read more...


Choose the right bias for your text generator


Tags: natural language generation

A suitable corpus, or dataset of text documents, is one of the main ingredients of many natural language generation projects. There are many general-purpose corpora available for download on the internet. For example, the natural language processing library nltk has a built-in download module, through which you can access various standard datasets. Anyone can download these datasets for free, which makes them a great NLP resource for programmers and language researchers.

Read more...


A quick overview of my PhD research project


Tags: natural language generation, research, projects

I'm working as a PhD student at the University of Twente, where I research natural language generation for adapative games. People often ask me what my research is about, so I figured I should write a blogpost to explain a bit more about my research field.

A natural language is a language that is used by humans, such as English, Dutch or Japanese, as opposed to the formal languages of mathematics and logics. Note that I use this very informal definition, since I'm a computer scientist and not a linguist. If you ask a linguist to define a natural language, you get probably many different answers, such as the definition from Wikipedia.

Read more...