0%
BuildHow a chatbot works, by building oneabout 41 min, 12 steps

A toy language model in blocks

Build a complete next-word generator, inspect its counts, and vary sampling without confusing word counts with understanding.

Begin with text you can inspect completely

In Level 1 you read a next-word table that someone had made up. In this lesson you build a real one: a program counts which word followed which in a piece of text, then uses those counts to write new text. It is a language model small enough to understand completely.

The text a model learns from is called a corpus. Ours is short enough to count by hand:

Text
the cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps

It has fifteen words. Look at every word that comes straight after "the": cat three times and dog twice. So of the five times something followed "the", cat was 3 of 5 (60%) and dog was 2 of 5 (40%). These counts describe this text only. They do not mean cats are more common than dogs, or that a sentence about a cat is more likely to be true.

Build with me · 1

Choose how much context to remember

A context is the preceding text used to predict what comes next. With context one, this toy looks at the last word. It has not learned any followers merely because you created the table.

The name patterns is a programming name for that table. The number one is a modelling choice. Later you can compare it with two-word context while keeping the source text fixed.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
sayOne-word context table created

From Words, add make word patterns called patterns with context 1. Add a status message underneath.

What to look for

The table is created and the status prints, with no learned sentence yet.

Make it yours

Explain why increasing context does not itself supply more training text.

Build with me · 2

Count the followers in a tiny corpus

The learns block reads the corpus word by word. Every word except the last is followed by another word, and the table counts each of those pairs.

The next-word table shows the counts you worked out above: after the, cat three times and dog twice.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
show the next-word table forpatterns

Connect patterns learns the patterns in and type the supplied corpus in its text socket. Add show the next-word table.

What to look for

The table for the contains cat with count three and dog with count two.

Make it yours

Count the followers of cat yourself, then compare with the table.

Greedy generation first

The simplest way to write with the table is greedy generation: always take the word that followed most often, like the greedy choice in Level 1. The next four stages build it one piece at a time.

Build with me · 3

Ask for the most frequent follower

Greedy selection takes the most frequent follower. Cat has three of five observations after the, so the guess is cat and the observed share is 60 out of 100. A tie is broken alphabetically in this implementation.

The share is a property of these counts. It is not a confidence that a claim about a cat is true. There is no fact-checking step in this procedure.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
sayAfter thethen
patternsguesses the word afterthe
sayObserved sharethen
out of 100, how surepatternsis afterthe

Put patterns guesses the word after the inside labelled output. Add the corresponding word-chance value underneath.

What to look for

After the is cat, and Observed share is 60.

Make it yours

Try context dog. Count which follower wins or ties before running.

Build with me · 4

Store the growing text

The seed is the starting text you give the generator. Store it as line because that variable will grow later. The guessing block reads line but does not modify it. We store the returned word separately so the intermediate value remains visible.

This is the familiar pattern from classifiers: supply an input, receive a model output, then decide what the surrounding program should do with it.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
setwordto
patternsguesses the word after
line
saySeedthen
line
sayNext wordthen
word

Set line to the, then set word to a guess reading the line variable. Print both names before adding the joining step.

What to look for

Seed remains the; Next word is cat.

Make it yours

Change the seed to cat and explain why line still does not grow until you explicitly append the returned word.

Build with me · 5

Join the word with a separating space

The inner join combines one space with word. The outer join combines the old line with that spaced word. Assigning the result back to line updates the text used for the next prediction.

Without the space, the result would be thecat, which is a different word and likely an unseen context. Without assignment back to line, a temporary joined value would be calculated but the stored seed would remain unchanged.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
setwordto
patternsguesses the word after
line
setlineto
join
line
and
join and
word
sayText after one additionthen
line

Nest join space and word inside join line and that result. Put the whole expression inside set line. The space field must contain one actual space.

What to look for

The stored text becomes the cat.

Make it yours

Remove the space, observe the concatenation, and Undo. Explain how a tiny formatting change can also change which words the table looks up later.

Build with me · 6

Repeat next-word prediction with updated context

The loop repeats guess, then append, eight times. It reads the updated line on every turn. With one-word context, the model uses only its last word even though line stores the complete generated text.

The seed contains one word and the loop adds eight, so the result has nine words when every context has a follower. Greedy selection always gives the same result for this corpus and seed, so running it again does not produce new wording.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
repeat8times
setwordto
patternsguesses the word after
line
setlineto
join
line
and
join and
word
sayGenerated textthen
line

Place the word assignment and line update inside repeat 8. Keep the initial seed above and the final output below the loop.

What to look for

The supplied greedy result is the cat sleeps the cat sleeps the cat sleeps.

Make it yours

Print line inside the loop temporarily to watch it grow. Keep the final outside print too so you can compare intermediate and final state.

The greedy line is stuck in a loop. After "the", cat always wins. After "cat", sleeps always wins. After "sleeps", the always wins, and the cycle starts again.

Sample instead of always taking the winner

Sampling draws the next word at random, but not evenly: a word that followed more often is more likely to be drawn. Think of the bag of tickets from Level 1. After "the", the bag holds three cat tickets and two dog tickets.

Build with me · 7

Draw each word from the counts

A sampled next word is drawn at random, weighted by the counts. At temperature one and top-p one, the weights are the observed counts, so cat is favoured but dog can still follow the.

Randomness does not guarantee different output on every run. A context with only one follower still has one candidate. Compare several generated lines, not a single lucky or repetitive result.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
repeat8times
setwordto
a random word after
line
frompatternswith temperature1and top-p1
setlineto
join
line
and
join and
word
sayGenerated textthen
line

Replace the greedy value with random word after line from patterns. Set temperature and top-p to 1, leaving the loop structure unchanged.

What to look for

A nine-word line follows observed transitions. The actual sequence can vary between runs.

Make it yours

Run several times and count whether cat or dog follows the first the. A few trials need not match exactly a 3-to-2 ratio.

Temperature: sharpen or flatten the odds

Temperature changes how strongly the common words are favoured. This program does it by changing the counts before the draw: each count is raised to the power 1 ÷ temperature. At temperature 0.5 the power is 2, which means squaring each count. At temperature 2 the power is one half, which means taking each count's square root. Here is what that does to the two candidates after "the":

TemperatureWhat happens to the counts 3 and 2Chance of catChance of dog
1They stay 3 and 23 of 5, 60%2 of 5, 40%
0.5, lowEach is squared: 3 × 3 = 9 and 2 × 2 = 49 of 13, about 69%4 of 13, about 31%
2, highEach is square-rooted: about 1.73 and 1.41about 55%about 45%

A low temperature makes the favourite even more likely, so the output repeats more. A high temperature evens the chances out, so less common words appear more often. Temperature changes how the next word is chosen. It never changes the counts that were learned, and it does not make the output more or less true.

Build with me · 8

Generate at a low and a high temperature

This stage runs the generator at temperature 0.5, which, as the table above shows, gives cat about 69% of the draws after the instead of 60%. It changes the draw, not the learned counts.

Keep top-p at one while comparing temperatures, so you change only one setting at a time and can say which setting caused a difference.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
repeat8times
setwordto
a random word after
line
frompatternswith temperature0.5and top-p1
setlineto
join
line
and
join and
word
sayGenerated textthen
line

Use temperature 0.5 first, then 2, keeping the same corpus, seed, top-p, and loop count.

What to look for

Lower temperature tends to favour frequent followers more strongly; any individual random run can still repeat or differ.

Make it yours

Record several lines at each setting. Explain a tendency rather than promising that a specific next word must appear.

Top-p: keep only a shortlist

Top-p works differently. It lines the candidates up from most to least common, then adds them to a shortlist one at a time until their shares add up to at least the top-p value. Only words on the shortlist can be drawn.

After "the", cat alone has a share of 0.6. With top-p 0.5, cat's 0.6 already reaches 0.5, so the shortlist is just cat, and dog can never follow "the". With top-p 1, the shares must add up to 1, so both cat and dog stay on the shortlist.

Build with me · 9

Generate with top-p 0.5

This stage sets top-p to 0.5. As worked out above, cat alone already reaches that share after the, so dog is left off the shortlist there.

This is a different operation from temperature: top-p decides which candidates are allowed at all, while temperature changes the chances of the ones that are allowed.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
repeat8times
setwordto
a random word after
line
frompatternswith temperature1and top-p0.5
setlineto
join
line
and
join and
word
sayGenerated textthen
line

Keep temperature at 1 and change only top-p to 0.5. Inspect the table after the and explain which candidate reaches the threshold.

What to look for

After the, the shortlist contains cat alone. The generated path becomes more constrained.

Make it yours

Compare top-p 0.5 and 1 while keeping temperature fixed. Say which candidates became available, not merely that the model became creative.

Many real language models have the same two settings, though each system applies them in its own way.

Longer contexts

So far the table has looked only at the last word. The next two stages make it look at the last two words instead, and deal with what happens when a pair of words never appeared in the corpus.

Build with me · 10

Supply enough words for a longer context

Context two distinguishes pairs such as the cat and the dog. The seed must contain enough words, so this version starts with the cat. Pass the whole growing line to prediction; the implementation takes its final two words.

Longer context can separate useful situations but makes observations sparser. A short corpus may contain a pair once or not at all, causing copying or a missing continuation.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of2words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe cat
repeat8times
setwordto
a random word after
line
frompatternswith temperature1and top-p1
setlineto
join
line
and
join and
word
sayGenerated textthen
line

Change context to 2 in the make-word-patterns block and the seed to the cat. Rebuild and rerun the whole program so the learned table matches the new context setting.

What to look for

Generation uses two-word contexts. Some choices may be highly repetitive or run out of observed continuation.

Make it yours

Try a two-word seed absent from the corpus. Observe the empty continuation before building a guard for it.

Build with me · 11

Stop cleanly when no follower exists

A keep-going-while block repeats its body for as long as its condition is true, rather than a fixed number of times. This loop counts down the remaining additions. If the model returns empty text, the first branch prints a clear explanation and sets remaining to zero, ending the while loop. Otherwise, it appends the word and reduces the remaining count by one.

Both branches make progress toward stopping. This is essential for a while loop: its condition can stay true forever if nothing changes the value it tests.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of2words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe cat
setremainingto8
keep going while
remaining
>0
setwordto
a random word after
line
frompatternswith temperature1and top-p1
if
word
isempty
then
sayNo continuation for this context
setremainingto0
otherwise
setlineto
join
line
and
join and
word
changeremainingby-1
sayGenerated textthen
line

Use remaining = 8 before a keep-going-while block. Put the empty-word check inside, with a decrement in the successful branch and a zero assignment in the dead-end branch.

What to look for

The generator produces at most eight additions and exits cleanly if it encounters an unknown context.

Make it yours

Use a deliberately absent two-word seed. Confirm that it reports a missing continuation once rather than appending spaces repeatedly.

An empty result means the table never saw that context. It tells you something about the data, not that the computer forgot English.

Finally, move the corpus into a file, so you can replace it with any text you like.

Build with me · 12

Read editable training text from Files

The final version reads story.txt from the article's Files tab. The file contains text, while patterns is the table built from it while the program runs. Editing the file requires another Run to rebuild the table.

Use your own short text or material you are allowed to use. Compare contexts one, two, and three with suitable seeds, changing one choice at a time. A longer file adds evidence; a longer context needs that evidence to cover more exact word sequences.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of2words
patternslearns the patterns in
the text in filestory.txt
setlinetothe cat
setremainingto8
keep going while
remaining
>0
setwordto
a random word after
line
frompatternswith temperature1and top-p1
if
word
isempty
then
sayNo continuation for this context
setremainingto0
otherwise
setlineto
join
line
and
join and
word
changeremainingby-1
sayGenerated textthen
line

Inspect story.txt in Files, then place the file-text value in the learns-text socket. Build the same guarded generation loop underneath.

What to look for

The supplied file reproduces the teaching corpus. Your edited corpus produces its own measured counts and continuations.

Make it yours

Write a short fictional story with repeated phrases. Test a known seed and an absent seed, then explain both outcomes using the table.

When you compare settings, change one thing at a time and keep the corpus fixed. Then you can say whether a difference came from the data, the context length, or the sampling settings, instead of calling every change "more creative".

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make word patterns calledpatternswith a context of2words
patternslearns the patterns in
the text in filestory.txt
setlinetothe cat
setremainingto8
keep going while
remaining
>0
setwordto
a random word after
line
frompatternswith temperature1and top-p1
if
word
isempty
then
sayNo continuation for this context
setremainingto0
otherwise
setlineto
join
line
and
join and
word
changeremainingby-1
sayGenerated textthen
line
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import jsonimport random  def _words_in(value):    letters = []    for character in str(value).lower():        letters.append(character if character.isalnum() else " ")    return [word for word in "".join(letters).split() if word]  def read_text(path):    """Everything in a text file, as one long string."""    with open(path) as handle:        return handle.read()  class WordPatterns:    """Counts which word tends to follow which. That is the whole trick behind    next word prediction, just with far fewer words than a real model.     The context is how many words back it looks. With a context of 1 it asks    "what usually comes after 'the'". With a context of 2 it asks "what usually    comes after 'the white'", which is a better question and needs more text to    answer."""     def __init__(self, name="patterns", context=1):        self.name = name        self.context = max(1, int(context))        # Keyed by a tuple of the last few words, so ("the",) and        # ("the", "white") are different questions with different answers.        self.follows = {}     def _context_of(self, words):        """The last few words, as the key this table is built on."""        if len(words) < self.context:            return None        return tuple(words[-self.context:])     def learn(self, text):        words = _words_in(text)        for i in range(len(words) - self.context):            key = tuple(words[i:i + self.context])            bucket = self.follows.setdefault(key, {})            following = words[i + self.context]            bucket[following] = bucket.get(following, 0) + 1        print("Learned patterns from " + str(len(words)) + " words.")     def next_word(self, text):        """The word that followed most often. On a tie, the earlier one in the        alphabet, so the same text always gives the same answer."""        bucket = self.follows.get(self._context_of(_words_in(text)))        if not bucket:            return ""        return min(bucket, key=lambda word: (-bucket[word], word))     def chance(self, text):        """Out of 100, how often that guess was the word that came next."""        bucket = self.follows.get(self._context_of(_words_in(text)))        if not bucket:            return 0.0        total = sum(bucket.values())        return round(100.0 * max(bucket.values()) / float(total), 1)     def sample(self, text, temperature=1.0, top_p=1.0):        """Picks one of the words that followed, at random, favouring the        common ones. This is what a real model does instead of always taking        the most likely word, and it is why a chatbot does not say the same        sentence twice.         top_p keeps only the shortlist of words that together account for that        share of the counts. temperature flattens or sharpens the odds: below 1        the common words get even more likely, above 1 the rare ones get a        look in."""        bucket = self.follows.get(self._context_of(_words_in(text)))        if not bucket:            return ""         # Commonest first, ties broken by the alphabet so a run is repeatable.        ranked = sorted(bucket.items(), key=lambda pair: (-pair[1], pair[0]))        total = float(sum(count for _, count in ranked))         # The shortlist: the fewest words whose share reaches top_p. Always at        # least one, or there would be nothing to choose from.        share = 0.0        shortlist = []        for word, count in ranked:            shortlist.append((word, count))            share += count / total            if share >= top_p:                break         # Raising each count to 1 / temperature is what sharpens or flattens        # them. The floor stops a temperature of 0 dividing by zero.        power = 1.0 / max(float(temperature), 0.01)        weights = [float(count) ** power for _, count in shortlist]         pick = random.random() * sum(weights)        running = 0.0        for (word, _), weight in zip(shortlist, weights):            running += weight            if pick <= running:                return word        return shortlist[-1][0]     def show(self, limit=8):        print(self.name + " learned " + str(len(self.follows)) + " starting words.")        for key in list(self.follows)[:limit]:            bucket = self.follows[key]            best = min(bucket, key=lambda other: (-bucket[other], other))            print("  after '" + " ".join(key) + "' comes '" + best + "'")     def table(self, limit=40):        """Opens the next-word table panel.         There is nothing useful to print in a terminal here, so this emits one        marked line that the workspace reads and draws as a table. The runner        takes the line back out of the output, so you never see it."""        rows = []        common = sorted(            self.follows.items(),            key=lambda pair: (-sum(pair[1].values()), pair[0]),        )        for key, bucket in common[:limit]:            total = sum(bucket.values())            followers = sorted(                bucket.items(), key=lambda pair: (-pair[1], pair[0])            )            rows.append(                {                    "context": " ".join(key),                    "count": total,                    "followers": [                        {                            "word": word,                            "count": count,                            "share": round(count / float(total), 4),                        }                        for word, count in followers                    ],                }            )        print("__TABLE__ " + json.dumps({"name": self.name, "contexts": rows}))  def new_word_patterns(name="patterns", context=1):    return WordPatterns(name, context)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  patterns = new_word_patterns("patterns", 2)patterns.learn(read_text("story.txt"))line = "the cat"remaining = 8while (remaining > 0):    word = patterns.sample(line, 1, 1)    if (word == ""):        print("No continuation for this context")        remaining = 0    else:        line = (str(line) + str((str(" ") + str(word))))        remaining = remaining + -1print("Generated text", line)

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in