A toy language model in blocks
Build a complete next-word generator, inspect its counts, and vary sampling without confusing word counts with understanding.
Begin with text you can inspect completely
In Level 1 you read a next-word table that someone had made up. In this lesson you build a real one: a program counts which word followed which in a piece of text, then uses those counts to write new text. It is a language model small enough to understand completely.
The text a model learns from is called a corpus. Ours is short enough to count by hand:
the cat sleeps the cat plays the dog sleeps the dog runs the cat sleepsIt has fifteen words. Look at every word that comes straight after "the": cat three times and dog twice. So of the five times something followed "the", cat was 3 of 5 (60%) and dog was 2 of 5 (40%). These counts describe this text only. They do not mean cats are more common than dogs, or that a sentence about a cat is more likely to be true.
Build with me · 1
Choose how much context to remember
A context is the preceding text used to predict what comes next. With context one, this toy looks at the last word. It has not learned any followers merely because you created the table.
The name patterns is a programming name for that table. The number one is a modelling choice. Later you can compare it with two-word context while keeping the source text fixed.
From Words, add make word patterns called patterns with context 1. Add a status message underneath.
What to look for
The table is created and the status prints, with no learned sentence yet.
Make it yours
Explain why increasing context does not itself supply more training text.
Build with me · 2
Count the followers in a tiny corpus
The learns block reads the corpus word by word. Every word except the last is followed by another word, and the table counts each of those pairs.
The next-word table shows the counts you worked out above: after the, cat three times and dog twice.
Connect patterns learns the patterns in and type the supplied corpus in its text socket. Add show the next-word table.
What to look for
The table for the contains cat with count three and dog with count two.
Make it yours
Count the followers of cat yourself, then compare with the table.
Greedy generation first
The simplest way to write with the table is greedy generation: always take the word that followed most often, like the greedy choice in Level 1. The next four stages build it one piece at a time.
Build with me · 3
Ask for the most frequent follower
Greedy selection takes the most frequent follower. Cat has three of five observations after the, so the guess is cat and the observed share is 60 out of 100. A tie is broken alphabetically in this implementation.
The share is a property of these counts. It is not a confidence that a claim about a cat is true. There is no fact-checking step in this procedure.
Put patterns guesses the word after the inside labelled output. Add the corresponding word-chance value underneath.
What to look for
After the is cat, and Observed share is 60.
Make it yours
Try context dog. Count which follower wins or ties before running.
Build with me · 4
Store the growing text
The seed is the starting text you give the generator. Store it as line because that variable will grow later. The guessing block reads line but does not modify it. We store the returned word separately so the intermediate value remains visible.
This is the familiar pattern from classifiers: supply an input, receive a model output, then decide what the surrounding program should do with it.
Set line to the, then set word to a guess reading the line variable. Print both names before adding the joining step.
What to look for
Seed remains the; Next word is cat.
Make it yours
Change the seed to cat and explain why line still does not grow until you explicitly append the returned word.
Build with me · 5
Join the word with a separating space
The inner join combines one space with word. The outer join combines the old line with that spaced word. Assigning the result back to line updates the text used for the next prediction.
Without the space, the result would be thecat, which is a different word and likely an unseen context. Without assignment back to line, a temporary joined value would be calculated but the stored seed would remain unchanged.
Nest join space and word inside join line and that result. Put the whole expression inside set line. The space field must contain one actual space.
What to look for
The stored text becomes the cat.
Make it yours
Remove the space, observe the concatenation, and Undo. Explain how a tiny formatting change can also change which words the table looks up later.
Build with me · 6
Repeat next-word prediction with updated context
The loop repeats guess, then append, eight times. It reads the updated line on every turn. With one-word context, the model uses only its last word even though line stores the complete generated text.
The seed contains one word and the loop adds eight, so the result has nine words when every context has a follower. Greedy selection always gives the same result for this corpus and seed, so running it again does not produce new wording.
Place the word assignment and line update inside repeat 8. Keep the initial seed above and the final output below the loop.
What to look for
The supplied greedy result is the cat sleeps the cat sleeps the cat sleeps.
Make it yours
Print line inside the loop temporarily to watch it grow. Keep the final outside print too so you can compare intermediate and final state.
The greedy line is stuck in a loop. After "the", cat always wins. After "cat", sleeps always wins. After "sleeps", the always wins, and the cycle starts again.
Sample instead of always taking the winner
Sampling draws the next word at random, but not evenly: a word that followed more often is more likely to be drawn. Think of the bag of tickets from Level 1. After "the", the bag holds three cat tickets and two dog tickets.
Build with me · 7
Draw each word from the counts
A sampled next word is drawn at random, weighted by the counts. At temperature one and top-p one, the weights are the observed counts, so cat is favoured but dog can still follow the.
Randomness does not guarantee different output on every run. A context with only one follower still has one candidate. Compare several generated lines, not a single lucky or repetitive result.
Replace the greedy value with random word after line from patterns. Set temperature and top-p to 1, leaving the loop structure unchanged.
What to look for
A nine-word line follows observed transitions. The actual sequence can vary between runs.
Make it yours
Run several times and count whether cat or dog follows the first the. A few trials need not match exactly a 3-to-2 ratio.
Temperature: sharpen or flatten the odds
Temperature changes how strongly the common words are favoured. This program does it by changing the counts before the draw: each count is raised to the power 1 ÷ temperature. At temperature 0.5 the power is 2, which means squaring each count. At temperature 2 the power is one half, which means taking each count's square root. Here is what that does to the two candidates after "the":
| Temperature | What happens to the counts 3 and 2 | Chance of cat | Chance of dog |
|---|---|---|---|
| 1 | They stay 3 and 2 | 3 of 5, 60% | 2 of 5, 40% |
| 0.5, low | Each is squared: 3 × 3 = 9 and 2 × 2 = 4 | 9 of 13, about 69% | 4 of 13, about 31% |
| 2, high | Each is square-rooted: about 1.73 and 1.41 | about 55% | about 45% |
A low temperature makes the favourite even more likely, so the output repeats more. A high temperature evens the chances out, so less common words appear more often. Temperature changes how the next word is chosen. It never changes the counts that were learned, and it does not make the output more or less true.
Build with me · 8
Generate at a low and a high temperature
This stage runs the generator at temperature 0.5, which, as the table above shows, gives cat about 69% of the draws after the instead of 60%. It changes the draw, not the learned counts.
Keep top-p at one while comparing temperatures, so you change only one setting at a time and can say which setting caused a difference.
Use temperature 0.5 first, then 2, keeping the same corpus, seed, top-p, and loop count.
What to look for
Lower temperature tends to favour frequent followers more strongly; any individual random run can still repeat or differ.
Make it yours
Record several lines at each setting. Explain a tendency rather than promising that a specific next word must appear.
Top-p: keep only a shortlist
Top-p works differently. It lines the candidates up from most to least common, then adds them to a shortlist one at a time until their shares add up to at least the top-p value. Only words on the shortlist can be drawn.
After "the", cat alone has a share of 0.6. With top-p 0.5, cat's 0.6 already reaches 0.5, so the shortlist is just cat, and dog can never follow "the". With top-p 1, the shares must add up to 1, so both cat and dog stay on the shortlist.
Build with me · 9
Generate with top-p 0.5
This stage sets top-p to 0.5. As worked out above, cat alone already reaches that share after the, so dog is left off the shortlist there.
This is a different operation from temperature: top-p decides which candidates are allowed at all, while temperature changes the chances of the ones that are allowed.
Keep temperature at 1 and change only top-p to 0.5. Inspect the table after the and explain which candidate reaches the threshold.
What to look for
After the, the shortlist contains cat alone. The generated path becomes more constrained.
Make it yours
Compare top-p 0.5 and 1 while keeping temperature fixed. Say which candidates became available, not merely that the model became creative.
Many real language models have the same two settings, though each system applies them in its own way.
Longer contexts
So far the table has looked only at the last word. The next two stages make it look at the last two words instead, and deal with what happens when a pair of words never appeared in the corpus.
Build with me · 10
Supply enough words for a longer context
Context two distinguishes pairs such as the cat and the dog. The seed must contain enough words, so this version starts with the cat. Pass the whole growing line to prediction; the implementation takes its final two words.
Longer context can separate useful situations but makes observations sparser. A short corpus may contain a pair once or not at all, causing copying or a missing continuation.
Change context to 2 in the make-word-patterns block and the seed to the cat. Rebuild and rerun the whole program so the learned table matches the new context setting.
What to look for
Generation uses two-word contexts. Some choices may be highly repetitive or run out of observed continuation.
Make it yours
Try a two-word seed absent from the corpus. Observe the empty continuation before building a guard for it.
Build with me · 11
Stop cleanly when no follower exists
A keep-going-while block repeats its body for as long as its condition is true, rather than a fixed number of times. This loop counts down the remaining additions. If the model returns empty text, the first branch prints a clear explanation and sets remaining to zero, ending the while loop. Otherwise, it appends the word and reduces the remaining count by one.
Both branches make progress toward stopping. This is essential for a while loop: its condition can stay true forever if nothing changes the value it tests.
Use remaining = 8 before a keep-going-while block. Put the empty-word check inside, with a decrement in the successful branch and a zero assignment in the dead-end branch.
What to look for
The generator produces at most eight additions and exits cleanly if it encounters an unknown context.
Make it yours
Use a deliberately absent two-word seed. Confirm that it reports a missing continuation once rather than appending spaces repeatedly.
An empty result means the table never saw that context. It tells you something about the data, not that the computer forgot English.
Finally, move the corpus into a file, so you can replace it with any text you like.
Build with me · 12
Read editable training text from Files
The final version reads story.txt from the article's Files tab. The file contains text, while patterns is the table built from it while the program runs. Editing the file requires another Run to rebuild the table.
Use your own short text or material you are allowed to use. Compare contexts one, two, and three with suitable seeds, changing one choice at a time. A longer file adds evidence; a longer context needs that evidence to cover more exact word sequences.
Inspect story.txt in Files, then place the file-text value in the learns-text socket. Build the same guarded generation loop underneath.
What to look for
The supplied file reproduces the teaching corpus. Your edited corpus produces its own measured counts and continuations.
Make it yours
Write a short fictional story with repeated phrases. Test a known seed and an absent seed, then explain both outcomes using the table.
When you compare settings, change one thing at a time and keep the corpus fixed. Then you can say whether a difference came from the data, the context length, or the sampling settings, instead of calling every change "more creative".
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import jsonimport random def _words_in(value): letters = [] for character in str(value).lower(): letters.append(character if character.isalnum() else " ") return [word for word in "".join(letters).split() if word] def read_text(path): """Everything in a text file, as one long string.""" with open(path) as handle: return handle.read() class WordPatterns: """Counts which word tends to follow which. That is the whole trick behind next word prediction, just with far fewer words than a real model. The context is how many words back it looks. With a context of 1 it asks "what usually comes after 'the'". With a context of 2 it asks "what usually comes after 'the white'", which is a better question and needs more text to answer.""" def __init__(self, name="patterns", context=1): self.name = name self.context = max(1, int(context)) # Keyed by a tuple of the last few words, so ("the",) and # ("the", "white") are different questions with different answers. self.follows = {} def _context_of(self, words): """The last few words, as the key this table is built on.""" if len(words) < self.context: return None return tuple(words[-self.context:]) def learn(self, text): words = _words_in(text) for i in range(len(words) - self.context): key = tuple(words[i:i + self.context]) bucket = self.follows.setdefault(key, {}) following = words[i + self.context] bucket[following] = bucket.get(following, 0) + 1 print("Learned patterns from " + str(len(words)) + " words.") def next_word(self, text): """The word that followed most often. On a tie, the earlier one in the alphabet, so the same text always gives the same answer.""" bucket = self.follows.get(self._context_of(_words_in(text))) if not bucket: return "" return min(bucket, key=lambda word: (-bucket[word], word)) def chance(self, text): """Out of 100, how often that guess was the word that came next.""" bucket = self.follows.get(self._context_of(_words_in(text))) if not bucket: return 0.0 total = sum(bucket.values()) return round(100.0 * max(bucket.values()) / float(total), 1) def sample(self, text, temperature=1.0, top_p=1.0): """Picks one of the words that followed, at random, favouring the common ones. This is what a real model does instead of always taking the most likely word, and it is why a chatbot does not say the same sentence twice. top_p keeps only the shortlist of words that together account for that share of the counts. temperature flattens or sharpens the odds: below 1 the common words get even more likely, above 1 the rare ones get a look in.""" bucket = self.follows.get(self._context_of(_words_in(text))) if not bucket: return "" # Commonest first, ties broken by the alphabet so a run is repeatable. ranked = sorted(bucket.items(), key=lambda pair: (-pair[1], pair[0])) total = float(sum(count for _, count in ranked)) # The shortlist: the fewest words whose share reaches top_p. Always at # least one, or there would be nothing to choose from. share = 0.0 shortlist = [] for word, count in ranked: shortlist.append((word, count)) share += count / total if share >= top_p: break # Raising each count to 1 / temperature is what sharpens or flattens # them. The floor stops a temperature of 0 dividing by zero. power = 1.0 / max(float(temperature), 0.01) weights = [float(count) ** power for _, count in shortlist] pick = random.random() * sum(weights) running = 0.0 for (word, _), weight in zip(shortlist, weights): running += weight if pick <= running: return word return shortlist[-1][0] def show(self, limit=8): print(self.name + " learned " + str(len(self.follows)) + " starting words.") for key in list(self.follows)[:limit]: bucket = self.follows[key] best = min(bucket, key=lambda other: (-bucket[other], other)) print(" after '" + " ".join(key) + "' comes '" + best + "'") def table(self, limit=40): """Opens the next-word table panel. There is nothing useful to print in a terminal here, so this emits one marked line that the workspace reads and draws as a table. The runner takes the line back out of the output, so you never see it.""" rows = [] common = sorted( self.follows.items(), key=lambda pair: (-sum(pair[1].values()), pair[0]), ) for key, bucket in common[:limit]: total = sum(bucket.values()) followers = sorted( bucket.items(), key=lambda pair: (-pair[1], pair[0]) ) rows.append( { "context": " ".join(key), "count": total, "followers": [ { "word": word, "count": count, "share": round(count / float(total), 4), } for word, count in followers ], } ) print("__TABLE__ " + json.dumps({"name": self.name, "contexts": rows})) def new_word_patterns(name="patterns", context=1): return WordPatterns(name, context) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- patterns = new_word_patterns("patterns", 2)patterns.learn(read_text("story.txt"))line = "the cat"remaining = 8while (remaining > 0): word = patterns.sample(line, 1, 1) if (word == ""): print("No continuation for this context") remaining = 0 else: line = (str(line) + str((str(" ") + str(word)))) remaining = remaining + -1print("Generated text", line)Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.