From word counts to ChatGPT
Explain what a modern language model shares with your generator and where the simple count-table analogy stops.
Keep the useful part of the analogy
Your generator reads a context, scores the possible next words, picks one, adds it to the text, and repeats. A chatbot writes its answers in the same way: one piece at a time, adding each chosen piece to the context before making the next choice. Generating like this, where each output becomes part of the next input, is called autoregressive generation. That shared loop is the useful connection between your toy and ChatGPT.
Build with me · 1
Separate input context from chosen output
The context is the text the model uses to make its choice. The chosen word is one continuation produced from it. A modern language model makes a related next-unit prediction, normally over tokens that may be words, word pieces, punctuation, or other encoded pieces.
This toy splits words with its own rule. It does not reproduce a modern tokenizer, so do not use its word counts as exact token counts for another system.
Store the context separately from the returned word and print both values.
What to look for
Context is the and Chosen word is cat.
Make it yours
Explain the shared role of context and the difference between this toy's words and a model's tokens.
Because a real model works in tokens, it can stumble when asked to count the letters in a word: its pieces do not always line up with single letters. It can still handle spelling using patterns it learned and, in some applications, tools.
Build with me · 2
See how one output becomes later input
Each chosen word is appended before the next prediction. This is the autoregressive loop: the growing output becomes part of the context used for later choices.
The toy still sees only the final one word, because its context setting limits the lookup. A modern model can use thousands of earlier tokens at once.
Put a labelled line output inside the generation loop so each growing prefix remains visible.
What to look for
The prefixes progress through the cat, the cat sleeps, and the cat sleeps the.
Make it yours
Identify which words the toy actually uses on each turn. Explain why storing a long line does not automatically mean the toy uses every word.
Why a bigger table is not enough
With a one-word context, your corpus had several examples of what follows "the". With a long context, almost every exact sequence of words appears once or never, even in an enormous amount of text. A table of counts for every possible long context would be almost entirely empty.
Build with me · 3
Find a context the table never observed
The exact pair the cat has observed followers. Bright moon does not, so the lookup returns empty text. The longer the context, the more of them are missing, even when a corpus has many words.
A table can only answer for contexts it has seen. The paragraphs below explain what a chatbot does instead of keeping a table.
Create context-two patterns, then print predictions for one seen and one absent pair.
What to look for
Known pair produces a follower. Unseen pair has no word after its label.
Make it yours
Explain why simply making an exact-context table longer cannot guarantee coverage of every sentence someone might type.
Modern chatbots are neural language models. Instead of a table, they use a neural network: a long chain of calculations controlled by billions of adjustable numbers, the parameters, which training sets. Inside, each token is turned into a list of numbers, and tokens that are used in similar ways end up with similar lists. These learned lists are called representations. Because "cat" and "kitten" end up with similar numbers, what the model learned about one can help with the other. That lets it give sensible scores to a sequence it never saw exactly, which is generalisation, although it can still repeat text it memorised.
Most modern language models use a design called a transformer. Its key step, called attention, lets the calculation for each token look back over the other tokens in the context and weigh which ones matter most for the next choice. You do not need to build one to see the boundary: the network calculates its scores from learned numbers; it does not look anything up in a table of exact word sequences.
Training and a running conversation
A chatbot's first and largest stage of training is called pretraining. The model reads an enormous amount of text and keeps guessing the next token. Each guess is checked against the real next token, and the parameters are adjusted a little: the guess, check, adjust loop from Level 1, repeated an enormous number of times. That is how it picks up grammar, style, and many facts. Further training then teaches it to follow instructions and answer helpfully, using example conversations and feedback from people. Different systems combine these stages in different ways.
Build with me · 4
Do not mistake a corpus for a helpful assistant
A corpus containing requests does not by itself make this toy follow requests. It counts which word followed explain and has no follower for why. There is no representation of the user's goal, no learned instruction-following behaviour, and no tool call.
Real assistants get their helpfulness from the further training described above, not from simply reading more requests.
Run the tiny request-like corpus and inspect the two continuations. Compare the actual output with what a helpful explanation would require.
What to look for
After explain is gravity because the tie is alphabetical; after why has no continuation.
Make it yours
Describe three things a helpful explanation needs that this count-table program never implemented.
During an ordinary conversation, nothing is adjusted. The model uses its parameters exactly as they are, together with the context it is given. Typing a message does not retrain it. A product may store your earlier conversation, or some saved notes, and add them to the context later; that is the app doing the remembering, not the model being trained.
Build with me · 5
Show that prediction leaves training fixed
Sampling reads the stored counts. It does not add the generated words back into training. The table before and after these predictions is the same.
The same is true of a chatbot answering you: using the model does not change it.
Show the table before and after three samples, without another learns block between them.
What to look for
Both table displays have the same cat and dog counts after the, even if sampled words differ.
Make it yours
Explain how a conversation can use earlier text as context without changing learned model weights.
Build with me · 6
Make an explicit model update visible
This time a learns block explicitly adds observations. It contributes three more dog followers after the, taking dog from two to five while cat stays at three. The model's stored counts changed, so its preferred follower changes.
A neural network is trained by adjusting its parameters rather than by adding counts, but the distinction is the same: training changes the model; using it to predict does not. Do not claim that a normal prediction performs the training step shown here.
Place the extra learns statement after the first corpus and before the new table inspection.
What to look for
The table records five dogs and three cats after the, and the greedy choice is dog.
Make it yours
Compare this deliberate learning update with the unchanged before/after tables from the previous stage.
Where facts and tools fit
Parameters can hold factual associations, so it would be wrong to say a language model knows no facts. The limit is that producing likely-sounding text does not check a claim, name a source, or guarantee up-to-date information.
An application can add documents, search results, or the output of a tool to the context, and the model can use them in its answer. Whether that actually happened depends on the app and what it did, not on how confident the answer sounds. A citation must point to a real source that supports the claim; a citation-shaped string is not proof. This is why Level 1 taught you to check specific claims against their sources.
Sampling does not evaluate truth
Temperature and top-p change how the next token is chosen, though real systems may apply them differently from your toy. A lower temperature can make answers more repeatable, but it does not check dates, calculations, or quotations. A higher temperature can vary the wording, and it can also bring in less suitable continuations.
Build with me · 7
Vary wording while leaving evidence untouched
Temperature changes selection weights. The underlying count table remains the same, and no source lookup or truth evaluation appears anywhere in the stack. A more varied answer is not automatically more false, and a repeated answer is not automatically true.
Modern assistants can have additional training and tools, but the responsibility to check a factual claim against evidence is not replaced by a sampling slider.
Ask for one low-temperature and one high-temperature sample while keeping the corpus fixed. Inspect the unchanged table.
What to look for
Outputs may differ or coincide by chance. Both are selected from the same observed candidates.
Make it yours
Point to the block that would verify a factual claim. There is none in this program; explain why that matters when interpreting fluent output.
The last stage puts the shared idea and its limits side by side.
Build with me · 8
Keep the connection and its limits together
A strong explanation names both halves. Choosing the next piece from the context and repeating is shared with a chatbot. Learned representations, attention across the context, the scale of training, instruction-following, and tools go far beyond this table.
Run the final short generator, then explain its steps from the actual growing prefixes. Use the final messages to check that your analogy has an explicit boundary.
What to look for
Five prefixes show repeated generation. The closing messages separate the shared idea from missing mechanisms.
Make it yours
Explain temperature to a beginner using one observed prefix, then add one sentence saying why it is not a truthfulness control.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- import jsonimport random def _words_in(value): letters = [] for character in str(value).lower(): letters.append(character if character.isalnum() else " ") return [word for word in "".join(letters).split() if word] class WordPatterns: """Counts which word tends to follow which. That is the whole trick behind next word prediction, just with far fewer words than a real model. The context is how many words back it looks. With a context of 1 it asks "what usually comes after 'the'". With a context of 2 it asks "what usually comes after 'the white'", which is a better question and needs more text to answer.""" def __init__(self, name="patterns", context=1): self.name = name self.context = max(1, int(context)) # Keyed by a tuple of the last few words, so ("the",) and # ("the", "white") are different questions with different answers. self.follows = {} def _context_of(self, words): """The last few words, as the key this table is built on.""" if len(words) < self.context: return None return tuple(words[-self.context:]) def learn(self, text): words = _words_in(text) for i in range(len(words) - self.context): key = tuple(words[i:i + self.context]) bucket = self.follows.setdefault(key, {}) following = words[i + self.context] bucket[following] = bucket.get(following, 0) + 1 print("Learned patterns from " + str(len(words)) + " words.") def next_word(self, text): """The word that followed most often. On a tie, the earlier one in the alphabet, so the same text always gives the same answer.""" bucket = self.follows.get(self._context_of(_words_in(text))) if not bucket: return "" return min(bucket, key=lambda word: (-bucket[word], word)) def chance(self, text): """Out of 100, how often that guess was the word that came next.""" bucket = self.follows.get(self._context_of(_words_in(text))) if not bucket: return 0.0 total = sum(bucket.values()) return round(100.0 * max(bucket.values()) / float(total), 1) def sample(self, text, temperature=1.0, top_p=1.0): """Picks one of the words that followed, at random, favouring the common ones. This is what a real model does instead of always taking the most likely word, and it is why a chatbot does not say the same sentence twice. top_p keeps only the shortlist of words that together account for that share of the counts. temperature flattens or sharpens the odds: below 1 the common words get even more likely, above 1 the rare ones get a look in.""" bucket = self.follows.get(self._context_of(_words_in(text))) if not bucket: return "" # Commonest first, ties broken by the alphabet so a run is repeatable. ranked = sorted(bucket.items(), key=lambda pair: (-pair[1], pair[0])) total = float(sum(count for _, count in ranked)) # The shortlist: the fewest words whose share reaches top_p. Always at # least one, or there would be nothing to choose from. share = 0.0 shortlist = [] for word, count in ranked: shortlist.append((word, count)) share += count / total if share >= top_p: break # Raising each count to 1 / temperature is what sharpens or flattens # them. The floor stops a temperature of 0 dividing by zero. power = 1.0 / max(float(temperature), 0.01) weights = [float(count) ** power for _, count in shortlist] pick = random.random() * sum(weights) running = 0.0 for (word, _), weight in zip(shortlist, weights): running += weight if pick <= running: return word return shortlist[-1][0] def show(self, limit=8): print(self.name + " learned " + str(len(self.follows)) + " starting words.") for key in list(self.follows)[:limit]: bucket = self.follows[key] best = min(bucket, key=lambda other: (-bucket[other], other)) print(" after '" + " ".join(key) + "' comes '" + best + "'") def table(self, limit=40): """Opens the next-word table panel. There is nothing useful to print in a terminal here, so this emits one marked line that the workspace reads and draws as a table. The runner takes the line back out of the output, so you never see it.""" rows = [] common = sorted( self.follows.items(), key=lambda pair: (-sum(pair[1].values()), pair[0]), ) for key, bucket in common[:limit]: total = sum(bucket.values()) followers = sorted( bucket.items(), key=lambda pair: (-pair[1], pair[0]) ) rows.append( { "context": " ".join(key), "count": total, "followers": [ { "word": word, "count": count, "share": round(count / float(total), 4), } for word, count in followers ], } ) print("__TABLE__ " + json.dumps({"name": self.name, "contexts": rows})) def new_word_patterns(name="patterns", context=1): return WordPatterns(name, context) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- patterns = new_word_patterns("patterns", 1)patterns.learn("the cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps")line = "the"for _ in range(int(5)): word = patterns.sample(line, 1, 1) line = (str(line) + str((str(" ") + str(word)))) print("Generated prefix", line)print("Shared idea: choose a next unit using context and continue")print("Limits: this toy has no neural representations, instruction training, retrieval, or factual verification")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.