0%
BuildHow a chatbot works, by building oneabout 29 min, 8 steps

From word counts to ChatGPT

Explain what a modern language model shares with your generator and where the simple count-table analogy stops.

Keep the useful part of the analogy

Your generator reads a context, scores the possible next words, picks one, adds it to the text, and repeats. A chatbot writes its answers in the same way: one piece at a time, adding each chosen piece to the context before making the next choice. Generating like this, where each output becomes part of the next input, is called autoregressive generation. That shared loop is the useful connection between your toy and ChatGPT.

Build with me · 1

Separate input context from chosen output

The context is the text the model uses to make its choice. The chosen word is one continuation produced from it. A modern language model makes a related next-unit prediction, normally over tokens that may be words, word pieces, punctuation, or other encoded pieces.

This toy splits words with its own rule. It does not reproduce a modern tokenizer, so do not use its word counts as exact token counts for another system.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setcontexttothe
setchosento
patternsguesses the word after
context
sayContextthen
context
sayChosen wordthen
chosen

Store the context separately from the returned word and print both values.

What to look for

Context is the and Chosen word is cat.

Make it yours

Explain the shared role of context and the difference between this toy's words and a model's tokens.

Because a real model works in tokens, it can stumble when asked to count the letters in a word: its pieces do not always line up with single letters. It can still handle spelling using patterns it learned and, in some applications, tools.

Build with me · 2

See how one output becomes later input

Each chosen word is appended before the next prediction. This is the autoregressive loop: the growing output becomes part of the context used for later choices.

The toy still sees only the final one word, because its context setting limits the lookup. A modern model can use thousands of earlier tokens at once.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
repeat3times
setwordto
patternsguesses the word after
line
setlineto
join
line
and
join and
word
sayExtended contextthen
line

Put a labelled line output inside the generation loop so each growing prefix remains visible.

What to look for

The prefixes progress through the cat, the cat sleeps, and the cat sleeps the.

Make it yours

Identify which words the toy actually uses on each turn. Explain why storing a long line does not automatically mean the toy uses every word.

Why a bigger table is not enough

With a one-word context, your corpus had several examples of what follows "the". With a long context, almost every exact sequence of words appears once or never, even in an enormous amount of text. A table of counts for every possible long context would be almost entirely empty.

Build with me · 3

Find a context the table never observed

The exact pair the cat has observed followers. Bright moon does not, so the lookup returns empty text. The longer the context, the more of them are missing, even when a corpus has many words.

A table can only answer for contexts it has seen. The paragraphs below explain what a chatbot does instead of keeping a table.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of2words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
sayKnown pairthen
patternsguesses the word afterthe cat
sayUnseen pairthen
patternsguesses the word afterbright moon

Create context-two patterns, then print predictions for one seen and one absent pair.

What to look for

Known pair produces a follower. Unseen pair has no word after its label.

Make it yours

Explain why simply making an exact-context table longer cannot guarantee coverage of every sentence someone might type.

Modern chatbots are neural language models. Instead of a table, they use a neural network: a long chain of calculations controlled by billions of adjustable numbers, the parameters, which training sets. Inside, each token is turned into a list of numbers, and tokens that are used in similar ways end up with similar lists. These learned lists are called representations. Because "cat" and "kitten" end up with similar numbers, what the model learned about one can help with the other. That lets it give sensible scores to a sequence it never saw exactly, which is generalisation, although it can still repeat text it memorised.

Most modern language models use a design called a transformer. Its key step, called attention, lets the calculation for each token look back over the other tokens in the context and weigh which ones matter most for the next choice. You do not need to build one to see the boundary: the network calculates its scores from learned numbers; it does not look anything up in a table of exact word sequences.

Training and a running conversation

A chatbot's first and largest stage of training is called pretraining. The model reads an enormous amount of text and keeps guessing the next token. Each guess is checked against the real next token, and the parameters are adjusted a little: the guess, check, adjust loop from Level 1, repeated an enormous number of times. That is how it picks up grammar, style, and many facts. Further training then teaches it to follow instructions and answer helpfully, using example conversations and feedback from people. Different systems combine these stages in different ways.

Build with me · 4

Do not mistake a corpus for a helpful assistant

A corpus containing requests does not by itself make this toy follow requests. It counts which word followed explain and has no follower for why. There is no representation of the user's goal, no learned instruction-following behaviour, and no tool call.

Real assistants get their helpfulness from the further training described above, not from simply reading more requests.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inplease explain gravity please explain magnets
sayAfter explainthen
patternsguesses the word afterexplain
sayAfter whythen
patternsguesses the word afterwhy

Run the tiny request-like corpus and inspect the two continuations. Compare the actual output with what a helpful explanation would require.

What to look for

After explain is gravity because the tie is alphabetical; after why has no continuation.

Make it yours

Describe three things a helpful explanation needs that this count-table program never implemented.

During an ordinary conversation, nothing is adjusted. The model uses its parameters exactly as they are, together with the context it is given. Typing a message does not retrain it. A product may store your earlier conversation, or some saved notes, and add them to the context later; that is the app doing the remembering, not the model being trained.

Build with me · 5

Show that prediction leaves training fixed

Sampling reads the stored counts. It does not add the generated words back into training. The table before and after these predictions is the same.

The same is true of a chatbot answering you: using the model does not change it.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
show the next-word table forpatterns
repeat3times
saySamplethen
a random word afterthefrompatternswith temperature1and top-p1
show the next-word table forpatterns

Show the table before and after three samples, without another learns block between them.

What to look for

Both table displays have the same cat and dog counts after the, even if sampled words differ.

Make it yours

Explain how a conversation can use earlier text as context without changing learned model weights.

Build with me · 6

Make an explicit model update visible

This time a learns block explicitly adds observations. It contributes three more dog followers after the, taking dog from two to five while cat stays at three. The model's stored counts changed, so its preferred follower changes.

A neural network is trained by adjusting its parameters rather than by adding counts, but the distinction is the same: training changes the model; using it to predict does not. Do not claim that a normal prediction performs the training step shown here.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
patternslearns the patterns inthe dog the dog the dog
show the next-word table forpatterns
sayNew greedy choicethen
patternsguesses the word afterthe

Place the extra learns statement after the first corpus and before the new table inspection.

What to look for

The table records five dogs and three cats after the, and the greedy choice is dog.

Make it yours

Compare this deliberate learning update with the unchanged before/after tables from the previous stage.

Where facts and tools fit

Parameters can hold factual associations, so it would be wrong to say a language model knows no facts. The limit is that producing likely-sounding text does not check a claim, name a source, or guarantee up-to-date information.

An application can add documents, search results, or the output of a tool to the context, and the model can use them in its answer. Whether that actually happened depends on the app and what it did, not on how confident the answer sounds. A citation must point to a real source that supports the claim; a citation-shaped string is not proof. This is why Level 1 taught you to check specific claims against their sources.

Sampling does not evaluate truth

Temperature and top-p change how the next token is chosen, though real systems may apply them differently from your toy. A lower temperature can make answers more repeatable, but it does not check dates, calculations, or quotations. A higher temperature can vary the wording, and it can also bring in less suitable continuations.

Build with me · 7

Vary wording while leaving evidence untouched

Temperature changes selection weights. The underlying count table remains the same, and no source lookup or truth evaluation appears anywhere in the stack. A more varied answer is not automatically more false, and a repeated answer is not automatically true.

Modern assistants can have additional training and tools, but the responsibility to check a factual claim against evidence is not replaced by a sampling slider.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
sayLow temperaturethen
a random word afterthefrompatternswith temperature0.3and top-p1
sayHigh temperaturethen
a random word afterthefrompatternswith temperature2and top-p1
show the next-word table forpatterns

Ask for one low-temperature and one high-temperature sample while keeping the corpus fixed. Inspect the unchanged table.

What to look for

Outputs may differ or coincide by chance. Both are selected from the same observed candidates.

Make it yours

Point to the block that would verify a factual claim. There is none in this program; explain why that matters when interpreting fluent output.

The last stage puts the shared idea and its limits side by side.

Build with me · 8

Keep the connection and its limits together

A strong explanation names both halves. Choosing the next piece from the context and repeating is shared with a chatbot. Learned representations, attention across the context, the scale of training, instruction-following, and tools go far beyond this table.

Blocks at this stageWorked example
make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
repeat5times
setwordto
a random word after
line
frompatternswith temperature1and top-p1
setlineto
join
line
and
join and
word
sayGenerated prefixthen
line
sayShared idea: choose a next unit using context and continue
sayLimits: this toy has no neural representations, instruction training, retrieval, or factual verification

Run the final short generator, then explain its steps from the actual growing prefixes. Use the final messages to check that your analogy has an explicit boundary.

What to look for

Five prefixes show repeated generation. The closing messages separate the shared idea from missing mechanisms.

Make it yours

Explain temperature to a beginner using one observed prefix, then add one sentence saying why it is not a truthfulness control.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make word patterns calledpatternswith a context of1words
patternslearns the patterns inthe cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps
setlinetothe
repeat5times
setwordto
a random word after
line
frompatternswith temperature1and top-p1
setlineto
join
line
and
join and
word
sayGenerated prefixthen
line
sayShared idea: choose a next unit using context and continue
sayLimits: this toy has no neural representations, instruction training, retrieval, or factual verification
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  import jsonimport random  def _words_in(value):    letters = []    for character in str(value).lower():        letters.append(character if character.isalnum() else " ")    return [word for word in "".join(letters).split() if word]  class WordPatterns:    """Counts which word tends to follow which. That is the whole trick behind    next word prediction, just with far fewer words than a real model.     The context is how many words back it looks. With a context of 1 it asks    "what usually comes after 'the'". With a context of 2 it asks "what usually    comes after 'the white'", which is a better question and needs more text to    answer."""     def __init__(self, name="patterns", context=1):        self.name = name        self.context = max(1, int(context))        # Keyed by a tuple of the last few words, so ("the",) and        # ("the", "white") are different questions with different answers.        self.follows = {}     def _context_of(self, words):        """The last few words, as the key this table is built on."""        if len(words) < self.context:            return None        return tuple(words[-self.context:])     def learn(self, text):        words = _words_in(text)        for i in range(len(words) - self.context):            key = tuple(words[i:i + self.context])            bucket = self.follows.setdefault(key, {})            following = words[i + self.context]            bucket[following] = bucket.get(following, 0) + 1        print("Learned patterns from " + str(len(words)) + " words.")     def next_word(self, text):        """The word that followed most often. On a tie, the earlier one in the        alphabet, so the same text always gives the same answer."""        bucket = self.follows.get(self._context_of(_words_in(text)))        if not bucket:            return ""        return min(bucket, key=lambda word: (-bucket[word], word))     def chance(self, text):        """Out of 100, how often that guess was the word that came next."""        bucket = self.follows.get(self._context_of(_words_in(text)))        if not bucket:            return 0.0        total = sum(bucket.values())        return round(100.0 * max(bucket.values()) / float(total), 1)     def sample(self, text, temperature=1.0, top_p=1.0):        """Picks one of the words that followed, at random, favouring the        common ones. This is what a real model does instead of always taking        the most likely word, and it is why a chatbot does not say the same        sentence twice.         top_p keeps only the shortlist of words that together account for that        share of the counts. temperature flattens or sharpens the odds: below 1        the common words get even more likely, above 1 the rare ones get a        look in."""        bucket = self.follows.get(self._context_of(_words_in(text)))        if not bucket:            return ""         # Commonest first, ties broken by the alphabet so a run is repeatable.        ranked = sorted(bucket.items(), key=lambda pair: (-pair[1], pair[0]))        total = float(sum(count for _, count in ranked))         # The shortlist: the fewest words whose share reaches top_p. Always at        # least one, or there would be nothing to choose from.        share = 0.0        shortlist = []        for word, count in ranked:            shortlist.append((word, count))            share += count / total            if share >= top_p:                break         # Raising each count to 1 / temperature is what sharpens or flattens        # them. The floor stops a temperature of 0 dividing by zero.        power = 1.0 / max(float(temperature), 0.01)        weights = [float(count) ** power for _, count in shortlist]         pick = random.random() * sum(weights)        running = 0.0        for (word, _), weight in zip(shortlist, weights):            running += weight            if pick <= running:                return word        return shortlist[-1][0]     def show(self, limit=8):        print(self.name + " learned " + str(len(self.follows)) + " starting words.")        for key in list(self.follows)[:limit]:            bucket = self.follows[key]            best = min(bucket, key=lambda other: (-bucket[other], other))            print("  after '" + " ".join(key) + "' comes '" + best + "'")     def table(self, limit=40):        """Opens the next-word table panel.         There is nothing useful to print in a terminal here, so this emits one        marked line that the workspace reads and draws as a table. The runner        takes the line back out of the output, so you never see it."""        rows = []        common = sorted(            self.follows.items(),            key=lambda pair: (-sum(pair[1].values()), pair[0]),        )        for key, bucket in common[:limit]:            total = sum(bucket.values())            followers = sorted(                bucket.items(), key=lambda pair: (-pair[1], pair[0])            )            rows.append(                {                    "context": " ".join(key),                    "count": total,                    "followers": [                        {                            "word": word,                            "count": count,                            "share": round(count / float(total), 4),                        }                        for word, count in followers                    ],                }            )        print("__TABLE__ " + json.dumps({"name": self.name, "contexts": rows}))  def new_word_patterns(name="patterns", context=1):    return WordPatterns(name, context)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  patterns = new_word_patterns("patterns", 1)patterns.learn("the cat sleeps the cat plays the dog sleeps the dog runs the cat sleeps")line = "the"for _ in range(int(5)):    word = patterns.sample(line, 1, 1)    line = (str(line) + str((str(" ") + str(word))))    print("Generated prefix", line)print("Shared idea: choose a next unit using context and continue")print("Limits: this toy has no neural representations, instruction training, retrieval, or factual verification")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in