0%
ReadingInside a chatbot8 min, about 641 words

Autocomplete that went to school

Trace how a language model turns a conversation into a sequence of generated tokens.

Begin with a familiar completion

Type “Please close the” and several endings are plausible: door, window, gate. The previous words limit what a sensible continuation looks like. A language model learns much richer relationships than this small example, but generating a continuation gives us a useful starting point.

A language model assigns scores to possible continuations of text. A chatbot application builds a conversation around a model. The application can also supply instructions, documents, search results, or tools. Keep those two things separate as we trace a response.

What is a token?

The model processes text in pieces called tokens. A token may be a whole word, part of a word, punctuation, or another text fragment. The exact split depends on the model's tokenizer, the part of the system that cuts text into tokens. For an illustration, imagine “unhelpful” split into “un”, “help”, and “ful”. This is an example split, not a promise about a particular product.

Tokens let a model handle words it has not encountered as complete units. They also explain why counting characters or spelling backward can be awkward: the units inside the model are not necessarily the letters you see.

Trace one response

Suppose the conversation ends with: “Finish this sentence: The dog chased the”.

  1. The application assembles the context: instructions and the conversation so far, perhaps with retrieved material (text the application looked up, such as a search result or part of a file).
  2. The text is cut into tokens, and each token is turned into numbers the model can calculate with.
  3. The trained model calculates scores for possible next tokens.
  4. A rule for picking one of the candidates selects one, perhaps “ball”. This picking rule is called a decoding procedure.
  5. That selected token becomes part of the context for the next step.
  6. The process repeats until the response ends or another stopping condition is met.

The model does not normally write a hidden complete answer first and reveal it one word at a time. Its next choices depend on the continuation already generated. If it starts with an incorrect assumption, later fluent text can build on that assumption.

Why two runs can differ

If the picking rule always takes the highest-scoring token, generation is called greedy. Other picking rules sample: they choose at random, but give higher-scoring candidates a better chance. Sampling can produce different responses to the same input. Temperature and other settings affect this selection; the next reading works through a numerical example.

Different responses can also result from different context, instructions, tools, or model versions. Variation is not evidence that the system learned something permanently from the first question.

What training contributed

During training, the learned numbers inside the model (its parameters) were adjusted using a very large amount of text and a goal, such as predicting the next token well. During an ordinary conversation, those learned parameters are only used to calculate outputs; they are not changed. Supplying “Here is my class timetable” gives the current response useful context; it does not, by itself, rewrite the model's training history.

Modern language models also receive forms of training intended to improve instruction-following and other behaviour. A next-token description explains an important mechanism, but it is not a complete explanation of every capability or limitation.

Where facts enter

A model may reproduce information learned during training. The application may also retrieve current sources or run calculations. A plain generated sentence does not tell you which of those happened. Look for actual tool use and check the supporting source when accuracy matters.

Remember the useful distinction: the model generates text from its supplied context and learned parameters; the whole application may do additional work. That is why “a chatbot is always searching” and “a chatbot can never search” are both poor descriptions.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in