Watch a model choose the next word
Use a small worked table to understand greedy selection, sampling, temperature, and repeated generation.
You can follow this without another website
We will use an illustrative next-word table. Real models operate on tokens, but whole words make the first example easier to read. These numbers are a teaching example, not measurements from a live model.
Context: “For breakfast I poured milk on my”.
| Next word | Probability in this example |
|---|---|
| cereal | 0.60 |
| oats | 0.25 |
| shoes | 0.10 |
| homework | 0.05 |
A probability of 0.60 means 60 chances out of 100. Add the four numbers and you get exactly 1, because one of the candidates will be chosen. The whole list of candidates with their chances, adding up to 1, is called a distribution. It describes what the model expects to come next in this made-up example. It is not evidence that a real breakfast happened.
Choose greedily
A greedy procedure takes cereal because 0.60 is the largest number. Repeat the exact same procedure on the exact same table and it takes cereal again. Greedy does not mean “knows the correct answer”; it means “chooses the highest-scoring candidate”.
Choose by sampling
Imagine a bag containing 60 cereal tickets, 25 oats tickets, 10 shoes tickets, and 5 homework tickets. Draw one ticket. Cereal is most likely, but oats is possible. Over many independent draws the proportions tend toward the table; ten draws need not contain exactly six cereals.
This is the intuition behind sampling. It can make text less repetitive, but can also choose an odd continuation. It is not a factual verification method.
Change the temperature
Many chatbots have a setting called temperature that reshapes the distribution before a candidate is picked. A lower temperature makes the likely candidates even more likely. A higher temperature spreads the chances more evenly.
Here is the same made-up table at a low and a high temperature. Each column still adds up to 1:
| Next word | Low temperature | Original | High temperature |
|---|---|---|---|
| cereal | 0.83 | 0.60 | 0.43 |
| oats | 0.14 | 0.25 | 0.28 |
| shoes | 0.02 | 0.10 | 0.17 |
| homework | 0.01 | 0.05 | 0.12 |
At the low temperature, cereal is picked almost every time and homework almost never. At the high temperature, the odd continuations come up far more often: shoes rises from 1 draw in 10 to about 1 in 6. The order never changes. Temperature changes how spread out the chances are. It does not add knowledge the model lacked, and it does not make any answer truer.
Continue one step further
After selecting cereal, the context becomes “For breakfast I poured milk on my cereal”. The next table could favour punctuation or a phrase such as “before school”. After selecting shoes, the next table is different, because the model is continuing a different sentence.
This is why the start of a generated answer influences everything that follows. A model can continue an unlikely or false premise with grammatically smooth text.
Connect this to your own use
For a fixed formatting task, you usually want dependable structure. For brainstorming, several varied candidates may be useful. In both cases you still judge the result against the task. A low-temperature answer can be false; a high-temperature answer can be correct.
If you later open a next-token visualizer, read the context first, then the candidate list, then the setting. Change only one setting before comparing results. You already have the idea needed to interpret the display; the visualizer is optional enrichment.
Check your understanding
If the sampler selects oats, has it malfunctioned? No: oats has positive probability in the example. If cereal is selected six times in a row, does that prove greedy selection? No: sampling can also produce that sequence. What we can conclude depends on the procedure, not just a small set of outputs.
Count a tiny model you can inspect completely
Take four training sentences: “we build robots”, “we build games”, “we build robots”, and “we test games”. After the context “we build”, robots appears twice and games once. A count-based teaching model would assign robots two of the three observed continuations and games one of three. After “we test”, only games has been observed.
Every number has a source you can point to. Counting a repeated sentence twice changes the distribution; removing it changes the counts. Such a tiny table is not how a large neural chatbot stores all its knowledge, but it exposes the roles of examples, context, continuation scores, and choice.
Add your own sentence beginning “we build”. Recount before predicting which next word is favoured. Later, the blocks version will build and display this real count table in the article's mini editor.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.