Guess, check, adjust
Follow a numerical training example and distinguish changing a model from using it.
A model small enough to inspect
We want to predict a delivery time from distance. To keep the example simple, suppose the model is:
predicted minutes = minutes per kilometre × distance
The adjustable number is minutes per kilometre. We call an adjustable numerical setting a parameter. Real deliveries have more causes, but this small model lets us see a training step clearly.
Our one training example is a 2-kilometre delivery that took 10 minutes. Start with the parameter equal to 3. The model predicts 3 × 2 = 6 minutes. It is four minutes short.
Give the error a score
Choose squared error as our loss: subtract the true answer from the prediction, then multiply the difference by itself.
- Prediction: 6.
- True answer: 10.
- Difference:
6 - 10 = -4. - Squared error:
(-4) × (-4) = 16.
Squaring makes both overestimates and underestimates costly. A different task can use a different loss; loss is not always squared error and is not the same thing as accuracy.
Adjust and compare
Try a parameter of 4. The prediction becomes 8 and the loss becomes 4. Try 5. The prediction becomes 10 and the loss becomes 0. On this one example, 5 fits exactly.
| Parameter | Prediction for 2 km | Squared error |
|---|---|---|
| 3 | 6 | 16 |
| 4 | 8 | 4 |
| 5 | 10 | 0 |
| 6 | 12 | 4 |
This table is a hand-operated version of “guess, check, adjust”. Actual training algorithms choose updates systematically rather than trying every possible number. For many models, including neural networks, an algorithm works out how changing each parameter would change the loss, then takes a small step toward lower loss. (Parameters are also called weights, especially in neural networks.)
The learning rate controls the size of those steps. Too small can make progress slow; too large can make training unstable. We will inspect those effects later, after we have a working model.
Why one correct example is not enough
A second delivery, 4 kilometres in 18 minutes, is not perfectly explained by 5 minutes per kilometre: the model predicts 20. Training on several examples seeks a useful compromise under the chosen loss. Fitting every training example exactly is not always possible or desirable.
We also need examples kept out of training on purpose, often called held-out examples. A model can fit the examples it saw and still fail on new ones. The training loss answers “How well does the model fit the examples it learned from?” Testing on held-out examples asks whether it generalises: whether it still works on examples it has never seen.
Training versus using the result
After training, suppose we keep the parameter at 5. Asking for a 3-kilometre prediction computes 5 × 3 = 15. That is prediction, also called inference. The act of asking did not update the parameter.
Similarly, an ordinary chatbot conversation supplies new context without necessarily changing the model's weights. A product may store information or later use data in a separate training process; that is a different operation from generating the next response.
Not every learning algorithm uses this exact loop. The nearest-example classifier later in the course stores examples and compares distances. The broader principle is still to derive behaviour from data and evaluate it honestly.
Work through two examples together
Return to the two observed deliveries: 2 kilometres took 10 minutes, and 4 kilometres took 18. We now score a proposed parameter on both examples. A parameter that fixes the first delivery can make the second worse, so calculate both before deciding.
| Minutes per kilometre | Prediction at 2 km | Prediction at 4 km | Two squared errors | Mean loss |
|---|---|---|---|---|
| 4 | 8 | 16 | 4 and 4 | 4 |
| 4.5 | 9 | 18 | 1 and 0 | 0.5 |
| 5 | 10 | 20 | 0 and 4 | 2 |
For the middle row, the differences are 9 - 10 = -1 and 18 - 18 = 0. Squaring gives one and zero; their average is (1 + 0) / 2 = 0.5. Among these three candidates, 4.5 is best by this loss, even though it does not exactly match the first delivery. We have compared three choices, not proved that 4.5 is the best possible value.
Notice what stays fixed during this comparison: the recorded distances and actual times. We change the model's parameter, not the answers in the dataset. Changing a correct label to agree with a prediction would hide the mistake instead of teaching the model.
What to remember
You should be able to point to the input, the known answer, the prediction, the loss, and the adjustable setting in this example. Those roles will reappear in photos, text, and handwritten digits, even when there are many more numbers.
Separate the three moments in one training example
Imagine a toy learner that predicts whether a drawn shape is a circle. For one training drawing it gives circle a score of 0.3, while the attached target is circle. First there is a prediction from the current parameters. Second, the loss measures how far that prediction is from the target. Third, an update changes the parameters in the direction that should reduce the loss.
The true label belongs to the training example; it is not produced by the model's guess. If the drawing was mislabelled square, training receives the wrong target. Repeating the update does not repair that target automatically.
Next imagine showing the trained model a new drawing whose label you have not supplied. It can predict, but nobody can say whether that prediction is right without an answer to compare it with. Write these three verbs separately: train changes the parameters, predict produces scores, evaluate compares predictions with known answers. You will perform each action in Level 2.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.