0%
ReadingUnder the hood14 min, about 875 words

The Examples Are the Lesson

An optional closer look at examples, labels, coverage, and why a dataset shapes a model.

One example has two roles

For a supervised image classifier, an example pairs an input with an answer: a photograph and the label mug. The photograph provides evidence; the label tells the training procedure what output is desired. If the label says glass instead, the procedure receives a contradictory instruction.

Consider four examples:

ExampleVisible objectLabel
AMug with a handlemug
BDifferent mug with a handlemug
CDrinking glassglass
DAnother drinking glassmug

Before changing a model setting, inspect D. Either its label is wrong or our definition of mug differs from what the table suggests. A shared written labeling rule helps resolve the disagreement.

More is not always more information

Fifty almost identical frames of the same mug show very little variation. Fifty photos across objects, positions, cameras, and lighting cover more cases. Quantity can help, but it does not automatically repair systematic mistakes or missing conditions.

Likewise, a thousand examples of one class and ten of another can encourage a system that favours the common class. Compare class counts and inspect errors separately. Level 2 turns this into an actual training workflow.

The accidental clue

Suppose every mug photo includes your hand and every glass photo is taken on a table. The hand predicts the label in your dataset. The model has no direct access to your intention that it should look for a handle instead.

To investigate, test a mug on the table and a glass held in your hand. If the predictions follow the hand, you have evidence of a shortcut. Make training examples that break that accidental relationship, then retrain and evaluate again.

What a missing example means

An absent condition does not guarantee failure: models can generalise. But you should not assume dependable performance on conditions you never evaluated. If your camera model is intended for dim rooms, include dim-room testing. If a new condition is very different, label the result as uncertain and investigate it.

A pretrained image model also brings patterns learned before your own small dataset. Your examples adapt that prior representation. “It learns from your photos” does not mean every visual feature was first learned from those photos alone.

Training is not the only information source in an AI product

A trained model's parameters reflect its learning process, but a deployed application can also supply a current document, retrieve a webpage, or call another tool. Keep these sources distinct when explaining an answer. The useful question is “What evidence was available for this result, and how was it checked?”

Design a collection that separates the label from the background

For a mug-versus-glass classifier, imagine collecting twenty mug images on a blue mat and twenty glass images on a wooden desk. Both classes have the same count, yet background and label are perfectly connected. The model may exploit the surface instead of the object.

Plan a replacement collection as a grid: mugs on the blue mat, mugs on the desk, glasses on the blue mat, and glasses on the desk. Vary angle and lighting within each combination. This breaks one shortcut while keeping the labels meaningful. It still does not prove that every other shortcut has disappeared.

Reserve a separate object or recording session for evaluation before taking the training photos. Consecutive video frames of one mug are not twenty independent mugs. Moving one frame into the test pile can make performance look stronger because nearly identical frames remain in training.

After collection, inspect each pile. Count labels, check that the object is visible, remove genuinely unusable captures, and correct labels using what the image depicts. Do not remove a valid difficult example merely because the model gets it wrong. That difficulty is evidence about the task you want to teach.

A dataset review you can repeat

Inspect a few examples from every label. Check ambiguous cases against the labeling rule. Count examples and look for repeated near-duplicates. Ask which conditions are absent. Set aside independent evaluation examples before adjusting the model. This review often reveals a more useful next action than changing a training setting at random.

Work through a data change before talking about a smarter model

Use a tiny invented classification task: decide whether a message asks a question. A collection containing only messages ending in a question mark may teach that punctuation is the answer. “Could you explain this” lacks the mark but is still a question; “I wondered, why?” may require a task-specific rule about quotations and context.

Write the label definition, then add cases that challenge the shortcut. Keep some separate for development so you can observe whether the change travels beyond the exact examples you added. If two people cannot apply the label rule consistently, resolve that disagreement before treating it as a fitting problem.

Later blocks and Python will make these operations executable: create a dataset, add input-label pairs, fit, predict, and compare with held-out labels. The programming syntax changes how you express the steps. It does not remove the need to decide what the examples mean.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in