Why you hold data back
Separate training, development, and final testing so that improvement does not turn your final test into another source of training decisions.
The question your test should answer
You want to know whether the model can classify mugs and glasses it has not seen, not whether it recognises the same desk scene again. So first decide what job the model is for. A model for one fixed camera on one desk has a different job from a model for anyone's phone in any kitchen.
A useful test uses photos that match that job and were not used in training. Testing on training photos usually makes the model look better than it really is, because training has already adjusted it to those exact photos and their labels.
Here is a comparison. A student who practises the same twenty arithmetic questions over and over can answer those twenty. Whether they can answer new questions of the same kind is a different question, and only new questions can answer it. Models do not learn the way people do, but the point about the test holds.
Three jobs, used the same way every time
You gave your photos these jobs when you set up the lab. Here is what each job is for, and what you are allowed to change because of it:
| Job | You use it to | What you may change because of it |
|---|---|---|
| Training | Train the model | The learned numbers, during training |
| Development | Compare versions and study mistakes | Your photos and training settings |
| Final test | Measure the version you already chose | Only your report |
Photos kept out of training on purpose are often called held-out photos. Many tools call the development set a validation set.
You may check the development set as often as you like; that is what makes it useful for comparing versions. But each time its results lead you to change something, it helps shape the model, so it can no longer act as an untouched test.
That is why the final test stays closed until you have picked a version. If you look at its mistakes and then change the model, the final test has become one more development set. Say so honestly, and collect fresh photos if you need a fresh final score.
Split by what should be new
Split physical objects before you take photos. All photos of one mug belong to one job. So do photos taken in one quick burst, and any copies or edited versions of a photo. Otherwise the test contains near-copies of training photos and rewards the model for recognising them.
With enough objects, use several per class for training, at least one different object per class for development, and different objects again for the final test. More objects make a stronger test. No split percentage makes a tiny collection reliable.
When information from the test photos sneaks into training, even through a near-copy or a duplicate, it is called data leakage. Leakage makes a score look better than the model deserves. Never copy photos between jobs to improve a score.
Recognising overfitting
Overfitting means the model has learned details that are special to its training photos and do not carry over to new ones. The warning sign is a big gap: the model fits its training photos very well (a very low training loss) but does clearly worse on development photos. A model that does well on new examples it never trained on is said to generalise.
A gap does not prove overfitting on its own. Different conditions between the sets, wrong labels, or a very small development set can also cause one. Look at the photos it got wrong before deciding.
If the model does badly on both (a training loss that stays high, and poor development results), look elsewhere: labels that disagree, photos that do not show the object clearly, or too little training. Scores that hover near 50% on ordinary photos point the same way: the model has not found a clear difference, and labels applied with two different rules are a common reason. More rounds are not an automatic fix. If moving the same object from one background to another changes the prediction, you may have found a background shortcut. Test that idea by photographing both classes on both backgrounds.
A confident wrong answer shows that confidence is not a guarantee. On its own, it does not tell you which shortcut the model learned.
Build with me · 2
Write the split down
A split you cannot describe is a split you cannot defend. Write it down while you still remember which object went where. The two most important lines are the last two: which set you will use to choose changes, and which set stays closed until you have chosen. Being able to answer those two questions matters more than any particular percentage split.
Split record (example) Training: mugs A, B, C and glasses A, B, C (desk and blue mat) Development: mug D and glass D (kitchen) Final test: mug E and glass E (another room), kept closed Chooses improvements: Development Untouched until the end: Final testIn Notes, write a split record like the example, using your own objects or sessions. If you used the shape demonstration, write that its drawings were split for you: 12, 4, and 4 per class.
What to look for
Someone else could read your record, check that no object appears in two jobs, and tell which set chooses improvements.
Make it yours
Add a line for any photo you left out on purpose, such as both objects in one frame, with the reason. A rule for leaving photos out, written before you see any predictions, keeps the test honest.
Full reference solution
The whole Level 2 workflow on one page. Each reading covers part of it; this list shows where that part fits. The numbers in the readings are examples, and your own results will differ.
1. Name the two classes and write a one-sentence rule for each.2. Reserve objects or sessions for Development and Final test before collecting.3. In Photos, choose Training and a label, then add varied photos of both classes.4. Add reserved photos under Development and Final test, never under Training.5. In Notes, record label rules, objects, conditions, the split, and a run name.6. In Train, check both counts, keep 20 rounds for the first run, and press Train model.7. Download project and Download model, with matching run names.8. In Test & export, predict Development and read every ID, actual label, prediction, and score.9. Work out accuracy yourself, read the confusion matrix, and compare with the Always- baseline.10. Write one hypothesis. Change only the Training photos (or one stated setting), retrain, and predict the same Development collection.11. Choose a version using Development results. Open its saved model if needed.12. Predict Final test once, after choosing, and report its fraction, conditions, and limits.13. For photos with no known answer, use Unlabelled targets, predict, and Download predictions.csv.14. Keep project, model, notes, and predictions file together under one run name. None of them replaces another.Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.