Fix the data and retrain
Use development mistakes to choose one controlled change, retrain, and keep an honest experiment log.
Start with a question you can investigate
In Test it honestly you wrote down a possible cause for your model's mistakes, a hypothesis. This reading tests one properly: change one thing, train again, and compare on the same development photos.
Open your development results again. Your model should still be loaded; if you refreshed the page, press Open model in Test & export, or train again with the same photos. If you need your run 1 photos and notes back, Open project restores them from your download.
A good hypothesis links an observed failure to one specific change. "Make the model smarter" is not specific. "Add glass and mug photos in dim light, and keep everything else the same" gives you a comparison you can interpret.
Build with me · 1
Turn a repeated mistake into one testable hypothesis
A mistake is an observation: something you saw. A possible reason for it is a hypothesis. Keep them on separate lines in your log. If the development mistakes are mostly dark glasses, and the training set has only bright glasses, missing dim-light photos is a sensible guess. Noticing both facts does not prove it.
Look at the failed photos themselves. Check their labels first; a wrong label is the cheapest fix. Then compare their objects, angles, crop, background, and lighting with the training photos. A photo whose handle was cut off by the square crop may fail simply because the evidence is missing. A well-framed photo that fails may point to a shortcut, or to a kind of object that training never showed.
Then choose one change to test. For the lighting idea, take new dim-light training photos of both classes, and keep the labels and training settings the same. Never copy the failed development photos into Training. That teaches the model the exact photos you are using to judge it, and the comparison stops meaning anything.
Observation: several dim glasses were called mugTraining check: every training glass is brightly litHypothesis: the missing dim-light photos matterChange: add NEW dim-light training photos of both classesComparison: same development photos, same settingsIn Notes, write Observation, Hypothesis, and One change on separate lines. Use IDs to point at the group of mistakes. If the model made no development mistakes, test a sensible new condition instead, using a separately named extra set of development photos, rather than inventing a failure.
What to look for
Your planned change can be carried out without touching the final-test photos and without copying development photos into Training.
Make it yours
Write a different explanation for the same mistakes. Describe what evidence would tell your two explanations apart.
Make one deliberate change
Keep the task and your label rules the same while you make the change. If you discover a training photo with the wrong label, fix it before the comparison and write the fix in your log.
For a background shortcut, photograph each background with both classes. Deleting every difficult photo can raise the score while making the task less like real use. Add the conditions the model will actually meet, with labels you can defend.
Build with me · 2
Train again, then ask the same questions
Adding photos changes the project. It does not change the model you already trained. The lab reminds you of this in Train: "Your examples or class names changed after training. This is still the earlier model." Press Train model again to make run 2, then predict the same Development collection with it.
A new training run replaces the model in the lab. So before you train, download run 1's project and model and give them matching names, as you did in Your first training run. If run 2 turns out worse, those files are your way back.
Keep the comparison fair. If you change the lighting, the class counts, the number of rounds, and the development photos all at once, you cannot tell which change mattered. Even with one change, a small set of photos limits what a single comparison can prove.
Run 1: original training photos development 13/20Run 2: + new dim-light TRAINING photos development 16/20Same 20 development photos. Same labels. Same 20 rounds.(Example numbers.)Download run 1's project and model. Make your one change to the Training photos, keep 20 rounds, and press Train model. Predict the same Development collection. Write run 2's fraction next to run 1's in Notes, and list which photos were fixed and which became wrong.
What to look for
Both runs were scored on the same development photos, so the denominators match. Your log records the change you actually made, even if the result got worse.
Make it yours
If accuracy did not change, compare the individual results. One photo fixed and another newly wrong cancel out in the total, while telling you something useful.
Keep a log of every run
A short table in Notes is enough:
| Run | The one change | Development result | What still goes wrong |
|---|---|---|---|
| 1 | First collection | 13/20 | Several dim-light glasses |
| 2 | Add both classes in dim light | 16/20 | Side-on narrow mugs |
| 3 | Add training views of narrow mugs | 17/20 | Three unrelated mistakes |
These are example numbers, not results to expect. Run 2 improves by three photos; run 3 by one. With a small set, a one-photo change may be luck rather than a real improvement. Repeat a promising experiment when you can, and collect more development photos if the decision matters.
What if the score gets worse?
Check the labels, compare the same photos, and look for regressions: photos the old model got right that the new one gets wrong. A change can fix one group while breaking another. Keep both models until you have decided which suits the job better. A worse result is still useful if your log says exactly what was tested.
If the training loss is low but development results are poor, look for overfitting, a mismatch between training and development conditions, leakage, or labelling errors. If both are poor, look for unclear labels, photos that do not show the object well, too little training, or poor data quality. Neither pattern points to one guaranteed fix.
The other training settings
Three settings control how training runs. You have met one already: the number of rounds, or epochs, where each round is one pass through all training photos. The batch size is how many photos the model looks at before each nudge to its learned numbers. The learning rate is how big each nudge is. The lab fixes the batch size at 16 and the learning rate at 0.001, and lets you change only the rounds.
For now, keep 20 rounds while you change photos, so each comparison changes one thing. Later you might compare 10 and 50 rounds on the same development photos. Without a chart of results after every round, do not claim the model "peaked at round 30" from the final score alone. Level 4 draws those charts.
Select, then test
Stop when the model is good enough for your stated job, or when your planned changes stop helping on development photos. Choose one version using the evidence you have. Only then open the final test, once.
Build with me · 3
Run the final test once, on the version you chose
The final test answers a different question from development: how well does the version you already chose do on photos that played no part in choosing it? Make the choice from development results first. If the chosen version is not the one in the lab, load it with Open model. Then predict the Final test collection.
Before it shows final-test results, the lab asks you to confirm that you have finished choosing. Press Evaluate selected model to go ahead, or Keep it reserved to wait. The dialog is only a reminder. Software cannot keep a test honest; your choices do. If you have already looked at these photos' mistakes and changed the model because of them, they have become development photos, and giving them a new name does not make them fresh.
Record the full fraction, the class counts, the variety of objects, and the limits of the test. If the result is disappointing, report it anyway. You can keep improving the model, but a new final score then needs a new, untouched set of photos. Never report only your best final-test attempt.
Choose a version using Development results ↓Open model, if that version is not the one loaded ↓Predict Final test once ↓Report the result, what was tested, and the limitsChoose your version, then in Test & export select Final test and press Predict collection. When the reminder appears, press Evaluate selected model. Copy the fraction and what the test covered into Notes.
What to look for
Your notes keep the development result that chose the version apart from the final result measured afterwards. Every planned final-test photo is included.
Make it yours
Write one situation your test does not cover, such as a different camera, unfamiliar materials, or both objects in the same frame. A precise limit makes the result more useful.
Write up what you did
Report the training collection, the change you chose, the development comparison, the final-test count and score, and the limits. Keep the chosen project, split record, final results, and log together. A reader should be able to tell how you chose the model apart from how you tested it.
Full reference solution
The whole Level 2 workflow on one page. Each reading covers part of it; this list shows where that part fits. The numbers in the readings are examples, and your own results will differ.
1. Name the two classes and write a one-sentence rule for each.2. Reserve objects or sessions for Development and Final test before collecting.3. In Photos, choose Training and a label, then add varied photos of both classes.4. Add reserved photos under Development and Final test, never under Training.5. In Notes, record label rules, objects, conditions, the split, and a run name.6. In Train, check both counts, keep 20 rounds for the first run, and press Train model.7. Download project and Download model, with matching run names.8. In Test & export, predict Development and read every ID, actual label, prediction, and score.9. Work out accuracy yourself, read the confusion matrix, and compare with the Always- baseline.10. Write one hypothesis. Change only the Training photos (or one stated setting), retrain, and predict the same Development collection.11. Choose a version using Development results. Open its saved model if needed.12. Predict Final test once, after choosing, and report its fraction, conditions, and limits.13. For photos with no known answer, use Unlabelled targets, predict, and Download predictions.csv.14. Keep project, model, notes, and predictions file together under one run name. None of them replaces another.Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.