0%
BuildFind out where it breaksabout 25 min, 2 steps

Test it honestly

Build a labelled development set, record every prediction, calculate accuracy, and reserve a separate final test.

Prepare a development set

Your first development photos were a taste. Now you will build a full development set and measure the model properly. Your model from Your first training run should still be loaded. If you refreshed the page, press Open model in Test & export and choose your downloaded model, or press Train model again with the same photos.

Use the development objects you reserved in Setting up your image lab. Plan, for example, ten usable mug photos and ten glass photos, with several views and ordinary lighting changes.

Choose the whole set before you see a single prediction. Do not drop a hard photo because the model gets it wrong. You may leave out a broken file or a photo outside the task (both objects in one frame, say), but decide that with a written rule, without looking at whether the model was right.

Build with me · 1

Add the whole development set and predict it once

Give your files names you can recognise later, such as dev01.jpg, dev02.jpg, and so on. The lab turns each filename into the photo's ID (dev01). An ID is a stable name for one example. It is not its answer.

Press the correct label before adding each photo, using your written rule. You decide the true answer before the model gets a say. The lab then keeps, for every photo, its ID, the actual label you gave it, the model's prediction, both scores, and whether the prediction was correct. Some things it cannot record, such as "dim room" or "side view". Those go in Notes, next to the ID:

IDActual labelCondition, in Notes
dev01mugside view
dev02glassdim room

After you predict, three things appear above the results: an Accuracy line, a line starting Always- (the baseline), and a small table. The next stage explains accuracy. The table gets a reading of its own straight after this one, and the baseline line is explained later, in Class balance, and what accuracy hides, so you can leave both for now.

Your image lab at this stageWorked example
Text
Photos  These photos are for: Development  mug:   dev01 to dev10  glass: dev11 to dev20Test & export  Predict this collection: Development (20)  Predict collection

In Photos, choose Development and add the rest of your planned development photos, with the correct label pressed each time. Check the IDs under the thumbnails. In Test & export, predict the Development collection once and scroll through every result, including the wrong ones.

What to look for

One result per development photo, each with an actual label and a correct or incorrect mark. Your training counts are unchanged.

Make it yours

Pick your two most unusual development photos and write their conditions in Notes before you look at their results.

Build with me · 2

Work out the accuracy yourself

Accuracy is the share of predictions that were correct. Count every photo you evaluated. Count the ones where the prediction matches the actual label. Divide the second number by the first, then multiply by 100 to get a percentage. The number you divide by, the bottom of the fraction, is called the denominator. Here it is 20.

In the example, the model got 8 of 10 mugs and 5 of 10 glasses right: 13 correct out of 20, which is 65%. The lab would show this as Accuracy: 13/20 (65.0%). The scores on individual photos play no part in this sum. A score of 95% on one photo is not 95% accuracy.

With 20 photos, each photo is worth 5 percentage points (100 divided by 20). So a change from 65% to 70% is one photo changing its answer, which is too little to prove a real improvement. Always report the fraction, not just the percentage: 13/20 (65%), with the number of different objects and the conditions you tested. Twenty photos of one mug test far less than twenty photos of ten different mugs.

These numbers are made up for the example. Yours will differ; never change your collection to match them.

Your image lab at this stageWorked example
Text
Example evaluation  actual mugs:    8 correct out of 10  actual glasses: 5 correct out of 10  correct  = 8 + 5 = 13  total    = 10 + 10 = 20  accuracy = 13 / 20 = 0.65, which is 65%

Count the correct results in your own Development predictions by hand, and check that your count matches the lab's Accuracy line. In Notes, write the fraction, the percentage, the number of different objects, and a short description of the conditions.

What to look for

Your hand count matches the lab. Every planned development photo is included, so hard photos have not quietly dropped out of the denominator.

Make it yours

Work out how many percentage points one photo is worth in your own set: 100 divided by the number of photos. Explain why a small test moves in such big steps.

Look for a useful next question

In the example, the model got 8 of the 10 mugs right but only 5 of the 10 glasses. The glasses are the weaker class. Look at the five glass photos it got wrong. Are they all dark? All narrow? Does their background match the mug photos in training?

Write down a possible cause, such as "Our training set has no glass photos in dim light." A possible cause that you can test is called a hypothesis. It is not a proven cause. The next reading sorts mistakes into a table that makes patterns like this easier to see, and Fix the data and retrain shows how to test a hypothesis.

Protect the final test

Because you will use these results to choose improvements, call them development results in your notes. Leave the final-test objects alone. After you have chosen your best version, you will predict the final test once and report that result separately.

You have finished this reading when another person could rebuild your score from your records. A screenshot of one confident prediction is not enough.

Full reference solution

The whole Level 2 workflow on one page. Each reading covers part of it; this list shows where that part fits. The numbers in the readings are examples, and your own results will differ.

Text
1. Name the two classes and write a one-sentence rule for each.2. Reserve objects or sessions for Development and Final test before collecting.3. In Photos, choose Training and a label, then add varied photos of both classes.4. Add reserved photos under Development and Final test, never under Training.5. In Notes, record label rules, objects, conditions, the split, and a run name.6. In Train, check both counts, keep 20 rounds for the first run, and press Train model.7. Download project and Download model, with matching run names.8. In Test & export, predict Development and read every ID, actual label, prediction, and score.9. Work out accuracy yourself, read the confusion matrix, and compare with the Always- baseline.10. Write one hypothesis. Change only the Training photos (or one stated setting), retrain, and predict the same Development collection.11. Choose a version using Development results. Open its saved model if needed.12. Predict Final test once, after choosing, and report its fraction, conditions, and limits.13. For photos with no known answer, use Unlabelled targets, predict, and Download predictions.csv.14. Keep project, model, notes, and predictions file together under one run name. None of them replaces another.

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in