0%
BuildMake it betterabout 12 min, 2 steps

Class balance, and what accuracy hides

Compare a model with a simple baseline, inspect each class, and connect error measures to the intended task.

A model that never looks at the picture

Imagine a labelled set of 200 dog photos and 5 cat photos. A program that ignores the pictures and always answers dog gets 200 of 205 right. That is 200 ÷ 205 × 100, about 97.6% accuracy, and it finds none of the cats.

This does not mean every model trained on uneven data ends up always saying dog. It means a high accuracy can be reached without learning anything about images, so a number on its own is no reason to celebrate. You need something to compare it with.

A baseline is a deliberately simple method used for that comparison. The majority-class baseline finds the most common label among the training photos and predicts it for every photo. Choose that label from the training data only. Never look at the final-test answers to pick whichever constant answer scores best there.

Work the baseline out before you train, as soon as your photos are split. Then it is a bar written down in advance that the model has to clear. Worked out afterwards, next to a score you already have, it tends to become a number you explain away. The arithmetic is the same either way; the order is what keeps you honest.

Build with me · 1

Find the baseline in your own results

The lab works out this baseline for you. When you train, it notes which class has more training photos; if the counts are equal, it picks the first class. After each prediction it shows a line such as Always-mug baseline: 10/20. That is the score you would get by answering mug for every photo in the collection you just predicted.

The baseline and your model are scored on the same photos, so you can compare them directly. On 10 mugs and 10 glasses, always saying mug gets 10 right: 50%. The example model got 13 right: 65%. So the model found three more correct answers than a program that never looks at the photo. That is a real but modest gain on 20 photos. It does not prove that every future user will see a 15-point improvement, so report the size of the test and its conditions alongside it.

Your image lab at this stageWorked example
Text
Training photos: 50 mugs, 48 glasses, so the majority label is mugBaseline: answer mug for every development photoDevelopment set: 10 mugs + 10 glassesBaseline score: 10/20 = 50%Trained model:  13/20 = 65%

Predict your Development collection and find the line starting Always- under Accuracy. In Train, check which class has more training photos. Count how many development photos have that label, and confirm the baseline's top number.

What to look for

The baseline and the model are divided by the same total. The baseline's label came from the training counts, not from trying labels on final-test answers.

Make it yours

Work out what an always-the-other-class program would score on your development set. Treat it as a way to understand the table, not as a new baseline chosen after seeing the answers.

When the baseline wins

If your model scores at or below the baseline, something basic is wrong. Check the labels, the photos, and the split before anything else. A complicated model is not automatically a useful one. A baseline also catches broken setups: a ten-class digit model that scores about the same as always guessing one digit may not be receiving the inputs you think it is.

Balance the collection without hiding the real job

For your two-class training collection, similar counts make missing coverage easier to notice. Still, fifty nearly identical mug photos and fifty varied glass photos are not equally informative. Count different objects and conditions as well as files.

Evaluation has a different purpose. A development set with equal numbers of each class makes it easy to compare classes. A test meant to show how the model will do in real use should match the mix of photos it will meet there, or explain why it does not. Never quietly rebalance a final test to get a better score.

Accuracy for each class

The overall percentage mixes all classes together. Split it by class and more of the story appears. In The confusion matrix in plain words you measured recall: of the real members of a class, how many did the model find? A class's recall is simply the accuracy measured on that class alone.

In the example, the model found 8 of 10 mugs (80%) and 5 of 10 glasses (50%). The overall 65% hides the fact that it finds glasses no better than tossing a coin. For the dog and cat set, the always-dog program finds 100% of the dogs and 0% of the cats: the whole problem in two numbers.

Always report the counts too. "80% accurate on 10 of each class" means something different from "80% accurate on 90 mugs and 10 glasses".

Build with me · 2

Read the score for each class

Each row of the lab's table holds one actual class. Divide that row's diagonal cell by the row total and you have the score for that class. Put the class scores next to the overall score and the baseline, and you can see at a glance whether one class is being left behind.

Your image lab at this stageWorked example
Text
From the table (one row per actual class):  mug row:   8 correct out of 10 = 80%  glass row: 5 correct out of 10 = 50%Overall:     13 correct out of 20 = 65%

From your own Development table, work out the score for each class. In Notes, write both class scores, the overall score, the baseline, and the number of photos in each class.

What to look for

Two class scores whose correct counts add up to the overall correct count. If one class is much weaker, you can name it.

Make it yours

If one class is weaker, write one hypothesis about why, and which new training photos would test it. That is the starting point for your next run.

Decide what the mistakes cost

For a spam filter, losing a real message and letting spam through have different costs. For a photo-sorting tool, a photo in the wrong folder may be easy to fix. The acceptable trade-off belongs to the job, not to a universal accuracy target.

Some applications use a confidence threshold: accept the model's answer automatically only when its top score is above a chosen level, such as 90%, and send the rest to a person to check. Raising the threshold means fewer automatic decisions and more unsure cases. It does not make the accepted answers correct. Measure both: how accurate the accepted answers are, and how many were sent for review.

Before you call a model good, state the baseline, the score for each class, the mistakes that matter most, and the range of situations you tested. Those details turn a percentage into evidence.

Full reference solution

The whole Level 2 workflow on one page. Each reading covers part of it; this list shows where that part fits. The numbers in the readings are examples, and your own results will differ.

Text
1. Name the two classes and write a one-sentence rule for each.2. Reserve objects or sessions for Development and Final test before collecting.3. In Photos, choose Training and a label, then add varied photos of both classes.4. Add reserved photos under Development and Final test, never under Training.5. In Notes, record label rules, objects, conditions, the split, and a run name.6. In Train, check both counts, keep 20 rounds for the first run, and press Train model.7. Download project and Download model, with matching run names.8. In Test & export, predict Development and read every ID, actual label, prediction, and score.9. Work out accuracy yourself, read the confusion matrix, and compare with the Always- baseline.10. Write one hypothesis. Change only the Training photos (or one stated setting), retrain, and predict the same Development collection.11. Choose a version using Development results. Open its saved model if needed.12. Predict Final test once, after choosing, and report its fraction, conditions, and limits.13. For photos with no known answer, use Unlabelled targets, predict, and Download predictions.csv.14. Keep project, model, notes, and predictions file together under one run name. None of them replaces another.

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in