0%
BuildFind out where it breaksabout 14 min, 2 steps

The confusion matrix in plain words

Turn individual predictions into a table that reveals which classes are being confused.

Build the table from individual results

A confusion matrix is a table that counts each combination of actual label and predicted label. It is a summary of your list of results, not another model. In this course, each row is an actual class and each column is a predicted class. Other tools sometimes swap them, so always read the labels on the table first.

Start with four empty cells. Go through your results one by one. A real mug predicted as mug adds one to the mug row, mug column. A real glass predicted as mug adds one to the glass row, mug column. Keep going until every result has been counted exactly once.

The example run from Test it honestly becomes:

Actual ↓ / Predicted →mugglassTotal
mug8210
glass5510
Total13720

Read across one row at a time. Of 10 real mugs, 8 were called mug and 2 were called glass. Of 10 real glasses, 5 were called mug and 5 were called glass. Now read down the mug column instead: the model answered mug 13 times, and 8 of those answers were right.

Find the correct answers and the mistakes

The diagonal runs from the top-left cell to the bottom-right cell. Those cells pair each class with itself, so they hold the correct answers: 8 + 5 = 13. Divided by all 20 photos, that is the 65% accuracy you calculated before.

The other two cells are the mistakes: 2 mugs called glass, and 5 glasses called mug. Start investigating with the bigger group. The table tells you which photos to look at; it does not tell you why they went wrong.

Check the total. If your table adds up to 19 but you evaluated 20 photos, find the missing one before working out any percentage.

Build with me · 1

Find your own table in the lab

Test & export draws this table for you, under the Accuracy line, with the caption "Actual rows; predicted columns". It is built from the same list of results you scroll through below it, so you can check every cell by hand. It is counted again from scratch each time you predict a collection.

Your image lab at this stageWorked example
Text
Accuracy: 13/20 (65.0%)Actual rows; predicted columns  Actual    mug   glass  mug         8       2  glass       5       5

Predict your Development collection. Choose one cell of the table, then find each result it counts in the list below it: for the glass row, mug column, every result whose actual label is glass and whose prediction is mug. Then add all four cells and check that they equal the number of photos evaluated.

What to look for

The diagonal cells add up to the number correct. The other two cells add up to the number wrong. Together they count every result exactly once.

Make it yours

Find the cell with the most mistakes in your own table and look at two of its photos. Write in Notes one thing they have in common.

Give the two kinds of mistake names

For a two-class task, choose one class to be the "yes" answer and call it positive. The other class is negative. Here, let mug be positive. Positive does not mean good or desirable; it is just the class you chose to ask about.

  • True positive: a mug predicted as mug. There are 8.
  • False negative: a mug predicted as glass. There are 2. The model missed a mug.
  • False positive: a glass predicted as mug. There are 5. The model raised a false mug alarm.
  • True negative: a glass predicted as glass. There are 5.

If you made glass the positive class instead, the names would swap but no prediction would change. Always say which class you chose.

Two questions the table can answer

Accuracy asks one question about all the photos at once. The table lets you ask two sharper questions about one class. Keep mug as the positive class and follow each question step by step.

Of the real mugs, how many did the model find? Read across the mug row. There are 8 + 2 = 10 real mugs, and the model called 8 of them mug. It found 8 out of 10: 8 ÷ 10 = 0.8, or 80%. This measure is called recall. Its denominator is the number of real mugs, the row total.

When the model answered mug, how often was it right? Read down the mug column. The model answered mug 8 + 5 = 13 times, and 8 of those were real mugs. That is 8 out of 13: 8 ÷ 13 ≈ 0.615, or about 61.5%. This measure is called precision. Its denominator is the number of mug answers, the column total.

Both fractions have the same 8 on top, but they divide by different numbers, so they answer different questions. This model finds most mugs (recall 80%), but when it says mug it is wrong more than a third of the time (precision about 61.5%), because it also called five glasses mug. If you forget which is which, say the question in words first, then find the row or column that holds its answer.

Build with me · 2

Work out precision and recall for your own model

Now do the same with your own numbers. The lab does not calculate precision or recall for you, so read them off your table: recall from the row of your positive class, precision from its column. If a denominator is zero (for example, the model never answered mug at all), that measure cannot be calculated for this set. Write "not defined here" rather than inventing a number.

Your image lab at this stageWorked example
Text
Mug as the positive class (example numbers)  recall    = mugs found / real mugs            = 8 / 10  = 80%  precision = correct mug answers / mug answers = 8 / 13, about 61.5%

In Notes, write which class you chose as positive. Write one sentence describing a false positive and one describing a false negative in your project. Then calculate recall and precision for that class from your own Development table.

What to look for

Your recall divides by the real members of the class, a row total. Your precision divides by the model's answers of that class, a column total.

Make it yours

Decide which mistake would be more annoying if your model sorted real photos for someone: a glass filed as mug, or a mug filed as glass. That is a decision about consequences. No score can make it for you.

Why the difference matters

Think of a spam filter, with spam as the positive class. A false positive sends a real message to the spam folder, where you might never see it. A false negative lets spam into your inbox. Each lowers accuracy by the same amount, but they can matter very differently.

The table also shows when a model ignores a class altogether. If a spam filter's "predicted spam" column is all zeros, it has never caught a single spam message, even though its accuracy can look high when most messages are not spam. Class balance, and what accuracy hides returns to this.

For the mug project, a mistake may be harmless. Elsewhere, one kind of mistake can matter much more than the other, and a single overall score cannot tell you which. With ten digit classes in Level 4, the same idea becomes a ten-by-ten table: the cell in row 4, column 9 counts real fours predicted as nines. You read it the same way.

Check your understanding

Suppose one of the wrongly predicted glasses becomes correct. The glass row, mug column drops from 5 to 4, and the glass row, glass column rises from 5 to 6. The glass row still holds 10 photos. Accuracy becomes 14 out of 20, or 70%. Mug precision becomes 8 out of 12, about 66.7%, because the model now answers mug 12 times instead of 13. Mug recall stays at 8 out of 10, because no real mug changed. Work through each number yourself before moving on.

Full reference solution

The whole Level 2 workflow on one page. Each reading covers part of it; this list shows where that part fits. The numbers in the readings are examples, and your own results will differ.

Text
1. Name the two classes and write a one-sentence rule for each.2. Reserve objects or sessions for Development and Final test before collecting.3. In Photos, choose Training and a label, then add varied photos of both classes.4. Add reserved photos under Development and Final test, never under Training.5. In Notes, record label rules, objects, conditions, the split, and a run name.6. In Train, check both counts, keep 20 rounds for the first run, and press Train model.7. Download project and Download model, with matching run names.8. In Test & export, predict Development and read every ID, actual label, prediction, and score.9. Work out accuracy yourself, read the confusion matrix, and compare with the Always- baseline.10. Write one hypothesis. Change only the Training photos (or one stated setting), retrain, and predict the same Development collection.11. Choose a version using Development results. Open its saved model if needed.12. Predict Final test once, after choosing, and report its fraction, conditions, and limits.13. For photos with no known answer, use Unlabelled targets, predict, and Download predictions.csv.14. Keep project, model, notes, and predictions file together under one run name. None of them replaces another.

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in