0%
BuildTrain a classifier in the browserabout 28 min, 8 steps

Reading a confusion matrix properly

Build the ten-by-ten table of true digits against predicted digits, read single cells and rows, and look at the images the model got wrong.

Work here, beside the explanation

In Level 2 you built a confusion matrix by hand for two labels, mug and glass: a two-by-two table with the true label down the side and the prediction across the top. With ten digits the table becomes ten by ten, but you read it exactly the same way. An accuracy of 0.98 tells you how many answers were right; the matrix tells you which digits were mistaken for which. You will build it from your model's real development predictions, read single cells, and then look at the actual images the model got wrong.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Create predictions with their matching labels

The program starts with the same loading, splitting, and fitting lines as the previous lesson, so it works on its own after a fresh page load. Then it predicts every development image. zip pairs each true label with the prediction in the same position, as in the pairs lesson, and list(...) turns the pairs into a list you can print. Each pair is written true label first, prediction second. The two arrays must stay in the same order, because position is the only thing connecting a prediction to its image.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)predictions = model.predict(X_dev)print("Predictions:", len(predictions))print("True labels:", len(y_dev))print("First five pairs:", list(zip(y_dev[:5], predictions[:5])))

Run and read the first five pairs aloud as "true 9, predicted 9" and so on.

What to look for

Both counts equal 360. Each pair refers to one development image.

Make it yours

Choose another five-row slice and keep both arrays on identical bounds.

Build with me · 2

2. Count all true-to-predicted pairs

confusion_matrix(y_dev, predictions, labels=np.arange(10)) counts every pair into a ten-by-ten table. Row number 3 is for images that truly show a 3; column number 7 is for images the model called 7. So the cell in row 3, column 7 counts true 3s that were predicted as 7. np.arange(10) makes the numbers 0 to 9, which fixes the rows and columns in digit order. The diagonal runs from the top left corner to the bottom right; those cells have the same row and column, so they count correct answers. Every other cell counts one particular kind of mistake. Adding up every cell gives the number of images checked.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix predictions = model.predict(X_dev)matrix = confusion_matrix(y_dev, predictions, labels=np.arange(10))print(matrix)print("Shape:", matrix.shape)print("Counted examples:", matrix.sum())

Run, then find the diagonal from the top left to the bottom right. Look for any non-zero numbers off it.

What to look for

The matrix is ten by ten, and its cells add up to 360. In our run almost every count sits on the diagonal, with a handful of 1s elsewhere.

Make it yours

Pick one off-diagonal cell and translate its row and column into a complete sentence.

Build with me · 3

3. Read one cell and one row

To read one cell, give the row first and the column second: matrix[8, 3] counts true 8s that the model called 3. It does not count true 3s called 8; that is the cell matrix[3, 8], a mistake in the other direction. matrix[8] is the whole of row 8, and .sum() adds it up to give the number of true 8s, whatever the model said. matrix[8, 8] is the diagonal cell in that row: the 8s the model got right.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix predictions = model.predict(X_dev)matrix = confusion_matrix(y_dev, predictions, labels=np.arange(10))true_digit = 8predicted_digit = 3print("True", true_digit, "predicted", predicted_digit, ":", matrix[true_digit, predicted_digit])print("All true", true_digit, "examples:", matrix[true_digit].sum())print("Correct true", true_digit, "examples:", matrix[true_digit, true_digit])

Run and check the selected cell against the printed full matrix from the previous stage.

What to look for

In our run no 8 was called a 3, so that cell is 0. There are 35 true 8s, and 33 of them were recognised.

Make it yours

Choose two other digits and compare both directions without assuming the counts are equal.

Build with me · 4

4. Recover accuracy from the diagonal

np.trace(matrix) adds up the diagonal, which is every correct answer, each counted once. matrix.sum() adds up every cell, which is every image checked. Correct divided by total is accuracy, so the table and the score from the previous lesson must agree. This is the same arithmetic you did for mug and glass in Level 2, just with ten diagonal cells instead of two.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix predictions = model.predict(X_dev)matrix = confusion_matrix(y_dev, predictions, labels=np.arange(10))correct = int(np.trace(matrix))total = int(matrix.sum())print("Correct:", correct, "of", total)print("Accuracy:", round(correct / total, 4))print("Agrees with model score:", np.isclose(correct / total, model.score(X_dev, y_dev)))

Run and compare the count with the one from your first complete classifier.

What to look for

Agrees with model score is True.

Make it yours

Print the number of errors as total - correct and identify where those errors live in the matrix.

Build with me · 5

5. Measure recall for each true class

Recall, from Level 2, asks: of the images that truly show this digit, how many did the model find? For each digit it is the diagonal cell divided by the row total. The line recall = recalled / actual if actual else 0.0 is a one-line if: it divides when actual is not zero, and otherwise uses 0.0, because dividing by zero would stop the program. Always read the two counts as well as the ratio: 33 of 35 is a small number of images, and one more mistake would move the ratio noticeably.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix predictions = model.predict(X_dev)matrix = confusion_matrix(y_dev, predictions, labels=np.arange(10))for digit in range(10):    actual = int(matrix[digit].sum())    recalled = int(matrix[digit, digit])    recall = recalled / actual if actual else 0.0    print("Digit", digit, "correct", recalled, "of", actual, "recall", round(recall, 3))

Run and find the digit with the lowest recall. Say it as a sentence with both counts, such as "33 of the 35 true 8s were found".

What to look for

Ten lines print, one per digit. In our run digit 8 has the lowest recall, 33 of 35.

Make it yours

Change the display precision and observe that rounding changes presentation rather than the underlying count.

Build with me · 6

6. Find a frequent off-diagonal error

To find the most common mistake, the program makes a copy of the matrix and sets its diagonal to zero with np.fill_diagonal, so only mistakes are left. Working on a copy keeps the original matrix intact. np.argmax finds the position of the largest remaining number, but it counts positions as if the table were one long row. np.unravel_index converts that single position back into a row and a column, which are unpacked into true_digit and predicted_digit. If several cells tie for the largest count, this reports only the first one it finds.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix predictions = model.predict(X_dev)matrix = confusion_matrix(y_dev, predictions, labels=np.arange(10))errors = matrix.copy()np.fill_diagonal(errors, 0)if errors.max() == 0:    print("No development mistakes in this run.")else:    true_digit, predicted_digit = np.unravel_index(np.argmax(errors), errors.shape)    print("One most frequent confusion: true", true_digit, "predicted", predicted_digit)    print("Count:", errors[true_digit, predicted_digit])

Run and translate the reported direction into a sentence.

What to look for

In our run every mistake happened only once, so the largest count is 1 and the program reports the first such cell, a true 2 predicted as 7.

Make it yours

Look at the full matrix from stage 2 and list every other cell that also has the largest count, so you do not treat one of several ties as special.

Build with me · 7

7. Inspect actual mistaken inputs

Counts tell you where mistakes are; pictures help you understand them. predictions != y_dev is True wherever the model was wrong, and np.where(...)[0] turns that into a list of positions, as in the arrays lesson. The loop takes the first three positions and draws each image with its true and predicted digit in the title. Some mistakes are digits a person would hesitate over too; others may show an unusual stroke. A mistake does not prove the label is wrong: look before you decide.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix predictions = model.predict(X_dev)matrix = confusion_matrix(y_dev, predictions, labels=np.arange(10))import matplotlib.pyplot as plt wrong = np.where(predictions != y_dev)[0]print("Mistake count:", len(wrong))for position in wrong[:3]:    print("Development row", position, "true", y_dev[position], "predicted", predictions[position])    plt.figure(figsize=(2.5, 2.5))    plt.imshow(X_dev[position].reshape(8, 8), cmap="gray", vmin=0, vmax=1)    plt.title("True " + str(y_dev[position]) + ", predicted " + str(predictions[position]))    plt.axis("off")    plt.tight_layout()    plt.show()

Run and describe what you can directly observe in one error image. Separate observation from your hypothesis about its cause.

What to look for

Three real mistakes appear. In our run the first is development row 34, a 4 that was predicted as 7.

Make it yours

Increase the display limit modestly if you need more evidence; keep the final test untouched.

Build with me · 8

8. Add a labelled visual matrix

ConfusionMatrixDisplay draws the same counts as a coloured grid, with the axis labels written on, so a reader cannot mix up rows and columns. cmap="Blues" picks the colours, darker for bigger counts; values_format="d" prints each count as a whole number. The title says these are development predictions. A picture is good for spotting patterns quickly; keep the exact counts too, because they are what you check claims against.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix predictions = model.predict(X_dev)matrix = confusion_matrix(y_dev, predictions, labels=np.arange(10))import matplotlib.pyplot as pltfrom sklearn.metrics import ConfusionMatrixDisplay print("Correct:", int(np.trace(matrix)), "of", int(matrix.sum()))print("Development accuracy:", round(model.score(X_dev, y_dev), 4))display = ConfusionMatrixDisplay(confusion_matrix=matrix, display_labels=np.arange(10))display.plot(cmap="Blues", colorbar=False, values_format="d")plt.title("Development predictions; rows are true labels")plt.tight_layout()plt.show()

Build the complete display and explain a diagonal and an off-diagonal entry using the axis labels.

What to look for

A labelled ten-by-ten matrix appears in Output, alongside its measured overall accuracy.

Make it yours

Write two sentences: one numerical observation and one possible next development experiment supported by it.

Read counts before stories

A confusion matrix puts true digits in rows and predicted digits in columns. The diagonal counts correct answers; every other cell counts one kind of mistake, in one direction. All the cells add up to the number of images checked, and the diagonal divided by that total is the accuracy. A row total is what you divide by for recall (how many of the true 8s were found). A column total is what you divide by for precision (how many of the model's "8" answers were really 8s), exactly as in Level 2. Look at the actual images behind a mistake before claiming why it happened.

The exercise that follows uses its own fixed starter and split, so its counts differ a little from this lesson's.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix predictions = model.predict(X_dev)matrix = confusion_matrix(y_dev, predictions, labels=np.arange(10))import matplotlib.pyplot as pltfrom sklearn.metrics import ConfusionMatrixDisplay print("Correct:", int(np.trace(matrix)), "of", int(matrix.sum()))print("Development accuracy:", round(model.score(X_dev, y_dev), 4))display = ConfusionMatrixDisplay(confusion_matrix=matrix, display_labels=np.arange(10))display.plot(cmap="Blues", colorbar=False, values_format="d")plt.title("Development predictions; rows are true labels")plt.tight_layout()plt.show()

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in