0%
BuildShip itabout 27 min, 8 steps

Improving a model, honestly

Check the data preparation, look at the mistakes, compare with a baseline, try one change such as shifted copies, and only then use the final test.

Work here, beside the explanation

When a model makes mistakes, the tempting fix is a bigger model. Usually the better first move is to check the simple things: is the data prepared correctly, what do the mistakes look like, is every digit well covered, and how does the model compare with a baseline? Only then try one planned change and measure it fairly. This lesson walks through that order with the closest-example model, tries one change, called data augmentation, and finishes by using the final test once.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Check the data preparation first

A disappointing score can come from a broken pipeline, the chain of steps between the raw data and the model, instead of from the model itself: pixels divided by 16 twice, columns swapped, labels in a different order. Training longer never fixes that. So check the cheap things first: that training and development images have the same number of columns, the same 0 to 1 range, and the same ten classes, and that the model's accuracy looks sensible. A program can run without any error and still be feeding the model the wrong numbers.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)print("Training shape:", X_train.shape)print("Development shape:", X_dev.shape)print("Ranges:", (X_train.min(), X_train.max()), (X_dev.min(), X_dev.max()))print("Classes:", model.classes_)print("Development accuracy:", model.score(X_dev, y_dev))

Run and confirm that training and development data are prepared the same way.

What to look for

Both sets have 64 columns and a range of 0.0 to 1.0, the classes are 0 to 9, and development accuracy is about 0.981.

Make it yours

Name one preparation mistake that these checks would not catch, such as pixels in the wrong order.

Build with me · 2

2. Look at the mistakes

The program finds the development images the model got wrong, as in the confusion matrix lesson, and draws the first two with their true and predicted digits. Looking comes before fixing. Some mistakes are digits a person would hesitate over; some show a stroke style the training images lack. When the model disagrees with a label, the label is not automatically wrong: check the image before deciding. Never change a development or final-test label just because your model disagrees with it.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import matplotlib.pyplot as plt predictions = model.predict(X_dev)wrong = np.where(predictions != y_dev)[0]print("Development mistakes:", len(wrong))for position in wrong[:2]:    plt.figure(figsize=(2.5, 2.5))    plt.imshow(X_dev[position].reshape(8, 8), cmap="gray", vmin=0, vmax=1)    plt.title("True " + str(y_dev[position]) + "; predicted " + str(predictions[position]))    plt.axis("off")    plt.tight_layout()    plt.show()

Run and describe what you can actually see in one mistaken image, before guessing why it happened.

What to look for

In our run the model makes 7 mistakes on development images, and the first two are drawn.

Make it yours

Write one guess about a cause that you could test by adding new training examples, not by changing labels.

Build with me · 3

3. Count each digit

A digit with very few training examples is a common cause of mistakes. np.unique(y_train, return_counts=True) returns two arrays: every distinct label, and how many times each appears. dict(zip(...)) pairs them into a dictionary, label to count, so they print readably; .tolist() turns NumPy's numbers into ordinary ones first. Balanced counts are good, but they do not show variety: 100 nearly identical 8s teach less than 100 different ones.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)train_labels, train_counts = np.unique(y_train, return_counts=True)dev_labels, dev_counts = np.unique(y_dev, return_counts=True)print("Training counts:", dict(zip(train_labels.tolist(), train_counts.tolist())))print("Development counts:", dict(zip(dev_labels.tolist(), dev_counts.tolist())))

Run and compare the counts in the two parts.

What to look for

Every digit has about 104 to 109 training images and 35 to 37 development images, so no digit is short of examples.

Make it yours

Name one kind of variety these counts cannot show, such as how many different people wrote the digits.

Build with me · 4

4. Compare with the baseline again

The always-one-digit baseline from Module 2 shows how much of the score comes from simply guessing the most common label. Both models are scored on the same development images. Because the digits are evenly balanced, the baseline scores only about 10 percent, so the model's 98 percent is clearly earned. On a lopsided dataset, where one label is very common, a high accuracy can be much less impressive than it looks, which is why the habit is worth keeping.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.dummy import DummyClassifier constant = DummyClassifier(strategy="most_frequent")constant.fit(X_train, y_train)print("Constant-label development:", constant.score(X_dev, y_dev))print("KNN development:", model.score(X_dev, y_dev))

Run and compare the two scores.

What to look for

The constant baseline scores about 0.103 and the closest-example model about 0.981.

Make it yours

Describe a task where missing a rare label would matter more than overall accuracy, such as spotting a rare disease.

Build with me · 5

5. Try one change: shifted copies

Data augmentation means making extra training examples by changing existing ones in ways that keep their label. Here every training image is copied and moved one column to the right. X_train.reshape(-1, 8, 8) turns the table back into a stack of 8 by 8 images; -1 means "work out this number yourself", here 1,077. np.zeros_like makes a stack of blank images the same size. The next line copies every column except the last into the stack one place to the right, so the new first column stays blank and the old last column is dropped. np.concatenate puts the originals and the shifted copies into one longer table, and does the same for the labels. Only training images are augmented; development and final-test images stay untouched, so the check stays fair. Doubling the rows does not double the real evidence: each copy is the same handwriting moved slightly.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)# Shift only training images; keep related copies out of development and final data.training_images = X_train.reshape(-1, 8, 8)shifted = np.zeros_like(training_images)shifted[:, :, 1:] = training_images[:, :, :-1]X_augmented = np.concatenate([X_train, shifted.reshape(-1, 64)])y_augmented = np.concatenate([y_train, y_train])print("Original training rows:", len(y_train))print("Augmented training rows:", len(y_augmented))print("Development unchanged:", len(y_dev))print("Final test unchanged:", len(y_final))

Run and check which part grew and which parts stayed the same.

What to look for

Training grows from 1,077 to 2,154 rows; development and final test stay at 360 each.

Make it yours

Explain why rotating or mirroring images could change what some digits mean, for example a 6 and a 9.

Build with me · 6

6. Look at a shifted image

Before trusting a change, look at what it actually did. The picture shows one training image and its shifted copy side by side. On an eight-pixel-wide image a one-column move is large, and the right edge of a stroke can be cut off. The copy keeps the original's label, which is only right if it still clearly shows the same digit.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)# Shift only training images; keep related copies out of development and final data.training_images = X_train.reshape(-1, 8, 8)shifted = np.zeros_like(training_images)shifted[:, :, 1:] = training_images[:, :, :-1]X_augmented = np.concatenate([X_train, shifted.reshape(-1, 64)])y_augmented = np.concatenate([y_train, y_train])import matplotlib.pyplot as plt figure, axes = plt.subplots(1, 2, figsize=(5, 3))axes[0].imshow(training_images[0], cmap="gray", vmin=0, vmax=1)axes[0].set_title("Original; label " + str(y_train[0]))axes[1].imshow(shifted[0], cmap="gray", vmin=0, vmax=1)axes[1].set_title("Training-only shift")for axis in axes:    axis.axis("off")plt.tight_layout()plt.show()

Run and compare the original with its shifted copy.

What to look for

Two pictures appear: a training digit and the same digit moved one column to the right.

Make it yours

Change which training image is shown and check that the shift still looks reasonable.

Build with me · 7

7. Measure the change

A fresh closest-example model stores the augmented training table, and both models are scored on the same, unchanged development images. The result could be better, the same, or worse, and each is useful to know. In our run the change fixes exactly one development image. One image out of 360 is a very small difference, and a different split could erase it.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)# Shift only training images; keep related copies out of development and final data.training_images = X_train.reshape(-1, 8, 8)shifted = np.zeros_like(training_images)shifted[:, :, 1:] = training_images[:, :, :-1]X_augmented = np.concatenate([X_train, shifted.reshape(-1, 64)])y_augmented = np.concatenate([y_train, y_train])augmented_model = KNeighborsClassifier(n_neighbors=3)augmented_model.fit(X_augmented, y_augmented)print("Original development accuracy:", model.score(X_dev, y_dev))print("Augmented development accuracy:", augmented_model.score(X_dev, y_dev))print("Changed only the training collection with the documented shift.")

Run and write the difference as a number of images, not just a percentage.

What to look for

In our run development accuracy goes from about 0.981 (353 of 360) to about 0.983 (354 of 360).

Make it yours

If the score had got worse, what would you check first: the shifted images, or the model?

Build with me · 8

8. Choose, then use the final test once

Now the rule decides: keep the augmented version only if its development accuracy is higher, otherwise keep the original. selected_model = augmented_model if use_augmented else model is the one-line if from the arrays lesson. Only after the choice is made is the chosen model scored on the final test, the images that have been set aside since the start. That final score is your honest estimate for this procedure. If you now go back and change something because of it, the final test has become development data, and a fresh honest estimate would need new images.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)# Shift only training images; keep related copies out of development and final data.training_images = X_train.reshape(-1, 8, 8)shifted = np.zeros_like(training_images)shifted[:, :, 1:] = training_images[:, :, :-1]X_augmented = np.concatenate([X_train, shifted.reshape(-1, 64)])y_augmented = np.concatenate([y_train, y_train])from sklearn.metrics import confusion_matriximport json augmented_model = KNeighborsClassifier(n_neighbors=3)augmented_model.fit(X_augmented, y_augmented)original_dev = float(model.score(X_dev, y_dev))augmented_dev = float(augmented_model.score(X_dev, y_dev))use_augmented = augmented_dev > original_devselected_model = augmented_model if use_augmented else modelselected_name = "shift augmentation" if use_augmented else "original baseline"final_predictions = selected_model.predict(X_final)final_accuracy = float(np.mean(final_predictions == y_final))report = {    "selection_rule": "higher development accuracy; retain baseline on a tie",    "baseline_development": original_dev, "augmented_development": augmented_dev,    "selected": selected_name, "final_count": len(y_final),    "final_accuracy": final_accuracy,}print(json.dumps(report, indent=2))print("Selected model final confusion matrix:")print(confusion_matrix(y_final, final_predictions, labels=np.arange(10)))with open("improvement-report.json", "w") as handle:    json.dump(report, handle, indent=2)

Run once, after you understand the rule. Record the final result as it is, even if it disappoints.

What to look for

The report shows the rule, both development scores, the chosen version, and a final-test accuracy of about 0.986 on 360 images, followed by the final confusion matrix.

Make it yours

Write one limitation of this result and one next experiment, and say what new data that experiment would need.

More data and more experiments still need care

More data helps only if it is good data: a bigger collection can still be lopsided, full of near duplicates, or labelled inconsistently. Each extra experiment you run on the same development images also makes it a little more likely that a winner won by luck. Keep a record of every experiment, including the ones that did not help, and keep your conclusions to what you actually measured.

What runs in this page

Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)# Shift only training images; keep related copies out of development and final data.training_images = X_train.reshape(-1, 8, 8)shifted = np.zeros_like(training_images)shifted[:, :, 1:] = training_images[:, :, :-1]X_augmented = np.concatenate([X_train, shifted.reshape(-1, 64)])y_augmented = np.concatenate([y_train, y_train])from sklearn.metrics import confusion_matriximport json augmented_model = KNeighborsClassifier(n_neighbors=3)augmented_model.fit(X_augmented, y_augmented)original_dev = float(model.score(X_dev, y_dev))augmented_dev = float(augmented_model.score(X_dev, y_dev))use_augmented = augmented_dev > original_devselected_model = augmented_model if use_augmented else modelselected_name = "shift augmentation" if use_augmented else "original baseline"final_predictions = selected_model.predict(X_final)final_accuracy = float(np.mean(final_predictions == y_final))report = {    "selection_rule": "higher development accuracy; retain baseline on a tie",    "baseline_development": original_dev, "augmented_development": augmented_dev,    "selected": selected_name, "final_count": len(y_final),    "final_accuracy": final_accuracy,}print(json.dumps(report, indent=2))print("Selected model final confusion matrix:")print(confusion_matrix(y_final, final_predictions, labels=np.arange(10)))with open("improvement-report.json", "w") as handle:    json.dump(report, handle, indent=2)

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in