Your first complete digit classifier
Load all 1,797 digits, keep some aside, train the closest-example model, and count how often it is right, one step at a time.
Work here, beside the explanation
In the previous lesson you worked out by hand how the closest-example model decides. Now you use it for real: load all 1,797 digit images, keep some of them aside for checking, give the rest to the model, and count how often it names the right digit. These are the same steps as the make-a-model, train, and predict blocks you used in Level 3, written in Python. Every stage adds one step to the same program.
Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.
Build with me · 1
1. Load inputs and their labels
These lines are the same loading lines as in the arrays lesson. X is a table with one row per image and one column per pixel: 1,797 rows of 64 brightness values. y holds the matching correct digit for each row, so y[0] is the answer for the image in row X[0]. Using a capital X for the table of inputs and a small y for the answers is a naming habit that machine learning programmers share; Python does not require it. Dividing by 16 puts every brightness on a 0 to 1 scale. Nothing has been learned yet: this only loads the examples.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetprint("Input shape:", X.shape)print("Label shape:", y.shape)print("Brightness range:", X.min(), X.max())Run and check that X has as many rows as y has labels.
What to look for
X has shape (1797, 64), y has shape (1797,), and brightness ranges from 0.0 to 1.0.
Make it yours
Print y[5] and describe which row of X it belongs to.
Build with me · 2
2. Reserve the final test first
In Level 2 you kept photos aside so you could test the model honestly on examples it never learned from. train_test_split does that job in one line. It shuffles the examples and divides the inputs and their labels together, so every image keeps its own answer. It returns four things, which the left side unpacks into four names, as in the pairs lesson: the inputs you keep working with, the inputs set aside, then their labels in the same order. The settings in brackets each have one plain job:
test_size=0.2sets aside 20 percent, one fifth, as the final test. You will not look at the model's answers on these until your choices are finished.random_state=42fixes the shuffle, so every run splits the examples in the same way and your numbers can be compared. The value 42 has no special meaning; any fixed number would do.stratify=ykeeps the same mix of digits in both parts, so the final test is not accidentally short of 7s.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)print("Development pool:", len(y_pool))print("Reserved final test:", len(y_final))Run and check that the two counts add up to 1,797.
What to look for
The pool has 1,437 examples and the final test has 360.
Make it yours
Change test_size to 0.1 and predict the new counts, then change it back to 0.2 so your numbers match the rest of the lesson.
Build with me · 3
3. Split the pool into training and development
A second split divides the pool into two more parts, giving three in total, the same three roles as in Level 2:
- Training (
X_train,y_train): the examples the model learns from. For this model, the examples it stores. - Development (
X_dev,y_dev): examples you use again and again to check the model and compare choices. - Final test (
X_final,y_final): kept aside until the end, used once to report how the finished model does.
test_size=0.25 takes a quarter of the pool for development. A quarter of the 80 percent pool is 20 percent of all the images, which leaves 60 percent for training. The comment line starting with # is a note for people reading the program; Python ignores it.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)print("Training:", len(y_train))print("Development:", len(y_dev))print("Final test:", len(y_final))print("Total:", len(y_train) + len(y_dev) + len(y_final))Run and write down the three counts with the job of each part.
What to look for
There are 1,077 training examples, 360 development examples, and 360 final-test examples, 1,797 in total.
Make it yours
Explain in your own words why looking at final-test results while you are still choosing settings would make the final test less honest.
Build with me · 4
4. Create the model object
This is the same line you used on the four-pixel images. KNeighborsClassifier is the closest-example model, and n_neighbors=3 sets k to 3, so the three closest stored images will vote on each new one. Creating the model is like dragging a make-a-model block into place: it exists, it knows its settings, but it has seen no examples yet. Asking it to predict now would give an error saying the model is not fitted.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)print(model)Run and read the printed description.
What to look for
The output is KNeighborsClassifier(n_neighbors=3).
Make it yours
Change the 3 to 1 or 5 if you are curious, but set it back to 3 so the rest of the lesson matches.
Build with me · 5
5. Fit only the training portion
model.fit(X_train, y_train) hands the model the training images with their labels. As you saw by hand, fitting this model just means storing the examples so they can be compared later; there is nothing to tune. Only training data goes in: the development and final-test images stay out of the model's store, which is what makes checking on them fair. After fitting, model.classes_ lists every label the model has seen. In this library, names that end with an underscore, like classes_, only exist after fit has run.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)print("Fitted on:", len(y_train), "labelled examples")print("Known classes:", model.classes_)Run and check the list of known classes.
What to look for
The model reports 1,077 training examples and classes 0 to 9.
Make it yours
Explain why creating the model and fitting it are two separate lines, using the make-a-model and train blocks as a comparison.
Build with me · 6
6. Predict a small unseen batch
predict receives images only, never their answers. For each of the first five development images it finds the three closest training images and returns the winning vote. X_dev[:5] is a slice, a batch of five rows, and the answer comes back as five labels in the same order. The true labels are printed underneath only so you can compare; the model never saw them.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)predictions = model.predict(X_dev[:5])print("Predictions:", predictions)print("True labels:", y_dev[:5])Line up each prediction with the true label in the same position.
What to look for
Five predictions and five true labels print. For these five images they are the same: 9, 4, 2, 8, 1.
Make it yours
Look at X_dev[5:10] and y_dev[5:10] instead, keeping both slices the same.
Build with me · 7
7. Measure every development prediction
Now predict all 360 development images and count the correct ones. predictions == y_dev compares the two arrays position by position and gives True where they match and False where they do not. np.sum adds those up, counting each True as 1, so the result is the number correct. int(...) turns NumPy's number into an ordinary whole number for printing. Accuracy is correct divided by total, exactly as in Level 2. The library can also do this in one call, model.score(X_dev, y_dev); the last line checks that both give the same answer. np.isclose means "equal, allowing for tiny rounding differences".
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)predictions = model.predict(X_dev)correct = int(np.sum(predictions == y_dev))accuracy = correct / len(y_dev)print("Correct:", correct, "of", len(y_dev))print("Development accuracy:", round(accuracy, 4))print("Score agrees:", np.isclose(accuracy, model.score(X_dev, y_dev)))Run and write down the count and the accuracy.
What to look for
In our run the model got 353 of 360 development images right, an accuracy of about 0.98. Score agrees is True.
Make it yours
Add a line that prints the accuracy as a percentage, using round(100 * accuracy, 1).
Build with me · 8
8. Compare with a simple baseline
A score means more next to a simple comparison, the baseline from Level 2. DummyClassifier(strategy="most_frequent") is a model that never looks at the pixels: during fitting it finds the most common label in the training data, and afterwards it predicts that same label for every image. Because the ten digits appear about equally often, always guessing one digit is right only about one time in ten. The closest-example model's score is far above that, which shows it really is using the pixels. The final test is still untouched, because this lesson is for learning and comparing.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.dummy import DummyClassifier baseline = DummyClassifier(strategy="most_frequent")baseline.fit(X_train, y_train)predictions = model.predict(X_dev)correct = int(np.sum(predictions == y_dev))print("KNN correct:", correct, "of", len(y_dev))print("KNN development accuracy:", round(model.score(X_dev, y_dev), 4))print("Constant-label development accuracy:", round(baseline.score(X_dev, y_dev), 4))print("Final-test examples still reserved:", len(y_final))Run and compare the two accuracies.
What to look for
KNN scores about 0.98 on development; the constant baseline scores about 0.10.
Make it yours
Explain why a baseline of about 0.10 is what you would expect for ten digits that appear about equally often.
Check what each object means
The dataset gives you pairs: a row of 64 pixels and its correct digit. The split keeps those pairs together while dividing them into training, development, and final-test parts. Creating the model sets its settings. Fitting gives it the training examples. Predicting uses it on images it did not store. Scoring compares its predictions with the correct labels.
If you see an error that the model is "not fitted", check that fit ran on the same model you are asking to predict. If an error says it expected a two-dimensional array, you probably passed one image on its own, such as X_dev[0]; wrap it as a batch of one with X_dev[:1] instead.
The exercise that follows uses its own fixed starter and split so that everyone's answers can be checked the same way. Use its settings when you complete it; its counts differ slightly from this lesson's.
What runs in this page
Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.
The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.
Full reference solution
This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.dummy import DummyClassifier baseline = DummyClassifier(strategy="most_frequent")baseline.fit(X_train, y_train)predictions = model.predict(X_dev)correct = int(np.sum(predictions == y_dev))print("KNN correct:", correct, "of", len(y_dev))print("KNN development accuracy:", round(model.score(X_dev, y_dev), 4))print("Constant-label development accuracy:", round(baseline.score(X_dev, y_dev), 4))print("Final-test examples still reserved:", len(y_final))Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.