0%
BuildTrain a classifier in the browserabout 28 min, 8 steps

Build a recognizer for your own digit

Draw a digit as eight rows of symbols, turn it into the 64 numbers the model expects, and ask the closest-example model what it is.

Work here, beside the explanation

So far the model has only seen digits from the dataset. Now you give it one you made yourself. You will draw a digit as eight rows of symbols, turn the symbols into the 64 numbers the model expects, look at exactly what the model will receive, and ask the closest-example model what it thinks. Decide which digit you mean before you ask: the model is allowed to be wrong, and the only way to know is to have written your answer down first.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Know what the model expects

A model cannot tell whether numbers were prepared sensibly; it just uses whatever it is given. So before drawing, pin down exactly what this model was trained on. People call this the model's input contract: 64 brightness values for an eight-by-eight image, listed row by row from the top left, with the digit bright on a dark background, on a scale from 0 to 1. model.n_features_in_ confirms that the model expects 64 values, and model.classes_ lists the ten answers it can give. Neither can check what the values mean: an upside-down drawing with 64 values would still be accepted.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)print("Model input columns:", model.n_features_in_)print("Class mapping:", model.classes_)print("Expected image: 8 rows, 8 columns, bright strokes on a dark background")print("Prepared brightness: 0 to 1, flattened row by row")

Run and read the four lines. They are the rules your drawing must follow.

What to look for

The model expects 64 input columns and knows the classes 0 to 9.

Make it yours

Decide which digit you will draw, and write it down now.

Build with me · 2

2. Draw the digit as text

Each string in the list is one row of the image, from top to bottom. A dot means dark background, a hash means a bright stroke, and a plus sign means half bright. There must be eight rows of exactly eight characters. Drawing with text means every pixel is visible and you need no camera or painting program. The starter drawing is a seven, drawn with strokes two pixels wide like the sevens in the dataset.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train) rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]for row in rows:    print(row)

Run to print the grid. Then edit the rows into your own digit, keeping every row eight characters long.

What to look for

Eight rows print exactly as you typed them.

Make it yours

Draw a different digit, and put the digit you meant in a comment above the rows, such as # I meant a 4.

Build with me · 3

3. Check the drawing before using it

Mistakes in the drawing should stop the program with a clear message, not reach the model silently. The first check uses any, from the pairs lesson: if the list does not have 8 rows, or any row is not 8 characters long, the program raises an error with a message saying what to fix. The second check goes through every character of every row and stops if it finds a symbol other than a dot, plus, or hash. These checks cannot tell whether your drawing looks like the digit you meant; they only catch broken input.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train) rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only ., +, and #.")print("Drawing dimensions and symbols are valid.")

Run with a correct grid, then delete one character from a row and run again to see the message. Put the character back afterwards.

What to look for

A correct grid prints that the drawing is valid; a broken row stops the program with the message about eight rows of eight characters.

Make it yours

Type a letter such as x into one row and confirm that it is rejected rather than quietly treated as some brightness.

Build with me · 4

4. Turn the symbols into numbers

The dictionary brightness says what each symbol is worth: 0.0, 0.5, or 1.0. The next line is a list comprehension inside another one, as the pairs lesson promised: for each row, it makes a list of the brightness of each symbol, and it does that for every row, giving eight lists of eight numbers. np.array(...) turns that into an eight-by-eight array. The values are already on the 0 to 1 scale, so they must not be divided by 16 again. The blank check stops an empty drawing. Finally reshape(1, 64) makes a batch of one image with 64 values, the shape the model expects.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train) rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)print("Numerical image:")print(image)print("Image shape:", image.shape)print("Input batch:", one_input.shape)print("Range:", one_input.min(), one_input.max())

Run and match a few symbols in your drawing with their numbers in the printed grid.

What to look for

The image has shape (8, 8), the model input has shape (1, 64), and the values run from 0.0 to 1.0.

Make it yours

Replace a few # symbols at the edge of a stroke with + and find the 0.5 values they create.

Build with me · 5

5. Look at what the model will see

Before asking for a prediction, draw the exact numbers the model will receive, not the text version. Setting vmin=0 and vmax=1 fixes the grey scale so that 0 is black and 1 is white, the same as for the dataset images you plotted in the arrays lesson. Eight by eight pixels is very coarse, so a shape that is clear as text can look blocky or ambiguous here. Nothing is hidden: the grid you drew is already the size the model uses.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train) rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)import matplotlib.pyplot as plt plt.figure(figsize=(3, 3))plt.imshow(one_input.reshape(8, 8), cmap="gray", vmin=0, vmax=1)plt.title("My prepared input")plt.axis("off")plt.tight_layout()plt.show()

Run and compare the picture with the digit you meant.

What to look for

An eight-by-eight picture of your drawing appears.

Make it yours

Move your digit so it sits in the middle of the grid, as the dataset digits do, and look again. That changes the input, not the model.

Build with me · 6

6. Ask the model

model.predict_proba(one_input) asks the closest-example model for its answer as shares of the vote, one number per digit. With three neighbours voting, each share is 0, one third, two thirds, or 1. [0] takes the answers for the first, and only, image in the batch. np.argsort lists positions from the smallest share to the largest, and [::-1] reverses the list so the biggest share comes first. model.classes_ turns that position into the digit it stands for. The last line is a reminder that only you know which digit you meant.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train) rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)probabilities = model.predict_proba(one_input)[0]order = np.argsort(probabilities)[::-1]prediction = int(model.classes_[order[0]])print("Predicted digit:", prediction)print("True intended digit: a choice you must record yourself")

Run and compare the predicted digit with the one you wrote down.

What to look for

For the starter drawing the model predicts 7.

Make it yours

Try this one-pixel-thin seven instead: "..####..", ".....#..", "....#...", "...#....", "..#.....", "..#.....", "..#.....", "........". In our run the model calls it an 8. The dataset's sevens have thick strokes, so a thin seven is closer to some stored eights than to any stored seven.

Build with me · 7

7. See how split the vote was

The top three digits and their vote shares show whether the three neighbours agreed. A share of 1.0 means all three voted the same way; two thirds means one neighbour disagreed. The threshold rule prints a different message when the winning share is below 0.8. That rule only changes what the program says; it does not change the model or its answer, and a unanimous vote can still be wrong on a drawing unlike the training images. The model also has no way to answer none of these: a scribble or a letter still gets a digit.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train) rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)probabilities = model.predict_proba(one_input)[0]order = np.argsort(probabilities)[::-1]for position in order[:3]:    print("Digit", int(model.classes_[position]), "vote support", round(float(probabilities[position]), 3))threshold = 0.8if probabilities[order[0]] < threshold:    print("Low support under our chosen display rule; inspect the input.")else:    print("Above our chosen display threshold; correctness is still not guaranteed.")

Run and read both the vote shares and the message.

What to look for

For the starter seven, 7 has a vote share of 1.0 and the message says it is above the chosen threshold.

Make it yours

Change the threshold, and explain which output changes and which does not.

Build with me · 8

8. Save a record of the drawing

A record lets you reproduce a surprising result later instead of relying on memory. The dictionary stores the exact rows, the digit you meant, the prediction, and a description of the model and the preparation. class_scores is a dictionary built in one line, like a list comprehension with curly brackets: for each digit and its vote share, store the share under the digit's name. json writes it all to my-digit.json, as in the pairs lesson, and the picture shows both what you meant and what the model said. Set intended_digit before running.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train) rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)import jsonimport matplotlib.pyplot as plt intended_digit = 7probabilities = model.predict_proba(one_input)[0]order = np.argsort(probabilities)[::-1]prediction = int(model.classes_[order[0]])record = {    "rows": rows, "intended_digit": intended_digit, "prediction": prediction,    "model": "KNN k=3", "dataset": "sklearn digits 8x8",    "preparation": "symbols directly to 0..1, row-major 64 features",    "purpose": "development probe",    "class_scores": {str(int(label)): float(score) for label, score in zip(model.classes_, probabilities)},}print(json.dumps(record, indent=2))with open("my-digit.json", "w") as handle:    json.dump(record, handle, indent=2)plt.figure(figsize=(3, 3))plt.imshow(image, cmap="gray", vmin=0, vmax=1)plt.title("Intended " + str(intended_digit) + "; predicted " + str(prediction))plt.axis("off")plt.tight_layout()plt.show()

Set intended_digit to the digit you drew, run, and download my-digit.json if you want to keep it.

What to look for

A my-digit.json file is offered for download, and the picture title shows the digit you meant and the predicted digit.

Make it yours

Draw several different digits, save a record for each, and keep the ones the model got wrong as well as the ones it got right.

From a photo to the same input

A photograph or a large drawing needs more preparation before this model could use it. It would need to be turned grey, made bright on a dark background, cropped to the digit, shrunk to eight by eight pixels, and scaled to 0 to 1. Each of those steps changes the numbers, so look at the final eight-by-eight grid before trusting the answer.

Your own drawings can differ from the dataset in stroke thickness, position, or style, as the thin seven showed. A high score on the dataset does not promise the same on every kind of handwriting. In Module 3 you will give your drawings to the neural network as well and compare the two models.

What runs in this page

Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train) rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)import jsonimport matplotlib.pyplot as plt intended_digit = 7probabilities = model.predict_proba(one_input)[0]order = np.argsort(probabilities)[::-1]prediction = int(model.classes_[order[0]])record = {    "rows": rows, "intended_digit": intended_digit, "prediction": prediction,    "model": "KNN k=3", "dataset": "sklearn digits 8x8",    "preparation": "symbols directly to 0..1, row-major 64 features",    "purpose": "development probe",    "class_scores": {str(int(label)): float(score) for label, score in zip(model.classes_, probabilities)},}print(json.dumps(record, indent=2))with open("my-digit.json", "w") as handle:    json.dump(record, handle, indent=2)plt.figure(figsize=(3, 3))plt.imshow(image, cmap="gray", vmin=0, vmax=1)plt.title("Intended " + str(intended_digit) + "; predicted " + str(prediction))plt.axis("off")plt.tight_layout()plt.show()

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in