0%
BuildJust enough Python, on real digitsabout 29 min, 9 steps

Arrays, shapes, and pixels

Inspect the real digit dataset and learn the NumPy operations every following model example uses.

Work here, beside the explanation

A model receives numbers, not the word 'picture'. The digit dataset represents every image as an eight-by-eight array of brightness values, plus a label saying which digit was written. Inspect the same example in several shapes so that a later reshape, scale, or prediction is an operation you understand rather than a line copied by habit.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Read the dataset's three views

An array's shape says how many values it holds along each direction. The dataset keeps the same images in three forms. digits.images has shape (1797, 8, 8): 1,797 images, each 8 rows of 8 pixels. digits.data has shape (1797, 64): the same images, each laid out as one row of 64 values, which is called flattened. y has shape (1797,): one label per image. All three are in the same order, so image number 5 in one form is image number 5 in the others.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetprint("Images:", digits.images.shape)print("Flat inputs:", digits.data.shape)print("Labels:", y.shape)

Run and say what each number in each shape counts.

What to look for

The shapes are (1797, 8, 8), (1797, 64), and (1797,).

Make it yours

Change no data yet; print the second image's true label with y[1].

Build with me · 2

2. Look at one numerical image

Selecting one example removes the example dimension, leaving an eight-by-eight grid. Each entry measures brightness, not a probability and not a class number. A value near zero is dark; a value near sixteen is bright. The label is a separate integer. The index selects both image and label, preserving their pairing. Printing a grid is an inspection step, not training.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetindex = 0image = digits.images[index]print("True label:", y[index])print(image)print("One image:", image.shape)

Run and look for brighter values outlining the digit whose label is printed.

What to look for

Index zero has label 0 and an eight-by-eight brightness grid.

Make it yours

Choose an index from 0 to 1796. Keep image and label on the same index.

Build with me · 3

3. Make a readable text preview

A nested loop turns bright pixels into # and darker pixels into dots. The conditional expression chooses one value: use # if the condition is true, otherwise use a dot. This threshold is for display only; it discards shades. The real model will still receive the original scaled brightness values. The outer loop ends each row with a newline, keeping the image's spatial arrangement visible.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetindex = 0print("True label:", y[index])for row in digits.images[index]:    for pixel in row:        symbol = "#" if pixel >= 8 else "."        print(symbol, end="")    print()

Run and compare the preview with the numerical grid from the previous stage.

What to look for

Eight lines of eight characters form a coarse zero for index zero.

Make it yours

Choose a display threshold of 4 or 12 and explain why the apparent stroke thickness changes even though the underlying image did not.

Build with me · 4

4. Flatten without losing order

Reshape reorganises dimensions without changing the number of values. The first eight flattened entries are the first image row, followed by the next row. Flattening does not average the image and does not discard all but one brightness. The order becomes part of the model's input contract: feeding the same values in a scrambled order would describe a different image. The equality check compares with the dataset's provided flattened row.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetgrid = digits.images[0]flat = grid.reshape(64)print("Grid shape:", grid.shape)print("Flat shape:", flat.shape)print("First grid row:", grid[0])print("First eight flat values:", flat[:8])print("Matches packaged row:", np.array_equal(flat, digits.data[0]))

Run and compare the first row with the first eight values.

What to look for

The shapes are (8, 8) and (64,), and the packaged row comparison is True.

Make it yours

Reshape the flat vector back to (8, 8) and check it equals the original grid.

Build with me · 5

5. Keep a batch dimension for one image

A model expects a table: one row per image, one column per pixel. Even for a single image, it wants a table with one row, called a batch of one. X[0] picks one image as a plain row of 64 values, shape (64,). X[:1] is a slice, as in the lists lesson, so it keeps the table form: one row of 64, shape (1, 64). X[0].reshape(1, 64) turns the plain row into the same one-row table. The trailing comma in (64,) is just how Python writes a shape with a single number.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetprint("One vector:", X[0].shape)print("One-row batch:", X[:1].shape)print("Another one-row batch:", X[0].reshape(1, 64).shape)print("Five-row batch:", X[:5].shape)

Run and identify which outputs describe tables suitable for a classifier.

What to look for

The shapes are (64,), (1, 64), (1, 64), and (5, 64).

Make it yours

Select a different single-image batch with X[10:11] and explain why its length is one.

Build with me · 6

6. Scale according to the actual source

The original brightness range is zero through sixteen, so dividing once by sixteen maps it to zero through one. NumPy applies division to every entry. The divisor comes from this dataset, not from a universal image rule. MNIST images commonly use 0-255 and require their own preparation. Training, development, final-test, and later drawings must follow the same scale for a particular model. Dividing an already normalised drawing again is a silent input error, not extra cleaning.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetraw = digits.dataprint("Original range:", raw.min(), raw.max())print("Scaled range:", X.min(), X.max())print("Sample values:", np.array([0, 8, 16]) / 16.0)print("Storage type:", X.dtype)

Run and check both minima and maxima. Explain why the sample midpoint becomes 0.5.

What to look for

Original range is 0 to 16 and scaled range is 0 to 1.

Make it yours

Print the range after an accidental second division and describe why that would change a future model's input.

Build with me · 7

7. Select positions with a condition

y == wanted compares every label with 8 at once and gives an array of True and False, one per image. np.where(...) returns the positions where the answer is True; the [0] after it picks the list of positions out of what np.where returns. Using those positions on y shows only 8s, which confirms the positions are right. Later you will use exactly this to find the images a model got wrong, with the condition predictions != y_dev.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetwanted = 8positions = np.where(y == wanted)[0]print("How many:", len(positions))print("First matching positions:", positions[:5])print("Their labels:", y[positions[:5]])

Run and confirm that every displayed selected label matches wanted.

What to look for

The chosen positions all have true label 8; the count comes from the real dataset.

Make it yours

Choose another digit and compare its count without changing the input-label ordering.

Build with me · 8

8. Choose the strongest score

These scores are made up, to show how to pick a winner. Each row belongs to one image, and each column to one possible answer. np.argmax(scores, axis=1) finds, in each row, the position of the largest score; axis=1 means look along each row. It returns positions, 1 and 0 here, not the answers themselves. classes[winner_positions] turns positions into the labels they stand for: position 1 means 5, position 0 means 2. Models report their answers in exactly this way, as a position in a list of classes, so keep the two apart. Scores that add up to one look like chances, but a high score is not a promise that the answer is right.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npscores = np.array([[0.1, 0.7, 0.2], [0.6, 0.3, 0.1]])classes = np.array([2, 5, 8])winner_positions = np.argmax(scores, axis=1)print("Winning column positions:", winner_positions)print("Winning class labels:", classes[winner_positions])print("Row totals:", scores.sum(axis=1))

Predict the two winning columns and then map them to labels before running.

What to look for

Winning positions are [1, 0], winning labels are [5, 2], and both rows sum to one.

Make it yours

Change the example scores while keeping each row total equal to one. Predict which labels will win.

Build with me · 9

9. Inspect a complete prepared example

This stage brings the checks together for one image: its label, its shape as a batch of one, and its brightness range. An assert line stops the program with an error if its condition is False, so a wrong shape or range cannot slip through unnoticed. Matplotlib, imported as plt, is the library for drawing pictures and charts. plt.imshow draws the 64 values as an 8 by 8 grey picture; vmin=0 and vmax=1 fix 0 as black and 1 as white, so two pictures can be compared fairly. No model is involved: checking the input comes before any training.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetimport matplotlib.pyplot as plt index = 0one_batch = X[index:index + 1]print("Label:", y[index])print("Batch shape:", one_batch.shape)print("Range:", one_batch.min(), one_batch.max())assert one_batch.shape == (1, 64)assert 0 <= one_batch.min() <= one_batch.max() <= 1plt.figure(figsize=(3, 3))plt.imshow(one_batch.reshape(8, 8), cmap="gray", vmin=0, vmax=1)plt.title("Prepared input; true label " + str(y[index]))plt.axis("off")plt.tight_layout()plt.show()

Run and inspect the image in Output. Change index to view several examples, including an ambiguous one.

What to look for

The selected example appears as a real plotted image with shape (1, 64) and values within 0-1.

Make it yours

Choose your own example index, keeping the title's true label tied to that same row.

Keep the pixel contract with the model

An image can be viewed as an eight-by-eight grid or as one row of 64 features. Reshaping changes how the same values are arranged; it does not resize a different image or learn from it. The model needs the same feature order and brightness scale for training and prediction. Inspect shape, numerical range, and a rendered example whenever data crosses that boundary. A correctly shaped array can still carry an inverted or otherwise wrong image.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetimport matplotlib.pyplot as plt index = 0one_batch = X[index:index + 1]print("Label:", y[index])print("Batch shape:", one_batch.shape)print("Range:", one_batch.min(), one_batch.max())assert one_batch.shape == (1, 64)assert 0 <= one_batch.min() <= one_batch.max() <= 1plt.figure(figsize=(3, 3))plt.imshow(one_batch.reshape(8, 8), cmap="gray", vmin=0, vmax=1)plt.title("Prepared input; true label " + str(y[index]))plt.axis("off")plt.tight_layout()plt.show()

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in