0%
BuildBuild a complete neural-network projectOptional reading.about 35 min, 8 steps

Optional: the same workflow in MNIST and Keras words

An optional guide for reading tutorials elsewhere: run each part of your network and see the matching MNIST and Keras vocabulary.

Work here, beside the explanation

This reading is optional. It is for when you want to follow a tutorial or course elsewhere, most of which use a bigger digit dataset called MNIST and a different library called Keras. The words in those tutorials look unfamiliar, but almost every one of them names something you have already done here. Each stage runs a piece of the network you already trained and puts the matching Keras word next to it. Nothing in the rest of this level depends on this page.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Distinguish two digit datasets

Both collections hold handwritten digits, but they are not interchangeable. Ours has 1,797 images of eight by eight pixels, with brightness from 0 to 16. MNIST has 60,000 training images and 10,000 test images of 28 by 28 pixels, with brightness from 0 to 255. Twenty-eight times twenty-eight is 784, so an MNIST model expects 784 numbers per image where ours expects 64. A model trained on one cannot read the other's images, and a network for MNIST has many more weights. This course uses the small collection so that all training fits comfortably in a browser.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetprint("This runnable dataset:", len(y), "images")print("Image shape:", digits.images[0].shape)print("Features per image:", X.shape[1])print("Original brightness:", digits.data.min(), "to", digits.data.max())print("Comparison only: MNIST images are 28 by 28, or", 28 * 28, "features.")

Run and separate the numbers measured from our data from the MNIST numbers the program simply prints for comparison.

What to look for

Our data has 64 numbers per image; the MNIST comparison line shows 784.

Make it yours

Explain why a model that expects 64 numbers cannot be given 784, even though both describe a handwritten digit.

Build with me · 2

2. Prepare numbers and protect evaluation

In both libraries the preparation is the same idea: scale the brightness to 0 to 1, then keep training, development, and final-test images apart. For our data that means dividing by 16. For MNIST it means dividing by 255, because 255 is its brightest value. MNIST comes already split into training and test images; tutorials often take some of the training images aside as a validation set, which is what this course calls the development set.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)print("Training:", X_train.shape)print("Development:", X_dev.shape)print("Final:", X_final.shape)print("Scaled training range:", X_train.min(), X_train.max())

Run and say which of the three parts is allowed to change the network's weights.

What to look for

Only the training images are used to adjust weights; development images are for checking and choosing; the final images stay aside.

Make it yours

Describe the job of each of the three parts without using their variable names.

Build with me · 3

3. Specify a dense architecture

The architecture is the shape of the network: how many layers and how many units in each. Ours is 64 inputs, one hidden layer of 32 units with ReLU, and 10 outputs. A Keras tutorial writes the same kind of network as a list of layers. Flatten turns an image grid into one long row of pixels, which our data already is. Dense(32, activation="relu") is a hidden layer of 32 units, each connected to every input; dense just means fully connected. Dense(10, activation="softmax") is the output layer, one unit per digit, with softmax turning scores into chances. The program prints the Keras words only for comparison; no Keras code runs here.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}print("Input units:", X_train.shape[1])print("Hidden units:", model.hidden_layer_sizes)print("Output classes:", np.arange(10))print("Keras vocabulary comparison: Flatten, Dense(32, relu), Dense(10, softmax)")

Run and draw the network as 64, then 32, then 10.

What to look for

The printed numbers describe 64 inputs, 32 hidden units, and 10 output classes.

Make it yours

Work out the number of weights and biases for a hidden layer of 64 units, as you did in the neural network lesson.

Build with me · 4

4. Connect configuration with compile

In Keras, a step called compile chooses three things before training: the optimiser (the rule for nudging weights, such as Adam), the loss (how wrongness is measured, cross-entropy for digits), and the metrics to report while training, such as accuracy. compile does not train anything. scikit-learn takes the same choices as settings when the model is created: solver is the optimiser, learning_rate_init its step size, and cross-entropy is always the loss for this kind of classifier. Keras tutorials also say sparse categorical cross-entropy; the word sparse only means the labels are plain digits such as 3, like ours.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}print("Optimiser:", model.solver)print("Initial learning rate:", model.learning_rate_init)print("Batch size:", model.batch_size)print("Classification loss used for reporting: cross-entropy")

Run and label each printed line as an update rule, a step size, a batch size, or a measurement.

What to look for

The network uses Adam, a learning rate of 0.003, and batches of 64 images.

Make it yours

Explain why choosing to report accuracy does not tell the network how to change its weights.

Build with me · 5

5. Perform the first real epoch

Keras trains with model.fit(...), giving it the number of epochs to run. Here partial_fit runs one epoch, as in the previous lessons. Both do the same job: pass through the training images and nudge the weights. Measuring on development images afterwards only checks the network; it is not more training.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}model.partial_fit(X_train, y_train, classes=np.arange(10))print("Training accuracy:", round(model.score(X_train, y_train), 3))print("Development accuracy:", round(model.score(X_dev, y_dev), 3))

Run and wait for both accuracies.

What to look for

In our run the network scores about 0.53 on its training images and about 0.56 on development after one epoch.

Make it yours

Explain the difference between a network that has only been created and one that has been trained for one epoch.

Build with me · 6

6. Record a complete training history

Keras's fit returns a history object holding the loss and metrics after every epoch. The dictionary here does the same job by hand: after each of 20 epochs it appends the training loss, the development loss, and the development accuracy. log_loss calculates the cross-entropy loss from the neural network lesson over a whole set of images at once.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))

Run and read the early and late lines.

What to look for

Twenty lines print. In our run the development loss falls from about 1.82 to about 0.16, and development accuracy rises to about 0.96.

Make it yours

Run fewer epochs and note that you changed the experiment, so its numbers are not directly comparable.

Build with me · 7

7. Inspect a prediction as a distribution

Keras's model.predict returns ten chances per image, just like predict_proba here. The largest chance, found with np.argmax, gives the predicted digit, and model.classes_ turns its position into the digit. Chances that add up to 1 are not a guarantee: a network can put a high chance on a wrong answer.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))probabilities = model.predict_proba(X_dev[:1])[0]winner = int(model.classes_[np.argmax(probabilities)])print("Class scores:", np.round(probabilities, 3))print("Predicted:", winner, "true:", int(y_dev[0]))print("Score sum:", probabilities.sum())

Run and compare the predicted digit with the true one.

What to look for

Ten chances print. In our run nearly all the chance, about 0.99, goes to 9, and the true label is 9.

Make it yours

Look at another development image and see whether its chances are as concentrated.

Build with me · 8

8. Read the actual learning curves

Tutorials often plot training loss and validation loss against the epoch number. The plot here is the same, drawn from your own history. Training loss shows how well the network fits the images it learns from. Development loss shows how well it does on images it does not learn from. If the development loss starts rising while the training loss keeps falling, the network is starting to memorise its training images; the next lesson looks at that closely.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import matplotlib.pyplot as plt plt.figure(figsize=(6, 3))plt.plot(history["epoch"], history["train_loss"], label="Training")plt.plot(history["epoch"], history["dev_loss"], label="Development")plt.xlabel("Epoch")plt.ylabel("Cross-entropy loss")plt.legend()plt.tight_layout()plt.show()print("Final development accuracy:", round(model.score(X_dev, y_dev), 4))print("Reserved final-test count:", len(y_final))

Run and describe what the two curves do in your run.

What to look for

A plot with a training curve and a development curve appears, followed by the final development accuracy.

Make it yours

Write the matching words in your own notes: create the network, compile (choose optimiser and loss), fit (train), evaluate (score), predict, save.

The Keras words, one line each

  • Flatten: turn an image grid into one row of pixels. No learned numbers.
  • Dense: a layer of units, each connected to every input, with weights and biases.
  • compile: choose the optimiser, loss, and reported metrics. Trains nothing.
  • fit: train, changing the weights. Usually given a number of epochs and a validation set.
  • evaluate: score the network on labelled images.
  • predict: get chances for new images, no labels needed.
  • save: write the trained network to a file, in Keras's own format.
  • EarlyStopping: stop training when the validation loss stops improving, and optionally go back to the best epoch.

If you follow an MNIST tutorial later, keep the same habits as here: set the final test aside first, choose settings using validation results only, and scale the images by their own brightest value.

What runs in this page

Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.

This course trains on the 1,797 small digit images that come with scikit-learn: eight by eight pixels, brightness 0 to 16. Many tutorials elsewhere use MNIST, a larger collection of 70,000 digit images that are 28 by 28 pixels with brightness 0 to 255, and a library called Keras. The ideas you learn here carry over, but the image sizes and scales differ, so a model trained on one cannot read the other's images.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import matplotlib.pyplot as plt plt.figure(figsize=(6, 3))plt.plot(history["epoch"], history["train_loss"], label="Training")plt.plot(history["epoch"], history["dev_loss"], label="Development")plt.xlabel("Epoch")plt.ylabel("Cross-entropy loss")plt.legend()plt.tight_layout()plt.show()print("Final development accuracy:", round(model.score(X_dev, y_dev), 4))print("Reserved final-test count:", len(y_final))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in