0%
BuildTrain a classifier in the browserabout 30 min, 8 steps

A small neural network in the browser

Train a real neural network on the digits with the same split as the closest-example model, read its settings and loss curve, and compare the two fairly.

Work here, beside the explanation

You now know what happens inside a neural network. In this lesson you train a real one on the digits, using exactly the same training and development images as your closest-example model, so the two can be compared fairly. The library's name for it is MLPClassifier. MLP stands for multilayer perceptron, an older name for the layered network of units you built by hand in the previous lesson. Training takes a few seconds in the browser; the first run of a visit also loads the libraries.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Check both models get the same data

Before training anything, check that the network will see the same data as the closest-example model did: the same 1,077 training images, the same 360 development images, and the same 0 to 1 brightness scale. If the two models were given different data, a difference in their scores might come from the data instead of the model, and the comparison would tell you nothing. The final test stays aside, as before.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)print("Training shape:", X_train.shape)print("Development shape:", X_dev.shape)print("Input range:", X_train.min(), X_train.max())print("Final test reserved:", len(y_final))

Run and compare the counts with the ones from your first complete classifier.

What to look for

Training has shape (1077, 64), development (360, 64), brightness runs from 0.0 to 1.0, and 360 final-test images are reserved.

Make it yours

Write one sentence describing what goes into the model and what comes out: how many numbers per image, on what scale, and which ten answers are possible.

Build with me · 2

2. Create a fresh neural network

The settings in brackets describe the network and how it will be trained. Each one maps to an idea from the previous lesson:

  • hidden_layer_sizes=(32,): one hidden layer of 32 units. It is written as a tuple, and a tuple with one item needs the trailing comma. (32, 16) would mean two hidden layers.
  • activation="relu": the hidden units use ReLU.
  • solver="adam": Adam is the rule that decides each step when the weights are nudged.
  • learning_rate_init=0.003: the starting step size, the learning rate.
  • batch_size=64: the network works through the training images 64 at a time, making one round of nudges per group of 64.
  • max_iter=100: at most 100 passes through all the training images. One full pass is called an epoch.
  • random_state=42: the weights start as small random numbers; fixing the seed gives the same starting numbers every run.
  • tol=0.0001: stop early if the loss stops improving by at least this much.

Creating the model only records these settings. No weights exist until you call fit.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifier model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, max_iter=100,    random_state=42, tol=0.0001,)print(model)

Run and match each printed setting to its line in the list above.

What to look for

The printed description repeats the settings. There is no accuracy yet, because nothing has been trained.

Make it yours

Pick one setting you would like to test later, such as the number of hidden units, but leave all of them unchanged for now.

Build with me · 3

3. Fit the training examples

fit gives every weight and bias a small random starting value, then runs the training loop from the previous lesson. For each batch of 64 training images it works out the loss and the gradients, nudges every weight, and moves on to the next batch. After the last batch it starts another pass. model.n_iter_ reports how many passes ran, and model.loss_curve_ is a list holding the training loss after each pass.

With these settings the training reaches the 100-pass limit, and Python prints a ConvergenceWarning. That is a note, not an error. It means the loss was still creeping down when the limit stopped it. The network is trained and usable. Whether more passes would actually help is a question for the development results, not something to fix by raising the limit without a reason.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifier model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, max_iter=100,    random_state=42, tol=0.0001,)model.fit(X_train, y_train)print("Epochs performed:", model.n_iter_)print("Training loss points:", len(model.loss_curve_))

Run and wait for training to finish. Write down how many passes ran.

What to look for

Epochs performed: 100 and Training loss points: 100. A ConvergenceWarning also appears in the output.

Make it yours

Explain in your own words why the warning does not mean the program failed.

Build with me · 4

4. Inspect the learned layer shapes

After fitting, model.coefs_ holds the weights as one table per layer, and model.intercepts_ holds the biases. The first weight table has 64 rows, one per pixel, and 32 columns, one per hidden unit, so each column is one hidden unit's 64 weights. The second table has 32 rows, one per hidden unit, and 10 columns, one per digit. The loop uses enumerate and zip from the pairs lesson to number the layers from 1 and to walk the weight tables and bias lists side by side. .size counts how many numbers an array holds, and the last line adds up the sizes of all four arrays.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifier model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, max_iter=100,    random_state=42, tol=0.0001,)model.fit(X_train, y_train)for layer, (weights, biases) in enumerate(zip(model.coefs_, model.intercepts_), start=1):    print("Layer", layer, "weights", weights.shape, "biases", biases.shape)print("Total learned numbers:", sum(w.size + b.size for w, b in zip(model.coefs_, model.intercepts_)))

Run and compare the total with the 2,410 you counted by hand in the previous lesson.

What to look for

Layer 1 has weights (64, 32) and biases (32,); layer 2 has weights (32, 10) and biases (10,). The total is 2,410 learned numbers.

Make it yours

Predict every shape for a network with 16 hidden units before trying it in a separate run.

Build with me · 5

5. Compare training and development

Accuracy on the training images tells you how well the network fits the examples it learned from. Accuracy on the development images tells you how well it does on images it never trained on, which is what you actually care about. The difference between the two is called the gap. A network with thousands of adjustable numbers can often fit its training images perfectly. A large gap is a sign that it has partly memorised them instead of learning patterns that carry over, which Level 2 called overfitting. The first five predictions are printed as a quick check that the program is wired up correctly.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifier model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, max_iter=100,    random_state=42, tol=0.0001,)model.fit(X_train, y_train)print("Training accuracy:", round(model.score(X_train, y_train), 4))print("Development accuracy:", round(model.score(X_dev, y_dev), 4))print("First development predictions:", model.predict(X_dev[:5]))print("Corresponding true labels:", y_dev[:5])

Run and write down both accuracies with the correct label next to each.

What to look for

In our run training accuracy is 1.0 and development accuracy is about 0.978, and the first five predictions match their true labels.

Make it yours

Work out the gap yourself, then write one sentence about this network that stays within what these two numbers show.

Build with me · 6

6. Plot the actual training loss

The loss curve shows the training loss after each pass. It usually falls quickly at first and then flattens out. It is measured on the training images only, so it shows the network getting better at the images it has seen; it cannot show how it does on new ones. A later lesson records the development loss after every pass as well, so the two curves can be compared. np.arange(1, len(model.loss_curve_) + 1) makes the pass numbers 1, 2, 3, and so on for the horizontal axis. The library's loss also includes a small extra term that discourages very large weights, so its numbers are close to, but not exactly, the cross-entropy you calculated by hand.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifier model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, max_iter=100,    random_state=42, tol=0.0001,)model.fit(X_train, y_train)import matplotlib.pyplot as plt epochs = np.arange(1, len(model.loss_curve_) + 1)print("First loss:", model.loss_curve_[0])print("Last loss:", model.loss_curve_[-1])plt.figure(figsize=(6, 3))plt.plot(epochs, model.loss_curve_)plt.xlabel("Epoch")plt.ylabel("Training objective")plt.title("This run's training curve")plt.tight_layout()plt.show()

Run and look at the start, the steep part, and the flat end of the curve.

What to look for

In our run the curve starts near 2.1 and falls to about 0.008 by the hundredth pass.

Make it yours

Describe the shape of the curve in one sentence, and say what it cannot tell you about new images.

Build with me · 7

7. Compare it with the closest-example model

Now compare the two models fairly: the same training images and the same development images. The closest-example model stores examples and lets the nearest ones vote; the network has adjusted 2,410 numbers. In our run the closest-example model scores slightly higher on development. That is a genuine result, not a mistake in your code. On 1,797 small, neat images, looking up similar stored images works very well, and a more complicated model is not automatically a better one. Keep the result either way; a comparison is only honest if you would have reported the other outcome too.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifier model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, max_iter=100,    random_state=42, tol=0.0001,)model.fit(X_train, y_train)from sklearn.neighbors import KNeighborsClassifier nearest = KNeighborsClassifier(n_neighbors=3)nearest.fit(X_train, y_train)print("KNN development:", round(nearest.score(X_dev, y_dev), 4))print("Network development:", round(model.score(X_dev, y_dev), 4))print("Network training:", round(model.score(X_train, y_train), 4))

Run and compare the two development accuracies.

What to look for

In our run the closest-example model scores about 0.981 on development and the network about 0.978, with the network at 1.0 on its training images.

Make it yours

Add one more comparison that is not accuracy, such as how many learned numbers each model has. The closest-example model has none; it keeps all 1,077 images instead.

Build with me · 8

8. Save a record of the run

A saved record lets you, or anyone checking your work, see exactly what was run and what it scored. The dictionary gathers the data, the split, the settings, the number of passes, and both measured development scores, and json writes it to a file, as in the pairs lesson. sklearn.__version__ records the library version, because a different version can give slightly different numbers. float(...) and int(...) convert NumPy's numbers into ordinary ones that JSON can store. The network's confusion matrix is printed too; read it as in the confusion matrix lesson. This file records results only. It does not contain the trained weights; saving the model itself comes in the last module.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifier model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, max_iter=100,    random_state=42, tol=0.0001,)model.fit(X_train, y_train)from sklearn.neighbors import KNeighborsClassifierfrom sklearn.metrics import confusion_matriximport sklearnimport json nearest = KNeighborsClassifier(n_neighbors=3)nearest.fit(X_train, y_train)report = {    "dataset": "sklearn digits 8x8", "scale_divisor": 16, "seed": 42,    "train_count": len(y_train), "development_count": len(y_dev),    "hidden_units": 32, "max_epochs": 100, "actual_epochs": int(model.n_iter_),    "network_development_accuracy": float(model.score(X_dev, y_dev)),    "knn_development_accuracy": float(nearest.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}print(json.dumps(report, indent=2))print("Network development confusion matrix:")print(confusion_matrix(y_dev, model.predict(X_dev), labels=np.arange(10)))with open("network-report.json", "w") as handle:    json.dump(report, handle, indent=2)print("Saved network-report.json")

Run, read the printed record, and download network-report.json if you want to keep it.

What to look for

A network-report.json file appears with this run's measured scores, and the network's ten-by-ten confusion matrix prints.

Make it yours

Add a "next_step" entry to the dictionary with one sentence about what you would test next, keeping every other entry as it is.

What this network showed

The network has one hidden layer of 32 units and 2,410 learned numbers. It used all 100 passes it was allowed, which is why the ConvergenceWarning appeared. It fits its training images perfectly and does very slightly worse than the closest-example model on development images. Both of those are facts about this data, this split, and these settings, measured in this run.

The exercise that follows uses its own starter and split, and a slightly larger network, so its numbers differ a little from this lesson's. Use its settings when you complete it.

What runs in this page

Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifier model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, max_iter=100,    random_state=42, tol=0.0001,)model.fit(X_train, y_train)from sklearn.neighbors import KNeighborsClassifierfrom sklearn.metrics import confusion_matriximport sklearnimport json nearest = KNeighborsClassifier(n_neighbors=3)nearest.fit(X_train, y_train)report = {    "dataset": "sklearn digits 8x8", "scale_divisor": 16, "seed": 42,    "train_count": len(y_train), "development_count": len(y_dev),    "hidden_units": 32, "max_epochs": 100, "actual_epochs": int(model.n_iter_),    "network_development_accuracy": float(model.score(X_dev, y_dev)),    "knn_development_accuracy": float(nearest.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}print(json.dumps(report, indent=2))print("Network development confusion matrix:")print(confusion_matrix(y_dev, model.predict(X_dev), labels=np.arange(10)))with open("network-report.json", "w") as handle:    json.dump(report, handle, indent=2)print("Saved network-report.json")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in