0%
BuildShip itabout 29 min, 8 steps

Save and reload your digit recognizer

Save your trained network's learned numbers and instructions to a file, load them back, predict with them yourself, and prove nothing changed.

Work here, beside the explanation

Everything a trained network has learned lives in the computer's memory, and it disappears when the page closes. To use it tomorrow without training again, you save its learned numbers, all 2,410 weights and biases, to a file, together with the facts needed to use them correctly. Then you load the file and make predictions using only what was saved. You will write that prediction step yourself; it is the same calculation you did by hand in the neural network lesson. Finally you prove that the reloaded network gives exactly the same answers as the original.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Find what the network learned

The setup trains the network for 20 epochs, as in the epochs lesson. What it learned is not the accuracy or the settings; it is the numbers inside it. model.coefs_ holds the two weight tables and model.intercepts_ the two bias lists, as you saw when you inspected layer shapes. model.classes_ records which output column stands for which digit. A file holding only the settings, such as 32 hidden units, could build a network of the right shape, but its weights would be random again and it would have to be retrained.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))print("Number of weight matrices:", len(model.coefs_))print("Weight shapes:", [array.shape for array in model.coefs_])print("Bias shapes:", [array.shape for array in model.intercepts_])print("Classes:", model.classes_)

Run and list what must be saved so the network can predict without training again.

What to look for

Two weight tables with shapes (64, 32) and (32, 10), two bias lists with shapes (32,) and (10,), and the ten classes 0 to 9.

Make it yours

Explain why a file that only says hidden_units=32 would not keep what the network learned.

Build with me · 2

2. Gather the numbers and the instructions

The learned numbers are only useful with instructions for using them. The dictionary payload holds both. The first group of entries describes the input the network expects: a name for this file format, the image size, the order the pixels are flattened in (row by row, which people call row-major), the number the original pixels must be divided by, the prepared range, and bright strokes on a dark background. Then come the activations (ReLU for the hidden layer, softmax for the output), the class order, and the weights and biases themselves. JSON can only store plain lists and numbers, so .tolist() turns each NumPy table into ordinary nested lists; the list comprehensions do that for every layer.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import json payload = {    "format": "code-guide-dense-v1",    "dataset": "sklearn digits 8x8",    "input_shape": [8, 8], "flatten_order": "row-major", "raw_pixel_divisor": 16,    "prepared_range": [0, 1], "stroke": "bright-on-dark",    "hidden_activation": "relu", "output_activation": "softmax",    "classes": [int(label) for label in model.classes_],    "weights": [weights.tolist() for weights in model.coefs_],    "biases": [bias.tolist() for bias in model.intercepts_],    "training": {"seed": 42, "epochs": 20, "train_count": len(y_train)},}print("Format:", payload["format"])print("Input shape:", payload["input_shape"])print("Original-pixel divisor:", payload["raw_pixel_divisor"])print("Prepared range:", payload["prepared_range"])print("Class order:", payload["classes"])

Run and read each printed field. Say what would go wrong if it were missing.

What to look for

The format name, an 8 by 8 input shape, a divisor of 16, a prepared range of 0 to 1, and the classes 0 to 9 print.

Make it yours

Add a "name" entry describing your model, leaving every other entry as it is.

Build with me · 3

3. Save the model to a file

json.dump writes the whole payload into digit-model.json as text, the same way you saved settings in the pairs lesson. This file holds the real learned numbers, not a screenshot and not a promise to retrain. The last line counts them using .size, as in the small network lesson. After the run the editor offers the file for download. Keep a downloaded copy: the browser's copy disappears with the page.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import json payload = {    "format": "code-guide-dense-v1",    "dataset": "sklearn digits 8x8",    "input_shape": [8, 8], "flatten_order": "row-major", "raw_pixel_divisor": 16,    "prepared_range": [0, 1], "stroke": "bright-on-dark",    "hidden_activation": "relu", "output_activation": "softmax",    "classes": [int(label) for label in model.classes_],    "weights": [weights.tolist() for weights in model.coefs_],    "biases": [bias.tolist() for bias in model.intercepts_],    "training": {"seed": 42, "epochs": 20, "train_count": len(y_train)},}with open("digit-model.json", "w") as handle:    json.dump(payload, handle)print("Saved digit-model.json")print("Learned numbers:", sum(w.size + b.size for w, b in zip(model.coefs_, model.intercepts_)))

Run and download digit-model.json.

What to look for

Saved digit-model.json, holding 2,410 learned numbers.

Make it yours

Open the downloaded file in any text editor and find the word relu and the start of the first weight list.

Build with me · 4

4. Load the file back

json.load reads the file back into a Python dictionary, here called restored. At this point it is just a dictionary of lists and numbers. It is not a scikit-learn model, and nothing in Python knows how to predict with it yet. np.asarray(...) turns the first weight list back into a NumPy table so its shape can be checked against the original.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import json payload = {    "format": "code-guide-dense-v1",    "dataset": "sklearn digits 8x8",    "input_shape": [8, 8], "flatten_order": "row-major", "raw_pixel_divisor": 16,    "prepared_range": [0, 1], "stroke": "bright-on-dark",    "hidden_activation": "relu", "output_activation": "softmax",    "classes": [int(label) for label in model.classes_],    "weights": [weights.tolist() for weights in model.coefs_],    "biases": [bias.tolist() for bias in model.intercepts_],    "training": {"seed": 42, "epochs": 20, "train_count": len(y_train)},}with open("digit-model.json", "w") as handle:    json.dump(payload, handle)with open("digit-model.json") as handle:    restored = json.load(handle)print("Loaded format:", restored["format"])print("First weight matrix:", np.asarray(restored["weights"][0]).shape)print("Saved epochs:", restored["training"]["epochs"])

Run and compare the restored table's shape with the one from stage 1.

What to look for

The loaded format name prints, the first weight table has shape (64, 32), and the saved epoch count is 20.

Make it yours

Look at a small part of the downloaded JSON and match it with the printed fields.

Build with me · 5

5. Predict using only the saved numbers

predict_saved does the calculation from the neural network lesson, using nothing but the file. Read it in four parts:

  1. Two checks: that this really is the course's model format, and that the inputs are a batch with 64 values per image. values.ndim is the number of dimensions, 2 for a table. A real public app would check more, such as the value range.
  2. Turn the saved lists back into NumPy tables with np.asarray(..., dtype=float).
  3. For each layer: values @ weight + bias, the weighted sum for every unit at once, then ReLU with np.maximum(0, ...) for every layer except the last.
  4. Softmax on the last layer's scores, then np.argmax picks the biggest chance in each row and classes turns its position into a digit. axis=1 means work along each row, one image at a time, and keepdims=True keeps each row's maximum and total lined up with that row.

No training happens inside this function; the weights never change.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import json payload = {    "format": "code-guide-dense-v1",    "dataset": "sklearn digits 8x8",    "input_shape": [8, 8], "flatten_order": "row-major", "raw_pixel_divisor": 16,    "prepared_range": [0, 1], "stroke": "bright-on-dark",    "hidden_activation": "relu", "output_activation": "softmax",    "classes": [int(label) for label in model.classes_],    "weights": [weights.tolist() for weights in model.coefs_],    "biases": [bias.tolist() for bias in model.intercepts_],    "training": {"seed": 42, "epochs": 20, "train_count": len(y_train)},}with open("digit-model.json", "w") as handle:    json.dump(payload, handle)with open("digit-model.json") as handle:    restored = json.load(handle) def predict_saved(payload, prepared_inputs):    if payload.get("format") != "code-guide-dense-v1":        raise ValueError("This is not the supported course model format.")    values = np.asarray(prepared_inputs, dtype=float)    if values.ndim != 2 or values.shape[1] != 64:        raise ValueError("Provide a batch with 64 prepared features per row.")    weights = [np.asarray(layer, dtype=float) for layer in payload["weights"]]    biases = [np.asarray(layer, dtype=float) for layer in payload["biases"]]    classes = np.asarray(payload["classes"])    for index, (weight, bias) in enumerate(zip(weights, biases)):        values = values @ weight + bias        if index < len(weights) - 1:            values = np.maximum(0, values)    exponentials = np.exp(values - values.max(axis=1, keepdims=True))    probabilities = exponentials / exponentials.sum(axis=1, keepdims=True)    predictions = classes[np.argmax(probabilities, axis=1)]    return predictions, probabilitiespredictions, probabilities = predict_saved(restored, X_dev[:5])print("Restored predictions:", predictions)print("Probability row totals:", probabilities.sum(axis=1))

Run and follow the steps of the function with the neural network lesson open beside it.

What to look for

Five predictions print, in our run 9, 4, 2, 8, 1, and each image's chances add up to 1.

Make it yours

Explain why passing these checks does not show that the model is accurate on new handwriting.

Build with me · 6

6. Prove the reloaded model matches the original

A file that opens is not proof that it works. This stage compares the original network and the reloaded one on all 360 development images. np.array_equal checks that every predicted digit is identical. np.allclose checks that every chance is equal apart from tiny rounding differences; rtol and atol say how tiny. The assert lines stop the program if either check fails, so you cannot miss a problem. This proves the file kept the network's behaviour. It is not a new test of how good the network is.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import json payload = {    "format": "code-guide-dense-v1",    "dataset": "sklearn digits 8x8",    "input_shape": [8, 8], "flatten_order": "row-major", "raw_pixel_divisor": 16,    "prepared_range": [0, 1], "stroke": "bright-on-dark",    "hidden_activation": "relu", "output_activation": "softmax",    "classes": [int(label) for label in model.classes_],    "weights": [weights.tolist() for weights in model.coefs_],    "biases": [bias.tolist() for bias in model.intercepts_],    "training": {"seed": 42, "epochs": 20, "train_count": len(y_train)},}with open("digit-model.json", "w") as handle:    json.dump(payload, handle)with open("digit-model.json") as handle:    restored = json.load(handle) def predict_saved(payload, prepared_inputs):    if payload.get("format") != "code-guide-dense-v1":        raise ValueError("This is not the supported course model format.")    values = np.asarray(prepared_inputs, dtype=float)    if values.ndim != 2 or values.shape[1] != 64:        raise ValueError("Provide a batch with 64 prepared features per row.")    weights = [np.asarray(layer, dtype=float) for layer in payload["weights"]]    biases = [np.asarray(layer, dtype=float) for layer in payload["biases"]]    classes = np.asarray(payload["classes"])    for index, (weight, bias) in enumerate(zip(weights, biases)):        values = values @ weight + bias        if index < len(weights) - 1:            values = np.maximum(0, values)    exponentials = np.exp(values - values.max(axis=1, keepdims=True))    probabilities = exponentials / exponentials.sum(axis=1, keepdims=True)    predictions = classes[np.argmax(probabilities, axis=1)]    return predictions, probabilitiesrestored_predictions, restored_probabilities = predict_saved(restored, X_dev)original_predictions = model.predict(X_dev)original_probabilities = model.predict_proba(X_dev)print("All labels identical:", np.array_equal(restored_predictions, original_predictions))print("Scores numerically match:", np.allclose(restored_probabilities, original_probabilities, rtol=1e-8, atol=1e-10))print("Largest score difference:", np.max(np.abs(restored_probabilities - original_probabilities)))assert np.array_equal(restored_predictions, original_predictions)assert np.allclose(restored_probabilities, original_probabilities, rtol=1e-8, atol=1e-10)

Run and make sure both checks pass before trusting the downloaded file.

What to look for

All labels identical: True, Scores numerically match: True, and the largest difference is 0.0 or a tiny number close to it.

Make it yours

Say what you would investigate first if these checks failed after someone changed the file format.

Build with me · 7

7. Use the saved model on your own drawing

Your drawing is prepared exactly as in the drawing lessons, giving a batch of one image on the 0 to 1 scale. predict_saved uses it directly. It must not be divided by 16 again, because the symbols were already turned into 0 to 1 values. This shows the saved file is useful beyond the dataset. A correct answer on your drawing is still one example, not an accuracy.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import json payload = {    "format": "code-guide-dense-v1",    "dataset": "sklearn digits 8x8",    "input_shape": [8, 8], "flatten_order": "row-major", "raw_pixel_divisor": 16,    "prepared_range": [0, 1], "stroke": "bright-on-dark",    "hidden_activation": "relu", "output_activation": "softmax",    "classes": [int(label) for label in model.classes_],    "weights": [weights.tolist() for weights in model.coefs_],    "biases": [bias.tolist() for bias in model.intercepts_],    "training": {"seed": 42, "epochs": 20, "train_count": len(y_train)},}with open("digit-model.json", "w") as handle:    json.dump(payload, handle)with open("digit-model.json") as handle:    restored = json.load(handle) def predict_saved(payload, prepared_inputs):    if payload.get("format") != "code-guide-dense-v1":        raise ValueError("This is not the supported course model format.")    values = np.asarray(prepared_inputs, dtype=float)    if values.ndim != 2 or values.shape[1] != 64:        raise ValueError("Provide a batch with 64 prepared features per row.")    weights = [np.asarray(layer, dtype=float) for layer in payload["weights"]]    biases = [np.asarray(layer, dtype=float) for layer in payload["biases"]]    classes = np.asarray(payload["classes"])    for index, (weight, bias) in enumerate(zip(weights, biases)):        values = values @ weight + bias        if index < len(weights) - 1:            values = np.maximum(0, values)    exponentials = np.exp(values - values.max(axis=1, keepdims=True))    probabilities = exponentials / exponentials.sum(axis=1, keepdims=True)    predictions = classes[np.argmax(probabilities, axis=1)]    return predictions, probabilities rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)predictions, probabilities = predict_saved(restored, one_input)print("Saved-model prediction:", int(predictions[0]))print("Class scores:", np.round(probabilities[0], 3))

Edit the rows into a digit of your choice, run, and compare the answer with the digit you meant.

What to look for

In our run the saved network reads the starter seven as 7, with a chance of about 0.84.

Make it yours

Use some + symbols for half-bright pixels and see how the chances change.

Build with me · 8

8. Load a saved model without training

The last program shows how the file is used on another day. os.path.exists("digit-model.json") asks whether that file is present. If you added your downloaded model in the Files tab, the first branch loads it and goes straight to predicting: no training at all. If there is no file, the else branch runs the whole training program from earlier stages, indented so it belongs to the branch, then saves the model so the program still works. The printed message tells you which branch ran. The file holds what is needed to predict, not everything needed to carry on training, such as Adam's step history.

Python at this stageWorked example
PythonHover over a line to see an explanation
import jsonimport osimport numpy as np def predict_saved(payload, prepared_inputs):    if payload.get("format") != "code-guide-dense-v1":        raise ValueError("This is not the supported course model format.")    values = np.asarray(prepared_inputs, dtype=float)    if values.ndim != 2 or values.shape[1] != 64:        raise ValueError("Provide a batch with 64 prepared features per row.")    weights = [np.asarray(layer, dtype=float) for layer in payload["weights"]]    biases = [np.asarray(layer, dtype=float) for layer in payload["biases"]]    classes = np.asarray(payload["classes"])    for index, (weight, bias) in enumerate(zip(weights, biases)):        values = values @ weight + bias        if index < len(weights) - 1:            values = np.maximum(0, values)    exponentials = np.exp(values - values.max(axis=1, keepdims=True))    probabilities = exponentials / exponentials.sum(axis=1, keepdims=True)    predictions = classes[np.argmax(probabilities, axis=1)]    return predictions, probabilities # Files uploaded as digit-model.json take this path without any training.if os.path.exists("digit-model.json"):    with open("digit-model.json") as handle:        restored = json.load(handle)    print("Loaded the supplied model; no fitting was performed.")else:    # A fresh lesson can still run end to end and produce its first artifact.    from sklearn.datasets import load_digits    import numpy as np        digits = load_digits()    X = digits.data.astype("float64") / 16.0    y = digits.target    from sklearn.model_selection import train_test_split        # Reserve the final test before trying model settings.    X_pool, X_final, y_pool, y_final = train_test_split(        X, y, test_size=0.2, random_state=42, stratify=y    )    X_train, X_dev, y_train, y_dev = train_test_split(        X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool    )    from sklearn.neural_network import MLPClassifier    from sklearn.metrics import log_loss        model = MLPClassifier(        hidden_layer_sizes=(32,), activation="relu", solver="adam",        learning_rate_init=0.003, batch_size=64, random_state=42,    )    history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}    for epoch in range(1, 21):        model.partial_fit(X_train, y_train, classes=np.arange(10))        train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))        dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))        dev_accuracy = model.score(X_dev, y_dev)        history["epoch"].append(epoch)        history["train_loss"].append(train_loss)        history["dev_loss"].append(dev_loss)        history["dev_accuracy"].append(dev_accuracy)        print("Epoch", epoch, "train loss", round(train_loss, 4),              "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))    payload = {        "format": "code-guide-dense-v1",        "dataset": "sklearn digits 8x8",        "input_shape": [8, 8], "flatten_order": "row-major", "raw_pixel_divisor": 16,        "prepared_range": [0, 1], "stroke": "bright-on-dark",        "hidden_activation": "relu", "output_activation": "softmax",        "classes": [int(label) for label in model.classes_],        "weights": [weights.tolist() for weights in model.coefs_],        "biases": [bias.tolist() for bias in model.intercepts_],        "training": {"seed": 42, "epochs": 20, "train_count": len(y_train)},    }    with open("digit-model.json", "w") as handle:        json.dump(payload, handle)    restored = payload    print("No model file was supplied; trained and saved a first model.") rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)predictions, probabilities = predict_saved(restored, one_input)print("Prediction:", int(predictions[0]))print("Class scores:", np.round(probabilities[0], 3))

First download digit-model.json from stage 3. Add it in the Files tab with that exact name, run this stage, and check the message says no fitting was performed.

What to look for

With the file supplied, the program loads it and predicts without training. Without it, the program trains, saves a first model, and then predicts.

Make it yours

Keep the model file, this program, and your drawing together, so someone else could reproduce your prediction.

Saving is something to test

Before saving, be sure which network you are saving. If you chose an earlier epoch in the epochs lesson, rebuild that epoch before saving it; saving the last network while describing an earlier one is a mismatch that is easy to miss.

The loader here only understands this course's network: dense layers, ReLU, then softmax, on 8 by 8 images. A different kind of network, image size, or class list needs a matching format and loader. Other libraries have their own file formats, and renaming a file does not convert one into another.

This JSON file is for making predictions. Continuing an interrupted training run exactly would also need the optimiser's state, which this file does not keep.

What runs in this page

Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
import jsonimport osimport numpy as np def predict_saved(payload, prepared_inputs):    if payload.get("format") != "code-guide-dense-v1":        raise ValueError("This is not the supported course model format.")    values = np.asarray(prepared_inputs, dtype=float)    if values.ndim != 2 or values.shape[1] != 64:        raise ValueError("Provide a batch with 64 prepared features per row.")    weights = [np.asarray(layer, dtype=float) for layer in payload["weights"]]    biases = [np.asarray(layer, dtype=float) for layer in payload["biases"]]    classes = np.asarray(payload["classes"])    for index, (weight, bias) in enumerate(zip(weights, biases)):        values = values @ weight + bias        if index < len(weights) - 1:            values = np.maximum(0, values)    exponentials = np.exp(values - values.max(axis=1, keepdims=True))    probabilities = exponentials / exponentials.sum(axis=1, keepdims=True)    predictions = classes[np.argmax(probabilities, axis=1)]    return predictions, probabilities # Files uploaded as digit-model.json take this path without any training.if os.path.exists("digit-model.json"):    with open("digit-model.json") as handle:        restored = json.load(handle)    print("Loaded the supplied model; no fitting was performed.")else:    # A fresh lesson can still run end to end and produce its first artifact.    from sklearn.datasets import load_digits    import numpy as np        digits = load_digits()    X = digits.data.astype("float64") / 16.0    y = digits.target    from sklearn.model_selection import train_test_split        # Reserve the final test before trying model settings.    X_pool, X_final, y_pool, y_final = train_test_split(        X, y, test_size=0.2, random_state=42, stratify=y    )    X_train, X_dev, y_train, y_dev = train_test_split(        X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool    )    from sklearn.neural_network import MLPClassifier    from sklearn.metrics import log_loss        model = MLPClassifier(        hidden_layer_sizes=(32,), activation="relu", solver="adam",        learning_rate_init=0.003, batch_size=64, random_state=42,    )    history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}    for epoch in range(1, 21):        model.partial_fit(X_train, y_train, classes=np.arange(10))        train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))        dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))        dev_accuracy = model.score(X_dev, y_dev)        history["epoch"].append(epoch)        history["train_loss"].append(train_loss)        history["dev_loss"].append(dev_loss)        history["dev_accuracy"].append(dev_accuracy)        print("Epoch", epoch, "train loss", round(train_loss, 4),              "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))    payload = {        "format": "code-guide-dense-v1",        "dataset": "sklearn digits 8x8",        "input_shape": [8, 8], "flatten_order": "row-major", "raw_pixel_divisor": 16,        "prepared_range": [0, 1], "stroke": "bright-on-dark",        "hidden_activation": "relu", "output_activation": "softmax",        "classes": [int(label) for label in model.classes_],        "weights": [weights.tolist() for weights in model.coefs_],        "biases": [bias.tolist() for bias in model.intercepts_],        "training": {"seed": 42, "epochs": 20, "train_count": len(y_train)},    }    with open("digit-model.json", "w") as handle:        json.dump(payload, handle)    restored = payload    print("No model file was supplied; trained and saved a first model.") rows = [    "..####..",    "..####..",    "....##..",    "...##...",    "...##...",    "..##....",    "..##....",    "........",]if len(rows) != 8 or any(len(row) != 8 for row in rows):    raise ValueError("Use exactly eight rows of eight characters.")if any(character not in ".+#" for row in rows for character in row):    raise ValueError("Use only . for dark, + for half-bright, and # for bright.")brightness = {".": 0.0, "+": 0.5, "#": 1.0}image = np.array([[brightness[character] for character in row] for row in rows])if image.max() == 0:    raise ValueError("The drawing is blank. Add a digit before predicting.")one_input = image.reshape(1, 64)predictions, probabilities = predict_saved(restored, one_input)print("Prediction:", int(predictions[0]))print("Class scores:", np.round(probabilities[0], 3))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in