0%
BuildUnder the hoodabout 26 min, 8 steps

Trace a neural network, and compare its MNIST counterpart

Inspect a real browser network's forward calculation and understand how a 28-by-28 MNIST architecture changes its dimensions.

Work here, beside the explanation

Trace one image through a complete dense network and compare the browser model's shapes with a typical MNIST architecture. The required training in this article is the real eight-by-eight browser network. The 28-by-28, 128-hidden-unit MNIST example is an explicitly labelled shape and parameter comparison, so understanding it never depends on leaving the page or running unsupported TensorFlow code.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Keep the batch dimension

The grid representation has two image dimensions; the model table has a batch dimension plus a flattened feature dimension. A one-image batch must still have one row. The true label belongs to the selected example but is not supplied to predict. A Keras MNIST input could begin as (1, 28, 28) before Flatten, whereas our prepared browser input already has shape (1, 64). The different numbers describe different datasets, not different definitions of a batch.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)one_image = X_dev[0].reshape(8, 8)one_batch = X_dev[:1]print("One grid:", one_image.shape)print("One model-input batch:", one_batch.shape)print("Corresponding true label:", y_dev[0])

Run and explain every dimension in the two printed shapes.

What to look for

The grid is (8, 8) and the one-row model batch is (1, 64).

Make it yours

Select another development row while keeping image and true label aligned.

Build with me · 2

2. Verify that flattening learns nothing

Flattening rearranges dimensions without learning weights, averaging pixels, or removing values. The following dense weights refer to feature positions, so flattening order must stay consistent at training and inference. A dense layer does not automatically understand two-dimensional neighbourhoods; it receives the chosen vector representation. Convolutional architectures use image structure differently, but the dense network here keeps the full calculation inspectable for beginners.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)one_grid = X_dev[0].reshape(8, 8)flat = one_grid.reshape(1, 64)print("Flattened shape:", flat.shape)print("Values preserved:", np.array_equal(flat, X_dev[:1]))print("Learned parameters in reshape:", 0)

Run and compare the flattened values with the original model row.

What to look for

Values preserved is True and reshape has no learned parameters.

Make it yours

Explain why shuffling the flattened positions would preserve shape while changing the image meaning.

Build with me · 3

3. Count browser-network parameters

Every hidden unit receives all 64 inputs, with one weight per input and one bias. The output layer receives 32 hidden values per class and adds one bias per output. Counting bias terms is essential. The resulting total describes adjustable numerical parameters, not separate human-readable concepts. Training does not promise that a particular unit becomes a neatly labelled 'loop detector' or 'vertical-stroke detector'; such interpretations require evidence.

Python at this stageWorked example
PythonHover over a line to see an explanation
inputs = 64hidden = 32outputs = 10hidden_count = inputs * hidden + hiddenoutput_count = hidden * outputs + outputsprint("Hidden parameters:", hidden_count)print("Output parameters:", output_count)print("Total:", hidden_count + output_count)

Calculate both layer totals before running.

What to look for

The hidden layer has 2,080 parameters, the output has 330, and total is 2,410.

Make it yours

Double the hidden width and predict the new total.

Build with me · 4

4. Compare a typical MNIST architecture

A common simple MNIST design flattens 784 pixels, uses 128 ReLU hidden units, and produces ten softmax outputs. The hidden layer contains 784 times 128 weights plus 128 biases, or 100,480 parameters. The output adds 1,290, giving 101,770. This arithmetic compares architecture size. It is not measured MNIST accuracy, an executed Keras training run, or a reason to assume the larger network is better for every task.

Python at this stageWorked example
PythonHover over a line to see an explanation
image_rows = 28image_columns = 28inputs = image_rows * image_columnshidden = 128outputs = 10hidden_count = inputs * hidden + hiddenoutput_count = hidden * outputs + outputsprint("MNIST comparison only; no MNIST model is fitted in this stage.")print("Shapes:", (1, 28, 28), (1, inputs), (1, hidden), (1, outputs))print("Hidden parameters:", hidden_count)print("Output parameters:", output_count)print("Total:", hidden_count + output_count)

Run and trace the four printed shapes from input grid to class scores.

What to look for

The comparison total is 101,770 parameters.

Make it yours

Choose a different hidden width and recompute the parameter count while leaving the 28-by-28 input contract explicit.

Build with me · 5

5. Fit the browser model and inspect its first layer

This is a real model trained in the browser. Matrix multiplication calculates every hidden unit's weighted sum for the selected input batch. The bias vector is added across the units, and maximum with zero applies ReLU. The shape changes from (1, 64) to (1, 32). You are inspecting learned numerical combinations, not assigning semantic names to units without evidence. Parameters remain fixed during this forward inspection.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))one_batch = X_dev[:1]hidden_raw = one_batch @ model.coefs_[0] + model.intercepts_[0]hidden = np.maximum(0, hidden_raw)print("Input batch:", one_batch.shape)print("Hidden weighted sums:", hidden_raw.shape)print("Hidden activations:", hidden.shape)print("First six activations:", hidden[0, :6])

Run and connect the two layer shapes with the earlier parameter calculation.

What to look for

The actual hidden activation has shape (1, 32), with nonnegative values after ReLU.

Make it yours

Count how many hidden activations are zero for this input and compare with another input.

Build with me · 6

6. Reconstruct the output distribution

The second weight matrix transforms 32 hidden values into ten logits. Subtracting each row's largest logit stabilises exponentials, and dividing by the row sum implements softmax. The reconstructed distribution should match predict_proba because it applies the same learned weights and activations. This check makes the model's prediction path concrete. A score distribution summing to one still does not guarantee confidence calibration or recognition of out-of-distribution inputs.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))one_batch = X_dev[:1]hidden = np.maximum(0, one_batch @ model.coefs_[0] + model.intercepts_[0])logits = hidden @ model.coefs_[1] + model.intercepts_[1]exponentials = np.exp(logits - logits.max(axis=1, keepdims=True))probabilities = exponentials / exponentials.sum(axis=1, keepdims=True)print("Output shape:", probabilities.shape)print("Class support:", np.round(probabilities[0], 3))print("Matches library:", np.allclose(probabilities, model.predict_proba(one_batch)))

Run and require the comparison with the library to pass.

What to look for

Output has shape (1, 10), and Matches library is True.

Make it yours

Explain why argmax returns a column position that must be interpreted through the class mapping.

Build with me · 7

7. Separate fitting, validation, and prediction

The training loop updates weights from training pairs, while development loss and accuracy only measure current behaviour. History stores measurements; the model object stores fitted parameters. In Keras, compile configures the learning procedure, fit updates the model, and a History object records results. Those roles remain distinct here despite different method names. A decreasing training objective does not replace development selection or a reserved final estimate.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))print("Training examples that updated weights:", len(y_train))print("Development examples used only for measurements:", len(y_dev))print("Final examples still reserved:", len(y_final))print("Last development accuracy:", history["dev_accuracy"][-1])print("One prediction without supplying a label:", model.predict(X_dev[:1])[0])

Run and identify which information lives in history and which lives in model.

What to look for

Data roles and a real label-free prediction are displayed without evaluating the final test.

Make it yours

Explain what would change if development examples were accidentally passed into partial_fit.

Build with me · 8

8. Verify the complete forward path

The final check reconstructs five predictions directly from learned arrays. It verifies the input shape, hidden representation, output distribution, and label mapping together. A finished recognizer also preserves the exact preparation, selected weights, environment, and evaluation record. The saving article's JSON loader applies this same forward calculation after reload. That is why reproducing a fresh-session prediction without retraining is a meaningful completion check.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))batch = X_dev[:5]hidden = np.maximum(0, batch @ model.coefs_[0] + model.intercepts_[0])logits = hidden @ model.coefs_[1] + model.intercepts_[1]exponentials = np.exp(logits - logits.max(axis=1, keepdims=True))probabilities = exponentials / exponentials.sum(axis=1, keepdims=True)predictions = model.classes_[np.argmax(probabilities, axis=1)]assert np.allclose(probabilities, model.predict_proba(batch))assert np.array_equal(predictions, model.predict(batch))print("Shape path:", batch.shape, hidden.shape, probabilities.shape)print("Predictions:", predictions)print("True labels:", y_dev[:5])print("Forward calculation matches the fitted model.")

Build and run the complete trace, then explain every shape and why the two assertions should pass.

What to look for

The path is (5, 64) to (5, 32) to (5, 10), and reconstructed predictions match the library.

Make it yours

Describe the corresponding MNIST path separately: (batch, 28, 28), (batch, 784), (batch, 128), (batch, 10). Do not reuse the browser weights for those incompatible shapes.

What the complete recognizer must preserve

Keep the selected model artifact, image-preparation procedure, class mapping, environment record, and evaluation report. Our course's JSON artifact supports the stated dense ReLU/softmax architecture; a Keras artifact uses its own compatible format. In either case, verify that a fresh run can load the artifact and classify a prepared input without fitting again.

The ability to trace dimensions and identify what each operation learns is more durable than memorising a particular library's method names. Improving robustness then becomes a measured development task, with representative inputs and honest evaluation, rather than a search for a more impressive layer name.

What runs in this page

Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.

This course trains on the 1,797 small digit images that come with scikit-learn: eight by eight pixels, brightness 0 to 16. Many tutorials elsewhere use MNIST, a larger collection of 70,000 digit images that are 28 by 28 pixels with brightness 0 to 255, and a library called Keras. The ideas you learn here carry over, but the image sizes and scales differ, so a model trained on one cannot read the other's images.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))batch = X_dev[:5]hidden = np.maximum(0, batch @ model.coefs_[0] + model.intercepts_[0])logits = hidden @ model.coefs_[1] + model.intercepts_[1]exponentials = np.exp(logits - logits.max(axis=1, keepdims=True))probabilities = exponentials / exponentials.sum(axis=1, keepdims=True)predictions = model.classes_[np.argmax(probabilities, axis=1)]assert np.allclose(probabilities, model.predict_proba(batch))assert np.array_equal(predictions, model.predict(batch))print("Shape path:", batch.shape, hidden.shape, probabilities.shape)print("Predictions:", predictions)print("True labels:", y_dev[:5])print("Forward calculation matches the fitted model.")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Worth a look

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in