What a neural network is, step by step
Build one unit of a neural network by hand, turn scores into chances, measure the loss, and nudge a weight to make the loss smaller.
Work here, beside the explanation
The closest-example model answers by comparing a new image with stored ones. A neural network works differently. During training it adjusts a large set of numbers until those numbers, combined with the pixels, give a high score to the right digit. This lesson builds the smallest piece of a network by hand with made-up numbers, then shows how the pieces combine and how training changes them. No real digits are involved yet; the next lesson trains a real network.
Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.
The building block: a weighted vote
A network is made of small identical pieces called units (sometimes called neurons). One unit does something you can do on a calculator. It takes some input numbers, for example pixel brightnesses. It multiplies each input by its own weight, a number that says how much that input matters and in which direction. It adds the results up.
Think of a judge scoring a dish: saltiness might count a little in its favour, sweetness might count against, and the judge's personal weights decide the total. The weights are exactly the kind of adjustable numbers Level 1 called parameters. Training a network means finding good weights.
Build with me · 1
1. Give inputs different influences
Here a unit has two inputs, 0.5 and 1.0, and two weights, 2 and -1. Each input is multiplied by its own weight to give its contribution. The first input is 0.5 with weight 2, so it contributes 1.0. The second input is 1.0 with weight -1, so it contributes -1.0. A positive weight means that input pushes the total up; a negative weight means it pushes the total down; a weight near zero means the input hardly matters. The bias line is used in the next stage.
inputs = [0.5, 1.0]weights = [2.0, -1.0]bias = 0.2first = inputs[0] * weights[0]second = inputs[1] * weights[1]print("First contribution:", first)print("Second contribution:", second)Work out both contributions on paper, then run.
What to look for
First contribution: 1.0 and Second contribution: -1.0.
Make it yours
Change the second weight to 1.0 and predict its new contribution before running.
Build with me · 2
2. Add the bias
After adding the contributions, the unit adds one more number, its bias. The bias is a starting nudge that does not depend on the inputs, like a head start or a handicap. It lets a unit lean towards a high or low total even when the inputs are small. Here the two contributions cancel out (1.0 plus -1.0 is 0), so the whole total, called the weighted sum, is just the bias, 0.2. Training adjusts biases as well as weights; both count as the unit's learned numbers.
inputs = [0.5, 1.0]weights = [2.0, -1.0]bias = 0.2weighted_sum = inputs[0] * weights[0] + inputs[1] * weights[1] + biasprint("Weighted sum plus bias:", weighted_sum)Run and trace the sum: 1.0, plus -1.0, plus 0.2.
What to look for
The result is 0.2.
Make it yours
Change the bias to -0.5 and predict whether the total becomes positive or negative.
Build with me · 3
3. Apply ReLU
The last thing a unit does is pass its total through a simple rule. The most common rule is called ReLU: keep the total if it is positive, and replace it with zero if it is negative. In Python that is just max(0, total). You can read it as: the unit reports how strongly its pattern is present, or stays silent. The result is called the unit's activation. This bend at zero matters more than it looks. Without it, stacking many units would add up to one big weighted sum, which can only describe very simple patterns. With it, many units together can describe the curved and broken shapes that handwriting needs.
inputs = [0.5, 1.0]weights = [2.0, -1.0]bias = 0.2weighted_sum = inputs[0] * weights[0] + inputs[1] * weights[1] + biasactivation = max(0, weighted_sum)print("Before ReLU:", weighted_sum)print("After ReLU:", activation)print("ReLU of -0.7:", max(0, -0.7))Run and compare a positive total with the negative example -0.7.
What to look for
After ReLU, 0.2 stays about 0.2, and -0.7 becomes 0.
Make it yours
Try several bias values and find where the output switches from zero to positive.
Build with me · 4
4. Express the same unit with an array
Real units have many inputs: a unit reading a digit has 64. Writing out 64 multiplications would be silly, so NumPy has a shortcut. With two arrays of the same length, inputs @ weights multiplies each pair of matching items and adds the results. That operation is called a dot product, and it is exactly stages 1 and 2 without the bias. np.maximum(0, ...) is ReLU for arrays. The calculation has not changed; it is only shorter to write.
import numpy as npinputs = np.array([0.5, 1.0])weights = np.array([2.0, -1.0])bias = 0.2weighted_sum = inputs @ weights + biasactivation = np.maximum(0, weighted_sum)print("Dot product:", inputs @ weights)print("Activated output:", activation)Run and check that the answer matches the long version from stage 3.
What to look for
The dot product is 0.0 and the activated output is about 0.2, as before.
Make it yours
Add a third input and a third weight to both arrays and predict the new dot product. The two arrays must keep the same length.
From one unit to a whole network
A network arranges units in layers. In the network you will train next, a hidden layer of 32 units each reads all 64 pixels. Each hidden unit has its own 64 weights, so each can learn to respond to a different pattern, such as a stroke across the top or a loop at the bottom. An output layer of 10 units then reads the 32 hidden activations and produces one score per digit, 0 to 9. The digit with the highest score is the network's answer.
Nobody chooses what each hidden unit looks for. Training sets all the weights, and the useful patterns emerge from the examples.
Build with me · 5
5. Turn three scores into a distribution
The output layer gives one raw score per answer, called a logit. Here there are three answers, the digits 2, 5, and 8, with scores 1, 2, and -1. Raw scores can be any size and can be negative, so they are turned into chances with a recipe called softmax:
- Subtract the largest score from every score. This is only a safety step that keeps the numbers small; it does not change the final result.
- Apply
np.exp, which raises the number e (about 2.718) to the power of each score. You do not need to know e. What matters is that the result is always positive, and each extra point of score makes the result about 2.7 times bigger, so the leading answer pulls ahead. - Divide each result by their total, so the chances add up to 1.
The list of chances is called a distribution, the same word Level 1 used for next-word chances. The digit with the largest chance is the prediction. np.argmax finds its position, and classes turns that position into the digit it stands for.
import numpy as nplogits = np.array([1.0, 2.0, -1.0])shifted = logits - logits.max()exponentials = np.exp(shifted)probabilities = exponentials / exponentials.sum()classes = np.array([2, 5, 8])print("Probabilities:", np.round(probabilities, 3))print("Total:", probabilities.sum())print("Predicted class:", classes[np.argmax(probabilities)])Predict which digit wins before running.
What to look for
The chances are about 0.259, 0.705, and 0.035. They add up to 1, and the predicted digit is 5.
Make it yours
Raise the first score from 1.0 to 3.0 and predict which digit wins now.
Measuring how wrong it was
In Level 1 you scored a delivery-time guess with a loss: a number that is 0 for a perfect answer and grows as the answer gets worse. A digit network needs the same kind of number. Its answer is a set of chances rather than a single number, so its loss asks one question: how much chance did the network give to the right digit?
Build with me · 6
6. Measure support for the correct label
For one training image, look at the chance the network gave to the correct digit. The loss is -math.log(chance). math.log is the natural logarithm, a function you do not need to calculate by hand; the table this program prints shows everything you need. When the chance for the right answer is high, such as 0.9, the loss is small. As the chance falls towards zero, the loss grows fast. The minus sign only makes the result positive. This loss is called cross-entropy. It rewards the network for being confident and right, and punishes it hardest for being confident and wrong. Accuracy only asks which digit won; loss also notices how sure the network was, which is why the two can move differently.
import mathfor correct_class_probability in [0.9, 0.5, 0.2, 0.01]: loss = -math.log(correct_class_probability) print("Correct-class support", correct_class_probability, "loss", round(loss, 4))Predict the order of the four losses, smallest to largest, before running.
What to look for
The losses are about 0.105, 0.693, 1.609, and 4.605. A chance of 0.01 for the right answer costs about 44 times as much loss as a chance of 0.9.
Make it yours
Add 0.51 and 0.99 to the list. Both would make the correct digit win in a two-way choice, yet their losses are about 0.673 and 0.010.
Learning: nudging each weight downhill
In Level 1 you trained a delivery model by trying different values of one parameter and keeping the one with the smallest loss: guess, check, adjust. A digit network has thousands of weights, far too many to try values by hand. Instead, for every weight the computer works out a gradient: which direction to nudge that weight to make the loss smaller, and how strongly the loss reacts. Then it moves every weight a small step that way. The size of the step is the learning rate. Repeat that over many examples and the loss goes down.
Build with me · 7
7. Follow one gradient update
This is the smallest possible model: one weight, one input of 2.0, and a correct answer of 1.0. The weight starts at 0, so the prediction is 0 times 2, which is 0, and the squared-error loss from Level 1 is (0 minus 1) squared, which is 1. The gradient line is a formula that comes from a branch of mathematics called calculus; you do not need to work it out, only to read what it tells you. Its sign says which way to move the weight: negative means increase it. Its size says how steeply the loss changes. The update line moves the weight a small step in the direction that lowers the loss: the learning rate, 0.1, times the gradient, -4, gives -0.4, and subtracting -0.4 from 0 gives a new weight of 0.4. With the new weight the prediction is 0.8 and the loss drops from 1 to about 0.04.
weight = 0.0input_value = 2.0target = 1.0learning_rate = 0.1prediction = weight * input_valueloss_before = (prediction - target) ** 2gradient = 2 * (prediction - target) * input_valueweight = weight - learning_rate * gradientloss_after = (weight * input_value - target) ** 2print("Gradient:", gradient)print("New weight:", weight)print("Loss before:", loss_before)print("Loss after:", loss_after)Run and follow the numbers: gradient, new weight, loss before, loss after.
What to look for
Gradient: -4.0, New weight: 0.4, the loss falls from 1.0 to about 0.04.
Make it yours
Change the learning rate to 0.5 and run. The weight jumps to 2.0, far past the best value of 0.5, and the loss rises to 9. A bigger step is not always better.
Build with me · 8
8. Count a full digit network's parameters
Now count the learned numbers in the network you will train next. Each of the 32 hidden units has 64 weights, one per pixel, plus one bias: 64 times 32 plus 32 is 2,080. Each of the 10 output units has 32 weights, one per hidden unit, plus one bias: 32 times 10 plus 10 is 330. Together that is 2,410 numbers for training to adjust. A bigger network has more numbers, which lets it describe more complicated patterns but also takes longer to train and can memorise its training examples. Only measuring on development data tells you whether bigger helps.
input_count = 64hidden_units = 32class_count = 10hidden_parameters = input_count * hidden_units + hidden_unitsoutput_parameters = hidden_units * class_count + class_countprint("Hidden weights and biases:", hidden_parameters)print("Output weights and biases:", output_parameters)print("Total parameters:", hidden_parameters + output_parameters)Check the two multiplications by hand, then run.
What to look for
The hidden layer has 2,080 parameters, the output layer 330, and the total is 2,410.
Make it yours
Set hidden_units to 64 and confirm the total becomes 4,810. Say which parts of the count changed and which stayed the same.
Put the pieces together
A unit multiplies its inputs by weights, adds a bias, and passes the total through ReLU. A hidden layer of units reads the pixels; an output layer turns the hidden activations into one score per digit; softmax turns the scores into chances that add up to 1. The loss measures how little chance went to the right digit. Training works out, for every weight and bias, which way to nudge it to lower the loss, and takes a small step. It repeats that over the training images many times.
In the next lesson a library does all of this for you, with 2,410 learned numbers instead of the handful here. You will recognise its settings: the number of hidden units, ReLU, the learning rate. Two new names will appear there as well. Backpropagation is the efficient method for working out the gradient of every weight at once, and Adam is a popular rule for choosing each step. Both do the job you just did by hand for one weight.
What runs in this page
Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.
The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.
Full reference solution
This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.
input_count = 64hidden_units = 32class_count = 10hidden_parameters = input_count * hidden_units + hidden_unitsoutput_parameters = hidden_units * class_count + class_countprint("Hidden weights and biases:", hidden_parameters)print("Output weights and biases:", output_parameters)print("Total parameters:", hidden_parameters + output_parameters)Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.