0%
BuildBuild a complete neural-network projectabout 28 min, 8 steps

Epochs, batches, and reading a loss curve

Measure the training and development loss after every epoch, draw both curves, read what they show, and choose how many epochs to train.

Work here, beside the explanation

In the previous lesson you trained a network one pass at a time and watched its accuracy climb. Now you measure its loss after every pass as well, on the training images and on the development images, and draw both as curves. Reading those two curves tells you whether the network is still learning, has learned enough, or has started to memorise its training images. At the end you use the curves to choose how many passes to train for.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Count the work in an epoch

An epoch is one pass through all the training images. The network does not change its weights once per epoch, though. It works through the images in groups called batches, and nudges its weights once per batch. With 1,077 training images and batches of 64, sixteen full batches cover 1,024 images and one last, smaller batch holds the remaining 53. That is 17 rounds of nudges per epoch. math.ceil rounds a division up to the next whole number, because the leftover 53 images still need a batch of their own. The % sign gives the remainder of a division: 1,077 divided by 64 leaves 53.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)import math batch_size = 64updates_per_epoch = math.ceil(len(y_train) / batch_size)print("Training examples:", len(y_train))print("Batch size:", batch_size)print("Batches per full epoch:", updates_per_epoch)print("Final partial batch:", len(y_train) % batch_size)

Work out 1,077 divided by 64 on paper, rounding up, before running.

What to look for

17 batches per epoch, with 53 images in the final batch.

Make it yours

Change the batch size to 32 and then 128, and predict the number of batches each time.

Build with me · 2

2. Initialise a history before training

The setup creates the network and a dictionary called history with four empty lists: epoch numbers, training loss, development loss, and development accuracy. The lists start empty because nothing has been measured yet. Keeping four clearly named lists stops you mixing up training loss with development loss later. The network is created once, here, outside any loop. A new experiment should start with a new network; another pass on an existing network continues its training.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}print("History:", history)print("Model created; no epochs run yet.")

Run and name the four kinds of measurement the history will hold.

What to look for

The history prints as four empty lists, and the message says no epochs have run yet.

Make it yours

Add a variable holding a name for this experiment, such as width-32-baseline, so you can tell runs apart later.

Build with me · 3

3. Measure one actual epoch

After one pass of partial_fit, the program measures the network three ways. model.predict_proba(X_train) gives the ten chances for every training image, and log_loss turns those chances and the true labels into one number: the average cross-entropy loss from the neural network lesson, over all the images. The same is done for the development images. labels=np.arange(10) tells log_loss that the answers are the digits 0 to 9, in order. Only the partial_fit line changes the network; the other lines only measure it. Accuracy is a different measure: it only counts which digit won, while loss also cares how sure the network was.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}model.partial_fit(X_train, y_train, classes=np.arange(10))train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))print("Training cross-entropy:", train_loss)print("Development cross-entropy:", dev_loss)print("Development accuracy:", model.score(X_dev, y_dev))

Run and point to the one line that changes the network.

What to look for

In our run, after one epoch the training loss is about 1.84, the development loss about 1.82, and development accuracy about 0.56.

Make it yours

Compare the loss numbers with the accuracy. They are different kinds of number: loss has no fixed top, and lower is better.

Build with me · 4

4. Append measurements after every pass

Now the loop runs 20 passes. After each one it measures the three numbers and appends each to its list in history, so position 0 of every list describes epoch 1, position 1 describes epoch 2, and so on. The lists must stay lined up like this, or the plot in the next stage would pair the wrong numbers. The printed line lets you watch progress; the lists keep the numbers after the output has scrolled away.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))

Run and watch the losses fall from line to line.

What to look for

Twenty lines print. In our run the training loss falls from about 1.84 to about 0.10, the development loss from about 1.82 to about 0.16, and development accuracy rises to about 0.96.

Make it yours

Change the number of epochs and record it as a separate experiment.

Build with me · 5

5. Plot comparable loss measurements

The plot draws both curves against the epoch number, using the same loss formula so they can be compared directly. Here is how to read two loss curves:

  • Both falling: the network is still learning patterns that also work on new images. More epochs may help.
  • Training falling, development flat or rising: the network is getting better at its own training images but not at new ones. It has started to memorise, which Level 2 called overfitting. More epochs make things worse.
  • Both high and flat, or shooting up: the network is not learning at all, often because the learning rate is too big.

A single wobble is normal; look for a trend that lasts several epochs. The final test stays out of this plot.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import matplotlib.pyplot as plt plt.figure(figsize=(6, 3))plt.plot(history["epoch"], history["train_loss"], label="Training")plt.plot(history["epoch"], history["dev_loss"], label="Development")plt.xlabel("Epoch")plt.ylabel("Cross-entropy loss")plt.legend()plt.tight_layout()plt.show()

Run and decide which of the three patterns your curves show.

What to look for

A plot with a training curve and a development curve appears. In our run both are still falling at epoch 20, with the development curve a little above the training curve.

Make it yours

Write one observation the plot supports and one question it cannot answer on its own.

Build with me · 6

6. Select a candidate epoch numerically

To choose the number of passes by a clear rule, find the epoch with the lowest development loss. np.argmin(history["dev_loss"]) returns the position of the smallest value in that list, and the same position in the other lists gives its epoch number and accuracy. Choosing by development accuracy instead could pick a different epoch, so say which rule you used. One thing trips people up: finding the best epoch in the history does not rewind the network. After the loop, model still holds the weights from the last epoch.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))best_position = int(np.argmin(history["dev_loss"]))best_epoch = history["epoch"][best_position]print("Lowest development loss at epoch:", best_epoch)print("Development loss there:", history["dev_loss"][best_position])print("Development accuracy there:", history["dev_accuracy"][best_position])print("Current model is still epoch:", history["epoch"][-1])

Run and compare the best epoch with the epoch the model is currently at.

What to look for

In our run the lowest development loss is at epoch 20, the last one, because the network was still improving when training stopped. So here the best epoch and the current model happen to match.

Make it yours

Change learning_rate_init=0.003 to 0.03 in the setup and the loop to 40 epochs, then run. In our run the bigger steps make the development loss bottom out at epoch 19 and then creep up, so the best epoch comes well before the end and the current model is no longer the best one.

Build with me · 7

7. Rebuild the selected epoch from a fresh start

When the best epoch comes before the end, you need a network trained for exactly that many passes. The simplest way is to build a fresh network with the same settings and seed and train it for best_epoch passes. With the same data and seed it follows the same path, so it arrives at the same weights. Training the old network for more passes would not work: that continues training instead of going back. Bigger projects save a copy of the weights after every epoch, called a checkpoint, so they can go back without retraining.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))best_position = int(np.argmin(history["dev_loss"]))best_epoch = history["epoch"][best_position]selected = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)for epoch in range(best_epoch):    selected.partial_fit(X_train, y_train, classes=np.arange(10))print("Rebuilt epochs:", best_epoch)print("Selected development loss:", log_loss(y_dev, selected.predict_proba(X_dev), labels=np.arange(10)))print("Selected development accuracy:", selected.score(X_dev, y_dev))

Run and compare the rebuilt development loss with the value recorded for the selected epoch.

What to look for

The rebuilt network's development loss and accuracy match the history's numbers for the selected epoch.

Make it yours

Explain how checkpoints would save time if each epoch took an hour instead of a second.

Build with me · 8

8. Save the complete evidence

The report keeps the whole history, not just the best number, together with the rule used to pick an epoch and the settings. "final_test_used": False states openly that the final test has not been touched yet. The file explains how and why a number of epochs was chosen. It does not contain the network's weights; saving those comes in the next module.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import matplotlib.pyplot as plt plt.figure(figsize=(6, 3))plt.plot(history["epoch"], history["train_loss"], label="Training")plt.plot(history["epoch"], history["dev_loss"], label="Development")plt.xlabel("Epoch")plt.ylabel("Cross-entropy loss")plt.legend()plt.tight_layout()plt.show()import json best_position = int(np.argmin(history["dev_loss"]))report = {    "selection_rule": "minimum development cross-entropy",    "selected_epoch": int(history["epoch"][best_position]),    "training_count": len(y_train), "development_count": len(y_dev),    "batch_size": 64, "seed": 42, "history": history,    "final_test_used": False,}with open("training-history.json", "w") as handle:    json.dump(report, handle, indent=2)print("Selected epoch:", report["selected_epoch"])print("Saved training-history.json; the current object remains the last epoch.")

Run, download training-history.json, and find the selected epoch inside it.

What to look for

training-history.json is offered for download, with twenty recorded epochs and the selection rule.

Make it yours

Add a "next_step" entry saying what you would try next, based on what the curves showed.

Choose the weights you actually use

Each epoch nudges the weights 17 times, once per batch of 64 training images. The 360 development images are only measured, never trained on. The lowest development loss picks a candidate number of epochs, but the network in memory stays at the last epoch unless you rebuild the chosen one or saved checkpoints on the way. Whatever you choose, write down the epoch and the rule, and save the weights of that exact network. A history file alone cannot make predictions.

What runs in this page

Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier(    hidden_layer_sizes=(32,), activation="relu", solver="adam",    learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(1, 21):    model.partial_fit(X_train, y_train, classes=np.arange(10))    train_loss = log_loss(y_train, model.predict_proba(X_train), labels=np.arange(10))    dev_loss = log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10))    dev_accuracy = model.score(X_dev, y_dev)    history["epoch"].append(epoch)    history["train_loss"].append(train_loss)    history["dev_loss"].append(dev_loss)    history["dev_accuracy"].append(dev_accuracy)    print("Epoch", epoch, "train loss", round(train_loss, 4),          "dev loss", round(dev_loss, 4), "dev accuracy", round(dev_accuracy, 3))import matplotlib.pyplot as plt plt.figure(figsize=(6, 3))plt.plot(history["epoch"], history["train_loss"], label="Training")plt.plot(history["epoch"], history["dev_loss"], label="Development")plt.xlabel("Epoch")plt.ylabel("Cross-entropy loss")plt.legend()plt.tight_layout()plt.show()import json best_position = int(np.argmin(history["dev_loss"]))report = {    "selection_rule": "minimum development cross-entropy",    "selected_epoch": int(history["epoch"][best_position]),    "training_count": len(y_train), "development_count": len(y_dev),    "batch_size": 64, "seed": 42, "history": history,    "final_test_used": False,}with open("training-history.json", "w") as handle:    json.dump(report, handle, indent=2)print("Selected epoch:", report["selected_epoch"])print("Saved training-history.json; the current object remains the last epoch.")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in