Change one thing and keep the evidence
Compare three network widths fairly: write the plan first, change only one setting, choose with a rule, and keep every result.
Work here, beside the explanation
How many hidden units should the network have? You cannot know by thinking about it; you find out by measuring. This lesson turns that question into a small, fair experiment. You write the plan down first, change only the number of hidden units, train every version the same way, pick a winner with a rule you chose in advance, and keep every result, including the losers. This is the same one-change-at-a-time habit Level 2 used for photos, now with network settings.
Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.
Build with me · 1
1. Write the experiment before running it
Writing the plan before seeing any results stops you from quietly changing the question to suit the answer. The plan is a dictionary: the question, the one setting that changes (the number of hidden units, also called the width), the three widths to try, everything that stays fixed, and how the winner will be chosen. If two widths tie on accuracy, the plan already says to prefer the smaller one, because it is smaller and faster for the same result. plan.items() hands back each key with its value, as in the pairs lesson, and key + ":" joins the key and a colon into one piece of text.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)plan = { "question": "Does hidden width change development accuracy after 20 epochs?", "changed_setting": "hidden units", "candidates": [16, 32, 64], "fixed": "data, split, seed, scale, activation, optimiser, learning rate, batch size, epochs", "selection_metric": "development accuracy",}for key, value in plan.items(): print(key + ":", value)Run and read the fixed list. Everything on it must stay the same for every width.
What to look for
The plan prints: one changed setting, three candidate widths, a list of fixed conditions, and development accuracy as the way to choose.
Make it yours
Write down which width you expect to win before running anything else.
Build with me · 2
2. Train one fresh baseline candidate
Start with the width you already know, 32, as the baseline to compare against. The network is created fresh and trained for exactly 20 passes with partial_fit, like the previous lessons. Using exactly 20 passes for every width, instead of fit with a limit it might stop short of, keeps the training the same for every candidate. The program measures development accuracy, which the plan uses to choose, and development loss as extra evidence.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss model = MLPClassifier( hidden_layer_sizes=(32,), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42,)history = {"epoch": [], "train_loss": [], "dev_loss": [], "dev_accuracy": []}for epoch in range(20): model.partial_fit(X_train, y_train, classes=np.arange(10))print("Width:", 32)print("Development accuracy:", model.score(X_dev, y_dev))print("Development loss:", log_loss(y_dev, model.predict_proba(X_dev), labels=np.arange(10)))Run and write down both numbers for width 32.
What to look for
In our run width 32 reaches a development accuracy of about 0.961 and a development loss of about 0.163.
Make it yours
Keep this result in your notes whatever happens next.
Build with me · 3
3. Reconstruct for every candidate
This is the whole experiment in one program. The outer loop goes through the widths 16, 32, and 64. For each width it creates a brand new network, so every candidate starts from scratch; reusing one trained network would give the later widths a head start. The inner loop trains that network for 20 passes. The results go into a small dictionary per width, and every dictionary is appended to the results list, including weak ones. All three use the same training and development images, so the only difference between them is the width.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss results = []for width in [16, 32, 64]: candidate = MLPClassifier( hidden_layer_sizes=(width,), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42, ) for epoch in range(20): candidate.partial_fit(X_train, y_train, classes=np.arange(10)) result = { "hidden_units": width, "epochs": 20, "development_accuracy": float(candidate.score(X_dev, y_dev)), "development_loss": float(log_loss(y_dev, candidate.predict_proba(X_dev), labels=np.arange(10))), } results.append(result) print(result)print("Candidates measured:", len(results))Run and trace the two loops: three widths, and 20 passes for each.
What to look for
Three results print. In our run width 16 reaches about 0.936, width 32 about 0.961, and width 64 also about 0.961.
Make it yours
Before running, predict how long a width of 128 would take compared with 64, then try it in a separate run.
Build with me · 4
4. Add compute size to the comparison
Accuracy is not the only thing that matters. A wider network has more learned numbers, which takes more memory and more time to train and use. The loop works out each width's number of learned numbers with the formula from the neural network lesson: 64 inputs times the width, plus the width's biases, plus the width times 10 outputs, plus 10 output biases. The count is a rough guide to size, not a measured speed.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss results = []for width in [16, 32, 64]: candidate = MLPClassifier( hidden_layer_sizes=(width,), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42, ) for epoch in range(20): candidate.partial_fit(X_train, y_train, classes=np.arange(10)) result = { "hidden_units": width, "epochs": 20, "development_accuracy": float(candidate.score(X_dev, y_dev)), "development_loss": float(log_loss(y_dev, candidate.predict_proba(X_dev), labels=np.arange(10))), } results.append(result) print(result)for result in results: width = result["hidden_units"] parameters = 64 * width + width + width * 10 + 10 print("Width", width, "parameters", parameters, "development accuracy", result["development_accuracy"])Run and compare accuracy with size for each width.
What to look for
Width 16 has 1,210 learned numbers, width 32 has 2,410, and width 64 has 4,810, printed next to their development accuracies.
Make it yours
Decide how you would choose if a bigger network were only one image more accurate but twice as slow.
Build with me · 5
5. Apply the declared selection rule
Now apply the rule from the plan. max(results, key=...) picks the record with the largest key, as in the pairs lesson. This key is a pair: (accuracy, -width). Python compares pairs one item at a time: first the accuracies, and only if those are exactly equal, the second item. Because the second item is minus the width, a smaller width gives a larger number there, so an exact tie goes to the smaller network. The winner is chosen using development images, so its development score is not an unbiased final result.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss results = []for width in [16, 32, 64]: candidate = MLPClassifier( hidden_layer_sizes=(width,), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42, ) for epoch in range(20): candidate.partial_fit(X_train, y_train, classes=np.arange(10)) result = { "hidden_units": width, "epochs": 20, "development_accuracy": float(candidate.score(X_dev, y_dev)), "development_loss": float(log_loss(y_dev, candidate.predict_proba(X_dev), labels=np.arange(10))), } results.append(result) print(result)selected = max(results, key=lambda row: (row["development_accuracy"], -row["hidden_units"]))print("Selected configuration:", selected)print("Rule: highest development accuracy, then smaller width on an exact tie")Run and check the selected record against the three results from stage 3.
What to look for
In our run widths 32 and 64 tie on development accuracy, so the rule picks 32, the smaller one.
Make it yours
Explain why choosing whichever measure a favourite width happens to win on would not be a fair comparison.
Build with me · 6
6. Record a limit of one-seed evidence
A difference in accuracy sounds bigger as a percentage than it often is. The program finds the best and worst candidates and turns the gap between their accuracies into a number of images: accuracy is a fraction of 360 development images, so multiplying the gap by 360 gives the number of images it represents. A different random starting point, set by random_state, could shuffle a close ranking. That is worth saying in any report.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss results = []for width in [16, 32, 64]: candidate = MLPClassifier( hidden_layer_sizes=(width,), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42, ) for epoch in range(20): candidate.partial_fit(X_train, y_train, classes=np.arange(10)) result = { "hidden_units": width, "epochs": 20, "development_accuracy": float(candidate.score(X_dev, y_dev)), "development_loss": float(log_loss(y_dev, candidate.predict_proba(X_dev), labels=np.arange(10))), } results.append(result) print(result)best = max(results, key=lambda row: row["development_accuracy"])worst = min(results, key=lambda row: row["development_accuracy"])difference = best["development_accuracy"] - worst["development_accuracy"]print("Observed accuracy spread:", difference)print("Equivalent development examples:", difference * len(y_dev))print("This is one split and one initialisation seed, not a certainty about every run.")Run and turn the spread into a number of images in your own words.
What to look for
In our run the spread between best and worst is about 0.025, which is about 9 development images.
Make it yours
Plan a follow-up that repeats the comparison with two or three different random_state values, and say what it would tell you.
Build with me · 7
7. Retrain the selected configuration
The results list holds numbers and settings, not the trained networks themselves. To use the winner you rebuild it: create a network with the chosen width and the same settings and seed, and train it for the same 20 passes. With everything the same, it reproduces the development score from the table. Do not mix the development images into its training now; that would make it a different network from the one you measured.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss results = []for width in [16, 32, 64]: candidate = MLPClassifier( hidden_layer_sizes=(width,), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42, ) for epoch in range(20): candidate.partial_fit(X_train, y_train, classes=np.arange(10)) result = { "hidden_units": width, "epochs": 20, "development_accuracy": float(candidate.score(X_dev, y_dev)), "development_loss": float(log_loss(y_dev, candidate.predict_proba(X_dev), labels=np.arange(10))), } results.append(result) print(result)selected = max(results, key=lambda row: (row["development_accuracy"], -row["hidden_units"]))chosen_model = MLPClassifier( hidden_layer_sizes=(selected["hidden_units"],), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42,)for epoch in range(20): chosen_model.partial_fit(X_train, y_train, classes=np.arange(10))print("Chosen width:", selected["hidden_units"])print("Reproduced development accuracy:", chosen_model.score(X_dev, y_dev))print("Final test remains reserved:", len(y_final))Run and compare the rebuilt development accuracy with the selected record.
What to look for
The rebuilt width-32 network matches its development accuracy from the table, and the final test is still reserved.
Make it yours
Say what you would need to save to use this trained network tomorrow without training it again.
Build with me · 8
8. Save the complete comparison table
The report saves every candidate, the fixed conditions, the selection rule, the chosen width, and the library version to width-experiment.json. Keeping the losing widths matters: they are the evidence that the winner was chosen fairly. "final_test_used": False states that the final test has not been used. A finding that a bigger network is not worth its cost is a useful result, not a failure.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss results = []for width in [16, 32, 64]: candidate = MLPClassifier( hidden_layer_sizes=(width,), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42, ) for epoch in range(20): candidate.partial_fit(X_train, y_train, classes=np.arange(10)) result = { "hidden_units": width, "epochs": 20, "development_accuracy": float(candidate.score(X_dev, y_dev)), "development_loss": float(log_loss(y_dev, candidate.predict_proba(X_dev), labels=np.arange(10))), } results.append(result) print(result)import jsonimport sklearn selected = max(results, key=lambda row: (row["development_accuracy"], -row["hidden_units"]))report = { "question": "Compare one hidden-layer width after 20 epochs", "dataset": "sklearn digits 8x8", "seed": 42, "scale_divisor": 16, "train_count": len(y_train), "development_count": len(y_dev), "selection_rule": "highest development accuracy, smaller width on exact tie", "results": results, "selected": selected, "sklearn_version": sklearn.__version__, "final_test_used": False,}with open("width-experiment.json", "w") as handle: json.dump(report, handle, indent=2)print("Selected:", selected)print("Saved all candidates in width-experiment.json")Run, download width-experiment.json, and write a two-sentence conclusion that stays within what these numbers show.
What to look for
width-experiment.json is offered for download, with all three candidates and the selected width.
Make it yours
Add a "next_experiment" entry that changes one different setting, such as the learning rate, with every other setting back at its baseline.
Read a controlled comparison
Each candidate was a fresh network, trained on the same images for the same 20 passes, with only the width changed. That is what makes the comparison fair. With 360 development images, a one-image difference moves accuracy by about 0.003, so a close ranking can change with a different random start. Repeating the comparison with a few seeds shows whether a winner stays a winner.
Trying many settings on the same development images can still fit your choices to those particular images, even though the final test stays hidden. Keep the final test until the procedure is fixed and the model chosen.
Changing the number of epochs, adding a layer, or changing the learning rate are good separate questions. Test each on its own, with everything else back at the baseline.
What runs in this page
Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.
The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.
Full reference solution
This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neural_network import MLPClassifierfrom sklearn.metrics import log_loss results = []for width in [16, 32, 64]: candidate = MLPClassifier( hidden_layer_sizes=(width,), activation="relu", solver="adam", learning_rate_init=0.003, batch_size=64, random_state=42, ) for epoch in range(20): candidate.partial_fit(X_train, y_train, classes=np.arange(10)) result = { "hidden_units": width, "epochs": 20, "development_accuracy": float(candidate.score(X_dev, y_dev)), "development_loss": float(log_loss(y_dev, candidate.predict_proba(X_dev), labels=np.arange(10))), } results.append(result) print(result)import jsonimport sklearn selected = max(results, key=lambda row: (row["development_accuracy"], -row["hidden_units"]))report = { "question": "Compare one hidden-layer width after 20 epochs", "dataset": "sklearn digits 8x8", "seed": 42, "scale_divisor": 16, "train_count": len(y_train), "development_count": len(y_dev), "selection_rule": "highest development accuracy, smaller width on exact tie", "results": results, "selected": selected, "sklearn_version": sklearn.__version__, "final_test_used": False,}with open("width-experiment.json", "w") as handle: json.dump(report, handle, indent=2)print("Selected:", selected)print("Saved all candidates in width-experiment.json")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.