0%
BuildShip itabout 26 min, 8 steps

Writing the method writeup

Write a short report that lets someone repeat your work: data, split, model and settings, how you chose, final score, and limits.

Work here, beside the explanation

A method writeup is a short report that lets someone else understand and repeat what you did: what the data was, how you split it, what model you trained with which settings, how you made your choices, what it scored, and where the result stops being evidence. Round 4 asks for one. This lesson builds a writeup for the closest-example model from numbers the program actually measures, so every claim in the text can be traced to a line of code. When you write your own, describe the model you really submitted, using your own measured numbers.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Say what the task and the input are

Start with the job and the data. The record dictionary, built in the setup, collects the facts as the program measures them, so the writeup cannot drift from the truth. This stage prints the task, the dataset, its size, the image shape, and the preparation. A reader needs these to rebuild the same input: 8 by 8 images divided by 16 are a different input from 28 by 28 MNIST images or photographs of handwriting.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}print("Task:", record["task"])print("Data:", record["dataset"], record["examples"], "images", record["image_shape"])print("Preparation: raw brightness divided by", record["scale_divisor"])

Run and turn the printed facts into two plain sentences.

What to look for

The task, 1,797 images of 8 by 8 pixels, and division by 16 are printed.

Make it yours

Add a sentence naming one kind of input this project has not been tested on, such as photographs of handwriting.

Build with me · 2

2. Say how the data was split

Give counts, not just "we tested it". Say which images trained the model, which ones you used to make choices, and which were kept aside until the end. If you checked a set again and again while making choices, call it development data, whatever its variable name was. Counts let a reader check that nothing was left out.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}print("Training:", record["train_count"])print("Development:", record["development_count"])print("Reserved final test:", record["final_count"])print("Split:", record["split"], "seed", record["split_seed"])

Run and write each part's job next to its count.

What to look for

1,077 training, 360 development, and 360 final-test images, split with seed 42.

Make it yours

If you knew who wrote each digit, how would you split so the final test contains only new writers?

Build with me · 3

3. Name the model and its settings

"An AI model" tells a reader nothing. Name the exact model, its settings, the library version, and the preparation. For the closest-example model that is k, the number of neighbours, and the scaling of the pixels. For a network it would be the hidden layer sizes, activation, learning rate, batch size, and number of epochs. Only describe what you actually ran.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}print("Model:", record["model"], "neighbors", record["neighbors"])print("Library version:", record["sklearn_version"])print("Input features:", X_train.shape[1])print("Known labels:", model.classes_)

Run and write one sentence naming the model, k, and the library version.

What to look for

KNeighborsClassifier with 3 neighbours, the library version, and 64 input features are printed.

Make it yours

List the settings you would report for your own network instead.

Build with me · 4

4. Say how you made your choices

Say what you compared and how you picked. This program trained one fixed model, so it says exactly that and does not pretend to have searched for the best one. If you compared widths, values of k, or numbers of epochs, list every candidate you tried, the measure you used to choose, and the result for each, including the losers. The development score that guided the choice is not the final score.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}record["selection"] = "Fixed teaching baseline k=3; no candidate search in this program"print("Selection:", record["selection"])print("Development accuracy:", record["development_accuracy"])print("A score used to choose settings is development evidence.")

Run and write a sentence about choices that matches what this program really did.

What to look for

The selection line says a fixed baseline was used with no search, next to the development accuracy of about 0.981.

Make it yours

Rewrite the selection line to describe an experiment you actually ran and kept a record of, such as the width comparison.

Build with me · 5

5. Report the final test

With the procedure fixed, the model is scored once on the final-test images, which have not been used for anything until now. Report the number correct and the total, not just a percentage, so a reader can check the division. This is the honest estimate for this procedure on this data. If you changed anything because of this number, the final test would stop being final.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}final_predictions = model.predict(X_final)correct = int(np.sum(final_predictions == y_final))record["final_correct"] = correctrecord["final_accuracy"] = correct / len(y_final)print("Final correct:", correct, "of", len(y_final))print("Final accuracy:", record["final_accuracy"])

Run once, then copy the exact count into your notes instead of a rounded percentage.

What to look for

In our run the model gets 355 of 360 final-test images right, an accuracy of about 0.986.

Make it yours

Say what extra evidence you would need to claim it works on photographs of real handwriting.

Build with me · 6

6. Show which digits it struggles with

One accuracy number can hide a weak digit. The confusion matrix and the recall for each digit, how many of the true images of that digit were found, show where the mistakes are. Rows are true digits and columns are predictions, as before. With only 35 or so images per digit, one more mistake moves a digit's recall by about 3 percentage points, so give the counts as well.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}from sklearn.metrics import confusion_matrix final_predictions = model.predict(X_final)matrix = confusion_matrix(y_final, final_predictions, labels=np.arange(10))print(matrix)for digit in range(10):    support = int(matrix[digit].sum())    recall = float(matrix[digit, digit] / support) if support else 0.0    print("Digit", digit, "support", support, "recall", round(recall, 3))

Run and pick one number from the output that a reader should know about.

What to look for

Most digits have a recall of 1.0. In our run digit 8 is lowest, about 0.914 (32 of 35), then digit 9 at about 0.944.

Make it yours

Write that observation as a sentence with its counts, such as "32 of the 35 true 8s were recognised".

Build with me · 7

7. Say where the result stops

A limitation is not an apology; it marks the edge of what the evidence covers: this dataset, this image size, one split, one set of writers. A good next step follows from something you observed, and says what you would measure. "Use a bigger model" is not a next step unless the evidence points there. If the next experiment uses the mistakes you have seen, it will need new final-test images to be judged fairly.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}limitations = [    "Evaluated packaged 8x8 digits, not arbitrary photographs or all handwriting styles.",    "One fixed split; results can change with the evaluation population.",    "No claim that neighbour vote support is calibrated real-world confidence.",]next_experiment = "Collect a planned labelled set of my own pixel drawings before inspecting predictions."for limitation in limitations:    print("Limit:", limitation)print("Next experiment:", next_experiment)

Run and adapt the limitations to the model and data you actually used.

What to look for

The limitations and one measurable next experiment are printed.

Make it yours

Replace the next experiment with one that follows from your own error analysis.

Build with me · 8

8. Put the writeup together

The final program writes the report's four paragraphs using the measured numbers from the record, so the text and the numbers cannot disagree. It saves the text to method-writeup.txt and the numbers behind it to method-evidence.json. Your own writeup should describe your own model, split, choices, results, and limits in your own words, keeping the exact numbers.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}final_predictions = model.predict(X_final)correct = int(np.sum(final_predictions == y_final))record["final_correct"] = correctrecord["final_accuracy"] = correct / len(y_final)paragraphs = [    "I built a classifier for prepared handwritten digits 0 through 9 using "    + str(record["examples"]) + " packaged eight-by-eight images. I flattened each image row by row "    + "and divided its raw brightness values by 16.",    "I used a stratified split with seed 42: " + str(record["train_count"]) + " training, "    + str(record["development_count"]) + " development, and " + str(record["final_count"])    + " reserved final-test examples. I fitted a fixed three-neighbour KNN baseline using scikit-learn "    + record["sklearn_version"] + ". This program did not search candidate settings.",    "The development accuracy was " + str(round(record["development_accuracy"], 4))    + ". After fixing the procedure, it correctly predicted " + str(correct) + " of "    + str(record["final_count"]) + " final-test images, accuracy "    + str(round(record["final_accuracy"], 4)) + ". No predictions were manually corrected.",    "These results describe this packaged-data evaluation, not arbitrary photographs or every handwriting style. "    + "My next experiment would collect a planned set of personal pixel drawings before viewing predictions, "    + "with separate development and final roles.",]report = "\n\n".join(paragraphs)print(report)with open("method-writeup.txt", "w") as handle:    handle.write(report + "\n")with open("method-evidence.json", "w") as handle:    json.dump(record, handle, indent=2)

Run, download both files, and check each number in the text against the program's output.

What to look for

method-writeup.txt and method-evidence.json are offered for download, containing this run's measured results.

Make it yours

Rewrite the paragraphs in your own voice, keeping every count and setting exactly as measured.

A useful writeup answers six questions

  1. What was the task, and what data did you use?
  2. How was the data split, with counts?
  3. Exactly what model and settings did you train?
  4. How did you choose between options?
  5. What did the final test show, with counts?
  6. Where does the evidence stop, and what would you test next?

For your own network, also give the image preparation, the network's settings, how you chose the number of epochs, and the name of your saved model file. Say whether your own drawings were quick probes or a planned test. If you corrected any predictions by hand, report that separately from the model's own score.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import jsonimport sklearn record = {    "task": "classify one prepared handwritten digit as 0 through 9",    "dataset": "sklearn digits", "examples": len(y), "image_shape": [8, 8],    "train_count": len(y_train), "development_count": len(y_dev), "final_count": len(y_final),    "split_seed": 42, "split": "stratified; approximately 60/20/20",    "model": "KNeighborsClassifier", "neighbors": 3, "scale_divisor": 16,    "development_accuracy": float(model.score(X_dev, y_dev)),    "sklearn_version": sklearn.__version__,}final_predictions = model.predict(X_final)correct = int(np.sum(final_predictions == y_final))record["final_correct"] = correctrecord["final_accuracy"] = correct / len(y_final)paragraphs = [    "I built a classifier for prepared handwritten digits 0 through 9 using "    + str(record["examples"]) + " packaged eight-by-eight images. I flattened each image row by row "    + "and divided its raw brightness values by 16.",    "I used a stratified split with seed 42: " + str(record["train_count"]) + " training, "    + str(record["development_count"]) + " development, and " + str(record["final_count"])    + " reserved final-test examples. I fitted a fixed three-neighbour KNN baseline using scikit-learn "    + record["sklearn_version"] + ". This program did not search candidate settings.",    "The development accuracy was " + str(round(record["development_accuracy"], 4))    + ". After fixing the procedure, it correctly predicted " + str(correct) + " of "    + str(record["final_count"]) + " final-test images, accuracy "    + str(round(record["final_accuracy"], 4)) + ". No predictions were manually corrected.",    "These results describe this packaged-data evaluation, not arbitrary photographs or every handwriting style. "    + "My next experiment would collect a planned set of personal pixel drawings before viewing predictions, "    + "with separate development and final roles.",]report = "\n\n".join(paragraphs)print(report)with open("method-writeup.txt", "w") as handle:    handle.write(report + "\n")with open("method-evidence.json", "w") as handle:    json.dump(record, handle, indent=2)

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in