0%
BuildShip itabout 28 min, 8 steps

Export digit predictions without losing IDs

Read a file of unlabelled digit images, keep every id with its own image, predict, and write and check an id,label answer file.

Work here, beside the explanation

In a competition round you receive images without their answers, each with an id, and you hand back a file that pairs every id with your model's predicted digit. A correct digit attached to the wrong id counts as wrong, so the whole skill is keeping every id with its own image from start to finish. The Files tab beside this lesson holds a small practice file, digit-inputs.csv, with six images in the same layout a round uses. You will read it, check it, predict, write the answer file, and read that back to confirm nothing got mixed up.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Look at the input file

Open digit-inputs.csv in the Files tab first and look at it. It is a CSV file, plain text with values separated by commas, as in Level 2. The first line is the header, the column names: id, then pixel_0 to pixel_63. Each later line is one image: its id, then its 64 brightness values on the original 0 to 16 scale. The program reads the whole file as text, and .splitlines() cuts it into a list of lines. Lines are long, so [:60] prints only the first 60 characters of each. The id is a name for the image, not a measurement, so it must never be given to the model as if it were a pixel.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csv with open("digit-inputs.csv") as handle:    lines = handle.read().splitlines()print(lines[0][:60], "...")print(lines[1][:60], "...")print("Data rows:", len(lines) - 1)

Run, then find the same two lines in the Files tab.

What to look for

The header starts id,pixel_0,pixel_1 and the first image's line starts practice-100,0,0,8,14. There are 6 data rows.

Make it yours

Count the commas you would expect in one line: one after the id, and one between each of the 64 pixels.

Build with me · 2

2. Read the file with the csv module

Python's csv module understands the CSV layout so you do not have to split lines on commas yourself. csv.DictReader reads each line into a dictionary whose keys come from the header, so row["id"] gives a row's id and row["pixel_0"] its first pixel. reader.fieldnames is the header as a list, and list(reader) collects every row. newline="" is the setting the csv module asks for when opening a file; it stops line endings being misread. The program then builds the header it expects and compares: if the columns were missing, renamed, or in a different order, the pixels would be fed to the model in the wrong order without any error. Note that every value arrives as text, such as "8", not the number 8.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csv with open("digit-inputs.csv", newline="") as handle:    reader = csv.DictReader(handle)    header = reader.fieldnames    table = list(reader)expected_header = ["id"] + ["pixel_" + str(i) for i in range(64)]print("Header matches:", header == expected_header)print("Rows read:", len(table))print("First id:", table[0]["id"])print("Its first pixel:", table[0]["pixel_0"])

Run and check that the header matches.

What to look for

Header matches: True, 6 rows read, first id practice-100, and its first pixel is the text 0.

Make it yours

Change one name in expected_header and run again to see the check report False. Then change it back.

Build with me · 3

3. Keep identifiers out of the feature table

Now the program splits each row into two lists that stay lined up. For every row it appends the id to ids and the 64 pixel values, turned from text into numbers with float, to rows, in the same turn of the loop. So position 0 of ids always belongs to position 0 of rows. It also stops with a clear message on an empty or repeated id, or on a pixel outside 0 to 16. The ids stay as text, so an id like p007 keeps its leading zeros. Finally np.asarray(rows, dtype=float) / 16.0 makes the table of pixels and scales it once, from 0 to 16 down to 0 to 1, to match the training images.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csvwith open("digit-inputs.csv", newline="") as handle:    reader = csv.DictReader(handle)    header = reader.fieldnames    table = list(reader)expected_header = ["id"] + ["pixel_" + str(i) for i in range(64)]if header != expected_header:    raise ValueError("Expected id followed by pixel_0 through pixel_63 in that order.")ids = []rows = []for row in table:    item_id = row["id"]    if not item_id or item_id in ids:        raise ValueError("Every input needs a unique, nonempty ID.")    values = [float(row["pixel_" + str(i)]) for i in range(64)]    if min(values) < 0 or max(values) > 16:        raise ValueError("This input format requires raw brightness from 0 to 16.")    ids.append(item_id)    rows.append(values)if not rows:    raise ValueError("The prediction file has no examples.")prepared = np.asarray(rows, dtype=float) / 16.0print("IDs:", ids)print("Prepared features:", prepared.shape)print("First ID:", ids[0])print("First five pixel values:", prepared[0, :5])

Run and check that the table has 64 columns, not 65.

What to look for

Six ids print, the prepared table has shape (6, 64), and the first five values of the first image are on the 0 to 1 scale.

Make it yours

Explain what would go wrong if the id were kept as a 65th column in the table.

Build with me · 4

4. Check the prepared batch

Just before the model sees the data, check once more that it is what the model expects: one row per id, 64 columns, every value a real number (np.isfinite is False for missing or broken numbers), and everything between 0 and 1. An assert line stops the program if its condition is False, so a problem is caught here instead of turning into a neat-looking file full of wrong answers. These checks cannot tell whether a row really shows the image its id names; only keeping the order, as the loop did, protects that.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csvwith open("digit-inputs.csv", newline="") as handle:    reader = csv.DictReader(handle)    header = reader.fieldnames    table = list(reader)expected_header = ["id"] + ["pixel_" + str(i) for i in range(64)]if header != expected_header:    raise ValueError("Expected id followed by pixel_0 through pixel_63 in that order.")ids = []rows = []for row in table:    item_id = row["id"]    if not item_id or item_id in ids:        raise ValueError("Every input needs a unique, nonempty ID.")    values = [float(row["pixel_" + str(i)]) for i in range(64)]    if min(values) < 0 or max(values) > 16:        raise ValueError("This input format requires raw brightness from 0 to 16.")    ids.append(item_id)    rows.append(values)if not rows:    raise ValueError("The prediction file has no examples.")prepared = np.asarray(rows, dtype=float) / 16.0print("Rows and IDs agree:", len(ids) == prepared.shape[0])print("Prepared range:", prepared.min(), prepared.max())print("Finite values:", np.isfinite(prepared).all())assert prepared.shape == (len(ids), 64)assert np.isfinite(prepared).all()assert prepared.min() >= 0 and prepared.max() <= 1

Run and read the three checks.

What to look for

Rows and IDs agree is True, the range is 0.0 to 1.0, and all values are finite. No assert stops the program.

Make it yours

Explain why dividing by 16 a second time would pass the range check yet still be wrong.

Build with me · 5

5. Predict in the original order

model.predict returns one digit per row, in the same order as the rows it was given. zip then pairs each id with the prediction in the same position, so every id is shown with its own answer. Do not sort, drop, or hand-edit predictions here: the file should contain exactly what the model said, so its score measures the model.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csvwith open("digit-inputs.csv", newline="") as handle:    reader = csv.DictReader(handle)    header = reader.fieldnames    table = list(reader)expected_header = ["id"] + ["pixel_" + str(i) for i in range(64)]if header != expected_header:    raise ValueError("Expected id followed by pixel_0 through pixel_63 in that order.")ids = []rows = []for row in table:    item_id = row["id"]    if not item_id or item_id in ids:        raise ValueError("Every input needs a unique, nonempty ID.")    values = [float(row["pixel_" + str(i)]) for i in range(64)]    if min(values) < 0 or max(values) > 16:        raise ValueError("This input format requires raw brightness from 0 to 16.")    ids.append(item_id)    rows.append(values)if not rows:    raise ValueError("The prediction file has no examples.")prepared = np.asarray(rows, dtype=float) / 16.0predictions = model.predict(prepared)for item_id, prediction in zip(ids, predictions):    print(item_id, "->", int(prediction))print("Count agrees:", len(predictions) == len(ids))

Run and count the id-to-digit pairs.

What to look for

Six lines pair each id with a digit. In our run the digits are 9, 4, 2, 8, 1, 1.

Make it yours

Swap the first two rows in digit-inputs.csv and run again. Each id should still come out with its own digit, because each id travels with its own pixels.

Build with me · 6

6. Write the answer file

csv.writer writes rows in CSV layout, adding the commas and line endings for you. The first row written is the header, id,label, the same two columns Level 2 used and the Round 4 challenge expects. Then there is one row per image: its id and its predicted digit. int(prediction) turns NumPy's number into a plain whole number. Opening the file with "w" creates it fresh each run. The last two lines read the finished file and print it, so you see exactly what was written. After the run the editor offers digit-predictions.csv for download.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csvwith open("digit-inputs.csv", newline="") as handle:    reader = csv.DictReader(handle)    header = reader.fieldnames    table = list(reader)expected_header = ["id"] + ["pixel_" + str(i) for i in range(64)]if header != expected_header:    raise ValueError("Expected id followed by pixel_0 through pixel_63 in that order.")ids = []rows = []for row in table:    item_id = row["id"]    if not item_id or item_id in ids:        raise ValueError("Every input needs a unique, nonempty ID.")    values = [float(row["pixel_" + str(i)]) for i in range(64)]    if min(values) < 0 or max(values) > 16:        raise ValueError("This input format requires raw brightness from 0 to 16.")    ids.append(item_id)    rows.append(values)if not rows:    raise ValueError("The prediction file has no examples.")prepared = np.asarray(rows, dtype=float) / 16.0predictions = model.predict(prepared)if len(predictions) != len(ids):    raise ValueError("Prediction count does not match IDs.")with open("digit-predictions.csv", "w", newline="") as handle:    writer = csv.writer(handle)    writer.writerow(["id", "label"])    for item_id, prediction in zip(ids, predictions):        writer.writerow([item_id, int(prediction)])with open("digit-predictions.csv") as handle:    print(handle.read())

Run and read the printed file before downloading it.

What to look for

digit-predictions.csv has one header line, id,label, and six data lines, each an id and a digit.

Make it yours

If a round asked for different column names, say which single line you would change.

Build with me · 7

7. Read the answer file back and check it

Checking the file you actually wrote is safer than trusting the variables that produced it. The program reads digit-predictions.csv back with DictReader and checks three things: the header is exactly id and label; the ids come back in the same order as they went in; and every label matches the model's prediction and is a digit from 0 to 9. The list comprehensions pull one column out of the rows so the whole column can be compared in one go. These checks prove the file is intact. They say nothing about whether the predictions are correct, because this file has no answers to compare with.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csvwith open("digit-inputs.csv", newline="") as handle:    reader = csv.DictReader(handle)    header = reader.fieldnames    table = list(reader)expected_header = ["id"] + ["pixel_" + str(i) for i in range(64)]if header != expected_header:    raise ValueError("Expected id followed by pixel_0 through pixel_63 in that order.")ids = []rows = []for row in table:    item_id = row["id"]    if not item_id or item_id in ids:        raise ValueError("Every input needs a unique, nonempty ID.")    values = [float(row["pixel_" + str(i)]) for i in range(64)]    if min(values) < 0 or max(values) > 16:        raise ValueError("This input format requires raw brightness from 0 to 16.")    ids.append(item_id)    rows.append(values)if not rows:    raise ValueError("The prediction file has no examples.")prepared = np.asarray(rows, dtype=float) / 16.0predictions = model.predict(prepared)if len(predictions) != len(ids):    raise ValueError("Prediction count does not match IDs.")with open("digit-predictions.csv", "w", newline="") as handle:    writer = csv.writer(handle)    writer.writerow(["id", "label"])    for item_id, prediction in zip(ids, predictions):        writer.writerow([item_id, int(prediction)])with open("digit-predictions.csv") as handle:    print(handle.read())with open("digit-predictions.csv", newline="") as handle:    check_reader = csv.DictReader(handle)    if check_reader.fieldnames != ["id", "label"]:        raise ValueError("Wrong output header.")    exported = list(check_reader)assert [row["id"] for row in exported] == idsassert [int(row["label"]) for row in exported] == [int(value) for value in predictions]assert all(0 <= int(row["label"]) <= 9 for row in exported)print("Read-back checks passed for", len(exported), "rows.")

Run and make sure every check passes before treating the file as ready.

What to look for

Read-back checks passed for 6 rows.

Make it yours

Name one thing these checks cannot prove, such as how accurate the model is on these images.

Build with me · 8

8. Keep the input and the answers together

A predictions file is only useful if someone can tell which input and which model produced it. The final program writes a short note, prediction-run.json, naming the input file, the output file, the model and how it was trained, how many rows there were, and the first and last ids. Download it with the two CSV files and keep the three together. There is no accuracy in the note: these six images came with no answers, so nothing can be scored here.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csvwith open("digit-inputs.csv", newline="") as handle:    reader = csv.DictReader(handle)    header = reader.fieldnames    table = list(reader)expected_header = ["id"] + ["pixel_" + str(i) for i in range(64)]if header != expected_header:    raise ValueError("Expected id followed by pixel_0 through pixel_63 in that order.")ids = []rows = []for row in table:    item_id = row["id"]    if not item_id or item_id in ids:        raise ValueError("Every input needs a unique, nonempty ID.")    values = [float(row["pixel_" + str(i)]) for i in range(64)]    if min(values) < 0 or max(values) > 16:        raise ValueError("This input format requires raw brightness from 0 to 16.")    ids.append(item_id)    rows.append(values)if not rows:    raise ValueError("The prediction file has no examples.")prepared = np.asarray(rows, dtype=float) / 16.0predictions = model.predict(prepared)if len(predictions) != len(ids):    raise ValueError("Prediction count does not match IDs.")with open("digit-predictions.csv", "w", newline="") as handle:    writer = csv.writer(handle)    writer.writerow(["id", "label"])    for item_id, prediction in zip(ids, predictions):        writer.writerow([item_id, int(prediction)])with open("digit-predictions.csv") as handle:    print(handle.read())import json run_note = {    "input_file": "digit-inputs.csv", "output_file": "digit-predictions.csv",    "model": "closest example, k=3", "trained_on": "course training split, seed 42",    "rows": len(ids), "first_id": ids[0], "last_id": ids[-1],}with open("prediction-run.json", "w") as handle:    json.dump(run_note, handle, indent=2)print(json.dumps(run_note, indent=2))

Run and download prediction-run.json and digit-predictions.csv.

What to look for

The note prints with 6 rows, first id practice-100, and last id practice-105.

Make it yours

Add a "date" entry to the note, written by hand as text, so you can tell runs apart later.

Using a real round's file

When a round gives you its own file, add it in the Files tab and change the file name in the open line. Then check the round's instructions: its column names, its brightness scale, and its image size may differ from this practice file. The Round 4 challenge, for example, names its pixel columns p0 to p63. Change the expected header to match, and never turn a 28 by 28 image into an 8 by 8 one just by renaming columns.

Before predicting a whole unfamiliar file, draw one prepared row as a picture, as in the arrays lesson. A file can pass every check and still hold images that are inverted or scaled differently.

What runs in this page

Everything here, including training, runs inside your browser. The first Run of a visit loads Python and its libraries, which can take a little while; wait for the loading message to finish before deciding something is wrong. Training speed depends on your device. Closing the page stops an unfinished run, so download any file you want to keep.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)import csvwith open("digit-inputs.csv", newline="") as handle:    reader = csv.DictReader(handle)    header = reader.fieldnames    table = list(reader)expected_header = ["id"] + ["pixel_" + str(i) for i in range(64)]if header != expected_header:    raise ValueError("Expected id followed by pixel_0 through pixel_63 in that order.")ids = []rows = []for row in table:    item_id = row["id"]    if not item_id or item_id in ids:        raise ValueError("Every input needs a unique, nonempty ID.")    values = [float(row["pixel_" + str(i)]) for i in range(64)]    if min(values) < 0 or max(values) > 16:        raise ValueError("This input format requires raw brightness from 0 to 16.")    ids.append(item_id)    rows.append(values)if not rows:    raise ValueError("The prediction file has no examples.")prepared = np.asarray(rows, dtype=float) / 16.0predictions = model.predict(prepared)if len(predictions) != len(ids):    raise ValueError("Prediction count does not match IDs.")with open("digit-predictions.csv", "w", newline="") as handle:    writer = csv.writer(handle)    writer.writerow(["id", "label"])    for item_id, prediction in zip(ids, predictions):        writer.writerow([item_id, int(prediction)])with open("digit-predictions.csv") as handle:    print(handle.read())import json run_note = {    "input_file": "digit-inputs.csv", "output_file": "digit-predictions.csv",    "model": "closest example, k=3", "trained_on": "course training split, seed 42",    "rows": len(ids), "first_id": ids[0], "last_id": ids[-1],}with open("prediction-run.json", "w") as handle:    json.dump(run_note, handle, indent=2)print(json.dumps(run_note, indent=2))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in