0%
BuildJust enough Python, on real digitsabout 36 min, 11 steps

Pairs, dictionaries, and files

Learn the small Python tools the model programs use: unpacking, zip, enumerate, dictionaries, one-line lists, any, f-strings, files, and picking the best record.

Work here, beside the explanation

The digit programs in the next modules use a handful of Python tools that the first four lessons did not cover. Each one is small, and each one saves you from writing a longer loop. Learn them here on tiny examples you can check by eye, so that when you meet them inside a model program they look familiar instead of mysterious. Nothing in this lesson involves a model.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Keep two values together as a pair

Round brackets with a comma between the values make a tuple, Python's name for a small, fixed group of values. Here the pair holds a true label and a predicted label for one example. You can read one item by position, just like a list: pair[0] is the first. The line true_label, predicted = pair is called unpacking. Python takes the two values out of the pair and gives each its own name, in order: the first value goes to the first name, the second to the second. The number of names on the left must match the number of values.

Python at this stageWorked example
PythonHover over a line to see an explanation
pair = (3, 8)print("The pair:", pair)print("First item:", pair[0])true_label, predicted = pairprint("True label:", true_label)print("Predicted:", predicted)

Predict all four printed lines, then run.

What to look for

The pair prints as (3, 8). True label is 3 and Predicted is 8.

Make it yours

Swap the two names on the left of the unpacking line and run again. The values still arrive in order, so the names now hold the other value.

Build with me · 2

2. Receive several values from one function

A function can hand back more than one value: write them after return with commas between them. Python packs them into a tuple, and the line that calls the function unpacks them straight into names. You will meet this pattern in the next module. The tool that splits data into a training part and a testing part returns four values at once, and the program catches them with four names on the left of a single = sign.

Python at this stageWorked example
PythonHover over a line to see an explanation
def lowest_and_highest(values):    return min(values), max(values) low, high = lowest_and_highest([4, 9, 2, 7])print("Lowest:", low)print("Highest:", high)

Run it, then call the function with a list of your own numbers.

What to look for

Lowest: 2 and Highest: 9.

Make it yours

Return a third value, len(values), without adding a third name on the left. Read the error Python gives when the counts do not match, then add the third name to fix it.

Build with me · 3

3. Walk through two lists together with zip

In the lists lesson you compared predictions with true labels using positions: for i in range(len(truth)). zip does the same job more directly. It walks through both lists at once and hands you one pair per turn: the first item of each list, then the second of each, and so on. The for line unpacks each pair into two names. Zip pairs items purely by position, so the two lists must describe the same examples in the same order. It cannot check that for you.

Python at this stageWorked example
PythonHover over a line to see an explanation
truth = [3, 8, 2, 6]predictions = [3, 1, 2, 6]correct = 0for true_label, predicted in zip(truth, predictions):    print("true", true_label, "predicted", predicted)    if true_label == predicted:        correct += 1print("Correct:", correct, "of", len(truth))

Before running, find the one position where the two lists disagree.

What to look for

Four lines, one per example, then Correct: 3 of 4.

Make it yours

Change the wrong prediction so it matches, predict the new total, and run again.

Build with me · 4

4. Number the items with enumerate

Sometimes you need an item and also where it sits, for example to report which image in a batch was misread. enumerate gives you both as a pair on every turn: the position, counting from 0 as always in Python, and the value. The for line unpacks the pair into position and label.

Python at this stageWorked example
PythonHover over a line to see an explanation
labels = [3, 8, 2, 8]for position, label in enumerate(labels):    print("Position", position, "holds label", label)

Run and check that the first position is 0, not 1.

What to look for

Four lines, positions 0 to 3, each with its label.

Make it yours

Add an if inside the loop so that only the positions holding an 8 are printed.

Build with me · 5

5. Look things up in a dictionary

A dictionary stores pairs of a key and a value, the way a printed dictionary pairs a word with its meaning. Curly brackets hold the pairs, and a colon joins each key to its value. Here each key is a symbol you might type in a drawing, and each value is the brightness that symbol stands for. brightness["#"] looks up the value stored under the key "#". A list finds things by position; a dictionary finds them by key. in asks whether a key exists. Asking for a key that is not there stops the program with a KeyError.

Python at this stageWorked example
PythonHover over a line to see an explanation
brightness = {".": 0.0, "+": 0.5, "#": 1.0}print("Whole dictionary:", brightness)print("Brightness of #:", brightness["#"])print("Brightness of .:", brightness["."])print("Is + a key?", "+" in brightness)print("Is x a key?", "x" in brightness)

Predict each printed value, then run.

What to look for

The # symbol gives 1.0, the dot gives 0.0, + is a key, and x is not.

Make it yours

Add a line brightness["*"] = 0.75 before the prints to store a new key, and print the whole dictionary again. Then try brightness["x"] and read the KeyError.

Build with me · 6

6. Count with a dictionary

A dictionary is a natural place to keep a count for each label. The loop visits every label. counts.get(label, 0) means "the count so far for this label, or 0 if we have not seen it yet". Adding 1 and storing the result under the same key updates that one count. At the end, counts.items() hands back every key with its value as a pair, which the second loop unpacks. In the functions lesson you counted one label at a time; this counts every label in one pass through the list.

Python at this stageWorked example
PythonHover over a line to see an explanation
labels = [3, 8, 2, 8, 3, 8, 2]counts = {}for label in labels:    counts[label] = counts.get(label, 0) + 1print("Counts:", counts)for label, count in counts.items():    print("Label", label, "appears", count, "times")

Count the 3s, 8s, and 2s yourself before running.

What to look for

Label 3 appears 2 times, label 8 appears 3 times, label 2 appears 2 times.

Make it yours

Add more labels to the list, including a new one such as 5, and check that the counts update.

Build with me · 7

7. Build a list in one line

Programs often make a new list by doing the same thing to every item of an old list. The first part uses the loop you already know: start with an empty list and append each result. The line after it does exactly the same work in one line. This is called a list comprehension. Read it left to right as "value divided by 16, for each value in raw". The last line combines a comprehension with the dictionary: "the brightness of each symbol, for each symbol in the row". A string can be looped over one character at a time, so each symbol is one character. Later you will see one comprehension inside another, which turns a whole drawing into numbers one row at a time. It is this same idea used twice.

Python at this stageWorked example
PythonHover over a line to see an explanation
raw = [0, 4, 8, 16]scaled_by_loop = []for value in raw:    scaled_by_loop.append(value / 16)scaled = [value / 16 for value in raw]print("Loop version:", scaled_by_loop)print("One-line version:", scaled)row = ".#+"brightness = {".": 0.0, "+": 0.5, "#": 1.0}print("Symbols to numbers:", [brightness[symbol] for symbol in row])

Check that the loop version and the one-line version print the same list.

What to look for

Both lists are [0.0, 0.25, 0.5, 1.0]. The symbols become [0.0, 1.0, 0.5].

Make it yours

Change the row of symbols to a pattern of your own and predict its numbers before running.

Build with me · 8

8. Ask whether any item breaks a rule

any(...) answers the question "is this true for at least one item?". all(...) answers "is this true for every item?". Inside the brackets is a condition written in the same style as a comprehension: "the length of the row is not 4, for each row in rows". Programs use these checks to test data before it reaches a model, for example that every row of a drawing has the right width.

Python at this stageWorked example
PythonHover over a line to see an explanation
rows = ["..##", ".#.#", "####"]print("Any row not 4 long?", any(len(row) != 4 for row in rows))print("All rows 4 long?", all(len(row) == 4 for row in rows))rows.append("#")print("After adding a short row:", any(len(row) != 4 for row in rows))

Predict the three answers before running.

What to look for

False, then True. After the short row is added, the any check becomes True.

Make it yours

Change the short row to four characters and confirm the check goes back to False.

Build with me · 9

9. Put values inside text with an f-string

Putting the letter f straight before the opening quote makes an f-string. Inside it, anything in curly brackets is worked out and its value is written into the text. That includes a calculation such as correct / total. F-strings are handy when a line must have an exact shape, such as p001,7 in a predictions file, where the spaces that print adds between values would get in the way.

Python at this stageWorked example
PythonHover over a line to see an explanation
name = "Ada"correct = 17total = 20print(f"{name} got {correct} of {total} right.")print(f"Accuracy: {correct / total}")row_id = "p001"prediction = 7print(f"{row_id},{prediction}")

Run and compare each printed line with the code that made it.

What to look for

Ada got 17 of 20 right., then Accuracy: 0.85, then p001,7.

Make it yours

Write an f-string that prints a fictional name of your choice and a score out of 20.

Build with me · 10

10. Save a dictionary to a file and read it back

A running program forgets everything when it ends, so anything worth keeping goes into a file. open("settings.json", "w") opens a file for writing (the w), creating it if it does not exist. The with line gives the open file a name, handle, for the indented lines below it, and closes the file properly when they finish. JSON is a common text format for dictionaries and lists. json.dump writes the dictionary into the file as text; indent=2 just makes that text easier to read. The second with block opens the same file for reading (no w), and json.load turns the text back into a dictionary. After the run, the editor offers the file for download.

Python at this stageWorked example
PythonHover over a line to see an explanation
import json settings = {"name": "first run", "hidden_units": 32, "epochs": 20}with open("settings.json", "w") as handle:    json.dump(settings, handle, indent=2)with open("settings.json") as handle:    loaded = json.load(handle)print("Read back:", loaded)print("Hidden units:", loaded["hidden_units"])print("Same as saved?", loaded == settings)

Run, then find settings.json in the output area and download it if you like.

What to look for

The dictionary read back is the same as the one saved, hidden units is 32, and Same as saved? is True.

Make it yours

Change one setting, run again, and open the downloaded file to see your dictionary written as plain text.

Build with me · 11

11. Pick the best record with a rule

max finds the largest item. With a list of dictionaries, Python has to be told what "largest" should mean, and key= supplies that rule. lambda result: result["accuracy"] is a tiny function written in one line: given one record, it returns that record's accuracy. max applies the rule to every record and returns the whole record whose rule gave the biggest number. You will use exactly this later to choose between model settings you have measured.

Python at this stageWorked example
PythonHover over a line to see an explanation
results = [    {"width": 16, "accuracy": 0.95},    {"width": 32, "accuracy": 0.97},    {"width": 64, "accuracy": 0.96},]best = max(results, key=lambda result: result["accuracy"])print("Best record:", best)print("Its width:", best["width"])

Predict which record wins before running.

What to look for

The record with width 32 and accuracy 0.97 wins.

Make it yours

Use min with the rule result["width"] to pick the smallest width instead.

What you can now read

These tools appear in the model programs from here on. Unpacking catches the several values a function returns, such as the four parts of a data split. zip and enumerate keep predictions lined up with the right examples. Dictionaries hold settings, counts, and results by name. Comprehensions turn a drawing into numbers. any and all check data before a model sees it. F-strings write lines with an exact shape. Files and JSON keep results after the page closes. key= picks the best of several measured results.

When one of these turns up inside a longer program, find the small example here that matches it and read the long line the same way.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
results = [    {"width": 16, "accuracy": 0.95},    {"width": 32, "accuracy": 0.97},    {"width": 64, "accuracy": 0.96},]best = max(results, key=lambda result: result["accuracy"])print("Best record:", best)print("Its width:", best["width"])

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in