0%
BuildUnder the hoodabout 25 min, 8 steps

Precision and recall in plain words

Calculate precision, recall, and F1 from a worked confusion table and handle absent classes honestly.

Work here, beside the explanation

Precision and recall answer different questions about the same decisions. Build a fictional urgent-message example from exact counts, calculate both measures manually, verify them with a library, and connect the denominators to a digit confusion matrix. The example's numbers are deliberately constructed, not the output of an invented trained model.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Name all four groups

Positive means urgent in this task. A true positive is a correctly flagged urgent message; a false negative is an urgent message missed; a false positive is a routine message flagged; a true negative is a routine message correctly unflagged. Positive is the class of interest, not a judgement that the situation is good. Giving the four groups concrete names is the safest way to avoid swapping denominators later.

Python at this stageWorked example
PythonHover over a line to see an explanation
true_positive = 16false_negative = 4false_positive = 8true_negative = 72print("Correct urgent flags:", true_positive)print("Urgent messages missed:", false_negative)print("Routine messages wrongly flagged:", false_positive)print("Routine messages correctly left alone:", true_negative)print("Total:", true_positive + false_negative + false_positive + true_negative)

Run and verify that the groups sum to 100 messages.

What to look for

There are 20 actual urgent messages and 80 routine messages, totalling 100.

Make it yours

Invent another internally consistent set of four nonnegative counts and keep track of what each group means.

Build with me · 2

2. Ask about positive predictions

Precision asks: among messages the system called urgent, how many truly were urgent? The denominator is every positive prediction, including false alarms. Here sixteen correct flags out of twenty-four total flags gives about 0.667. The missed urgent messages do not belong in this denominator because they were not flagged. Precision is relevant to the reliability or review burden of the flags the system produces.

Python at this stageWorked example
PythonHover over a line to see an explanation
true_positive = 16false_negative = 4false_positive = 8true_negative = 72positive_predictions = true_positive + false_positiveprecision = true_positive / positive_predictionsprint("Correct flags:", true_positive)print("All flags:", positive_predictions)print("Precision:", precision)

Say the question in words before reading the formula, then run.

What to look for

Precision is 16/24, approximately 0.667.

Make it yours

Increase only false positives and predict what happens to precision and to the number of actual urgent messages.

Build with me · 3

3. Ask about actual positives

Recall asks: among all truly urgent messages, how many did the system find? Its denominator contains true positives plus false negatives. The numerator is the same sixteen as precision, but the denominator is twenty rather than twenty-four. Recall therefore equals 0.8. The two values differ because they describe different populations, not because one formula is more correct than the other.

Python at this stageWorked example
PythonHover over a line to see an explanation
true_positive = 16false_negative = 4false_positive = 8true_negative = 72actual_positives = true_positive + false_negativerecall = true_positive / actual_positivesprint("Urgent messages found:", true_positive)print("All urgent messages:", actual_positives)print("Recall:", recall)

Run and contrast the two denominators using their plain-language meanings.

What to look for

Recall is 16/20, or 0.8.

Make it yours

Increase only missed positives and explain how recall changes while the existing flagged-message precision stays the same.

Build with me · 4

4. Combine the two with F1

F1 is the harmonic mean of precision and recall. The equivalent count formula exposes which groups it includes: true positives, false positives, and false negatives. True negatives do not appear. F1 can summarise a task where both false alarms and missed positives matter, but it does not encode every application cost, calibration, or population shift. Report component counts when readers need to understand what changed.

Python at this stageWorked example
PythonHover over a line to see an explanation
true_positive = 16false_negative = 4false_positive = 8true_negative = 72precision = true_positive / (true_positive + false_positive)recall = true_positive / (true_positive + false_negative)f1 = 2 * precision * recall / (precision + recall)f1_from_counts = 2 * true_positive / (2 * true_positive + false_positive + false_negative)print("Precision:", precision)print("Recall:", recall)print("F1:", f1)print("Count formula:", f1_from_counts)

Run and verify that the two F1 calculations agree.

What to look for

Both formulas produce about 0.7273.

Make it yours

Change only true negatives and explain why F1 stays unchanged even though overall accuracy can change.

Build with me · 5

5. Verify the library on explicit labels

The label arrays reconstruct the same fictional counts. The explicit positive label tells the metric which class is urgent. With labels ordered [0, 1], the matrix is [[true negatives, false positives], [false negatives, true positives]]. Different class ordering changes where those cells appear, so always read labels and axis conventions. Library functions are convenient once you can independently check what they should calculate.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.metrics import precision_score, recall_score, f1_score, confusion_matrix # A fully specified fictional message example: 20 urgent and 80 routine.truth = np.array([1] * 20 + [0] * 80)predictions = np.array([1] * 16 + [0] * 4 + [1] * 8 + [0] * 72)print("Confusion matrix, rows true / columns predicted:")print(confusion_matrix(truth, predictions, labels=[0, 1]))print("Precision:", precision_score(truth, predictions, pos_label=1))print("Recall:", recall_score(truth, predictions, pos_label=1))print("F1:", f1_score(truth, predictions, pos_label=1))

Run and compare all three measurements with the manual calculations.

What to look for

The matrix is [[72, 8], [4, 16]], and the metrics match the manual values.

Make it yours

Change the class names in your notes and explain how you would select the corresponding positive label.

Build with me · 6

6. Handle an undefined denominator honestly

A model that never flags anything makes zero positive predictions. Precision divides zero by zero and is undefined, not perfect. If actual positives exist, recall is zero because none were found. Conversely, an evaluation containing no actual positives has undefined recall. Some libraries substitute a configured number and may warn; that convention must be explained rather than presented as measured success. Here None makes the missing denominator visible.

Python at this stageWorked example
PythonHover over a line to see an explanation
def ratio_or_none(numerator, denominator):    return numerator / denominator if denominator else None true_positive = 0false_positive = 0false_negative = 20precision = ratio_or_none(true_positive, true_positive + false_positive)recall = ratio_or_none(true_positive, true_positive + false_negative)print("Precision:", precision, "(None means undefined: no positive predictions)")print("Recall:", recall)

Run and read the explanatory message with the result.

What to look for

Precision is None, explicitly undefined, while recall is 0.0.

Make it yours

Construct a case with no actual positives and identify which denominator becomes zero.

Build with me · 7

7. Read digit precision and recall from a matrix

Treat one digit as positive and all other digits as negative. Under rows=true and columns=predicted, its row total counts actual positives and its column total counts positive predictions. The diagonal cell is the shared true-positive numerator. This connects binary metric language to the ten-class table without losing which examples each denominator includes. Always report support counts when interpreting a small class-specific ratio.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix matrix = confusion_matrix(y_dev, model.predict(X_dev), labels=np.arange(10))digit = 8true_positive = int(matrix[digit, digit])actual = int(matrix[digit].sum())predicted = int(matrix[:, digit].sum())print("Digit:", digit)print("Correct:", true_positive, "actual:", actual, "predicted:", predicted)print("Recall:", true_positive / actual if actual else None)print("Precision:", true_positive / predicted if predicted else None)

Run for digit eight and check the denominators against the matrix convention.

What to look for

The precision and recall are computed from actual development predictions for the chosen digit.

Make it yours

Choose a second digit and compare the counts, not only the rounded ratios.

Build with me · 8

8. Compare multiclass averaging rules

Macro averaging gives each class equal weight; weighted averaging gives classes weight proportional to their true support. These can differ substantially under imbalance. The classification report keeps class-specific values and counts visible alongside aggregates. The explicit zero_division setting is a reporting convention and is disclosed in output. A better representation can improve precision and recall together; the common threshold tradeoff concerns changing a decision rule on fixed scores, not a law forbidding overall model improvement.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import precision_score, recall_score, f1_score, classification_report predictions = model.predict(X_dev)for average in ["macro", "weighted"]:    print(average, "precision", precision_score(y_dev, predictions, average=average, zero_division=0))    print(average, "recall", recall_score(y_dev, predictions, average=average, zero_division=0))    print(average, "F1", f1_score(y_dev, predictions, average=average, zero_division=0))print(classification_report(y_dev, predictions, zero_division=0))print("Undefined metrics use 0 for this display; inspect class support before interpreting them.")

Run and compare the two averaging methods with the per-class table.

What to look for

Real multiclass precision, recall, F1, and support are reported with the averaging and undefined-value conventions stated.

Make it yours

Choose the metric set that answers a concrete task question, and explain what that set still leaves unmeasured.

A threshold is part of the procedure

Raising a positive threshold makes fewer examples qualify on the same score vector. That can trade missed positives for fewer false alarms, but precision may fluctuate on a finite sample. Choose thresholds using development examples and the consequences of errors, then evaluate the entire chosen procedure on final data.

Precision and recall are not substitutes for task context. Keep counts, class support, population conditions, and the role of manual review visible when interpreting any combined score.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import precision_score, recall_score, f1_score, classification_report predictions = model.predict(X_dev)for average in ["macro", "weighted"]:    print(average, "precision", precision_score(y_dev, predictions, average=average, zero_division=0))    print(average, "recall", recall_score(y_dev, predictions, average=average, zero_division=0))    print(average, "F1", f1_score(y_dev, predictions, average=average, zero_division=0))print(classification_report(y_dev, predictions, zero_division=0))print("Undefined metrics use 0 for this display; inspect class support before interpreting them.")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in