Precision and recall in plain words
Calculate precision, recall, and F1 from a worked confusion table and handle absent classes honestly.
Work here, beside the explanation
Precision and recall answer different questions about the same decisions. Build a fictional urgent-message example from exact counts, calculate both measures manually, verify them with a library, and connect the denominators to a digit confusion matrix. The example's numbers are deliberately constructed, not the output of an invented trained model.
Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.
Build with me · 1
1. Name all four groups
Positive means urgent in this task. A true positive is a correctly flagged urgent message; a false negative is an urgent message missed; a false positive is a routine message flagged; a true negative is a routine message correctly unflagged. Positive is the class of interest, not a judgement that the situation is good. Giving the four groups concrete names is the safest way to avoid swapping denominators later.
true_positive = 16false_negative = 4false_positive = 8true_negative = 72print("Correct urgent flags:", true_positive)print("Urgent messages missed:", false_negative)print("Routine messages wrongly flagged:", false_positive)print("Routine messages correctly left alone:", true_negative)print("Total:", true_positive + false_negative + false_positive + true_negative)Run and verify that the groups sum to 100 messages.
What to look for
There are 20 actual urgent messages and 80 routine messages, totalling 100.
Make it yours
Invent another internally consistent set of four nonnegative counts and keep track of what each group means.
Build with me · 2
2. Ask about positive predictions
Precision asks: among messages the system called urgent, how many truly were urgent? The denominator is every positive prediction, including false alarms. Here sixteen correct flags out of twenty-four total flags gives about 0.667. The missed urgent messages do not belong in this denominator because they were not flagged. Precision is relevant to the reliability or review burden of the flags the system produces.
true_positive = 16false_negative = 4false_positive = 8true_negative = 72positive_predictions = true_positive + false_positiveprecision = true_positive / positive_predictionsprint("Correct flags:", true_positive)print("All flags:", positive_predictions)print("Precision:", precision)Say the question in words before reading the formula, then run.
What to look for
Precision is 16/24, approximately 0.667.
Make it yours
Increase only false positives and predict what happens to precision and to the number of actual urgent messages.
Build with me · 3
3. Ask about actual positives
Recall asks: among all truly urgent messages, how many did the system find? Its denominator contains true positives plus false negatives. The numerator is the same sixteen as precision, but the denominator is twenty rather than twenty-four. Recall therefore equals 0.8. The two values differ because they describe different populations, not because one formula is more correct than the other.
true_positive = 16false_negative = 4false_positive = 8true_negative = 72actual_positives = true_positive + false_negativerecall = true_positive / actual_positivesprint("Urgent messages found:", true_positive)print("All urgent messages:", actual_positives)print("Recall:", recall)Run and contrast the two denominators using their plain-language meanings.
What to look for
Recall is 16/20, or 0.8.
Make it yours
Increase only missed positives and explain how recall changes while the existing flagged-message precision stays the same.
Build with me · 4
4. Combine the two with F1
F1 is the harmonic mean of precision and recall. The equivalent count formula exposes which groups it includes: true positives, false positives, and false negatives. True negatives do not appear. F1 can summarise a task where both false alarms and missed positives matter, but it does not encode every application cost, calibration, or population shift. Report component counts when readers need to understand what changed.
true_positive = 16false_negative = 4false_positive = 8true_negative = 72precision = true_positive / (true_positive + false_positive)recall = true_positive / (true_positive + false_negative)f1 = 2 * precision * recall / (precision + recall)f1_from_counts = 2 * true_positive / (2 * true_positive + false_positive + false_negative)print("Precision:", precision)print("Recall:", recall)print("F1:", f1)print("Count formula:", f1_from_counts)Run and verify that the two F1 calculations agree.
What to look for
Both formulas produce about 0.7273.
Make it yours
Change only true negatives and explain why F1 stays unchanged even though overall accuracy can change.
Build with me · 5
5. Verify the library on explicit labels
The label arrays reconstruct the same fictional counts. The explicit positive label tells the metric which class is urgent. With labels ordered [0, 1], the matrix is [[true negatives, false positives], [false negatives, true positives]]. Different class ordering changes where those cells appear, so always read labels and axis conventions. Library functions are convenient once you can independently check what they should calculate.
import numpy as npfrom sklearn.metrics import precision_score, recall_score, f1_score, confusion_matrix # A fully specified fictional message example: 20 urgent and 80 routine.truth = np.array([1] * 20 + [0] * 80)predictions = np.array([1] * 16 + [0] * 4 + [1] * 8 + [0] * 72)print("Confusion matrix, rows true / columns predicted:")print(confusion_matrix(truth, predictions, labels=[0, 1]))print("Precision:", precision_score(truth, predictions, pos_label=1))print("Recall:", recall_score(truth, predictions, pos_label=1))print("F1:", f1_score(truth, predictions, pos_label=1))Run and compare all three measurements with the manual calculations.
What to look for
The matrix is [[72, 8], [4, 16]], and the metrics match the manual values.
Make it yours
Change the class names in your notes and explain how you would select the corresponding positive label.
Build with me · 6
6. Handle an undefined denominator honestly
A model that never flags anything makes zero positive predictions. Precision divides zero by zero and is undefined, not perfect. If actual positives exist, recall is zero because none were found. Conversely, an evaluation containing no actual positives has undefined recall. Some libraries substitute a configured number and may warn; that convention must be explained rather than presented as measured success. Here None makes the missing denominator visible.
def ratio_or_none(numerator, denominator): return numerator / denominator if denominator else None true_positive = 0false_positive = 0false_negative = 20precision = ratio_or_none(true_positive, true_positive + false_positive)recall = ratio_or_none(true_positive, true_positive + false_negative)print("Precision:", precision, "(None means undefined: no positive predictions)")print("Recall:", recall)Run and read the explanatory message with the result.
What to look for
Precision is None, explicitly undefined, while recall is 0.0.
Make it yours
Construct a case with no actual positives and identify which denominator becomes zero.
Build with me · 7
7. Read digit precision and recall from a matrix
Treat one digit as positive and all other digits as negative. Under rows=true and columns=predicted, its row total counts actual positives and its column total counts positive predictions. The diagonal cell is the shared true-positive numerator. This connects binary metric language to the ten-class table without losing which examples each denominator includes. Always report support counts when interpreting a small class-specific ratio.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import confusion_matrix matrix = confusion_matrix(y_dev, model.predict(X_dev), labels=np.arange(10))digit = 8true_positive = int(matrix[digit, digit])actual = int(matrix[digit].sum())predicted = int(matrix[:, digit].sum())print("Digit:", digit)print("Correct:", true_positive, "actual:", actual, "predicted:", predicted)print("Recall:", true_positive / actual if actual else None)print("Precision:", true_positive / predicted if predicted else None)Run for digit eight and check the denominators against the matrix convention.
What to look for
The precision and recall are computed from actual development predictions for the chosen digit.
Make it yours
Choose a second digit and compare the counts, not only the rounded ratios.
Build with me · 8
8. Compare multiclass averaging rules
Macro averaging gives each class equal weight; weighted averaging gives classes weight proportional to their true support. These can differ substantially under imbalance. The classification report keeps class-specific values and counts visible alongside aggregates. The explicit zero_division setting is a reporting convention and is disclosed in output. A better representation can improve precision and recall together; the common threshold tradeoff concerns changing a decision rule on fixed scores, not a law forbidding overall model improvement.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import precision_score, recall_score, f1_score, classification_report predictions = model.predict(X_dev)for average in ["macro", "weighted"]: print(average, "precision", precision_score(y_dev, predictions, average=average, zero_division=0)) print(average, "recall", recall_score(y_dev, predictions, average=average, zero_division=0)) print(average, "F1", f1_score(y_dev, predictions, average=average, zero_division=0))print(classification_report(y_dev, predictions, zero_division=0))print("Undefined metrics use 0 for this display; inspect class support before interpreting them.")Run and compare the two averaging methods with the per-class table.
What to look for
Real multiclass precision, recall, F1, and support are reported with the averaging and undefined-value conventions stated.
Make it yours
Choose the metric set that answers a concrete task question, and explain what that set still leaves unmeasured.
A threshold is part of the procedure
Raising a positive threshold makes fewer examples qualify on the same score vector. That can trade missed positives for fewer false alarms, but precision may fluctuate on a finite sample. Choose thresholds using development examples and the consequences of errors, then evaluate the entire chosen procedure on final data.
Precision and recall are not substitutes for task context. Keep counts, class support, population conditions, and the role of manual review visible when interpreting any combined score.
The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.
Full reference solution
This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split( X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.metrics import precision_score, recall_score, f1_score, classification_report predictions = model.predict(X_dev)for average in ["macro", "weighted"]: print(average, "precision", precision_score(y_dev, predictions, average=average, zero_division=0)) print(average, "recall", recall_score(y_dev, predictions, average=average, zero_division=0)) print(average, "F1", f1_score(y_dev, predictions, average=average, zero_division=0))print(classification_report(y_dev, predictions, zero_division=0))print("Undefined metrics use 0 for this display; inspect class support before interpreting them.")Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.