0%
BuildUnder the hoodabout 25 min, 8 steps

The accuracy trap and the baseline you must beat

Expose a high-accuracy classifier that misses the important class and decide which additional evidence to report.

Work here, beside the explanation

A high accuracy can be correct arithmetic and weak evidence for the actual task. Reconstruct a fictional rare-event example, compare a constant prediction with a detector that catches positives, and inspect which counts one percentage hides. Then apply the same questions to the real digit dataset. These invented transaction labels teach metric reasoning; they are not a deployed fraud system or financial advice.

Each numbered stage below shows a complete program. Try this stage copies it into the editor beside the article, including every line it needs from earlier stages, so it works even after you reload the page. Read the program first, predict what it will print, then press Run. Loading a stage replaces what is in the editor; Undo brings your own version back.

Build with me · 1

1. Count the actual class distribution

The example contains 995 ordinary transactions and five fraudulent ones. Ordinary is represented by zero and fraud by one. Before computing a model score, inspect which class dominates the evaluation. Accuracy gives every example equal weight, so a common class can overwhelm the contribution of a rare class. The distribution alone does not tell you the real cost of either kind of error.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix # Fictional evaluation: 995 ordinary transactions, then five fraudulent ones.truth = np.array([0] * 995 + [1] * 5)baseline_predictions = np.zeros(len(truth), dtype=int)labels, counts = np.unique(truth, return_counts=True)print("Class counts:", dict(zip(labels.tolist(), counts.tolist())))print("Total:", len(truth))

Run and calculate the fraud proportion from the counts.

What to look for

There are 1,000 examples and fraud accounts for 0.5%.

Make it yours

Choose another fictional class balance and predict how a constant-majority accuracy would change.

Build with me · 2

2. Score a constant ordinary prediction

Predicting ordinary for every row gets 995 decisions right and misses all five fraud cases. The 99.5% accuracy is numerically correct. Calling it proof of useful fraud detection would be misleading because fraud recall is zero. For an honest real baseline, the constant class must be chosen from training data, not selected by inspecting evaluation labels. This fictional example assumes ordinary is also the training majority.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix # Fictional evaluation: 995 ordinary transactions, then five fraudulent ones.truth = np.array([0] * 995 + [1] * 5)baseline_predictions = np.zeros(len(truth), dtype=int)print("Accuracy:", accuracy_score(truth, baseline_predictions))print("Fraud recall:", recall_score(truth, baseline_predictions, pos_label=1))print("Predicted fraud count:", int(baseline_predictions.sum()))

Run and explain how the high accuracy and zero fraud recall can both be true.

What to look for

Accuracy is 0.995 and fraud recall is 0.0.

Make it yours

State which task question this baseline fails despite its high overall score.

Build with me · 3

3. Construct a candidate with explicit errors

The candidate catches four fraudulent transactions, misses one, and incorrectly flags twenty ordinary transactions. The remaining 975 ordinary transactions are correctly unflagged. The four groups still total 1,000. Writing the table prevents hiding the new false alarms behind improved recall or hiding the detected frauds behind reduced accuracy. Both effects occurred and belong in the comparison.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix # Fictional evaluation: 995 ordinary transactions, then five fraudulent ones.truth = np.array([0] * 995 + [1] * 5)baseline_predictions = np.zeros(len(truth), dtype=int)# Candidate flags 20 ordinary cases and catches four of the five frauds.candidate_predictions = np.array([1] * 20 + [0] * 975 + [1] * 4 + [0])print(confusion_matrix(truth, candidate_predictions, labels=[0, 1]))print("True negatives:", 975)print("False positives:", 20)print("False negatives:", 1)print("True positives:", 4)

Run and map each matrix cell to one of the four named groups.

What to look for

The matrix is [[975, 20], [1, 4]] under true-row, predicted-column ordering.

Make it yours

Check that the true-class row totals still match the original distribution.

Build with me · 4

4. Compare several measures

The candidate's accuracy is 97.9%, below the constant baseline. Its fraud recall is 80%, and its precision is four correct flags among twenty-four, about 16.7%. Whether that is useful depends on what a flag triggers, how costly review is, and how costly a missed fraud is. Neither the lower accuracy nor the higher recall alone decides suitability. The baseline makes no positive predictions, so its precision requires an explicit undefined-value convention.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix # Fictional evaluation: 995 ordinary transactions, then five fraudulent ones.truth = np.array([0] * 995 + [1] * 5)baseline_predictions = np.zeros(len(truth), dtype=int)# Candidate flags 20 ordinary cases and catches four of the five frauds.candidate_predictions = np.array([1] * 20 + [0] * 975 + [1] * 4 + [0])for name, predictions in [("constant ordinary", baseline_predictions), ("candidate", candidate_predictions)]:    print(name)    print("  accuracy", accuracy_score(truth, predictions))    print("  fraud precision", precision_score(truth, predictions, pos_label=1, zero_division=0))    print("  fraud recall", recall_score(truth, predictions, pos_label=1))print("The baseline precision is undefined; 0 above is a declared display replacement.")

Run and explain one improvement and one cost of the candidate.

What to look for

The candidate has accuracy 0.979, recall 0.8, and precision about 0.167.

Make it yours

Describe an application condition that could make twenty false alarms acceptable, and one that could make them unacceptable.

Build with me · 5

5. Make a hypothetical cost rule explicit

These costs are invented teaching choices, not measured money or an established policy. The calculation makes a preference explicit: a missed positive is assigned one hundred units and a false alarm one. Under that assumption, catching four positives can outweigh the twenty false alarms. Changing the assumptions may change the decision. Real deployments require evidence about consequences, capacity, affected people, and review, not arbitrary numbers copied from a tutorial.

Python at this stageWorked example
PythonHover over a line to see an explanation
import numpy as npfrom sklearn.metrics import accuracy_score, precision_score, recall_score, confusion_matrix # Fictional evaluation: 995 ordinary transactions, then five fraudulent ones.truth = np.array([0] * 995 + [1] * 5)baseline_predictions = np.zeros(len(truth), dtype=int)# Candidate flags 20 ordinary cases and catches four of the five frauds.candidate_predictions = np.array([1] * 20 + [0] * 975 + [1] * 4 + [0])miss_cost = 100false_alarm_cost = 1for name, predictions in [("constant ordinary", baseline_predictions), ("candidate", candidate_predictions)]:    false_negatives = int(np.sum((truth == 1) & (predictions == 0)))    false_positives = int(np.sum((truth == 0) & (predictions == 1)))    cost = miss_cost * false_negatives + false_alarm_cost * false_positives    print(name, "hypothetical cost units", cost)

Run and calculate each total independently from the confusion counts.

What to look for

The baseline has 500 hypothetical units and the candidate has 120 under the stated assumptions.

Make it yours

Choose a different fictional false-alarm cost and see when the ranking changes.

Build with me · 6

6. Reconstruct a smaller message example

In this fictional hundred-message example, ten are spam and ninety ordinary. The candidate catches eight spam messages, misses two, and flags six ordinary messages. Its 92% accuracy improves on a 90% constant-ordinary baseline, but that two-percentage-point change hides a much larger change in spam detection plus six new false alarms. Reconstructing counts exposes the operational meaning that a single percentage conceals.

Python at this stageWorked example
PythonHover over a line to see an explanation
true_positive = 8false_negative = 2false_positive = 6true_negative = 84total = true_positive + false_negative + false_positive + true_negativeprint("Total:", total)print("Accuracy:", (true_positive + true_negative) / total)print("Spam recall:", true_positive / (true_positive + false_negative))print("Spam precision:", true_positive / (true_positive + false_positive))print("Always-ordinary accuracy:", (true_negative + false_positive) / total)

Run and check that the four groups add to one hundred.

What to look for

Accuracy is 0.92, spam recall 0.8, precision about 0.571, and baseline accuracy 0.9.

Make it yours

Write a short comparison that mentions both detected spam and false alarms.

Build with me · 7

7. Inspect the real digit distribution

Digits are far more balanced than the rare-event example, so overall accuracy is informative. Balance still does not guarantee equal performance on each class or handwriting style. The constant baseline chooses its label from training, then is evaluated on development. This procedure avoids choosing a constant after seeing evaluation answers. For your own drawings, a planned count per digit is more interpretable than repeatedly redrawing until the recognizer succeeds.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.dummy import DummyClassifier labels, counts = np.unique(y_dev, return_counts=True)baseline = DummyClassifier(strategy="most_frequent")baseline.fit(X_train, y_train)print("Development class counts:", dict(zip(labels.tolist(), counts.tolist())))print("Training-chosen constant baseline:", baseline.score(X_dev, y_dev))print("KNN development:", model.score(X_dev, y_dev))

Run and compare the actual class counts and baseline with the fictional imbalance.

What to look for

The real development distribution and both actual model scores are shown.

Make it yours

Name a subgroup that digit labels alone do not measure, such as writer or capture conditions.

Build with me · 8

8. Report accuracy with context

A useful report keeps the evaluation count, class distribution, baseline method, split role, and important class-specific errors alongside accuracy. If a person reviews or corrects outputs, report that workload and separate the combined process from the model alone. High scores can also reflect duplicate leakage, unrepresentative easy inputs, or repeated tuning on evaluation data. Clear boundaries and records help readers assess what a percentage really supports.

Python at this stageWorked example
PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.dummy import DummyClassifierfrom sklearn.metrics import classification_report, confusion_matrix baseline = DummyClassifier(strategy="most_frequent")baseline.fit(X_train, y_train)predictions = model.predict(X_dev)print("Evaluation role: development; count:", len(y_dev))print("Baseline: most frequent training class; score:", baseline.score(X_dev, y_dev))print("KNN accuracy:", model.score(X_dev, y_dev))print("Rows true, columns predicted:")print(confusion_matrix(y_dev, predictions, labels=np.arange(10)))print(classification_report(y_dev, predictions, zero_division=0))print("Inspect per-class support; undefined metrics use 0 only for display.")

Run and write one supported conclusion plus one limitation from the complete evidence.

What to look for

Actual confusion counts and per-class metrics accompany the model's development accuracy.

Make it yours

Choose a metric set for a concrete use case and explain why one overall score would omit something important.

Ask what the score could hide

No single metric determines the correct tradeoff without the application context. Define the task, identify which errors matter, choose measurements that answer those questions, and preserve an honest evaluation protocol. A classroom recognizer and a consequential decision system require different evidence even when both can print an accuracy percentage.

The complete reference is folded away below. Compare it with your work after trying the steps; changing a personal choice such as a name, a colour, or a display threshold can produce a different valid program.

Full reference solution

This is the final complete program built in the walkthrough. All its setup is included. Personal choices may differ in your own version; model scores are measured when you run, not promises about a future dataset.

PythonHover over a line to see an explanation
from sklearn.datasets import load_digitsimport numpy as np digits = load_digits()X = digits.data.astype("float64") / 16.0y = digits.targetfrom sklearn.model_selection import train_test_split # Reserve the final test before trying model settings.X_pool, X_final, y_pool, y_final = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)X_train, X_dev, y_train, y_dev = train_test_split(    X_pool, y_pool, test_size=0.25, random_state=42, stratify=y_pool)from sklearn.neighbors import KNeighborsClassifier model = KNeighborsClassifier(n_neighbors=3)model.fit(X_train, y_train)from sklearn.dummy import DummyClassifierfrom sklearn.metrics import classification_report, confusion_matrix baseline = DummyClassifier(strategy="most_frequent")baseline.fit(X_train, y_train)predictions = model.predict(X_dev)print("Evaluation role: development; count:", len(y_dev))print("Baseline: most frequent training class; score:", baseline.score(X_dev, y_dev))print("KNN accuracy:", model.score(X_dev, y_dev))print("Rows true, columns predicted:")print(confusion_matrix(y_dev, predictions, labels=np.arange(10)))print(classification_report(y_dev, predictions, zero_division=0))print("Inspect per-class support; undefined metrics use 0 only for display.")

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in