Two spam filters
Follow one message through a written rule and a learned spam classifier.
One message, two ways to decide
A spam filter decides whether a message belongs in an inbox or a spam folder. This example makes the previous lesson concrete. We will build a rule-based filter on paper, then describe what changes when a model learns from examples.
Our message is: “Hi, can you confirm the invoice total before Friday? The PDF is attached.” The sender is not in the address book. We do not yet know whether it is spam. A request involving money is a reason to check, not proof of fraud.
Follow the written rules
Imagine the author of the filter chose these instructions:
| Condition | Change to score |
|---|---|
| Contains the phrase FREE MONEY | Add 2 |
| Contains more than four exclamation marks | Add 1 |
| Sender is not in the address book | Add 1 |
Start the score at zero for each new message. Check each condition once. Put the message in spam if the final score is at least 3.
Our message contains neither the phrase nor the exclamation marks. The sender is unknown, so the score becomes 1. Because 1 is below 3, this filter keeps it in the inbox.
Every part of the reasoning is visible. Changing the threshold from 3 to 1 would block this message, but also many legitimate messages from new people. Fixing one missed message can create other mistakes.
What happens with a learned filter?
First collect messages labeled by people as spam or legitimate mail. The training procedure looks for features that help distinguish the two groups. A feature is information used for a prediction, such as words in the subject, the sender's history, or characteristics of an attachment.
The procedure might find that some combinations are associated with spam. That does not mean every payment request is spam. It means the examples provide evidence of a statistical pattern. Which features matter depends on the dataset and model; the combinations here are illustrative, not claims about a particular mail provider.
For the same invoice message, a trained model produces a label or score from its learned behaviour. A product then decides what to do with that output. A written threshold might place high-scoring messages in spam and send uncertain ones for extra checking.
Two mistakes to distinguish
A false positive is legitimate mail treated as spam. A false negative is spam allowed into the inbox. Which mistake is more costly depends on the use: losing a school deadline and allowing a phishing link are both problems, but of different kinds.
Test a filter on labeled messages it did not train on. Count both error types. Reporting one overall percentage can hide that the filter rarely catches a dangerous category or disproportionately blocks mail in a less common language.
Why the two approaches are often combined
An explicit “allow this trusted sender” rule is easy to understand. A learned model can detect combinations too varied for a short manually maintained list. Combining them does not eliminate errors; it gives the system different tools for different parts of the job.
New spellings such as “FR33” can defeat a narrow phrase rule. A learned filter might handle them if its training or representation supports that variation, but there is no guarantee. Neither approach is a mind reading a sender's intentions.
A small change to think through
If we add the sender to the address book, the example rule score falls from 1 to 0. We can predict that exactly. We cannot say what an unspecified learned model will do without examining or testing it. That distinction is the lesson: explicit instructions give us a direct trace, while learned behaviour must also be evaluated with evidence.
Keep the rule fixed while you inspect its mistakes
Use this toy rule: if a message contains “free”, send it to spam; otherwise keep it. “Free concert at school” goes to spam, “Claim your free prize” goes to spam, and “Urgent: collect your reward” stays in the inbox. The computer has followed the rule correctly in all three cases, even when the decision is undesirable. Correct execution and a good method are different things.
Now imagine a labelled training collection. The unwanted messages need examples with and without the word free; legitimate messages need the same variety. If all school messages lack free, the data never challenges that shortcut. A learned filter may reproduce it rather than escape it.
Choose one additional legitimate example and one additional unwanted example that would expose the shortcut. Explain the target label for each before discussing a prediction. Later, you will enter exactly this kind of input-label pair using blocks.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.