Bias comes from the data
Trace how data collection, labels, and design choices can create unequal errors.
A pattern is not a judgment of fairness
Suppose a club wants a model to identify good applications. It trains on previous decisions: accepted applications receive one label, rejected applications another. What does the model learn to reproduce? The historical decisions. If those decisions were unfair or measured the wrong thing, good imitation can preserve the problem.
The model does not need human feelings or intentions to cause unequal outcomes. Conversely, saying “the data did it” can hide choices made by people. The target, the labels, the features (the pieces of information the model is given), the loss, the cut-off score at which the system says yes (its decision threshold), and the way the system is used all matter.
Follow an image example
Imagine a hand-gesture recognizer trained mainly on well-lit photographs from one group of volunteers. It may work less reliably on other hands, lighting, camera positions, or backgrounds. We need measurements rather than assumptions about exactly which factor explains the gap.
Construct an illustrative test:
| Test group | Correct | Total | Accuracy |
|---|---|---|---|
| Familiar capture conditions | 90 | 100 | 90% |
| Different capture conditions | 50 | 100 | 50% |
Combined, it gets 140 of 200 right, or 70%. That single number hides a large difference. Reporting only 70% prevents the reader from seeing who or what the model struggles with.
Where the mismatch can enter
Collection: some conditions may be rare or absent. Labels: different annotators may apply different definitions. Features: a background or location may accidentally stand in for the label. Objective, meaning what training tries to maximise: aiming only for the best average accuracy may neglect a small group. Use: a model tested for a low-stakes demonstration may be unsuitable for a decision with serious consequences.
Deleting a sensitive piece of information from the data does not automatically remove it. Other features can act as proxies: stand-ins that carry much of the same information. For instance, a home neighbourhood can reveal a lot about someone's background even when that background is never recorded directly.
What a builder can do
- Define the actual task and the people affected by mistakes.
- Examine how examples and labels were obtained.
- Test on relevant groups and conditions, with enough examples to make the comparison meaningful.
- Inspect error types rather than only total accuracy.
- Improve the collection or design, then evaluate again on suitable held-out data (examples kept out of training).
- Describe limitations and provide a way to handle or challenge bad results where the application needs one.
There is no universal fairness number that resolves every tradeoff. A small sample can also exaggerate apparent differences. The responsible response is to investigate and report clearly, not to announce that one score proves a model is fair.
Apply it as a user
If an AI-generated image collection shows only one narrow kind of person in a profession, ask what representation is missing. If a recommendation or screening decision affects someone, ask what it is based on and whether a human review route exists. Do not confuse a technical-looking output with a neutral decision.
Carry this into your own project
When your recognizer works for your handwriting and fails for a friend's, do not immediately blame the friend. Compare the inputs, check how the images were prepared before training, and examine the training examples. The same habit, asking whose examples and conditions are represented, applies from a small classroom model to larger systems.
Inspect coverage before naming a fix
Suppose a fictional handwriting collection contains many neat, large digits from one classroom and very few small, slanted digits from another. A model can score well on a test drawn from the first classroom while failing the second. Counting all correct answers together may hide the unequal performance.
First state the intended users and inspect how the collection represents them. Then evaluate relevant groups with adequate labelled examples. If one group has only two test examples, report that the estimate is fragile; do not turn a single failure into a precise population conclusion.
A possible improvement is collecting useful examples from missing conditions, with appropriate permissions and consistent labels. Evening out the numbers per group does not, by itself, repair an unclear target or a harmful use. The point is to connect the evidence, intended use, and proposed change rather than assuming a technical adjustment settles every concern.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.