Kramizo
Log inSign up free
Home › Kramizo AI Literacy › Bias and fairness in AI systems
Kramizo · · AI Literacy · Revision Notes

Bias and fairness in AI systems

2,216 words · Last updated October 2026

⚡
Ready to practise? Test yourself on Bias and fairness in AI systems with instantly-marked questions.
Practice now →

What you'll learn

  • The two different things the word bias means, and why confusing them causes arguments
  • The five stages at which bias enters a system — only one of which is the dataset
  • Historical bias: how a perfectly accurate model of an unfair world reproduces the unfairness
  • Proxy variables, and why deleting a sensitive field does not make a system fair
  • Feedback loops: how a deployed system changes the data it will later be trained on
  • Why there is more than one reasonable definition of fairness, and why they can be impossible to satisfy at once
  • Why you may have to collect sensitive data in order to prove you are not discriminating

Key terms and definitions

Term Meaning
Statistical bias A systematic difference between what a system estimates and the true value
Social bias A system treating groups of people unequally in ways regarded as unjust
Historical bias Bias present because the data accurately records a world that was already unfair
Proxy variable A field that stands in for a sensitive characteristic without naming it, such as a postcode
Fairness through unawareness The mistaken belief that removing a sensitive field makes a system fair
Feedback loop A system's own outputs shaping the data used to judge or retrain it
Aggregation bias Using one model for everyone when groups genuinely differ
Allocation harm Being denied something — a loan, a job, a place
Representation harm Being depicted or categorised demeaningly, even with nothing allocated
Fairness metric A specific, measurable definition of what fair treatment means

Core concepts

Two meanings of one word

In statistics, bias is a neutral technical term: a systematic error in one direction. A scale that always reads 200 g heavy is biased, and nothing about justice is involved.

In ordinary use, bias means unfair treatment of people.

Both meanings are live in discussions of AI, and they come apart in a way that matters: a system can be statistically excellent and socially unjust at the same time. A model that predicts a biased world very accurately has low statistical bias and may do serious harm. Keeping the two apart is the first step to arguing about this clearly.

Bias does not only come from the data

"Biased data" is the usual explanation, and it is one of five places bias enters. Naming the stage is what makes a criticism useful:

1 — Defining the problem. Someone chose what to predict. A company wanting to find good employees cannot measure "good employee", so it picks something it can measure — perhaps who stayed three years, or who got high performance ratings. If those ratings were awarded unfairly, the target itself carries the unfairness, and no amount of careful modelling fixes it.

2 — Choosing who is in the data. A sample that covers one group well and another thinly produces a system that works well for the first group.

3 — Applying the labels. Labels record human judgements. Two people marking the same cases may disagree systematically, and whichever view dominated the labelling becomes the model's view.

4 — Selecting features. What a model is allowed to notice is a decision. Including a feature closely tied to a protected characteristic has consequences; so does excluding one that would have explained a difference innocently.

5 — Deploying the output. The same prediction can be used to offer help or to withhold it. A model flagging pupils at risk of falling behind could direct support towards them or quietly lower expectations of them. The model is identical; the fairness is not.

Historical bias: accurate and unjust

This is the hardest idea in the topic and the one most worth getting.

Suppose a company trains a model on a decade of its own hiring decisions, and for that decade it promoted men more often than women for reasons unrelated to performance. The model learns the pattern faithfully. It is accurate, in the sense that it predicts the historical decisions well. It is also unfair, because the pattern it learned was unfair.

There is no modelling error here to find. The system has done its job correctly. The unfairness came in with the examples, and it emerges looking like objectivity — which is what makes it dangerous. A spreadsheet of past decisions does not announce which of them were unjust.

Proxy variables and why deleting a field fails

The obvious fix is to remove the sensitive field: do not tell the model anyone's sex or ethnicity, and it cannot discriminate.

This does not work, and the reason is proxy variables — fields that carry much of the same information without naming it:

  • A postcode can track ethnicity and income closely, because of where people live
  • The name of a school can track social background
  • A first name can indicate sex or ethnic background
  • Gaps in employment history can track parental leave, and so sex
  • Membership of certain clubs or societies can track sex or religion

A model is good at exactly this: finding combinations of available features that predict the target. If sex predicts the historical outcome and sex is absent but correlated with five present features, the model reconstructs the effect from those five.

This mistake has a name — fairness through unawareness — and it fails twice over. It does not remove the bias, and by deleting the field it removes your ability to measure whether the system is biased.

Feedback loops: the system shapes its own evidence

A model's predictions often influence the world the next batch of data comes from.

The clearest example is predictive policing. A model trained on recorded crime sends more patrols to a particular area. More patrols mean more offences are recorded there — not necessarily more committed. That record becomes next year's training data, which confirms the original prediction.

The loop is self-reinforcing, and from inside it the system looks increasingly well validated. Its own effects are being read as evidence that it was right.

The same shape appears elsewhere:

  • A hiring model trained on who succeeded at a company, where who succeeded depended on who was hired and supported
  • A recommendation system judged on what users clicked, where users could only click what it chose to show them
  • A credit model whose refusals mean no repayment history is ever generated for those applicants

The warning sign is a system evaluated on data its own decisions helped produce.

Fairness has more than one definition

People often talk as if fairness were a single property a system either has or lacks. It is not. Here are three reasonable definitions of a hiring tool being fair:

  • Equal selection rates — it recommends the same proportion of applicants from each group
  • Equal error rates — it is wrong about equally many people in each group
  • Equal treatment of equals — two applicants with the same relevant qualities get the same outcome

Each sounds obviously right. The difficulty is that they cannot generally all hold at once: when the groups differ in the underlying data, satisfying one definition mathematically forces you to violate another. This is a proven result, not a limitation of current technology.

The consequence is important and often resisted: choosing a fairness definition is a value judgement, not a technical decision. It belongs to the people affected and those accountable, not to whoever is writing the code. A system described simply as "fair" has not told you which definition was chosen or what was given up.

Who is harmed, and how

Two kinds of harm are worth distinguishing:

  • Allocation harms — somebody does not get something: a loan, a job, a place at a school, a benefit
  • Representation harms — somebody is depicted or categorised demeaningly: an image generator that produces only one kind of person for "nurse", a translation system that assigns genders to professions, an autocomplete that finishes a group's name with an insult

Allocation harms are easier to count and attract most attention. Representation harms shape expectations at scale and are easy to dismiss as trivial by anyone they do not describe.

You cannot fix what you do not measure

A final, awkward conclusion. To know whether a system treats groups differently, you have to measure its behaviour by group — which means holding data about people's characteristics.

Organisations are often reluctant to collect that data, sometimes for good privacy reasons and sometimes because not knowing is comfortable. But a single overall accuracy figure cannot reveal a gap, and "we don't record ethnicity, so we can't be biased" has the logic backwards: not recording it means you cannot tell.

Worked examples

Example 1: Explaining a failed fix (4 marks)

A company removes the sex of applicants from its hiring model's input and states the model is now fair. Explain why this is wrong.

  • Other features can act as proxies, carrying much the same information — employment gaps, school attended, club membership (1 mark)
  • A model finds whatever combination of available features predicts the target, so it reconstructs the effect (1 mark)
  • This is fairness through unawareness, and it does not remove the bias (1 mark)
  • Removing the field also removes the ability to measure whether the model treats the groups differently (1 mark)

Example 2: Accurate and unfair (3 marks)

A model trained on ten years of a company's promotion decisions reproduces a pattern of promoting men more often. Explain how the model can be both accurate and unfair.

  • It is accurate in that it predicts the historical decisions correctly (1 mark)
  • Those decisions were themselves unfair, so the pattern it learned is unfair — this is historical bias (1 mark)
  • There is no modelling error to find; the unfairness entered with the data and now looks like objectivity (1 mark)

Example 3: Identifying a loop (4 marks)

A model sends extra patrols to areas it predicts will have more crime. A year later its predictions match the recorded data closely, and the developers call this validation. Explain the flaw.

  • More patrols in an area means more offences are recorded there, not necessarily more committed (1 mark)
  • That record becomes the next round of training and evaluation data (1 mark)
  • The system is being judged against data its own decisions helped produce, so the agreement is a feedback loop rather than evidence (1 mark)
  • Validation requires data not influenced by the system's own deployment (1 mark)

Common mistakes and how to avoid them

  • Treating "biased data" as the whole story. Name the stage: the target chosen, the sample, the labels, the features, or how the output is used.
  • Assuming accurate means fair. A faithful model of an unfair history is both accurate and unjust.
  • Thinking deleting a sensitive field fixes anything. Proxies survive, and you lose the ability to measure.
  • Saying a system is "fair" without saying which definition. The main definitions conflict mathematically.
  • Treating the fairness choice as technical. It is a value judgement for those accountable and affected.
  • Dismissing representation harms because nothing was allocated.
  • Claiming "we don't collect ethnicity so we can't be biased". Not measuring is not the same as not happening.
  • Calling a model objective because it is mathematical. The maths is neutral; the target, sample, labels and use are all choices.

Using this in practice

When you meet a claim that a system is fair:

  1. Which definition? Equal selection rates, equal error rates, or equal treatment of equals — and what was traded away?
  2. Measured how, and broken down by what? An overall figure cannot show a gap.
  3. Where did the target come from? If the thing predicted was itself awarded unfairly, nothing downstream repairs it.
  4. Could its own decisions be producing its evidence?
  5. Who is affected, and can they appeal? A fair-looking system with no route to challenge it is not one.

Quick revision summary

  • Statistical bias is systematic error; social bias is unjust treatment — a system can have little of the first and plenty of the second
  • Bias enters at five stages: the target chosen, who is sampled, how labels are applied, which features are used, and how the output is deployed
  • Historical bias means an accurate model of an unfair world is itself unfair, with no modelling error to find
  • Deleting a sensitive field fails: proxies carry the information, and you lose the ability to measure — this is fairness through unawareness
  • Feedback loops let a system's own decisions generate the evidence that it was right
  • The main fairness definitions conflict mathematically; choosing between them is a value judgement, not a technical one
  • Distinguish allocation harms from representation harms; both matter
  • You cannot detect a gap you do not measure, so fairness work usually requires the very data people hesitate to collect

Bias and fairness in AI systems: common questions

What are the most common mistakes in Bias and fairness in AI systems?

Treating "biased data" as the whole story: Name the stage: the target chosen, the sample, the labels, the features, or how the output is used. Assuming accurate means fair: A faithful model of an unfair history is both accurate and unjust. Thinking deleting a sensitive field fixes anything: Proxies survive, and you lose the ability to measure.

Where can I practise Bias and fairness in AI systems questions for free?

Kramizo has free Kramizo AI Literacy practice questions on Bias and fairness in AI systems, each marked instantly with a full explanation. No card is required.

Free for students

Lock in Bias and fairness in AI systems with real exam questions.

Free instantly-marked Kramizo AI Literacy practice — 45 questions a day, no card required.

Try a question →See practice bank