Kramizo
Log inSign up free
HomeAQA GCSE MathematicsComparing distributions using statistical measures and diagrams
AQA · GCSE · Mathematics · Revision Notes

Comparing distributions using statistical measures and diagrams

1,704 words · Last updated September 2026

Ready to practise? Test yourself on Comparing distributions using statistical measures and diagrams with instantly-marked questions.
Practice now →
Quick answer

Distributionthe pattern of how a set of values is spread out.

A fair comparison needs two statements: one about the average and one about the spread. Either alone scores half the marks.

Comparing Distributions — AQA GCSE Maths Revision Notes

What you'll learn

This topic covers comparing two sets of data fairly. By the end of this guide you should be able to choose an appropriate average and an appropriate measure of spread, calculate both, and write a comparison that earns full marks.

You should also be able to compare distributions shown as box plots, cumulative frequency curves or frequency polygons, decide when the median is a better choice than the mean, and explain the effect an outlier has on each measure.

The organising idea is that a fair comparison always needs two statements: one about average and one about spread. Two groups can share an identical mean while one is wildly inconsistent and the other completely predictable — a class averaging 50% where everyone scored between 48 and 52 is a very different class from one averaging 50% where marks ran from 10 to 90. Neither figure alone describes the data. Examiners mark this topic accordingly: one statement about the centre, one about the spread, and both written in context.

Key terms and definitions

Distribution — the pattern of how a set of values is spread out.

Average (measure of central tendency) — the mean, median or mode.

Spread (measure of dispersion) — the range or the interquartile range.

Range — largest value minus smallest value.

Interquartile range (IQR) — the upper quartile minus the lower quartile, covering the middle half.

Outlier — a value far away from the rest of the data.

Skewed — a distribution with values bunched at one end and a tail at the other.

In context — written in terms of what the data actually measures, not in bare numbers.

Core concepts

The two-statement rule

A comparison must say something about the centre and something about the spread.

"Group A had a higher median mark" describes the centre only. "Group A had a smaller interquartile range" describes the spread only. Either alone scores half the marks available; together they score full.

The two say genuinely different things: the first is about which group did better, the second about how consistent each group was.

Choosing the average

The median is the middle value, and it is the safe choice when the data is skewed or contains outliers, because extreme values barely move it.

The mean uses every value, which makes it a good summary of a symmetrical distribution with no unusual values — but that same property makes it vulnerable, since one extreme figure drags it noticeably.

The mode is the most common value. It is the only average available for categorical data such as favourite colour, and it is useful when the question asks what is most typical.

A worked illustration: for the salaries 18, 20, 21, 22 and 200 thousand, the median is 21 and the mean is 56.2. The median describes what most people earn; the mean describes nobody.

Choosing the measure of spread

The range is quick to work out but depends entirely on the two most extreme values, so a single outlier distorts it completely.

The interquartile range ignores the top and bottom quarters and measures the middle half, which makes it far more robust. It is the better choice whenever outliers are present, and it is the measure paired with the median.

For the salaries above, the range is 182 but the interquartile range is only about 2 — and the interquartile range is plainly the better description of the group.

A smaller spread means more consistent data. That is the phrase to use.

Pairing the measures

Choose the pair that suits the data, and keep them together.

Median with interquartile range when the data is skewed or contains outliers.

Mean with range when the data is symmetrical and well behaved.

Mixing them — comparing one group's mean with another's median — is not a fair comparison at all, and questions sometimes include this as the thing to criticise.

Comparing from box plots

Box plots are built for this. The line inside the box gives the median, and the width of the box gives the interquartile range.

So the comparison is visual: which box sits further along the scale, and which box is narrower.

A box plot does not show the mean, and it does not show how many values there are, so neither can be part of the comparison.

Comparing from other displays

Cumulative frequency curves give the median and quartiles by reading at half, a quarter and three-quarters of the total frequency, and the comparison then proceeds exactly as with box plots. A steeper curve indicates data bunched more tightly.

Frequency polygons drawn on the same axes show which distribution peaks further along and which is more tightly clustered, though exact values are harder to read.

Stem-and-leaf diagrams keep the original values, so exact medians and ranges can be calculated. A back-to-back diagram is designed for comparing two sets.

Writing the comparison

Three things earn the marks: compare, use figures, and write in context.

Compare rather than describe. "Group A's median was 52 and Group B's was 47" lists two facts; "Group A's median was higher, at 52 compared with 47" compares them.

Quote the figures. A comparison without numbers rarely scores full marks.

Write in context. The subject of the sentence should be what was measured — reaction times, test scores, plant heights — not "the data" or "the box plot". Saying "the second box plot is further right" describes the diagram rather than the situation, and earns nothing.

A model answer has two sentences: "The boys' median time was faster, at 12.4 seconds compared with 13.1 seconds. The boys' times were also more consistent, since their interquartile range was 1.2 seconds compared with 2.6 seconds."

Worked examples

Example 1: Choosing the right measures

A set of house prices in a street is 180, 190, 195, 200 and 950 thousand pounds. Which average and which measure of spread describe this street best, and why?

The value of 950 is an outlier, far above the rest.

The mean is 343 thousand, which is higher than four of the five houses, so it describes none of them well. The median of 195 sits among the typical values and represents the street far better.

The range is 770, driven entirely by that one house. The interquartile range of 190 to 200, so 10, describes the bulk of the street properly.

So the median and interquartile range are the appropriate pair, because the data contains an outlier.

Example 2: A comparison from box plots

Two classes sat the same test. Class A has a median of 58 and an interquartile range of 14. Class B has a median of 52 and an interquartile range of 22. Compare the two classes.

On average, Class A scored higher: its median was 58 marks compared with 52 marks for Class B.

Class A's marks were also more consistent, since its interquartile range was 14 marks compared with 22 marks for Class B.

Both statements name the measure, quote the figures, and refer to marks rather than to the diagram — which is what full marks require.

Example 3: The effect of an outlier

A set of seven times has a mean of 15 seconds and a median of 14 seconds. An eighth runner finishes in 60 seconds. Describe the effect on each average.

The mean rises sharply. The original total was 105 seconds; adding 60 gives 165 over 8 runners, so the mean becomes about 20.6 seconds — higher than every one of the original seven times.

The median barely moves, shifting only to the next value along, because it depends on position rather than size.

This is exactly why the median is preferred when outliers are present, and why a question mentioning one unusual value is usually steering you towards it.

Common mistakes and how to avoid them

Giving only one statement. Compare an average and a spread. One without the other scores half.

Describing instead of comparing. Use comparative words: higher, lower, more consistent, more varied.

Leaving out the figures. Quote both values whenever you compare.

Writing about the diagram. Refer to what was measured, not to boxes, bars or lines.

Comparing a mean with a median. Use the same measure for both groups.

Saying a smaller range means a higher average. Spread says nothing about the centre; they are independent.

Using the mean when there is an obvious outlier. Choose the median, and say why.

Exam technique for "Comparing Distributions"

Write two sentences, one for the average and one for the spread. Structuring the answer that way guarantees you address both marks.

Name the measure you are using in each sentence — "the median mark", "the interquartile range" — so the examiner can see the comparison is like for like.

Quote both figures with their units in every sentence.

Make the subject of each sentence the thing being measured, not the diagram.

When asked which average is most appropriate, name it and give the reason, usually the presence or absence of outliers.

Check that your two statements are genuinely about different things: one about which is higher, one about which is more consistent.

Quick revision summary

A fair comparison needs two statements: one about the average and one about the spread. Either alone scores half the marks.

Use the median when the data is skewed or contains outliers, the mean when it is symmetrical, and the mode for categorical data or the most typical value.

Pair the median with the interquartile range, and the mean with the range. Never compare one group's mean with another's median.

A smaller spread means more consistent data. Spread says nothing about the centre.

On a box plot, the line gives the median and the box width gives the interquartile range. It shows neither the mean nor the number of values.

Write in context with figures, and use comparative language: "the boys' median time was faster, at 12.4 seconds compared with 13.1 seconds".

An outlier moves the mean and the range a great deal, and the median and interquartile range hardly at all.

Comparing distributions using statistical measures and diagrams: common questions

What is Distribution?

Distribution — the pattern of how a set of values is spread out.

What do you need to know about Comparing distributions using statistical measures and diagrams for AQA GCSE Mathematics?

A fair comparison needs two statements: one about the average and one about the spread. Either alone scores half the marks.

What are the most common mistakes in Comparing distributions using statistical measures and diagrams?

Giving only one statement: Compare an average and a spread. One without the other scores half. Describing instead of comparing: Use comparative words: higher, lower, more consistent, more varied. Leaving out the figures: Quote both values whenever you compare.

Where can I practise Comparing distributions using statistical measures and diagrams questions for free?

Kramizo has free AQA GCSE Mathematics practice questions on Comparing distributions using statistical measures and diagrams, each marked instantly with a full explanation. No card is required.

Free for GCSE students

Lock in Comparing distributions using statistical measures and diagrams with real exam questions.

Free instantly-marked AQA GCSE Mathematics practice — 45 questions a day, no card required.

Try a question →See practice bank