Kramizo
Log inSign up free
HomePearson Edexcel International IGCSE MathematicsStatistics and Probability
Pearson Edexcel International · IGCSE · Mathematics · Revision Notes

Statistics and Probability

2,015 words · Last updated July 2026

Ready to practise? Test yourself on Statistics and Probability with instantly-marked questions.
Practice now →
Quick answer

Statistics involves calculating averages (mean, median, mode) and spread (range, IQR) from datasets. Present data using frequency tables, histograms, stem-and-leaf diagrams, scatter graphs and cumulative frequency curves. Use cumulative frequency to find medians and quartiles for large grouped datasets. Probability ranges from 0 to 1, with basic probability calculated as favourable outcomes over total outcomes. For combined events, use tree diagrams, multiplying along branches and adding separate routes. Apply conditional probability when events are dependent. Two-way tables and Venn diagrams organise data for probability calculations involving two categories or sets.

What you'll learn

Statistics and Probability forms a significant component of your Pearson Edexcel International IGCSE Mathematics examination, accounting for approximately 25% of the paper. This topic covers the collection, presentation, analysis and interpretation of data, alongside the fundamental principles of calculating probabilities. You'll develop skills in working with frequency tables, cumulative frequency diagrams, scatter graphs, and probability trees, all of which are regularly examined.

Key terms and definitions

Mean — the sum of all values divided by the number of values; the arithmetic average

Median — the middle value when data is arranged in ascending order; for an even number of values, it's the mean of the two central values

Mode — the value that appears most frequently in a dataset

Range — the difference between the highest and lowest values in a dataset

Cumulative frequency — a running total of frequencies; used to find medians and quartiles from grouped data

Interquartile range (IQR) — the difference between the upper quartile (Q₃) and lower quartile (Q₁); a measure of spread that excludes extreme values

Mutually exclusive events — events that cannot occur at the same time; if one happens, the other cannot

Independent events — events where the outcome of one does not affect the probability of the other

Core concepts

Averages and measures of spread

The three measures of central tendency each have specific applications:

Calculating the mean:

  • From a list: add all values and divide by how many there are
  • From a frequency table: multiply each value by its frequency, sum these products, then divide by the total frequency
  • Formula: x̄ = Σfx / Σf

Finding the median:

  • For n values, the median position is (n+1)/2
  • For grouped data, use cumulative frequency and the position n/2

Identifying the mode:

  • For discrete data: the most frequent value
  • For grouped data: the modal class (the class interval with the highest frequency)

Measures of spread:

  • Range shows the full spread but is affected by outliers
  • Interquartile range is more robust, showing the spread of the middle 50% of data
  • Lower quartile (Q₁) is at position n/4
  • Upper quartile (Q₃) is at position 3n/4

Data presentation methods

Frequency tables:

  • Discrete data uses exact values
  • Continuous data requires class intervals
  • Class intervals must not overlap (e.g., 0 ≤ x < 10, 10 ≤ x < 20)
  • The modal class is identified but not the exact mode

Stem-and-leaf diagrams:

  • Must include a key (e.g., 2|7 represents 27)
  • Ordered form required for finding median
  • Back-to-back diagrams compare two datasets

Scatter graphs:

  • Plot two variables to identify correlation (relationship)
  • Positive correlation: as one variable increases, so does the other
  • Negative correlation: as one variable increases, the other decreases
  • No correlation: no clear relationship
  • Line of best fit: drawn by eye to pass through the mean point, with roughly equal numbers of points on each side
  • Used for interpolation (within data range) but extrapolation (beyond data) is unreliable

Histograms:

  • Used for grouped continuous data
  • Frequency density = frequency ÷ class width
  • Vertical axis must be frequency density
  • Area of each bar represents frequency

Pie charts:

  • Angle for each sector = (frequency ÷ total frequency) × 360°
  • Used for categorical data to show proportions

Cumulative frequency

This method handles large grouped datasets efficiently:

Drawing a cumulative frequency curve:

  1. Create a cumulative frequency column (running total)
  2. Plot cumulative frequency against the upper class boundary
  3. Join points with a smooth curve
  4. Always starts at (lowest boundary, 0)

Reading from the curve:

  • Median: read across from n/2 on the vertical axis
  • Lower quartile (Q₁): read from n/4
  • Upper quartile (Q₃): read from 3n/4
  • IQR = Q₃ - Q₁

Box plots:

  • Display five key values: minimum, Q₁, median, Q₃, maximum
  • The box shows the IQR; the line inside marks the median
  • Whiskers extend to minimum and maximum values
  • Useful for comparing distributions

Basic probability

Probability measures the likelihood of an event occurring, expressed as a fraction, decimal or percentage.

Probability scale:

  • 0 = impossible
  • 1 = certain
  • All probabilities satisfy 0 ≤ P(event) ≤ 1

Basic probability formula: P(event) = number of favourable outcomes / total number of possible outcomes

Important rules:

  • P(event happens) + P(event doesn't happen) = 1
  • P(A or B) = P(A) + P(B) for mutually exclusive events
  • For independent events: P(A and B) = P(A) × P(B)

Expected outcomes: Expected frequency = probability × number of trials

Two-way tables and Venn diagrams

Two-way tables:

  • Organise data by two categories
  • Always complete row and column totals
  • Use algebra when information is incomplete
  • Extract probabilities by reading relevant cells

Venn diagrams:

  • Circles represent sets
  • Overlapping regions show elements in both sets
  • Rectangle represents the universal set
  • Fill in from the centre outwards
  • For probability: divide each region by the total

Set notation:

  • n(A) means the number of elements in set A
  • A ∩ B means A and B (intersection)
  • A ∪ B means A or B or both (union)
  • A' means not A (complement)

Tree diagrams

Tree diagrams show combined probabilities for two or more events.

Constructing tree diagrams:

  • Each branch represents a possible outcome
  • Write probabilities on each branch
  • Probabilities from the same point must sum to 1
  • For independent events, second-set probabilities stay the same
  • For dependent events (without replacement), second-set probabilities change

Using tree diagrams:

  • Multiply along branches for AND probabilities
  • Add separate routes for OR probabilities
  • List all outcomes systematically

Conditional probability:

  • When one event affects another
  • Drawing without replacement creates dependence
  • Adjust probabilities on second branches accordingly

Worked examples

Example 1: Cumulative frequency and quartiles

The table shows the masses of 80 students.

Mass (m kg) Frequency
40 ≤ m < 50 8
50 ≤ m < 60 22
60 ≤ m < 70 31
70 ≤ m < 80 14
80 ≤ m < 90 5

(a) Complete a cumulative frequency table and draw a cumulative frequency curve. (b) Use your graph to estimate the median and interquartile range.

Solution:

(a) Cumulative frequency table:

Mass (m kg) Frequency Upper boundary Cumulative frequency
40 ≤ m < 50 8 50 8
50 ≤ m < 60 22 60 30
60 ≤ m < 70 31 70 61
70 ≤ m < 80 14 80 75
80 ≤ m < 90 5 90 80

Plot points: (50, 8), (60, 30), (70, 61), (80, 75), (90, 80) and (40, 0) Draw a smooth curve through these points. [2 marks]

(b)

  • n = 80
  • Median position = 80/2 = 40th value → read across from 40 → median ≈ 64 kg [1 mark]
  • Q₁ position = 80/4 = 20th value → read from 20 → Q₁ ≈ 57 kg [1 mark]
  • Q₃ position = 3×80/4 = 60th value → read from 60 → Q₃ ≈ 69 kg [1 mark]
  • IQR = 69 - 57 = 12 kg [1 mark]

Example 2: Probability with tree diagrams

A bag contains 5 red counters and 3 blue counters. A counter is taken at random and not replaced. A second counter is then taken at random.

(a) Draw a tree diagram showing all possible outcomes. (b) Calculate the probability that both counters are the same colour.

Solution:

(a) Tree diagram:

                   Red (4/7)
         Red (5/8)
                   Blue (3/7)
Start
                   Red (5/7)
         Blue (3/8)
                   Blue (2/7)

First branches: P(Red) = 5/8, P(Blue) = 3/8 Second branches depend on first selection (without replacement) [2 marks]

(b) P(both same colour) = P(RR) + P(BB)

  • P(RR) = 5/8 × 4/7 = 20/56
  • P(BB) = 3/8 × 2/7 = 6/56
  • P(both same) = 20/56 + 6/56 = 26/56 = 13/28 [3 marks]

Example 3: Scatter graphs and correlation

Eight students recorded their test scores in Mathematics and Science:

Maths 45 62 71 54 80 38 67 75
Science 52 68 75 60 84 44 70 79

(a) Describe the correlation between Maths and Science scores. (b) A student scored 58 in Maths but missed the Science test. Estimate their Science score.

Solution:

(a) Plot the points on a scatter graph. The points show a clear upward trend from bottom-left to top-right. Strong positive correlation — students who score higher in Maths tend to score higher in Science. [2 marks]

(b) Draw a line of best fit through the points, ensuring roughly equal numbers above and below. Read from 58 on the horizontal axis up to the line, then across to the vertical axis. Estimated Science score ≈ 64 marks (accept 62-66) [2 marks]

Common mistakes and how to avoid them

  • Confusing mean and median — remember the median is the middle value when ordered, not an average of all values. For grouped data, you cannot calculate an exact mean without knowing individual values.

  • Using class midpoints incorrectly — when estimating the mean from grouped data, multiply each class midpoint by its frequency. Students often forget the midpoint step or use boundaries instead.

  • Probability errors exceeding 1 — always check your answer is between 0 and 1. If calculating P(A or B), only add probabilities when events are mutually exclusive.

  • Tree diagram multiplication mistakes — multiply probabilities along branches for combined events (AND), but add separate pathways for alternatives (OR). Write fractions rather than decimals to avoid rounding errors.

  • Forgetting to adjust probabilities without replacement — when items aren't replaced, both the numerator and denominator change for the second event. A bag with 5 red and 3 blue has P(red then red) = 5/8 × 4/7, not 5/8 × 5/8.

  • Plotting cumulative frequency incorrectly — always plot against the upper class boundary, never the midpoint. The curve should be smooth and always increasing, starting from zero.

Exam technique for "Statistics and Probability"

  • Command word "estimate" indicates you should use a graph or cumulative frequency curve. Show clearly on your diagram where you're reading values, using dotted lines. Examiners award method marks even if your curve is slightly inaccurate.

  • Show full working for probability calculations — write out the calculation in full (e.g., P(A and B) = 5/8 × 3/7 = 15/56) even for simple fractions. A correct answer with no working may only score partial marks if the final answer is wrong due to arithmetic error.

  • In scatter graph questions, "draw a line of best fit" means use a ruler and ensure your line goes through or near the mean point (mean of x-values, mean of y-values). Balance points above and below the line.

  • For grouped data questions, make sure you understand whether the question asks for the modal class (just state the interval) or an estimated mean (requires calculation using midpoints). Read carefully — these are different skills worth different marks.

Quick revision summary

Statistics involves calculating averages (mean, median, mode) and spread (range, IQR) from datasets. Present data using frequency tables, histograms, stem-and-leaf diagrams, scatter graphs and cumulative frequency curves. Use cumulative frequency to find medians and quartiles for large grouped datasets. Probability ranges from 0 to 1, with basic probability calculated as favourable outcomes over total outcomes. For combined events, use tree diagrams, multiplying along branches and adding separate routes. Apply conditional probability when events are dependent. Two-way tables and Venn diagrams organise data for probability calculations involving two categories or sets.

Statistics and Probability: common questions

What do you need to know about Statistics and Probability for Pearson Edexcel International IGCSE Mathematics?

Statistics involves calculating averages (mean, median, mode) and spread (range, IQR) from datasets. Present data using frequency tables, histograms, stem-and-leaf diagrams, scatter graphs and cumulative frequency curves. Use cumulative frequency to find medians and quartiles for large grouped datasets. Probability ranges from 0 to 1, with basic probability calculated as favourable outcomes over total outcomes. For combined events, use tree diagrams, multiplying along branches and adding separate routes. Apply conditional probability when events are dependent. Two-way tables and Venn diagrams organise data for probability calculations involving two categories or sets.

What are the most common mistakes in Statistics and Probability?

Confusing mean and median: remember the median is the middle value when ordered, not an average of all values. For grouped data, you cannot calculate an exact mean without knowing individual values. Using class midpoints incorrectly: when estimating the mean from grouped data, multiply each class midpoint by its frequency. Students often forget the midpoint step or use boundaries instead. Probability errors exceeding 1: always check your answer is between 0 and 1. If calculating P(A or B), only add probabilities when events are mutually exclusive.

Where can I practise Statistics and Probability questions for free?

Kramizo has free Pearson Edexcel International IGCSE Mathematics practice questions on Statistics and Probability, each marked instantly with a full explanation. No card is required.

Free for IGCSE students

Lock in Statistics and Probability with real exam questions.

Free instantly-marked Pearson Edexcel International IGCSE Mathematics practice — 45 questions a day, no card required.

Try a question →See practice bank