Kramizo
Log inSign up free
HomeAQA GCSE StatisticsRepresenting Data: Scatter Diagrams and Correlation
AQA · GCSE · Statistics · Revision Notes

Representing Data: Scatter Diagrams and Correlation

2,177 words · Last updated July 2026

Ready to practise? Test yourself on Representing Data: Scatter Diagrams and Correlation with instantly-marked questions.
Practice now →
Quick answer

Correlationa measure of the relationship between two variables, describing whether they tend to increase or decrease together

Scatter diagrams plot bivariate data to investigate relationships between variables. Correlation describes both direction (positive/negative/zero) and strength (strong/moderate/weak). Lines of best fit summarise trends and enable predictions through interpolation (reliable, within data range) and extrapolation (less reliable, outside data range). Always draw lines through the mean point with equal scatter above and below, ignoring outliers. Remember that correlation never proves causation. In exams, use precise terminology, show all construction lines, and label diagrams fully for maximum marks.

What you'll learn

This revision guide covers scatter diagrams and correlation as required by the AQA GCSE Statistics specification. You'll learn how to construct and interpret scatter diagrams, identify types of correlation, draw and use lines of best fit, and make predictions using these graphical representations. Understanding these concepts is essential for analysing bivariate data and forms a core part of Paper 1.

Key terms and definitions

Bivariate data — data collected on two variables for each individual or item, allowing investigation of possible relationships between the variables

Scatter diagram (scatter graph) — a graph plotting one variable against another on Cartesian axes, with each point representing paired values from bivariate data

Correlation — a measure of the relationship between two variables, describing whether they tend to increase or decrease together

Positive correlation — when one variable tends to increase as the other increases, shown by points forming an upward trend from left to right

Negative correlation — when one variable tends to decrease as the other increases, shown by points forming a downward trend from left to right

Zero (no) correlation — when there is no linear relationship between two variables, with points scattered randomly across the diagram

Line of best fit — a straight line drawn through the middle of scattered points to represent the trend in the data, passing as close as possible to all points with roughly equal numbers above and below

Outlier — a data point that lies well away from the general pattern of the other points on a scatter diagram

Core concepts

Understanding bivariate data and scatter diagrams

Bivariate data involves measuring two different variables for the same set of individuals or items. The purpose of collecting such data is to investigate whether a relationship exists between the variables.

When constructing a scatter diagram:

  • Plot the independent variable (explanatory variable) on the horizontal x-axis
  • Plot the dependent variable (response variable) on the vertical y-axis
  • Each point represents one pair of measurements
  • Label both axes clearly with variable names and units
  • Use an appropriate scale that makes good use of the graph paper
  • Give the diagram a title that identifies both variables

The independent variable is the one you think might influence or explain the other. For example, when investigating the relationship between hours of revision and exam score, hours of revision is independent (you choose this) and exam score is dependent (this responds to revision time).

Types and strength of correlation

Correlation describes both the direction and strength of a linear relationship between two variables.

Direction of correlation:

  • Positive correlation: as one variable increases, the other tends to increase. Examples: height and shoe size; temperature and ice cream sales
  • Negative correlation: as one variable increases, the other tends to decrease. Examples: age of car and value; distance from equator and average temperature
  • Zero/no correlation: no apparent linear relationship between the variables. Examples: shoe size and IQ; favourite colour and height

Strength of correlation:

The strength describes how closely the points cluster around an imaginary straight line:

  • Strong correlation: points lie close to a straight line with little scatter
  • Moderate correlation: points show a clear trend but with noticeable scatter
  • Weak correlation: points show some trend but with considerable scatter
  • Zero correlation: no discernible linear pattern

You must be able to distinguish between these different strengths when describing correlation in exam questions. Use precise terminology: "strong positive correlation" is better than "positive correlation."

Drawing lines of best fit

A line of best fit summarises the trend in bivariate data and allows predictions to be made.

Rules for drawing a line of best fit:

  • Use a transparent ruler and draw a single straight line
  • The line should pass through or near as many points as possible
  • Roughly equal numbers of points should lie above and below the line
  • The line does not need to pass through the origin (0,0)
  • The line should pass through the mean point (x̄, ȳ) where x̄ is the mean of all x-values and ȳ is the mean of all y-values
  • Ignore any outliers when drawing the line — they should not influence its position
  • Extend the line across the full range of the data

In the exam, you should draw your line carefully in pencil so it can be corrected if necessary. The line should be long enough to span the data range but not extend far beyond unless specifically required for extrapolation.

Identifying and dealing with outliers

An outlier is a point that does not fit the general pattern shown by other data points. On a scatter diagram, outliers lie far from the line of best fit.

When identifying outliers:

  • Look for points that are isolated from the main cluster
  • An outlier may have an unusually high or low x-value, y-value, or both
  • State clearly which point is the outlier by giving its coordinates
  • Suggest a possible reason for the outlier (e.g., measurement error, recording error, or a genuine unusual case)

Dealing with outliers:

  • Do not include outliers when drawing the line of best fit
  • If asked whether to remove an outlier, consider whether it represents an error or genuine extreme case
  • Outliers should usually be investigated rather than automatically removed

Using lines of best fit for interpolation and extrapolation

Once a line of best fit is drawn, it can be used to make predictions.

Interpolation — making predictions within the range of the observed data. This is generally reliable because you are predicting within known values. To interpolate:

  • Locate the known value on the appropriate axis
  • Draw a line to the line of best fit (use a ruler and make it neat)
  • Draw a line from this point to the other axis to read off the predicted value
  • Show your working clearly on the diagram

Extrapolation — making predictions outside the range of the observed data. This is less reliable because:

  • The relationship may not continue in the same linear way beyond the data range
  • Unknown factors may influence the variables at extreme values
  • You cannot be certain the trend continues

Always state when you are extrapolating and acknowledge that predictions are less reliable. Exam questions may ask you to comment on the reliability of a prediction, requiring you to identify whether interpolation or extrapolation has been used.

Correlation and causation

A crucial concept in statistics is understanding that correlation does not imply causation.

Just because two variables are correlated does not mean that one causes the other. There are several possible explanations for correlation:

  • Direct causation: one variable directly affects the other (e.g., hours of study may directly cause higher exam scores)
  • Reverse causation: the causation works in the opposite direction to what you assumed
  • Common cause (confounding variable): a third variable affects both measured variables (e.g., age affects both height and reading ability in children, creating correlation between height and reading ability)
  • Coincidence: the correlation is purely by chance, especially with small datasets

In exam answers, avoid stating that correlation proves causation. Use phrases like "suggests a relationship" or "may be associated with" rather than "causes."

Worked examples

Example 1: Constructing and interpreting a scatter diagram

Question: The table shows the average daily temperature (°C) and number of hot drinks sold at a café over 8 days.

Temperature (°C) 8 12 15 18 20 22 25 28
Hot drinks sold 87 76 68 62 55 48 41 35

(a) Draw a scatter diagram to represent this data. (3 marks) (b) Describe the correlation. (2 marks) (c) Draw a line of best fit. (1 mark)

Solution:

(a) [For full marks, the scatter diagram must have:]

  • Both axes labelled correctly with "Temperature (°C)" and "Hot drinks sold" ✓
  • Appropriate scales chosen (e.g., x-axis 0-30°C, y-axis 30-90 drinks) ✓
  • All 8 points plotted correctly as crosses or small dots ✓

(b) The scatter diagram shows strong negative correlation. ✓ As temperature increases, the number of hot drinks sold decreases. ✓

(c) [Line of best fit should:]

  • Be drawn with a ruler as a straight line
  • Pass through the general middle of points with roughly equal numbers above and below
  • Extend across the range of temperatures shown
  • Have negative gradient matching the trend

Example 2: Using a line of best fit for predictions

Question: A scatter diagram shows the relationship between hours of training per week and race time (in minutes) for 10 athletes. A line of best fit has been drawn.

(a) Use the line of best fit to estimate the race time for an athlete who trains 6 hours per week. (2 marks) (b) An athlete who trains 15 hours per week has a race time of 42 minutes. Explain why using the line of best fit to predict this would be unreliable. (2 marks)

Solution:

(a) [Assuming the graph shows this relationship] Draw a vertical line from 6 hours on the x-axis to the line of best fit ✓ Draw a horizontal line to the y-axis to read the race time of approximately 48 minutes ✓

(b) This would be extrapolation because 15 hours is outside the range of the data collected. ✓ The linear relationship may not continue beyond the observed data, making predictions unreliable. ✓

Example 3: Identifying outliers and commenting on correlation

Question: A scatter diagram shows the relationship between engine size (litres) and fuel consumption (miles per gallon) for 12 cars. One car with a 2.0 litre engine has fuel consumption of 55 mpg, while all other cars with similar engine sizes have fuel consumption between 30-35 mpg.

(a) Identify the outlier and suggest a reason for it. (2 marks) (b) Should this outlier be included when drawing the line of best fit? Explain your answer. (2 marks)

Solution:

(a) The outlier is the car with engine size 2.0 litres and fuel consumption 55 mpg. ✓ Possible reasons: this could be a hybrid or electric vehicle / measurement error / data recording error. ✓

(b) No, this outlier should not be included when drawing the line of best fit. ✓ It does not follow the same pattern as the other data points and would distort the line, making it unrepresentative of the general relationship. ✓

Common mistakes and how to avoid them

  • Confusing which variable goes on which axis: Always place the independent (explanatory) variable on the x-axis and the dependent (response) variable on the y-axis. Read the context carefully to determine which is which.

  • Drawing lines of best fit through the origin: The line of best fit does not need to pass through (0,0) unless the data pattern clearly requires this. Follow the trend of the actual data points.

  • Describing correlation vaguely: Avoid imprecise language like "the correlation is good" or "there is some correlation." Use the correct terminology: strong/moderate/weak combined with positive/negative/zero.

  • Joining points dot-to-dot: A line of best fit is a single straight line showing the trend, not a line connecting all the points. Scatter diagrams show trends, not exact relationships.

  • Stating that correlation proves causation: Even strong correlation does not prove that one variable causes changes in the other. Always consider alternative explanations including confounding variables.

  • Forgetting to show working for interpolation/extrapolation: When using a line of best fit to make predictions, draw clear construction lines on your diagram so the examiner can follow your method and award method marks even if your answer is slightly inaccurate.

Exam technique for "Representing Data: Scatter Diagrams and Correlation"

  • "Describe the correlation" requires both strength (strong/moderate/weak) and direction (positive/negative/zero). Full marks typically require both elements clearly stated (2 marks).

  • When drawing scatter diagrams, accuracy is essential: use a sharp pencil, plot points precisely using crosses or neat dots, and label axes fully including units. Expect 3 marks for a well-constructed diagram.

  • For interpolation/extrapolation questions, always show construction lines on the graph even if you can read the value directly. This demonstrates your method and secures method marks (usually 1 mark for correct method, 1 mark for correct answer).

  • "Explain why this prediction may be unreliable" expects you to identify extrapolation and state that the relationship may not continue outside the data range. Two clear points score 2 marks.

Quick revision summary

Scatter diagrams plot bivariate data to investigate relationships between variables. Correlation describes both direction (positive/negative/zero) and strength (strong/moderate/weak). Lines of best fit summarise trends and enable predictions through interpolation (reliable, within data range) and extrapolation (less reliable, outside data range). Always draw lines through the mean point with equal scatter above and below, ignoring outliers. Remember that correlation never proves causation. In exams, use precise terminology, show all construction lines, and label diagrams fully for maximum marks.

Representing Data: Scatter Diagrams and Correlation: common questions

What is Correlation?

Correlation — a measure of the relationship between two variables, describing whether they tend to increase or decrease together

What do you need to know about Representing Data: Scatter Diagrams and Correlation for AQA GCSE Statistics?

Scatter diagrams plot bivariate data to investigate relationships between variables. Correlation describes both direction (positive/negative/zero) and strength (strong/moderate/weak). Lines of best fit summarise trends and enable predictions through interpolation (reliable, within data range) and extrapolation (less reliable, outside data range). Always draw lines through the mean point with equal scatter above and below, ignoring outliers. Remember that correlation never proves causation. In exams, use precise terminology, show all construction lines, and label diagrams fully for maximum marks.

What are the most common mistakes in Representing Data: Scatter Diagrams and Correlation?

Confusing which variable goes on which axis: Always place the independent (explanatory) variable on the x-axis and the dependent (response) variable on the y-axis. Read the context carefully to determine which is which. Drawing lines of best fit through the origin: The line of best fit does not need to pass through (0,0) unless the data pattern clearly requires this. Follow the trend of the actual data points. Describing correlation vaguely: Avoid imprecise language like "the correlation is good" or "there is some correlation." Use the correct terminology: strong/moderate/weak combined with positive/negative/zero.

Where can I practise Representing Data: Scatter Diagrams and Correlation questions for free?

Kramizo has free AQA GCSE Statistics practice questions on Representing Data: Scatter Diagrams and Correlation, each marked instantly with a full explanation. No card is required.

Free for GCSE students

Lock in Representing Data: Scatter Diagrams and Correlation with real exam questions.

Free instantly-marked AQA GCSE Statistics practice — 45 questions a day, no card required.

Try a question →See practice bank