What you'll learn
This revision guide covers scatter diagrams and correlation as required by the AQA GCSE Statistics specification. You'll learn how to construct and interpret scatter diagrams, identify types of correlation, draw and use lines of best fit, and make predictions using these graphical representations. Understanding these concepts is essential for analysing bivariate data and forms a core part of Paper 1.
Key terms and definitions
Bivariate data — data collected on two variables for each individual or item, allowing investigation of possible relationships between the variables
Scatter diagram (scatter graph) — a graph plotting one variable against another on Cartesian axes, with each point representing paired values from bivariate data
Correlation — a measure of the relationship between two variables, describing whether they tend to increase or decrease together
Positive correlation — when one variable tends to increase as the other increases, shown by points forming an upward trend from left to right
Negative correlation — when one variable tends to decrease as the other increases, shown by points forming a downward trend from left to right
Zero (no) correlation — when there is no linear relationship between two variables, with points scattered randomly across the diagram
Line of best fit — a straight line drawn through the middle of scattered points to represent the trend in the data, passing as close as possible to all points with roughly equal numbers above and below
Outlier — a data point that lies well away from the general pattern of the other points on a scatter diagram
Core concepts
Understanding bivariate data and scatter diagrams
Bivariate data involves measuring two different variables for the same set of individuals or items. The purpose of collecting such data is to investigate whether a relationship exists between the variables.
When constructing a scatter diagram:
- Plot the independent variable (explanatory variable) on the horizontal x-axis
- Plot the dependent variable (response variable) on the vertical y-axis
- Each point represents one pair of measurements
- Label both axes clearly with variable names and units
- Use an appropriate scale that makes good use of the graph paper
- Give the diagram a title that identifies both variables
The independent variable is the one you think might influence or explain the other. For example, when investigating the relationship between hours of revision and exam score, hours of revision is independent (you choose this) and exam score is dependent (this responds to revision time).
Types and strength of correlation
Correlation describes both the direction and strength of a linear relationship between two variables.
Direction of correlation:
- Positive correlation: as one variable increases, the other tends to increase. Examples: height and shoe size; temperature and ice cream sales
- Negative correlation: as one variable increases, the other tends to decrease. Examples: age of car and value; distance from equator and average temperature
- Zero/no correlation: no apparent linear relationship between the variables. Examples: shoe size and IQ; favourite colour and height
Strength of correlation:
The strength describes how closely the points cluster around an imaginary straight line:
- Strong correlation: points lie close to a straight line with little scatter
- Moderate correlation: points show a clear trend but with noticeable scatter
- Weak correlation: points show some trend but with considerable scatter
- Zero correlation: no discernible linear pattern
You must be able to distinguish between these different strengths when describing correlation in exam questions. Use precise terminology: "strong positive correlation" is better than "positive correlation."
Drawing lines of best fit
A line of best fit summarises the trend in bivariate data and allows predictions to be made.
Rules for drawing a line of best fit:
- Use a transparent ruler and draw a single straight line
- The line should pass through or near as many points as possible
- Roughly equal numbers of points should lie above and below the line
- The line does not need to pass through the origin (0,0)
- The line should pass through the mean point (x̄, ȳ) where x̄ is the mean of all x-values and ȳ is the mean of all y-values
- Ignore any outliers when drawing the line — they should not influence its position
- Extend the line across the full range of the data
In the exam, you should draw your line carefully in pencil so it can be corrected if necessary. The line should be long enough to span the data range but not extend far beyond unless specifically required for extrapolation.
Identifying and dealing with outliers
An outlier is a point that does not fit the general pattern shown by other data points. On a scatter diagram, outliers lie far from the line of best fit.
When identifying outliers:
- Look for points that are isolated from the main cluster
- An outlier may have an unusually high or low x-value, y-value, or both
- State clearly which point is the outlier by giving its coordinates
- Suggest a possible reason for the outlier (e.g., measurement error, recording error, or a genuine unusual case)
Dealing with outliers:
- Do not include outliers when drawing the line of best fit
- If asked whether to remove an outlier, consider whether it represents an error or genuine extreme case
- Outliers should usually be investigated rather than automatically removed
Using lines of best fit for interpolation and extrapolation
Once a line of best fit is drawn, it can be used to make predictions.
Interpolation — making predictions within the range of the observed data. This is generally reliable because you are predicting within known values. To interpolate:
- Locate the known value on the appropriate axis
- Draw a line to the line of best fit (use a ruler and make it neat)
- Draw a line from this point to the other axis to read off the predicted value
- Show your working clearly on the diagram
Extrapolation — making predictions outside the range of the observed data. This is less reliable because:
- The relationship may not continue in the same linear way beyond the data range
- Unknown factors may influence the variables at extreme values
- You cannot be certain the trend continues
Always state when you are extrapolating and acknowledge that predictions are less reliable. Exam questions may ask you to comment on the reliability of a prediction, requiring you to identify whether interpolation or extrapolation has been used.
Correlation and causation
A crucial concept in statistics is understanding that correlation does not imply causation.
Just because two variables are correlated does not mean that one causes the other. There are several possible explanations for correlation:
- Direct causation: one variable directly affects the other (e.g., hours of study may directly cause higher exam scores)
- Reverse causation: the causation works in the opposite direction to what you assumed
- Common cause (confounding variable): a third variable affects both measured variables (e.g., age affects both height and reading ability in children, creating correlation between height and reading ability)
- Coincidence: the correlation is purely by chance, especially with small datasets
In exam answers, avoid stating that correlation proves causation. Use phrases like "suggests a relationship" or "may be associated with" rather than "causes."
Worked examples
Example 1: Constructing and interpreting a scatter diagram
Question: The table shows the average daily temperature (°C) and number of hot drinks sold at a café over 8 days.
| Temperature (°C) | 8 | 12 | 15 | 18 | 20 | 22 | 25 | 28 |
|---|---|---|---|---|---|---|---|---|
| Hot drinks sold | 87 | 76 | 68 | 62 | 55 | 48 | 41 | 35 |
(a) Draw a scatter diagram to represent this data. (3 marks) (b) Describe the correlation. (2 marks) (c) Draw a line of best fit. (1 mark)
Solution:
(a) [For full marks, the scatter diagram must have:]
- Both axes labelled correctly with "Temperature (°C)" and "Hot drinks sold" ✓
- Appropriate scales chosen (e.g., x-axis 0-30°C, y-axis 30-90 drinks) ✓
- All 8 points plotted correctly as crosses or small dots ✓
(b) The scatter diagram shows strong negative correlation. ✓ As temperature increases, the number of hot drinks sold decreases. ✓
(c) [Line of best fit should:]
- Be drawn with a ruler as a straight line
- Pass through the general middle of points with roughly equal numbers above and below
- Extend across the range of temperatures shown
- Have negative gradient matching the trend
Example 2: Using a line of best fit for predictions
Question: A scatter diagram shows the relationship between hours of training per week and race time (in minutes) for 10 athletes. A line of best fit has been drawn.
(a) Use the line of best fit to estimate the race time for an athlete who trains 6 hours per week. (2 marks) (b) An athlete who trains 15 hours per week has a race time of 42 minutes. Explain why using the line of best fit to predict this would be unreliable. (2 marks)
Solution:
(a) [Assuming the graph shows this relationship] Draw a vertical line from 6 hours on the x-axis to the line of best fit ✓ Draw a horizontal line to the y-axis to read the race time of approximately 48 minutes ✓
(b) This would be extrapolation because 15 hours is outside the range of the data collected. ✓ The linear relationship may not continue beyond the observed data, making predictions unreliable. ✓
Example 3: Identifying outliers and commenting on correlation
Question: A scatter diagram shows the relationship between engine size (litres) and fuel consumption (miles per gallon) for 12 cars. One car with a 2.0 litre engine has fuel consumption of 55 mpg, while all other cars with similar engine sizes have fuel consumption between 30-35 mpg.
(a) Identify the outlier and suggest a reason for it. (2 marks) (b) Should this outlier be included when drawing the line of best fit? Explain your answer. (2 marks)
Solution:
(a) The outlier is the car with engine size 2.0 litres and fuel consumption 55 mpg. ✓ Possible reasons: this could be a hybrid or electric vehicle / measurement error / data recording error. ✓
(b) No, this outlier should not be included when drawing the line of best fit. ✓ It does not follow the same pattern as the other data points and would distort the line, making it unrepresentative of the general relationship. ✓
Common mistakes and how to avoid them
Confusing which variable goes on which axis: Always place the independent (explanatory) variable on the x-axis and the dependent (response) variable on the y-axis. Read the context carefully to determine which is which.
Drawing lines of best fit through the origin: The line of best fit does not need to pass through (0,0) unless the data pattern clearly requires this. Follow the trend of the actual data points.
Describing correlation vaguely: Avoid imprecise language like "the correlation is good" or "there is some correlation." Use the correct terminology: strong/moderate/weak combined with positive/negative/zero.
Joining points dot-to-dot: A line of best fit is a single straight line showing the trend, not a line connecting all the points. Scatter diagrams show trends, not exact relationships.
Stating that correlation proves causation: Even strong correlation does not prove that one variable causes changes in the other. Always consider alternative explanations including confounding variables.
Forgetting to show working for interpolation/extrapolation: When using a line of best fit to make predictions, draw clear construction lines on your diagram so the examiner can follow your method and award method marks even if your answer is slightly inaccurate.
Exam technique for "Representing Data: Scatter Diagrams and Correlation"
"Describe the correlation" requires both strength (strong/moderate/weak) and direction (positive/negative/zero). Full marks typically require both elements clearly stated (2 marks).
When drawing scatter diagrams, accuracy is essential: use a sharp pencil, plot points precisely using crosses or neat dots, and label axes fully including units. Expect 3 marks for a well-constructed diagram.
For interpolation/extrapolation questions, always show construction lines on the graph even if you can read the value directly. This demonstrates your method and secures method marks (usually 1 mark for correct method, 1 mark for correct answer).
"Explain why this prediction may be unreliable" expects you to identify extrapolation and state that the relationship may not continue outside the data range. Two clear points score 2 marks.
Quick revision summary
Scatter diagrams plot bivariate data to investigate relationships between variables. Correlation describes both direction (positive/negative/zero) and strength (strong/moderate/weak). Lines of best fit summarise trends and enable predictions through interpolation (reliable, within data range) and extrapolation (less reliable, outside data range). Always draw lines through the mean point with equal scatter above and below, ignoring outliers. Remember that correlation never proves causation. In exams, use precise terminology, show all construction lines, and label diagrams fully for maximum marks.