Scatter Graphs and Lines of Best Fit — AQA GCSE Maths Revision Notes
What you'll learn
This topic covers scatter graphs: plotting two variables against each other to see whether they are related. By the end of this guide you should be able to plot one accurately, describe the correlation it shows, and draw a line of best fit.
You should also be able to use that line to make predictions, explain why some predictions are unreliable, identify outliers, and explain why a relationship between two variables does not prove that one causes the other.
The organising idea is that a scatter graph shows a relationship, not a rule. The points never lie perfectly on a line, because real data has variation in it, so everything the graph gives you is an estimate rather than a fact. That is why the line is called a line of best fit rather than simply "the line", why a prediction made inside the data range is trustworthy and one made outside it is not, and why even a strong pattern cannot prove that one quantity causes the other. Holding on to "relationship, not rule" answers most of the interpretation questions in this topic.
Key terms and definitions
Scatter graph — a graph plotting paired values of two variables as points.
Bivariate data — data in which each item has two measurements, such as a person's height and mass.
Correlation — the relationship between the two variables.
Positive correlation — as one variable increases, so does the other.
Negative correlation — as one variable increases, the other decreases.
Line of best fit — a single straight line following the trend of the points.
Interpolation — estimating a value inside the range of the data.
Extrapolation — estimating a value beyond the range of the data.
Outlier — a point that does not follow the general pattern.
Causation — one variable actually causing the change in the other.
Core concepts
Plotting
Each item of data supplies two values, which together give one point. Plot each with a small neat cross, and do not join the points.
Choose scales that spread the data across the available grid. A cramped plot makes the correlation difficult to see and difficult to mark.
Describing correlation
A description needs two parts: the type and the strength.
The type is positive if the points rise from left to right, negative if they fall, and there is no correlation if they show no consistent direction.
The strength depends on how closely the points cluster around a straight line. Tightly clustered points show strong correlation; widely scattered ones show weak correlation.
So a full description reads "strong positive correlation" or "weak negative correlation" — never just "positive".
Add the context to gain the final mark: "there is strong positive correlation, so taller students tend to have a greater arm span".
Note the word tends. Correlation describes a general tendency, not something true of every individual point.
Drawing a line of best fit
The line is a single straight line, drawn with a ruler, following the trend of the points.
Aim for roughly equal numbers of points above and below it, and keep it running through the bulk of the data.
It does not have to pass through the origin, and forcing it there is a common error that distorts every prediction made from it.
It is usually a good line if it passes close to the mean point — the point whose coordinates are the mean of each variable — and some questions ask you to plot that point first.
Ignore outliers when positioning the line. A single stray point should not drag the line away from the pattern the rest of the data shows.
Using the line to predict
To estimate a value, read from the known value across or up to the line, then across or down to the other axis.
Draw the reading lines on the graph. Marks are awarded for a correct reading line even if the value read off is slightly out.
Interpolation and extrapolation
Interpolation means estimating inside the range of the data collected, and it is reasonably reliable because the pattern has been observed there.
Extrapolation means estimating outside that range, and it is unreliable, because there is no evidence the trend continues beyond the data.
This is a standard exam question, and the answer needs the reason, not just the word. "The estimate is unreliable because 90 kg is outside the range of the data collected, so the trend may not continue that far" earns the mark; "because it is extrapolation" often does not.
Extrapolation can produce obviously absurd results, which is the clearest illustration of why it fails: a line relating age to height in children, extended far enough, predicts impossible heights for adults.
The mean point
The point whose coordinates are the mean of each variable is called the mean point, and a good line of best fit passes through it or very close to it.
If ten students have a mean revision time of 5 hours and a mean score of 58 marks, the mean point is at 5 hours and 58 marks, and the line should pass through there.
Some questions ask you to calculate and plot it before drawing the line. Doing so turns a judgement by eye into something much more reliable, because the line then has a fixed point to pivot around.
The equation of the line of best fit
On the higher tier you may be asked for the equation of the line, in the form y = mx + c.
Read the gradient from two well-separated points on the line — not two data points, which is the usual error, since the data points do not lie exactly on it.
The intercept is where the line crosses the vertical axis, read directly if the axis starts at zero.
Both numbers mean something in context, and questions ask for that. For a line relating revision hours to score, the gradient gives the marks gained per extra hour of revision, and the intercept gives the predicted score for no revision at all. Interpreting the intercept requires care: it is only meaningful when zero lies within or close to the range of the data.
Outliers
An outlier is a point clearly away from the pattern. It may be a genuine unusual case or a recording error.
Identify it by its coordinates, not by pointing: "the point at (45, 12) is an outlier".
Outliers are ignored when drawing the line of best fit, but they are not deleted from the data — a question may ask you to suggest why one occurred.
Correlation is not causation
Two variables can correlate strongly without either causing the other.
Ice cream sales and swimming pool visits rise together, but neither causes the other; hot weather causes both. A third factor of this kind lies behind many strong correlations.
When a question asks whether one thing causes another, the answer is that correlation alone cannot show it. Say what else might explain the link.
Worked examples
Example 1: Describing what a scatter graph shows
A scatter graph plots hours of revision against test score for 20 students. The points rise steadily from bottom left to top right and lie fairly close to a straight line. Describe the correlation and what it means.
The points rise, so the correlation is positive. They lie close to a line, so it is strong.
The full description is strong positive correlation.
In context: students who revised for longer tended to score higher marks.
The word "tended" matters, since individual students will not all follow the pattern.
Example 2: Interpolation and extrapolation
The revision data covers 0 to 10 hours. Explain why an estimate of the score for 4 hours is more reliable than one for 25 hours.
Four hours lies within the range of the collected data, so the line of best fit is supported by evidence there. This is interpolation and it is reliable.
Twenty-five hours lies well beyond the data, so there is no evidence the trend continues. This is extrapolation and it is unreliable — in reality the scores must level off, since the test has a maximum mark.
Both halves of that answer are needed: inside or outside the range, and what follows from it.
Example 3: Correlation and causation
A study finds strong positive correlation between the number of firefighters sent to a fire and the damage caused. Does sending more firefighters cause more damage?
No. The correlation is genuine but neither variable causes the other.
Larger fires cause both — they require more firefighters and they cause more damage. The size of the fire is the third factor explaining the link.
An answer that names the third factor demonstrates the point properly, which is what the question is testing.
Common mistakes and how to avoid them
Joining the points. A scatter graph shows separate points; only the line of best fit is drawn.
Describing the type but not the strength. Say "strong positive", not just "positive".
Forcing the line through the origin. It goes where the data goes, and zero is not always meaningful.
Letting an outlier pull the line. Ignore it when drawing, but do not delete it.
Answering "extrapolation" without a reason. Say the value is outside the range of the data, so the trend may not continue.
Claiming one variable causes the other. Correlation shows a relationship, never causation. Suggest a third factor.
Forgetting the context. Name the actual variables when describing correlation.
Exam technique for "Scatter Graphs"
Plot points as small crosses and use a ruler for the line of best fit. Accuracy marks depend on both.
Describe correlation with three things: type, strength, and what it means for the variables in question.
Draw the reading lines when making a prediction, and leave them on the graph. They earn credit independently of the value.
When judging reliability, state whether the value lies inside or outside the range of the data, and say what follows.
Identify an outlier by its coordinates.
If asked whether one thing causes another, say that correlation cannot show causation and suggest a plausible third factor.
Quick revision summary
A scatter graph shows a relationship, not a rule — every value taken from it is an estimate.
Describe correlation by type and strength: strong or weak, positive or negative, or none. Add what it means for the variables.
The line of best fit is a single straight line with roughly equal numbers of points either side. It need not pass through the origin, and outliers are ignored when drawing it. It should pass close to the mean point.
Interpolation — predicting inside the data range — is reliable. Extrapolation — predicting beyond it — is unreliable, because the trend may not continue. Give that reason, not just the word.
An outlier is a point away from the pattern; name it by its coordinates and consider why it occurred.
Correlation is not causation. A strong relationship may be explained by a third factor affecting both.