What you'll learn
Collecting data is the easy part; the marks are in what you do with it afterwards, and most Internal Assessments lose ground by describing findings rather than analysing them. This topic covers how to organise raw data, the difference between describing and interpreting, tabulation and frequency distributions, percentages and measures of central tendency, choosing an appropriate form of presentation, thematic analysis of qualitative data, cross-tabulation, and the discipline of matching your claims to what the evidence can actually support. It also covers where findings end and discussion begins, which is a distinction the mark scheme depends on.
Key terms and definitions
Raw data — the responses exactly as collected, before any processing.
Coding — assigning categories or numbers to responses so they can be counted. Essential for open-ended answers.
Tabulation — arranging data into tables so patterns become visible.
Frequency — the number of cases falling into a category.
Frequency distribution — a table showing how cases are spread across all categories of a variable.
Percentage — a frequency expressed out of one hundred, allowing comparison between groups of different sizes.
Mean — the arithmetic average: the total divided by the number of cases. Sensitive to extreme values.
Median — the middle value when cases are arranged in order. Unaffected by extremes, so preferable for skewed data such as income.
Mode — the most frequently occurring value. The only average usable with categories that have no numerical order.
Range — the difference between the highest and lowest values; a simple measure of spread.
Cross-tabulation — a table showing two variables together, so a relationship between them becomes visible.
Thematic analysis — identifying recurring themes across qualitative responses and organising the discussion around them.
Findings — what the data shows.
Discussion — what the findings mean in relation to the research question and the literature.
Interpretation — explaining the significance of a result rather than restating it.
Core concepts
From raw data to something analysable
Raw responses are not findings. The first task is to convert them into a form that can be counted and compared.
For closed questions this is straightforward: count how many respondents selected each option and record the frequencies. For open-ended responses and interview transcripts, you must code — read through, identify recurring categories, assign each response to one, and then count. Coding decisions should be stated in the methodology, because a reader needs to know how you decided what counted as which category.
Convert frequencies to percentages when comparing groups of unequal size. Twelve respondents out of twenty and eighteen out of forty are not comparable until expressed as 60 per cent and 45 per cent. Report the actual number alongside the percentage where the sample is small, because "50 per cent" from a sample of four overstates the finding considerably.
Tabulation
A table is usually the clearest way to present quantitative findings, and it costs nothing to produce accurately.
| Reason given for non-participation | Frequency | Percentage |
|---|---|---|
| Family responsibilities after school | 14 | 35% |
| Transport difficulties | 11 | 27.5% |
| Cost of equipment or fees | 9 | 22.5% |
| No interest in activities offered | 6 | 15% |
| Total | 40 | 100% |
Every table needs a number, a clear title, labelled columns, and totals that add correctly. Percentages should sum to 100 and be rounded consistently. Tables that do not add up are a straightforward and avoidable loss of marks.
Measures of central tendency
Choosing the right average is examinable, and the reasoning matters more than the calculation.
Use the mean for numerical data without extreme values — average number of hours studied, for instance. Use the median where a few extreme values would distort the mean, which is why income is normally reported as a median. Use the mode for categorical data, since there is no meaningful arithmetic average of "Hindu, Christian, Muslim".
A worked illustration: for the values 2, 3, 3, 4, 5, 5, 6 and 40, the mean is 8.5 — a figure higher than all but one of the observations, and therefore misleading. The median is 4.5 and the mode is 3 and 5. Reporting the mean alone here would misrepresent the data, and saying so is exactly the kind of evaluative comment that earns credit.
Choosing a form of presentation
Different data suits different presentation, and selecting appropriately is itself assessed.
A bar chart compares frequencies across separate categories. A pie chart shows how a whole divides into parts, and works only when categories are mutually exclusive and sum to the whole; it becomes unreadable beyond about six segments. A line graph shows change over time. A histogram shows the distribution of continuous data grouped into intervals. A table suits precise values and multiple variables at once.
Whatever you choose, each display needs a number, a title stating what it shows, labelled axes or segments, and a stated source. A chart without labels communicates nothing, and a chart duplicating a table on the same page adds length without adding information.
Analysing qualitative data
Interview and open-response data is analysed thematically rather than counted.
Read through the responses first without coding, to get a sense of the whole. Then identify recurring ideas and group them into themes. Name each theme in a way that describes what it contains. Then write the analysis theme by theme, supporting each with a short direct quotation from a respondent.
Quotations should be brief, chosen because they capture a theme clearly, and followed by your comment on what they show. A quotation left to speak for itself is description; a quotation plus interpretation is analysis. Respondents must not be identifiable, so use labels such as Respondent A rather than names.
Describing versus interpreting
This is the distinction that most affects the mark.
Description states what the data shows: "35 per cent of respondents gave family responsibilities as their reason for not participating."
Interpretation explains what it means: "Family responsibilities was the most frequently given reason, and taken with the interview finding that these duties fall disproportionately on female respondents, this suggests that non-participation reflects household obligations rather than lack of interest — which is what the school's current approach assumes."
Most weak IAs report percentage after percentage with no interpretation. The remedy is to ask, after every finding, so what? — what does this tell us about the research question, does it agree with the literature, and does anything about it surprise?
Cross-tabulation
Placing two variables together often reveals what a single distribution conceals. If 35 per cent give family responsibilities as their reason, cross-tabulating by sex may show that figure is 55 per cent among female respondents and 12 per cent among male respondents — a far more interesting finding than the overall figure, and one that only appears when the variables are examined together.
Note carefully that an association is not a cause. Two variables moving together may both be produced by a third. Claiming causation from a cross-tabulation is one of the most common analytical errors.
Matching claims to evidence
Your conclusions must be scaled to what your data can support. A study of forty students at one school supports statements about those forty students and, cautiously, about that school. It does not support statements about the country or the region.
Language matters here: suggests, indicates, and among respondents in this study are defensible; proves, shows conclusively and all young people are not. Overclaiming is penalised in the discussion and again in the conclusion.
Worked examples
Example 1: Choosing an average
Question: "A researcher collects weekly income data from ten households, nine of which report between $200 and $400 while one reports $5,000. Which measure of central tendency should be used, and why?" (6 marks)
Outline. The median should be used. The single very high value is an extreme that would pull the mean far above the figure typical of the group, producing an average higher than nine of the ten households actually earn and therefore misrepresenting the sample. The median, being the middle value when the cases are ordered, is unaffected by how extreme the highest value is and reports the typical household accurately. Add, for the higher mark, that reporting both together with a note on the outlier would be more transparent still.
Example 2: Moving from description to interpretation
Question: "Explain the difference between describing and interpreting findings, using an example." (8 marks)
Outline. Description reports what the data shows without comment: sixty per cent of respondents reported using Creole at home and Standard English at school. Interpretation explains the significance: this indicates that respondents are moving along the Creole continuum according to setting rather than lacking competence in either variety, which supports the view in the literature that code-switching is a skill rather than a deficiency, and suggests the school's corrective approach to Creole may be misdirected. The distinction is between restating the number and explaining what it means for the research question. Note that a report consisting only of description cannot reach the higher mark bands however accurate its figures.
Example 3: Reading a cross-tabulation
Question: "A study finds that 55 per cent of female respondents and 12 per cent of male respondents cite family responsibilities as a barrier to participation. Comment on this finding." (8 marks)
Outline. State the pattern: the barrier is reported more than four times as often by female respondents. Interpret: this suggests household duties are distributed unevenly by sex within the sample, so a barrier that appears to be about time is substantially about gendered expectations. Note the implication: an intervention scheduling activities later in the day would help only if it addressed the underlying distribution of responsibilities. Then qualify properly: the association does not establish cause, the finding describes this sample and cannot be generalised beyond it, and self-reported reasons may be shaped by what respondents consider acceptable to say. Pattern, interpretation, implication, qualification is a reliable four-step structure for any question of this kind.
Common mistakes and how to avoid them
Presenting findings without interpreting them. The commonest fault by a wide margin. After every figure, ask so what?
Tables and percentages that do not add up. Easily avoided and easily penalised. Check totals and round consistently.
Using the mean where the median is appropriate. Extreme values distort the mean; income and similar skewed data need the median.
Duplicating the same data in a table and a chart. Choose one. Repetition adds length, not marks.
Charts without titles, labels or sources. An unlabelled display communicates nothing and cannot be credited.
Claiming causation from an association. Cross-tabulations show relationships, not causes.
Overclaiming from a small sample. Findings from forty students at one school describe that school. Say so.
Confusing findings with discussion. Findings present the data; discussion explains it against the research question and literature. Merging them obscures both.
How this links to your Internal Assessment
Presentation of findings, analysis and discussion are separately assessed sections, and their boundary matters. Findings should present the data with tables and displays and minimal comment; discussion should interpret it against the research question and the literature reviewed earlier.
Two practical habits earn marks reliably. Refer back to your research question explicitly when interpreting each finding, so the link is visible rather than implied. And where a finding contradicts your expectation or the literature, say so and explore it rather than passing over it — an unexpected result honestly discussed is worth more than a tidy one asserted.
Exam technique for analysis and presentation of findings
Know which average suits which data, and be ready to justify the choice. Mean for numerical data without extremes, median for skewed data, mode for categories.
Practise the move from description to interpretation. Questions frequently give you a figure and ask you to comment; restating the figure earns almost nothing.
Use the four-step structure for any "comment on this finding" question: state the pattern, interpret it, draw the implication, then qualify what the evidence can support.
Know the presentation forms and their uses. Bar chart for comparison across categories, pie chart for parts of a whole, line graph for change over time, histogram for grouped continuous data, table for precise values.
Qualify every claim. Suggests and among respondents in this study are defensible; proves and all young people are not.
Quick revision summary
Raw data must be coded and tabulated before it becomes findings, with open responses assigned to categories whose definitions are stated. Convert frequencies to percentages when comparing groups of different sizes, and report actual numbers alongside percentages where samples are small. Choose the mean for numerical data without extremes, the median where extreme values would distort it, and the mode for categorical data. Present with tables for precise values, bar charts for comparison, pie charts for parts of a whole, line graphs for change over time and histograms for grouped continuous data — each numbered, titled, labelled and sourced. Analyse qualitative data thematically, supporting each theme with a short anonymised quotation followed by your comment. The decisive skill is moving from description, which restates what the data shows, to interpretation, which explains what it means for the research question. Cross-tabulation reveals relationships a single distribution conceals, but an association is not a cause. Scale every claim to the evidence: a study of forty students at one school describes that school, and language such as suggests is defensible where proves is not.