Kramizo
Log inSign up free
Home › Kramizo AI Literacy › How machine learning learns from data
Kramizo · · AI Literacy · Revision Notes

How machine learning learns from data

2,068 words · Last updated October 2026

⚡
Ready to practise? Test yourself on How machine learning learns from data with instantly-marked questions.
Practice now →

What you'll learn

  • What training data is, and why it sets the limit on what a model can do
  • The difference between supervised learning from labelled examples and unsupervised learning that finds structure
  • What a trained model actually contains — and why it is not a stored copy of its examples
  • What features are, and how the choice of features shapes what a system can notice
  • Why data is split into training and test sets, and what goes wrong without that split
  • Overfitting, generalisation and concept drift: the three things that decide whether a model works outside the lab

Key terms and definitions

Term Meaning
Training data The examples a model learns from
Label The correct answer attached to a training example, such as "spam" or "not spam"
Supervised learning Learning from examples that carry labels
Unsupervised learning Finding structure in data that has no labels, such as grouping similar customers
Feature A measurable property of an example that the model uses, such as a word count or a pixel value
Parameter (weight) One of the adjustable numbers inside a model that training sets
Training The process of adjusting parameters so the model's outputs match the examples better
Inference Using an already-trained model to produce an output for new input
Test set Data held back from training, used to estimate how the model will do on input it has never seen
Overfitting Fitting the training data so closely that performance on new data gets worse
Generalisation Performing well on data the model was not trained on — the actual goal
Concept drift The world changing after training, so the learned patterns stop matching reality

Core concepts

Learning from examples instead of rules

A rule-based spam filter needs somebody to write the rules: block messages containing certain phrases, block senders on a list. Every rule is a decision a person made in advance, and every new trick spammers invent needs a new rule.

A machine learning spam filter is given something different: a large collection of emails, each marked spam or not spam. Nobody tells it what matters. It works out for itself which properties of a message predict the label — and it finds combinations no person would have thought to write down.

This is the whole shift, and everything else in this topic follows from it: the examples do the work the rules used to do. That has a consequence worth stating plainly, because it explains most AI failures you will meet:

A model can only learn what its training data contains.

Supervised and unsupervised learning

Supervised learning uses labelled examples. Each one comes with the answer attached:

  • photographs labelled with what is in them
  • past loan applications labelled repaid or defaulted
  • recordings labelled with the words that were spoken

The model's job is to produce the label for examples it has not seen. Labels usually cost money and time, because people have to supply them.

Unsupervised learning has no labels. The system is given data and asked to find structure in it:

  • grouping customers with similar buying habits, without being told what the groups should be
  • spotting unusual transactions that do not resemble the rest
  • organising documents by topic

The useful question when you meet a system is: was somebody required to supply the right answers? If so, it is supervised, and the quality of those answers limits everything.

What a trained model actually contains

This is the point most often misunderstood, and getting it right explains several behaviours that otherwise look mysterious.

A trained model is a large set of numbers — its parameters, sometimes called weights — together with the structure that says how to combine them. Training adjusts those numbers so that the model's outputs match the training examples more closely.

A trained model does not contain a copy of its training data, and it cannot look anything up in it. The examples shaped the numbers and were then left behind.

That single fact explains a lot:

  • A model cannot cite which example it learned something from, because it has no index of examples
  • It cannot report what it does not know, because there is no list of what went in
  • Two facts that appeared equally often in training are held with equal confidence, whether or not one is true
  • Deleting a fact from a trained model is genuinely hard, because the fact is not stored in one place

Features: deciding what the model can notice

A feature is a measurable property used to describe an example. For an email: how many links it contains, whether the sender is known, the length of the subject line. For a photograph: the values of the pixels.

Features matter because a model cannot use information it was never given. A system predicting exam results from attendance and past grades knows nothing about a student's home circumstances — not because it has decided to ignore them, but because they are not among its features.

In older machine learning, people chose the features by hand. Deep learning changed this: given raw data and enough of it, the network learns useful features for itself, which is why it handles images and audio so much better than earlier methods. The trade-off is that nobody specified those learned features, so nobody can straightforwardly say what the model is using.

Training, testing and why the split matters

Imagine judging a student by letting them see the exam paper during revision and then setting that exact paper. A high mark would prove they could reproduce those questions, not that they understood the subject.

Machine learning has the same trap, so data is split:

  • the training set is used to adjust the parameters
  • the test set is held back and used once, to estimate performance on unseen input

A score on the training set is nearly meaningless on its own. The only number that tells you anything about real use is the one from data the model never trained on.

Overfitting and generalisation

Overfitting is fitting the training examples so closely that the model captures their accidents as well as their patterns.

The signature is unmistakable once you know it: high accuracy on training data, much lower accuracy on test data. The model has effectively memorised rather than learned, and memorising is a failure, because every real use involves input it has not seen.

A classic illustration: a model trained to spot a disease in chest X-rays learns that images from one hospital's machine are more likely to be positive, because that hospital treated more severe cases. It scores well in testing and fails elsewhere, having learned a fact about the equipment rather than about the disease.

Generalisation is the opposite and the actual goal — performing well on data the model has never seen. Everything in practice is a trade-off between fitting the examples you have and generalising to the ones you do not.

Data quality decides capability

Three consequences follow, and each matters more than model size:

  • Gaps become blind spots. A face recognition system trained mostly on one group of people works less well on others. Nothing in the method causes this; the training data does.
  • Errors in labels become errors in behaviour. If some examples were labelled wrongly, the model learns those mistakes as though they were the pattern.
  • More of the same data does not fix a gap. Doubling a dataset that was missing a group still misses that group. The fix is different data, not more data.

Concept drift: models go stale

A model learns the patterns in data gathered at one moment. The world then moves on.

A spam filter trained on last year's spam meets this year's tactics. A model predicting shopping habits meets a change in prices. A model trained on one city's traffic meets a new road layout. Performance decays without anything inside the model changing — this is concept drift, and it is why deployed systems are monitored and retrained rather than finished.

Worked examples

Example 1: Diagnosing a result (4 marks)

A team reports that their model is 99% accurate on its training data but 62% accurate on held-back test data. Name the problem and explain what has happened.

  • The problem is overfitting (1 mark)
  • The model has fitted the accidents of the training examples as well as their general patterns (1 mark)
  • It has effectively memorised those examples rather than learning a pattern that transfers (1 mark)
  • The test score is the one that matters, because every real use involves unseen input (1 mark)

Example 2: Explaining a limitation (3 marks)

A model predicting which students need extra support uses attendance and previous grades. A teacher asks why it never flags students with difficult home circumstances. Explain.

  • Home circumstances are not among the features the model was given (1 mark)
  • A model cannot use information that was never supplied to it (1 mark)
  • The omission is a decision about what data to collect, not something the model chose (1 mark)

Example 3: Evaluating a proposal (4 marks)

A face recognition system works less well for some groups of users. An engineer proposes collecting twice as much data from the existing sources. Evaluate this proposal.

  • The weakness comes from which groups are under-represented, not from the total quantity of data (1 mark)
  • Doubling the same sources keeps the same proportions, so the gap remains (1 mark)
  • The fix is data that covers the under-represented groups — different data rather than more data (1 mark)
  • Performance should then be reported per group, since a single overall score can hide a gap (1 mark)

Common mistakes and how to avoid them

  • Saying a model "stores" its training data. It holds adjusted numbers. The examples shaped them and were left behind, which is why it cannot cite a source.
  • Quoting training accuracy as the result. Only performance on held-back data says anything about real use.
  • Treating overfitting as "trying too hard". It is a specific, diagnosable gap between training and test performance.
  • Thinking more data always helps. More of the same data cannot fill a gap in coverage.
  • Assuming a deployed model stays accurate. Concept drift degrades it even when nothing in the model changes.
  • Confusing training with inference. Training happens once and is expensive; inference is using the finished model and happens every time you type something.

Using this in practice

When you are told a system "learns", four questions will tell you most of what matters:

  1. What was it trained on? Whose data, collected when, covering whom.
  2. Was it labelled, and by whom? Labels carry the judgement of whoever applied them.
  3. How was it tested? On held-back data, or on the data it trained on?
  4. When was it last retrained? A model that has not been updated has been drifting since the day it shipped.

These are also the questions to ask about a tool a school or an employer wants to adopt, and the answers are often more revealing than any accuracy figure.

Quick revision summary

  • Machine learning replaces written rules with examples; a model can only learn what its data contains
  • Supervised learning uses labelled examples; unsupervised learning finds structure in unlabelled data
  • A trained model is adjustable numbers, not a stored copy of its examples — so it cannot cite or look anything up
  • Features decide what a model can notice; information never supplied cannot be used
  • Keep a test set back: training accuracy proves almost nothing
  • Overfitting = high training accuracy, low test accuracy; generalisation to unseen data is the goal
  • Gaps and label errors in data become blind spots and mistakes in behaviour, and more of the same data cannot fix a gap
  • Concept drift means deployed models decay as the world changes, so they need monitoring and retraining

How machine learning learns from data: common questions

What are the most common mistakes in How machine learning learns from data?

Saying a model "stores" its training data: It holds adjusted numbers. The examples shaped them and were left behind, which is why it cannot cite a source. Quoting training accuracy as the result: Only performance on held-back data says anything about real use. Treating overfitting as "trying too hard": It is a specific, diagnosable gap between training and test performance.

Where can I practise How machine learning learns from data questions for free?

Kramizo has free Kramizo AI Literacy practice questions on How machine learning learns from data, each marked instantly with a full explanation. No card is required.

Free for students

Lock in How machine learning learns from data with real exam questions.

Free instantly-marked Kramizo AI Literacy practice — 45 questions a day, no card required.

Try a question →See practice bank