What you'll learn
- What a token is, and why a model reads text in pieces rather than letters or words
- The single mechanism behind every chatbot: predicting what comes next
- Why the same prompt can produce a different answer each time
- What a context window is, and why long conversations start losing the beginning
- Why these systems are unreliable at arithmetic and at counting letters
- Why fluent writing is not evidence of truth, and why the model cannot tell when it is wrong
Key terms and definitions
| Term | Meaning |
|---|---|
| Large language model (LLM) | A model trained on very large quantities of text to predict what text comes next |
| Token | The unit of text a model works in — roughly a short word or a piece of a longer one |
| Tokenisation | Splitting text into tokens before the model processes it |
| Next-token prediction | Producing a probability for every possible next token, then choosing one |
| Sampling | Choosing the next token from those probabilities rather than always taking the likeliest |
| Temperature | A setting controlling how adventurous that choice is; low is predictable, high is varied |
| Context window | The maximum amount of text the model can take into account at once |
| Prompt | The text you supply, which becomes part of the context the prediction is based on |
| Pre-training | The long, expensive stage in which the model learns language patterns from text |
| Fine-tuning | Further training that adjusts a pre-trained model towards particular behaviour |
Core concepts
Text arrives as tokens
A model does not read characters, and it does not quite read words. Text is first broken into tokens — common short words tend to be one token, longer or unusual words are split into several, and spaces and punctuation are included.
So "understanding" might arrive as under, stand, ing. The model never sees the letters as letters; it sees identifiers for the pieces.
This sounds like an implementation detail and it explains a family of odd failures:
- Asking how many times the letter r appears in a word is awkward, because the model is not looking at letters
- Reversing a word, or spelling it out, is harder than its general fluency suggests
- Rhyme and wordplay are less reliable than the writing quality implies
- Unusual names and technical terms split into more pieces and are handled less consistently
One mechanism: what comes next
Here is the whole thing. Given the text so far, the model produces a probability for every token it could produce next. Then one is chosen, appended, and the process repeats with the longer text as input.
That is all a chatbot does. There is no separate step where it looks up a fact, checks a claim, or decides what it believes. Writing an essay, answering a question and inventing a story are the same operation applied over and over.
Two consequences follow immediately, and most of this course rests on them:
- The model is optimised for text that is likely, not text that is true. Those usually coincide, because accurate text is common in writing. When they come apart, likelihood wins.
- A right answer and a wrong one are produced the same way. There is no internal signal that distinguishes them, which is why the system cannot warn you.
Why answers vary between runs
If the model always took the single likeliest token, it would give the same answer every time — and its writing would be noticeably flat and repetitive. Instead the next token is usually sampled from the probabilities, so a less likely token is sometimes chosen.
Temperature controls this. Low temperature makes output predictable and repetitive; high temperature makes it varied and more prone to drift away from the point.
This is why asking the same question twice can give two different answers, and why neither is more authoritative than the other. It also means that getting the same answer twice is not confirmation — ask in the same way and you sample from much the same distribution.
The context window
The context window is the amount of text the model can attend to at once: your prompt, the conversation so far, and any documents you have pasted in, all counted in tokens.
Three things follow:
- Nothing outside the window exists for the model. In a long conversation, the earliest exchanges fall out, which is why a chatbot can appear to forget what you agreed at the start.
- Between conversations there is normally no memory at all. A new chat begins with nothing, unless the product deliberately stores notes and feeds them back in.
- Pasting a document does not teach the model anything. It puts text in the window for this conversation. The model is unchanged, and nothing is remembered afterwards.
The last point is worth being clear about, because "I trained it on my notes" is a common and wrong description of pasting notes into a chat.
Training and using are different stages
Pre-training is where the model learns language patterns from an enormous quantity of text. It is slow and expensive, happens once, and fixes what the model knows. After it, the model's parameters stop changing.
That has a consequence people run into constantly: a knowledge cut-off. The model's sense of the world ends when its training data ended. It does not know that time has passed, so asked about something recent it will often answer confidently from an older state of affairs rather than saying it cannot know.
Fine-tuning adjusts a pre-trained model towards particular behaviour — being helpful in conversation, following instructions, declining certain requests. It shapes how the model responds far more than what it knows.
When you type a question, neither stage is happening. You are doing inference: the finished model predicts tokens. Nothing you type changes it.
Why arithmetic is unreliable
A calculator implements addition. A language model predicts text, and nothing inside it performs arithmetic.
For small sums the right answer is overwhelmingly the likeliest continuation, so it appears correct. For longer calculations, unusual numbers or multi-step problems, the likeliest-looking continuation can simply be wrong — and it will be laid out in exactly the confident form a correct answer takes.
This is why many tools now hand calculations to an actual calculator behind the scenes. Where that is not happening, arithmetic from a language model needs checking.
Fluency is not truth
The hardest habit this topic asks for is separating how good the writing is from whether it is right.
Fluent, well-organised, confident prose is what the model is best at, because that is what it was trained to produce. The signals people normally use to judge reliability — coherence, appropriate vocabulary, a confident tone, a plausible structure — are precisely the signals a language model generates whether or not the content is accurate.
So the usual instincts point the wrong way. A model does not sound hesitant when it is on thin ground, because its sense of likelihood is about text, not about truth.
Worked examples
Example 1: Explaining varied output (3 marks)
A student asks a chatbot the same question twice and receives two different answers. Explain why, and state whether one of them is more trustworthy.
- The next token is sampled from a probability distribution rather than always taken as the single likeliest (1 mark)
- So repeated runs can follow different paths and produce different wording or content (1 mark)
- Neither answer is more authoritative; both were produced by the same process and both need checking (1 mark)
Example 2: Diagnosing forgetfulness (3 marks)
During a long conversation, a chatbot starts contradicting something agreed near the beginning. Explain what has happened.
- The model can only attend to text inside its context window (1 mark)
- As the conversation grows, the earliest exchanges fall outside that window (1 mark)
- The model is not recalling and rejecting the earlier agreement — from its point of view it is no longer there (1 mark)
Example 3: Judging a claim (4 marks)
A user says "I trained the AI on my revision notes by pasting them into the chat." Explain why this description is wrong, and give the accurate version.
- Pasting text places it in the context window for that conversation only (1 mark)
- Training adjusts a model's parameters, which inference does not do (1 mark)
- The model is unchanged, and the notes are gone once the conversation ends (1 mark)
- Accurately: the notes were supplied as context, not used for training (1 mark)
Common mistakes and how to avoid them
- Saying the model "looks up" an answer. There is no lookup. It predicts tokens.
- Treating a confident tone as a reliability signal. Confidence is a feature of the writing style, not a measure of accuracy.
- Thinking the same answer twice confirms it. Asking the same way samples from the same distribution.
- Saying you "trained" a model by pasting text. That is context, not training.
- Expecting it to know what it does not know. There is no list of what went into training, so it cannot report the gaps.
- Assuming it can count letters because it writes well. It works in tokens, not characters.
- Describing a knowledge cut-off as forgetting. It never knew; training ended.
Using this in practice
Knowing the mechanism changes how you use the tool:
- Give it context rather than expecting recall. If something matters, paste it in; do not rely on the model to have retained it.
- Ask for reasoning you can inspect, so you are checking a visible chain rather than trusting a conclusion.
- Verify anything specific — numbers, dates, names, references. These are exactly where likely-looking text and true text come apart.
- Re-ask differently, not identically. A genuinely different phrasing is a weak test; the same phrasing is no test at all.
- Use a calculator for calculations, and a search for anything recent.
Quick revision summary
- Text is split into tokens; the model never sees letters as letters, which explains its trouble with spelling, counting and rhyme
- Everything a chatbot does is next-token prediction: produce probabilities, choose one, repeat
- It is optimised for likely text, not true text, and a right answer is produced by the same process as a wrong one
- Output varies because tokens are sampled; temperature sets how varied, and repetition is not confirmation
- The context window bounds what it can consider — long chats lose the start, and new chats normally begin with nothing
- Pre-training fixes what it knows and creates a knowledge cut-off; fine-tuning shapes how it responds
- Pasting text is context, not training — the model is unchanged
- Arithmetic is unreliable because nothing inside the model calculates
- Fluency is not truth, and the model has no internal signal telling it when it is wrong