Kramizo
Log inSign up free
Home › Kramizo AI Literacy › Checking and verifying AI output
Kramizo · · AI Literacy · Revision Notes

Checking and verifying AI output

1,964 words · Last updated October 2026

⚡
Ready to practise? Test yourself on Checking and verifying AI output with instantly-marked questions.
Practice now →

What you'll learn

  • The difference between an output being plausible and being verified
  • Lateral reading: checking the source rather than only the claim
  • How to check a citation properly — existence is not the same as support
  • Why triangulation fails when three sources all trace back to one
  • Internal consistency checks: recalculating rather than re-reading
  • How to design a spot-check when you cannot verify everything
  • When the cost of verification means you should not have used AI for the task at all

Key terms and definitions

Term Meaning
Plausibility Looking like the kind of thing that would be true
Verification Establishing from independent evidence that something is true
Triangulation Confirming a claim against two or more genuinely independent sources
Circular reporting Several sources appearing to agree because all repeat one original
Lateral reading Leaving the claim to investigate the source before judging the claim
Primary source The original record or study
Secondary source Something describing or interpreting a primary source
Misattribution A real source cited for a claim it does not actually make
Internal consistency Whether the parts of an answer agree with one another
Spot-check Verifying a deliberately chosen sample when checking everything is impractical

Core concepts

Plausible is not verified

The single most common failure in using AI is accepting output because nothing about it looked wrong.

Plausibility is a judgement about whether something resembles the truth. It draws on your existing expectations — and a generative model is, in effect, a machine for producing exactly that resemblance. Plausibility is therefore the one test these systems are guaranteed to pass, including when they are wrong.

Verification requires information from outside the output: a source, a calculation you did yourself, a document you can open.

The practical form of this is a question worth asking of anything important:

What evidence, from outside this answer, tells me it is right?

If the honest answer is "it sounds right", you have not checked it. You have recognised it.

Lateral reading: check the source, not just the claim

Faced with a claim and a source, most people read down — they examine the claim more closely and decide whether it feels sound.

Lateral reading means leaving the claim alone and investigating the source instead:

  • Does this organisation exist, and what is it?
  • Who funds it, and does it have an interest in the conclusion?
  • Is this a research body, a campaign group, or a company selling something?
  • Is this a primary source, or is it describing someone else's work?

This is quicker than evaluating a claim on its merits, and far more reliable when the subject is outside your knowledge — which is exactly when you most need help.

Checking a citation properly

A citation has two separate things that can be wrong, and most people only check the first:

1 — Does it exist? Search for the title, the authors, the journal. A fabricated reference typically fails here.

2 — Does it say what was claimed? This is the step people skip. A real paper can be cited for a conclusion it does not reach, or reaches with heavy qualifications. This is misattribution, and it is in some ways worse than a fabrication, because the reference survives a casual check.

So: find the source, then find the claim inside the source. A reference you have located but not read is not yet evidence.

Triangulation, and circular reporting

Confirming a claim against several sources is sound in principle. The failure mode is circular reporting: three pages agree because all three copied one original, which may itself have been wrong.

Signs that your three sources are really one:

  • The wording is suspiciously similar
  • They all cite the same single study, or none cites anything
  • They appeared within a short period after one publication
  • Following each one's own references leads to the same place

Genuine triangulation needs independent origins. One primary source you have read beats five pages describing it.

Internal consistency: recalculate, do not re-read

Some errors can be caught without any external source, because the answer disagrees with itself:

  • Figures in a table that do not add to the stated total
  • A percentage that does not match the raw numbers given
  • A conclusion that does not follow from the stated steps
  • A date in the text contradicting a date in a timeline
  • A quantity in the wrong units

The important discipline here is recalculating rather than re-reading. Reading a calculation tends to confirm it, because following somebody's steps is not the same as doing the work. Do the arithmetic yourself, on paper or with a calculator, and compare the answer.

This is a cheap check that requires no sources, and it catches a useful proportion of errors in anything numerical.

Spot-checking when you cannot check everything

Faced with fifty rows of data or a long report, full verification may be unrealistic. A spot-check is legitimate — provided it is designed rather than casual.

A casual spot-check looks at the first two rows. A designed one samples where errors are likeliest and where they would matter most:

  • The extremes — the largest and smallest values, where mistakes show up
  • Anything surprising — a figure that would change your conclusion if wrong
  • A random few, so you are not only looking where you expect problems
  • Every total or derived figure, since these depend on everything else

And interpret the result honestly. If a sample of six contains one error, you have not found one error — you have evidence that the whole set has an error rate, and the rest needs checking.

Why asking the model to check itself is weak

"Check your answer for mistakes" produces a reply that looks like checking. It is the same system, with the same gaps, predicting what a checking response looks like — and it will often "find" a problem because that is what a critique is expected to contain, or confirm the work because agreement is the cooperative move.

It is not useless: it sometimes surfaces a genuine inconsistency, and asking for the reasoning gives you something to inspect. But it is not verification, because no information from outside has entered the process.

A better version of the same idea is to ask it to do the task a second time independently, without showing it the first attempt, and compare. Two independent attempts that differ is useful information. One attempt reviewing itself is not.

When verification costs more than the task

A conclusion people resist: for some tasks, checking the output properly would take longer than doing the work yourself.

If you must verify every figure in a generated table against the original data, you have read the original data anyway — and the AI step added time. This is not an argument against using AI; it is an argument for noticing which tasks it genuinely helps with:

  • Good fit: drafting, restructuring, explaining a concept you can then test against a textbook, generating options you will judge
  • Poor fit: producing specific facts, figures or citations you will then have to confirm one by one

The decision rule is: if verifying would be as much work as the task, do the task.

Record what you checked

For anything submitted or acted on, keep a note of which claims you verified and how. It takes seconds, prevents the same checking twice, and means that when something is challenged you can say precisely what you confirmed and what you did not.

Worked examples

Example 1: Diagnosing a citation problem (4 marks)

A student finds that a reference an AI gave does exist, so accepts the claim. The paper, when read, does not support it. Name the problem and explain how to avoid it.

  • This is misattribution: a real source cited for a claim it does not make (1 mark)
  • Checking existence alone cannot detect it, because the reference survives that check (1 mark)
  • The claim has to be located inside the source, not just the source located (1 mark)
  • A reference found but not read is not yet evidence (1 mark)

Example 2: Apparent agreement (3 marks)

Three websites agree on a statistic. Explain why this may not confirm it, and what would.

  • They may all be repeating one original, which is circular reporting (1 mark)
  • Signs include near-identical wording and each citing the same single study (1 mark)
  • Genuine confirmation needs sources with independent origins; one primary source, read, beats several describing it (1 mark)

Example 3: Designing a spot-check (4 marks)

An AI produces a 50-row table of figures. Full checking is impractical. Describe a spot-check and how to interpret it.

  • Check the extreme values, where errors are most visible (1 mark)
  • Check any figure that would change the conclusion if it were wrong, and recalculate every total (1 mark)
  • Include a few chosen at random, so the sample is not only where problems are expected (1 mark)
  • If the sample contains an error, treat it as evidence of a rate across the table rather than as one isolated mistake (1 mark)

Common mistakes and how to avoid them

  • Accepting output because nothing looked wrong. Plausibility is the one test the system always passes.
  • Checking a citation exists and stopping. Find the claim inside the source.
  • Counting three pages as three sources. Follow their references; they may be one.
  • Re-reading a calculation. Do it yourself and compare answers.
  • Asking the model to check its own work and treating that as verification. No outside information has entered.
  • Spot-checking the first few rows. Sample the extremes, the surprises and the totals.
  • Treating one error found as the only error. It is a rate, not an incident.
  • Verifying everything in a task where that is the whole job. Then do the job.

Using this in practice

A workable order for anything that matters:

  1. Ask what outside evidence supports it. If the answer is "it sounds right", start checking.
  2. Recalculate the numbers — totals, percentages, units.
  3. Check internal consistency before going to any source; it is free.
  4. Locate each source, then locate the claim inside it.
  5. Read laterally on any source you do not recognise.
  6. Check whether your sources are independent.
  7. Record what you checked, and what you did not.

Quick revision summary

  • Plausible is not verified: plausibility is the one test a generative system is guaranteed to pass
  • Ask what evidence from outside the answer supports it; "it sounds right" is recognition, not checking
  • Lateral reading investigates the source rather than the claim, and is faster and more reliable outside your own knowledge
  • A citation can exist and still be misattributed — find the claim inside the source
  • Circular reporting makes three sources look like confirmation when they are one; triangulation needs independent origins
  • Catch what you can for free with internal consistency, and recalculate rather than re-read
  • Design a spot-check around extremes, surprises, totals and a random few; one error found is a rate, not an incident
  • Asking a model to check itself adds no outside information; two independent attempts compared is a better test
  • If verifying would be as much work as the task, do the task

Checking and verifying AI output: common questions

What are the most common mistakes in Checking and verifying AI output?

Accepting output because nothing looked wrong: Plausibility is the one test the system always passes. Checking a citation exists and stopping: Find the claim inside the source. Counting three pages as three sources: Follow their references; they may be one.

Where can I practise Checking and verifying AI output questions for free?

Kramizo has free Kramizo AI Literacy practice questions on Checking and verifying AI output, each marked instantly with a full explanation. No card is required.

Free for students

Lock in Checking and verifying AI output with real exam questions.

Free instantly-marked Kramizo AI Literacy practice — 45 questions a day, no card required.

Try a question →See practice bank