Prove It · Module 1: Why it lies
Where it fails most
Errors are not evenly spread. They cluster around specific kinds of question, and knowing the clusters tells you where to look hardest.
Specifics: exact dates, figures, quotations, section numbers, names of people and papers. The model has seen the shape of these and reproduces the shape. Recent events: anything after its training cut-off, which it may not know it does not know. Niche topics: the less it has read about something, the more it fills in. Long chains of reasoning: each step is probably right; the chain may not be. Anything you have asked it to confirm: models agree with the premise of a question more often than they should.
Conversely, it is most reliable on widely-documented general knowledge, on transforming text you have given it (summarising, reformatting, rewording), and on explaining well-established concepts. Those are the tasks where fluency and accuracy mostly travel together.
| Higher risk | Lower risk |
|---|---|
| Exact dates, figures, quotations, references | Explaining a well-known concept |
| Recent events and current office-holders | Summarising or reformatting text you supplied |
| Niche or specialist topics | Widely-documented general knowledge |
| Long multi-step reasoning | Single-step tasks with a clear answer |
| Questions that contain an assumption | Open questions with no premise built in |
In practice. A consultant used AI to draft a market overview. The general framing was accurate and useful. Every specific market-size figure was wrong, several by an order of magnitude, and two of the four cited reports did not exist. The structure was the safe part. The numbers were the dangerous part, and they were the part the client would have quoted.
At your desk. Take a piece of AI output you have used recently. Highlight every specific: every date, figure, name, quotation and reference. Count them. Those are the checks you owe.
Write it down.
- Specifics in my sample:
- How many I had actually checked:
- The one that worried me: