How to check AI output before you trust it
A better brief gets you a better draft. It does not tell you whether the draft is right. One in 277 peer-reviewed papers now cites work that was never written, and the checking nobody taught you is the part of the job that is left.

Kingsley Ijomah
Founder, Codehance
A better brief gets you a better draft. It does not tell you whether the draft is right, and that is the part almost nobody has been taught.
I notice this most when I sit with teams who are already good at this. They have stopped writing bad prompts. They get useful output on the first or second attempt. And then they take that output and do something with it, and the checking step is either missing or is a quick read for tone. When I ask what they actually verified, the honest answer is usually that it looked fine.
That is not laziness. Nothing in the experience of using these tools tells you that checking is your job now. The output arrives formatted, confident and finished. Every signal the interface gives you says the work is complete, and the one thing it cannot tell you is whether any of it is true.
There is a study I keep coming back to on this. Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real uses of these tools in their actual jobs. Two findings sit next to each other and explain most of what I see. The more confidence someone had in the tool, the less critical thinking they applied. The more confidence they had in themselves, the more they applied.
That is not the obvious result. Trusting the machine and trusting yourself pull in opposite directions, which means the person most at risk is not the anxious sceptic checking everything twice. It is the person who has had a good run.
The same study found something more useful still: for the people doing this well, the thinking had not disappeared, it had moved. It shifted from producing the work to verifying it, integrating it, and deciding what to do with it. That is the job now, and almost nobody has been shown how to do it.
Reading output like an editor
An author and an editor read the same page completely differently. The author reads to see whether it says what they meant. The editor reads to find what is wrong with it. When you use one of these tools you are not the author any more, and continuing to read like one is the single most expensive habit in this whole field.
The specific failure has a name in the safety literature: automation complacency. When a system is right most of the time, reviewers stop genuinely reviewing. They skim, they pattern-match, they approve. The reason it is worse here than with older automation is that these errors do not look like errors. A broken machine grinds or stops. This one produces something well formed and wrong, in exactly the register of the parts that were fine.
So the editorial pass has to be deliberate, and it is mostly one move: check the things that can be checked, and be honest that the rest is unverified.
In practice that means going after specifics rather than reading for flow. Names, numbers, dates, quotations, citations, version numbers, any claim about what a document says. Those are checkable, and they are exactly where invention lives, because they are the parts that had to be recalled rather than constructed.
It also means reading for the thing that is missing. Fabrication is easy to imagine and easy to look for. Omission is harder, because a summary that quietly drops the one clause that mattered reads perfectly. When you have given it a document, the question is not only whether everything it said is in there. It is whether everything that mattered came out.
And there is a limit worth stating plainly. If you could not have written it, you cannot review it. Output you are not qualified to evaluate is not a draft you can use, whatever it looks like. That is not a reason never to ask about unfamiliar things. It is a reason to treat what comes back as something to learn from and then confirm elsewhere, rather than something to forward.
Where did that come from
The numbers here are worse than most people realise, and unlike almost everything else in this field, they are moving in the wrong direction.
In 2026 The Lancet published an audit of roughly two and a half million biomedical papers, looking for references to work that does not exist. About one paper in 277 published in the first seven weeks of 2026 cited something that was never written. The year before it was one in 458. In 2023 it was one in 2,828.
Sit with that trend for a second. These are peer reviewed papers, in a field with more checking than almost any other, and fabricated references are getting through at a rate that has roughly quadrupled in three years. The audit verified around 97 million references to find them, and turned up 4,406 fake citations across 2,810 papers. More than 98% of the affected articles had seen no publisher action at the time of the audit.
Law is no better. Stanford's RegLab put over 800,000 verifiable legal questions to public-facing models and found hallucination rates of 58% for GPT-4, 69% for GPT-3.5 and 88% for Llama 2. Those were 2023-era models and the newer ones are better, which is worth saying. They are not fixed, and filings containing invented cases keep reaching courts, filed by firms with everything to lose by it.
Two things follow, and they are both practical.
Turning on web search reduces this and does not remove it. When Stanford researchers audited four generative search engines, engines that do retrieve real pages and cite them, they found only 51.5% of generated sentences were fully supported by the citations attached to them, and only 74.5% of citations actually supported the sentence they were attached to. A citation is not proof. Opening it is proof.
The less written about your subject, the more it invents. This follows directly from the mechanism. The OpenAI and Georgia Tech work on why models hallucinate makes the point precisely: facts that appear rarely in training data produce unavoidable errors, while recurring regularities do not. A well documented framework has been written about ten thousand times. Your industry's niche regulation has not. Where the data is thin, the machine fills the gap rather than reporting one.
Turn that into a working habit. Before you accept a specific, ask yourself how much has plausibly been written about it. A famous case, a common library, a mainstream standard: probably sound, still worth a glance. A small company, a recent event, a local rule, anything from a corner of the world that gets written about less: assume it is invented until you have seen it somewhere else.
Provenance is the whole discipline in one question. Where did that come from? If the answer is a page you can open, you have something. If the answer is that it sounds right, you have nothing yet.
When a human has to look
Some things you check yourself. Some things need somebody else, and the difference is worth deciding in advance rather than in the moment.
And the evidence here is not the comfortable result you might expect. A meta-analysis in Nature Human Behaviour pooled 370 effect sizes from 106 experiments on humans alone, AI alone, and the two combined. On average, the combination performed significantly worse than whichever of the two was better on its own. Adding a human to the loop is not automatically an improvement. Quite often it makes things worse.
The detail underneath that is what makes it useful. Combining worked when each side brought something the other lacked: on bird identification, humans scored 81%, the system 73%, and together they reached 90%. It failed when the human could not really do the task. On spotting fake hotel reviews the system managed 73% on its own and the pairing dropped to 69%, because the humans were guessing. The split ran deeper too, with combinations underperforming on decision tasks and outperforming on creative ones.
So review is not a safeguard you can add. It is a safeguard you can add when the reviewer can actually do the job, and the failure mode when they cannot has a name I would put on a wall: accountability theatre. If the reviewer cannot realistically evaluate what they are looking at, or there is too much of it, or they have no authority to send it back, then the review is a signature rather than a check. Everyone involved believes the work has been checked. Nobody has checked it.
That is not hypothetical. In a Wolters Kluwer survey of 230 US banking professionals, automation bias, defined as reviewers deferring to algorithmic output rather than exercising independent judgement, was picked by 34% as the single greatest human-centred risk to their AI safety framework. It came out ahead of misaligned incentives, the skills gap and shadow AI. Not the models. The people nominally watching them.
So set the threshold on purpose. The questions I would ask about any piece of work before it leaves your hands:
Does it leave the building? Anything a customer, a regulator, a client or a board will see has a different standard than something that helps you think.
Is it hard to reverse? A sent email, a published figure, a committed number in a plan. Reversibility is a better test than importance, because it survives being wrong.
Would being wrong be expensive, or embarrassing, or both? Say the number out loud in your own steering meeting and imagine somebody asking where it came from.
Am I the right person to be checking this? If the answer is no, saying so is the check.
And one that matters more the better things go: who is checking the work when it has been right for a month? Complacency is what a good run does to a team, and it arrives without announcing itself. If your process depends on nobody getting comfortable, it is not a process.
See it for yourself
This one is less comfortable than the other two, which is why it is worth doing.
Find the last thing you accepted from one of these tools and passed on. A summary you forwarded, a paragraph that went into a document, a number that ended up in a deck. Something that has already left your hands.
Read it as an editor rather than an author. Underline every specific in it: names, numbers, dates, quotations, references, any claim about what a document says. Then check them. Actually open the things, rather than deciding they look plausible.
Two questions at the end, and the second is the one that matters.
Was anything wrong? And if something had been wrong, would the process you used at the time have caught it?
Most people find the work was fine and the process would not have caught it. That is the useful result, because it tells you that what you have been relying on is a good run rather than a method.
What comes next
If you have read the three posts in this series, you have done Layer 1. That is worth naming, because there is a whole industry telling you that Layer 1 is either trivial or the entire subject, and it is neither.
What that buys you is not a technique. It is prediction. The failures arrive as something you were already watching for, rather than as the thing that caught you out in front of people whose opinion of you matters.
That is the working knowledge of somebody who uses these tools well. For most people it is also enough, and there is nothing second-rate about stopping here.
If it turns out not to be enough, the map is in the three layers. Layer 2 is building things other people use, where the model is one component among many and most of the work is everything around it. Layer 3 is making the models themselves, which is a different profession.
But that is a decision for later, and it is one you are now equipped to make rather than guess at. For now the useful thing is smaller and immediate. The next time something arrives well written and complete, and you feel the pull to accept it because it reads well, you will recognise the feeling for what it is.
Sources
Every figure in this post is linked below so you can check it yourself, which is the only respectable way to publish an article of this kind.
- Lee, Sarkar, Tankelevitch, Drosos, Rintel, Banks and Wilson, The Impact of Generative AI on Critical Thinking, Microsoft Research and Carnegie Mellon, CHI 2025. The 319 knowledge workers, 936 use cases, and the confidence findings.
- Topaz et al., Fabricated citations: an audit across 2.5 million biomedical papers, The Lancet, 2026. One in 277, one in 458, one in 2,828, and the 4,406 fake citations. Retraction Watch summary.
- Dahl, Magesh, Suzgun and Ho, Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models, Stanford RegLab, 2024. The 58%, 69% and 88% figures, from over 800,000 legal questions put to 2023-era models.
- Liu, Zhang and Liang, Evaluating Verifiability in Generative Search Engines, EMNLP Findings 2023. The 51.5% and 74.5% citation figures.
- Kalai, Nachum, Vempala and Zhang, Why Language Models Hallucinate, OpenAI and Georgia Tech, 2025. Why sparse training data produces unavoidable errors.
- Vaccaro, Almaatouq and Malone, When combinations of humans and AI are useful, Nature Human Behaviour, 2024. The 370 effect sizes, the 81/73/90 bird result, the 73/69 hotel review result, and the finding that combinations underperform on average.
- Wolters Kluwer survey of 230 US bank professionals, reported by American Banker, 2026. The 34% automation bias figure.
Want to build this, not just read about it?
The free 3-day Codehance challenge teaches the architecture-first method hands-on. No coding background needed.
Start the free challenge