A step-by-step method for verifying a summary, report or letter drafted by AI: fixing what you check, breaking it into claims, finding the evidence, giving honest verdicts, looking for omissions and keeping a record that stands up later.
Reading an AI-drafted document and deciding it looks right is not a check. It leaves no record, it depends on who happened to read it, and it is almost guaranteed to miss what was left out. This guide sets out a method you can follow by hand on a single document, and scale up with automation once the method is settled.
Step 1: Fix exactly what you are checking
Before anything else, record the exact versions of the output and of every source document it was produced from, plus what produced it: the system or model, the date and the instruction it was given.
Take a digest of each file. A digest, such as SHA-256, is a short fingerprint of a document's exact content; any change to the document will, with very high probability, produce a different digest. Recording it lets anyone confirm later that the check was done against these exact documents and that neither has changed since.
Step 2: Break the output into claims
Split the document into atomic claims: short statements that each say one checkable thing. "The policy covers flood damage and excludes subsidence" is two claims, because one half can be right while the other is wrong. This is the approach the research on factuality uses, because a single sentence frequently mixes supported and unsupported facts (Min et al., EMNLP 2023).
Leave out what cannot be checked against a source, such as recommendations, opinions and courtesies, but note that you left them out. Mark attributions carefully: "the report states that profits rose" is a claim about the report, not about profits.
Step 3: Find the passage each claim rests on
For every claim, search the sources for the passage that bears on it and record it word for word, with its location. This is the step that turns a check into evidence: anyone reading the result can see what each verdict rests on without taking your word for it.
Step 4: Give each claim one of four verdicts
- Supported: the source says so. Record the passage.
- Contradicted: the source says otherwise. Record the passage that contradicts it.
- Unsupported: the source does not address it either way. This is not the same as wrong; the claim may be true but it did not come from your source, which in a regulated document is a finding in itself.
- Disputed: the checkers disagree. Leave it for a person who knows the subject to decide, and record their decision.
Resist the temptation to collapse these into right and wrong. An unsupported claim marked wrong overstates the problem; one marked right hides it.
Step 5: Look for what was left out
Now reverse the direction. Go through the source and list the facts that matter for the purpose of the document: the exclusions, conditions, dates, amounts, risks and caveats a reader would need. Check each one against the output.
Report omissions as their own list, beside the claim results, never averaged into them. This is the step most checks skip, and it is where much of the risk sits: in one clinical study, AI-drafted notes omitted far more sentences than they invented (Asgari et al., npj Digital Medicine, 2025).
Step 6: Report every count
Summarise with the counts, not a single score: how many claims were checked, and how many were supported, contradicted, unsupported and disputed; how many material facts were listed from the source, and how many were missing from the output. State what was out of scope. A reader should be able to compute any percentage they want from what you report, and see exactly what it would leave out.
Step 7: Keep the record
Store the digests, the claims, the verdicts, the source passages, the omissions, who or what performed each check, and the method and date. That record is what you produce when a regulator, an auditor, a client or a court asks how you know the document was right.
Scaling it up with automation
Doing this by hand is slow, which is why it is rarely done. Each step can be automated, with three rules that the research supports.
- Never let the model that wrote the output check it. Model evaluators tend to favour their own outputs (Panickssery et al., NeurIPS 2024).
- Use more than one independent checker. A panel of models from different families was more reliable and less biased than a single large judge (Verga et al., 2024).
- Send every disagreement to a person. Automated checkers are good on average and still wrong on individual cases; disputes are where a human adds most.
Start by checking a handful of documents by hand and with automation side by side. Where the two agree consistently, let the automation carry the volume and keep people on the disputes and the omissions.
A checklist for each document
- Output and source versions fixed, with digests recorded
- Every checkable statement split into a single claim
- A source passage recorded for every supported or contradicted claim
- Unsupported kept separate from contradicted
- Disputed claims decided by a person
- Material facts listed from the source and checked for omission
- Every count reported, with the method and date
To see what the finished record looks like, open the sample Evidence File: one invented source, four outputs, and a contradiction, an unsupported claim and an omission planted among them.
The Accuracy Evidence File applies this method to up to 30 of your AI outputs in three weeks at a fixed price, using independent checking models and your reviewer on every dispute.