Skip to content
AI Accuracy

An accuracy score is not evidence: how to measure AI output so the result stands up

, Chairman & CTO, ETT1 October 20268 min readFirst published on ettgroup.ai

"Our AI is 95% accurate" answers none of the questions an auditor will ask. Measured against what? Counting what? And what about the facts it left out? Here is what the research says a defensible measurement looks like.

Ask a vendor how accurate their AI is and you will usually get a percentage. Ask what the percentage counted, against what, and who did the counting, and the conversation gets shorter. A single score sounds like evidence. On its own, it is not.

What does an accuracy score actually count?

Usually one of three different things, and they are rarely labelled. It might be the share of answers that matched a reference answer, the share of individual statements that a checker found support for, or a reviewer's overall impression of a sample. Each measures something different, and a figure that does not say which it is cannot be compared with anything, including last quarter's figure from the same system.

Why count claims rather than answers?

Because one answer can be partly right. The research on long-form factuality settled on breaking text into small, checkable statements and checking each one. FActScore, a widely cited method, splits a generation into atomic facts and reports the percentage supported by a reliable source (Min et al., EMNLP 2023).

Its findings show why answer-level scoring misleads. Checked this way, only 58% of the facts in ChatGPT's biographies of people were supported. And in 40% of cases a single sentence mixed supported and unsupported facts, so marking whole sentences, let alone whole answers, as right or wrong hides exactly the errors that matter.

Why does the denominator matter so much?

A percentage is a numerator over a denominator, and the denominator is where the judgement hides. Ninety-five per cent of what? Of claims the checker could find a source for, or of all claims? If unverifiable claims are quietly dropped from the count, the score rises while nothing about the output improves. A defensible report states every count: claims checked, supported, contradicted, unsupported and disputed, so a reader can do the division themselves.

What does a single score leave out entirely?

Omissions. A score built on checking what the output says cannot see what it failed to say. The authors of FActScore are explicit that it measures precision, not recall: a response can have every fact supported and still miss a significant piece of information.

That is not a theoretical gap. In a clinical study where 50 doctors reviewed 450 AI-drafted consultation notes, omissions were more common than hallucinations: 1,712 omitted sentences against 191 hallucinated ones, although hallucinations were more often judged major (Asgari et al., npj Digital Medicine, 2025). A separate study of AI-drafted emergency department summaries found that 47% of GPT-4 summaries omitted clinically relevant information (Williams et al., PLOS Digital Health, 2025).

Omissions need their own count, reported beside accuracy and never averaged into it. Blend the two and a summary that leaves out half the material can score well, because everything it does say is true.

Can an AI check AI?

Yes, with care, and the research is useful here. Automated checkers can be good: Google DeepMind's SAFE agreed with crowdsourced human annotators on 72% of around 16,000 facts and was right in 76% of a sample of the disagreements, at more than twenty times lower cost (Wei et al., NeurIPS 2024).

But a single AI judge brings biases. In the study that popularised using a model as a judge, the judge gave a consistent verdict only 65% of the time when two answers were simply presented in the opposite order (Zheng et al., NeurIPS 2023). Models also tend to favour their own writing: evaluators that can recognise their own outputs score them more highly (Panickssery et al., NeurIPS 2024).

Two practical conclusions follow. Never let a model mark its own work. And prefer a panel of independent checkers from different model families over one large judge: a panel of three smaller models outperformed a single large judge and showed less bias, at a fraction of the cost (Verga et al., 2024). Where the panel disagrees, a person decides.

What a defensible measurement includes

  • Claim-level checking, with the source passage behind every verdict recorded
  • Four verdicts, not two: supported, contradicted, unsupported and disputed
  • Every count stated, so the reader can see the denominator
  • Omissions measured separately and never averaged into accuracy
  • Independent checkers, never the model that wrote the output, with disagreements left for a person
  • The method, thresholds and model versions written down, because a result is a property of a particular setup

Key takeaways

  • A single accuracy score cannot be compared or audited without its definition and denominator
  • Check claims, not answers: one sentence often mixes right and wrong
  • Omissions are common and invisible to a fact-check; count them separately
  • Use independent checkers and keep disputed claims for people

The Accuracy Evidence File measures a set of your AI outputs this way, in three weeks at a fixed price, and gives you the register to show for it.

AI AccuracyAccuracy Evidence FileEvaluation

Want this record for your own AI outputs?

The Accuracy Evidence File checks up to 30 of your AI outputs against their sources, claim by claim, with every omission listed and a certificate for each. Three weeks, fixed price.

See the Accuracy Evidence File