Zum Inhalt springenSkip to content

If you have AI output checked, know what your tool does not find

Two classes of checking program, one task: they agree on arithmetic. On what is absent, they diverge.

Andreas Ehstand ·

Why this matters

Public authorities, schools and offices now let software draft their text: decisions, quotes, reports, contract summaries. The same requirement usually applies — the output is checked again, often by a second program and finally by a trained reviewer. Anyone setting up that review faces a practical question: what should the search be aimed at, and how far can the tool be trusted?

We measured that, using two very different classes of checking program on the same task.

What was measured

The basis is a collection of German-language administrative and office texts with deliberately planted errors: wrong totals and percentages, wrong deadlines, invented rules, sentences that contradict their own figure in the same paragraph — and missing information: something that appears in the source material and simply does not appear in the output under review. The collection also contains passages and whole tasks with no errors at all, which is the only way to count how often a program flags something correct.

The same deliberately hard version went to two classes: five inexpensive language models — programs that read and write text — and four current top-tier models from four different providers.

The result in two sentences

Both classes work as a first-pass filter: the inexpensive ones flagged 126 of 150 planted errors, the strong ones 116 of 120 — 84.0 against 96.7 percent, with not a single answer failing to arrive. The difference between the classes does not sit in the overall score but in one error type — and it happens to be the expensive one.

Bar chart. All planted errors found: five inexpensive models 126 of 150 (84 percent), four top-tier models 116 of 120 (97 percent). Missing information at the three hardest spots: 4 of 15 (27 percent) against 10 of 12 (83 percent). False alarms per run: 8 against under 3.
Chart: own measurement, one pass per program.

Where the differences sit

When the text contains a number to check against, almost any program finds the error: a total that does not add up, a percentage that is wrong, a sentence contradicting its own figure. There is an anchor to pull on.

When something is absent, there is nothing to check against — the missing deadline or condition has to be carried in from the source material. This is exactly where the classes part company:

  • At the three spots the inexpensive programs found hardest, they caught the error in 4 out of 15 opportunities. The strong programs, at those same three spots, reached 10 out of 12.

A second, independent run with six top models confirmed the picture: at the same two spots, four of six programs left at least one item behind.

The second figure belongs with the first: false alarms, meaning passages where something correct was marked as an error. Here too the classes differ — on average eight false alarms per run against under three, and on the error-free tasks the strong class stood at zero. A tool that flags everything just in case finds everything and is worthless all the same. A hit rate without a false-alarm figure is not a statement.

The one sentence this comes down to

Missing information — conditions that do not appear in the output under review at all — is the error type where checking programs differ most: stronger programs largely find them, weaker ones regularly leave them behind. The pattern among the weaker programs has a name: Counter-Number Blindness.

Why this is expensive in practice

In administrative work this is the costly half. A wrong total surfaces at the next invoice. A condition that was never in the text never surfaces at all — until somebody enforces it.

And there is a second problem, larger than the first: from the outside, no office can see which model is working inside its tool. The interface is the same, the tone is equally confident, and with the inexpensive tools the price per document checked is no higher either. Anyone reading the label "AI-checked" does not know whether their tool finds what is absent. Call it Model Blind Flight.

Why the trained reviewer is not made redundant

The job gets sharper. Because the label says nothing about how good the tool is at missing information, a review needs two fixed questions, identical every time:

  1. What is in the source material that is absent from the output?
  2. Which of the machine's flags are not errors at all?

The first question covers precisely the gap weaker tools leave open. The second saves the most time: rejecting what was flagged is faster than searching from scratch. A review then stops being a second reading and becomes a targeted search — the difference between a control step that costs time and one that saves money.

What an office can do tomorrow

Two small steps:

  • Do not trust the label

    "Checked with artificial intelligence" says nothing about whether the tool finds missing information.

  • Measure your own tool on a known task

    Take one of your own cases, plant errors in it yourself — including at least one condition you have removed from the output — and run the tool over it. What it finds in that run is a first data point - it is no guarantee for other cases. What it leaves behind stays the trained reviewer's work.

Context: one pass per program, built as a starting point rather than a verdict on individual tools. An independent, reproducible reference distribution — a set of results that someone outside could regenerate step by step — does not yet exist for this question. That is what comes next.

This note is part 1 of the series Overlooked moments between people and AI — eleven parts, all in full on one page.

Andreas Ehstand, Augmanitai

How to cite

Andreas Ehstand: If you have AI output checked, know what your tool does not find. augmanitai.com, 11 September 2026. https://augmanitai.com/en/notiz-gegenzahl-blindheit

Quoting with attribution to Andreas Ehstand and AUGMANITAI is welcome. The measured values are also available in machine-readable form: messungen/gegenzahl-blindheit.json · Page last updated: 12 September 2026.

In other languages

A short version of this measurement, with all the figures, is available in eight further languages:

Unabhängige Forschungs-Publikation — über diese Website wird nichts verkauft.Independent research publication — nothing is sold through this website.

ImpressumLegal notice · DatenschutzPrivacy · Haftungs-Hinweise (Disclaimer)Disclaimer

Hinweis zu künstlicher IntelligenzNote on artificial intelligence

Texte, Bilder und Filme dieser Seite sind mit Unterstützung künstlicher Intelligenz entstanden. Jeder veröffentlichte Inhalt wurde von einem Menschen geprüft und überarbeitet; die redaktionelle Verantwortung trägt Andreas Ehstand.Texts, images and films on this site were created with the support of artificial intelligence. Every published item has undergone human review and editing; editorial responsibility is held by Andreas Ehstand.

Page catalog & data accessJSON