Overlooked moments between people and AI
A series in eleven parts: does the reviewer notice when something in an AI answer is wrong — or missing?
Andreas Ehstand ·
Anyone who signs off drafts from artificial intelligence (AI) knows moments that had no words until now: trust that arrives too fast, the crack after the first error found by chance, the calm of the final click. This series gives those moments clear names and explains them through examples.
All eleven parts are here in full. Each part stands on its own; the series opens with a measurement and two key figures, then come eight moments from the lexicon.
- What your tool does not find
- The second number
- Advance-Trust Unease
- What the reviewer does not see
- Floor Crack
- Trust Loop
- Final-Grip Calm
- Smoothness Watch
- Second Eye
- Path Trust
- Responsibility Pendulum
Part 1What your tool does not find
Two classes of checking program, one task: they agree on arithmetic. On what is absent, they diverge.
At the three spots the inexpensive programs found hardest, they caught the error in 4 out of 15 opportunities. The strong programs, at those same three spots, reached 10 out of 12.
Missing information — conditions that do not appear in the checked output at all — is the kind of error on which checking programs differ most: stronger programs find most of it, weaker ones regularly leave it behind. The pattern in the weaker programs has a name: counter-figure blindness.
Part 2The second number
Time saved is half the measurement.
When firms roll out AI, they report one number: hours saved. A second number belongs next to it: do staff still find known errors reliably?
A comparison from competitive sport: performance is measured together with load. A plan that raises results while quietly wearing the athlete down is called overtraining. Applied to the office: measuring only hours saved misses the second half.
The second number is the catch rate. An announced test task contains a known error, say an invented source. How many reviewers find it? That share belongs next to hours saved in the same report. Count false alarms too: correct results that are wrongly flagged.
If the catch rate falls while hours saved rise, that is a reason to look at the review more closely. If both rise, that can be a good sign, provided tasks and review conditions are comparable. The two numbers alone do not prove a general improvement.
How to set up the catch rate in a week (business-wissen.de, in German) →
Part 3Advance-Trust Unease
Sometimes we trust an AI answer faster than we check the evidence behind it.
You know the feeling. A colleague gave three good tips, so you take the fourth as fact right away — and mid-nod a quiet thought pipes up: "Wait, do I actually know this, or do I just believe it because he said it?"
That can happen with AI too. After several correct answers, we wave the next one through before we have looked at its evidence. I already trust the answer, and I notice that I have not checked its basis yet.
I call that quiet warning Advance-Trust Unease: a trust that does not trust its own speed.
Two things it is not. It is not a sign that the trust is wrong; trust may grow faster than evidence. And it is not a sign that the machine has earned it. It is simply the moment you still know there is a gap. Whoever silences that moment instead of listening to it turns advance trust into blind flight.
Part 4What the reviewer does not see
Picture two AI systems with the same error rate. One can still be more dangerous.
The first writes its errors clumsily, the second fluently and with confidence. If reviewers miss the fluent ones more often, more of them reach the customer. Same error rate, different errors at the customer.
The system's error rate alone does not show how many errors reviewers notice. You need two numbers, always together: how many of the wrong outputs does the reviewer recognise as wrong? And how many correct ones does the reviewer wrongly classify as wrong? The first number alone is not enough to judge the review. Whoever doubts everything also finds every error.
And both numbers need comparison values for similar tasks. Otherwise nobody knows whether, say, 80 percent is good or bad.
In a public call for input on AI evaluation procedures I submitted a proposal: evaluation reports should show how many known errors the reviewers detected and how many correct outputs they wrongly flagged. If you read or write such reports: ask for the second number.
Part 5Floor Crack
One wrong number, found by chance. Suddenly you doubt the whole document.
Imagine this: a colleague spots a false figure in an AI summary you had already signed off. What shakes you is not just the error. It is the thought that the error was found only by chance — and that nobody knows how many others slipped through.
The lexicon entry describes it like this: Floor Crack — the quiet feeling that all stored answers suddenly seem less certain when one key statement turns out to be wrong; it is not the error itself that lingers, but knowing I found it only by chance.
The entry also draws a line: a floor crack does not mean the whole foundation was worthless. A crack is a measuring point, not a collapse.
My suggestion: at every sign-off, note what was checked, what stays open, and who checks next.
Part 6Trust Loop
Because the AI seems honest, you trust it more. Because you trust it more, you check less.
Imagine this: after several usable drafts you only skim the next one and sign it off out of habit. You think: the previous drafts were usable.
The lexicon entry describes the Trust Loop like this: Because the AI seems honest you trust it more; because you trust it more you check less — until in the end you believe everything blindly, exactly when checking would matter most.
The entry calls for the machine to recognise and point out reduced questioning.
My suggestion: ask two questions before each sign-off. What is in the source material that is absent from the output? Which passages flagged as errors turn out to be correct on review? The same two questions, even on the twentieth good draft.
Part 7Final-Grip Calm
If nothing goes out without your click, you feel safe. But when did you last really check before clicking?
Imagine this: the artificial intelligence (AI) may only send once you approve. You skim the draft, you click, it goes. You could stop it from being sent — just knowing this makes handing work over feel easier.
The lexicon entry describes it like this: The reassuring certainty that in the end YOU still press the last button and could stop everything — just knowing this makes it easy to hand a lot over to the AI, even if you almost never need that veto.
That is Final-Grip Calm. The entry draws a line: feeling calm does not mean you are still in control of everything.
My suggestion: for one week, count how often your last click follows a fresh check and how often it does not. Pay particular attention to the approvals given without checking again.
Part 8Smoothness Watch
The result looks perfect. No rough edge, no open question. That is exactly when a quiet suspicion can stir.
Imagine this: a draft from an artificial intelligence (AI) sounds so flawless that you suspect an error — and you check whether the suspicion is even justified.
The lexicon entry describes it like this: The suspicion that wakes up when a result comes back suspiciously smooth and without any roughness at all — the feeling "that was too easy", louder than with a result that shows visible uncertainty.
That is Smoothness Watch. And the entry draws a line: it does not mean that smooth results are automatically wrong.
My suggestion: when a result feels too smooth, start by comparing one passage with the source material. If the passage matches the source, only that passage has been checked. If it does not match, investigate the difference before passing the text on.
Part 9Second Eye
I find it especially helpful when a checking program finds exactly the thing I was not looking at.
Imagine this: while proofreading a draft from an artificial intelligence (AI), a second program points to a missing condition that you had overlooked yourself. This finding does not feel like a stranger's remark. It feels like your own look, arriving late.
The lexicon entry describes it like this: The AI looks exactly where you're not looking, and what it finds doesn't feel foreign — it feels like your own gaze that just took a detour.
That is the Second Eye. The entry draws a line: the second check does not replace your own review.
My suggestion: give the second program one job only: what is in the source material that is absent from the AI draft? The measurement in part 1 shows how differently checking programs find exactly such information. For me, that is a reason to ask specifically about missing information during a second review.
Part 10Path Trust
An explanation accompanying the recommendation is reassuring. That is exactly its risk.
Imagine this: the recommendation from an artificial intelligence (AI) comes with three sentences of explanation. You already feel safe before you have checked a single one of the steps.
The lexicon entry describes it like this: You trust the AI more when it shows you HOW it got there — a bare result with no reasoning stays foreign to you.
That is Path Trust: trust that grows not from the result but from the visible path toward it.
My suggestion: check one step in the explanation against the source material. An explanation can read fluently and still skip a step. The question to ask: is a necessary step missing?
Part 11Responsibility Pendulum
Once a wrong answer has gone out, the question of blame starts to swing — and it does not stop on its own.
Imagine this: a wrong answer from an artificial intelligence (AI) was signed off and reached the customer. Your attention swings back and forth: to the AI that wrote the text; to the person or program that assigned the task; and to you, who approved it and are responsible for the whole. It never settles.
The lexicon entry describes it like this: After a mistake your gaze swings back and forth: was the executing part at fault, the distributor, or you as the one responsible for the whole? — and it never comes to rest at any one spot.
That is the Responsibility Pendulum.
My suggestion: at each sign-off, record what was checked, what stayed open and who checks next. The note records what had been checked at that point.