Blog

I caught it on review one. I wouldn't have on thirty.

When organizations mandate human review of every AI output, the volume itself becomes the failure mode.

Samip Shah Aug 22, 2026 7 min read
ai-oversight ai-review-fatigue automation-complacency cognitive-overload human-in-the-loop workplace-ai
A woman in glasses gazes at a laptop screen with focused, concentrated attention, hand raised to her chin
Photo by Shixart1985 via Wikimedia Commons (CC BY 2.0).

I sat with the R program Haiku had built for me, read it line by line, and knew in ten minutes it was not going to fly.

In November 2025, I had a customer requirement to perform risk analysis on the R packages installed in their environment. I used Claude Haiku. I asked it to create a detailed plan first, which it did. Then I asked it to take the necessary actions and deliver the final product. Haiku produced an R program. I sat with it. Read it line by line. Ten minutes in, I knew the output was not convincing. It felt half-baked. I was not happy with the deliverable. The program stayed with me. It did not ship.

What made that catch possible was one condition I rarely have at nine in the morning with a full review queue: I had the time and the attention to read it carefully. No pressure from what was coming next. No AI review fatigue narrowing the eye. That is what active human oversight looks like when it works, and in a workflow that mandates sign-off on every AI output, it is the rarest resource in the room.

The question I kept coming back to was not whether the output was good, but whether I would have caught it on review number thirty.

The incident worked because I was doing AI output review with full attention. That sounds like a baseline, not a special circumstance. But anyone who works inside a team that has mandated human sign-off on every AI-generated deliverable knows what the queue looks like by mid-morning.

The outputs arrive polished, well-formatted, and confident. Each one looks much like the last. The attention you bring to item five is not the same attention you brought to item one. The variable the incident exposed was not the quality of what Haiku produced. The variable was my attention. The question I could not shake afterward was whether I would have caught the same half-baked R program if it was the thirtieth item in the queue instead of the first.

Cruise control on a long highway does the same thing to a driver that a review queue does to a reviewer.

Imagine a driver who has been on a straight highway for an hour at 100 km/h. The car is managing it. Every kilometer looks like the last. The driver is still behind the wheel, eyes open, technically present. But the active steering is gone. The road is not asking for input. This is not carelessness. It is how human attention responds to low-variation, high-volume repetition.

A review queue full of AI output creates the same condition. Each item arrives formatted, fluent, and professionally worded. None of them announce themselves as wrong. The reviewer is still there, cursor blinking. Automation complacency settles in at the same pace it does on a flat highway, for the same reason: the signal-to-noise ratio of the incoming items is too low to keep the eye sharp enough to catch the one that diverges.

That is the automation paradox. The better the AI gets at producing outputs that look right, the harder it becomes to catch the one that is not.

Four findings say reviewer attention under volume is a structural design problem, not a motivation problem.

This is not a discipline problem that training or reminders can fix. A 2021 study found that evaluating every explanation requires substantial cognitive effort, which humans are averse to. When you apply that to a review queue, the cognitive overload AI workflows create is not a character flaw in the reviewer. It is a documented response to a real cost. Telling reviewers to concentrate harder does not reduce that cost; it only redirects the blame.

The mechanism that turns cognitive cost into passive acceptance is named directly in a 2025 article: when an automated system produces confident recommendations, humans tend to over-accept them and reduce independent checking, especially under time pressure or cognitive load. Volume and time pressure are the precise conditions of every mandated-review queue. Automation complacency under those conditions is not a sign of weak reviewers. It is a sign of normal ones placed inside a design that produces the failure mode it was built to prevent.

What happens to review quality as the queue grows? A second 2025 research article is direct: intervention quality degrades under high workload and limited transparency. More reviews in a session does not produce better reviews. It produces the opposite.

At the organizational level, the rubber stamp is already the default design. The same article notes: In many deployed AI systems, oversight is implemented as a single review step or a nominal approval interface, offering limited visibility into system behaviour and limited intervention authority. The policy says "human review required." The implementation is a checkbox. Those are not the same thing.

The objection that matters is that human review is the regulation, not a choice among options.

The EU AI Act mandates human oversight of AI for high-risk systems. That is a real constraint, not a conservative reflex. The argument here is not against that requirement. Unchecked model output in high-stakes domains is worse than even the weakest form of review.

The counterpoint is about what counts as meeting the requirement. Does reviewing AI output prevent errors when the reviewer has processed thirty items in ninety minutes? A reviewer who approves every output because each one looks polished has not provided oversight. They have provided a signature. A signature is not a judgment.

The regulation was written assuming the human would think. It was not written assuming volume would make thinking optional. Human-in-the-loop fatigue is the gap between what the policy text requires and what a reviewer can deliver at item twenty-six. Organizations that treat the presence of a reviewer as equivalent to the presence of a review have confused two different things.

The phone deal Gemini invented sounded authentic enough that I almost acted on it before going back to check.

I was planning to buy a new phone. I gave Gemini my requirements and intended use and asked it to find the best deal. The AI confidently recommended a phone model that never existed, complete with model numbers, specifications, and price comparisons that sounded fully authentic. I almost acted on it. Only after checking the original source did I discover the recommendation was completely fabricated. I cite this as my personal example of what an AI hallucination is: not deception, but a confident prediction that diverged from reality.

The parallel to the reviewer problem is exact. The output that fails a review does not announce itself as flawed. It looks like every other output in the queue. I caught the fabrication only because I made a deliberate decision to go back and check. That was an active step, not a passive glance. The check was a choice. A passive review of that recommendation would have produced the same outcome: a thumbs-up and a wrong decision.

Misreading this as "review less" is exactly the wrong conclusion, and it is also the easiest one to reach.

The easiest misapplication is reading this as an argument for skipping AI output review entirely. It is not. Shipping unchecked output under the label of efficiency is the failure mode this post is warning against, not a fix for it.

The three structural changes that work are what I use in my own workflow now. Active sampling means reviewing a deliberate subset at full attention rather than every output at degraded attention - ten items reviewed carefully deliver more real oversight than thirty reviewed on autopilot. Threshold-based automated testing lets automated checks handle high-volume, rule-bound verification so human attention stays reserved for judgment calls the rules cannot cover. Isolated validation steps separate the review moment from the production queue so the reviewer is not evaluating while under throughput pressure. Human oversight of AI only works when the oversight is not the same motion as the production step.

If the review queue never empties, the question worth asking is whether it is still oversight or just a signature on a moving belt.

The next time the review queue is open in front of you, the question is not whether the policy requires sign-off. It probably does, and the queue probably never fully empties. The real question is whether item twenty-four got a judgment or just a stamp. Those are not the same thing, even when they produce the same click on the same checkbox. That difference is what makes oversight worth having.

← All posts