Blog

I counted twenty minutes saved and called it ROI.

The verification-tax formula turns gross time saved into net hours recovered.

Samip Shah Aug 11, 2026 6 min read
ai-productivity ai-roi hidden-cost-ai time-savings verification-tax
A man typing on a MacBook Air laptop at a wooden desk, with a notebook and smartphone nearby
Photo by Alejandro Escamilla via Wikimedia Commons (CC0).

I sat there for twenty minutes typing and deleting the same line.

A customer was agitated about a delay. The delay was not from my side. It was from the customer's own IT department, but the customer was not ready to understand that. I had to write the reply. Senior, calm, no blame, move things forward. I sat there for twenty minutes. Type and remove. Type and remove. Every sentence had a second meaning. I would draft a line, read it back, hear it land badly, and delete it. So I asked AI to write it. 30 seconds. The response was as if AI could read my mind. It understood the dilemma, the tone I needed, the part I could not say. I edited two phrases and sent it.

Twenty minutes to thirty seconds. I wrote that number down and filed it as AI ROI. The saving was real. But calling it ROI without running the rest of the calculation is the same mistake most vendor slide decks make. The headline is one row of a four-row formula, and vendors sell the one row.

Twenty minutes to thirty seconds is the headline; the formula is what the vendor skips.

If that saving is real, why do so many AI tool evaluations end in a business case that cannot close?

The vendor gives you one number and omits four. The gross time saved is measured from task start to AI response. What the vendor does not count is the prompt setup time before the task reaches the model. The model execution time, which is rarely zero on a complex ask. The human audit pass, which for any output going outside your building is not optional. And the question of whether the hours you recover go back into the same queue or toward work that creates value.

Call those three middle items the verification tax. The headline minus the verification tax is net hours recovered. That is the number AI productivity savings lives or dies on.

Outsourcing a task does not eliminate the coordination overhead.

When a company outsources a function, the gross saving is the headcount reduction. The net saving is what remains after contract management, quality review, and rework cycles are subtracted. No finance team takes the gross figure and signs the contract. They want the net.

AI works the same way. The gross saving on the customer email was twenty minutes. The net saving was close to twenty minutes, because prompt setup took seconds, execution took thirty seconds, and the audit was two phrases edited before I sent it. That is a high-ROI use case, and the formula explains why.

The risk is the reverse scenario. Imagine a team that feeds a batch deviation log to AI and receives a polished root-cause summary that confidently misidentifies the source. They spend two hours verifying. The headline saving was thirty minutes. The verification tax was four times that.

The gap between a 55.8-percent headline and a 1.4-percent reality is where the hours go.

The highest controlled figure in the research literature comes from a study on GitHub Copilot: The treatment group, with access to the AI pair programmer, completed the task 55.8% faster than the control group. That is the ceiling. Pre-practiced prompts, a constrained coding task, zero domain-verification overhead. It is the number that ends up on the vendor's opening slide.

Strip away the lab conditions and move into a live operation. A field study of AI-assisted customer-service agents found that Access to AI assistance increases worker productivity, as measured by issues resolved per hour, by 15% on average, with substantial heterogeneity across workers. Fifteen percent. More than two-thirds of the ceiling is already gone. The average hides what the gap between 55.8 and 15 already shows: most of the headline saving disappears before it reaches the worker's timesheet.

Scale to the population level and the number drops further. A working paper from NBER found that respondents report time savings equivalent to 1.4 percent of total work hours. That is the real-world net after setup, execution, and review across industries and task types.

55.8 percent to 15 percent to 1.4 percent. That journey is not a failure of the technology. It is the verification tax made visible in aggregate.

Here is how to calculate the true ROI of an AI tool. Take the raw time the AI saves on a given task. Subtract prompt setup time, model execution wait, and the audit pass. The result is net hours recovered. Multiply by your loaded hourly rate. If that figure exceeds total tool cost with a 25 percent buffer for ramp time and occasional rework, the tool pays for itself. If it does not, it is a productivity theatre line item, not a return on AI investment.

If it saved you time, that is enough, and for low-stakes tasks, it probably is.

The objection is fair. For a personal task with no audit requirement, drafting a party invitation, summarising a podcast for your own notes, the verification tax is near-zero. The headline saving and the net saving are the same number. Reaching for a four-row formula before writing an invitation is not rigour. It is wasted effort.

Conceded.

The formula matters where the stakes raise the audit requirement. A deliverable going to a customer. A document touching a regulated process. A financial estimate feeding a business case. These are not low-stakes tasks, and the hidden cost of AI tools shows up precisely here. The worker who skips verification to claim the full headline saving is not measuring AI ROI. They are measuring speed while offloading risk to whoever reads the output next.

When the audit tier shrinks to near-zero, the ROI ceiling finally becomes reachable.

In October 2025, I started using an agentic AI tool, Claude. For the first time, the tool could perform the activity on my behalf based on a single instruction, rather than describing the steps. The shift gave me a working analogy: generative AI is like a help-desk employee who guides you on what needs to be done; agentic AI is like a consultant who understands the request, analyses it, and performs the activity to deliver the outcome.

With a generative tool, the output is prose or a plan, and the audit pass covers accuracy, tone, and fitness for purpose. With a well-scoped agentic action, the output is a completed task. The verification pass narrows to confirming the outcome is correct.

The formula still applies. The audit-time variable shrinks, and net hours recovered move closer to the headline figure. This is the upper bound of the AI tool cost evaluation formula for knowledge workers: design tasks where the output is checkable in a single pass, and the ceiling becomes reachable.

The worst way to use this formula is to skip the verification line.

Someone will read this framework, see that audit time is the largest cost item in their calculation, and conclude the solution is to audit less.

That is optimising the wrong variable.

The NIST AI Risk Management Framework is explicit: AI risk management efforts should prioritize the minimization of potential negative impacts, and may need to include human intervention in cases where the AI system cannot detect or correct errors. In pharma IT, that human intervention is non-negotiable. The same principle holds anywhere the output leaves your desk and someone acts on it.

The formula is a diagnostic, not a mandate to cut. When audit time is too high, the right move is to narrow the scope of the task, not to shorten the review. A cost-neutral tool that forces rushed audits does not carry neutral AI ROI. It carries a hidden negative one.

The formula is the easy part.

Running the four-row calculation takes ten minutes. The harder question it surfaces is not whether the tool pays for itself. That is arithmetic.

The question is what you do with the hours it returns. If they go back into the same queue, the return on AI investment is real but narrow. If they go to the work only you can do, the work that requires judgment, relationships, and context no model holds, that is where the multiplier lives.

Do not bank those hours. Point them at the work the formula cannot measure.

← All posts