Three keywords, a question mark, enter, and the answer told me nothing I did not already know.
In February 2024, I opened ChatGPT for the first time. I treated it the way I had been treating Google for twenty years. Three or four keywords, a question mark, enter. The answer came back fluent, generic, useless. Not wrong, just the kind of text you could find on any Wikipedia introduction to the subject. I closed the tab. For about a week, I told anyone who would listen that AI was overhyped. My exact conclusion: "AI is not as good as Google." I walked off.
That conclusion was wrong. But the failure was mine, not the tool's.
The problem was not the tool, it was which job I handed it.
Here is the question that failure raises: if the answer was fluent, why was it useless?
Because fluent is not the same as useful. What I was asking ChatGPT to do in February 2024 was synthesise something from nothing, to take a few keywords and construct an answer on my behalf. That is one kind of job. The other kind is taking something I have already built and finding every flaw in it. Those two jobs are not equivalent, and handing AI the first one is where the trouble starts.
When you ask AI to generate your first draft, you are delegating synthesis. Synthesis is the part that trains judgment. The mental work of sorting what matters from what does not, deciding which argument is stronger, choosing which detail to drop, that is not a side effect of writing. That is the work. When AI does it for you, you get a polished document and a quieter mind. The document looks done. The mind did not do the work.
Most people delegating thinking to AI today are making exactly this move. They produce a better-looking output faster, and the cost is invisible because it sits in future performance, not in today's deliverable.
Asking AI to author your first draft is like hiring a driver who has never been to your destination and calling it a shortcut.
Think about driving a route you have never driven before. Every wrong turn, every recalculation, every moment you recognise a landmark and adjust, that is how you learn the road. You will not need the GPS the third time.
Now imagine handing the keys to someone with a navigation system but no actual knowledge of your destination. The car arrives somewhere close. Next week, you need to make the same trip. You are no better off. You did not learn the route.
The AI draft that looks done is the car parked in the approximate neighbourhood. It got you there faster. But you did not learn the road, and the next trip starts from the same zero.
Research shows that when you let AI go first, your brain steps back and your judgment goes with it.
There is a name for what happens when you hand off cognitive work to a tool. Researchers call it cognitive offloading. It is not laziness. It is a predictable response to having a capable assistant.
Ethan Mollick's essay on choosing to stay human puts the mechanism plainly: By short-circuiting effort, you short-circuit learning. The learning is not in the output. The learning is the cognitive struggle that produced the output. AI removes the struggle before it can do its work.
The deeper problem is what happens once the polished output is in front of you. That surface signals "done" and, in the same motion, suppresses the skepticism that would catch what is wrong with it. In his essay on the jagged frontier, Mollick notes: People really can go on autopilot when using AI, falling asleep at the wheel and failing to notice AI mistakes.
"Falling asleep at the wheel" is apt. It is not a decision to stop checking. The authoritativeness of the output pre-empts the check. You read the document, it looks right, and you move on. AI judgment erosion does not announce itself. It happens in the margin between "that looks good" and "let me think about whether that is right."
Review fatigue is a structural outcome of delegating synthesis. The doubt that catches mistakes is part of the generation process. Receive the polished result instead, and you skip straight to approval.
If AI measurably lifts output quality, the argument for going second looks like perfectionism, not strategy.
The objection here is fair. The NBER working paper on AI at work is a real field experiment with a real number: Access to the tool increases productivity, as measured by issues resolved per hour, by 14% on average. That is not a vendor press release. The gain is real.
Concede the number. But look at what it measures. Issues resolved per hour in a single task cycle. It does not measure what the person can resolve without the tool six months from now, or whether their judgment on harder problems improved.
The 14% productivity gain does not compound. The skill attrition it enables does compound, in the other direction. AI making you worse at your job long term is not dramatic. It is what happens when a tool substitutes for the cognitive work that builds capability. You produce more today while building less for tomorrow.
The counterpoint gets the short-term result right and the long-term trajectory wrong.
Same model, same transcript, different role, and the second answer told me which paragraph I needed to fix.
In August 2024, six months after I closed that tab, I came back to ChatGPT. This time I stopped searching and started talking. I had a 40-minute stakeholder interview transcript with the head of clinical operations at a pharma company, and I was writing a requirements document for a clinical data review feature. The bland prompt, "Summarise this stakeholder interview transcript," became a structured ask: You are a consultant at a pharma IT company. Below is the transcript. Pull out the three things she clearly wants, the two things she said no to, and the one thing she contradicted herself about. Same model. Same transcript. The second answer told me which paragraph I needed to rewrite before showing the document to anyone.
The only variable was the context I supplied. Mollick's essay on two paths to prompting names the principle: push the AI to a more interesting corner of its knowledge, resulting in you getting more unique answers. The richer input got a genuinely different answer, one that surfaced a contradiction I had missed.
That is the execution model. Draft the messy core yourself first. The act of drafting is where you form the judgment about what the document needs to say. Then shift AI from generation to friction: give it your draft and ask it to find the contradictions, the gaps, the claims a skeptical reader would reject. Use AI for edge cases, not the judgment call at the centre. That is how to use AI as a sparring partner, not a ghostwriter.
The three-step model breaks the moment you use it to confirm what you already decided.
Here is the failure mode worth naming before you adopt this model as a habit.
The execution model inversion only works if the friction in step two is genuine. The failure mode is writing a weak draft, handing it to AI, accepting the polished return, and calling that a sparring-partner workflow. That is a mirror that knows how to format.
If your draft pre-answers every question AI might raise, the AI will find the flaws you planted, not the ones you missed. Genuine friction requires a draft that exposes your actual uncertainty. A draft that hides the uncertainty behind clean prose gives AI nothing real to push against.
The concrete test: if AI's feedback never surprises you, the workflow has drifted back to delegation. When AI flags something you were confident about, that is the model working. No surprise means you wrote a brief for a verdict you had already reached.
The question is not whether AI makes you faster today, but whether it makes you sharper six months from now.
Every tool you use either builds or borrows from the judgment you bring to the next task. Using AI as a generation engine borrows. Using it as a friction engine builds.
The choice is not between AI and no AI. It is between two roles you give it, and only one of those roles compounds.