My developer friend sent an entire codebase and the model reviewed the wrong product.
I was having coffee with an old friend from my development days when he told me about an experiment that had turned him off LLMs entirely. He had shared an entire multi-product code repository with a model and asked it to do a code review. The model came back with comments. His colleague read them, stopped, and pointed out that the review described a different product in the same repository, not the one my friend was working on.
The model had every file it could have wanted. The entire codebase was right there. And that was precisely the problem. This is AI context overload in its plainest form: the model does not run out of context, it loses the thread inside it.
If the model lost the thread in his session, it can lose it in yours.
My friend's experience is not unusual. It is what happens any time a professional copies an entire email chain into a prompt, pastes every section of a long report, or uploads a full dataset and asks what the key issue is. The model receives all of it. It responds to all of it. But it could not really figure out its place and the scope as to where it should be looking for.
That description has stayed with me. Not because it is dramatic, but because it is exact. The model read the repository. It produced coherent, technically fluent comments. They landed on the wrong product.
It's a double-edged sword. The same long-context capability that feels powerful, the sense that you can hand the model everything and let it sort through the pile, is the one that makes its attention scatter. More input is not a cleaner signal.
More information is not a cleaner signal; it is a bigger haystack.
Picture a forensic accountant handed a complete twenty-year company archive and asked to find the discrepancy in last quarter's filings. The answer is in there somewhere. But the volume is not the aid, it is the obstacle. That accountant needs the right drawer opened to the right year, not the whole warehouse.
An LLM processing a wide context works on a similar principle. The model reads the input in sequence. The further a relevant detail sits from the edges of what you sent, the harder it becomes for the model to weight it correctly. Hand over a hundred files when the answer lives in three, and you have made your own question harder to answer.
What happens when you feed an LLM too much information is not a crash. It is a quiet drift.
The research shows that a larger context window does not mean the model reads all of it.
This is the mechanism behind AI context overload. The numbers are stark. A 2024 NeurIPS benchmark tested popular models against tasks requiring reasoning across very long documents. The finding: popular LLMs effectively utilize only 10-20% of the context. Researchers call this context rot, the decline in output quality as the input grows without a corresponding gain in precision. At most a fifth of what you included was doing real work.
The mechanism was studied in a paper from Stanford-affiliated researchers that gave the problem its technical name. Their finding: performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models. That is the lost-in-the-middle problem, where information buried in the centre of a long prompt is processed less reliably than content at the start or end.
A separate study published at ACL 2024 padded the same reasoning tasks with increasing amounts of irrelevant text and measured the effect on performance. The result: notable degradation in LLMs' reasoning performance at much shorter input lengths than their technical maximum. The context window was nowhere near full. The performance had already dropped.
The fix the research points toward is not a smarter model. It is a more disciplined prompt. Anthropic's guidance on prompt engineering for business performance states: sometimes Claude performs better on complex tasks if you break the task down into multiple prompts corresponding to each step. Fewer things per prompt means each thing receives real attention.
Bigger context windows are real, and they do not solve the attention problem.
The obvious objection is that context windows keep growing. 100K tokens, 200K, 1M. If the window is that large, does any of this still hold?
The concession is fair. Larger windows do help in specific workflows. Legal discovery where completeness is the requirement. Full-document translation where no sentence can be dropped. When the task genuinely needs the whole document, a larger context window is the right tool.
But the research finding is about how the model allocates attention once the input is inside the window, not about where the ceiling sits. The model does not distribute attention evenly. It weights the edges and under-weights the middle. A larger window does not change that mechanism, it just moves the ceiling on a pattern that was already there before models could handle a million tokens.
Most readers arrive at this post with one question: does more context always help? The answer is no. A larger haystack does not make the needle easier to find.
So he narrowed it. One product. Same model. The second answer told him.
When my friend described the code-review failure over that coffee, I suggested he try one thing before writing off the model. Reduce the input to the relevant product only. Strip the rest of the repository out. Same model, same task, narrower context.
He said he had not thought of that.
A few days later he called back.
Same model. Same task. Narrower context. The second answer told him exactly which functions needed attention and why, in the product he was working on, not the one sitting three files over.
What changed was not the model's capability. What changed was the scope of what went in. That is scope reduction: sending what is relevant and leaving out what is not. The review comments landed where they needed to. He said it had worked well, and he would not forget the lesson.
The habit breaks the moment you strip context the model actually needed.
There is a failure mode on the other side of this. A reader who takes this post as permission to strip everything will find it. The model returns an answer that misses something, and the instinct is to blame the model again.
The habit is not "send less." It is "send what is relevant and leave out what is not."
Before I send a prompt now, I run one question: what is in here that the model does not need to know? Scope reduction at its most practical takes thirty seconds. That question is also a form of prompt engineering. Send the email threads that contain the actual dispute, not the six-month chain. Send the contract clauses in question, not the whole document. The model's attention is not unlimited. Your job is to point it at the right thing.
Before your next session, ask yourself what you are sending that the model does not need.
The next time an AI answer seems off, before you retry with a better model or a longer prompt, try a shorter input instead. Ask yourself what you included that the model did not need to know. Take it out. Run it again.
That one question, asked before every session, is worth more than any model upgrade.