Your Context Window Is Not Memory
Every frontier model now advertises a million tokens. Published results show the usable fraction is far smaller, and it moves with the shape of your question. Why bigger windows and higher-recall retrieval both aim the wrong way, and how to measure your own ceiling.

Ask a model to copy a list of words back to you. No reasoning, no retrieval, no ambiguity — just reproduce what you were given. Now make the list longer.
It gets worse at it.
That result comes from Chroma's July 2025 technical report, which ran 18 frontier models — Claude 4, GPT-4.1, Gemini 2.5, Qwen3 and others — through a battery of deliberately trivial tasks at increasing input lengths. The repeated-words task has no difficulty in it at all. Models still degraded: under-generating, misplacing words, and inventing words that were never in the input.
Hold that next to the thing every model provider put on its spec sheet in 2026: a one-million-token context window.
Both are true, and the gap between them is where most agent systems quietly fail. A context window is a capacity number, like the size of a disk. Effective context — how much a model can actually use without getting worse — is something else entirely, it is much smaller, and it is not a single number. It moves with the shape of your task.
This matters more than it used to, because the two things engineers reflexively do when an AI system underperforms both push in the wrong direction. Give it more context. Retrieve more aggressively. Each of those is a reasonable instinct, and each one, done naively, makes the measured problem worse.
The number on the box
Start with what the benchmarks actually say, because the headline figures and the useful figures are not the same figures.
The test everyone quotes is needle-in-a-haystack: hide one sentence in a large body of text and ask the model to find it. Frontier models are now excellent at it. On one 2026 compilation of published results, Gemini 3 Deep Think retrieves a single needle at one million tokens with 99% accuracy.
Now change one variable — ask for eight needles instead of one — and models on the same context length, all advertising the same window, fall apart unevenly.
A roughly 60 to 70 percent gap between advertised and effective context is now the ordinary finding across RULER, MRCR and needle-style suites. But treating that as a single discount rate — "assume you really have 300K" — is still wrong, because the discount is not a property of the model alone. It is a property of the model and the question you are asking of it.
Three axes, not one dial
Across the research, degradation tracks three independent variables. They are rarely separated, and separating them is what makes the problem tractable.
Take them one at a time, because each has a consequence that contradicts standard practice.
Length hurts on its own
The comfortable assumption is that long-context failure is really a retrieval failure: the model did badly because the right information was buried. Fix retrieval and the problem goes away.
In October 2025, Du and colleagues tested that directly in Context Length Alone Hurts LLM Performance Despite Perfect Retrieval. They removed every confound they could think of. The relevant evidence was perfectly retrievable. Irrelevant tokens were masked so the model attended only to what mattered. The evidence sat immediately before the question. Across five models on maths, QA and coding, performance still fell by 13.9% to 85% as input length grew.
There is no retrieval architecture that removes that. The tokens being present is itself the cost.
Multiplicity collapses faster than length
Most published long-context marketing rests on finding one thing. Almost no real question is shaped that way. "Why did margin fall in the north region last quarter" needs the margin figures, the regional split, the quarter boundaries, the one-off costs, and whatever changed in the pricing policy — five or six retrievals that then have to be reconciled.
The eight-needle numbers above are the closest published proxy, and the drop from one needle to eight is far steeper than the drop from 100K tokens to 1M. If you are choosing what to optimise, reducing the number of things the model must hold simultaneously beats reducing the number of tokens.
Indirection is the axis nobody measures
This is the most under-appreciated of the three. NoLiMa, published at ICML 2025 by Modarressi and colleagues, makes one change to the standard test: it removes literal word overlap between the question and the answer, so the model has to reason associatively rather than pattern-match on shared vocabulary.
Of 13 models that all advertise at least 128K context, 11 fell below half their short-context baseline at just 32,000 tokens. GPT-4o, one of the best performers, went from 99.3% on short inputs to 69.7% at 32K and 56% at 128K.
Read that against the needle-in-a-haystack results and the picture resolves. Those tests are easy partly because the needle usually shares wording with the question. Strip that away — which is what happens the moment a user asks a question in their own words about a document written in someone else's — and the usable window shrinks by an order of magnitude.
The part that should change how you build
Everything so far is a measurement problem. Here is where it becomes an architecture problem, and where the standard playbook turns out to be working against itself.
Similarity search manufactures the worst possible context
Chroma's distractor experiments found two things that fit together badly. First, a single distractor — content that is topically related but does not answer the question — measurably reduced accuracy against a needle-only baseline, with four distractors degrading it further. Second, performance degraded faster with length when the semantic similarity between the question and the answer was lower.
Now consider what a vector database does. It embeds the query, finds the nearest neighbours, and returns the top k. By construction, results two through ten are the chunks that look most like the answer without being it. That is the textbook definition of a distractor.
Retrieval-augmented generation, in its most common form, is a machine for finding one correct chunk and surrounding it with the most convincing wrong ones available in the corpus. Raising k to improve recall increases the distractor count on every query. The lever everyone reaches for when a RAG system misses an answer makes the context measurably harder for the model to use.
This does not mean retrieval is wrong. It means precision matters more than recall in a way most pipelines do not reflect, and that reranking is not a nice-to-have — it is the component that converts a distractor generator into a context builder.
Coherent prose is harder than shuffled text
The strangest result in the Chroma report is also the most immediately actionable. They tested the same needles in two kinds of haystack: one a coherent, logically flowing document, the other the same content shuffled so local coherence was destroyed. The shuffled version consistently performed better.
The authors are careful here, and so should we be: they state plainly that they do not explain the mechanism and point to interpretability research as future work. But the engineering implication stands on its own. If you are assembling context, you are not writing a document for a human reader. Narrative flow, smooth transitions and topical grouping are not helping and may be hurting. Structured, labelled, discrete facts are a better container than well-written prose.
Most context-assembly code does the opposite. It concatenates retrieved passages into something that reads like a report, because that feels like good engineering hygiene.
Position still matters, seven years on
Liu and colleagues established the U-shape in Lost in the Middle (TACL, 2023): accuracy is highest when the relevant information sits at the beginning or end of the input and degrades significantly when the model has to reach into the middle. Three years and several model generations later, the effect has softened but not disappeared.
The practical form is boring and widely ignored: put the instruction last. A system prompt at the top and a question at the bottom places both at the strong ends of the curve. Burying the actual task in the middle of a large retrieved payload puts it exactly where attention is weakest.
Agents make this the default trajectory
A single long-context call is a choice. An agent loop is a context-growth machine that runs whether you chose it or not.
Every step appends: the model's reasoning, the tool call, the tool's response, the error, the retry. A loop that starts at 4,000 tokens and adds 1,500 per step is at 34,000 by step twenty — through the knee of the NoLiMa curve — without anyone deciding anything. The agent does not get a warning. It just gets worse, and the symptom is not an error message but a plan that drifts, an instruction from step three that stops being honoured, and a tool call with an argument that was correct eleven steps ago.
In February 2026, Zeng, Huang and He built LOCA-bench specifically to measure this: agents under controlled, extreme context growth with task semantics held fixed. Their finding is the one that matters for anyone shipping agents, and it cuts both ways. Agent performance degrades as environment state grows — and context management techniques substantially improve success rates.
That second half is important enough to say loudly, because it is the strongest argument against reading this article as fatalism.
The counterargument, taken seriously
Three honest objections, and what survives them.
"The models are fixing it." Partly true, and the trend is real. Single-needle retrieval at a million tokens is close to solved. But the gap is between nominal and effective context, and both have been rising together. The eight-needle spread above — 76% to 24.5% at the same advertised window — is 2026 data, not 2023 data. The discount changed size; it did not close.
"This is a benchmark artefact." The fairest version of this objection is that synthetic needle tasks do not resemble real work. It is a reasonable concern, and it is why the Du result matters most of the three: maths, QA and coding tasks, with every confound stripped out, still degrading on length alone. The effect is visible across synthetic retrieval, associative reasoning, agent trajectories and ordinary task benchmarks. That is a wide enough net.
"Good engineering removes the problem." This is the one with the most evidence behind it, and it is the reason the article does not end in gloom. Compaction, structured note-taking and sub-agent isolation all work. Anthropic reports context editing cutting token consumption by 84% in a 100-turn web search evaluation while enabling workflows that would otherwise have failed outright. Du and colleagues found that simply having the model recite the retrieved evidence before answering recovered up to 4% on RULER.
But notice what those techniques have in common. None of them makes the model better at long context. Every one of them is a way of keeping the model out of long context. That distinction is the whole argument: the ceiling is a model property, and engineering is how you stay underneath it rather than how you raise it.
It is worth being precise about the uncertainty here. Chroma's authors explicitly decline to explain the mechanism. Attention dilution is a plausible story and not a demonstrated one. What is established is the behaviour, repeatedly, across labs and model families. For engineering purposes the behaviour is sufficient; for predicting how fast it will improve, it is not.
A context budget
Treat context as a scarce resource with a quality cost, not a buffer with a size limit. Four rules, in the order they pay off.
| Rule | What it means in code | Why, specifically |
|---|---|---|
| Measure your own ceiling | Build 30–50 real questions with known answers. Run them at 4K, 16K, 64K and 128K of realistic padding. Plot accuracy. | Your task sits somewhere on three axes. No published benchmark describes it. The curve takes an afternoon and is the only number that applies to you. |
| Spend precision, not recall | Lower k. Rerank hard. Drop anything below a relevance floor rather than filling the quota. | Every extra chunk is a probable distractor, and distractors cost more than filler. |
| Structure over prose | Labelled fields, explicit delimiters, discrete facts. Stop concatenating passages into a readable narrative. | Local coherence measurably hurt in controlled tests. You are building an input, not a document. |
| Bound the loop | Compact on a token threshold. Push tool output to a scratchpad and re-read on demand. Isolate sub-tasks in their own windows. | Agents walk into the degraded regime by default. Something has to walk them back out. |
The measurement step is the one teams skip, and it is the one that makes the rest concrete. You are not looking for a pass or fail. You are looking for the length at which your accuracy starts bending, because that number becomes a real budget: a compaction threshold, a retrieval ceiling, a point at which a task must be split rather than continued.
What a measured ceiling changes
Once you have the curve, several arguments that are normally settled by intuition become arithmetic.
- Whether to split a task. If accuracy bends at 30K and a job reliably needs 80K of state, that is not a prompt to tune. It is two jobs with a checkpoint between them.
- How much to retrieve. If accuracy at k=20 is below accuracy at k=5 on your own questions — and it often is — the recall-maximising instinct has been costing you answers.
- Which model to use. The eight-needle spread says model choice matters far more for multi-fact work than for simple lookup. Benchmark the shape of question you actually ask.
- When the context window stops being the constraint. Below your bend point, buying a larger window buys nothing. That is usually a surprise, and usually a saving.
Where this is heading
Two things are likely to stay true for a while.
The first is that nominal and effective context will keep rising together without converging. Both numbers have grown by more than an order of magnitude in three years and the discount between them has stayed stubbornly large. A model that could use its full advertised window on multi-fact associative questions would be a genuinely different kind of system, and nothing in the current results suggests one is imminent.
The second is that context management is becoming a durable layer of the stack rather than a workaround. Compaction, eviction policies, scratchpad memory and sub-agent isolation are being studied, benchmarked and standardised — the same path configuration management and caching took. Work that looks like plumbing today tends to look like architecture in hindsight.
The practical consequence for anyone building now: the question to ask of an AI system is not "how big is the context window." It is "at what input length does this system stop being right, and what happens when it gets there." The first question has a number on a spec sheet. The second has an answer only you can measure, and it is the one that determines whether the thing works.
Frequently asked
What is context rot?
The measurable decline in output quality as input length grows, well before the context window is full. It has been reproduced across 18 frontier models on tasks as simple as copying a word list, and it is distinct from — though often confused with — the lost-in-the-middle position effect, where accuracy depends on where the relevant information sits rather than how much surrounds it.
What is effective context, and how is it different from the context window?
The context window is the maximum input the API will accept. Effective context is how much the model can use before accuracy starts falling, and it is typically a small fraction of the window. Benchmark suites commonly show a 60–70% gap, but it is not a fixed discount: it moves with input length, the number of facts that must be combined, and how indirectly the answer is worded relative to the question.
Does a larger context window remove the need for RAG?
No, and in many cases it inverts the question. Putting a whole corpus in the window places you at the far end of the length axis, where even perfectly retrievable information degrades. Retrieval remains valuable precisely because it keeps the input short — but only if it is tuned for precision. A high-recall pipeline floods the context with near-miss chunks, which measurably cost more accuracy than neutral filler.
Why do my agents get worse the longer they run?
Because an agent loop appends on every step, so context grows whether or not anyone intended it. A loop adding 1,500 tokens per step crosses 30K before step twenty. The symptoms are not errors but drift: earlier instructions quietly stop being honoured, and tool arguments reflect a state from several steps ago. Compaction, external scratchpads and isolated sub-agent windows all address it; a bigger window does not.
How do I find my system's effective context limit?
Assemble 30 to 50 real questions with known-correct answers, then run them at several input lengths using realistic padding from your own corpus rather than synthetic filler. Plot accuracy against length and find where the curve bends. That bend is your budget — the compaction threshold, the retrieval ceiling, and the point at which a task should be split instead of continued.
Is this getting better with newer models?
On easy retrieval, substantially. Single-needle lookup at a million tokens is close to solved. On multi-fact and associative work the picture is uneven: published eight-needle results at the same context length range from 76% to 24.5% across current frontier models. Nominal and effective context have risen together without the gap closing, so design against measurement rather than against the trend.
The short version
A context window is a capacity limit, not a working memory. What a model can actually use degrades along three independent axes — how long the input is, how many facts must be combined, and how indirectly the answer is expressed — and those axes compound in exactly the combination that describes ordinary business questions.
The two standard responses both aim the wrong way. A bigger window moves you further along the axis that hurts even under perfect retrieval. Higher-recall search fills the context with the near-misses that cost more than noise does. What works is subtraction: retrieve less and more precisely, structure rather than narrate, keep the instruction at the end, and bound the loop before it wanders into the part of the curve where the model is quietly wrong.
None of that makes a model better at long context. It keeps the model out of it, which is a different and more achievable goal.
If you are building something where the answer depends on combining several facts out of a large corpus, the effective-context curve for your own questions is worth an afternoon before it is worth a model upgrade. Tell us what you are building and we will work through where your ceiling actually sits. Our writing on what agent loops cost as they grow and how to know when AI output is right covers the two problems this one tends to arrive with.
Frequently asked questions
What is context rot?
The measurable decline in output quality as input length grows, well before the context window is full. It has been reproduced across 18 frontier models on tasks as simple as copying a word list, and it is distinct from — though often confused with — the lost-in-the-middle position effect, where accuracy depends on where the relevant information sits rather than how much surrounds it.
What is effective context, and how is it different from the context window?
The context window is the maximum input the API will accept. Effective context is how much the model can use before accuracy starts falling, and it is typically a small fraction of the window. Benchmark suites commonly show a 60–70% gap, but it is not a fixed discount: it moves with input length, the number of facts that must be combined, and how indirectly the answer is worded relative to the question.
Does a larger context window remove the need for RAG?
No, and in many cases it inverts the question. Putting a whole corpus in the window places you at the far end of the length axis, where even perfectly retrievable information degrades. Retrieval remains valuable precisely because it keeps the input short — but only if it is tuned for precision. A high-recall pipeline floods the context with near-miss chunks, which measurably cost more accuracy than neutral filler.
Why do my agents get worse the longer they run?
Because an agent loop appends on every step, so context grows whether or not anyone intended it. A loop adding 1,500 tokens per step crosses 30K before step twenty. The symptoms are not errors but drift: earlier instructions quietly stop being honoured, and tool arguments reflect a state from several steps ago. Compaction, external scratchpads and isolated sub-agent windows all address it; a bigger window does not.
How do I find my system's effective context limit?
Assemble 30 to 50 real questions with known-correct answers, then run them at several input lengths using realistic padding from your own corpus rather than synthetic filler. Plot accuracy against length and find where the curve bends. That bend is your budget — the compaction threshold, the retrieval ceiling, and the point at which a task should be split instead of continued.
Is this getting better with newer models?
On easy retrieval, substantially. Single-needle lookup at a million tokens is close to solved. On multi-fact and associative work the picture is uneven: published eight-needle results at the same context length range from 76% to 24.5% across current frontier models. Nominal and effective context have risen together without the gap closing, so design against measurement rather than against the trend.
Related Posts
AI Agent Cost Per Task: Why Your Bill Is a Reliability Problem
Token cost grows with the square of the steps, the chance of needing another attempt grows exponentially with them, and the human who cleans up afterwards costs more than both. Here is the model we use to decide whether an agent is viable, with the arithmetic shown.
Your Model Is Not a Function
A thousand identical requests at temperature zero produced eighty different answers. The usual explanation about floating point is only half of it — the other half is that your output depends on how many strangers were using the server at the same moment.
Your Agent Has a Permissions Problem, Not a Prompt Problem
Prompt injection is a confused deputy attack, described in 1988 and unfixable at the model layer for the same reason it was unfixable then. The containment architecture already exists, most of it is standardised, and almost nobody ships it.
