The short version
A document that fits in the context window is not a document the model reads evenly. Research since 2023 has measured accuracy falling as the input grows, and falling furthest for information in the middle. A request for "any issues" across a whole bundle gets a fluent list, and the list does not say what it missed.
What it is
The request is simple: put the entire trial bundle, the lease with its schedules or the data room export into one prompt, and ask for the issues. The answer comes back organised and confident. What it leaves out does not announce itself, because nothing tells the reader that a clause was never weighed.
One vendor says so itself. Anthropic's documentation on context windows, checked on 1 October 2026, says that "more context isn't automatically better. As token count grows, accuracy and recall degrade, a phenomenon known as context rot."
A sighting
The measured version comes from two studies.
Lost in the Middle (Nelson F. Liu and others; arXiv, July 2023; published in Transactions of the Association for Computational Linguistics in 2024) gave models a question and a set of documents, one of which held the answer, and moved that document through the set. The authors found that "performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts, even for explicitly long-context models." In one setting, with the answer in the middle of twenty documents, GPT-3.5-Turbo did worse than it did with no documents at all (56.1% closed-book). The models were those of 2023, including GPT-3.5-Turbo and Claude 1.3, and twenty documents came to about 4,000 tokens, far smaller than a modern bundle.
Context Rot (Chroma, 14 July 2025, by Kelly Hong, Anton Troynikov and Jeff Huber) tested 18 models, including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3. Its conclusion: "models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows." The tasks were controlled tests, such as finding a sentence planted in long text, answering from a long chat history and copying repeated words, not legal documents.
Neither study tested a bundle. The example below is made up to show the pattern. A 300-page lease with its schedules goes in whole, with the instruction "List any issues for the tenant." The answer deals fully with the rent review and repair covenants near the front and with the break clause in the last schedule. It says nothing about an alienation restriction on page 140. Nothing in the answer suggests that page 140 was skipped.
What to say back
"Did the tool read the whole bundle evenly, or give a confident answer about the parts it weighted? Ask it about something known to be in the middle before relying on what it says about the rest."
Origin
- 6 July 2023. Lost in the Middle is posted on arXiv; it is published in the Transactions of the Association for Computational Linguistics in 2024.
- 14 July 2025. Chroma publishes Context Rot, testing 18 models.
- Now. Context windows have grown: on 1 October 2026, Anthropic's documentation listed models with a window of 1M tokens. The same page says that more context is not automatically better. The site's methodology piece, Context Architecture, suggests a working margin and ways to split a large document; the Claude Context Window guide covers one tool.