Skip to content
Back to Errors

Context rot

The matter chat kept open for weeks, which drifts from what it was told at the start. In Chroma's tests, adding tokens around an unchanged task was enough to bring performance down.

  • First recorded: July 2023
  • Where it shows up: matter chats · drafting sessions · long project threads
  • Status: measured

The short version

One conversation is kept open for a matter and used for weeks. Its first message set the ground rules: English law, acting for the tenant, these defined terms. By the end, the answers no longer follow them. Chroma's researchers kept each task fixed and changed only how many tokens surrounded it; the added length alone brought performance down.

What it is

"Just upload the whole bundle." puts a large input in at once. A long conversation builds one up a message at a time. Every question, pasted letter and draft reply is added to what the model reads on the next turn.

Anthropic's documentation on context windows, checked on 1 October 2026, describes this for requests to its API: "each user message and assistant response accumulates within the context window, and previous turns are preserved completely." The same page says that chat interfaces such as claude.ai "can also manage the context window on a rolling 'first in, first out' basis", and it describes compaction, which "automatically summarizes earlier parts of the conversation". Depending on the tool, an instruction from the first day may by the third week be a small part of a very long input, a line in a summary, or no longer there.

As the token count grows, the page says, "accuracy and recall degrade, a phenomenon known as context rot", and this "makes curating what's in context just as important as how much space is available".

A sighting

Context Rot (Kelly Hong, Anton Troynikov and Jeff Huber; Chroma, 14 July 2025) ran controlled tests on 18 models, GPT-4.1, Claude 4, Gemini 2.5 and Qwen3 among them. It begins from a common assumption, that "the model should handle the 10,000th token just as reliably as the 100th", and reports that "in practice, this assumption does not hold."

One of its tests is close to a matter chat. Built on the LongMemEval benchmark, it gave each model a conversation between a user and an assistant and asked a question about one part of it. The question came in two forms: a focused prompt of about 300 tokens holding only the relevant exchanges, and the full history of about 113,000 tokens, most of it irrelevant to the question. All the models scored markedly higher on the focused version, and "The Claude models exhibit the most pronounced gap between focused and full prompt performance."

Lost in the Middle (Nelson F. Liu and others; arXiv, July 2023; Transactions of the Association for Computational Linguistics, 2024) measured position rather than length. Its abstract reports performance "often highest when relevant information occurs at the beginning or end of the input context". An instruction in the first message is therefore not in the weakest position. What a matter chat adds, week after week, is length.

Both studies used controlled tasks, not legal work, and newer models have come out since. The example below is made up to show the pattern. A conversation opened for a commercial lease dispute begins: "English law. The client is the tenant." Six weeks and forty pasted letters later, asked to summarise the remedies for unpaid rent, it answers as though the client were the landlord, and describes a procedure that is not the procedure in England and Wales.

Why it happens

The Chroma report locates the problem in presentation as well as presence: "Whether relevant information is present in a model's context is not all that matters; what matters more is how that information is presented." A matter chat leaves its opening instruction as written and changes its setting: six weeks on, it heads a long run of letters and drafts.

The judicial guidance does not discuss long conversations. It does warn that the "view" of the law held by the currently available models "is often based heavily on US and historic law, although some do purport to be able to distinguish between that and the law of England and Wales" (Artificial Intelligence (AI): Guidance for Judicial Office Holders, 31 October 2025, section 3, part I). The inference is this page's own, not the guidance's: an answer that has stopped following an instruction about jurisdiction may fall back on the law the guidance describes.

What to say back

"How long has this conversation been running, and is the instruction about the law and the client still doing any work? Ask the same question in a new conversation that opens with those instructions, and compare the two answers."

Origin

  • 6 July 2023. Lost in the Middle is posted on arXiv; it is published in the Transactions of the Association for Computational Linguistics in 2024.
  • 14 July 2025. Chroma publishes Context Rot: 18 models, with one test on long chat histories.
  • 31 October 2025. The current judicial guidance notes that models' view of the law leans on United States and historic law (section 3, part I).
  • Now. On 1 October 2026, Anthropic's documentation listed models with a context window of 1M tokens, in which turns accumulate or, in chat interfaces, leave first in, first out, and used the name context rot for the decline. In Chroma's test, about 113,000 tokens of chat history put every model below its focused score. The site's methodology piece, Context Architecture, suggests a working margin, and the Claude Context Window guide covers one tool's limits.

Sources

  1. Hong, Troynikov and Huber, Context Rot: How Increasing Input Tokens Impacts LLM Performance (Chroma),
  2. Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni and Liang, Lost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics, vol. 12 (2024), 157–173
  3. Lost in the Middle, arXiv preprint 2307.03172,
  4. Anthropic, Context windows (checked 1 October 2026),
  5. Courts and Tribunals Judiciary, Artificial Intelligence (AI): Guidance for Judicial Office Holders,