1 million tokens: why it changes the game (without changing everything)
Anthropic expands Claude to 1 million tokens. Is context rot finally solved? A closer look at what really changes.
Key takeaways
- Context rot — declining AI performance in long conversations — is finally under control: Claude Opus 4.6 loses only 14% between 256K and 1M tokens
- At 1 million tokens, Opus 4.6 (78.3%) far outperforms GPT-5.4 (36.6%) and Gemini 3.1 Pro (25.9%)
- In practice, long conversations go off track less often. Context rot is managed, not eliminated
AI-generated summary

You may have seen the news: Anthropic has expanded Claude's context window to 1 million tokens. One million. Five times the previous size. But the real story is not the size of the thing — yes, I know, everyone says that. It is that, for the first time, an AI can actually USE all that context without losing the plot.
Context rot: what is this mess?
For anyone just joining us: context rot is THE problem that has been holding AI models back for a while. Basically, the more information you give an AI, the less capable it becomes. Ironic, right?
Imagine talking to someone for three hours straight. Eventually, they start forgetting what you said at the beginning, confusing details and mixing things up. AI models did exactly that, only worse. Beyond 100,000 tokens — roughly 75,000 words — performance collapsed. Vendors proudly announced windows of 200K, 500K or 1 million tokens, but in practice, it was fool's gold.
The eight-needle test
To measure whether an AI forgets, researchers use a clever benchmark called the eight-needle test: eight needles in a haystack.
The idea? Give the AI lots of similar tasks, such as writing eight poems about dogs, scattered through an ENORMOUS conversation. Then ask it to retrieve the third poem word for word. If it succeeds at 100K tokens but fails at 500K, context rot has struck.
This is where Anthropic's new figures get interesting.
The numbers are impressive
Long context retrieval
MRCR v2 — 8-needle test
See Opus 4.6's orange line? It starts at 91.9% with 256K tokens and ends at 78.3% with 1 million. That is a loss of only about 14 points over an additional 750,000 tokens. Compare that with GPT-5.4 falling from 79.3% to 36.6%, or Gemini 3.1 Pro dropping from 59.1% to 25.9%.
The wildest part? Opus 4.6 at 1 MILLION tokens (78.3%) performs almost as well as GPT-5.4 at 256K (79.3%). Read that again. An AI with four times the context performs JUST AS WELL as its competitor with a quarter of it. Pretty impressive!
For perspective, Opus 4.5, the previous version, scored around 27% on the same test. From 27% to 78.3%: almost three times the score in a single update. Hats off.
Great, but what actually changes?
Let's come back down to earth for a moment. A million tokens is fantastic for developers working with huge codebases. I use it daily with Claude Code for my clients' websites. If you are interested, I have described my AI services which cover exactly this. But what about everyone else?
What really changes: you can analyse MUCH longer documents at once, up to 600 PDF pages or 600 images. Long conversations go off track less often. And the pricing rate stays the same whether you use 9,000 or 900,000 tokens.
What does NOT change: if you use AI for short tasks such as writing an email or summarising an article, you will see no difference. Context rot is not DEAD, just better controlled. Degradation still happens, but much more gradually: roughly 2% for each additional 100,000 tokens, compared with around 15–20% before.
The 2% rule
According to the benchmarks, Opus 4.6 loses roughly 2% effectiveness per 100,000 tokens. It is a rule of thumb, but a useful reference point. We have gone from a waterslide to a gentle slope.
Instead of having to reset your conversation every 100–120K tokens, as previously recommended, you now have MUCH more room. Still, if you can start fresh, you might as well. Why accept any degradation, however small, when you don't need to?
If you want to understand how I use these AI tools every day for my clients, my automation and AI training services cover precisely this kind of subject.
A final word
Let's be honest: this is probably Anthropic's most significant update in recent months. More than loops or new tools. A huge context window that ACTUALLY WORKS changes things for everyone using AI daily.
Context rot is still around, but it has just taken a serious hit. And that is a big deal.
Source: Chase AI — analysis of Anthropic benchmarks
Translated from the original French article. Publication dates, examples and figures refer to that original version.


