The Open-Weights Race Just Got a 10M-Token Problem
Meta's Llama 4 Scout shipped with a 10M-token context window, and every open lab is now racing to match it. Here's why context length became the new frontier battleground.
For two years, the open-weights race was measured in benchmark points. This week it started being measured in tokens — ten million of them.
Meta's Llama 4 Scout landed with a context window large enough to hold roughly 7,500 pages of text in a single pass, and the rest of the ecosystem responded within 48 hours with their own long-context claims.
#Why context became the battleground
Reasoning benchmarks have saturated. The top open and closed models now sit within a few points of each other on math and code suites, so labs need a new axis to compete on. Context is that axis.
Long context unlocks workloads that short-context models simply cannot touch:
- Whole-repository code agents that read the codebase in one pass
- Multi-hour meeting and video summarization
- Legal and scientific document review across hundreds of files
- Persistent agent memory without external retrieval
#The catch nobody mentions
A 10M-token window is a capability claim, not a quality guarantee. Attention quality degrades unevenly across long contexts, and most public evaluations still test retrieval at 128K or less.
#What to watch next
Expect DeepSeek and Qwen to respond within the quarter, and expect inference providers to start pricing long-context requests differently. The economics of attention are about to become the economics of the whole stack.
// share this dispatch
// the signal
One email. The week's sharpest AI analysis.
Every Friday: the dispatches that mattered, the models that shipped, and the one chart you need to see. No spam, no filler — unsubscribe anytime.
Join 2,400+ researchers and engineers. Unsubscribe anytime.