Skip to main content
TOWARD//AGI
OPEN-WEIGHTS-CLOSING-THE-GAP
Trends#open-weights#frontier#benchmarks

The Open-Weights Gap Is Closing — And It's Closing Fast

A year ago, open-weights models trailed frontier closed models by a wide margin. Now the gap is shrinking quarter over quarter. Here's what changed.

Mira ChenToward AGI Editorial2 min read

Twelve months ago, the conventional wisdom was clear: open-weights models were good for their size, but they could not compete with the frontier closed models on reasoning, coding, or instruction following.

That conventional wisdom is now outdated.

#The convergence chart

If you plot benchmark scores over time, two lines are visible: one for the best closed models (GPT, Claude, Gemini) and one for the best open-weights models (Llama, Qwen, DeepSeek, Mistral). A year ago, the gap was significant. Today, it is shrinking quarter over quarter.

On several widely-used benchmarks — coding suites, math reasoning, multilingual understanding — the top open-weights models now sit within striking distance of the closed frontier. On a few narrow tasks, they have already matched or surpassed it.

#What changed

Three forces are driving the convergence:

1. Architecture innovation went open. Techniques like Mixture of Experts, which were initially proprietary, are now standard in open-weights releases. The labs publishing openly are not trailing — they are often leading.

2. Inference cost became a competitive advantage. When open models can be self-hosted at a fraction of closed-model API pricing, the economic calculus shifts. Teams that optimize for cost-quality tradeoffs increasingly land on open weights.

3. The evaluation bar moved. As benchmarks saturate, the differentiating factors become reliability, controllability, and deployability — areas where open models have structural advantages.

#What this means for labs

The closed-model moat was never just about model weights. It was about training data, infrastructure, and iteration speed. But as open models close the quality gap, the value proposition of "closed but slightly better" becomes harder to sustain.

Expect the next competitive frontier to shift toward:

  • Specialized models fine-tuned for specific verticals.
  • Inference infrastructure — who can serve the model fastest and cheapest.
  • Trust and verifiability — open weights allow inspection that closed models cannot.

#The honest caveat

The gap is closing, but it has not closed. On the hardest reasoning tasks, long-horizon agentic work, and multimodal understanding, the frontier closed models still hold an edge. The question is how long that edge persists.

If the current trajectory holds, we are looking at rough parity on most benchmarks within 12-18 months. The labs that plan for that world — rather than assuming the gap will remain — will be best positioned.

#The bottom line

The open-weights ecosystem is no longer the "budget option." It is a genuine alternative to closed frontier models for an increasing range of workloads. The gap is closing, and the pace of closure is accelerating.

That changes the strategy for everyone building on top of these models.

// share this dispatch

// the signal

One email. The week's sharpest AI analysis.

Every Friday: the dispatches that mattered, the models that shipped, and the one chart you need to see. No spam, no filler — unsubscribe anytime.

Join 2,400+ researchers and engineers. Unsubscribe anytime.