Inference Cost Is the New Benchmark
Benchmark leaderboards are saturating. The metric that now decides model selection is cost per million tokens at acceptable quality — and it's reshaping how labs build.
Ask any team shipping an AI product which model they use, and the answer increasingly starts with a dollar sign. The leaderboard era is ending; the unit-economics era has begun.
#The saturation problem
Top models now cluster within a few points of each other on MMLU-class suites, math olympiads, and coding benchmarks. When quality differences shrink below the noise of prompt variation, price and latency become the tiebreakers.
#The new selection stack
Teams we talk to now evaluate models on a four-factor stack:
- Quality floor — does it clear the minimum bar for the task?
- Cost per million tokens — at production batch sizes, not list price
- Latency distribution — p95, not average
- Context economics — how the price scales as context grows
#What this does to labs
When cost is the battleground, architecture decisions become pricing decisions. Sparse MoE, speculative decoding, and aggressive quantization are no longer research curiosities — they are the product.
#The takeaway
Expect marketing to shift from "state of the art" to "state of the art per dollar." The labs that win the next two years will be the ones that treat inference efficiency as a first-class research objective.
// share this dispatch
// the signal
One email. The week's sharpest AI analysis.
Every Friday: the dispatches that mattered, the models that shipped, and the one chart you need to see. No spam, no filler — unsubscribe anytime.
Join 2,400+ researchers and engineers. Unsubscribe anytime.