Three Kinds of Smart: Why 'Best Model' Is a Category Error
Raw intelligence, SWE effectiveness, and flash-tier quality for real-world work are three different competitions — and the leaderboard keeps conflating them. A map of who actually holds what.
Every model launch arrives with the same claim: number one on the chart. But "the chart" is three different competitions wearing one scoreboard. Conflating them produces bad procurement decisions, bad product choices, and bad journalism. So let's separate them.
#Axis one: raw intelligence
The olympiad axis — frontier reasoning, novel problem solving, the top of the curve on hard evals. This is the axis the labs' marketing lives on, and it's real: a small set of models genuinely does things nothing else can. For research, strategy, and genuinely novel problems, raw intelligence is the constraint and the flagship price is defensible.
But it's also the axis with the least commercial surface area. Most deployed workloads never touch the top of the curve. Raw intelligence is the headline; the other two axes are the economy.
#Axis two: coding and SWE effectiveness
Software engineering is its own competition, and it rewards a different profile: long-horizon task decomposition, tool use, self-correction, code-base navigation, and — critically — failure cost. A coding agent that's right 92% of the time on real repos is not 8% worse than one that's right 94%; on a 20-step task, that gap compounds into the difference between "reviews the diff" and "rebuilds the feature."
This is why SWE benchmarks keep reordering the leaderboard. Raw intelligence helps, but the effective stack is scaffolding, context management, and tool ergonomics as much as weights. It's also why dedicated coding tiers and agentic products — Claude Code, Devin, Codex-class systems — can beat a "smarter" general model on the actual job.
#Axis three: flash quality for real-world applications
The most under-covered axis: how good is the cheap/fast tier at the work that actually runs in production? Extraction, triage, drafting, classification, routing, short answers, structured output. Here the metrics are p95 latency, cost per million tokens, and quality floor — not ceiling.
This is the axis where the market share is, and it's the axis where Google's position is badly underestimated. The "Google is falling behind" narrative keeps being written from axis one. On axis three, Gemini Flash-class models pair near-frontier quality with the industry's best cost and latency profile, a giant context window, native multimodality, and a distribution channel through Cloud, Workspace, and Android that nobody else can replicate. Companies counting Google out are reading the wrong scoreboard.
#The contenders, honestly assessed
- Google — dominant on axis three, competitive on axes one and two, and owns the deployment surface (devices, workspace, cloud). The most complete position in the field.
- Meta — a stronger contender than its press coverage suggests. The open-weight Llama line is the substrate of the regional and on-device market; for applications that live on hardware — phones, glasses, edge — Meta's strategy compounds in ways API providers can't match.
- OpenAI — still sets the reference on axis one and holds the developer mindshare, but is now defending price at the bottom of the market.
- Anthropic — owns the axis-two reputation for long-horizon agentic work; the question is whether that extends downward into volume tiers.
- xAI — well-capitalized and genuinely competitive; treats compute as the strategy. A real contender, not a vanity project.
- Open-source and Eastern labs — DeepSeek, Qwen, and friends have already won their segment of axis three. Their role is structurally important: regional privacy, self-hosting, and affordable speed.
#The geography problem
Here's the part the leaderboards never show. Individual Eastern providers will struggle to penetrate Western engineering markets — not primarily because of capability, but because of compute access, latency to Western infrastructure, and enterprise trust/compliance posture. The majority of engineering spend sits in the Western sphere, and serving it well means serving it near it.
Open weights are the workaround: when the weights ship, jurisdiction becomes a deployment decision rather than a vendor risk. That's why Llama and DeepSeek derivatives keep showing up in the radar's fine-tune long tail — the ecosystem is doing locally what the providers can't do remotely.
// share this dispatch
// the signal
One email. The week's sharpest AI analysis.
Every Friday: the dispatches that mattered, the models that shipped, and the one chart you need to see. No spam, no filler — unsubscribe anytime.
Join 2,400+ researchers and engineers. Unsubscribe anytime.