Amplifying/agent-intelligence

Models Index

The coding models that matter

Ranked by the Amplifying Coding Score, our composite of capability and human-preference signals, with adoption, price, and release date alongside. Click any column to re-sort. The formula and every source are at the bottom of the page.

Top of the board

100 / 98.7

GPT-5.6 Sol edges Claude Opus 5 on the Amplifying Coding Score; Grok 4.6 leads the top tier on price at $2/$6.

Volume kings

23.2T / 11.6T

Stealth Ox Alpha and DeepSeek V4 Flash out-token every frontier model on OpenRouter, weekly.

August riser

+210%

Gemini 3.7 Flash week-over-week token growth: near-frontier coding at Flash prices.

17 models · click a column to sort° partial score · † secondhand · ‡ July token snapshot
#ModelWeightsHome harness
1GPT-5.6 SolOpenAIHighest Coding Index score of any model right now.100.078.3xhigh effort15671.6T14.9%$2 · $1050% promoJul 9, 2026ClosedCodex
2Claude Opus 5Anthropic#1 on AA's Intelligence Index; Claude Code + Opus 5 also tops their Coding Agent Index.98.778.0max effort15661.61T2.58T/wk in OpenRouter's programming category3.5%$5 · $25Jul 24, 2026ClosedClaude Code
3Claude Fable 5AnthropicThe hard-task tier: #2 on intelligence, priciest mainstream model per token.94.176.5max effort1566~450B8%$10 · $50Jun 9, 2026ClosedClaude Code
4Grok 4.6xAICheapest model in the 76+ coding tier; AA flags it for cost efficiency.93.976.8high effort1563~729BGrok latest family alias on OpenRouter$2 · $6Aug 12, 2026Closed
5GPT-5.6 TerraOpenAICoding-skewed 5.6 sibling: top-5 coding despite mid-pack general intelligence.93.276.7max effort1562155B2%$2 · $12Jul 9, 2026ClosedCodex
6Kimi K3Moonshot AIThe open-weights frontier for coding, within ~2 points of Opus 5.91.676.2max effort15621.51T$3 · $15Jul 16, 2026OpenKimi Code CLI
7Gemini 3.7 FlashGoogleAugust's fastest riser: near-frontier coding at Flash prices.89.476.1high effort15572.58T+210% week over week$0.375 · $1.87575% promoAug 13, 2026ClosedGemini CLI / Antigravity
8GPT-5.5OpenAIThe previous OpenAI flagship, still top-10 while the 5.6 family displaces it.87.374.9xhigh effort1561599B7.1%$5 · $30Apr 24, 2026ClosedCodex
9GLM 5.3Z.aiTop-10 coding at a fraction of frontier pricing; 5.2 weights are open, 5.3 Flash weights promised.86.674.8max effort1560972BGLM 5.2 adds another 3.22T/wk$1.4 · $4.4Aug 14, 2026PartialZCode
10Qwen 3.8 MaxAlibabaAlibaba's frontier bid; the 2.4T-parameter open-weights sibling shipped Aug 12.77.571.8preview1560n/aopen Qwen3.8-27B sibling: 193B/wk$2 · $6Aug 3, 2026PartialQwen Code
11Muse Spark 1.2MetaMeta's return to the coding table, ahead of Sonnet 5 on the Coding Index.70.872.21540165B$1.25 · $4.25Aug 5, 2026ClosedMuse Code
12DeepSeek V4 ProDeepSeekDeepSeek's reasoning flagship, refreshed mid-August.64.468.815501.84TApril build; the Aug 12 refresh is ramping$1.12 · $3.37Aug 12, 2026Partial
13Claude Sonnet 5AnthropicThe workhorse mid-tier for agentic coding at a fifth of Fable pricing.60.971.515201.19T3.6%$2 · $10Jun 30, 2026ClosedClaude Code
14Gemini 3.1 ProGoogleGoogle's Pro tier, now behind its own 3.7 Flash on both AA indexes.56.968.81531175B$2 · $12Feb 19, 2026ClosedGemini CLI / Antigravity
15GPT-5.6 LunaOpenAIThe volume play: near-Sonnet coding scores at commodity pricing.39.371.514654.16T0.4%$0.2 · $1.2Jul 9, 2026ClosedCodex
16DeepSeek V4 FlashDeepSeekThe volume king of coding: usage spiked ~570% after the July 31 agent-tuned retrain.32.469.1146611.6T#1 in OpenRouter's programming category$0.05 · $0.1approx; tieredJul 31, 2026Partial
17MiniMax M3MiniMaxOpen-weights long-horizon agent model with a 1M context at commodity cost.7.858.614851.66T$0.3 · $1.2approx; tieredMay 31, 2026Open

Unranked, not unimportant

The score needs a benchmark listing, and the boards run behind the market. These models matter on evidence the boards can't see yet, so they get their own shelf instead of an n/a at the bottom of the table.

GLM 5.3 Flash (Ox Alpha)

Aug 20, 2026

Z.ai (revealed Aug 26)

23.2T/wk

#1 on all of OpenRouter in its first week

The story of August: launched stealth and free as 'Ox Alpha' with a 1M context, out-tokened every named model on OpenRouter inside a week, then revealed as Z.ai's GLM 5.3 Flash. Weights promised.

Why unranked · Too new for the boards: no AA or arena listing yet. The adoption evidence is the strongest of any model on this page.

Composer 2.5

May 18, 2026

Cursor

$0.5 · $2.5

Cursor pricing

Editor-native speed model and Cursor's default; vendor-reported near-Opus SWE-bench at a tenth of the cost.

Why unranked · Not on the public boards: Cursor doesn't expose it outside its own products, so only vendor-reported numbers exist.

Claude Haiku 4.5

Oct 15, 2025

Anthropic

248B/wk

OpenRouter routed tokens

Cheap, fast Claude tier for high-volume subagent and edit loops; aging against the 2026 field.

Why unranked · Aged off the current boards rather than never listed; still a workhorse in subagent fleets.

The August story

A free stealth model out-tokened everyone, then took its mask off

"Ox Alpha" appeared on OpenRouter on August 20 with no named maker, a 1M-token context, and a $0 preview price, and became the #1 model on the entire platform within a week at over 23T tokens. On August 26, Z.ai confirmed to Bloomberg it's a GLM-series model, now listed as GLM 5.3 Flash. Two lessons for anyone watching this market: free plus capable moves developer volume faster than any launch campaign, and the open-ecosystem harnesses (OpenCode was the launch partner) are where that volume lands first.

Volume watch

Huge on OpenRouter, not (yet) coding-quality leaders. If your product serves agent platforms, these move more tokens than most of the table above.

MiMo-V2.5Xiaomi9.92T/wk#3 on OpenRouter; Pro variant scores 60.2 on the Coding Index (secondhand)
Hy3Tencent7.16T/wkopen weights; Coding Index 58.8 (secondhand)
Nemotron 3 Ultra (free)NVIDIA5.4T/wkfree tier absorbing enormous routed volume
GLM 5.2Z.ai3.22T/wkopen weights (MIT); Coding Index 68.8; superseded by 5.3 but still huge
Solar Pro 4Upstage573B/wkAug 12 release; a sleeper climbing the rankings

What businesses actually pay for

A third adoption signal, from a different wallet: share of model API spend among US businesses on Ramp (July 2026). The headline is the lag: prior-generation models still carry 46% of spend while the current frontier ramps, and Anthropic holds 63.4% of the pool to OpenAI's 36%.

Claude Opus 4.8
28%-2.9 pt
GPT-5.6 Sol
14.9%+14.9 pt
Claude Sonnet 4.6
8.3%-4.4 pt
Claude Fable 5
8%+6.2 pt
GPT-5.5
7.1%-4.7 pt
Claude Opus 4.6
6.9%-6.4 pt
Claude Sonnet 5
3.6%+3.6 pt
Claude Opus 5
3.5%+3.5 pt
GPT-5.4
2.4%-1.3 pt
GPT-5.6 Terra
2%+2 pt

Emerald = Anthropic, grey = OpenAI. Delta vs the prior month. Source: Ramp AI Index. Card-visible model API spend only: Google (billed through cloud contracts) and open-weights models (no per-model spend line) are invisible to this measure by construction.

How to read this table

The Amplifying Coding Score blends two independent signals: 60% Artificial Analysis Coding Index (Terminal-Bench v2.1 + SciCode, a capability benchmark) and 40% arena coding Elo (OpenLM's Arena+ coding column, Aug 22 snapshot, an LLM-judged preference signal). Each is min-max normalized within this cohort, so 100 is the cohort frontier and 0 the cohort floor, not absolute quality. A model missing one component is scored on the other alone and marked °; missing both means no score and no rank.

Coding Index is quoted at the effort setting noted per row. Scores marked † reached us through public mirrors of the AA board rather than the live dashboard. Composer 2.5 and Ox Alpha aren't on either board yet, so they carry no score rather than a guessed one.

Tokens per week is OpenRouter's routed volume, which counts only bring-your-own-key traffic through OpenRouter. First-party usage (Claude Code on Anthropic's API, Codex on OpenAI's, Gemini CLI on Google's) never appears there, so frontier-lab volumes are heavily understated. Figures marked ‡ are from a July 22 community snapshot because the model has since left the visible top lists.

Biz spend is the model's share of API spend among US businesses on Ramp's card data (July 2026, per the Ramp AI Index). A dash means invisible to that measure, not zero: Google bills through cloud contracts and open-weights models have no per-model spend line. It doesn't feed the score; it's a who-pays signal next to the who-routes signal.

Tokenizers differ per provider, so cross-model token totals are directional, not strictly comparable. One more directional fact worth knowing: Chinese-lab models have out-tokened US models on OpenRouter for 15+ straight weeks (34.25T vs 9.17T in the week of Aug 3–9, per public trackers).

Curated, not exhaustive. This is the set we'd put in a harness-and-model test matrix today, plus the volume outliers you'd be wrong to ignore. It updates when the landscape moves. Sources: Artificial Analysis, OpenRouter rankings, OpenRouter models API, and recent release coverage.

Models are half the story

The harness running the model changes the result. The Coding Agents Index tracks the harnesses themselves: adoption, attributed commits, merge rates, and pricing.

Coding Agents Index