Models Index
The coding models that matter
Ranked by the Amplifying Coding Score, our composite of capability and human-preference signals, with adoption, price, and release date alongside. Click any column to re-sort. The formula and every source are at the bottom of the page.
Top of the board
100 / 98.7
GPT-5.6 Sol edges Claude Opus 5 on the Amplifying Coding Score; Grok 4.6 leads the top tier on price at $2/$6.
Volume kings
23.2T / 11.6T
Stealth Ox Alpha and DeepSeek V4 Flash out-token every frontier model on OpenRouter, weekly.
August riser
+210%
Gemini 3.7 Flash week-over-week token growth: near-frontier coding at Flash prices.
| # | Model | Weights | Home harness | |||||||
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-5.6 SolOpenAIHighest Coding Index score of any model right now. | 100.0 | 78.3xhigh effort | 1567 | 1.6T | 14.9% | $2 · $1050% promo | Jul 9, 2026 | Closed | Codex |
| 2 | Claude Opus 5Anthropic#1 on AA's Intelligence Index; Claude Code + Opus 5 also tops their Coding Agent Index. | 98.7 | 78.0max effort | 1566 | 1.61T2.58T/wk in OpenRouter's programming category | 3.5% | $5 · $25 | Jul 24, 2026 | Closed | Claude Code |
| 3 | Claude Fable 5AnthropicThe hard-task tier: #2 on intelligence, priciest mainstream model per token. | 94.1 | 76.5max effort | 1566 | ~450B ‡ | 8% | $10 · $50 | Jun 9, 2026 | Closed | Claude Code |
| 4 | Grok 4.6xAICheapest model in the 76+ coding tier; AA flags it for cost efficiency. | 93.9 | 76.8high effort | 1563 | ~729BGrok latest family alias on OpenRouter | — | $2 · $6 | Aug 12, 2026 | Closed | — |
| 5 | GPT-5.6 TerraOpenAICoding-skewed 5.6 sibling: top-5 coding despite mid-pack general intelligence. | 93.2 | 76.7max effort | 1562 | 155B ‡ | 2% | $2 · $12 | Jul 9, 2026 | Closed | Codex |
| 6 | Kimi K3Moonshot AIThe open-weights frontier for coding, within ~2 points of Opus 5. | 91.6 | 76.2max effort | 1562 | 1.51T | — | $3 · $15 | Jul 16, 2026 | Open | Kimi Code CLI |
| 7 | Gemini 3.7 FlashGoogleAugust's fastest riser: near-frontier coding at Flash prices. | 89.4 | 76.1high effort | 1557 | 2.58T+210% week over week | — | $0.375 · $1.87575% promo | Aug 13, 2026 | Closed | Gemini CLI / Antigravity |
| 8 | GPT-5.5OpenAIThe previous OpenAI flagship, still top-10 while the 5.6 family displaces it. | 87.3 | 74.9xhigh effort | 1561 | 599B ‡ | 7.1% | $5 · $30 | Apr 24, 2026 | Closed | Codex |
| 9 | GLM 5.3Z.aiTop-10 coding at a fraction of frontier pricing; 5.2 weights are open, 5.3 Flash weights promised. | 86.6 | 74.8 †max effort | 1560 | 972BGLM 5.2 adds another 3.22T/wk | — | $1.4 · $4.4 | Aug 14, 2026 | Partial | ZCode |
| 10 | Qwen 3.8 MaxAlibabaAlibaba's frontier bid; the 2.4T-parameter open-weights sibling shipped Aug 12. | 77.5 | 71.8 †preview | 1560 | n/aopen Qwen3.8-27B sibling: 193B/wk | — | $2 · $6 | Aug 3, 2026 | Partial | Qwen Code |
| 11 | Muse Spark 1.2MetaMeta's return to the coding table, ahead of Sonnet 5 on the Coding Index. | 70.8 | 72.2 † | 1540 | 165B | — | $1.25 · $4.25 | Aug 5, 2026 | Closed | Muse Code |
| 12 | DeepSeek V4 ProDeepSeekDeepSeek's reasoning flagship, refreshed mid-August. | 64.4 | 68.8 † | 1550 | 1.84TApril build; the Aug 12 refresh is ramping | — | $1.12 · $3.37 | Aug 12, 2026 | Partial | — |
| 13 | Claude Sonnet 5AnthropicThe workhorse mid-tier for agentic coding at a fifth of Fable pricing. | 60.9 | 71.5 | 1520 | 1.19T | 3.6% | $2 · $10 | Jun 30, 2026 | Closed | Claude Code |
| 14 | Gemini 3.1 ProGoogleGoogle's Pro tier, now behind its own 3.7 Flash on both AA indexes. | 56.9 | 68.8 † | 1531 | 175B ‡ | — | $2 · $12 | Feb 19, 2026 | Closed | Gemini CLI / Antigravity |
| 15 | GPT-5.6 LunaOpenAIThe volume play: near-Sonnet coding scores at commodity pricing. | 39.3 | 71.5 † | 1465 | 4.16T | 0.4% | $0.2 · $1.2 | Jul 9, 2026 | Closed | Codex |
| 16 | DeepSeek V4 FlashDeepSeekThe volume king of coding: usage spiked ~570% after the July 31 agent-tuned retrain. | 32.4 | 69.1 † | 1466 | 11.6T#1 in OpenRouter's programming category | — | $0.05 · $0.1approx; tiered | Jul 31, 2026 | Partial | — |
| 17 | MiniMax M3MiniMaxOpen-weights long-horizon agent model with a 1M context at commodity cost. | 7.8 | 58.6 † | 1485 | 1.66T | — | $0.3 · $1.2approx; tiered | May 31, 2026 | Open | — |
Unranked, not unimportant
The score needs a benchmark listing, and the boards run behind the market. These models matter on evidence the boards can't see yet, so they get their own shelf instead of an n/a at the bottom of the table.
GLM 5.3 Flash (Ox Alpha)
Aug 20, 2026Z.ai (revealed Aug 26)
23.2T/wk
#1 on all of OpenRouter in its first week
The story of August: launched stealth and free as 'Ox Alpha' with a 1M context, out-tokened every named model on OpenRouter inside a week, then revealed as Z.ai's GLM 5.3 Flash. Weights promised.
Why unranked · Too new for the boards: no AA or arena listing yet. The adoption evidence is the strongest of any model on this page.
Composer 2.5
May 18, 2026Cursor
$0.5 · $2.5
Cursor pricing
Editor-native speed model and Cursor's default; vendor-reported near-Opus SWE-bench at a tenth of the cost.
Why unranked · Not on the public boards: Cursor doesn't expose it outside its own products, so only vendor-reported numbers exist.
Claude Haiku 4.5
Oct 15, 2025Anthropic
248B/wk
OpenRouter routed tokens
Cheap, fast Claude tier for high-volume subagent and edit loops; aging against the 2026 field.
Why unranked · Aged off the current boards rather than never listed; still a workhorse in subagent fleets.
The August story
A free stealth model out-tokened everyone, then took its mask off
"Ox Alpha" appeared on OpenRouter on August 20 with no named maker, a 1M-token context, and a $0 preview price, and became the #1 model on the entire platform within a week at over 23T tokens. On August 26, Z.ai confirmed to Bloomberg it's a GLM-series model, now listed as GLM 5.3 Flash. Two lessons for anyone watching this market: free plus capable moves developer volume faster than any launch campaign, and the open-ecosystem harnesses (OpenCode was the launch partner) are where that volume lands first.
Volume watch
Huge on OpenRouter, not (yet) coding-quality leaders. If your product serves agent platforms, these move more tokens than most of the table above.
| MiMo-V2.5 | Xiaomi | 9.92T/wk | #3 on OpenRouter; Pro variant scores 60.2 on the Coding Index (secondhand) |
| Hy3 | Tencent | 7.16T/wk | open weights; Coding Index 58.8 (secondhand) |
| Nemotron 3 Ultra (free) | NVIDIA | 5.4T/wk | free tier absorbing enormous routed volume |
| GLM 5.2 | Z.ai | 3.22T/wk | open weights (MIT); Coding Index 68.8; superseded by 5.3 but still huge |
| Solar Pro 4 | Upstage | 573B/wk | Aug 12 release; a sleeper climbing the rankings |
What businesses actually pay for
A third adoption signal, from a different wallet: share of model API spend among US businesses on Ramp (July 2026). The headline is the lag: prior-generation models still carry 46% of spend while the current frontier ramps, and Anthropic holds 63.4% of the pool to OpenAI's 36%.
Emerald = Anthropic, grey = OpenAI. Delta vs the prior month. Source: Ramp AI Index. Card-visible model API spend only: Google (billed through cloud contracts) and open-weights models (no per-model spend line) are invisible to this measure by construction.
How to read this table
The Amplifying Coding Score blends two independent signals: 60% Artificial Analysis Coding Index (Terminal-Bench v2.1 + SciCode, a capability benchmark) and 40% arena coding Elo (OpenLM's Arena+ coding column, Aug 22 snapshot, an LLM-judged preference signal). Each is min-max normalized within this cohort, so 100 is the cohort frontier and 0 the cohort floor, not absolute quality. A model missing one component is scored on the other alone and marked °; missing both means no score and no rank.
Coding Index is quoted at the effort setting noted per row. Scores marked † reached us through public mirrors of the AA board rather than the live dashboard. Composer 2.5 and Ox Alpha aren't on either board yet, so they carry no score rather than a guessed one.
Tokens per week is OpenRouter's routed volume, which counts only bring-your-own-key traffic through OpenRouter. First-party usage (Claude Code on Anthropic's API, Codex on OpenAI's, Gemini CLI on Google's) never appears there, so frontier-lab volumes are heavily understated. Figures marked ‡ are from a July 22 community snapshot because the model has since left the visible top lists.
Biz spend is the model's share of API spend among US businesses on Ramp's card data (July 2026, per the Ramp AI Index). A dash means invisible to that measure, not zero: Google bills through cloud contracts and open-weights models have no per-model spend line. It doesn't feed the score; it's a who-pays signal next to the who-routes signal.
Tokenizers differ per provider, so cross-model token totals are directional, not strictly comparable. One more directional fact worth knowing: Chinese-lab models have out-tokened US models on OpenRouter for 15+ straight weeks (34.25T vs 9.17T in the week of Aug 3–9, per public trackers).
Curated, not exhaustive. This is the set we'd put in a harness-and-model test matrix today, plus the volume outliers you'd be wrong to ignore. It updates when the landscape moves. Sources: Artificial Analysis, OpenRouter rankings, OpenRouter models API, and recent release coverage.
Models are half the story
The harness running the model changes the result. The Coding Agents Index tracks the harnesses themselves: adoption, attributed commits, merge rates, and pricing.