Leaderboard / large language models
LLM reference table
We do not run a language-model harness. Every benchmark figure on this page is a third party's measurement, named and dated, and every specification and price is the vendor's own published figure. Where a vendor does not publish a number we leave the cell empty and say so rather than sourcing it from an aggregator. This board is a sourced reference table, not an LLM Tier ranking. It has no composite index because we have not earned the right to publish one: we hold no model-level run records, so we could not show you the arithmetic.
The model layer and the product boards are separate measurements. A product built on a strong model does not inherit that model's standing here, and neither product board's composite is derived from anything on this page.
Top of Arena WebDev
1691 Elo
Claude Opus 5 (max effort), 1691 Elo, Arena, 21 Aug 2026.
Best independently measured agentic coding score
89.5%
GPT-5.6 Sol at xhigh effort, 89.5 percent pass@1 on Terminal-Bench v2.1, ahead of Claude Opus 5 at max effort on 89.1 percent and Grok 4.6 at high effort on 88.4 percent.
Widest published context window
1,048,576 tokens
1,048,576 tokens, published by Moonshot for Kimi K3. Several models are documented at "1,000,000"; we print the largest figure a vendor actually states.
Cheapest published output token
No single answer
This has no single answer. DeepSeek-V4-Flash is 0.66 dollars per million output tokens off-peak and 1.32 at peak. GPT-5.6 Luna is a flat 1.20. So DeepSeek is cheapest for most of the day and Luna is cheapest during DeepSeek's peak window, which runs 01:00 to 04:00 and 06:00 to 10:00 UTC. A cheapest-output crown that flips on the clock is not a crown, so we publish both.
26 models, sourced roster
Sorted by Arena WebDev Elo, descending. Rows with no published Elo sort last, alphabetically by vendor. That ordering is Arena's measurement, not ours.
| Model | Vendor | Status / tier | Licence | Context window | Max output | Price in $/MTok | Price out $/MTok | Modality | Arena WebDev | AA index v4.1.1 | Terminal-Bench v2.1 | Release note |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | Current · flagship | Proprietary | 1,000,000 | 128,000 | 5.00 | 25.00 | text and image in, text out | 1691rank 1, 8,116 votes | 63 at max effort, 63 at xhigh, 61 at high | 89.1% at max effort | Release date not stated by the vendor; training data cutoff May 2026. Arena also lists the high-effort variant separately: 1663 Elo, rank 4, 8659 votes. |
| Kimi K3 | Moonshot AI | Current · flagship | Modified MIT, published by Moonshot as the Kimi K3 License; open weights | 1,048,576 | Not stated on the vendor pricing page | 3.00 on a cache miss, 0.30 on a cache hit | 15.00 | Multimodal reasoning with thinking always on, per secondary reporting | 1674rank 2, 4,546 votes | not measured | not measured | Released 16 July 2026; full open weights 27 July 2026. Kimi K3 always reasons and exposes a reasoning_effort setting of low, high or max, defaulting to max, which is why Arena lists it as kimi-k3-max (max effort). A third-party host lists the same weights at 2.60 / 13.00. That is the host's price, not Moonshot's, and we publish Moonshot's. |
| Qwen3.8-Max | Alibaba | Current · flagship | Proprietary | 1,048,576, charged at one flat rate across the whole window with no length tiering | 131,000 | 2.00 (batch inference 1.65; context-caching discount available) | 6.00, covering chain of thought plus answer (batch inference 4.951) | text, image and video in, text out | 1669Prelimrank 3, 3,212 votes | not measured | not measured | Previewed 19 July 2026, global API access via Alibaba Cloud Model Studio 3 August 2026. There is no independently measured coding-benchmark score for this model, so we publish none, and its Arena Elo carries Arena's own Preliminary flag on 3,212 votes. It is also the only flagship on this board with no 200k price step, which makes it unusually cheap on very long prompts. |
| Grok 4.6 | xAI | Current · flagship | Proprietary | 500,000 | No stated text output limit | 2.00 below 200k prompt tokens, 4.00 above (cached 0.50 / 1.00) | 6.00 below 200k prompt tokens, 12.00 above | The xAI models page lists text in and text out; secondary reporting says image input is also supported, which we have not confirmed | 1629Prelimrank 5, 1,523 votes | 61 per secondary reporting of the index | 88.4% at high effort | 12 August 2026 per secondary reporting; not dated on xAI's models page. Arena files xAI's models under the lab name SpaceXAI; the vendor documentation domain is docs.x.ai and we publish the vendor as xAI. See Grok 4.3 below for the headline counter-example on this board. |
| Claude Fable 5 | Anthropic | Current · flagship | Proprietary | 1,000,000 | 128,000 (300,000 on the Batch API with the output-300k-2026-03-24 beta header) | 10.00 | 50.00 | text and image in, text out | 1626rank 6, 8,382 votes | 62 (Adaptive Reasoning, Max Effort) | Not in the published top three | Available from 9 June 2026. |
| GPT-5.6 Sol | OpenAI | Current · flagship | Proprietary | Not published on OpenAI's pricing page | Not published on OpenAI's pricing page | 4.00 (cached input 0.40) | 20.00 | text in, text out per the pricing table | 1619rank 7, 9,499 votes | 61 at max effort | 89.5% at xhigh effort, the published leader | Limited preview 26 June 2026, public launch 9 July 2026. Arena entry uses xhigh effort in the Codex harness. See the pricing warning box above: Arena lists this model at the older GPT-5.5 price. |
| Gemini 3.7 Flash | Current · workhorse | Proprietary | Not stated on the fetched pages | Not stated | 0.75 through 31 December 2026, then 1.50 from 1 January 2027 | 3.75 through 31 December 2026, then 7.50 from 1 January 2027 | Not stated per model on the fetched page | 1587Prelimrank 10, 2,552 votes | not measured | not measured | 13 August 2026 per secondary reporting; Google's own pages mark it stable but do not date it. Arena entry uses high effort and carries Arena's Preliminary flag. Google publishes DeepSWE v1.1 rising from 49.0 to 65.3 percent and GDM-MRCR v2 at 128k rising from 91.8 to 97.0 percent versus the prior Flash. Those are Google's own figures on Google's own evaluation and are labelled vendor-claimed, not independent. | |
| DeepSeek-V4-Pro | DeepSeek | Current · flagship | MIT | 1,000,000 | Not stated on the pricing page | 0.66 off-peak, 1.32 at peak (cache hit 0.022 off-peak, 0.044 peak) | 1.98 off-peak, 3.96 at peak | text in, text out; a separate deepseek-v4-flash-vision-exp handles vision | 1582rank 12, 2,869 votes | not measured | not measured | Version string DeepSeek-V4-Pro-0813 implies 13 August 2026; not dated in prose. Arena entry uses high effort on the 20260813 build. DeepSeek has no single price: off-peak is half of peak, and peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC. Any value comparison must state which rate it used. |
| DeepSeek-V4-Flash | DeepSeek | Current · budget | MIT | 1,000,000 | Not stated | 0.22 off-peak, 0.44 at peak (cache hit 0.007 / 0.014) | 0.66 off-peak, 1.32 at peak | text in, text out | 1579rank 13, 3,537 votes | not measured | not measured | Version string DeepSeek-V4-Flash-0731 implies 31 July 2026. Arena entry uses high effort. DeepSeek has no single price: off-peak is half of peak, and peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC. Any value comparison must state which rate it used. See the cheapest-output callout above for why this row and GPT-5.6 Luna cannot both hold that crown. |
| Claude Sonnet 5 | Anthropic | Current · workhorse | Proprietary | 1,000,000 | 128,000 | 2.00 | 10.00 | text and image in, text out | 1539rank 22, 7,267 votes | not measured | not measured | Release date not stated by the vendor. Arena effort level: high. |
| Gemini 3.6 Flash | Current · workhorse | Proprietary | Not stated | Not stated | 0.75, rising to 1.50 on 1 January 2027 | 3.75, rising to 7.50 on 1 January 2027 | Not stated | 1539rank 21, 6,464 votes | not measured | not measured | Not dated on Google's pages; marked stable. Arena entry uses high effort. | |
| Muse Spark 1.1 | Meta | Current · flagship | Proprietary, closed weights | 1,000,000 | Not specified in Meta's announcement | Not published by Meta | Not published by Meta | text, images, video and PDFs in, text out | 1539rank 20, 6,204 votes | not measured | not measured | Developed by Meta Superintelligence Labs. Muse Spark launched 8 April 2026 as Meta's first closed-weight model; 1.1 released 9 July 2026. Meta's own announcement publishes no per-token price and no numerical benchmark score, so both cells are empty here. Price is the single most-quoted number in secondary coverage and the single number Meta does not state. Arena also lists a muse-spark-1.2 at 1534 Elo; we have not sourced 1.2 from Meta and so do not list it as a row. |
| GPT-5.6 Terra | OpenAI | Current · workhorse | Proprietary | Not published | Not published | 2.00 (cached input 0.20) | 12.00 | text in, text out | 1520rank 27, 5,677 votes | not measured | not measured | Launched 9 July 2026. Arena entry uses xhigh effort in the Codex harness. |
| GPT-5.6 Luna | OpenAI | Current · budget | Proprietary | Not published | Not published | 0.20 (cached input 0.02) | 1.20 | text in, text out | 1518rank 29, 5,742 votes | not measured | not measured | Launched 9 July 2026. Arena entry uses xhigh effort in the Codex harness. Flat 1.20 output price. See the DeepSeek-V4-Flash row for the peak and off-peak comparison, since neither is cheapest all day. |
| Gemini 3.5 Flash | Current · workhorse | Proprietary | Not stated | Not stated | 1.50 | 9.00 | Not stated | 1499rank 34, 7,559 votes | not measured | not measured | Not dated; marked stable. Arena entry uses high effort. This middle rung is worse value than both newer Flash models, which cost 0.75 / 3.75. Inside one vendor's line, newer here is both better and cheaper. | |
| Gemini 3.5 Flash-Lite | Current · budget | Proprietary | Not stated | Not stated | 0.30 | 2.50 | Not stated | 1449rank 47, 228 votes | not measured | not measured | Not dated; marked stable. Arena's interval on this row is plus or minus 43 points on only 228 votes (1449 ± 43), far too wide to rank against its neighbours. | |
| Gemini 3.1 Pro Preview | Preview · flagship | Proprietary | Not stated on Google's models or pricing pages | Not stated | 2.00 for prompts up to 200k tokens, 4.00 above 200k (batch 1.00 / 2.00) | 12.00 up to 200k, 18.00 above 200k (batch 6.00 / 9.00) | Not stated per model on the fetched page | 1446rank 48, 21,573 votes | not measured | not measured | Still carries preview status. This is the only Pro-tier Gemini 3 text model on Google's own list; every stable Gemini 3 text model is a Flash or Flash-Lite. | |
| GPT-5.3 Codex | OpenAI | Current · coding | Proprietary | Not published | Not published | 1.75 (cached input 0.175) | 14.00 | text in, text out, coding specialised | 1408rank 61, 2,526 votes | not measured | not measured | Release date not published. Arena entry runs in the Codex harness. |
| Grok 4.3 | xAI | Current · workhorse | Proprietary | 1,000,000 | Not stated | 1.25 below 200k prompt tokens, 2.50 above (cached 0.20 / 0.40) | 2.50 below 200k, 5.00 above | text in, text out | 1357rank 86, 12,942 votes | not measured | not measured | Not dated on xAI's models page. HEADLINE COUNTER-EXAMPLE ON THIS BOARD: Grok 4.3 has double the context window of xAI's newer flagship (Grok 4.6) at roughly 40 percent of its output price, and sits 272 Elo points lower on Arena WebDev. Newer, wider and cheaper are three different axes and no vendor wins all three at once. |
| Amazon Nova 2 Lite | Amazon | Current · workhorse | Proprietary | 1,000,000 | 65,536 | Not published on the Nova 2 user guide, which defers to Amazon Bedrock Pricing; the Bedrock pricing page carries no Nova 2 row | Not published | text, images, video and documents in, text out | not present | not measured | not measured | Release date not stated on the Nova 2 user guide. Nova 2 is Amazon's current generation and its whole roster is Nova 2 Lite, Nova 2 Sonic and Nova Multimodal Embeddings. Nova Premier, Nova Pro, Nova Lite and Nova Micro are all the superseded Nova 1 generation. Every Nova 2 model is 1,000,000 tokens in and 65,536 out. |
| Claude Haiku 4.5 | Anthropic | Current · budget | Proprietary | 200,000 | 64,000 | 1.00 | 5.00 | text and image in, text out | not present | not measured | not measured | Model ID pins the snapshot to 1 October 2025. |
| Command A+ | Cohere | Current · flagship | Apache 2.0 | Not published for this model | Not published | Not published. cohere.com/pricing carries only legacy Command rows and has no Command A+ row at all. | Not published | text and image in, text out | not present | not measured | not measured | 20 May 2026. 218 billion total parameters with 25 billion active in a sparse mixture of experts, 48 languages, native citation generation. Cohere's flagship is three months old and its own public pricing page does not list it. We publish the model with no price rather than sourcing one from an aggregator. |
| Gemini 3.1 Flash-Lite | Current · budget | Proprietary | Not stated | Not stated | 0.25 for text, image and video; 0.50 for audio | 1.50 | text, image, video and audio in per the pricing tiers | not present | not measured | not measured | Not dated; marked stable. Google's own model list describes this as frontier class; we attribute that claim to Google and do not restate it as fact. | |
| Llama 4 Maverick | Meta | Open weight · workhorse | Open weights | Previously published here as 1,000,000; not re-verified in this pass | Not re-verified | No vendor list price exists | No vendor list price exists | Not re-verified | not present | not measured | not measured | 5 April 2025, alongside Llama 4 Scout. No new Llama model has shipped since. We previously published 0.22 / 0.85 for this model. That figure was unsourceable by construction, not merely stale: these are open weights and there is no vendor list price at all, only per-host prices. We have removed it. Llama is no longer Meta's frontier: Meta's frontier since April 2026 is Muse Spark, which is closed and paid. |
| Mistral Medium 3.5 | Mistral AI | Current · flagship | Modified MIT | Not listed on Mistral's models overview page | Not listed | Not published on Mistral's docs page | Not published on Mistral's docs page | Multimodal per Mistral's own description | not present | not measured | not measured | Version string v26.04 implies April 2026; not dated in prose. Mistral's current premier flagship, described by Mistral as frontier class and optimised for agentic and coding work. |
| Mistral Small 4 | Mistral AI | Open weight · workhorse | Apache 2.0 | Not listed on Mistral's docs page | Not listed | Not published on Mistral's docs page | Not published on Mistral's docs page | Unified instruct, reasoning and coding, with multimodal understanding, per secondary reporting | not present | not measured | not measured | 16 March 2026 per secondary reporting; version string v26.03 is consistent. |
Benchmark provenance
What each figure measures, who runs it, and its date.
Arena WebDev (arena.ai)
Human pairwise preference Elo on web-app building tasks. The closest public analogue to what our product boards measure.
Run by: arena.ai
Board dated 21 Aug 2026, 603,789 votes, 118 models. Some rows are flagged Preliminary by Arena on low vote counts; we carry that flag through.
Artificial Analysis Intelligence Index v4.1.1
A normalised composite of nine evaluations.
Run by: Artificial Analysis, independent
Artificial Analysis publishes no page-level revision date, and publishes separate rows per reasoning effort level.
Terminal-Bench v2.1
89 curated terminal tasks, pass@1 averaged over 3 repeats.
Run by: Governed by the Laude Institute with Stanford researchers and an open-source community; independently evaluated by Artificial Analysis
Always cited with the effort level alongside the score.
A pricing error we found on Arena, not on a vendor page
Arena publishes a price column beside its Elo. We do not use it. Arena lists GPT-5.6 Sol at 5 / 30 dollars per million tokens; OpenAI's own pricing page lists 4 / 20 for Sol and 5 / 30 for GPT-5.5, so Arena has carried the older model's price onto the newer one. It also lists Gemini 3.7 Flash output at 3.57 where Google publishes 3.75. We take Elo, vote counts, licence and the Preliminary flag from Arena, and we take every price, context window and output limit from the vendor. This is the same failure we have found on price-aggregator sites, and it is worth knowing that a benchmark operator is not immune to it.
Four caveats worth reading before you compare rows
- Reasoning effort changes both the score and the price, so one row per model name understates the variation. Claude Opus 5 is published by Artificial Analysis at three effort levels with three different scores, and it sits at two different Arena Elos. Where we hold per-effort figures we show them; where we do not, assume the number belongs to one specific effort setting we cannot name.
- Claude models from 4.7 onward use a tokenizer that produces roughly 30 percent more tokens for the same text than the earlier line. Any cross-vendor cost-per-task comparison that multiplies a price by a token count is wrong in Anthropic's favour on tokens and against it on price. We publish per-token prices and deliberately do not publish a cost-per-task figure.
- Superlatives are computed per row, never per vendor. A vendor's newest model is not reliably its widest or its cheapest, and this board contains counter-examples in both directions.
- Open weights have no vendor list price by construction. Where a model is open weights we publish the vendor's own hosted API price if it has one, and otherwise no price at all. A price from a third-party host is a host's price, not the model's.
Superseded, including everything this board used to show (12 rows)
To show the price ladder and to make our own previous board auditable, rather than quietly deleting it. Six of these rows are the exact eight we used to publish minus two duplicates already covered by a newer legacy row from the same vendor.
| Model | Vendor | Licence | Context | Output | In $/MTok | Out $/MTok | Arena WebDev | Note |
|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.7 | Anthropic | Proprietary | 1,000,000 | 128,000 | 5.00 | 25.00 | 1558rank 15 | Anthropic itself files this model under Legacy models and recommends migrating. Arena also lists a high-effort variant separately: 1557 Elo, rank 16. |
| Claude Opus 4.8 | Anthropic | Proprietary | 1,000,000 | 128,000 | 5.00 | 25.00 | 1539rank 19 | Anthropic itself files this model under Legacy models and recommends migrating. Arena also lists a high-effort variant separately: 1563 Elo, rank 14. |
| Claude Opus 4.6 | Anthropic | Proprietary | 1,000,000 | 128,000 | 5.00 | 25.00 | 1536rank 23 | Anthropic itself files this model under Legacy models and recommends migrating. Arena also lists a high-effort variant separately: 1546 Elo, rank 18. |
| Claude Sonnet 4.6 | Anthropic | Proprietary | 1,000,000 | 128,000 | 3.00 | 15.00 | 1522rank 25 | Anthropic itself files this model under Legacy models and recommends migrating. |
| Qwen3 Max | Alibaba | Proprietary | 262,144 as originally published; understates the real current window by roughly four times | - | 2.50 list, limited-time 50 percent discount | 7.50 list, limited-time 50 percent discount | 1517rank 30 | Was our number 6 on the old board, published here at 262,144 context. Arena and Alibaba both now identify the current build as the 3.7-Max release, which Alibaba lists at 2.50 / 7.50 list with a limited-time 50 percent discount. |
| GPT-5.5 | OpenAI | Proprietary | - | - | 5.00 (cached 0.50) | 30.00 | 1508rank 33 | This is the row whose price aggregators and Arena wrongly carried onto GPT-5.6 Sol (see the warning box above). |
| Claude Opus 4.5 | Anthropic | Proprietary | 200,000 | - | 5.00 | 25.00 | 1468rank 42 | Was our index leader at 89.10 under the old, single merged board. Still purchasable. Sits below Opus 4.8, 4.7 and 4.6, all three of which the vendor itself files as legacy, which is why the old board was three generations behind, not one. Arena also lists a high-effort (32k) variant separately: 1494 Elo, rank 35. |
| Gemini 3 Pro | Proprietary | 1,000,000 | - | 2.00 | 12.00 | 1438rank 52 | Was our number 3 on the old board, and the row our previous "widest context" callout used to rest on. | |
| Claude Sonnet 4.5 | Anthropic | Proprietary | 200,000 | - | 3.00 | 15.00 | 1386rank 74 | Was our number 4 on the old, single merged board. |
| DeepSeek V3.2 | DeepSeek | MIT | 128,000 as we published it | - | 0.28 | 0.42 | 1360rank 82 | Was our number 7 on the old board, listed on Arena as deepseek-v3.2-thinking. We published 0.28 / 0.42 for it, and 0.42 matches no DeepSeek rate, current or off-peak, that we can find on the vendor's pricing page. |
| GPT-5.1 | OpenAI | Proprietary | 400,000 | - | 1.25 | 10.00 | 1341rank 89 | Was our number 2 on the old, single merged board. |
| Grok 4 | xAI | Proprietary | 256,000 | - | 3.00 | 15.00 | not present | Was our number 5 on the old board. Not present in Arena's published WebDev list under this label. |
Why we still publish a model page at all
Almost every product in the app builder and coding agent indexes is a harness around a frontier model. When a model ships an update, the products built on it change behaviour without shipping anything themselves, so it is worth knowing what those models actually are, what they cost, and how they compare on a measurement someone else, not us, actually ran.
What this page is not: a benchmark we operate, a composite we compute, or a ranking of our own. Read the methodology page for the full coverage boundary, including which vendors and models we left out of this pass and why.