Leaderboard / large language models

LLM reference table

The models underneath most of the products in the app builder index. A sourced specification and benchmark table, not an LLM Tier ranking.

We do not run a language-model harness. Every benchmark figure on this page is a third party's measurement, named and dated, and every specification and price is the vendor's own published figure. Where a vendor does not publish a number we leave the cell empty and say so rather than sourcing it from an aggregator. This board is a sourced reference table, not an LLM Tier ranking. It has no composite index because we have not earned the right to publish one: we hold no model-level run records, so we could not show you the arithmetic.

The model layer and the product boards are separate measurements. A product built on a strong model does not inherit that model's standing here, and neither product board's composite is derived from anything on this page.

Top of Arena WebDev

1691 Elo

Claude Opus 5 (max effort), 1691 Elo, Arena, 21 Aug 2026.

Best independently measured agentic coding score

89.5%

GPT-5.6 Sol at xhigh effort, 89.5 percent pass@1 on Terminal-Bench v2.1, ahead of Claude Opus 5 at max effort on 89.1 percent and Grok 4.6 at high effort on 88.4 percent.

Widest published context window

1,048,576 tokens

1,048,576 tokens, published by Moonshot for Kimi K3. Several models are documented at "1,000,000"; we print the largest figure a vendor actually states.

Cheapest published output token

No single answer

This has no single answer. DeepSeek-V4-Flash is 0.66 dollars per million output tokens off-peak and 1.32 at peak. GPT-5.6 Luna is a flat 1.20. So DeepSeek is cheapest for most of the day and Luna is cheapest during DeepSeek's peak window, which runs 01:00 to 04:00 and 06:00 to 10:00 UTC. A cheapest-output crown that flips on the clock is not a crown, so we publish both.

26 models, sourced roster

Sorted by Arena WebDev Elo, descending. Rows with no published Elo sort last, alphabetically by vendor. That ordering is Arena's measurement, not ours.

LLM Tier sourced specification and benchmark table
ModelVendorStatus / tierLicenceContext windowMax outputPrice in $/MTokPrice out $/MTokModalityArena WebDevAA index v4.1.1Terminal-Bench v2.1Release note
Claude Opus 5AnthropicCurrent · flagshipProprietary1,000,000128,0005.0025.00text and image in, text out
1691rank 1, 8,116 votes
63 at max effort, 63 at xhigh, 61 at high89.1% at max effortRelease date not stated by the vendor; training data cutoff May 2026. Arena also lists the high-effort variant separately: 1663 Elo, rank 4, 8659 votes.
Kimi K3Moonshot AICurrent · flagshipModified MIT, published by Moonshot as the Kimi K3 License; open weights1,048,576Not stated on the vendor pricing page3.00 on a cache miss, 0.30 on a cache hit15.00Multimodal reasoning with thinking always on, per secondary reporting
1674rank 2, 4,546 votes
not measurednot measuredReleased 16 July 2026; full open weights 27 July 2026. Kimi K3 always reasons and exposes a reasoning_effort setting of low, high or max, defaulting to max, which is why Arena lists it as kimi-k3-max (max effort). A third-party host lists the same weights at 2.60 / 13.00. That is the host's price, not Moonshot's, and we publish Moonshot's.
Qwen3.8-MaxAlibabaCurrent · flagshipProprietary1,048,576, charged at one flat rate across the whole window with no length tiering131,0002.00 (batch inference 1.65; context-caching discount available)6.00, covering chain of thought plus answer (batch inference 4.951)text, image and video in, text out
1669Prelimrank 3, 3,212 votes
not measurednot measuredPreviewed 19 July 2026, global API access via Alibaba Cloud Model Studio 3 August 2026. There is no independently measured coding-benchmark score for this model, so we publish none, and its Arena Elo carries Arena's own Preliminary flag on 3,212 votes. It is also the only flagship on this board with no 200k price step, which makes it unusually cheap on very long prompts.
Grok 4.6xAICurrent · flagshipProprietary500,000No stated text output limit2.00 below 200k prompt tokens, 4.00 above (cached 0.50 / 1.00)6.00 below 200k prompt tokens, 12.00 aboveThe xAI models page lists text in and text out; secondary reporting says image input is also supported, which we have not confirmed
1629Prelimrank 5, 1,523 votes
61 per secondary reporting of the index88.4% at high effort12 August 2026 per secondary reporting; not dated on xAI's models page. Arena files xAI's models under the lab name SpaceXAI; the vendor documentation domain is docs.x.ai and we publish the vendor as xAI. See Grok 4.3 below for the headline counter-example on this board.
Claude Fable 5AnthropicCurrent · flagshipProprietary1,000,000128,000 (300,000 on the Batch API with the output-300k-2026-03-24 beta header)10.0050.00text and image in, text out
1626rank 6, 8,382 votes
62 (Adaptive Reasoning, Max Effort)Not in the published top threeAvailable from 9 June 2026.
GPT-5.6 SolOpenAICurrent · flagshipProprietaryNot published on OpenAI's pricing pageNot published on OpenAI's pricing page4.00 (cached input 0.40)20.00text in, text out per the pricing table
1619rank 7, 9,499 votes
61 at max effort89.5% at xhigh effort, the published leaderLimited preview 26 June 2026, public launch 9 July 2026. Arena entry uses xhigh effort in the Codex harness. See the pricing warning box above: Arena lists this model at the older GPT-5.5 price.
Gemini 3.7 FlashGoogleCurrent · workhorseProprietaryNot stated on the fetched pagesNot stated0.75 through 31 December 2026, then 1.50 from 1 January 20273.75 through 31 December 2026, then 7.50 from 1 January 2027Not stated per model on the fetched page
1587Prelimrank 10, 2,552 votes
not measurednot measured13 August 2026 per secondary reporting; Google's own pages mark it stable but do not date it. Arena entry uses high effort and carries Arena's Preliminary flag. Google publishes DeepSWE v1.1 rising from 49.0 to 65.3 percent and GDM-MRCR v2 at 128k rising from 91.8 to 97.0 percent versus the prior Flash. Those are Google's own figures on Google's own evaluation and are labelled vendor-claimed, not independent.
DeepSeek-V4-ProDeepSeekCurrent · flagshipMIT1,000,000Not stated on the pricing page0.66 off-peak, 1.32 at peak (cache hit 0.022 off-peak, 0.044 peak)1.98 off-peak, 3.96 at peaktext in, text out; a separate deepseek-v4-flash-vision-exp handles vision
1582rank 12, 2,869 votes
not measurednot measuredVersion string DeepSeek-V4-Pro-0813 implies 13 August 2026; not dated in prose. Arena entry uses high effort on the 20260813 build. DeepSeek has no single price: off-peak is half of peak, and peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC. Any value comparison must state which rate it used.
DeepSeek-V4-FlashDeepSeekCurrent · budgetMIT1,000,000Not stated0.22 off-peak, 0.44 at peak (cache hit 0.007 / 0.014)0.66 off-peak, 1.32 at peaktext in, text out
1579rank 13, 3,537 votes
not measurednot measuredVersion string DeepSeek-V4-Flash-0731 implies 31 July 2026. Arena entry uses high effort. DeepSeek has no single price: off-peak is half of peak, and peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC. Any value comparison must state which rate it used. See the cheapest-output callout above for why this row and GPT-5.6 Luna cannot both hold that crown.
Claude Sonnet 5AnthropicCurrent · workhorseProprietary1,000,000128,0002.0010.00text and image in, text out
1539rank 22, 7,267 votes
not measurednot measuredRelease date not stated by the vendor. Arena effort level: high.
Gemini 3.6 FlashGoogleCurrent · workhorseProprietaryNot statedNot stated0.75, rising to 1.50 on 1 January 20273.75, rising to 7.50 on 1 January 2027Not stated
1539rank 21, 6,464 votes
not measurednot measuredNot dated on Google's pages; marked stable. Arena entry uses high effort.
Muse Spark 1.1MetaCurrent · flagshipProprietary, closed weights1,000,000Not specified in Meta's announcementNot published by MetaNot published by Metatext, images, video and PDFs in, text out
1539rank 20, 6,204 votes
not measurednot measuredDeveloped by Meta Superintelligence Labs. Muse Spark launched 8 April 2026 as Meta's first closed-weight model; 1.1 released 9 July 2026. Meta's own announcement publishes no per-token price and no numerical benchmark score, so both cells are empty here. Price is the single most-quoted number in secondary coverage and the single number Meta does not state. Arena also lists a muse-spark-1.2 at 1534 Elo; we have not sourced 1.2 from Meta and so do not list it as a row.
GPT-5.6 TerraOpenAICurrent · workhorseProprietaryNot publishedNot published2.00 (cached input 0.20)12.00text in, text out
1520rank 27, 5,677 votes
not measurednot measuredLaunched 9 July 2026. Arena entry uses xhigh effort in the Codex harness.
GPT-5.6 LunaOpenAICurrent · budgetProprietaryNot publishedNot published0.20 (cached input 0.02)1.20text in, text out
1518rank 29, 5,742 votes
not measurednot measuredLaunched 9 July 2026. Arena entry uses xhigh effort in the Codex harness. Flat 1.20 output price. See the DeepSeek-V4-Flash row for the peak and off-peak comparison, since neither is cheapest all day.
Gemini 3.5 FlashGoogleCurrent · workhorseProprietaryNot statedNot stated1.509.00Not stated
1499rank 34, 7,559 votes
not measurednot measuredNot dated; marked stable. Arena entry uses high effort. This middle rung is worse value than both newer Flash models, which cost 0.75 / 3.75. Inside one vendor's line, newer here is both better and cheaper.
Gemini 3.5 Flash-LiteGoogleCurrent · budgetProprietaryNot statedNot stated0.302.50Not stated
1449rank 47, 228 votes
not measurednot measuredNot dated; marked stable. Arena's interval on this row is plus or minus 43 points on only 228 votes (1449 ± 43), far too wide to rank against its neighbours.
Gemini 3.1 Pro PreviewGooglePreview · flagshipProprietaryNot stated on Google's models or pricing pagesNot stated2.00 for prompts up to 200k tokens, 4.00 above 200k (batch 1.00 / 2.00)12.00 up to 200k, 18.00 above 200k (batch 6.00 / 9.00)Not stated per model on the fetched page
1446rank 48, 21,573 votes
not measurednot measuredStill carries preview status. This is the only Pro-tier Gemini 3 text model on Google's own list; every stable Gemini 3 text model is a Flash or Flash-Lite.
GPT-5.3 CodexOpenAICurrent · codingProprietaryNot publishedNot published1.75 (cached input 0.175)14.00text in, text out, coding specialised
1408rank 61, 2,526 votes
not measurednot measuredRelease date not published. Arena entry runs in the Codex harness.
Grok 4.3xAICurrent · workhorseProprietary1,000,000Not stated1.25 below 200k prompt tokens, 2.50 above (cached 0.20 / 0.40)2.50 below 200k, 5.00 abovetext in, text out
1357rank 86, 12,942 votes
not measurednot measuredNot dated on xAI's models page. HEADLINE COUNTER-EXAMPLE ON THIS BOARD: Grok 4.3 has double the context window of xAI's newer flagship (Grok 4.6) at roughly 40 percent of its output price, and sits 272 Elo points lower on Arena WebDev. Newer, wider and cheaper are three different axes and no vendor wins all three at once.
Amazon Nova 2 LiteAmazonCurrent · workhorseProprietary1,000,00065,536Not published on the Nova 2 user guide, which defers to Amazon Bedrock Pricing; the Bedrock pricing page carries no Nova 2 rowNot publishedtext, images, video and documents in, text outnot presentnot measurednot measuredRelease date not stated on the Nova 2 user guide. Nova 2 is Amazon's current generation and its whole roster is Nova 2 Lite, Nova 2 Sonic and Nova Multimodal Embeddings. Nova Premier, Nova Pro, Nova Lite and Nova Micro are all the superseded Nova 1 generation. Every Nova 2 model is 1,000,000 tokens in and 65,536 out.
Claude Haiku 4.5AnthropicCurrent · budgetProprietary200,00064,0001.005.00text and image in, text outnot presentnot measurednot measuredModel ID pins the snapshot to 1 October 2025.
Command A+CohereCurrent · flagshipApache 2.0Not published for this modelNot publishedNot published. cohere.com/pricing carries only legacy Command rows and has no Command A+ row at all.Not publishedtext and image in, text outnot presentnot measurednot measured20 May 2026. 218 billion total parameters with 25 billion active in a sparse mixture of experts, 48 languages, native citation generation. Cohere's flagship is three months old and its own public pricing page does not list it. We publish the model with no price rather than sourcing one from an aggregator.
Gemini 3.1 Flash-LiteGoogleCurrent · budgetProprietaryNot statedNot stated0.25 for text, image and video; 0.50 for audio1.50text, image, video and audio in per the pricing tiersnot presentnot measurednot measuredNot dated; marked stable. Google's own model list describes this as frontier class; we attribute that claim to Google and do not restate it as fact.
Llama 4 MaverickMetaOpen weight · workhorseOpen weightsPreviously published here as 1,000,000; not re-verified in this passNot re-verifiedNo vendor list price existsNo vendor list price existsNot re-verifiednot presentnot measurednot measured5 April 2025, alongside Llama 4 Scout. No new Llama model has shipped since. We previously published 0.22 / 0.85 for this model. That figure was unsourceable by construction, not merely stale: these are open weights and there is no vendor list price at all, only per-host prices. We have removed it. Llama is no longer Meta's frontier: Meta's frontier since April 2026 is Muse Spark, which is closed and paid.
Mistral Medium 3.5Mistral AICurrent · flagshipModified MITNot listed on Mistral's models overview pageNot listedNot published on Mistral's docs pageNot published on Mistral's docs pageMultimodal per Mistral's own descriptionnot presentnot measurednot measuredVersion string v26.04 implies April 2026; not dated in prose. Mistral's current premier flagship, described by Mistral as frontier class and optimised for agentic and coding work.
Mistral Small 4Mistral AIOpen weight · workhorseApache 2.0Not listed on Mistral's docs pageNot listedNot published on Mistral's docs pageNot published on Mistral's docs pageUnified instruct, reasoning and coding, with multimodal understanding, per secondary reportingnot presentnot measurednot measured16 March 2026 per secondary reporting; version string v26.03 is consistent.

Benchmark provenance

What each figure measures, who runs it, and its date.

Arena WebDev (arena.ai)

Human pairwise preference Elo on web-app building tasks. The closest public analogue to what our product boards measure.

Run by: arena.ai

Board dated 21 Aug 2026, 603,789 votes, 118 models. Some rows are flagged Preliminary by Arena on low vote counts; we carry that flag through.

Artificial Analysis Intelligence Index v4.1.1

A normalised composite of nine evaluations.

Run by: Artificial Analysis, independent

Artificial Analysis publishes no page-level revision date, and publishes separate rows per reasoning effort level.

Terminal-Bench v2.1

89 curated terminal tasks, pass@1 averaged over 3 repeats.

Run by: Governed by the Laude Institute with Stanford researchers and an open-source community; independently evaluated by Artificial Analysis

Always cited with the effort level alongside the score.

A pricing error we found on Arena, not on a vendor page

Arena publishes a price column beside its Elo. We do not use it. Arena lists GPT-5.6 Sol at 5 / 30 dollars per million tokens; OpenAI's own pricing page lists 4 / 20 for Sol and 5 / 30 for GPT-5.5, so Arena has carried the older model's price onto the newer one. It also lists Gemini 3.7 Flash output at 3.57 where Google publishes 3.75. We take Elo, vote counts, licence and the Preliminary flag from Arena, and we take every price, context window and output limit from the vendor. This is the same failure we have found on price-aggregator sites, and it is worth knowing that a benchmark operator is not immune to it.

Four caveats worth reading before you compare rows

  • Reasoning effort changes both the score and the price, so one row per model name understates the variation. Claude Opus 5 is published by Artificial Analysis at three effort levels with three different scores, and it sits at two different Arena Elos. Where we hold per-effort figures we show them; where we do not, assume the number belongs to one specific effort setting we cannot name.
  • Claude models from 4.7 onward use a tokenizer that produces roughly 30 percent more tokens for the same text than the earlier line. Any cross-vendor cost-per-task comparison that multiplies a price by a token count is wrong in Anthropic's favour on tokens and against it on price. We publish per-token prices and deliberately do not publish a cost-per-task figure.
  • Superlatives are computed per row, never per vendor. A vendor's newest model is not reliably its widest or its cheapest, and this board contains counter-examples in both directions.
  • Open weights have no vendor list price by construction. Where a model is open weights we publish the vendor's own hosted API price if it has one, and otherwise no price at all. A price from a third-party host is a host's price, not the model's.

Superseded, including everything this board used to show (12 rows)

To show the price ladder and to make our own previous board auditable, rather than quietly deleting it. Six of these rows are the exact eight we used to publish minus two duplicates already covered by a newer legacy row from the same vendor.

Superseded and legacy language models
ModelVendorLicenceContextOutputIn $/MTokOut $/MTokArena WebDevNote
Claude Opus 4.7AnthropicProprietary1,000,000128,0005.0025.00
1558rank 15
Anthropic itself files this model under Legacy models and recommends migrating. Arena also lists a high-effort variant separately: 1557 Elo, rank 16.
Claude Opus 4.8AnthropicProprietary1,000,000128,0005.0025.00
1539rank 19
Anthropic itself files this model under Legacy models and recommends migrating. Arena also lists a high-effort variant separately: 1563 Elo, rank 14.
Claude Opus 4.6AnthropicProprietary1,000,000128,0005.0025.00
1536rank 23
Anthropic itself files this model under Legacy models and recommends migrating. Arena also lists a high-effort variant separately: 1546 Elo, rank 18.
Claude Sonnet 4.6AnthropicProprietary1,000,000128,0003.0015.00
1522rank 25
Anthropic itself files this model under Legacy models and recommends migrating.
Qwen3 MaxAlibabaProprietary262,144 as originally published; understates the real current window by roughly four times-2.50 list, limited-time 50 percent discount7.50 list, limited-time 50 percent discount
1517rank 30
Was our number 6 on the old board, published here at 262,144 context. Arena and Alibaba both now identify the current build as the 3.7-Max release, which Alibaba lists at 2.50 / 7.50 list with a limited-time 50 percent discount.
GPT-5.5OpenAIProprietary--5.00 (cached 0.50)30.00
1508rank 33
This is the row whose price aggregators and Arena wrongly carried onto GPT-5.6 Sol (see the warning box above).
Claude Opus 4.5AnthropicProprietary200,000-5.0025.00
1468rank 42
Was our index leader at 89.10 under the old, single merged board. Still purchasable. Sits below Opus 4.8, 4.7 and 4.6, all three of which the vendor itself files as legacy, which is why the old board was three generations behind, not one. Arena also lists a high-effort (32k) variant separately: 1494 Elo, rank 35.
Gemini 3 ProGoogleProprietary1,000,000-2.0012.00
1438rank 52
Was our number 3 on the old board, and the row our previous "widest context" callout used to rest on.
Claude Sonnet 4.5AnthropicProprietary200,000-3.0015.00
1386rank 74
Was our number 4 on the old, single merged board.
DeepSeek V3.2DeepSeekMIT128,000 as we published it-0.280.42
1360rank 82
Was our number 7 on the old board, listed on Arena as deepseek-v3.2-thinking. We published 0.28 / 0.42 for it, and 0.42 matches no DeepSeek rate, current or off-peak, that we can find on the vendor's pricing page.
GPT-5.1OpenAIProprietary400,000-1.2510.00
1341rank 89
Was our number 2 on the old, single merged board.
Grok 4xAIProprietary256,000-3.0015.00not presentWas our number 5 on the old board. Not present in Arena's published WebDev list under this label.

Why we still publish a model page at all

Almost every product in the app builder and coding agent indexes is a harness around a frontier model. When a model ships an update, the products built on it change behaviour without shipping anything themselves, so it is worth knowing what those models actually are, what they cost, and how they compare on a measurement someone else, not us, actually ran.

What this page is not: a benchmark we operate, a composite we compute, or a ranking of our own. Read the methodology page for the full coverage boundary, including which vendors and models we left out of this pass and why.