Leaderboard / coding agents / cycle 04
Coding agent leaderboard 2026
Coding agents
17
This board only, see /leaderboard/app-builders for the other
Prompt executions
765
Published individually at /api/runs.json
Index leader
Claude Code
Composite 79.72
Fastest median turn
31.4 s per prompt
Zed - 4.7 min across the full nine prompt specification, derived
Full measurement grid
Index, turn latency percentiles, outcome split and all nine subscores. No code ownership column: it is not scored on this board. Delta is not shown for cycle 04, see below.
| #Tool | Lab index | Reader index | Delta | Turn latency p10/med/p90 | p90 / median | p90 / p10 | Full spec (derived) | Reliability % | Runs pass/part/fail | Agent | Scal. | SEO | API | Integr. | Design | Value | HTML no JS | LCP ms | CLS | Entry price |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
79.72 | 85.23 | withheld | 39.9/65.5/100.7 | 1.54x | 2.52x | 9.8 min | 91.1% | 41/3/1 | 95.0 | 86.0 | 62.0 | 78.0 | 76.0 | 68.0 | 80.0 | 26.0 KB | 1,680 | 0.030 | USD 17 / month, billed annually | |
77.92 | 82.93 | withheld | 28.7/53.4/82.7 | 1.55x | 2.88x | 8.0 min | 86.7% | 39/3/3 | 92.0 | 82.0 | 58.0 | 74.0 | 80.0 | 66.0 | 80.0 | 24.0 KB | 1,720 | 0.030 | USD 20 / month | |
75.67 | 78.57 | withheld | 49.1/72.4/140.4 | 1.94x | 2.86x | 10.9 min | 93.3% | 42/2/1 | 84.0 | 80.0 | 60.0 | 72.0 | 74.0 | 65.0 | 78.0 | 22.0 KB | 1,800 | 0.040 | USD 20 / user / month | |
75.63 | 2 voting | withheld | 44.9/71.6/110.0 | 1.54x | 2.45x | 10.7 min | 88.9% | 40/3/2 | 86.0 | 84.0 | 58.0 | 74.0 | 76.0 | 62.0 | 77.0 | 22.0 KB | 1,800 | 0.040 | USD 20 / month, flat for a team up to 50 seats | |
73.35 | 2 voting | withheld | 24.2/39.7/64.4 | 1.62x | 2.66x | 6.0 min | 77.8% | 35/6/4 | 85.0 | 74.0 | 56.0 | 72.0 | 68.0 | 60.0 | 77.0 | 20.0 KB | 1,840 | 0.040 | USD 20 / month | |
71.48 | 64.41 | withheld | 84.8/161.7/350.8 | 2.17x | 4.14x | 24.3 min | 82.2% | 37/5/3 | 81.0 | 82.0 | 56.0 | 76.0 | 78.0 | 62.0 | 76.0 | 19.0 KB | 1,940 | 0.050 | USD 20 / month | |
71.37 | 1 voting | withheld | 45.0/78.7/123.1 | 1.56x | 2.74x | 11.8 min | 73.3% | 33/7/5 | 83.0 | 76.0 | 58.0 | 70.0 | 72.0 | 70.0 | 90.0 | 22.0 KB | 1,820 | 0.040 | No cost | |
71.13 | 79.79 | withheld | 19.8/31.4/54.5 | 1.74x | 2.75x | 4.7 min | 71.1% | 32/5/8 | 77.0 | 70.0 | 55.0 | 66.0 | 64.0 | 64.0 | 82.0 | 19.0 KB | 1,780 | 0.040 | USD 10 / month | |
70.66 | 77.04 | withheld | 47.4/95.1/201.7 | 2.12x | 4.26x | 14.3 min | 80.0% | 36/8/1 | 82.0 | 72.0 | 57.0 | 70.0 | 72.0 | 63.0 | 88.0 | 21.0 KB | 1,860 | 0.040 | No licence fee, bring your own model key | |
68.04 | 1 voting | withheld | 84.4/146.5/292.7 | 2.00x | 3.47x | 22.0 min | 75.6% | 34/7/4 | 80.0 | 74.0 | 55.0 | 66.0 | 70.0 | 62.0 | 88.0 | 21.0 KB | 1,860 | 0.040 | No cost on the free tier | |
66.64 | 2 voting | withheld | 230.2/353.6/651.4 | 1.84x | 2.83x | 53.0 min | 71.1% | 32/7/6 | 78.0 | 80.0 | 54.0 | 72.0 | 74.0 | 60.0 | 74.0 | 20.0 KB | 1,900 | 0.050 | USD 20 / month | |
66.06 | 2 voting | withheld | 66.3/99.3/189.8 | 1.91x | 2.86x | 14.9 min | 77.8% | 35/6/4 | 70.0 | 76.0 | 52.0 | 68.0 | 70.0 | 56.0 | 70.0 | 16.0 KB | 2,000 | 0.050 | USD 30 / month, pooled across a team of up to 30 users | |
65.88 | 71.33 | withheld | 22.5/39.4/75.7 | 1.92x | 3.36x | 5.9 min | 68.9% | 31/7/7 | 72.0 | 64.0 | 54.0 | 62.0 | 58.0 | 52.0 | 86.0 | 17.0 KB | 1,920 | 0.050 | No licence fee, bring your own model key | |
64.51 | 1 voting | withheld | 36.1/59.8/110.7 | 1.85x | 3.07x | 9.0 min | 68.9% | 31/7/7 | 76.0 | 68.0 | 55.0 | 64.0 | 62.0 | 58.0 | 52.0 | 18.0 KB | 1,900 | 0.050 | USD 100 / month minimum subscription | |
63.93 | 2 voting | withheld | 81.4/144.8/269.1 | 1.86x | 3.31x | 21.7 min | 64.4% | 29/9/7 | 74.0 | 70.0 | 56.0 | 68.0 | 66.0 | 58.0 | 86.0 | 18.0 KB | 1,980 | 0.050 | No licence fee, self hosted | |
61.00 | 2 voting | withheld | 38.2/96.7/167.9 | 1.74x | 4.40x | 14.5 min | 51.1% | 23/9/13 | 73.0 | 66.0 | 54.0 | 60.0 | 66.0 | 66.0 | 84.0 | 18.0 KB | 1,960 | 0.050 | USD 3 / month, billed monthly | |
59.89 | 1 voting | withheld | 54.3/96.4/202.9 | 2.10x | 3.74x | 14.5 min | 53.3% | 24/11/10 | 71.0 | 66.0 | 53.0 | 62.0 | 60.0 | 56.0 | 84.0 | 17.0 KB | 1,980 | 0.050 | No licence fee, self hosted up to 10 users |
Rank deltas are not shown for cycle 04 because the roster was re-based from a single merged ranking into two class boards, so this cycle's ranks and the previous cycle's are not on the same scale. Deltas resume in cycle 05.
Turn latency columns are median wall clock seconds for one prompt, timed from submission to the agent's returned diff or completion signal; there is no deployment step on this board. The full spec column derives what nine of those medians would add up to and is not separately measured. Runs are split pass, partial and fail across the 45 recorded executions per product. Reliability is the strict pass share. HTML no JS is real text bytes with scripts disabled, where the prompt produced a public page at all.
Read a single axis instead
The composite compresses nine different questions into one number. Each axis page shows both boards.
Turn latency
Median wall clock seconds for one prompt, not for a whole build. The test specification is nine prompts, run five times per product, so a single row here is the middle value of 45 recorded prompt executions.
Reliability
Strict pass rate across 765 recorded prompt executions.
SEO and GEO
Server rendered HTML bytes with JavaScript disabled, LCP, CLS and structured data output.
API and MCP
REST surface, webhook delivery, background jobs and Model Context Protocol support.
Design output
Blind scored visual quality of the generated admin panel and marketing surface.
Value
Index points delivered per euro of entry price. The two boards price different things, so the two value scores are not comparable across boards.
Notes on reading this table
Products within about one and a half index points of each other should be read as tied. Five runs per product gives a usable median and a rough sense of spread, not a tight confidence interval, and the honest thing to do with a small sample is to say so rather than to present a false ordering.
This table holds only the coding agents: products that edit a repository, terminal or IDE you already own and return a diff or a pull request rather than a deployed URL. Full-stack app builders, which own the whole pipeline from prompt to a hosted deployed application, are a separate board with their own ten axis weights at /leaderboard/app-builders. The two are never merged into one ordering on this site: an index score from one board is not comparable to an index score from the other.
Code ownership is not scored on this board because a coding agent edits a repository you already own, so the property was never the vendor's to grant and a full mark measures nothing. The other nine weights are rescaled proportionally so they still sum to 100; their relative importance to each other is unchanged.
Reliability at 91.1% at the top of the table means roughly one execution in eight still needed intervention or failed outright, for the best product in the cohort. That is the state of the field in August 2026, and it is the single most useful thing to know before planning a delivery date around one of these tools.