Leaderboard / coding agents / cycle 04

Coding agent leaderboard 2026

Every measurement we hold on 17 coding agents, in one grid. 5 runs of the 9 prompt vibeOps specification per product. You point this class at an existing repository, terminal or IDE and it changes the code; the output is a diff, not a deployment. The full-stack app builder board at /leaderboard/app-builders is a separate measurement with its own weights. Scroll sideways for the full axis set. The first column stays put.
Last run: 15 Aug 2026 09:11 UTCCadence: roughly every seven days, next date announced once it is scheduled

Coding agents

17

This board only, see /leaderboard/app-builders for the other

Prompt executions

765

Published individually at /api/runs.json

Index leader

Claude Code

Composite 79.72

Fastest median turn

31.4 s per prompt

Zed - 4.7 min across the full nine prompt specification, derived

Full measurement grid

Index, turn latency percentiles, outcome split and all nine subscores. No code ownership column: it is not scored on this board. Delta is not shown for cycle 04, see below.

Full Coding agents leaderboard with per axis subscores and raw measurements
#ToolLab indexReader indexDeltaTurn latency p10/med/p90p90 / medianp90 / p10Full spec (derived)Reliability %Runs pass/part/failAgentScal.SEOAPIIntegr.DesignValueHTML no JSLCP msCLSEntry price
79.72
85.23withheld
39.9/65.5/100.7
1.54x2.52x9.8 min91.1%
41/3/1
95.086.062.078.076.068.080.026.0 KB1,6800.030USD 17 / month, billed annually
77.92
82.93withheld
28.7/53.4/82.7
1.55x2.88x8.0 min86.7%
39/3/3
92.082.058.074.080.066.080.024.0 KB1,7200.030USD 20 / month
75.67
78.57withheld
49.1/72.4/140.4
1.94x2.86x10.9 min93.3%
42/2/1
84.080.060.072.074.065.078.022.0 KB1,8000.040USD 20 / user / month
75.63
2 votingwithheld
44.9/71.6/110.0
1.54x2.45x10.7 min88.9%
40/3/2
86.084.058.074.076.062.077.022.0 KB1,8000.040USD 20 / month, flat for a team up to 50 seats
73.35
2 votingwithheld
24.2/39.7/64.4
1.62x2.66x6.0 min77.8%
35/6/4
85.074.056.072.068.060.077.020.0 KB1,8400.040USD 20 / month
71.48
64.41withheld
84.8/161.7/350.8
2.17x4.14x24.3 min82.2%
37/5/3
81.082.056.076.078.062.076.019.0 KB1,9400.050USD 20 / month
71.37
1 votingwithheld
45.0/78.7/123.1
1.56x2.74x11.8 min73.3%
33/7/5
83.076.058.070.072.070.090.022.0 KB1,8200.040No cost
71.13
79.79withheld
19.8/31.4/54.5
1.74x2.75x4.7 min71.1%
32/5/8
77.070.055.066.064.064.082.019.0 KB1,7800.040USD 10 / month
70.66
77.04withheld
47.4/95.1/201.7
2.12x4.26x14.3 min80.0%
36/8/1
82.072.057.070.072.063.088.021.0 KB1,8600.040No licence fee, bring your own model key
68.04
1 votingwithheld
84.4/146.5/292.7
2.00x3.47x22.0 min75.6%
34/7/4
80.074.055.066.070.062.088.021.0 KB1,8600.040No cost on the free tier
66.64
2 votingwithheld
230.2/353.6/651.4
1.84x2.83x53.0 min71.1%
32/7/6
78.080.054.072.074.060.074.020.0 KB1,9000.050USD 20 / month
66.06
2 votingwithheld
66.3/99.3/189.8
1.91x2.86x14.9 min77.8%
35/6/4
70.076.052.068.070.056.070.016.0 KB2,0000.050USD 30 / month, pooled across a team of up to 30 users
65.88
71.33withheld
22.5/39.4/75.7
1.92x3.36x5.9 min68.9%
31/7/7
72.064.054.062.058.052.086.017.0 KB1,9200.050No licence fee, bring your own model key
64.51
1 votingwithheld
36.1/59.8/110.7
1.85x3.07x9.0 min68.9%
31/7/7
76.068.055.064.062.058.052.018.0 KB1,9000.050USD 100 / month minimum subscription
63.93
2 votingwithheld
81.4/144.8/269.1
1.86x3.31x21.7 min64.4%
29/9/7
74.070.056.068.066.058.086.018.0 KB1,9800.050No licence fee, self hosted
61.00
2 votingwithheld
38.2/96.7/167.9
1.74x4.40x14.5 min51.1%
23/9/13
73.066.054.060.066.066.084.018.0 KB1,9600.050USD 3 / month, billed monthly
59.89
1 votingwithheld
54.3/96.4/202.9
2.10x3.74x14.5 min53.3%
24/11/10
71.066.053.062.060.056.084.017.0 KB1,9800.050No licence fee, self hosted up to 10 users

Rank deltas are not shown for cycle 04 because the roster was re-based from a single merged ranking into two class boards, so this cycle's ranks and the previous cycle's are not on the same scale. Deltas resume in cycle 05.

Turn latency columns are median wall clock seconds for one prompt, timed from submission to the agent's returned diff or completion signal; there is no deployment step on this board. The full spec column derives what nine of those medians would add up to and is not separately measured. Runs are split pass, partial and fail across the 45 recorded executions per product. Reliability is the strict pass share. HTML no JS is real text bytes with scripts disabled, where the prompt produced a public page at all.

Read a single axis instead

The composite compresses nine different questions into one number. Each axis page shows both boards.

Notes on reading this table

Products within about one and a half index points of each other should be read as tied. Five runs per product gives a usable median and a rough sense of spread, not a tight confidence interval, and the honest thing to do with a small sample is to say so rather than to present a false ordering.

This table holds only the coding agents: products that edit a repository, terminal or IDE you already own and return a diff or a pull request rather than a deployed URL. Full-stack app builders, which own the whole pipeline from prompt to a hosted deployed application, are a separate board with their own ten axis weights at /leaderboard/app-builders. The two are never merged into one ordering on this site: an index score from one board is not comparable to an index score from the other.

Code ownership is not scored on this board because a coding agent edits a repository you already own, so the property was never the vendor's to grant and a full mark measures nothing. The other nine weights are rescaled proportionally so they still sum to 100; their relative importance to each other is unchanged.

Reliability at 91.1% at the top of the table means roughly one execution in eight still needed intervention or failed outright, for the best product in the cohort. That is the state of the field in August 2026, and it is the single most useful thing to know before planning a delivery date around one of these tools.