Axis leaderboard / Reliability / two boards

Reliability

Strict pass rate across 1,665 recorded prompt executions. Shown as two separate tables, one per board: an index score, subscore or rank on one board is never comparable to the other.

Full-stack app builders: ranked by reliability

Ordered by axis subscore, highest first, within this board's own cohort.

Reliability, Full-stack app builders
#ToolReliabilityPass / partial / failExecutionsComposite
93.3
42/2/1
4589.48
2Firebase Studio logoFirebase StudioSunsets Mar 2027
86.7
39/5/1
4580.92
86.7
39/4/2
4575.47
84.4
38/6/1
4578.31
84.4
38/7/0
4573.05
82.2
37/6/2
4580.54
80.0
36/8/1
4574.68
80.0
36/3/6
4568.68
77.8
35/8/2
4569.97
73.3
33/7/5
4569.48
68.9
31/8/6
4568.66
66.7
30/12/3
4562.08
64.4
29/7/9
4565.59
60.0
27/7/11
4562.12
60.0
27/8/10
4561.06
55.6
25/10/10
4567.49
55.6
25/10/10
4559.85
51.1
23/11/11
4558.11
48.9
22/9/14
4558.53
48.9
22/10/13
4556.99

Coding agents: ranked by reliability

Ordered by axis subscore, highest first, within this board's own cohort.

Reliability, Coding agents
#ToolReliabilityPass / partial / failExecutionsComposite
93.3
42/2/1
4575.67
91.1
41/3/1
4579.72
88.9
40/3/2
4575.63
86.7
39/3/3
4577.92
82.2
37/5/3
4571.48
80.0
36/8/1
4570.66
77.8
35/6/4
4573.35
77.8
35/6/4
4566.06
75.6
34/7/4
4568.04
73.3
33/7/5
4571.37
11Zed logoZed
71.1
32/5/8
4571.13
71.1
32/7/6
4566.64
68.9
31/7/7
4565.88
68.9
31/7/7
4564.51
64.4
29/9/7
4563.93
53.3
24/11/10
4559.89
51.1
23/9/13
4561.00

How the reliability axis is measured

Full-stack app builders. Strict pass rate over all 45 recorded prompt executions per product. Partial completions and outright failures both cost. Executions running past 20 minutes are flagged for manual review and scored on the outcome the reviewer recorded; the duration is always retained and always counts toward the turn latency percentiles.

Coding agents. Strict pass rate over all 45 recorded prompt executions per product. Partial completions and outright failures both cost. Executions running past 20 minutes are flagged for manual review and scored on the outcome the reviewer recorded; the duration is always retained and always counts toward the turn latency percentiles.

This axis carries 18 of the 100 index points on the builders board and 18.56 of the 100 index points on the coding agent board. Each product completed 5 runs of 9 prompts in this cycle, and the underlying per run data behind every figure on this page is downloadable at /api/runs.json under a CC BY 4.0 licence.

A high score on one axis is not a recommendation. Read it against the composite column on the right and against the axes that matter for the work you have in front of you. Reliability at 93.3% for the leader of the builders cohort is the context for everything else on the page.

Other axis leaderboards

Reliability (2026 measurements) | LLM Tier