Axis leaderboard / Reliability / two boards
Reliability
Full-stack app builders: ranked by reliability
Ordered by axis subscore, highest first, within this board's own cohort.
| #Tool | Reliability | Pass / partial / fail | Executions | Composite |
|---|---|---|---|---|
93.3 | 42/2/1 | 45 | 89.48 | |
86.7 | 39/5/1 | 45 | 80.92 | |
86.7 | 39/4/2 | 45 | 75.47 | |
84.4 | 38/6/1 | 45 | 78.31 | |
84.4 | 38/7/0 | 45 | 73.05 | |
82.2 | 37/6/2 | 45 | 80.54 | |
80.0 | 36/8/1 | 45 | 74.68 | |
80.0 | 36/3/6 | 45 | 68.68 | |
77.8 | 35/8/2 | 45 | 69.97 | |
73.3 | 33/7/5 | 45 | 69.48 | |
68.9 | 31/8/6 | 45 | 68.66 | |
66.7 | 30/12/3 | 45 | 62.08 | |
64.4 | 29/7/9 | 45 | 65.59 | |
60.0 | 27/7/11 | 45 | 62.12 | |
60.0 | 27/8/10 | 45 | 61.06 | |
55.6 | 25/10/10 | 45 | 67.49 | |
55.6 | 25/10/10 | 45 | 59.85 | |
51.1 | 23/11/11 | 45 | 58.11 | |
48.9 | 22/9/14 | 45 | 58.53 | |
48.9 | 22/10/13 | 45 | 56.99 |
Coding agents: ranked by reliability
Ordered by axis subscore, highest first, within this board's own cohort.
| #Tool | Reliability | Pass / partial / fail | Executions | Composite |
|---|---|---|---|---|
93.3 | 42/2/1 | 45 | 75.67 | |
91.1 | 41/3/1 | 45 | 79.72 | |
88.9 | 40/3/2 | 45 | 75.63 | |
86.7 | 39/3/3 | 45 | 77.92 | |
82.2 | 37/5/3 | 45 | 71.48 | |
80.0 | 36/8/1 | 45 | 70.66 | |
77.8 | 35/6/4 | 45 | 73.35 | |
77.8 | 35/6/4 | 45 | 66.06 | |
75.6 | 34/7/4 | 45 | 68.04 | |
73.3 | 33/7/5 | 45 | 71.37 | |
71.1 | 32/5/8 | 45 | 71.13 | |
71.1 | 32/7/6 | 45 | 66.64 | |
68.9 | 31/7/7 | 45 | 65.88 | |
68.9 | 31/7/7 | 45 | 64.51 | |
64.4 | 29/9/7 | 45 | 63.93 | |
53.3 | 24/11/10 | 45 | 59.89 | |
51.1 | 23/9/13 | 45 | 61.00 |
How the reliability axis is measured
Full-stack app builders. Strict pass rate over all 45 recorded prompt executions per product. Partial completions and outright failures both cost. Executions running past 20 minutes are flagged for manual review and scored on the outcome the reviewer recorded; the duration is always retained and always counts toward the turn latency percentiles.
Coding agents. Strict pass rate over all 45 recorded prompt executions per product. Partial completions and outright failures both cost. Executions running past 20 minutes are flagged for manual review and scored on the outcome the reviewer recorded; the duration is always retained and always counts toward the turn latency percentiles.
This axis carries 18 of the 100 index points on the builders board and 18.56 of the 100 index points on the coding agent board. Each product completed 5 runs of 9 prompts in this cycle, and the underlying per run data behind every figure on this page is downloadable at /api/runs.json under a CC BY 4.0 licence.
A high score on one axis is not a recommendation. Read it against the composite column on the right and against the axes that matter for the work you have in front of you. Reliability at 93.3% for the leader of the builders cohort is the context for everything else on the page.