Axis leaderboard / Turn latency score / two boards
Turn latency
Full-stack app builders: ranked by turn latency score
Ordered by median seconds per prompt, fastest first, within this board's own cohort.
| #Tool | Turn latency score | p10 / median / p90 | p90 / median | p90 / p10 | Full spec (derived) | Composite |
|---|---|---|---|---|---|---|
100.0 | 44.9/65.2/104.1 | 1.60x | 2.32x | 9.8 min | 78.31 | |
81.5 | 46.3/80.0/143.0 | 1.79x | 3.09x | 12.0 min | 58.53 | |
79.4 | 51.8/82.1/136.0 | 1.66x | 2.63x | 12.3 min | 74.68 | |
77.8 | 54.4/83.8/149.9 | 1.79x | 2.76x | 12.6 min | 67.49 | |
76.5 | 52.6/85.2/133.1 | 1.56x | 2.53x | 12.8 min | 80.54 | |
74.9 | 49.4/87.0/137.1 | 1.58x | 2.78x | 13.1 min | 58.11 | |
71.6 | 46.1/91.1/174.2 | 1.91x | 3.78x | 13.7 min | 68.66 | |
69.3 | 58.8/94.1/161.4 | 1.72x | 2.74x | 14.1 min | 73.05 | |
69.3 | 55.2/94.1/146.8 | 1.56x | 2.66x | 14.1 min | 69.97 | |
67.9 | 53.5/96.0/186.6 | 1.94x | 3.49x | 14.4 min | 65.59 | |
65.0 | 65.8/100.3/161.3 | 1.61x | 2.45x | 15.0 min | 80.92 | |
61.3 | 59.2/106.3/167.6 | 1.58x | 2.83x | 15.9 min | 89.48 | |
60.8 | 53.3/107.2/217.0 | 2.02x | 4.07x | 16.1 min | 62.12 | |
60.1 | 65.1/108.4/176.3 | 1.63x | 2.71x | 16.3 min | 69.48 | |
57.5 | 60.2/113.4/193.1 | 1.70x | 3.21x | 17.0 min | 61.06 | |
57.3 | 59.6/113.8/231.5 | 2.03x | 3.88x | 17.1 min | 59.85 | |
48.0 | 92.8/135.7/188.8 | 1.39x | 2.03x | 20.4 min | 75.47 | |
47.6 | 72.2/137.1/249.0 | 1.82x | 3.45x | 20.6 min | 56.99 | |
45.1 | 81.7/144.6/309.3 | 2.14x | 3.79x | 21.7 min | 68.68 | |
43.3 | 76.0/150.6/295.5 | 1.96x | 3.89x | 22.6 min | 62.08 |
Coding agents: ranked by turn latency score
Ordered by median seconds per prompt, fastest first, within this board's own cohort.
| #Tool | Turn latency score | p10 / median / p90 | p90 / median | p90 / p10 | Full spec (derived) | Composite |
|---|---|---|---|---|---|---|
100.0 | 19.8/31.4/54.5 | 1.74x | 2.75x | 4.7 min | 71.13 | |
79.7 | 22.5/39.4/75.7 | 1.92x | 3.36x | 5.9 min | 65.88 | |
79.1 | 24.2/39.7/64.4 | 1.62x | 2.66x | 6.0 min | 73.35 | |
58.8 | 28.7/53.4/82.7 | 1.55x | 2.88x | 8.0 min | 77.92 | |
52.5 | 36.1/59.8/110.7 | 1.85x | 3.07x | 9.0 min | 64.51 | |
47.9 | 39.9/65.5/100.7 | 1.54x | 2.52x | 9.8 min | 79.72 | |
43.9 | 44.9/71.6/110.0 | 1.54x | 2.45x | 10.7 min | 75.63 | |
43.4 | 49.1/72.4/140.4 | 1.94x | 2.86x | 10.9 min | 75.67 | |
39.9 | 45.0/78.7/123.1 | 1.56x | 2.74x | 11.8 min | 71.37 | |
33.0 | 47.4/95.1/201.7 | 2.12x | 4.26x | 14.3 min | 70.66 | |
32.6 | 54.3/96.4/202.9 | 2.10x | 3.74x | 14.5 min | 59.89 | |
32.5 | 38.2/96.7/167.9 | 1.74x | 4.40x | 14.5 min | 61.00 | |
31.6 | 66.3/99.3/189.8 | 1.91x | 2.86x | 14.9 min | 66.06 | |
21.7 | 81.4/144.8/269.1 | 1.86x | 3.31x | 21.7 min | 63.93 | |
21.4 | 84.4/146.5/292.7 | 2.00x | 3.47x | 22.0 min | 68.04 | |
19.4 | 84.8/161.7/350.8 | 2.17x | 4.14x | 24.3 min | 71.48 | |
8.9 | 230.2/353.6/651.4 | 1.84x | 2.83x | 53.0 min | 66.64 |
How the turn latency score axis is measured
Full-stack app builders. Timed by the harness, not by the product interface, from request submission to a 200 response on the deployed URL. We publish p10, median and p90 over all 45 recorded prompt executions per product, regardless of outcome. The subscore is the fastest median in this board's cohort divided by the product's median, expressed as a percentage.
Coding agents. Timed by the harness, not by the product interface, from request submission to the agent's returned diff or completion signal for that prompt. There is no deployment step: aider, cline and claude-code, among others, never produce a hosted URL. We publish p10, median and p90 over all 45 recorded prompt executions per product, regardless of outcome. The subscore is the fastest median in this board's cohort divided by the product's median, expressed as a percentage.
This axis carries 7 of the 100 index points on the builders board and 7.22 of the 100 index points on the coding agent board. Each product completed 5 runs of 9 prompts in this cycle, and the underlying per run data behind every figure on this page is downloadable at /api/runs.json under a CC BY 4.0 licence.
A high score on one axis is not a recommendation. Read it against the composite column on the right and against the axes that matter for the work you have in front of you. Reliability at 84.4% for the leader of the builders cohort is the context for everything else on the page.
How the percentiles are computed. p10, median and p90 are linear-interpolated over all 45 recorded executions per product, regardless of outcome: a fail that ran long still counts, nothing is truncated or conditioned on success. The full disclosure, including the cycle-04 outcome-mix statistics and one named inconsistency we corrected in the reliability rule rather than in the data, is on the methodology page.