Axis leaderboard / Turn latency score / two boards

Turn latency

Median wall clock seconds for one prompt, not for a whole build. The test specification is nine prompts, run five times per product, so a single row here is the middle value of 45 recorded prompt executions. Shown as two separate tables, one per board: an index score, subscore or rank on one board is never comparable to the other.

Full-stack app builders: ranked by turn latency score

Ordered by median seconds per prompt, fastest first, within this board's own cohort.

Turn latency score, Full-stack app builders
#ToolTurn latency scorep10 / median / p90p90 / medianp90 / p10Full spec (derived)Composite
100.0
44.9/65.2/104.1
1.60x2.32x9.8 min78.31
81.5
46.3/80.0/143.0
1.79x3.09x12.0 min58.53
79.4
51.8/82.1/136.0
1.66x2.63x12.3 min74.68
77.8
54.4/83.8/149.9
1.79x2.76x12.6 min67.49
76.5
52.6/85.2/133.1
1.56x2.53x12.8 min80.54
74.9
49.4/87.0/137.1
1.58x2.78x13.1 min58.11
71.6
46.1/91.1/174.2
1.91x3.78x13.7 min68.66
69.3
58.8/94.1/161.4
1.72x2.74x14.1 min73.05
69.3
55.2/94.1/146.8
1.56x2.66x14.1 min69.97
67.9
53.5/96.0/186.6
1.94x3.49x14.4 min65.59
11Firebase Studio logoFirebase StudioSunsets Mar 2027
65.0
65.8/100.3/161.3
1.61x2.45x15.0 min80.92
61.3
59.2/106.3/167.6
1.58x2.83x15.9 min89.48
60.8
53.3/107.2/217.0
2.02x4.07x16.1 min62.12
60.1
65.1/108.4/176.3
1.63x2.71x16.3 min69.48
57.5
60.2/113.4/193.1
1.70x3.21x17.0 min61.06
57.3
59.6/113.8/231.5
2.03x3.88x17.1 min59.85
48.0
92.8/135.7/188.8
1.39x2.03x20.4 min75.47
47.6
72.2/137.1/249.0
1.82x3.45x20.6 min56.99
45.1
81.7/144.6/309.3
2.14x3.79x21.7 min68.68
43.3
76.0/150.6/295.5
1.96x3.89x22.6 min62.08

Coding agents: ranked by turn latency score

Ordered by median seconds per prompt, fastest first, within this board's own cohort.

Turn latency score, Coding agents
#ToolTurn latency scorep10 / median / p90p90 / medianp90 / p10Full spec (derived)Composite
100.0
19.8/31.4/54.5
1.74x2.75x4.7 min71.13
79.7
22.5/39.4/75.7
1.92x3.36x5.9 min65.88
79.1
24.2/39.7/64.4
1.62x2.66x6.0 min73.35
58.8
28.7/53.4/82.7
1.55x2.88x8.0 min77.92
52.5
36.1/59.8/110.7
1.85x3.07x9.0 min64.51
47.9
39.9/65.5/100.7
1.54x2.52x9.8 min79.72
43.9
44.9/71.6/110.0
1.54x2.45x10.7 min75.63
43.4
49.1/72.4/140.4
1.94x2.86x10.9 min75.67
39.9
45.0/78.7/123.1
1.56x2.74x11.8 min71.37
33.0
47.4/95.1/201.7
2.12x4.26x14.3 min70.66
32.6
54.3/96.4/202.9
2.10x3.74x14.5 min59.89
32.5
38.2/96.7/167.9
1.74x4.40x14.5 min61.00
31.6
66.3/99.3/189.8
1.91x2.86x14.9 min66.06
21.7
81.4/144.8/269.1
1.86x3.31x21.7 min63.93
21.4
84.4/146.5/292.7
2.00x3.47x22.0 min68.04
19.4
84.8/161.7/350.8
2.17x4.14x24.3 min71.48
8.9
230.2/353.6/651.4
1.84x2.83x53.0 min66.64

How the turn latency score axis is measured

Full-stack app builders. Timed by the harness, not by the product interface, from request submission to a 200 response on the deployed URL. We publish p10, median and p90 over all 45 recorded prompt executions per product, regardless of outcome. The subscore is the fastest median in this board's cohort divided by the product's median, expressed as a percentage.

Coding agents. Timed by the harness, not by the product interface, from request submission to the agent's returned diff or completion signal for that prompt. There is no deployment step: aider, cline and claude-code, among others, never produce a hosted URL. We publish p10, median and p90 over all 45 recorded prompt executions per product, regardless of outcome. The subscore is the fastest median in this board's cohort divided by the product's median, expressed as a percentage.

This axis carries 7 of the 100 index points on the builders board and 7.22 of the 100 index points on the coding agent board. Each product completed 5 runs of 9 prompts in this cycle, and the underlying per run data behind every figure on this page is downloadable at /api/runs.json under a CC BY 4.0 licence.

A high score on one axis is not a recommendation. Read it against the composite column on the right and against the axes that matter for the work you have in front of you. Reliability at 84.4% for the leader of the builders cohort is the context for everything else on the page.

How the percentiles are computed. p10, median and p90 are linear-interpolated over all 45 recorded executions per product, regardless of outcome: a fail that ran long still counts, nothing is truncated or conditioned on success. The full disclosure, including the cycle-04 outcome-mix statistics and one named inconsistency we corrected in the reliability rule rather than in the data, is on the methodology page.

Other axis leaderboards

Turn latency (2026 measurements) | LLM Tier