Axis leaderboard / Design output / two boards

Design output

Blind scored visual quality of the generated admin panel and marketing surface. Shown as two separate tables, one per board: an index score, subscore or rank on one board is never comparable to the other.

Full-stack app builders: ranked by design output

Ordered by axis subscore, highest first, within this board's own cohort.

Design output, Full-stack app builders
#ToolDesign outputCLSComposite
94.0
0.04078.31
92.0
0.06080.54
86.0
0.02089.48
85.0
0.09074.68
84.0
0.05068.66
83.0
0.05069.97
82.0
0.07058.53
80.0
0.08073.05
80.0
0.09056.99
79.0
0.07058.11
78.0
0.06069.48
77.0
0.11061.06
75.0
0.16062.12
74.0
0.08059.85
73.0
0.12075.47
16Firebase Studio logoFirebase StudioSunsets Mar 2027
72.0
0.05080.92
72.0
0.04067.49
71.0
0.14068.68
70.0
0.07065.59
68.0
0.11062.08

Coding agents: ranked by design output

Ordered by axis subscore, highest first, within this board's own cohort.

Design output, Coding agents
#ToolDesign outputCLSComposite
70.0
0.04071.37
68.0
0.03079.72
66.0
0.03077.92
66.0
0.05061.00
65.0
0.04075.67
64.0
0.04071.13
63.0
0.04070.66
62.0
0.04075.63
62.0
0.05071.48
62.0
0.04068.04
11Amp logoAmp
60.0
0.04073.35
60.0
0.05066.64
58.0
0.05064.51
58.0
0.05063.93
56.0
0.05066.06
56.0
0.05059.89
52.0
0.05065.88

How the design output axis is measured

Full-stack app builders. Blind review of the prompt 7 and prompt 8 artifacts by three reviewers who do not know which product produced which screenshot. Scored on hierarchy, typographic control, spacing discipline, state coverage and responsive behaviour at 390, 768 and 1440 pixels.

Coding agents. Blind review of the prompt 7 and prompt 8 artifacts by three reviewers who do not know which product produced which screenshot. Scored on hierarchy, typographic control, spacing discipline, state coverage and responsive behaviour at 390, 768 and 1440 pixels.

This axis carries 7 of the 100 index points on the builders board and 7.22 of the 100 index points on the coding agent board. Each product completed 5 runs of 9 prompts in this cycle, and the underlying per run data behind every figure on this page is downloadable at /api/runs.json under a CC BY 4.0 licence.

A high score on one axis is not a recommendation. Read it against the composite column on the right and against the axes that matter for the work you have in front of you. Reliability at 84.4% for the leader of the builders cohort is the context for everything else on the page.

Other axis leaderboards