Axis leaderboard / Design output / two boards
Design output
Full-stack app builders: ranked by design output
Ordered by axis subscore, highest first, within this board's own cohort.
| #Tool | Design output | CLS | Composite |
|---|---|---|---|
94.0 | 0.040 | 78.31 | |
92.0 | 0.060 | 80.54 | |
86.0 | 0.020 | 89.48 | |
85.0 | 0.090 | 74.68 | |
84.0 | 0.050 | 68.66 | |
83.0 | 0.050 | 69.97 | |
82.0 | 0.070 | 58.53 | |
80.0 | 0.080 | 73.05 | |
80.0 | 0.090 | 56.99 | |
79.0 | 0.070 | 58.11 | |
78.0 | 0.060 | 69.48 | |
77.0 | 0.110 | 61.06 | |
75.0 | 0.160 | 62.12 | |
74.0 | 0.080 | 59.85 | |
73.0 | 0.120 | 75.47 | |
72.0 | 0.050 | 80.92 | |
72.0 | 0.040 | 67.49 | |
71.0 | 0.140 | 68.68 | |
70.0 | 0.070 | 65.59 | |
68.0 | 0.110 | 62.08 |
Coding agents: ranked by design output
Ordered by axis subscore, highest first, within this board's own cohort.
| #Tool | Design output | CLS | Composite |
|---|---|---|---|
70.0 | 0.040 | 71.37 | |
68.0 | 0.030 | 79.72 | |
66.0 | 0.030 | 77.92 | |
66.0 | 0.050 | 61.00 | |
65.0 | 0.040 | 75.67 | |
64.0 | 0.040 | 71.13 | |
63.0 | 0.040 | 70.66 | |
62.0 | 0.040 | 75.63 | |
62.0 | 0.050 | 71.48 | |
62.0 | 0.040 | 68.04 | |
60.0 | 0.040 | 73.35 | |
60.0 | 0.050 | 66.64 | |
58.0 | 0.050 | 64.51 | |
58.0 | 0.050 | 63.93 | |
56.0 | 0.050 | 66.06 | |
56.0 | 0.050 | 59.89 | |
52.0 | 0.050 | 65.88 |
How the design output axis is measured
Full-stack app builders. Blind review of the prompt 7 and prompt 8 artifacts by three reviewers who do not know which product produced which screenshot. Scored on hierarchy, typographic control, spacing discipline, state coverage and responsive behaviour at 390, 768 and 1440 pixels.
Coding agents. Blind review of the prompt 7 and prompt 8 artifacts by three reviewers who do not know which product produced which screenshot. Scored on hierarchy, typographic control, spacing discipline, state coverage and responsive behaviour at 390, 768 and 1440 pixels.
This axis carries 7 of the 100 index points on the builders board and 7.22 of the 100 index points on the coding agent board. Each product completed 5 runs of 9 prompts in this cycle, and the underlying per run data behind every figure on this page is downloadable at /api/runs.json under a CC BY 4.0 licence.
A high score on one axis is not a recommendation. Read it against the composite column on the right and against the axes that matter for the work you have in front of you. Reliability at 84.4% for the leader of the builders cohort is the context for everything else on the page.