Axis leaderboard / Design output / weight 7 of 100

Best AI app builders for design output

Blind scored visual quality of the generated admin panel and marketing surface. Visual quality of what the product produces unprompted, both the admin surface and the public pages.

Ranked by design output

Ordered by axis subscore, highest first.

Best AI app builders for design output
#ToolDesign outputCLSComposite
94.0
0.04074.69
92.0
0.06077.87
86.0
0.02087.96
86.0
0.06059.33
85.0
0.09071.94
84.0
0.05066.17
83.0
0.05067.16
82.0
0.07056.55
80.0
0.08070.54
80.0
0.09055.27
79.0
0.07055.45
78.0
0.06067.30
77.0
0.11058.78
76.0
0.04070.50
75.0
0.16060.16
74.0
0.08058.17
73.0
0.12073.42
72.0
0.05078.06
72.0
0.04064.77
72.0
0.13055.67
71.0
0.14066.79
70.0
0.04071.91
70.0
0.07063.18
68.0
0.03080.30
68.0
0.11060.10
66.0
0.03078.65
66.0
0.05062.25
65.0
0.04076.31
29Zed logoZed
64.0
0.04072.17
63.0
0.04071.51
62.0
0.04075.92
62.0
0.05071.52
62.0
0.04068.69
34Amp logoAmp
60.0
0.04074.04
60.0
0.05066.36
60.0
0.06056.79
58.0
0.05066.89
58.0
0.05065.18
56.0
0.05067.13
56.0
0.05061.04
52.0
0.05067.27

How the design output axis is measured

Blind review of the prompt 7 and prompt 8 artifacts by three reviewers who do not know which product produced which screenshot. Scored on hierarchy, typographic control, spacing discipline, state coverage and responsive behaviour at 390, 768 and 1440 pixels.

This axis carries 7 of the 100 index points. Each product completed 5 runs of 9 prompts in this cycle, and the underlying per run data behind every figure on this page is downloadable at /api/runs.json under a CC BY 4.0 licence.

A high score on one axis is not a recommendation. Read it against the composite column on the right and against the axes that matter for the work you have in front of you. Reliability at 84.4% for the leader of this cohort is the context for everything else on the page.

Other axis leaderboards