Leaderboard / app builders / cycle 04

AI app builder leaderboard 2026

Every measurement we hold on 41 products, in one grid. 5 runs of the 9 prompt vibeOps specification per product. Scroll sideways for the full axis set. The first column stays put.
Last run: 14 Aug 2026 17:40 UTCNext scheduled run: 21 Aug 2026 09:00 UTC

Products

41

Full stack builders and coding agents

Prompt executions

1,845

Published individually at /api/runs.json

Index leader

Totalum

Composite 87.96

Fastest median

31.4 s

Zed

Full measurement grid

Index, delta, speed percentiles, outcome split and all ten subscores.

Full AI app builder leaderboard with per axis subscores and raw measurements
#ToolLab indexReader indexDeltaSpeed p10/med/p90Reliability %Runs pass/part/failAgentScal.SEOAPIIntegr.DesignValueCode own.HTML no JSLCP msCLSFrom EUR
87.96
72.29 4
59.2/106.3/167.6
93.3%
42/2/1
94.092.095.097.088.086.084.096.041.0 KB1,4800.02029.00
80.30
85.23 1
39.9/65.5/100.7
91.1%
41/3/1
95.086.062.078.076.068.080.099.026.0 KB1,6800.03017.00
78.65
82.93 1
28.7/53.4/82.7
86.7%
39/3/3
92.082.058.074.080.066.082.099.024.0 KB1,7200.03019.00
78.06
74.74 1
65.8/100.3/161.3
86.7%
39/5/1
79.087.073.079.090.072.082.084.022.0 KB1,9800.0500.00
77.87
82.79 1
52.6/85.2/133.1
82.2%
37/6/2
89.078.075.066.086.092.074.079.018.0 KB2,1400.06021.00
76.31
78.570
49.1/72.4/140.4
93.3%
42/2/1
84.080.060.072.074.065.078.097.022.0 KB1,8000.04018.00
75.92
2 voting0
44.9/71.6/110.0
88.9%
40/3/2
86.084.058.074.076.062.070.097.022.0 KB1,8000.04045.00
74.69
80.68 1
44.9/65.2/104.1
84.4%
38/6/1
80.072.076.060.066.094.068.086.027.0 KB1,6200.04018.00
74.04
2 voting 1
24.2/39.7/64.4
77.8%
35/6/4
85.074.056.072.068.060.076.098.020.0 KB1,8400.0400.00
73.42
67.02 1
92.8/135.7/188.8
86.7%
39/4/2
83.080.062.071.077.073.066.082.09.0 KB2,8800.12022.00
11Zed logoZed
72.17
79.79 1
19.8/31.4/54.5
71.1%
32/5/8
77.070.055.066.064.064.086.099.019.0 KB1,7800.04010.00
71.94
77.49 1
51.8/82.1/136.0
80.0%
36/8/1
81.070.066.058.072.085.074.088.012.0 KB2,4600.09018.00
71.91
1 voting 1
45.0/78.7/123.1
73.3%
33/7/5
83.076.058.070.072.070.086.096.022.0 KB1,8200.0400.00
71.52
64.41 1
84.8/161.7/350.8
82.2%
37/5/3
81.082.056.076.078.062.064.093.019.0 KB1,9400.05036.00
71.51
77.04 1
47.4/95.1/201.7
80.0%
36/8/1
82.072.057.070.072.063.088.099.021.0 KB1,8600.0400.00
70.54
72.48 1
58.8/94.1/161.4
84.4%
38/7/0
75.068.064.066.075.080.076.055.014.0 KB2,3200.08018.00
70.50
70.80 1
51.5/88.7/142.4
60.0%
27/8/10
75.084.065.082.072.076.078.088.012.0 KB1,8400.04025.00
68.69
1 voting 1
84.4/146.5/292.7
75.6%
34/7/4
80.074.055.066.070.062.084.096.021.0 KB1,8600.0400.00
67.30
1 voting 1
65.1/108.4/176.3
73.3%
33/7/5
74.068.066.058.070.078.070.074.017.0 KB2,1800.06019.00
67.27
71.33 1
22.5/39.4/75.7
68.9%
31/7/7
72.064.054.062.058.052.094.099.017.0 KB1,9200.0500.00
67.16
71.63 1
55.2/94.1/146.8
77.8%
35/8/2
71.062.068.057.066.083.070.070.018.0 KB1,9600.05022.00
67.13
2 voting 1
66.3/99.3/189.8
77.8%
35/6/4
70.076.052.068.070.056.074.095.016.0 KB2,0000.05030.00
66.89
1 voting 1
36.1/59.8/110.7
68.9%
31/7/7
76.068.055.064.062.058.080.097.018.0 KB1,9000.05019.00
66.79
2 voting 1
81.7/144.6/309.3
80.0%
36/3/6
79.066.058.061.063.071.069.072.08.0 KB3,1200.14019.00
66.36
2 voting 1
230.2/353.6/651.4
71.1%
32/7/6
78.080.054.072.074.060.052.094.020.0 KB1,9000.05018.00
66.17
1 voting 1
46.1/91.1/174.2
68.9%
31/8/6
72.066.070.055.064.084.072.062.019.0 KB2,0200.05020.00
65.18
2 voting 1
81.4/144.8/269.1
64.4%
29/9/7
74.070.056.068.066.058.090.099.018.0 KB1,9800.0500.00
64.77
72.53 1
54.4/83.8/149.9
55.6%
25/10/10
68.064.072.060.062.072.092.098.020.0 KB1,8800.0400.00
63.18
- 1
53.5/96.0/186.6
64.4%
29/7/9
69.060.058.062.068.070.076.080.013.0 KB2,2400.07015.00
62.25
2 voting 1
38.2/96.7/167.9
51.1%
23/9/13
73.066.054.060.066.066.088.096.018.0 KB1,9600.0503.00
61.04
1 voting 1
54.3/96.4/202.9
53.3%
24/11/10
71.066.053.062.060.056.084.098.017.0 KB1,9800.0500.00
60.16
1 voting 1
53.3/107.2/217.0
60.0%
27/7/11
70.061.057.052.059.075.071.058.07.0 KB3,2600.16017.00
60.10
- 1
76.0/150.6/295.5
66.7%
30/12/3
70.074.048.052.046.068.076.058.06.0 KB2,8400.1100.00
59.33
2 voting 1
53.4/90.3/168.5
53.3%
24/8/13
64.056.064.050.054.086.074.076.016.0 KB2,0600.06016.00
58.78
1 voting 1
60.2/113.4/193.1
60.0%
27/8/10
68.058.060.047.055.077.072.048.010.0 KB2,7400.11016.00
58.17
2 voting 1
59.6/113.8/231.5
55.6%
25/10/10
60.052.062.048.054.074.090.097.015.0 KB2,3200.0800.00
56.79
- 1
69.6/120.9/243.9
53.3%
24/7/14
65.058.050.054.056.060.078.090.015.0 KB2,1400.06014.00
56.55
2 voting 1
46.3/80.0/143.0
48.9%
22/9/14
63.054.061.044.052.082.078.052.014.0 KB2,1200.07012.00
55.67
1 voting 1
61.1/105.2/186.8
53.3%
24/12/9
66.055.055.044.052.072.074.050.06.0 KB3,0400.13020.00
55.45
2 voting 1
49.4/87.0/137.1
51.1%
23/11/11
62.052.063.041.049.079.075.045.015.0 KB2,2600.07015.00
55.27
- 1
72.2/137.1/249.0
48.9%
22/10/13
66.058.042.048.060.080.074.066.08.0 KB2,4600.09025.00

Speed columns are wall clock seconds per prompt, measured by the harness from submission to the first 200 response on the deployed URL. Runs are split pass, partial and fail across the 45 recorded executions per product. Reliability is the strict pass share. HTML no JS is real text bytes with scripts disabled.

Read a single axis instead

The composite compresses ten different questions into one number.

Notes on reading this table

Products within about one and a half index points of each other should be read as tied. Five runs per product gives a usable median and a rough sense of spread, not a tight confidence interval, and the honest thing to do with a small sample is to say so rather than to present a false ordering.

The segment column separates full stack builders, which own the whole pipeline from prompt to hosting, from coding agents, which work inside a development environment you control. Both are in the same table because both were given the same specification and both are chosen for the same job, but the coding agents tend to trade speed for operational completeness. That trade shows clearly in the speed and scalability columns.

Reliability at 93.3% at the top of the table means roughly one execution in eight still needed intervention or failed outright, for the best product in the cohort. That is the state of the field in August 2026, and it is the single most useful thing to know before planning a delivery date around one of these tools.