Measurement lab / cycle 04 / August 2026
The AI app builder and coding agent benchmark, published in full
An independent lab gives every product the same written job, runs it repeatedly, and publishes every run, including the ones that failed.
The job: build a real business application with a public API, webhooks, background jobs, a staff-gated admin panel and a marketing page that works without JavaScript. We call this specification vibeOps; its full prompt list and pass criteria are on the methodology page.
Which board is yours
Two different products doing two different jobs. Never merged into one ordering.
Full-stack app builders
You describe the app. The tool builds it and hosts it. You get a working URL.
20 measured, including Totalum, Firebase Studio, Lovable, v0.
See the full-stack app builder leaderboard
Coding agents
You already have a codebase. The tool works in it and hands back a code change.
17 measured, including Claude Code, Cursor, Kiro, Augment Code.
See the coding agent leaderboard
The two are never ranked against each other: one builds and hosts a new application, the other edits code you already own, so one ordering across both would not be measuring one thing. Read why.
Also on this site, not a third leaderboard
LLM reference table
The language models most of these products are built on. Sourced vendor specifications and attributed third party benchmarks, no ranking of our own: we do not run a model harness.
Products measured
37
20 builders, 17 coding agents, plus 26 language models in a sourced reference table
Prompt executions
1,665
Published cohort, every one of them published row by row
Fastest full sequence, derived
4.7 min
Zed (Coding agents). Derived from a 31.4s median for one of the nine prompts, not a measured end to end session.
Highest reliability
93.3%
Totalum (Full-stack app builders). Strict pass rate over its recorded prompt executions.
How to read a number on this page
Score out of 100
Each board's composite index is a weighted sum of that board's own axis subscores. Multiply each subscore by its published weight and divide by 100, and you get the number in the table. No curve, no cohort adjustment, no editorial nudge.
Axis
The name this site gives to one separately measured property, such as reliability or design output. A board has nine or ten of them; the weight tables below list every one.
Lab index and reader index
Lab index: produced only by our recorded runs. Reader index: produced only by reader submitted scores. The reader index is a weighted mean of reader submitted axis scores using the same published weights as the lab index, renormalised over the axes readers have actually scored. It is published only once a product has at least three distinct voters covering at least half of the index weight. It is never blended into the lab index. Below that threshold the cell shows how many people have voted so far instead of a number.
Turn latency score, and turn latency
Two different things share a similar name, on purpose kept distinct here. Turn latency score is the 0 to 100 axis figure inside the composite: the fastest median in a board's own cohort divided by the product's median. Turn latency, printed in seconds in the tables and tiles, is the median wall clock time for one prompt out of the nine in the specification, never the time to build a whole app. A "full spec, derived" figure next to it is nine of those medians added together, not a separately measured session.
When this was last measured
Cycle 04, last run 15 Aug 2026. Cadence is roughly every seven days; the next date is announced once it is scheduled.
The limits
- Five runs gives a usable median, not a confident ordering: read two products within about one and a half index points of each other as tied.
- Rank deltas are not shown this cycle: the roster was just re-based from one merged ranking into two class boards, so last cycle's ranks and this cycle's are not on the same scale.
- The full specification minutes figure is derived (nine medians added together), never a separately measured end to end session.
+Axis weights, both boardsEvery axis, every weight, unchanged from /methodology
Full-stack app builders weight table
| Axis | Weight | % |
|---|---|---|
| Agent performance | 18 | |
| Reliability | 18 | |
| Scalability | 13 | |
| SEO and GEO | 12 | |
| API and MCP | 9 | |
| Integrations | 8 | |
| Design output | 7 | |
| Turn latency score | 7 | |
| Value | 5 | |
| Code ownership | 3 |
Coding agents weight table
| Axis | Weight | % |
|---|---|---|
| Agent performance | 18.55 | |
| Reliability | 18.56 | |
| Scalability | 13.4 | |
| SEO and GEO | 12.37 | |
| API and MCP | 9.28 | |
| Integrations | 8.25 | |
| Design output | 7.22 | |
| Turn latency score | 7.22 | |
| Value | 5.15 |
Code ownership is not scored on this board because a coding agent edits a repository you already own, so the property was never the vendor's to grant and a full mark measures nothing. The other nine weights are rescaled proportionally so they still sum to 100; their relative importance to each other is unchanged.
Composite index: 20 full-stack app builders
Sorted by index score descending. Click a product for the raw run data behind its row.
| RankTool | Lab index | Reader index | Turn latency (s) | Full spec (derived) | Reliability % | SEO score | Delta |
|---|---|---|---|---|---|---|---|
89.48 | 72.29 | 106.3 | 15.9 min | 93.3% | 95.0 | withheld | |
80.92 | 74.74 | 100.3 | 15.0 min | 86.7% | 73.0 | withheld | |
80.54 | 82.79 | 85.2 | 12.8 min | 82.2% | 75.0 | withheld | |
78.31 | 80.68 | 65.2 | 9.8 min | 84.4% | 76.0 | withheld | |
75.47 | 67.02 | 135.7 | 20.4 min | 86.7% | 62.0 | withheld | |
74.68 | 77.49 | 82.1 | 12.3 min | 80.0% | 66.0 | withheld | |
73.05 | 72.48 | 94.1 | 14.1 min | 84.4% | 64.0 | withheld | |
69.97 | 71.63 | 94.1 | 14.1 min | 77.8% | 68.0 | withheld | |
69.48 | 1 voting | 108.4 | 16.3 min | 73.3% | 66.0 | withheld | |
68.68 | 2 voting | 144.6 | 21.7 min | 80.0% | 58.0 | withheld | |
68.66 | 1 voting | 91.1 | 13.7 min | 68.9% | 70.0 | withheld | |
67.49 | 72.53 | 83.8 | 12.6 min | 55.6% | 72.0 | withheld | |
65.59 | - | 96.0 | 14.4 min | 64.4% | 58.0 | withheld | |
62.12 | 1 voting | 107.2 | 16.1 min | 60.0% | 57.0 | withheld | |
62.08 | - | 150.6 | 22.6 min | 66.7% | 48.0 | withheld | |
61.06 | 1 voting | 113.4 | 17.0 min | 60.0% | 60.0 | withheld | |
59.85 | 2 voting | 113.8 | 17.1 min | 55.6% | 62.0 | withheld | |
58.53 | 2 voting | 80.0 | 12.0 min | 48.9% | 61.0 | withheld | |
58.11 | 2 voting | 87.0 | 13.1 min | 51.1% | 63.0 | withheld | |
56.99 | - | 137.1 | 20.6 min | 48.9% | 42.0 | withheld |
Rank deltas are withheld for cycle 04 because the roster was re-based from a single merged ranking into two class boards, so this cycle's ranks and the previous cycle's are not on the same scale. Deltas resume in cycle 05.
Turn latency is the median wall clock seconds for one prompt, not for a whole build; the full spec column derives what nine of those would add up to and is not a separately measured figure. Reliability is the share of the 45 recorded executions that met the published pass criterion with no operator intervention. SEO score is the 0 to 100 subscore for the crawlability of the public pages each product generates.
Composite index: 17 coding agents
Sorted by index score descending. Click a product for the raw run data behind its row.
| RankTool | Lab index | Reader index | Turn latency (s) | Full spec (derived) | Reliability % | SEO score | Delta |
|---|---|---|---|---|---|---|---|
79.72 | 85.23 | 65.5 | 9.8 min | 91.1% | 62.0 | withheld | |
77.92 | 82.93 | 53.4 | 8.0 min | 86.7% | 58.0 | withheld | |
75.67 | 78.57 | 72.4 | 10.9 min | 93.3% | 60.0 | withheld | |
75.63 | 2 voting | 71.6 | 10.7 min | 88.9% | 58.0 | withheld | |
73.35 | 2 voting | 39.7 | 6.0 min | 77.8% | 56.0 | withheld | |
71.48 | 64.41 | 161.7 | 24.3 min | 82.2% | 56.0 | withheld | |
71.37 | 1 voting | 78.7 | 11.8 min | 73.3% | 58.0 | withheld | |
71.13 | 79.79 | 31.4 | 4.7 min | 71.1% | 55.0 | withheld | |
70.66 | 77.04 | 95.1 | 14.3 min | 80.0% | 57.0 | withheld | |
68.04 | 1 voting | 146.5 | 22.0 min | 75.6% | 55.0 | withheld | |
66.64 | 2 voting | 353.6 | 53.0 min | 71.1% | 54.0 | withheld | |
66.06 | 2 voting | 99.3 | 14.9 min | 77.8% | 52.0 | withheld | |
65.88 | 71.33 | 39.4 | 5.9 min | 68.9% | 54.0 | withheld | |
64.51 | 1 voting | 59.8 | 9.0 min | 68.9% | 55.0 | withheld | |
63.93 | 2 voting | 144.8 | 21.7 min | 64.4% | 56.0 | withheld | |
61.00 | 2 voting | 96.7 | 14.5 min | 51.1% | 54.0 | withheld | |
59.89 | 1 voting | 96.4 | 14.5 min | 53.3% | 53.0 | withheld |
Rank deltas are withheld for cycle 04 because the roster was re-based from a single merged ranking into two class boards, so this cycle's ranks and the previous cycle's are not on the same scale. Deltas resume in cycle 05.
Turn latency is the median wall clock seconds for one prompt, not for a whole build; the full spec column derives what nine of those would add up to and is not a separately measured figure. Reliability is the share of the 45 recorded executions that met the published pass criterion with no operator intervention. SEO score is the 0 to 100 subscore for the crawlability of the public pages each product generates.
Single axis leaderboards
The composite hides trade offs. These do not. Each shows both boards as two separate tables.
Turn latency score
Turn latency
Median wall clock seconds for one prompt, not for a whole build. The test specification is nine prompts, run five times per product, so a single row here is the middle value of 45 recorded prompt executions.
Reliability
Reliability
Strict pass rate across 1,665 recorded prompt executions.
SEO and GEO
SEO and GEO
Server rendered HTML bytes with JavaScript disabled, LCP, CLS and structured data output.
API and MCP
API and MCP
REST surface, webhook delivery, background jobs and Model Context Protocol support.
Design output
Design output
Blind scored visual quality of the generated admin panel and marketing surface.
Value
Value
Index points delivered per euro of entry price. The two boards price different things, so the two value scores are not comparable across boards.
The LLM reference table
The models most of these products are built on. Sourced facts and attributed third party benchmarks, no composite index of our own.
Anthropic
Claude Opus 5
1691 Elo
Context 1,000,000
Moonshot AI
Kimi K3
1674 Elo
Context 1,048,576
Alibaba
Qwen3.8-Max
1669 Elo
Context 1,048,576, charged at one flat rate across the whole window with no length tiering
xAI
Grok 4.6
1629 Elo
Context 500,000
Head to head comparisons
Curated pairs only, built from the same run data as the index. Every pair stays within one board.
Lab notes
news
957 reader rows carried a date from before this domain existed
09 Sep 2026
news
Eight articles carried a publication date from before this domain was registered
23 Aug 2026
news
Entry prices are now published in the vendor's own currency, not a euro conversion that tracked no single rate
22 Aug 2026
news
The model board carried a composite index with no run behind it. It is gone.
22 Aug 2026
How this benchmark works, and why it is two boards
Most published comparisons of AI app builders are written from a free trial and a screenshot. Someone signs up, types a prompt about a to do list, likes what appears, and writes five hundred words about the future of software. That process cannot separate a product that is genuinely good from a product that is good at the first ninety seconds, and it produces rankings that change with the mood of the writer rather than with the behaviour of the software. This site exists because we wanted the other kind of comparison: one specification, applied identically, repeated enough times to see variance, with every measurement published so the conclusion can be checked.
The specification is called vibeOps. It describes a multi tenant SaaS product: a public REST API with key based authentication and rate limiting, outbound webhooks that are signed and retried, background jobs that survive a redeploy, an internal admin panel gated to staff accounts, and a public marketing surface that has to render without JavaScript. It is deliberately not a landing page and not a to do list. It is the shape of work people actually pay for, and it forces every product to touch data modelling, authentication, asynchronous work and an operations surface rather than just layout.
Both classes are given the same nine prompts, but a merged ranking across them was never measuring one thing: the two products start from a different state (a cold empty project versus a warm existing repository), produce a different unit of work (a deployed URL versus a diff), and are judged by a different success criterion (does the app work versus do the existing tests still pass). Turn latency score, value, integrations and code ownership together carry 23 of the 100 index points on the builders board, and none of those four survives a merge across the two classes unchanged. That is why this site publishes two boards, each with its own weight table, its own rank column starting at 1, and no page that mixes them into one ordering.
We issue 9 prompts in a fixed order. Each prompt is phrased precisely, in the language a working engineer would use, and each has a published pass criterion written before the run. The criteria are concrete rather than aesthetic: a cross tenant read must be refused, a rate limiter must return 429 with a retry hint, a webhook signature must verify against the stored secret, a job must be idempotent when replayed and must still be queued after a restart. The prompts are never rephrased to help a product that is struggling, and a product that asks a clarifying question gets an answer without being told what it forgot.
Every product runs the full sequence 5 times, which gives 45 recorded executions per product and 1,665 across the published cohort this cycle. Repetition is the part most benchmarks skip, and it is the part that matters most, because these systems are not deterministic. The same prompt to the same product on the same day can produce a clean result once and a stalled agent the next time. Publishing a single build time from a single lucky run is not a measurement, it is an anecdote. So for every product we publish the tenth percentile, the median and the ninetieth percentile of per-prompt turn latency over all 45 recorded executions regardless of outcome, and we publish the outcome of every individual execution as pass, partial or fail.
Turn latency score is worth less than most readers expect, 7 of 100 points on the builders board and 7.22 of 100 on the coding agent board, and the seconds figure behind it measures one prompt, not a whole build: a product that returns a 200 response in thirty seconds and leaves you to write the queue yourself has not saved you an afternoon. Code ownership is worth only 3 points on the builders board. Code ownership is not scored at all on the coding agent board, because a coding agent edits a repository you already own, so the property was never the vendor's to grant and a full mark measures nothing. The other nine weights are rescaled proportionally so they still sum to 100; their relative importance to each other is unchanged.
Each board's composite is a plain weighted sum with no curve, no normalisation against a moving cohort mean and no editorial adjustment. Take the subscores from any product page, multiply each by its board's published weight, divide by one hundred, and you get the number in the index column to two decimal places. Where two products on the same board sit within about one and a half index points of each other, treat them as tied: five runs is enough for a usable median and not enough for a confident ordering at that distance. When a number here is wrong, we would rather hear it with evidence and correct it in public, which has already happened twice. The whole dataset, including every individual run, is available under a CC BY 4.0 licence for anyone who wants to check the arithmetic or disagree with the weights. Derived, not separately measured. It is nine times the median prompt latency, converted to minutes. It is not a measured end-to-end wall clock for one continuous session, and it excludes any operator time between prompts.