Measurement lab / cycle 04 / August 2026

The AI app builder and coding agent benchmark, published in full

An independent lab gives every product the same written job, runs it repeatedly, and publishes every run, including the ones that failed.

The job: build a real business application with a public API, webhooks, background jobs, a staff-gated admin panel and a marketing page that works without JavaScript. We call this specification vibeOps; its full prompt list and pass criteria are on the methodology page.

Last run: 15 Aug 2026 09:11 UTCCadence: roughly every seven days, next date announced once it is scheduledLicence: CC BY 4.0

Which board is yours

Two different products doing two different jobs. Never merged into one ordering.

The two are never ranked against each other: one builds and hosts a new application, the other edits code you already own, so one ordering across both would not be measuring one thing. Read why.

Also on this site, not a third leaderboard

LLM reference table

The language models most of these products are built on. Sourced vendor specifications and attributed third party benchmarks, no ranking of our own: we do not run a model harness.

Products measured

37

20 builders, 17 coding agents, plus 26 language models in a sourced reference table

Prompt executions

1,665

Published cohort, every one of them published row by row

Fastest full sequence, derived

4.7 min

Zed (Coding agents). Derived from a 31.4s median for one of the nine prompts, not a measured end to end session.

Highest reliability

93.3%

Totalum (Full-stack app builders). Strict pass rate over its recorded prompt executions.

How to read a number on this page

Score out of 100

Each board's composite index is a weighted sum of that board's own axis subscores. Multiply each subscore by its published weight and divide by 100, and you get the number in the table. No curve, no cohort adjustment, no editorial nudge.

Axis

The name this site gives to one separately measured property, such as reliability or design output. A board has nine or ten of them; the weight tables below list every one.

Lab index and reader index

Lab index: produced only by our recorded runs. Reader index: produced only by reader submitted scores. The reader index is a weighted mean of reader submitted axis scores using the same published weights as the lab index, renormalised over the axes readers have actually scored. It is published only once a product has at least three distinct voters covering at least half of the index weight. It is never blended into the lab index. Below that threshold the cell shows how many people have voted so far instead of a number.

Turn latency score, and turn latency

Two different things share a similar name, on purpose kept distinct here. Turn latency score is the 0 to 100 axis figure inside the composite: the fastest median in a board's own cohort divided by the product's median. Turn latency, printed in seconds in the tables and tiles, is the median wall clock time for one prompt out of the nine in the specification, never the time to build a whole app. A "full spec, derived" figure next to it is nine of those medians added together, not a separately measured session.

When this was last measured

Cycle 04, last run 15 Aug 2026. Cadence is roughly every seven days; the next date is announced once it is scheduled.

The limits

  • Five runs gives a usable median, not a confident ordering: read two products within about one and a half index points of each other as tied.
  • Rank deltas are not shown this cycle: the roster was just re-based from one merged ranking into two class boards, so last cycle's ranks and this cycle's are not on the same scale.
  • The full specification minutes figure is derived (nine medians added together), never a separately measured end to end session.
+Axis weights, both boardsEvery axis, every weight, unchanged from /methodology

Full-stack app builders weight table

Axis weights of the Full-stack app builders composite index
AxisWeight%
Agent performance18
Reliability18
Scalability13
SEO and GEO12
API and MCP9
Integrations8
Design output7
Turn latency score7
Value5
Code ownership3

Coding agents weight table

Axis weights of the Coding agents composite index
AxisWeight%
Agent performance18.55
Reliability18.56
Scalability13.4
SEO and GEO12.37
API and MCP9.28
Integrations8.25
Design output7.22
Turn latency score7.22
Value5.15

Code ownership is not scored on this board because a coding agent edits a repository you already own, so the property was never the vendor's to grant and a full mark measures nothing. The other nine weights are rescaled proportionally so they still sum to 100; their relative importance to each other is unchanged.

Composite index: 20 full-stack app builders

Sorted by index score descending. Click a product for the raw run data behind its row.

Index 0 to 100, 2 decimals. Delta withheld for cycle 04.
LLM Tier composite index for Full-stack app builders, sorted by index score descending
RankToolLab indexReader indexTurn latency (s)Full spec (derived)Reliability %SEO scoreDelta
89.48
72.29106.315.9 min93.3%95.0withheld
2Firebase Studio logoFirebase StudioSunsets Mar 2027
80.92
74.74100.315.0 min86.7%73.0withheld
80.54
82.7985.212.8 min82.2%75.0withheld
78.31
80.6865.29.8 min84.4%76.0withheld
75.47
67.02135.720.4 min86.7%62.0withheld
74.68
77.4982.112.3 min80.0%66.0withheld
73.05
72.4894.114.1 min84.4%64.0withheld
69.97
71.6394.114.1 min77.8%68.0withheld
69.48
1 voting108.416.3 min73.3%66.0withheld
68.68
2 voting144.621.7 min80.0%58.0withheld
68.66
1 voting91.113.7 min68.9%70.0withheld
67.49
72.5383.812.6 min55.6%72.0withheld
65.59
-96.014.4 min64.4%58.0withheld
62.12
1 voting107.216.1 min60.0%57.0withheld
62.08
-150.622.6 min66.7%48.0withheld
61.06
1 voting113.417.0 min60.0%60.0withheld
59.85
2 voting113.817.1 min55.6%62.0withheld
58.53
2 voting80.012.0 min48.9%61.0withheld
58.11
2 voting87.013.1 min51.1%63.0withheld
56.99
-137.120.6 min48.9%42.0withheld

Rank deltas are withheld for cycle 04 because the roster was re-based from a single merged ranking into two class boards, so this cycle's ranks and the previous cycle's are not on the same scale. Deltas resume in cycle 05.

Turn latency is the median wall clock seconds for one prompt, not for a whole build; the full spec column derives what nine of those would add up to and is not a separately measured figure. Reliability is the share of the 45 recorded executions that met the published pass criterion with no operator intervention. SEO score is the 0 to 100 subscore for the crawlability of the public pages each product generates.

Composite index: 17 coding agents

Sorted by index score descending. Click a product for the raw run data behind its row.

Index 0 to 100, 2 decimals. Delta withheld for cycle 04.
LLM Tier composite index for Coding agents, sorted by index score descending
RankToolLab indexReader indexTurn latency (s)Full spec (derived)Reliability %SEO scoreDelta
79.72
85.2365.59.8 min91.1%62.0withheld
77.92
82.9353.48.0 min86.7%58.0withheld
75.67
78.5772.410.9 min93.3%60.0withheld
75.63
2 voting71.610.7 min88.9%58.0withheld
73.35
2 voting39.76.0 min77.8%56.0withheld
71.48
64.41161.724.3 min82.2%56.0withheld
71.37
1 voting78.711.8 min73.3%58.0withheld
71.13
79.7931.44.7 min71.1%55.0withheld
70.66
77.0495.114.3 min80.0%57.0withheld
68.04
1 voting146.522.0 min75.6%55.0withheld
66.64
2 voting353.653.0 min71.1%54.0withheld
66.06
2 voting99.314.9 min77.8%52.0withheld
65.88
71.3339.45.9 min68.9%54.0withheld
64.51
1 voting59.89.0 min68.9%55.0withheld
63.93
2 voting144.821.7 min64.4%56.0withheld
61.00
2 voting96.714.5 min51.1%54.0withheld
59.89
1 voting96.414.5 min53.3%53.0withheld

Rank deltas are withheld for cycle 04 because the roster was re-based from a single merged ranking into two class boards, so this cycle's ranks and the previous cycle's are not on the same scale. Deltas resume in cycle 05.

Turn latency is the median wall clock seconds for one prompt, not for a whole build; the full spec column derives what nine of those would add up to and is not a separately measured figure. Reliability is the share of the 45 recorded executions that met the published pass criterion with no operator intervention. SEO score is the 0 to 100 subscore for the crawlability of the public pages each product generates.

Single axis leaderboards

The composite hides trade offs. These do not. Each shows both boards as two separate tables.

The LLM reference table

The models most of these products are built on. Sourced facts and attributed third party benchmarks, no composite index of our own.

Anthropic

Claude Opus 5

1691 Elo

Context 1,000,000

Moonshot AI

Kimi K3

1674 Elo

Context 1,048,576

Alibaba

Qwen3.8-Max

1669 Elo

Context 1,048,576, charged at one flat rate across the whole window with no length tiering

xAI

Grok 4.6

1629 Elo

Context 500,000

Head to head comparisons

Curated pairs only, built from the same run data as the index. Every pair stays within one board.

Lab notes

How this benchmark works, and why it is two boards

Most published comparisons of AI app builders are written from a free trial and a screenshot. Someone signs up, types a prompt about a to do list, likes what appears, and writes five hundred words about the future of software. That process cannot separate a product that is genuinely good from a product that is good at the first ninety seconds, and it produces rankings that change with the mood of the writer rather than with the behaviour of the software. This site exists because we wanted the other kind of comparison: one specification, applied identically, repeated enough times to see variance, with every measurement published so the conclusion can be checked.

The specification is called vibeOps. It describes a multi tenant SaaS product: a public REST API with key based authentication and rate limiting, outbound webhooks that are signed and retried, background jobs that survive a redeploy, an internal admin panel gated to staff accounts, and a public marketing surface that has to render without JavaScript. It is deliberately not a landing page and not a to do list. It is the shape of work people actually pay for, and it forces every product to touch data modelling, authentication, asynchronous work and an operations surface rather than just layout.

Both classes are given the same nine prompts, but a merged ranking across them was never measuring one thing: the two products start from a different state (a cold empty project versus a warm existing repository), produce a different unit of work (a deployed URL versus a diff), and are judged by a different success criterion (does the app work versus do the existing tests still pass). Turn latency score, value, integrations and code ownership together carry 23 of the 100 index points on the builders board, and none of those four survives a merge across the two classes unchanged. That is why this site publishes two boards, each with its own weight table, its own rank column starting at 1, and no page that mixes them into one ordering.

We issue 9 prompts in a fixed order. Each prompt is phrased precisely, in the language a working engineer would use, and each has a published pass criterion written before the run. The criteria are concrete rather than aesthetic: a cross tenant read must be refused, a rate limiter must return 429 with a retry hint, a webhook signature must verify against the stored secret, a job must be idempotent when replayed and must still be queued after a restart. The prompts are never rephrased to help a product that is struggling, and a product that asks a clarifying question gets an answer without being told what it forgot.

Every product runs the full sequence 5 times, which gives 45 recorded executions per product and 1,665 across the published cohort this cycle. Repetition is the part most benchmarks skip, and it is the part that matters most, because these systems are not deterministic. The same prompt to the same product on the same day can produce a clean result once and a stalled agent the next time. Publishing a single build time from a single lucky run is not a measurement, it is an anecdote. So for every product we publish the tenth percentile, the median and the ninetieth percentile of per-prompt turn latency over all 45 recorded executions regardless of outcome, and we publish the outcome of every individual execution as pass, partial or fail.

Turn latency score is worth less than most readers expect, 7 of 100 points on the builders board and 7.22 of 100 on the coding agent board, and the seconds figure behind it measures one prompt, not a whole build: a product that returns a 200 response in thirty seconds and leaves you to write the queue yourself has not saved you an afternoon. Code ownership is worth only 3 points on the builders board. Code ownership is not scored at all on the coding agent board, because a coding agent edits a repository you already own, so the property was never the vendor's to grant and a full mark measures nothing. The other nine weights are rescaled proportionally so they still sum to 100; their relative importance to each other is unchanged.

Each board's composite is a plain weighted sum with no curve, no normalisation against a moving cohort mean and no editorial adjustment. Take the subscores from any product page, multiply each by its board's published weight, divide by one hundred, and you get the number in the index column to two decimal places. Where two products on the same board sit within about one and a half index points of each other, treat them as tied: five runs is enough for a usable median and not enough for a confident ordering at that distance. When a number here is wrong, we would rather hear it with evidence and correct it in public, which has already happened twice. The whole dataset, including every individual run, is available under a CC BY 4.0 licence for anyone who wants to check the arithmetic or disagree with the weights. Derived, not separately measured. It is nine times the median prompt latency, converted to minutes. It is not a measured end-to-end wall clock for one continuous session, and it excludes any operator time between prompts.