The lab

About LLM Tier

LLM Tier is an independent measurement lab. We build the same application with every AI app builder we cover, we record what happens on every attempt, and we publish the raw numbers rather than an opinion.
Products measured 41Runs recorded 1,845Last run 14 Aug 2026

Test spec

vibeOps

One multi tenant SaaS build, identical for every product

Prompts per run

9

Fixed wording, fixed order, fixed pass criteria

Runs per product

5

We publish the full spread, not only the mean

Licence

CC-BY-4.0

Reuse the numbers with attribution

What we do

Measurement first, commentary second

Most coverage of AI app builders is a screenshot of a landing page and a paragraph of enthusiasm. That tells you nothing about whether the tool can hold a tenancy boundary across nine prompts, whether it produces a page a crawler can read with JavaScript turned off, or whether the second attempt looks anything like the first. So we built a harness instead.

The harness drives one specification, vibeOps, a multi tenant SaaS with an API, webhooks, background jobs and an admin panel. Every product gets the same nine prompts in the same order, 5 times. We time each prompt, we mark it pass, partial or fail against criteria written before the run started, and we keep the artefact so we can measure the shipped output: HTML bytes with JavaScript disabled, largest contentful paint, layout shift, what the API surface actually exposes.

Those measurements roll into ten weighted axes and a single composite index from 0 to 100 with two decimals. The weights are published on the methodology page, the subscores are published on every product page, and the arithmetic is deliberately simple enough that you can recompute any index by hand from the numbers we show you. If our sum does not match yours, one of us has made a mistake and we want to hear about it.

What we will not do

We do not rate products with stars. A star rating compresses a dozen unrelated properties into one blunt symbol and invites the reader to skip the detail. We publish the detail. We do not accept payment for a position, a review or a mention. We do not show a vendor its results before publication, because a preview is an invitation to negotiate. We do not quietly restate an old measurement as a current one, which is why every product page carries a last run date and a next scheduled run date pulled from the database rather than typed into the template.

Where the numbers can be wrong

A benchmark is a model of reality and every model leaks. Our runs are a snapshot of a moving target: these products ship weekly, and a result from last month may describe software that no longer exists. Our spec rewards the kind of application it describes, so a product tuned for marketing sites will look worse here than it would in a marketing site test. Our pass criteria are judgements, written down in advance and published, but judgements nonetheless. We would rather state all of this plainly than imply a precision we do not have.

Corrections

If a figure is wrong, tell us. Send the product, the figure, the page it appears on and the evidence to lab@llm-tier.com or through the contact form. We re run the affected measurement rather than editing a number in place, we update the row, and we record the change in the lab notes so the history stays visible.

Common questions

Who runs LLM Tier?

LLM Tier is an independent measurement lab. The lab is funded by data licensing and by readers. No vendor pays for placement, and no vendor sees a result before publication.

Do vendors pay to appear in the index?

No. There is no paid inclusion, no paid ranking and no sponsored row. A product enters the index when it can complete the public test spec, and it is measured on the same terms as every other product.

Can I reuse the numbers?

Yes. Every figure on this site is published under CC-BY-4.0. Use the JSON and CSV endpoints on the data page, keep the attribution, and link back to the page you took the figure from.

How do I challenge a measurement?

Write to lab@llm-tier.com or use the contact form with the product, the exact figure and the evidence. Corrections are applied to the affected rows and noted in the lab notes.

Reach the lab