# LLM Tier > Independent measurement lab for AI app builders, coding agents and LLMs. Every product is driven through the same published build spec, vibeOps, 5 times with 9 fixed prompts per run. We publish the raw per run outcomes, the ten weighted subscores and a composite index from 0 to 100 with two decimals. Licence: CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). Cite as: LLM Tier, https://www.llm-tier.com. Contact: lab@llm-tier.com Scoring: composite index 0 to 100, two decimals. There are no star ratings anywhere on this site. ## Core rankings - [App builder leaderboard](https://www.llm-tier.com/leaderboard/app-builders): full index with speed spread, reliability and per axis subscores for every measured builder. - [LLM leaderboard](https://www.llm-tier.com/leaderboard/llms): model index with context window and token pricing. - [Full comparison matrix](https://www.llm-tier.com/matrix): every measured builder against all ten axes in one grid, winner marked per column. - [Homepage index table](https://www.llm-tier.com): rank, index, speed, reliability, SEO score and rank delta. - [Methodology](https://www.llm-tier.com/methodology): the ten axes, their weights, the nine prompts and the pass criteria. - [Open data](https://www.llm-tier.com/data): endpoints, schema and citation guidance. ## Machine readable data - [Rankings JSON](https://www.llm-tier.com/api/rankings.json): all builders and models with subscores and licence metadata. - [Raw runs JSON](https://www.llm-tier.com/api/runs.json): one row per prompt attempt, filter with ?tool=slug. - [Rankings CSV](https://www.llm-tier.com/data/rankings.csv): flat export with a licence column on every row. ## Measured builders - [Totalum](https://www.llm-tier.com/builders/totalum): index 87.96, rank 1, reliability 93.3 percent, median run 106.3 seconds. Reviews: https://www.llm-tier.com/builders/totalum/reviews - [Claude Code](https://www.llm-tier.com/builders/claude-code): index 80.30, rank 2, reliability 91.1 percent, median run 65.5 seconds. Reviews: https://www.llm-tier.com/builders/claude-code/reviews - [Cursor](https://www.llm-tier.com/builders/cursor): index 78.65, rank 3, reliability 86.7 percent, median run 53.4 seconds. Reviews: https://www.llm-tier.com/builders/cursor/reviews - [Firebase Studio](https://www.llm-tier.com/builders/firebase-studio): index 78.06, rank 4, reliability 86.7 percent, median run 100.3 seconds. Reviews: https://www.llm-tier.com/builders/firebase-studio/reviews - [Lovable](https://www.llm-tier.com/builders/lovable): index 77.87, rank 5, reliability 82.2 percent, median run 85.2 seconds. Reviews: https://www.llm-tier.com/builders/lovable/reviews - [Kiro](https://www.llm-tier.com/builders/kiro): index 76.31, rank 6, reliability 93.3 percent, median run 72.4 seconds. Reviews: https://www.llm-tier.com/builders/kiro/reviews - [Augment Code](https://www.llm-tier.com/builders/augment-code): index 75.92, rank 7, reliability 88.9 percent, median run 71.6 seconds. Reviews: https://www.llm-tier.com/builders/augment-code/reviews - [v0](https://www.llm-tier.com/builders/v0): index 74.69, rank 8, reliability 84.4 percent, median run 65.2 seconds. Reviews: https://www.llm-tier.com/builders/v0/reviews - [Amp](https://www.llm-tier.com/builders/amp): index 74.04, rank 9, reliability 77.8 percent, median run 39.7 seconds. Reviews: https://www.llm-tier.com/builders/amp/reviews - [Replit](https://www.llm-tier.com/builders/replit): index 73.42, rank 10, reliability 86.7 percent, median run 135.7 seconds. Reviews: https://www.llm-tier.com/builders/replit/reviews - [Zed](https://www.llm-tier.com/builders/zed): index 72.17, rank 11, reliability 71.1 percent, median run 31.4 seconds. Reviews: https://www.llm-tier.com/builders/zed/reviews - [Bolt.new](https://www.llm-tier.com/builders/bolt-new): index 71.94, rank 12, reliability 80 percent, median run 82.1 seconds. Reviews: https://www.llm-tier.com/builders/bolt-new/reviews - [Antigravity](https://www.llm-tier.com/builders/antigravity): index 71.91, rank 13, reliability 73.3 percent, median run 78.7 seconds. Reviews: https://www.llm-tier.com/builders/antigravity/reviews - [Factory](https://www.llm-tier.com/builders/factory): index 71.52, rank 14, reliability 82.2 percent, median run 161.7 seconds. Reviews: https://www.llm-tier.com/builders/factory/reviews - [Cline](https://www.llm-tier.com/builders/cline): index 71.51, rank 15, reliability 80 percent, median run 95.1 seconds. Reviews: https://www.llm-tier.com/builders/cline/reviews - [Base44](https://www.llm-tier.com/builders/base44): index 70.54, rank 16, reliability 84.4 percent, median run 94.1 seconds. Reviews: https://www.llm-tier.com/builders/base44/reviews - [Convex Chef](https://www.llm-tier.com/builders/convex-chef): index 70.50, rank 17, reliability 60 percent, median run 88.7 seconds. Reviews: https://www.llm-tier.com/builders/convex-chef/reviews - [Jules](https://www.llm-tier.com/builders/jules): index 68.69, rank 18, reliability 75.6 percent, median run 146.5 seconds. Reviews: https://www.llm-tier.com/builders/jules/reviews - [Softgen](https://www.llm-tier.com/builders/softgen): index 67.30, rank 19, reliability 73.3 percent, median run 108.4 seconds. Reviews: https://www.llm-tier.com/builders/softgen/reviews - [Aider](https://www.llm-tier.com/builders/aider): index 67.27, rank 20, reliability 68.9 percent, median run 39.4 seconds. Reviews: https://www.llm-tier.com/builders/aider/reviews - [Anything](https://www.llm-tier.com/builders/anything): index 67.16, rank 21, reliability 77.8 percent, median run 94.1 seconds. Reviews: https://www.llm-tier.com/builders/anything/reviews - [Qodo](https://www.llm-tier.com/builders/qodo): index 67.13, rank 22, reliability 77.8 percent, median run 99.3 seconds. Reviews: https://www.llm-tier.com/builders/qodo/reviews - [Codebuff](https://www.llm-tier.com/builders/codebuff): index 66.89, rank 23, reliability 68.9 percent, median run 59.8 seconds. Reviews: https://www.llm-tier.com/builders/codebuff/reviews - [Emergent](https://www.llm-tier.com/builders/emergent): index 66.79, rank 24, reliability 80 percent, median run 144.6 seconds. Reviews: https://www.llm-tier.com/builders/emergent/reviews - [Devin](https://www.llm-tier.com/builders/devin): index 66.36, rank 25, reliability 71.1 percent, median run 353.6 seconds. Reviews: https://www.llm-tier.com/builders/devin/reviews - [Floot](https://www.llm-tier.com/builders/floot): index 66.17, rank 26, reliability 68.9 percent, median run 91.1 seconds. Reviews: https://www.llm-tier.com/builders/floot/reviews - [OpenHands](https://www.llm-tier.com/builders/openhands): index 65.18, rank 27, reliability 64.4 percent, median run 144.8 seconds. Reviews: https://www.llm-tier.com/builders/openhands/reviews - [Dyad](https://www.llm-tier.com/builders/dyad): index 64.77, rank 28, reliability 55.6 percent, median run 83.8 seconds. Reviews: https://www.llm-tier.com/builders/dyad/reviews - [Bind AI](https://www.llm-tier.com/builders/bind-ai): index 63.18, rank 29, reliability 64.4 percent, median run 96 seconds. Reviews: https://www.llm-tier.com/builders/bind-ai/reviews - [Trae](https://www.llm-tier.com/builders/trae): index 62.25, rank 30, reliability 51.1 percent, median run 96.7 seconds. Reviews: https://www.llm-tier.com/builders/trae/reviews - [Roomote](https://www.llm-tier.com/builders/roomote): index 61.04, rank 31, reliability 53.3 percent, median run 96.4 seconds. Reviews: https://www.llm-tier.com/builders/roomote/reviews - [Rocket.new](https://www.llm-tier.com/builders/rocket-new): index 60.16, rank 32, reliability 60 percent, median run 107.2 seconds. Reviews: https://www.llm-tier.com/builders/rocket-new/reviews - [Caffeine](https://www.llm-tier.com/builders/caffeine): index 60.10, rank 33, reliability 66.7 percent, median run 150.6 seconds. Reviews: https://www.llm-tier.com/builders/caffeine/reviews - [Tesslate](https://www.llm-tier.com/builders/tesslate): index 59.33, rank 34, reliability 53.3 percent, median run 90.3 seconds. Reviews: https://www.llm-tier.com/builders/tesslate/reviews - [Zite](https://www.llm-tier.com/builders/zite): index 58.78, rank 35, reliability 60 percent, median run 113.4 seconds. Reviews: https://www.llm-tier.com/builders/zite/reviews - [bolt.diy](https://www.llm-tier.com/builders/bolt-diy): index 58.17, rank 36, reliability 55.6 percent, median run 113.8 seconds. Reviews: https://www.llm-tier.com/builders/bolt-diy/reviews - [Pazi](https://www.llm-tier.com/builders/pazi): index 56.79, rank 37, reliability 53.3 percent, median run 120.9 seconds. Reviews: https://www.llm-tier.com/builders/pazi/reviews - [YouWare](https://www.llm-tier.com/builders/youware): index 56.55, rank 38, reliability 48.9 percent, median run 80 seconds. Reviews: https://www.llm-tier.com/builders/youware/reviews - [Mocha](https://www.llm-tier.com/builders/mocha): index 55.67, rank 39, reliability 53.3 percent, median run 105.2 seconds. Reviews: https://www.llm-tier.com/builders/mocha/reviews - [Trickle](https://www.llm-tier.com/builders/trickle): index 55.45, rank 40, reliability 51.1 percent, median run 87 seconds. Reviews: https://www.llm-tier.com/builders/trickle/reviews - [CatDoes](https://www.llm-tier.com/builders/catdoes): index 55.27, rank 41, reliability 48.9 percent, median run 137.1 seconds. Reviews: https://www.llm-tier.com/builders/catdoes/reviews ## Axis leaderboards - [Fastest AI app builders](https://www.llm-tier.com/leaderboard/speed): Median wall clock seconds per vibeOps prompt, with the p10 to p90 spread. - [Most reliable AI app builders](https://www.llm-tier.com/leaderboard/reliability): Strict pass rate across 495 recorded prompt executions. - [Best AI app builders for SEO and GEO](https://www.llm-tier.com/leaderboard/seo-geo): Server rendered HTML bytes with JavaScript disabled, LCP, CLS and structured data output. - [Best AI app builders for API and MCP work](https://www.llm-tier.com/leaderboard/api-mcp): REST surface, webhook delivery, background jobs and Model Context Protocol support. - [Best AI app builders for design output](https://www.llm-tier.com/leaderboard/design): Blind scored visual quality of the generated admin panel and marketing surface. - [Best value AI app builders](https://www.llm-tier.com/leaderboard/value): Index points delivered per euro of entry price. ## Lab notes - [This cycle's composite recompute, and why it barely moved most rows](https://www.llm-tier.com/news/this-cycles-composite-recompute): We re derived every composite index from the ten published axis weights and the stored subscores, row by row, and checked the arithmetic against the methodology page. Here is what changed and what that process actually verifies. - [Windsurf is gone from the index](https://www.llm-tier.com/news/windsurf-is-gone-from-the-index): windsurf.com now redirects to devin.ai/desktop. The product is discontinued, so it has been removed from the roster, the raw runs, every leaderboard, every compare pair and the matrix, not just flagged as inactive. - [August 2026 index update: three products move, one regression](https://www.llm-tier.com/news/august-2026-index-update): The fourth full run of the vibeOps specification against 11 products. Two products improved their reliability materially, one regressed on tenancy isolation. - [Why we publish the spread and not the mean](https://www.llm-tier.com/news/why-we-publish-the-spread-not-the-mean): A single average build time hides the thing you actually care about: how bad the slow runs get. Here is the p10 to p90 data for every product in the index. - [Server rendering is the whole GEO story right now](https://www.llm-tier.com/news/server-rendering-is-the-geo-story): We disabled JavaScript on every generated marketing page and counted the bytes of real text. The gap between the top and the bottom of the cohort is nearly seven to one. - [How we score agent performance without reading the chat log](https://www.llm-tier.com/news/how-we-score-agent-performance): Agent performance is 18 points of the index, the joint largest weight. Here is exactly how a prompt earns 0, 50 or 100. - [Reproducing any row in the index by hand](https://www.llm-tier.com/news/reproducing-the-index-by-hand): The composite is a weighted sum, nothing more. Take the ten subscores off any product page, multiply by the published weights and you get the number in the table. - [What this benchmark cannot tell you](https://www.llm-tier.com/news/what-a-benchmark-cannot-tell-you): A composite index is a compression of a complicated thing into one number. Here are the limits of ours, stated plainly. ## Optional - [About the lab](https://www.llm-tier.com/about): who measures this and how corrections work. - [Full text export](https://www.llm-tier.com/llms-full.txt): every published figure in one plain text file. - [Terms](https://www.llm-tier.com/terms) and [Privacy](https://www.llm-tier.com/privacy).