news / 22 Aug 2026 / 3 min read

The model board carried a composite index with no run behind it. It is gone.

The eight rows on /leaderboard/llms carried an Index column that looked exactly like the composite this lab actually measures on the other two boards. It was not: no per-model run record, subscore or recomputed field sat behind any of the eight numbers, and Arena's own leaderboard shows all eight rows were two to three generations stale. The index is removed, not refreshed. What replaces it is a 26-row sourced specification-and-benchmark table, with the twelve superseded rows kept in a separate table rather than deleted.

By Daniel Osei | Updated 22 Aug 2026

What we found

The language model board at /leaderboard/llms published eight rows, each carrying an Index figure in the same visual position and the same numeric format as the composite index this lab computes for full-stack app builders and coding agents. That placement was misleading by itself: our composite is a weighted sum of subscores derived from recorded run executions, published with the weights and the formula next to it, reproducible by anyone with a calculator. Nothing of the kind sat behind the model board's Index column. There is no model-level run table, no per-model subscore, no weight sheet, and no recompute path for any of the eight numbers. We looked, because we had to check before writing this, and there is nothing there.

The eight rows were also stale, not just unsupported. We pulled Arena's WebDev leaderboard, at arena.ai/leaderboard/code/webdev, as it stood on 21 August 2026: 603,789 votes across 118 ranked models. Mapped against that board, the eight models we had published sat at ranks 35, 42, 52, 74, 82, 86 and 89, two to three model generations behind current releases from their own vendors. Arena's own top seven models were not represented anywhere on our board at all. A reader landing on our page had no way to know any of that, because the page did not cite Arena, or anything else, for the number it was showing.

Why removal, not a refresh

We considered recomputing the eight numbers against fresher inputs and leaving the column in place. We rejected that. A refreshed number in the same column would still imply a harness behind it, and there has never been one: this lab has no model-level run record to point to, for any model, in any cycle. Presenting a figure as measured when it was never measured is the same defect whether the figure is current or three generations stale. The honest fix is to stop publishing a number we cannot show the work for, not to publish a newer version of the same unsupported number.

What replaces it

The board is rebuilt as a sourced specification-and-benchmark table. Every cell is now one of three things: a vendor-stated fact (context window, price, licence, modality), a named third-party benchmark with its own methodology and its own attribution (Arena WebDev Elo, the Artificial Analysis Intelligence Index v4.1.1, Terminal-Bench v2.1), or an explicit blank where nothing sourced is available. None of it is our own estimate, average, or aggregate score, and there is no Index or Rank column anywhere on the page.

The rebuilt board carries 26 rows across 11 vendors, ordered by Arena WebDev Elo, which is Arena's measurement and is labelled as such, not ours. A second, visually separate table on the same page holds twelve superseded rows, including every model the old board used to show, so the previous board is auditable rather than quietly deleted. Full attribution, benchmark provenance, a pricing-error warning carried over from Arena's own listing, and the coverage gap this table does not close, including several open-weight and proprietary models ahead of our listed flagships that we have not sourced yet, are all stated directly on the page and on /methodology.

We would rather publish 26 rows we can source completely than 40 with an estimate in half the cells. If you can point us at a primary source for a model or a price we have marked not published, the contact address is on the contact page.

More lab notes