Teardowns
Teardown: The Leaderboard Is a Product. You Are Not the User.

In the months before Llama 4 launched, researchers identified 27 private variants of Meta models being tested on Chatbot Arena, the public leaderboard that a large share of the industry treats as the scoreboard for AI. Private variants are not cheating under the Arena's rules. That is the point. The mechanic is legal, available to any lab with the capacity, and invisible in the final ranking, where only the best-performing variant needs to surface.
This is the second teardown in a series that takes apart one AI product mechanic at a time. The mechanic today is the public leaderboard, and the thesis is falsifiable: a leaderboard is a product whose paying users are the labs being ranked, so its mechanics optimize for lab participation, and rank converts into a marketing signal rather than a shipping signal. If leaderboard rank reliably predicted deployed product outcomes across domains, this thesis would be wrong.
The mechanics serve the participants, not the readers
The Leaderboard Illusion study documents three mechanics worth reading as product decisions. First, best-of-N disclosure: providers can test many private variants and let only the winner appear, which statistically inflates the visible score. Second, unequal data access: the study estimates Google and OpenAI each received about a fifth of all Arena data, while dozens of open-weight models shared a much smaller pool. Third, asymmetric retention: models can be silently deprecated, and the study argues this hits open-weight providers hardest.
The fair reading requires the other side. LMArena's official response disputes the framing and some numbers, noting that its testing policy is open to any provider with capacity and that open models received a larger share of battles than the study's method counted, 40.9 percent by the Arena's own April 2025 figures. The dispute itself is instructive: both sides have incentives. The study's lead authors work at a competing lab. The platform is grading its own fairness. Neither fact settles the argument, and a product manager should read both documents the way they read a vendor benchmark, which is to say with the incentive diagram drawn first.
What neither side disputes is the mechanic. Private variant testing with selective disclosure exists, is unequally exercised, and shapes the number at the top of the page. A ranking produced under those rules is a legitimate answer to "which lab played this venue best." It is not an answer to "which model should process your documents."
I have watched the aggregate number lie before
My interpretation comes from operating a much smaller scoreboard. Our document-automation portfolio once gated releases on an offline eval score, our own private leaderboard of one. It had the same failure mode in miniature: the number was real, the improvement was real, and it had near-zero correlation with case outcomes. A candidate that beat the suite lost on three of the five highest-volume document segments, and we only caught it by replaying twelve weeks of production traffic. The score was answering a question nobody was asking.
The fix was not a better leaderboard. It was changing what the number gates. Replay agreement against production, segmented, became the condition for shipping. The offline score kept existing as a fast development signal, which is the job it was always qualified for. The sanitized brief is published in the shadow-mode eval decision doc, and you can play the same gate decisions yourself.
The gate audit, in five questions
The artifact I use before trusting any evaluation, public or private, with a decision:
The gate audit, in five questions
- Selection: who decides which attempts get scored, and can losers stay invisible?
- Distribution: does the eval's traffic look like your traffic, segment by segment?
- Stakes: what does a win actually permit: attention, budget, or shipping to users?
- Symmetry: do the same rules bind every participant, including retention and retries?
- Accountability: who owns the consequences when the number is high and the product is wrong?
Run the Arena through it and the verdict is not "corrupt." It is "miscast." A community preference poll with open participation is a genuinely useful scouting instrument, and the Arena's own response commits to better transparency. The failure happens downstream, in the teams that let a scouting instrument make shipping decisions because the number was public and the font was large.
The counterargument is that nothing better exists
The strongest objection: public leaderboards are the only evaluation most teams can afford, imperfect rankings beat vibes, and the Arena at least measures real human preference at scale. All true. The boundary condition of my thesis is the decision being gated. For choosing which three models to shortlist, rank is cheap and adequate. For deciding what ships to your users, on your traffic, with your failure costs, no external scoreboard can carry the weight, because none of them contains your distribution or your consequences.
What would change my conclusion: evidence that leaderboard rank predicts deployed outcomes on specific product workloads better than small domain evals do. The study most likely to produce that evidence would itself have to pass the audit above.
What to do with this
- Reclassify every external benchmark in your decision docs as a scouting signal, and write down what it is allowed to gate.
- Before adopting any eval as a gate, run the five-question audit and attach it to the decision.
- Build one replay corpus from your own traffic, however small. Twelve weeks of production beats two million battles that are not yours.
Leaderboards will keep being products, and labs will keep being their best customers. That is fine, as long as your release pipeline knows the difference between a venue and a gate.
Previous teardown: AI search ships answers, the missing feature is doubt.
Sources. The Leaderboard Illusion, Singh et al., arXiv 2504.20879, April 2025 (accessed August 17, 2026); lead authors are affiliated with Cohere, a competing model provider. LMArena's response, April 2025 (accessed August 17, 2026); the platform disputes parts of the study's methodology, including the open-model data share. Operator anchor: the shadow-mode replay gate documented in the eval decision doc, owner-verified. The platform has continued revising its methodology since this exchange, rebranding as Arena in early 2026; the mechanics analyzed here are as documented in the April 2025 study and response.


