Every month brings a chart with a new model at the top, and every board that buys AI is eventually shown one. April 2025 was the month the charts came under scrutiny. Meta's Llama 4 arrived with a leaderboard result that turned out not to describe the model anyone could use, and a paper from a group of academic and industry researchers showed the practice behind it was widespread. This briefing sets out what happened, what the main leaderboards actually measure, and how a business should read one.
What Meta released, and what it ranked
On 5 April Meta published two open-weight models, Llama 4 Scout and Llama 4 Maverick, and previewed a third, Behemoth, still in training. Scout has 109 billion parameters with 17 billion active across 16 experts and claims a ten-million-token context; Maverick has 400 billion with 17 billion active across 128 experts; Behemoth was described at nearly two trillion parameters with 288 billion active.[1] The models ship under Meta's own licence rather than an open-source one, with a cap for companies above 700 million monthly users and, as with Llama 3.2, a restriction on multimodal use for companies domiciled in the EU.[2]
The launch materials pointed to LMArena, where a Maverick model sat second with a score of 1417. Within days it emerged that the version on the leaderboard was an "experimental chat version" optimised for conversationality, not the weights Meta had released. LMArena said publicly that Meta's interpretation of its policy did not match what it expected from providers and changed its rules so that the evaluated model must be the released one. When the public Maverick was tested it ranked around thirty-second, below models many months older. Meta denied training on the test set and said it had experimented with custom variants.[3]
The Leaderboard Illusion
- 27
- Private Llama 4 variants Meta tested on the Arena before the public release, per the paper [4]
- 39.6%
- Share of all Arena comparison data received by Google and OpenAI between them, per the paper's estimate [4]
- ≈1417
- The Arena score of the experimental Maverick build that did not match the released weights [3]
On 29 April a paper titled The Leaderboard Illusion, from researchers at Cohere Labs, Princeton, Stanford, the Allen Institute, MIT and Waterloo among others, set out how Chatbot Arena had come to work. Some providers were allowed to test many private variants and publish only the best, which is a form of selection that inflates a score without improving a model. Meta had tested 27 variants before Llama 4. Proprietary models were sampled far more often than open ones, so Google and OpenAI between them received roughly 40% of all the comparison data, and the authors estimated that 205 of 243 public models had been quietly retired from the comparison pool. Access to that data, they showed, measurably improves a model's Arena score.[4] LMArena disputed parts of the analysis and defended its policies, and the exchange was public.[5]
None of this makes the Arena worthless. It remains the largest public measure of which answers people prefer, and preference is a real thing to measure. The paper's contribution is to say what the number is: a measure shaped by who tests most, who is shown most, and what style people reward, as well as by capability.
Why the static benchmarks stopped working
The Arena's problems are the mirror image of the older benchmarks'. A fixed test with public answers ends up in training data, and a score of 90% on a saturated exam separates nothing. The field's response by early 2025 was to write harder tests and hold the answers back. Epoch AI's FrontierMath, launched in November 2024, used unpublished research-level problems.[6] Humanity's Last Exam, published in January 2025, assembled 2,500 questions from subject experts across dozens of fields, chosen because frontier models failed them.[7] The ARC Prize's ARC-AGI-2, launched in March, was designed so that ordinary people solve it and models do not.[8] Each of these bought time, and each will saturate in turn.
How to read one
- Check the model on the board is the model on sale. The Llama 4 episode was exactly this. Providers now have to say, but a chart from before April 2025 may not.
- Ask what the score measures. Arena is preference in a chat window; a coding benchmark is pass rate on a fixed set of repository tasks; neither is your invoicing workflow.
- Prefer boards that hold answers back and publish their method, and treat a score quoted by the vendor as the vendor's.
- Look at the spread, not the rank. When the top five sit within a few points, the ranking is noise and the price and the licence are the decision.
- Run your own. Twenty representative tasks from your own work, scored by your own people, beat any public board for the question of whether a model will do your job.


