Your Benchmark Leaderboard Is Measuring the Test, Not the Model. Stanford Checked 56 of Them.
A Stanford HAI team borrowed a century-old psychometric method and asked whether AI benchmarks measure what they claim. Safety tests disagree with each other, "reasoning" and "knowledge" benchmarks barely differ, and one bias benchmark correlates more strongly with reasoning tests.

Every model purchase, every agent platform evaluation, every procurement spreadsheet starts from the same place: a benchmark number. A Stanford HAI team has now asked the question that should have been asked years ago. Do those numbers measure what they claim to measure? For 56 capability and safety benchmarks, the answer is: often, no.
The paper is "What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks" (arXiv:2609.08812, submitted 8 September 2026), from Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, and Angelina Wang. Stanford HAI announced the work on 5 October 2026, with the blunt framing that benchmark scores influence which AI models get funded, bought, and regulated.
The method is borrowed from the social sciences, where psychometricians solved this problem decades ago for tests of human abilities. It rests on two ideas. Convergent validity: tests that claim to measure the same concept should agree with each other. Discriminant validity: tests that claim to measure different concepts should disagree more. The team labeled benchmarks with substantively similar purported concepts to a shared "assigned concept", then checked whether model rankings on benchmarks sharing a concept correlate more strongly than rankings on benchmarks with different concepts. They ran the analogous check at the item level using item response theory. The dataset of model outputs and scores, at item and benchmark level, was released to support future work.
Finding 1: Safety benchmarks do not agree with each other
The most consequential finding. Correlations between model rankings on benchmarks that share the same assigned safety concept are often weak. In plain terms: two benchmarks that both claim to measure, say, refusal or bias will rank your candidate models differently, and there is no principled way to know which ranking to trust.
This matters more than the capability findings because safety benchmarks feed deployment decisions. If one safety test clears a model for production and an ostensibly equivalent test does not, the choice of benchmark, not the model, determines what ships. The paper suggests these safety concepts may be conceptualized inconsistently across benchmarks, which is the polite phrasing. The impolite phrasing is that the field has been arguing about safety scores without agreeing on what safety means.
Finding 2: Capability benchmarks do not discriminate
For assigned capability concepts like reasoning and knowledge, model rankings are often as strongly correlated among benchmarks with the same concept as between benchmarks with different concepts. If a "reasoning" benchmark ranks models almost identically to a "knowledge" benchmark, the categories are not telling you anything distinct. Your leaderboard may be measuring one underlying thing, probably general model quality, while wearing five different costumes.
This is the finding that should make evaluation-minded practitioners wince. We have built an entire model-selection apparatus, leaderboards, procurement rubrics, press-release score comparisons, on the premise that these labels carve reality at its joints. The paper says they mostly do not.
Finding 3: Sometimes the test format matters more than the concept
In some cases, benchmarks that share design elements, like the format their scores take, correlate more strongly than benchmarks that share the same assigned concept. The instrument is doing the measuring, not the model. If two benchmarks agree with each other mainly because they are built the same way, their agreement is evidence of shared construction, not shared validity.
For practitioners, this reframes a common shortcut: "three benchmarks agree, so the result is robust." If all three use the same question format and scoring scheme, the agreement tells you about the format, not the model.
Finding 4: The bias benchmark that measures reasoning
The paper's exhibit A: BBQ-accuracy, a benchmark assigned the concept of bias, correlates more strongly with benchmarks labeled reasoning than with benchmarks sharing its own assigned concept. Read that again. A test that purports to measure bias behaves statistically like a reasoning test. Teams using BBQ scores to compare models on bias may be, unknowingly, comparing them on reasoning instead.
This is the falsifiable center of the paper. It is not a philosophical claim about what bias "really is". It is a correlation result on 53 models, checkable by anyone with the released dataset. Some individual benchmarks correlate more strongly with benchmarks assigned a different concept than with their own concept's benchmarks. That is not a subtle statistical artifact. It is the test flunking its own label.
What this means for data and AI practitioners
- Stop selecting models on a single benchmark. If safety tests disagree with each other and capability labels barely discriminate, one number cannot carry a procurement decision. Use benchmark suites, and weight benchmarks whose rankings diverge, since agreement may just be shared format.
- Audit what your evals actually measure before you trust them. This is the convergent/discriminant check, and you can run a cheap version: rank your candidate models on every benchmark you care about, then check whether tests that claim to measure the same thing rank them similarly. If they do not, your "suite" is not measuring what you think.
- Beware the confident press release. Model vendors report the benchmark where they look best. A leaderboard position on a test that correlates more with reasoning than with its own label is not evidence of progress on the labeled capability.
- Regulators should read this before writing rules that reference benchmarks. Stanford HAI's announcement says it plainly: benchmark scores influence what gets funded, bought, and regulated. Rules that name specific benchmarks inherit this paper's findings.
- The 100-line eval harness still applies. Our agent eval harness tutorial made the case for fixed case sets and regression tracking. That advice gets stronger, not weaker: your own task-grounded evals measure your task, while public benchmarks may be measuring their own format.
In favour
The method is a genuine contribution. Convergent and discriminant validity are the standard tools for exactly this question, and nobody had applied them systematically at this scale: 56 benchmarks, 53 models, both benchmark-level and item-level analysis. Releasing the full dataset of item-level outputs and scores means the findings are checkable, and the BBQ example is the kind of concrete, falsifiable claim that separates a real audit from a hot take. The timing matters too: the work landed while regulators and procurement teams are actively building processes on top of benchmark numbers.
Against
The "assigned concepts" are the researchers' own labels, and labeling is itself a judgment call. If a benchmark was mislabeled, some of the weak convergent validity could be misclassification rather than mismeasurement. Rank-correlation methods are also coarse: two benchmarks can measure the same concept through different difficulty profiles and produce different rankings without either being wrong. And this is an arXiv preprint from September 2026, not yet peer reviewed. The abstract does not report effect sizes or significance levels, so I cannot tell you how much of "often weak" is a little weak versus dramatically weak. The item-level IRT analysis is doing a lot of the argumentative work, and I have not read its details.
What would make me wrong
Four things would change my read. First, if the full paper shows the weak correlations are small and concentrated in a few benchmarks, this becomes a cleanup job rather than a crisis. Second, if re-labeling a handful of mis-assigned benchmarks restores convergent validity, the problem is taxonomy, not measurement. Third, if item-level IRT analysis shows the items discriminate well even when benchmark-level rankings do not, the headline finding softens. Fourth, if a replication on a different model cohort reproduces the BBQ result specifically, the crisis reading hardens. Until then, the prudent assumption is that your favorite benchmark measures something adjacent to its label.
Sources
- Desai, M. et al. "What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks." arXiv:2609.08812 (8 Sep 2026). https://arxiv.org/abs/2609.08812
- Stanford HAI announcement coverage: https://theagenttimes.com/articles/stanford-hai-studies-find-ai-benchmarks-often-fail-to-measur-d81f303f
Related on Everyday Data Science: the pandas/Polars/DuckDB benchmark we ran ourselves, the tabular foundation models deep dive with its TabArena Elo scoreboard, and the 100-line agent eval harness for building evals you can actually trust.
Which benchmark do you trust the most for your own work? What would it take for you to stop trusting it?
About the writer
Data Scientist & AI Researcher
Data scientist and AI researcher at Pace University. I coined Artificial Frictional Unemployment, and built the first machine learning model for crop yield prediction in Sierra Leone. Author of Understanding Agentic AI. I write about agentic systems and applied ML, with a bias toward what actually works, and who gets left out when it doesn't.
Found this useful? Passing it on to someone who builds is the best way to help the publication grow.
Built something worth sharing? Write it up for us →