Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks

The paper adapts convergent and discriminant validity from the social sciences into a procedure for interrogating whether AI benchmarks measure the concepts they claim to measure, applying it to 56 capability and safety benchmarks across 53 models. Benchmarks are labelled with a shared assigned concept (for example refusal, over-refusal, bias, toxicity, reasoning) and the authors test whether model rankings on same-concept benchmarks correlate more strongly than rankings on different-concept benchmarks, repeating the analysis at item level with item response theory. Correlations between model rankings on benchmarks assigned the same safety concept are often weak, which the authors read as inconsistent conceptualisation across benchmarks.

Publisher

arXiv (University of Michigan; Stanford University; Yale University; Microsoft Research; Abridge; Cornell Tech)

Published

8 Sept 2026

Added

today

Key Findings

  • 56 benchmarks (capability and safety) evaluated across 53 models
  • Same-concept safety benchmarks (refusal, over-refusal, bias, toxicity) often show weak rank correlations with one another
  • Capability-concept benchmarks often correlate as strongly across different concepts as within a concept
  • Item-level analysis using item response theory corroborates the ranking-level result
  • Safety benchmarks examined include XSTest, OR-Bench, WildGuard, HarmBench, SORRY-Bench, SALAD-Bench, BBQ, TruthfulQA, RealToxicityPrompts and WMDP

Methodology Notes

Preprint, v1 posted 8 September 2026 (17 pages plus 28 pages of appendix); affiliations from the PDF title block: University of Michigan, Stanford, Yale, Microsoft Research, Abridge, Cornell Tech. Correlational validity analysis over public benchmark results; none of the 56 benchmarks is a conversational mental-health or crisis benchmark, so the payload is methodological. Verified at the arXiv abstract page (HTTP 200; title, eleven authors and date matched); the bench beat also read the PDF. Missed by the 09-08 and later arXiv passes.

Authors

Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang

Tags

benchmark-validityconvergent-validityirtrefusalover-refusalmicrosoft-researchstanfordcoverage-miss

Cite This

APA

Meera Desai et al. (2026). What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks. arXiv (University of Michigan; Stanford University; Yale University; Microsoft Research; Abridge; Cornell Tech). https://arxiv.org/abs/2609.08812