What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
The paper adapts convergent and discriminant validity from the social sciences into a procedure for interrogating whether AI benchmarks measure the concepts they claim to measure, applying it to 56 capability and safety benchmarks across 53 models. Benchmarks are labelled with a shared assigned concept (for example refusal, over-refusal, bias, toxicity, reasoning) and the authors test whether model rankings on same-concept benchmarks correlate more strongly than rankings on different-concept benchmarks, repeating the analysis at item level with item response theory. Correlations between model rankings on benchmarks assigned the same safety concept are often weak, which the authors read as inconsistent conceptualisation across benchmarks.
Publisher
arXiv (University of Michigan; Stanford University; Yale University; Microsoft Research; Abridge; Cornell Tech)
Published
8 Sept 2026
Added
today
Key Findings
- 56 benchmarks (capability and safety) evaluated across 53 models
- Same-concept safety benchmarks (refusal, over-refusal, bias, toxicity) often show weak rank correlations with one another
- Capability-concept benchmarks often correlate as strongly across different concepts as within a concept
- Item-level analysis using item response theory corroborates the ranking-level result
- Safety benchmarks examined include XSTest, OR-Bench, WildGuard, HarmBench, SORRY-Bench, SALAD-Bench, BBQ, TruthfulQA, RealToxicityPrompts and WMDP
Methodology Notes
Preprint, v1 posted 8 September 2026 (17 pages plus 28 pages of appendix); affiliations from the PDF title block: University of Michigan, Stanford, Yale, Microsoft Research, Abridge, Cornell Tech. Correlational validity analysis over public benchmark results; none of the 56 benchmarks is a conversational mental-health or crisis benchmark, so the payload is methodological. Verified at the arXiv abstract page (HTTP 200; title, eleven authors and date matched); the bench beat also read the PDF. Missed by the 09-08 and later arXiv passes.
Sources
Authors
Meera Desai, Sang T. Truong, Hanna Wallach, Alex Chouldechova, A. Feder Cooper, Jean Garcia-Gathright, Daniel E. Ho, Abigail Z. Jacobs, Sanmi Koyejo, Nicholas Pangakis, Angelina Wang
Tags
Cite This
APA
Meera Desai et al. (2026). What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks. arXiv (University of Michigan; Stanford University; Yale University; Microsoft Research; Abridge; Cornell Tech). https://arxiv.org/abs/2609.08812