Skip to main content
Benchmark / dataset Credible

RealityTest: How People Probe AI Identity and Whether Models Disclose It

A multimodal, multilingual benchmark testing whether AI systems disclose their identity when users ask, built on data about how people actually raise the question rather than on machine-generated probes. The released dataset holds 3,152 identity-probing queries collected from roughly 750 participants across 49 countries and five languages, in text and speech scenarios, and 17 text models and 6 speech models are evaluated against it.

Publisher

arXiv (AI Security Institute, UK Department for Science, Innovation and Technology; University of Oxford)

Published

29 May 2026

Added

today

Key Findings

  • Only 31% of people ask about identity directly in ambiguous scenarios, and the questions people actually ask are far more varied than machine-generated query sets
  • Disclosure behaviour varies substantially across the 17 text and 6 speech models tested
  • A single suppression instruction drives disclosure below 30% even in the best-performing models
  • How the question is phrased and the surrounding conversational context matter more for whether disclosure happens than which model is being tested
  • The authors conclude that safety evaluations built on narrow or synthetic query sets risk mischaracterising how models behave in realistic deployment

Methodology Notes

Human-grounded data collection from approximately 750 participants across 49 countries and five languages, covering both text and speech scenarios; 17 text models and 6 speech models evaluated. Nine pages, four figures, CC BY 4.0, primary listing cs.CL. Not peer reviewed. Date note: the AI Security Institute's own listing page dates this 2026-06-08, which is its posting date; the arXiv v1 submission date of 2026-05-29 is used here. Curator-verified at the arXiv abs page and through the arXiv API record (title, all five authors, submission date, licence and full abstract read directly).

Sources

Authors

Anna Gausen, Sarenne Wallbridge, Bessie O'Dell, Christopher Summerfield, Hannah Rose Kirk

Tags

ai-disclosureidentity-probingmultilingualspeech-modelsuk-aisieu-ai-act-article-50

Cite This

APA

Anna Gausen et al. (2026). RealityTest: How People Probe AI Identity and Whether Models Disclose It. arXiv (AI Security Institute, UK Department for Science, Innovation and Technology; University of Oxford). https://arxiv.org/abs/2606.00168