RealityTest: How People Probe AI Identity and Whether Models Disclose It
A multimodal, multilingual benchmark testing whether AI systems disclose their identity when users ask, built on data about how people actually raise the question rather than on machine-generated probes. The released dataset holds 3,152 identity-probing queries collected from roughly 750 participants across 49 countries and five languages, in text and speech scenarios, and 17 text models and 6 speech models are evaluated against it.
Publisher
arXiv (AI Security Institute, UK Department for Science, Innovation and Technology; University of Oxford)
Published
29 May 2026
Added
2 weeks ago
Key Findings
- Only 31% of people ask about identity directly in ambiguous scenarios, and the questions people actually ask are far more varied than machine-generated query sets
- Disclosure behaviour varies substantially across the 17 text and 6 speech models tested
- A single suppression instruction drives disclosure below 30% even in the best-performing models
- How the question is phrased and the surrounding conversational context matter more for whether disclosure happens than which model is being tested
- The authors conclude that safety evaluations built on narrow or synthetic query sets risk mischaracterising how models behave in realistic deployment
Methodology Notes
Human-grounded data collection from approximately 750 participants across 49 countries and five languages, covering both text and speech scenarios; 17 text models and 6 speech models evaluated. Nine pages, four figures, CC BY 4.0, primary listing cs.CL. Not peer reviewed. Date note: the AI Security Institute's own listing page dates this 2026-06-08, which is its posting date; the arXiv v1 submission date of 2026-05-29 is used here. Curator-verified at the arXiv abs page and through the arXiv API record (title, all five authors, submission date, licence and full abstract read directly).
Sources
arXiv abstract page (2606.00168)(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Anna Gausen, Sarenne Wallbridge, Bessie O'Dell, Christopher Summerfield, Hannah Rose Kirk
Tags
Cite This
APA
Anna Gausen et al. (2026). RealityTest: How People Probe AI Identity and Whether Models Disclose It. arXiv (AI Security Institute, UK Department for Science, Innovation and Technology; University of Oxford). https://arxiv.org/abs/2606.00168
Related Insights
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
arXiv (National Taiwan University; University of Bamberg) · 12 Aug 2026
Guidelines on the implementation of the transparency obligations for certain AI systems under Article 50 of the AI Act
European Commission (EU AI Office) · 20 Jul 2026
People readily follow personal advice from AI but it does not improve their well-being
arXiv (UK AI Security Institute; Limbic AI) · 19 Nov 2025
When Do LLM Preferences Predict Downstream Behavior?
AI Security Institute (UK Department for Science, Innovation and Technology) · 26 Aug 2026
One-shot emergency psychiatric triage across 15 frontier AI chatbots
arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI) · 28 Apr 2026
Privacy assurances and professional-boundary warnings in generative AI mental health chatbots: a randomized vignette experiment on calibrated trust, overreliance risk, and professional help-seeking intentions
Frontiers in Psychology (Frontiers Media) · 1 Sept 2026