Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness
Introduces a browser extension that flags concerning chatbot behaviour (overconfidence, sycophancy, anthropomorphism, persuasive influence and related classes) inline in ChatGPT and Claude conversations, grounded in evidence spans, and evaluates it in a two-week field study with 45 frequent chatbot users. Participants found the nudges useful and minimally disruptive and nearly all reported greater awareness of AI harms, but awareness alone did not produce discernible behaviour change.
Publisher
arXiv (Carnegie Mellon University)
Published
22 Sept 2026
Added
today
DOI
—
Key Findings
- Two-week field deployment with N=45 Prolific-recruited frequent chatbot users (from 419 screened), using ChatGPT or Claude in their normal workflows
- Nearly all participants reported increased awareness of potential AI harms; the authors found no discernible behavioural change from awareness alone
- Participants used chatbots weekly most often for personal or life advice (75.6%) and work (68.9%), then casual conversation (62.2%), career advice (55.6%) and school tasks (35.6%)
- The named issue taxonomy covered 96.4% of 111 human annotations in a coverage check; only 1.4% of LLM-identified issues in the field study fell in 'other'
- Design recommendations centre on relevance, calibration and user control of nudges
Methodology Notes
Browser extension with adapters for ChatGPT and Claude; review model routed to a different provider (gpt-5-mini reviewing Claude, claude-sonnet-4.6 reviewing ChatGPT) or user-selected; interaction logs, surveys and per-nudge feedback; mixed-methods analysis. v1 posted 2026-09-22; project site open-reflection.com.
Sources
arXiv abstract (v1)(opens in a new tab) (primary)
Authors
Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith
Tags
Cite This
APA
Varshini Elangovan et al. (2026). Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness. arXiv (Carnegie Mellon University). https://arxiv.org/abs/2609.26865