Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness

Introduces a browser extension that flags concerning chatbot behaviour (overconfidence, sycophancy, anthropomorphism, persuasive influence and related classes) inline in ChatGPT and Claude conversations, grounded in evidence spans, and evaluates it in a two-week field study with 45 frequent chatbot users. Participants found the nudges useful and minimally disruptive and nearly all reported greater awareness of AI harms, but awareness alone did not produce discernible behaviour change.

Publisher

arXiv (Carnegie Mellon University)

Published

22 Sept 2026

Added

today

DOI

Key Findings

  • Two-week field deployment with N=45 Prolific-recruited frequent chatbot users (from 419 screened), using ChatGPT or Claude in their normal workflows
  • Nearly all participants reported increased awareness of potential AI harms; the authors found no discernible behavioural change from awareness alone
  • Participants used chatbots weekly most often for personal or life advice (75.6%) and work (68.9%), then casual conversation (62.2%), career advice (55.6%) and school tasks (35.6%)
  • The named issue taxonomy covered 96.4% of 111 human annotations in a coverage check; only 1.4% of LLM-identified issues in the field study fell in 'other'
  • Design recommendations centre on relevance, calibration and user control of nudges

Methodology Notes

Browser extension with adapters for ChatGPT and Claude; review model routed to a different provider (gpt-5-mini reviewing Claude, claude-sonnet-4.6 reviewing ChatGPT) or user-selected; interaction logs, surveys and per-nudge feedback; mixed-methods analysis. v1 posted 2026-09-22; project site open-reflection.com.

Authors

Varshini Elangovan, James Wedgwood, Chhavi Yadav, William Agnew, Sauvik Das, Virginia Smith

Tags

nudgesbrowser-extensionfield-studysycophancyanthropomorphism

Cite This

APA

Varshini Elangovan et al. (2026). Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness. arXiv (Carnegie Mellon University). https://arxiv.org/abs/2609.26865