Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
A preprint examining how large language models handle psychological distress when it is entangled with delusional beliefs, using matched multi-turn simulations across clinically grounded personas and six LLMs. The study isolates the effect of delusional framing by pairing each delusional conversation with a distress-only control, finding that models detect distress at similar rates regardless of framing but sharply fail to intervene once distress is embedded in delusion.
Key Findings
- Identifies a 'recognition-intervention gap': models detect user distress at comparable rates whether or not delusional framing is present, but safety interventions are suppressed by up to 4.5x when distress is entangled with delusion
- The failure tracks the model's accumulated acceptance of the user's premises over the conversation, rather than simple emotional validation
- Prompting models to explicitly assess user distress backfires under delusional framing, worsening rather than improving intervention rates
- Only delusion-aware prompting with explicit response guidance closes the gap, and this remains dependent on a delusion classifier that is itself unreliable on the most vulnerable models
- Argues delusional framing should be treated as a distinct risk signal that overrides ordinary conversational accommodation
Methodology Notes
Matched multi-turn simulation study across six LLMs and clinically grounded personas, pairing delusional and distress-only conditions to isolate framing effects. Preprint, not yet peer-reviewed.
Sources
arXiv abstract (primary)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Andrew Aquilina, Chetna Nihalani, Vasudha Varadarajan, Nathan S. Fishbein, Yu-Ru Lin, Maarten Sap
Tags
Cite This
APA
Andrew Aquilina et al. (2026). Lost in Delusion: Examining LLM Safety Under User Delusions and Distress. arXiv preprint. https://arxiv.org/abs/2606.00975