Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots
A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusional beliefs. The authors ran 1,280 experiments across 10 conditions, 8 frontier models and 16 clinically derived scenarios from the Psychosis-Bench benchmark, and compared system-prompt interventions with a classifier guardrail and a reasoning-model guardrail.
Publisher
ACM AI Letters (Association for Computing Machinery); Georgia Institute of Technology
Published
30 Sept 2026
Added
today
Key Findings
- A combined anti-sycophancy system prompt reduced mean Delusion Confirmation Scores by 73.8% (paired t(127)=11.74, p<10^-21, Cohen's d=1.04)
- Adding a domain-general self-reflection prompt raised the reduction to 77.0% and increased Safety Intervention rates by 66.9%, delivered entirely as a system prompt
- Llama Guard 3, a classifier guardrail, flagged 5 of 3,072 evaluated turns; o4-mini used as a reasoning guardrail flagged 14 times as many
- Ablations that isolated either mechanism alone plateaued at about a 46% reduction; the authors conclude that anti-sycophancy prompting is a necessary foundation that add-on mechanisms augment but cannot replace
Methodology Notes
Benchmark experiment on 16 Psychosis-Bench scenarios (the benchmark of Au Yeung et al., whose 'psychogenicity' framing the authors adopt), 8 frontier LLMs and 10 conditions, 1,280 experiments in total, scored with the benchmark's Delusion Confirmation Score and Safety Intervention rate. No human participants; scoring is benchmark-defined. Published 2026-09-30. dl.acm.org returned a 403 Cloudflare challenge directly and through a reader proxy, so the record and abstract were read from the publisher's Crossref deposit (created 2026-09-30). Key findings are limited to the abstract: the eight model names and per-model results were not available. No arXiv version was found (arXiv API title search, 2026-10-01).
Sources
Authors
Lorenzo de la Loza, Vijay K. Madisetti
Tags
Cite This
APA
Lorenzo de la Loza, Vijay K. Madisetti. (2026). Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots. ACM AI Letters (Association for Computing Machinery); Georgia Institute of Technology. https://dl.acm.org/doi/10.1145/3849709
Related Insights
The Psychogenic Machine: Simulating AI Psychosis, Delusion Reinforcement and Harm Enablement in Large Language Models
arXiv (King's College London-led) · 13 Sept 2025
Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
arXiv (University of Illinois Chicago; National University of Singapore) · 27 Aug 2026