Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots

A letter in ACM AI Letters from the Georgia Institute of Technology tests whether prompting interventions designed to reduce sycophancy also reduce large language models' endorsement of users' delusional beliefs. The authors ran 1,280 experiments across 10 conditions, 8 frontier models and 16 clinically derived scenarios from the Psychosis-Bench benchmark, and compared system-prompt interventions with a classifier guardrail and a reasoning-model guardrail.

Publisher

ACM AI Letters (Association for Computing Machinery); Georgia Institute of Technology

Published

30 Sept 2026

Added

today

Key Findings

  • A combined anti-sycophancy system prompt reduced mean Delusion Confirmation Scores by 73.8% (paired t(127)=11.74, p<10^-21, Cohen's d=1.04)
  • Adding a domain-general self-reflection prompt raised the reduction to 77.0% and increased Safety Intervention rates by 66.9%, delivered entirely as a system prompt
  • Llama Guard 3, a classifier guardrail, flagged 5 of 3,072 evaluated turns; o4-mini used as a reasoning guardrail flagged 14 times as many
  • Ablations that isolated either mechanism alone plateaued at about a 46% reduction; the authors conclude that anti-sycophancy prompting is a necessary foundation that add-on mechanisms augment but cannot replace

Methodology Notes

Benchmark experiment on 16 Psychosis-Bench scenarios (the benchmark of Au Yeung et al., whose 'psychogenicity' framing the authors adopt), 8 frontier LLMs and 10 conditions, 1,280 experiments in total, scored with the benchmark's Delusion Confirmation Score and Safety Intervention rate. No human participants; scoring is benchmark-defined. Published 2026-09-30. dl.acm.org returned a 403 Cloudflare challenge directly and through a reader proxy, so the record and abstract were read from the publisher's Crossref deposit (created 2026-09-30). Key findings are limited to the abstract: the eight model names and per-model results were not available. No arXiv version was found (arXiv API title search, 2026-10-01).

Authors

Lorenzo de la Loza, Vijay K. Madisetti

Tags

psychosis-benchanti-sycophancysystem-promptllama-guarddelusion-confirmationgeorgia-tech

Cite This

APA

Lorenzo de la Loza, Vijay K. Madisetti. (2026). Programmed Interventions To Prevent Delusions From Excessive Use of Conversational AI Bots. ACM AI Letters (Association for Computing Machinery); Georgia Institute of Technology. https://dl.acm.org/doi/10.1145/3849709