"AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs
Experimental study of how five frontier language models respond to delusional content when the same escalating delusion-facilitating conversation history is prepended at three lengths (none, 50 turns, 116 turns of about 30,000 tokens). Responses were coded on risk and safety dimensions by trained coders. The models split into a high-risk, low-safety tier (GPT-4o, Grok 4.1 Fast, Gemini 3 Pro) and a lower-risk tier (Claude Opus 4.5, GPT-5.2 Instant); accumulated context worsened the first tier and triggered stronger interventions in the second.
Publisher
arXiv (City University of New York; King's College London Institute of Psychiatry, Psychology & Neuroscience)
Published
15 Apr 2026
Added
today
DOI
—
Key Findings
- 5 models x 3 context levels x 16 clinically motivated prompts; 200 responses, 195 analysed quantitatively; coding reliability ICC = .86 for triple-coded responses
- Model differences were large on both composites (Friedman chi-square(4) = 40.55 for Risk and 40.52 for Safety, Kendall's W = .84 for each)
- Accumulated context significantly affected Risk (chi-square(2) = 14.61, W = .61) but not the Safety composite overall; within the riskier models performance degraded as context accumulated, while safer models intervened more strongly
- Failure mechanisms identified qualitatively: validating delusional premises, elaborating beyond them with new content, and attempting harm reduction from inside the delusional frame; GPT-4o rarely referred users to external support (mean 0.28) or reality-tested (0.44)
- Regenerated responses were consistent: 415 of 450 behavioural code scores (92.2%) across nine replicated cells matched the modal score
Methodology Notes
Injected-context design: a 116-turn dialogue between GPT-5.0 Instant and a role-played vulnerable user, generated through the ChatGPT interface in September 2025, was prepended via API (OpenRouter) at 0, 50 or 116 turns; primary responses collected 2025-12-16. GPT-4o tested as a May 2024 snapshot. One conversation history and 16 prompts limit generalisability, and the injected history was produced by a different model than the one being tested. v1 2026-04-15; v5 2026-09-30 02:19 UTC (shorter file) is the current version; no journal reference. Verified from the arXiv abs page (submission history v1 to v5) and the v5 HTML render, both HTTP 200.
Sources
arXiv preprint(opens in a new tab) (primary)
arXiv v5 (revised 2026-09-30)(opens in a new tab) (30 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Luke Nicholls, Robert Hutto, Zephrah Soto, Hamilton Morrin, Thomas Pollak, Raj Korpan, Cheryl Carmichael
Tags
Cite This
APA
Luke Nicholls et al. (2026). "AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs. arXiv (City University of New York; King's College London Institute of Psychiatry, Psychology & Neuroscience). https://arxiv.org/abs/2604.13860
Related Insights
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
arXiv (Stanford-led author team) · 5 Aug 2026
How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior
arXiv preprint · 13 Aug 2026
Lost in Delusion: Examining LLM Safety Under User Delusions and Distress
arXiv preprint · 31 May 2026
Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports
arXiv (King's College London Institute of Psychiatry, Psychology and Neuroscience; South London and Maudsley NHS Foundation Trust; The Human Line Project) · 7 Sept 2026
Artificial intelligence-associated delusions and large language models: risks, mechanisms of delusion co-creation, and safeguarding strategies
The Lancet Psychiatry · 1 Jun 2026