Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Conference paper from Apart Research with co-authors at the London School of Economics and the Stanford Institute for Human-Centered AI, testing whether automated judges can stand in for human raters when scoring how models respond to users showing signs of psychosis. The authors worked with clinical psychologists and psychiatrists to operationalise seven binary safety criteria, built a human-consensus gold-standard dataset over responses from four frontier models, then measured how closely single LLM judges and a three-model jury agreed with that consensus. Agreement was substantial for the best judge and comparable for the jury, but varied widely by criterion.
Publisher
Proceedings of the IASEAI Conference (published by AAAI)
Published
15 Jul 2026
Added
3 days ago
DOI
—
Key Findings
- Seven binary criteria were defined for responses to a user in psychosis: Stigmatizes, Validates Delusion, Embellishes, Challenges, No Referral, Provides Non-Referral Advice, and Continues Conversation
- Two human raters produced 448 binary judgments (7 criteria x 16 stimuli x 4 responder models) with substantial agreement between them (Cohen's kappa = .80), forming the gold standard
- Against that consensus, agreement was substantial for Gemini 2.5 Pro (kappa = .75) and Qwen-32B (kappa = .68) and moderate for Kimi K2 (kappa = .56); the three-model jury scored kappa = .74, slightly below the single best judge
- Per-criterion agreement ranged from .34 to 1.00 — highest for the concrete 'No Referral' criterion and lowest for the interpretive 'Embellishes' criterion — so the aggregate figure hides criteria that are not reliably auto-judgeable
- Judge-to-judge agreement was weaker than judge-to-human on the worst pairing (Gemini x Kimi kappa = .51), which the authors treat as evidence that judge selection matters as much as the jury mechanism
Methodology Notes
Single-turn stimulus study, not a multi-turn simulation. Nineteen stimuli were built from published clinical-psychology vignettes of professionally diagnosed patients and rewritten into first-person user messages using Claude Sonnet 4; three were held out to calibrate raters and refine criteria, leaving 16 for the studies. Responses were collected from GPT-4o, Claude Sonnet 4, DeepSeek-v3-0324 and Llama-3.1-405B-Instruct-FP8; judges were Gemini 2.5 Pro, Qwen-32B-FP8 and Kimi-K2-Instruct, deliberately disjoint from the responder set to avoid self-preference. Judge results are averaged over 25 seeds. Limitations the authors state: the stimulus set is small and curated so per-criterion kappa may not hold at scale, the dataset contains only psychosis presentations with no controls, and — the material caveat against the title's 'clinically-validated' — the two human raters received clinician guidance but had no clinical training or professional experience themselves, so the gold standard is a guided lay consensus rather than a clinician consensus. Peer-reviewed conference proceedings, IASEAI'26 Main Track, pages 610-624, volume 2(1). Verification route: the OJS landing page was fetched through a text proxy for the title, author affiliations, page range and the 2026-07-15 publication date, and the full PDF was downloaded from ojs.aaai.org (HTTP 200, 420,997 bytes) and read; all figures above were extracted from its methods and results sections.
Sources
IASEAI'26 Proceedings article page (primary)
Full paper PDF (IASEAI'26 proceedings)
Archived snapshot (Wayback Machine) — preserved against link rot
Topics
Authors
May Lynn Reese, Markela Zeneli, Mindy Ng, Jacob Haimes, Andreea Damien, Elizabeth Stade
Tags
Cite This
APA
May Lynn Reese et al. (2026). Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis. Proceedings of the IASEAI Conference (published by AAAI). https://ojs.aaai.org/index.php/IASEAI/article/view/43055
Related Insights
Independent Clinical Evaluation of General-Purpose LLM Responses to Signals of Suicide Risk
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
arXiv (Stanford-led author team) · 5 Aug 2026
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
ACM (Proceedings of FAccT 2026) · 25 Jun 2026
Conversational AI and Emerging Psychosis: A Simulation Study of Potentially Iatrogenic Response Patterns
PsyArXiv (Universidad Francisco de Vitoria; Durham University) · 6 Aug 2026
A clinically validated framework for auditing AI chatbot behavior in mental health interactions
Nature Medicine · 7 Aug 2026
A Framework for Evidence-Based Psychotherapy with AI (EBP-AI)
Journal of Psychopathology and Clinical Science (American Psychological Association) · 17 Aug 2026