Skip to main content
Peer-reviewed Authoritative

Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis

Conference paper from Apart Research with co-authors at the London School of Economics and the Stanford Institute for Human-Centered AI, testing whether automated judges can stand in for human raters when scoring how models respond to users showing signs of psychosis. The authors worked with clinical psychologists and psychiatrists to operationalise seven binary safety criteria, built a human-consensus gold-standard dataset over responses from four frontier models, then measured how closely single LLM judges and a three-model jury agreed with that consensus. Agreement was substantial for the best judge and comparable for the jury, but varied widely by criterion.

Publisher

Proceedings of the IASEAI Conference (published by AAAI)

Published

15 Jul 2026

Added

3 days ago

DOI

Key Findings

  • Seven binary criteria were defined for responses to a user in psychosis: Stigmatizes, Validates Delusion, Embellishes, Challenges, No Referral, Provides Non-Referral Advice, and Continues Conversation
  • Two human raters produced 448 binary judgments (7 criteria x 16 stimuli x 4 responder models) with substantial agreement between them (Cohen's kappa = .80), forming the gold standard
  • Against that consensus, agreement was substantial for Gemini 2.5 Pro (kappa = .75) and Qwen-32B (kappa = .68) and moderate for Kimi K2 (kappa = .56); the three-model jury scored kappa = .74, slightly below the single best judge
  • Per-criterion agreement ranged from .34 to 1.00 — highest for the concrete 'No Referral' criterion and lowest for the interpretive 'Embellishes' criterion — so the aggregate figure hides criteria that are not reliably auto-judgeable
  • Judge-to-judge agreement was weaker than judge-to-human on the worst pairing (Gemini x Kimi kappa = .51), which the authors treat as evidence that judge selection matters as much as the jury mechanism

Methodology Notes

Single-turn stimulus study, not a multi-turn simulation. Nineteen stimuli were built from published clinical-psychology vignettes of professionally diagnosed patients and rewritten into first-person user messages using Claude Sonnet 4; three were held out to calibrate raters and refine criteria, leaving 16 for the studies. Responses were collected from GPT-4o, Claude Sonnet 4, DeepSeek-v3-0324 and Llama-3.1-405B-Instruct-FP8; judges were Gemini 2.5 Pro, Qwen-32B-FP8 and Kimi-K2-Instruct, deliberately disjoint from the responder set to avoid self-preference. Judge results are averaged over 25 seeds. Limitations the authors state: the stimulus set is small and curated so per-criterion kappa may not hold at scale, the dataset contains only psychosis presentations with no controls, and — the material caveat against the title's 'clinically-validated' — the two human raters received clinician guidance but had no clinical training or professional experience themselves, so the gold standard is a guided lay consensus rather than a clinician consensus. Peer-reviewed conference proceedings, IASEAI'26 Main Track, pages 610-624, volume 2(1). Verification route: the OJS landing page was fetched through a text proxy for the title, author affiliations, page range and the 2026-07-15 publication date, and the full PDF was downloaded from ojs.aaai.org (HTTP 200, 420,997 bytes) and read; all figures above were extracted from its methods and results sections.

Authors

May Lynn Reese, Markela Zeneli, Mindy Ng, Jacob Haimes, Andreea Damien, Elizabeth Stade

Tags

iaseaillm-as-judgellm-as-jurypsychosisapart-researchstanford-haicodebookinter-rater-reliability

Cite This

APA

May Lynn Reese et al. (2026). Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis. Proceedings of the IASEAI Conference (published by AAAI). https://ojs.aaai.org/index.php/IASEAI/article/view/43055