aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI
A fully-crossed study in which three frontier models act as both generators and judges of 3,000 mental-health, companion and coaching messages, with a psychologist's ratings as the anchor. The authors report that disagreement between the judges is structured rather than random and concentrates on the safety-critical metrics, then distil a per-metric psychologist-corrected target into a small frozen local model, released as aipsy-judge-1.0 under Apache-2.0.
Key Findings
- One judge (Gemini) is a systematic outlier: the most lenient of the three, carrying a +0.99 self-preference premium on its own generations, flagging roughly a third to a fifth as many tail failures as the others, and rating a self-harm response containing means information as exemplary
- Inter-judge agreement is lowest on empathy (alpha 0.24), the metric the authors argue sycophancy hides inside
- The binary crisis-detection flag is the one safety-critical signal the judges agree on (alpha 0.80), and they err toward over-flagging, which is the safe direction for a triage screen
- Equal-weight averaging across judges, the canonical remedy, blends one judge's leniency and tail-blindness into the composite safety score rather than cancelling it
- The distilled local judge tracks the psychologist-corrected target better than its base model on the composite (ICC 0.64 to 0.75) and on crisis detection (kappa 0.65 to 0.82)
Methodology Notes
The generators and judges are gpt-5.4-mini, claude-sonnet-4-6 and gemini-2.5-flash; the released model is an Apache-2.0 fine-tune of Gemma-4-26B-A4B. The anchoring instrument, aipsy-bench, is described as open and frozen but is the authors' own, so the psychologist correction rests on a single rater's ratings within a bench this team also built. Keido Labs is a small independent shop with no institutional affiliation stated on the paper, which is why the credibility grade here is preliminary rather than credible. Announcement lag matters for anyone reconstructing the timeline: the arXiv submission date is 2026-07-13 but the paper was only announced in the 2026-08-25 batch, and the HuggingFace model was last modified 2026-07-04. Not peer reviewed. Curator-verified: the arXiv abs page and PDF were fetched and read, and the HuggingFace model keidolabs/aipsy-judge-1.0 is public, ungated and Apache-2.0.
Sources
arXiv abstract page (2608.24899) (primary)
Released judge model (HuggingFace, Apache-2.0)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Michael Keeman, Anastasia Keeman
Tags
Cite This
APA
Michael Keeman, Anastasia Keeman (2026). aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI. arXiv (Keido Labs). https://arxiv.org/abs/2608.24899
Related Insights
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026
Using LLM-as-a-Judge/Jury to Advance Scalable, Clinically-Validated Safety Evaluations of Model Responses to Users Demonstrating Psychosis
Proceedings of the IASEAI Conference (published by AAAI) · 15 Jul 2026
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026