Language-model ratings of depression reflect the rater more than the patient
Pre-registered study of whether accurate language-model raters agree about individuals when screening for depression. 880 raters (11 open models crossed with prompt wordings, conversation views, answer-reading rules and score constructions) were applied to 189 DAIC-WOZ clinical interviews against the eight-item PHQ-8, with a locked confirmation on 86 new E-DAIC interviews. Model choice explained three times as much score variance as stable differences between participants, and accurate raters still disagreed on screening decisions for two in five participants.
Publisher
arXiv (Icahn School of Medicine at Mount Sinai; James J. Peters VA Medical Center; Berkman Klein Center, Harvard University)
Published
6 Oct 2026
Added
today
DOI
—
Key Findings
- Model choice explained 30.0% of summed-symptom score variance; stable participant differences explained 10.5%.
- Two randomly drawn raters with AUC of at least 0.70 disagreed on screening decisions for 40% of participants on average; AUC for PHQ-8 of 10 or more ranged from 0.34 to 0.88 (median 0.75), with 39% of raters below 0.70.
- Instructing raters to score only explicitly reported symptoms lowered the median share flagged by 27 percentage points in 10 of 11 models.
- Internal consistency (median Cronbach's alpha 0.91) was unrelated to validity across raters.
- A locked analysis of 86 new interviews reproduced the pre-registered findings; exploratory recalibration with 40 labelled participants raised accuracy and halved disagreement but still left about one participant in five decided differently.
Methodology Notes
Pre-registered on OSF (osf.io/8qg2m) before any model output was compared with the criterion; discovery set DAIC-WOZ (189 semi-structured interviews with an operator-controlled virtual interviewer, USC Institute for Creative Technologies, available under data-use agreement); confirmation set 86 E-DAIC interviews with an autonomous interviewer and automatic transcripts. Single author; funded in part by a NARSAD Young Investigator Grant. Preprint v1 submitted 2026-10-06; not peer reviewed. Verified on the arXiv abstract page and the HTML full text (affiliations, registration, corpus description and every number read from the text); the OSF registration URL was not fetched.
Authors
Baihan Lin
Tags
Cite This
APA
Baihan Lin. (2026). Language-model ratings of depression reflect the rater more than the patient. arXiv (Icahn School of Medicine at Mount Sinai; James J. Peters VA Medical Center; Berkman Klein Center, Harvard University). https://arxiv.org/abs/2610.08501
Related Insights
The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7
Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026) · 1 Jul 2026
Generative Large Language Models in Mental Health Care Settings: Systematic Review and Meta-Analysis
JMIR AI (JMIR Publications) · 31 Aug 2026