Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Language-model ratings of depression reflect the rater more than the patient

Pre-registered study of whether accurate language-model raters agree about individuals when screening for depression. 880 raters (11 open models crossed with prompt wordings, conversation views, answer-reading rules and score constructions) were applied to 189 DAIC-WOZ clinical interviews against the eight-item PHQ-8, with a locked confirmation on 86 new E-DAIC interviews. Model choice explained three times as much score variance as stable differences between participants, and accurate raters still disagreed on screening decisions for two in five participants.

Publisher

arXiv (Icahn School of Medicine at Mount Sinai; James J. Peters VA Medical Center; Berkman Klein Center, Harvard University)

Published

6 Oct 2026

Added

today

DOI

—

Key Findings

  • Model choice explained 30.0% of summed-symptom score variance; stable participant differences explained 10.5%.
  • Two randomly drawn raters with AUC of at least 0.70 disagreed on screening decisions for 40% of participants on average; AUC for PHQ-8 of 10 or more ranged from 0.34 to 0.88 (median 0.75), with 39% of raters below 0.70.
  • Instructing raters to score only explicitly reported symptoms lowered the median share flagged by 27 percentage points in 10 of 11 models.
  • Internal consistency (median Cronbach's alpha 0.91) was unrelated to validity across raters.
  • A locked analysis of 86 new interviews reproduced the pre-registered findings; exploratory recalibration with 40 labelled participants raised accuracy and halved disagreement but still left about one participant in five decided differently.

Methodology Notes

Pre-registered on OSF (osf.io/8qg2m) before any model output was compared with the criterion; discovery set DAIC-WOZ (189 semi-structured interviews with an operator-controlled virtual interviewer, USC Institute for Creative Technologies, available under data-use agreement); confirmation set 86 E-DAIC interviews with an autonomous interviewer and automatic transcripts. Single author; funded in part by a NARSAD Young Investigator Grant. Preprint v1 submitted 2026-10-06; not peer reviewed. Verified on the arXiv abstract page and the HTML full text (affiliations, registration, corpus description and every number read from the text); the OSF registration URL was not fetched.

Authors

Baihan Lin

Tags

depression-screeningphq-8daic-wozllm-raterspre-registeredmount-sinai

Cite This

APA

Baihan Lin. (2026). Language-model ratings of depression reflect the rater more than the patient. arXiv (Icahn School of Medicine at Mount Sinai; James J. Peters VA Medical Center; Berkman Klein Center, Harvard University). https://arxiv.org/abs/2610.08501