Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated

Shows that the near-perfect depression-prediction results common in LLM mental-health papers are largely an artefact of predicting an instrument's score from language the same instrument elicited. Participants completed both a structured diagnostic depression interview and a life-history interview, and the same models predicted DSM-5 criteria from each. The mirrored condition is near-perfect and the non-mirrored condition collapses, but against an independent measure the mirror advantage disappears entirely.

Publisher

arXiv (Washington University in St. Louis; Southern Methodist University)

Published

7 Aug 2025

Added

today

Key Findings

  • N = 110 participants completed both a structured diagnostic depression interview (Mirror) and a life-history interview (Non-Mirror); interrater ICC across the ten criterion items ranged 0.84 to 0.99, median 0.93.
  • Mirror condition: GPT-4 Turbo F1 0.94, Jaccard 0.86, R-squared 0.80, accuracy 0.97, Pearson r 0.90; GPT-4o F1 0.92; LLaMA-3-70B F1 0.86.
  • Non-Mirror condition collapses for GPT-4 Turbo to F1 0.52, Jaccard 0.34 and R-squared 0.27, while Pearson r on total symptom burden remains 0.57.
  • Against an independent criterion, self-reported PHQ-9, the Mirror advantage vanishes: GPT-4 Turbo Mirror r = 0.53 against Non-Mirror r = 0.55 (PHQ-9 itself correlates with the interview score at r = 0.58).
  • Confidence-based posterior filtering helps almost only where performance was weak: Non-Mirror R-squared rises 0.27 to 0.35 (+29.6%) against Mirror 0.80 to 0.81 (+1.3%).
  • The paper names the contaminated prior results it corrects, including reported correlations of r = 0.73 (n = 393) and r = 0.70 (n = 963).

Methodology Notes

Single sample of 110, three models, English only; an APA-format manuscript preprint with no journal venue. v1 posted 2025-08-07; v2 2025-10-17; v3 announced 2026-09-11 and read for this record. published_date is the v1 date. Abstract page fetched and read.

Authors

Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns

Tags

arxivcriterion-contaminationdepressionphq-9measurement-validitywustl

Cite This

APA

Tong Li et al. (2025). "Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated. arXiv (Washington University in St. Louis; Southern Methodist University). https://arxiv.org/abs/2508.05830