"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated
Shows that the near-perfect depression-prediction results common in LLM mental-health papers are largely an artefact of predicting an instrument's score from language the same instrument elicited. Participants completed both a structured diagnostic depression interview and a life-history interview, and the same models predicted DSM-5 criteria from each. The mirrored condition is near-perfect and the non-mirrored condition collapses, but against an independent measure the mirror advantage disappears entirely.
Publisher
arXiv (Washington University in St. Louis; Southern Methodist University)
Published
7 Aug 2025
Added
today
Key Findings
- N = 110 participants completed both a structured diagnostic depression interview (Mirror) and a life-history interview (Non-Mirror); interrater ICC across the ten criterion items ranged 0.84 to 0.99, median 0.93.
- Mirror condition: GPT-4 Turbo F1 0.94, Jaccard 0.86, R-squared 0.80, accuracy 0.97, Pearson r 0.90; GPT-4o F1 0.92; LLaMA-3-70B F1 0.86.
- Non-Mirror condition collapses for GPT-4 Turbo to F1 0.52, Jaccard 0.34 and R-squared 0.27, while Pearson r on total symptom burden remains 0.57.
- Against an independent criterion, self-reported PHQ-9, the Mirror advantage vanishes: GPT-4 Turbo Mirror r = 0.53 against Non-Mirror r = 0.55 (PHQ-9 itself correlates with the interview score at r = 0.58).
- Confidence-based posterior filtering helps almost only where performance was weak: Non-Mirror R-squared rises 0.27 to 0.35 (+29.6%) against Mirror 0.80 to 0.81 (+1.3%).
- The paper names the contaminated prior results it corrects, including reported correlations of r = 0.73 (n = 393) and r = 0.70 (n = 963).
Methodology Notes
Single sample of 110, three models, English only; an APA-format manuscript preprint with no journal venue. v1 posted 2025-08-07; v2 2025-10-17; v3 announced 2026-09-11 and read for this record. published_date is the v1 date. Abstract page fetched and read.
Sources
Authors
Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns
Tags
Cite This
APA
Tong Li et al. (2025). "Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated. arXiv (Washington University in St. Louis; Southern Methodist University). https://arxiv.org/abs/2508.05830
Related Insights
Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media
Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026) · 1 Jul 2026
Mapping the Content and Consistency of Suicidal Thoughts with Large Language Models: A Two-Study Longitudinal Investigation Among Adolescents and Young Adults
PsyArXiv (Harvard University; Yale University; McLean Hospital; University of Denver) · 31 Aug 2026