Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7

Tests whether language models used as simulated patients produce psychometrically valid responses to the PHQ-9 and GAD-7 depression and anxiety scales. Four open-weight models (Llama-3.1-8B-Instruct, Mistral-7B-Instruct, Llama-3.3-70B-Instruct, GPT-OSS-120B) completed the scales under varied prompts, sampling temperatures and situational scenarios, and the responses were compared with human baseline data using multi-group confirmatory factor analysis, differential item functioning and measurement-invariance testing. The models showed exceptionally high internal consistency but a factor structure that did not match human responses, no scalar invariance, pervasive item-level bias and strong prompt fragility, which the authors call a reliability illusion.

Publisher

Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026)

Published

1 Jul 2026

Added

today

Key Findings

  • LLM-generated responses reached Cronbach's alpha of 0.86 to 0.96, often exceeding human baselines
  • Confirmatory factor analysis showed poor structural fit for the LLM data (RMSEA 0.124 to 0.364, CFI as low as 0.724) against a human baseline of CFI 0.972 and RMSEA 0.047 on a sample of N = 6,000
  • No model achieved even partial scalar measurement invariance against human responses on either scale, and configural fit was already inadequate for some models
  • Differential item functioning was pervasive and small prompt changes disrupted metric and scalar invariance, indicating stereotype-driven semantic matching rather than a stable latent construct
  • The authors conclude that high internal consistency in synthetic respondents should not be read as validity when pre-validating scales or simulating patients

Methodology Notes

ACL Anthology 2026.clpsych-1.7, DOI 10.18653/v1/2026.clpsych-1.7, pages 88-99, CLPsych 2026 (July 2026; month precision). Authors at the University of Florida and the University of Mississippi (PDF title block). Four open-weight models only, no proprietary models; the human comparison data are existing baselines rather than a matched sample. Findings concern scale responses, so the transfer to free-text simulated-patient dialogue is by argument, not measured here.

Authors

Qian Shen, Yu Han

Tags

clpsych-2026synthetic-patientspsychometricsphq-9gad-7measurement-invariancesimulated-users

Cite This

APA

Qian Shen, Yu Han. (2026). The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7. Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026). https://aclanthology.org/2026.clpsych-1.7/