The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7
Tests whether language models used as simulated patients produce psychometrically valid responses to the PHQ-9 and GAD-7 depression and anxiety scales. Four open-weight models (Llama-3.1-8B-Instruct, Mistral-7B-Instruct, Llama-3.3-70B-Instruct, GPT-OSS-120B) completed the scales under varied prompts, sampling temperatures and situational scenarios, and the responses were compared with human baseline data using multi-group confirmatory factor analysis, differential item functioning and measurement-invariance testing. The models showed exceptionally high internal consistency but a factor structure that did not match human responses, no scalar invariance, pervasive item-level bias and strong prompt fragility, which the authors call a reliability illusion.
Publisher
Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026)
Published
1 Jul 2026
Added
today
Key Findings
- LLM-generated responses reached Cronbach's alpha of 0.86 to 0.96, often exceeding human baselines
- Confirmatory factor analysis showed poor structural fit for the LLM data (RMSEA 0.124 to 0.364, CFI as low as 0.724) against a human baseline of CFI 0.972 and RMSEA 0.047 on a sample of N = 6,000
- No model achieved even partial scalar measurement invariance against human responses on either scale, and configural fit was already inadequate for some models
- Differential item functioning was pervasive and small prompt changes disrupted metric and scalar invariance, indicating stereotype-driven semantic matching rather than a stable latent construct
- The authors conclude that high internal consistency in synthetic respondents should not be read as validity when pre-validating scales or simulating patients
Methodology Notes
ACL Anthology 2026.clpsych-1.7, DOI 10.18653/v1/2026.clpsych-1.7, pages 88-99, CLPsych 2026 (July 2026; month precision). Authors at the University of Florida and the University of Mississippi (PDF title block). Four open-weight models only, no proprietary models; the human comparison data are existing baselines rather than a matched sample. Findings concern scale responses, so the transfer to free-text simulated-patient dialogue is by argument, not measured here.
Sources
ACL Anthology(opens in a new tab) (primary)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Qian Shen, Yu Han
Tags
Cite This
APA
Qian Shen, Yu Han. (2026). The Reliability Illusion in Synthetic Patients: Psychometric Misalignment of Open-weight LLMs on PHQ-9 and GAD-7. Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026). https://aclanthology.org/2026.clpsych-1.7/
Related Insights
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
JMIR AI · 29 Jun 2026
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research) · 31 Aug 2026
One-shot emergency psychiatric triage across 15 frontier AI chatbots
arXiv (Max Planck UCL Centre for Computational Psychiatry and Ageing Research; UK AI Security Institute; University of Oxford; Microsoft AI) · 28 Apr 2026
The Attachment Index: Auditing Attachment Language Cues and Relational Safety Risks in Human-LLM Dialogue
Association for Computational Linguistics (Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology, CLPsych 2026) · 1 Jul 2026
Large Language Model–Based Behavioral Activation Chatbot for Young People With Depression Using Artificial Users and Clinical Experts: Mixed Methods Evaluation
JMIR Mental Health (JMIR Publications) · 1 Sept 2026