Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage
A blinded, paired benchmark of three consumer assistants answering patient- and caregiver-facing questions about late-life depression. Ninety questions covering six geriatric-psychiatry domains, stratified evenly across low, moderate and high clinical risk, were put independently to ChatGPT (GPT-5.5 Instant), Gemini (Gemini 3.5 Flash) and Doubao (Seed2.0 Pro). Two psychiatrists rated anonymised outputs against prespecified item-level reference standards, with a third adjudicating clinically important disagreements.
Publisher
Frontiers in Psychiatry; Guilin Medical University; Beijing Union University; Xiangnan University; Zhejiang University of Science and Technology
Published
11 Sept 2026
Added
today
Key Findings
- Clinically acceptable responses: 78.9% for ChatGPT, 72.2% for Gemini, 60.0% for Doubao (P < 0.001).
- Major safety errors: 5.6% for ChatGPT, 10.0% for Gemini, 16.7% for Doubao (P = 0.015; FDR q = 0.023).
- Complete geriatric-specific appropriateness was reached by only 64.4%, 55.6% and 43.3% of responses respectively.
- Performance declined substantially as the clinical-risk stratum rose, so the models are weakest on the questions that matter most.
- On the caregiver-centred subset, acceptability was 76.7%, 70.0% and 53.3%.
- 270 primary-round responses plus a 30-question retest subset in fresh conversations, 360 responses in total, used to assess test-retest consistency.
Methodology Notes
Blinded paired benchmarking on vignette-style questions rather than real patients; the three products were queried through their consumer interfaces, not APIs, so the results describe the shipped experience. Two independent psychiatrist raters with third-rater adjudication. Question framing is Chinese-clinical-context and the models include one Chinese frontier system. Published in Frontiers in Psychiatry 2026-09-11, CC BY 4.0, DOI 10.3389/fpsyt.2026.1956736; metadata and full structured abstract verified from the Crossref record.
Sources
Topics
Authors
Wei Xiao, Huanyu Zhang, Xiaoyi Chen, Jun Cai, Xuchen Luo, Jiehua Deng
Tags
Cite This
APA
Wei Xiao et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry; Guilin Medical University; Beijing Union University; Xiangnan University; Zhejiang University of Science and Technology. https://doi.org/10.3389/fpsyt.2026.1956736