Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage

A blinded, paired benchmark of three consumer assistants answering patient- and caregiver-facing questions about late-life depression. Ninety questions covering six geriatric-psychiatry domains, stratified evenly across low, moderate and high clinical risk, were put independently to ChatGPT (GPT-5.5 Instant), Gemini (Gemini 3.5 Flash) and Doubao (Seed2.0 Pro). Two psychiatrists rated anonymised outputs against prespecified item-level reference standards, with a third adjudicating clinically important disagreements.

Publisher

Frontiers in Psychiatry; Guilin Medical University; Beijing Union University; Xiangnan University; Zhejiang University of Science and Technology

Published

11 Sept 2026

Added

today

Key Findings

  • Clinically acceptable responses: 78.9% for ChatGPT, 72.2% for Gemini, 60.0% for Doubao (P < 0.001).
  • Major safety errors: 5.6% for ChatGPT, 10.0% for Gemini, 16.7% for Doubao (P = 0.015; FDR q = 0.023).
  • Complete geriatric-specific appropriateness was reached by only 64.4%, 55.6% and 43.3% of responses respectively.
  • Performance declined substantially as the clinical-risk stratum rose, so the models are weakest on the questions that matter most.
  • On the caregiver-centred subset, acceptability was 76.7%, 70.0% and 53.3%.
  • 270 primary-round responses plus a 30-question retest subset in fresh conversations, 360 responses in total, used to assess test-retest consistency.

Methodology Notes

Blinded paired benchmarking on vignette-style questions rather than real patients; the three products were queried through their consumer interfaces, not APIs, so the results describe the shipped experience. Two independent psychiatrist raters with third-rater adjudication. Question framing is Chinese-clinical-context and the models include one Chinese frontier system. Published in Frontiers in Psychiatry 2026-09-11, CC BY 4.0, DOI 10.3389/fpsyt.2026.1956736; metadata and full structured abstract verified from the Crossref record.

Authors

Wei Xiao, Huanyu Zhang, Xiaoyi Chen, Jun Cai, Xuchen Luo, Jiehua Deng

Tags

late-life-depressionolder-adultsbenchmarkdoubaoblinded-ratingfrontiers

Cite This

APA

Wei Xiao et al. (2026). Large language models for late-life depression: a blinded benchmark of clinical safety, geriatric appropriateness, and triage. Frontiers in Psychiatry; Guilin Medical University; Beijing Union University; Xiangnan University; Zhejiang University of Science and Technology. https://doi.org/10.3389/fpsyt.2026.1956736