When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
Proposes Med-Stress, a stress test of whether large language models keep a correct diagnosis under escalating conversational pressure, and finds across nine frontier models a clear dissociation between medical knowledge and robustness: high initial diagnostic capability does not imply belief stability, leaving large knowledge-robustness gaps for several models. Two mitigations are proposed, an inference-time role-based epistemic defence and resilience-oriented fine-tuning, the latter nearly eliminating belief change.
Publisher
Association for Computational Linguistics (Proceedings of ACL 2026, Long Papers); Harbin Institute of Technology
Published
1 Jul 2026
Added
today
Key Findings
- Despite strong medical benchmark accuracy, LLMs can abandon an initially correct diagnosis under escalating pressure in clinical dialogue, a multi-turn sycophancy failure.
- Across nine frontier LLMs, high initial diagnostic capability did not imply high belief stability; the benchmarks beat's read of the PDF records GPT-4o and Claude Sonnet 4 among the highest on initial capability and substantially lower on stability, with vanilla misbelief rates ranging from about 6% to 74% by model.
- Resilience-oriented fine-tuning (R-FT) nearly eliminated belief change, and the inference-time defence (RBED) gave partial protection.
- The test corpus was filtered by a GPT-4o verifier with a 20% expert-review sample.
Methodology Notes
Automated stress test with escalating-pressure dialogues over medical diagnosis cases, nine frontier models (closed and open weights), plus two mitigation methods; no patients. Published in the ACL 2026 Long Papers volume (July 2026), pages 8720 to 8764, DOI 10.18653/v1/2026.acl-long.395. Verified at the ACL Anthology on 2026-09-15; model-level figures are from the benchmarks beat's PDF read and should be checked against the paper before quotation.
Sources
Authors
Boyu Xiao, Xiuqi Tian, Xuwen Song, Haochun Wang, Guanchun Song, Sendong Zhao, Bing Qin
Tags
Cite This
APA
Boyu Xiao et al. (2026). When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure. Association for Computational Linguistics (Proceedings of ACL 2026, Long Papers); Harbin Institute of Technology. https://aclanthology.org/2026.acl-long.395/
Related Insights
Too agreeable to be accurate? Sycophancy and diagnostic instability of large language models in medical diagnosis
Artificial Intelligence in Medicine (Elsevier); University of Toronto · 4 Sept 2026
Measuring LLM Sycophancy under Sustained Multi-Turn Pressure
arXiv (Texas A&M University; University of Cincinnati) · 8 Sept 2026