Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
Tests what a model does when the evidence in its context is wrong: Cochrane-derived clinical comparison questions are rewritten so the real intervention is replaced by a counterfactual term, including toxic substances, and the model is measured on whether it backs off to an uncertainty label or synthesises the fake evidence into a confident answer. Faithfulness to supplied context and safety pull in opposite directions here, and no model resolves the conflict.
Publisher
Association for Computational Linguistics (Findings of ACL 2026); University of Texas at Austin; Northeastern University; MD Anderson Cancer Center; Georgia Tech
Published
1 Jul 2026
Added
today
Key Findings
- 809 question-evidence tuples built from 203 filtered questions over 100 PubMed-referenced Cochrane systematic reviews citing 329 randomised trials; four counterfactual types (nonce, medical, non-medical, toxic) with 200 items in the toxic category.
- Nine models evaluated: Gemini-2.5-flash, GPT-5-mini, Llama-3.1-8B-Instruct, Llama-3.1-405B-Instruct, Llama-4-Maverick-17B-128E-Instruct, OLMo-3-7B-Instruct, OLMo-3-7B-Think, Qwen2.5-7B-Instruct and HuatuoGPT-o1-7B.
- Under the most adversarial prompt, no model's mean change in uncertainty rate exceeded 0.13, regardless of size or training paradigm; Llama-3.1-405B behaved like the 7B models.
- With counterfactual evidence present, up to 80% of outputs show no awareness or hesitation even when the substituted intervention is a poisonous substance.
- Label distribution of the source questions: 26.1% Higher, 46.3% No Difference, 27.6% Lower.
- Generation validation reached 92.86% accuracy on 70 sampled instances and reasoning-trace classification 90.00% on 70.
Methodology Notes
Counterfactuals are synthetic, generated by GPT-5-mini and labelled by Claude Sonnet 4.5, each human-validated on a 70-item sample; English only. The free-form prompt condition, which is closest to real use, is where models were least likely to express uncertainty. ACL 2026 was held 2 to 7 July 2026 in San Diego; the Anthology bib gives month and year only, so published_date is set to 2026-07-01 and the true precision is month. Findings of ACL 2026 pp. 37053-37081, DOI 10.18653/v1/2026.findings-acl.1847; metadata verified at Crossref.
Sources
Topics
Authors
Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C. Wallace, Junyi Jessy Li
Tags
Cite This
APA
Kaijie Mo et al. (2026). Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence. Association for Computational Linguistics (Findings of ACL 2026); University of Texas at Austin; Northeastern University; MD Anderson Cancer Center; Georgia Tech. https://aclanthology.org/2026.findings-acl.1847/