Skip to main content
Benchmark / dataset Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence

Tests what a model does when the evidence in its context is wrong: Cochrane-derived clinical comparison questions are rewritten so the real intervention is replaced by a counterfactual term, including toxic substances, and the model is measured on whether it backs off to an uncertainty label or synthesises the fake evidence into a confident answer. Faithfulness to supplied context and safety pull in opposite directions here, and no model resolves the conflict.

Publisher

Association for Computational Linguistics (Findings of ACL 2026); University of Texas at Austin; Northeastern University; MD Anderson Cancer Center; Georgia Tech

Published

1 Jul 2026

Added

today

Key Findings

  • 809 question-evidence tuples built from 203 filtered questions over 100 PubMed-referenced Cochrane systematic reviews citing 329 randomised trials; four counterfactual types (nonce, medical, non-medical, toxic) with 200 items in the toxic category.
  • Nine models evaluated: Gemini-2.5-flash, GPT-5-mini, Llama-3.1-8B-Instruct, Llama-3.1-405B-Instruct, Llama-4-Maverick-17B-128E-Instruct, OLMo-3-7B-Instruct, OLMo-3-7B-Think, Qwen2.5-7B-Instruct and HuatuoGPT-o1-7B.
  • Under the most adversarial prompt, no model's mean change in uncertainty rate exceeded 0.13, regardless of size or training paradigm; Llama-3.1-405B behaved like the 7B models.
  • With counterfactual evidence present, up to 80% of outputs show no awareness or hesitation even when the substituted intervention is a poisonous substance.
  • Label distribution of the source questions: 26.1% Higher, 46.3% No Difference, 27.6% Lower.
  • Generation validation reached 92.86% accuracy on 70 sampled instances and reasoning-trace classification 90.00% on 70.

Methodology Notes

Counterfactuals are synthetic, generated by GPT-5-mini and labelled by Claude Sonnet 4.5, each human-validated on a 70-item sample; English only. The free-form prompt condition, which is closest to real use, is where models were least likely to express uncertainty. ACL 2026 was held 2 to 7 July 2026 in San Diego; the Anthology bib gives month and year only, so published_date is set to 2026-07-01 and the true precision is month. Findings of ACL 2026 pp. 37053-37081, DOI 10.18653/v1/2026.findings-acl.1847; metadata verified at Crossref.

Authors

Kaijie Mo, Siddhartha Venkatayogi, Chantal Shaib, Ramez Kouzy, Wei Xu, Byron C. Wallace, Junyi Jessy Li

Tags

acl-2026benchmarkcochranecounterfactual-evidencefaithfulnesstoxic

Cite This

APA

Kaijie Mo et al. (2026). Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence. Association for Computational Linguistics (Findings of ACL 2026); University of Texas at Austin; Northeastern University; MD Anderson Cancer Center; Georgia Tech. https://aclanthology.org/2026.findings-acl.1847/