MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
Benchmark measuring whether LLMs keep a correct medical judgment when misleading context is injected into questions they otherwise answer correctly. It spans medical reasoning, agentic capability and patient-journey evaluation, with misleading injections varied by content-corruption type and provenance framing. A clinical panel reviewed outputs for potential harm. The paper is accepted (poster) to the NeurIPS 2026 Evaluations and Datasets Track.
Publisher
bioRxiv (University of Oxford-led consortium); accepted to NeurIPS 2026 Evaluations and Datasets Track
Published
28 May 2026
Added
today
Key Findings
- 10,932 medical question items and 48,889 misleading context-option pairs built from 5 source datasets
- Across 11 model configurations, mean accuracy fell from 71.1% on original questions to 38.0% under focused misleading context (51.5% attack success)
- Authority-framed falsehoods reached 69.5% attack success and exception-poisoning claims 64.1%
- A 14-member clinical panel from 7 countries identified serious potential harm in 38.2% of reviewed cases
Methodology Notes
Injections vary on 5 content-corruption types x 3 provenance framings. Mostly multiple-choice items derived from existing datasets; patient-journey subset is the participant-facing slice. bioRxiv v1 posted 2026-05-28 (DOI 10.64898/2026.05.25.727671, api.biorxiv.org 200); arXiv 2606.12291 v1 2026-06-10, v2 2026-06-15 under the title 'Measuring Epistemic Resilience of LLMs Under Misleading Medical Context'. NeurIPS 2026 acceptance from neurips.cc/static/virtual/data/neurips-2026-orals-posters.json (Accept (poster), Evaluations_and_Datasets_Track, OpenReview id IWp0p0ZBCj; OpenReview itself walled).
Sources
bioRxiv preprint(opens in a new tab) (primary)
arXiv version(opens in a new tab) (15 Jun 2026)
Code and data (GitHub)(opens in a new tab)
NeurIPS 2026 OpenReview forum(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Hongjian Zhou, Xinyu Zou, Jinge Wu, Sean Wu, Junchi Yu, Bradley Max Segal, Tobias Erich Niebuhr, Sara Amro, Michael Petrus, Sheikh Momin, Alexandra M. Cardoso Pinto, Rachel Niesen, David A. Clifton
Tags
Cite This
APA
Hongjian Zhou et al. (2026). MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context. bioRxiv (University of Oxford-led consortium); accepted to NeurIPS 2026 Evaluations and Datasets Track. https://www.biorxiv.org/content/10.64898/2026.05.25.727671v1