Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions
Adversarial benchmark of 585 cancer-related patient questions containing physician-verified false presuppositions, built after three hematology-oncology physicians reviewed real patient questions and found that large language models answer accurately but rarely notice or correct the false premise. No frontier model tested, including GPT-5, Gemini 2.5 Pro and Claude 4 Sonnet, corrected the presupposition more than 43% of the time. A companion 150-question set without false presuppositions (Cancer-Myth-NFP) shows that prompt-based mitigations raise correction rates at the cost of falsely flagging premises in questions that have none.
Publisher
arXiv (Thomas Lord Department of Computer Science and Keck School of Medicine, University of Southern California)
Published
15 Apr 2025
Added
today
Key Findings
- Three hematology-oncology physicians evaluated model answers to real cancer-patient questions and found responses generally accurate but frequently failing to recognise false presuppositions embedded in the question
- On the 585-question Cancer-Myth set, no frontier model corrected the false presupposition more than 43% of the time
- Precautionary prompting with GEPA optimisation raised accuracy on Cancer-Myth to 80% but misidentified presuppositions in 41% of the 150 Cancer-Myth-NFP questions and cut performance on other medical benchmarks by about 10% relative
- Dataset, code and a project website are public (Hugging Face dataset Cancer-Myth/Cancer-Myth; GitHub Bill1235813/cancer-myth)
Methodology Notes
Expert-verified adversarial dataset construction from real patient questions with physician review; frontier-model evaluation on presupposition correction; mitigation study with a matched no-false-presupposition control set. Preprint, marked 'under review' on the v3 PDF (2025-10-30); no venue stated. Published on arXiv 2025-04-15 (v1); curator read the abstract page and the v3 PDF title block on 2026-09-17.
Topics
Authors
Zhu, Wang Bill, Chen, Tianqi, Yu, Xinyan Velocity, Lin, Ching Ying, Law, Jade, Jizzini, Mazen, Nieva, Jorge J., Liu, Ruishan, Jia, Robin
Tags
Cite This
APA
Zhu, Wang Bill et al. (2025). Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions. arXiv (Thomas Lord Department of Computer Science and Keck School of Medicine, University of Southern California). https://arxiv.org/abs/2504.11373
Related Insights
When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
arXiv (University of Oxford) · 11 Sept 2026
Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy
arXiv (Virginia Tech) · 2 Aug 2026
Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations
arXiv (accepted to Machine Learning for Healthcare, MLHC 2026); Northeastern University · 14 Jul 2026