Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions

Adversarial benchmark of 585 cancer-related patient questions containing physician-verified false presuppositions, built after three hematology-oncology physicians reviewed real patient questions and found that large language models answer accurately but rarely notice or correct the false premise. No frontier model tested, including GPT-5, Gemini 2.5 Pro and Claude 4 Sonnet, corrected the presupposition more than 43% of the time. A companion 150-question set without false presuppositions (Cancer-Myth-NFP) shows that prompt-based mitigations raise correction rates at the cost of falsely flagging premises in questions that have none.

Publisher

arXiv (Thomas Lord Department of Computer Science and Keck School of Medicine, University of Southern California)

Published

15 Apr 2025

Added

today

Key Findings

  • Three hematology-oncology physicians evaluated model answers to real cancer-patient questions and found responses generally accurate but frequently failing to recognise false presuppositions embedded in the question
  • On the 585-question Cancer-Myth set, no frontier model corrected the false presupposition more than 43% of the time
  • Precautionary prompting with GEPA optimisation raised accuracy on Cancer-Myth to 80% but misidentified presuppositions in 41% of the 150 Cancer-Myth-NFP questions and cut performance on other medical benchmarks by about 10% relative
  • Dataset, code and a project website are public (Hugging Face dataset Cancer-Myth/Cancer-Myth; GitHub Bill1235813/cancer-myth)

Methodology Notes

Expert-verified adversarial dataset construction from real patient questions with physician review; frontier-model evaluation on presupposition correction; mitigation study with a matched no-false-presupposition control set. Preprint, marked 'under review' on the v3 PDF (2025-10-30); no venue stated. Published on arXiv 2025-04-15 (v1); curator read the abstract page and the v3 PDF title block on 2026-09-17.

Authors

Zhu, Wang Bill, Chen, Tianqi, Yu, Xinyan Velocity, Lin, Ching Ying, Law, Jade, Jizzini, Mazen, Nieva, Jorge J., Liu, Ruishan, Jia, Robin

Tags

cancerfalse-presuppositionmedical-qauscbenchmark

Cite This

APA

Zhu, Wang Bill et al. (2025). Cancer-Myth: Evaluating Large Language Models on Patient Questions with False Presuppositions. arXiv (Thomas Lord Department of Computer Science and Keck School of Medicine, University of Southern California). https://arxiv.org/abs/2504.11373