When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation
Tests whether rubric-based evaluation, the dominant approach for grading LLMs in medicine, detects clinically relevant hallucinations. After a controlled study on MedHallu showing more specific rubrics separate correct from hallucinated answers better, the authors build a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that produces matched correct and error-injected responses. Across HealthBench, HealthBench Professional and LiveMedBench, injected clinically relevant errors are frequently missed and often leave rubric scores unchanged; rubrics work when they explicitly check facts and fail on unanticipated errors, and a retrieval-based factuality check recovers some misses.
Publisher
arXiv (University of Oxford)
Published
11 Sept 2026
Added
today
Key Findings
- Clinically relevant hallucinations injected into otherwise correct responses were often missed by the rubrics of HealthBench, HealthBench Professional and LiveMedBench, frequently leaving scores unchanged.
- Rubrics were most effective when explicitly checking facts and least effective for additional or unexpected errors they did not anticipate.
- In a controlled MedHallu setting, more specific rubrics distinguished correct from hallucinated responses better than generic ones.
- A preliminary retrieval-based factuality check recovered some rubric-blind errors, suggesting a complementary layer.
- The authors conclude rubric scores alone are insufficient to establish clinical reliability and may undermine clinician trust if relied on for deployment decisions.
Methodology Notes
Evaluation-methodology study: taxonomy of medical hallucination types, clinician-validated error-injection pipeline producing matched response pairs, applied to three rubric-scored medical benchmarks; details of clinician validation and model coverage are in the paper body, not the abstract. arXiv v1 2026-09-11 (14 September partition); all seven authors at the University of Oxford per the PDF title block; no venue stated.
Sources
Authors
Griffin Farrow, Lily Sijia Li, Jack Johnson, Tingyan Wang, Philip Torr, William Bolton, Fabio J. Fehr
Tags
Cite This
APA
Griffin Farrow et al. (2026). When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation. arXiv (University of Oxford). https://arxiv.org/abs/2609.12718
Related Insights
Scaling Clinical Judgment to Evaluate Medical AI
arXiv (Harvard Medical School; Beth Israel Deaconess Medical Center; Stanford; Massachusetts General Hospital; University of Alberta; MIT; Erasmus MC; University of Maryland) · 11 Sept 2026
HealthBench-Psych: A Mental Health Subset of OpenAI's HealthBench
arXiv (Division of Digital Psychiatry, Beth Israel Deaconess Medical Center) · 25 Aug 2026