Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation

Tests whether rubric-based evaluation, the dominant approach for grading LLMs in medicine, detects clinically relevant hallucinations. After a controlled study on MedHallu showing more specific rubrics separate correct from hallucinated answers better, the authors build a taxonomy of medical hallucination types and a clinician-validated error-injection pipeline that produces matched correct and error-injected responses. Across HealthBench, HealthBench Professional and LiveMedBench, injected clinically relevant errors are frequently missed and often leave rubric scores unchanged; rubrics work when they explicitly check facts and fail on unanticipated errors, and a retrieval-based factuality check recovers some misses.

Publisher

arXiv (University of Oxford)

Published

11 Sept 2026

Added

today

Key Findings

  • Clinically relevant hallucinations injected into otherwise correct responses were often missed by the rubrics of HealthBench, HealthBench Professional and LiveMedBench, frequently leaving scores unchanged.
  • Rubrics were most effective when explicitly checking facts and least effective for additional or unexpected errors they did not anticipate.
  • In a controlled MedHallu setting, more specific rubrics distinguished correct from hallucinated responses better than generic ones.
  • A preliminary retrieval-based factuality check recovered some rubric-blind errors, suggesting a complementary layer.
  • The authors conclude rubric scores alone are insufficient to establish clinical reliability and may undermine clinician trust if relied on for deployment decisions.

Methodology Notes

Evaluation-methodology study: taxonomy of medical hallucination types, clinician-validated error-injection pipeline producing matched response pairs, applied to three rubric-scored medical benchmarks; details of clinician validation and model coverage are in the paper body, not the abstract. arXiv v1 2026-09-11 (14 September partition); all seven authors at the University of Oxford per the PDF title block; no venue stated.

Authors

Griffin Farrow, Lily Sijia Li, Jack Johnson, Tingyan Wang, Philip Torr, William Bolton, Fabio J. Fehr

Tags

rubric-evaluationhallucinationhealthbenchlivemedbenchoxforderror-injection

Cite This

APA

Griffin Farrow et al. (2026). When Rubrics Fail: Hallucinations Reveal Blind Spots in Medical AI Evaluation. arXiv (University of Oxford). https://arxiv.org/abs/2609.12718