Scaling Clinical Judgment to Evaluate Medical AI
Introduces PrecepTron, a 32-billion-parameter model fine-tuned with low-rank adaptation on a small number of physician examples to grade open-ended clinical-reasoning responses at physician level, and releases GRAND-ROUNDS, a benchmark of 9,217 scores by 11 physicians across seven studies. Frontier models used as judges often disagree with physicians and with each other, whereas the fine-tuned grader scores consistently with physicians across tasks; the authors use it to reproduce headline findings from five influential JAMA, Science and Nature Medicine studies without new human grading and to ask new questions such as diagnostic accuracy when cases arrive piecemeal.
Publisher
arXiv (Harvard Medical School; Beth Israel Deaconess Medical Center; Stanford; Massachusetts General Hospital; University of Alberta; MIT; Erasmus MC; University of Maryland)
Published
11 Sept 2026
Added
today
Key Findings
- GRAND-ROUNDS releases 9,217 physician scores by 11 physicians across seven studies, described as a large-scale physician-annotated benchmark for evaluating open-ended medical AI responses.
- Frontier LLMs in typical LLM-as-a-judge configurations often disagree with physicians and with each other on clinical-reasoning grading.
- PrecepTron, a 32B model with LoRA fine-tuning on a small number of physician-graded cases, achieves physician-level consistent scoring across tasks.
- The grader reproduced headline findings from five influential studies of LLMs in clinical care (JAMA, Science, Nature Medicine) without new human grading.
- Using it, the authors measure how frontier LLM diagnostic accuracy changes when clinical information is provided piecemeal, down to token by token; all code, data and labels are released.
Methodology Notes
Methods and resource paper: physician annotation corpus assembled from seven prior studies, LoRA fine-tuning of a 32B open model as a grader, agreement analyses against physician scores and against frontier LLM judges, plus reproduction and new experiments. arXiv v1 2026-09-11 (announced in the 14 September partition). Affiliations from the PDF title block; no venue stated. Clinical reasoning tasks rather than patient-facing conversation, so the payload is evaluation methodology.
Sources
Authors
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur, Byron Crowe, Anthony M. Pettinato, Aashna P. Shah, Adrian D. Haimovich, Liam G. McCoy, Daniel Restrepo, Ethan Goh, Jonathan H. Chen, Laura Zwaan, Katherine E. Goodman, Daniel J. Morgan, Raja-Elie E. Abdulnour, Adam Rodman, Arjun K. Manrai
Tags
Cite This
APA
Thomas A. Buckley et al. (2026). Scaling Clinical Judgment to Evaluate Medical AI. arXiv (Harvard Medical School; Beth Israel Deaconess Medical Center; Stanford; Massachusetts General Hospital; University of Alberta; MIT; Erasmus MC; University of Maryland). https://arxiv.org/abs/2609.12822