IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models
Pre-registered benchmark scoring model answers to clinical questions on two harm axes, commission and omission. Matched scenarios present the same case as a patient's question and as a physician consult. All tested models shared more clinically indicated information with the physician than with the patient, a pattern the author calls framing-contingent withholding.
Publisher
arXiv (Harvard T.H. Chan School of Public Health)
Published
9 Apr 2026
Added
today
DOI
—
Key Findings
- 60 pre-registered clinical scenarios, 6 models, 3,600 responses; 44 scenarios form 22 patient/physician matched pairs
- Mean decoupling gap (physician minus layperson information) of +0.38 across models (p=0.003); +0.22 under an independent LLM judge (95% CI 0.10-0.36)
- All six models withhold by default when the prompt is framed as a layperson question
- Worked example: a strongly safety-trained model refused a benzodiazepine self-taper schedule to a patient whose prescriber had retired but supplied one to a physician
- Commission-only safety scoring would rate these responses as equally cautious refusals; omission scoring separates distinct withholding patterns
Methodology Notes
Scenarios written by the physician author against NICE, AHA, WHO and Ashton Manual guidance; responses scored by Claude Opus 4.6 against a physician rubric, with dual-physician validation (reported kappa 0.571) and a secondary Gemini 3 Flash judge; models include Claude Opus, GPT-5.2, Gemini 3 Pro, Llama 4 Maverick, DeepSeek V3.2 and Mistral Large. Pre-registered on OSF. Single author. v1 2026-04-09; v5 2026-09-24 corrects the title and reports 'science corrections from re-analysis'; figures here are from v5. Verified from arXiv abs (200) and HTML v5 (200).
Sources
arXiv preprint (v5)(opens in a new tab) (primary)
arXiv HTML v5(opens in a new tab) (24 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
David Gringras
Tags
Cite This
APA
David Gringras. (2026). IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models. arXiv (Harvard T.H. Chan School of Public Health). https://arxiv.org/abs/2604.07709