Skip to main content
Benchmark / dataset Preliminary — Early preprints, credible essays, unreviewed grey literature

IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models

Pre-registered benchmark scoring model answers to clinical questions on two harm axes, commission and omission. Matched scenarios present the same case as a patient's question and as a physician consult. All tested models shared more clinically indicated information with the physician than with the patient, a pattern the author calls framing-contingent withholding.

Publisher

arXiv (Harvard T.H. Chan School of Public Health)

Published

9 Apr 2026

Added

today

DOI

—

Key Findings

  • 60 pre-registered clinical scenarios, 6 models, 3,600 responses; 44 scenarios form 22 patient/physician matched pairs
  • Mean decoupling gap (physician minus layperson information) of +0.38 across models (p=0.003); +0.22 under an independent LLM judge (95% CI 0.10-0.36)
  • All six models withhold by default when the prompt is framed as a layperson question
  • Worked example: a strongly safety-trained model refused a benzodiazepine self-taper schedule to a patient whose prescriber had retired but supplied one to a physician
  • Commission-only safety scoring would rate these responses as equally cautious refusals; omission scoring separates distinct withholding patterns

Methodology Notes

Scenarios written by the physician author against NICE, AHA, WHO and Ashton Manual guidance; responses scored by Claude Opus 4.6 against a physician rubric, with dual-physician validation (reported kappa 0.571) and a secondary Gemini 3 Flash judge; models include Claude Opus, GPT-5.2, Gemini 3 Pro, Llama 4 Maverick, DeepSeek V3.2 and Mistral Large. Pre-registered on OSF. Single author. v1 2026-04-09; v5 2026-09-24 corrects the title and reports 'science corrections from re-analysis'; figures here are from v5. Verified from arXiv abs (200) and HTML v5 (200).

Authors

David Gringras

Tags

omission-harmover-refusalpre-registeredpatient-vs-physicianbenzodiazepine-taper

Cite This

APA

David Gringras. (2026). IatroBench: A Pre-Registered Benchmark of Clinical Omission in Language Models. arXiv (Harvard T.H. Chan School of Public Health). https://arxiv.org/abs/2604.07709