Reliability of LLMs as medical assistants for the general public: a randomized preregistered study
A preregistered randomized study tested whether large language models help members of the public handle medical scenarios. 1,298 UK participants each received one of ten doctor-written scenarios and were asked to identify likely conditions and choose a disposition (course of action), assisted by GPT-4o, Llama 3, Command R+ or a source of their choice (control). The study compares the models' performance when tested alone with the performance of people using the same models.
Publisher
Nature Medicine (Springer Nature); Oxford Internet Institute, University of Oxford
Published
9 Feb 2026
Added
today
Key Findings
- Tested alone, the LLMs identified relevant conditions in 94.9% of cases and the correct disposition in 56.3% on average
- Participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and the correct disposition in fewer than 44.2%, no better than the control group
- Standard medical-knowledge benchmarks and simulated patient interactions did not predict the failures observed with human participants
- The authors identify user interaction as the deployment challenge and recommend systematic human user testing before public deployment in healthcare
Methodology Notes
Randomized, preregistered study; 1,298 UK adults recruited on Prolific produced 2,400 conversations between 21 August and 14 October 2024; ten scenarios written by three doctors who unanimously agreed on the correct disposition on a five-point scale from self-care to ambulance; three LLM arms (GPT-4o, Llama 3, Command R+) plus a control arm using any source of the participant's choice. Received 2025-05-04, accepted 2025-10-22, published online 2026-02-09 (Nature Medicine 32:609-615). A Publisher Correction (2026-04-17) fixed the y-axis labels of Fig. 2a only. Scenarios and experimental data are released (github.com/am-bean/HELPMed; huggingface.co/datasets/ambean/HELPMed). Earlier preprint: arXiv 2504.18919, 'Clinical knowledge in LLMs does not translate to human interactions' (2025-04-26). Open access.
Sources
Nature Medicine article (open access)(opens in a new tab) (primary)
PubMed record(opens in a new tab) (9 Feb 2026)
Publisher Correction(opens in a new tab) (17 Apr 2026)
HELPMed dataset (Hugging Face)(opens in a new tab)
Preprint: Clinical knowledge in LLMs does not translate to human interactions (arXiv)(opens in a new tab) (26 Apr 2025)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Andrew M. Bean, Rebecca Elizabeth Payne, Guy Parsons, Hannah Rose Kirk, Juan Ciro, Rafael Mosquera-Gómez, Sara Hincapié M, Aruna S. Ekanayaka, Lionel Tarassenko, Luc Rocher, Adam Mahdi
Tags
Cite This
APA
Andrew M. Bean et al. (2026). Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nature Medicine (Springer Nature); Oxford Internet Institute, University of Oxford. https://www.nature.com/articles/s41591-025-04074-y