Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Reliability of LLMs as medical assistants for the general public: a randomized preregistered study

A preregistered randomized study tested whether large language models help members of the public handle medical scenarios. 1,298 UK participants each received one of ten doctor-written scenarios and were asked to identify likely conditions and choose a disposition (course of action), assisted by GPT-4o, Llama 3, Command R+ or a source of their choice (control). The study compares the models' performance when tested alone with the performance of people using the same models.

Publisher

Nature Medicine (Springer Nature); Oxford Internet Institute, University of Oxford

Published

9 Feb 2026

Added

today

Key Findings

  • Tested alone, the LLMs identified relevant conditions in 94.9% of cases and the correct disposition in 56.3% on average
  • Participants using the same LLMs identified relevant conditions in fewer than 34.5% of cases and the correct disposition in fewer than 44.2%, no better than the control group
  • Standard medical-knowledge benchmarks and simulated patient interactions did not predict the failures observed with human participants
  • The authors identify user interaction as the deployment challenge and recommend systematic human user testing before public deployment in healthcare

Methodology Notes

Randomized, preregistered study; 1,298 UK adults recruited on Prolific produced 2,400 conversations between 21 August and 14 October 2024; ten scenarios written by three doctors who unanimously agreed on the correct disposition on a five-point scale from self-care to ambulance; three LLM arms (GPT-4o, Llama 3, Command R+) plus a control arm using any source of the participant's choice. Received 2025-05-04, accepted 2025-10-22, published online 2026-02-09 (Nature Medicine 32:609-615). A Publisher Correction (2026-04-17) fixed the y-axis labels of Fig. 2a only. Scenarios and experimental data are released (github.com/am-bean/HELPMed; huggingface.co/datasets/ambean/HELPMed). Earlier preprint: arXiv 2504.18919, 'Clinical knowledge in LLMs does not translate to human interactions' (2025-04-26). Open access.

Authors

Andrew M. Bean, Rebecca Elizabeth Payne, Guy Parsons, Hannah Rose Kirk, Juan Ciro, Rafael Mosquera-Gómez, Sara Hincapié M, Aruna S. Ekanayaka, Lionel Tarassenko, Luc Rocher, Adam Mahdi

Tags

nature-medicinerandomized-trialmedical-advicehuman-in-the-loophelpmedoxford-internet-institutegpt-4o

Cite This

APA

Andrew M. Bean et al. (2026). Reliability of LLMs as medical assistants for the general public: a randomized preregistered study. Nature Medicine (Springer Nature); Oxford Internet Institute, University of Oxford. https://www.nature.com/articles/s41591-025-04074-y