The widening evaluation gap in medical large language model research 2023 to 2026
A bibliometric analysis of 11,628 PubMed records on medical large language models from January 2023 to June 2026 across fourteen clinical domains, asking whether clinical evidence keeps pace with the models it evaluates. Only 2.5% of studies used a randomised, controlled or prospective design; the lag between a study's newest named model and its publication widened from 1.33 to 6.08 quarters; randomised trials evaluated models a median 4.6 quarters older than other designs and 62% of them evaluated a discontinued model family.
Publisher
arXiv (XU Exponential University of Applied Sciences; Abu Dhabi University; Zayed University)
Published
10 Sept 2026
Added
today
Key Findings
- 11,628 PubMed records from January 2023 to June 2026 across fourteen clinical domains, a 45-fold growth; 2.5% used a randomised, controlled or prospective design.
- Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters.
- Against a counterfactual holding model composition fixed, migration to newer systems offset only 56% of the drift (95% CI 50 to 65).
- Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), and 62% of randomised trials evaluated a discontinued model family.
- The authors conclude that rigour and currency are in tension and that the tension reflects model selection rather than research timelines.
Methodology Notes
Bibliometric study of PubMed records with model-name extraction and counterfactual drift modelling; the fourteen domains and the model-release calendar used are in the paper body. arXiv v1 2026-09-10; affiliations from the PDF title block (Potsdam, Abu Dhabi); no venue stated. No participant-facing data; the payload is a caveat on the medical-LLM evidence base.
Sources
Authors
Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif
Tags
Cite This
APA
Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif. (2026). The widening evaluation gap in medical large language model research 2023 to 2026. arXiv (XU Exponential University of Applied Sciences; Abu Dhabi University; Zayed University). https://arxiv.org/abs/2609.11770