Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

The widening evaluation gap in medical large language model research 2023 to 2026

A bibliometric analysis of 11,628 PubMed records on medical large language models from January 2023 to June 2026 across fourteen clinical domains, asking whether clinical evidence keeps pace with the models it evaluates. Only 2.5% of studies used a randomised, controlled or prospective design; the lag between a study's newest named model and its publication widened from 1.33 to 6.08 quarters; randomised trials evaluated models a median 4.6 quarters older than other designs and 62% of them evaluated a discontinued model family.

Publisher

arXiv (XU Exponential University of Applied Sciences; Abu Dhabi University; Zayed University)

Published

10 Sept 2026

Added

today

Key Findings

  • 11,628 PubMed records from January 2023 to June 2026 across fourteen clinical domains, a 45-fold growth; 2.5% used a randomised, controlled or prospective design.
  • Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters.
  • Against a counterfactual holding model composition fixed, migration to newer systems offset only 56% of the drift (95% CI 50 to 65).
  • Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), and 62% of randomised trials evaluated a discontinued model family.
  • The authors conclude that rigour and currency are in tension and that the tension reflects model selection rather than research timelines.

Methodology Notes

Bibliometric study of PubMed records with model-name extraction and counterfactual drift modelling; the fourteen domains and the model-release calendar used are in the paper body. arXiv v1 2026-09-10; affiliations from the PDF title block (Potsdam, Abu Dhabi); no venue stated. No participant-facing data; the payload is a caveat on the medical-LLM evidence base.

Authors

Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif

Tags

bibliometricsevaluation-lagmedical-llmrandomised-trialsmodel-discontinuation

Cite This

APA

Raad Bin Tareaf, Murad Al-Rajab, Samia Loucif. (2026). The widening evaluation gap in medical large language model research 2023 to 2026. arXiv (XU Exponential University of Applied Sciences; Abu Dhabi University; Zayed University). https://arxiv.org/abs/2609.11770