Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Audit of whether model performance measured through developer APIs reflects the behaviour of the consumer chat interfaces people actually use. Sends identical prompts to ChatGPT, Claude and Gemini through the API and through the chat interface across seven systems and nine benchmarks spanning general capability, social bias (BBQ) and sycophancy (AITA items from ELEPHANT), using fresh anonymous accounts with personalisation disabled.

Publisher

arXiv (Stanford University)

Published

8 Sept 2026

Added

today

Key Findings

  • API evaluations score 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test-retest agreement than the corresponding interface evaluations on average.
  • For ChatGPT, the difference between API and interface access exceeds the API-only difference between GPT 5.3 and GPT 5.4, so switching access surface can degrade measured performance as much as downgrading a model generation.
  • Varying system prompts, sampling parameters and reasoning settings through the API shifts behaviour in some cases but does not reliably eliminate the gap.
  • Data were collected through five ChatGPT Enterprise, three Claude Pro and three Google AI Pro accounts, 200 sampled items per benchmark (100 flipped AITA pairs), five runs per item.

Methodology Notes

Paired API and interface runs on identical prompts; benchmarks include BBQ, AITA-NTA from ELEPHANT and AA-Omniscience; interface runs through anonymous consumer accounts with chat history and memory disabled; no user population. arXiv 2609.08861 version 1, 8 September 2026 (announced 9 September); not peer reviewed. Authors are at Stanford (RegLab and computer science).

Authors

Jennifer Wang, Joachim Baumann, Daniel E. Ho, Sanmi Koyejo

Tags

context-validityapi-vs-interfacestanfordbenchmark-validity

Cite This

APA

Jennifer Wang et al. (2026). API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces. arXiv (Stanford University). https://arxiv.org/abs/2609.08861