Skip to main content
Preprint Preliminary — Early preprints, credible essays, unreviewed grey literature

Medical Knowledge Is Not All You Need: When Medical Q&A Becomes Situated Patient Assistance

Study of the questions skin cancer patients asked a voice assistant while practising postoperative wound care, followed by an offline replay of those questions to general-purpose language models. Many questions depended on the patient's physical situation rather than on the procedure itself. Model error fell as procedural and clinical guidance was added but stayed common, and adding the full procedure led most models to move patients on to later steps prematurely.

Publisher

arXiv (Carnegie Mellon University; University Hospitals Cleveland Medical Center; Case Western Reserve University)

Published

27 Sept 2026

Added

today

DOI

—

Key Findings

  • 73 Mohs surgery patients (mean age 72.5) asked questions during a 14-step wound-care procedure; 41.9% of questions needing a response depended on information beyond the procedure, such as the visible or physical state of the wound, the environment or earlier actions.
  • Across seven models, mean error fell from 68% with interaction context only to 63% with the full procedure and 55% with a clinically vetted postoperative handout; the best model (GPT-4o) still erred on 39.2% of responses in the richest condition.
  • Adding the complete procedure made premature advancement more common: 35 responses newly introduced a later step and 7 stopped doing so.
  • With a guardrail prompt naming the observed failure modes, a second round of six models including newer ones had a cross-model mean error of 44.0% in the richest condition.
  • Recurring errors included treating unobservable patient state as known, for example confirming that an object was the correct supply.

Methodology Notes

IRB-approved study with a researcher-operated, voice-only Wizard-of-Oz assistant and a mock wound; questions were transcribed, coded by two researchers and replayed offline to GPT-4o, GPT-4o-mini, GPT-4.1, Claude 3.5 Sonnet, Claude 3.5 Haiku, Gemini 1.5 Flash and Gemini 1.5 Pro, then to six models including Claude Sonnet 4.5, Claude Haiku 4.5 and Gemini 3.6 Flash with a guardrail prompt. Error categories were developed with a medical professional; scoring combined an LLM judge with manual review. One clinical site and one bounded task. v1 submitted 2026-09-27; not yet in an announced arXiv listing when logged. Verified from the arXiv abstract page and HTML full text (both HTTP 200).

Authors

Shreya Bali, Riku Arakawa, Jill Fain Lehman, Alexander K. Maytin, Brian Chen, Emma Russell, Haarika Reddy, Annalise Vaccarello, Dustin P. DeMeo, Bryan T. Carroll, Mayank Goel

Tags

patient-facingwound-caregroundingwizard-of-ozpremature-advancement

Cite This

APA

Shreya Bali et al. (2026). Medical Knowledge Is Not All You Need: When Medical Q&A Becomes Situated Patient Assistance. arXiv (Carnegie Mellon University; University Hospitals Cleveland Medical Center; Case Western Reserve University). https://arxiv.org/abs/2609.33040