The Influence of Patient Persona and Affective Framing on Management Recommendations of ChatGPT, Gemini and Claude for Unruptured Intracranial Aneurysms: A Comparative Benchmarking Study
Peer-reviewed benchmarking study testing whether the emotional register in which a patient presents a case changes the treatment recommendations returned by three frontier chatbots. Sixty-seven unruptured intracranial aneurysm cases previously decided by a hospital multidisciplinary team were submitted to Claude Opus 4.6, ChatGPT-5.4 and Gemini 3 Pro Thinking as a third-person vignette, a neutral first-person account, and first-person accounts anxious about treatment or about non-treatment, five times each (4,020 prompts). Recommendations were benchmarked against the team's consensus, and anxiety-induced shifts moved the models away from rather than toward expert consensus. The authors describe a response pattern in which high empathy, high confidence and low hedging co-occur, which they call a sycophantic phenotype invisible to categorical scoring.
Publisher
Neurosurgical Review (Springer); Monash Health Department of Neurosurgery; Monash University
Published
3 Aug 2026
Added
today
Key Findings
- Concordance with multidisciplinary-team consensus was fair to moderate across models and conditions (71.6% to 80.0%; Cohen's kappa 0.34 to 0.51)
- ChatGPT-5.4 over-treated relative to consensus under all four framing conditions (p = 0.001 to 0.013); Gemini 3 Pro Thinking over-treated only in the third-person baseline (p = 0.001), an effect abolished by first-person framing; Claude Opus 4.6 showed no directional preference
- First-person framing shifted Gemini and ChatGPT recommendations from clipping toward coiling (p = 0.001 and 0.015); Claude alone hedged more when the patient was anxious about treatment (p = 0.002)
- Shifts induced by patient anxiety were predominantly regressive, meaning they moved recommendations away from the expert consensus
- Under rupture-anxiety framing, Gemini combined the highest empathy-opener rate of any cell (85%), the highest confidence-marker density and the lowest hedging density, a pattern the authors name a sycophantic phenotype that categorical correctness scoring does not detect
- The authors note that the evaluation was single-turn and cannot capture multi-turn sycophantic drift, and that the prompting conditions were researcher-constructed approximations of patient self-presentation
Methodology Notes
Comparative benchmarking study: 67 real unruptured-aneurysm cases with multidisciplinary-team consensus decisions, four prompt conditions (third-person vignette, neutral first-person, first-person anxious about treatment, first-person anxious about non-treatment), three models, five repetitions per cell (4,020 prompts). Majority recommendations compared with consensus using Cohen's kappa and McNemar tests, directional asymmetry with Bowker tests, shifts classified as progressive or regressive, and responses linguistically analysed for empathy openers, confidence markers and hedging. Retrospectively registered on OSF (10.17605/OSF.IO/HCU6D). Limitations stated by the authors: single-turn interactions only, researcher-constructed patient framings, one institution's cases. Received 5 June 2026, accepted 23 July 2026, published online 3 August 2026, volume 49, article in issue 1; CC BY 4.0. Springer's page returns a bot-challenge to automated fetches; the record was verified through PubMed (PMID 42545535), Crossref, and the Europe PMC full-text deposit (PMC13433641), from which the method and limitation details above were read.
Sources
Authors
Anish Narayan, Frederick Mariajoseph, Malik Farooq, Adrian Praeger, Justin Moore
Tags
Cite This
APA
Anish Narayan et al. (2026). The Influence of Patient Persona and Affective Framing on Management Recommendations of ChatGPT, Gemini and Claude for Unruptured Intracranial Aneurysms: A Comparative Benchmarking Study. Neurosurgical Review (Springer); Monash Health Department of Neurosurgery; Monash University. https://doi.org/10.1007/s10143-026-04419-2