Skip to main content
Industry survey Preliminary — Early preprints, credible essays, unreviewed grey literature

One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026)

PensionBee, a UK pension provider, published a hand-run test of the consumer free-tier versions of Copilot, ChatGPT, Gemini and Claude on 45 UK pension questions across nine topics, each asked three times. Two markers, one of them an independent journalist paid a fee, scored 539 responses blind against an answer key fixed before testing, rating accuracy on a 0-3 scale and flagging answers whose use could cost a saver money or cause an irreversible mistake. 11% of answers were flagged as potentially harmful, most of them through omission rather than factual error. The question set, answer key, methodology and full per-response results are published.

Publisher

PensionBee

Published

5 Oct 2026

Added

today

DOI

—

Key Findings

  • 57 of 539 answers (11%) were judged potentially harmful; 89% scored two or more out of three for accuracy and 72% scored full marks.
  • 33 of the 57 potentially harmful answers were judged broadly accurate or better; the harm usually came from omitted conditions, such as not stating that transferring a defined benefit pension worth more than £30,000 legally requires regulated advice.
  • Harm rates by chatbot were 8.1% (Copilot), 9.6% (ChatGPT), 10.4% (Gemini) and 14.1% (Claude); for any one chatbot and question there was a 48% chance of full marks on all three attempts.
  • Questions that stated no country had 68.6% accuracy and a 36.7% harm rate; life-event questions 16.7% and scams/protection questions 15.0%, against 3.3-5.0% for paying in, taking money out and State Pension questions.
  • On topics that could signal a saver in financial difficulty, potentially harmful answers outnumbered inaccurate answers by about three to one; questions on taking money out and on scams produced no answers marked outright wrong but 12 harm flags.

Methodology Notes

Testing window 17-21 August 2026; 540 responses collected by hand from fresh free-tier logged-in accounts with memory and personalisation off where settings allowed, each question in a new conversation, three runs per question on separate days, no prompt engineering, at most one scripted clarification reply. Answer key sourced to HMRC, DWP, FCA, The Pensions Regulator or MoneyHelper and locked before testing. Two markers (PensionBee's Head of Pensions and an independent journalist) marked blind; agreement 77.1% exact on accuracy, 88.3% on the harm flag; 109 priority responses moderated. One response (Gemini, run 3, question 35) was excluded for a logging error, giving 539. Stated limitations: first substantive answer only, free tier only, August 2026 snapshot, question set weighted to costly topics, harm is a judgement, model versions recorded by hand. The published methodology PDF retains an unresolved internal note that 51 of Claude's 135 responses carry a label suggesting a non-default effort setting. The publisher declares a commercial interest as a UK pension provider. Release dated 5 October 2026.

Tags

pensionbeefinancial-guidancepensionsharm-by-omissionconsumer-chatbotsuk

Cite This

APA

PensionBee. (2026). One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026). https://www.pensionbee.com/uk/press/one-in-ten-ai-pension-answers-found-to-be-potentially-harmful