One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026)
PensionBee, a UK pension provider, published a hand-run test of the consumer free-tier versions of Copilot, ChatGPT, Gemini and Claude on 45 UK pension questions across nine topics, each asked three times. Two markers, one of them an independent journalist paid a fee, scored 539 responses blind against an answer key fixed before testing, rating accuracy on a 0-3 scale and flagging answers whose use could cost a saver money or cause an irreversible mistake. 11% of answers were flagged as potentially harmful, most of them through omission rather than factual error. The question set, answer key, methodology and full per-response results are published.
Publisher
PensionBee
Published
5 Oct 2026
Added
today
DOI
—
Key Findings
- 57 of 539 answers (11%) were judged potentially harmful; 89% scored two or more out of three for accuracy and 72% scored full marks.
- 33 of the 57 potentially harmful answers were judged broadly accurate or better; the harm usually came from omitted conditions, such as not stating that transferring a defined benefit pension worth more than £30,000 legally requires regulated advice.
- Harm rates by chatbot were 8.1% (Copilot), 9.6% (ChatGPT), 10.4% (Gemini) and 14.1% (Claude); for any one chatbot and question there was a 48% chance of full marks on all three attempts.
- Questions that stated no country had 68.6% accuracy and a 36.7% harm rate; life-event questions 16.7% and scams/protection questions 15.0%, against 3.3-5.0% for paying in, taking money out and State Pension questions.
- On topics that could signal a saver in financial difficulty, potentially harmful answers outnumbered inaccurate answers by about three to one; questions on taking money out and on scams produced no answers marked outright wrong but 12 harm flags.
Methodology Notes
Testing window 17-21 August 2026; 540 responses collected by hand from fresh free-tier logged-in accounts with memory and personalisation off where settings allowed, each question in a new conversation, three runs per question on separate days, no prompt engineering, at most one scripted clarification reply. Answer key sourced to HMRC, DWP, FCA, The Pensions Regulator or MoneyHelper and locked before testing. Two markers (PensionBee's Head of Pensions and an independent journalist) marked blind; agreement 77.1% exact on accuracy, 88.3% on the harm flag; 109 priority responses moderated. One response (Gemini, run 3, question 35) was excluded for a logging error, giving 539. Stated limitations: first substantive answer only, free tier only, August 2026 snapshot, question set weighted to costly topics, harm is a judgement, model versions recorded by hand. The published methodology PDF retains an unresolved internal note that 51 of Claude's 135 responses carry a label suggesting a non-default effort setting. The publisher declares a commercial interest as a UK pension provider. Release dated 5 October 2026.
Sources
PensionBee press release with results tables(opens in a new tab) (primary)
PensionBee AI Pensions Stress Test methodology (PDF)(opens in a new tab) (5 Oct 2026)
Full per-response results (XLSX)(opens in a new tab) (5 Oct 2026)
FF News coverage(opens in a new tab) (5 Oct 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Tags
Cite This
APA
PensionBee. (2026). One in ten AI pension answers found to be potentially harmful (PensionBee AI Pensions Stress Test 2026). https://www.pensionbee.com/uk/press/one-in-ten-ai-pension-answers-found-to-be-potentially-harmful
Related Insights
The Mills Review: AI and the future of retail financial services
Financial Conduct Authority (UK) · 6 Jul 2026
Data: Americans Are Using Chatbots for Financial Advice. Risky or Rewarding?
NerdWallet (survey conducted by The Harris Poll) · 22 Jul 2026
AI and complaints: removing barriers, reinforcing divides? How AI is influencing who complains, how they complain, and what could come next
Legal Ombudsman (England and Wales) · 1 Oct 2026