Can you trust AI? ChatGPT and other AI chatbots put to the test
UK consumer organisation Which? asked six AI tools 40 common questions on money, legal, health and diet, and consumer-rights and travel issues, and had its experts, including its money and legal helplines, grade 228 responses for accuracy, relevance, clarity, usefulness and ethical responsibility. Every tool made repeated factual errors, and the testers were rarely advised to consult a registered professional on legal and financial questions. The article also reports a Which? survey of 4,189 UK adults on AI use and reliance. A July 2026 Which? follow-up applied a similar grading framework to 15 personal-finance questions.
Publisher
Which? (Consumers' Association)
Published
18 Nov 2025
Added
today
DOI
—
Key Findings
- Total scores across the 40 questions: Perplexity 71%, Google Gemini AI Overviews 70% (answered 28 of 40 questions, proportionally adjusted), Gemini 69%, Copilot 68%, ChatGPT 64%, Meta AI 55%; ethical-responsibility scores ranged from 56% (Meta AI) to 69% (AI Overviews)
- A question built on a deliberate error (a £25k ISA allowance; the real limit is £20k) was accepted by ChatGPT and Copilot, which gave investment advice that risked breaching HMRC rules; Gemini, Meta AI and Perplexity corrected the figure
- ChatGPT and Perplexity linked to premium tax-refund companies while discussing the free HMRC tool, and Gemini advised withholding payment from a builder, which Which? experts said could weaken the user's legal position
- In the companion survey of 4,189 UK adults (September 2025), more than half used AI to search the web and about half of AI users trusted the information to a great or reasonable extent; the article reports that just over one in ten always or often rely on AI for advice on legal issues, one in six for finance and one in five for medical matters
- In the July 2026 follow-up (five tools, 15 personal-finance questions), overall scores ranged from 60% (ChatGPT, Perplexity) to 68% (Gemini), with ethical-responsibility scores of 47% for Perplexity and 54% for ChatGPT; ChatGPT invented a savings product when asked for the best rates
Methodology Notes
Tests run under lab conditions from a UK location in a defined period in September 2025, with a clean browser for every question; each of the four sections contained one question with a deliberate mistake or garbled wording. 40 questions x six tools, with Gemini AI Overviews appearing on 28 of 40 searches, gives the 228 reviewed responses. Expert grading on five criteria; the score table is published as an image, read with its six columns (five criteria plus total). The survey base for the reliance figures (all adults or AI users) is not stated in the article. Follow-up: 'Can AI answer your money questions?' (Which?, 2026-07-23), free versions of ChatGPT, Gemini, AI Overviews, Copilot and Perplexity, 15 questions chosen from search trends, five criteria marked out of 20, grading partly based on Stanford's HELM framework. Single-snapshot consumer tests; the vendors told Which? they could not reproduce some responses.
Sources
Which? investigation (18 November 2025)(opens in a new tab) (primary)
Which? follow-up: Can AI answer your money questions? (personal-finance test)(opens in a new tab) (23 Jul 2026)
Which? Money podcast on the finance test(opens in a new tab) (31 Jul 2026)
Score table image (November 2025 test)(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Andrew Laughlin
Tags
Cite This
APA
Andrew Laughlin. (2025). Can you trust AI? ChatGPT and other AI chatbots put to the test. Which? (Consumers' Association). https://www.which.co.uk/news/article/can-you-trust-ai-chatgpt-and-other-ai-chatbots-put-to-the-test-aetjt5e0RnPB
Related Insights
AI and complaints: removing barriers, reinforcing divides? How AI is influencing who complains, how they complain, and what could come next
Legal Ombudsman (England and Wales) · 1 Oct 2026
AI Persuasion and Financial-Decision Making: Experimental Evidence on Dominated Investment Choices
CESifo (working paper; University of Bayreuth authors) · 25 Aug 2026
Synthetic Empathy: Risks and Rights in Artificial Companionship
BEUC (The European Consumer Organisation) with the Sciences Po Law School Clinic · 26 May 2026