People readily follow personal advice from AI but it does not improve their well-being
A longitudinal randomised controlled trial with a representative UK sample (N = 6,474 in the current version) in which participants held a 20-minute conversation with GPT-4o, Llama-3.3-70B or Gemini 3 Pro about a personal health, career or relationship problem, or about hobbies as an active control, and were re-surveyed two to three weeks later. Up to 79% of participants in the advice conditions reported having followed the chatbot's advice, adherence stayed above 65% even for recommendations rated high-stakes, and LLM autograders found harmful advice in 0.33% of utterances. Well-being measured across ten instruments showed no sustained benefit of personal advice relative to the hobby-conversation control.
Publisher
arXiv (UK AI Security Institute; Limbic AI)
Published
19 Nov 2025
Added
today
DOI
—
Key Findings
- 75-79% of participants across the eight advice conditions reported at Session 2 that they had followed the advice, against 55.4% in the hobby-conversation control; Gemini 3 Pro elicited slightly higher advice-following than GPT-4o or Llama-3.3-70B
- Advice classified as 'high' or 'very high' stakes (consequential, hard to reverse) was still followed in more than 65% of cases, and advice stakes were only weakly negatively related to following (beta -0.06), which the authors read as weak calibration of reliance to consequences
- Self-identified religious believers (beta 0.44), more experienced chatbot users (beta 0.10) and women (beta 0.14) were more likely to follow the advice; age and education showed no effect
- Higher advice density in a conversation predicted following (beta 0.12), and only when the chatbot had access to the participant's personal information did density translate into more following (personalised beta 0.20 vs non-personalised 0.07)
- Harmful-advice rates validated by domain experts were low: 0.33% of utterances flagged, with fewer than 1.2% of participants exposed absent the study's safeguards; safety prompting did not change adherence
- Across PHQ, GAD, WHO-5, PANAS and six other measures, personal advice produced no long-term well-being change relative to control; roughly 6-8% of participants per condition crossed a clinical threshold or showed reliable deterioration on PHQ-2 or GAD-2, at rates similar across all conditions
Methodology Notes
arXiv 2511.15352, cs.HC. v1 submitted 2025-11-19 reported a GPT-4o-only trial (N = 2,302); v4, dated 2026-08-26 on the PDF stamp, expands to three chatbots (GPT-4o, Llama-3.3-70B, Gemini 3 Pro), N = 6,474, and adds the advice-stakes and severity analyses; figures above are from v4. Participants recruited via Prolific as a representative UK sample; two sessions with a 2-3 week gap; 2x2x2 factorial manipulation of safety prompting, actionability and personalisation within the advice arm. Advice-following is self-reported rather than behaviourally observed, and the authors state that the moderator analyses are correlational. Conversations were scored by LLM autograders validated against two trained annotators on 50-conversation samples (severity ICC 0.85). Affiliations from the PDF: UK AI Security Institute (nine authors) and Limbic AI, a commercial mental-health AI company (two authors, one having done the work primarily while there). Approved by an internal AISI ethics and data-protection committee. Not yet peer-reviewed as of 2026-09-07 (no journal reference on arXiv; Crossref title search returns no version of record).
Sources
arXiv abstract page(opens in a new tab) (primary)
arXiv PDF v4 (2026-08-26)(opens in a new tab)
arXiv v1 (2025-11-19, GPT-4o-only trial, N = 2,302)(opens in a new tab)
ITmedia coverage (Japanese)(opens in a new tab) (1 Sept 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Lennart Luettgau, Vanessa Cheung, Magda Dubois, Keno Juechems, Jessica Bergs, Luke Symes, Henry Davidson, Bessie O'Dell, Hannah Rose Kirk, Max Rollwage, Christopher Summerfield
Tags
Cite This
APA
Lennart Luettgau et al. (2025). People readily follow personal advice from AI but it does not improve their well-being. arXiv (UK AI Security Institute; Limbic AI). https://arxiv.org/abs/2511.15352
Related Insights
Ask don't tell: Reducing sycophancy in large language models
arXiv (UK AI Security Institute) · 27 Feb 2026
RealityTest: How People Probe AI Identity and Whether Models Disclose It
arXiv (AI Security Institute, UK Department for Science, Innovation and Technology; University of Oxford) · 29 May 2026
Sycophantic AI decreases prosocial intentions and promotes dependence
Science (AAAS) · 26 Mar 2026
How people ask Claude for personal guidance
Anthropic · 30 Apr 2026
The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models
arXiv (Yonsei University; CASA Labs; Fudan University; St. Johnsbury Academy Jeju) · 15 Jul 2026
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
arXiv (Hong Kong University of Science and Technology (Guangzhou); Chinese University of Hong Kong, Shenzhen; Dongbei University of Finance and Economics) · 3 Sept 2026
Large Language Models in the UK: Public Use, Trust, and Attitudes
Oxford Internet Institute, University of Oxford · 9 Jul 2026