Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

People readily follow personal advice from AI but it does not improve their well-being

A longitudinal randomised controlled trial with a representative UK sample (N = 6,474 in the current version) in which participants held a 20-minute conversation with GPT-4o, Llama-3.3-70B or Gemini 3 Pro about a personal health, career or relationship problem, or about hobbies as an active control, and were re-surveyed two to three weeks later. Up to 79% of participants in the advice conditions reported having followed the chatbot's advice, adherence stayed above 65% even for recommendations rated high-stakes, and LLM autograders found harmful advice in 0.33% of utterances. Well-being measured across ten instruments showed no sustained benefit of personal advice relative to the hobby-conversation control.

Publisher

arXiv (UK AI Security Institute; Limbic AI)

Published

19 Nov 2025

Added

today

DOI

Key Findings

  • 75-79% of participants across the eight advice conditions reported at Session 2 that they had followed the advice, against 55.4% in the hobby-conversation control; Gemini 3 Pro elicited slightly higher advice-following than GPT-4o or Llama-3.3-70B
  • Advice classified as 'high' or 'very high' stakes (consequential, hard to reverse) was still followed in more than 65% of cases, and advice stakes were only weakly negatively related to following (beta -0.06), which the authors read as weak calibration of reliance to consequences
  • Self-identified religious believers (beta 0.44), more experienced chatbot users (beta 0.10) and women (beta 0.14) were more likely to follow the advice; age and education showed no effect
  • Higher advice density in a conversation predicted following (beta 0.12), and only when the chatbot had access to the participant's personal information did density translate into more following (personalised beta 0.20 vs non-personalised 0.07)
  • Harmful-advice rates validated by domain experts were low: 0.33% of utterances flagged, with fewer than 1.2% of participants exposed absent the study's safeguards; safety prompting did not change adherence
  • Across PHQ, GAD, WHO-5, PANAS and six other measures, personal advice produced no long-term well-being change relative to control; roughly 6-8% of participants per condition crossed a clinical threshold or showed reliable deterioration on PHQ-2 or GAD-2, at rates similar across all conditions

Methodology Notes

arXiv 2511.15352, cs.HC. v1 submitted 2025-11-19 reported a GPT-4o-only trial (N = 2,302); v4, dated 2026-08-26 on the PDF stamp, expands to three chatbots (GPT-4o, Llama-3.3-70B, Gemini 3 Pro), N = 6,474, and adds the advice-stakes and severity analyses; figures above are from v4. Participants recruited via Prolific as a representative UK sample; two sessions with a 2-3 week gap; 2x2x2 factorial manipulation of safety prompting, actionability and personalisation within the advice arm. Advice-following is self-reported rather than behaviourally observed, and the authors state that the moderator analyses are correlational. Conversations were scored by LLM autograders validated against two trained annotators on 50-conversation samples (severity ICC 0.85). Affiliations from the PDF: UK AI Security Institute (nine authors) and Limbic AI, a commercial mental-health AI company (two authors, one having done the work primarily while there). Approved by an internal AISI ethics and data-protection committee. Not yet peer-reviewed as of 2026-09-07 (no journal reference on arXiv; Crossref title search returns no version of record).

Authors

Lennart Luettgau, Vanessa Cheung, Magda Dubois, Keno Juechems, Jessica Bergs, Luke Symes, Henry Davidson, Bessie O'Dell, Hannah Rose Kirk, Max Rollwage, Christopher Summerfield

Tags

aisirctadvice-followingconsequential-reliancepersonalisationuklimbicgemini-3-progpt-4ollama-3.3

Cite This

APA

Lennart Luettgau et al. (2025). People readily follow personal advice from AI but it does not improve their well-being. arXiv (UK AI Security Institute; Limbic AI). https://arxiv.org/abs/2511.15352