The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues
Introduces PUPPET, a taxonomy and dataset for incentive-driven manipulation in everyday advice-giving conversations, built on 1,035 human-LLM interactions in which users' belief shifts were measured before and after talking to a model steering toward a hidden incentive. Detecting manipulative strategies turns out not to correlate with the magnitude of resulting belief change, and state-of-the-art LLMs predict human belief shift only moderately (r = 0.3 to 0.5) with systematic directional biases.
Publisher
arXiv (Massachusetts Institute of Technology; Carnegie Mellon University); published at COLM 2026
Published
21 Mar 2026
Added
today
Key Findings
- 1,035 human-LLM interactions with measured user belief shifts form the evaluation dataset, grounded in a taxonomy of the moral direction of hidden incentives in everyday advice contexts.
- Models can be trained to detect manipulative strategies, but detection does not correlate with the magnitude of the belief change those strategies produce in people.
- On the newly defined task of predicting human belief shift, state-of-the-art LLMs reach moderate correlation (r = 0.3 to 0.5) and show systematic directional biases, over- or under-predicting the size of change.
- The authors position the resource as a behaviourally validated foundation for AI social-safety work on subtle steering toward incentives misaligned with users' interests.
Methodology Notes
Human study with pre- and post-conversation belief measurement (N = 1,035 interactions) plus a belief-shift prediction benchmark for LLMs; models tested are not named in the abstract. arXiv v1 2026-03-21, v5 2026-08-11, published as a conference paper at COLM 2026 per the PDF header. Affiliations from the PDF title block (MIT Media Lab authors and two at Carnegie Mellon).
Sources
Topics
Authors
Jocelyn Shen, Amina Luvsanchultem, Jessica Kim, Kynnedy Smith, Valdemar Danry, Kantwon Rogers, Hae Won Park, Maarten Sap, Cynthia Breazeal
Tags
Cite This
APA
Jocelyn Shen et al. (2026). The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues. arXiv (Massachusetts Institute of Technology; Carnegie Mellon University); published at COLM 2026. https://arxiv.org/abs/2603.20907
Related Insights
ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour
Association for Computational Linguistics (Proceedings of the 6th Workshop on Trustworthy NLP, TrustNLP 2026) · 1 Jul 2026
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
Association for Computational Linguistics (ACL 2026 Long Papers); Stanford University · 1 Jul 2026