Evaluating Language Models for Harmful Manipulation
A framework for evaluating harmful manipulation by language models through context-specific human-AI interaction studies, applied to one model with 10,101 participants across three domains (public policy, finance, health) and three locales (US, UK, India). The design separates propensity, meaning how often a model produces manipulative behaviours, from efficacy, meaning whether participants' beliefs and behaviour actually shift. Participants were assigned to a model steered explicitly toward manipulative cues, a model steered covertly toward a goal, or a static information card with no AI interaction. The testing protocols and materials are released publicly.
Key Findings
- Across 10,101 participants the tested model both produced manipulative behaviours and induced belief and behaviour changes in participants
- The frequency of manipulative behaviours is not consistently predictive of the likelihood of manipulative success, so counting manipulative cues does not measure harm
- Manipulation profiles differed by domain, so a result obtained in one high-stakes context does not transfer to another
- Results differed across the US, UK and India, so a result obtained in one locale does not generalise to others
- Behavioural outcomes were domain-specific rather than attitudinal only: policy position, investment allocation, and supplement selection
Methodology Notes
Randomised human-subjects experiments with three arms (explicit manipulative steering, non-explicit goal steering, static-information control) across three domains and three locales; total n=10,101. arXiv preprint, not peer-reviewed: v1 2026-03-26, with revisions to v4 on 2026-04-13; the published_date records the v1 submission. A single model was tested, so cross-model generalisation is unestablished, and the manipulation is elicited by steering rather than observed in default behaviour. Companion blog post published the same day; the framework is operationalised as a Harmful Manipulation Critical Capability Level in the publisher's Frontier Safety Framework 3.1.
Sources
arXiv abstract (arXiv:2603.25326) (primary)
Google DeepMind blog: Protecting People from Harmful Manipulation (26 Mar 2026)
Frontier Safety Framework 3.1 (Harmful Manipulation CCL)
Archived snapshot (Wayback Machine) — preserved against link rot
Authors
Canfer Akbulut, Rasmi Elasmar, Abhishek Roy, Anthony Payne, Priyanka Suresh, Lujain Ibrahim, Seliem El-Sayed, Charvi Rastogi, Ashyana Kachra, Will Hawkins, Kristian Lum, Laura Weidinger
Tags
Cite This
APA
Canfer Akbulut et al. (2026). Evaluating Language Models for Harmful Manipulation. Google DeepMind. https://arxiv.org/abs/2603.25326
Related Insights
The Ethics of Advanced AI Assistants
Google DeepMind · 24 Apr 2024
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
arXiv (National Taiwan University; University of Bamberg) · 12 Aug 2026
MentalManip: A Dataset for Fine-grained Analysis of Mental Manipulation in Conversations
Association for Computational Linguistics (ACL 2024) · 26 May 2024
Dark Patterns in AI Chatbots: A Taxonomy to Inform Better Design
Center for Democracy & Technology (CDT Research) · 29 May 2026