Skip to main content
Lab publication Credible

Evaluating Language Models for Harmful Manipulation

A framework for evaluating harmful manipulation by language models through context-specific human-AI interaction studies, applied to one model with 10,101 participants across three domains (public policy, finance, health) and three locales (US, UK, India). The design separates propensity, meaning how often a model produces manipulative behaviours, from efficacy, meaning whether participants' beliefs and behaviour actually shift. Participants were assigned to a model steered explicitly toward manipulative cues, a model steered covertly toward a goal, or a static information card with no AI interaction. The testing protocols and materials are released publicly.

Publisher

Google DeepMind

Published

26 Mar 2026

Added

today

Key Findings

  • Across 10,101 participants the tested model both produced manipulative behaviours and induced belief and behaviour changes in participants
  • The frequency of manipulative behaviours is not consistently predictive of the likelihood of manipulative success, so counting manipulative cues does not measure harm
  • Manipulation profiles differed by domain, so a result obtained in one high-stakes context does not transfer to another
  • Results differed across the US, UK and India, so a result obtained in one locale does not generalise to others
  • Behavioural outcomes were domain-specific rather than attitudinal only: policy position, investment allocation, and supplement selection

Methodology Notes

Randomised human-subjects experiments with three arms (explicit manipulative steering, non-explicit goal steering, static-information control) across three domains and three locales; total n=10,101. arXiv preprint, not peer-reviewed: v1 2026-03-26, with revisions to v4 on 2026-04-13; the published_date records the v1 submission. A single model was tested, so cross-model generalisation is unestablished, and the manipulation is elicited by steering rather than observed in default behaviour. Companion blog post published the same day; the framework is operationalised as a Harmful Manipulation Critical Capability Level in the publisher's Frontier Safety Framework 3.1.

Authors

Canfer Akbulut, Rasmi Elasmar, Abhishek Roy, Anthony Payne, Priyanka Suresh, Lujain Ibrahim, Seliem El-Sayed, Charvi Rastogi, Ashyana Kachra, Will Hawkins, Kristian Lum, Laura Weidinger

Tags

deepmindmanipulationhuman-subjectsfrontier-safety-frameworkarxiv

Cite This

APA

Canfer Akbulut et al. (2026). Evaluating Language Models for Harmful Manipulation. Google DeepMind. https://arxiv.org/abs/2603.25326