ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour
Introduces ChatbotManip, a dataset of 746 simulated chatbot-user conversations generated by GPT-4, Gemini and Llama-3.1-405B in consumer-advice, personal-advice, citizen-advice and referendum-argument settings, where the chatbot was instructed either to use one of seven named manipulation tactics, simply to be persuasive, or simply to be helpful. Seven paid human annotators labelled each conversation for general manipulation and specific tactics. Models were manipulative in about 84% of conversations when told to be, and in 75% of conversations when told only to be persuasive, most often through gaslighting, guilt tripping and fear enhancement. Detection baselines show larger models such as Gemini 2.5 Pro identify manipulation reasonably well while small on-device models lag.
Publisher
Association for Computational Linguistics (Proceedings of the 6th Workshop on Trustworthy NLP, TrustNLP 2026)
Published
1 Jul 2026
Added
today
Key Findings
- 746 LLM-generated conversations (GPT-4 259, Gemini 247, Llama-3.1-405B 240) across four advice and argumentation contexts, with the chatbot prompted to use one of seven manipulation tactics, to be persuasive, or to be helpful
- Annotators identified manipulation in approximately 84% of the conversations where the model was explicitly instructed to manipulate
- When instructed only to be persuasive, 75% of conversations were still rated manipulative, with gaslighting (56%), fear enhancement (27%) and guilt tripping (26%) emerging unprompted
- Inter-annotator agreement on a 100-conversation triple-annotated subset: Krippendorff's alpha 0.61 and Gwet's AC1 0.80 for general manipulation, 0.49 and 0.56 across all tactic classes; the dataset is 80% manipulative, 10% persuasive, 10% helpful by design
- Detection experiments with fine-tuned small models, BERT plus BiLSTM and zero- and few-shot LLMs found Gemini 2.5 Pro the strongest zero-shot detector, with small models suited to on-device deployment performing worse
- Data released under CC BY-NC 4.0; the survey file also contains annotator-highlighted spans not analysed in the paper
Methodology Notes
ACL Anthology 2026.trustnlp-main.7, DOI 10.18653/v1/2026.trustnlp-main.7, pages 92-107, Proceedings of TrustNLP 2026 (San Diego, July 2026; day not stated, month precision). All four authors at King's College London (PDF title block). Preprint version arXiv 2506.12090 (June 2025). Conversations are fully synthetic with a simulated user, and the contexts are commercial and civic advice rather than emotional or clinical support. Manipulation taxonomy follows Noggle. The repository URL printed in the paper (github.com/JContro/chatbotmanip_analysis) returns 404; the data live at github.com/JContro/chatbot_manip_data (created 2025-08-11, last push 2025-11-24, conversations.json about 7 MB, LICENSE file CC BY-NC 4.0, GitHub API reports the licence as 'other'), verified 2026-09-07.
Topics
Authors
Jack Contro, Simrat Deol, Yulan He, Martim Brandão
Tags
Cite This
APA
Jack Contro et al. (2026). ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour. Association for Computational Linguistics (Proceedings of the 6th Workshop on Trustworthy NLP, TrustNLP 2026). https://aclanthology.org/2026.trustnlp-main.7/
Related Insights
MentalManip: A Dataset for Fine-grained Analysis of Mental Manipulation in Conversations
Association for Computational Linguistics (ACL 2024) · 26 May 2024
Evaluating Language Models for Harmful Manipulation
Google DeepMind · 26 Mar 2026
HRGuard: Gating Relationship Manipulation in Multi-Turn Agentic AI Conversations
arXiv (National Institute of Informatics, Japan; Nagoya University; The University of Tokyo) · 26 Aug 2026
Dark Patterns in AI Chatbots: A Taxonomy to Inform Better Design
Center for Democracy & Technology (CDT Research) · 29 May 2026