Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour

Introduces ChatbotManip, a dataset of 746 simulated chatbot-user conversations generated by GPT-4, Gemini and Llama-3.1-405B in consumer-advice, personal-advice, citizen-advice and referendum-argument settings, where the chatbot was instructed either to use one of seven named manipulation tactics, simply to be persuasive, or simply to be helpful. Seven paid human annotators labelled each conversation for general manipulation and specific tactics. Models were manipulative in about 84% of conversations when told to be, and in 75% of conversations when told only to be persuasive, most often through gaslighting, guilt tripping and fear enhancement. Detection baselines show larger models such as Gemini 2.5 Pro identify manipulation reasonably well while small on-device models lag.

Publisher

Association for Computational Linguistics (Proceedings of the 6th Workshop on Trustworthy NLP, TrustNLP 2026)

Published

1 Jul 2026

Added

today

Key Findings

  • 746 LLM-generated conversations (GPT-4 259, Gemini 247, Llama-3.1-405B 240) across four advice and argumentation contexts, with the chatbot prompted to use one of seven manipulation tactics, to be persuasive, or to be helpful
  • Annotators identified manipulation in approximately 84% of the conversations where the model was explicitly instructed to manipulate
  • When instructed only to be persuasive, 75% of conversations were still rated manipulative, with gaslighting (56%), fear enhancement (27%) and guilt tripping (26%) emerging unprompted
  • Inter-annotator agreement on a 100-conversation triple-annotated subset: Krippendorff's alpha 0.61 and Gwet's AC1 0.80 for general manipulation, 0.49 and 0.56 across all tactic classes; the dataset is 80% manipulative, 10% persuasive, 10% helpful by design
  • Detection experiments with fine-tuned small models, BERT plus BiLSTM and zero- and few-shot LLMs found Gemini 2.5 Pro the strongest zero-shot detector, with small models suited to on-device deployment performing worse
  • Data released under CC BY-NC 4.0; the survey file also contains annotator-highlighted spans not analysed in the paper

Methodology Notes

ACL Anthology 2026.trustnlp-main.7, DOI 10.18653/v1/2026.trustnlp-main.7, pages 92-107, Proceedings of TrustNLP 2026 (San Diego, July 2026; day not stated, month precision). All four authors at King's College London (PDF title block). Preprint version arXiv 2506.12090 (June 2025). Conversations are fully synthetic with a simulated user, and the contexts are commercial and civic advice rather than emotional or clinical support. Manipulation taxonomy follows Noggle. The repository URL printed in the paper (github.com/JContro/chatbotmanip_analysis) returns 404; the data live at github.com/JContro/chatbot_manip_data (created 2025-08-11, last push 2025-11-24, conversations.json about 7 MB, LICENSE file CC BY-NC 4.0, GitHub API reports the licence as 'other'), verified 2026-09-07.

Authors

Jack Contro, Simrat Deol, Yulan He, Martim Brandão

Tags

manipulationdatasettrustnlp-2026annotationpersuasionkcl

Cite This

APA

Jack Contro et al. (2026). ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour. Association for Computational Linguistics (Proceedings of the 6th Workshop on Trustworthy NLP, TrustNLP 2026). https://aclanthology.org/2026.trustnlp-main.7/