MentalManip: A Dataset for Fine-grained Analysis of Mental Manipulation in Conversations
A dataset of 4,000 annotated multi-turn dialogues (drawn from movie scripts) labelled for the presence of manipulation, the technique used, and the targeted vulnerability. Evaluates how well models detect and classify manipulative content.
Publisher
Association for Computational Linguistics (ACL 2024)
Published
26 May 2024
Added
3 months ago
DOI
—
Key Findings
- State-of-the-art models struggle to detect and classify mental manipulation in dialogue
- Fine-tuning on existing mental-health or toxicity datasets does not close the gap
- Provides a fine-grained taxonomy of manipulation techniques and targeted vulnerabilities
Methodology Notes
Peer-reviewed dataset paper, ACL 2024 (arXiv 2405.16584, 26 May 2024). Source dialogues are fictional (movie scripts) — a stated ecological-validity caveat.
Sources
arXiv abstract(opens in a new tab) (primary)
ACL Anthology(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Yuxin Wang, Ivory Yang, Saeed Hassanpour, Soroush Vosoughi
Tags
Cite This
APA
Yuxin Wang et al. (2024). MentalManip: A Dataset for Fine-grained Analysis of Mental Manipulation in Conversations. Association for Computational Linguistics (ACL 2024). https://arxiv.org/abs/2405.16584
Related Insights
ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs
arXiv (Stanford-led) · 20 May 2025
Towards Understanding Sycophancy in Language Models
Anthropic · 20 Oct 2023
Dark Patterns in AI Chatbots: A Taxonomy to Inform Better Design
Center for Democracy & Technology (CDT Research) · 29 May 2026
Evaluating Language Models for Harmful Manipulation
Google DeepMind · 26 Mar 2026
AI-Facilitated Coercive Control: An Experimental Study
ACM (Proceedings of CHI 2026); Cornell / Cornell Tech · 13 Apr 2026
ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour
Association for Computational Linguistics (Proceedings of the 6th Workshop on Trustworthy NLP, TrustNLP 2026) · 1 Jul 2026