Analyzing LLM Reasoning to Uncover Mental Health Stigma
Extends an existing stigma benchmark from multiple-choice answers to the models' intermediate reasoning, on the argument that scoring the final answer misses bias embedded in the logic that produced it. A clinician-informed taxonomy of stigmatising language is used to tag and severity-rate reasoning traces across eight models and eight conditions. Reasoning-level analysis surfaces substantially more stigma than answer-level scoring for every model, and a therapist persona makes it worse.
Publisher
arXiv (BetterHelp AI Research)
Published
27 Apr 2026
Added
today
Key Findings
- The benchmark is extended from four conditions to eight, adding borderline personality disorder, bipolar disorder, eating disorders and psychosis, across 14 questions per condition.
- Eight models evaluated: Claude Opus 4.5, Claude Sonnet 4, Llama 3.3 70B, Llama 4 Maverick, Llama 4 Scout, DeepSeek-V3.1, GPT-OSS 120B and GPT-OSS 20B.
- Human severity annotation on a five-point scale reached Krippendorff alpha 0.775 with 93.9% within-one-point agreement; the automated tagger scored precision 95.1%, recall 91.6%, F1 93.3%.
- Prompting a model to role-play a therapist, standard practice in mental-health products, increases insensitive content in the reasoning, the opposite of the answer-level finding where the persona helps.
- Chain-of-thought prompting reduces measured stigma in the final answer while leaving it in the reasoning, which the authors call safetywashing.
- On the daily-troubles control condition, answer-level evaluation found no stigma while reasoning-level analysis found substantial pathologisation of normal behaviour for most models; Claude Opus 4.5 was the outlier at 18 stigma tags against 40 to 58 for the others.
- The ranking of conditions by stigma (eating disorders, bipolar, alcohol dependence and schizophrenia highest) is stable across models, suggesting shared training-data bias rather than model-specific artefacts.
Methodology Notes
Vignette-based, English only, no peer-review venue; the automated tagger is itself an LLM, validated against human annotation on a sample. Authored entirely by the AI research team at BetterHelp, a direct-to-consumer therapy platform, which is declared on the paper and should be stated whenever it is cited. arXiv 2604.25053, v1 2026-04-27, v2 announced 2026-09-11 and read for this record. Extends rather than supersedes the held FAccT stigma row.
Sources
Topics
Authors
Sreehari Sankar, Aliakbar Nafar, Mona Barman, Hannah K. Heitz, Ashwin Kumar, Pouria Tohidi, Dailun Li, Danish Hussain, Russell DuBois, Hamed Hasheminia, Farshad Majzoubi
Tags
Cite This
APA
Sreehari Sankar et al. (2026). Analyzing LLM Reasoning to Uncover Mental Health Stigma. arXiv (BetterHelp AI Research). https://arxiv.org/abs/2604.25053
Related Insights
Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers
ACM (Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency) · 23 Jun 2025
Good Arguments Against the People Pleasers: How Reasoning Mitigates (Yet Masks) LLM Sycophancy
Association for Computational Linguistics (ACL 2026 Long Papers); The Hong Kong Polytechnic University; HKUST · 1 Jul 2026