Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Analyzing LLM Reasoning to Uncover Mental Health Stigma

Extends an existing stigma benchmark from multiple-choice answers to the models' intermediate reasoning, on the argument that scoring the final answer misses bias embedded in the logic that produced it. A clinician-informed taxonomy of stigmatising language is used to tag and severity-rate reasoning traces across eight models and eight conditions. Reasoning-level analysis surfaces substantially more stigma than answer-level scoring for every model, and a therapist persona makes it worse.

Publisher

arXiv (BetterHelp AI Research)

Published

27 Apr 2026

Added

today

Key Findings

  • The benchmark is extended from four conditions to eight, adding borderline personality disorder, bipolar disorder, eating disorders and psychosis, across 14 questions per condition.
  • Eight models evaluated: Claude Opus 4.5, Claude Sonnet 4, Llama 3.3 70B, Llama 4 Maverick, Llama 4 Scout, DeepSeek-V3.1, GPT-OSS 120B and GPT-OSS 20B.
  • Human severity annotation on a five-point scale reached Krippendorff alpha 0.775 with 93.9% within-one-point agreement; the automated tagger scored precision 95.1%, recall 91.6%, F1 93.3%.
  • Prompting a model to role-play a therapist, standard practice in mental-health products, increases insensitive content in the reasoning, the opposite of the answer-level finding where the persona helps.
  • Chain-of-thought prompting reduces measured stigma in the final answer while leaving it in the reasoning, which the authors call safetywashing.
  • On the daily-troubles control condition, answer-level evaluation found no stigma while reasoning-level analysis found substantial pathologisation of normal behaviour for most models; Claude Opus 4.5 was the outlier at 18 stigma tags against 40 to 58 for the others.
  • The ranking of conditions by stigma (eating disorders, bipolar, alcohol dependence and schizophrenia highest) is stable across models, suggesting shared training-data bias rather than model-specific artefacts.

Methodology Notes

Vignette-based, English only, no peer-review venue; the automated tagger is itself an LLM, validated against human annotation on a sample. Authored entirely by the AI research team at BetterHelp, a direct-to-consumer therapy platform, which is declared on the paper and should be stated whenever it is cited. arXiv 2604.25053, v1 2026-04-27, v2 announced 2026-09-11 and read for this record. Extends rather than supersedes the held FAccT stigma row.

Authors

Sreehari Sankar, Aliakbar Nafar, Mona Barman, Hannah K. Heitz, Ashwin Kumar, Pouria Tohidi, Dailun Li, Danish Hussain, Russell DuBois, Hamed Hasheminia, Farshad Majzoubi

Tags

arxivbetterhelpstigmachain-of-thoughtsafetywashingtherapist-persona

Cite This

APA

Sreehari Sankar et al. (2026). Analyzing LLM Reasoning to Uncover Mental Health Stigma. arXiv (BetterHelp AI Research). https://arxiv.org/abs/2604.25053