Skip to main content
Peer-reviewed Authoritative — Peer-reviewed venues, standards bodies, regulators, official government publications

Beyond the Response: Examining Reasoning and Execution Fidelity in Large Language Models for Mental Health

Process-level observational audit of five state-of-the-art large language models on 75 high-risk mental-health prompts spanning depression, anxiety, trauma and emotional distress. The study treats self-generated chain-of-thought as an inspectable planning artifact, codes the commitments it states (empathy, guidance, boundary setting, problem framing) and checks with lexical and natural-language-inference alignment whether the final response upholds them, across nearly 400 planning steps.

Publisher

Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26) (ACM); Vanderbilt University

Published

25 Jun 2026

Added

today

Key Findings

  • Models differ systematically in which commitments they prioritise during planning
  • Recurrent cases were observed in which safety-relevant commitments stated during planning were omitted from the response: dropped disclaimers, compressed guidance, and generic emotional reassurance substituted for specific support
  • Linguistic instability during planning, such as increased hedging and lexical variability, co-occurred with commitment violations
  • Execution fidelity, the consistency between stated plan and delivered response, is proposed as an audit dimension beyond output accuracy or perceived empathy

Methodology Notes

Observational audit; five LLMs; 75 high-risk mental-health prompts; about 400 chain-of-thought planning steps coded qualitatively and checked with lexical analysis and NLI-based alignment. Several given names are recorded only as initials in the Semantic Scholar record. FAccT 2026 proceedings dated 2026-06-25; no preprint located. The ACM Digital Library blocks automated access; verified from the Crossref record and the Semantic Scholar DOI record (full abstract) on 2026-09-17.

Authors

Qadir, Sarvech, Ni, C., Vaidya, M., Ryu, Hye-Young, Mulvaney, S., Kantarcioglu, M., Novak, L., Malin, Bradley A., Rose, S., Yin, Z.

Tags

facct-2026chain-of-thoughtexecution-fidelityprocess-auditvanderbilt

Cite This

APA

Qadir, Sarvech et al. (2026). Beyond the Response: Examining Reasoning and Execution Fidelity in Large Language Models for Mental Health. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26) (ACM); Vanderbilt University. https://doi.org/10.1145/3805689.3806456