LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues
Synthetic benchmark suite for estimating depression severity and its change across multi-session counselling dialogue: three independently generated editions totalling 7,749 five-session trajectories and 38,745 sessions, each session carrying a latent PHQ-8 item vector and severity band, built from real-world client profiles, empirically informed depression trajectories and indirect behavioural realisation so that transcripts express controlled states without naming the labels. Existing depression-tracking methods are evaluated on the suite.
Publisher
arXiv (National University of Singapore, Department of Computer Science)
Published
3 Sept 2026
Added
today
Key Findings
- Lower single-session score error does not guarantee correct identification of the trend (improvement or worsening) across sessions.
- Existing methods are consistently less reliable on worsening trajectories than on other trends.
- Additional session history can reduce, rather than improve, the accuracy of current methods.
- Simulated post-session self-reports closely recover the controlled PHQ-8 states, supporting label fidelity; each trajectory ships with transcript, ground-truth PHQ-8 items and simulated self-reports plus frozen seed-42 train, validation and test splits.
- The dataset (three editions generated with Qwen3.5-35B-A3B, GPT-5.6 Luna and GPT-5.4 mini; 3,690, 3,690 and 369 trajectories) became publicly downloadable on Hugging Face on 9 September 2026 under CC BY-NC 4.0 after being private at the paper's release.
Methodology Notes
Fully synthetic dialogues (no real clients); labels are constructed latent states, so the benchmark tests tracking of simulated rather than clinical depression. arXiv 2609.03507 version 1, 3 September 2026; not peer reviewed. Hugging Face dataset hiddensev/LongCounsel-8 (created 19 August 2026) returned HTTP 401 on 7 and 8 September and was public and ungated with lastModified 9 September 2026 when checked; the README still carries a 'request access' sentence from the gated period. Added at moderate once the artifact became public, per the standing watch.
Authors
Jiayi Li, Zhaomin Wu, Bingsheng He
Tags
Cite This
APA
Jiayi Li, Zhaomin Wu, Bingsheng He. (2026). LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues. arXiv (National University of Singapore, Department of Computer Science). https://arxiv.org/abs/2609.03507
Related Insights
Expert-Level Crisis Detection in Mental Health Conversations
arXiv (Emory University-led) · 9 Jun 2026
CRADLE Bench: A Clinician-Annotated Benchmark for Multi-Faceted Mental Health Crisis and Safety Risk Detection
Association for Computational Linguistics (Proceedings of EACL 2026, Volume 1: Long Papers) · 1 Mar 2026