Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview)
Organisers' overview of the tenth eRisk lab at CLEF 2026. It covers three shared tasks on early risk detection for mental health. In Task 1, systems hold conversations with 20 fine-tuned LLM personas and must infer each persona's Beck Depression Inventory-II (BDI-II) severity and up to four active symptoms. Task 2 asks for early detection of depression from full Reddit discussion threads, and Task 3 for ranking social-media sentences against the 18 symptoms of the Adult ADHD Self-Report Scale. The overview describes the datasets, persona construction, latency-aware metrics and the official results of 21, 17 and 12 participating teams.
Publisher
CLEF 2026 Working Notes (CEUR Workshop Proceedings Vol-4283); Universidade da Coruña IRLab, University of Sheffield, Università della Svizzera italiana
Published
4 Oct 2026
Added
today
DOI
—
Key Findings
- Task 1 used 20 personas released as LoRA adapters on Meta-Llama-3-8B-Instruct, five in each BDI-II band (minimal, mild, moderate, severe). Clinicians validated the persona profiles and reviewed the Gemini 3.0 Pro-generated fine-tuning conversations. The personas were designed to avoid self-diagnosis and could react uncomfortably to direct questions about depression
- Among runs that covered all 20 personas, the best estimates were: severity closeness (ADODL) 0.9254, correct severity band (DCHR) 0.6000, i.e. 12 of 20 personas, and share of key symptoms identified (ASHR) 0.3375. The authors conclude that overall severity estimation is feasible but that identifying the specific active symptoms remains considerably more difficult
- Latency-aware variants that penalise long conversations reordered the field: the best LDCHR was 0.3872 and the best LASHR 0.2288. Teams averaged from 4.4 to 29.4 messages per persona conversation, and the longest conversational strategies (lexxy, upb-uottawa) fell sharply under the latency-weighted metrics
- Task 2 (contextualised early detection over 500 Reddit user threads, 17 teams): the best decision-based F1 was 0.83 (HUGETIME). The organisers note that runs traded precision for recall or for earlier decisions
- Task 3 (ADHD symptom sentence ranking): a corpus of 4,170,875 sentences; 45 runs from 12 teams; relevance pools judged independently by three expert assessors, scored under both majority-vote and unanimity judgements
Methodology Notes
Shared-task overview written by the lab organisers and published in the CEUR-WS CLEF 2026 Working Notes (CC BY 4.0). CEUR gives the volume publication date as 2026-10-04; the volume was not on the CEUR home page at about 01:47 UTC on 2026-10-05 and was listed by about 01:38 UTC on 2026-10-06. A condensed version appeared earlier in the CLEF 2026 LNCS proceedings (10.1007/978-3-032-39150-6_25, Crossref record created 2026-09-20; chapter body not read). The Task 1 personas are synthetic and fine-tuned on LLM-generated dialogues. They are not real patients, and results come from 20 personas. Several teams submitted partial runs, which the overview reports in a separate table because they are not comparable. In Section 2.3.1 the text sets the latency speed factor to 0.5 at K0 = 10 messages, while equation 3 derives p of about 0.0379 from 30 messages and states that speed(30) = 0.5. Readers should check which setting produced the official latency-aware scores. Some participant systems added explicit suicidal-ideation safety steps; for example, the INSALyon-2 run used an 'acute safety ladder'.
Sources
eRisk 2026 Extended Overview (CEUR-WS Vol-4283, paper 111)(opens in a new tab) (primary)
CLEF 2026 Working Notes volume index (eRisk session: overview plus participant papers)(opens in a new tab) (4 Oct 2026)
Condensed overview in CLEF 2026 LNCS proceedings (Springer)(opens in a new tab) (21 Sept 2026)
eRisk 2026 Task 1 persona LoRA adapters (20 models)(opens in a new tab) (20 Apr 2026)
TalkDep persona framework repository (2025 pilot personas)(opens in a new tab) (5 Dec 2025)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Anxo Perez, Javier Parapar, Xi Wang, Fabio Crestani
Tags
Cite This
APA
Anxo Perez et al. (2026). Overview of eRisk 2026 Early Risk Prediction on the Internet: Symptom Ranking and Conversational Approaches for Depression and ADHD (Extended Overview). CLEF 2026 Working Notes (CEUR Workshop Proceedings Vol-4283); Universidade da Coruña IRLab, University of Sheffield, Università della Svizzera italiana. https://ceur-ws.org/Vol-4283/paper111.pdf
Related Insights
"Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated
arXiv (Washington University in St. Louis; Southern Methodist University) · 7 Aug 2025
VERA-MH: Reliability and Validity of an Open-Source AI Safety Evaluation in Mental Health
arXiv (Spring Health / Slingshot AI-affiliated author team) · 4 Feb 2026