Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment
Audits 31 pre-specified NLP techniques from seven methodological families (model scaling, synthetic data, loss reweighting, ensembling, structured prediction, threshold tuning and LLM methods) in roughly 300 controlled experiments on 1,635 clinician-annotated social-media posts with author-disjoint partitions, for three coupled outputs: four-level suicide risk, evidence spans and 24 clinical risk and protective factors. Reports which techniques hold up under severe class imbalance and small author-level data and describes the resulting task-grounded system.
Publisher
arXiv (Indian AI Research Organisation; Ahmedabad University; University of Maryland, Baltimore County)
Published
7 Sept 2026
Added
today
Key Findings
- Only 5 of 31 technique comparisons produced reliable gains in this small-data, high-imbalance regime; the authors found no prior audit of the standard playbook here.
- Factor prediction reformulated as entailment between each post and its codebook definitions, with a class-balanced ensemble and score rescaling, outperformed direct classification; deployment-consistent calibration (matching validation thresholds to test-time ensemble scores) gave the largest improvement to the factor system.
- Risk predictions condition a seven-model evidence-tagger ensemble, evidence restricts symbolic risk rules, and a difficult risk class is routed separately; the factor predictor stays independent because risk evidence adds no factor signal.
- Final scores: 0.8203 for risk, 0.7953 for evidence and 0.7045 for factors; the work was shaped by the IEEE Big Data Cup, where the team ranked third.
Methodology Notes
Controlled ablation study on a clinician-annotated social-media dataset (1,635 posts, author-disjoint splits); competition-derived; no deployment or user evaluation. arXiv 2609.07766 version 1, 7 September 2026 (announced 9 September); 21 pages; not peer reviewed.
Sources
arXiv preprint(opens in a new tab) (primary)
PDF (21 pages)(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Shlok Shelat, Shrey Salvi, Souvik Roy, Manas Gaur, Amit Sheth
Tags
Cite This
APA
Shlok Shelat et al. (2026). Bag of Tricks or Bag of Myths? Reducing Modeling Complexity with Task Knowledge in Explainable Suicide Risk Assessment. arXiv (Indian AI Research Organisation; Ahmedabad University; University of Maryland, Baltimore County). https://arxiv.org/abs/2609.07766
Related Insights
Exploring generalizability and explainability of LLMs in classifying clinically rated suicidal ideation using heterogeneous data
npj Digital Medicine (Springer Nature) · 5 Sept 2026
Evaluating Large Language Models in Crisis Detection: A Real-World Benchmark from Psychological Support Hotlines (PsyCrisisBench)
arXiv (Chinese research team) · 2 Jun 2025