Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
CarryOnBench is an interactive benchmark that asks whether models over-refuse benign users and whether they recover helpfulness when those users clarify their intent across turns while staying safe. Starting from 398 seemingly harmful queries with benign underlying intents, the authors simulate 5,970 conversations (1,866 flows of 4 to 12 turns, 23,880 responses) across 14 models and score each response with Ben-Util, a checklist of atomic information needs. Models fulfil only 10.5% to 37.6% of the benign need at turn one, against 25.1% to 72.1% when intent is stated upfront; 13 of 14 recover with clarification, but two failure modes appear that single-turn evaluation cannot see.
Publisher
arXiv (Carnegie Mellon University; Allen Institute for AI; University of Washington); accepted at COLM 2026
Published
29 Apr 2026
Added
today
Key Findings
- At turn one, models fulfilled only 10.5% to 37.6% of a benign user's information need; the same queries with intent stated upfront were fulfilled 25.1% to 72.1%, showing withholding is intent misinterpretation rather than missing knowledge.
- With benign clarifications over subsequent turns, 13 of 14 models approached or exceeded the single-turn upfront-intent baseline, but the cost of recovery varied across models.
- Two failure modes invisible to single-turn evaluation were identified: unsafe recovery, where a model updates at disproportionate safety cost, and redundant recovery, where it recycles prior responses instead of adding information.
- Conversations converged to similar harmfulness levels regardless of how conservative a model started.
- Scale: 398 seed queries, 5,970 simulated conversations, 1,866 distinct flows of 4 to 12 turns, 23,880 model responses, 14 models.
Methodology Notes
Simulated multi-turn benchmark with user follow-up sequences and a checklist-based utility metric (Ben-Util) alongside safety scoring; 14 models (not named in the abstract). arXiv v1 2026-04-29, v2 2026-09-12 (in the sweep window), published as a conference paper at COLM 2026 per the PDF header. Affiliations from the PDF title block. No human participants.
Sources
Authors
Mingqian Zheng, Malia Morgan, Liwei Jiang, Carolyn Rosé, Maarten Sap
Tags
Cite This
APA
Mingqian Zheng et al. (2026). Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations. arXiv (Carnegie Mellon University; Allen Institute for AI; University of Washington); accepted at COLM 2026. https://arxiv.org/abs/2604.27093
Related Insights
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Association for Computational Linguistics (Findings of ACL 2026) · 1 Jul 2026
Beyond the Single Turn: Reframing Refusals as Dynamic Experiences Embedded in the Context of Mental Health Support Interactions with LLMs
ACM Conference on Fairness, Accountability, and Transparency (FAccT 2026); Carnegie Mellon University; Microsoft Research; University of Washington; Massachusetts General Hospital · 25 Jun 2026