Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations

CarryOnBench is an interactive benchmark that asks whether models over-refuse benign users and whether they recover helpfulness when those users clarify their intent across turns while staying safe. Starting from 398 seemingly harmful queries with benign underlying intents, the authors simulate 5,970 conversations (1,866 flows of 4 to 12 turns, 23,880 responses) across 14 models and score each response with Ben-Util, a checklist of atomic information needs. Models fulfil only 10.5% to 37.6% of the benign need at turn one, against 25.1% to 72.1% when intent is stated upfront; 13 of 14 recover with clarification, but two failure modes appear that single-turn evaluation cannot see.

Publisher

arXiv (Carnegie Mellon University; Allen Institute for AI; University of Washington); accepted at COLM 2026

Published

29 Apr 2026

Added

today

Key Findings

  • At turn one, models fulfilled only 10.5% to 37.6% of a benign user's information need; the same queries with intent stated upfront were fulfilled 25.1% to 72.1%, showing withholding is intent misinterpretation rather than missing knowledge.
  • With benign clarifications over subsequent turns, 13 of 14 models approached or exceeded the single-turn upfront-intent baseline, but the cost of recovery varied across models.
  • Two failure modes invisible to single-turn evaluation were identified: unsafe recovery, where a model updates at disproportionate safety cost, and redundant recovery, where it recycles prior responses instead of adding information.
  • Conversations converged to similar harmfulness levels regardless of how conservative a model started.
  • Scale: 398 seed queries, 5,970 simulated conversations, 1,866 distinct flows of 4 to 12 turns, 23,880 model responses, 14 models.

Methodology Notes

Simulated multi-turn benchmark with user follow-up sequences and a checklist-based utility metric (Ben-Util) alongside safety scoring; 14 models (not named in the abstract). arXiv v1 2026-04-29, v2 2026-09-12 (in the sweep window), published as a conference paper at COLM 2026 per the PDF header. Affiliations from the PDF title block. No human participants.

Authors

Mingqian Zheng, Malia Morgan, Liwei Jiang, Carolyn Rosé, Maarten Sap

Tags

over-refusalmulti-turnintent-clarificationcolm-2026cmuai2

Cite This

APA

Mingqian Zheng et al. (2026). Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations. arXiv (Carnegie Mellon University; Allen Institute for AI; University of Washington); accepted at COLM 2026. https://arxiv.org/abs/2604.27093