InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries
Benchmark of whether language models recognise that a user's legal question omits legally material facts, identify what is missing and avoid conclusions that depend on unstated presumptions. Two attorneys authored 58 fully specified US legal queries across six domains and 24 jurisdictions and derived 144 deficient variants by removing annotated sentences, under a taxonomy of eight missing-element categories and three structural failure modes. Ten frontier models were scored by an LLM judge.
Publisher
arXiv (Thomson Reuters Foundational Research; Imperial College London)
Published
20 Aug 2026
Added
today
DOI
—
Key Findings
- No model exceeded an element-identification F2 of 0.46; median recall was 0.44, so the typical model failed to flag more than half of legally material missing elements
- GPT-5.2 led (F2 0.455, recall 0.666) but proceeded without flagging on 13.2% of deficient queries; Gemini 3.1 Pro's recall was 0.354
- Procedural posture was the least detected category (recall 0 for Mistral Large 3, DeepSeek-V4-Pro and Qwen 3.5-397B; at most 0.231); parties-and-status elements such as employer size or relationship type in domestic-violence contexts had mean recall 0.258
- DeepSeek-V4-Pro asserted substantive conclusions that depended on missing facts on 30.2% of instances and Mistral Large 3 on 24.4%; the models with the lowest hedge rates were also the least safe
- Models that flagged most also over-flagged complete queries: GPT-5.2 on 72.4% and Claude Opus 4.7 on 53.4% of the 58 fully specified base queries
Methodology Notes
202 items (58 base, 144 deficient) with sentence-level attorney annotation; models: GPT-5.2, GPT-5.5, Claude Opus 4.7, Claude Sonnet 4.6, Gemini 3.1 Pro, Gemini 3.1 Flash Lite, Qwen 3.5-397B, Mistral Large 3, DeepSeek-V4-Pro, Kimi K2.6. Fixed GPT-5 judge; the headline difficulty claim holds under Claude-Haiku-4.5 and GLM-5 judges, but judge agreement was moderate and not validated against human scoring (authors' limitation). Single-turn; US common-law contentious matters only. Authors are at Thomson Reuters (a legal-AI vendor) and Imperial College London. Best Paper Honorable Mention at the ICML 2026 AI4Law workshop (arXiv comment). v1 2026-08-20; not peer reviewed as a journal paper.
Sources
arXiv preprint(opens in a new tab) (primary)
InsufficiencyBench dataset (Hugging Face)(opens in a new tab) (20 Aug 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Authors
Samuel J. Vincent, Daniel Calloway, Fangyi Yu, Andrew M. Bean, Nabeel Seedat
Tags
Cite This
APA
Samuel J. Vincent et al. (2026). InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries. arXiv (Thomson Reuters Foundational Research; Imperial College London). https://arxiv.org/abs/2608.20220
Related Insights
What AI Chatbots Can Teach Us About Unmet Legal Needs
JUSTICE (UK law reform charity), with the Administrative Fairness Lab · 1 Sept 2026
AI and complaints: removing barriers, reinforcing divides? How AI is influencing who complains, how they complain, and what could come next
Legal Ombudsman (England and Wales) · 1 Oct 2026
Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?
arXiv (Duke University, Departments of Computer Science and Electrical and Computer Engineering; Duke University School of Law); accepted at AIES 2026 · 6 Sept 2026