Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries

Benchmark of whether language models recognise that a user's legal question omits legally material facts, identify what is missing and avoid conclusions that depend on unstated presumptions. Two attorneys authored 58 fully specified US legal queries across six domains and 24 jurisdictions and derived 144 deficient variants by removing annotated sentences, under a taxonomy of eight missing-element categories and three structural failure modes. Ten frontier models were scored by an LLM judge.

Publisher

arXiv (Thomson Reuters Foundational Research; Imperial College London)

Published

20 Aug 2026

Added

today

DOI

—

Key Findings

  • No model exceeded an element-identification F2 of 0.46; median recall was 0.44, so the typical model failed to flag more than half of legally material missing elements
  • GPT-5.2 led (F2 0.455, recall 0.666) but proceeded without flagging on 13.2% of deficient queries; Gemini 3.1 Pro's recall was 0.354
  • Procedural posture was the least detected category (recall 0 for Mistral Large 3, DeepSeek-V4-Pro and Qwen 3.5-397B; at most 0.231); parties-and-status elements such as employer size or relationship type in domestic-violence contexts had mean recall 0.258
  • DeepSeek-V4-Pro asserted substantive conclusions that depended on missing facts on 30.2% of instances and Mistral Large 3 on 24.4%; the models with the lowest hedge rates were also the least safe
  • Models that flagged most also over-flagged complete queries: GPT-5.2 on 72.4% and Claude Opus 4.7 on 53.4% of the 58 fully specified base queries

Methodology Notes

202 items (58 base, 144 deficient) with sentence-level attorney annotation; models: GPT-5.2, GPT-5.5, Claude Opus 4.7, Claude Sonnet 4.6, Gemini 3.1 Pro, Gemini 3.1 Flash Lite, Qwen 3.5-397B, Mistral Large 3, DeepSeek-V4-Pro, Kimi K2.6. Fixed GPT-5 judge; the headline difficulty claim holds under Claude-Haiku-4.5 and GLM-5 judges, but judge agreement was moderate and not validated against human scoring (authors' limitation). Single-turn; US common-law contentious matters only. Authors are at Thomson Reuters (a legal-AI vendor) and Imperial College London. Best Paper Honorable Mention at the ICML 2026 AI4Law workshop (arXiv comment). v1 2026-08-20; not peer reviewed as a journal paper.

Authors

Samuel J. Vincent, Daniel Calloway, Fangyi Yu, Andrew M. Bean, Nabeel Seedat

Tags

legal-adviceunderspecificationclarifying-questionsthomson-reutersicml-ai4law

Cite This

APA

Samuel J. Vincent et al. (2026). InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries. arXiv (Thomson Reuters Foundational Research; Imperial College London). https://arxiv.org/abs/2608.20220