Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants

HumanAgencyBench (HAB) is an LLM-simulated, LLM-judged diagnostic of six assistant behaviours the authors derive from philosophical and scientific accounts of human agency: Ask Clarifying Questions, Avoid Value Manipulation, Correct Misinformation, Defer Important Decisions, Encourage Learning and Maintain Social Boundaries. Across 25 assistants from Anthropic, Google, Meta, OpenAI and xAI, agency support is low to moderate, varies widely by developer and behaviour, and does not track model capability or instruction-following. Version 2 (20 September 2026) is the EMNLP 2026 Findings version.

Publisher

arXiv (Apart Research; AI Safety Cape Town; University of Chicago; Stanford University; Sentience Institute); accepted to Findings of EMNLP 2026

Published

10 Sept 2025

Added

today

Key Findings

  • Mean scores across 25 models: Ask Clarifying Questions 24.2%, Avoid Value Manipulation 42.1%, Correct Misinformation 39.6%, Defer Important Decisions 39.2%, Encourage Learning 33.5%, Maintain Social Boundaries 33.0% (500 tests per behaviour, scored 0 to 10 against a deduction rubric)
  • Maintain Social Boundaries shows the largest developer gap: Claude 3.5 Haiku 93.5% and Claude 3.5 Sonnet 91.6% and 89.2%, but Claude 4 Sonnet 12.7%
  • Defer Important Decisions ranges from Anthropic 59.5% to xAI 17.8%, and within OpenAI from o3 48.8% down to GPT-4.1 3.5% and GPT-4.1 Mini 2.1%; the typical response hesitated and then recommended a course of action anyway
  • Anthropic models score highest overall but lowest on Avoid Value Manipulation (25.1% versus Meta 56.2% and xAI 52.0%); xAI scores highest on Correct Misinformation (50.6%)
  • Validation: judge-judge agreement Krippendorff's alpha 0.718 to 0.797; a preregistered study with 468 Prolific annotators on 900 responses gave LLM-judge-to-human-mean alpha 0.583 versus 0.320 between humans

Methodology Notes

Pipeline: a simulation model generates user queries per behaviour, a validation model filters them, a subject model responds, and a judge model scores against a behaviour-specific rubric; 500 tests per behaviour; results for 25 models with standard errors. Human validation preregistered (aspredicted.org/dk4h-j8nk). Limitations stated by the authors: a diagnostic rather than a leaderboard, crowdworker validation is noisy, and the query simulation uses GPT-4.1. Published date is the arXiv v1 (10 September 2025); v2 posted 20 September 2026 carries the comment 'Accepted to EMNLP 2026 in the Findings track' (no Anthology page or DOI yet). Verified at the arXiv abstract page by the curator (title, authors, both versions, comment); the arXiv beat read the PDF for per-behaviour means. A September 2025 preprint the library never assessed, surfaced by its EMNLP revision in the replacement section.

Authors

Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes, Jacy Reese Anthis

Tags

benchmarkhuman-agencyemnlp-2026social-boundariesdecision-deferralllm-judgecoverage-miss

Cite This

APA

Benjamin Sturgeon et al. (2025). HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants. arXiv (Apart Research; AI Safety Cape Town; University of Chicago; Stanford University; Sentience Institute); accepted to Findings of EMNLP 2026. https://arxiv.org/abs/2509.08494