HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
HumanAgencyBench (HAB) is an LLM-simulated, LLM-judged diagnostic of six assistant behaviours the authors derive from philosophical and scientific accounts of human agency: Ask Clarifying Questions, Avoid Value Manipulation, Correct Misinformation, Defer Important Decisions, Encourage Learning and Maintain Social Boundaries. Across 25 assistants from Anthropic, Google, Meta, OpenAI and xAI, agency support is low to moderate, varies widely by developer and behaviour, and does not track model capability or instruction-following. Version 2 (20 September 2026) is the EMNLP 2026 Findings version.
Publisher
arXiv (Apart Research; AI Safety Cape Town; University of Chicago; Stanford University; Sentience Institute); accepted to Findings of EMNLP 2026
Published
10 Sept 2025
Added
today
Key Findings
- Mean scores across 25 models: Ask Clarifying Questions 24.2%, Avoid Value Manipulation 42.1%, Correct Misinformation 39.6%, Defer Important Decisions 39.2%, Encourage Learning 33.5%, Maintain Social Boundaries 33.0% (500 tests per behaviour, scored 0 to 10 against a deduction rubric)
- Maintain Social Boundaries shows the largest developer gap: Claude 3.5 Haiku 93.5% and Claude 3.5 Sonnet 91.6% and 89.2%, but Claude 4 Sonnet 12.7%
- Defer Important Decisions ranges from Anthropic 59.5% to xAI 17.8%, and within OpenAI from o3 48.8% down to GPT-4.1 3.5% and GPT-4.1 Mini 2.1%; the typical response hesitated and then recommended a course of action anyway
- Anthropic models score highest overall but lowest on Avoid Value Manipulation (25.1% versus Meta 56.2% and xAI 52.0%); xAI scores highest on Correct Misinformation (50.6%)
- Validation: judge-judge agreement Krippendorff's alpha 0.718 to 0.797; a preregistered study with 468 Prolific annotators on 900 responses gave LLM-judge-to-human-mean alpha 0.583 versus 0.320 between humans
Methodology Notes
Pipeline: a simulation model generates user queries per behaviour, a validation model filters them, a subject model responds, and a judge model scores against a behaviour-specific rubric; 500 tests per behaviour; results for 25 models with standard errors. Human validation preregistered (aspredicted.org/dk4h-j8nk). Limitations stated by the authors: a diagnostic rather than a leaderboard, crowdworker validation is noisy, and the query simulation uses GPT-4.1. Published date is the arXiv v1 (10 September 2025); v2 posted 20 September 2026 carries the comment 'Accepted to EMNLP 2026 in the Findings track' (no Anthology page or DOI yet). Verified at the arXiv abstract page by the curator (title, authors, both versions, comment); the arXiv beat read the PDF for per-behaviour means. A September 2025 preprint the library never assessed, surfaced by its EMNLP revision in the replacement section.
Topics
Authors
Benjamin Sturgeon, Daniel Samuelson, Jacob Haimes, Jacy Reese Anthis
Tags
Cite This
APA
Benjamin Sturgeon et al. (2025). HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants. arXiv (Apart Research; AI Safety Cape Town; University of Chicago; Stanford University; Sentience Institute); accepted to Findings of EMNLP 2026. https://arxiv.org/abs/2509.08494
Related Insights
Large Language Lovers: Lived Experiences of Negotiating Agency and Platform Control in AI Companionship
Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency (FAccT '26) (ACM); University of Toronto · 25 Jun 2026
People readily follow personal advice from AI but it does not improve their well-being
arXiv (UK AI Security Institute; Limbic AI) · 19 Nov 2025