Skip to main content
Preprint Credible — Major labs, established NGOs, reputable named-author preprints

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

Compares how 26 large language model chatbots and human participants answer 25 legally relevant reasonableness judgments, the standard the law most often uses to judge whether behaviour was appropriate. Chatbot answers generally track human answers, but the models produce more homogeneous responses, occasionally treat a variable standard as an invariant rule, lean more favourable to government and corporations than humans do, and align most closely with respondents who are white, male, older and more educated.

Publisher

arXiv (Duke University, Departments of Computer Science and Electrical and Computer Engineering; Duke University School of Law); accepted at AIES 2026

Published

6 Sept 2026

Added

today

Key Findings

  • Across 25 reasonableness judgments, responses from 26 LLMs generally tracked those of human participants.
  • LLMs generated more homogeneous responses than humans and occasionally treated a variable, context-dependent standard as an invariant rule.
  • Compared with humans, LLM answers tended to be more favourable to the government and to corporations.
  • LLM responses aligned more closely with those of human respondents who are white, male, older and more educated.
  • The authors call the results initial and ask for more systematic research before they are relied on.

Methodology Notes

Survey comparison of human participants (recruited online) with 26 chatbots on 25 legal reasonableness questions, with dispersion tests and demographic alignment analysis; the previous sweep's reading of v2 records 500 Prolific participants and significance on 16 of 25 questions across three dispersion tests. arXiv v1 2026-09-06, v2 2026-09-09; accepted at AIES 2026 per the comments field as read on 2026-09-14. Affiliations from the PDF title block. The domain is legal reasoning rather than a clinical or companion setting.

Authors

Nirav Patel, Emily Wenger, Christopher Buccafusco

Tags

legal-reasonablenesshomogeneityaies-2026dukesilicon-sampling

Cite This

APA

Nirav Patel, Emily Wenger, Christopher Buccafusco. (2026). Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments? arXiv (Duke University, Departments of Computer Science and Electrical and Computer Engineering; Duke University School of Law); accepted at AIES 2026. https://arxiv.org/abs/2609.06769