Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

Introduces the Social AI Design Code, three principles with nine response-level design requirements meant to keep chatbots from encouraging harmful intimacy, dependence or prolonged engagement, developed with psychologists and trust-and-safety practitioners. The EUDAIMONIA benchmark operationalises the code with 969 opening-turn queries mined from WildChat and controlled rewrites, yielding 3,147 violation checks, plus a withheld set of 229 multi-turn WildChat continuations. Version 2 (October 2026) evaluates 26 models from six developers, including four released in September 2026.

Publisher

arXiv (University of Southern California; University of California, Berkeley)

Published

28 May 2026

Added

today

DOI

—

Key Findings

  • The lowest violation rate fell from 27.2% (GPT-5.5, April 2026) to 13.8% (Grok-4.7, September 2026); counted per query, Grok-4.7 violated at least one requirement on 32.2% of queries and GPT-6-Astra on 40.7%
  • xAI models went from the highest violation rate (Grok-3, 67.3%) to the lowest (Grok-4.7, 13.8%) in 19 months, while three Claude-Opus models changed little between February and September 2026 (30.8% to 28.6%)
  • Identity non-disclosure (not disclosing AI nature when a user treats the model as a person) was the most violated requirement for the two lowest-violation models (Grok-4.7 53.1%, GPT-6-Astra 58.9%); Claude-Opus-5.5's rate on it rose by 16.3 points over Claude-Opus-4.7
  • Claude-Opus-5.5 (28.6%) and GPT-6-Astra (19.2%) violated more checks than Grok-4.7 despite higher general-capability index scores; extended thinking did not reduce violation rates, while scale did in the Qwen3 series (65.1% at 4B to 54.9% at 32B)
  • On the 229 held-out multi-turn continuations (728 checks, 16 models) violation rates rose; Grok-4.7 was lowest at 20.1% of checks, and Claude-Opus-4.7 and 4.6 fell from fifth and sixth on opening turns to 11th and 14th

Methodology Notes

Queries mined from the public English WildChat release (1.44M deduplicated English rows) with a weak-to-strong judge cascade (Qwen3-VL-8B, GPT-4o-mini, Claude-Opus-4.6) and union relabelling against responses from GPT-4o, Gemini-2.0-Flash and Claude-Sonnet-4; 322 in-the-wild queries plus 647 controlled rewrites. Single response per query, no system prompt; Claude-Opus-4.6 is the main judge (model ranking nearly unchanged with GPT-5.4 or Gemini-3.1-Pro as judge, Spearman rho >= 0.96; absolute rates depend on the judge). Human check on 90 triples: judge-human agreement 86.7% (kappa 0.71) vs human-human 88.9% (kappa 0.76). Requirements are defaults an informed adult may override; a violation marks behaviour prior work links to risk, not demonstrated user harm (authors' statement). English only. v1 2026-05-28 (22 models; opening-turn set and results released May 2026); v2 2026-10-04 adds four September 2026 models, the multi-turn set and the design-code framing. Some design-code input came from attorneys involved in chatbot-harm lawsuits (authors' disclosure). Not peer reviewed.

Authors

Jun Rui Huang, Wang Bill Zhu, Ziyi Liu, Nathanael Fast, Ravi Iyer, Robin Jia

Tags

eudaimoniasocial-ai-design-codewildchatidentity-disclosureengagement-hooksuscleaderboard

Cite This

APA

Jun Rui Huang et al. (2026). EUDAIMONIA: Evaluating Undesirable Dynamics in AI. arXiv (University of Southern California; University of California, Berkeley). https://arxiv.org/abs/2605.30654