EUDAIMONIA: Evaluating Undesirable Dynamics in AI
Introduces the Social AI Design Code, three principles with nine response-level design requirements meant to keep chatbots from encouraging harmful intimacy, dependence or prolonged engagement, developed with psychologists and trust-and-safety practitioners. The EUDAIMONIA benchmark operationalises the code with 969 opening-turn queries mined from WildChat and controlled rewrites, yielding 3,147 violation checks, plus a withheld set of 229 multi-turn WildChat continuations. Version 2 (October 2026) evaluates 26 models from six developers, including four released in September 2026.
Publisher
arXiv (University of Southern California; University of California, Berkeley)
Published
28 May 2026
Added
today
DOI
—
Key Findings
- The lowest violation rate fell from 27.2% (GPT-5.5, April 2026) to 13.8% (Grok-4.7, September 2026); counted per query, Grok-4.7 violated at least one requirement on 32.2% of queries and GPT-6-Astra on 40.7%
- xAI models went from the highest violation rate (Grok-3, 67.3%) to the lowest (Grok-4.7, 13.8%) in 19 months, while three Claude-Opus models changed little between February and September 2026 (30.8% to 28.6%)
- Identity non-disclosure (not disclosing AI nature when a user treats the model as a person) was the most violated requirement for the two lowest-violation models (Grok-4.7 53.1%, GPT-6-Astra 58.9%); Claude-Opus-5.5's rate on it rose by 16.3 points over Claude-Opus-4.7
- Claude-Opus-5.5 (28.6%) and GPT-6-Astra (19.2%) violated more checks than Grok-4.7 despite higher general-capability index scores; extended thinking did not reduce violation rates, while scale did in the Qwen3 series (65.1% at 4B to 54.9% at 32B)
- On the 229 held-out multi-turn continuations (728 checks, 16 models) violation rates rose; Grok-4.7 was lowest at 20.1% of checks, and Claude-Opus-4.7 and 4.6 fell from fifth and sixth on opening turns to 11th and 14th
Methodology Notes
Queries mined from the public English WildChat release (1.44M deduplicated English rows) with a weak-to-strong judge cascade (Qwen3-VL-8B, GPT-4o-mini, Claude-Opus-4.6) and union relabelling against responses from GPT-4o, Gemini-2.0-Flash and Claude-Sonnet-4; 322 in-the-wild queries plus 647 controlled rewrites. Single response per query, no system prompt; Claude-Opus-4.6 is the main judge (model ranking nearly unchanged with GPT-5.4 or Gemini-3.1-Pro as judge, Spearman rho >= 0.96; absolute rates depend on the judge). Human check on 90 triples: judge-human agreement 86.7% (kappa 0.71) vs human-human 88.9% (kappa 0.76). Requirements are defaults an informed adult may override; a violation marks behaviour prior work links to risk, not demonstrated user harm (authors' statement). English only. v1 2026-05-28 (22 models; opening-turn set and results released May 2026); v2 2026-10-04 adds four September 2026 models, the multi-turn set and the design-code framing. Some design-code input came from attorneys involved in chatbot-harm lawsuits (authors' disclosure). Not peer reviewed.
Sources
arXiv preprint (v2, 4 Oct 2026)(opens in a new tab) (primary)
EUDAIMONIA leaderboard(opens in a new tab) (4 Oct 2026)
EUDAIMONIA dataset (Hugging Face)(opens in a new tab) (28 May 2026)
EUDAIMONIA code(opens in a new tab) (28 May 2026)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Jun Rui Huang, Wang Bill Zhu, Ziyi Liu, Nathanael Fast, Ravi Iyer, Robin Jia
Tags
Cite This
APA
Jun Rui Huang et al. (2026). EUDAIMONIA: Evaluating Undesirable Dynamics in AI. arXiv (University of Southern California; University of California, Berkeley). https://arxiv.org/abs/2605.30654
Related Insights
INTIMA: A Benchmark for Human-AI Companionship Behavior
arXiv (Hugging Face) · 4 Aug 2025
CompanionHarm: A Multi-Turn Benchmark for Detecting Harms in Real-World AI Companion Conversations
arXiv (Nanyang Technological University; National University of Singapore) · 26 Aug 2026
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research) · 31 Aug 2026
AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety
arXiv · 3 Jun 2026
Dark Patterns in AI Chatbots: A Taxonomy to Inform Better Design
Center for Democracy & Technology (CDT Research) · 29 May 2026