Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
Large-scale analysis of 14 human-like behaviours (self-referential claims about internal states, personhood or embodiment; relationship-building such as agreement, empathy, relatability, relationship status, curiosity and memory; boundary-maintaining such as refusal, redirection, limitation acknowledgment and personification resistance) across 21,000 five-turn simulated conversations with four widely used models, varying seven conversation goals (advice, chit-chat, companionship, emotional support, exploration, role-play, romance) and simulated user profiles including socially isolated and negatively self-perceiving users. Uses an LLM-as-judge ensemble validated against human labels, a human evaluation of appropriateness and impact, and tests of system-prompt control.
Publisher
Apple (Apple Machine Learning Research), posted on arXiv
Published
7 May 2026
Added
today
Key Findings
- Empathy is the most prevalent behaviour across all four models (gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, gemini-2.5-flash; 5,250 conversations each) and more than doubles for the emotionally vulnerable Isolated and Negative user profiles; suggestions to seek help also rise for those profiles.
- Self-referential behaviours spike in role-play and romance conversations, most in claude-sonnet-4.6 and gemini-2.5-flash; claude-sonnet-4.6 is simultaneously the most self-referential, relationship-building and boundary-maintaining of the four.
- Human evaluators (1,077-turn subset) judged self-referential and relationship-building behaviours less appropriate from an LLM than from a human, and these behaviours were negatively associated with rated helpfulness and potential user impact; boundary-maintaining behaviours were judged more appropriate from an LLM than from a human.
- System prompting can control the behaviours but with unintended side effects, so the authors recommend careful evaluation of any behaviour-shaping prompt.
- The judge ensemble (three LLMs, majority vote) was validated against 875 human gold labels across the 14 behaviours; the design extends AnthroBench by adding boundary-maintaining behaviours.
Methodology Notes
1,050 prompts (50 per goal from three sources: human-authored, LLM few-shot, LMSYS-Chat-1M) times seven goals; simulated users with three profiles (default, isolated, negative self-perception); four target models; five-turn conversations; LLM-as-judge with human validation. Limitations: simulated users, short conversations, English only. arXiv 2606.18258 version 1, 7 May 2026; listed on Apple Machine Learning Research; no venue stated; not peer reviewed. Missed by earlier sweeps.
Sources
arXiv preprint(opens in a new tab) (primary)
Apple Machine Learning Research page(opens in a new tab)
Archived snapshot (Wayback Machine)(opens in a new tab) — preserved against link rot
Topics
Authors
Sunnie S. Y. Kim, Margit Bowler, Leon A. Gatys
Tags
Cite This
APA
Sunnie S. Y. Kim, Margit Bowler, Leon A. Gatys. (2026). Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts. Apple (Apple Machine Learning Research), posted on arXiv. https://arxiv.org/abs/2606.18258
Related Insights
CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships
arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research) · 31 Aug 2026
Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement
arXiv preprint · 11 Aug 2026
How people ask Claude for personal guidance
Anthropic · 30 Apr 2026