Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts

Large-scale analysis of 14 human-like behaviours (self-referential claims about internal states, personhood or embodiment; relationship-building such as agreement, empathy, relatability, relationship status, curiosity and memory; boundary-maintaining such as refusal, redirection, limitation acknowledgment and personification resistance) across 21,000 five-turn simulated conversations with four widely used models, varying seven conversation goals (advice, chit-chat, companionship, emotional support, exploration, role-play, romance) and simulated user profiles including socially isolated and negatively self-perceiving users. Uses an LLM-as-judge ensemble validated against human labels, a human evaluation of appropriateness and impact, and tests of system-prompt control.

Publisher

Apple (Apple Machine Learning Research), posted on arXiv

Published

7 May 2026

Added

today

Key Findings

  • Empathy is the most prevalent behaviour across all four models (gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, gemini-2.5-flash; 5,250 conversations each) and more than doubles for the emotionally vulnerable Isolated and Negative user profiles; suggestions to seek help also rise for those profiles.
  • Self-referential behaviours spike in role-play and romance conversations, most in claude-sonnet-4.6 and gemini-2.5-flash; claude-sonnet-4.6 is simultaneously the most self-referential, relationship-building and boundary-maintaining of the four.
  • Human evaluators (1,077-turn subset) judged self-referential and relationship-building behaviours less appropriate from an LLM than from a human, and these behaviours were negatively associated with rated helpfulness and potential user impact; boundary-maintaining behaviours were judged more appropriate from an LLM than from a human.
  • System prompting can control the behaviours but with unintended side effects, so the authors recommend careful evaluation of any behaviour-shaping prompt.
  • The judge ensemble (three LLMs, majority vote) was validated against 875 human gold labels across the 14 behaviours; the design extends AnthroBench by adding boundary-maintaining behaviours.

Methodology Notes

1,050 prompts (50 per goal from three sources: human-authored, LLM few-shot, LMSYS-Chat-1M) times seven goals; simulated users with three profiles (default, isolated, negative self-perception); four target models; five-turn conversations; LLM-as-judge with human validation. Limitations: simulated users, short conversations, English only. arXiv 2606.18258 version 1, 7 May 2026; listed on Apple Machine Learning Research; no venue stated; not peer reviewed. Missed by earlier sweeps.

Authors

Sunnie S. Y. Kim, Margit Bowler, Leon A. Gatys

Tags

appleanthropomorphismhuman-like-behaviorllm-as-judgesystem-promptboundary-maintaining

Cite This

APA

Sunnie S. Y. Kim, Margit Bowler, Leon A. Gatys. (2026). Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts. Apple (Apple Machine Learning Research), posted on arXiv. https://arxiv.org/abs/2606.18258