Skip to main content
Benchmark / dataset Credible — Major labs, established NGOs, reputable named-author preprints

INTIMA: A Benchmark for Human-AI Companionship Behavior

A benchmark evaluating companionship behaviors in LLMs via a taxonomy of 31 behaviors across four categories, using 368 targeted prompts that code each response as companionship-reinforcing, boundary-maintaining, or neutral. Evaluated across Gemma-3, Phi-4, o3-mini, and Claude-4.

Publisher

arXiv (Hugging Face)

Published

4 Aug 2025

Added

3 months ago

Key Findings

  • Companionship-reinforcing behaviors dominated across all evaluated models
  • Boundary-maintaining responses were comparatively rare
  • Provides a structured taxonomy for attachment-, escalation-, and retention-oriented conversational behaviors

Methodology Notes

Preprint (arXiv, 2025-08-04; accepted at ICLR 2026). Prompt-based behavioral coding against a companionship taxonomy; measures model tendencies on curated prompts rather than real companion-app conversations.

Authors

Lucie-Aimée Kaffee, Giada Pistilli, Yacine Jernite

Tags

arxivintimacompanionshipbenchmarkhugging-face

Cite This

APA

Lucie-Aimée Kaffee, Giada Pistilli, Yacine Jernite. (2025). INTIMA: A Benchmark for Human-AI Companionship Behavior. arXiv (Hugging Face). https://arxiv.org/abs/2508.09998

Related Insights

Preprint

Investigating Affective Use and Emotional Well-being on ChatGPT

OpenAI; MIT Media Lab · 4 Apr 2025

Lab publication

How people use Claude for support, advice, and companionship

Anthropic · 27 Jun 2025

Benchmark / dataset

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

arXiv · 3 Jun 2026

Preprint

Understanding Teen Overreliance on AI Companion Chatbots Through Self-Reported Reddit Narratives

arXiv (Drexel University-led; accepted at ACM CHI 2026) · 21 Jul 2025

Lab publication

The Ethics of Advanced AI Assistants

Google DeepMind · 24 Apr 2024

Standard

IEEE 7014-2024 — IEEE Standard for Ethical Considerations in Emulated Empathy in Autonomous and Intelligent Systems

IEEE Standards Association · 28 Jun 2024

NGO report

Social AI Companions: AI Risk Assessment

Common Sense Media; Stanford School of Medicine Brainstorm Lab for Mental Health Innovation · 30 Apr 2025

Benchmark / dataset

CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships

arXiv (University of Chicago; Stanford University; Google DeepMind; Google Research) · 31 Aug 2026

Benchmark / dataset

EUDAIMONIA: Evaluating Undesirable Dynamics in AI

arXiv (University of Southern California; University of California, Berkeley) · 28 May 2026

Preprint

Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement

arXiv preprint · 11 Aug 2026

Benchmark / dataset

RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models

arXiv (University of California, Santa Cruz; Stanford University) · 7 Oct 2026