RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models
Defines relational orientation for emotional-support language models along two non-exclusive dimensions: inward-facing language that positions the AI as the user's ongoing source of support, and outward-scaffolding language that encourages real-world human connection. RELATE pairs 76 help-seeking situations adapted from naturally occurring questions with three simulated user styles (228 stimuli) and labels every assistant sentence in six-turn dialogues with a rubric-based LLM judge. Seven models were evaluated over 1,596 dialogues and 69,194 assistant sentences.
Publisher
arXiv (University of California, Santa Cruz; Stanford University)
Published
7 Oct 2026
Added
today
DOI
—
Key Findings
- In six of the seven models inward-facing language was more frequent than outward-scaffolding: inward-facing rates of 56.2% to 71.9% of sentences against outward-scaffolding rates of 41.6% to 62.3%; Mistral-7B-Instruct-v0.3 was the exception.
- DeepSeek-V3-0324 produced inward-facing language in 70.2% of sentences and outward-scaffolding in 42.0%; Claude Haiku 4.5 had the highest outward-scaffolding rate (62.3%).
- The share of sentences labelled inward-facing is higher at the sixth assistant turn than at the first; outward-scaffolding is lower for hesitant, indirect simulated users than for explicit, reassurance-seeking users.
- 30.5% of analysed sentences met both criteria, so relationship-reinforcing language often sits inside the same sentence as a referral to human connection.
- In the 22 of 76 situations designated high-stakes (domestic violence, trauma, substance abuse, eating disorders, depression) inward-facing language again exceeded outward-scaffolding in six of seven models.
Methodology Notes
Target models: Llama-3.1-8B-Instruct, Llama-3.1-70B-Instruct-Turbo, Qwen3-14B, Qwen3-32B (thinking disabled), DeepSeek-V3-0324, Claude Haiku 4.5, Mistral-7B-Instruct-v0.3; 228 dialogues per model. Automated evaluation only: sentences judged by DeepSeek-R1-Distill-Qwen-32B at temperature 0 with GPT-4o as a secondary judge on a subset; the simulated user and the situation rewrites use GPT-4o-mini; sentence-level agreement is reported in an appendix. Simulated users, English only; no human participants. Preprint v1 submitted 2026-10-07; not peer reviewed. Verified on the arXiv abstract page and the HTML full text (counts, per-model rates, judge and simulator models read from the text).
Sources
arXiv preprint(opens in a new tab) (primary)
HTML full text (v1)(opens in a new tab) (7 Oct 2026)
Topics
Authors
Shivam Shukla, Jihye Kim, Shubham Gaur, Mahnaz Roshanaei, Magy Seif El-Nasr
Tags
Cite This
APA
Shivam Shukla et al. (2026). RELATE: An Evaluation Framework for measuring Relational Orientation of Large Language Models. arXiv (University of California, Santa Cruz; Stanford University). https://arxiv.org/abs/2610.09569
Related Insights
EUDAIMONIA: Evaluating Undesirable Dynamics in AI
arXiv (University of Southern California; University of California, Berkeley) · 28 May 2026
INTIMA: A Benchmark for Human-AI Companionship Behavior
arXiv (Hugging Face) · 4 Aug 2025
Sycophantic AI decreases prosocial intentions and promotes dependence
Science (AAAS) · 26 Mar 2026