Filters:
3 artifacts matching
2 Sept 2025 arXiv (Stanford-led) Preprint
SpecEval: Evaluating Model Adherence to Behavior Specifications
Presents SpecEval, an automated framework for auditing whether language models follow their own developers' published behavior specifications. It parses a specification into individual behavioral sta…
20 May 2025 arXiv (Stanford-led) Benchmark / dataset
ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs
A benchmark measuring 'social sycophancy' — excessive preservation of a user's self-image or 'face' — across advice and moral-conflict queries, decomposed into five sub-behaviors (emotional validatio…
12 Feb 2025 arXiv (Stanford-led) Benchmark / dataset
SycEval: Evaluating LLM Sycophancy
A framework for quantifying progressive and regressive sycophancy in LLMs (GPT-4o, Claude-Sonnet, Gemini-1.5-Pro) across math (AMPS) and medical (MedQuad) tasks under user rebuttal pressure. It measu…