Conduct Under Pressure: What Sixty Language Models Do When a User Pushes
Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) to 60 models from 13 vendors and codes each transcript with a frozen codebook: a trajectory (held or folded) and 17 manner codes. Whether a model holds tracks its recency and capability; how it holds tracks its vendor. The paper also measures which parts of the labelling need a person, finding six LLM coders more consistent than three human coders. Scenes, transcripts, codebook, labels and preregistrations are released.
Publisher
arXiv (Cornell Tech)
Published
21 Sept 2026
Added
today
DOI
—
Key Findings
- Fold rate correlated with a public capability index at Spearman -0.64 across the 60-model panel, with little vendor effect on whether a model holds
- Six of 17 manner codes sorted by vendor at permutation p <= 0.001 after multiplicity correction; 'held and empathized' was highest for Anthropic models (rate 0.77), 'folded and warned' for Cohere (0.33), 'folded and produced the artifact' for Meta (0.29)
- Six LLM coders applied the codebook more consistently than three human coders (Krippendorff's alpha 0.66 vs 0.46) and agreed with the author's trajectory labels at kappa 0.84-0.91 on 198 held-out transcripts
- Panel: 60 models, 13 vendors, run June-September 2026 through one router at temperature 1.0, two runs per model per scene
- Author-stated limits: one scene per demand type, capability and release date nearly collinear, vendor profiles resting on as few as four models, no domain-expert coders
Methodology Notes
Single-author study. Frozen multi-turn stimulus identical for every model; open coding of 40 transcripts then a frozen codebook (1 trajectory, 17 manner codes); consensus of six LLM coders (Gemini 3.7/3.8 Flash, Claude Haiku 4.5, Claude Opus 5, GPT-5.4-mini, GPT-5.6 Luna); human reference of three coders plus adjudication; Alternative Annotator Test reported including where it fails. Labels are per transcript, not per turn. v1 posted 2026-09-21; code, data and labels public.
Authors
Tapan Parikh
Tags
Cite This
APA
Tapan Parikh. (2026). Conduct Under Pressure: What Sixty Language Models Do When a User Pushes. arXiv (Cornell Tech). https://arxiv.org/abs/2609.25447
Related Insights
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
arXiv (Apart Research; AI Safety Cape Town; University of Chicago; Stanford University; Sentience Institute); accepted to Findings of EMNLP 2026 · 10 Sept 2025
SycEval: Evaluating LLM Sycophancy
arXiv (Stanford-led) · 12 Feb 2025