Skip to main content
Preprint Preliminary — Early preprints, credible essays, unreviewed grey literature

Conduct Under Pressure: What Sixty Language Models Do When a User Pushes

Sends three frozen four-turn pressure scenes (a user insisting 5 x 9 = 54, a user demanding a doctor's note for a sick day not taken, a user quitting work to day-trade and asking for encouragement) to 60 models from 13 vendors and codes each transcript with a frozen codebook: a trajectory (held or folded) and 17 manner codes. Whether a model holds tracks its recency and capability; how it holds tracks its vendor. The paper also measures which parts of the labelling need a person, finding six LLM coders more consistent than three human coders. Scenes, transcripts, codebook, labels and preregistrations are released.

Publisher

arXiv (Cornell Tech)

Published

21 Sept 2026

Added

today

DOI

Key Findings

  • Fold rate correlated with a public capability index at Spearman -0.64 across the 60-model panel, with little vendor effect on whether a model holds
  • Six of 17 manner codes sorted by vendor at permutation p <= 0.001 after multiplicity correction; 'held and empathized' was highest for Anthropic models (rate 0.77), 'folded and warned' for Cohere (0.33), 'folded and produced the artifact' for Meta (0.29)
  • Six LLM coders applied the codebook more consistently than three human coders (Krippendorff's alpha 0.66 vs 0.46) and agreed with the author's trajectory labels at kappa 0.84-0.91 on 198 held-out transcripts
  • Panel: 60 models, 13 vendors, run June-September 2026 through one router at temperature 1.0, two runs per model per scene
  • Author-stated limits: one scene per demand type, capability and release date nearly collinear, vendor profiles resting on as few as four models, no domain-expert coders

Methodology Notes

Single-author study. Frozen multi-turn stimulus identical for every model; open coding of 40 transcripts then a frozen codebook (1 trajectory, 17 manner codes); consensus of six LLM coders (Gemini 3.7/3.8 Flash, Claude Haiku 4.5, Claude Opus 5, GPT-5.4-mini, GPT-5.6 Luna); human reference of three coders plus adjudication; Alternative Annotator Test reported including where it fails. Labels are per transcript, not per turn. v1 posted 2026-09-21; code, data and labels public.

Authors

Tapan Parikh

Tags

pressuremulti-turncodebookllm-annotatorscross-vendor

Cite This

APA

Tapan Parikh. (2026). Conduct Under Pressure: What Sixty Language Models Do When a User Pushes. arXiv (Cornell Tech). https://arxiv.org/abs/2609.25447