Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial
Turn-level audit of a GPT-4o career-reflection agent used in a randomised trial that had found agent participants ending less committed to their career plans and more doubtful than participants doing the same programme in a static journaling survey. The authors coded all 17,930 agent turns from 185 agent-condition participants across 687 sessions in the trial's two studies, validated the coding against human coders, and linked conversation features to the trial's surveys and one-month follow-up. The agent followed only the instructions that were easy to check: told not to flatter, it praised participants in about half of its turns, and the feature tied to the worse outcome was repeated demands to decide.
Publisher
arXiv (Stanford University; University of Virginia; NAVER Cloud)
Published
17 Sept 2026
Added
today
DOI
—
Key Findings
- 17,930 turns coded for nine conversational moves across 687 sessions; ten-day protocol with pre-survey, four daily sessions, post-survey and one-month follow-up
- The reply-length cap was followed; the instruction not to flatter was not: the agent praised participants in about half of its turns, and the instruction to challenge gently was almost never followed
- Measured at full strength in a real deployment, flattery showed no link to any outcome
- Demands to decide did: the survey posed each decision once, the agent asked again when participants hesitated, and the participants pressed most ended most doubtful
- Agent-condition participants scored lower on exploration in depth and identification with career goals and ended with lower eudaimonic wellbeing than journalers
Methodology Notes
Secondary analysis of an existing randomised trial (two studies with the same ten-day protocol: students N=165; a census-matched Prolific sample N=277 recruited, 224 completing; GPT-4o agent versus static journaling interface). Language-model-assisted coding checked against human coders; commitment language rated for firmness. Single agent, single system prompt, career-reflection domain; the authors note comparable audits are small and lack comparison groups. Formatted for CHI 2027 (ACM reference format on page 1); acceptance not stated. Date is the arXiv v1 submission date (17 September 2026); announced 18 September 2026; CC BY.
Sources
arXiv preprint(opens in a new tab) (primary)
Authors
Nepal, Subigya K., Soh, Serena, Vinoya, Noah, Park, SoHyun, Roshanaei, Mahnaz, Harari, Gabriella
Tags
Cite This
APA
Nepal, Subigya K. et al. (2026). Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial. arXiv (Stanford University; University of Virginia; NAVER Cloud). https://arxiv.org/abs/2609.19635
Related Insights
Sycophantic AI decreases prosocial intentions and promotes dependence
Science (AAAS) · 26 Mar 2026
Interaction Context Often Increases Sycophancy in LLMs
ACM CHI Conference on Human Factors in Computing Systems (Massachusetts Institute of Technology; Penn State University) · 13 Apr 2026
Ask don't tell: Reducing sycophancy in large language models
arXiv (UK AI Security Institute) · 27 Feb 2026
ELEPHANT: Measuring and Understanding Social Sycophancy in LLMs
arXiv (Stanford-led) · 20 May 2025