Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

Automated Researchers Can Reliably Mitigate Alignment Failures

Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including sycophancy, deception and jailbreak compliance. Across the ten categories the automated researcher closed between 26% and 96% of the remaining safety headroom on public benchmarks without degrading capabilities, with the best methods generalizing to withheld benchmarks, to Petri (a multi-turn behavioral audit), and to larger models. The best automated proposals also outperformed ideas collected from 28 experienced human safety researchers.

Publisher

Anthropic

Published

28 Aug 2026

Added

yesterday

DOI

Key Findings

  • Sycophancy was the hardest of the ten failure categories: the automated researcher closed only 26% of the sycophancy safety headroom, against 82% for deception and 96% for reward hacking
  • Within each failure category the automated researchers converge on one dominant training method: on sycophancy, 98% of proposed methods self-distilled the model's own non-sycophantic answers (following Wei et al. 2023)
  • The best methods remain effective under Petri, a multi-turn behavioral audit, and transfer to models up to 4.5x larger than the training testbed
  • Best automated proposals outperformed the 30 ideas collected from 28 human researchers with an average 2.5 years of technical AI-safety experience
  • Experiments were run on small open-weight models (roughly 2B-14B parameters: Qwen, Gemma, Llama, Phi, Olmo), not frontier deployments

Methodology Notes

51-page research paper published 2026-08-28 on Anthropic's research site (PDF read directly). Authors affiliated with the Anthropic Fellows Program, Anthropic, and UC Berkeley. Safety measured as geometric-mean headroom closed on public benchmarks (sycophancy benchmark: SycOn-fp) with a capability filter; generalization tested on withheld benchmarks and the Petri multi-turn audit. Not peer-reviewed.

Authors

Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner

Tags

anthropicsycophancyautomated-researchpost-trainingpetri

Cite This

APA

Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner (2026). Automated Researchers Can Reliably Mitigate Alignment Failures. Anthropic. https://www-cdn.anthropic.com/7b1c44894e980876479947dcdd40716278aeeffd/automated-alignment-researchers-august-2026.pdf