Skip to main content
Lab publication Credible — Major labs, established NGOs, reputable named-author preprints

Automated Researchers Can Reliably Mitigate Alignment Failures

Anthropic study of whether automated alignment researchers — Claude autonomously proposing and running post-training interventions — can mitigate ten benchmark-measurable alignment failures including sycophancy, deception and jailbreak compliance. Across the ten categories the automated researcher closed between 26% and 96% of the remaining safety headroom on public benchmarks without degrading capabilities, with the best methods generalizing to withheld benchmarks, to Petri (a multi-turn behavioral audit), and to larger models. The best automated proposals also outperformed ideas collected from 28 experienced human safety researchers.

Publisher

Anthropic

Published

28 Aug 2026

Added

1 week ago

DOI

Key Findings

  • Sycophancy was the hardest of the ten failure categories: the automated researcher closed only 26% of the sycophancy safety headroom, against 82% for deception and 96% for reward hacking
  • Within each failure category the automated researchers converge on one dominant training method: on sycophancy, 98% of proposed methods self-distilled the model's own non-sycophantic answers (following Wei et al. 2023)
  • The best methods remain effective under Petri, a multi-turn behavioral audit, and transfer to models up to 4.5x larger than the training testbed
  • Best automated proposals outperformed the 30 ideas collected from 28 human researchers with an average 2.5 years of technical AI-safety experience
  • Experiments were run on small open-weight models (roughly 2B-14B parameters: Qwen, Gemma, Llama, Phi, Olmo), not frontier deployments

Methodology Notes

51-page research paper published 2026-08-28 on Anthropic's research site (PDF read directly). Authors affiliated with the Anthropic Fellows Program, Anthropic, and UC Berkeley. Safety measured as geometric-mean headroom closed on public benchmarks (sycophancy benchmark: SycOn-fp) with a capability filter; generalization tested on withheld benchmarks and the Petri multi-turn audit. Not peer-reviewed.

Authors

Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner

Tags

anthropicsycophancyautomated-researchpost-trainingpetri

Cite This

APA

Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner. (2026). Automated Researchers Can Reliably Mitigate Alignment Failures. Anthropic. https://www-cdn.anthropic.com/7b1c44894e980876479947dcdd40716278aeeffd/automated-alignment-researchers-august-2026.pdf