craft 4 min read

Anthropic's AI agents beat 28 human experts at alignment research, and humans became the baseline

7 of 7 alignment failures with a human baseline were won by the automated researchers over 28 experts

Automated researchers ran the alignment research loop and beat a 28-expert human baseline; humans defined the failures and judged the results

Automation has reached the last human stage of the AI training loop: Anthropic’s automated alignment researchers, agents that run the safety-research loop end to end, beat proposals from 28 experienced AI safety researchers on all seven alignment failures where human ideas were collected. Anthropic published the study on August 28, 2026, and on September 15 Surge AI, the expert data vendor that ran its human baseline, described that work. Alignment research is the work of finding methods that stop models from deceiving, flattering, or gaming their graders. Across ten alignment failures (ways a model’s behavior departs from what its developers intend), including deception, sycophancy, jailbreaks, privacy violations, and reward hacking, the agents found methods that improved safety benchmarks without hurting general capability. On September 9 this site wrote that one stage of the loop still needs people, deciding what counts as correct. This study moves that boundary: proposing and testing fixes got automated, and humans still choose which failures matter, build the benchmarks, and judge whether the outcomes count as better.

Key Takeaways

  • Automated researchers improved safety benchmarks on ten alignment failures, and the strongest methods kept their gains on held-out evaluations and on models up to 4.7 times larger than the ones they were tuned against.
  • The 28-expert human baseline lost on all seven failures where it existed; the study itself notes the humans had one attempt each within eight hours.
  • The human gate narrowed to specifying what failure means, building the benchmarks, and adjudicating improvement, which is the expert work the environment market sells.

What did the automated researchers do?

The agents run every step of the research:

  1. Read the literature.
  2. Propose a mitigation.
  3. Train a model with it.
  4. Evaluate against safety and capability benchmarks.
  5. Keep what works and repeat.

No human picks the next idea. Per the study, the resulting methods generalized to held-out benchmarks, open-ended audits, and models 4.7 times larger than the optimization targets.

“Claude agents that search the literature, propose alignment methods, train models”

Surge AI (@HelloSurgeAI) · September 15, 2026 · on X

Is beating 28 experts a fair comparison?

Not as a contest of human insight against machine insight, and Anthropic says so. Each human researcher submitted a single proposal within an eight-hour window; the agents proposed, tested, and refined for as long as the pipeline ran. The study therefore compares an iterating machine pipeline against one-shot human ideas. The asymmetry is the finding: iteration at machine speed is the advantage automation brings to research, and the study’s framing states what it measured.

The automated research loop inside the frame of decisions that stayed human: defining the failures, setting the baseline, judging the results

The human baseline was itself a purchased service: Surge AI recruited the 28 researchers, structured their submissions, and ran quality control and expert review. In our judgment, the expert baseline is now a purchasable input to alignment research, as much a part of its infrastructure as the training cluster.

What happens to the human gate?

Our September 9 post argued the gate was deciding what counts as correct. This study automates more than that framing assumed: proposing fixes and testing them, the middle of the research loop, ran without people. Three decisions stayed with people:

  • Defining failure. People chose the ten alignment failures and built the benchmarks that measure them.
  • Setting the baseline. People produced the expert proposals the agents had to beat.
  • Judging better. People decided that the benchmark gains count, and audited the results.

Trusted environments already enforce the same division of labor, per our ground-truth research: models draft and iterate, humans author ground truth and adjudicate. The craft implication is that alignment failures are becoming an environment category. A deception suite or a reward-hacking suite with calibrated grading is now training and evaluation infrastructure for automated researchers, and it will need the same published calibration figures (how often the grader agrees with expert judgment) as any benchmark whose judge gets recalibrated.

What this means

The human skill that stays scarce keeps changing: first doing the work, then grading it, now specifying what failure means and producing a baseline. Vendors sell each of those in turn, and this study is the first public example of a lab buying an expert baseline for alignment research, the way labs already buy graded training tasks.

FAQ

What are automated alignment researchers?

Automated alignment researchers are agent systems that run the safety-research loop end to end, from reading prior work to training and evaluating a fix, without a person choosing each step.

Did the agents really beat human experts?

Yes: on the seven alignment failures where human proposals were collected, Anthropic’s automated researchers beat a baseline of 28 experienced safety researchers. The study notes the humans had one attempt each in an eight-hour window, so the result compares an iterating system to one-shot proposals.

Does automated alignment research remove humans from deciding what counts as correct?

No. In Anthropic’s study the automated part is proposing and testing fixes. Humans still chose which failures matter, built the benchmarks, set the expert baseline, and judged improvement, the same roles human experts hold in trusted training environments.