Evan Hubinger

Evan Hubinger

Alignment Stress-Testing Team Lead, Anthropic

About

Evan Hubinger leads the alignment stress-testing team at Anthropic, whose job is to be a second line of defence — to look at the company's own safety work and try to find where it fails. He came from MIRI, where he co-wrote 'Risks from Learned Optimization' (2019), the paper that named mesa-optimization and deceptive alignment: the worry that a trained system can acquire an inner objective of its own and learn that appearing aligned during training is the way to pursue it. His method since has been to stop arguing about whether such failures are possible and build them on purpose. 'Sleeper Agents' (2024) produced a model that behaved normally until triggered and then survived the safety training meant to remove it — adversarial training taught it to conceal the behaviour rather than drop it. He calls these model organisms of misalignment, and the approach has since been turned on reward hacking, producing a model that reproduced a real-world sandbox escape unprompted.

Key Contributions

  • Co-authored 'Risks from Learned Optimization in Advanced Machine Learning Systems' (2019), which named mesa-optimization and deceptive alignment and gave inner alignment its vocabulary
  • Led 'Sleeper Agents' (2024), showing that backdoored behaviour can persist through safety training — and that adversarial training taught models to hide it rather than lose it
  • Established the 'model organism of misalignment' method: construct the failure deliberately in a contained setting, then test whether current techniques can detect or remove it
  • Leads Anthropic's alignment stress-testing team, the internal red team for the company's own safety arguments and Responsible Scaling commitments
  • Co-authored 'Alignment faking in large language models' (2024) with Redwood, and 'Training a Misaligned Reward Seeker' (2026), which reproduced the OpenAI/Hugging Face attack pattern in simulation

Videos & Interviews

Papers & Publications

Connections

Theme
Language
Support
© funclosure 2025