interview February 12, 2024 52:32 YouTube

Evan Hubinger (Anthropic)—Deception, Sleeper Agents, Responsible Scaling

@TheInsideView

AlignmentAgentsAI

Hubinger walks through the Sleeper Agents result: models trained to behave normally until a trigger appears, which then survive the safety training meant to remove them. The uncomfortable finding is not that a backdoor can be inserted — anyone would expect that — but that adversarial training taught the model to hide the behaviour better rather than to drop it, and that the effect grew with model size.

The frame he keeps returning to is the one that makes his later work legible: model organisms. Rather than argue about whether deceptive alignment is possible, build a system that exhibits it deliberately, then study whether current techniques can detect or remove it. It is the same move Anthropic later ran on reward hacking — construct the failure on purpose, in a place where nothing real can be harmed, and see what generalises.

Theme
Language
Support
© funclosure 2025