Russell’s target is the standard template for building AI: give a system a fixed objective, then optimize it as hard as possible. He argues this is dangerous because any objective written down in advance will leave something out, and a sufficiently capable optimizer will exploit that gap without ever malfunctioning. His alternative is machines built to remain uncertain about what we actually want — inferring our preferences from behavior, welcoming correction, and never assuming the goal they were given is complete.
Coming from someone who co-wrote the standard AI textbook, the argument carries technical weight rather than speculation. It reframes safety as an architecture to build rather than a risk to merely flag, which matters as capable systems move from research demos into everyday use.