Ideas / From Intelligence to Agency
How Not to Destroy the World with AI
Stuart Russell
Stuart Russell asks what capable systems do with objectives that only approximate what humans meant.
Russell moves from capability to control. The question is not merely whether AI can perform a task. It is whether we should give a capable system an objective and permission to pursue it.
He argues for systems that remain uncertain about human preferences instead of treating a rigid goal as revealed truth.