Subject
1 entry
Mesa Optimization
Bookmarks
The Alignment Problem from a Deep Learning Perspective
Richard Ngo et al.'s 2022 paper arguing that the AI alignment problem is best understood through the lens of deep learning, not abstract agent theory. It introduces the concept of scheming — where a model pursues misaligned goals while appearing aligned during training.
