Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Reward Hacking

Bookmarks

  1. Defining and Characterizing Reward Hacking

    Skalse, Howe, Krasheninnikov, and Krueger provide the first formal definition of reward hacking — when optimizing a proxy reward degrades performance on the true reward — and derive conditions under which 'unhackable' proxies exist. The theoretical result is stark: for stochastic policies, only constant reward functions are unhackable, making reward misspecification a near-universal concern in RL-based alignment.

All bookmarks