Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Overoptimization

Bookmarks

  1. Reward Model Ensembles Help Mitigate Overoptimization

    Coste et al. (2022) show that ensembling multiple reward models substantially reduces reward overoptimization during RLHF, where a policy learns to exploit artifacts in a single reward model rather than truly improving. This is a practical mitigation for one of the central failure modes in aligning language models with human preferences.

All bookmarks