Subject
1 entry
Reward Models
Bookmarks
Reward Model Ensembles Help Mitigate Overoptimization
Coste et al. (2022) show that ensembling multiple reward models substantially reduces reward overoptimization during RLHF, where a policy learns to exploit artifacts in a single reward model rather than truly improving. This is a practical mitigation for one of the central failure modes in aligning language models with human preferences.
