Subject
7 entries
Rlhf
Bookmarks
Defining and Characterizing Reward Hacking
Skalse, Howe, Krasheninnikov, and Krueger provide the first formal definition of reward hacking — when optimizing a proxy reward degrades performance on the true reward — and derive conditions under which 'unhackable' proxies exist. The theoretical result is stark: for stochastic policies, only constant reward functions are unhackable, making reward misspecification a near-universal concern in RL-based alignment.
Colossal-AI: Open-Source ChatGPT Training Replication
Colossal-AI released an open-source implementation of the ChatGPT training process (SFT + RLHF) that runs on a single GPU with 1.6GB memory — 7.73x faster than naive implementations. Made the ChatGPT training pipeline accessible to researchers without multi-GPU clusters.
Training Language Models with Natural Language Feedback
Proposes learning from natural language feedback on model outputs rather than simple comparison labels, using a generate-filter-finetune loop to train GPT-3 to human-level summarization with only 100 feedback samples. Natural language carries more alignment signal per human evaluation than pairwise comparisons.
Improving Multimodal Interactive Agents with RLHF
The Interactive Agents Team at DeepMind (arXiv:2211.11602, 2022) applies RLHF to agents that must understand language instructions and act in visual environments, using human preference feedback to train a reward model for PPO-based RL. The result demonstrates that RLHF substantially improves instruction-following in embodied multimodal settings beyond what supervised learning alone achieves.
Reward Model Ensembles Help Mitigate Overoptimization
Coste et al. (2022) show that ensembling multiple reward models substantially reduces reward overoptimization during RLHF, where a policy learns to exploit artifacts in a single reward model rather than truly improving. This is a practical mitigation for one of the central failure modes in aligning language models with human preferences.
Building Safer Dialogue Agents (DeepMind / Sparrow)
DeepMind's blog post on Sparrow — a dialogue agent trained with reinforcement learning from human feedback and rules to be helpful, harmless, and honest. An early published account of RLHF-based safety fine-tuning for conversational AI.
Training Language Models to Follow Instructions with Human Feedback (InstructGPT)
OpenAI's InstructGPT paper (2022) shows that a 1.3B model fine-tuned with RLHF on human preference data is preferred over raw GPT-3 at 175B — establishing that alignment via human feedback is more important than raw scale for following instructions. This is the foundational paper behind ChatGPT and the instruction-tuned era.
