Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Embodied Agents

Bookmarks

  1. Improving Multimodal Interactive Agents with RLHF

    The Interactive Agents Team at DeepMind (arXiv:2211.11602, 2022) applies RLHF to agents that must understand language instructions and act in visual environments, using human preference feedback to train a reward model for PPO-based RL. The result demonstrates that RLHF substantially improves instruction-following in embodied multimodal settings beyond what supervised learning alone achieves.

All bookmarks