Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Natural Language Feedback

Bookmarks

  1. Training Language Models with Natural Language Feedback

    Proposes learning from natural language feedback on model outputs rather than simple comparison labels, using a generate-filter-finetune loop to train GPT-3 to human-level summarization with only 100 feedback samples. Natural language carries more alignment signal per human evaluation than pairwise comparisons.

All bookmarks