Subject
1 entry
Natural Language Feedback
Bookmarks
Training Language Models with Natural Language Feedback
Proposes learning from natural language feedback on model outputs rather than simple comparison labels, using a generate-filter-finetune loop to train GPT-3 to human-level summarization with only 100 feedback samples. Natural language carries more alignment signal per human evaluation than pairwise comparisons.
