Skip to main content
Ryan Orban

Ryan Orban

Subject
3 entries

Instruction Tuning

Bookmarks

  1. Scaling Instruction-Finetuned Language Models

    The FLAN-T5 and Flan-PaLM paper (2022) showing that instruction finetuning across 1.8K tasks improves performance on zero-shot, few-shot, and chain-of-thought prompting across multiple model families. Established instruction finetuning as a general-purpose method and released Flan-T5 checkpoints that shaped the open-source LLM ecosystem.

  2. Large Language Models Are Human-Level Prompt Engineers

    Large Language Models Are Human-Level Prompt Engineers (APE) introduces an automated method for generating and selecting optimal prompts using LLMs themselves, matching or beating human-crafted instructions on 19 of 24 NLP tasks. It reframes prompt engineering as a program search problem, making manual iteration unnecessary.

  3. Multitask Prompted Training Enables Zero-Shot Task Generalization

    T0 shows that training a language model on 2000+ diverse human-authored prompts across 170+ NLP tasks dramatically improves zero-shot generalization to unseen tasks. An 11B T0 model outperformed 175B GPT-3 zero-shot — proving prompt diversity during training matters more than raw scale for generalization.

All bookmarks