Subject
7 entries
Few Shot Learning
Bookmarks
Reframing Instructional Prompts to GPTk's Language
ACL 2022 Findings paper showing that manually reframing instructional prompts — decomposing tasks, itemizing steps, adding positive examples — yields 6–12% performance gains on GPT-2 and GPT-3. The key insight is that models respond better to concrete, step-by-step instructions than to long abstract descriptions.
AlexaTM 20B: Amazon's Few-Shot Language Model
AlexaTM 20B is Amazon's 20B parameter seq2seq language model that outperforms PaLM 540B on one-shot summarization and sets state-of-the-art on multilingual translation — while training on one-fifth of GPT-3's carbon footprint. A strong argument for encoder-decoder architecture in few-shot settings.
Differentiable Prompt Makes Pre-trained Language Models Better Few-Shot Learners
DifferentiablePrompt (DPT) replaces discrete token prompts with optimized continuous embeddings, enabling gradient-based prompt tuning for few-shot learning. Published at ICLR 2022, it established that soft prompts can match full fine-tuning with far fewer parameters.
DART: Differentiable Prompt Makes Pre-Trained Language Models Better Few-Shot Learners
DART (Differentiable pRompT) trains prompt templates end-to-end via backpropagation, treating prompts as learnable continuous vectors rather than fixed text. It makes small pre-trained language models competitive few-shot learners without scaling to GPT-3 sizes — an important stepping stone between hand-crafted prompts and full fine-tuning.
Flamingo: A Visual Language Model for Few-Shot Learning
DeepMind's Flamingo (2022) bridges a frozen vision encoder and a frozen large language model with cross-attention layers, enabling powerful few-shot vision-language capabilities without retraining either component. It set new few-shot records on image captioning and VQA benchmarks by treating visual inputs as just another type of context for an LM.
Neural Instrument Cloning from Very Few Samples
Research on cloning musical instrument sounds using neural audio synthesis from very few samples — few-shot timbre transfer. Relevant to AI music generation tools and the question of how much training data audio models need.
