Skip to main content
Ryan Orban

Ryan Orban

Subject
3 entries

Zero Shot

Bookmarks

  1. ASIF: Coupled Data Turns Unimodal Models to Multimodal Without Training

    Norelli, Fumero, Maiorca, Rodolà et al. (Sapienza University, arXiv:2210.01738, 2022) show that any two unimodal models can be composed into a zero-shot multimodal system by finding approximate shared nearest neighbors across their embedding spaces, given only a small set of coupled pairs. The result challenges the assumption that multimodal capability requires joint training.

  2. What Language Model Architecture and Pretraining Objective Work Best for Zero-Shot Generalization?

    Wang, Roberts, Scao et al. (BigScience Architecture Group) conduct a large-scale comparison of model architectures (causal decoder, non-causal decoder, encoder-decoder) and pretraining objectives (autoregressive, masked LM) for zero-shot generalization. The key finding: causal decoders + autoregressive LM win at zero-shot; non-causal decoders + masked LM + multitask fine-tuning win overall.

  3. Multitask Prompted Training Enables Zero-Shot Task Generalization

    T0 shows that training a language model on 2000+ diverse human-authored prompts across 170+ NLP tasks dramatically improves zero-shot generalization to unseen tasks. An 11B T0 model outperformed 175B GPT-3 zero-shot — proving prompt diversity during training matters more than raw scale for generalization.

All bookmarks