Skip to main content
Ryan Orban

Ryan Orban

Subject
2 entries

Pre Training

Bookmarks

  1. Pre-Trained Models: Past, Present and Future

    Comprehensive survey of large-scale pre-trained models (PTMs) tracing the evolution from BERT and GPT through four research frontiers: architecture, contextual use, efficiency, and interpretability. Required reading for understanding how self-supervised pre-training became the unified backbone of modern AI.

  2. Unifying Language Learning Paradigms (UL2)

    Google Research's UL2 paper proposes Mixture-of-Denoisers (MoD), a unified pre-training objective that combines span corruption, prefix LM, and causal LM into a single framework — and separates architecture choices from pre-training objectives, which were previously conflated. The insight that objective and architecture are orthogonal opened the door to mixing paradigms that were previously treated as distinct camps.

All bookmarks