Subject
2 entries
Pre Training
Bookmarks
Pre-Trained Models: Past, Present and Future
Comprehensive survey of large-scale pre-trained models (PTMs) tracing the evolution from BERT and GPT through four research frontiers: architecture, contextual use, efficiency, and interpretability. Required reading for understanding how self-supervised pre-training became the unified backbone of modern AI.
Unifying Language Learning Paradigms (UL2)
Google Research's UL2 paper proposes Mixture-of-Denoisers (MoD), a unified pre-training objective that combines span corruption, prefix LM, and causal LM into a single framework — and separates architecture choices from pre-training objectives, which were previously conflated. The insight that objective and architecture are orthogonal opened the door to mixing paradigms that were previously treated as distinct camps.
