Subject
1 entry
Mixture of Experts
Bookmarks
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Switch Transformers (Fedus, Zoph, Shazeer 2021) scales language models to 1.6 trillion parameters using a simplified sparse Mixture of Experts architecture that routes each token to exactly one expert. It's the paper that made sparse MoE practical at scale and laid the architecture foundation for models like Mixtral and GPT-4.
