Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Mixture of Experts

Bookmarks

  1. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

    Switch Transformers (Fedus, Zoph, Shazeer 2021) scales language models to 1.6 trillion parameters using a simplified sparse Mixture of Experts architecture that routes each token to exactly one expert. It's the paper that made sparse MoE practical at scale and laid the architecture foundation for models like Mixtral and GPT-4.

All bookmarks