Subject
4 entries
Multilingual
Bookmarks
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
The BLOOM paper introducing a 176-billion parameter open-access multilingual language model trained by the BigScience collaborative on 46 natural languages and 13 programming languages. It demonstrated that a community-organized research effort could produce a frontier-scale LLM without proprietary infrastructure.
AlexaTM 20B: Amazon's Few-Shot Language Model
AlexaTM 20B is Amazon's 20B parameter seq2seq language model that outperforms PaLM 540B on one-shot summarization and sets state-of-the-art on multilingual translation — while training on one-fifth of GPT-3's carbon footprint. A strong argument for encoder-decoder architecture in few-shot settings.
What Language Model to Train if You Have One Million GPU Hours?
An ablation study by the BigScience group comparing architectural choices and training setups for large multilingual language models targeting 100B+ parameters within a fixed 1M A100 GPU-hour budget. It shows that careful architecture and training setup decisions at the 1.3B scale transfer predictably to larger models, making principled design tractable even at extreme scale.
BLOOM: Open Multilingual Large Language Model
BLOOM is the first open, multilingual large language model trained transparently by a global coalition of AI researchers — 176B parameters, 46 languages, trained on the Jean Zay supercomputer in France. A direct counterpoint to GPT-3's closed access.
