Skip to main content
Ryan Orban

Ryan Orban

Subject
2 entries

Large Models

Bookmarks

  1. How Hugging Face Accelerate Runs Very Large Models

    Hugging Face's technical guide to running very large models using Accelerate — covers device_map, model parallelism across GPUs, CPU offloading, and the mechanics of loading models that don't fit in a single GPU's VRAM. Essential reading for anyone self-hosting large LLMs.

  2. Alpa: Automated Distributed Training for Large Models

    Alpa is a system for automatically parallelizing large neural network training across distributed hardware — finding optimal parallelism strategies without manual configuration. From a Berkeley/CMU research collaboration, it targets the challenge of scaling models beyond single-GPU memory.

All bookmarks