Subject
14 entries
Scaling
Bookmarks
PaLM: Scaling Language Modeling with Pathways
PaLM is Google's 540B parameter language model trained across 6144 TPU v4 chips using the Pathways distributed training system. It achieved breakthrough performance on multi-step reasoning and BIG-Bench, and documented discontinuous capability gains at scale — capabilities that emerged suddenly with more compute.
RAG at Planet Scale
Arcus describes their multi-tiered RAG approach for handling massive external data corpora at planet scale — one of the largest RAG deployments of 2023. The key innovation is tiered retrieval that narrows the candidate pool progressively rather than searching the full index directly.
Scaling Instruction-Finetuned Language Models
The FLAN-T5 and Flan-PaLM paper (2022) showing that instruction finetuning across 1.8K tasks improves performance on zero-shot, few-shot, and chain-of-thought prompting across multiple model families. Established instruction finetuning as a general-purpose method and released Flan-T5 checkpoints that shaped the open-source LLM ecosystem.
The Cost of Training NLP Models: A Concise Overview
AI21 Labs' 2020 overview of the financial and compute cost of training large NLP models, tracing the exponential growth in training expenditure from BERT through GPT-3. An early quantitative lens on the economics of scale in language model development.
What Language Model to Train if You Have One Million GPU Hours?
An ablation study by the BigScience group comparing architectural choices and training setups for large multilingual language models targeting 100B+ parameters within a fixed 1M A100 GPU-hour budget. It shows that careful architecture and training setup decisions at the 1.3B scale transfer predictably to larger models, making principled design tractable even at extreme scale.
Notes on Ethereum L2 Solutions
Jin's 2021 notes on Ethereum L2 scaling solutions — covering optimistic rollups, ZK-rollups, state channels, and plasma. A solid survey of the L2 design space written before rollups became the dominant paradigm.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Switch Transformers (Fedus, Zoph, Shazeer 2021) scales language models to 1.6 trillion parameters using a simplified sparse Mixture of Experts architecture that routes each token to exactly one expert. It's the paper that made sparse MoE practical at scale and laid the architecture foundation for models like Mixtral and GPT-4.
The Bitter Lesson
Rich Sutton's 2019 essay arguing that the dominant lesson from 70 years of AI research is that general methods leveraging computation always win over human-engineered knowledge — a humbling argument against clever domain-specific tricks. One of the most cited and debated essays in ML.
Exploring Zero Knowledge: zkSync and the zkEVM
An explainer on zkSync and the zkEVM — how zero-knowledge proofs enable EVM-compatible rollups that give instant finality without trusting the operator. Written in early 2022 as zkEVM was the most technically ambitious open problem in Ethereum scaling.
Scaling SQL with Redis
David Cramer's post on using Redis to scale SQL databases — covering caching patterns, read replica offloading, and where Redis fits in a stack that can't abandon SQL entirely. Practical patterns from the Disqus/Sentry engineering blog at production scale.
High Performance at Massive Scale: Lessons Learned at Facebook
Summary of a 2009 Facebook engineering talk on high-performance systems at massive scale — covering their memcached deployment, MySQL sharding, and the operational realities of running at hundreds of millions of users. An early public window into big-company distributed systems practice.
A Rough Guide to Keeping Your Website Up Through High Traffic
Rainforest's 2012 guide to keeping a web application running through traffic spikes — covering CDN, caching, database connection pooling, and graceful degradation. Early practical devops writing from the era before managed auto-scaling became trivial.
What Now: Scaling the Sales Organization
PandoDaily on the mechanics of scaling a startup sales org — the transition from founder-led sales to a managed sales team. A 2012 take on what was then an emerging playbook for B2B SaaS companies.
