Skip to main content
Ryan Orban

Ryan Orban

Subject
29 entries

Transformers

Bookmarks

  1. PaLM: Scaling Language Modeling with Pathways

    PaLM is Google's 540B parameter language model trained across 6144 TPU v4 chips using the Pathways distributed training system. It achieved breakthrough performance on multi-step reasoning and BIG-Bench, and documented discontinuous capability gains at scale — capabilities that emerged suddenly with more compute.

  2. How Attention Got So Efficient: GQA, MLA, DSA explained

    A YouTube explainer on how attention mechanisms evolved from full multi-head attention to the efficient variants powering modern LLMs — covering GQA (Grouped Query Attention), MLA (Multi-head Latent Attention from DeepSeek), and DSA (Dynamic Sparse Attention). Useful for understanding why modern inference is faster and cheaper than it was two years ago.

  3. Flash Attention: Derived and Coded from First Principles with Triton

    A from-scratch implementation of Flash Attention using Triton (Python GPU programming), deriving the algorithm from first principles before coding it. Valuable for anyone wanting to understand how Flash Attention achieves its memory efficiency — learning by building rather than just using the library.

  4. A Visual Guide to Vision Transformers

    A visual guide to Vision Transformers (ViT) — explains how the transformer architecture is adapted for images, covering patch embeddings, position encodings, and attention in visual domains with diagrams. Good complement to the original ViT paper for building intuition.

  5. Timeline of Transformer Models and Large Language Models

    A visual timeline of transformer models and large language models from 2017 through the present, mapping the lineage of architectures from the original Attention Is All You Need paper through GPT-4 and beyond. A useful reference for understanding model genealogy.

  6. From Deep to Long Learning?

    Stanford Hazy Research argues the next frontier is moving from deep networks to networks that can process very long sequences — motivating state space models like Mamba as a shift away from transformer attention's O(n²) complexity. A prescient 2023 post about where sequence modeling was headed.

  7. Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation

    NeurIPS 2022 paper investigating how to train very deep Transformers by removing shortcut connections (residual paths), which typically cause rank collapse in attention layers. The work has implications for understanding how information propagates through depth in Transformer architectures.

  8. Human Motion Diffusion Model (MDM)

    Tevet et al. (Tel Aviv University, arXiv:2209.14916, 2022) apply diffusion models to human motion generation, producing MDM — a transformer-based denoiser that generates realistic motion sequences from text descriptions or action labels. It matters because it extends the generative power of diffusion to a structured temporal domain, enabling controllable motion editing that prior methods couldn't match.

  9. What Learning Algorithm Is In-Context Learning? Investigations with Linear Models

    ICLR 2023 paper showing that transformers trained on in-context learning tasks implicitly implement gradient descent and ridge regression on linear problems, with layers encoding weight vectors and moment matrices. Foundational theoretical work explaining ICL as implicit algorithm execution rather than pure pattern matching.

  10. In-Context Reinforcement Learning with Algorithm Distillation

    Algorithm Distillation trains a causal transformer on sequences of RL learning histories so the model can improve its policy entirely in-context without gradient updates. A key step toward meta-learning agents that get better at RL through experience rather than parameter updates.

  11. Wide Attention Is The Way Forward for Transformers

  12. Attention? Attention! (Lilian Weng)

    Lilian Weng's canonical blog post on attention mechanisms — written in 2018, it covers sequence-to-sequence attention, self-attention, and multi-head attention. Still the clearest single-page reference for understanding how attention works before reading the Transformer paper.

  13. How AI Transformers Mimic Parts of the Brain (Quanta Magazine)

    Quanta Magazine's September 2022 piece on emerging research showing that transformer attention patterns converge with how the brain processes language and vision — a surprising empirical finding with implications for both AI interpretability and computational neuroscience.

  14. Memorizing Transformers

    ICLR 2022 spotlight proposing memory-augmented transformers that use approximate k-NN lookup into stored (key, value) pairs at inference time, without weight updates. Scales to 262K token memory with consistent perplexity improvements — an early approach to giving language models dynamic, test-time-updateable knowledge stores.

  15. Explaining Transformer Model Predictions

    A practical comparison of SHAP, Transformers Interpret, and Ferret for explaining Hugging Face transformer predictions. Key takeaway: different methods give different results for the same prediction — all require careful interpretation.

  16. SLED: Efficient Long-Text Understanding with Short-Text Models

    SLED (Sliding-Encoder and Decoder) lets you apply short-context pretrained models to arbitrarily long documents by chunking input with overlap and fusing representations in the decoder. Competitive with specialized long-context models without the expensive custom pretraining.

  17. What Language Model to Train if You Have One Million GPU Hours?

    An ablation study by the BigScience group comparing architectural choices and training setups for large multilingual language models targeting 100B+ parameters within a fixed 1M A100 GPU-hour budget. It shows that careful architecture and training setup decisions at the 1.3B scale transfer predictably to larger models, making principled design tractable even at extreme scale.

  18. Jay Alammar — Visualizing Machine Learning

    Jay Alammar's blog is the go-to resource for visually understanding modern ML architectures — transformers, BERT, GPT, and more through hand-crafted diagrams. His illustrated explainers have become canonical references for practitioners who want intuition before equations.

  19. How fast can we perform a forward pass?

    A deep technical investigation into the theoretical and practical speed limits of a single transformer forward pass, accounting for compute, memory bandwidth, and hardware constraints. The answer matters for inference costs, latency SLAs, and understanding where efficiency gains are actually possible.

  20. Hopfield Networks is All You Need

    Ramsauer et al. introduce a modern continuous-state Hopfield network with exponential storage capacity and prove that its update rule is mathematically equivalent to transformer self-attention. This is the foundational paper connecting classical associative memory to the attention mechanism, explaining why transformers work through an energy-function lens.

  21. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

    Switch Transformers (Fedus, Zoph, Shazeer 2021) scales language models to 1.6 trillion parameters using a simplified sparse Mixture of Experts architecture that routes each token to exactly one expert. It's the paper that made sparse MoE practical at scale and laid the architecture foundation for models like Mixtral and GPT-4.

  22. Decision Transformer: Reinforcement Learning via Sequence Modeling

    Decision Transformer recasts offline reinforcement learning as a conditional sequence modeling problem, using a causally masked Transformer to generate actions conditioned on desired return, past states, and actions. It matches or exceeds model-free offline RL baselines on Atari, OpenAI Gym, and Key-to-Door without any value function or policy gradient computation.

  23. Transformer Inference Arithmetic

    Carol Chen (kipply) works through the arithmetic of transformer inference — compute vs. memory bandwidth, KV cache sizing, and how batch size shapes throughput/latency tradeoffs. Essential reference for anyone reasoning about LLM serving costs.

  24. Transformers for Software Engineers

    Nelson Elhage's explainer on transformers pitched at software engineers — treats the architecture as a data structure rather than mysterious ML magic. Grounding for anyone who codes but hasn't internalized what attention actually does.

  25. Will Transformers Take Over Artificial Intelligence?

    Quanta Magazine's 2022 look at whether transformer architectures will dominate all of AI — following their success in NLP and early incursion into image classification. A useful time-capsule of the moment when the transformer paradigm started feeling inevitable.

  26. Task-Specific Knowledge Distillation for BERT

    A tutorial on task-specific knowledge distillation for BERT using Hugging Face Transformers and Amazon SageMaker — compressing a 109M-parameter BERT-base into a 4M-parameter student with 90%+ performance retention. Demonstrates that you don't need a massive model in production if you can distill task knowledge from one.

  27. Financial Text Summarization with Hugging Face and Keras

    A tutorial on fine-tuning distilled BART for financial news summarization using Hugging Face Transformers with Keras and Amazon SageMaker — generating headline-length summaries from longer articles. A practical demonstration of seq2seq fine-tuning on domain-specific data.

  28. The Illustrated Transformer

    Jay Alammar's illustrated walkthrough of the Transformer architecture — the most widely cited visual explainer of attention, encoder-decoder structure, and multi-head attention. A must-read before diving into any BERT, GPT, or T5 paper.

  29. Generative Pretraining from Pixels (iGPT)

    The iGPT paper from OpenAI (ICML 2020) showing that a GPT-2-scale transformer trained to autoregressively predict pixels learns strong image representations — 96.3% accuracy on CIFAR-10 with a linear probe. It's a direct transposition of NLP pretraining ideas to the image domain, predating CLIP and DALL-E.

All bookmarks