Skip to main content
Ryan Orban

Ryan Orban

Subject
77 entries

Deep Learning

Bookmarks

  1. An Introduction to Variational Autoencoders

    The canonical tutorial on Variational Autoencoders by Kingma and Welling — the original VAE inventors. Covers the ELBO, reparameterization trick, and extensions to deeper generative models. Essential background for anyone working with latent variable models or modern diffusion/flow models.

  2. Why Deep Learning Works Unreasonably Well

    Part 3 of a 'How Models Learn' series on why deep learning works unreasonably well — addressing the apparent paradox that overparameterized models generalize when classical statistics says they shouldn't. Covers implicit regularization, loss landscape geometry, and the lottery ticket hypothesis.

  3. CubeCL: Multi-Platform GPU Compute for Rust

    CubeCL is a multi-platform GPU compute language extension for Rust — write one kernel that targets CUDA, ROCm, Vulkan/WebGPU, and Metal. The foundation for the Burn deep learning framework's backend abstraction.

  4. Neural Networks — 3Blue1Brown

    3Blue1Brown's neural networks playlist is the canonical visual introduction to how neural networks and backpropagation work — four episodes, starting from scratch and building to the chain rule. The most-watched math explainer for deep learning fundamentals.

  5. fast.ai

    fast.ai is Jeremy Howard and Rachel Thomas's free deep learning course and library, famous for teaching neural networks top-down — use them effectively first, understand the math later. Widely credited with democratizing deep learning education.

  6. Deep Learning Systems (CMU 10-414/714)

    10-414/714: Deep Learning Systems at CMU — a publicly available course on building the components of a deep learning framework from scratch, including automatic differentiation, optimization, and hardware acceleration. One of the best resources for understanding how frameworks like PyTorch actually work.

  7. A Visual Guide to Vision Transformers

    A visual guide to Vision Transformers (ViT) — explains how the transformer architecture is adapted for images, covering patch embeddings, position encodings, and attention in visual domains with diagrams. Good complement to the original ViT paper for building intuition.

  8. Building a Deep Learning Rig (Part 2)

    Part 2 of building a home deep learning rig: upgrading to a Threadripper 1920X to support three RTX 3090s with full PCIe bandwidth. Total cost €2,379. Key lesson: GPU peer-to-peer DDP failures were fixed by downgrading the NVIDIA driver, not hardware changes.

  9. Neural Networks from Scratch in Python

    Neural Networks from Scratch (NNFS) is a book by Harrison Kinsley and Daniel Kukiela that builds neural networks in pure Python with no frameworks — the go-to resource for understanding what's actually happening inside backpropagation and gradient descent.

  10. Google Research Deep Learning Tuning Playbook

    Google Research's Deep Learning Tuning Playbook — a systematic guide to maximizing model performance through hyperparameter optimization. Written by Braxton Osting and team, it covers the science and art of tuning learning rates, batch sizes, regularization, and the full training configuration.

  11. TorToiSe TTS: Architectural Design Document

    The architectural design document for TorToiSe TTS — James Betker's highly capable open-source voice cloning and text-to-speech system. Explains the multi-model pipeline combining autoregressive and diffusion components that made it state-of-the-art in 2022.

  12. Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation

    NeurIPS 2022 paper investigating how to train very deep Transformers by removing shortcut connections (residual paths), which typically cause rank collapse in attention layers. The work has implications for understanding how information propagates through depth in Transformer architectures.

  13. EVA: Exploring the Limits of Masked Visual Representation Learning at Scale

    EVA is a 1-billion-parameter vision foundation model from BAAI that achieves state-of-the-art on image classification, detection, and segmentation by pretraining a ViT to reconstruct masked CLIP features. Initializing CLIP's vision tower from EVA dramatically stabilizes training — an important practical finding for building large multimodal systems.

  14. fast.ai: From Deep Learning Foundations to Stable Diffusion

    fast.ai's Part 2 2022 course preview — the first two lessons of their deep learning foundations to Stable Diffusion curriculum, taught bottom-up from first principles. Jeremy Howard teaching diffusion models the way fast.ai teaches everything: by building it yourself.

  15. AI Winter Is Well On Its Way

    Filip Piekniewski's 2018 blog post predicting an AI winter due to deep learning's fundamental limitations — a contrarian view written before the GPT era proved the scaling hypothesis. Useful as a document of expert skepticism that turned out to be largely wrong.

  16. MIT 6.S898: Deep Learning (Fall 2022)

    MIT's 6.S898 Deep Learning course taught by Phillip Isola, covering the full modern deep learning stack from fundamentals through generative models and transformers. One of the cleaner academic deep learning curricula, with public materials.

  17. How Diffusion Models Work: The Math from Scratch

    AI Summer's mathematical walkthrough of how diffusion models work from scratch — covering the forward noising process, reverse denoising, DDPM training objective, and score matching. The most math-forward accessible introduction to the field.

  18. ML YouTube Courses

    A curated catalog of machine learning courses available on YouTube, maintained by DAIR.AI. Covers ML fundamentals, deep learning, NLP, computer vision, and specialized topics — a single index for free university-grade ML education.

  19. Attention? Attention! (Lilian Weng)

    Lilian Weng's canonical blog post on attention mechanisms — written in 2018, it covers sequence-to-sequence attention, self-attention, and multi-head attention. Still the clearest single-page reference for understanding how attention works before reading the Transformer paper.

  20. A Short Chronology of Deep Learning for Tabular Data

    Sebastian Raschka's chronological survey of deep learning approaches for tabular data — the domain where gradient boosted trees still dominate. A clear-eyed accounting of why deep learning hasn't won on structured data despite winning everywhere else.

  21. Deep Anomaly Detection with Self-Supervised Learning and Adversarial Training

    This paper combines self-supervised learning and adversarial training to improve deep anomaly detection, leveraging unlabeled normal data to learn representations that are robust to perturbations and more sensitive to out-of-distribution inputs. It matters because labeled anomaly data is rare in practice, making self-supervised approaches essential for real-world deployment.

  22. Topaz Labs: AI Image Enhancement

    Topaz Labs makes deep learning-powered photo and video enhancement software — noise reduction, sharpening, and upscaling. One of the earliest commercial successes applying neural networks to real-world creative workflows, predating the generative AI wave.

  23. The Alignment Problem from a Deep Learning Perspective

    The Alignment Problem from a Deep Learning Perspective (Ngo, Chan, Mindermann, 2022) — frames AI alignment as a problem of specification, robustness, and assurance in deep learning systems. An influential restatement of alignment concerns in terms of modern ML rather than AGI thought experiments.

  24. Probing Vision Transformers

    A research repository probing the internal representations and attention mechanisms of Vision Transformers (ViT, DeiT, DINO). Notable finding: self-supervised DINO produces more salient attention maps than supervised models, suggesting better spatial semantics from self-supervision.

  25. Awesome Colab Notebooks

    Curated collection of Google Colab notebooks for ML experiments — covering Stable Diffusion, GANs, NLP models, and more. A snapshot of what was freely runnable in browser-based GPU compute in 2022, before Hugging Face Spaces simplified deployment further.

  26. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

    ZeRO (Zero Redundancy Optimizer) eliminates memory redundancy in distributed training by partitioning optimizer states, gradients, and parameters across data-parallel processes rather than replicating them. It enables training models 8x larger than prior methods on the same hardware and forms the foundation of Microsoft's DeepSpeed library.

  27. The Annotated Transformer

    The Annotated Transformer walks through the 'Attention Is All You Need' paper with working PyTorch code alongside every equation — the canonical resource for understanding transformer architecture from first principles. Published by Harvard NLP.

  28. Colossal-AI: Distributed Deep Learning System

    Colossal-AI is an open-source distributed deep learning framework that makes training very large models more accessible — cutting GPU memory requirements by up to 10x versus standard PyTorch. One of several systems research projects responding to the GPU memory wall problem in 2022.

  29. Privacy-Preserving Machine Learning with Fully Homomorphic Encryption for Deep Neural Networks

    This paper demonstrates running deep neural network inference entirely on encrypted data using fully homomorphic encryption, so the server never sees plaintext inputs or outputs. It makes encrypted ML inference practical by combining FHE with approximation-friendly neural network architectures.

  30. Hopfield Networks is All You Need

    Ramsauer et al. introduce a modern continuous-state Hopfield network with exponential storage capacity and prove that its update rule is mathematically equivalent to transformer self-attention. This is the foundational paper connecting classical associative memory to the attention mechanism, explaining why transformers work through an energy-function lens.

  31. A Hitchhiker's Guide to Distributed Training of Deep Neural Networks

    Chahal, Grover, and Dey survey the algorithms and engineering techniques for distributed deep learning training, covering data parallelism, model parallelism, AllReduce strategies, gradient compression, and mixed precision. A practical reference for scaling training from single GPU to multi-node clusters.

  32. Proof of Learning (PoLe): Blockchain Consensus via Neural Network Training

    Proof of Learning (PoLe) repurposes the computation wasted on Proof-of-Work consensus by directing it toward training neural networks, making blockchain nodes useful ML workers. The cheating-prevention mechanism (Secure Mapping Layer) is the interesting technical contribution — a linear layer baked into the network architecture that makes it costly to fake training progress.

  33. Agatha: Smart Contract for DNN Computation

    Agatha is a system for verifiable DNN computation on Ethereum smart contracts, achieving native-speed off-chain inference with only 3% overhead via a graph-based pinpoint protocol. It bridges AI and blockchain by solving the mismatch between neural network computational graphs and the VM-based execution model that existing verification schemes assume.

  34. Deep Learning with PyTorch

    Manning's 2020 practical guide to deep learning with PyTorch by Stevens, Antiga, and Viehmann, covering tensors through CNNs, RNNs, generative models, and production deployment. The go-to book for practitioners who want to understand PyTorch from first principles rather than copy-paste patterns.

  35. Machine Learning in Finance: From Theory to Practice

    Springer 2020 textbook by Dixon, Halperin, and Bilokon bridging ML theory and quantitative finance practice — covering supervised learning, NLP for financial texts, RL for trading, and deep learning for derivatives pricing. The most rigorous academic treatment of ML applied to finance.

  36. Transformers for Software Engineers

    Nelson Elhage's explainer on transformers pitched at software engineers — treats the architecture as a data structure rather than mysterious ML magic. Grounding for anyone who codes but hasn't internalized what attention actually does.

  37. Word2Vec Explained

    A clear walkthrough of Word2Vec's intuition and mechanics — skip-gram vs. CBOW, negative sampling, and why the king-queen analogy works. Good primer before reading papers on later embedding methods.

  38. Will Transformers Take Over Artificial Intelligence?

    Quanta Magazine's 2022 look at whether transformer architectures will dominate all of AI — following their success in NLP and early incursion into image classification. A useful time-capsule of the moment when the transformer paradigm started feeling inevitable.

  39. Swarm Training

    Shawn Presser's Swarm Training explores distributed ML training across many commodity machines with low-bandwidth interconnects — democratizing large model training beyond clusters with expensive NVLink. Part of the broader open-source effort to train large models outside of big lab infrastructure.

  40. Dive into Deep Learning Compiler

    A free textbook companion to Apache TVM covering deep learning compiler design — how neural networks get transformed into optimized code for CPUs, GPUs, and specialized accelerators. Essential for anyone working on ML infrastructure or hardware-software co-design.

  41. DeZero Book: Build a Deep Learning Framework from Scratch

    DeZero is a book that builds a deep learning framework from scratch in pure Python — teaching automatic differentiation, computational graphs, and the internals of PyTorch/Chainer by implementing them. The most hands-on way to understand how deep learning frameworks actually work.

  42. Full Stack Deep Learning (Spring 2021)

    Full Stack Deep Learning is a free course bridging ML research and production deployment — covering the full pipeline from data management and model training to testing, monitoring, and team structures. The course for researchers who want to ship and engineers who want to understand ML.

  43. The Alignment Problem from a Deep Learning Perspective

    Richard Ngo et al.'s 2022 paper arguing that the AI alignment problem is best understood through the lens of deep learning, not abstract agent theory. It introduces the concept of scheming — where a model pursues misaligned goals while appearing aligned during training.

  44. Speech Driven Talking Head Generation via Attentional Landmarks Based Representation

    This paper introduces an attentional landmark-based representation for generating realistic talking head video from speech audio, using facial landmarks as a compact intermediate representation that bridges audio and visual domains. The approach decouples appearance generation from motion modeling, improving generalization across identities.

  45. Reproducible Deep Learning PhD Course

    Simone Scardapane's PhD course on Reproducible Deep Learning — covering Git, Docker, DVC, experiment tracking, and CI/CD for ML research. Addresses the reproducibility crisis in deep learning with practical tooling.

  46. Self-Supervised Learning: The Dark Matter of Intelligence

    Yann LeCun and Ishan Misra's Facebook AI blog post arguing that self-supervised learning — learning from unlabeled data — is the key to human-level AI, analogous to the dark matter that makes up most of the universe's mass. Published ahead of a wave of self-supervised breakthroughs.

  47. State-of-the-Art Image Generative Models (2021)

    Aran Komatsuzaki's March 2021 survey of state-of-the-art image generative models — covering BigGAN, VQVAE-2, DALL-E, CLIP, and diffusion models just as they were emerging. A historical snapshot of the field one year before diffusion models took over.

  48. LabML Neural Networks: Annotated Implementations

    LabML Neural Networks is a collection of PyTorch implementations of ML papers with line-by-line annotations — making research papers readable by walking through the actual code. Covers transformers, diffusion models, GANs, and RL algorithms side-by-side with the paper math.

  49. CS231n: Convolutional Neural Networks for Visual Recognition

    Stanford CS231n: Convolutional Neural Networks for Visual Recognition — Andrej Karpathy's course that became the de facto entry point into deep learning for computer vision. The lecture notes remain among the best written explanations of CNNs, backprop, and training practice.

  50. The Illustrated Transformer

    Jay Alammar's illustrated walkthrough of the Transformer architecture — the most widely cited visual explainer of attention, encoder-decoder structure, and multi-head attention. A must-read before diving into any BERT, GPT, or T5 paper.

  51. 10 Categories of Deep Recommendation Systems

    James Le's survey of 10 categories of deep learning-based recommendation systems — from MLP and autoencoder approaches through attention-based and graph neural network methods. A useful taxonomy for understanding how the field moved beyond matrix factorization.

  52. Generative Pretraining from Pixels (iGPT)

    The iGPT paper from OpenAI (ICML 2020) showing that a GPT-2-scale transformer trained to autoregressively predict pixels learns strong image representations — 96.3% accuracy on CIFAR-10 with a linear probe. It's a direct transposition of NLP pretraining ideas to the image domain, predating CLIP and DALL-E.

  53. Practical Deep Learning for Coders — fast.ai

    fast.ai's Practical Deep Learning for Coders — Jeremy Howard and Rachel Thomas's free course that inverted the standard pedagogy: start with working image classifiers, then learn the theory underneath. Democratized deep learning at a moment when most education assumed a PhD on-ramp.

  54. Yann LeCun: Making Facebook's AI Predict What Happens in Videos

    New Scientist interview with Yann LeCun on Facebook AI Research's goal to build models that predict what will happen in videos — covering what AI can and can't do in 2015, and LeCun's view on unsupervised learning as the key unsolved problem.

  55. Google Open-Sources Its Artificial Intelligence Engine

    Wired's coverage of Google open-sourcing TensorFlow on November 9, 2015 — the moment that made production-grade deep learning infrastructure freely available to the world. A historical inflection point in the commoditization of AI tooling.

  56. What a Deep Neural Network Thinks About Your Selfie

    Andrej Karpathy trained a VGGNet on 2 million Instagram selfies to learn what makes a selfie 'good' — using likes-per-follower as the quality signal. Beyond the entertainment value, it's a sharp demonstration of how supervised learning can proxy for human aesthetic judgment at scale.

  57. Visualizing Representations: Deep Learning and Human Beings

    Christopher Olah's essay on visualizing what neural networks actually learn — using dimensionality reduction to show how deep networks transform data into progressively more separable representations. One of the most important early pieces on deep learning interpretability.

  58. Zipfian Academy Partners with Skymind to Teach Deep Learning

    Zipfian Academy announces a partnership with Skymind to add deep learning to its data science curriculum in September 2014 — one of the first bootcamps to formally incorporate neural network training. Reflects the early excitement around deep learning reaching commercial viability.

  59. Recommending Music on Spotify with Deep Learning

    Sander Dieleman's write-up of his Spotify internship work — using convolutional neural networks on raw audio spectrograms to generate track embeddings for music recommendation. A landmark applied deep learning paper from 2014, before CNNs on audio became routine.

  60. 10 Tips for Better Deep Learning Models

    Laura Diane Hamilton's ten practical tips for improving deep learning model performance, covering data preparation, architecture choices, regularization, and training tricks. A snapshot of practitioner wisdom circa 2014, before the era of giant pretrained models made many of these tradeoffs less urgent.

  61. The Mission to Bring Google's AI to the Rest of the World

    Wired profiles Skymind, the startup founded by Adam Gibson to bring deep learning to enterprise Java shops — featuring Zipfian Academy's role in democratizing the field. A snapshot of 2014's optimism about deep learning accessibility before TensorFlow existed.

  62. The Mystery of Go, the Ancient Game That Computers Still Can't Win

    Wired's 2014 article arguing computers couldn't beat top Go players — written two years before DeepMind's AlphaGo defeated Lee Sedol. A remarkable historical document showing how quickly AI capabilities confounded expert predictions.

  63. ConvNetJS Demo: Classify Toy 2D Data

    Andrej Karpathy's ConvNetJS demo for classifying 2D toy datasets — a real-time browser visualization of a neural network training on user-drawn boundaries. One of the first compelling interactive deep learning visualizations.

  64. My Solution for the Galaxy Zoo Challenge — Sander Dieleman

    Sander Dieleman's winning solution for the Galaxy Zoo Kaggle challenge using convolutional neural networks to classify galaxy morphology from images. A landmark result showing CNNs achieving human-level performance on a citizen science dataset.

  65. Introduction to Deep Learning on Hadoop (Hadoop Summit 2014)

    A Hadoop Summit 2014 session proposal on deep learning at Hadoop scale — from the team behind DL4J (DeepLearning4J), a Java-native deep learning framework designed to run on Hadoop/Spark clusters. A snapshot of the moment distributed deep learning was being invented.

  66. NIPS 2012 Paper #1338

    A NIPS 2012 (NeurIPS 2012) conference paper — saved in the same timeframe as deep learning and distributed neural network research. Likely related to neural architecture or distributed/large-scale training given the surrounding bookmarks.

  67. Multilayer Perceptron with Jobman

    Deeplearning.net tutorial on training a multilayer perceptron using Jobman — a job scheduling system from the Montreal LISA Lab for running batches of hyperparameter experiments. Shows the infrastructure behind systematic deep learning research in 2014.

  68. Computer Science: The Learning Machines

    Nature News's January 2014 feature on machine learning and deep neural networks going mainstream — the moment the field started reaching a broader scientific audience. Documents the Hinton-LeCun-Bengio wave before it was inevitable.

  69. OverFeat: Integrated Recognition, Localization, and Detection

    OverFeat is NYU CILVR lab's deep learning framework for object recognition, localization, and detection using convolutional networks — won the ImageNet 2013 localization task. An important 2014 artifact of Yann LeCun's group.

  70. ConvNetJS: Deep Learning in Your Browser

    Andrej Karpathy's JavaScript library for training neural networks entirely in the browser — no install, no GPU, instant demos. In 2014 it made deep learning tangible for anyone with a web browser.

  71. UFLDL Tutorial — Stanford

    Andrew Ng's Stanford UFLDL Tutorial — the primary self-study resource for deep learning before MOOCs existed. Teaches sparse autoencoders, PCA, CNNs, and deep belief networks with mandatory from-scratch implementation exercises.

  72. Regularizing Neural Networks with Dropout and DropConnect

    FastML's comparison of Dropout and DropConnect — two techniques for regularizing neural networks by randomly zeroing activations or weights during training. Clarifies that DropConnect's CIFAR-10 SOTA came from model ensembling, not the technique itself.

  73. Deep Learning Tutorials — DeepLearning.net

    The canonical deep learning tutorial site from the Montreal group — code-first walkthrough of core architectures (RBM, DBN, CNN, LSTM) using Theano. Written by the lab around Yoshua Bengio and became the standard reference before fast.ai existed.

  74. Deep Learning 101

    Introductory survey of deep learning for non-specialists — covers hierarchical representation learning, RBMs, autoencoders, and the four core obstacles. Written to help readers mentally filter hype from substance in 2013.

  75. Deep Learning 101 — Hacker News Discussion

    Hacker News discussion of the Deep Learning 101 article (markus.com) — saved alongside the original article. See the primary note for content.

  76. Recursive Deep Models for Semantic Compositionality

    Stanford's Recursive Neural Tensor Network paper and Sentiment Treebank dataset — Richard Socher's EMNLP 2013 work that used tree-structured recursive neural networks to predict fine-grained sentiment at every node of a parse tree. A landmark paper that pushed NLP models toward compositionality.

  77. Deep Support Vector Machines

    A video lecture on Deep Support Vector Machines from ROKS 2013 — hybrid architectures combining deep feature learning with SVM classification. A snapshot of the moment researchers explored whether SVMs and deep learning could coexist before end-to-end networks won out.

All bookmarks