Subject
9 entries
Theory
Bookmarks
Why Deep Learning Works Unreasonably Well
Part 3 of a 'How Models Learn' series on why deep learning works unreasonably well — addressing the apparent paradox that overparameterized models generalize when classical statistics says they shouldn't. Covers implicit regularization, loss landscape geometry, and the lottery ticket hypothesis.
Machine Learning and Complexity Theory
A theoretical computer scientist's reflection on the relationship between machine learning and computational complexity theory after a Dagstuhl seminar — noting that the two fields remain largely disconnected despite studying adjacent phenomena. Raises questions about whether complexity theory has useful things to say about why ML works.
What Learning Algorithm Is In-Context Learning? Investigations with Linear Models
ICLR 2023 paper showing that transformers trained on in-context learning tasks implicitly implement gradient descent and ridge regression on linear problems, with layers encoding weight vectors and moment matrices. Foundational theoretical work explaining ICL as implicit algorithm execution rather than pure pattern matching.
Notes on Theory of Distributed Systems
James Aspnes's freely-distributed lecture notes on the theory of distributed systems, covering fault tolerance, consensus, synchrony models, and randomized algorithms. A rigorous but accessible graduate reference that grounds distributed computing in formal models.
Hopfield Networks is All You Need
Ramsauer et al. introduce a modern continuous-state Hopfield network with exponential storage capacity and prove that its update rule is mathematically equivalent to transformer self-attention. This is the foundational paper connecting classical associative memory to the attention mechanism, explaining why transformers work through an energy-function lens.
The Overfitted Brain: Dreams Evolved to Assist Generalization
A 2020 paper proposing that dreams evolved as a biological regularization mechanism — the brain 'trains' on noisy, hallucinated data during sleep to prevent overfitting to waking experience. A striking bridge between ML theory and sleep neuroscience.
50 Years of Data Science
David Donoho's essay arguing that 'data science' is a real intellectual discipline distinct from statistics — tracing 50 years of statistical evolution toward greater empiricism, computation, and scale. A foundational text for anyone who wants to understand what data science actually is and where it came from.
The Remarkable k-means++
Larry Wasserman's Normal Deviate blog post on k-means++ — the 2007 initialization trick from Arthur and Vassilvitskii that gives k-means an O(log k) approximation guarantee and better convergence in practice.
The True Power of Regular Expressions
Nikita Popov's deep dive into what regular expressions can theoretically do — connecting regex to finite automata and formal language theory. Goes beyond syntax tutorials to explain why backtracking regex engines can be exponentially slow, and how to avoid it.
