Subject
8 entries
Bandits
Bookmarks
Multi-Armed Bandits and the Stitch Fix Experimentation Platform
Stitch Fix's blog on multi-armed bandits as an alternative to A/B testing — Thompson Sampling routes traffic toward better-performing arms dynamically, reducing wasted exposure. Strong motivation for when bandits beat traditional experimentation.
Optimism in the Face of Uncertainty: the UCB1 Algorithm
Jeremy Kun's accessible treatment of the UCB1 algorithm — the principle of 'optimism in the face of uncertainty' formalized as a bandit algorithm with proven regret bounds. Shows why adding a confidence bonus to estimated rewards elegantly solves the exploration-exploitation tradeoff.
Optimal Thompson Sampling: Asymptotic Analysis
Emilie Kaufmann's arXiv paper on the asymptotic optimality of Thompson Sampling for multi-armed bandits — the theoretical grounding that explains why Thompson Sampling works as well as it does empirically. Proves it achieves near-optimal regret bounds.
Multi-Armed Bandits
Cameron Davidson-Pilon's blog post on multi-armed bandits from a Bayesian perspective — draws on the same probabilistic programming intuition as his 'Bayesian Methods for Hackers' book. Frames bandits as the natural application of iterative belief updating.
Multi-Armed Bandit Experiments
Analytics blog post on multi-armed bandit experiments as a replacement for static A/B testing in website optimization — covers epsilon-greedy, UCB, and Thompson Sampling with practical framing for product teams.
A Book About Bandit Algorithms
John Myles White's free book on bandit algorithms for website optimization — covers epsilon-greedy, softmax, UCB, and Thompson Sampling with practical web application examples. An accessible bridge from theory to product experimentation.
Understanding Multi-Armed Bandit Algorithms
DataBozo's conceptual explanation of multi-armed bandit algorithms — epsilon-greedy, UCB, and Thompson Sampling compared from first principles. Aimed at practitioners who want to understand the tradeoffs before implementing.
Why Wikipedia's A/B Testing Is All Wrong (And How Contextual Bandits Can Fix It)
Synference blog post critiquing Wikipedia's A/B testing approach — argues that testing average effects ignores user heterogeneity, and contextual bandits could deliver personalized treatments rather than one-size-fits-all decisions.
