Skip to main content
Ryan Orban

Ryan Orban

Subject
8 entries

Bandits

Bookmarks

  1. Multi-Armed Bandits and the Stitch Fix Experimentation Platform

    Stitch Fix's blog on multi-armed bandits as an alternative to A/B testing — Thompson Sampling routes traffic toward better-performing arms dynamically, reducing wasted exposure. Strong motivation for when bandits beat traditional experimentation.

  2. Optimism in the Face of Uncertainty: the UCB1 Algorithm

    Jeremy Kun's accessible treatment of the UCB1 algorithm — the principle of 'optimism in the face of uncertainty' formalized as a bandit algorithm with proven regret bounds. Shows why adding a confidence bonus to estimated rewards elegantly solves the exploration-exploitation tradeoff.

  3. Optimal Thompson Sampling: Asymptotic Analysis

    Emilie Kaufmann's arXiv paper on the asymptotic optimality of Thompson Sampling for multi-armed bandits — the theoretical grounding that explains why Thompson Sampling works as well as it does empirically. Proves it achieves near-optimal regret bounds.

  4. Multi-Armed Bandits

    Cameron Davidson-Pilon's blog post on multi-armed bandits from a Bayesian perspective — draws on the same probabilistic programming intuition as his 'Bayesian Methods for Hackers' book. Frames bandits as the natural application of iterative belief updating.

  5. Multi-Armed Bandit Experiments

    Analytics blog post on multi-armed bandit experiments as a replacement for static A/B testing in website optimization — covers epsilon-greedy, UCB, and Thompson Sampling with practical framing for product teams.

  6. A Book About Bandit Algorithms

    John Myles White's free book on bandit algorithms for website optimization — covers epsilon-greedy, softmax, UCB, and Thompson Sampling with practical web application examples. An accessible bridge from theory to product experimentation.

  7. Understanding Multi-Armed Bandit Algorithms

    DataBozo's conceptual explanation of multi-armed bandit algorithms — epsilon-greedy, UCB, and Thompson Sampling compared from first principles. Aimed at practitioners who want to understand the tradeoffs before implementing.

  8. Why Wikipedia's A/B Testing Is All Wrong (And How Contextual Bandits Can Fix It)

    Synference blog post critiquing Wikipedia's A/B testing approach — argues that testing average effects ignores user heterogeneity, and contextual bandits could deliver personalized treatments rather than one-size-fits-all decisions.

All bookmarks