Skip to main content
Ryan Orban

Ryan Orban

Subject
8 entries

Distributed Computing

Bookmarks

  1. A Hitchhiker's Guide to Distributed Training of Deep Neural Networks

    Chahal, Grover, and Dey survey the algorithms and engineering techniques for distributed deep learning training, covering data parallelism, model parallelism, AllReduce strategies, gradient compression, and mixed precision. A practical reference for scaling training from single GPU to multi-node clusters.

  2. Secure Multiparty Computation (MPC)

    Yehuda Lindell's accessible overview of secure multiparty computation (MPC) — how parties can jointly compute on private inputs without revealing them. Traces the field from Yao's two-party garbled circuits through modern practical protocols used in industry.

  3. Data Science with Python and Dask

    Jesse Daniel's Manning 2019 book teaching Dask for parallel and out-of-core data science in Python — using familiar pandas-like DataFrames and numpy-like arrays across cores and machines. The go-to resource for scaling Python data science workflows beyond single-machine memory limits.

  4. Ponder: Pandas at Scale

    Ponder is a startup that makes Pandas run at scale without rewriting your code — a drop-in compatibility layer that runs standard Pandas operations on distributed backends. Targets the massive installed base of data scientists who know Pandas but hit its single-machine limits.

  5. Bayesian Machine Learning on Apache Spark

    Cloudera's engineering blog post on implementing Bayesian machine learning on Apache Spark — combining probabilistic inference with distributed computation. A technically ambitious combination that was ahead of most production ML stacks in 2014.

  6. Hadoop for Data Science

    Mortar Data's introduction to Hadoop for data scientists — when to use it, what the MapReduce programming model actually means, and how Pig Latin abstracts away the low-level boilerplate.

  7. Parameter Optimization with Zipline, PiCloud, StarCluster, and IPython Parallel

    Quantopian's blog post on running parameter optimization for trading strategies using Zipline backtester on PiCloud, StarCluster, and IPython Parallel — a 2013 example of cloud-distributed backtesting before it was a product. Shows the DIY infrastructure that Quantopian later packaged into their platform.

  8. CloudCamp / Hadoop

    CloudCamp wiki page on Hadoop — a community-maintained reference during the peak of the Hadoop hype cycle in 2012-2013. One of many informal educational resources that proliferated as enterprise interest in Hadoop outpaced formal documentation.

All bookmarks