Skip to main content
Ryan Orban

Ryan Orban

Subject
10 entries

Mapreduce

Bookmarks

  1. Model for Massively Parallel Computation — MapReduce Theory

    Grigory Yaroslavtsev's theoretical treatment of MapReduce as a model for massively parallel computation — covering the MRC complexity class and what it tells us about which problems can be solved efficiently at scale.

  2. Pig Not-So-Foreign Language: Paper Notes

    Bugra Akyildiz's notes on the 'Pig Latin: A Not-So-Foreign Language for Data Processing' paper by Yahoo! Research — explaining how Pig Latin compiles high-level data flow operations to MapReduce jobs. Captures why Pig was a meaningful step up from raw MapReduce for ETL work.

  3. Introduction to Recommendations with Map-Reduce and mrjob

    Tutorial on building item-based collaborative filtering recommendation systems using MapReduce and Yelp's mrjob Python library. Shows why distributed computation is necessary for large-scale similarity calculations.

  4. Streaming MapReduce with Summingbird

    Twitter's Summingbird library unifies batch (Hadoop/Scalding) and streaming (Storm) MapReduce into a single Scala API — write once, run on either backend. An early practical implementation of the Lambda Architecture's dual-speed compute model.

  5. Hadoop 2.0 & YARN: The Big Data Breakthrough

    ReadWrite's accessible overview of Hadoop 2.0 and YARN, explaining why the resource manager redesign was a bigger deal than an incremental release — it turned Hadoop from a MapReduce platform into a general-purpose cluster resource manager.

  6. MapReduce Patterns, Algorithms, and Use Cases

    Ilya Katsov's comprehensive taxonomy of MapReduce design patterns — from basic counting and filtering through complex join strategies and graph algorithms. The field guide for wringing correct and efficient computation out of the MapReduce model.

  7. Apache Hadoop: Best Practices and Anti-Patterns

    Yahoo's engineering blog guide on Hadoop best practices and anti-patterns from the team that ran the world's largest Hadoop clusters in 2010 — practical tuning advice from first-hand production experience at scale.

  8. Introduction to Big Data and Hadoop Ecosystem

    A beginner's introduction to the Apache Hadoop ecosystem cataloguing all major components — HDFS, MapReduce, Hive, HBase, Pig, Mahout, Sqoop, ZooKeeper, and more. A useful taxonomy of the 2012-era Hadoop stack.

  9. Beyond Hadoop: Next-Generation Big Data Architectures

    GigaOM's 2010 survey of Google's next-generation big data architectures — MPI, Pregel, Dremel, and Percolator — that were being developed as alternatives or complements to MapReduce. The paper that first systematically articulated why MapReduce wasn't sufficient for everything.

  10. YARN: Hadoop NextGen MapReduce

    YARN's official documentation from 2012 — the architecture that decoupled Hadoop cluster resource management from MapReduce. By separating resource negotiation into its own layer, YARN turned Hadoop from a MapReduce platform into a general-purpose cluster OS.

All bookmarks