Subject
10 entries
Mapreduce
Bookmarks
Model for Massively Parallel Computation — MapReduce Theory
Grigory Yaroslavtsev's theoretical treatment of MapReduce as a model for massively parallel computation — covering the MRC complexity class and what it tells us about which problems can be solved efficiently at scale.
Pig Not-So-Foreign Language: Paper Notes
Bugra Akyildiz's notes on the 'Pig Latin: A Not-So-Foreign Language for Data Processing' paper by Yahoo! Research — explaining how Pig Latin compiles high-level data flow operations to MapReduce jobs. Captures why Pig was a meaningful step up from raw MapReduce for ETL work.
Introduction to Recommendations with Map-Reduce and mrjob
Tutorial on building item-based collaborative filtering recommendation systems using MapReduce and Yelp's mrjob Python library. Shows why distributed computation is necessary for large-scale similarity calculations.
Streaming MapReduce with Summingbird
Twitter's Summingbird library unifies batch (Hadoop/Scalding) and streaming (Storm) MapReduce into a single Scala API — write once, run on either backend. An early practical implementation of the Lambda Architecture's dual-speed compute model.
Hadoop 2.0 & YARN: The Big Data Breakthrough
ReadWrite's accessible overview of Hadoop 2.0 and YARN, explaining why the resource manager redesign was a bigger deal than an incremental release — it turned Hadoop from a MapReduce platform into a general-purpose cluster resource manager.
MapReduce Patterns, Algorithms, and Use Cases
Ilya Katsov's comprehensive taxonomy of MapReduce design patterns — from basic counting and filtering through complex join strategies and graph algorithms. The field guide for wringing correct and efficient computation out of the MapReduce model.
Apache Hadoop: Best Practices and Anti-Patterns
Yahoo's engineering blog guide on Hadoop best practices and anti-patterns from the team that ran the world's largest Hadoop clusters in 2010 — practical tuning advice from first-hand production experience at scale.
Introduction to Big Data and Hadoop Ecosystem
A beginner's introduction to the Apache Hadoop ecosystem cataloguing all major components — HDFS, MapReduce, Hive, HBase, Pig, Mahout, Sqoop, ZooKeeper, and more. A useful taxonomy of the 2012-era Hadoop stack.
Beyond Hadoop: Next-Generation Big Data Architectures
GigaOM's 2010 survey of Google's next-generation big data architectures — MPI, Pregel, Dremel, and Percolator — that were being developed as alternatives or complements to MapReduce. The paper that first systematically articulated why MapReduce wasn't sufficient for everything.
YARN: Hadoop NextGen MapReduce
YARN's official documentation from 2012 — the architecture that decoupled Hadoop cluster resource management from MapReduce. By separating resource negotiation into its own layer, YARN turned Hadoop from a MapReduce platform into a general-purpose cluster OS.
