Subject
69 entries
Hadoop
Bookmarks
Hadoop, Python, and NoSQL Lead the Pack for Big Data Jobs
InfoWorld's 2014 analysis of job postings showing Hadoop, Python, and NoSQL as the top skills in big data job listings — a snapshot of the technology bets companies were making at the height of the big data boom.
Why Apache Spark is a Crossover Hit for Data Scientists
Cloudera's post on why Apache Spark resonated with data scientists in ways Hadoop MapReduce never did — the interactive REPL, Python support, and in-memory computation made it feel like a supercharged pandas rather than a distributed systems project.
Introduction to Deep Learning on Hadoop (Hadoop Summit 2014)
A Hadoop Summit 2014 session proposal on deep learning at Hadoop scale — from the team behind DL4J (DeepLearning4J), a Java-native deep learning framework designed to run on Hadoop/Spark clusters. A snapshot of the moment distributed deep learning was being invented.
Cloudera Rebuilding Machine Learning for Hadoop with Oryx
GigaOm's coverage of Cloudera launching Oryx — an open-source ML-on-Hadoop framework using the Lambda Architecture for batch retraining plus real-time serving. An early attempt to make production machine learning first-class on Hadoop.
Pig Not-So-Foreign Language: Paper Notes
Bugra Akyildiz's notes on the 'Pig Latin: A Not-So-Foreign Language for Data Processing' paper by Yahoo! Research — explaining how Pig Latin compiles high-level data flow operations to MapReduce jobs. Captures why Pig was a meaningful step up from raw MapReduce for ETL work.
Hadoop for Data Science
Mortar Data's introduction to Hadoop for data scientists — when to use it, what the MapReduce programming model actually means, and how Pig Latin abstracts away the low-level boilerplate.
Introduction to Recommendations with Map-Reduce and mrjob
Tutorial on building item-based collaborative filtering recommendation systems using MapReduce and Yelp's mrjob Python library. Shows why distributed computation is necessary for large-scale similarity calculations.
Hadoop Creator: Google Is Living a Few Years in the Future
Doug Cutting (Hadoop creator) on Google living years ahead in infrastructure — the observation that Google's internal systems consistently anticipate what the rest of the industry will need, and then the open-source community builds it later.
Rolling Average in Hive
Brent Ozar's walkthrough of computing rolling averages in Hive — a problem that looks like a simple SQL query but requires window functions or self-joins in Hive's then-limited SQL dialect. A practical data engineering puzzle from the Hadoop era.
Apache Hadoop 2 Is Now GA
Hortonworks announcement that Apache Hadoop 2 reached general availability in October 2013 — introducing YARN as the cluster resource manager. A landmark release that decoupled compute from MapReduce and made Hadoop a general-purpose cluster platform.
Git Data Mining with Hadoop
WANdisco's post on mining git repository data at scale using Hadoop — applying distributed batch processing to version control history to extract patterns across large codebases. An early example of treating code evolution as a data science problem.
Hadoop's Unsung Sweet Spot: Unstructured ETL
A LinkedIn discussion arguing that Hadoop's real sweet spot wasn't analytics but unstructured ETL — transforming messy, heterogeneous data into structured forms before loading into traditional data warehouses. A nuanced counterpoint to the 'Hadoop replaces SQL' narrative of 2013.
Big Data Ecosystem Map (2013)
A map of the big data ecosystem as it existed in 2013 — an attempt to catalog the dozens of Hadoop-adjacent tools, databases, analytics platforms, and services that had emerged. A useful historical artifact of the big data vendor explosion before consolidation.
Hadoop Meets SQL
IBM Big Data Hub's overview of SQL-on-Hadoop approaches in 2013 — Hive, Impala, and the broader push to make Hadoop queryable by the vast majority of analysts who knew SQL but not MapReduce. The SQL interface became the primary adoption driver for Hadoop in the enterprise.
Nutanix Engineer Explains Why His Appliance Makes Sense for Analytics
SiliconAngle interview with a Nutanix engineer making the case for why hyper-converged infrastructure is well-suited to analytics workloads in 2013. Nutanix was positioning its appliances as a natural home for Hadoop alongside traditional virtualization.
Implementing Hadoop for Big Data Projects
Inside Analysis piece on practical considerations for implementing Hadoop in enterprise big data projects circa 2013 — covering organizational readiness, hardware choices, and the gap between Hadoop's promise and production reality.
Nutanix Hadoop Solution Brief
Nutanix's Hadoop solution brief describing how NDFS (Nutanix Distributed File System) enables running Hadoop workloads on hyper-converged infrastructure. Marketing collateral saved during Ryan's time at Nutanix.
Yahoo! Spinning Continuous Computing with YARN
Yahoo's 2013 exploration of using YARN as a substrate for continuous/streaming computation beyond batch MapReduce. An early signal that the Hadoop ecosystem was trying to absorb real-time processing use cases before Apache Spark and Flink fully took over.
Netflix Genie: Hadoop Platform-as-a-Service
Netflix open-sourced Genie in mid-2013 — a REST-based Hadoop Platform-as-a-Service that abstracted job submission across multiple Hadoop clusters. A key piece of Netflix's data platform that became an influential pattern for multi-cluster job routing.
Nutanix Joins Hortonworks Certified Technology Partner Program
Nutanix joined the Hortonworks Certified Technology Partner Program in June 2013, validating that Nutanix's NDFS could run certified HDP (Hortonworks Data Platform) workloads. A business development milestone for Nutanix's Hadoop positioning.
Hadoop 2.0 & YARN: The Big Data Breakthrough
ReadWrite's accessible overview of Hadoop 2.0 and YARN, explaining why the resource manager redesign was a bigger deal than an incremental release — it turned Hadoop from a MapReduce platform into a general-purpose cluster resource manager.
Hadoop on VMware: Another Workload Conquered?
EMC's Chuck Hollis on running Hadoop workloads on VMware virtualization — the argument that bare-metal Hadoop deployments could be replaced by virtualized clusters with manageable performance trade-offs. A 2013 salvo in the bare-metal vs. virtualization debate for big data workloads.
The Buzz and Fuzz on SSD in Hadoop
Hadoopsphere's analysis of SSD adoption in Hadoop clusters — the trade-off between SSD's dramatically faster random I/O and its higher cost per GB compared to spinning disk. An early look at how flash storage would eventually reshape big data infrastructure.
Nobody Ever Got Fired for Using Hadoop on a Cluster
Steve Loughran's riff on the classic 'nobody got fired for buying IBM' adapted for the big data era: Hadoop became the safe enterprise choice for data infrastructure, which brought both legitimacy and the mediocrity that comes with default choices.
Is There Room for SSDs in the Hadoop Framework?
StorageTuning blog's technical assessment of where SSDs fit in Hadoop's storage hierarchy — evaluating specific bottlenecks (NameNode metadata, shuffle I/O, random reads) where flash storage provides measurable gains over spinning disk.
Hadoop Innovations at BayLISA
BayLISA meetup event featuring talks on Hadoop innovations, including from Ryan Orban (the vault owner), Eric Sammer, and Alan Gates. A snapshot of the Bay Area Hadoop community circa 2013 and the kind of practitioner knowledge-sharing that drove the ecosystem forward.
Cloudera Developer Class Links
Cloudera's developer training class resource links page — a collection of documentation, tutorials, and reference materials for the Cloudera Hadoop Developer certification course. Saved as a training reference during the 2013 era of Hadoop skills development.
Get Started: Ambari for Provisioning, Managing and Monitoring Hadoop
Hortonworks' getting-started guide for Apache Ambari — the web-based Hadoop cluster provisioning and monitoring tool that aimed to make Hadoop operations accessible without deep Linux/Hadoop expertise. An important piece of 2013 Hadoop ecosystem tooling.
Best Practices for Virtualizing Hadoop
Hadoop Summit presentation on best practices for running Hadoop on VMware HVE (Hadoop Virtualization Extensions) with Hortonworks HDP — addressing the core tension between virtualization flexibility and Hadoop's data locality requirements.
MapReduce Patterns, Algorithms, and Use Cases
Ilya Katsov's comprehensive taxonomy of MapReduce design patterns — from basic counting and filtering through complex join strategies and graph algorithms. The field guide for wringing correct and efficient computation out of the MapReduce model.
Apache Hadoop: Best Practices and Anti-Patterns
Yahoo's engineering blog guide on Hadoop best practices and anti-patterns from the team that ran the world's largest Hadoop clusters in 2010 — practical tuning advice from first-hand production experience at scale.
Welcome to Berkeley: Where Hadoop Isn't Nearly Fast Enough
GigaOM's April 2013 profile of UC Berkeley's AMPLab making the case that Hadoop is too slow for interactive and iterative workloads — the academic origin story of Apache Spark, Shark (early Spark SQL), and Mesos. A snapshot of the moment Spark was about to go mainstream.
Future of Apache Hive — SQL PASS BA 2013
Carter Shanklin's 2013 talk on the future of Apache Hive — covering the push to make Hive's HiveQL a proper SQL dialect with ACID semantics, better query planning, and sub-second latency. A snapshot of the war between SQL-on-Hadoop and traditional data warehouses.
Natural Language Processing with Apache Hadoop and Python
Cloudera's 2010 post (bookmarked in 2013) on running natural language processing pipelines with Apache Hadoop and Python using Hadoop Streaming — an early recipe for scaling NLP beyond single-machine limits using commodity clusters.
Hadoop Illuminated — Free Open-Source Hadoop Book
Hadoop Illuminated is a free, open-source book on Apache Hadoop — a community-maintained guide covering HDFS, MapReduce, Hive, Pig, and the broader Hadoop ecosystem. One of the better free learning resources at a time when the Hadoop ecosystem was evolving faster than formal textbooks.
Apache Incubator Giraph — Distributed Graph Processing on Hadoop
Apache Giraph is a graph processing framework built on Hadoop — an open-source implementation of Google's Pregel model for iterative graph algorithms at scale. The Apache answer to graph-scale problems like PageRank, community detection, and shortest paths on billion-node graphs.
BioInformatics: A Data Deluge with Hadoop to the Rescue
Datanami on using Apache Hadoop for bioinformatics data pipelines — how genomic sequencing data had outpaced traditional computational biology infrastructure and why Hadoop's distributed file system and MapReduce were being adopted to handle the deluge.
Towards an Elastic Elephant: Enabling Hadoop for the Cloud
VMware's CTO blog on making Hadoop work in elastic cloud environments — the fundamental tension between Hadoop's static cluster model and cloud infrastructure's on-demand scaling. A 2013 take on a problem that took years to fully solve.
YARN: Yet Another Resource Negotiator
Apache YARN (Yet Another Resource Negotiator) is the cluster resource management layer introduced in Hadoop 2 — the architectural change that turned Hadoop from a MapReduce system into a general-purpose distributed compute platform.
Working with Big Data: Assembling Your Toolkit
Acquia's guide to assembling a big data toolkit in 2013 — walking through the layered stack from storage to analysis. Practical survey of the Hadoop ecosystem components a team would actually need to evaluate.
CloudCamp / Hadoop
CloudCamp wiki page on Hadoop — a community-maintained reference during the peak of the Hadoop hype cycle in 2012-2013. One of many informal educational resources that proliferated as enterprise interest in Hadoop outpaced formal documentation.
Big Data Explained
BlazeClan's explainer on big data for a general technical audience — covering the 4 Vs framework and why traditional databases couldn't handle the scale. A typical 2013 entry-level primer during peak big data hype.
The Big Challenge of Big Data and Hadoop Integration
Cloud Computing Journal on the integration challenges of adding Hadoop to existing enterprise data infrastructure — connecting it to RDBMS systems, BI tools, and existing ETL pipelines. The 2013 reality check on big data adoption.
Introduction to Big Data and Hadoop Ecosystem
A beginner's introduction to the Apache Hadoop ecosystem cataloguing all major components — HDFS, MapReduce, Hive, HBase, Pig, Mahout, Sqoop, ZooKeeper, and more. A useful taxonomy of the 2012-era Hadoop stack.
Big Data: Hadoop Distributions Compared
Comparison of the main commercial Hadoop distributions in 2013 — Cloudera CDH, Hortonworks HDP, and MapR — when the market was consolidating around a handful of vendors each taking different bets on what enterprises needed.
Beyond Hadoop: Next-Generation Big Data Architectures
GigaOM's 2010 survey of Google's next-generation big data architectures — MPI, Pregel, Dremel, and Percolator — that were being developed as alternatives or complements to MapReduce. The paper that first systematically articulated why MapReduce wasn't sufficient for everything.
Proprietary Hadoop Is a Losing Strategy
ReadWrite's 2013 argument that commercial Hadoop vendors who added proprietary lock-in would lose to those who contributed everything upstream — a prescient thesis that proved partly right, partly wrong over the following decade.
Parquet: Columnar Storage for Hadoop
The original announcement of Apache Parquet, the columnar storage format Twitter and Cloudera jointly released for Hadoop in 2013. Parquet became the dominant format for analytical workloads in the Hadoop/Spark ecosystem and remains ubiquitous in modern data lakes.
Analyzing Big Data with Twitter — UC Berkeley iSchool
UC Berkeley iSchool's 2012 course on analyzing big data with Twitter — an early academic offering that bridged social media data and distributed computing tooling before data science programs existed at most universities.
5 Reasons Why the Future of Hadoop Is Real-Time
GigaOM's 2013 argument for why Hadoop was evolving toward real-time processing — covering YARN, Storm, and in-memory frameworks as the drivers. A snapshot of the moment when batch-only Hadoop started feeling inadequate.
Cloudera's Support Team Shares Some Basic Hardware Recommendations
Cloudera's 2010 hardware recommendations for Hadoop clusters — still relevant in 2013 when bookmarked. The canonical guidance on disk, RAM, and CPU specs for commodity Hadoop nodes before cloud deployments became dominant.
Hortonworks Joins OpenStack Foundation
Hortonworks joining the OpenStack Foundation in early 2013 — a sign that the big data and cloud infrastructure worlds were converging. The tweet called 'virtualized Hadoop' the future, which captures exactly the direction that HDP-on-OpenStack was heading.
A Programmer's Guide to Big Data: 12 Tools to Know
GigaOM's 2012 reference guide to 12 big data tools a programmer should know — covering Hadoop, Pig, Hive, HBase, Storm, and emerging alternatives. A snapshot of the Hadoop ecosystem at its peak complexity, before Spark simplified much of it.
Hidden Markov Models on Hadoop — Isabel Drost
Isabel Drost's slides on Hidden Markov Models and Hadoop — covering how HMMs can be implemented and trained at scale using MapReduce. A 2012-era reference for scaling sequence models before deep learning displaced them.
HDFS Has Won: De Facto Standard for Centralized Data Storage
A 2012 claim that HDFS had emerged as the de facto standard for centralized big data storage — a snapshot of a moment when Hadoop's dominance seemed settled. Written just before the data lake era that HDFS would define, and before object storage (S3) would eventually displace it.
5 Trends That Are Changing How We Do Big Data
GigaOM's 2012 survey of five trends reshaping big data: real-time processing rising against batch, NoSQL maturity, cloud-based data platforms, open-source ecosystem growth, and the shift from data collection to data monetization. A useful snapshot of where the field was heading at Hadoop's peak.
Analyzing Human Genomes with Hadoop
Cloudera's 2009 blog post (bookmarked in 2012) showing how MapReduce and Hadoop can process human genome sequences at scale — an early example of big data infrastructure being applied to life sciences problems that were previously computationally intractable.
Facebook Seeks Next-Generation Big Data Tools
InformationWeek's 2012 report on Facebook's efforts to move beyond Hadoop for its next-generation big data infrastructure — one of the first major signals that the tech industry's heaviest Hadoop user was already looking past it.
New Tools Simplify the Power of Hadoop
InformationWeek's 2012 survey of tools making Hadoop more approachable for enterprise practitioners — the abstraction layers (Hive, Pig, HBase) and commercial distributions (Cloudera, Hortonworks) that were papering over Hadoop's low-level complexity. Captures the ecosystem's maturation moment.
Big Ideas: Demystifying Hadoop
A 'Big Ideas: Demystifying Hadoop' YouTube explainer from 2012 — one of many educational resources that emerged as Hadoop moved from niche to mainstream. Aimed at explaining the MapReduce paradigm and HDFS to practitioners who hadn't yet had to deal with data at scale.
Nutanix Complete Hadoop Appliance
Ryan's own SlideShare presentation for a Nutanix webinar on the Complete Hadoop Appliance — Nutanix's pitch for running Hadoop workloads on hyper-converged infrastructure instead of dedicated Hadoop clusters. This was Ryan's role at Nutanix circa 2012.
How Is Big Data Faring in the Enterprise?
A ZDNet survey of how big data technologies were actually landing in enterprise environments in mid-2012 — a useful reality check during peak Hadoop hype. Most large organizations were experimenting but few had production big data systems delivering measurable business value.
Cleversafe and Hadoop: Combining Next Generation Storage with Big Data Analytics
Cleversafe's 2012 integration with Hadoop connected dispersed object storage (using information dispersal algorithms rather than replication) to the Hadoop analytics ecosystem. Later acquired by IBM and became IBM Cloud Object Storage.
YARN: Hadoop NextGen MapReduce
YARN's official documentation from 2012 — the architecture that decoupled Hadoop cluster resource management from MapReduce. By separating resource negotiation into its own layer, YARN turned Hadoop from a MapReduce platform into a general-purpose cluster OS.
Why the Days Are Numbered for Hadoop As We Know It
A 2012 GigaOM piece arguing that Hadoop's architecture had fundamental limitations that would force it to evolve or be displaced — written at the peak of Hadoop hype. Prescient in identifying YARN and the multi-framework future, but underestimated how long it would take for cloud-native alternatives to win.
Twitter to Open Source Hadoop-Like Tool (Storm)
GigaOM coverage of Twitter's plans to open-source Storm — their real-time stream processing system. The moment the Hadoop-for-streaming gap became a major industry conversation, and Nathan Marz's Storm became the answer.
Apache Incubator Giraph
Apache Giraph's incubator homepage from 2012 — the open-source implementation of Google's Pregel bulk-synchronous-parallel graph processing model. Bookmarked during an early phase of the Hadoop ecosystem expansion into graph workloads.
Apache Hive
Apache Hive's homepage from 2012 — the SQL-on-Hadoop layer that made big data accessible to analysts who knew SQL but not Java MapReduce. Hive translated HiveQL queries into MapReduce jobs, trading latency for familiarity.
Apache HBase
Apache HBase's homepage from 2012 — the open-source implementation of Google Bigtable that added random-read/write access to Hadoop's otherwise write-once HDFS. HBase filled the gap MapReduce couldn't: low-latency lookups on data stored across a distributed cluster.
