Subject
105 entries
Big Data
Bookmarks
Data Science with Python and Dask
Jesse Daniel's Manning 2019 book teaching Dask for parallel and out-of-core data science in Python — using familiar pandas-like DataFrames and numpy-like arrays across cores and machines. The go-to resource for scaling Python data science workflows beyond single-machine memory limits.
Algorithms and Data Structures for Massive Datasets
Manning 2021 textbook covering algorithms and data structures built for massive datasets — Bloom filters, HyperLogLog, Count-Min Sketch, LSH, and streaming algorithms. Practical treatment of how to handle data that won't fit in memory or where exact answers are too expensive.
Insight Data Engineering Ecosystem Map
Insight Data Science's map of the data engineering ecosystem circa 2017 — a diagram organizing dozens of tools by pipeline stage (ingest, store, process, visualize). A widely shared snapshot of the Hadoop/Spark era's explosion of competing infrastructure tools.
Big Data Was Supposed to Fix Education. It Didn't. It's Time for 'Small Data.'
A 2016 Washington Post argument that big data initiatives in K-12 education have failed to improve learning outcomes — and that the alternative is 'small data': teachers knowing individual students qualitatively rather than tracking them at population scale.
Is It Pokemon or Big Data?
A quiz that asks whether a name belongs to a Pokemon or a Big Data tool — hilariously difficult because the Hadoop ecosystem's naming conventions are indistinguishable from fictional creature names. A sharp cultural critique disguised as a game.
Big Data: The Power of Petabytes (Genomics Edition)
Nature's supplement on big data in genomics, covering the arrival of the $1,000 genome and what it means for medicine — the data storage, analysis, and clinical interpretation challenges of a world where genome sequencing is routine. A milestone in the convergence of biology and data infrastructure.
Top Mistakes Developers Make When Using Python for Big Data Analytics
A practical rundown of the top Python performance mistakes for big data workloads — covering generator vs. list comprehension choices, pandas anti-patterns, and when to reach for NumPy. Still relevant since Python's core performance traps haven't changed.
Databricks Debuts Apache Spark MOOCs on BerkeleyX and edX
Databricks launches two MOOCs on Apache Spark through BerkeleyX and edX in December 2014, making distributed data processing education freely accessible. Marks the moment Spark was trying to become the standard for big data — before it succeeded.
Model for Massively Parallel Computation — MapReduce Theory
Grigory Yaroslavtsev's theoretical treatment of MapReduce as a model for massively parallel computation — covering the MRC complexity class and what it tells us about which problems can be solved efficiently at scale.
For Big Data Scientists, Hurdle to Insights Is Janitor Work
The New York Times article that popularized the term 'janitor work' for data cleaning — reporting that data scientists spend 50-80% of their time on data preparation rather than analysis. Validation from a mainstream outlet that this unglamorous reality was the actual job.
Hadoop, Python, and NoSQL Lead the Pack for Big Data Jobs
InfoWorld's 2014 analysis of job postings showing Hadoop, Python, and NoSQL as the top skills in big data job listings — a snapshot of the technology bets companies were making at the height of the big data boom.
Why Apache Spark is a Crossover Hit for Data Scientists
Cloudera's post on why Apache Spark resonated with data scientists in ways Hadoop MapReduce never did — the interactive REPL, Python support, and in-memory computation made it feel like a supercharged pandas rather than a distributed systems project.
Random Sampling from Very Large Files
Practical techniques for taking random samples from large files without loading them into memory — covering Unix tools (shuf, awk) and reservoir sampling. Essential for working with data too large for pandas to read in one shot.
Introduction to Deep Learning on Hadoop (Hadoop Summit 2014)
A Hadoop Summit 2014 session proposal on deep learning at Hadoop scale — from the team behind DL4J (DeepLearning4J), a Java-native deep learning framework designed to run on Hadoop/Spark clusters. A snapshot of the moment distributed deep learning was being invented.
Spark, Storm and Real-Time Analytics
An early comparison of Apache Storm and Spark Streaming for real-time analytics — covering their different processing models, latency characteristics, and use cases. A snapshot of the stream processing options available before Kafka Streams, Flink, and other tools matured.
The Next Big Thing You Missed: Keen IO's Plot to Beat Google at Big Data
Wired's 2014 profile of Keen IO — a startup offering analytics-as-an-API to let developers add event tracking without building their own data infrastructure. A bet that the analytics pipeline problem was common enough to be sold as a service.
Hadoop for Data Science
Mortar Data's introduction to Hadoop for data scientists — when to use it, what the MapReduce programming model actually means, and how Pig Latin abstracts away the low-level boilerplate.
Deploying Storm on GCE
Tutorial on deploying Apache Storm on Google Compute Engine — a setup guide for real-time stream processing at a time when cloud deployments of Storm were uncommon. GCE was a relatively new platform and Storm was the dominant real-time processing framework before Flink/Spark Streaming.
Drawing Inferences From Very Large Datasets
Econometrician Dave Giles on why large datasets make standard p-value thresholds useless — with N in the millions, almost any null hypothesis rejects, regardless of practical importance. A necessary corrective for data scientists drowning in statistical significance.
VentureBeat Data Architecture Diagram (Screenshot)
Screenshot of a data analysis architecture workflow diagram from VentureBeat, saved alongside the Twitter share of the same image. See the adjacent note on the data analysis architecture diagram.
Data Analysis Architecture Workflow Diagram
VentureBeat/DataBeat data analysis architecture diagram circulated on Twitter in 2013 — a workflow schematic showing the layers of a big data pipeline from collection through analysis to visualization. Snapshot of how practitioners were thinking about data infrastructure.
Hadoop Creator: Google Is Living a Few Years in the Future
Doug Cutting (Hadoop creator) on Google living years ahead in infrastructure — the observation that Google's internal systems consistently anticipate what the rest of the industry will need, and then the open-source community builds it later.
How Companies Are Using Spark
Strata/O'Reilly coverage of how companies were adopting Apache Spark in 2013, early in the engine's rise to ubiquity. A snapshot of early enterprise Spark use cases before it displaced Hadoop MapReduce as the default.
Presto: Interacting with Petabytes of Data at Facebook
Hacker News discussion on Facebook's newly open-sourced Presto SQL engine, capable of querying petabytes of data interactively. A watershed moment — before Presto, interactive SQL at Facebook scale wasn't possible.
Apache Hadoop 2 Is Now GA
Hortonworks announcement that Apache Hadoop 2 reached general availability in October 2013 — introducing YARN as the cluster resource manager. A landmark release that decoupled compute from MapReduce and made Hadoop a general-purpose cluster platform.
IBM Wants to Use Big Data to Predict Heart Disease Long Before It Strikes
VentureBeat on IBM's initiative to use big data and predictive models to forecast heart disease years before symptoms appear. A 2013 case study in applying ML to healthcare at scale — ambitious but illustrative of where the field thought it was going.
Why Recommendation Engines Are About to Get Much Better
Coverage of advances in recommendation engine technology in 2013 — driven by larger datasets, better collaborative filtering, and contextual signals. The moment when personalization was transitioning from a luxury to an expectation.
Timely Dataflow: An Introduction
An introduction to Timely Dataflow, the distributed computation model developed at Microsoft Research that unified batch and streaming processing through a novel timestamp-based progress tracking system. A technically significant but underappreciated alternative to the Spark/Storm paradigm.
Git Data Mining with Hadoop
WANdisco's post on mining git repository data at scale using Hadoop — applying distributed batch processing to version control history to extract patterns across large codebases. An early example of treating code evolution as a data science problem.
INFORMS Narrows Big Data Skills Gap
INFORMS (the operations research professional society) launching continuing education courses to address the big data skills gap in 2013 — a telling sign that demand for analytics talent had outpaced formal education pipelines. The gap was real, but the institutional response came well after the bootcamp ecosystem had already mobilized.
8 Awesome Books on Algorithms & Big Data
A 2013 roundup of eight books on algorithms and big data — a snapshot of the canonical reading list practitioners were recommending at the start of the data science hiring boom. Most of these titles held up as long-term references.
Streaming MapReduce with Summingbird
Twitter's Summingbird library unifies batch (Hadoop/Scalding) and streaming (Storm) MapReduce into a single Scala API — write once, run on either backend. An early practical implementation of the Lambda Architecture's dual-speed compute model.
Interactive Big Data Analysis Using Approximate Answers
O'Reilly Strata coverage of approximate query processing for interactive big data analysis — using sketches and sampling to get fast approximate answers over large datasets when exact computation is too slow. An important design pattern for data products that prioritize responsiveness over precision.
Jeremy Howard on the Big Data Obsession
Jeremy Howard's Quora answer on why the 'big data' obsession was somewhat misplaced — arguing that algorithms and predictive modeling matter more than raw data volume, and that the real value was in applying machine learning, not just collecting more data. A contrarian view from someone who knew ML deeply before the hype peaked.
Hadoop's Unsung Sweet Spot: Unstructured ETL
A LinkedIn discussion arguing that Hadoop's real sweet spot wasn't analytics but unstructured ETL — transforming messy, heterogeneous data into structured forms before loading into traditional data warehouses. A nuanced counterpoint to the 'Hadoop replaces SQL' narrative of 2013.
Nate Silver Gets Real About Big Data
ReadWrite's coverage of Nate Silver pushing back on big data hype — arguing that more data doesn't automatically improve predictions, and that statistical rigor and good models matter more than raw volume. Silver had credibility from his 2012 election forecasting success.
Why Data Virtualization Is Good for Big Data Analytics
Data-Informed's case for data virtualization in big data analytics — querying data in-place across Hadoop, relational databases, and other sources without physical ETL. A precursor to the 'data fabric' and 'data mesh' concepts that would emerge years later.
Big Data Ecosystem Map (2013)
A map of the big data ecosystem as it existed in 2013 — an attempt to catalog the dozens of Hadoop-adjacent tools, databases, analytics platforms, and services that had emerged. A useful historical artifact of the big data vendor explosion before consolidation.
Ex-Yahoo CEO Backs Genomics Big Data Startup Bina
FierceBiotechIT covering Bina Technologies, a genomics big data startup backed by ex-Yahoo CEO Scott Thompson, building hardware-accelerated pipelines for processing whole-genome sequencing data. A 2013 marker of when genomics data volumes began requiring big data infrastructure at clinical scale.
Introduction to Apache Kafka (TriHUG, July 2013)
TriHUG July 2013 talk introducing Apache Kafka — the distributed log system LinkedIn built and open-sourced. Caught at the moment Kafka was still an unfamiliar tool to most data engineers, before it became the de-facto streaming backbone of the modern data stack.
Implementing Hadoop for Big Data Projects
Inside Analysis piece on practical considerations for implementing Hadoop in enterprise big data projects circa 2013 — covering organizational readiness, hardware choices, and the gap between Hadoop's promise and production reality.
Twitter Visualizes Billions of Tweets in Interactive 3D Maps
The Verge's coverage of Twitter's interactive 3D tweet visualization showing billions of geotagged tweets rendered as a globe. An early example of browser-based 3D data visualization at social media scale using WebGL.
Don't MAWK AWK – the Fastest and Most Elegant Big Data Munging Language
Brendan O'Connor's defense of AWK as a fast, elegant big data munging language — arguing it beats Python for many common structured text processing tasks. A counterpoint to the idea that awk is an obsolete curiosity.
Facebook Unveils Presto for 250 PB Data Warehouse
GigaOm's coverage of Facebook unveiling Presto, their distributed SQL query engine for interactive queries against a 250 petabyte data warehouse. Presto addressed the core limitation of Hive — batch latency — by using a pipelined execution model that avoided writing intermediate results to disk.
14 Big Data Startups You're Going To Be Hearing About
Business Insider's 2013 list of 14 big data startups to watch — a market snapshot of the venture-backed companies building the infrastructure, analytics, and tooling layers of the emerging big data stack. A time capsule of which companies the industry thought would matter.
Idiot's Guide to Big Data
Mediasmiths' accessible overview of big data concepts for non-technical audiences, covering the 3 Vs (volume, velocity, variety), typical use cases, and the tools landscape circa 2013. A period document capturing how 'big data' was being explained to business decision-makers.
Hadoop 2.0 & YARN: The Big Data Breakthrough
ReadWrite's accessible overview of Hadoop 2.0 and YARN, explaining why the resource manager redesign was a bigger deal than an incremental release — it turned Hadoop from a MapReduce platform into a general-purpose cluster resource manager.
Hadoop on VMware: Another Workload Conquered?
EMC's Chuck Hollis on running Hadoop workloads on VMware virtualization — the argument that bare-metal Hadoop deployments could be replaced by virtualized clusters with manageable performance trade-offs. A 2013 salvo in the bare-metal vs. virtualization debate for big data workloads.
The Buzz and Fuzz on SSD in Hadoop
Hadoopsphere's analysis of SSD adoption in Hadoop clusters — the trade-off between SSD's dramatically faster random I/O and its higher cost per GB compared to spinning disk. An early look at how flash storage would eventually reshape big data infrastructure.
Nobody Ever Got Fired for Using Hadoop on a Cluster
Steve Loughran's riff on the classic 'nobody got fired for buying IBM' adapted for the big data era: Hadoop became the safe enterprise choice for data infrastructure, which brought both legitimacy and the mediocrity that comes with default choices.
The Value of Big Data Isn't the Data
HBR argument that the value of big data comes from the questions you ask and the actions you take, not from data accumulation itself. An early and important counterpoint to the data-as-moat hype of 2013.
Google's Self-Driving Car Generates ~1GB Per Second
Bill Gross's 2013 tweet with an infographic showing Google's self-driving car generating ~1GB of sensor data per second — a striking illustration of the gap between human-scale perception and machine-scale data requirements for autonomous driving.
MapReduce Patterns, Algorithms, and Use Cases
Ilya Katsov's comprehensive taxonomy of MapReduce design patterns — from basic counting and filtering through complex join strategies and graph algorithms. The field guide for wringing correct and efficient computation out of the MapReduce model.
Pundits: Stop Sounding Ignorant About Data
Andrew McAfee's HBR post calling out common misunderstandings in big data media coverage — conflating correlation with causation, treating anecdotes as data, and misusing statistical concepts. A 2013 data literacy argument aimed at commentators rather than practitioners.
Welcome to Berkeley: Where Hadoop Isn't Nearly Fast Enough
GigaOM's April 2013 profile of UC Berkeley's AMPLab making the case that Hadoop is too slow for interactive and iterative workloads — the academic origin story of Apache Spark, Shark (early Spark SQL), and Mesos. A snapshot of the moment Spark was about to go mainstream.
Future of Apache Hive — SQL PASS BA 2013
Carter Shanklin's 2013 talk on the future of Apache Hive — covering the push to make Hive's HiveQL a proper SQL dialect with ACID semantics, better query planning, and sub-second latency. A snapshot of the war between SQL-on-Hadoop and traditional data warehouses.
Natural Language Processing with Apache Hadoop and Python
Cloudera's 2010 post (bookmarked in 2013) on running natural language processing pipelines with Apache Hadoop and Python using Hadoop Streaming — an early recipe for scaling NLP beyond single-machine limits using commodity clusters.
Hadoop Illuminated — Free Open-Source Hadoop Book
Hadoop Illuminated is a free, open-source book on Apache Hadoop — a community-maintained guide covering HDFS, MapReduce, Hive, Pig, and the broader Hadoop ecosystem. One of the better free learning resources at a time when the Hadoop ecosystem was evolving faster than formal textbooks.
Automated Science, Deep Data, and the Paradox of Information
Oscillatory Thoughts blog post on the paradox of automated science and deep data — arguing that as data volumes grow and pattern detection automates, we risk producing empirical findings we can't explain or falsify. A philosophical challenge to data-driven science.
Doctors Use Big Data to Improve Cancer Treatments
Mashable on how doctors were using big data to improve cancer treatment outcomes in 2013 — early coverage of precision oncology using genomic data and clinical records. A snapshot of the medical establishment's first serious engagement with large-scale data-driven treatment.
Apache Incubator Giraph — Distributed Graph Processing on Hadoop
Apache Giraph is a graph processing framework built on Hadoop — an open-source implementation of Google's Pregel model for iterative graph algorithms at scale. The Apache answer to graph-scale problems like PageRank, community detection, and shortest paths on billion-node graphs.
Locality-Sensitive Hashing
Locality-sensitive hashing (LSH) is a family of algorithms for approximate nearest-neighbor search — hashing high-dimensional vectors so that similar items hash to the same bucket with high probability. The practical solution to similarity search at scale when exact methods are too slow.
Big Data and the Topologist
Low Dimensional Topology blog post on what topologists bring to big data — specifically topological data analysis (TDA) and the Mapper algorithm for finding shape in high-dimensional datasets. An unusual perspective on why geometric intuition matters for data analysis.
BioInformatics: A Data Deluge with Hadoop to the Rescue
Datanami on using Apache Hadoop for bioinformatics data pipelines — how genomic sequencing data had outpaced traditional computational biology infrastructure and why Hadoop's distributed file system and MapReduce were being adopted to handle the deluge.
Towards an Elastic Elephant: Enabling Hadoop for the Cloud
VMware's CTO blog on making Hadoop work in elastic cloud environments — the fundamental tension between Hadoop's static cluster model and cloud infrastructure's on-demand scaling. A 2013 take on a problem that took years to fully solve.
Analyzing Billions of Credit Card Transactions with Low-Latency Cloud Insights
High Scalability's case study of serving low-latency insights over billions of credit card transactions in the cloud. A concrete example of the architectural patterns needed when analytical query latency directly affects user experience in a financial context.
Working with Big Data: Assembling Your Toolkit
Acquia's guide to assembling a big data toolkit in 2013 — walking through the layered stack from storage to analysis. Practical survey of the Hadoop ecosystem components a team would actually need to evaluate.
CloudCamp / Hadoop
CloudCamp wiki page on Hadoop — a community-maintained reference during the peak of the Hadoop hype cycle in 2012-2013. One of many informal educational resources that proliferated as enterprise interest in Hadoop outpaced formal documentation.
Big Data Explained
BlazeClan's explainer on big data for a general technical audience — covering the 4 Vs framework and why traditional databases couldn't handle the scale. A typical 2013 entry-level primer during peak big data hype.
Walmart Makes Big Data Part of Its DNA
Walmart's 2013 integration of big data and social media analytics into retail operations — one of the first major brick-and-mortar retailers to make data infrastructure a competitive differentiator rather than just a reporting layer.
The Utilization Gap: Big Data's Biggest Challenge
Forbes on the gap between data collection capability and actual data use in 2013 — organizations were investing heavily in Hadoop and data warehouses while most of the collected data sat unanalyzed. The insight that technology was not the bottleneck; talent and culture were.
The Big Challenge of Big Data and Hadoop Integration
Cloud Computing Journal on the integration challenges of adding Hadoop to existing enterprise data infrastructure — connecting it to RDBMS systems, BI tools, and existing ETL pipelines. The 2013 reality check on big data adoption.
Probabilistic Data Structures for Web Analytics and Data Mining
The Highly Scalable Blog's comprehensive survey of probabilistic data structures for web analytics — Bloom filters, HyperLogLog, Count-Min sketch, and MinHash explained with their trade-offs. The standard reference for understanding when to trade exactness for speed and memory.
Introduction to Big Data and Hadoop Ecosystem
A beginner's introduction to the Apache Hadoop ecosystem cataloguing all major components — HDFS, MapReduce, Hive, HBase, Pig, Mahout, Sqoop, ZooKeeper, and more. A useful taxonomy of the 2012-era Hadoop stack.
Big Data: Hadoop Distributions Compared
Comparison of the main commercial Hadoop distributions in 2013 — Cloudera CDH, Hortonworks HDP, and MapR — when the market was consolidating around a handful of vendors each taking different bets on what enterprises needed.
Beyond Hadoop: Next-Generation Big Data Architectures
GigaOM's 2010 survey of Google's next-generation big data architectures — MPI, Pregel, Dremel, and Percolator — that were being developed as alternatives or complements to MapReduce. The paper that first systematically articulated why MapReduce wasn't sufficient for everything.
Parquet: Columnar Storage for Hadoop
The original announcement of Apache Parquet, the columnar storage format Twitter and Cloudera jointly released for Hadoop in 2013. Parquet became the dominant format for analytical workloads in the Hadoop/Spark ecosystem and remains ubiquitous in modern data lakes.
Analyzing Big Data with Twitter — UC Berkeley iSchool
UC Berkeley iSchool's 2012 course on analyzing big data with Twitter — an early academic offering that bridged social media data and distributed computing tooling before data science programs existed at most universities.
5 Reasons Why the Future of Hadoop Is Real-Time
GigaOM's 2013 argument for why Hadoop was evolving toward real-time processing — covering YARN, Storm, and in-memory frameworks as the drivers. A snapshot of the moment when batch-only Hadoop started feeling inadequate.
Cloudera's Support Team Shares Some Basic Hardware Recommendations
Cloudera's 2010 hardware recommendations for Hadoop clusters — still relevant in 2013 when bookmarked. The canonical guidance on disk, RAM, and CPU specs for commodity Hadoop nodes before cloud deployments became dominant.
Hortonworks Joins OpenStack Foundation
Hortonworks joining the OpenStack Foundation in early 2013 — a sign that the big data and cloud infrastructure worlds were converging. The tweet called 'virtualized Hadoop' the future, which captures exactly the direction that HDP-on-OpenStack was heading.
A Programmer's Guide to Big Data: 12 Tools to Know
GigaOM's 2012 reference guide to 12 big data tools a programmer should know — covering Hadoop, Pig, Hive, HBase, Storm, and emerging alternatives. A snapshot of the Hadoop ecosystem at its peak complexity, before Spark simplified much of it.
Blaze: A Python Compiler for Big Data
Continuum Analytics' announcement of Blaze — a Python compiler and array expression system designed to scale NumPy-style computations beyond in-memory datasets. An early attempt to bring Python's scientific computing ecosystem to big data before Spark/Dask became dominant.
UC Berkeley Course: Analyzing Big Data with Twitter
UC Berkeley's I290 course on analyzing big data with Twitter — lecture videos posted publicly covering Hadoop, NLP, streaming analytics, and social network analysis using Twitter's data firehose. One of the first openly-published university courses on the emerging data science field.
Shark: Real-time Queries and Analytics for Big Data
O'Reilly Strata article on Shark — the precursor to Spark SQL that brought real-time interactive queries to Hadoop/Spark in 2012. Part of the wave of tools (Impala, Shark, Drill) that challenged Hive's batch-query dominance.
HDFS Has Won: De Facto Standard for Centralized Data Storage
A 2012 claim that HDFS had emerged as the de facto standard for centralized big data storage — a snapshot of a moment when Hadoop's dominance seemed settled. Written just before the data lake era that HDFS would define, and before object storage (S3) would eventually displace it.
5 Trends That Are Changing How We Do Big Data
GigaOM's 2012 survey of five trends reshaping big data: real-time processing rising against batch, NoSQL maturity, cloud-based data platforms, open-source ecosystem growth, and the shift from data collection to data monetization. A useful snapshot of where the field was heading at Hadoop's peak.
Analyzing Human Genomes with Hadoop
Cloudera's 2009 blog post (bookmarked in 2012) showing how MapReduce and Hadoop can process human genome sequences at scale — an early example of big data infrastructure being applied to life sciences problems that were previously computationally intractable.
Facebook Seeks Next-Generation Big Data Tools
InformationWeek's 2012 report on Facebook's efforts to move beyond Hadoop for its next-generation big data infrastructure — one of the first major signals that the tech industry's heaviest Hadoop user was already looking past it.
New Tools Simplify the Power of Hadoop
InformationWeek's 2012 survey of tools making Hadoop more approachable for enterprise practitioners — the abstraction layers (Hive, Pig, HBase) and commercial distributions (Cloudera, Hortonworks) that were papering over Hadoop's low-level complexity. Captures the ecosystem's maturation moment.
Big Ideas: Demystifying Hadoop
A 'Big Ideas: Demystifying Hadoop' YouTube explainer from 2012 — one of many educational resources that emerged as Hadoop moved from niche to mainstream. Aimed at explaining the MapReduce paradigm and HDFS to practitioners who hadn't yet had to deal with data at scale.
Nutanix Complete Hadoop Appliance
Ryan's own SlideShare presentation for a Nutanix webinar on the Complete Hadoop Appliance — Nutanix's pitch for running Hadoop workloads on hyper-converged infrastructure instead of dedicated Hadoop clusters. This was Ryan's role at Nutanix circa 2012.
How Is Big Data Faring in the Enterprise?
A ZDNet survey of how big data technologies were actually landing in enterprise environments in mid-2012 — a useful reality check during peak Hadoop hype. Most large organizations were experimenting but few had production big data systems delivering measurable business value.
The Perfect Milk Machine: How Big Data Transformed the Dairy Industry
Alexis Madrigal's Atlantic piece on how the dairy industry used decades of genetic and performance data to engineer Holstein cows into radically more efficient milk producers. The best early example of big data optimization applied to a non-tech domain.
Cleversafe and Hadoop: Combining Next Generation Storage with Big Data Analytics
Cleversafe's 2012 integration with Hadoop connected dispersed object storage (using information dispersal algorithms rather than replication) to the Hadoop analytics ecosystem. Later acquired by IBM and became IBM Cloud Object Storage.
YARN: Hadoop NextGen MapReduce
YARN's official documentation from 2012 — the architecture that decoupled Hadoop cluster resource management from MapReduce. By separating resource negotiation into its own layer, YARN turned Hadoop from a MapReduce platform into a general-purpose cluster OS.
Why the Days Are Numbered for Hadoop As We Know It
A 2012 GigaOM piece arguing that Hadoop's architecture had fundamental limitations that would force it to evolve or be displaced — written at the peak of Hadoop hype. Prescient in identifying YARN and the multi-framework future, but underestimated how long it would take for cloud-native alternatives to win.
Twitter to Open Source Hadoop-Like Tool (Storm)
GigaOM coverage of Twitter's plans to open-source Storm — their real-time stream processing system. The moment the Hadoop-for-streaming gap became a major industry conversation, and Nathan Marz's Storm became the answer.
Large-Scale Graph Computing at Google (Pregel)
Google Research's 2009 blog post introducing Pregel — their internal system for large-scale graph computation using a bulk-synchronous-parallel model. The post that launched the graph processing systems category and eventually spawned Apache Giraph, GraphX, and the whole vertex-centric computing tradition.
Google Opens BigQuery Data Analytics to All
GigaOM coverage of Google opening BigQuery to all developers in 2012 — the productization of Dremel as a cloud service. The moment when interactive SQL over petabyte datasets became a commercial product rather than a Google-internal tool.
Dremel: Interactive Analysis of Web-Scale Datasets
Google's Dremel paper — the system that enabled sub-second SQL queries over petabyte datasets via columnar storage and a multi-level serving tree. The direct precursor to BigQuery, and the inspiration behind Apache Parquet's nested record encoding.
Apache Incubator Giraph
Apache Giraph's incubator homepage from 2012 — the open-source implementation of Google's Pregel bulk-synchronous-parallel graph processing model. Bookmarked during an early phase of the Hadoop ecosystem expansion into graph workloads.
Apache Hive
Apache Hive's homepage from 2012 — the SQL-on-Hadoop layer that made big data accessible to analysts who knew SQL but not Java MapReduce. Hive translated HiveQL queries into MapReduce jobs, trading latency for familiarity.
Apache HBase
Apache HBase's homepage from 2012 — the open-source implementation of Google Bigtable that added random-read/write access to Hadoop's otherwise write-once HDFS. HBase filled the gap MapReduce couldn't: low-latency lookups on data stored across a distributed cluster.
How Big Data Is Going to Change Entrepreneurship
Summary of the 2012 Stanford entrepreneurship conference on Big Data — panelists argued data was growing faster than Moore's Law, creating the next oil economy, with advertising and insurance as the most immediately impacted sectors. An early articulation of what became the data economy thesis.
