Subject
6 entries
Apache Spark
Bookmarks
Databricks Debuts Apache Spark MOOCs on BerkeleyX and edX
Databricks launches two MOOCs on Apache Spark through BerkeleyX and edX in December 2014, making distributed data processing education freely accessible. Marks the moment Spark was trying to become the standard for big data — before it succeeded.
Bayesian Machine Learning on Apache Spark
Cloudera's engineering blog post on implementing Bayesian machine learning on Apache Spark — combining probabilistic inference with distributed computation. A technically ambitious combination that was ahead of most production ML stacks in 2014.
Why Apache Spark is a Crossover Hit for Data Scientists
Cloudera's post on why Apache Spark resonated with data scientists in ways Hadoop MapReduce never did — the interactive REPL, Python support, and in-memory computation made it feel like a supercharged pandas rather than a distributed systems project.
Spark, Storm and Real-Time Analytics
An early comparison of Apache Storm and Spark Streaming for real-time analytics — covering their different processing models, latency characteristics, and use cases. A snapshot of the stream processing options available before Kafka Streams, Flink, and other tools matured.
How Companies Are Using Spark
Strata/O'Reilly coverage of how companies were adopting Apache Spark in 2013, early in the engine's rise to ubiquity. A snapshot of early enterprise Spark use cases before it displaced Hadoop MapReduce as the default.
Welcome to Berkeley: Where Hadoop Isn't Nearly Fast Enough
GigaOM's April 2013 profile of UC Berkeley's AMPLab making the case that Hadoop is too slow for interactive and iterative workloads — the academic origin story of Apache Spark, Shark (early Spark SQL), and Mesos. A snapshot of the moment Spark was about to go mainstream.
