Skip to main content
Ryan Orban

Ryan Orban

Subject
6 entries

Cloudera

Bookmarks

  1. Cloudera Rebuilding Machine Learning for Hadoop with Oryx

    GigaOm's coverage of Cloudera launching Oryx — an open-source ML-on-Hadoop framework using the Lambda Architecture for batch retraining plus real-time serving. An early attempt to make production machine learning first-class on Hadoop.

  2. Cloudera Developer Class Links

    Cloudera's developer training class resource links page — a collection of documentation, tutorials, and reference materials for the Cloudera Hadoop Developer certification course. Saved as a training reference during the 2013 era of Hadoop skills development.

  3. Big Data: Hadoop Distributions Compared

    Comparison of the main commercial Hadoop distributions in 2013 — Cloudera CDH, Hortonworks HDP, and MapR — when the market was consolidating around a handful of vendors each taking different bets on what enterprises needed.

  4. Proprietary Hadoop Is a Losing Strategy

    ReadWrite's 2013 argument that commercial Hadoop vendors who added proprietary lock-in would lose to those who contributed everything upstream — a prescient thesis that proved partly right, partly wrong over the following decade.

  5. Cloudera's Support Team Shares Some Basic Hardware Recommendations

    Cloudera's 2010 hardware recommendations for Hadoop clusters — still relevant in 2013 when bookmarked. The canonical guidance on disk, RAM, and CPU specs for commodity Hadoop nodes before cloud deployments became dominant.

  6. Analyzing Human Genomes with Hadoop

    Cloudera's 2009 blog post (bookmarked in 2012) showing how MapReduce and Hadoop can process human genome sequences at scale — an early example of big data infrastructure being applied to life sciences problems that were previously computationally intractable.

All bookmarks