Subject
5 entries
Apache
Bookmarks
Apache Incubator Giraph — Distributed Graph Processing on Hadoop
Apache Giraph is a graph processing framework built on Hadoop — an open-source implementation of Google's Pregel model for iterative graph algorithms at scale. The Apache answer to graph-scale problems like PageRank, community detection, and shortest paths on billion-node graphs.
YARN: Yet Another Resource Negotiator
Apache YARN (Yet Another Resource Negotiator) is the cluster resource management layer introduced in Hadoop 2 — the architectural change that turned Hadoop from a MapReduce system into a general-purpose distributed compute platform.
Apache Incubator Giraph
Apache Giraph's incubator homepage from 2012 — the open-source implementation of Google's Pregel bulk-synchronous-parallel graph processing model. Bookmarked during an early phase of the Hadoop ecosystem expansion into graph workloads.
Apache Hive
Apache Hive's homepage from 2012 — the SQL-on-Hadoop layer that made big data accessible to analysts who knew SQL but not Java MapReduce. Hive translated HiveQL queries into MapReduce jobs, trading latency for familiarity.
Apache HBase
Apache HBase's homepage from 2012 — the open-source implementation of Google Bigtable that added random-read/write access to Hadoop's otherwise write-once HDFS. HBase filled the gap MapReduce couldn't: low-latency lookups on data stored across a distributed cluster.
