Subject
3 entries
Hive
Bookmarks
Rolling Average in Hive
Brent Ozar's walkthrough of computing rolling averages in Hive — a problem that looks like a simple SQL query but requires window functions or self-joins in Hive's then-limited SQL dialect. A practical data engineering puzzle from the Hadoop era.
Hadoop Meets SQL
IBM Big Data Hub's overview of SQL-on-Hadoop approaches in 2013 — Hive, Impala, and the broader push to make Hadoop queryable by the vast majority of analysts who knew SQL but not MapReduce. The SQL interface became the primary adoption driver for Hadoop in the enterprise.
Future of Apache Hive — SQL PASS BA 2013
Carter Shanklin's 2013 talk on the future of Apache Hive — covering the push to make Hive's HiveQL a proper SQL dialect with ACID semantics, better query planning, and sub-second latency. A snapshot of the war between SQL-on-Hadoop and traditional data warehouses.
