Subject
4 entries
Data Warehouse
Bookmarks
The Lazy Analyst's Guide to Amazon Redshift
Periscope Data's practical guide to Amazon Redshift — covering the distribution and sort key mechanics that determine query performance, plus common gotchas for analysts who know SQL but not columnar databases. Still one of the clearest explanations of why Redshift behaves differently from Postgres.
Building Analytics at 500px
A first-person account of building 500px's analytics infrastructure from scratch — Amazon Redshift data warehouse, Luigi ETL, Periscope BI. The 20% evangelism rule and 'don't bake your own ETL' lesson make it one of the most practical early data engineering retrospectives.
Hadoop Meets SQL
IBM Big Data Hub's overview of SQL-on-Hadoop approaches in 2013 — Hive, Impala, and the broader push to make Hadoop queryable by the vast majority of analysts who knew SQL but not MapReduce. The SQL interface became the primary adoption driver for Hadoop in the enterprise.
Apache Hive
Apache Hive's homepage from 2012 — the SQL-on-Hadoop layer that made big data accessible to analysts who knew SQL but not Java MapReduce. Hive translated HiveQL queries into MapReduce jobs, trading latency for familiarity.
