Subject
2 entries
Columnar Storage
Bookmarks
Parquet: Columnar Storage for Hadoop
The original announcement of Apache Parquet, the columnar storage format Twitter and Cloudera jointly released for Hadoop in 2013. Parquet became the dominant format for analytical workloads in the Hadoop/Spark ecosystem and remains ubiquitous in modern data lakes.
Dremel: Interactive Analysis of Web-Scale Datasets
Google's Dremel paper — the system that enabled sub-second SQL queries over petabyte datasets via columnar storage and a multi-level serving tree. The direct precursor to BigQuery, and the inspiration behind Apache Parquet's nested record encoding.
