Skip to main content
Ryan Orban

Ryan Orban

Subject
2 entries

Columnar Storage

Bookmarks

  1. Parquet: Columnar Storage for Hadoop

    The original announcement of Apache Parquet, the columnar storage format Twitter and Cloudera jointly released for Hadoop in 2013. Parquet became the dominant format for analytical workloads in the Hadoop/Spark ecosystem and remains ubiquitous in modern data lakes.

  2. Dremel: Interactive Analysis of Web-Scale Datasets

    Google's Dremel paper — the system that enabled sub-second SQL queries over petabyte datasets via columnar storage and a multi-level serving tree. The direct precursor to BigQuery, and the inspiration behind Apache Parquet's nested record encoding.

All bookmarks