Skip to main content
Ryan Orban

Ryan Orban

Subject
2 entries

Probabilistic Data Structures

Bookmarks

  1. HyperLogLog in Pure SQL

    Periscope Data's post implementing HyperLogLog in pure SQL — a probabilistic cardinality estimator that counts distinct values using a fixed amount of memory regardless of dataset size. Clever engineering that demonstrates how probabilistic algorithms can be embedded in SQL-only environments.

  2. Probabilistic Data Structures for Web Analytics and Data Mining

    The Highly Scalable Blog's comprehensive survey of probabilistic data structures for web analytics — Bloom filters, HyperLogLog, Count-Min sketch, and MinHash explained with their trade-offs. The standard reference for understanding when to trade exactness for speed and memory.

All bookmarks