Subject
2 entries
Probabilistic Data Structures
Bookmarks
HyperLogLog in Pure SQL
Periscope Data's post implementing HyperLogLog in pure SQL — a probabilistic cardinality estimator that counts distinct values using a fixed amount of memory regardless of dataset size. Clever engineering that demonstrates how probabilistic algorithms can be embedded in SQL-only environments.
Probabilistic Data Structures for Web Analytics and Data Mining
The Highly Scalable Blog's comprehensive survey of probabilistic data structures for web analytics — Bloom filters, HyperLogLog, Count-Min sketch, and MinHash explained with their trade-offs. The standard reference for understanding when to trade exactness for speed and memory.
