Skip to main content
Ryan Orban

Ryan Orban

Subject
5 entries

Data Mining

Bookmarks

  1. KDD 2022 Keynote

    Keynote slides from the ACM SIGKDD 2022 conference on knowledge discovery and data mining. Content covers advances in large-scale ML, graph learning, or responsible AI — exact topic depends on which keynote this slide deck is from.

  2. Data Mining for Co-Location Patterns

    Guoqing Zhou's CRC Press book on co-location pattern mining, a spatial data mining problem concerned with finding features that frequently appear together in geographic proximity. A specialized but valuable topic for anyone working with geospatial data analysis.

  3. Git Data Mining with Hadoop

    WANdisco's post on mining git repository data at scale using Hadoop — applying distributed batch processing to version control history to extract patterns across large codebases. An early example of treating code evolution as a data science problem.

  4. Mining of Massive Datasets (Stanford)

    The Stanford textbook by Rajaraman and Ullman on algorithms for mining massive datasets — locality-sensitive hashing, PageRank, collaborative filtering, stream algorithms, and more. Freely available online and a standard reference for large-scale data algorithms.

  5. Probabilistic Data Structures for Web Analytics and Data Mining

    The Highly Scalable Blog's comprehensive survey of probabilistic data structures for web analytics — Bloom filters, HyperLogLog, Count-Min sketch, and MinHash explained with their trade-offs. The standard reference for understanding when to trade exactness for speed and memory.

All bookmarks