Skip to main content
Ryan Orban

Ryan Orban

Subject
3 entries

Data Quality

Bookmarks

  1. Beware of Unreliable Data in Model Evaluation

    Cleanlab's case study showing that noisy test data leads to suboptimal prompt selection for LLMs — you can choose the wrong prompt because your evaluation data contains labeling errors. A practical warning about data quality in LLM evaluation pipelines.

  2. cleanlab 2.0: Automatically Find Errors in ML Datasets

    cleanlab 2.0 is an open-source Python framework for automatically finding and fixing errors in ML datasets — mislabeled examples, out-of-distribution samples, near-duplicates. Built on the 'confident learning' statistical framework for label noise estimation.

  3. Data Cleaning: Problems and Current Approaches (Berkeley/UNECE)

    Joe Hellerstein's Berkeley paper on data cleaning for UNECE — a systematic treatment of the data quality problem from a database research perspective. The academic foundation for what practitioners know as the most time-consuming part of data science work.

All bookmarks