Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Ml Ops

Bookmarks

  1. Beware of Unreliable Data in Model Evaluation

    Cleanlab's case study showing that noisy test data leads to suboptimal prompt selection for LLMs — you can choose the wrong prompt because your evaluation data contains labeling errors. A practical warning about data quality in LLM evaluation pipelines.

All bookmarks