Subject
1 entry
Ml Ops
Bookmarks
Beware of Unreliable Data in Model Evaluation
Cleanlab's case study showing that noisy test data leads to suboptimal prompt selection for LLMs — you can choose the wrong prompt because your evaluation data contains labeling errors. A practical warning about data quality in LLM evaluation pipelines.
