Skip to main content
Ryan Orban

Ryan Orban

Subject
6 entries

Datasets

Bookmarks

  1. ROOTS Search Tool — BigScience

    A Hugging Face Space for searching ROOTS — the massive multilingual dataset used to train BLOOM, the BigScience open LLM. Lets researchers trace which training documents a model might have learned from.

  2. PromptSource: An Integrated Development Environment and Repository for Natural Language Prompts

    PromptSource is an IDE and community repository for creating, sharing, and iterating on natural language prompts that map dataset examples to input-output pairs for language model training and evaluation. With over 2,000 prompts for ~170 datasets, it provided the infrastructure behind the T0 family of models and multitask prompted training research.

  3. Where Can I Find Large Datasets Open to the Public?

    Quora thread on where to find large public datasets — a community-curated reference from 2014 when open data sources were less centralized than today. The answers pointed to government portals, academic repositories, and early Kaggle.

  4. Where Can I Find Large Datasets Open to the Public?

    A 2013 Quora thread aggregating large public datasets (≥1 GB) for machine learning and data science research. A community-curated snapshot of the open data landscape before Kaggle, HuggingFace Datasets, and government open data portals became the primary discovery mechanisms.

  5. Some Datasets Available on the Web

    Data Wrangling Blog's curated list of publicly available datasets for machine learning and data analysis. An early community resource for finding training data before Kaggle and HuggingFace centralized dataset discovery.

  6. Many Downloadable Twitter Archives Available for Researchers

    DataScholars on downloadable Twitter archives available for academic researchers — a 2013 directory of public Twitter datasets for NLP, social network analysis, and computational social science. A snapshot of open Twitter data before the API became restrictive.

All bookmarks