Subject
17 entries
Open Data
Bookmarks
FiveThirtyEight Uber TLC FOIL Response Data
FiveThirtyEight's public release of 6 months of Uber NYC pickup data obtained via FOIL request from the NYC Taxi and Limousine Commission — one of the most significant data journalism transparency moments in ride-share's early history. Made possible by public records law applied to a tech company's regulatory filings.
Dat Project — Open Data Infrastructure
Wired profile of Max Ogden and the Dat Project — a decentralized, versioned data sharing protocol designed to make scientific and public datasets as easy to share and update as code on GitHub. An ambitious vision for open data infrastructure that was ahead of its time.
Where Can I Find Large Datasets Open to the Public?
Quora thread on where to find large public datasets — a community-curated reference from 2014 when open data sources were less centralized than today. The answers pointed to government portals, academic repositories, and early Kaggle.
Strava Global Heatmap
Strava's global heatmap aggregates GPS tracks from millions of runs and rides to show where athletes move around the world — a stunning geographic visualization that also accidentally revealed military base locations. Privacy implications aside, technically a landmark crowdsourced geospatial visualization.
DataDonors — Connecting Data Scientists with Nonprofits
DataDonors is a platform connecting data scientists with nonprofits that need analytical help — a talent-sharing model for applying quantitative skills to social good. Launched in the same 2014 wave as Bayes Impact and similar data-for-good initiatives.
The Rise of OpenStreetMap
The Next Web's 2014 profile of OpenStreetMap's rise as a challenger to Google Maps — covering how volunteer-driven mapping achieved coverage quality sufficient for Foursquare, Apple, and Mapbox to rely on it. A case study in open data displacing a proprietary incumbent.
Happy Healthy Hungry: San Francisco Data-Driven Narrative
Jay Oh-en's 'Happy Healthy Hungry' IPython notebook — a data-driven narrative about San Francisco restaurant health inspections shared at the Zipfian Academy graduation. One of the early examples of a published, storytelling-oriented data science notebook.
Defining Open Data
Open Knowledge Foundation's definition of open data — data that can be freely used, reused, and redistributed by anyone. Foundational framing for a movement that was gaining momentum in 2013 as governments began opening datasets.
SLEEP — Open Data Protocols
SLEEP (Seekable Lightweight Error-free Efficient Protocol) is an open binary data storage format with built-in indexing for random access and integrity verification. Created as part of the dat project ecosystem for distributing open datasets efficiently.
Git (and Github) for Data
Open Knowledge Foundation's 2013 post on using Git and GitHub for data versioning — exploring whether the version control model that worked for code could work for datasets. An early articulation of what became the data-as-code movement.
Where Can I Find Large Datasets Open to the Public?
A 2013 Quora thread aggregating large public datasets (≥1 GB) for machine learning and data science research. A community-curated snapshot of the open data landscape before Kaggle, HuggingFace Datasets, and government open data portals became the primary discovery mechanisms.
Some Datasets Available on the Web
Data Wrangling Blog's curated list of publicly available datasets for machine learning and data analysis. An early community resource for finding training data before Kaggle and HuggingFace centralized dataset discovery.
London Calling: Winning the Data Olympics
Mozilla OpenNews writeup on data journalism techniques used during the 2012 London Olympics, covering how journalists used open data and visualization to create compelling stories. An early example of the data journalism craft crystallizing around concrete, time-pressured work.
Crunching NYC Subway Data: A New Yorker's Busiest Stations
A data analysis post combining Node.js and SQL to parse MTA turnstile data and identify New York City's busiest subway stations. An early example of civic data journalism: public transit data as a lens on urban density and movement.
Quandl — Intelligent Search for Numerical Data
Quandl is a search engine for numerical data — financial series, economic indicators, and alternative datasets unified behind a single API. In 2013 it was the go-to free source for quantitative researchers who needed structured time-series data without Bloomberg Terminal access.
Cracking Oakland's Code
East Bay Express feature on Oakland's early open data and civic technology efforts — how local developers and city staff were beginning to use public data to address urban problems. An early document of the civic tech movement in the Bay Area.
Hackathon Aims to Make Oakland More Open
Code for Oakland 2012 brought together developers and city officials to build tools using open government data — part of the early Code for America ecosystem that helped establish civic tech as a legitimate software practice.
