San Francisco
[email protected]
github.com/orban
linkedin.com/in/ryanorban
Fifteen years of systems graded from the outside: whether the graduate got hired, whether the cluster held, whether the agent passes twice.
The record
now
Agent evaluation
Independent · open source
An agent that passes once has not passed.
Cadea
Founder
Secure workspace chatbots: RBAC-aware retrieval and agentic pipelines over regulated data. The prototypes worked. The buyers were about two years out. I returned the remaining capital and wrote up what I’d learned about red-teaming, data governance, and evals instead of raising again.
Tribe AI
CTO
Product and infrastructure for a network of 150+ senior ML practitioners. Made scoping and delivery repeatable instead of re-invented per engagement. Ran tech lead duty on client work: LLM prototyping, recommenders, data platforms.
Placement.com
Head of Data Science
Ranking, matching, and funnel analytics for a job marketplace: Elasticsearch with Learning to Rank, BERT embeddings, Bayesian A/B tests. Every one of them graded on a single outcome — whether the person got hired.
Away
Sabbatical
Three years overland through North, Central, and South America.
Zipfian → Galvanize
Founder, then CTO
Bootstrapped the first immersive data science program in the US: ten people, $1M+ of revenue in year one, 91% placement, graduates into Tesla, Facebook, and Google. Galvanize acquired it, and I stayed on as CTO to run engineering, curriculum, and platform across campuses.
Nutanix
Sr. Systems Engineer
Distributed systems first, then the field. Designed and deployed $100M+ of virtualization clusters for Army, Air Force, and intelligence customers — rooms where “it usually works” is not an answer anyone accepts.
Building now
Python
Moirai
Evaluation research
Run one agent on one task enough times and it branches. Moirai lines those runs up, finds the branch points, tests which choices actually predicted success, and exports the survivors as preference data. Run against SWE-rebench:
- Runs analyzed
- 12,854
- Mixed-outcome tasks
- 1,096
- Preference pairs
- 11,006
TypeScript
Cerberus
Statistical CI
A green test run on a stochastic system is one sample, not a verdict. Cerberus treats every invocation as a Bernoulli trial, runs a sequential probability ratio test to decide when it has seen enough, and reports a Wilson interval against your release bar — as an exit code your CI already understands.
Other projects
Solar
A solar-powered server in San Francisco hosts sequoia.garden and reports on itself: state of charge by coulomb counting off an INA228 shunt, plus voltage, load average, and temperature. I did the solar install, the current-sensing circuit, and the hardened Raspberry Pi it runs on, and built the software stack with Camille Teicheira — after Low-tech Magazine's solar-powered site. Sequoia Fabrica is a volunteer-run community workshop.
Scraper
Bay Area event discovery: several sources scraped into one searchable calendar. A static snapshot ships with this site.
Web
Strangers name public playlists after feelings. This plays them back anonymously — 725 stations curated down from 13,190 distinct names across 22,000 results, with the overlaps counted where several people independently landed on the same name and the same songs.
Map
Every SF 311 complaint, ranked and mapped as though the city's worst-reviewed corners were attractions.
Sculpture
An interactive light sculpture — 14 by 9 by 9 feet of stainless steel, acrylic, and 3,500 LEDs that visitors drive from buttons inside it. Built with Taylor Dean Harrison and Henry Richardson; Burning Man Honoraria grant, 2016.
What I’m looking for
A team with a system that matters and a measurement problem they can’t keep ignoring: agent evaluation, observability, memory, and the infrastructure that makes any of it legible. Or an argument about any of the above.
The most useful first message is the thing you currently can’t measure.
