Skip to main content
Ryan Orban

Ryan Orban

San Francisco
[email protected]
github.com/orban
linkedin.com/in/ryanorban

Fifteen years of systems graded from the outside: whether the graduate got hired, whether the cluster held, whether the agent passes twice.

The record

  1. now

    Agent evaluation

    Independent · open source

    An agent that passes once has not passed.

  2. Cadea

    Founder

    Secure workspace chatbots: RBAC-aware retrieval and agentic pipelines over regulated data. The prototypes worked. The buyers were about two years out. I returned the remaining capital and wrote up what I’d learned about red-teaming, data governance, and evals instead of raising again.

  3. Tribe AI

    CTO

    Product and infrastructure for a network of 150+ senior ML practitioners. Made scoping and delivery repeatable instead of re-invented per engagement. Ran tech lead duty on client work: LLM prototyping, recommenders, data platforms.

  4. Placement.com

    Head of Data Science

    Ranking, matching, and funnel analytics for a job marketplace: Elasticsearch with Learning to Rank, BERT embeddings, Bayesian A/B tests. Every one of them graded on a single outcome — whether the person got hired.

  5. Away

    Sabbatical

    Three years overland through North, Central, and South America.

  6. Zipfian → Galvanize

    Founder, then CTO

    Bootstrapped the first immersive data science program in the US: ten people, $1M+ of revenue in year one, 91% placement, graduates into Tesla, Facebook, and Google. Galvanize acquired it, and I stayed on as CTO to run engineering, curriculum, and platform across campuses.

  7. Nutanix

    Sr. Systems Engineer

    Distributed systems first, then the field. Designed and deployed $100M+ of virtualization clusters for Army, Air Force, and intelligence customers — rooms where “it usually works” is not an answer anyone accepts.

Building now

  1. Python

    Moirai

    Evaluation research

    Run one agent on one task enough times and it branches. Moirai lines those runs up, finds the branch points, tests which choices actually predicted success, and exports the survivors as preference data. Run against SWE-rebench:

    Runs analyzed
    12,854
    Mixed-outcome tasks
    1,096
    Preference pairs
    11,006
  2. TypeScript

    Cerberus

    Statistical CI

    A green test run on a stochastic system is one sample, not a verdict. Cerberus treats every invocation as a Bernoulli trial, runs a sequential probability ratio test to decide when it has seen enough, and reports a Wilson interval against your release bar — as an exit code your CI already understands.

Other projects

  1. Solar

    A solar-powered server in San Francisco hosts sequoia.garden and reports on itself: state of charge by coulomb counting off an INA228 shunt, plus voltage, load average, and temperature. I did the solar install, the current-sensing circuit, and the hardened Raspberry Pi it runs on, and built the software stack with Camille Teicheira — after Low-tech Magazine's solar-powered site. Sequoia Fabrica is a volunteer-run community workshop.

  2. Scraper

    Bay Area event discovery: several sources scraped into one searchable calendar. A static snapshot ships with this site.

  3. Web

    Strangers name public playlists after feelings. This plays them back anonymously — 725 stations curated down from 13,190 distinct names across 22,000 results, with the overlaps counted where several people independently landed on the same name and the same songs.

  4. Map

    Every SF 311 complaint, ranked and mapped as though the city's worst-reviewed corners were attractions.

  5. Sculpture

    An interactive light sculpture — 14 by 9 by 9 feet of stainless steel, acrylic, and 3,500 LEDs that visitors drive from buttons inside it. Built with Taylor Dean Harrison and Henry Richardson; Burning Man Honoraria grant, 2016.

What I’m looking for

Email

A team with a system that matters and a measurement problem they can’t keep ignoring: agent evaluation, observability, memory, and the infrastructure that makes any of it legible. Or an argument about any of the above.

The most useful first message is the thing you currently can’t measure.