Skip to main content
Ryan Orban

Ryan Orban

Subject
2 entries

Anthropic

Bookmarks

  1. Simple probes can catch sleeper agents

    Anthropic research showing that simple linear probes trained on internal activations can reliably detect 'sleeper agent' backdoors in LLMs — models trained to behave well normally but activate malicious behavior on a trigger. Evidence that activation-space representations of deceptive intent are linearly separable.

  2. Anthropic (April 2022)

    Anthropic's job postings page saved in April 2022 — when the company was roughly one year old, pre-Claude, and actively hiring its founding team. A historical snapshot of Anthropic's early public presence before it became widely known.

All bookmarks