Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Sleeper Agents

Bookmarks

  1. Simple probes can catch sleeper agents

    Anthropic research showing that simple linear probes trained on internal activations can reliably detect 'sleeper agent' backdoors in LLMs — models trained to behave well normally but activate malicious behavior on a trigger. Evidence that activation-space representations of deceptive intent are linearly separable.

All bookmarks