Subject
2 entries
Anthropic
Bookmarks
Simple probes can catch sleeper agents
Anthropic research showing that simple linear probes trained on internal activations can reliably detect 'sleeper agent' backdoors in LLMs — models trained to behave well normally but activate malicious behavior on a trigger. Evidence that activation-space representations of deceptive intent are linearly separable.
Anthropic (April 2022)
Anthropic's job postings page saved in April 2022 — when the company was roughly one year old, pre-Claude, and actively hiring its founding team. A historical snapshot of Anthropic's early public presence before it became widely known.
