Skip to main content
Ryan Orban

Ryan Orban

Subject
1 entry

Deception

Bookmarks

  1. Reducing LLM Deception with Self-Other Overlap Fine-Tuning

    Self-Other Overlap (SOO) fine-tuning reduces deceptive behavior in LLMs by aligning internal activations for self-referential and other-referential prompts — cutting deceptive responses from 73% to 17% on Mistral-7B with minimal capability loss. A representation-level approach rather than behavioral supervision.

All bookmarks