Subject
1 entry
Deception
Bookmarks
Reducing LLM Deception with Self-Other Overlap Fine-Tuning
Self-Other Overlap (SOO) fine-tuning reduces deceptive behavior in LLMs by aligning internal activations for self-referential and other-referential prompts — cutting deceptive responses from 73% to 17% on Mistral-7B with minimal capability loss. A representation-level approach rather than behavioral supervision.
