Skip to main content
Ryan Orban

Ryan Orban

Subject
2 entries

Sparse Autoencoders

Bookmarks

  1. Neuronpedia Gemma Scope: interactive mechanistic interpretability for Gemma 2

    Neuronpedia's Gemma Scope microscope lets you scan Gemma 2's internal features using Sparse Autoencoders — activating features, steering behavior, browsing SAE decompositions. Interactive mechanistic interpretability for a production model.

  2. ARENA Chapter 1.3.1: Toy Models of Superposition and SAEs

    ARENA Chapter 1.3.1: an interactive curriculum chapter on Toy Models of Superposition and Sparse Autoencoders — building from the theoretical model to hands-on SAE implementation. Part of the ARENA AI safety education program.

All bookmarks