Subject
2 entries
Sparse Autoencoders
Bookmarks
Neuronpedia Gemma Scope: interactive mechanistic interpretability for Gemma 2
Neuronpedia's Gemma Scope microscope lets you scan Gemma 2's internal features using Sparse Autoencoders — activating features, steering behavior, browsing SAE decompositions. Interactive mechanistic interpretability for a production model.
ARENA Chapter 1.3.1: Toy Models of Superposition and SAEs
ARENA Chapter 1.3.1: an interactive curriculum chapter on Toy Models of Superposition and Sparse Autoencoders — building from the theoretical model to hands-on SAE implementation. Part of the ARENA AI safety education program.
