Sparse Autoencoder Training

Train and analyze Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features for mechanistic interpretability research.

Orchestra Research Updated 10.4k repo stars

File contents

Orchestra-Research/AI-Research-SKILLs/tree/main/04-mechanistic-interpretability/saelens commit 773a52944b

Frequently asked questions

npx skillmds@latest add orchestra-research/sparse-autoencoder-training