Sparse autoencoders aim to decompose dense activations into sparse combinations of more interpretable features.
Why It Matters
Feature extraction can provide building blocks for probing, monitoring, and targeted interventions.
Extract interpretable features from model activations using sparse representations.
Sparse autoencoders aim to decompose dense activations into sparse combinations of more interpretable features.
Feature extraction can provide building blocks for probing, monitoring, and targeted interventions.