Sparse Autoencoders & Feature Extraction

Extract interpretable features from model activations using sparse representations.

●●●●● Complexity

Sparse autoencoders aim to decompose dense activations into sparse combinations of more interpretable features.

Why It Matters

Feature extraction can provide building blocks for probing, monitoring, and targeted interventions.