Inference

Inference-time decomposition of activations

SAE LLM Mechanistic interpretability transformer architecture activations inference

Scalable, cross-model alternative to SAEs for mechanistic interpretability