SAE
SAE stitching
SAE LLM Mechanistic interpretability transformer architectureBalancing novel and reconstruction latents by smoothly interpolating between differently-sized SAEs
Latent mechanistic interpretability
Mechanistic interpretability Benchmarking SAEThis project seeks to provide a benchmark for evaluating the extent to which a model is mechanistically interpretable
Inference-time decomposition of activations
SAE LLM Mechanistic interpretability transformer architecture activations inferenceScalable, cross-model alternative to SAEs for mechanistic interpretability
Meta SAE Dashboard
SAE LLM Mechanistic interpretability transformer architectureAn interactive dashboard of the meta-SAE decompositions
BatchTopK SAEs
SAE LLM Mechanistic interpretability transformer architectureSparse autoencoder training technique achieving better reconstruction, at the same sparsity, for less compute
Stitching Sparse Autoencoders of Different Sizes
SAE Stitching Stitching SAE Sparsity Autoencoders Latents Mechanistic InterpretabilityPatrick Leask and Noura Al Moubayed introduce SAE stitching, a new method for mechanistic intepretability, in a poster at NeurIPS 2024.
BatchTopK Sparse Autoencoders
BatchTopK SAE Sparsity Autoencoders Mechanistic Interpretability ArchitecturePatrick Leask contributes to BatchTopK, a new SAE architecture introduced in a NeurIPS'24 poster.
