SAE

SAE stitching

SAE LLM Mechanistic interpretability transformer architecture

Balancing novel and reconstruction latents by smoothly interpolating between differently-sized SAEs

Latent mechanistic interpretability

Mechanistic interpretability Benchmarking SAE

This project seeks to provide a benchmark for evaluating the extent to which a model is mechanistically interpretable

Inference-time decomposition of activations

SAE LLM Mechanistic interpretability transformer architecture activations inference

Scalable, cross-model alternative to SAEs for mechanistic interpretability

Meta SAE Dashboard

SAE LLM Mechanistic interpretability transformer architecture

An interactive dashboard of the meta-SAE decompositions

BatchTopK SAEs

SAE LLM Mechanistic interpretability transformer architecture

Sparse autoencoder training technique achieving better reconstruction, at the same sparsity, for less compute

Stitching Sparse Autoencoders of Different Sizes

SAE Stitching Stitching SAE Sparsity Autoencoders Latents Mechanistic Interpretability

Patrick Leask and Noura Al Moubayed introduce SAE stitching, a new method for mechanistic intepretability, in a poster at NeurIPS 2024.

BatchTopK Sparse Autoencoders

BatchTopK SAE Sparsity Autoencoders Mechanistic Interpretability Architecture

Patrick Leask contributes to BatchTopK, a new SAE architecture introduced in a NeurIPS'24 poster.