LLM
SAE stitching
SAE LLM Mechanistic interpretability transformer architectureBalancing novel and reconstruction latents by smoothly interpolating between differently-sized SAEs
Inference-time decomposition of activations
SAE LLM Mechanistic interpretability transformer architecture activations inferenceScalable, cross-model alternative to SAEs for mechanistic interpretability
Tangible LLMs: Exploring Tangible Sense-Making For Trustworthy Large Language Models
LLM Tangibility Participatory design Design EmbodimentGoda Klumbytė and Claude Draude co-organise a workshop on tangibility for understandability at TEI ‘25
The Resource Debate in Machine Translation and Large Language Models
LLM Translation ResourcePaolo Caffoni contributes to a reference entry on the language resource debate in LLMs to the Handbuch Soziale Praktiken und Digitale Alltagswelten
Meta SAE Dashboard
SAE LLM Mechanistic interpretability transformer architectureAn interactive dashboard of the meta-SAE decompositions
BatchTopK SAEs
SAE LLM Mechanistic interpretability transformer architectureSparse autoencoder training technique achieving better reconstruction, at the same sparsity, for less compute
