LLM

SAE stitching

SAE LLM Mechanistic interpretability transformer architecture

Balancing novel and reconstruction latents by smoothly interpolating between differently-sized SAEs

Inference-time decomposition of activations

SAE LLM Mechanistic interpretability transformer architecture activations inference

Scalable, cross-model alternative to SAEs for mechanistic interpretability

Tangible LLMs: Exploring Tangible Sense-Making For Trustworthy Large Language Models

LLM Tangibility Participatory design Design Embodiment

Goda Klumbytė and Claude Draude co-organise a workshop on tangibility for understandability at TEI ‘25

The Resource Debate in Machine Translation and Large Language Models

LLM Translation Resource

Paolo Caffoni contributes to a reference entry on the language resource debate in LLMs to the Handbuch Soziale Praktiken und Digitale Alltagswelten

Meta SAE Dashboard

SAE LLM Mechanistic interpretability transformer architecture

An interactive dashboard of the meta-SAE decompositions

BatchTopK SAEs

SAE LLM Mechanistic interpretability transformer architecture

Sparse autoencoder training technique achieving better reconstruction, at the same sparsity, for less compute