Mechanistic interpretability

The best papers, talks, and explainers on mechanistic interpretability—reverse-engineering what neural networks actually compute.

Browse the full interactive library →

Verbalizable Representations Form a Global Workspace in Language Models

Wes Gurnee et al.

Anthropic's interpretability team identified a small, evolving buffer of concepts that a language model can report, hold, and reason with — functionally similar to a cognitive global workspace — giving researchers a new structural handle on what a model is silently processing before it says anything.

Advanced~2 hr read2026

Transformer Circuits

Anthropic / community

The home of mechanistic interpretability research, publishing detailed analyses of how transformer models represent and process information internally.

Advanced

Distill

Distill

Pioneering interactive journal for ML interpretability and visualization, setting the standard for making neural network internals understandable.

Advanced

Scaling Interpretability

Anthropic

Anthropic researchers explain mechanistic interpretability—reading the millions of concepts represented inside a production model like Claude—as a path to understanding and steering AI behavior.

Intermediate~55 min watch2024