Skip to content
Explainers

Explainers

Interactive, visual explainers of how transformers work: the residual stream, attention and the logit lens, with real numbers you can step through.

InteractiveLive

A transformer block, one matrix at a time

Follow four tokens through a single transformer block from the residual stream's point of view. Every matrix is shown with real numbers, computed live, so you can trace one token's row from embedding to next-token probabilities.

  • Transformers
  • Residual stream
  • Attention
NotebookLive

The logit lens in PyTorch

Watch GPT-2 build its prediction layer by layer. A runnable walkthrough that decodes the residual stream after every block, shows why ln_final matters, and plots top-k and rank heatmaps.

  • Interpretability
  • Logit lens
  • PyTorch