Skip to content
InteractiveLive

A transformer block, one matrix at a time

Follow four tokens through a single transformer block from the residual stream's point of view. Every matrix is shown with real numbers, computed live, so you can trace one token's row from embedding to next-token probabilities.

Published
Topics
  • Transformers
  • Residual stream
  • Attention
  • Interactive

The ten steps

The residual stream (d_model numbers per token) is the one thing every part of the block reads from and writes to. The walkthrough moves through the block in the order the computation happens:

  • Input: the residual stream begins
  • LayerNorm: normalize before attention
  • Project into queries, keys and values
  • Split d_model into heads
  • Attention: who looks at whom
  • Mix values, merge heads, project back
  • Add back into the stream
  • Normalize, then widen and squeeze (the MLP)
  • Add again: the output leaves the block
  • Unembed: from hidden state to next-token probabilities

How to read it

Every row is one token. Follow a row across every matrix to see what happens to that token, and step back and forth to watch each operation change the stream. The model is kept tiny (d_model = 8) so that every number fits on screen.

It pairs with attnlab's attention lab: once you know what a head computes here, the heatmaps there show what real heads in GPT-2 do with it.

More from Sarvabhaum

Web appLive

attnlab

A set of labs, walked in order, that run real TransformerLens models behind a web page. Type a prompt, pick a model, and see how it is tokenized, where every head attends, and when the model knows its answer.

  • Mechanistic interpretability
  • TransformerLens
  • Attention
NotebookLive

The logit lens in PyTorch

Watch GPT-2 build its prediction layer by layer. A runnable walkthrough that decodes the residual stream after every block, shows why ln_final matters, and plots top-k and rank heatmaps.

  • Interpretability
  • Logit lens
  • PyTorch