Skip to content
InteractiveLive

One transformer block, neuron by neuron

A tiny real transformer (d = 8, 2 heads, an 8 → 32 → 8 MLP, 4 blocks) runs on the sentence you type. Every circle holds the number it actually computed, and hovering one shows the weighted lines that feed it.

Published
Topics
  • Transformers
  • Neural networks
  • Attention
  • MLP
  • Interactive

A transformer drawn as a neural network

Most transformer diagrams are boxes and arrows. This one goes back to the picture everyone learns first: circles for neurons and lines for weights. Each step shows two views side by side. The neuron view draws one token's vector flowing through a layer, and the matrix view shows the same computation for the whole sentence at once, as X @ W.

Three ideas carry the walkthrough. Every linear layer in the block (Q, K, V, O and both MLP layers) is a small fully connected network applied to each token separately, with the same weights. Attention is the only place where tokens exchange information. And the residual stream is each token's running vector: sublayers read from it and add their result back, never overwriting it.

The sixteen steps

  • Neuron view and matrix view: one vector per token, stacked into X
  • Tokens become vectors: the embedding lookup and sine-cosine positions
  • The residual stream as parallel highways, one lane per token
  • LayerNorm: rescale a token's vector before reading it
  • Three linear layers: queries, keys and values
  • Split into heads
  • Scores, causal mask and softmax: which earlier tokens to listen to
  • Mix the values with the attention weights
  • Concatenate the heads and project with W_O
  • Add back to the highway
  • MLP, part 1: widen 8 → 32 and apply GELU
  • MLP, part 2: project back to 8
  • Add again: the block's output
  • Stacking blocks
  • From the stream to a next-token prediction
  • How it is done at scale

What you can do with it

  • Type your own sentence; the model is rebuilt and every number recomputed in your browser
  • Pick a focus token, one of the 4 blocks and one of the 2 attention heads
  • Hover any neuron to see its incoming weights and the dot product that produced its value
  • Click any part of the block map to jump straight to that step
  • Re-roll the random weights and watch the same structure produce different numbers

From the toy to a real model

Each step ends with a note on what changes at scale: fused QKV projections, grouped-query attention, FlashAttention, the KV cache, RMSNorm, SwiGLU and RoPE. The last step lists every operation in the block with its tensor shapes, and has a calculator for parameters, training FLOPs per token and attention-matrix memory for any width, depth, head count, vocabulary and context length.

It pairs with the matrix-by-matrix residual stream explainer, which follows the same block with the focus on shapes rather than neurons, and with attnlab, where you can see what real attention heads in GPT-2 do.

Questions

Is a transformer a neural network?
Yes. Every linear layer in a transformer block is an ordinary fully connected layer, and the MLP is a classic two-layer feed-forward network. What is new is attention, where the connection strengths between tokens are computed from the input instead of being fixed learned weights.
Is the model in the explainer real?
Yes, but tiny and untrained. It has 8-dimensional vectors, 2 attention heads, an 8 → 32 → 8 MLP and 4 blocks, with random weights. Every value is genuinely computed from your sentence; only the predictions are meaningless, because nothing has been trained.
What does the MLP in a transformer do?
It runs on each token separately: it widens the vector to four times its size, applies GELU, and projects it back. Each hidden neuron detects an input pattern (its column of W_in) and writes an output vector when active (its row of W_out). It holds about two thirds of each block's parameters.
Why does a transformer use LayerNorm?
So each sublayer sees inputs of a stable size however large the residual stream grows. LayerNorm subtracts a token's mean and divides by its standard deviation across that token's own numbers, then applies a learned scale and shift. Many recent models use RMSNorm, which skips the mean.
How many parameters does one transformer block have?
About 12·d² for a model of width d with a 4× MLP: 4d² in attention (Q, K, V and O) and 8d² in the MLP. The explainer's calculator works this out, plus training FLOPs per token, for any model size you enter.

Credits

One transformer block, neuron by neuron stands on open-source work and published research. Thank you to everyone behind it.

Background

More from Sarvabhaum

Web appLive

attnlab

A set of labs, walked in order, that run real TransformerLens models behind a web page. Type a prompt, pick a model, and see how it is tokenized, where every head attends, and when the model knows its answer.

  • Mechanistic interpretability
  • TransformerLens
  • Attention
InteractiveLive

Training a transformer on eight GPUs

One concrete configuration, followed from start to finish: a 1.27B-parameter model with Llama 2 7B's layer shape, trained on a single 8-GPU node split 2-way tensor × 2-way pipeline × 2-way data parallel, then post-trained and deployed.

  • Distributed training
  • Parallelism
  • LLM training
InteractiveLive

A transformer block, one matrix at a time

Follow four tokens through a single transformer block from the residual stream's point of view. Every matrix is shown with real numbers, computed live, so you can trace one token's row from embedding to next-token probabilities.

  • Transformers
  • Residual stream
  • Attention