Skip to content
Web appLive

attnlab: an interactive playground for transformer interpretability

A set of labs, walked in order, that run real TransformerLens models behind a web page. Type a prompt, pick a model, and see how it is tokenized, where every head attends, and when the model knows its answer.

Published
Topics
  • Mechanistic interpretability
  • TransformerLens
  • Attention
  • Logit lens
  • Tokenizers

The labs

  1. Step 01Live

    Tokenizer lab

    What does the model actually read?

    Nine tokenizers, no model loaded: GPT-2, GPT-NeoX/Pythia, BLOOM, Qwen 2.5, Llama 2, GPT-4's cl100k, GPT-4o's o200k, BERT and XLM-R.

    • Inspect token chips as text, ids, bytes or raw vocabulary strings, with the merge that made each one
    • Compare one text across up to six tokenizers, stacked
    • The same FLORES+ sentences in 32 languages, with each tokenizer's premium over English (under GPT-2, Burmese costs ×16)
    • Step through real BPE merges one at a time, checked against the real tokenizer
    • A gallery of 16 quirks: leading spaces, digits, NFC/NFD, homoglyphs, glitch tokens, special-token injection
    Open the tokenizer lab
  2. Step 02Live

    Attention patterns

    Where does each token look?

    Per-layer, per-head attention heatmaps from a real HookedTransformer forward pass, rendered live as you type.

    • Read a head from the destination (a softmax row) or from the source (where attention lands)
    • Hover a token to link it across the token strip and the heatmap; click to pin
    • Devanagari and other non-Latin scripts shown as real characters, with byte-fragment provenance
    • An exact memory and FLOP cost panel for the current prompt
    • Permalinks: model, prompt, layer, head and direction all live in the URL
    Open the attention patterns
  3. Step 03Planned

    Induction heads

    How does a model copy what it has seen before?

    Repeated random tokens, per-head induction scores, and the loss drop in the second half of the sequence.

    • Scoring heads by their attention diagonal
    • Why the loss falls on repeated text
    • Composition between layers
  4. Step 04Live

    Logit lens

    When does the model know the answer?

    Decode the residual stream after the embedding and after every attention and MLP sub-layer, and watch the prediction form. One forward pass; every view is instant.

    • Colour by top-1 probability, P(actual next), rank, entropy or KL to the output
    • Two lenses side by side: the model's own ln_final and plain normalization, and why the second fails
    • Trajectory: one position through every layer, with rank, probability and logit-difference curves
    • Direct logit attribution: the output logit split exactly into embeddings, every head and every MLP (on IOI it finds the name movers L9H6 and L9H9 unprompted)
    • Five correctness checks recomputed on every run
    Open the logit lens
  5. Step 05Planned

    Ablation & attribution

    Which components actually matter?

    Zero- or mean-ablate heads and measure the change in loss and logits.

    • Zero vs mean ablation
    • Activation patching
    • Why direct attribution isn't causation

Why it exists

Learning mechanistic interpretability usually means a loop of opening a Colab notebook, waiting for a model to load, editing a cell and re-running it. attnlab removes the loop. It was inspired by ARENA 3.0 chapter 1.2 (Intro to Mech Interp) and pins the same TransformerLens version, so its numbers reproduce the notebooks exactly.

Each lab answers one question and hands its text and model to the next, so the labs read as one course rather than a pile of widgets.

Models

Small models, chosen so that every forward pass finishes in well under a second on a CPU. Adding a model is a data change, not a code change.

  • Attn-Only 2L (54M): ARENA's own induction-head demo model, with no MLPs
  • GPT-2 Small (163M, 12 layers × 12 heads): the ARENA default
  • Pythia 160M (12 × 12): EleutherAI, trained on the Pile
  • GPT-2 Medium (406M, 24 × 16): more composition structure
  • Qwen3 0.6B base (752M, 28 × 16): a 2025 architecture with RMSNorm, rotary embeddings and grouped-query attention, trained on 119 languages

How it works

A React and TypeScript front end draws everything on canvas. Behind it, a FastAPI server runs TransformerLens HookedTransformer models from a memory-budgeted model cache and refuses work rather than running out of memory.

Attention patterns travel in a compact binary format (triangle-packed, square-root companded, 8-bit) instead of raw float32. For gpt2-small at 512 tokens, raw attention would be 151 MB; nobody's browser needs that.

Questions

What is attnlab?
attnlab is a free, open-source web app for mechanistic interpretability. It runs TransformerLens models such as GPT-2, Pythia and Qwen3 on a server and lets you explore their tokenizers, attention patterns and logit lens from your browser.
Do I need a GPU or Python to use it?
No. Everything runs on the server; you only need a browser. If you want to run it yourself, the source is on GitHub and every model runs on two CPU cores.
Will the numbers match the ARENA notebooks?
Yes. attnlab pins the same TransformerLens version as ARENA 3.0, and the logit lens is tested to reproduce its reference notebook number for number.
Which tokenizers can I compare?
GPT-2, GPT-NeoX/Pythia, BLOOM, Qwen 2.5, Llama 2, GPT-4's cl100k, GPT-4o's o200k, BERT and XLM-R, across 32 languages.

Credits

attnlab stands on open-source work and published research. Thank you to everyone behind it.

Inspired by

  • ARENA 3.0, chapter 1.2: Intro to Mech Interp; attnlab turns its notebooks into labs and reproduces their numbers
  • CircuitsVis: the destination/source reading of attention patterns

Built with

Models and data

Background

More from Sarvabhaum

InteractiveLive

A transformer block, one matrix at a time

Follow four tokens through a single transformer block from the residual stream's point of view. Every matrix is shown with real numbers, computed live, so you can trace one token's row from embedding to next-token probabilities.

  • Transformers
  • Residual stream
  • Attention
NotebookLive

The logit lens in PyTorch

Watch GPT-2 build its prediction layer by layer. A runnable walkthrough that decodes the residual stream after every block, shows why ln_final matters, and plots top-k and rank heatmaps.

  • Interpretability
  • Logit lens
  • PyTorch