Skip to content
Interpretability, in the browser

See inside
language models.

Open tools and interactive explainers for mechanistic interpretability. Type a prompt, pick a real model, and look at what every token, attention head and layer is doing, without opening a notebook.

Free · Open source · Real models, real numbers

Attention patterndestination → source
<bos>·Mr·Dursley·was·Mr·Dursley<bos>·Mr·Dursley·was·Mr·Dursley
An induction head, illustrated: on the second Mr Dursley, each token looks back at what came after it last time. attnlab draws these for every head of a real model.
Tools

Run real models. Look inside.

Free, open-source interpretability tools that run real language models behind a web page: tokenizers, attention patterns and the logit lens.

Web appLive

attnlab

A set of labs, walked in order, that run real TransformerLens models behind a web page. Type a prompt, pick a model, and see how it is tokenized, where every head attends, and when the model knows its answer.

  1. 01Tokenizer labWhat does the model actually read?Live
  2. 02Attention patternsWhere does each token look?Live
  3. 03Induction headsHow does a model copy what it has seen before?Planned
  4. 04Logit lensWhen does the model know the answer?Live
  5. 05Ablation & attributionWhich components actually matter?Planned
Explainers

How transformers work, one step at a time.

Interactive, visual explainers of how transformers work: the residual stream, attention and the logit lens, with real numbers you can step through.

All explainers
InteractiveLive

A transformer block, one matrix at a time

Follow four tokens through a single transformer block from the residual stream's point of view. Every matrix is shown with real numbers, computed live, so you can trace one token's row from embedding to next-token probabilities.

  • Transformers
  • Residual stream
  • Attention
NotebookLive

The logit lens in PyTorch

Watch GPT-2 build its prediction layer by layer. A runnable walkthrough that decodes the residual stream after every block, shows why ln_final matters, and plots top-k and rank heatmaps.

  • Interpretability
  • Logit lens
  • PyTorch

More writing on transformers and NLP lives on Nirajan Paudel’s site.

In progress

What's next

Planned labs, built in the open. Each one answers a single question and links to the step before it.

  • attnlab

    Induction heads

    How does a model copy what it has seen before?

    Repeated random tokens, per-head induction scores, and the loss drop in the second half of the sequence.

  • attnlab

    Ablation & attribution

    Which components actually matter?

    Zero- or mean-ablate heads and measure the change in loss and logits.

About

Sarvabhaum AI

Sarvabhaum AI makes the inside of language models something you can look at. The tools run small, real models such as GPT-2, Pythia and Qwen3 on a server we host, and the explainers show every matrix with real numbers, so each idea can be checked rather than taken on trust.

Everything here is free and open source. New labs and explainers are added as they are finished.

It stands on other people’s open work, above all TransformerLens and the ARENA curriculum. Every page lists the libraries, models, data and papers it builds on.