Skip to content
Web appLive

Speculative Decoding Playground: watch a small model draft and a large one verify

A small draft model guesses a few tokens, a large target model checks them all in one forward pass, and the text comes out exactly as the large model alone would write it. The playground runs that loop on real models and lets you inspect every round.

Published
Topics
  • Inference
  • Speculative decoding
  • Sampling
  • KV cache

What's in it

  1. Step 01Live

    Run any prompt

    Which tokens did the draft get right?

    Run speculative decoding on your own prompt with real models, in sampling or greedy mode, and see the output colored by where each token came from.

    • Accepted draft tokens, corrections from the target and bonus tokens, each marked
    • γ (draft length), temperature, top-k, top-p and repetition penalty, applied identically to both models
    • Recorded runs, including a larger Qwen3 1.7B / 0.6B pair, that work without the model server
    Open the run any prompt
  2. Step 02Live

    Round inspector

    What happened inside one round?

    Every round broken into five steps, draft, verify, accept or reject, correction or bonus, and commit with cache rollback, with live numbers and the lines of code that ran highlighted.

    • Both models' probability distributions at every drafted position
    • The acceptance test p/q, the random draw that decided it, and the residual distribution a correction is sampled from
    • Tensor shapes and how far each KV cache is rolled back
    Open the round inspector
  3. Step 03Live

    Speed

    Was it actually faster?

    Where the time went in each run, checked against the target model decoding on its own with the same seed.

    • Draft and target forward-pass timings per round
    • A side-by-side baseline run of the target alone
    • An honest answer when the small pair is not faster, and why
    Open the speed
  4. Step 04Live

    Two labs

    Why is it exact, and when does it pay off?

    Drag two toy distributions to see why accepting with probability p/q reproduces the target exactly, then explore the speedup as a function of acceptance rate, cost ratio and γ.

    • Lab 1: rejection sampling on a five-token vocabulary, simulated
    • Lab 2: expected tokens per round and wall-clock speedup, from Leviathan et al. (2023)
    • Pitfalls: the mistakes that keep the output looking fine but stop it being the target's
    Open the two labs

Why it exists

Speculative decoding is easy to describe and easy to get subtly wrong. A bug in the acceptance test, the residual distribution or the cache rollback still produces fluent text; it just isn't the large model's text any more. Reading the papers rarely shows where those mistakes hide.

The playground runs a complete reference implementation of about a hundred lines, shows that code next to every step, and is tested to emit exactly the same tokens as the reference and, in greedy mode, exactly the target model's own output.

How it works

A Python server loads a draft and a target model that share a tokenizer, runs the decoding loop with read-only snapshots of every intermediate value, and sends the whole trace to a plain HTML and JavaScript page with no build step.

The live server runs SmolLM2-360M as the target and SmolLM2-135M as the draft. At this size every forward pass costs about the same fixed overhead, so the draft is not much cheaper than the target and there is little speedup; the page measures and explains this. The gain appears when the target is large and memory-bound, which the second lab models.

Questions

What is speculative decoding?
A way to make a large language model generate text in fewer slow steps. A small draft model proposes several tokens, the large model scores all of them in one forward pass, and the tokens it agrees with are kept. With the right acceptance rule the output has exactly the large model's distribution.
Is the output really identical to the large model's?
In greedy mode it is token for token the same, and the playground's tests check this. In sampling mode it follows exactly the same probability distribution: each draft token is accepted with probability p/q, and after a rejection the replacement is drawn from the normalized difference between the two distributions.
Why isn't the small pair faster?
With models of a few hundred million parameters, a forward pass is dominated by fixed overhead rather than by reading weights, so the draft costs nearly as much as the target. Speculative decoding pays off when the target has billions of parameters and each pass is limited by memory bandwidth.
What do γ and α mean?
γ (gamma) is how many tokens the draft model guesses per round. α (alpha) is the acceptance rate, how often a guess is kept. Predictable text such as counting or code gives a high α; creative text at high temperature gives a low one.

Credits

Speculative Decoding Playground stands on open-source work and published research. Thank you to everyone behind it.

Built with

Models

  • SmolLM2: Hugging Face; the live 360M / 135M pair
  • Qwen3: Qwen team, Alibaba; the recorded 1.7B / 0.6B runs

More from Sarvabhaum

Web appLive

attnlab

A set of labs, walked in order, that run real TransformerLens models behind a web page. Type a prompt, pick a model, and see how it is tokenized, where every head attends, and when the model knows its answer.

  • Mechanistic interpretability
  • TransformerLens
  • Attention
InteractiveLive

One transformer block, neuron by neuron

A tiny real transformer (d = 8, 2 heads, an 8 → 32 → 8 MLP, 4 blocks) runs on the sentence you type. Every circle holds the number it actually computed, and hovering one shows the weighted lines that feed it.

  • Transformers
  • Neural networks
  • Attention
InteractiveLive

Training a transformer on eight GPUs

One concrete configuration, followed from start to finish: a 1.27B-parameter model with Llama 2 7B's layer shape, trained on a single 8-GPU node split 2-way tensor × 2-way pipeline × 2-way data parallel, then post-trained and deployed.

  • Distributed training
  • Parallelism
  • LLM training