- What is speculative decoding?
- A way to make a large language model generate text in fewer slow steps. A small draft model proposes several tokens, the large model scores all of them in one forward pass, and the tokens it agrees with are kept. With the right acceptance rule the output has exactly the large model's distribution.
- Is the output really identical to the large model's?
- In greedy mode it is token for token the same, and the playground's tests check this. In sampling mode it follows exactly the same probability distribution: each draft token is accepted with probability p/q, and after a rejection the replacement is drawn from the normalized difference between the two distributions.
- Why isn't the small pair faster?
- With models of a few hundred million parameters, a forward pass is dominated by fixed overhead rather than by reading weights, so the draft costs nearly as much as the target. Speculative decoding pays off when the target has billions of parameters and each pass is limited by memory bandwidth.
- What do γ and α mean?
- γ (gamma) is how many tokens the draft model guesses per round. α (alpha) is the acceptance rate, how often a guess is kept. Predictable text such as counting or code gives a high α; creative text at high temperature gives a low one.