- Is a transformer a neural network?
- Yes. Every linear layer in a transformer block is an ordinary fully connected layer, and the MLP is a classic two-layer feed-forward network. What is new is attention, where the connection strengths between tokens are computed from the input instead of being fixed learned weights.
- Is the model in the explainer real?
- Yes, but tiny and untrained. It has 8-dimensional vectors, 2 attention heads, an 8 → 32 → 8 MLP and 4 blocks, with random weights. Every value is genuinely computed from your sentence; only the predictions are meaningless, because nothing has been trained.
- What does the MLP in a transformer do?
- It runs on each token separately: it widens the vector to four times its size, applies GELU, and projects it back. Each hidden neuron detects an input pattern (its column of W_in) and writes an output vector when active (its row of W_out). It holds about two thirds of each block's parameters.
- Why does a transformer use LayerNorm?
- So each sublayer sees inputs of a stable size however large the residual stream grows. LayerNorm subtracts a token's mean and divides by its standard deviation across that token's own numbers, then applies a learned scale and shift. Many recent models use RMSNorm, which skips the mean.
- How many parameters does one transformer block have?
- About 12·d² for a model of width d with a 4× MLP: 4d² in attention (Q, K, V and O) and 8d² in the MLP. The explainer's calculator works this out, plus training FLOPs per token, for any model size you enter.