- What is the difference between data, tensor and pipeline parallelism?
- Data parallelism gives every GPU a full copy of the model and a different slice of the batch. Tensor parallelism splits each weight matrix across GPUs, so they cooperate on every layer. Pipeline parallelism puts different groups of layers on different GPUs and streams micro-batches through them. Large runs combine all three.
- How much GPU memory does training a model take?
- With bf16 matmuls and AdamW, about 16 bytes per parameter for weights, gradients and optimizer state, before any activations: roughly 20 GB for a 1.27B model and over 1 TB for a 70B model. That is why model state has to be sharded.
- What is the difference between ZeRO and FSDP?
- Both shard model state across data-parallel GPUs. ZeRO stage 1 shards optimizer state, stage 2 also shards gradients, and stage 3 also shards the weights themselves. FSDP is PyTorch's implementation of the stage 3 approach.
- How long does a training step take on 8 GPUs?
- For the example (1.27B parameters, 32,768 tokens per step), about 240 TFLOP per step. Eight H100s at around 40% utilization finish it in roughly 0.08 seconds, about 430,000 tokens per second.