DeepSpeed ZeRO Stages: Which Training State Is Sharded?
ZeRO-1 shards optimizer state, ZeRO-2 also shards gradients, and ZeRO-3 also shards model parameters across data-parallel GPUs. Test the lowest stage that lets the workload fit at acceptable throughput. A higher stage saves more model-state memory but can require more communication.
These stages change where training state is stored. They do not automatically reduce activation memory.
A memory example with stated assumptions
Suppose training uses FP16 weights and gradients plus FP32 Adam master weights and two optimizer moments. This allocates 2 bytes for weights, 2 for gradients, and 12 for optimizer state per parameter: 16 bytes in total.
For 7.5 billion parameters on eight GPUs, the model-state calculation is:
| Stage | State split across GPUs | Model-state GB per GPU |
|---|---|---|
| Ordinary data parallelism | None | 120 |
| ZeRO-1 | Optimizer state | 41.25 |
| ZeRO-2 | Optimizer state and gradients | 28.125 |
| ZeRO-3 | All three components | 15 |
These are decimal GB. They exclude activations, temporary gathered parameters, allocator overhead, and runtime buffers. A training stack that stores different precisions needs different arithmetic.
The ZeRO paper explains the state partitions and analyzes their communication. ZeRO-3 reconstructs needed parameters through all-gather operations; the paper’s full-sharding analysis reaches about 1.5 times standard data-parallel communication volume. That ratio is an analysis under its assumptions, not a guarantee for every network and configuration.
Select from the memory profile
If parameters fit and optimizer states cause the failure, test stage 1 or 2. If full replicated parameters cannot fit, test stage 3. If activations dominate, combine the chosen stage with activation checkpointing or a smaller training batch.
DeepSpeed also supports CPU and NVMe state offload. Check the current ZeRO configuration and measure transfer cost on the actual host and storage. Offload capacity can help, but slow transfers can reduce throughput.
The ZeRO section of the LLM Engineering Guide provides an executable calculation for other parameter counts and GPU groups.