LLM Parallelism: How Do TP, PP, DP, and EP Differ?
Tensor parallelism splits weight operations. Pipeline parallelism splits layers. Serving data parallelism runs independent model replicas. Expert parallelism distributes the experts of a mixture-of-experts model. Choose from model fit, traffic, and the links between GPUs.
These are four common approaches. Large deployments can combine them, and long-context training may also use sequence or context parallelism.
What does each GPU store?
| Approach | Model placement | Communication to account for |
|---|---|---|
| Tensor parallelism, TP | Pieces of weight matrices | Frequent collective operations within model execution |
| Pipeline parallelism, PP | Consecutive layer groups | Activations passed between stages |
| Data parallelism, DP | Complete serving replicas | No cross-replica model computation for independent requests |
| Expert parallelism, EP | Different MoE experts | Tokens routed to experts, often through all-to-all exchange |
Training DP differs from serving replicas: training combines gradients, and ZeRO or FSDP can shard training state. MoE systems may also coordinate shared components even when some work is data-parallel.
The Llama 3 report describes combined parallelism for large-model training and pipeline-parallel inference. Its deployment shows why model placement and topology must be considered together; it does not establish a universal configuration.
Choose the smallest group that meets the workload
If a serving model fits on one GPU with sufficient KV-cache capacity, start by testing independent replicas. They can increase total throughput without adding collectives to each request.
If the model needs several GPUs, test TP within a node with fast GPU links. More shards can reduce per-GPU memory but add communication. If the model spans nodes, compare PP and combinations of TP and PP; uneven stage times and idle pipeline periods can reduce utilization.
For an MoE model, measure expert placement, token balance, and interconnect traffic before adding EP. Available expert memory is only one part of the decision.
Keep prompts, output lengths, precision, and latency thresholds constant in the comparison. Report throughput for the complete serving group; dividing by GPU count does not predict how another topology will scale.
The parallelism section of the LLM Engineering Guide illustrates the four placements.