LLM Parallelism: How Do TP, PP, DP, and EP Differ?

Tensor parallelism splits weight operations. Pipeline parallelism splits layers. Serving data parallelism runs independent model replicas. Expert parallelism distributes the experts of a mixture-of-experts model. Choose from model fit, traffic, and the links between GPUs.

These are four common approaches. Large deployments can combine them, and long-context training may also use sequence or context parallelism.

What does each GPU store?

ApproachModel placementCommunication to account for
Tensor parallelism, TPPieces of weight matricesFrequent collective operations within model execution
Pipeline parallelism, PPConsecutive layer groupsActivations passed between stages
Data parallelism, DPComplete serving replicasNo cross-replica model computation for independent requests
Expert parallelism, EPDifferent MoE expertsTokens routed to experts, often through all-to-all exchange

Training DP differs from serving replicas: training combines gradients, and ZeRO or FSDP can shard training state. MoE systems may also coordinate shared components even when some work is data-parallel.

The Llama 3 report describes combined parallelism for large-model training and pipeline-parallel inference. Its deployment shows why model placement and topology must be considered together; it does not establish a universal configuration.

Choose the smallest group that meets the workload

If a serving model fits on one GPU with sufficient KV-cache capacity, start by testing independent replicas. They can increase total throughput without adding collectives to each request.

If the model needs several GPUs, test TP within a node with fast GPU links. More shards can reduce per-GPU memory but add communication. If the model spans nodes, compare PP and combinations of TP and PP; uneven stage times and idle pipeline periods can reduce utilization.

For an MoE model, measure expert placement, token balance, and interconnect traffic before adding EP. Available expert memory is only one part of the decision.

Keep prompts, output lengths, precision, and latency thresholds constant in the comparison. Report throughput for the complete serving group; dividing by GPU count does not predict how another topology will scale.

The parallelism section of the LLM Engineering Guide illustrates the four placements.