Track 02 · Intermediate · ~30 min
When one GPU is not enough.
Models outgrow a single GPU, and traffic outgrows a single replica. Learn the five ways to split a model, how a smart router sends each request to the GPU that already remembers its prompt, and why the network becomes part of the compute path.
Will it fit?
Pick a model and a GPU, then raise the tensor-parallel degree. Each GPU holds a slice of the weights, and whatever is left becomes KV cache. That leftover is why going from 1 to 2 GPUs can more than double throughput.
Llama-3.1-70B on 4 × H100 80GB
Five ways to split a model.
Each strategy cuts along a different axis: inside a layer (TP), between layers (PP), whole copies (DP), across experts (EP), or along the sequence (CP). Pick one and step through what moves where.
Tensor parallelism · split inside each layer
TP · Tensor
Split each weight matrix across GPUs.
Talks: 2 all-reduces per layer, every step.
Use when: Model won’t fit, or you need lower latency. Stay within NVLink.
PP · Pipeline
Consecutive layers on different GPUs.
Talks: Activations at stage boundaries only.
Use when: Too big for one node and inter-node links are slow. Watch bubbles.
DP · Data
Independent replicas serve different requests.
Talks: None between replicas.
Use when: You need throughput. Pair with cache-aware routing.
EP · Expert
MoE experts live on different GPUs.
Talks: All-to-all dispatch + combine per MoE layer.
Use when: MoE models, usually with DP attention (“wide EP”).
CP · Context
One request’s context is sharded across GPUs.
Talks: Query gather + partial-output merge.
Use when: Very long contexts: DCP for decode, PCP for prefill.
Route to the GPU that remembers.
Four vLLM replicas each keep a prefix cache of the prompts they’ve seen. A request whose prefix is cached skips most of its prefill. Switch routing policies, turn on a hot prompt, and watch hit rate, TTFT and queues react.
Same traffic. Four routing brains.
Why round-robin fails
LLM requests are expensive and uneven, replicas hold cached state, and prefill/decode stress different hardware. Spreading requests evenly spreads cache misses evenly.
Good vs bad signals
Good: queue depth, KV-cache usage, prefix overlap, predicted latency. Bad: CPU or GPU utilization and connection counts.
In the real world
llm-d’s endpoint picker cut mean TTFT ~3× and served ~50–100% more QPS within a 2 s P95 TTFT SLO versus baseline routing. See it on Kubernetes →
Staff answer
“Inference scheduling isn’t CPU-style load balancing. I’d score replicas on prefix-cache overlap, queue depth and KV-cache pressure, then break ties by load, so we’re sticky until saturated. The cheapest-looking routing decision can otherwise create expensive recomputation.”
GPUs talking: collectives and links.
Split models must exchange data every step. NCCL provides the collective operations; the interconnect decides how long they take. Step through each collective, then compare links.
All-reduce
Used by tensor parallelism: partial results are summed after attention and after the MLP, twice per layer.
The bandwidth ladder
Time to move 3.3 GB, the KV cache of a 10K-token prompt on Llama-3.1-70B (≈ 320 KiB/token). Log scale: every gridline step is ×10.
Per-direction peak rates; real transfers reach less. HBM is the GPU’s own memory, shown for scale.
Inside a node
NVLink/NVSwitch joins 8 GPUs at hundreds of GB/s. Keep tensor parallelism here: it all-reduces twice per layer, every step.
Across nodes
InfiniBand or RoCE with RDMA (GPU memory → NIC with no CPU copy). Roughly 10× slower than NVLink, so use PP, EP or DP across nodes.
Staff answer
“In multi-node inference the network is part of the compute path: TP, PP and MoE all move tensors every step. I treat bandwidth, latency, RDMA and topology as first-class constraints, and when throughput drops going from one node to two, I check NCCL collectives before blaming the model.”
Pick the right split.
Scenario questions like the ones interviewers ask. For drag-and-drop matching, visit the playground.
1. Llama-3.1-70B in BF16 (~141 GB) must serve a chat product with low latency on one 8×H100 node. Best starting layout?
2. A 405B model doesn’t fit on one node. Nodes connect over 400 Gb/s InfiniBand. How do you split it?
3. Why is plain TP a poor fit for DeepSeek-V3’s MLA attention?
4. Round robin gave a 30% prefix hit rate. Prefix-hash routing gave 90%, but P99 TTFT exploded when one system prompt went viral. The fix?
5. Which NCCL collective does MoE token dispatch use?
0 / 5 answered