inference atlas

Track 05 · Advanced · ~25 min

The unit of serving becomes a workflow.

Modern workloads aren’t one prompt and one answer. Mixture-of-experts routes every token across GPUs, multimodal models chain encoders and LLMs, RL puts inference inside the training loop, and agents turn one request into dozens of cached turns.

5.1Advancedsimulator

Send tokens to experts.

Each token’s router picks its top-2 of 8 experts, spread over 4 GPUs. Every MoE layer waits for the slowest GPU. Skew the traffic toward code and watch one GPU become the straggler, then switch on EPLB.

Simulator · 8 experts · top-2 routing · expert parallelism over 4 GPUs

Lockstep means the busiest GPU sets the pace.

Send a token
Router scores · last token
E1
E2
E3
E4
E5
E6
E7
E8
GPU 0
E10.0
E20.0
GPU 1
E30.0
E40.0
GPU 2
E50.0
E60.0
GPU 3
E70.0
E80.0
Tokens routed0
GPU load imbalance1.00× max/mean
Step time vs perfectly balanced1.00×
EPLB replicas0
balanced
Tokens are routed to their top-2 experts. Press a token button, or watch the auto-stream.

What

MoE layers hold many expert FFNs, but each token activates only a few. Memory scales with total parameters and compute with active ones: DeepSeek-R1 uses 37B of 671B per token.

Small batches

Tokens per expert ≈ batch × top-k ÷ experts. DeepSeek-V3 sends each token to 8 of 256 experts, so a 256-token decode batch gives each expert about 8 tokens: a bandwidth-bound GEMM.

EPLB

vLLM’s expert-parallel load balancer records per-expert load, periodically computes a new placement with redundant copies of hot experts, and shuffles weights without a restart. Redundant experts cost HBM that could hold KV.

Staff answer

“MoE serving cuts active compute per token but adds all-to-all token routing. Because every rank must join each MoE layer’s collective, the group moves at the pace of its slowest rank, so expert load balance, DP-aware routing and prefill scheduling decide throughput. I’d measure the max/mean tokens per rank per layer before and after EPLB.”

5.2Advanced

A picture is worth hundreds of tokens.

A user uploads a screenshot of a failing pod and asks what’s wrong. Follow the image through preprocessing, the vision encoder and into the LLM, then send it again and watch the caches kick in.

Walkthrough · vision-language serving

Image → patches → embeddings → LLM

turn 1
Uploadscreenshot.png + “what’s wrong?”
Preprocessdecode · resize · split into patchesCPU · separate process
Vision encoderViT turns patches into embeddingsGPU · encoder cache
Image tokens~576 for a 336 px image (24×24 patches)
LLM prefillimage tokens + text tokens in one sequence
block hash includes image hash
Decodestreams the answer
OOMKilled: raise the memory limit
Any-to-any (vLLM-Omni):Thinker · reasoning LLM→Talker · audio codes→Code2Wav · vocoderEach stage gets its own GPUs, memory share and metrics.
Step 1/8
The user uploads a screenshot of a failing pod and asks “what’s wrong?”. The request carries an image and some text.

Off-loop preprocessing

Decoding, resizing and cropping images runs in a separate process, with a cache for repeated inputs, so the GPU loop never waits on it.

Encoder cache

Vision embeddings are kept after the encoder runs, so a long prefill can be chunked across steps without re-running the encoder.

Multimodal prefix caching

Image hashes join token IDs in block hashes, so multi-turn chats about the same image hit the cache. Client-side re-encoding breaks it: normalize media first.

Staff answer

“Multimodal serving chains heterogeneous stages: preprocessing on CPU, vision or audio encoders, then the LLM. I’d treat the encoder as cacheable, separately scaled work, via E/PD disaggregation at scale, and pick per-stage metrics, because TTFT and TPOT don’t describe a vocoder or a diffusion stage.”

5.3Advanced

Inference inside the training loop.

Reasoning models are post-trained with reinforcement learning: generate attempts, score them, update the weights, repeat. The generation half is an inference workload, and it must never starve the trainer.

Walkthrough · RL post-training loop

Rollout → reward → train → sync weights

Rollout enginesvLLM generates many long samples from the current policyinference workload
Environment · rewardunit tests, math checkers, verifiers or a reward model score each rolloutCPU / sandbox / RM
Trainercomputes gradients from rewarded rollouts and updates the policytraining GPUs
Weight syncnew weights pushed to every rollout engine, no restartsNCCL broadcast · sleep mode
Rollout GPUs
generategenerate
Trainer GPUs
traintrain

Synchronous: each side waits for the other. Rollout GPUs busy 72%, trainer busy 24% (hatched = idle).

Step 1/7
RL post-training: the policy generates attempts, a reward scores them, and the trainer updates the weights. Repeat thousands to millions of times.

Fast weight sync

New policy weights must reach rollout engines without restarts: modular weight-sync APIs, NCCL broadcast, and sleep mode to hand GPU memory between trainer and engine.

The silent bug

After a weight update, cached KV was computed with the old weights. Forget to reset the prefix cache and rollouts become subtly off-policy, with no error anywhere.

Batch invariance

Floating-point reductions change with batch shape, so a request’s logprobs can shift with its batchmates. RL importance ratios then drift for no visible reason.

Staff answer

“RL infra has two competing workloads, training and rollout generation. I treat rollouts as a scalable inference workload, overlap them with training where the algorithm allows, and make weight sync, prefix-cache invalidation and batch-invariant logprobs part of the contract, so training is never starved and never silently off-policy.”

5.4Advancedsimulator

An agent is a session, not a request.

An SRE agent investigates “why is checkout failing?”. Every turn resends the whole growing context plus one new tool result. Route turns stickily or load-balance them, add a shared KV pool, and compare what you pay.

Simulator · one agent session · 3 vLLM pods

Cached prefix vs new tokens, turn by turn

Routing
1. plan the investigationsystem + tools + task
pod-1
TTFT 573 ms
2. prometheus_queryerror-rate metrics
pod-1
TTFT 160 ms
3. kubectl logs checkoutstack traces
pod-1
TTFT 253 ms
4. kubectl describe podevents · OOMKilled
pod-1
TTFT 140 ms
5. git diff HEAD~1recent change
pod-1
TTFT 213 ms
6. run_teststest output
pod-1
TTFT 120 ms
7. kubectl rollout undorollback result
pod-1
TTFT 80 ms
8. final answersummary for the user
pod-1
TTFT 67 ms
cached prefix (reused on the same pod)fetched from shared KV poolnew prefill (computed)output tokens
Input reused (cache + pool)85%
Prefill tokens computed19K
Sum of TTFTs1.61 s
Pod hops0
sticky
Session-sticky: every turn lands on the pod that already holds the conversation, so only the new tool result is prefilled. Input reuse is 85%, the AgentX pattern (over 96% on long real sessions).

Real agent traffic

SemiAnalysis’ AgentX traces: a median of 43 turns per session, 142K input and 444 output tokens per request, and a prefix-cache hit rate above 96%. 44% of sessions spawn subagents.

What’s scarce

With 96% cached input, a turn costs its decode, a small prefill and the memory to keep its KV resident. KV capacity, not FLOPs, becomes the bottleneck.

Bitter lessons

On agent traces, session-sticky routing beat balancing by queue, tokens or KV, and pipeline parallelism lost because warm turns add too few new tokens to fill a pipeline.

Staff answer

“Agentic serving changes the unit from one model call to a multi-step session. I’d route by session ID, sticky until a replica saturates, keep KV warm in tiers between turns, cap long fresh prefills so they don’t block short turns, and measure tokens per GPU-second at a P90 interactivity floor, reporting cached, uncached and output tokens separately.”

5.5Quiz

Check yourself.

Frontier-workload questions from real staff loops.

1. DeepSeek-R1 has 671B parameters, 37B active per token. What scales with the 671B?

2. In a 32-rank wide-EP decode group, one rank starts a long prefill. What happens?

3. In a multi-turn image chat, turn 2 about the same image is as slow as turn 1. Likely cause?

4. After a weight update, RL rollouts are subtly off-policy but nothing errors. First suspect?

5. Which routing policy serves long agentic sessions best?

0 / 5 answered