Track 06 · Advanced · ~45 min · from the vLLM Office Hours
Inside the engine.
The staff-level layer: how vLLM’s V1 engine schedules every GPU step, why CPU overhead became the enemy, how to size KV and GPUs on a whiteboard, and how wide EP, KV connectors and llm-d fit together. Every lab maps to questions in the 63-question bank.
Two processes and a busy loop.
With Llama-8B on an H100 a forward pass takes about 5 ms, so HTTP handling, tokenization and scheduling were eating a large share of every step. V1 moved that work into a separate process and made the EngineCore a tight loop. Step through it, then compare sync and async scheduling.
API server ⇄ ZeroMQ ⇄ EngineCore → GPU workers
4 steps take 28.8 ms · GPU busy 69% (hatched = GPU idle, waiting on the CPU). With a small model on a fast GPU, CPU overhead is a big slice of every step, which is why V1’s process split and async scheduling paid off.
Why split processes
V1 delivered up to 1.7× V0’s throughput with almost the same kernels, so nearly all of the gain was CPU overhead removed. Workers cache request state and receive only per-step diffs.
Spot a CPU-bound server
Gaps between kernels in a profiler trace, step time well above kernel time. It’s worst with small models on fast GPUs.
Model Runner V2 (2026)
A rewrite with a stable state table, GPU-native input prep, async-first design and a Triton sampler: Qwen3-0.6B on one GB200 went from 16K to 25K output tokens/s.
Staff answer
“Async scheduling plans step N+1 while step N runs, so the GPU never waits on the CPU. The catch is that the scheduler doesn’t yet know step N’s results (stop tokens, accepted drafts, grammar state), so it schedules optimistically and corrects a step later. Any hidden .item() or .cpu() sync breaks it; I’d find those in an Nsight trace.”
Be the scheduler, one step at a time.
Each step vLLM hands out a token budget as a map {request: tokens}. Running requests go first, then waiting ones are admitted while budget, sequence slots and KV blocks last. Out of blocks? Someone gets preempted. Change the knobs and step through six requests.
Step 0
scheduled = { } // press StepThe loop
Running requests first (so decodes aren’t starved), then waiting ones while budget, --max-num-seqs and KV blocks remain. Chunked prefill falls out of the budget for free.
Preemption
If a running request can’t get a new block, the scheduler preempts the most recently admitted one, frees its blocks and recomputes it later. With prefix caching, the recompute is often a cache hit.
Prefix hits
Full blocks are hashed. Freed blocks sit in an LRU free queue and stay reusable until evicted, so a later request with the same prefix skips those tokens.
Staff answer
“The V1 scheduler has no separate prefill or decode phase: each step it assigns tokens per request under a budget. The knobs are --max-num-batched-tokens, --max-num-seqs and --long-prefill-token-threshold. Speculative drafts are just extra tokens for a request, rolled back on rejection by moving its computed-token count.”
When the CPU can’t launch fast enough.
Every kernel costs the CPU a few microseconds to launch. When kernels are tiny, as in decode on a small model, the GPU waits between them. CUDA graphs record a sequence of launches once and replay it. Compare the modes and shrink the kernels.
Eager launches
Size it on a whiteboard.
Two calculators interviewers love: KV bytes and max concurrency for a model on a GPU, and the full “how many GPUs does this chat service need?” derivation, with every step shown so you can say it out loud.
KV cache & concurrency
Pick a preset or type your own architecture. The formula is the one interviewers expect you to derive out loud.
Eight ranks, one heartbeat.
In wide expert parallelism every rank must join every MoE all-to-all, so each step lasts as long as the slowest rank. Drop a long prefill onto one rank, then fix it with the real vLLM levers: prefill scheduling intervals, EPLB and dual-batch overlap.
Everyone waits at the all-to-all.
--prefill-schedule-interval.Why wide EP
MLA stores one shared latent, so TP can’t split KV. DP attention + EP gives every rank its own requests’ KV. Result: about 2.2k tokens/s per H200 for DeepSeek decode, up from ~1.5k.
The levers
--prefill-schedule-interval: align prefills across ranks--enable-eplb: replicate hot experts--enable-dbo: overlap all-to-all with compute
Failure domain
One sick GPU stalls the whole group: every MoE layer needs every rank. You need fast health detection, routing away at the EPP, and eventually elastic EP.
Staff answer
“Expert-parallel ranks run in lockstep because every MoE all-to-all needs all ranks, so imbalance becomes idle time. I’d disaggregate prefill away from wide-EP decode, align any prefills with a schedule interval, use EPLB for skewed experts, and turn on DBO once all-to-all dominates the profile, checking small-batch latency didn’t regress.”
How KV actually moves.
One connector interface serves P/D transfer, CPU offloading and remote caches. It has a scheduler side (what can be loaded instead of computed?) and a worker side (start loads, wait per layer, save). Follow a decode instance pulling a prefilled request’s KV over NIXL.
Scheduler → worker → NIXL → prefill instance
Tune the endpoint picker.
llm-d’s EPP runs filters, then weighted scorers. Pick a request type, toggle filters and move the weights. Watch the winner change and the reasons appear.
Which pod gets this request?
score = wP·prefix + wQ·(1 − queue/8) + wK·(1 − KV)
| Pod | Role | Healthy | Adapters | Prefix match | Queue | KV used | Score | Note |
|---|---|---|---|---|---|---|---|---|
| decode-1 | decode | yes | base | 90% | 6 | 72% | 2.33 | ✓ picked |
| decode-2 | decode | yes | base, sql-v2 | 10% | 2 | 35% | 1.60 | |
| decode-3 | decode | NO | base | 90% | 0 | 5% | 3.75 | filtered: unhealthy |
| decode-4 | decode | yes | base | 0% | 1 | 20% | 1.68 | |
| prefill-1 | prefill | yes | base, sql-v2 | 90% | 0 | 10% | 3.70 | filtered: wrong role |
Build a vllm serve command.
Choose a model and features. The command updates live, every flag is explained, and the builder warns about combinations that don’t make sense. It’s a quick way to connect the concepts to real flags.
Choose
Model ids and flags are real vLLM options. Pin versions in production.
vllm serve meta-llama/Llama-3.1-70B-Instruct \ --tensor-parallel-size 4 \ --max-model-len 32768 \ --gpu-memory-utilization 0.90 \ --max-num-batched-tokens 8192 \ --max-num-seqs 256 \ --kv-cache-dtype fp8
--tensor-parallel-size — Shard every layer across 4 GPUs over NVLink. Pick the smallest TP with enough KV headroom.--max-model-len — Longest prompt + output accepted. It bounds per-request KV.--gpu-memory-utilization — Fraction of GPU memory vLLM claims. Whatever is left after weights and activations becomes KV blocks.--max-num-batched-tokens — Per-step token budget: the latency ↔ throughput dial (chunked prefill).--max-num-seqs — Max sequences per step. It caps batch size and concurrent KV demand.--kv-cache-dtype — Halves KV bytes per token, roughly doubling concurrency. Validate long-context accuracy.Numbers worth memorizing.
The figures interviewers expect you to quote or estimate. Tap a card to flip it. Try to say what each number measures before you look.
Staff-level check.
Harder questions drawn from the vLLM Office Hours. For all 63 questions with model answers, open the interview bank.
1. V1 gave up to 1.7× the throughput of V0 with almost the same kernels. Where did the gain come from?
2. In V1’s scheduler, how is chunked prefill implemented?
3. Why do piecewise CUDA graphs run attention eagerly?
4. Scenario D: 50 req/s, 500 output tokens, 40 ms TPOT. How many sequences are in flight at peak?
5. In the KV connector API, why is get_num_new_matched_tokens a scheduler-side call?
6. P99 TTFT jumped after a deploy while NCCL and GPU health look fine. Best first move?
0 / 6 answered