inference atlas

Track 06 · Advanced · ~45 min · from the vLLM Office Hours

Inside the engine.

The staff-level layer: how vLLM’s V1 engine schedules every GPU step, why CPU overhead became the enemy, how to size KV and GPUs on a whiteboard, and how wide EP, KV connectors and llm-d fit together. Every lab maps to questions in the 63-question bank.

6.1AdvancedQ7 · Q11 · Q12

Two processes and a busy loop.

With Llama-8B on an H100 a forward pass takes about 5 ms, so HTTP handling, tokenization and scheduling were eating a large share of every step. V1 moved that work into a separate process and made the EngineCore a tight loop. Step through it, then compare sync and async scheduling.

Walkthrough · vLLM V1 architecture

API server ⇄ ZeroMQ ⇄ EngineCore → GPU workers

Process 1 · API server (frontend)
HTTP · OpenAI APIFastAPI server on :8000
Tokenizer+ multimodal preprocessing
Detokenizer + streamingSSE back to the client
requests →
ZeroMQ
← outputs (diffs)
Process 2 · EngineCore (busy loop)
schedule()token budget → {req: tokens}
execute_model()forward pass on all workers
update_from_output()stops · frees finished blocks
sample()next token per sequence
KV cache managerblocks · prefix hashes · LRU
GPU workers (TP=4)
rank 0rank 1rank 2rank 3
CPU vs GPU timeline · 4 steps
CPU
sched 1update+sched 2update+sched 3update+sched 4
GPU
step 1step 2step 3step 4

4 steps take 28.8 ms · GPU busy 69% (hatched = GPU idle, waiting on the CPU). With a small model on a fast GPU, CPU overhead is a big slice of every step, which is why V1’s process split and async scheduling paid off.

Step 1/9
A request arrives over HTTP at the API server process, an OpenAI-compatible FastAPI server.

Why split processes

V1 delivered up to 1.7× V0’s throughput with almost the same kernels, so nearly all of the gain was CPU overhead removed. Workers cache request state and receive only per-step diffs.

Spot a CPU-bound server

Gaps between kernels in a profiler trace, step time well above kernel time. It’s worst with small models on fast GPUs.

Model Runner V2 (2026)

A rewrite with a stable state table, GPU-native input prep, async-first design and a Triton sampler: Qwen3-0.6B on one GB200 went from 16K to 25K output tokens/s.

Staff answer

“Async scheduling plans step N+1 while step N runs, so the GPU never waits on the CPU. The catch is that the scheduler doesn’t yet know step N’s results (stop tokens, accepted drafts, grammar state), so it schedules optimistically and corrects a step later. Any hidden .item() or .cpu() sync breaks it; I’d find those in an Nsight trace.”

6.2Advancedflagship lab · Q8 · Q9

Be the scheduler, one step at a time.

Each step vLLM hands out a token budget as a map {request: tokens}. Running requests go first, then waiting ones are admitted while budget, sequence slots and KV blocks last. Out of blocks? Someone gets preempted. Change the knobs and step through six requests.

Simulator · 24 KV blocks × 16 tokens · 6 requests

Step 0

--max-num-batched-tokens
--max-num-seqs
--long-prefill-token-threshold
Waiting queue
A
arrives after step 0
B
arrives after step 0
C
arrives after step 1
D
arrives after step 2
E
arrives after step 4
F
arrives after step 5
Running
nothing running
KV block pool
ownedfreed but still cached (evictable)free
This step’s budget
scheduled = { } // press Step
Step0
Budget used0 / 128
Free KV blocks24 / 24
Preemptions0
Finished0 / 6
TTFT (steps)—

The loop

Running requests first (so decodes aren’t starved), then waiting ones while budget, --max-num-seqs and KV blocks remain. Chunked prefill falls out of the budget for free.

Preemption

If a running request can’t get a new block, the scheduler preempts the most recently admitted one, frees its blocks and recomputes it later. With prefix caching, the recompute is often a cache hit.

Prefix hits

Full blocks are hashed. Freed blocks sit in an LRU free queue and stay reusable until evicted, so a later request with the same prefix skips those tokens.

Staff answer

“The V1 scheduler has no separate prefill or decode phase: each step it assigns tokens per request under a budget. The knobs are --max-num-batched-tokens, --max-num-seqs and --long-prefill-token-threshold. Speculative drafts are just extra tokens for a request, rolled back on rejection by moving its computed-token count.”

6.3AdvancedQ10

When the CPU can’t launch fast enough.

Every kernel costs the CPU a few microseconds to launch. When kernels are tiny, as in decode on a small model, the GPU waits between them. CUDA graphs record a sequence of launches once and replay it. Compare the modes and shrink the kernels.

Lab · one decode step · 3 layers × 5 kernels

Eager launches

CPU
GPU
0 µs80 µs
CPU kernel launch (~5 µs)graph replay launchGPU kernelattention kernel
Step time80 µs
GPU busy82%
vs eager1.00× faster
Eager launches
Kernels (4 µs) are shorter than a launch (5 µs), so the GPU finishes each one and waits for the CPU. That’s the CPU-bound decode problem.
6.4AdvancedQ2 · Scenario D

Size it on a whiteboard.

Two calculators interviewers love: KV bytes and max concurrency for a model on a GPU, and the full “how many GPUs does this chat service need?” derivation, with every step shown so you can say it out loud.

KV cache & concurrency

Pick a preset or type your own architecture. The formula is the one interviewers expect you to derive out loud.

Model preset
Parameters (B)
Layers
KV heads
Head dim
KV dtype
Weight dtype
GPU memory (GB)
GPUs (TP)
gpu-memory-utilization
Avg context per request (tokens)
1
KV bytes per token
2 × 80 layers × 8 KV heads × 128 dims × 2 B
320 KiB per token
2
Per GPU under TP
KV heads split across min(TP, heads) = 4 GPUs
80.0 KiB per token per GPU
3
Free memory per GPU for KV
80 GB × 0.9 − 35.3 GB weights − ~2 GB activations
34.7 GB
4
Max concurrency (what vLLM logs at startup)
34.7 GB ÷ (8,192 tokens × 80.0 KiB)
≈ 51 concurrent requests
5
One 128K-token context
131,072 × 320 KiB
40.0 GiB of KV for a single request
Answer: 51 requests at 8,192 tokens on 4 × 80 GB. FP8 KV would roughly double it.
6.5AdvancedQ20–Q24

Eight ranks, one heartbeat.

In wide expert parallelism every rank must join every MoE all-to-all, so each step lasts as long as the slowest rank. Drop a long prefill onto one rank, then fix it with the real vLLM levers: prefill scheduling intervals, EPLB and dual-batch overlap.

Lab · one decode step across 8 EP ranks (DeepSeek-style, DP attention + EP)

Everyone waits at the all-to-all.

rank 0
rank 1
rank 2
rank 3
rank 4
rank 5
rank 6
rank 7
computeall-to-all (dispatch + combine)idle, waiting for the slowest rank
This step64.0 ms
Avg over 4 steps64.0 ms
Idle rank-time63%
Throughput vs ideal28%
wide EP
A long prefill on rank 3 makes this step last until rank 3 finishes: the other seven ranks sit idle at the all-to-all. Prefills arrive on different ranks in different steps, so every step stalls like this. Try --prefill-schedule-interval.

Why wide EP

MLA stores one shared latent, so TP can’t split KV. DP attention + EP gives every rank its own requests’ KV. Result: about 2.2k tokens/s per H200 for DeepSeek decode, up from ~1.5k.

The levers

  • --prefill-schedule-interval: align prefills across ranks
  • --enable-eplb: replicate hot experts
  • --enable-dbo: overlap all-to-all with compute

Failure domain

One sick GPU stalls the whole group: every MoE layer needs every rank. You need fast health detection, routing away at the EPP, and eventually elastic EP.

Staff answer

“Expert-parallel ranks run in lockstep because every MoE all-to-all needs all ranks, so imbalance becomes idle time. I’d disaggregate prefill away from wide-EP decode, align any prefills with a schedule interval, use EPLB for skewed experts, and turn on DBO once all-to-all dominates the profile, checking small-batch latency didn’t regress.”

6.6AdvancedQ26

How KV actually moves.

One connector interface serves P/D transfer, CPU offloading and remote caches. It has a scheduler side (what can be loaded instead of computed?) and a worker side (start loads, wait per layer, save). Follow a decode instance pulling a prefilled request’s KV over NIXL.

Sequence · KVConnector with NixlConnector · decode side

Scheduler → worker → NIXL → prefill instance

Schedulerdecode instance · CPU
Workerdecode GPU
NIXLtransfer library
Prefill instanceholds the KV
get_num_new_matched_tokens() · → 8,192 loadable
allocate · update_state_after_alloc()
build_connector_meta() · load plan for this step
start_load_kv() · async
RDMA read of KV blocks · decode pulls
KV lands layer by layer
wait_for_layer_load(i) · per layer
get_finished() · transfers done
request_finished() → free · timeout as backstop
Step 1/9
The scheduler asks the connector how many of this request’s tokens can be loaded instead of computed. It must know this before spending budget or allocating blocks.
6.7AdvancedQ28 · Q29

Tune the endpoint picker.

llm-d’s EPP runs filters, then weighted scorers. Pick a request type, toggle filters and move the weights. Watch the winner change and the reasons appear.

Playground · 5 vLLM pods behind one InferencePool

Which pod gets this request?

Filters
Scorer weights
 

score = wP·prefix + wQ·(1 − queue/8) + wK·(1 − KV)

PodRoleHealthyAdaptersPrefix matchQueueKV usedScoreNote
decode-1decodeyesbase90%672%2.33✓ picked
decode-2decodeyesbase, sql-v210%235%1.60
decode-3decodeNObase90%05%3.75filtered: unhealthy
decode-4decodeyesbase0%120%1.68
prefill-1prefillyesbase, sql-v290%010%3.70filtered: wrong role
Support bot turn
decode-1 wins with 2.33. It has 90% of this prompt cached, which outweighs its longer queue. Sticky until saturated. Try raising the queue weight.
6.8Advancedhands-on

Build a vllm serve command.

Choose a model and features. The command updates live, every flag is explained, and the builder warns about combinations that don’t make sense. It’s a quick way to connect the concepts to real flags.

Choose

Model ids and flags are real vLLM options. Pin versions in production.

Model
GPU
Tensor parallel
Data parallel
Max context
Token budget / step
Max sequences
Speculative decoding
Command
vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --tensor-parallel-size 4 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90 \
  --max-num-batched-tokens 8192 \
  --max-num-seqs 256 \
  --kv-cache-dtype fp8
--tensor-parallel-size — Shard every layer across 4 GPUs over NVLink. Pick the smallest TP with enough KV headroom.
--max-model-len — Longest prompt + output accepted. It bounds per-request KV.
--gpu-memory-utilization — Fraction of GPU memory vLLM claims. Whatever is left after weights and activations becomes KV blocks.
--max-num-batched-tokens — Per-step token budget: the latency ↔ throughput dial (chunked prefill).
--max-num-seqs — Max sequences per step. It caps batch size and concurrent KV demand.
--kv-cache-dtype — Halves KV bytes per token, roughly doubling concurrency. Validate long-context accuracy.
6.9Flashcards

Numbers worth memorizing.

The figures interviewers expect you to quote or estimate. Tap a card to flip it. Try to say what each number measures before you look.

6.10Quiz

Staff-level check.

Harder questions drawn from the vLLM Office Hours. For all 63 questions with model answers, open the interview bank.

1. V1 gave up to 1.7× the throughput of V0 with almost the same kernels. Where did the gain come from?

2. In V1’s scheduler, how is chunked prefill implemented?

3. Why do piecewise CUDA graphs run attention eagerly?

4. Scenario D: 50 req/s, 500 output tokens, 40 ms TPOT. How many sequences are in flight at peak?

5. In the KV connector API, why is get_num_new_matched_tokens a scheduler-side call?

6. P99 TTFT jumped after a deploy while NCCL and GPU health look fine. Best first move?

0 / 6 answered