inference atlas

Practice · interview-ready

Say it like a staff engineer.

Interviewers don’t want definitions. They want to hear you reason from workload → bottleneck → architecture → trade-off. Practise the 20 core topics out loud, then work through 63 vLLM questions and 4 system-design scenarios, rating yourself as you go. Your ratings stay in this browser.

20 core topics, out loud.

Read the prompt, answer out loud (aim for 60 seconds), then reveal. Practice mode picks a random topic and starts a timer.

01 · FOUNDATIONS

Prefill vs decode

“Walk me through what happens between a prompt arriving and tokens streaming back. Where do TTFT and ITL come from?”

Answer out loud first, then reveal
What it is
Prefill runs the whole prompt through the model in one parallel pass and builds the KV cache. Decode then generates one token at a time.
Why it matters
Prefill is compute-bound and sets time to first token; decode is memory-bandwidth-bound and sets inter-token latency. They need different tuning.
Scenario
“The first response takes 5 seconds” → look at queueing and prefill. “The first word is quick but the answer takes forever” → look at decode throughput and ITL.
Staff answer
“Prefill processes the input and builds the KV cache, so it’s primarily compute-bound and affects TTFT. Decode generates tokens autoregressively and is more memory-bandwidth and KV-cache sensitive, affecting inter-token latency. That’s why production systems often optimize or scale the two phases independently.”
Play the lab →
02 · FOUNDATIONS

KV cache

“What is the KV cache, why does it exist, and what problem does it create?”

Answer out loud first, then reveal
What it is
The stored Key and Value tensors of every previous token at every layer, so decode doesn’t recompute the whole context for each new token.
Why it matters
Without it, decode work grows quadratically. With it, the cost moves into GPU memory: weights + KV + runtime must fit.
Scenario
An 80 GB GPU running a 60 GB model has ~20 GB left for KV and runtime. Long contexts and thousands of users exhaust that quickly.
Staff answer
“KV cache stores attention Key/Value states from previously processed tokens so decode doesn’t recompute the entire context for every new token. It significantly improves decode efficiency, but it becomes a major GPU-memory constraint as context length and concurrency increase.”
Play the lab →
03 · FOUNDATIONS

Continuous batching

“Why does continuous batching beat static batching for LLMs, and what new problem does it introduce?”

Answer out loud first, then reveal
What it is
The batch is re-formed every forward step: finished sequences leave and waiting ones join immediately.
Why it matters
Output lengths vary wildly; static batches idle until their longest member finishes. Continuous batching keeps every slot busy.
Scenario
100 users with mixed answer lengths: when one request finishes, its slot is refilled at the very next step instead of waiting for the batch.
Staff answer
“Continuous batching dynamically admits and removes sequences during generation instead of waiting for a fixed batch to complete. It improves GPU utilization and throughput under variable request lengths; the interference it introduces from long prefills is bounded with chunked prefill.”
Play the lab →
04 · SCALE OUT

Tensor parallelism

“A 200B model doesn’t fit on one GPU. How does tensor parallelism help, and what does it cost?”

Answer out loud first, then reveal
What it is
Shards each layer’s matrices across GPUs; each computes a slice and they combine results with collectives.
Why it matters
Larger models fit and each GPU reads only 1/N of the weights per step, but every layer needs an all-reduce, twice, on every step.
Scenario
TP=8 on one 8×H100 node forms one model replica. Stretch it across nodes and throughput collapses on the slower link.
Staff answer
“Tensor parallelism shards individual tensor operations across GPUs. It’s useful when a model or layer is too large for one GPU, but it introduces collective communication overhead, so GPU interconnect bandwidth and topology become important.”
Play the lab →
05 · SCALE OUT

Pipeline parallelism

“When would you choose pipeline parallelism over tensor parallelism?”

Answer out loud first, then reveal
What it is
Partitions consecutive layers across devices; micro-batches flow through the stages.
Why it matters
Reduces per-device memory with communication only at stage boundaries, which tolerates slower inter-node links. The cost is pipeline bubbles.
Scenario
A 60-layer model on 3 GPUs: layers 1–20, 21–40, 41–60. With too few micro-batches, stages idle.
Staff answer
“Pipeline parallelism partitions model layers across devices. It reduces per-device memory requirements, but introduces pipeline scheduling and bubble overhead, so micro-batching is used to improve utilization. I’d use it across nodes when the inter-node link is too slow for TP.”
Play the lab →
06 · SCALE OUT

Data parallelism

“How is data parallelism different from tensor parallelism in serving?”

Answer out loud first, then reveal
What it is
Multiple complete replicas of the model, each serving different requests.
Why it matters
Scales throughput nearly linearly with no cross-replica traffic, but each replica must fit on its GPUs, and naive balancing ignores KV caches.
Scenario
100 requests across 3 replicas ≈ 33 each. TP is one model split across GPUs; DP is many copies.
Staff answer
“Data parallelism replicates the model across devices and distributes independent requests across replicas. It scales throughput well but requires each replica to fit on its assigned GPU set, and it needs cache-aware routing to keep prefix hits.”
Play the lab →
07 · SCALE OUT

Expert parallelism

“How do you serve a mixture-of-experts model across GPUs?”

Answer out loud first, then reveal
What it is
Experts live on different GPUs; a router sends each token to its top-k experts.
Why it matters
Huge capacity with little active compute per token, but tokens must travel: all-to-all dispatch and combine every MoE layer.
Scenario
A token is routed to experts 7 and 31, which sit on two different GPUs. The network is now on the critical path.
Staff answer
“Expert parallelism distributes MoE experts across GPUs and routes tokens to the required experts. The key infrastructure challenge is efficient token routing and all-to-all communication, plus expert load balance, since ranks move in lockstep.”
Play the lab →
08 · OPTIMIZE

KV cache management

“How do you stop the KV cache from exhausting GPU memory under load?”

Answer out loud first, then reveal
What it is
Treat KV as a managed memory pool: allocation, eviction, offloading, sharing and prefix reuse.
Why it matters
KV capacity directly limits concurrency. When it runs out, requests queue or get preempted and recomputed.
Scenario
80 GB GPU, 60 GB model: only ~20 GB remains for KV. Thousands of concurrent users exhaust it quickly.
Staff answer
“KV-cache management is fundamentally a memory-management problem. I care about allocation, fragmentation, eviction, offloading and reuse, because KV capacity directly limits concurrency and can be the reason requests queue or fail.”
Play the lab →
09 · OPTIMIZE

Prefix caching + routing

“Every request starts with the same 5,000-token system prompt. What do you do?”

Answer out loud first, then reveal
What it is
Compute the shared prefix once, cache its KV, reuse it, and route requests to workers that already hold it.
Why it matters
Removes redundant prefill work, cutting TTFT and GPU cost. Routing must preserve locality or the cache is wasted.
Scenario
Enterprise support bot: system prompt + company docs + user question. The first 5,000 tokens are identical across requests.
Staff answer
“Prefix caching reuses KV state for shared prompt prefixes. Prefix-aware routing then sends requests to workers that already hold the relevant cache, reducing redundant prefill and improving TTFT and GPU efficiency.”
Play the lab →
10 · OPTIMIZE

Disaggregated inference

“Why would you separate prefill and decode onto different GPU pools?”

Answer out loud first, then reveal
What it is
Prefill pool computes KV, transfers it, and a decode pool generates tokens.
Why it matters
The phases have different resource profiles; separating them lets each scale independently and stops prefills from stalling decodes.
Scenario
Morning: many long prompts → scale prefill. Later: long answers → scale decode.
Staff answer
“Disaggregated inference separates prefill and decode onto independently managed worker pools. The goal is to match compute resources to each phase and scale them independently, with KV transfer becoming a key networking consideration.”
Play the lab →
11 · OPTIMIZE

Scheduling

“How is inference scheduling different from ordinary load balancing?”

Answer out loud first, then reveal
What it is
Choosing which worker handles a request using load, KV locality, queue length, model, topology and SLOs.
Why it matters
The “least loaded” worker may have to recompute a prefix another worker already holds.
Scenario
GPU 1: 30% load, no relevant KV. GPU 2: 50% load, has the KV. GPU 2 can be the better choice.
Staff answer
“Inference scheduling isn’t CPU-style load balancing. I’d consider queueing delay, KV-cache locality, model variant, GPU capacity and SLOs, because the cheapest-looking routing decision can create expensive recomputation.”
Play the lab →
12 · PLATFORM

GPU scheduling

“How does Kubernetes place an 8-GPU inference pod, and what can go wrong?”

Answer out loud first, then reveal
What it is
Pods request nvidia.com/gpu; the scheduler places them on nodes with capacity.
Why it matters
Not every GPU is equal: type, memory, NVLink topology, NUMA, MIG and network all matter for performance.
Scenario
A TP=4 pod stays Pending with 6 free GPUs because they’re spread across nodes.
Staff answer
“GPU scheduling is placing inference workloads on nodes with the right accelerator capacity and topology, not simply finding an available GPU. For distributed inference, placement and GPU-to-GPU and network topology directly affect latency and throughput.”
Play the lab →
13 · PLATFORM

Autoscaling

“GPU utilization is 70%, but 500 requests are queued and TTFT is 8 s. What should drive autoscaling?”

Answer out loud first, then reveal
What it is
Scale on inference signals: queue length, KV usage, active sequences, tokens/s, TTFT.
Why it matters
CPU barely moves in LLM serving; GPU utilization saturates early and hides overload.
Scenario
GPU at 70%, queue at 500, TTFT at 8 s: clearly overloaded, and the HPA on CPU never fires.
Staff answer
“For LLM serving I don’t rely solely on CPU or GPU utilization for autoscaling. Queue depth, active sequences, KV-cache pressure, TTFT and token throughput are better indicators of user-visible saturation. I also plan for cold start.”
Play the lab →
14 · PLATFORM

Multi-node networking

“Why does networking matter so much for multi-node inference?”

Answer out loud first, then reveal
What it is
Distributed models move tensor or expert data between nodes every step.
Why it matters
Slow links leave expensive GPUs idle waiting on the network.
Scenario
Two 8-GPU nodes: GPU compute → wait → network → wait. Bandwidth, latency, RDMA, InfiniBand/RoCE, topology and congestion all matter.
Staff answer
“In multi-node inference, networking becomes part of the compute path because distributed collectives move large amounts of tensor or expert data. I treat bandwidth, latency, RDMA and topology as first-class performance constraints.”
Play the lab →
15 · PLATFORM

NCCL

“Throughput dropped after moving from one node to two. What do you investigate?”

Answer out loud first, then reveal
What it is
NVIDIA’s collective communication library: all-reduce, all-gather, reduce-scatter, broadcast, all-to-all.
Why it matters
TP, PP and MoE rely on collectives; their performance can become the limiting factor.
Scenario
1 node → 2 nodes, throughput falls: check NCCL performance, network bandwidth, RDMA and GPU topology.
Staff answer
“NCCL provides optimized collective communication primitives between NVIDIA GPUs. In distributed inference, NCCL performance can become the limiting factor because TP, PP or MoE workloads can be communication-heavy.”
Play the lab →
16 · PLATFORM

Observability

“What would you put on the dashboard for an LLM serving platform?”

Answer out loud first, then reveal
What it is
Infrastructure metrics (GPU, memory, power, network, NCCL) + inference metrics (TTFT, ITL, tokens/s, queue, active sequences, KV usage, batch size) + model metrics (tokens in/out, context length, errors).
Why it matters
Only the combination tells you whether you’re compute-, memory-, communication- or scheduling-bound.
Scenario
P99 TTFT jumps: correlate queue depth, prefix hit rate, KV usage and GPU/NCCL health to find which one moved.
Staff answer
“For LLM serving, observability has to connect infrastructure metrics with inference metrics. I correlate GPU utilization and memory with queue depth, TTFT, inter-token latency, token throughput and KV pressure, so I can tell whether the bottleneck is compute, memory, networking or scheduling.”
Play the lab →
17 · FRONTIER

MoE serving

“What makes mixture-of-experts models different to serve?”

Answer out loud first, then reveal
What it is
Only a subset of experts runs per token; experts may live on different GPUs.
Why it matters
Large capacity without activating every parameter, but token routing crosses the network and expert load can be uneven.
Scenario
Token → router → experts 7 and 31 on other GPUs → network → result.
Staff answer
“MoE serving reduces active compute per token by routing tokens to a subset of experts, but introduces distributed token routing and all-to-all communication. The main challenge is balancing expert load while keeping communication overhead low.”
Play the lab →
18 · FRONTIER

Multimodal serving

“A user uploads a screenshot and asks what’s wrong. What does the serving pipeline look like?”

Answer out loud first, then reveal
What it is
Image → vision encoder → embeddings → LLM → answer; audio and video add more stages.
Why it matters
Different modalities need different compute, so stages must be orchestrated and cached.
Scenario
A Kubernetes error screenshot: preprocess on CPU, encode on GPU, then LLM prefill over image + text tokens.
Staff answer
“Multimodal serving combines components such as vision encoders, audio encoders and language models. The infrastructure challenge is coordinating these stages efficiently while managing heterogeneous compute, latency and reusable context like encoder caches.”
Play the lab →
19 · FRONTIER

RL / rollout infrastructure

“How would you design infrastructure for RL post-training of a reasoning model?”

Answer out loud first, then reveal
What it is
Model generates rollouts → environment scores them → trainer updates weights → repeat.
Why it matters
Training GPUs shouldn’t sit idle waiting for rollout generation; weights must sync fast.
Scenario
Millions of rollouts: separate (or time-share) training and rollout clusters with fast weight sync and cache resets.
Staff answer
“RL infrastructure has two competing workloads: training and rollout generation. I’d treat rollout serving as a scalable inference workload and design the system so training isn’t starved by rollout latency or capacity constraints.”
Play the lab →
20 · FRONTIER

Agentic serving

“How does serving agents differ from serving a chatbot?”

Answer out loud first, then reveal
What it is
The unit of work is a multi-step workflow: LLM → tool → LLM → database → … → answer.
Why it matters
Latency, cost, retries, state, concurrency, context growth, failures, timeouts and security now span many calls.
Scenario
“Investigate why my Kubernetes service is failing”: Prometheus, logs, Kubernetes API, Git, then reasoning, across dozens of cached turns.
Staff answer
“Agentic serving changes the serving unit from a single model invocation to a multi-step workload involving models, tools, APIs and state. At the infrastructure layer I focus on orchestration, concurrency, state and context management, failure isolation, retries, observability and cost controls.”
Play the lab →

The vLLM question bank.

63 staff-level questions and 4 system-design scenarios with model answers, drawn from the vLLM Office Hours. Rate each one; filter to your weak spots before an interview.

0 / 67 rated✓ 0 nailed~ 0 shaky✗ 0 missed

Inference fundamentals

Most serving questions reduce to one fact: prefill is compute-bound, decode is memory-bandwidth-bound, and nearly every vLLM feature trades between the two.

Q1Why is prefill compute-bound and decode memory-bound, and why does that drive serving design?

Prefill pushes every prompt token through the model in one pass, so the weight matmuls are large and arithmetic intensity is high. Decode emits one token per sequence per step, so each step streams all the weights and that sequence's KV cache from HBM to do very little math.

Batching fixes half of this. Weight reads are shared by every sequence in the batch, so intensity grows roughly with batch size. KV reads are private to each sequence, so decode attention stays bandwidth-bound at any batch size.

Go deeper

Do the roofline out loud. An H100 SXM offers about 989 dense BF16 TFLOPS and 3.35 TB/s of HBM, a ridge near 300 FLOPs per byte. A decode weight GEMM does roughly B FLOPs per byte at batch B, so it stays bandwidth-bound until the batch holds a few hundred tokens.

Likely follow-up

Which vLLM features exist because of this split? Continuous batching and chunked prefill grow the decode batch. Speculative decoding spends idle decode compute to emit several tokens per weight read. P/D disaggregation stops compute-heavy prefills from stalling bandwidth-heavy decodes.

How did you do?
Q2How do you size the KV cache, and what does the number tell you?

KV bytes per token = 2 (K and V) × layers × KV heads × head dim × bytes per element. Llama 3.1 70B has 80 layers, 8 KV heads and head dim 128, so BF16 costs about 320 KiB per token. One 128K-token context therefore needs about 40 GiB.

Maximum concurrency ≈ free KV memory ÷ (average context × bytes per token). vLLM does this at startup: it profiles a forward pass, claims --gpu-memory-utilization of the GPU (default 0.9), gives what is left after weights and activations to KV blocks, and logs the resulting maximum concurrency.

Go deeper

Name each lever and its cost. FP8 KV halves bytes per token but needs scales and an accuracy check. GQA and MLA cut what is stored: DeepSeek-V3's MLA keeps a 576-value latent per layer, about 69 KiB per token in BF16. Under TP that latent is replicated on every rank, which pushed DeepSeek serving towards data-parallel attention.

Likely follow-up

What happens when KV runs out? The scheduler preempts running requests, frees their blocks and recomputes them later. With prefix caching on, freed blocks stay cached until evicted, so the recompute is often a cache hit. The KV offloading connector can also reload preempted requests from CPU memory.

How did you do?
Q3What problem does PagedAttention solve, and what else does the block abstraction buy?

Earlier servers reserved one contiguous KV region per request, sized for the maximum length. The vLLM paper found 60–80% of KV memory wasted that way. PagedAttention stores KV in fixed-size blocks (16 tokens by default) reached through a per-request block table, cutting waste to under 4%.

The block table then became the base for most later features:

  • Sharing: parallel samples and common prefixes point at the same physical blocks, with copy-on-write.
  • Prefix caching: full blocks are hashed, so a new request can reuse any cached prefix.
  • Movement: CPU offloading, P/D transfer and remote KV stores all move KV block by block.
  • Hybrid models: one allocator can hold attention KV, sliding-window KV and Mamba state if their pages share a size.
Go deeper

Block size is a trade-off. Small blocks cut fragmentation and make prefix hits finer, but they also make each transfer tiny. Before vLLM 0.12, one block's KV was split per layer (and often per K and V) into 8–72 KB pieces, which crippled CPU offloading. One contiguous block across all layers (0.44–2.5 MB) raised offload throughput by an order of magnitude.

Likely follow-up

What does paging cost? An extra indirection in the attention kernel, which gathers KV through the block table. Hybrid models also force odd block sizes, such as 672-token attention blocks sized to match a Mamba page.

How did you do?
Q4What is continuous batching, and what new problem does it create?

Continuous (iteration-level) batching reschedules after every forward step: finished sequences leave and waiting ones join at once. Static batching holds the whole batch until its longest sequence ends, idling the GPU for everyone else.

The new problem is interference. A long prompt admitted into a running batch slows that step for every decoding request, which shows up as inter-token latency spikes. vLLM answers with chunked prefill under a per-step token budget, a cap on prefill tokens per request per step, and, at scale, P/D disaggregation.

Go deeper

The token budget (--max-num-batched-tokens) is the main latency–throughput dial. A larger budget finishes prefills sooner, helping TTFT and throughput. A smaller budget bounds the step time that decodes see, helping ITL.

Likely follow-up

How would you tune it for a chat product with a P99 ITL target? Measure P99 step time against the budget and pick the largest budget that meets the target. Then stop long prompts from starving short ones with --long-prefill-token-threshold (see Agentic workloads).

How did you do?
Q5Which latency and throughput metrics belong on the SLO dashboard?
  • TTFT: queueing plus prefill; how long the user waits before anything appears.
  • ITL / TPOT: time between output tokens; ITL is the per-token distribution, TPOT the per-request mean.
  • End-to-end latency: TTFT + TPOT × (output tokens − 1).
  • Throughput: output and total tokens per second per GPU, and tokens per dollar.
  • Goodput: requests per second that meet every SLO at once, the target DistServe optimizes.
Go deeper

Report percentiles at a stated load, never averages at an unstated one. Better, sweep concurrency and plot throughput per GPU against per-user interactivity (tokens/s/user). The vLLM team's AgentX results report exactly that: the highest-throughput configuration that keeps P90 interactivity above 50 tokens/s/user.

Likely follow-up

Why is total tokens/s misleading for agentic traffic? It counts cached input tokens, which cost little. At a 96% prefix-hit rate, total tokens look huge while GPU time goes mostly to decode, so report cached, uncached and output tokens separately.

How did you do?
Q6Why can going from one GPU to two more than double throughput?

Weights are a fixed cost, and the KV cache gets whatever memory is left. In the January 2025 session, Llama 70B went from TP=1 to TP=2 and gained 13.9× more KV cache blocks and 3.9× more token throughput. More KV allows bigger batches, which raises decode arithmetic intensity.

Go deeper

The gain flattens once KV stops being the constraint, while TP's cost keeps growing. Megatron-style TP needs two all-reduces per layer, after attention's output projection and the MLP's down projection, on every decode step. Pick the smallest TP that gives enough KV headroom for your target concurrency and context, then scale out with replicas.

Likely follow-up

When would you use pipeline parallelism instead? When the model does not fit in one node and the inter-node link is slow. PP only sends activations at stage boundaries, but it does not cut per-token latency and needs enough micro-batches to keep stages busy.

How did you do?

vLLM V1 engine internals

V1 (2025) and Model Runner V2 (2026) tell the same story twice: as GPUs got faster, CPU work around each step became the bottleneck, so vLLM moved it off the critical path.

Q7Why did V1 split the API server and the EngineCore into separate processes?

Because CPU overhead, not the GPU, had become the limit. With Llama-8B on an H100 a forward pass takes about 5 ms, so HTTP handling, scheduling, input prep, detokenization and streaming ate a large share of each step. V1 runs an isolated EngineCore busy loop that only schedules and executes, while tokenization, multimodal preprocessing, detokenization and streaming overlap in the frontend process over ZeroMQ.

Tensor parallelism changed too. V0 kept the scheduler and worker 0 in one process, an asymmetric design. V1 workers cache request state and receive only per-step diffs, so the scheduler lives apart and every worker runs the same code on one GPU or many.

Go deeper

V1 delivered up to 1.7× the throughput of V0 with almost the same kernels, so the gain is nearly all CPU overhead removed. Show how you would spot a CPU-bound server: gaps between kernels in a profiler trace, step time well above kernel time, worst with small models on fast GPUs.

Likely follow-up

What does the multiprocess design cost? Serialization and IPC on every step, which is why only diffs are sent. It also adds failure modes, such as a dead worker process, and makes debugging span processes.

How did you do?
Q8How does the V1 scheduler decide what runs each step?

It hands out a token budget. The decision is a map of {request_id: num_tokens}, and prompt tokens and generated tokens are treated alike, so there is no separate prefill or decode phase. A decoding request asks for one token; a prefilling request asks for its remaining prompt, clipped to the budget. Chunked prefill falls out of this for free, and prefix caching simply lowers the tokens a request still needs.

Go deeper

Walk the loop. Running requests are served first, so decodes are not starved, then waiting requests are admitted while budget and KV blocks remain. If a running request cannot get a new block, the scheduler preempts one to make room. The main knobs are --max-num-batched-tokens, --max-num-seqs and --long-prefill-token-threshold.

Likely follow-up

How does speculative decoding fit this model? Draft tokens are just extra tokens scheduled for that request in the step. Rejected ones are rolled back by moving the request's computed-token count back.

How did you do?
Q9How does prefix caching work in V1, and why could it be on by default when V0's could not?

Only full blocks are cached. Each block's hash covers its parent block's hash, its own token IDs and extra keys (LoRA ID, multimodal input hashes, cache salt), so one hash identifies the whole prefix up to that block. A new request walks its block hashes and reuses the longest cached run.

Freed blocks join a doubly linked free queue in LRU order. Eviction pops the head, and a cache hit can pull a block from the middle in O(1). V0's version cost enough CPU to lower throughput at low hit rates, so it shipped off. V1 made eviction constant-time and cut Python object churn: under 1% throughput loss at a 0% hit rate and several-fold gains at high hit rates.

Go deeper

Security is part of the design. SHA-256 is the default hash, with xxHash as a faster option. A per-request cache_salt is mixed into the first block's hash, so only requests with the same salt share blocks, which blocks timing attacks that probe another tenant's prompts.

Likely follow-up

What kills the hit rate in production? Anything that changes early tokens: a timestamp or user ID at the top of the system prompt, reordered tool definitions, random request IDs. Context truncation shifts every block boundary. So does routing that scatters one session across replicas.

How did you do?
Q10What are piecewise CUDA graphs, and why did vLLM settle on FULL_AND_PIECEWISE?

CUDA graphs record a sequence of kernel launches once and replay it, removing per-kernel CPU launch cost, which matters when a decode step lasts a few milliseconds. Graphs need static shapes, and attention's per-step metadata varies. Piecewise graphs capture everything except attention and run attention eagerly between the captured pieces.

Full graphs also capture attention, which works for uniform decode-only batches. FULL_AND_PIECEWISE uses full graphs for decode-only steps and piecewise graphs for mixed steps. For hybrid models whose Mamba kernels are written in Triton, which has high launch overhead, it recovered V0 performance and gave up to 91% more throughput on granite-4.0-h-tiny.

Go deeper

Graphs are captured for a set of padded batch sizes. Each batch pads up to the next captured size, wasting some compute, and more capture sizes cost memory and startup time. torch.compile runs first and fuses ops such as SiLU-multiply-quantize, so fewer hand-written kernels are needed.

Likely follow-up

Why do graphs clash with some kernels? A captured launch grid replays as recorded. A grid sized for a bigger workload wastes work on every replay, so the Triton attention backend moved to persistent kernels with a fixed grid that reads its work from GPU memory.

How did you do?
Q11What is async scheduling, and what makes it hard?

Without it, the CPU schedules step N+1 only after step N's tokens come back, so the GPU idles during scheduling and input prep. Async scheduling plans step N+1 while step N still runs. The catch is that the scheduler does not yet know step N's results: which requests hit a stop token, how many draft tokens were accepted, what grammar state follows.

It must schedule optimistically and correct a step later, and every feature has to tolerate that lag. In V1, async was retrofitted, so many features needed awkward special cases. Model Runner V2 makes async the core assumption: GPU-side input prep consumes rejection-sampling results directly, with no CPU–GPU sync.

Go deeper

Name what breaks: stop conditions, speculative acceptance counts, structured-output state, logprobs, and any .item() or .cpu() call inside the step.

Likely follow-up

How would you find a hidden sync? Look for device-to-host copies and stream synchronizations inside the step in an Nsight or torch profiler trace. PyTorch's CUDA sync debug mode can raise an error or a warning at each one.

How did you do?
Q12What is Model Runner V2, and why rebuild the model runner a year after V1?

MRV2 (March 2026) is a ground-up rewrite of the component that turns scheduler output into GPU inputs, runs the model and samples. V1's runner had grown past 6,700 lines in one file. It coupled persistent request state to per-step inputs, had async scheduling bolted on, and did input prep and sampling as many small CPU operations.

  • Stable state table: each live request owns a fixed row, and a gather builds each step's ordered inputs.
  • GPU-native input prep: Triton kernels build input_ids, positions, query_start_loc and seq_lens on the device.
  • Async-first: designed for zero CPU–GPU sync, including with speculative decoding.
  • Triton sampler: Gumbel-max sampling without materializing softmax; top-k logprobs computed only for the selected candidates.
  • ModelState: model-specific logic (multimodal embeddings, attention metadata, CUDA graph capture) is isolated; the largest file is under 1,300 lines.

Qwen3-0.6B on one GB200 went from 16K to 25K output tokens/s (+56%). GLM-4.7-FP8 with MTP on 4× GB200 saw 6.3% lower TPOT.

Go deeper

MRV2 is opt-in (VLLM_USE_V2_MODEL_RUNNER=1). As of v0.18.0 it lacked linear-attention models, EPLB, DBO, logits processors, LoRA and draft methods beyond EAGLE, EAGLE-3 and MTP. A staff answer adds how you would adopt it: gate by that feature matrix, shadow traffic on both runners, compare accuracy and latency, then switch.

Likely follow-up

Why Gumbel-max sampling? Taking the argmax of logits/T plus Gumbel noise samples exactly from softmax(logits/T). The kernel never normalizes a vocabulary-sized softmax, and its stateless per-request RNG keeps sampling reproducible.

How did you do?

KV cache management

The KV cache is now a tiered, shared resource rather than a per-GPU buffer: GPU memory, then CPU DRAM, then distributed pools and disk, with policies deciding what stays warm.

Q13How does vLLM's hybrid KV cache manager handle models that mix attention types?

Layers of the same type (full attention, sliding window, Mamba or linear attention) form a KV cache group, and all groups draw from one shared block pool with a single page size. Memory then flows between layer types on demand instead of being split up front. V0 instead preallocated Mamba state per sequence, sized by max_num_seqs, so a wrong guess meant either OOM or low concurrency.

A uniform page is hard for Mamba. On Nemotron-Nano-12B-v2 one sequence's Mamba state is about 2.57 MiB, while a 16-token attention block is about 64 KiB. vLLM grows the attention block size until the pages match (672 tokens, for example) and pads Mamba pages slightly. It later decoupled the kernel's block size from the manager's, so TRT-LLM attention kernels still run on Blackwell.

Go deeper

Tell the war story. Attention and Mamba views shared one tensor with different layouts, so writing a block through one view corrupted another block in the other view; re-striding the Mamba state fixed it. DeepSeek V4 later exposed layout fragmentation again (92 separate tensors), fixed by packing each block into one contiguous allocation, which also saved about 10% of KV memory with the FP4 indexer.

Likely follow-up

Why do hybrids win at long context? A Mamba state is fixed-size, while attention KV grows with every token. At 128K tokens one attention layer's KV is roughly 200× a Mamba state, so only the few attention layers pay the growing cost.

How did you do?
Q14Walk me through the KV offloading connector. Why copy with DMA instead of a kernel?

The connector API lets vLLM ask an external store for KV before scheduling a request and hand it newly computed KV afterwards. Since v0.9.0 it is asynchronous, so loads and stores overlap model compute. The offloading connector (v0.11.0) adds a pluggable backend whose core is one transfer function; the CPU backend uses cudaMemcpyAsync, which runs on the GPU's DMA engines.

A copy kernel wins on tiny blocks but competes with the model for SMs. Once blocks were 0.5–2 MB, DMA reached 83.4 GB/s bidirectional versus 68.5 GB/s for the kernel. End to end, DMA gave 5.5–15% more throughput on the worst-case model and up to 32% on Llama-3.1-8B; at a 0% hit rate the kernel was 6% slower than no offloading at all.

Go deeper

Quote the payoff and say which one matters. Loading from CPU cut single-request TTFT 2–22× depending on prompt size. With 10,000 concurrent 512-token prefills, throughput rose up to 9× while TTFT improved only 2×, so the main win is throughput. Enable it with --kv_offloading_backend native --kv_offloading_size <GB>.

Likely follow-up

When does CPU offloading not help? When there is neither reuse nor preemption to save. Also when PCIe or host memory bandwidth is already busy, and, as of early 2026, for hybrid models, which the connector did not yet optimize.

How did you do?
Q15Reload or recompute: how do you decide?

Compare transfer time with prefill time. Llama-3.1-8B stores 128 KiB of KV per token, so 10K tokens is about 1.3 GB, roughly 26 ms at the 50 GB/s one-way rate measured on H100. Prefilling those 10K tokens costs about 1.6×10¹⁴ FLOPs (2 × 8B × 10K), roughly 0.3 s at a realistic 500 TFLOPS. Reload wins by about 10×, and the gap widens with context because attention compute grows quadratically while transfer grows linearly.

Go deeper

The ratio depends on KV bytes per FLOP. Models with a lot of KV per parameter (full multi-head attention, small models) favour recompute sooner. Remote tiers change the maths again: LMCache measured loads from S3 Express at 22–32% lower TTFT than a full prefill.

Likely follow-up

And for short prompts? Fixed costs dominate: lookup, scheduling and transfer setup. Below a few hundred tokens, recompute is usually simpler and just as fast.

How did you do?
Q16What do LMCache and Mooncake add beyond local CPU offloading?

They make KV reusable across engines and extend capacity past one machine. LMCache is a KV layer outside the engine for cross-query reuse and P/D transfer. It moves KV in 256-token chunks instead of 16–64 KB pages, reaching 400 Gbps from CPU where small transfers manage about 88 Gbps. It adds zero-copy writes, layer-by-layer pipelining with a one-layer GPU buffer, and lookup, move, pin and compress APIs for routers.

Mooncake Store, integrated in 2026, is a distributed KV pool. In standalone-store mode an external Mooncake client owns each node's CPU memory and disk, even on CPU-only nodes, and vLLM workers just request blocks. Routers such as Dynamo and llm-d can then send a request to any instance and still get a hit.

Go deeper

Bring the field data. LMCache reported 2.3–14× higher throughput than plain vLLM, and 1.9–8.1× lower TTFT at low QPS. In enterprise deployments, truncating long contexts dropped the prefix hit ratio from about 85% to 45%, and coding and RAG workloads reached about 50% hits from reused context, not just shared system prompts.

Likely follow-up

So can routing ignore locality now? No. A remote hit still moves KV, and prefetched blocks hold GPU capacity while in flight; on AgentX traces, session-sticky routing beat load balancing for exactly this reason. A shared pool is also one more stateful distributed system to run.

How did you do?
Q17How does prefix caching work for hybrid models, where the state is not per token?

A Mamba or linear-attention layer keeps one recurrent state that is valid only at one position. Reusing a prefix needs a snapshot of that state exactly at the prefix boundary, and snapshotting every token is too costly. vLLM combines two retention policies:

  • Interval-based: keep the prompt-end state at every turn, so the next turn or a forked subagent can extend it.
  • Marconi-style selective: when a prefix is seen a second time with no checkpoint, recompute the state and save one at that boundary for later requests.
Go deeper

Explain why intervals alone miss hits: shared prefixes often end mid-turn, for example a common system prompt plus tool definitions. Sliding-window layers are the easier case; only the last window's blocks matter, and blocks outside it can be freed.

Likely follow-up

What changes for P/D disaggregation with hybrids? The transfer must carry the Mamba convolution and SSM state as well as attention KV. vLLM extended its NIXL connector in 2026 to move both through separate descriptor views.

How did you do?

Parallelism and distributed inference

The right parallelism follows three questions: does the model fit, where is the bottleneck (compute, bandwidth or communication), and how does the attention architecture store KV.

Q18How does tensor parallelism shard a transformer layer, and what does it cost?

Megatron-style TP splits the QKV and MLP up/gate projections by columns, so each GPU owns a slice of heads or features. Element-wise ops such as SiLU run on those slices. The attention output and MLP down projections split by rows, each GPU produces a partial sum, and an all-reduce combines them: two all-reduces per layer.

Each GPU reads only 1/N of the weights per step, which effectively multiplies memory bandwidth and cuts decode latency. The price is an all-reduce on every layer of every step, which needs NVLink-class links; decode messages are small, so latency rather than bandwidth dominates.

Go deeper

Watch the KV heads. With 8 KV heads, TP above 8 forces vLLM to replicate KV heads across ranks, wasting KV memory. For small decode messages, vLLM uses its own intra-node all-reduce kernel instead of NCCL, and newer paths write directly into peers' memory through symmetric-memory buffers.

Likely follow-up

How else would you cut TP communication cost? Fuse the all-reduce with the RMSNorm or quantization that follows it, or fuse the GEMM with a reduce-scatter; Kimi K3's GEMM + reduce-scatter fusion is a 2026 example.

How did you do?
Q19How do you choose among TP, PP, DP, EP and context parallelism?

Work through it in order:

  1. Fit: the smallest TP inside one node that holds the weights plus enough KV for your target concurrency and context.
  2. Too big for a node: PP across nodes when inter-node links are slow; stretch TP across nodes only on a fast fabric such as an NVL72 domain.
  3. Scale out: data-parallel replicas, which need no cross-replica traffic, behind a cache-aware router.
  4. MoE: expert parallelism for the experts, usually with data-parallel attention (DEP), especially for MLA models.
  5. Long context: decode context parallelism (DCP) for long-context decode, prefill context parallelism (PCP) for long prompts.
Go deeper

The 2026 lesson is that parallelism follows the architecture and the topology. Kimi K3 did best with DCP8 on 8-GPU nodes, but wide EP (DEP16) scaled better on NVL72 once each rank held more than 3 requests. For DeepSeek V4, DCP only matched DEP, so DEP became the default.

Likely follow-up

How would you settle it empirically? Sweep each candidate across concurrency and compare throughput per GPU at your interactivity SLO. The Pareto frontier decides, not peak throughput.

How did you do?
Q20Why is TP a poor fit for MLA models, and what is wide EP?

MLA compresses K and V into one shared latent, in effect a single KV head, so TP cannot split the cache by heads and every rank stores the full latent. Data-parallel attention fixes this: each rank runs attention only for its own requests with its own KV, so KV capacity grows with the number of ranks.

The MoE layers then use expert parallelism. Each token is sent all-to-all to the ranks holding its routed experts and combined afterwards; DeepSeek-R1 activates only 37B of its 671B parameters per token. Wide EP combines EP with DP across many GPUs (--enable-expert-parallel). vLLM supports DeepEP's high-throughput and low-latency all-to-all kernels, Perplexity's kernels and an NCCL all-gather/reduce-scatter path.

Go deeper

Quote the result: 2.2k tokens/s per H200 on a CoreWeave InfiniBand cluster, up from about 1.5k, thanks to kernel fusions and dual-batch overlap. Under TP, DeepSeek-V3 on H200 left 34 GB per GPU for KV, but every rank had to store the same latents, so effective capacity was far lower.

Likely follow-up

What is the downside of wide EP? Every rank in the group must step together, idle ranks run dummy passes, and imbalance becomes idle time. The failure domain also grows, since one GPU can stall the group; elastic EP is on the roadmap for that reason.

How did you do?
Q21What is dual-batch overlap, and when does it help?

DBO splits a step's batch into two micro-batches and overlaps one micro-batch's MoE all-to-all with the other's compute, following DeepSeek's micro-batching design. With --enable-dbo, ranks first all-reduce to agree that splitting is worth it (threshold --dbo-decode-token-threshold). Worker threads then run each micro-batch, each yielding while its dispatch or combine is in flight.

It helps when communication dominates, as with high EP degrees, where profiles showed dispatch and combine outweighing the small decode compute. It hurts tiny batches, because splitting halves GEMMs that are already small, hence the threshold.

Go deeper

DBO, together with kernel fusions such as SiLU-multiply-quantize, took DeepSeek decode from about 1.5k to 2.2k tokens/s per H200.

Likely follow-up

How would you prove it is helping? Compare profiles before and after: all-to-all should move from exposed gaps to overlapped regions. Then compare TPOT at fixed concurrency, and check small-batch latency has not regressed.

How did you do?
Q22What is EPLB, and how does it rebalance experts without a restart?

Real traffic routes tokens unevenly across experts, so in wide EP some ranks idle while others overload, and lockstep means the busiest rank sets the pace. vLLM implements DeepSeek's expert-parallel load balancer (--enable-eplb). Each MoE forward pass records per-expert load, and a sliding window aggregates it across ranks. At each rebalance interval, EPLB computes a new logical-to-physical expert map, optionally with redundant copies of hot experts, and shuffles weights while serving.

Go deeper

Name the costs. Redundant experts take HBM that could hold KV. Weight shuffles use interconnect bandwidth and can blip latency. A short window thrashes and a long one goes stale; the hierarchical policy balances within node groups first to limit cross-node traffic.

Likely follow-up

How do you measure imbalance? The ratio of maximum to mean tokens per rank for each MoE layer, plus step-time spread across ranks, before and after rebalancing.

How did you do?
Q23What are decode and prefill context parallelism, and when do they beat TP?

Both shard along the sequence. DCP gives each rank 1/N of every request's KV, so long-context decode attention, which is memory-bound and grows with context, is split across ranks without replicating MLA latents. Each decode layer then needs a query gather before attention and a merge of partial outputs afterwards. PCP instead shards a long prompt across ranks: for DeepSeek V4, PCP8 prefilled a 32K prompt 2.65× faster than TP8.

For Kimi K3, DCP8 beat TP8 on decode latency and concurrency. vLLM fused DCP's communication into the attention kernels through symmetric-memory buffers, bypassing NCCL and cutting about 13% per layer. PCP still replicates decode-side state, so it suits dedicated prefill workers.

Go deeper

Know where it loses. At multi-node DCP sizes the collectives outweigh the savings, and on NVL72 DEP16 beat DCP8 once each rank held more than 3 requests. DeepSeek V4's compressed sparse attention has three sublayers to partition, so DCP only matched DEP there.

Likely follow-up

Why can partial attention outputs be merged exactly? Each rank returns its output with a log-sum-exp. Rescaling by the log-sum-exps and summing gives exact softmax attention, the same online-softmax maths FlashAttention uses across tiles.

How did you do?
Q24Why do expert-parallel ranks run in lockstep, and what follows?

Every MoE layer's all-to-all needs all ranks at once, because a token on one rank may need an expert on another. The group moves at the pace of its slowest rank, and ranks with no work run dummy passes just to join the collectives.

  • Prefill interference: one long, compute-bound prefill on any rank delays decode on every rank, which is why P/D disaggregation pays off more for EP deployments.
  • Prefill cadence: prefills landing on different steps stall the group repeatedly, so --prefill-schedule-interval admits prefills only every Nth step, on a counter aligned across ranks.
  • Imbalance: hot experts or uneven requests per DP rank turn directly into idle time, so EPLB and DP-aware routing matter.
Likely follow-up

What else does disaggregation buy an EP deployment? Each pool can use the DeepEP kernel suited to it: high-throughput all-to-all for prefill, low-latency for decode.

How did you do?

Disaggregated serving and llm-d

Disaggregation buys latency control, not raw throughput; the throughput win comes from specializing and rate-matching each pool, and from routing requests to where their KV already lives.

Q25Why disaggregate prefill and decode, and when is it not worth it?

The vLLM docs give two reasons: tune TTFT and ITL separately, and control tail ITL, because prefills can no longer interrupt decodes. They are also blunt that disaggregated prefill "DOES NOT improve throughput" on its own. The throughput gains seen at scale come from specializing each pool, for example PCP prefill workers and wide-EP decode workers with low-latency all-to-all kernels.

It is rarely worth it for short prompts, small fleets where a dedicated pool sits idle, or networks without RDMA-class bandwidth for KV transfer. It pays most for long-prompt workloads and MoE deployments, where one prefill stalls a whole lockstep EP group.

Go deeper

Price the transfer. Llama 70B stores about 320 KiB of KV per token, so a 10K-token prompt moves about 3.3 GB. At 50 GB/s (a 400 Gb/s NIC) that adds about 65 ms to TTFT, unless transfer is pipelined layer by layer during prefill.

Likely follow-up

What happens when the KV reaches the decode side? In vLLM's NIXL connector, the decode instance allocates blocks and pulls the KV over RDMA while the request waits. It then tells the prefill instance to free its copy; a timeout frees it if the decode side never arrives.

How did you do?
Q26How does the KV connector API work, and where does NIXL fit?

One interface serves P/D transfer, CPU offloading and remote caches, and it has two halves:

  • Scheduler side: get_num_new_matched_tokens says how many tokens can be loaded instead of computed, update_state_after_alloc reacts once blocks exist, build_connector_meta tells workers what to load and save this step, and request_finished manages transfers still in flight.
  • Worker side: start_load_kv begins async loads, wait_for_layer_load blocks per layer so transfer can pipeline with compute, save_kv_layer and wait_for_save handle stores, and get_finished reports completed transfers.

In-tree connectors include NIXL, LMCache, Mooncake, MoRI-IO (ROCm), FlexKV, the offloading connector and MultiConnector, which composes several, such as NIXL for P/D plus CPU offloading. NIXL, NVIDIA's inference transfer library, hides the transport (RDMA, NVLink, storage) behind one asynchronous read/write API.

Go deeper

Mention the operational edges. Since late 2025 the NIXL handshake checks a hash of the vLLM version, connector version and relevant config, so prefill and decode pools must upgrade together. Different TP sizes on the two sides mean re-sharding KV heads during transfer, which the roadmap lists as ongoing work.

Likely follow-up

Why is the lookup a scheduler-side call? The scheduler must know how many tokens are already available before it spends budget and allocates blocks. Otherwise it would schedule compute that the transfer makes unnecessary.

How did you do?
Q27How do you choose the prefill-to-decode ratio?

Rate-match the pools: prefill capacity in requests per second must equal decode capacity at your SLOs, or one pool queues while the other idles. The vLLM team uses two phases:

  1. Saturation profiling: benchmark prefill-only and decode-only deployments separately, sweeping parallelism (TP or wide EP) and size (8, 16, 32 GPUs) until throughput saturates.
  2. P/D sweep: derive the ratio from those saturation points, then sweep concurrency on the combined deployment.
Go deeper

Show the arithmetic with labelled, illustrative numbers. Suppose a prefill instance sustains 80K prompt tokens/s and the average uncached prompt is 8K tokens: 10 requests/s. A decode instance holding 256 sequences at 40 ms TPOT emits 6,400 tokens/s; at 500 output tokens that is 12.8 requests/s. So you need about 1.3 prefill instances per decode instance.

Likely follow-up

What if the prefix hit rate doubles? Uncached prompt tokens fall, prefill capacity roughly doubles, and the ratio swings towards decode. That is why agentic traffic, with hit rates above 96%, is decode-heavy, and why each pool should autoscale on its own.

How did you do?
Q28Why does round-robin load balancing fail for LLM serving?

The llm-d design lays out four reasons:

  • Expensive, uneven requests: RAG has long inputs and short outputs, reasoning the reverse. Imbalance compounds: an overloaded replica's ITL rises, requests stay longer, and load grows further.
  • Cached state: sending a multi-turn or agentic request to the replica holding its prefix skips most of its prefill.
  • Different phases: compute-bound prefill and bandwidth-bound decode favour specialized replicas.
  • Different QoS: code completion needs milliseconds, while batch summarization tolerates hours.
Go deeper

Name good and bad routing signals. Good: waiting-queue depth, KV cache utilization, prefix overlap with each replica's cache, predicted latency. Bad: CPU or GPU utilization and connection counts, which say little about LLM load.

Likely follow-up

What goes wrong with pure prefix-affinity routing? A popular prefix, such as a shared system prompt, piles onto one replica. You need load-aware tie-breaking, the idea behind llm-d's "sticky until saturated" routing.

How did you do?
Q29How does llm-d's inference scheduler pick an endpoint?

llm-d plugs into Kubernetes' Gateway API Inference Extension. For each request the gateway, such as Envoy, calls an Endpoint Picker (EPP). The EPP runs filters (healthy pods, prefill or decode role) and then weighted scorers (prefix-cache affinity, KV cache utilization, queue depth) and returns the best vLLM pod. Prefix awareness can be approximate, from the router's own history, or precise, from KV cache events that vLLM emits as blocks are stored and evicted.

On two 8×H100 nodes with a long-input, short-output multi-turn benchmark, llm-d cut mean TTFT about 3× at 4 QPS (Llama 4 Scout, 2 replicas). It served about 50% more QPS within a P95 TTFT of 2 s (Scout, 4 replicas) and 2× the QPS within that SLO (Llama 3.1 70B, 4 replicas).

Go deeper

The design keeps the gateway generic and puts LLM-specific judgment in the EPP, so scorers can evolve: 2026 releases added predicted-latency scheduling, token-aware sticky routing and router high availability.

Likely follow-up

Why not route inside one big vLLM deployment instead? The EPP works across heterogeneous pods, versions and hardware, and plugs into Kubernetes rollouts, autoscaling and multi-model routing.

How did you do?
Q30llm-d, Dynamo, Ray Serve LLM or the vLLM production stack: how would you choose?

All four run vLLM underneath; they differ in the platform they assume and what they optimize.

StackBuilt onStrengthsPick it when
llm-dKubernetes, Gateway API Inference ExtensionCache- and load-aware scheduling, P/D and wide-EP "well-lit paths"You run Kubernetes and want upstream-aligned parts
NVIDIA DynamoOrchestration above vLLM, SGLang or TensorRT-LLMKV-aware routing, KV Block Manager tiers, Planner that matches load dynamicallyYou want multiple engines behind one layer on NVIDIA fleets
Ray Serve LLMRay and KubeRayP/D, data-parallel attention, prefix-affinity routing; ties into Ray data and RLYour platform already runs on Ray
vLLM production stackHelm charts on KubernetesReference router, LMCache integration and observabilityYou need a simpler, quick-start deployment
Go deeper

Give the staff criteria: the platform you already run, hardware vendors, whether you need several engines, your team's operating capacity and each project's upstream velocity. Keep an OpenAI-compatible API and engine-neutral metrics at the edge so you can switch later.

Likely follow-up

When would you build your own router? For unusual policy: per-tenant quotas and priority classes, cross-region placement or cost-aware routing across model sizes. Even then, start by adding scorers to an EPP rather than writing a new proxy.

How did you do?

Speculative decoding

Speculative decoding trades spare compute for fewer memory-bound steps, so it shines at low load and can hurt at high load; the 2026 shift is from sequential to parallel drafting.

Q31How does speculative decoding speed up generation without changing the output?

Decode is memory-bound, so scoring k+1 positions in one target forward pass costs about the same as generating one token. A cheap drafter proposes k tokens and the target checks them all in one pass. Rejection sampling accepts a prefix of the drafts and samples a correction at the first rejection, and the acceptance rule makes the output distribution identical to sampling from the target alone.

Go deeper

Quantify it. With per-token acceptance rate α and k drafts, the expected tokens per verification step is (1 − α^(k+1)) / (1 − α). At α = 0.8 and k = 4 that is about 3.4 tokens per step, and each extra draft token adds less, which is why k stays small. The real speedup divides that by one plus the drafter's cost relative to a target step.

Likely follow-up

What changes at temperature 0? A draft token is accepted only if it equals the target's argmax, which is simpler and still matches plain greedy decoding, apart from floating-point differences between batch shapes.

How did you do?
Q32Compare draft models, n-gram lookup, EAGLE-3, MTP and the 2026 parallel drafters.
MethodHow it draftsBest forWatch out for
Draft modelA small LM with the same vocabulary proposes tokens one by oneGeneral text when a matching small model existsNeeds the same tokenizer; its own weights and KV
n-gram (prompt lookup)Copies what followed a matching n-gram in the promptSummarization, code edits, RAG that quotes sourcesNothing to copy, nothing gained
EAGLE-3A one-layer head reads hidden features from several target layersChat, math, RAGTied to one target model; weak outside its training data
MTPThe model's own multi-token prediction headsModels that ship them, such as DeepSeek and GLMDraft depth fixed by the model
P-EAGLE, DFlash, DSparkA whole block of drafts in one forward passLonger drafts and larger draftersMore memory; results vary by workload
Go deeper

Explain why parallel drafting matters. Sequential drafters pay one forward pass per draft token, so they must stay tiny and k needs retuning as load changes. DFlash projects the target's hidden states into the drafter's KV cache and drafts by block diffusion; DSpark adds a correction head and a confidence head that drops weak drafts before verification.

Likely follow-up

Why not tree-based drafting? Trees help at low concurrency, but at higher request rates they slow generation significantly, and vLLM did not support them as of mid-2025.

How did you do?
Q33When does speculative decoding hurt, and how would you gate it in production?

It spends compute to save bandwidth, so it loses once batches are large enough to be compute-bound. In 2024 tests on Llama3-70B with 4×H100, a draft model was 1.5× faster at QPS 1 on ShareGPT but 1.4× slower at high QPS. n-gram drafting went from 2.8× faster to 1.8× slower on CNN/DailyMail. Content matters too: an EAGLE-3 head trained on chat data reached 2.1× on math and RAG but did poorly on German-to-English translation.

  • Measure: draft acceptance rate, per-position acceptance and mean acceptance length, all exported since v0.9.1.
  • Tune by load: longer drafts at low request rates, shorter at high ones; translation-like traffic may want k of 0 or 1, RAG 5 or more.
  • Scope it: enable it on latency-sensitive routes and the decode pool, not on batch pools.
  • Test at real load: compare TPOT and throughput per GPU at your target concurrency, never only at QPS 1.
Go deeper

vLLM's research direction here is dynamic speculation, which shortens drafts as load rises, less so when acceptance is high.

Likely follow-up

Acceptance is 70% but there is no speedup. Why? The drafter is too expensive relative to the target, or its TP communication dominates (a draft can run at TP=1 beside a TP=4 target). Other suspects: the drafter is not captured in CUDA graphs, or the batch is already compute-bound.

How did you do?
Q34How do you get a good drafter for your own fine-tuned model?

Train one against your target on traffic that looks like yours. Speculators, a Red Hat library in the vLLM project, standardizes the loop: extract target hidden states offline with vLLM, train an EAGLE-3, P-EAGLE, DFlash or DSpark drafter, and save it in a Hugging Face-compatible format. vllm serve <speculator> then reads the config and turns speculation on, or you pass --speculative-config with the method and num_speculative_tokens.

Go deeper

Treat the drafter as an artifact coupled to its target. Retrain when the target changes, version the two together, measure acceptance per task category on held-out production-like prompts, and roll out behind a canary that watches acceptance and TPOT.

Likely follow-up

Does quantizing the target break the drafter? It shifts the target's distribution slightly, so acceptance can drop. Measure it, and retrain on the quantized target's hidden states if it does.

How did you do?
Q35How does speculative decoding interact with the scheduler, KV cache and CUDA graphs?
  • Scheduler: a request's k drafts are extra tokens in the step's budget, and rejected positions are rolled back afterwards.
  • KV cache: blocks must have room for k+1 new tokens; KV written for rejected drafts is simply overwritten later.
  • CUDA graphs: captured sizes must cover (k+1) × batch tokens. EAGLE gained CUDA graph support in v0.9.1, and Model Runner V2 can capture several draft passes in one graph.
  • Async scheduling: the CPU does not know how many drafts were accepted until the step ends, so MRV2 lets GPU-side input prep consume the rejection-sampling results directly.
Go deeper

Speculation stacks with disaggregation. The vLLM team's GLM-5.2-NVFP4 deployment on 24 B300s combined P/D disaggregation, MTP speculation and Model Runner V2 to cut mean TPOT from 40 ms to 17 ms.

Likely follow-up

What does it cost in memory? Separate drafters need weights and their own KV cache, and parallel drafters are larger still. That memory comes out of the target's KV budget, lowering maximum concurrency.

How did you do?

Quantization and kernels

Pick precision by the bottleneck at your operating point, prove accuracy on your own tasks, and prefer portable kernels unless a vendor kernel earns its maintenance cost.

Q36Weight-only or weight-and-activation quantization: when do you use each?

Match the scheme to the bottleneck.

  • Weight-only (W4A16, W8A16; GPTQ, AWQ): 4-bit weights cut weight memory about 4×, which is what memory-bound decode and small batches need. A mixed-input GEMM dequantizes weights inside the kernel and does the maths in 16-bit.
  • Weights and activations (W8A8 in FP8 or INT8): runs on low-precision tensor cores, doubling peak FLOPs over BF16, which helps compute-bound prefill and large batches. FP8 needs Hopper or Ada; Blackwell adds FP4 formats such as NVFP4 and MXFP4.

Hopper needed a new mixed-input kernel. Marlin, built for Ampere, used older mma instructions and lost about 37% of Hopper's peak compared with wgmma. Machete (CUTLASS 3.5.1) pre-shuffles weights offline for 128-bit shared-memory loads, keeps weights in registers through a transposed product, and uses TMA with warp specialization. It served Llama 3.1 70B on one H100 at 5 requests/s with TTFT under 250 ms and TPOT under 100 ms.

Go deeper

Say which wins where. At batch 1, W4A16 moves the fewest bytes and is fastest. At large batches the work is compute-bound, the 16-bit maths is half FP8's rate, and dequantization adds overhead, so FP8 W8A8 usually wins. More than 20% of vLLM deployments already used quantization in 2024.

Likely follow-up

What changes for MoE? Expert weights dominate memory, so MoE uses block-scaled FP8 or MXFP4 expert kernels, and per-expert batches are small, which keeps expert GEMMs bandwidth-bound.

How did you do?
Q37How do you validate a quantized model before it ships?

Pass two gates on your own workload: accuracy and performance. Red Hat's validated-model process is a good template: capacity-planning runs with GuideLLM, accuracy runs with the Language Model Evaluation Harness, and serving with vLLM across accelerators.

  • Accuracy: compare against the BF16 baseline on suites that match real use (reasoning, maths, code, long context, languages), plus product checks such as tool-call validity and JSON-schema adherence.
  • Performance: benchmark at your real prompt and output lengths and concurrency, because a scheme that wins at batch 1 can lose at batch 256.
  • Calibration: produce checkpoints with versioned LLM Compressor recipes (GPTQ, SmoothQuant, AWQ and others), calibrated on data that resembles production traffic.
Go deeper

Averages hide regressions. Break results out by task and context length, check KV cache quantization separately from weight quantization, and set explicit recovery thresholds with the product owner before running anything.

Likely follow-up

Who should own this pipeline? The platform team: versioned recipes, reproducible calibration, automated eval gates in CI, and registry metadata recording scheme, calibration set and scores for every checkpoint.

How did you do?
Q38What does an FP8 KV cache buy, and what can go wrong?

It halves KV bytes per token, roughly doubling the concurrency or context that fits, and halves the bytes each decode step reads, which speeds memory-bound attention. The risks are accuracy and support: quantization error accumulates over long contexts, scales must be calibrated or computed on the fly, and backends and GPUs differ in what they support.

Go deeper

Validate on long-context tasks such as needle-in-a-haystack and long-document QA, not only short benchmarks. Check each attention backend you run, because FP8 KV paths are backend-specific code.

Likely follow-up

FP8 KV or CPU offloading? They solve different problems and stack. FP8 fits more active sequences on the GPU and speeds decode; offloading keeps inactive prefixes warm off the GPU.

How did you do?
Q39Why build a Triton attention backend when FlashAttention exists?

For portability and maintainability. Hand-tuning kernels for every GPU generation and vendor does not scale, while one Triton kernel of about 800 lines, against roughly 70,000 for FlashAttention 3, runs unchanged on NVIDIA, AMD and Intel. On H100 it reached 100.7% of FA3's performance on long decode requests, and on MI300 it was about 5.8× faster than earlier implementations. It is the default on ROCm and the fallback everywhere else.

  • Q-blocks: group all query heads that share a KV head, plus several query tokens, into one work item so tl.dot gets large enough tiles.
  • Parallel tiled softmax ("3D kernel"): split the KV traversal across instances for decode, then reduce in a second kernel; heuristics decide when the extra launch pays.
  • Persistent kernels: launch a fixed grid sized to the GPU and pull work from metadata, so CUDA graph replays do not repeat wasted waves.
Go deeper

Explain the wave problem. A captured grid replays at its recorded size; if it spilled into a second wave of SMs, that waste repeats on every replay, even when the real workload is smaller.

Likely follow-up

What is Helion? PyTorch's higher-level kernel language, in effect tiled PyTorch. An experimental paged-attention kernel took 133 lines in Helion against 295 in Triton.

How did you do?
Q40How does vLLM support non-NVIDIA hardware without forking?

Through hardware plugins. A vendor implements vLLM's platform interface (device, attention backend, communication, custom ops) in its own package, such as vllm-ascend for Huawei Ascend NPUs, the Tenstorrent plugin or vllm-metal for Apple Silicon, while the core stays shared. AMD ROCm lives in-tree with its own kernels and the Triton attention backend. By 2024 vLLM already ran on NVIDIA, AMD, Google TPU, AWS Inferentia and Trainium, Intel Gaudi and GPUs, and x86, ARM and PowerPC CPUs.

Go deeper

The cost is the feature matrix. Some optimizations exist only as vendor kernels, so a model × hardware × feature combination can silently fall back to a slow path. Define a supported matrix, gate releases on accuracy and performance tests per backend, and route each model to hardware where it is validated.

Likely follow-up

How would you evaluate a new accelerator? Run your top models at your SLOs through the vendor's plugin. Check how fast new models get support, parity on prefix caching, speculation, structured outputs and P/D, whether the plugin keeps pace with vLLM releases, and cost per token at SLO.

How did you do?

Model architectures: MoE, MLA, sparse and hybrid attention, multimodal

Each new architecture moved a different bottleneck, and the engine had to change its memory layout, kernels or process model to follow.

Q41What makes MoE models different to serve?

Memory scales with total parameters while compute scales with active ones: DeepSeek-R1 activates 37B of its 671B parameters per token. Every expert must be resident somewhere, yet each token touches only a few. So per-expert batches are small: tokens per expert ≈ batch × top-k ÷ number of experts. DeepSeek-V3 sends each token to 8 of 256 routed experts, so a 256-token decode batch gives each expert about 8 tokens, a bandwidth-bound GEMM over full expert weights.

The responses follow: large global batches through wide EP with DP attention, grouped-GEMM expert kernels, block-scaled FP8 or MXFP4 expert weights, and EPLB for skewed routing.

Go deeper

Contrast TP and EP for experts. TP splits every expert across GPUs, giving many tiny GEMMs plus an all-reduce; EP keeps whole experts on each GPU and moves tokens all-to-all instead. EP wins at scale, while TP is simpler on one node. vLLM's modular MoE layer keeps the all-to-all backend separate from the expert compute kernel, so each can change on its own.

Likely follow-up

Why do prefill and decode diverge even more for MoE? Prefill brings many tokens, so every expert gets a healthy, compute-bound batch; decode brings few, so it is bandwidth-bound and dominated by communication. That is why DeepEP ships separate high-throughput and low-latency kernels, and why P/D pays off.

How did you do?
Q42What is DeepSeek Sparse Attention, and what did it change in vLLM?

DeepSeek-V3.2 adds a lightning indexer in front of MLA. For each query token the indexer scores earlier tokens with its own small key cache and keeps the top 2,048; sparse MLA then attends only to those, or to everything when the context is shorter. Beyond 2K tokens, attention cost per token stops growing, apart from the indexer's cheap scan.

  • A second cache: indexer keys live in their own buffer, separate from MLA's, and this model requires 64-token blocks.
  • A new FP8 layout: each token's MLA entry is 656 bytes: 512 FP8 latent values, 16 bytes of scales and 128 bytes of unquantized BF16 RoPE values.
  • Batching: the indexer has separate prefill and decode paths, and batched prefill tracks each request's start and end so tokens never score another request's context; shorter results are padded with -1.
  • Kernels: DeepGEMM's indexer kernels, FlashMLA's sparse attention and a fused top-k.
Go deeper

Explain why this is a systems problem, not just a kernel. Continuous batching mixes requests of different lengths, and every token's top-k set is different, which complicates CUDA graphs and context parallelism. The launch recipe was P/D disaggregation over NIXL, with DP routing inside each prefill and decode instance.

Likely follow-up

Why is the indexer cheaper than attention? It keeps small FP8 keys and computes only scores, a ReLU-weighted dot product over a few heads, never a softmax-weighted sum over values. The expensive step runs on 2,048 tokens instead of the whole context.

How did you do?
Q43What does serving hybrid Mamba or linear-attention models require from the engine?

Hybrids such as Qwen3-Next, Nemotron Nano 2, MiniMax-Text-01 and Granite 4.0 keep a few full-attention layers and make the rest linear, because attention's KV grows linearly with context and its prefill quadratically. The engine needs three things:

  1. State management: a fixed-size recurrent state per sequence, held in the same block pool as attention KV (Q13).
  2. Launch overhead control: many Mamba and linear-attention kernels are Triton, with high CPU launch cost. V1 with piecewise graphs was slower than V0 at low concurrency; FULL_AND_PIECEWISE fixed it, with 2–18% more throughput on Nemotron-Nano-12B-v2 and up to 91% on granite-4.0-h-tiny.
  3. Feature adaptations: prefix caching needs state checkpoints (Q17), P/D must transfer the state, and speculative decoding must roll the state back, because rejected tokens have already updated it.
Go deeper

Tie it to workloads: RAG, agent loops and reasoning traces all grow context fast, which is where hybrids pay off.

Likely follow-up

Give an example of feature-matrix gating. As of v0.18.0, Model Runner V2 did not support linear-attention models such as Qwen3.5 and Nemotron 3 Super, so those deployments stayed on the V1 runner.

How did you do?
Q44How does vLLM keep multimodal models fast?

It treats preprocessing and the encoder as separate, cacheable work:

  • Off-loop preprocessing: image decoding, resizing and cropping run in a separate process, with a cache for repeated inputs, so the GPU loop never waits on them.
  • Encoder cache: vision embeddings are kept after the encoder runs, so a long text prefill can be chunked across steps without re-running the encoder. V0 had to process the image and text in one step.
  • Multimodal prefix caching: image hashes join token IDs in block hashes, so multi-turn chats about the same image hit the cache.

V1's speedup over V0 was larger on Qwen2-VL than on text models.

Go deeper

At scale the encoder becomes a stage of its own. Encode/prefill/decode (E/PD) disaggregation puts vision encoders on separate workers; llm-d has written up heterogeneous E/PD serving of Kimi-VL, with SGLang as the engine.

Likely follow-up

What breaks image prefix hits? Anything that changes the bytes: client-side resizing, re-encoding or metadata. Normalize media before it reaches the server.

How did you do?
Q45What is vLLM-Omni, and why does any-to-any serving need a stage graph?

Omni models are pipelines of different models. Qwen3-Omni has a Thinker (about 30B, MoE, with audio and SigLIP2 vision encoders) for reasoning and text, a Talker (about 3B, MoE) that turns the Thinker's hidden states into audio codec codes, and Code2Wav, a ConvNet vocoder. vLLM-Omni describes the model as a graph of stages, each with its own devices, memory share and parallelism, behind one OpenAI-compatible API (vllm serve <model> --omni).

Autoregressive stages reuse PagedAttention, continuous batching, CUDA graphs and the scheduler. A connector splits control from data: on one node, payloads move through shared memory and only a handle crosses the control plane; across nodes, RDMA or TCP. Chunks stream so later stages start before earlier ones finish. On Qwen3-TTS-1.7B it reached a real-time factor of 0.17 against 2.64 for Hugging Face Transformers.

Go deeper

Pick metrics per stage. TTFT and TPOT suit a text reasoner, not a vocoder or a diffusion model, where real-time factor and time to first audio chunk matter more. A text-only request runs only the Thinker.

Likely follow-up

Where does it get hard? Capacity balancing across N stages, the multi-stage version of the P:D ratio problem, plus per-request paths that skip stages and backpressure between stages.

How did you do?

Structured outputs, tool calling and semantic routing

Constraints guarantee syntax, not good answers; parsers and routers are part of the model contract and need their own evals.

Q46How does structured output decoding work in vLLM V1, and how do you pick a backend?

The engine compiles the constraint (choice list, JSON schema, regex, grammar or structural tags) into a grammar state machine. Every step it turns the allowed next tokens into a bitmask, applies it to the logits before sampling, and advances the grammar state with the chosen token. In V0, compiling a grammar stalled the whole engine; in V1 compilation is non-blocking, handled by the scheduler while other requests keep decoding.

  • XGrammar: caches compiled grammars; lowest TPOT for repeated schemas and long generations.
  • Guidance (llguidance): computes constraints token by token; faster TTFT on complex or changing schemas, suiting multi-tenant traffic.
  • auto (default): picks a backend per request.

With reused schemas, TPOT was only marginally higher than unconstrained generation.

Go deeper

Say where the cost lands: the first request with a new schema pays compilation in TTFT, mask computation takes CPU every step (XGrammar needs thread tuning, and too many threads hurt), and the mask depends on the last token, which complicates async scheduling and speculation. Also note the quality risk: forcing JSON before the model has reasoned can lower answer quality.

Likely follow-up

One tenant's huge schema slows its own requests. What do you do? Pre-register and cache schemas, cap schema complexity, prefer Guidance for one-off dynamic schemas, and isolate the tenant if CPU contention spreads.

How did you do?
Q47How does tool calling work in vLLM, and where does it break?

The chat template renders the tool definitions into the prompt, the model emits a call in its own format, and a model-specific parser (--tool-call-parser; more than 20 ship in tree, including Hermes, Mistral, Llama 3 JSON, Qwen and DeepSeek) converts it into OpenAI-style tool_calls. --enable-auto-tool-choice switches this on. The guarantee depends on tool_choice:

  • required or a named function: vLLM uses structured outputs, so calls always match the tool's JSON schema: valid, not necessarily sensible.
  • auto: the model decides, and arguments can occasionally be malformed unless strict constraints apply, which needs VLLM_ENFORCE_STRICT_TOOL_CALLING (on by default) and at least one tool marked strict: true.
  • none: tool calling is off.
Go deeper

Name the break points: a template and parser that drift apart after a model upgrade, partial JSON while streaming, reasoning text mixed with calls, parallel calls, and custom formats that need a --tool-parser-plugin. Pin model, template and parser versions together, and run a tool-call eval suite on every release.

Likely follow-up

How do you make agents robust to bad calls? Use strict schemas for critical tools, validate on the server, return structured errors so the model can retry, and track parse failures per model version as a regression metric.

How did you do?
Q48What is the vLLM Semantic Router, and when would you put one in front of your fleet?

It is a routing layer for mixture-of-models systems: it picks which model serves each request. Requests flow client → Envoy → an ExtProc filter running the Go router → the chosen backend. Version 0.1 (Iris, January 2026) extracts six signal types (domain, keyword, embedding similarity, factual, user feedback and preference), and a decision engine combines them with AND/OR rules and priorities. Plugins add semantic caching, jailbreak and PII detection, hallucination detection, system-prompt injection and header rewriting.

Use one when you run models of very different cost and quality, want safety checks at one chokepoint, or need to cut spend. In Red Hat's walkthrough, 86% of requests stayed on a local quantized Qwen model and 14% went to Gemini, with a P50 routing latency of 40 ms against 800–11,000 ms of inference.

Go deeper

Name the risks. Misrouting degrades quality silently, so each route needs evals and feedback signals. Classification adds to TTFT, semantic caches can return false hits, and routing policies need versioning and canaries like code. Keep one stable public model name so routes can change behind it.

Likely follow-up

How does this relate to llm-d's EPP? Different layers that compose. The semantic router chooses the model; the EPP chooses the replica of that model: Envoy → semantic router → inference gateway EPP → vLLM pod.

How did you do?

Agentic workloads

Agent traffic is long, multi-turn and almost entirely cached, so KV residency and routing locality, not FLOPs, decide cost; several textbook optimizations failed on it in 2026.

Q49What makes agentic traffic different from chat, and why does it matter?

SemiAnalysis's AgentX benchmark, built from real agentic coding traces, shows the shape:

  • Long sessions: a median of 43 turns per session.
  • Long inputs, short outputs: a median of 142K input tokens and 444 output tokens per request.
  • Heavy reuse: a prefix-cache hit rate above 96%.
  • Subagents: 44% of sessions spawn at least one, with a median of four among those.

Each turn appends the latest tool result and resends the whole context, so inputs keep growing while each turn adds only a short new prefill. The vLLM team names three challenges: keeping many sessions' KV warm between turns, executing long-context steps fast enough, and finding a P:D ratio when hit rates and context lengths vary by session.

Go deeper

Explain the economics. With over 96% of input cached, a turn costs its decode, a small prefill and the memory to keep its KV resident, so KV capacity becomes the scarce resource. On AgentX, vLLM reached up to 130K total tokens per GPU-second on DeepSeek V4 Pro and up to 376 tokens/s per user on MiniMax M3.

Likely follow-up

Which metric would you optimize? Tokens per GPU-second, or per dollar of TCO, at a P90 interactivity floor such as 50 tokens/s per user, with cached, uncached and output tokens reported separately.

How did you do?
Q50A long fresh prompt is stalling short agent turns. How do you fix it?

That is head-of-line blocking. With first-in, first-out chunked prefill, one long prefill can claim the entire token budget step after step, so short cached turns on the same rank never get scheduled. Cap the prefill tokens one request may take per step with --long-prefill-token-threshold: at 512, short turns join every step. On DeepSeek V4 Pro with B300s, this raised total tokens per GPU-second by up to 93% and P90 interactivity by about 2.3×.

The trade-off is a higher TTFT for the long request itself, so TTFT-sensitive deployments should use a larger threshold. In DEP deployments, also set --prefill-schedule-interval so prefills land on the same steps across ranks and the steps between are decode-only.

Go deeper

Offer the structural fix: route first turns, which need long fresh prefills, and later turns, which append a little to a cached prefix, to different pools. Each side can then use its own engine settings and parallelism, such as PCP for the cold prefills.

Likely follow-up

Why not use priority scheduling instead? Priority reorders the queue, but a long prefill that is already running still consumes budget every step. The cap bounds each request's share of a step, which is what protects short turns.

How did you do?
Q51Why did session-sticky routing beat load balancing for agents?

In aggregated DEP deployments the team saw large KV-usage imbalance across ranks and tried balancing by queue depth, running tokens or KV utilization. On AgentX traces, all three lost to simple session-aware sticky routing. Inter-turn gaps are short, so the next turn often arrives while its prefix is still resident on the previous GPU. Moving the session forces a KV fetch from the shared pool, and although the transfer overlaps compute, the prefetched blocks occupy GPU capacity, so the destination admits fewer sequences.

The queue looked balanced while the system served fewer concurrent requests. Route on the state already resident on each worker, not only on queued work.

Go deeper

Sketch the policy: stick by session ID until a KV-usage or queue threshold, then spill to the least-loaded rank with the best prefix overlap, and watch hit rate and P90 interactivity together. llm-d's "sticky until saturated" routing follows the same idea.

Likely follow-up

Where does the session ID come from? From the agent harness. vLLM's Q3 2026 roadmap adds Session-ID and Correlation-ID hints so the engine can recognize multi-turn agents and their subagents.

How did you do?
Q52Which plausible optimizations failed on agentic workloads, and why?

The vLLM team published three "bitter lessons":

  • Pipeline parallelism, including chunked PP: it scales well on long, fresh prompts, but warm agentic turns add only a few hundred to a few thousand new tokens. There is too little work to fill the pipeline, and bubbles eat the gain.
  • Decode context parallelism on DeepSeek V4: it works for pure and hybrid MLA models such as Kimi K3, but V4's attention has an indexer, a compressor and the main attention to partition. After heavy optimization it only matched DEP.
  • Load balancing: it lost to session-sticky routing (Q51).
Go deeper

Draw the meta-lesson: intuitions must survive end-to-end measurement on realistic traces. Synthetic, uniform benchmarks hide exactly these effects, so characterize the workload before investing in an optimization.

Likely follow-up

How could you have predicted the PP result? Measure the distribution of uncached prefill tokens per request first. If the median new prefill is a few thousand tokens, the pipeline-fill arithmetic already says PP cannot win on warm turns.

How did you do?
Q53What would you add to the serving stack for agents next?

Make the agent's structure explicit to the engine:

  • Agent hints: harnesses pass session structure, branch points, cache positions, expected tool-call latency and session lifecycle; the Q3 2026 roadmap starts with Session-ID and Correlation-ID.
  • Programmable KV cache: workloads choose how KV is prefetched, evicted or soft-pinned.
  • Session-based KV management: use the gap between turns to move retained KV towards the worker likely to serve the next turn.
  • Split routing: send first turns and later turns to differently configured pools.

The same roadmap targets production-ready distributed, multi-tier KV offloading and speculative decoding with acceptance lengths above 5 tokens.

Go deeper

Treat hints as untrusted input. They must be optional, capped per tenant so nobody pins the whole cache, and carry identifiers rather than content.

Likely follow-up

What is the risk of pinning? Tenants hoarding KV. Pinning needs quotas, TTLs, fair eviction under pressure and billing that reflects the memory held.

How did you do?

Production operations

Operating LLM serving means measuring with realistic load, scaling on LLM-native signals, and treating every engine upgrade as a model change.

Q54How do you benchmark an inference deployment credibly?
  • Realistic workload: traces or datasets that match your prompt and output length distributions and your prefix sharing; uniform random prompts hide cache and scheduling effects.
  • Load sweep: raise request rate or concurrency to saturation, and at each point report P50, P90 and P99 TTFT and ITL plus throughput per GPU; plot throughput against interactivity.
  • Controlled state: say whether the prefix cache is warm, keep versions, flags and hardware fixed, and record them.
  • Tools: vllm bench serve, GuideLLM for capacity planning, and public references such as SemiAnalysis's AgentX harness.
  • Accuracy beside speed: quantization, speculation and new kernels can all change outputs.
Go deeper

Name the common benchmark lies: speedups measured only at QPS 1, fixed-length outputs from ignoring end-of-sequence, single requests on an empty server, and mismatched quantization between contenders. Prefer open-loop arrivals at a request rate to fixed concurrency when you care about queueing, because a closed loop hides it.

Likely follow-up

How do you size a cluster from the results? Take throughput per GPU at the SLO point, divide peak demand by it, and add headroom for failures and bursts. Do it separately for prefill and decode pools if disaggregated.

How did you do?
Q55What signals should drive autoscaling, and why is it hard for LLMs?

Scale on LLM-native signals: waiting-queue depth (vllm:num_requests_waiting), KV cache usage (vllm:kv_cache_usage_perc), SLO attainment for TTFT and ITL, and preemption rate. GPU utilization misleads, because a GPU can be busy yet stalled on memory, and it says nothing about KV headroom. With disaggregation, scale each pool on its own signal: queue and TTFT for prefill, KV usage and ITL for decode.

Cold start is the hard part. A replica must load tens to hundreds of GB of weights, compile and capture CUDA graphs before it can serve, which takes minutes. Keep warm pools, reuse compile caches, stream weights from peers or fast storage, and scale ahead of predictable peaks; reducing cold start is on vLLM's Q3 2026 roadmap.

Go deeper

llm-d's autoscaling design goes further: measure each variant's capacity, derive a load function over request shapes and QoS classes, and compute the optimal mix of prefill, decode and latency-tolerant instances.

Likely follow-up

Why can scaling out briefly make latency worse? New replicas start with cold prefix caches, so traffic moved to them loses its hits. Ramp them in gradually and let cache-aware routing warm them.

How did you do?
Q56How would you roll out a new vLLM version to production safely?

vLLM moves fast: from v0.16 in February 2026 to v0.22 in June, a release every two to three weeks, and defaults change along the way.

  1. Pin everything: vLLM, PyTorch, CUDA and kernel libraries, in immutable images.
  2. Diff behaviour: read release notes for changed defaults and deprecations, and compare the effective engine config.
  3. Accuracy gate: per-model evals, including tool calls and structured outputs, against the current version.
  4. Performance gate: fixed benchmark scenarios per model at your SLO points, with regression thresholds.
  5. Canary, then progressive rollout by pool, with rollback images kept warm.
  6. Upgrade coupled parts together: prefill and decode pools (the NIXL handshake checks versions and config), drafters with their targets, LoRA adapters with their bases.
Go deeper

Mirror upstream's own discipline. vLLM's Q3 2026 roadmap targets a 30-minute CI signal, automatic quarantine of flaky tests per hardware backend and wider performance-regression coverage.

Likely follow-up

What triggers a rollback? Thresholds agreed in advance: error rate, P99 TTFT or ITL regression beyond an agreed percentage at matched load, an eval drop beyond tolerance, or crash loops. SLO breaches should roll back automatically.

How did you do?
Q57P99 TTFT jumped after a deploy. How do you debug it?

Split TTFT into its parts: queueing, KV load or transfer, prefill compute, and frontend work such as tokenization, image preprocessing and grammar compilation. Then check, roughly in order:

  1. Queueing: is the waiting queue longer? Fewer KV blocks after a memory change, or more preemptions, both cut capacity.
  2. Cache: did the prefix hit rate drop? Look for a template change, such as a timestamp at the top of the prompt, a tokenizer change, or routing that lost session stickiness.
  3. Scheduling: did token-budget or long-prefill defaults change, or did speculation turn on at high load?
  4. Frontend: is the API server CPU-bound on long prompts, images or new schemas?
  5. Disaggregation: is KV transfer slower, or has the P:D ratio drifted with traffic?
Go deeper

A P99-only regression usually means a subset: long prompts blocking short ones, one tenant's schemas, one bad node, or routing skew. Slice by prompt length, tenant, node and route before theorizing, and use end-to-end traces from router to engine.

Likely follow-up

How do you stop it recurring? Add the missing signal (hit rate, queue-time histogram), a canary check on it, and a benchmark scenario that reproduces the regression.

How did you do?
Q58How do you serve many fine-tuned variants cost-effectively?

Use multi-LoRA on shared base weights. vLLM batches requests for different adapters in the same step with batched LoRA kernels and loads adapters on demand, capped by --max-loras on the GPU, --max-cpu-loras in host memory and --max-lora-rank. The LoRA ID is part of every block hash, so adapters never share KV by mistake. vLLM has since added fused MoE LoRA kernels for MoE bases.

Go deeper

Know the limits. A long tail of adapters thrashes the per-batch adapter limit, so route by adapter affinity as you would by prefix. Very high-traffic variants are cheaper merged into their own deployment, and full fine-tunes cannot use LoRA at all.

Likely follow-up

And for many full models? Multiplex them. vLLM's sleep mode offloads or discards weights and frees KV so another model can wake on the same GPUs, at the cost of reload time.

How did you do?

System design scenarios

Staff design rounds reward clarifying the workload first, deriving routing, parallelism and memory from it, and naming what you would measure; all four scenarios reuse one reference layout.

Sc ADesign an inference platform for about 20 open models shared by many product teams

Clarify first: traffic per model (requests/s, prompt and output length distributions, prefix sharing), SLO tiers (interactive TTFT and ITL versus batch deadlines), the GPU inventory and vendors, isolation and compliance needs, and how often models change.

Design:

  • Edge: an Envoy gateway with auth, per-tenant quotas and priority classes behind one OpenAI-compatible API, with stable model aliases so backends can change.
  • Placement by traffic: the top few models get dedicated vLLM pools with cache-aware routing; long-tail fine-tunes share bases through multi-LoRA; rarely used models share GPUs through sleep mode or scale to zero.
  • Engine defaults: the smallest TP with KV headroom, prefix caching on, FP8 where evals pass, speculation only on interactive pools.
  • Routing: llm-d's EPP picks the pod; add a semantic router only if teams want automatic model choice.
  • Scaling and release: per-pool autoscaling on queue depth and KV usage, warm replicas for top models, and release gates for every engine or model change.

Trade-offs to name: dedicated pools isolate tenants but strand capacity, while shared pools raise utilization but need quotas and priorities. Multi-LoRA is cheap until adapter churn thrashes batches. P/D is worth it only for the one or two models with long prompts at scale.

What you would measure: SLO attainment per tier, GPU-hours per million tokens (cached and uncached separately), prefix hit rate, preemptions and cold-start time.

Likely follow-up: One tenant floods long prompts. What protects everyone else? Per-tenant token-rate quotas at the gateway, the long-prefill cap in the engine, a lower priority class for that tenant, and a separate pool if it persists.

How did you do?
Sc BServe a DeepSeek-class MoE model at high throughput across many GPUs

Clarify first: the interactivity target (tokens/s per user), prompt and context lengths, daily token volume, and whether the hardware is NVL72-class racks or 8-GPU nodes on InfiniBand.

Design:

  • Disaggregate: separate prefill and decode pools linked by NIXL over RDMA, rate-matched with saturation sweeps.
  • Decode pool: wide EP with data-parallel attention so MLA latents are not replicated, DeepEP low-latency all-to-all, dual-batch overlap, EPLB with a few redundant hot experts, and MTP speculation for latency.
  • Prefill pool: DeepEP high-throughput kernels, or PCP for very long prompts.
  • Precision: block-scaled FP8 weights and an FP8 latent cache, validated on long-context evals.
  • Routing: llm-d or Dynamo, cache-aware across DP ranks.

Trade-offs to name: a wider EP group gives more KV capacity and throughput, but more synchronization and a bigger failure domain. On NVL72-class systems, wide DEP scaled better than context parallelism once each rank held a few requests. The reference point is 2.2k tokens/s per H200 in a production-like llm-d deployment.

What you would measure: per-rank load imbalance, share of step time spent in all-to-all, P90 TPOT, and queue balance between the two pools.

Likely follow-up: One GPU in a 32-rank EP group fails. What happens? The whole group stalls, because every MoE layer needs every rank. You need fast health detection, routing away at the EPP, spare capacity to restart the group, and eventually elastic EP, which is on vLLM's roadmap.

How did you do?
Sc CServe an agentic coding assistant with 100K-token sessions at a fixed interactivity target

Clarify first: session length, turns and subagent fan-out, the interactivity floor, the cost target per completed task, the model's attention architecture, and data-residency rules.

Design:

  • Routing: session-sticky until a replica saturates, then spill to the best prefix overlap; the harness passes session IDs.
  • KV tiers: GPU, CPU offloading, a shared Mooncake or LMCache pool, then disk, with interval and selective retention for hybrid models.
  • Scheduling: cap long prefills, for example at 512 tokens per request per step, so short turns are never blocked; send first turns and later turns to differently configured pools.
  • Parallelism: follow the architecture, such as DCP for MLA-style models on 8-GPU nodes and DEP on NVL72.
  • P/D: size the pools with saturation sweeps, expecting decode-heavy ratios because most input is cached.

Trade-offs to name: stickiness against balance, retention storage against recompute, and the long request's TTFT against every short turn's latency.

What you would measure: hit rate by tier, tokens per GPU-second at a P90 interactivity floor, time to complete a task, and KV-pool bandwidth.

Likely follow-up: The CPU tier is full. What do you evict? Sessions least likely to return, judged by lifecycle hints and TTLs. Keep prompt-end checkpoints for active sessions, and consider FP8 KV to fit more before evicting.

How did you do?
Sc DBack-of-envelope: how many GPUs does this chat service need?

A chat service peaks at 50 requests/s with 2,000-token prompts, half of them a shared, cached system prompt, and 500-token outputs. The target is 40 ms TPOT on Llama 3.1 70B with FP8 weights and FP8 KV on H100 80 GB. These are practice assumptions, not measurements.

  1. Concurrency (Little's law): each request decodes for about 500 × 40 ms = 20 s, so about 50 × 20 = 1,000 sequences are in flight at peak.
  2. KV capacity: FP8 KV is 160 KiB per token. The shared prompt is stored once per replica, so each sequence holds about 1,250 unique tokens (about 195 MiB), or about 190 GiB for 1,000 sequences.
  3. Decode replicas: a TP=2 replica has 160 GB × 0.9 = 144 GB; minus 70 GB of weights and about 10 GB of activations, that leaves about 64 GB, roughly 300 sequences. Four replicas (8 GPUs) hold the working set.
  4. Step-time check: each GPU reads about 35 GB of weights and 47 GB of KV per step, because every sequence still reads the shared prefix unless cascade attention applies. That is about 25 ms at 3.35 TB/s, or 33 ms at 75% of peak: inside 40 ms, with little slack.
  5. Prefill compute: 50 × 1,000 uncached tokens × 2 × 70B ≈ 7 PFLOP/s. At about 900 effective FP8 TFLOPS per H100 (45% of peak), that is about 8 GPUs of compute.
  6. Combine: folding that prefill into the decode steps would push them well past 40 ms, so plan about 16 GPUs, either aggregated or split into about 8 prefill and 8 decode GPUs. Burst and N+1 headroom takes it to about 20–24 GPUs, three 8-GPU nodes.

What makes it a staff answer: stated assumptions; decode sized by KV and step time, prefill by FLOPs; noticing that a shared prefix saves memory but not bandwidth; and closing with "then I would confirm it with a load sweep at these shapes".

How did you do?

Staff-level judgment

These rounds probe judgment: choosing on evidence at your operating point, keeping options open, and turning engine work into a cost per token that leadership can weigh.

Q59vLLM, SGLang or TensorRT-LLM: how would you choose an engine?

On evidence at your operating point, not on reputation.

  • Model coverage: how fast new open models run well; vLLM shipped day-0 support for DeepSeek-V3.2's sparse attention.
  • Hardware breadth: vLLM covers NVIDIA and AMD GPUs, Google TPUs, AWS Trainium and Intel, plus vendor plugins; TensorRT-LLM is NVIDIA-only.
  • Performance at your SLO: run your own traces. vLLM's Q3 2026 roadmap claims TensorRT-level performance on top models, and public benchmarks such as SemiAnalysis's include it.
  • Ecosystem: llm-d, Dynamo, Ray Serve LLM and KServe build on vLLM; Dynamo can also run SGLang or TensorRT-LLM underneath.
  • Governance and backing: vLLM began at UC Berkeley's Sky Computing Lab in 2023, went to the Linux Foundation in July 2024 and became PyTorch Foundation-hosted in 2025. Its creators founded Inferact in January 2026 with $150M in seed funding.
Go deeper

Keep the engine swappable: an OpenAI-compatible API at the edge, engine-neutral metrics, and an orchestration layer that can host more than one engine. Weigh operating cost, meaning upgrades, debugging and community responsiveness, as heavily as peak numbers.

Likely follow-up

When would you not pick vLLM? When a specific model and hardware combination runs clearly better elsewhere at your SLO and the gain pays for running a second engine. Make it a measured exception, not a platform religion.

How did you do?
Q60How would you run your team's relationship with upstream vLLM?

Upstream by default; never fork.

  • Carry no patches you can avoid: each one taxes every upgrade, and vLLM ships a release every few weeks.
  • Engage through the SIGs: Core, Large-Scale Serving, RL, Model Performance and Quantization each have a Slack channel and tracking issues; bring an RFC before a big change.
  • Invest in CI: contribute tests and hardware coverage for the models and accelerators you depend on.
  • Track direction: the quarterly roadmap issues and the Office Hours.
Go deeper

The Q3 2026 roadmap names a new pressure: coding agents flooding maintainers with changes. Changes that land arrive with a design rationale, tests, benchmarks and an owner who will maintain them, and that is how a team earns influence upstream.

Likely follow-up

Upstream will not take a feature you need. Now what? Use an extension point instead of patching the core: a KV connector module, a tool-parser plugin, a hardware plugin or an out-of-tree model.

How did you do?
Q61How would you make the business case for disaggregation or wide EP?

Frame it as cost per million tokens at a fixed SLO, with a staged plan:

  1. Baseline: measure today's throughput-per-GPU versus interactivity curve and the cost per million tokens at your SLO.
  2. Target: model the gain from published references. DeepSeek decode on H200 rose from about 1.5k to 2.2k tokens/s per GPU through runtime and kernel work, so the same load needs about 32% fewer GPUs.
  3. Cost: engineering time, two pools to operate, RDMA networking for KV transfer, and new failure modes.
  4. Plan: pilot one high-volume model with success and kill criteria agreed up front, then roll out pool by pool.
Go deeper

Run a sensitivity analysis. The gain depends on prompt and output shapes and on hit rates, and agentic traffic shifts the P:D ratio. Re-benchmark every quarter, because engine releases move the baseline; sometimes a plain upgrade delivers the gain.

Likely follow-up

Why not just buy more GPUs? GPUs scale cost linearly, while efficiency work compounds across the fleet and cuts power and capacity risk. If the payback runs longer than the model or hardware refresh cycle, buying GPUs may well be right.

How did you do?
Q62How would you plan for heterogeneous accelerators?

Treat hardware as a validation and routing problem:

  • Validation matrix: model × hardware × feature, with accuracy and performance gates per cell.
  • Portable layers: vLLM's Triton attention backend (one kernel source across NVIDIA, AMD and Intel, default on ROCm), torch.compile and hardware plugins.
  • Placement: route each model or traffic class to the hardware where it is validated and cheapest per token at SLO; llm-d has written up serving across three GPU vendors in one cluster.
  • Trade-off: procurement leverage and supply resilience against the cost of running several driver, kernel and debugging stacks.
Likely follow-up

What breaks first? Feature gaps. A model's fast path may depend on vendor kernels, such as FlashMLA or DeepGEMM, so other hardware falls back to slower code. P/D across vendors also needs compatible KV layouts and transports.

How did you do?
Q63What does an inference engine need to support RL post-training?

RL generates rollouts from a policy that changes every step, so the engine becomes part of the training loop:

  • Fast weight sync from trainer to inference workers without restarts; vLLM's Q1 2026 roadmap lists modular weight-sync APIs for its RL SIG.
  • Memory handoff: sleep mode frees or offloads weights and KV so training and rollouts can share GPUs.
  • Cache invalidation: after a weight update, cached KV was computed with old weights, so the prefix cache must be reset.
  • Determinism and batch invariance: the same prompt should get the same logprobs whatever its batchmates; it is on vLLM's roadmap, and the Triton backend already supports it.
  • Scheduling: rollouts are long, batch-like generations; llm-d's co-operative time-slicing reports 40% more post-training experiments on the same GPUs.
Go deeper

Call out the silent bug: forgetting to reset the prefix cache after a weight update reuses stale KV, making rollouts subtly off-policy without any error.

Likely follow-up

Why does batch invariance matter? Floating-point reductions change with batch shape, so a request's logprobs can shift with its batchmates. RL importance ratios and eval reproducibility then drift for no visible reason.

How did you do?

Numbers worth memorizing.

Cover the right-hand columns and quote them from memory. For flip cards, use the vLLM page.

NumberWhat it measuresSource
~300 FLOPs per byteH100 SXM BF16 ridge point (989 TFLOPS ÷ 3.35 TB/s); decode sits far below itDerived (Q1)
320 KiB per tokenLlama 3.1 70B KV cache in BF16 (160 KiB in FP8); 128K tokens ≈ 40 GiBDerived (Q2)
60–80% vs under 4%KV memory wasted by earlier servers vs with PagedAttentionvLLM launch blog
16 tokensDefault KV block sizeKV offloading blog
13.9× and 3.9×KV blocks and throughput gained going from TP=1 to TP=2 on Llama 70BDistributed inference recap
Up to 1.7×V1 throughput over V0V1 blog
Under 1%V1 prefix caching's throughput cost at a 0% hit rateV1 blog
16K → 25K tokens/sModel Runner V2 vs V1, Qwen3-0.6B on one GB200 (+56%)MRV2 blog
2–22× and up to 9×TTFT cut and throughput gain from CPU KV offloading, Llama-3.1-8B on H100KV offloading blog
83.4 vs 68.5 GB/sBidirectional GPU–CPU copy rate, DMA vs a copy kernel, H100KV offloading blog
2.57 MiB vs 64 KiBOne Mamba state vs one 16-token attention block, Nemotron-Nano-12B-v2Hybrid models blog
Up to 91%Throughput gain from FULL_AND_PIECEWISE CUDA graphs, granite-4.0-h-tinyHybrid models blog
37B of 671BDeepSeek-R1 parameters active per tokenWide-EP blog
2.2k tokens/s per H200DeepSeek wide-EP decode throughput, up from ~1.5kWide-EP blog
2.65×PCP8 vs TP8 prefill speed on a 32K prompt, DeepSeek V4vLLM x AgentX
~3×, +50%, 2×llm-d cache-aware routing: lower mean TTFT, and more QPS within SLO in two setupsllm-d announcement
1.5× and 2.8× faster; 1.4× and 1.8× slowerDraft-model and n-gram speculation on Llama3-70B at QPS 1 vs high QPSSpeculative decoding blog
Up to 2.1×EAGLE-3 speedup on maths and RAG, Llama 3.3 70BEAGLE-3 article
5 requests/s, TTFT < 250 ms, TPOT < 100 msLlama 3.1 70B W4A16 on one H100 with MacheteMachete article
~800 vs ~70,000 lines; 100.7%Triton attention vs FlashAttention 3 code size, and its speed relative to FA3 on long decodeTriton backend blog
2,048 tokensTokens each query attends to under DeepSeek Sparse AttentionDeepSeek-V3.2 blog
43 turns, 142K in, 444 out, >96% hitsMedian AgentX session shape and prefix-cache hit ratevLLM x AgentX
+93% and ~2.3×Tokens per GPU-second and P90 interactivity from a 512-token long-prefill cap, DeepSeek V4 Pro on B300vLLM x AgentX
85% → 45%Prefix hit ratio after truncating long contexts, in LMCache deploymentsLMCache paper
0.17 vs 2.64Real-time factor, vLLM-Omni vs Hugging Face Transformers on Qwen3-TTS-1.7BvLLM-Omni architecture

Rewatch by topic.

The vLLM Office Hours sessions behind the bank, newest first. Each row links the recording or recap.

Show the session map
DateSessionUse it for
Sep 3, 2026#57 vLLM-Omni project update and demosOmni-modal serving (Q45)
Jul 9, 2026#53 llm-d update and wide EP for agentic workloadsAgentic serving and wide EP (Q19–Q24, Q27, Q49–Q51)
Jun 25, 2026#52 vLLM Semantic Router: safer, faster, multi-model inferenceModel routing and safety (Q48)
Jun 11, 2026#51 vLLM v0.22, Speculators update, accelerating sparse MLAParallel drafting and sparse MLA (Q32, Q42)
May 28, 2026#50 GenAI with vLLM on Intel CPUsCPU inference and v0.21 (Q40)
May 18, 2026#49 Latest trends in AI agent applications and vLLMAgent workloads (Q49, Q53)
Apr 30, 2026#48 vLLM project and tool calling updatev0.20 and tool calling (Q47)
Apr 16, 2026#47 LLM Compressor updateQuantization tooling (Q37)
Apr 9, 2026#46 Intro to vLLM-OmniOmni-modal serving (Q45)
Feb 26, 2026#44 vLLM v0.16.0 release updateRelease cadence and open Q&A (Q56)
Feb 12, 2026#43 vLLM Triton backend deep divePortable attention kernels (Q10, Q39, Q62)
Jan 29, 2026#42 Deep dive into the CPU offloading connectorKV offloading (Q2, Q3, Q14, Q15)
Jan 15, 2026#40 Intro to SpeculatorsTraining draft models (Q32, Q34)
Dec 18, 2025#38 vLLM 2025 retrospective and 2026 roadmapProject direction (Q60)
Nov 6, 2025#36 Live from the Zürich vLLM meetupCommunity talks from Red Hat, IBM and Mistral AI
Oct 23, 2025#35 How to build and contribute to vLLMWorking upstream (Q60)
Oct 9, 2025#34 AI-powered vLLM Semantic RouterModel routing (Q48)
Sep 25, 2025#33 Hybrid models as first-class citizensHybrid KV management and CUDA graphs (Q10, Q13, Q17, Q43)
Aug 28, 2025#31 vLLM and LLM Compressor updateQuantization tooling (Q37)
Jun 12, 2025#27 Intro to llm-dDistributed serving on Kubernetes (Q1, Q28, Q29, Q55)
May 8, 2025#25 Structured outputs in vLLMGrammar-constrained decoding (Q46)
Apr 2025#24 Performance optimization of vLLM on Google TPUsTPU serving
Mar 27, 2025#22 Intro to vLLM V1V1 engine internals (Q4, Q7–Q9)
Feb 27, 2025DeepSeek and vLLMMLA, MoE and DeepSeek kernels (Q20, Q41)
Feb 6, 2025#19 Multimodal LLMs with vLLM V1Multimodal serving (Q44)
Jan 23, 2025Distributed inference with vLLMTP, PP and scaling (Q1, Q6, Q18, Q19)
Dec 19, 20242024 retrospective and 2025 roadmapProject history and hardware breadth (Q40)
Dec 5, 2024Machete, a mixed-input GEMM kernel for HopperWeight-only quantization kernels (Q36)
Nov 14, 2024Disaggregated prefill and KV cache storageP/D and external KV stores (Q16, Q25)
Oct 2024Speculative decoding in vLLMSpeculation basics and limits (Q31, Q33)
Sep 19, 2024Advanced techniques for maximizing vLLM performancev0.6.0–0.6.1 performance work
Aug 8, 2024Multimodal models in vLLMEarly multimodal support
Jul 25, 2024Model quantization for efficient vLLM inferenceQuantization basics (Q36)
Jul 9, 2024FP8 quantization deep diveFP8 weights and activations (Q38)
Jun 20, 2024One-year anniversary updateFP8, speculative decoding, vision API
Jun 5, 2024Office hours Q&AQuantization, 70B GPU usage, vLLM vs TGI
May 15, 2024First office hoursOpen Q&A with the maintainers