inference atlas

Track 04 · Intermediate · ~35 min

Kubernetes, but it understands tokens.

A plain Kubernetes Service balances connections. LLM serving needs a platform that knows about prompts, queues, KV caches and GPU topology. Walk a request through llm-d + vLLM, read a real pod spec, pack GPUs, tune an autoscaler and debug a live incident.

4.1Intermediatecore idea

Follow one request through llm-d.

Compare a naive Deployment + Service with llm-d: an Envoy gateway, an endpoint picker that scores pods on prefix cache, queue and KV usage, and separate prefill and decode pools linked by NIXL. Click any box to learn what it is.

Walkthrough · Kubernetes serving stack

llm-d: inference-aware routing on Kubernetes

GPU node A · 8× H100role=prefill
⇣ KV blocks · NIXL over RDMA ⇣
GPU node B · 8× H100role=decode

Client

Your app, agent or SDK calling an OpenAI-compatible API (/v1/chat/completions). It has no idea how many GPUs sit behind the endpoint.

Step 1/11
A user asks the support bot a question. The app sends POST /v1/chat/completions with model: "support-bot".

Gateway API Inference Extension

Standard Kubernetes APIs for inference routing: an InferencePool groups model-server pods, and an Endpoint Picker (EPP) chooses one pod per request through Envoy’s ext-proc.

llm-d

A Kubernetes-native distributed inference stack on vLLM: an inference scheduler (EPP scorers), P/D disaggregation, wide-EP “well-lit paths”, KV-cache-aware routing and autoscaling.

Two routers, two jobs

A semantic router picks which model (cheap vs strong). The EPP picks which replica of that model. They compose: Envoy → semantic router → EPP → vLLM pod.

Staff answer

“I keep the gateway generic and put LLM-specific judgment in the endpoint picker: filter by health, role and loaded adapters, then score on prefix overlap, queue depth and KV utilization. That runs across heterogeneous pods and plugs into Kubernetes rollouts and autoscaling, which one big monolithic deployment can’t do.”

4.2Intermediate

Read a vLLM pod spec line by line.

Click any highlighted line of this Deployment to see what it controls and which interview topic it connects to. The diagram shows the processes running inside the pod.

apiVersion: apps/v1kind: Deploymentmetadata: name: llama-70b-decode labels:spec: template: spec: nodeSelector: containers: - name: vllm args: ports: resources: limits: readinessProbe: volumeMounts:
Inside the pod
API server processHTTP · tokenize · detokenize · stream (OpenAI API on :8000)
EngineCore · schedulerbusy loop: token budget → {request: tokens} each step
EngineCore · KV cache managerpaged blocks · prefix hash · LRU free queue
KV connector · NIXLloads/saves KV blocks to other pods or tiers
4 GPU workers (TP ranks 0–3)NVLink all-reduce twice per layer
GPU0GPU1GPU2GPU3

Tensor parallelism = 4

This pod shards every layer across its 4 GPUs over NVLink. Must match the GPU limit below. Pick the smallest TP with enough KV headroom.

4.3Intermediatedrag & drop

Be the scheduler: pack the GPUs.

Drag each pod onto a node, or tap a pod and then tap a node. Tensor-parallel pods need all their GPUs on one NVLink node, and some need a specific GPU type. Place everything without stranding GPUs.

Pending pods
node-a8× H100 80GB · NVLink
node-b8× H100 80GB · NVLink · 4 GPUs used by a training job
node-c4× L4 24GB · small models
Tip: place the biggest pods first. That’s what bin-packing heuristics do.

What K8s knows

The NVIDIA device plugin / GPU Operator exposes nvidia.com/gpu as a countable resource, with node labels for GPU model. The default scheduler counts GPUs; it doesn’t see NVLink islands.

What you add

nodeSelector or affinity for GPU type, topology-aware placement, gang scheduling for multi-node groups (e.g. LeaderWorkerSet, Kueue, Volcano), and MIG partitions for small models.

Fragmentation

Four free GPUs spread over four nodes can’t host one TP=4 pod. Big jobs first, or a descheduler, keeps whole nodes free.

Staff answer

“GPU scheduling is placement on the right accelerator type and topology, not just finding a free GPU. For TP I need all ranks inside one NVLink domain; for multi-node groups I need gang scheduling so a half-placed group doesn’t sit on idle GPUs. Poor placement shows up directly as latency and throughput loss.”

4.4Intermediatesimulator

Scale on the right signal.

A morning ramp and a lunchtime spike hit a vLLM deployment. Pick the metric your autoscaler watches, set how long a new replica takes to become ready, and see who blows the TTFT SLO.

Simulator · 2 hours · each replica serves ~10 req/s

Same traffic, different autoscaling signal

Autoscaling signal
Demand vs ready capacity (req/s)
P95 TTFT (s) · SLO 2 s
Replicas
Minutes over TTFT SLO10 / 120
GPU-hours (8 GPUs/replica)151
Peak replicas16
Peak queue1.1K req
Queue depth · 6 min cold start
Waiting requests measure backlog directly, so desired replicas follow load plus backlog. Most of the remaining SLO misses come from cold start: try a warm pool.

Good signals

vllm:num_requests_waiting, vllm:kv_cache_usage_perc, TTFT/ITL SLO attainment, preemption rate. With P/D: queue and TTFT for prefill, KV usage and ITL for decode.

Misleading signals

CPU barely moves. GPU utilization pegs near 100% at moderate load, so it can’t tell “busy” from “drowning”.

Cold start

A new replica loads tens to hundreds of GB of weights and captures CUDA graphs: minutes. Use warm pools, compile caches, fast weight streaming, and scale ahead of known peaks.

Staff answer

“For LLM serving I don’t autoscale on CPU or raw GPU utilization. Queue depth, KV-cache pressure, TTFT and token throughput track user-visible saturation. I also budget for cold start: new replicas have cold prefix caches, so I ramp them in and let cache-aware routing warm them.”

4.5Advancedgame

You’re on call. What broke?

A live dashboard for a vLLM fleet. Start an incident, watch which panels move, then pick the root cause. Each answer explains the tell-tale signals, the way you’d narrate it in an interview.

Dashboard · vllm-prod · live

All systems normal

score 0 / 0
TTFT p990.6 s
ITL p9932 ms
Output tokens/s8.6K
Requests waiting6
KV cache usage62%
Prefix hit rate85%
GPU utilization69%
Preemptions / min0
NCCL bus bandwidth383 GB/s
On call
Watch the baseline for a moment: every panel has a faint line at its normal level. Then press Start an incident.
4.6Advanced

llm-d, Dynamo, Ray Serve, KServe, production-stack?

All of them can run vLLM underneath. They differ in the platform they assume and what they optimize. Answer three questions for a starting recommendation, then read the trade-offs.

StackBuilt onStrengthsPick it when
llm-dKubernetes, Gateway API Inference ExtensionCache- and load-aware scheduling (EPP), P/D and wide-EP “well-lit paths”You run Kubernetes and want upstream-aligned parts
NVIDIA DynamoOrchestration above vLLM, SGLang or TensorRT-LLMKV-aware routing, KV Block Manager tiers, Planner for dynamic P/DYou want several engines behind one layer on NVIDIA fleets
Ray Serve LLMRay and KubeRayP/D, data-parallel attention, prefix-affinity routing; ties into Ray Data and RLYour platform already runs on Ray
vLLM production-stackHelm charts on KubernetesReference router, LMCache integration, observabilityYou need a simpler quick-start deployment
KServeKubernetes serving CRDsStandard InferenceService API, vLLM runtime, multi-framework model servingYou serve many model types (not only LLMs) behind one API

Quick chooser

A starting point, not a verdict. Keep an OpenAI-compatible API and engine-neutral metrics at the edge so you can switch later.

Do you already run Kubernetes for serving?
Do you need several engines (vLLM + TensorRT-LLM / SGLang)?
Is your platform (data, training, RL) built on Ray?
Do you just need a simple, quick-start deployment?
llm-d. Kubernetes-native, upstream-aligned (Gateway API Inference Extension), cache- and load-aware scheduling, and well-lit paths for P/D and wide EP.
4.7Quiz

Check yourself.

Platform questions staff interviews love.

1. In llm-d, which component decides which vLLM pod serves a given request?

2. Your HPA targets 70% CPU and never scales, even while TTFT is 8 s. Why?

3. A TP=4 pod stays Pending although the cluster has 6 free H100s. Most likely?

4. Why mount a memory-backed volume at /dev/shm for a TP vLLM pod?

5. During an incident, NCCL bandwidth drops, ITL rises, throughput halves and GPU utilization FALLS. Likely cause?

0 / 5 answered