Compute-bound?
- GPU FLOPs saturated?
- Long prompts in prefill?
- Big batches?
“TTFT is high even when the queue is empty.”
AI infrastructure · interactive field guide
Type a prompt and watch it become tokens, fill a KV cache, and stream back one token at a time. Then follow that same request through GPUs, routers, vLLM and Kubernetes until you can explain every hop in an interview.
Tokenizer, timings and answers are a teaching simulation. The KV size uses Llama-3.1-8B in BF16 (128 KiB per token).
Go top to bottom, or jump in. Each page mixes narrated animations, simulators you can break, and drag-and-drop checks. Mark lessons as understood and your progress is saved in this browser.
No background needed. What actually happens when you press Enter.
Scaling, optimizing and running it on a cluster.
Split models across GPUs with TP, PP, DP, EP and CP, route requests intelligently, and see NCCL move data.
Play with: will-it-fit · router simulator · collectives
PagedAttention, prefix caching, chunked prefill, P/D disaggregation, speculative decoding, quantization and KV tiers.
Play with: memory pager · hash chain · token budget
vLLM and llm-d on Kubernetes: gateway, endpoint picker, pods, GPU scheduling, autoscaling and observability.
Play with: request walk · YAML explorer · incident game
Frontier workloads, the vLLM engine itself, and keeping GPU clusters healthy.
Mixture-of-experts, multimodal pipelines, RL rollout loops and agentic sessions with 96% cache hits.
Play with: expert router · agent session sim
Engine internals from the vLLM Office Hours: scheduler, CUDA graphs, KV connectors, wide EP, capacity maths.
Play with: scheduler stepper · GPU calculator · serve builder
Keep GPUs scheduled, patched and serving: the four states of GPU capacity, why pods stay Pending, quota gates, and node rotation without an outage.
Play with: Pending debugger · quota gates · node-rotation simulator
Test yourself and get interview-ready.
Staff-level answers for 20 core topics, 63 vLLM questions with self-rating, and numbers to memorize.
Play with: practice mode · flashcards
Drag-and-drop checks: build the request path, sort concepts, diagnose bottlenecks, match fixes, explore the map.
Play with: 5 drag-and-drop games
Every term in plain English, with a link to the lab where you can see it working.
Play with: search
Every layer of the stack does one job for your request. Press play to trace it down the stack, or click any layer.
An interview framework that shows you reason from workload → bottleneck → architecture → trade-off instead of listing buzzwords. Practise it in the diagnosis game.
“TTFT is high even when the queue is empty.”
“Stream slows as more users join; preemptions climb.”
“Throughput dropped when the model moved from 1 node to 2.”
“GPUs look fine but P99 TTFT spikes and hit rate fell.”
“The traffic changed shape, not the system.”