inference atlas

AI infrastructure · interactive field guide

Watch an LLM
system think.

Type a prompt and watch it become tokens, fill a KV cache, and stream back one token at a time. Then follow that same request through GPUs, routers, vLLM and Kubernetes until you can explain every hop in an interview.

Try it: watch your own prompt

Try it · your prompt, step by step

Type anything. Watch prefill fill the KV cache, then decode read it.

idle
Time to first token—
Inter-token latency—
KV cache size0 B
Token passes (work)0
Tokensprompt → answer
Why#29829
·is#18783
·my#67377
·Kub#97823
ernetes#43929
·pod#21266
·stuck#97129
·in#82688
·Pending#63298
?#47258
The#73736
·pod#21266
·is#18783
·Pending#63298
·because#89769
·no#82664
·node#80211
·has#7953
·a#88398
·free#45657
·GPU#73211
⟨eos⟩#2
Model4 layers
KV cacheK | V per layer
Attentionwhat the query reads
GPU compute (math)
Memory bandwidth (reading weights + KV)
KeyValueQuery / attention readOutput token
Arrive
A request arrives: “Why is my Kubernetes pod stuck in Pending?”. A GPU can’t read text, so the words must become numbers first.

Tokenizer, timings and answers are a teaching simulation. The KV size uses Llama-3.1-8B in BF16 (128 KiB per token).

Your route through the atlas.

Go top to bottom, or jump in. Each page mixes narrated animations, simulators you can break, and drag-and-drop checks. Mark lessons as understood and your progress is saved in this browser.

One request, top to bottom.

Every layer of the stack does one job for your request. Press play to trace it down the stack, or click any layer.

Trace · the full serving stack

Where does my request go?

tracing
Layer 1/11
Workload — Everything starts with the traffic shape: how long are prompts and answers, how much is shared, and how many turns does a session take?

Something is slow? Ask five questions.

An interview framework that shows you reason from workload → bottleneck → architecture → trade-off instead of listing buzzwords. Practise it in the diagnosis game.

01

Compute-bound?

  • GPU FLOPs saturated?
  • Long prompts in prefill?
  • Big batches?

“TTFT is high even when the queue is empty.”

02

Memory-bound?

  • KV cache full?
  • Preemptions?
  • HBM bandwidth in decode?

“Stream slows as more users join; preemptions climb.”

03

Communication-bound?

  • NCCL all-reduce time?
  • All-to-all for experts?
  • Cross-node links?

“Throughput dropped when the model moved from 1 node to 2.”

04

Scheduling-bound?

  • Queue depth?
  • Bad placement?
  • Lost KV locality?

“GPUs look fine but P99 TTFT spikes and hit rate fell.”

05

Workload-bound?

  • Context length?
  • Output length?
  • Agent loops, multimodal, RL?

“The traffic changed shape, not the system.”