Practice · drag and drop · ~20 min
Learn it with your hands.
Five games that turn the atlas into muscle memory. Drag a card onto a target, or tap a card and then tap a target (works on phones and with the keyboard). Every drop explains itself, right or wrong.
Build the request path.
Put the eight hops a chat request takes in order, from the user’s app to the streamed answer. Get them all right and the request runs through your pipeline.
Which layer does it live in?
Sort 20 concepts into the five layers of the atlas. Each drop is checked instantly, with a one-line reason and a link to the lab.
Diagnose the bottleneck.
The staff interview framework: is it compute, memory, communication, scheduling or workload? Drag each symptom to its most likely category.
Match the problem to the fix.
Eight production problems, eight techniques. Drop each fix onto the problem it solves best.
The same 6K-token system prompt is recomputed for every request.
A 70B model in BF16 doesn’t fit on one 80 GB GPU.
Long document prompts make everyone’s token stream stutter.
Decode is slow at low load while GPU compute sits idle.
A viral system prompt overloads the one replica that caches it.
Long contexts fill the KV cache and cause preemptions.
The HPA never scales, although TTFT is 8 s.
One MoE rank is always the straggler.
The complete map.
Every concept in the atlas and how they connect. Click a node to light up its neighbours and read what it is. Drag nodes to untangle the map your own way. Follow the link to its lab.
Pick any node
32 concepts, 42 connections. Click a node to see how it connects; drag nodes to rearrange. Nodes you’ve opened count toward your scoreboard.
Explored: 0 / 32
Practice lab.
400 hands-on exercises across every track: tune, predict, calculate, fix, order, diagnose, respond to incidents, build commands, classify and recall. Raise your confidence to 90%.
Coverage map
Every lesson in the atlas, coloured by your best scores on the exercises that practise it. Pick one to list them.
- Not started
- Learning
- Solid
- Mastered
- No exercises yet
01Foundations50 exercises
02Scale out50 exercises
03Optimize50 exercises
04Kubernetes platform50 exercises
05Frontier serving50 exercises
06Inside vLLM50 exercises
07Kubernetes for GPU workloads50 exercises
08Interview kit50 exercises
Exercises
1–24 of 400 exercises
- A 36,000-word contract in tokens and KV01 · Tokens01 · GPU memory budgetNew
- Which token cost does it cut?01 · Tokens01 · Prefill → KV → Decode+1New
- Same characters, twice the tokens01 · Tokens01 · TTFT · ITL · E2E+1New
- Fix the context budget01 · Tokens01 · GPU memory budgetNew
- One request, from text to its last token01 · Prefill → KV → Decode01 · Tokens+1New
- Switch the KV cache off01 · Prefill → KV → Decode01 · Attention (Q·K·V)New
- The fastest a decode step can be01 · Prefill → KV → Decode01 · GPU memory budget+1New
- Buy GPUs with twice the FLOPs?01 · Prefill → KV → Decode01 · TTFT · ITL · E2E+1New
- Cache it, or use it once?01 · Attention (Q·K·V)01 · Prefill → KV → DecodeNew
- One attention layer, one decode step01 · Attention (Q·K·V)01 · Prefill → KV → DecodeNew
- 32 KV heads to 8: what changes?01 · Attention (Q·K·V)01 · GPU memory budgetNew
- Fix the attention pseudo-code01 · Attention (Q·K·V)01 · Prefill → KV → DecodeNew
- TTFT, ITL or neither?01 · TTFT · ITL · E2E01 · Prefill → KV → Decode+1New
- TTFT 20 s, GPUs at 70%01 · TTFT · ITL · E2E01 · Continuous batchingNew
- Measure latency at a stated load01 · TTFT · ITL · E2E01 · Continuous batchingNew
- “It feels slow” after a context-limit release01 · TTFT · ITL · E2E01 · GPU memory budget+1New
- Eight requests, four slots01 · Continuous batching01 · TTFT · ITL · E2ENew
- How big is the batch, really?01 · Continuous batching01 · TTFT · ITL · E2ENew
- Pick the batch: throughput vs per-user speed01 · Continuous batching01 · TTFT · ITL · E2E+1New
- Llama-3.1-70B on four H100s: how many 16K conversations?01 · GPU memory budget02 · Will it fit?New
- Fit Llama-3.1-8B on one L401 · GPU memory budget06 · vllm serve builderNew
- Fix a long-context serve script01 · GPU memory budget06 · vllm serve builderNew
- How vLLM turns GPU memory into KV capacity01 · GPU memory budget06 · Capacity mathsNew
- Tokens, prefill, decode and attention: one-sentence answers01 · Prefill → KV → Decode01 · Tokens+1New