inference atlas

Track 01 · Beginner · ~25 min

How one request becomes an answer.

No background needed. By the end you’ll be able to explain what a token is, why an LLM works in two very different phases, what the KV cache stores, and why GPU memory, not maths, usually decides how many users you can serve.

1.1Beginner

Models read tokens, not words.

Before anything else, text is chopped into tokens, pieces from a fixed vocabulary, each with an integer ID. Every cost in serving is counted in tokens: compute, memory, latency and price.

Lab · tokenizer

Type something. Watch it split.

Kub92545ernetes43929·schedul55251es17165·infer92053ence742·pods39635·onto17059·GPU73211·nodes11944.40973
Characters51
Tokens11
Chars / token4.6
KV cache to hold it1.38 MiB (Llama-3.1-8B)

What

A tokenizer (usually a BPE variant) maps text to IDs from a vocabulary of roughly 30K–260K pieces (see the table below). Common words are one token; rare words, code and numbers split into several.

Rule of thumb

English averages about 4 characters ≈ 1 token, or ¾ of a word. Code and non-English text use more tokens per character.

Why infra cares

Prompt tokens drive prefill compute, and every token held in context costs KV cache memory. Context length is a memory budget.

This demo splits long words with a simple heuristic. Real BPE tokenizers learn their merges from data.

Four ways to split text

BPE

Training starts from single characters and keeps merging the most frequent adjacent pair into a new token until the vocabulary is full. Encoding replays those merges in order. With byte fallback, a character outside the vocabulary becomes raw bytes instead of “unknown”.

Used by Llama 2, Gemma (via SentencePiece)

Byte-level BPE

The same merging, over the 256 possible byte values instead of characters, so any input can be encoded and nothing is “unknown”. A rare character or an emoji can take several tokens.

Used by GPT-2 onward, Llama 3 and 4, Qwen, DeepSeek, gpt-oss

WordPiece

Like BPE, but each merge is the one that most raises the likelihood of the training data, not simply the most frequent pair. Encoding takes the longest matching piece first; pieces inside a word start with ##.

Used by BERT-style encoders, still common behind embedding and reranking models

Unigram

Works the other way round: start from a large candidate vocabulary and prune the pieces the training data needs least, until it reaches the target size. Each input gets its most probable split.

Used by T5, ALBERT, XLNet, XLM-R (via SentencePiece)

SentencePiece and tiktoken are libraries, not algorithms. SentencePiece trains BPE or Unigram straight from raw text and marks spaces with ▁; tiktoken is a fast byte-level BPE encoder.

Vocabulary sizes at a glance

ModelTokenizerVocab size
BERT-baseWordPiece30,522
Llama 2 7BBPE (SentencePiece, byte fallback)32,000
T5-baseUnigram (SentencePiece)32,128
GPT-2Byte-level BPE50,257
Llama 3.1 8BByte-level BPE128,256
DeepSeek-V3Byte-level BPE129,280
Qwen3-8BByte-level BPE151,936
gpt-oss-20bByte-level BPE201,088
Llama 4 ScoutByte-level BPE202,048
Gemma 3 4BBPE (SentencePiece, byte fallback)262,208

Vocab size is vocab_size in each model’s published config.json. It counts embedding rows, which can include padding and reserved slots, so the tokenizer itself may hold slightly fewer pieces.

Bigger isn’t free

A larger vocabulary packs text into fewer tokens: less prefill and KV per request. But the embedding and output layers grow with it. In Qwen3-0.6B the 151,936 × 1,024 embedding table is about a quarter of all weights, and every decode step scores all 151,936 entries.

Counts don’t transfer

A token count belongs to one tokenizer. The same prompt is a different length on another model, so context limits, quotas and cost estimates must use the served model’s tokenizer.

In production

what this looks like on a live dashboard
“Context length exceeded” errors, though the client checked the length
The client counted with a different tokenizer, or left out the chat template’s tokens. Count with the served model’s tokenizer, after applying its template.
Prompt tokens per request jump while the request rate is flat
Traffic shifted toward code, JSON or non-English text, which take more tokens per character. Prefill time and KV memory per request rise with it.
Many answers stop at max_tokens after a model upgrade
The chat template or stop tokens don’t match the new model, so its end-of-turn token is never treated as a stop.
TTFT climbs while the GPUs sit partly idle
The API server’s CPU may be saturated tokenizing very long prompts. Check the frontend’s CPU, and add API server processes or replicas.
Read more · scenario: The answers that never ended
  1. Alert

    After a model upgrade, end-to-end latency p90 jumped from about 9 s to 50 s and output tokens per day nearly tripled. TTFT and ITL looked normal.

  2. Look

    35% of requests now finished with finished_reason="length", up from 1%: they stopped at the 2,048-token max_tokens cap instead of ending on their own.

  3. Cause

    The server still used the old model’s stop tokens. The new model ends a turn with a different special token (Llama 3 Instruct uses <|eot_id|>), so generation ran on to the cap.

  4. Fix

    Roll back, then load the chat template and stop token IDs from the new model’s own tokenizer and generation config.

  5. Guardrail

    Alert on the share of finished_reason="length" and on output tokens per request, per model version, and canary new models on a slice of traffic first.

Illustrative: a composite scenario with rounded numbers, not a real incident.

1.2Beginnercore idea

Prefill fills the KV cache. Decode reads it.

Step through one request. Watch the prompt go through the model in parallel, leave Keys and Values in the cache, then watch decode produce one token per step by reading that cache. Then switch the cache off and compare the work counter.

Lab · one request, step by step

Watch the cache fill in prefill, then get read in decode.

idle
Time to first token—
Inter-token latency—
KV cache size0 B
Token passes (work)0
Tokensprompt → answer
Explain#66560
·Kub#97823
ernetes#43929
·network#94427
ing#18807
·simply#1805
Every#86672
·pod#21266
·gets#3882
·its#28847
·own#73053
·IP#34474
·address#60705
.#40973
Model4 layers
KV cacheK | V per layer
Attentionwhat the query reads
GPU compute (math)
Memory bandwidth (reading weights + KV)
KeyValueQuery / attention readOutput token
Arrive
A request arrives: “Explain Kubernetes networking simply”. A GPU can’t read text, so the words must become numbers first.

What

Prefill processes the whole prompt at once and builds the KV cache. Decode generates tokens one at a time, each reading the cache and appending to it.

Why it matters

They stress different hardware. Prefill is compute-bound and sets TTFT. Decode is memory-bandwidth-bound and sets ITL. Production systems tune, and sometimes scale, them separately.

Scenario

“The first response takes 5 seconds” → look at prefill / TTFT. “The first word is quick but the answer takes forever” → look at decode / ITL.

Staff answer

“Prefill processes the input and builds the KV cache, so it’s primarily compute-bound and affects TTFT. Decode generates tokens autoregressively and is more memory-bandwidth and KV-cache sensitive, affecting inter-token latency. That’s why production systems often optimize or scale the two phases independently.”

In production

what this looks like on a live dashboard
TTFT p95 rises while ITL stays flat
Prefill or queueing: longer prompts, or requests waiting for a slot. Check prompt length and vllm:num_requests_waiting.
ITL rises while TTFT stays flat
Decode is slowing: bigger batches or longer contexts mean more bytes to read on every step.
Streams pause mid-answer, then resume
Requests were preempted when the KV cache filled, and their KV is being recomputed. Watch vllm:num_preemptions_total.
“GPU utilization 100%”, yet another replica still cuts latency
That metric is the share of time any kernel is running, not how much of the GPU’s compute is in use. Decode is bound by memory bandwidth, so judge headroom by latency and tokens/s.
Read more · scenario: The retrieval launch
  1. Alert

    The day retrieval (RAG) shipped, TTFT p95 went from 0.4 s to 2.6 s. ITL didn’t move.

  2. Look

    Prompts had grown from about 600 to 7,000 tokens per request, and vllm:num_requests_waiting stayed above zero at peak.

  3. Cause

    Each request now carried more than 10× the prefill work. Prefill is compute-bound and runs before the first token, so TTFT grew while decode was untouched.

  4. Fix

    Cap retrieved context (fewer, reranked chunks), reuse the shared system prompt with prefix caching, and add prefill capacity, or split prefill from decode if long prompts stay the norm.

  5. Guardrail

    Track TTFT per prompt-length bucket, with prompt tokens per request on the same dashboard.

Illustrative: a composite scenario with rounded numbers, not a real incident.

1.3Beginner

Why keep the Keys and Values?

Attention is how a token looks back at earlier tokens. Click a word and see which earlier words it reads. Every one of those reads needs the earlier token’s Key and Value, which is exactly what the KV cache keeps.

Lab · attention explorer

Click a word to see what it attends to.

causal: only looks back
1 · QueryThe token “it” makes a query vector: what am I looking for?
2 · ScoresQuery · Key for each of the 9 visible tokens (itself included). Tokens after it are hidden: the model can’t see the future.
3 · SoftmaxScores become weights that sum to 100%. Top: service 52% · pod 18% · it 8%
4 · Mix values“it” is ambiguous: the pod or the service? Attention puts most weight on service — the model resolves the reference by reading earlier Keys.

Query

“What am I looking for?” The current token makes a query vector.

Key

“What do I contain?” Every earlier token has a key. Query·Key gives a relevance score.

Value

“What do I pass on if chosen?” The output is the scores (after softmax) times each earlier token’s value.

So

Queries are used once, but Keys and Values are reused by every future token. Cache K and V, never Q. That’s the KV cache.

In production

what this looks like on a live dashboard
ITL creeps up as conversations get longer
Each new token attends to every earlier one, so every decode step reads more K and V. Long chat histories slow every step.
On long prompts, TTFT grows faster than prompt length
Prefill attention compares every token with every earlier token, so that part grows with the square of the length. Doubling a long prompt more than doubles prefill.
One very long request slows its whole batch
Its K and V reads can outweigh the rest of the batch’s in every shared step. Give long-context traffic its own pool.
The same prompt at temperature 0 gives different answers under load
Attention and matmul kernels add numbers in a different order at different batch sizes, so results shift slightly and can flip a token. Expected, unless you run batch-invariant kernels at some cost in speed.
Read more · scenario: Slower every hour
  1. Alert

    A support assistant’s ITL p95 drifted from 25 ms to 60 ms over the day, at a steady request rate.

  2. Look

    The assistant resends the whole conversation every turn. Average context per request grew from 2K to 20K tokens as sessions lengthened, and vllm:kv_cache_usage_perc sat near 100%.

  3. Cause

    Every decode step reads the K and V of every earlier token. Ten times the context meant far more memory reads per step, and fewer sequences fit in the cache.

  4. Fix

    Cap history per session (keep recent turns, summarize older ones) and route long sessions to a pool sized for them.

  5. Guardrail

    Plot ITL by context-length bucket, and alert when p95 context per request drifts.

Illustrative: a composite scenario with rounded numbers, not a real incident.

1.4Beginner

The anatomy of “it feels slow”.

Users feel two numbers: how long until the first token appears (TTFT), and how fast the rest streams (ITL / TPOT). Drag the sliders or pick a user complaint and see which part of the timeline grows.

Lab · latency timeline

End-to-end = TTFT + TPOT × (output tokens − 1)

decode · 300 tokens
TTFT 462 ms
done 6.76 s
TTFT462 ms
TPOT / ITL21 ms
Tokens/s per user47
End-to-end6.76 s
Healthy: TTFT under 2 s and more than about 22 tokens/s per user (faster than people read). Try a complaint above.

TTFT

Time to first token = queueing + prefill (+ any KV transfer). Grows with prompt length and load.

ITL / TPOT

Inter-token latency is the gap between tokens; TPOT is its per-request mean. Grows with batch size and context length.

Throughput vs goodput

Tokens/s per GPU measures efficiency. Goodput counts only requests that meet every SLO, which is what users actually get.

Staff answer

“I report TTFT and ITL as percentiles at a stated load, never averages at an unstated one. Then I plot throughput per GPU against per-user tokens/s, because the right configuration is the highest-throughput point that still meets the interactivity SLO.”

In production

what this looks like on a live dashboard
Average latency looks fine, users still complain
Averages hide the tail. Watch TTFT and ITL at p95 and p99, split by prompt length.
TTFT p99 jumps at peak, ITL steady
Requests are queueing for a slot (vllm:num_requests_waiting above zero). Add replicas or shed load; raising batch limits trades the queue for slower streams.
Tokens/s per GPU went up, and so did complaints
Bigger batches raise throughput per GPU but slow every user’s stream. Goodput, the requests that meet the SLO, went down.
Users wait seconds; server-side TTFT says milliseconds
Time is lost outside the engine: gateway, auth, network, or a proxy buffering the stream.
Read more · scenario: Fast server, slow users
  1. Alert

    After the API moved behind a new ingress, users reported a blank screen for 8–10 s. Server-side TTFT p95 was 400 ms.

  2. Look

    Tokens reached the client in one burst at the end, so client-measured TTFT matched end-to-end latency.

  3. Cause

    The new proxy buffered responses, so the streamed (server-sent events) answer arrived only once it was complete.

  4. Fix

    Turn off response buffering on streaming routes (for example proxy_buffering off in NGINX) and check idle timeouts on long streams.

  5. Guardrail

    Run a synthetic client that streams a request from outside the cluster, and alert on the gap between its TTFT and the server’s.

Illustrative: a composite scenario with rounded numbers, not a real incident.

1.5Beginner

Static vs continuous batching: a race.

A GPU serves many users per step. Static batching waits for the whole batch to finish; continuous batching refills a slot the moment a sequence ends. Same 14 requests, same GPU. Press play.

Lab · batching race · 4 slots per GPU step

Hatched = a paid-for GPU slot doing nothing.

Static batchingNext batch starts only when the longest request in this batch ends.
Finished0 / 14
Slot utilization—
Avg wait to start0.0 steps
R1
R5
R9
R13
R2
R6
R10
R14
R3
R7
R11
R4
R8
R12
all done · step 34
Continuous batchingA slot is refilled the step after its request ends.
Finished0 / 14
Slot utilization—
Avg wait to start0.0 steps
R1
R5
R8
R10
R2
R6
R7
R11
R13
R3
R12
R14
R4
R9
all done · step 23
generatingfinishedidle slot (waiting for the batch)

What

Continuous (iteration-level) batching re-decides the batch every forward step: finished sequences leave and waiting ones join immediately.

Why it matters

Output lengths vary wildly. Static batches idle while the longest request finishes, so GPU utilization and throughput suffer.

The new problem

A long prompt admitted into a running batch slows that step for everyone decoding. The fix is chunked prefill under a token budget.

Staff answer

“Continuous batching dynamically admits and removes sequences during generation instead of waiting for a fixed batch to complete. It improves GPU utilization and throughput while keeping latency good under variable request lengths; the trade-off it introduces is prefill interference, which chunked prefill bounds.”

In production

what this looks like on a live dashboard
Every user’s ITL spikes when a long prompt arrives
A big prefill is sharing steps with everyone’s decode. Chunked prefill under a per-step token budget (--max-num-batched-tokens) bounds the stall.
Requests queue while the GPU looks underused
A batch limit is binding: --max-num-seqs is too low, or the KV cache is full. Compare running and waiting requests with KV usage.
Throughput drops at peak while preemptions climb
More sequences were admitted than the KV cache can hold as they grow; preempted ones are recomputed, which wastes GPU time.
The load test looked great; production doesn’t
The test used fixed prompt and output lengths. Real traffic has a long tail, and the tail drives batching behavior.
Read more · scenario: The 2 a.m. ITL spike
  1. Alert

    Every night at 02:00, chat ITL p99 jumped from about 30 ms to 400 ms for an hour.

  2. Look

    Prompt tokens/s spiked at the same time: a batch summarization job sends 30K-token documents to the same pool, and someone had raised --max-num-batched-tokens to 32K to speed it up.

  3. Cause

    Whole 30K-token prefills now landed in single steps alongside the chat decodes. Such a step takes far longer than a decode step, and every streaming user waits for it.

  4. Fix

    Lower the per-step token budget so long prompts are chunked, and move offline jobs to their own pool or a lower priority.

  5. Guardrail

    Track ITL p99 per tenant or traffic class, with separate SLOs for batch and interactive traffic.

Illustrative: a composite scenario with rounded numbers, not a real incident.

1.6Beginnerbridge

The GPU memory budget.

A GPU’s memory holds the model weights, some runtime overhead and everyone’s KV cache. Change the model, context length and number of users, and watch when requests stop fitting.

Lab · one GPU, many users

Weights are fixed. KV cache grows with every user and token.

KV overflow
Weight precision
KV cache precision
weights 16.0 GB
KV needed 68.7 GB
doesn’t fit
90% · gpu-memory-utilization
H100 80GB = 80 GB
Weights16.0 GB
KV per token128 KiB
KV per user1.07 GB
Free for KV53.9 GB
Max users at this context50
KV/token = 2 × 32 layers × 8 KV heads × 128 dims × 2 B = 128 KiB
KV needed = 64 users × 8,192 tokens × 128 KiB = 68.7 GB vs free = 80 × 0.9 − 16.0 GB weights − 2.14 GB runtime = 53.9 GB
→ Only 50 users fit. The rest queue, or running requests get preempted and recomputed later.

What

KV bytes per token = 2 (K and V) × layers × KV heads × head dim × bytes per value. Multiply by context length and concurrent users.

Why it matters

When KV runs out, requests queue or get preempted. KV capacity, not FLOPs, usually caps concurrency.

Levers

  • GQA/MLA models store fewer KV heads
  • FP8 KV halves bytes
  • More GPUs (TP) free memory
  • Paging removes waste
Staff answer

“KV-cache management is fundamentally a memory-management problem. KV capacity directly limits concurrency, so I size it explicitly: bytes per token times context times users, against what’s left after weights. Then I pick levers: FP8 KV, more tensor parallelism for headroom, offloading, or prefix reuse.”

In production

what this looks like on a live dashboard
Preemptions climb and some streams stall mid-answer
The KV cache is full (vllm:kv_cache_usage_perc near 100%). Preempted requests are recomputed later.
The server won’t start: the KV cache can’t hold --max-model-len
Weights plus overhead leave no room for even one maximum-length sequence. Lower --max-model-len, raise --gpu-memory-utilization, or shard with tensor parallelism.
Out-of-memory errors once something else lands on the GPU
--gpu-memory-utilization is a fraction of the GPU’s total memory, so a second process (another model, a profiler) leaves less than the budget assumes.
Concurrency falls as long-context traffic grows
KV is per token: a request using 4× the context holds 4× the memory, so fewer requests fit.
Read more · scenario: The 128K upgrade
  1. Alert

    Hours after 128K context was enabled for an 8B model on one 80 GB GPU, the waiting queue grew at peak and preemptions climbed.

  2. Look

    vllm:kv_cache_usage_perc sat at 100%. Running requests fell from about 60 to under 10, with a handful of long-document requests holding most of the cache.

  3. Cause

    Llama-3.1-8B stores 128 KiB of KV per token (the lab’s number), so one 128K-token request holds 16 GiB. After weights and overhead, about 50 GiB is left for KV: room for only three full-length requests.

  4. Fix

    Give long-context traffic its own pool and --max-model-len, use an FP8 KV cache (--kv-cache-dtype fp8, half the bytes), or add tensor parallelism for more room.

  5. Guardrail

    Alert on KV usage above 90% and on any sustained preemptions, and require a KV budget (bytes per token × p95 context × target users) in context-length change reviews.

Illustrative: a composite scenario with rounded numbers, not a real incident.

1.7Quiz

Check yourself.

Five quick questions. Every answer explains itself, and wrong answers are how you learn fastest. For hands-on practice, try the drag-and-drop playground.

1. Users say: “the first word takes 5 seconds, then it streams quickly.” Where do you look first?

2. What does the KV cache store?

3. Why is decode usually memory-bandwidth-bound?

4. Continuous batching raises throughput mainly because…

5. Llama-2-7B needs about 4× more KV memory per token than Llama-3.1-8B. Why?

0 / 5 answered