Track 01 · Beginner · ~25 min
How one request becomes an answer.
No background needed. By the end you’ll be able to explain what a token is, why an LLM works in two very different phases, what the KV cache stores, and why GPU memory, not maths, usually decides how many users you can serve.
Models read tokens, not words.
Before anything else, text is chopped into tokens, pieces from a fixed vocabulary, each with an integer ID. Every cost in serving is counted in tokens: compute, memory, latency and price.
Type something. Watch it split.
What
A tokenizer (usually a BPE variant) maps text to IDs from a vocabulary of roughly 30K–260K pieces (see the table below). Common words are one token; rare words, code and numbers split into several.
Rule of thumb
English averages about 4 characters ≈ 1 token, or ¾ of a word. Code and non-English text use more tokens per character.
Why infra cares
Prompt tokens drive prefill compute, and every token held in context costs KV cache memory. Context length is a memory budget.
This demo splits long words with a simple heuristic. Real BPE tokenizers learn their merges from data.
Four ways to split text
BPE
Training starts from single characters and keeps merging the most frequent adjacent pair into a new token until the vocabulary is full. Encoding replays those merges in order. With byte fallback, a character outside the vocabulary becomes raw bytes instead of “unknown”.
Used by Llama 2, Gemma (via SentencePiece)
Byte-level BPE
The same merging, over the 256 possible byte values instead of characters, so any input can be encoded and nothing is “unknown”. A rare character or an emoji can take several tokens.
Used by GPT-2 onward, Llama 3 and 4, Qwen, DeepSeek, gpt-oss
WordPiece
Like BPE, but each merge is the one that most raises the likelihood of the training data, not simply the most frequent pair. Encoding takes the longest matching piece first; pieces inside a word start with ##.
Used by BERT-style encoders, still common behind embedding and reranking models
Unigram
Works the other way round: start from a large candidate vocabulary and prune the pieces the training data needs least, until it reaches the target size. Each input gets its most probable split.
Used by T5, ALBERT, XLNet, XLM-R (via SentencePiece)
SentencePiece and tiktoken are libraries, not algorithms. SentencePiece trains BPE or Unigram straight from raw text and marks spaces with ▁; tiktoken is a fast byte-level BPE encoder.
Vocabulary sizes at a glance
| Model | Tokenizer | Vocab size |
|---|---|---|
| BERT-base | WordPiece | 30,522 |
| Llama 2 7B | BPE (SentencePiece, byte fallback) | 32,000 |
| T5-base | Unigram (SentencePiece) | 32,128 |
| GPT-2 | Byte-level BPE | 50,257 |
| Llama 3.1 8B | Byte-level BPE | 128,256 |
| DeepSeek-V3 | Byte-level BPE | 129,280 |
| Qwen3-8B | Byte-level BPE | 151,936 |
| gpt-oss-20b | Byte-level BPE | 201,088 |
| Llama 4 Scout | Byte-level BPE | 202,048 |
| Gemma 3 4B | BPE (SentencePiece, byte fallback) | 262,208 |
Vocab size is vocab_size in each model’s published config.json. It counts embedding rows, which can include padding and reserved slots, so the tokenizer itself may hold slightly fewer pieces.
Bigger isn’t free
A larger vocabulary packs text into fewer tokens: less prefill and KV per request. But the embedding and output layers grow with it. In Qwen3-0.6B the 151,936 × 1,024 embedding table is about a quarter of all weights, and every decode step scores all 151,936 entries.
Counts don’t transfer
A token count belongs to one tokenizer. The same prompt is a different length on another model, so context limits, quotas and cost estimates must use the served model’s tokenizer.
In production
what this looks like on a live dashboard- “Context length exceeded” errors, though the client checked the length
- The client counted with a different tokenizer, or left out the chat template’s tokens. Count with the served model’s tokenizer, after applying its template.
- Prompt tokens per request jump while the request rate is flat
- Traffic shifted toward code, JSON or non-English text, which take more tokens per character. Prefill time and KV memory per request rise with it.
- Many answers stop at
max_tokensafter a model upgrade - The chat template or stop tokens don’t match the new model, so its end-of-turn token is never treated as a stop.
- TTFT climbs while the GPUs sit partly idle
- The API server’s CPU may be saturated tokenizing very long prompts. Check the frontend’s CPU, and add API server processes or replicas.
Read more · scenario: The answers that never ended
- Alert
After a model upgrade, end-to-end latency p90 jumped from about 9 s to 50 s and output tokens per day nearly tripled. TTFT and ITL looked normal.
- Look
35% of requests now finished with
finished_reason="length", up from 1%: they stopped at the 2,048-tokenmax_tokenscap instead of ending on their own. - Cause
The server still used the old model’s stop tokens. The new model ends a turn with a different special token (Llama 3 Instruct uses
<|eot_id|>), so generation ran on to the cap. - Fix
Roll back, then load the chat template and stop token IDs from the new model’s own tokenizer and generation config.
- Guardrail
Alert on the share of
finished_reason="length"and on output tokens per request, per model version, and canary new models on a slice of traffic first.
Illustrative: a composite scenario with rounded numbers, not a real incident.
Prefill fills the KV cache. Decode reads it.
Step through one request. Watch the prompt go through the model in parallel, leave Keys and Values in the cache, then watch decode produce one token per step by reading that cache. Then switch the cache off and compare the work counter.
Watch the cache fill in prefill, then get read in decode.
What
Prefill processes the whole prompt at once and builds the KV cache. Decode generates tokens one at a time, each reading the cache and appending to it.
Why it matters
They stress different hardware. Prefill is compute-bound and sets TTFT. Decode is memory-bandwidth-bound and sets ITL. Production systems tune, and sometimes scale, them separately.
Scenario
“The first response takes 5 seconds” → look at prefill / TTFT. “The first word is quick but the answer takes forever” → look at decode / ITL.
Staff answer
“Prefill processes the input and builds the KV cache, so it’s primarily compute-bound and affects TTFT. Decode generates tokens autoregressively and is more memory-bandwidth and KV-cache sensitive, affecting inter-token latency. That’s why production systems often optimize or scale the two phases independently.”
In production
what this looks like on a live dashboard- TTFT p95 rises while ITL stays flat
- Prefill or queueing: longer prompts, or requests waiting for a slot. Check prompt length and
vllm:num_requests_waiting. - ITL rises while TTFT stays flat
- Decode is slowing: bigger batches or longer contexts mean more bytes to read on every step.
- Streams pause mid-answer, then resume
- Requests were preempted when the KV cache filled, and their KV is being recomputed. Watch
vllm:num_preemptions_total. - “GPU utilization 100%”, yet another replica still cuts latency
- That metric is the share of time any kernel is running, not how much of the GPU’s compute is in use. Decode is bound by memory bandwidth, so judge headroom by latency and tokens/s.
Read more · scenario: The retrieval launch
- Alert
The day retrieval (RAG) shipped, TTFT p95 went from 0.4 s to 2.6 s. ITL didn’t move.
- Look
Prompts had grown from about 600 to 7,000 tokens per request, and
vllm:num_requests_waitingstayed above zero at peak. - Cause
Each request now carried more than 10× the prefill work. Prefill is compute-bound and runs before the first token, so TTFT grew while decode was untouched.
- Fix
Cap retrieved context (fewer, reranked chunks), reuse the shared system prompt with prefix caching, and add prefill capacity, or split prefill from decode if long prompts stay the norm.
- Guardrail
Track TTFT per prompt-length bucket, with prompt tokens per request on the same dashboard.
Illustrative: a composite scenario with rounded numbers, not a real incident.
Why keep the Keys and Values?
Attention is how a token looks back at earlier tokens. Click a word and see which earlier words it reads. Every one of those reads needs the earlier token’s Key and Value, which is exactly what the KV cache keeps.
Click a word to see what it attends to.
Query
“What am I looking for?” The current token makes a query vector.
Key
“What do I contain?” Every earlier token has a key. Query·Key gives a relevance score.
Value
“What do I pass on if chosen?” The output is the scores (after softmax) times each earlier token’s value.
So
Queries are used once, but Keys and Values are reused by every future token. Cache K and V, never Q. That’s the KV cache.
In production
what this looks like on a live dashboard- ITL creeps up as conversations get longer
- Each new token attends to every earlier one, so every decode step reads more K and V. Long chat histories slow every step.
- On long prompts, TTFT grows faster than prompt length
- Prefill attention compares every token with every earlier token, so that part grows with the square of the length. Doubling a long prompt more than doubles prefill.
- One very long request slows its whole batch
- Its K and V reads can outweigh the rest of the batch’s in every shared step. Give long-context traffic its own pool.
- The same prompt at temperature 0 gives different answers under load
- Attention and matmul kernels add numbers in a different order at different batch sizes, so results shift slightly and can flip a token. Expected, unless you run batch-invariant kernels at some cost in speed.
Read more · scenario: Slower every hour
- Alert
A support assistant’s ITL p95 drifted from 25 ms to 60 ms over the day, at a steady request rate.
- Look
The assistant resends the whole conversation every turn. Average context per request grew from 2K to 20K tokens as sessions lengthened, and
vllm:kv_cache_usage_percsat near 100%. - Cause
Every decode step reads the K and V of every earlier token. Ten times the context meant far more memory reads per step, and fewer sequences fit in the cache.
- Fix
Cap history per session (keep recent turns, summarize older ones) and route long sessions to a pool sized for them.
- Guardrail
Plot ITL by context-length bucket, and alert when p95 context per request drifts.
Illustrative: a composite scenario with rounded numbers, not a real incident.
The anatomy of “it feels slow”.
Users feel two numbers: how long until the first token appears (TTFT), and how fast the rest streams (ITL / TPOT). Drag the sliders or pick a user complaint and see which part of the timeline grows.
End-to-end = TTFT + TPOT × (output tokens − 1)
TTFT
Time to first token = queueing + prefill (+ any KV transfer). Grows with prompt length and load.
ITL / TPOT
Inter-token latency is the gap between tokens; TPOT is its per-request mean. Grows with batch size and context length.
Throughput vs goodput
Tokens/s per GPU measures efficiency. Goodput counts only requests that meet every SLO, which is what users actually get.
Staff answer
“I report TTFT and ITL as percentiles at a stated load, never averages at an unstated one. Then I plot throughput per GPU against per-user tokens/s, because the right configuration is the highest-throughput point that still meets the interactivity SLO.”
In production
what this looks like on a live dashboard- Average latency looks fine, users still complain
- Averages hide the tail. Watch TTFT and ITL at p95 and p99, split by prompt length.
- TTFT p99 jumps at peak, ITL steady
- Requests are queueing for a slot (
vllm:num_requests_waitingabove zero). Add replicas or shed load; raising batch limits trades the queue for slower streams. - Tokens/s per GPU went up, and so did complaints
- Bigger batches raise throughput per GPU but slow every user’s stream. Goodput, the requests that meet the SLO, went down.
- Users wait seconds; server-side TTFT says milliseconds
- Time is lost outside the engine: gateway, auth, network, or a proxy buffering the stream.
Read more · scenario: Fast server, slow users
- Alert
After the API moved behind a new ingress, users reported a blank screen for 8–10 s. Server-side TTFT p95 was 400 ms.
- Look
Tokens reached the client in one burst at the end, so client-measured TTFT matched end-to-end latency.
- Cause
The new proxy buffered responses, so the streamed (server-sent events) answer arrived only once it was complete.
- Fix
Turn off response buffering on streaming routes (for example
proxy_buffering offin NGINX) and check idle timeouts on long streams. - Guardrail
Run a synthetic client that streams a request from outside the cluster, and alert on the gap between its TTFT and the server’s.
Illustrative: a composite scenario with rounded numbers, not a real incident.
Static vs continuous batching: a race.
A GPU serves many users per step. Static batching waits for the whole batch to finish; continuous batching refills a slot the moment a sequence ends. Same 14 requests, same GPU. Press play.
Hatched = a paid-for GPU slot doing nothing.
What
Continuous (iteration-level) batching re-decides the batch every forward step: finished sequences leave and waiting ones join immediately.
Why it matters
Output lengths vary wildly. Static batches idle while the longest request finishes, so GPU utilization and throughput suffer.
The new problem
A long prompt admitted into a running batch slows that step for everyone decoding. The fix is chunked prefill under a token budget.
Staff answer
“Continuous batching dynamically admits and removes sequences during generation instead of waiting for a fixed batch to complete. It improves GPU utilization and throughput while keeping latency good under variable request lengths; the trade-off it introduces is prefill interference, which chunked prefill bounds.”
In production
what this looks like on a live dashboard- Every user’s ITL spikes when a long prompt arrives
- A big prefill is sharing steps with everyone’s decode. Chunked prefill under a per-step token budget (
--max-num-batched-tokens) bounds the stall. - Requests queue while the GPU looks underused
- A batch limit is binding:
--max-num-seqsis too low, or the KV cache is full. Compare running and waiting requests with KV usage. - Throughput drops at peak while preemptions climb
- More sequences were admitted than the KV cache can hold as they grow; preempted ones are recomputed, which wastes GPU time.
- The load test looked great; production doesn’t
- The test used fixed prompt and output lengths. Real traffic has a long tail, and the tail drives batching behavior.
Read more · scenario: The 2 a.m. ITL spike
- Alert
Every night at 02:00, chat ITL p99 jumped from about 30 ms to 400 ms for an hour.
- Look
Prompt tokens/s spiked at the same time: a batch summarization job sends 30K-token documents to the same pool, and someone had raised
--max-num-batched-tokensto 32K to speed it up. - Cause
Whole 30K-token prefills now landed in single steps alongside the chat decodes. Such a step takes far longer than a decode step, and every streaming user waits for it.
- Fix
Lower the per-step token budget so long prompts are chunked, and move offline jobs to their own pool or a lower priority.
- Guardrail
Track ITL p99 per tenant or traffic class, with separate SLOs for batch and interactive traffic.
Illustrative: a composite scenario with rounded numbers, not a real incident.
The GPU memory budget.
A GPU’s memory holds the model weights, some runtime overhead and everyone’s KV cache. Change the model, context length and number of users, and watch when requests stop fitting.
Weights are fixed. KV cache grows with every user and token.
KV needed = 64 users × 8,192 tokens × 128 KiB = 68.7 GB vs free = 80 × 0.9 − 16.0 GB weights − 2.14 GB runtime = 53.9 GB
→ Only 50 users fit. The rest queue, or running requests get preempted and recomputed later.
What
KV bytes per token = 2 (K and V) × layers × KV heads × head dim × bytes per value. Multiply by context length and concurrent users.
Why it matters
When KV runs out, requests queue or get preempted. KV capacity, not FLOPs, usually caps concurrency.
Levers
- GQA/MLA models store fewer KV heads
- FP8 KV halves bytes
- More GPUs (TP) free memory
- Paging removes waste
Staff answer
“KV-cache management is fundamentally a memory-management problem. KV capacity directly limits concurrency, so I size it explicitly: bytes per token times context times users, against what’s left after weights. Then I pick levers: FP8 KV, more tensor parallelism for headroom, offloading, or prefix reuse.”
In production
what this looks like on a live dashboard- Preemptions climb and some streams stall mid-answer
- The KV cache is full (
vllm:kv_cache_usage_percnear 100%). Preempted requests are recomputed later. - The server won’t start: the KV cache can’t hold
--max-model-len - Weights plus overhead leave no room for even one maximum-length sequence. Lower
--max-model-len, raise--gpu-memory-utilization, or shard with tensor parallelism. - Out-of-memory errors once something else lands on the GPU
--gpu-memory-utilizationis a fraction of the GPU’s total memory, so a second process (another model, a profiler) leaves less than the budget assumes.- Concurrency falls as long-context traffic grows
- KV is per token: a request using 4× the context holds 4× the memory, so fewer requests fit.
Read more · scenario: The 128K upgrade
- Alert
Hours after 128K context was enabled for an 8B model on one 80 GB GPU, the waiting queue grew at peak and preemptions climbed.
- Look
vllm:kv_cache_usage_percsat at 100%. Running requests fell from about 60 to under 10, with a handful of long-document requests holding most of the cache. - Cause
Llama-3.1-8B stores 128 KiB of KV per token (the lab’s number), so one 128K-token request holds 16 GiB. After weights and overhead, about 50 GiB is left for KV: room for only three full-length requests.
- Fix
Give long-context traffic its own pool and
--max-model-len, use an FP8 KV cache (--kv-cache-dtype fp8, half the bytes), or add tensor parallelism for more room. - Guardrail
Alert on KV usage above 90% and on any sustained preemptions, and require a KV budget (bytes per token × p95 context × target users) in context-length change reviews.
Illustrative: a composite scenario with rounded numbers, not a real incident.
Check yourself.
Five quick questions. Every answer explains itself, and wrong answers are how you learn fastest. For hands-on practice, try the drag-and-drop playground.
1. Users say: “the first word takes 5 seconds, then it streams quickly.” Where do you look first?
2. What does the KV cache store?
3. Why is decode usually memory-bandwidth-bound?
4. Continuous batching raises throughput mainly because…
5. Llama-2-7B needs about 4× more KV memory per token than Llama-3.1-8B. Why?
0 / 5 answered