inference atlas

Reference · plain English

Every term, decoded.

Short definitions for everything in the atlas, each with a “like…” analogy where it helps and a link to the lab where you can watch it working.

97 of 97 terms
Frontier

Agentic serving

Serving multi-turn agent sessions: long, 96%-cached contexts where KV residency and sticky routing dominate cost.

See it in the lab →
Network

All-reduce

Every GPU ends with the sum of everyone’s buffers. Used by TP after attention and MLP.

See it in the lab →
Network

All-to-all

Every GPU sends a different chunk to every other GPU at once. Used for MoE dispatch and combine.

See it in the lab →
Engine

Async scheduling

Planning step N+1 while step N runs on the GPU, so the GPU never waits on the CPU.

See it in the lab →
Basics

Attention (Q, K, V)

Each token’s query is compared with earlier tokens’ keys; the resulting weights mix their values. K and V are reused by every future token; Q isn’t.

See it in the lab →
Basics

Autoregressive

Each new token depends on all previous tokens, so decode steps can’t run in parallel for one sequence.

See it in the lab →
Frontier

Batch invariance

A prompt gets identical logprobs whatever else is in its batch. Needed for reproducible RL and evals.

See it in the lab →
Memory

cache_salt

A per-tenant value mixed into the first block hash so tenants never share KV blocks.

See it in the lab →
Serving

Cache-aware routing

Sending each request to the replica that already holds its prefix, balanced against load.

See it in the lab →
Operations

Capacity reservation

Capacity bought ahead for one instance type in one zone, sometimes for a fixed window (EC2 Capacity Blocks). Outside that box, reserved-only pods wait.

See it in the lab →
Operations

Capacity type

On-demand, spot or reserved. A pod or pool that allows only one type can only use offerings of that type.

See it in the lab →
Scheduling

Chunked prefill

Splitting a long prompt’s prefill across several steps so it shares the GPU with ongoing decodes.

See it in the lab →
Platform

ClusterQueue / cohort

A Kueue quota pool, often one per team. ClusterQueues in one cohort can borrow each other’s idle quota, and lenders can reclaim it.

See it in the lab →
Platform

Cold start

The minutes a new replica spends loading weights and capturing CUDA graphs before it can serve.

See it in the lab →
Operations

Consolidation

Karpenter deleting or replacing nodes to cut cost. consolidateAfter sets how long a node must be stable first (default 0s; Never turns it off).

See it in the lab →
Parallelism

Context parallelism (DCP / PCP)

Shard one long sequence’s KV (decode CP) or prompt (prefill CP) across GPUs; partial attention merged exactly.

See it in the lab →
Scheduling

Continuous batching

The batch is re-formed every step: finished sequences leave, waiting ones join immediately.

See it in the lab →
Operations

Cordon / drain

Cordon marks a node unschedulable; drain then evicts its pods through the Eviction API, which respects PDBs.

See it in the lab →
Engine

CUDA graphs

Record a sequence of kernel launches once and replay it, removing CPU launch overhead. vLLM uses FULL_AND_PIECEWISE.

See it in the lab →
Parallelism

Data parallelism (DP)

Independent model replicas serving different requests. For MoE, “DP attention” gives each rank its own requests’ KV.

See it in the lab →
Parallelism

DBO

Dual-batch overlap: splits a step into two micro-batches so one’s all-to-all overlaps the other’s compute.

See it in the lab →
Basics

Decode

Generating output one token per step; each step reads all weights plus the KV cache. Memory-bandwidth-bound; sets ITL.

See it in the lab →
Platform

Device plugin

A node agent that registers devices with the kubelet as allocatable resources such as nvidia.com/gpu. Until it runs, a Ready node advertises 0 GPUs.

See it in the lab →
Operations

do-not-disrupt

The karpenter.sh/do-not-disrupt pod annotation: acts like a single-pod blocking PDB, permanently or for a set duration.

See it in the lab →
Operations

Drift

A node no longer matches its NodePool or node class, for example after a new image. Karpenter replaces drifted nodes gracefully, within disruption budgets.

See it in the lab →
Optimization

EAGLE-3 / MTP

Drafters: EAGLE-3 is a small head reading the target’s hidden states; MTP uses the model’s own multi-token prediction heads.

See it in the lab →
Frontier

Encoder cache / E/PD

Keep vision embeddings so images aren’t re-encoded; at scale run encoders on their own workers (encode/prefill/decode).

See it in the lab →
Platform

Endpoint Picker (EPP)

Called by the gateway via ext-proc; runs filters (health, role, adapter) then scorers (prefix, queue, KV).

See it in the lab →
Parallelism

EPLB

Expert-parallel load balancer: replicates hot experts and reshuffles placement live so lockstep ranks stay balanced.

See it in the lab →
Parallelism

Expert parallelism (EP)

MoE experts spread across GPUs; tokens travel all-to-all to their experts and back.

See it in the lab →
Operations

expireAfter

Maximum node lifetime (default 720h). Expiry is forceful: draining starts without a pre-spun replacement.

See it in the lab →
Operations

Four states of GPU capacity

Enabled (a pool exists), allocatable (a node advertises GPUs), requested (pods ask for them), consumed (GPUs bound to running pods). Four numbers, four owners.

See it in the lab →
Memory

FP8 KV cache

Storing KV in 8 bits: half the bytes per token, roughly double the capacity. Validate long-context accuracy.

See it in the lab →
Platform

Gang scheduling

Place all pods of a multi-node group together or none (LeaderWorkerSet, Kueue, Volcano), so half-placed groups don’t idle GPUs.

See it in the lab →
Platform

Gateway API Inference Extension

Kubernetes APIs for inference routing: InferencePool groups model-server pods; an Endpoint Picker chooses one per request.

See it in the lab →
Metrics

Goodput

Requests per second that meet every SLO at once. Throughput that users actually benefit from.

See it in the lab →
Models

GQA

Grouped-query attention: several query heads share one K/V head, cutting KV memory (Llama-3.1: 8 KV heads vs 32 in Llama-2-7B).

See it in the lab →
Operations

Graceful vs forceful disruption

Graceful (Karpenter drift and consolidation) pre-spins a replacement node and waits for it. Forceful (expiration, interruption, repair) starts draining immediately.

See it in the lab →
Scheduling

Head-of-line blocking

One long request consumes every step’s budget so short requests behind it can’t start. Fixed with a per-request prefill cap.

See it in the lab →
Platform

HPA / KEDA

Kubernetes autoscalers. For LLMs, drive them with queue depth or KV usage, not CPU.

See it in the lab →
Network

InfiniBand / RoCE + RDMA

Node-to-node networks; RDMA lets a NIC read and write GPU memory without CPU copies.

See it in the lab →
Operations

Insufficient capacity (ICE)

The cloud can’t supply that instance type in that zone right now. Pods that allow other zones or types still land.

See it in the lab →
Metrics

Interactivity

Tokens per second per user. Benchmarks report throughput per GPU at a fixed interactivity floor (e.g. 50 tok/s/user).

See it in the lab →
Metrics

ITL / TPOT

Inter-token latency (per-token gaps) and time per output token (per-request mean). What users feel as streaming speed.

See it in the lab →
Platform

KServe

Kubernetes model-serving CRDs (InferenceService) with runtimes such as vLLM, for many model types behind one API.

See it in the lab →
Platform

Kueue

Job queueing and quota for Kubernetes: decides when a job is admitted, waits or is preempted. It doesn’t create nodes.

See it in the lab →
Basics

KV cache

The Key and Value vectors of every earlier token at every layer, kept in GPU memory so decode never recomputes them.

See it in the lab →
Serving

KV connector

vLLM’s interface for loading and saving KV outside the GPU: P/D transfer, CPU offload, remote caches.

See it in the lab →
Memory

KV offloading

Moving KV blocks to CPU memory or disk (DMA copies) so they can be reloaded instead of recomputed.

See it in the lab →
Metrics

Little’s law

In-flight requests = arrival rate × time in system. 50 req/s × 20 s decode ≈ 1,000 concurrent sequences.

See it in the lab →
Platform

llm-d

Kubernetes-native distributed inference on vLLM: inference scheduler (EPP), P/D and wide-EP “well-lit paths”, KV-aware routing.

See it in the lab →
Memory

LMCache / Mooncake

KV cache layers outside the engine: shared pools across pods and nodes for reuse and P/D transfer.

See it in the lab →
Platform

MIG

Multi-Instance GPU: partitions one GPU into isolated slices, useful for small models.

See it in the lab →
Frontier

Mixture-of-experts (MoE)

Many expert FFNs; each token activates a few. Memory ∝ total params, compute ∝ active (DeepSeek-R1: 37B of 671B).

See it in the lab →
Models

MLA

Multi-head latent attention (DeepSeek): stores one compressed latent per token instead of per-head K/V. Tiny KV, but TP can’t split it.

See it in the lab →
Engine

Model Runner V2

vLLM’s 2026 rewrite of the model runner: GPU-native input prep, async-first, Triton sampler (+56% on a small model).

See it in the lab →
Serving

Multi-LoRA

Many fine-tuned adapters on shared base weights, batched together; the LoRA ID is part of every block hash.

See it in the lab →
Operations

N+1 capacity

One spare node’s worth of capacity so a replacement can be Ready before the old node drains.

See it in the lab →
Network

NCCL

NVIDIA Collective Communications Library: all-reduce, all-gather, reduce-scatter, broadcast, all-to-all across GPUs.

See it in the lab →
Network

NIXL

NVIDIA Inference Xfer Library: async KV transfer over RDMA/NVLink/storage; vLLM’s NixlConnector uses it for P/D.

See it in the lab →
Operations

Node auto repair

Karpenter feature (alpha) that force-replaces nodes whose health conditions stay bad, e.g. AcceleratedHardwareReady false for 10 minutes. It pauses if over 20% of a pool is unhealthy.

See it in the lab →
Platform

Node class

The cloud launch settings a NodePool uses (EC2NodeClass on AWS): image, subnets, security groups, disks and capacity reservations.

See it in the lab →
Platform

NodeClaim

One concrete node Karpenter decided to launch (instance type, zone, capacity type). It’s initialized once the node is Ready, startup taints are gone and requested resources are registered.

See it in the lab →
Platform

NodePool

A Karpenter policy for new nodes: allowed instance types, zones and capacity types, plus taints, limits and disruption rules. Ready means valid config, not running nodes.

See it in the lab →
Platform

NodePool limits

A cap on the total CPU, memory or GPUs a NodePool may provision. Once it’s reached, Karpenter launches nothing more and pods stay Pending.

See it in the lab →
Platform

NVIDIA Dynamo

Orchestration above vLLM, SGLang or TensorRT-LLM: KV-aware routing, KV block manager tiers, planner for P/D.

See it in the lab →
Platform

nvidia.com/gpu

The extended resource the NVIDIA device plugin exposes; pods request whole GPUs through it.

See it in the lab →
Network

NVLink

Direct GPU-to-GPU links inside a node (hundreds of GB/s per GPU per direction).

See it in the lab →
Serving

P/D disaggregation

Separate prefill and decode GPU pools; KV moves between them. Better latency control; needs rate-matched pools.

See it in the lab →
Memory

PagedAttention

KV stored in fixed 16-token blocks mapped by a per-request block table; cut KV waste from 60–80% to under 4%.

See it in the lab →
Parallelism

Pipeline parallelism (PP)

Different layer ranges on different GPUs; micro-batches flow through; costs pipeline bubbles.

See it in the lab →
Operations

PodDisruptionBudget (PDB)

Limits how many replicas voluntary evictions may take down at once. Honoured by kubectl drain and Karpenter; can’t prevent involuntary failures.

See it in the lab →
Scheduling

Preemption

When KV blocks run out, the scheduler evicts a running request and recomputes it later.

See it in the lab →
Basics

Prefill

The phase that runs the whole prompt through the model in one parallel pass and writes the KV cache. Compute-bound; sets TTFT.

See it in the lab →
Memory

Prefix caching

Full KV blocks are hashed (chained through their parents) so later requests with the same prefix reuse them.

See it in the lab →
Optimization

Quantization

Fewer bits per weight/activation. W4A16 = 4-bit weights, 16-bit maths; W8A8 = 8-bit both (FP8 on Hopper).

See it in the lab →
Operations

Replacement-first

Maintenance rule: the new replica is Ready before the old one is evicted. Needs two or more spread replicas, a PDB and room for a surge node.

See it in the lab →
Platform

ResourceQuota

A namespace cap on requested resources, checked when pods are created. Over quota, pod creation is forbidden.

See it in the lab →
Frontier

RL rollouts

Samples generated by the current policy during RL post-training. An inference workload inside the training loop.

See it in the lab →
Metrics

Roofline / arithmetic intensity

FLOPs per byte moved. Below the GPU’s ridge (~300 FLOPs/byte on H100) you’re bandwidth-bound: decode lives there.

See it in the lab →
Platform

Semantic router

Chooses which model serves a request (cheap vs strong), with safety plugins. The EPP then picks the replica.

See it in the lab →
Optimization

Speculative decoding

A cheap drafter proposes k tokens; the target verifies them in one pass. Same output, more tokens per step at low load.

See it in the lab →
Serving

Sticky until saturated

Stay on the warm replica for a prefix or session until it’s overloaded, then spill to the next best.

See it in the lab →
Serving

Structured outputs

Grammar-constrained decoding: the allowed next tokens become a bitmask on the logits (XGrammar, llguidance).

See it in the lab →
Platform

Taint / toleration

A taint repels pods that don’t tolerate it. GPU nodes are usually tainted nvidia.com/gpu:NoSchedule so only GPU pods land there.

See it in the lab →
Parallelism

Tensor parallelism (TP)

Shard every layer’s matrices across GPUs; two all-reduces per layer per step. Keep it inside an NVLink node.

See it in the lab →
Operations

terminationGracePeriod (NodePool)

Upper bound on draining a node. After it, remaining pods are deleted even if a PDB or do-not-disrupt blocks them.

See it in the lab →
Basics

Token

A piece of text from the model’s fixed vocabulary, with an integer ID. Every cost in serving is counted in tokens.

See it in the lab →
Scheduling

Token budget

Max tokens per GPU step (--max-num-batched-tokens): decodes take 1 each, prefills take chunks of the rest.

See it in the lab →
Basics

Tokenizer (BPE)

Turns text into token IDs using merges learned from data (byte-pair encoding). English averages ~4 characters per token.

See it in the lab →
Serving

Tool calling

The chat template renders tools into the prompt; a model-specific parser turns the model’s call into OpenAI tool_calls.

See it in the lab →
Metrics

TTFT

Time to first token = queueing + prefill (+ KV transfer). What users feel as “thinking…”.

See it in the lab →
Engine

vLLM V1 / EngineCore

vLLM’s 2025 architecture: an API-server process plus an EngineCore busy loop (schedule → execute → sample → update).

See it in the lab →
Frontier

vLLM-Omni

Serving any-to-any models as a graph of stages (e.g. Thinker → Talker → vocoder), each with its own GPUs.

See it in the lab →
Platform

waitForPodsReady

A Kueue setting that evicts and requeues an admitted job whose pods don’t all become ready in time, so it stops holding quota.

See it in the lab →
Frontier

Weight sync

Pushing updated policy weights to rollout engines. Reset the prefix cache afterwards, or KV goes stale.

See it in the lab →
Parallelism

Wide EP

DP attention + expert parallelism across many GPUs, as used for DeepSeek-class models (~2.2k tok/s per H200 decode).

See it in the lab →