inference atlas

Track 07 · Advanced · ~40 min · lessons from on-call

GPUs rarely vanish. They get stuck.

Most GPU outages on Kubernetes aren’t broken hardware. A pod waits for a node that can never exist, a job holds quota it can’t use, or routine maintenance evicts every replica before a replacement is ready. Follow one GPU from pool → node → pod, debug a Pending pod, find the gate that blocks a job, then patch a GPU node without taking the model down.

7.1Intermediatestep player

From “we have a GPU pool” to “a GPU is doing work.”

A GPU passes through four states, and each one is a different number on a different dashboard. Step through one pod’s life and watch the counters. Then break the device plugin and see which number lies to you.

Step player · one GPU pod, end to end

Enabled → allocatable → requested → consumed

  1. NodePoolPolicy: GPU families, capacity type, taint, limits
  2. Pod asksnvidia.com/gpu: 4 + toleration
  3. NodeClaimAutoscaler picks an offering and launches it
  4. Node joinsBoots; kubelet Ready; CNI up
  5. Device pluginDriver + toolkit, then GPUs advertised
  6. Bound podScheduler binds; GPUs do work
Enabled pools0
Allocatable GPUs0
Requested GPUs0
Consumed GPUs0
GPU nodenone
Enabled
A platform team defines a GPU NodePool: which GPU families, on-demand or reserved, the nvidia.com/gpu taint and a limit. It reports Ready, yet no GPU exists. Enabled is a promise, not capacity.

Enabled ≠ allocated

A Ready NodePool is configuration, a promise. Only a NodeClaim and a registered node prove capacity exists, and only bound pods prove it’s used.

Node Ready ≠ GPU ready

The kubelet reports Ready before the driver, container toolkit and device plugin finish. Until the plugin registers, allocatable nvidia.com/gpu is 0.

Idle isn’t free

An empty GPU node bills until the NodePool’s disruption policy removes it. Karpenter’s default is consolidateAfter: 0s; GPU pools often wait longer to avoid churn, and pay for that wait.

Staff answer

“When someone says ‘we have no GPUs’, I ask which of four numbers they mean: enabled pools, allocatable GPUs on registered nodes, GPUs requested by pods, or GPUs bound to running pods. Each points to a different owner: platform config, autoscaler and cloud capacity, the workload, or the scheduler.”

7.2Advanceddebugger

Why is my GPU pod Pending?

A node autoscaler launches a node only if some offering survives every filter: the pod’s requests and tolerations, each NodePool’s requirements, zones, capacity type, and what the cloud can actually give you. Change the pod or the world, and see which filter rules out each pool.

Debugger · pod constraints ∩ NodePools ∩ cloud offerings

Which NodePool can host this pod?

The pod
GPUs per pod
Accelerator selector
Capacity type
Zone
The world
H100 reservation (zone a)
Sold out (insufficient capacity)
Five NodePools, eight filters
cpu-generalCPU onlyon-demand · spotzones a, b
GPUtaintlabelcap typeshapeEFAzonecloud

✗ GPU: offers no nvidia.com/gpu

gpu-l4-odL4 × 1 per nodeon-demandGPU taintzones a, b
GPUtaintlabelcap typeshapeEFAzonecloud

✗ label: selector wants L40S; pool offers L4

gpu-l40s-odL40S × 4 per nodeon-demandGPU taintzones a, b
GPUtaintlabelcap typeshapeEFAzonecloud

✓ launchable in zone a

gpu-h100-rsvH100 × 8 per nodereservedEFAreservation in zone a
GPUtaintlabelcap typeshapeEFAzonecloud

✗ label: selector wants L40S; pool offers H100

gpu-h100-odH100 × 8 per nodeon-demandEFAzones a, b
GPUtaintlabelcap typeshapeEFAzonecloud

✗ label: selector wants L40S; pool offers H100

launchable = pod ∩ NodePool requirements ∩ node class ∩ cloud offerings (zone · capacity type · quota · reservation)

Schedulable: the autoscaler launches a node from gpu-l40s-od in zone a

Only one pool qualifies. That’s fragile: one sold-out zone or an expired reservation and this pod waits.

Normal Nominated pod will run on NodeClaim gpu-l40s-od-x7k2p Normal Launched gpu-l40s-od: instance in zone a Normal Scheduled after node Ready + GPUs registered
Pending, for now
Pending isn’t always a bug: the pod waits a few minutes for the node to boot and register its GPUs. Check for a NodeClaim before paging anyone. Now break it: remove the toleration, pin a zone, or expire the reservation.

Scheduler, then autoscaler

The scheduler only places pods on nodes that exist; FailedScheduling says why each was rejected. The autoscaler then simulates new nodes from every NodePool. Read both messages.

Flexibility is capacity

Every hard constraint (one zone, one instance type, reserved only) shrinks the offering set. When one zone runs dry, flexible pods still land somewhere else.

Reservations have edges

A capacity reservation covers one instance type, in one zone, until an end time. Outside that box a reserved-only pod has nowhere to go, even while the pool shows Ready.

Staff answer

“I debug a Pending GPU pod as a set intersection. Pod first: does it request the GPU resource, tolerate the GPU taint, select a label that exists? Then each NodePool’s requirements and limits, then the node class, then the cloud: offering in that zone, quota, reservation state. Whatever empties the set is the fix. ‘The cloud has no capacity’ is only the answer once everything else has passed.”

7.3Advancedsimulator

Four gates between submit and running.

A GPU batch job can be stopped by four different budgets owned by four different teams. Kueue decides when a job may start, ResourceQuota caps a namespace, NodePool limits cap the autoscaler, and the cloud caps everything. Size the job and each budget, then watch where it stops.

Simulator · one 8-GPU-per-pod job

The job runs

  1. Quota admins · ClusterQueueKueue admission16 needed · 24 free
    Admitted within nominal quota (16 of 24 free).
  2. Namespace adminsNamespace ResourceQuota16 needed · 32 free
    Pods created (16 of 32 GPUs left in the namespace).
  3. Platform · FinOpsNodePool limit16 needed · 32 free
    Autoscaler may launch 16 GPUs (32 left under the NodePool limit).
  4. Capacity ownersCloud capacity16 needed · 32 free
    The cloud has 32 GPUs obtainable: all nodes launch and the job runs.
Running
All four budgets covered the job. In real clusters these are owned by different teams and change independently, which is why the job that ran yesterday is Pending today.

Kueue is admission, not capacity

Kueue keeps a job suspended until its ClusterQueue quota, plus anything it may borrow from the cohort, covers it. It doesn’t create nodes; the autoscaler still has to find them.

Admitted ≠ running

An admitted job whose pods can’t be placed holds quota that other teams could use. A pods-ready timeout, or an admission check that provisions capacity first, releases it.

Each gate has its own symptom

Kueue: workload pending, no pods yet. ResourceQuota: pod creation forbidden. NodePool limit: pods Pending, autoscaler refuses. Cloud: launch fails on capacity or quota.

Staff answer

“When a GPU job doesn’t start I walk the gates in order: is the workload admitted, were pods created, did the autoscaler try to launch, and what did the cloud say? Each gate has a different owner, so naming the gate is how the right team gets paged. I also watch quota held by admitted-but-unscheduled jobs; it’s invisible waste.”

7.4Advancedsimulatorcase file

Patch a GPU node without taking the model down.

GPU nodes must be rotated for AMI, kernel and driver patches. How you rotate decides whether users notice. Pick the rotation method, replicas, placement, disruption budget and spare capacity, then replay a two-hour maintenance window minute by minute.

Simulator · one model · node-1 is due for rotation at minute 0

Ready replicas during a node rotation

Rotation method
Replicas
Placement
GPU capacity
consolidateAfter (empty node)
Ready model replicas · GPU nodes up (minutes after rotation starts)
Event log
  1. t+0mRotation job cordons node-1 and evicts its pods
  2. t+0mreplica-1 evicted
  3. t+0mreplica-2 evicted
  4. t+0mnode-1 is empty; the NodePool waits consolidateAfter (30 min) before removing it
  5. t+0mNo free capacity: the reservation is full while node-1 still exists
  6. t+31mnode-1 terminated: its capacity slot is free
  7. t+31mAutoscaler launches node-2 for replica-1 (boot + GPU software: 8 min)
  8. t+39mnode-2 Ready, GPUs registered
  9. t+39mreplica-1 scheduled on node-2; loading the model (12 min)
  10. t+39mreplica-2 scheduled on node-2; loading the model (12 min)
  11. t+51mreplica-1 Ready on node-2
  12. t+51mreplica-2 Ready on node-2
  13. t+51mRotation complete: old node gone, every replica Ready
Outage · 0 replicas ready51 min
Degraded · below desired0 min
Rotationdone at 51 min
Peak GPU nodes1
Outage
The classic outage: every replica was evicted at once (no PDB, all on one node), and the reservation had no room for a replacement until node-1 was gone. node-1 sat empty for consolidateAfter (30 min), then termination freed the slot, the new node booted (8 min) and the model loaded (12 min). The outage is roughly the sum of those timers; a shorter consolidateAfter trims it but can’t remove it.

Graceful vs forceful

Karpenter’s drift and consolidation pre-spin a replacement node and wait for it before draining. Expiration, interruptions and an external drain evict first and find room later.

A PDB trades an outage for a stuck drain

Eviction honours a PodDisruptionBudget, so a drain blocks instead of taking the last replica. If there’s nowhere to reschedule, maintenance stalls. That’s the right failure; alert on it.

N+1, or a window

Replacement-first needs somewhere to go. If reserved GPUs are sized exactly to demand, you need one spare node, an on-demand fallback, or a planned maintenance window.

Node ready ≠ model ready

Pre-spinning a node doesn’t pre-load weights. Image pull, weight load and warm-up must finish before the next eviction, or you still drop to zero.

Staff answer

“For GPU serving maintenance my rule is replacement-first: the new replica is Ready before the old one is evicted. That takes three things together: at least two replicas spread across nodes, a PDB with maxUnavailable: 1, and room for one surge node. If the reservation has no spare, I’d rather the drain block and page someone than evict into a void. And I prefer the autoscaler’s graceful drift over a home-grown drain job.”

Case file · anonymized, illustrative

A scheduled node-rotation job picked a week-old GPU node, cordoned it and evicted both replicas of a model. They were packed on that one node with no PDB. The GPU pool was tied to a reservation with no spare slot, so no replacement could launch. The empty node lingered for the pool’s consolidateAfter window before termination freed the slot; then a new node booted and the model loaded. Users saw roughly the sum of those timers.

Fixes: spread + PDB, N+1 reserved capacity, a rotation job that checks for spare capacity before it evicts, and alerts on “evicted with nowhere to go”. Press Reproduce the outage above to replay it.

7.5Quiz

Check yourself.

Questions platform and SRE loops ask about GPU capacity and maintenance. Then flip the cards until the definitions are automatic.

1. A GPU NodePool shows Ready, yet the team says “we have no GPUs”. What does Ready prove?

2. A GPU node is Ready but allocatable nvidia.com/gpu is 0, and the pod stays Pending. Where do you look first?

3. Kueue admitted a 32-GPU job, but its pods are Pending because the NodePool limit is reached. What is the hidden cost?

4. Which Karpenter disruption methods pre-spin a replacement node before draining?

5. Two replicas share one node, there’s no PDB, and the GPU reservation has no spare slot. A rotation job drains the node. What happens?

6. You add a PDB (maxUnavailable: 1) but still have no spare capacity. What changes?

0 / 6 answered

Flashcards

Simulations are illustrative models of public Kubernetes, Kueue and Karpenter behaviour (Karpenter v1 disruption docs, Kubernetes disruption docs, Kueue overview; checked Sep 2026). Timings are examples, not measurements of any real environment. Event text is paraphrased; exact wording varies by version.