Track 07 · Advanced · ~40 min · lessons from on-call
GPUs rarely vanish. They get stuck.
Most GPU outages on Kubernetes aren’t broken hardware. A pod waits for a node that can never exist, a job holds quota it can’t use, or routine maintenance evicts every replica before a replacement is ready. Follow one GPU from pool → node → pod, debug a Pending pod, find the gate that blocks a job, then patch a GPU node without taking the model down.
From “we have a GPU pool” to “a GPU is doing work.”
A GPU passes through four states, and each one is a different number on a different dashboard. Step through one pod’s life and watch the counters. Then break the device plugin and see which number lies to you.
Enabled → allocatable → requested → consumed
- NodePoolPolicy: GPU families, capacity type, taint, limits
- Pod asks
nvidia.com/gpu: 4+ toleration - NodeClaimAutoscaler picks an offering and launches it
- Node joinsBoots; kubelet Ready; CNI up
- Device pluginDriver + toolkit, then GPUs advertised
- Bound podScheduler binds; GPUs do work
nvidia.com/gpu taint and a limit. It reports Ready, yet no GPU exists. Enabled is a promise, not capacity.Enabled ≠ allocated
A Ready NodePool is configuration, a promise. Only a NodeClaim and a registered node prove capacity exists, and only bound pods prove it’s used.
Node Ready ≠ GPU ready
The kubelet reports Ready before the driver, container toolkit and device plugin finish. Until the plugin registers, allocatable nvidia.com/gpu is 0.
Idle isn’t free
An empty GPU node bills until the NodePool’s disruption policy removes it. Karpenter’s default is consolidateAfter: 0s; GPU pools often wait longer to avoid churn, and pay for that wait.
Staff answer
“When someone says ‘we have no GPUs’, I ask which of four numbers they mean: enabled pools, allocatable GPUs on registered nodes, GPUs requested by pods, or GPUs bound to running pods. Each points to a different owner: platform config, autoscaler and cloud capacity, the workload, or the scheduler.”
Why is my GPU pod Pending?
A node autoscaler launches a node only if some offering survives every filter: the pod’s requests and tolerations, each NodePool’s requirements, zones, capacity type, and what the cloud can actually give you. Change the pod or the world, and see which filter rules out each pool.
Which NodePool can host this pod?
✗ GPU: offers no nvidia.com/gpu
✗ label: selector wants L40S; pool offers L4
✓ launchable in zone a
✗ label: selector wants L40S; pool offers H100
✗ label: selector wants L40S; pool offers H100
Schedulable: the autoscaler launches a node from gpu-l40s-od in zone a
Only one pool qualifies. That’s fragile: one sold-out zone or an expired reservation and this pod waits.
Scheduler, then autoscaler
The scheduler only places pods on nodes that exist; FailedScheduling says why each was rejected. The autoscaler then simulates new nodes from every NodePool. Read both messages.
Flexibility is capacity
Every hard constraint (one zone, one instance type, reserved only) shrinks the offering set. When one zone runs dry, flexible pods still land somewhere else.
Reservations have edges
A capacity reservation covers one instance type, in one zone, until an end time. Outside that box a reserved-only pod has nowhere to go, even while the pool shows Ready.
Staff answer
“I debug a Pending GPU pod as a set intersection. Pod first: does it request the GPU resource, tolerate the GPU taint, select a label that exists? Then each NodePool’s requirements and limits, then the node class, then the cloud: offering in that zone, quota, reservation state. Whatever empties the set is the fix. ‘The cloud has no capacity’ is only the answer once everything else has passed.”
Four gates between submit and running.
A GPU batch job can be stopped by four different budgets owned by four different teams. Kueue decides when a job may start, ResourceQuota caps a namespace, NodePool limits cap the autoscaler, and the cloud caps everything. Size the job and each budget, then watch where it stops.
The job runs
- Quota admins · ClusterQueueKueue admission16 needed · 24 free
Admitted within nominal quota (16 of 24 free). - Namespace adminsNamespace ResourceQuota16 needed · 32 free
Pods created (16 of 32 GPUs left in the namespace). - Platform · FinOpsNodePool limit16 needed · 32 free
Autoscaler may launch 16 GPUs (32 left under the NodePool limit). - Capacity ownersCloud capacity16 needed · 32 free
The cloud has 32 GPUs obtainable: all nodes launch and the job runs.
Kueue is admission, not capacity
Kueue keeps a job suspended until its ClusterQueue quota, plus anything it may borrow from the cohort, covers it. It doesn’t create nodes; the autoscaler still has to find them.
Admitted ≠ running
An admitted job whose pods can’t be placed holds quota that other teams could use. A pods-ready timeout, or an admission check that provisions capacity first, releases it.
Each gate has its own symptom
Kueue: workload pending, no pods yet. ResourceQuota: pod creation forbidden. NodePool limit: pods Pending, autoscaler refuses. Cloud: launch fails on capacity or quota.
Staff answer
“When a GPU job doesn’t start I walk the gates in order: is the workload admitted, were pods created, did the autoscaler try to launch, and what did the cloud say? Each gate has a different owner, so naming the gate is how the right team gets paged. I also watch quota held by admitted-but-unscheduled jobs; it’s invisible waste.”
Patch a GPU node without taking the model down.
GPU nodes must be rotated for AMI, kernel and driver patches. How you rotate decides whether users notice. Pick the rotation method, replicas, placement, disruption budget and spare capacity, then replay a two-hour maintenance window minute by minute.
Ready replicas during a node rotation
- t+0mRotation job cordons node-1 and evicts its pods
- t+0mreplica-1 evicted
- t+0mreplica-2 evicted
- t+0mnode-1 is empty; the NodePool waits consolidateAfter (30 min) before removing it
- t+0mNo free capacity: the reservation is full while node-1 still exists
- t+31mnode-1 terminated: its capacity slot is free
- t+31mAutoscaler launches node-2 for replica-1 (boot + GPU software: 8 min)
- t+39mnode-2 Ready, GPUs registered
- t+39mreplica-1 scheduled on node-2; loading the model (12 min)
- t+39mreplica-2 scheduled on node-2; loading the model (12 min)
- t+51mreplica-1 Ready on node-2
- t+51mreplica-2 Ready on node-2
- t+51mRotation complete: old node gone, every replica Ready
consolidateAfter (30 min), then termination freed the slot, the new node booted (8 min) and the model loaded (12 min). The outage is roughly the sum of those timers; a shorter consolidateAfter trims it but can’t remove it.Graceful vs forceful
Karpenter’s drift and consolidation pre-spin a replacement node and wait for it before draining. Expiration, interruptions and an external drain evict first and find room later.
A PDB trades an outage for a stuck drain
Eviction honours a PodDisruptionBudget, so a drain blocks instead of taking the last replica. If there’s nowhere to reschedule, maintenance stalls. That’s the right failure; alert on it.
N+1, or a window
Replacement-first needs somewhere to go. If reserved GPUs are sized exactly to demand, you need one spare node, an on-demand fallback, or a planned maintenance window.
Node ready ≠ model ready
Pre-spinning a node doesn’t pre-load weights. Image pull, weight load and warm-up must finish before the next eviction, or you still drop to zero.
Staff answer
“For GPU serving maintenance my rule is replacement-first: the new replica is Ready before the old one is evicted. That takes three things together: at least two replicas spread across nodes, a PDB with maxUnavailable: 1, and room for one surge node. If the reservation has no spare, I’d rather the drain block and page someone than evict into a void. And I prefer the autoscaler’s graceful drift over a home-grown drain job.”
Case file · anonymized, illustrative
A scheduled node-rotation job picked a week-old GPU node, cordoned it and evicted both replicas of a model. They were packed on that one node with no PDB. The GPU pool was tied to a reservation with no spare slot, so no replacement could launch. The empty node lingered for the pool’s consolidateAfter window before termination freed the slot; then a new node booted and the model loaded. Users saw roughly the sum of those timers.
Fixes: spread + PDB, N+1 reserved capacity, a rotation job that checks for spare capacity before it evicts, and alerts on “evicted with nowhere to go”. Press Reproduce the outage above to replay it.
Check yourself.
Questions platform and SRE loops ask about GPU capacity and maintenance. Then flip the cards until the definitions are automatic.
1. A GPU NodePool shows Ready, yet the team says “we have no GPUs”. What does Ready prove?
2. A GPU node is Ready but allocatable nvidia.com/gpu is 0, and the pod stays Pending. Where do you look first?
3. Kueue admitted a 32-GPU job, but its pods are Pending because the NodePool limit is reached. What is the hidden cost?
4. Which Karpenter disruption methods pre-spin a replacement node before draining?
5. Two replicas share one node, there’s no PDB, and the GPU reservation has no spare slot. A rotation job drains the node. What happens?
6. You add a PDB (maxUnavailable: 1) but still have no spare capacity. What changes?
0 / 6 answered
Flashcards
Simulations are illustrative models of public Kubernetes, Kueue and Karpenter behaviour (Karpenter v1 disruption docs, Kubernetes disruption docs, Kueue overview; checked Sep 2026). Timings are examples, not measurements of any real environment. Event text is paraphrased; exact wording varies by version.