Polyaxon v3 is coming →

Keep inference available during Kubernetes node drains

Set a PodDisruptionBudget around usable inference capacity, inspect blocked evictions, and account for replacement GPUs and model warmup.

September 29, 2026by Polyaxon
White Kubernetes wheel inside a blue hexagon on a pale blue background

Three inference replicas are serving traffic when a GPU node needs maintenance. Removing one replica may be acceptable. Removing two before the first replacement loads its model may leave the service unable to keep up.

A PodDisruptionBudget (PDB) lets Kubernetes limit API-initiated evictions of a selected group of Pods. The useful budget comes from the service's measured capacity under reduced availability, including the time needed to schedule and warm a replacement. It does not come from choosing a percentage that looks reasonable.

This article follows planned node maintenance. The broader Pod eviction guide covers node pressure, preemption, and failures that require different recovery decisions.

Define how many replicas must remain usable

Suppose a service normally runs three independent, one-GPU replicas. Load measurements show that two ready replicas can handle the expected traffic during maintenance. An initial budget is therefore minAvailable: 2.

This is an illustrative capacity decision, not a throughput claim. Establish it with your model, request lengths, concurrency, latency objective, and hardware. A replica with a listening HTTP socket but an unloaded model should not count as ready. Connect the model server's readiness behavior to actual serving capability; the startup-probe guide covers slow initialization.

Assuming the budget controller has reconciled the current state and there are no outstanding evictions:

Ready replicasMinimum requiredFurther eviction of a healthy replica
32One can be permitted
22Wait for capacity to recover
12Wait; the service is already below its budget

The PDB does not measure request latency or GPU utilization. Its health accounting uses Pod readiness. A bad readiness signal can therefore produce a misleading budget. PDB API reference.

Select only the intended serving group

Assume an existing ml-serving namespace and three replicas labeled ml.example.com/serving-group: recommender-prod. Their workload controller already owns replacement behavior. The label is an example your platform must add to the actual Pod template.

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: recommender-prod
  namespace: ml-serving
spec:
  minAvailable: 2
  unhealthyPodEvictionPolicy: AlwaysAllow
  selector:
    matchLabels:
      ml.example.com/serving-group: recommender-prod

Keep unrelated services, evaluation clients, and preview deployments out of this selector. In policy/v1, an empty selector matches every Pod in the namespace, so omitting the intended labels is consequential.

An integer minAvailable also makes the example usable with controllers for which Kubernetes cannot infer replica counts through a supported scale interface. Percentage budgets and maxUnavailable have additional controller requirements. For small services, percentage rounding deserves particular attention: maxUnavailable: "30%" can permit evicting the only replica because Kubernetes rounds that allowance up to one. Configuring a PDB.

Decide what should happen to an unhealthy replica

AlwaysAllow permits eviction of running Pods that are not ready even when the healthy-replica budget is already exhausted. Healthy Pods remain subject to the budget. This helps a broken replica avoid blocking a node drain indefinitely. The default IfHealthyBudget policy can retain an unhealthy running Pod while the application is below its desired health count. Unhealthy Pod eviction policy.

For inference, consider a replica that is still loading a large model. It is unready, but it may soon become useful. With AlwaysAllow, maintenance can interrupt that warmup. Agree on the maintenance window and replacement sequence so restarting slow initialization repeatedly does not prevent recovery.

The setting is a decision about unhealthy Pods during eviction. It does not declare them healthy or repair the application.

Inspect the budget before changing maintenance behavior

For this example, inspect both the selector and the current accounting:

kubectl get pods -n ml-serving \
  -l ml.example.com/serving-group=recommender-prod -o wide

kubectl get pdb recommender-prod -n ml-serving \
  -o jsonpath='{.status.currentHealthy}{"\n"}{.status.desiredHealthy}{"\n"}{.status.disruptionsAllowed}{"\n"}{.status.observedGeneration}{"\n"}{.metadata.generation}{"\n"}'

kubectl describe pdb recommender-prod -n ml-serving

The output separates current health, desired health, and allowed disruptions. Compare the observed and metadata generations when assessing whether status reflects the latest specification. If a drain is waiting, inspect the unready or replacement Pod's events before weakening the budget.

Replacement stateNext evidence
No replacement PodWorkload-controller behavior and events
Pending without a nodeGPU availability, placement constraints, and storage topology
Assigned but not runningImage pulls, volume mounts, and container startup
Running but unreadyModel load, health endpoint, dependencies, and warmup

A node being drained is unavailable for new scheduling. If all remaining eligible GPUs are occupied, evicting the first replica does not create a usable GPU elsewhere. The next eviction can remain blocked while its replacement waits. Reserve replacement capacity or agree on a maintenance sequence that can complete with the available nodes.

Keep maintenance and application rollout policies distinct

PDBs constrain the eviction path used by tools such as kubectl drain. Direct Pod deletion and involuntary disruptions do not provide the same protection. Deployment and StatefulSet rolling updates use their own update rules; unavailable rollout Pods still affect budget accounting, but the PDB does not govern the rollout itself. Kubernetes disruptions.

Avoid overlapping a model rollout with a node drain unless the combined loss of capacity has been planned. The GPU-constrained rollout guide handles the separate question of replacing an application version without spare GPUs.

Carry the serving policy into Polyaxon

For a replicated Polyaxon service, use its service runtime and environment labels to make the intended group explicit. This is a fragment to merge into an existing inference component with its container, resources, ports, and readiness configuration:

run:
  kind: service
  replicas: 3
  environment:
    labels:
      ml.example.com/serving-group: recommender-prod

The PDB is a separate Kubernetes resource in the agent's workload namespace. Confirm the resolved Pod labels, replica ownership, and replacement behavior before adopting the budget. Giving two different services the same label would make them share the same accounting, even if they serve different models.

Use a scheduling preset to retain the appropriate GPU pool and resource settings. Record the service revision, PDB revision, replica count, replacement startup time, and reduced-capacity measurements alongside your Polyaxon benchmark runs.

Training and notebooks need their own policy. A budget that blocks eviction of a single notebook may also block maintenance; it does not save the kernel's memory state. A restartable training job benefits from durable checkpoints and a defined retry owner. Choose those recovery contracts before applying a serving-style availability budget to every workload.