Spread inference replicas across Kubernetes zones
Use topology spread constraints for inference availability, understand minDomains and GPU capacity, and preserve placement intent in Polyaxon.
Three inference replicas are running, but all three occupy the same zone. Replica count alone has not protected the service from a zone outage.
Kubernetes topology spread constraints let you distribute matching Pods across node labels such as topology.kubernetes.io/zone. The scheduler evaluates that distribution alongside GPU, memory, affinity, and storage requirements. A hard spread rule can improve the intended placement while making a replacement wait when the required zone lacks capacity.
This is a different objective from keeping distributed training workers close together. Here, each replica serves independently. A tensor-parallel replica containing several tightly coupled workers needs a separate policy for its internal placement.
Define the replicas and the failure domains
Choose a label that identifies replicas of the same serving group. A generic label such as workload: inference can mix unrelated models into one count. A unique label for every Pod has the opposite problem: the scheduler cannot recognize its peers.
Topology spread counts matching Pods in the incoming Pod's namespace. The topologyKey names a node label, while labelSelector selects the Pods whose placement should be compared. Keep those two roles explicit. Topology spread documentation.
For a one-GPU model replica, inventory eligible GPU nodes in each zone before setting a strict three-zone policy. Existing CPU-only nodes do not make that zone useful if a node selector excludes them, and a zone containing eligible nodes may still have no free GPUs.
Request a bounded difference between zones
This native Deployment assumes Kubernetes 1.33 or later, an existing ml-serving namespace, and GPU nodes labeled ml.example.com/pool: inference. The platform must supply accurate zone labels, a compatible GPU device plugin, image access, and any required taint tolerations.
Replace the image with your model server. It must start through its image entrypoint, listen on port 8000, and implement /ready so readiness becomes true only when it can serve. The manifest describes replica placement; expose it through your existing Service or gateway configuration.
apiVersion: apps/v1
kind: Deployment
metadata:
name: zonal-model-api
namespace: ml-serving
spec:
replicas: 3
selector:
matchLabels:
ml.example.com/serving-group: zonal-model-api
template:
metadata:
labels:
ml.example.com/serving-group: zonal-model-api
spec:
nodeSelector:
ml.example.com/pool: inference
topologySpreadConstraints:
- maxSkew: 1
minDomains: 3
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
nodeAffinityPolicy: Honor
nodeTaintsPolicy: Honor
labelSelector:
matchLabels:
ml.example.com/serving-group: zonal-model-api
containers:
- name: model
image: registry.example.com/ml/model-api:approved
ports:
- containerPort: 8000
resources:
requests:
cpu: "4"
memory: 16Gi
limits:
nvidia.com/gpu: 1
readinessProbe:
httpGet:
path: /ready
port: 8000
periodSeconds: 5The two Honor settings make domain counting respect the workload's node selection and tolerated taints. Those node-inclusion controls are stable from Kubernetes 1.33. They do not reserve GPUs in a domain. Node-inclusion policy, feature history.
With DoNotSchedule, maxSkew: 1 limits the difference from the global minimum count. minDomains: 3 makes that minimum zero when fewer than three domains are eligible. It does not wait for three replicas to be placed as a group, and it does not provision a missing zone.
Work through a constrained GPU example
Assume three eligible zones and no other matching Pods. After two replicas start, consider this illustrative snapshot:
| Zone | Matching Pods | Free compatible GPUs |
|---|---|---|
| A | 1 | 2 |
| B | 1 | 1 |
| C | 0 | 0 |
The third replica cannot use a GPU in C. Placing it in A or B would produce a count of two against C's zero, exceeding the allowed skew. It can remain Pending even though the cluster has three free GPUs.
Removing C from eligibility does not necessarily unblock it: with only A and B eligible and minDomains: 3, the global minimum remains zero. After one matching Pod in each remaining domain, the hard rule still prevents another there.
This may be exactly the policy you want. It may also be the wrong tradeoff during an outage when two zones could serve more traffic. Decide whether preserving the spread requirement or restoring replica count takes priority before the incident.
For a preference that can yield to capacity constraints, use whenUnsatisfiable: ScheduleAnyway and remove minDomains, which is only valid with DoNotSchedule. The scheduler then scores spread rather than requiring it. A preference cannot guarantee zone diversity.
Inspect placement at both levels
These commands show the available zone labels, the matching replicas, and a waiting Pod's events:
kubectl get nodes -l ml.example.com/pool=inference \
-L topology.kubernetes.io/zone
kubectl get pods -n ml-serving \
-l ml.example.com/serving-group=zonal-model-api -o wide
kubectl describe pod PENDING_POD_NAME -n ml-servingJoin the Pod's node name to the node's zone label. A Pod count by itself does not reveal zone placement, and aggregate GPU utilization does not reveal the resources available to this workload.
Topology information is derived from existing nodes. A zone whose node pool has scaled to zero may not be visible to the scheduler; the autoscaler must also understand the intended topology and provide compatible capacity. Do not assume the spread rule alone can recover a missing domain. Topology spread limitations.
The example selector also counts both old and new replicas during a rollout. Inspect the final distribution after the old replicas disappear. If each revision must be spread independently, review matchLabelKeys: [pod-template-hash] for the Deployment rather than assuming a balanced combined count guarantees a balanced new revision. Spread during rolling updates.
Preserve the placement intent in Polyaxon
Polyaxon exposes node selectors, labels, and affinity in the environment specification. A supported way to express a preference for separating service replicas is Pod anti-affinity:
run:
kind: service
replicas: 3
environment:
labels:
ml.example.com/serving-group: polyaxon-model-api
nodeSelector:
ml.example.com/pool: inference
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
topologyKey: topology.kubernetes.io/zone
labelSelector:
matchLabels:
ml.example.com/serving-group: polyaxon-model-apiMerge this fragment into an existing service component with the model container, ports, resources, and probes. It uses documented Polyaxon node scheduling fields. Preferred anti-affinity influences placement among feasible nodes; it does not implement the exact maxSkew and minDomains contract above. Kubernetes affinity rules.
If the exact native spread contract is required, apply it through your platform's supported Pod-template or admission integration and confirm the resolved Pod contains it. The documented Polyaxon environment schema used here does not expose a topologySpreadConstraints field; adding an invented field to the fragment would not establish support.
Retain the chosen placement configuration in a scheduling preset where appropriate. Use Polyaxon load-test jobs to compare request latency, available serving capacity, cross-zone dependencies, and replacement startup time. Keep the model revision and traffic shape fixed so the result explains the placement decision.
Zone spread complements a disruption budget. The first influences where replicas start; the second constrains API-initiated eviction. Neither configures routing or rebalances existing Pods automatically when capacity changes. Plan replacement and rollout behavior around the same capacity assumptions.
If you created the native example solely for a disposable trial, remove that Deployment when finished:
kubectl delete deployment zonal-model-api -n ml-servingFor a controller-managed or GitOps-managed service, retire it through its owner so it is not recreated.