Align CPU and GPU resources with Kubernetes NUMA policies
Combine CPU Manager eligibility, full-core allocation, and Topology Manager policy, then compare CPU/GPU workloads with Polyaxon.
A GPU inference process also needs CPUs for tokenization, preprocessing, request handling, and host-to-device work. Two runs with the same CPU and GPU counts can behave differently when they receive different physical placements inside a large server.
On a NUMA machine, processors and devices have locality relationships. Kubernetes CPU Manager can give eligible containers exclusive CPU allocations, while Topology Manager coordinates locality hints from resource providers. These mechanisms are worth evaluating when measurements point to CPU contention or locality; they are not a reason to make every ML workload use the strictest node policy.
Start with the CPU requests and throttling guide if the problem is an undersized CPU budget. This article addresses the physical allocation behind that budget.
Separate the three decisions
| Decision | Responsible mechanism | Question it answers |
|---|---|---|
| Choose a node | Scheduler and workload placement constraints | Which node satisfies the Pod's resource and placement requirements? |
| Allocate exclusive CPUs | Kubelet CPU Manager with the static policy | Which CPUs can an eligible container use? |
| Align participating resources | Kubelet Topology Manager and resource providers | Can their allocations satisfy the requested NUMA policy? |
A scheduler placement can therefore be followed by a kubelet admission failure. Aggregate free CPU and GPU counts do not describe every possible arrangement within the node. Topology Manager documentation.
Device locality also depends on the device plugin supplying useful topology information. A policy cannot infer an unreported GPU connection, and it cannot prove that the application benefits from the resulting allocation.
Establish CPU Manager eligibility
With the static CPU Manager policy, exclusive allocation applies to containers in a Guaranteed Pod with whole-number CPU requests. Fractional requests continue to use the shared pool. For the container-level resource configuration used here, set equal CPU requests and limits and equal memory requests and limits on every relevant container so the resulting Pod qualifies. CPU management policies.
Check the resolved Pod, including init containers and sidecars. A main container configured with four CPUs does not establish the QoS class of the whole Pod when another container has missing or unequal settings. Pod QoS criteria.
CPU quantity is also not always a physical-core count. On an SMT machine with two hardware threads per core, four allocatable logical CPUs can correspond to two full physical cores. The full-pcpus-only option asks CPU Manager to allocate complete cores; requests that cannot satisfy that alignment can fail with SMTAlignmentError. Review auxiliary-container requests against the same policy. Resource-manager options.
Exclusive CPU allocation does not eliminate shared-cache, memory-bandwidth, storage, or network contention. System processes also need their own placement and reservation strategy.
Use a deliberately configured node pool
For a Linux pool running Kubernetes 1.33 or later, where full-pcpus-only is stable, a platform owner could use this kubelet configuration fragment for experiments with full-core allocation and strict locality:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
cpuManagerPolicy: static
cpuManagerPolicyOptions:
full-pcpus-only: "true"
topologyManagerPolicy: single-numa-node
topologyManagerScope: containerThis is not a complete kubelet configuration or a Kubernetes object to apply. Static CPU management requires nonzero CPU reservation for system or Kubernetes use. The owner must choose reservations and compatible settings for the actual machine, then follow the distribution's node-maintenance procedure when changing policy. CPU Manager retains allocation state, so changing policy is not a casual edit to a busy node. CPU Manager configuration.
The container scope evaluates alignment per container. A pod scope evaluates the Pod as a group and can be more restrictive. The single-numa-node policy rejects an allocation when the participating providers cannot satisfy a single-domain alignment; best-effort can accept a less favorable arrangement. Choose the tradeoff around the workload. Topology policies and scopes.
This fragment does not enable Memory Manager. Equal memory requests and limits establish part of the QoS contract, not a claim that memory pages are NUMA-bound. Keep memory-allocation policy as a separately reviewed configuration.
Explain a rejection with a physical inventory
Consider a container requesting four logical CPUs and one GPU. Assume a two-thread-per-core server, correct device topology hints, and otherwise compatible resources. This synthetic snapshot has enough resources in total:
| NUMA domain | Free allocatable logical CPUs, in full cores | Available GPUs |
|---|---|---|
| 0 | 2 | 1 |
| 1 | 8 | 0 |
There are ten free logical CPUs and a GPU, but the GPU's domain has only two CPUs available. A strict single-domain allocation cannot satisfy this request. The platform may need to free suitable CPUs, use another node, or select a less restrictive policy based on measured performance.
A Topology Manager rejection is not simply a Pending Pod waiting for the scheduler to move it. The failed Pod needs replacement through its workload controller or another recovery mechanism. Bound retries and retain the admission evidence so repeated attempts do not hide a persistent topology mismatch.
Compare the workload in Polyaxon
Use Polyaxon to keep the benchmark workload consistent across node profiles. The following component assumes the platform has labeled a correctly configured pool ml.example.com/cpu-profile: full-core-numa, installed the GPU device plugin, and provided a compatible agent and queue.
The image and benchmark module are reader-supplied. They must contain your model, fixed input workload, and result instrumentation. Replace the illustrative tag with an approved image digest for comparisons.
version: 1.1
kind: component
name: numa-inference-comparison
run:
kind: job
environment:
nodeSelector:
ml.example.com/cpu-profile: full-core-numa
container:
image: registry.example.com/ml/numa-inference:approved
command: ["python", "-m", "benchmark"]
resources:
requests:
cpu: "4"
memory: 16Gi
limits:
cpu: "4"
memory: 16Gi
nvidia.com/gpu: 1Save it as numa-comparison.yaml, adapt any injected container resources through your workload configuration, and submit it to the intended pool:
polyaxon run -p YOUR_PROJECT -f numa-comparison.yaml --queue AGENT/QUEUEThe node selector chooses the labeled pool; it does not turn on CPU Manager. Use node scheduling and scheduling presets to keep the approved profiles reusable. Resource scheduling supplies the CPU, memory, and GPU request.
For a second comparison, select the platform's other measured profile while holding model revision, hardware generation, request shape, concurrency, and resource quantities fixed. Record queue delay as well as execution time: a strict placement that runs faster but waits much longer may be a poor fit for short experiments.
Inspect the actual allocation
For a running comparison, replace the Pod and main-container names below:
kubectl get pod POD_NAME -n ml-team \
-o jsonpath='{.spec.nodeName}{"\n"}{.status.qosClass}{"\n"}'
kubectl describe pod POD_NAME -n ml-team
kubectl exec POD_NAME -n ml-team -c MAIN_CONTAINER -- \
python -c 'import os; print(sorted(os.sched_getaffinity(0)))'The Python command assumes a Linux container with Python installed. It reports the calling thread's allowed CPU IDs, not the GPU's NUMA attachment or application throughput. Have the node owner map those IDs to the machine topology and inspect admission events when the Pod never starts.
Retain the node profile, observed QoS, CPU set, relevant device placement, and workload measurements with the run using Polyaxon tracking. Compare latency distributions, preprocessing time, and GPU utilization rather than treating a smaller CPU set as proof of improvement. Stop any remaining benchmark runs when finished; keep node-policy changes under the platform owner's maintenance process.