Polyaxon v3 is coming →

Set GPU quotas for shared Kubernetes namespaces

Limit aggregate GPU requests with ResourceQuota, supply CPU and memory defaults, and align Polyaxon resource profiles and queues with namespace budgets.

October 2, 2026by Polyaxon

Suppose a namespace has three training Pods, each requesting one GPU. Its quota allows four. A new two-GPU Pod would take the total to five, so Kubernetes rejects its creation—even if another node has free GPUs.

ResourceQuota lets a platform team enforce a shared namespace budget. For ML workloads, the useful combination is a GPU quota, explicit CPU and memory requirements, and a queue that controls how much work is submitted at once.

White Kubernetes wheel inside a blue hexagon on a light blue gradient background.

The examples use an existing workload namespace named ml-team and NVIDIA GPUs exposed as nvidia.com/gpu through a device plugin. They assume exclusive GPU allocation. If you use GPU sharing or DRA, check what the requested resource counts. Applying namespace policies requires administrator permissions.

Set the namespace budget

Save this as ml-quota.yaml:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: ml-budget
  namespace: ml-team
spec:
  hard:
    requests.nvidia.com/gpu: "4"
    requests.cpu: "12"
    requests.memory: 48Gi
    limits.cpu: "24"
    limits.memory: 64Gi

An administrator can apply it with:

kubectl apply -f ml-quota.yaml
kubectl describe resourcequota ml-budget -n ml-team

The extended-resource quota key is requests.nvidia.com/gpu, even when a container declares its GPU only under limits. The quota counts requested devices across non-terminal Pods, including Pods still waiting for placement; low GPU activity does not release that budget. See the resource quota reference.

The CPU and memory fields constrain both aggregate requests and aggregate limits. Containers need the corresponding declarations, supplied explicitly or by defaults. The CPU and memory quota walkthrough explains that accounting.

These values are an example budget. They neither add nodes nor reserve capacity for the namespace.

Supply defaults without undersizing the trainer

A LimitRange can fill in missing CPU and memory fields. Save this as ml-defaults.yaml:

apiVersion: v1
kind: LimitRange
metadata:
  name: container-defaults
  namespace: ml-team
spec:
  limits:
    - type: Container
      defaultRequest:
        cpu: 100m
        memory: 128Mi
      default:
        cpu: "1"
        memory: 512Mi
kubectl apply -f ml-defaults.yaml

These are small fallback values, not a training profile. Set explicit resources for trainers, data loaders, and artifact helpers based on their needs.

Declare both CPU and memory requests and limits for the trainer: a request of 8Gi combined with an omitted limit could conflict with this default limit of 512Mi. Defaults apply when new Pods are admitted; they do not resize existing Pods. See LimitRange behavior.

Apply a resource profile in Polyaxon

Use a Polyaxon scheduling preset to share an explicit resource profile across training components. Save this as gpu-profile.yaml:

queue: research/gpu
patchStrategy: post_merge
runPatch:
  container:
    resources:
      requests:
        cpu: "2"
        memory: 8Gi
        nvidia.com/gpu: 1
      limits:
        cpu: "4"
        memory: 12Gi
        nvidia.com/gpu: 1

Apply it to your existing training component:

polyaxon run -f train.yaml -f gpu-profile.yaml

train.yaml must contain your working component, including its image, training command, code, and data connections. This preset merges the listed settings over the component's settings. Use a profile that fits your trainer and replace research/gpu with an allowed agent/queue targeting the workload namespace.

The preset sets both GPU fields to one because GPU requests and limits must match. Kubernetes also accepts a GPU limit alone and uses it as the request. The GPU scheduling documentation describes this rule.

Four main containers using this profile request eight CPUs and 32Gi of memory, with limits totaling 16 CPUs and 48Gi. That leaves room within the example quota for helpers, but inspect the whole resolved Pod, including sidecars and init containers, before choosing the namespace budget.

Polyaxon queues can control concurrency, resource quotas, and routing in the commercial editions. For independent one-GPU jobs, a concurrency of four is a useful starting point. It is not a four-GPU cap when operations request different GPU counts or create multiple workers.

The Kubernetes quota applies to every workload in the namespace: training jobs, notebooks, inference services, and work submitted outside Polyaxon. Projects and team spaces sharing that namespace share its budget. Coordinate all queues targeting it; a project or team-space boundary does not create a separate Kubernetes quota.

Find the reason a run did not start

For a native Job that exists without a Pod, inspect its events and the namespace policies. Replace the Job name below with the actual owning workload name:

kubectl describe job TRAINING_JOB_NAME -n ml-team
kubectl describe resourcequota -n ml-team
kubectl get limitrange -n ml-team -o yaml

A quota violation rejects creation with 403 Forbidden; a Job controller can report that failure in its events. Read which resource was requested, which quota was exceeded, and the reported usage. Do not reduce the GPU request unless the training program can actually run with fewer devices. The quota walkthrough shows rejection before a Pod exists.

If Polyaxon has not submitted the workload yet, inspect its queue and run status first. If a Pod exists but has no node, investigate resource fit and placement constraints. The Pending GPU guide follows those stages.

For a temporary policy exercise, remove only the objects created above:

kubectl delete resourcequota ml-budget -n ml-team
kubectl delete limitrange container-defaults -n ml-team

For a shared production namespace, maintain the quota alongside the Polyaxon resource profiles and queues that use it. Include long-lived notebooks and inference services in the budget before allocating the remainder to training.