Connect the cluster and define access
Install Polyaxon Agent (compute cluster) on the Kubernetes cluster and connect it to your Polyaxon control plane. The cluster needs working GPU drivers and device plugins, suitable storage, and network access for the workloads you plan to run.
Create projects for the teams and assign the roles they need. Keep responsibility for shared queue and preset settings with the platform administrators.
Set limits for teams and workflows
Give each team a queue with an explicit resource quota and concurrency limit. If interactive work needs a different dispatch priority from long-running experiments, give it a separate queue.
Set a default queue on each project and restrict the queues that project can use. For a large parameter sweep, also set workflow concurrency so its parallel work is bounded without changing the shared queue every time.
These are submission and dispatch controls, not a reservation of physical GPUs. Work may still wait in Kubernetes when compatible capacity is unavailable. Queue priority alone does not evict a running training job.
Manage queues and project restrictions · Limit workflow parallelism
Give teams reusable presets
Create presets for the workload shapes you actually support: a single-GPU experiment, a multi-GPU training run, or a notebook session. Include the appropriate resource requests, placement settings, queue, and connections.
Set project defaults where appropriate. A team can then change its training code or model inputs while reusing the cluster configuration maintained by the platform team.
Review preset merge behavior and project restrictions together. A convenient default is not a substitute for access control, Kubernetes policy, or resource quotas.
Create organization presets · Understand preset merge behavior
Explain why a job is waiting
Use the operation’s status timeline and logs to distinguish dispatch from execution. A run waiting on queue limits needs a different response from a dispatched workload that Kubernetes cannot place on a node.
When a run completes, inspect its metrics and outputs in the same operation record. Teams can compare results without losing the connection to the configuration used for execution.
If demand outgrows the cluster, the platform team still needs to add capacity or revise the allocation. Adjust queue limits using the workload requirements and observed behavior—not a promised utilization percentage.
When this setup fits
Use this approach when you operate Kubernetes GPU capacity and need teams to submit work through shared controls, with run records and reusable configuration.
It does not replace GPU procurement, cluster operations, or a scheduler-specific fair-share or preemption policy. If you need GPU partitioning, gang scheduling, or another placement capability, configure and validate the appropriate Kubernetes components separately.
For a rollout, submit representative notebook, training, and batch workloads. Check queue limits, project access, resource placement, waiting states, and output retrieval before opening the cluster to more teams.