Polyaxon v3 is coming →

Choosing GPU infrastructure for Polyaxon

Polyaxon does not replace the GPU cloud beneath your workloads. It connects to Kubernetes clusters and gives teams one layer for workload definitions, queues, scheduling policy, metadata, and lifecycle. The useful question is therefore not which provider has the longest GPU list, but which infrastructure boundary your team can operate reliably.

A provider shortlist, a defensible cluster architecture, and a representative workload you can use to validate capacity and operations before making a commitment.

Three layers, three different jobs

GPU clouds and Polyaxon are not interchangeable products. The cleanest architecture gives each layer one clear owner.

The provider supplies capacity

The infrastructure vendor owns physical GPUs and some combination of control planes, worker nodes, networking, storage, drivers, and hardware remediation.

Kubernetes exposes the substrate

Node pools, resource labels, storage classes, network policy, quotas, autoscaling, and device plugins turn provider capacity into schedulable resources.

Polyaxon owns the workload lifecycle

Agents connect clusters to shared queues, presets, connections, experiments, pipelines, services, registries, approvals, and operational metadata.

Kubernetes compatibility

Connect the Polyaxon Agent

CoreWeave CKS, Nebius Managed Kubernetes, Crusoe Managed Kubernetes, and Lambda Managed Kubernetes expose the cluster boundary Polyaxon expects. Validate supported versions, permissions, networking, storage, and provider-managed add-ons before installation.

Keep external capacity outside the scheduler

A Lambda On-Demand VM or Vast.ai container instance is not automatically a Polyaxon compute cluster. Use it through a deliberately operated Kubernetes layer or as external capacity with immutable inputs, versioned outputs, and one authoritative retry boundary.

GPU infrastructure comparison

This is a scope and ownership comparison, not a performance ranking. GPU availability, products, regions, and commercial terms change; validate them directly for the workload and date you intend to run.

Lambda

Direct Kubernetes target

GPU virtual machines and larger GPU clusters, with Managed Kubernetes, preinstalled Kubernetes, and Slurm paths depending on the offering.

Deploy a Polyaxon Agent into Lambda Managed Kubernetes or another supported Kubernetes cluster. A standalone On-Demand VM is not itself a Polyaxon compute cluster.

On-demand GPU instances plus multi-node cluster offerings intended for training and other tightly coupled workloads.

A straightforward path from a single GPU VM to preconfigured GPU clusters, with managed Kubernetes providing GPU, InfiniBand, and shared-storage support.

Confirm which product includes Kubernetes, the cluster size and term, regional capacity, network topology, storage behavior, and who owns upgrades or remediation.

CoreWeave

Direct Kubernetes target

CoreWeave Kubernetes Service runs managed Kubernetes control and data planes on dedicated bare-metal GPU nodes.

Install a Polyaxon Agent in a CKS cluster and map node pools or capacity classes to Polyaxon queues and presets.

Reserved, flex reservation, spot, and on-demand capacity models exposed through Kubernetes node pools.

Kubernetes is the primary infrastructure interface, with dedicated hardware, VPC isolation, provider-managed GPU components, HPC networking, and multiple storage choices.

Review capacity guarantees, spot interruption, supported cluster components, the provider-managed GPU Operator, image caching behavior, regions, and contract terms.

Nebius

Direct Kubernetes target

An AI cloud spanning GPU VMs and clusters, Managed Kubernetes, Soperator-based Slurm, storage, registries, and higher-level AI services.

Use Managed Service for Kubernetes as the direct Polyaxon target; keep serverless jobs, managed MLflow, or Soperator as separate service decisions.

GPU VMs and InfiniBand-connected GPU clusters alongside managed Kubernetes and Slurm-oriented infrastructure.

A broader cloud foundation around Kubernetes, including object storage, shared filesystems, container registries, IAM, audit logs, and AI-specific services.

Check regional service availability and maturity, GPU quota, storage placement, network topology, IAM integration, and whether overlapping managed services will duplicate Polyaxon capabilities.

Crusoe

Direct Kubernetes target

Infrastructure Cloud provides GPU VMs, Managed Kubernetes, Managed Slurm, storage, networking, and provider telemetry, alongside separate Managed AI services.

Deploy a Polyaxon Agent into Crusoe Managed Kubernetes and keep Polyaxon authoritative for workloads, application policy, and lifecycle metadata.

GPU infrastructure delivered as VMs or multi-node Kubernetes and Slurm clusters, with managed AI services as a separate abstraction.

A documented Kubernetes responsibility boundary, managed control plane, GPU and network operators, InfiniBand, health telemetry, and automated hardware remediation options.

Your team still owns workloads, Kubernetes RBAC, application secrets, quotas, scheduling rules, third-party operators, and much of application observability.

Vast.ai

Conditional target

A marketplace for dedicated GPU-backed Docker instances from datacenter and community providers, plus separate serverless endpoints.

Treat marketplace instances as external capacity unless your team deliberately builds and operates a supported Kubernetes cluster on top. Do not assume an instance can register as a Polyaxon Agent by itself.

On-demand, reserved, and interruptible marketplace rentals selected by GPU, host, location, price, and other filters.

Broad heterogeneous supply, custom container templates, per-instance SSH or Jupyter access, and flexible short-lived capacity discovery.

Provider heterogeneity increases validation work for trust, networking, storage durability, bandwidth, availability, interruption, multi-node topology, and repeatability.

Reference architectures

One managed Kubernetes cluster

Start with one provider-operated Kubernetes cluster, one Polyaxon Agent, explicit GPU node pools, durable artifact storage, and separate queues for reliable and interruptible capacity.

Multiple provider clusters

Connect one Agent per cluster and route workloads through queues and presets. Use this only when isolation, geography, compliance, or materially different capacity justifies the operational overhead.

External marketplace capacity

Keep an external instance provider outside the Polyaxon scheduling boundary, pass immutable inputs, and return versioned outputs with the external job identifier recorded on the run.

Evaluation plan

  • Define the workload envelope

    Record GPU memory, GPU count, single-node or multi-node topology, runtime, checkpoint size, data locality, network needs, expected duration, and interruption tolerance.

  • Separate hard gates from preferences

    Treat region, compliance, private networking, supported Kubernetes versions, accelerator topology, storage semantics, and capacity guarantees as gates before comparing convenience or price.

  • Map the ownership boundary

    Name who upgrades Kubernetes, manages GPU and network operators, replaces failed hardware, configures autoscaling, owns secrets, and responds when a job is pending or slow.

  • Run the same representative workload

    Validate image pulls, dataset staging, multi-node startup when needed, logs, metrics, checkpoint recovery, cancellation, artifact persistence, and cleanup on each finalist.

  • Model steady state and failure state

    Compare committed and burst capacity, idle resources, storage and egress, interrupted work, support response, quota lead time, and the engineering cost of operating the cluster.

Procurement checklist

  • Written confirmation of GPU availability, region, quota, provisioning lead time, and any capacity guarantee.
  • Supported Kubernetes versions, upgrade policy, control-plane responsibility, and node-pool scaling behavior.
  • GPU topology, RDMA or InfiniBand requirements, host health checks, and failed-hardware replacement process.
  • Object, block, shared, and local storage semantics for datasets, checkpoints, caches, and artifacts.
  • Private networking, ingress and egress controls, identity integration, auditability, and support-access policy.
  • Interruption notice and recovery expectations for spot, interruptible, marketplace, or otherwise reclaimable capacity.
  • A complete cost model covering idle reservations, active compute, storage, networking, support, and operational labor.

Sources

Recheck current regions, capacity, product maturity, supported versions, and commercial terms before procurement.

Design your cluster

We can map GPU topology, queues, storage, network policy, interruption, and multi-cluster requirements before you commit to capacity.