Build and run AI workloads on your infrastructure.
Practical guides for training, clusters, distributed compute, inference, accelerators, and model-specific recipes with Polyaxon.
Training and fine-tuning
Run training frameworks on Kubernetes with explicit resources, inputs, outputs, and run history.
Fine-tune with TRL on Polyaxon
Run an SFT or PEFT workload with explicit model, dataset, accelerator, secret, and checkpoint boundaries.
Fine-tune with Axolotl on Polyaxon
Run an Axolotl configuration as a tracked GPU workload with controlled secrets, outputs, and placement.
Train agents with Ray and RAGEN
Run a RAGEN reinforcement-learning workload on a tracked KubeRay cluster with explicit head, worker, GPU, and checkpoint policy.
Run Miles post-training on Polyaxon
Coordinate a Miles post-training recipe across Ray workers with durable model conversion, rollout, and checkpoint stages.
Infrastructure and clusters
Connect GPU clusters and validate scheduling, storage, networking, and multi-node communication.
Choose infrastructure
Compare provider fit, ownership, capacity, storage, networking, and commercial constraints.
Cloud Kubernetes
Validate GPU capacity exposed through a hyperscaler-managed Kubernetes service.
Validate an Amazon EKS GPU cluster with Polyaxon
Check GPU allocation, EFA, storage, DNS, and pod identity through the same Polyaxon queue production workloads will use.
Validate a GKE GPU cluster with Polyaxon
Verify a GKE GPU pool, its driver mode, placement policy, storage, networking, and workload identity through Polyaxon.
GPU clouds
Connect provider-managed GPU Kubernetes capacity without hiding its operating boundary.
Validate Lambda Managed Kubernetes with Polyaxon
Exercise Lambda GPU, storage, networking, and namespace policy through the same Polyaxon queue production workloads will use.
Validate Crusoe Managed Kubernetes with Polyaxon
Verify a Crusoe GPU node pool, storage class, network path, and shared-responsibility boundary through Polyaxon.
Validate Nebius Managed Kubernetes with Polyaxon
Verify a Nebius GPU node group, driver mode, storage, interconnect, and workload identity through Polyaxon.
Cluster validation
Test accelerator discovery and multi-node communication before accepting a pool.
Inference and serving
Deploy model servers with explicit resources, health checks, access controls, and evaluation workloads.
Serve models with SGLang on Polyaxon
Schedule an SGLang model server as a GPU-backed Polyaxon service and retain the runtime recipe with the endpoint.
Deploy NVIDIA Dynamo with Polyaxon
Release a DynamoGraphDeployment through a governed Polyaxon operation, then evaluate its operator-managed frontend and workers.
Serve models with vLLM on Polyaxon
Run a vLLM OpenAI-compatible endpoint as a tracked GPU service with typed model and serving inputs.
Deploy NVIDIA NIM on Polyaxon
Run an NVIDIA NIM container as a governed Polyaxon service with registry credentials, GPU resources, model cache, and endpoint validation.
Serve models with TensorRT-LLM
Track engine preparation and run `trtllm-serve` as a GPU-backed Polyaxon service with explicit topology and cache policy.
Distributed compute
Run framework-native clusters while Polyaxon manages placement, logs, metadata, and workload lifecycle.
Run Ray on Polyaxon
Provision a KubeRay cluster, submit a Ray entrypoint through Polyaxon, inspect the operation, and prepare the workload for production.
Run Dask on Polyaxon
Provision a Dask scheduler and worker group through Polyaxon, connect a client operation, add autoscaling, and prepare the cluster for production.
Model recipes
Connect model-specific requirements to the runtime, accelerator topology, and validation process.
Serve DeepSeek V4 on Polyaxon
Turn a DeepSeek V4 runtime recipe into a governed Polyaxon service with explicit Blackwell topology, cache, parsers, and validation.
Serve Qwen 3.6 on Polyaxon
Deploy Qwen 3.6 with SGLang on NVIDIA or AMD Kubernetes nodes and keep context, parser, cache, and topology choices reviewable.
Accelerators
Validate drivers, runtimes, resource discovery, scheduling, and workload performance on specialized hardware.
Run Polyaxon workloads on AMD GPUs
Expose AMD GPUs through Kubernetes, target them with Polyaxon presets, and validate ROCm before training or inference.
Run Polyaxon workloads on Tenstorrent
Connect a Kubernetes Tenstorrent pool to Polyaxon and keep device acquisition, runtime images, HugePages, and validation explicit.
Run Polyaxon workloads on Google Cloud TPUs
Schedule JAX, PyTorch/XLA, or supported inference workloads onto GKE TPU slices through Polyaxon.
Compare platforms
Review how Polyaxon differs from other platforms in scheduling, tracking, infrastructure ownership, and workload lifecycle.
Explore platform comparisons