Run ML and agent workflows on your infrastructure
Give teams a defined path from submitting work to reviewing results, with compute, configuration, and run history managed together.
Choose a workflow
Shared GPU clusters
Give each team its own queues and resource limits, with reusable presets for workloads on a shared Kubernetes cluster.
Train and fine-tune models
Package your training recipe, choose its compute, and compare the resulting metrics and checkpoints without separating them from the configuration that produced them.
Automate training and evaluation
Connect data preparation, training, and evaluation in a repeatable workflow. Define which results can move forward and where a person must approve the next operation.
Serve models on private infrastructure
Give an internal application a model endpoint backed by your Kubernetes compute, with an explicit model version, access path, and operating plan.
Batch inference and data processing
Process a dataset with a recorded model and configuration, save the results, and inspect the run without maintaining a permanent inference endpoint.
Run agent code on your infrastructure
Connect application-defined tools to a Polyaxon sandbox so an agent can run commands, work with files, and return results from your Kubernetes infrastructure.
Keep your framework and own the setup
Your training libraries, model server, or agent application define the task. Kubernetes supplies the execution environment. Polyaxon adds workload configuration, dispatch controls, tracking, and workflow orchestration around them.
Use the linked integrations for framework-specific configuration, and the product pages to understand the controls available at each stage. Each workflow also identifies the infrastructure and operational decisions your team needs to make.