The AI control plane that runs on your infrastructure
Run everything from GPU training to agent sessions on your own clusters — scheduled, tracked, and governed in one place. Self-hosted deployments keep your data inside the infrastructure you control.
Frameworks, specs, clients, and APIs feed the platform.
Organization, access, and policy controls admit every run.
Orchestrate runs, observe systems, compare results, and register assets.
Agents and queues translate product intent into cluster work.
Polyaxon runs across your clouds, clusters, accelerators, and storage.
AI-first infrastructure.
Training, services, and sandboxes share the same specs, queues, governance, and data connections — on your own Kubernetes and storage backends.

The full training loop.
Train new models, post-train open-source ones, and fine-tune with LoRA — from a single GPU to multi-node runs. Sweeps, DAGs, batch jobs, and agent executions ride the same queues and quotas.
Train models from scratch, or post-train open-source ones with LoRA and SFT — single GPU or multi-node.
PyTorch DDP, MPI, and Horovod gang-scheduled across your GPU nodes from a single spec.
Scale Ray, Dask, and Spark workloads for preprocessing, feature pipelines, and batch compute.
Stateful rollouts, reward loops, and environments for long-horizon training.
Grid, random, Bayesian, and Hyperband strategies with early stopping.
Multi-step pipelines, scheduled batch work, retries, and dependencies.
Schedule agent workloads like any run — queued, traced, in parallel, and governed.
Run Chromium instances — headed or headless — and process images or audio at scale.
One control plane, every feature included.
Run AI workloads across your infrastructure with tracking, observability, dashboards, logs, lineage, evaluation, automation, and governance built in.
Works with your stack.
Track experiments from the ML libraries your team already uses, run distributed jobs on Kubernetes, route AI gateway traffic, and connect artifacts, registries, and infrastructure across cloud, on-prem, and hybrid clusters.
Why use Polyaxon?
An opinionated AI engineering platform for teams that want full control over their infrastructure, data, execution, and governance.
Polyaxon covers ML engineering, AI training, agentic development, production monitoring, and infrastructure management in one platform.
The full AI engineering cycle
Polyaxon covers ML engineering, AI training, agentic development, production monitoring, and infrastructure management in one platform.
Built on an open-source core
Polyaxon is built on an Apache 2.0 core you can inspect, self-host, and extend through the free Community Edition.
Framework and runtime agnostic
Use PyTorch, TensorFlow, JAX, XGBoost, Scikit-learn, LLM frameworks, shell commands, Python, TypeScript, SSH, and tmux from the same control plane.
Reproducible by default
Every run can capture code, hyperparameters, dependencies, artifacts, prompts, traces, and lineage. Re-run or inspect work months later.
Scales without rewrites
Go from a single project to multi-cluster GPU fleets, agent pools, namespaces, and backend storage without changing the workflow model.
First-class CLI, SDKs, and APIs
Use CLI, REST, gRPC, Python, and TypeScript APIs to start sessions, inject code, run commands, manipulate files, use git, and call HTTP services.
Enterprise Control and Governance.
Polyaxon runs on your infrastructure - on-prem, in your VPC, or on any managed Kubernetes. Everything is designed around isolation, auditability, activity tracking, and team-level controls.
Architecture
Governance
Operations
- Logs, metrics, traces, and health checks
- Approvals and activity streams
- Queues, agents, connections, and namespaces
Git, HTTP, filesystem, and artifact workflows
Offering
Polyaxon scales with your operations.
Polyaxon(Cloud)
Fully Managed Control Plane with 1-click deployment, orchestration, logging, backups, upgrades and security. All workload, code, data, artifacts, and models stay on your clusters.
Start freePolyaxon(EE)
On premise Enterprise Edition Control Plane with all features, upgrades, security, monitoring, compliance, scale, and premium support.
Explore EnterpriseOpen source
Running on top of Polyaxon's open-source tools, Polyaxon Community Edition is our free version with all core features to get you started.
Deploy Community
Questions & Answers
Polyaxon is an open-source AI engineering control plane for ML workloads, AI training, agent sandboxes, coding sessions, observability, and model management on Kubernetes. Try it with the community edition or Polyaxon Cloud.
Get started with Polyaxon
Ready to give it a try?
Start with Polyaxon Cloud or deploy the Community Edition for free.
Learn more!
Get in touch to learn more about our Enterprise offering.

















































