Polyaxon v3 is coming →

Polyaxon Features

Build, monitor, deliver and scale machine learning experiments — we're the easiest way to go from an idea to a fully deployable model, bypassing all those infrastructure headaches.

Polyaxon Features

Polyaxon

One AI engineering control plane for training, agents, evaluation, observability, and infrastructure.

Track experiments, run workloads, manage models, compare results, route model calls, and operate production AI systems on your own Kubernetes infrastructure.

AI engineering control plane

Run training, inference, evaluations, notebooks, services, gateway traffic, and agent workflows across clusters from one place.

Experiment and run tracking

Capture parameters, metrics, logs, code versions, artifacts, visualizations, lineage, and runtime state for every run.

Workload orchestration

Compose jobs, services, DAGs, hooks, schedules, sweeps, distributed workloads, and reusable components.

Scheduling and resource policy

Route workloads through queues, priorities, concurrency rules, approvals, presets, namespaces, and agent tags.

Agents and sandboxes

Launch isolated execution environments with controlled filesystems, network access, tools, SSH, and runtime policies.

Prompt management

Version prompts, manage labels and releases, test in playgrounds, connect prompts to traces, and evaluate changes.

Registry and lineage

Version and promote models, artifacts, datasets, components, and prompts with lineage back to the work that produced them.

Evaluation

Run offline and online evaluations, compare outputs, score models and prompts, and promote approved versions.

Observability

Monitor runs, services, agents, and infrastructure with logs, traces, metrics, health checks, alerts, and dashboards.

AI gateway

Route model calls across providers, manage credentials, enforce policies, cache responses, and trace every request.

Governance and access control

Control organizations, teams, roles, service accounts, approvals, audit trails, and access to projects and assets.

APIs, CLI, and SDKs

Automate every workflow from the CLI, REST API, Python client, TypeScript client, CI/CD systems, and internal tools.

Gain more productivity and ship faster

Polyaxon provides an interactive workspace with notebooks, tensorboards, visualizations, and dashboards.

User management

Collaborate with the rest of your team, share and compare experiments and results.

User resources allocation

Manage your team's resources and parallelism, and set quotas.

Versioning and reproducibility

Reproducible results with a built-in version control for code and experiments.

Data autonomy

Maintain complete control of your data and persistence choices.

Hyperparameter search & optimization

Run group of experiments in parallel and in distributed way.

Maximizes Resource Utilization

Spin up or down, add more nodes, add more GPUs, and expand storage.

Cost Effective

Leverage commodity on-premise infrastructure or spot instances to reduce costs.

Powerful interface

Author jobs, experiments, and pipelines in Json, YAML, and Python.

Runs on any infrastructure

Deploy Polyaxon in the cloud, on-premises or in hybrid environments, including single laptop, container management platforms, or on Kubernetes.

Scalable

Polyaxon can easily scale horizontally by adding more nodes.

Modular

Extend Polyaxon's functionalities with custom plugins.

And so many more ...

Jobs, services, and sessions

Run batch jobs, long-running services, notebooks, terminals, dashboards, SSH sessions, and custom APIs.

Versatile Runtime and workloads

Schedule GPU jobs, distributed training, Ray, Dask, Spark, Kubeflow operators, and custom Kubernetes workloads.

Sweeps and optimization

Run grid search, random search, Bayesian-style optimization, mapping, early stopping, and parallel evaluations.

DAGs and reusable workflows

Build multi-step pipelines with dependencies, cached steps, typed inputs and outputs, retries, and failure handling.

Events, hooks, and schedules

Trigger downstream jobs, notifications, webhooks, evaluations, and alerts from platform state changes or schedules.

Queues and concurrency

Control fairness and utilization with queues, priorities, rate limits, concurrency rules, and per-workflow budgets.

Dashboards and comparison

Compare runs, inspect charts and tables, rank candidates, review diffs, and build project-level dashboards.

Logs, traces, and metrics

Stream logs, inspect traces, query metrics, follow live runtime state, and debug production behavior.

Prompt and model releases

Promote prompt versions, model versions, and artifacts across staging and production with labels and approvals.

Datasets and artifacts

Track datasets, files, model outputs, reports, and generated assets with metadata and lineage.

Caching and reproducibility

Avoid repeating expensive work while preserving code, inputs, outputs, environment, and lineage snapshots.

Connections and secrets

Manage registry credentials, data stores, environment variables, service accounts, and protected access to resources.

Hybrid and multi-cluster execution

Use agents, our open-source component, to run workloads across namespaces, clusters, regions, clouds, and private infrastructure.

Data autonomy

Keep code, data, models, logs, metrics, and connections in your own environment while using a managed control plane.

Audit and compliance

Track activity, enforce roles, review access, retain history, and support organization-level governance requirements.

Manual and automated controls

Start, stop, resume, restart, clone, approve, promote, and automate work through UI, CLI, API, and hooks.

Custom runtime environments

Define containers, presets, resources, mounts, node selectors, tolerations, sidecars, and init containers per workload.

Framework and provider choice

Use your preferred ML frameworks, LLM providers, storage systems, container registries, and Kubernetes operators.

Open extension points

Integrate with internal systems through webhooks, APIs, custom components, actions, and automation hooks.

Team management

Manage organizations, projects, teams, members, service accounts, SSO, quotas, and role-based permissions.

Cost and capacity controls

Measure usage, right-size workloads, route expensive jobs, use spot capacity, and control shared GPU resources.

Self-hosted and cloud options

Run open source, self-hosted, hybrid, or fully managed deployments without changing the workflow model.

Production AI operations

Operate model services, AI applications, agent sessions, evaluations, prompts, traces, and alerts in one system.

Logs

Stream, filter, and search logs from all operations.

SLAs

Enforce SLAs with TTL, timeout, and retries.

Custom Run Environments

Configure reusable presets or define per run environments.

Events

Send and subscribe to events mid-run or at the end of operations.

Rendering and Visualization

Log artifacts and custom visualizations.

Manual Triggers

Start, stop, resume, restart, and copy any job or service.

Projects

Organize your efforts and define directories of work.

Component Hub

Extract reusable modules and define typed inputs and outputs for your runtime.

Model Registry

Lock experiments and promote your work for production deployment.

Powerful Search & query language

Search by name, description, status, tags, metrics, parameters, artifact metadata, model versions, regex, custom fields, metrics, or configurations.

Cache Layer

Avoid running expensive computations several times.

Affinity and Routing

Runs are assigned to queues and they are only picked up by agents with matching tags, you can manage multiple namespaces and clusters.

Sharding

Parallelize your work with operation mapping and hyperparameter optimization.

Joining

Filter, aggregate, and annotate inputs/outputs/artifacts from multiple or parallel upstream runs.

Consensus

Define the success of your workflows based on early stopping strategies.

Concurrency

Enforce global concurrency, and split the quota on per queue or workflow level.

Priority

Define executions priority and enforce rate limits.

Versioning

Every run is versioned and has a unique hash, all outputs and artifacts are tracked.

Auto API, authZ, and authN

Every run will be checked prior to execution and will have a permissioned API.

Team Management

Invite users to your team and assign roles and permissions as appropriate.

Customization

Integrate with external systems with hooks and actions.