Polyaxon Features
Build, monitor, deliver and scale machine learning experiments — we're the easiest way to go from an idea to a fully deployable model, bypassing all those infrastructure headaches.
Polyaxon
One AI engineering control plane for training, agents, evaluation, observability, and infrastructure.
Track experiments, run workloads, manage models, compare results, route model calls, and operate production AI systems on your own Kubernetes infrastructure.
AI engineering control plane
Run training, inference, evaluations, notebooks, services, gateway traffic, and agent workflows across clusters from one place.
Experiment and run tracking
Capture parameters, metrics, logs, code versions, artifacts, visualizations, lineage, and runtime state for every run.
Workload orchestration
Compose jobs, services, DAGs, hooks, schedules, sweeps, distributed workloads, and reusable components.
Scheduling and resource policy
Route workloads through queues, priorities, concurrency rules, approvals, presets, namespaces, and agent tags.
Agents and sandboxes
Launch isolated execution environments with controlled filesystems, network access, tools, SSH, and runtime policies.
Prompt management
Version prompts, manage labels and releases, test in playgrounds, connect prompts to traces, and evaluate changes.
Registry and lineage
Version and promote models, artifacts, datasets, components, and prompts with lineage back to the work that produced them.
Evaluation
Run offline and online evaluations, compare outputs, score models and prompts, and promote approved versions.
Observability
Monitor runs, services, agents, and infrastructure with logs, traces, metrics, health checks, alerts, and dashboards.
AI gateway
Route model calls across providers, manage credentials, enforce policies, cache responses, and trace every request.
Governance and access control
Control organizations, teams, roles, service accounts, approvals, audit trails, and access to projects and assets.
APIs, CLI, and SDKs
Automate every workflow from the CLI, REST API, Python client, TypeScript client, CI/CD systems, and internal tools.
Gain more productivity and ship faster
Polyaxon provides an interactive workspace with notebooks, tensorboards, visualizations, and dashboards.
User management
Collaborate with the rest of your team, share and compare experiments and results.
User resources allocation
Manage your team's resources and parallelism, and set quotas.
Versioning and reproducibility
Reproducible results with a built-in version control for code and experiments.
Data autonomy
Maintain complete control of your data and persistence choices.
Hyperparameter search & optimization
Run group of experiments in parallel and in distributed way.
Maximizes Resource Utilization
Spin up or down, add more nodes, add more GPUs, and expand storage.
Cost Effective
Leverage commodity on-premise infrastructure or spot instances to reduce costs.
Powerful interface
Author jobs, experiments, and pipelines in Json, YAML, and Python.
Runs on any infrastructure
Deploy Polyaxon in the cloud, on-premises or in hybrid environments, including single laptop, container management platforms, or on Kubernetes.
Scalable
Polyaxon can easily scale horizontally by adding more nodes.
Modular
Extend Polyaxon's functionalities with custom plugins.
And so many more ...
Jobs, services, and sessions
Run batch jobs, long-running services, notebooks, terminals, dashboards, SSH sessions, and custom APIs.
Versatile Runtime and workloads
Schedule GPU jobs, distributed training, Ray, Dask, Spark, Kubeflow operators, and custom Kubernetes workloads.
Sweeps and optimization
Run grid search, random search, Bayesian-style optimization, mapping, early stopping, and parallel evaluations.
DAGs and reusable workflows
Build multi-step pipelines with dependencies, cached steps, typed inputs and outputs, retries, and failure handling.
Events, hooks, and schedules
Trigger downstream jobs, notifications, webhooks, evaluations, and alerts from platform state changes or schedules.
Queues and concurrency
Control fairness and utilization with queues, priorities, rate limits, concurrency rules, and per-workflow budgets.
Dashboards and comparison
Compare runs, inspect charts and tables, rank candidates, review diffs, and build project-level dashboards.
Logs, traces, and metrics
Stream logs, inspect traces, query metrics, follow live runtime state, and debug production behavior.
Prompt and model releases
Promote prompt versions, model versions, and artifacts across staging and production with labels and approvals.
Datasets and artifacts
Track datasets, files, model outputs, reports, and generated assets with metadata and lineage.
Caching and reproducibility
Avoid repeating expensive work while preserving code, inputs, outputs, environment, and lineage snapshots.
Connections and secrets
Manage registry credentials, data stores, environment variables, service accounts, and protected access to resources.
Hybrid and multi-cluster execution
Use agents, our open-source component, to run workloads across namespaces, clusters, regions, clouds, and private infrastructure.
Data autonomy
Keep code, data, models, logs, metrics, and connections in your own environment while using a managed control plane.
Audit and compliance
Track activity, enforce roles, review access, retain history, and support organization-level governance requirements.
Manual and automated controls
Start, stop, resume, restart, clone, approve, promote, and automate work through UI, CLI, API, and hooks.
Custom runtime environments
Define containers, presets, resources, mounts, node selectors, tolerations, sidecars, and init containers per workload.
Framework and provider choice
Use your preferred ML frameworks, LLM providers, storage systems, container registries, and Kubernetes operators.
Open extension points
Integrate with internal systems through webhooks, APIs, custom components, actions, and automation hooks.
Team management
Manage organizations, projects, teams, members, service accounts, SSO, quotas, and role-based permissions.
Cost and capacity controls
Measure usage, right-size workloads, route expensive jobs, use spot capacity, and control shared GPU resources.
Self-hosted and cloud options
Run open source, self-hosted, hybrid, or fully managed deployments without changing the workflow model.
Production AI operations
Operate model services, AI applications, agent sessions, evaluations, prompts, traces, and alerts in one system.
Logs
Stream, filter, and search logs from all operations.
SLAs
Enforce SLAs with TTL, timeout, and retries.
Custom Run Environments
Configure reusable presets or define per run environments.
Events
Send and subscribe to events mid-run or at the end of operations.
Rendering and Visualization
Log artifacts and custom visualizations.
Manual Triggers
Start, stop, resume, restart, and copy any job or service.
Projects
Organize your efforts and define directories of work.
Component Hub
Extract reusable modules and define typed inputs and outputs for your runtime.
Model Registry
Lock experiments and promote your work for production deployment.
Powerful Search & query language
Search by name, description, status, tags, metrics, parameters, artifact metadata, model versions, regex, custom fields, metrics, or configurations.
Cache Layer
Avoid running expensive computations several times.
Affinity and Routing
Runs are assigned to queues and they are only picked up by agents with matching tags, you can manage multiple namespaces and clusters.
Sharding
Parallelize your work with operation mapping and hyperparameter optimization.
Joining
Filter, aggregate, and annotate inputs/outputs/artifacts from multiple or parallel upstream runs.
Consensus
Define the success of your workflows based on early stopping strategies.
Concurrency
Enforce global concurrency, and split the quota on per queue or workflow level.
Priority
Define executions priority and enforce rate limits.
Versioning
Every run is versioned and has a unique hash, all outputs and artifacts are tracked.
Auto API, authZ, and authN
Every run will be checked prior to execution and will have a permissioned API.
Team Management
Invite users to your team and assign roles and permissions as appropriate.
Customization
Integrate with external systems with hooks and actions.