Polyaxon v3 is coming →

Learn how to run AI workloads in production.

Follow practical paths through GPU infrastructure, MLOps, sandboxes, evaluation, observability, and workload automation.

Schedule and operate shared GPU workloads.

Queues, placement, recovery, and utilization for shared GPU clusters.

Explore GPU orchestration

Make experiments and models reproducible.

Tracking, metadata, registries, and repeatable model development.

Explore MLOps & tracking

Develop in the same environment as your workloads.

Interactive environments, access, resource control, and automation.

Explore Sandboxes

Evaluate AI changes before release.

Metrics, judges, regression suites, and production feedback.

Explore LLM evaluation

Trace and evaluate agent behavior.

Model calls, tools, handoffs, outcomes, and evaluation across the agent lifecycle.

Explore Agents

Operate AI workloads on Kubernetes.

Architecture, storage, metrics, resource inspection, and failure diagnosis.

Explore Kubernetes for AI

Connect runtime signals to AI outcomes.

Traces, metrics, evaluations, and deployment context for diagnosing failures.

Explore AI observability

Operate LLM applications in production.

Prompt versions, model access, cost, evaluation, and release decisions.

Explore LLMOps

Test AI systems before failures reach users.

Authorized adversarial cases, evidence capture, and regression testing.

Explore AI red teaming

Automate repeatable AI workflows.

Dependencies, retries, schedules, hooks, and repeatable execution.

Explore Pipelines & automation

Link model versions to their evidence.

Models, runs, artifacts, lineage, and approval context in one path.

Explore Model registry & lineage