Polyaxon v3 is coming →

Run model endpoints on your Kubernetes clusters

Package a model server as a Polyaxon service, choose its resources and model inputs, and inspect its status and logs alongside the rest of your workloads.

Protected organization service access is documented for Polyaxon Cloud and self-hosted Enterprise. Your team operates the model server and its compute capacity.

From application request to model server

Your application

Calls the configured service URL with the required credentials.

Polyaxon service access

Routes requests through the service proxy with access controls enabled.

Model-server container

Runs on your connected Kubernetes compute.

Model
The artifact or revision you configure
Runtime
The server image and startup arguments
Resources
CPU, memory, and GPU requests
A protected-service deployment pattern. Your server handles inference; Kubernetes provides capacity and placement. Public exposure and additional ingress controls are deployment choices.

Keep the model server that fits your workload

Follow a trained model into an API

Specify the model and how the service obtains it

Expose an endpoint with deliberate access settings

Operate the service as a workload