Polyaxon v3 is coming →

Serve models on infrastructure you control

Give an internal application a model endpoint backed by your Kubernetes compute, with an explicit model version, access path, and operating plan.

This protected-service workflow uses Polyaxon Cloud or self-hosted Enterprise. Control-plane location, endpoint exposure, and outbound dependencies must match your deployment requirements.

An internal assistant with a dedicated model endpoint

An application team selects a language model and serving image. The platform team provides GPU capacity, model storage, and a protected service route. The application calls that endpoint using credentials permitted by the deployment.

Make the model, compute, and access path explicit

  1. Select the model

    Record its revision and the serving image.

  2. Start the service

    Load the model on compatible Kubernetes resources.

  3. Configure access

    Set authentication and the allowed network path.

  4. Inspect requests

    Check responses, errors, logs, and resource use.

Illustrative deployment, not a benchmark or an air-gap guarantee. Running inference on your cluster does not determine where every dependency or control-plane service runs.

Decide what must stay inside your environment

Turn the selected model into a service definition

Verify which applications can call it

Plan model changes before production traffic