Decide what must stay inside your environment
Start with the request path, model files, logs, and control plane—not just the location of the GPU. Polyaxon Cloud manages the control plane while workloads execute on connected compute. Self-hosted Enterprise also puts the control plane under your team's operation.
Review model downloads, image pulls, telemetry, and any external APIs used by the application. Place or mirror dependencies where your network policy requires them, and decide which request content may appear in logs. Private compute alone does not make the full application air-gapped.
Connected compute architecture · Self-hosted networking · Cloud data regions
Turn the selected model into a service definition
Choose a server compatible with the model and configure its image, model revision, resource requests, startup arguments, and storage connections. The vLLM integration provides the Polyaxon service configuration and client request pattern for an OpenAI-compatible endpoint.
Use scoped credentials for model access and a suitable cache when downloads are expensive. Confirm that the requested GPUs and memory are available at the selected destination. Wait for the model server to become ready rather than equating a started container with a usable endpoint.
Verify which applications can call it
Use the deployment's service URL and the required authentication rather than exposing the model-server port directly. Keep Polyaxon's access controls enabled unless another deliberately configured mechanism replaces them; isExternal disables the built-in checks.
Before handing the endpoint to the application, verify both an authorized request and a request that should be denied. Check the actual network route, TLS, token permissions, and API path. Ingress and application policies should agree about who can reach and use the service.
Service access configuration · Authenticated request example
Plan model changes before production traffic
Inspect response quality, error behavior, memory use, and latency under representative load. Retain the model revision, image, and configuration needed to start the previous service if a replacement fails. Define how clients move to a new endpoint and how existing requests finish.
A registry stage can record that a candidate is selected; it does not change a running service. Your team still owns the deployment and rollback procedure, gateway behavior, and compute capacity.
This fits teams that want to operate a model server on their own Kubernetes resources. If the application only needs periodic predictions or embeddings over stored data, use a batch workflow instead of keeping an endpoint running.
Model serving responsibilities · Model versions · Batch inference