Keep the model server that fits your workload
Run a serving container as a Polyaxon service, with its command, arguments, ports, and resource requests defined together. Use the vLLM integration for a compatible language-model API, or the FastAPI integration for an application-specific prediction service.
The model server owns inference behavior, including its supported models, request format, batching, and memory use. Polyaxon manages the service operation around it. Choose a compatible image and model revision and size resources for the actual workload.
Follow a trained model into an API
The FastAPI documentation example accepts the UUID of a training run, loads its model.joblib artifact into the service, and starts a container exposing port 8000. The application defines a prediction endpoint at /api/v1/predict.
A caller supplies four Iris measurements: sepal length and width, and petal length and width. The server loads the classifier, computes a prediction, and returns its class name and numeric value. Polyaxon provides the service operation and access route; FastAPI and the model implement the prediction behavior.
The same source model can also be loaded by a finite batch-scoring job. Choose the service when clients need to make requests, or the batch job when the work is a known dataset to process.
Follow the model-to-API example · Compare the batch workflow
Specify the model and how the service obtains it
Pass the model choice as an input and configure initialization or storage connections to make its files available to the server. A service can load an artifact from an earlier run, or use a serving engine's model-loading mechanism with the required credentials and cache.
Record the exact model revision, server image, and configuration used for each deployment. A registry version gives a selected output a name and source-run reference; your service definition still decides which artifact to load.
Load a tracked model into an API · Model and artifact registry · Configure connections
Expose an endpoint with deliberate access settings
Declare the service port and configure path rewriting if the application cannot handle Polyaxon's service URL prefix. With protected service access on Cloud or Enterprise, callers use their authenticated session or a valid token with the appropriate permissions.
Treat external exposure as a separate decision. The service runtime's isExternal option bypasses Polyaxon's built-in authentication and authorization; it is not required just to retrieve a service URL. Configure gateway exposure, TLS, and network policy for the clients that should reach the endpoint.
Service ports and access settings · Authenticated API access · Deployment networking
Operate the service as a workload
Inspect operation status and logs when a server cannot load its model, obtain storage access, or fit within its resource limits. Configure readiness, timeouts, and shutdown behavior, then measure the service under representative requests before relying on it for production traffic.
Your team owns capacity planning, server upgrades, and the procedure for replacing or rolling back a model. Starting a Polyaxon service does not automatically provide scale-to-zero, a traffic rollout strategy, or an inference latency guarantee.
For periodic predictions over a dataset, a batch job may be the better fit: it produces recorded outputs without maintaining a request-serving endpoint.