Inference articles
Browse Polyaxon articles about Inference.

Keep inference available during Kubernetes node drains
Set a PodDisruptionBudget around usable inference capacity, inspect blocked evictions, and account for replacement GPUs and model warmup.
Sep 29, 2026
Polyaxon
KubernetesInference
Serve multiple LoRA adapters with vLLM on Polyaxon
Deploy two task-specific LoRA adapters over one Qwen base model, select them by name, and measure quality and shared GPU capacity with Polyaxon.
Sep 27, 2026
Polyaxon
LLMOpsVllm
Spread inference replicas across Kubernetes zones
Use topology spread constraints for inference availability, understand minDomains and GPU capacity, and preserve placement intent in Polyaxon.
Sep 24, 2026
Polyaxon
KubernetesScheduling
Deploy Laya typed decisions on Polyaxon
Serve the open Laya typed-decisions checkpoint through Polyaxon, with a CPU classifier endpoint, explicit model scope, and a review path for uncertain classifications.
Sep 23, 2026
Polyaxon
LLMOpsInference
Serve a Jev-style Qwen classifier on Polyaxon
Deploy an open Qwen model behind a typed classification API on Polyaxon, then check the decision schema and calibration before routing real work.
Sep 23, 2026
Polyaxon
LLMOpsInference
Serve DiffusionGemma on Polyaxon
Adapt Google's Gemma-on-Kubernetes deployment choices to Polyaxon with a DiffusionGemma vLLM service, a batch evaluation job, and an explicit compatibility checklist.
Sep 23, 2026
Polyaxon
LLMOpsInference
Make vLLM prefix caching work for repeated prompts
Understand exact prefix reuse in vLLM, structure repeated context, account for replica routing, and compare cold and warm workloads with Polyaxon.
Sep 21, 2026
Polyaxon
LLMOpsInference
Optimize LLM inference with repeatable benchmarks
Compare inference optimizations against a fixed workload, latency limits, and quality checks, with benchmark configurations and results tracked in Polyaxon.
Sep 14, 2026
Polyaxon
LLMOpsInference
Control notebook and inference service lifetimes
Reuse Polyaxon's termination specification to bound notebook and inference service lifetimes, stop idle services, and distinguish inactivity from ongoing work.
Sep 13, 2026
Polyaxon
PolyaxonScheduling
Mount model artifacts with Kubernetes image volumes
Separate immutable model files from the serving image with Kubernetes image volumes, and choose when an OCI artifact fits better than a download or PVC.
Sep 10, 2026
Polyaxon
KubernetesStorage
Autoscale inference services using workload metrics
Use custom metrics with Kubernetes HPA for independent inference replicas, account for model startup, and distinguish desired replicas from usable capacity.
Aug 22, 2026
Polyaxon
KubernetesInference
Roll out a new model when every GPU is occupied
Plan GPU inference rollouts around surge capacity, temporary reduced availability, model warmup, and request draining with a worked Deployment scenario.
Aug 17, 2026
Polyaxon
KubernetesInference
Share CPU and memory across containers in a Pod
Use Pod-level resource budgets for cooperating containers, understand the remaining isolation boundaries, and evaluate the fit for ML services.
Aug 15, 2026
Polyaxon
KubernetesResources
Self-hosted vs. managed AI inference
Choose between self-hosted and managed AI inference using data policy, model access, latency, utilization, reliability, staffing, and cost per accepted outcome.
Jul 20, 2026
Polyaxon
InferenceInfrastructure
Serve DeepSeek V4 on Polyaxon
Configure a DeepSeek V4 Pro service on B200 GPUs and check model loading, reasoning output, tool calls, and resource use.
May 3, 2026
Polyaxon
LLMOpsInference
Compare serverless and dedicated LLM inference with Polyaxon
Use Polyaxon components, matrices, tracking, and run comparisons to evaluate serverless model APIs against dedicated inference services on quality, latency, and cost.
Mar 18, 2026
Polyaxon
LLMOpsInference
Serve Qwen 3.6 on Polyaxon
Deploy Qwen3.6-27B with SGLang, then check the model's API behavior, memory use, and response quality before sending it production traffic.
Mar 15, 2026
Polyaxon
LLMOpsInference
When CPU-accelerated AI inference makes sense
Evaluate CPU inference using model fit, precision, latency, throughput, memory, Kubernetes placement, and end-to-end cost instead of assuming every model needs a GPU.
Sep 9, 2025
Polyaxon
InferenceKubernetes