Polyaxon v3 is coming →

Inference articles

Browse Polyaxon articles about Inference.

Keep inference available during Kubernetes node drains

Keep inference available during Kubernetes node drains

Set a PodDisruptionBudget around usable inference capacity, inspect blocked evictions, and account for replacement GPUs and model warmup.

Sep 29, 2026

Polyaxon

KubernetesInference
Serve multiple LoRA adapters with vLLM on Polyaxon

Serve multiple LoRA adapters with vLLM on Polyaxon

Deploy two task-specific LoRA adapters over one Qwen base model, select them by name, and measure quality and shared GPU capacity with Polyaxon.

Sep 27, 2026

Polyaxon

LLMOpsVllm
Spread inference replicas across Kubernetes zones

Spread inference replicas across Kubernetes zones

Use topology spread constraints for inference availability, understand minDomains and GPU capacity, and preserve placement intent in Polyaxon.

Sep 24, 2026

Polyaxon

KubernetesScheduling
Deploy Laya typed decisions on Polyaxon

Deploy Laya typed decisions on Polyaxon

Serve the open Laya typed-decisions checkpoint through Polyaxon, with a CPU classifier endpoint, explicit model scope, and a review path for uncertain classifications.

Sep 23, 2026

Polyaxon

LLMOpsInference
Serve a Jev-style Qwen classifier on Polyaxon

Serve a Jev-style Qwen classifier on Polyaxon

Deploy an open Qwen model behind a typed classification API on Polyaxon, then check the decision schema and calibration before routing real work.

Sep 23, 2026

Polyaxon

LLMOpsInference
Serve DiffusionGemma on Polyaxon

Serve DiffusionGemma on Polyaxon

Adapt Google's Gemma-on-Kubernetes deployment choices to Polyaxon with a DiffusionGemma vLLM service, a batch evaluation job, and an explicit compatibility checklist.

Sep 23, 2026

Polyaxon

LLMOpsInference
Make vLLM prefix caching work for repeated prompts

Make vLLM prefix caching work for repeated prompts

Understand exact prefix reuse in vLLM, structure repeated context, account for replica routing, and compare cold and warm workloads with Polyaxon.

Sep 21, 2026

Polyaxon

LLMOpsInference
Optimize LLM inference with repeatable benchmarks

Optimize LLM inference with repeatable benchmarks

Compare inference optimizations against a fixed workload, latency limits, and quality checks, with benchmark configurations and results tracked in Polyaxon.

Sep 14, 2026

Polyaxon

LLMOpsInference
Control notebook and inference service lifetimes

Control notebook and inference service lifetimes

Reuse Polyaxon's termination specification to bound notebook and inference service lifetimes, stop idle services, and distinguish inactivity from ongoing work.

Sep 13, 2026

Polyaxon

PolyaxonScheduling
Mount model artifacts with Kubernetes image volumes

Mount model artifacts with Kubernetes image volumes

Separate immutable model files from the serving image with Kubernetes image volumes, and choose when an OCI artifact fits better than a download or PVC.

Sep 10, 2026

Polyaxon

KubernetesStorage
Autoscale inference services using workload metrics

Autoscale inference services using workload metrics

Use custom metrics with Kubernetes HPA for independent inference replicas, account for model startup, and distinguish desired replicas from usable capacity.

Aug 22, 2026

Polyaxon

KubernetesInference
Roll out a new model when every GPU is occupied

Roll out a new model when every GPU is occupied

Plan GPU inference rollouts around surge capacity, temporary reduced availability, model warmup, and request draining with a worked Deployment scenario.

Aug 17, 2026

Polyaxon

KubernetesInference
Share CPU and memory across containers in a Pod

Share CPU and memory across containers in a Pod

Use Pod-level resource budgets for cooperating containers, understand the remaining isolation boundaries, and evaluate the fit for ML services.

Aug 15, 2026

Polyaxon

KubernetesResources
Self-hosted vs. managed AI inference

Self-hosted vs. managed AI inference

Choose between self-hosted and managed AI inference using data policy, model access, latency, utilization, reliability, staffing, and cost per accepted outcome.

Jul 20, 2026

Polyaxon

InferenceInfrastructure
Serve DeepSeek V4 on Polyaxon

Serve DeepSeek V4 on Polyaxon

Configure a DeepSeek V4 Pro service on B200 GPUs and check model loading, reasoning output, tool calls, and resource use.

May 3, 2026

Polyaxon

LLMOpsInference
Compare serverless and dedicated LLM inference with Polyaxon

Compare serverless and dedicated LLM inference with Polyaxon

Use Polyaxon components, matrices, tracking, and run comparisons to evaluate serverless model APIs against dedicated inference services on quality, latency, and cost.

Mar 18, 2026

Polyaxon

LLMOpsInference
Serve Qwen 3.6 on Polyaxon

Serve Qwen 3.6 on Polyaxon

Deploy Qwen3.6-27B with SGLang, then check the model's API behavior, memory use, and response quality before sending it production traffic.

Mar 15, 2026

Polyaxon

LLMOpsInference
When CPU-accelerated AI inference makes sense

When CPU-accelerated AI inference makes sense

Evaluate CPU inference using model fit, precision, latency, throughput, memory, Kubernetes placement, and end-to-end cost instead of assuming every model needs a GPU.

Sep 9, 2025

Polyaxon

InferenceKubernetes