Kubernetes for AI
Understand how Kubernetes architecture, persistent storage, metrics, and resource inspection affect ML infrastructure.
4 guides · Suggested reading order
Start here
Kubernetes architecture for ML workloads
Understand control planes, workers, services, and storage.
Continue the learning path
Go deeper
GPU jobs stuck Pending on Kubernetes: a debugging guide
Diagnose GPU resource availability and the constraints that prevent placement.
GPU sharing on Kubernetes: MIG vs. time-slicing
Compare MIG instances and shared GPU access before configuring workloads.
Run batch LLM evaluations on Kubernetes
Design evaluation workers that tolerate retries and produce complete results.
Run Promptfoo evaluations on Kubernetes with Polyaxon
Package an evaluation suite and submit it as a Polyaxon job.
Put the foundations into practice
Explore practical guides and resources to take the next step.