Kubernetes monitoring for ML workloads
Monitor Kubernetes control planes, nodes, containers, schedulers, applications, and ML outcomes with useful correlations and controlled cardinality.
Kubernetes can report that a pod is running while the training process inside it is stalled. It can show a fully allocated GPU while the device spends most of its time waiting for data. It can restart an unhealthy container without explaining whether the model service still meets its latency target.
Useful Kubernetes monitoring connects platform health to workload progress and user outcomes. Build the measurement model in layers, then preserve the identifiers needed to move between them.
Monitor five layers
| Layer | Questions it should answer |
|---|---|
| Control plane | Can the cluster accept, schedule, and reconcile work? |
| Nodes and capacity | Is eligible CPU, memory, storage, and accelerator capacity healthy? |
| Kubernetes workloads | Are Jobs, Deployments, Pods, and containers reaching the expected state? |
| Application | Is training, batch processing, or inference behaving correctly? |
| ML outcome | Is the workload producing useful artifacts, evaluations, or predictions? |
A red control-plane alert can explain many workloads at once. A healthy cluster cannot prove that one model is converging or one inference service is accurate. Keep infrastructure and ML outcome metrics separate, then join them during investigation.
Kubernetes describes observability as collecting and analyzing metrics, logs, and traces in its observability documentation. Events and audit records add important state-change and security context.
Cover control plane and nodes
Monitor API request behavior, scheduler and controller errors, webhook latency, datastore health where accessible, and certificate or quota conditions. Managed Kubernetes may expose a different subset than a self-managed control plane; document what the provider retains and what your team must collect.
At the node layer, collect readiness, pressure conditions, allocatable and requested resources, filesystem capacity, network errors, container-runtime health, and accelerator signals. For GPU nodes, distinguish:
- Allocation: capacity assigned to workloads;
- Activity: whether devices are executing work;
- Throughput: whether the application is completing useful units.
These measurements answer different questions. Continue with GPU utilization metrics for a detailed model.
Observe workload state transitions
Track desired and available replicas, pending and unschedulable pods, container restarts, termination reasons, probe failures, evictions, image pulls, volume attachments, and Job completion. Preserve the controller hierarchy so an alert on a container leads back to its pod, workload, namespace, project, and owner.
Pending time should be segmented. A workload waiting for quota, a node selector, a volume, a missing image, or unavailable accelerators requires a different response. Alerting on “pending pods” without the scheduling reason creates noise.
Kubernetes' resource monitoring guide notes that the resource metrics pipeline is intentionally limited. It supports features such as autoscaling and kubectl top; a complete operational pipeline needs broader system and application telemetry.
Add application and ML progress
Instrument training samples per second, step duration, data-loader time, checkpoint duration, loss or evaluation summaries, and heartbeat age. For batch work, record items completed, remaining work, failures, and retry outcomes. For inference, record request rate, latency distributions, errors, queue depth, batch size, model revision, and resource saturation.
Do not turn every run UUID, prompt, dataset row, or user into a metric label. Use bounded dimensions for aggregate views and place high-cardinality identifiers in logs, traces, or run records. Keep sensitive inputs and outputs out of telemetry by default.
Correlate each signal with environment, cluster, namespace, workload, component, project, and deployment revision. Preserve timestamps consistently so a rollout, node event, and application regression can be compared.
Work an incident from symptom to evidence
Consider an illustrative training incident: the run is active, its Pod is Running, and a GPU is assigned, but training stops advancing. The timeline below describes evidence you would collect; it is not a measured Polyaxon incident or a benchmark.
| Point in the investigation | Evidence | What it changes |
|---|---|---|
| The workload starts | Pod scheduled, container ready, model initialization completes | Scheduling and initial startup succeeded |
| Progress stalls | Sample counter stops increasing and the training heartbeat grows old | The application is no longer making expected progress |
| Resource charts are inspected | GPU activity falls while the GPU remains assigned; no restart or node-pressure event appears | Allocation does not explain the stall, and a restart is not yet the leading hypothesis |
| Application logs are correlated | Repeated data-read timeouts begin at the same time as the throughput drop | The input path becomes a concrete hypothesis to investigate |
| Dependency evidence is checked | Storage request errors or network failures align with the affected workload | The team can investigate a specific dependency instead of changing GPU size blindly |
First scope the investigation to the actual Pod and namespace recorded on the run:
kubectl describe pod TRAINING_POD --namespace TRAINING_NAMESPACE
kubectl logs TRAINING_POD --namespace TRAINING_NAMESPACE --container TRAINING_CONTAINER --since=15m --timestampsInspect scheduling events, GPU resource requests, container state, and relevant log timestamps together. Do not infer GPU assignment from device activity alone. If the container restarted, inspect the previous container's logs and termination reason as well.
For the application chart, suppose your training code exports a Prometheus counter named training_samples_total with a bounded workload label. This is instrumentation you provide, not a metric automatically emitted by Kubernetes or Polyaxon:
sum by (workload) (
rate(training_samples_total{namespace="ml-team", workload="image-training"}[5m])
)Use a window appropriate to the scrape interval and expected progress cadence. A flat counter with fresh scrapes differs from an absent series; check scrape health before interpreting either. Model loading, evaluation, or an expected checkpoint phase can legitimately pause training progress, so include the workload phase in the investigation.
On NVIDIA clusters with DCGM Exporter, a device-activity chart may use:
DCGM_FI_DEV_GPU_UTILFilter that chart to the device assigned to the affected Pod using the labels your exporter actually supplies. DCGM Exporter's counter definitions distinguish percentage utilization from profiling ratios and other measurements. Available metrics and workload attribution depend on the device, partitioning mode, and exporter configuration; do not assume every GPU series has a usable Pod label.
Once data-read failures are confirmed, inspect endpoint reachability, storage errors, credentials, and data locality. If the evidence points to a slow remote input path, compare one controlled change, such as a documented local staging strategy, against the same code, data version, and resource allocation. Keep throughput, read latency, and error evidence for both runs. Low GPU activity by itself does not justify that change, and higher activity alone does not establish success.
The useful incident record links the workload's state transitions to its application logs and dependency telemetry, then records the resulting decision. Our OpenTelemetry Collector guide covers transporting those signals; load testing ML services applies a similar evidence model to serving capacity.
Design logs, events, and traces together
Collect container logs from every relevant node and handle multiline records, rotation, backpressure, and pod deletion. Set log levels by environment and prevent secrets or training data from entering logs.
Retain Kubernetes events long enough for incident investigation and export them if the cluster's default retention is insufficient. Enable and protect audit records appropriate to your security policy. Trace inference and pipeline services across queues, feature stores, model servers, and downstream APIs where distributed latency matters.
Define what happens when the telemetry destination fails. Bounded buffers, rate limits, and drop policies should protect node storage and application performance while making missing data visible.
Alert on symptoms with owners
Page on conditions that require timely action: unavailable serving capacity, repeated workload failure, control-plane impairment, sustained resource pressure, or missed pipeline objectives. Route warnings such as inefficient requests, stale artifacts, and rising telemetry cost to review queues rather than waking an operator immediately.
Every alert needs an owner, severity, evidence links, runbook, and expected response. Test alerts by creating controlled failures and verify both firing and recovery. Review alerts that never fire, always fire, or cannot lead to an action.
Connect Kubernetes to Polyaxon
Use Polyaxon tracking to preserve parameters, metrics, code and environment context for ML runs. Keep detailed logs, checkpoints, and reports in artifacts. Use pipelines to encode recurring monitoring validation and workload checks.
Export stable Polyaxon identifiers into the observability system and link back to the run for high-cardinality detail. Kubernetes monitoring then explains what the platform did, while Polyaxon explains which ML operation ran, what it produced, and whether the result was useful.