Polyaxon v3 is coming →

Tune checkpoint storage with Kubernetes VolumeAttributesClass

Change persistent-volume performance profiles, inspect the completed change, and compare checkpoint behavior in Polyaxon GPU jobs.

September 29, 2025by Polyaxon
White Kubernetes wheel in a blue hexagon above stylized blue waves

A training job spends more time saving checkpoints as the model grows. The volume still has free space, and adding GPUs would leave the same storage bottleneck in place. Before moving the data, check whether the volume's provisioned performance can be changed.

Kubernetes VolumeAttributesClass provides a declarative way to request supported changes to an existing volume's attributes. A platform team can offer a few named performance profiles, while training jobs keep mounting the same PersistentVolumeClaim (PVC).

The API reached general availability in Kubernetes 1.34. The examples below use storage.k8s.io/v1; the installed CSI driver and its controller components must support volume modification. Kubernetes API availability alone does not establish that support. Kubernetes GA announcement.

Identify the storage constraint

Start with a checkpoint timeline. Measure serialization, writing, flushing, and any upload to the final artifact store separately. A slow object-store upload or CPU-bound serialization step will not become faster just because a block volume has more provisioned throughput.

ObservationEvidence to collect before changing the volume
Writing large checkpoint files dominates the pauseApplication write duration, bytes written, volume throughput, and node storage limits
Many small writes dominateI/O operations, latency, and the checkpoint file layout
The job waits before it can mount the volumePVC, attachment, and topology events; use the volume troubleshooting guide
Local writing is fast but publication is slowNetwork transfer and artifact-store timings

These are diagnostic directions, not benchmark results. Keep the checkpoint format, model, dataset, and durability requirements fixed while comparing storage profiles.

Define two driver-specific profiles

A StorageClass describes provisioning. A VolumeAttributesClass describes attributes supported by a particular CSI driver. Its parameters are immutable: create a second class, then change the PVC's reference to select it. VolumeAttributesClass documentation.

This example is specifically for the AWS EBS CSI driver and gp3 volumes. It assumes a compatible driver installation with the required controller permissions and volume-modification support. The type, iops, and throughput parameters come from that driver; they are not portable Kubernetes parameter names. Throughput values here are in MiB/s. EBS CSI parameters, volume modification.

apiVersion: storage.k8s.io/v1
kind: VolumeAttributesClass
metadata:
  name: checkpoint-baseline
driverName: ebs.csi.aws.com
parameters:
  type: gp3
  iops: "3000"
  throughput: "125"
---
apiVersion: storage.k8s.io/v1
kind: VolumeAttributesClass
metadata:
  name: checkpoint-throughput
driverName: ebs.csi.aws.com
parameters:
  type: gp3
  iops: "6000"
  throughput: "250"

These are illustrative configurations, not measured recommendations. Have the storage owner check the volume, instance, account, and regional limits, as well as the cost of each profile. Changing a class does not bypass the provider's modification constraints or make repeated changes instantaneous. Amazon EBS Elastic Volumes.

Give the job a persistent claim

The following claim assumes an existing ml-team namespace and an ebs-gp3 StorageClass using ebs.csi.aws.com. For this example, that StorageClass should use WaitForFirstConsumer so provisioning accounts for the consuming Pod's placement. The class names above are cluster-scoped; the claim belongs to the workload namespace.

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: training-checkpoints
  namespace: ml-team
spec:
  accessModes:
    - ReadWriteOnce
  storageClassName: ebs-gp3
  volumeAttributesClassName: checkpoint-baseline
  resources:
    requests:
      storage: 200Gi

Create this claim once for the example. Do not replace an existing production claim to match it. If you already have a bound claim, first check that its CSI driver, volume type, and current settings support the intended transition.

The performance profile does not change the claim's access mode or move the volume to another availability zone. ReadWriteOnce is not shared storage for distributed workers on different nodes. See the persistent-volume guide when choosing storage for a distributed training topology.

Mount the claim in a Polyaxon GPU job

Polyaxon can schedule the training process on a GPU and mount this existing claim. Save the following as checkpoint-training.yaml. Replace the image with your own image containing a train module that accepts the shown arguments. The example requires an NVIDIA device plugin, a suitable GPU node, and a Polyaxon agent operating in the claim's namespace. Configure private-image access if needed.

version: 1.1
kind: component
name: checkpoint-training
run:
  kind: job
  volumes:
    - name: checkpoints
      persistentVolumeClaim:
        claimName: training-checkpoints
  container:
    image: registry.example.com/ml/checkpoint-training:approved
    command: ["python", "-m", "train"]
    args:
      - "--checkpoint-dir"
      - "/checkpoints/baseline-run"
    env:
      - name: STORAGE_PROFILE
        value: checkpoint-baseline
    resources:
      requests:
        cpu: "4"
        memory: 8Gi
      limits:
        cpu: "8"
        memory: 16Gi
        nvidia.com/gpu: 1
    volumeMounts:
      - name: checkpoints
        mountPath: /checkpoints

Submit it to your selected agent and queue:

polyaxon run -p YOUR_PROJECT -f checkpoint-training.yaml --queue AGENT/QUEUE

The GPU request and volume mount use documented resource scheduling and volume mounting capabilities. For repeated use, the platform team can expose the PVC as a data connection so authors request a connection instead of repeating the mount.

Have the training code record STORAGE_PROFILE, checkpoint timings, bytes written, image digest, and code revision with the run using Polyaxon tracking. The environment variable is a record label; it does not change the PVC. Keep it aligned with the observed storage configuration.

Request a change and inspect its completion

After collecting the baseline, the storage owner can request the second profile:

kubectl patch pvc training-checkpoints -n ml-team --type=merge \
  -p '{"spec":{"volumeAttributesClassName":"checkpoint-throughput"}}'

Keep the desired configuration in the repository or controller that owns the PVC so another reconciliation does not restore the old reference. Then inspect both the desired and current classes:

kubectl get pvc training-checkpoints -n ml-team \
  -o jsonpath='{.spec.volumeAttributesClassName}{"\n"}{.status.currentVolumeAttributesClassName}{"\n"}{.status.modifyVolumeStatus}{"\n"}'

kubectl describe pvc training-checkpoints -n ml-team

spec.volumeAttributesClassName is the request. status.currentVolumeAttributesClassName reports the class in use. The modification status can describe a pending, in-progress, or infeasible request. A successful patch response alone does not establish that the volume now has the requested performance. PVC API reference.

Wait for the change to complete and for any provider-side transition to settle before recording the next comparison. Repeat the training job with a separate checkpoint directory and the updated STORAGE_PROFILE label. Compare runs sequentially on this claim so concurrent writes do not obscure the effect.

Keep the profile that improves the whole run

Use the Polyaxon comparison view to evaluate checkpoint pause time alongside total runtime and model-quality checks. A faster write stage may make little difference if another stage dominates. A shorter run may also cost more if the storage change outweighs the saved GPU time.

Record the claim identity, applied class, measured interval, and provider cost assumptions with each comparison. If you return to the baseline profile, treat that as another asynchronous modification subject to the same provider constraints. For a notebook workspace or persistent preprocessing cache, use the same approach with the operation that matters to that workload: saving, scanning, or materializing data.

After the comparison, stop any remaining experiment runs and agree on the retained storage profile with the volume owner. Keep checkpoints according to your retention policy. Before deleting an example PVC, inspect its StorageClass and PV reclaim policy: deleting the claim can also delete the backing storage.