Polyaxon v3 is coming →

Verify credentials for cached private images on Kubernetes

Understand kubelet credential verification for cached images and apply it to shared Polyaxon training, notebook, and inference nodes.

August 28, 2025by Polyaxon
White package cube inside a blue hexagon

Two teams share a GPU node. One team pulls a private training image, leaving it in the node's cache. Later, another team requests the same image with imagePullPolicy: IfNotPresent. Whether that second workload can start should follow the intended image-access policy, even though downloading the layers again may be unnecessary.

Kubernetes has a kubelet feature for this boundary: KubeletEnsureSecretPulledImages. It is beta and enabled by default from Kubernetes 1.35. Its effective behavior depends on the node's credential-verification policy, including how the node treats preloaded images. Kubernetes image credential verification.

This is especially relevant for shared ML infrastructure. Training, notebooks, and inference services may use large private images, and keeping a warm cache can save substantial startup work. The cache and the permission to use an image need separate consideration.

Separate image identity from permission

An image digest identifies content. A registry credential establishes access. A pull policy influences when Kubernetes tries the registry. None of these replaces the others.

Pinning a training image by digest helps make runs reproducible, but it does not grant access to that image. Similarly, a cached copy is not evidence that every project using the node should have access to it. The original Kubernetes feature announcement explains the shared-node problem that motivated credential verification.

Consider an illustrative case: team Vision has a pull secret for registry.example.com/vision/train; team Language has no credential for that repository. Both may reach the same node, and there are no node-wide credentials granting Language access. Credential verification lets the kubelet check the second request even when Vision already populated the cache.

This is one part of the tenancy design. Pod authors' ability to select ServiceAccounts or reference Secrets in a shared namespace still needs appropriate control. A Polyaxon team name alone does not establish a Kubernetes credential boundary. Use the broader multi-tenancy guide alongside this node policy.

Choose the node policy deliberately

The relevant field is imagePullCredentialsVerificationPolicy in KubeletConfiguration. It is not a Pod field, registry setting, or Polyaxonfile option.

PolicyCached-image treatment
NeverVerifySkip credential verification for locally present images
NeverVerifyPreloadedImagesExempt images pulled outside the kubelet; the default policy
NeverVerifyAllowlistedImagesRestrict the preload exemption to the configured image allowlist
AlwaysVerifyRequire verification regardless of how the image reached the node

Use the exact NeverVerifyAllowlistedImages spelling from the Kubelet configuration API. The field's accepted values are also defined in the Kubernetes source.

For a node pool intended to verify all cached-image use, a platform owner could merge these fields into its existing kubelet configuration:

apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
  KubeletEnsureSecretPulledImages: true
imagePullCredentialsVerificationPolicy: AlwaysVerify

This is a configuration fragment, not a complete replacement file and not an object to submit with kubectl apply. Apply it through the node lifecycle mechanism supported by your distribution. On managed clusters, first confirm that the provider exposes this setting and that replacement nodes receive it too.

Verification can still reuse the cache

AlwaysVerify does not mean contacting the registry for every container start. If the locally available image was successfully pulled with the same credentials before, the kubelet can validate that history locally. Credentials without that successful history require a registry pull attempt.

That distinction also matters for revocation: previously successful credentials can continue to verify locally after the registry changes. Do not describe this feature as immediate registry revocation enforcement. New or rotated credentials follow a new verification path. Credential reuse and rotation.

Keep imagePullPolicy decisions explicit. Selecting AlwaysVerify is not the same configuration change as selecting imagePullPolicy: Always. Decide whether you require fresh registry interaction, digest-based reproducibility, cached startup, or a particular combination, then review the resulting behavior on your cluster.

Use the intended pull identity in Polyaxon

The workload still needs the correct image credentials. Polyaxon exposes the Kubernetes ServiceAccount through run.environment.serviceAccountName. A ServiceAccount can carry the pull-secret reference for the operations that use it. The private-image walkthrough covers creating that association.

Assume the agent's ml-team namespace already contains a vision-image-puller ServiceAccount with an imagePullSecrets reference to a Secret authorized for the private repository. Save this component as vision-training.yaml:

version: 1.1
kind: component
name: vision-training
run:
  kind: job
  environment:
    serviceAccountName: vision-image-puller
  container:
    image: registry.example.com/vision/train:approved
    imagePullPolicy: IfNotPresent
    command: ["python", "-m", "train"]
    resources:
      requests:
        cpu: "4"
        memory: 8Gi
      limits:
        nvidia.com/gpu: 1

Replace the illustrative image with an image containing your training module, preferably pinned by digest. The example assumes a GPU-capable agent and the required device plugin; it does not supply the application or registry credentials.

polyaxon run -p YOUR_PROJECT -f vision-training.yaml --queue AGENT/QUEUE

The selected agent and queue determine where the workload is submitted. The ServiceAccount must exist in that workload namespace and remain compatible with any Polyaxon auxiliary containers. The receiving node's kubelet applies the image-verification policy. Custom ServiceAccounts, queue scheduling.

A platform team can standardize the ServiceAccount through a scheduling preset. Apply the same credential discipline to notebook and inference images, including images used by init containers or sidecars. Presets simplify workload configuration; they do not configure kubelets or independently enforce a security boundary against arbitrary Pod authors.

Inspect the Pod that actually reached the node

When a run cannot start, first inspect its resolved Pod. Replace POD_NAME and the namespace below with those for the affected run:

kubectl get pod POD_NAME -n ml-team \
  -o jsonpath='{.spec.nodeName}{"\n"}{.spec.serviceAccountName}{"\n"}{.spec.imagePullSecrets[*].name}{"\n"}{.spec.containers[*].image}{"\n"}{.spec.containers[*].imagePullPolicy}{"\n"}'

kubectl describe pod POD_NAME -n ml-team

These commands identify placement, credential references, and image settings without printing Secret contents. Associate that evidence with the Polyaxon run ID and have the node owner confirm the effective kubelet configuration. An image-pull error can also come from a missing image, network failure, registry throttling, or an incorrect Secret; caching is only one part of the investigation.

For a rollout review, include a cached image with an authorized identity, the same image without that access, a fresh pull, and any deliberate preload exemption. Specify the expected outcome for each case before changing a shared node pool. Include autoscaled replacement nodes so the policy remains consistent as capacity changes.

Account for existing node caches

On first enablement, the kubelet has no pull history for images already on the node. Kubernetes treats those as preloaded. Under the default preload-exempt policy, that distinction affects verification. Initial enablement behavior.

Inventory intentional preloads and historical caches with the node owner. Choose the policy and rollout sequence around that inventory; an upgrade alone is not evidence that every cached private image now requires credentials.

Keep the effective node policy, workload ServiceAccount, image digest, and observed startup outcome with the run's operational record. That gives the next investigation enough context to distinguish an image-access failure from an application failure.