Polyaxon v3 is coming →

Finish Indexed Jobs after enough attempts succeed

Use Kubernetes Job success policies for chosen indexes or success counts, preserve durable results, and distinguish them from Polyaxon metric early stopping.

September 25, 2026by Polyaxon
Kubernetes wheel surrounded by connected gears on a teal background

Some parallel workloads are finished before every attempt completes. A search may need two acceptable solutions, or a distributed computation may define success through a specific set of leader indexes. Continuing every remaining Pod can spend resources after the application already has what it needs.

Kubernetes Indexed Jobs support successPolicy, stable since Kubernetes 1.33. It lets a Job's author define success through completed indexes rather than requiring all configured completions. Once a rule is satisfied, the controller terminates the remaining Pods. Success policy GA announcement.

Use this when a partial set of successful attempts fulfills the application contract. An evaluation split into required data shards usually needs every shard; the Indexed Job retry guide preserves that completeness requirement.

Define what a successful attempt proves

A successful Pod exit is a signal from the application. Kubernetes does not examine a model's quality, compare generated candidates, or confirm that a result reached durable storage.

For a search that needs any two acceptable outputs, each attempt should perform the full required computation, validate its candidate, publish a durable result, and only then exit successfully. Different shards of one incomplete dataset are not interchangeable attempts.

Application requirementSuitable completion contract
Every evaluation case must be accounted forRequire all expected shard results and reconcile completeness
Any two independently valid solutions are sufficientCount successful attempts after application validation
A specific leader subset determines successDefine that subset explicitly and ensure its output is complete
Choose the best score across all candidatesDo not stop at the first acceptable count unless that is the intended search policy

The stopping rule changes which evidence you collect. Early stopping is a resource decision and an experimental-design decision, not merely a controller setting.

Choose a success rule

A success policy is available for Jobs with completionMode: Indexed. It can require all indexes in a specified set, a count across the Job, or a count from a specified subset. Multiple rules are alternatives evaluated in order; satisfying one is enough. Job controller documentation.

For example:

successPolicy:
  rules:
    - succeededIndexes: "0-2"
      succeededCount: 2

This means two successful indexes from the set 0, 1, and 2. Index 3 would not contribute to that rule, even if it finished first. Remove succeededIndexes to count successful indexes from the entire Job; remove succeededCount to require all listed indexes.

Do not combine several rules to express an AND condition. Put the required subset and count into the appropriate rule, and perform any richer result validation in the application.

Follow the controller with a small example

This disposable fixture needs Kubernetes 1.33 or later, permission to create a namespace and Job, and access to the Python image. It runs no model and produces no candidate artifacts. Its only purpose is to make index completion and early termination observable.

Save it as success-policy-demo.yaml:

apiVersion: v1
kind: Namespace
metadata:
  name: success-policy-demo
---
apiVersion: batch/v1
kind: Job
metadata:
  name: enough-attempts
  namespace: success-policy-demo
spec:
  completionMode: Indexed
  completions: 4
  parallelism: 4
  backoffLimit: 4
  activeDeadlineSeconds: 300
  successPolicy:
    rules:
      - succeededIndexes: "0-2"
        succeededCount: 2
  template:
    spec:
      restartPolicy: Never
      containers:
        - name: worker
          image: python:3.12-slim
          command:
            - python
            - -c
            - |
              import json
              import os
              import time

              index = int(os.environ["JOB_COMPLETION_INDEX"])
              delay_seconds = [2, 4, 60, 90][index]
              print(json.dumps({"index": index, "event": "started"}), flush=True)
              time.sleep(delay_seconds)
              print(json.dumps({"index": index, "event": "finished"}), flush=True)
          resources:
            requests:
              cpu: 100m
              memory: 64Mi
            limits:
              cpu: 500m
              memory: 128Mi

The delays are synthetic. Scheduling, image pulls, and controller timing affect the actual order. If all Pods start promptly and no unrelated failure intervenes, indexes 0 and 1 are expected to satisfy the rule while 2 and 3 are still waiting. This is an expectation, not captured execution output. Use an approved image digest if you need a fixed fixture environment.

Create the example and inspect its progress:

kubectl apply -f success-policy-demo.yaml

kubectl get job enough-attempts -n success-policy-demo -o yaml

kubectl get pods -n success-policy-demo \
  -l batch.kubernetes.io/job-name=enough-attempts

kubectl logs -n success-policy-demo \
  -l batch.kubernetes.io/job-name=enough-attempts --prefix=true

Observe the decision and the completed shutdown

The Job can report SuccessCriteriaMet with reason SuccessPolicy before it reaches its final Complete condition. Kubernetes waits for the Job's Pods to terminate before reporting that terminal completion. A success decision is therefore not proof that every requested GPU has already been released. Success-policy lifecycle.

Failure and time bounds still matter. A success rule does not override an exhausted failure budget, a deadline, or a terminating failure policy. Inspect the actual conditions and events rather than treating one successful index as proof that the overall Job must succeed.

Choose which layer owns retries when integrating this Job into a larger workflow. Retrying the parent while the original Job is still active can duplicate expensive attempts and complicate result selection.

Commit results before reporting success

For a real workload, write a result receipt containing the immutable input revision, index, attempt identifier, validation result, and artifact location. Use a write protocol that can tolerate retries or duplicate execution, then expose only complete outputs to downstream consumers.

Keep temporary output paths distinct from committed ones. The controller may terminate another attempt while it is writing, so an object appearing in a directory is not by itself proof that it is usable. Preserve a summary of completed and stopped attempts with the selected results.

A reducer should verify the application contract from those receipts. Kubernetes' completion count answers a controller question; it does not replace model-quality checks or dataset completeness checks.

Choose the corresponding Polyaxon workflow

Polyaxon matrix operations are useful when every attempt should be an individually tracked run with its own parameters, metrics, and artifacts. For a search that can stop once one run meets a metric threshold, Polyaxon also provides metric early stopping.

Merge this matrix.earlyStopping setting into an existing search operation whose component logs the named metric:

matrix:
  earlyStopping:
    - kind: metric_early_stopping
      metric: accepted_result
      value: 0.5
      optimization: maximize

Here the application-defined convention is to emit accepted_result: 1 only after validation and durable result publication. A value of zero is insufficient. The threshold is illustrative; it is not a measured quality target.

Without a stopping policy, this condition ends the pipeline successfully and stops pending or running operations when the metric condition is met. Finish saving the result before logging the signal, because the emitting run is also subject to that stop behavior. Metric-based stopping is a distinct mechanism from counting two successful Kubernetes indexes.

If you specifically need the native subset/count contract, keep the Indexed Job as the execution unit and have your integration retain its final conditions and result receipts in the Polyaxon workflow. Avoid configuring independent stopping rules in both layers without deciding which outcome owns the final result.

After reviewing the disposable example, remove only its namespace:

kubectl delete namespace success-policy-demo