Reuse operations instead of coordinating scripts by hand
A Polyaxon DAG connects operations through dependencies and input/output references. A preparation job can produce the inputs for training; an evaluation job can consume the resulting model. Each operation keeps its own configuration, status, logs, and outputs.
The steps can use different runtimes, including jobs, distributed jobs, services, or nested workflows. The DAG coordinates them; the components and their containers define the actual work.

Make continuation a deliberate choice
Use a trigger to decide which upstream statuses permit the next operation. Add a condition when the decision depends on a recorded output, such as a validation score. A job succeeding is not the same as a model meeting your quality requirement.
For a human review, mark the relevant operation as requiring approval. Its dependent branch waits until that operation is approved. Define the promotion or publication step explicitly; Polyaxon does not decide what makes a model acceptable.
Schedule recurring work and bound parallel runs
Run an operation or DAG on a cron or interval schedule. Where runs should not overlap, configure the schedule to depend on its previous execution. A parameter sweep can explore multiple configurations while a workflow concurrency limit bounds its parallel work.
The documented automation example starts an experiment, runs a tuning process when loss exceeds a configured threshold, and collects the best results. It shows how conditions, sweeps, and output references fit together without requiring a separate script to submit every run.
Walk through a conditional DAG · Configure a schedule · Limit concurrency
Inspect failures before choosing how to recover
Follow the failed operation's logs, outputs, and lineage to determine whether the input, code, or environment needs to change. Resume an operation when it should reuse its prior artifacts, or restart it when you need a separate run record. Copy mode can load artifacts into a restarted run.
Your application must implement checkpoint loading for training to resume. Review the downstream dependencies before retrying work that writes to external systems: restarting an operation does not roll back those writes or make them safe to repeat.