Define what each operation produces
Start with preparation, training, and evaluation components that work independently. Give them explicit inputs and outputs: a dataset reference, a model artifact, and an evaluation result. Connect the operations through dependencies and output references in a DAG.
Keep the evaluation dataset and procedure under version control appropriate to those assets. A pipeline can repeat a process, but an unrecorded change to the data or scoring code can still invalidate the comparison.
Define dependencies and output references · Record experiment inputs and outputs
Distinguish a successful job from an acceptable model
An evaluator exiting successfully means that the evaluation ran; it does not mean that the model should be released. Configure the downstream operation to require successful upstream execution and a condition on the recorded result.
Define how missing or invalid results are handled. They must not be treated as a passing score. If the condition is not met, the promotion step should not run; inspect the evaluation record before changing the recipe or its criteria.
This is ordinary workflow orchestration around your evaluation code. It does not require or imply access to the private-beta GenAI evaluation product.
Require approval before the handoff
When review is required, configure the registration operation with isApproved: false. Its dependent branch waits for approval rather than continuing immediately after evaluation. A reviewer can inspect the candidate's metrics, artifacts, and source run before approving the next operation.
The registration step uses the model-registry API or CLI to associate selected artifacts with a version. Keep stage changes and deployment actions explicit. Registering a version does not replace a running endpoint or validate its production behavior.
Decide how repeats and failures should behave
Once the workflow is working, add a cron or interval schedule. Use the schedule's dependency on the previous execution when overlapping refreshes would be unsafe, and bound parallel jobs with workflow concurrency limits.
Choose recovery behavior for each step. Training may resume from retained artifacts if the application supports it; a data-writing or registration step needs to be safe to rerun. Inspect existing outputs and external side effects before restarting a failed operation.
Automate an established process first
This fits teams that already know how to prepare data, train, and evaluate a candidate and want those operations to run with recorded dependencies and controlled handoffs. Keep a one-off experiment as a job until the repetition or dependencies justify a pipeline.
The conditional DAG example demonstrates the orchestration primitives with an experiment, threshold-based tuning, and result collection. Use that as a starting point, then exercise both passing and failing evaluations and a pending approval before enabling recurring model updates.
Follow the conditional DAG example · Understand the workflow engine