Keep training choices separate from cluster settings
Use your framework's configuration for model architecture, precision, optimization, and data processing. Define the container, command, inputs, connections, and outputs in a Polyaxon component. Apply scheduling presets for the resource and placement settings maintained by your platform team.
The TRL integration is a concrete starting point: it uses a published container and exposes model and dataset choices as inputs. The Axolotl integration follows the same separation while retaining Axolotl's configuration format.
Pin the image and model, data, and code revisions for the runs you intend to compare. Referencing a model or dataset by name alone does not make it immutable.
Start with one job, then choose how to distribute it
Submit a single-node job with explicit accelerator requests and the connections needed to read data and save outputs. Inspect its status and logs before increasing the workload size.
For distributed training, choose the launcher or operator that matches your framework. Polyaxon's PyTorchJob integration requires the corresponding operator and custom resource definition on the cluster. Polyaxon manages the operation; the training framework and operator handle worker coordination and distributed execution.
Validate GPU memory requirements, communication between workers, and checkpoint recovery for your recipe. Adding GPU requests alone does not distribute a single-process training script.
Compare results and keep the useful checkpoints
Log training parameters and evaluation metrics, and write checkpoints to the configured run output storage. Use the comparison view to inspect changes across runs, then follow a candidate back to its recorded inputs and artifacts.
Use the same evaluation data and procedure when comparing model quality. Faster completion on different hardware or a lower training loss does not, by itself, establish a better model.
For interrupted work, Polyaxon can load artifacts from the earlier operation. Your training code must know which checkpoint to restore and how to resume from it.
Compare experiment results · Configure output storage · Resume training operations
Make the selected output traceable
Keep the chosen checkpoint, evaluation results, and source run together. Where model-registry access is enabled, register a version from that run and specify the artifacts it contains, so another workflow can refer to a named version.
Registration records a selection; it is not an independent quality or deployment check. Keep any approval criteria explicit, and validate the model with its intended serving environment before release.
When this workflow fits
Use this approach when your team owns its training code and Kubernetes compute and needs repeatable submissions, inspectable experiments, and retained outputs. It does not provide GPU capacity or replace framework-specific training and evaluation work.
Start with the linked TRL recipe for fine-tuning, or the TensorFlow quick start for a smaller managed-training example. Confirm one complete run—including output retrieval—before adding sweeps or multiple nodes.