Tune hyperparameters with native TPE in Polyaxon
Use completed trial results to guide hyperparameter choices with native TPE, run GPU training trials, and compare validation loss in Polyaxon.
You have enough GPU time for a few dozen training runs. Instead of choosing every learning rate in advance, you want the results of completed runs to influence what you try next.
Polyaxon 2.17 adds a native Tree-structured Parzen Estimator (TPE) implementation, replacing the Hyperopt-backed matrix. Set matrix.kind: tpe, define the metric to optimize, and Polyaxon generates suggestions, schedules training trials, and feeds their results into subsequent suggestions.
How TPE chooses what to try next
Grid search visits predefined combinations. Random search samples configurations independently of their scores. TPE starts with random samples, then uses the observed scores to guide later choices.
It separates better-performing trials from the remaining trials and models the parameter values in each group. It favors candidates that are more likely under the better-performing group's model. This is the idea behind the original TPE method; Polyaxon's native implementation models each parameter independently.
TPE is useful when individual trials are expensive and you can explore only part of the search space. Its suggestions still depend on the ranges you choose and the metric you report. Use a validation metric that matches your selection goal.
Define the search
The example trains a small PyTorch classifier on Fashion-MNIST. It tunes learning rate and AdamW weight decay while holding the model, training duration, split, and training seed fixed.
Use a Polyaxon 2.17-compatible deployment and CLI, a configured project and artifact store, and an NVIDIA GPU queue. The deployment needs its native TPE tuner. The containers need access to the package index and dataset downloads.
Save the following as tpe-study.yaml. The training component and program are supplied in the next two sections.
version: 1.1
kind: operation
name: fashion-mnist-tpe
tags: [native-tpe, fashion-mnist]
pathRef: ./train-component.yaml
queue: research/gpu
params:
epochs: 3
seed: 23
matrix:
kind: tpe
numRuns: 6
maxIterations: 3
concurrency: 2
seed: 19
metric:
name: validation_loss
optimization: minimize
params:
learning_rate:
kind: loguniform
value: [-9.21034, -4.60517] # ln(0.0001), ln(0.01)
weight_decay:
kind: loguniform
value: [-13.81551, -4.60517] # ln(0.000001), ln(0.01)loguniform uses natural-log bounds: these ranges sample learning rates of about 0.0001–0.01 and weight decay of about 0.000001–0.01. The parameter distributions reference describes how the bounds are encoded.
The budget controls have different roles:
| Setting | Effect in this study |
|---|---|
numRuns: 6 | Generate six suggestions per batch |
maxIterations: 3 | Initial batch at iteration 0, then iterations 1–3: up to 24 training trials |
concurrency: 2 | Permit up to two trial operations at once, subject to queue capacity |
matrix.seed: 19 | Control suggestion sampling |
params.seed: 23 | Keep the training seed fixed across trials |
The tuner uses completed results between batches. A larger batch provides more work to schedule in parallel, but commits more trials before incorporating fresh results. The initial batch here uses random sampling; later suggestions become adaptive as observations accumulate.
Schedule each trial on a GPU
Save this as train-component.yaml in the same folder:
version: 1.1
kind: component
name: fashion-mnist-train
inputs:
- name: learning_rate
type: float
- name: weight_decay
type: float
- name: epochs
type: int
value: 3
- name: seed
type: int
value: 23
outputs:
- name: validation_loss
type: float
run:
kind: job
termination:
timeout: 1200
maxRetries: 0
container:
image: pytorch/pytorch:2.7.1-cuda12.6-cudnn9-runtime
workingDir: "{{ globals.run_artifacts_path }}/uploads"
command: [bash, -c]
args:
- |
set -euo pipefail
python -m pip install --no-cache-dir polyaxon torchvision==0.22.1
python -m pip freeze
exec python train.py \
--learning-rate {{ learning_rate }} \
--weight-decay {{ weight_decay }} \
--epochs {{ epochs }} --seed {{ seed }}
resources:
requests:
cpu: "2"
memory: "4Gi"
nvidia.com/gpu: 1
limits:
cpu: "4"
memory: "8Gi"
nvidia.com/gpu: 1Each training trial requests one GPU. The study's queue must have access to GPU nodes with compatible NVIDIA drivers and device-plugin support. Replace research/gpu with your allowed agent/queue. The termination timeout bounds each trial to 20 minutes; retries are disabled for this example.
The PyTorch runtime image supplies PyTorch and CUDA. Runtime installation supplies the tracking client and matching torchvision version. Keep the resolved environment recorded in the logs; pin your reviewed client version and image digest when repeating the study.
Train and return the objective
Save this complete program as train.py beside the two YAML files:
import argparse
from pathlib import Path
import torch
from torch import nn
from torch.utils.data import DataLoader, random_split
from torchvision.datasets import FashionMNIST
from torchvision.transforms import ToTensor
from polyaxon import tracking
parser = argparse.ArgumentParser()
parser.add_argument("--learning-rate", type=float, required=True)
parser.add_argument("--weight-decay", type=float, required=True)
parser.add_argument("--epochs", type=int, default=3)
parser.add_argument("--seed", type=int, default=23)
args = parser.parse_args()
torch.manual_seed(args.seed)
device = torch.device("cuda")
tracking.init()
try:
data = FashionMNIST(
"/tmp/fashion-mnist", train=True, download=True, transform=ToTensor()
)
train_data, validation_data = random_split(
data, [50000, 10000], generator=torch.Generator().manual_seed(args.seed)
)
train_loader = DataLoader(
train_data, batch_size=256, shuffle=True, pin_memory=True
)
validation_loader = DataLoader(
validation_data, batch_size=256, pin_memory=True
)
model = nn.Sequential(
nn.Flatten(), nn.Linear(784, 128), nn.ReLU(), nn.Linear(128, 10)
).to(device)
optimizer = torch.optim.AdamW(
model.parameters(), lr=args.learning_rate, weight_decay=args.weight_decay
)
loss_fn = nn.CrossEntropyLoss()
for epoch in range(1, args.epochs + 1):
model.train()
for images, labels in train_loader:
images, labels = images.to(device), labels.to(device)
optimizer.zero_grad()
loss = loss_fn(model(images), labels)
loss.backward()
optimizer.step()
model.eval()
total_loss = 0.0
with torch.inference_mode():
for images, labels in validation_loader:
images, labels = images.to(device), labels.to(device)
total_loss += loss_fn(model(images), labels).item() * len(labels)
validation_loss = total_loss / len(validation_data)
tracking.log_metrics(step=epoch, validation_loss=validation_loss)
checkpoint = Path(tracking.get_outputs_path("model.pt"))
torch.save({"state_dict": model.state_dict(), "params": vars(args)}, checkpoint)
tracking.log_file_ref(path=str(checkpoint), name="model", is_input=False)
tracking.log_outputs(async_req=False, validation_loss=validation_loss)
finally:
tracking.end()The tracking API saves the validation curve and final objective with each run. The final output name must match matrix.metric.name. The tuner collects successful trials with that objective; missing results cannot guide the next suggestions.
This program trains a fresh model in every trial. It uses 50,000 examples for training and 10,000 for validation from Fashion-MNIST's training split, leaving the separate test split for final assessment.
Launch and compare the trials
Submit from the folder containing all three files, replacing the project name:
polyaxon run --project=acme/hpo -f tpe-study.yaml -uThe upload option supplies the code under uploads, matching the component's working directory. Polyaxon creates the study, tuner operations, and training trials. Open the training children in the comparison dashboard, sort by validation_loss, and inspect their parameter values and curves. Check failed trials separately.
Confirm promising settings with additional training seeds before making the final choice. Then evaluate that fixed choice on the held-out test split. For a larger model, keep this orchestration and replace the training component with your own trainer and validation metric.
Move an existing Hyperopt matrix to native TPE
On Polyaxon 2.17, replace matrix.kind: hyperopt with matrix.kind: tpe and review your parameter distributions against the TPE reference. If you set a custom tuner.hubRef, review that tuner as part of the upgrade. The default native tuner no longer depends on Hyperopt.
Keep the objective, data, training duration, and resource allocation consistent when assessing the new study. TPE guides the search using your results; the quality of those results still depends on the training and evaluation procedure.