Correct evaluation labels and rescore saved predictions
Keep corrected labels as a new dataset version, reuse predictions when model inputs are unchanged, and compare evaluation runs in Polyaxon.
You discover a wrong label in your evaluation dataset after running an expensive model evaluation. If only the expected answer changed, you can compare the predictions you already saved against the corrected labels. That avoids running the model again just to recalculate accuracy.
In Polyaxon, keep the corrected dataset and the new evaluation as separate records. You can then compare the scores and see which label version each run used.
A label correction changes how a prediction is scored
Suppose a support-ticket classifier predicted account for case-104. A reviewer discovers that its expected label should have been account, too:
| Evaluation | Saved prediction | Expected label | Result for this case |
|---|---|---|---|
Original labels, v7 | account | billing | Incorrect |
Corrected labels, v8 | account | account | Correct |
The model produced the same answer. The score changes because the reference label changed. Record that explanation alongside the new evaluation.
Decide what needs to run again
- Only held-out evaluation labels changed: save the corrected labels and run the scoring code again using the saved predictions.
- Input text, prompts, retrieval context, or model settings changed: run inference again, then score the new predictions.
- Training labels changed: review the affected training runs and whether the model needs retraining. Rescoring an evaluation does not address a training-data correction.
Before reusing predictions, check that the case IDs, input text, model, and inference settings are unchanged. Match predictions to labels by case ID, rather than relying on file order.
Keep both evaluations in Polyaxon
- Preserve the original labels and save the corrected labels as a new snapshot, such as
labels-v8. Dataset versioning covers saving and registering these snapshots. - Create a new scoring run that reads the saved predictions and corrected labels. Log both files as inputs, plus the metrics your evaluator calculates.
- Compare the original and corrected evaluation runs. Keep the model and prediction inputs fixed so the comparison shows the effect of the label correction.
This helper adds the input references, label version, and calculated accuracy to your existing scoring job. Call it between tracking.init() and tracking.end() with the files actually read by the evaluator:
from polyaxon import tracking
def record_rescore(labels_path, predictions_path, label_version, accuracy):
tracking.log_data_ref(
name="evaluation-labels", path=labels_path, is_input=True,
)
tracking.log_file_ref(
name="saved-predictions", path=predictions_path, is_input=True,
)
tracking.log_inputs(label_version=label_version)
tracking.log_metrics(accuracy=accuracy)The tracking API records these inputs and metrics with the run; its lineage view lets you inspect the file references. Reference logging identifies the files. Keep the actual snapshots and predictions in artifact storage so they remain available for rescoring.
If the original score supported a model release, revisit that decision using the corrected evaluation. Retain the old run to show what was known at the time and link the new result when recording the revised decision.