Polyaxon v3 is coming →

Correct evaluation labels and rescore saved predictions

Keep corrected labels as a new dataset version, reuse predictions when model inputs are unchanged, and compare evaluation runs in Polyaxon.

September 30, 2024by Polyaxon

You discover a wrong label in your evaluation dataset after running an expensive model evaluation. If only the expected answer changed, you can compare the predictions you already saved against the corrected labels. That avoids running the model again just to recalculate accuracy.

In Polyaxon, keep the corrected dataset and the new evaluation as separate records. You can then compare the scores and see which label version each run used.

Corrected labels and existing saved predictions feed a new evaluation score, which is compared with the original evaluation in Polyaxon.

A label correction changes how a prediction is scored

Suppose a support-ticket classifier predicted account for case-104. A reviewer discovers that its expected label should have been account, too:

EvaluationSaved predictionExpected labelResult for this case
Original labels, v7accountbillingIncorrect
Corrected labels, v8accountaccountCorrect

The model produced the same answer. The score changes because the reference label changed. Record that explanation alongside the new evaluation.

Decide what needs to run again

  • Only held-out evaluation labels changed: save the corrected labels and run the scoring code again using the saved predictions.
  • Input text, prompts, retrieval context, or model settings changed: run inference again, then score the new predictions.
  • Training labels changed: review the affected training runs and whether the model needs retraining. Rescoring an evaluation does not address a training-data correction.

Before reusing predictions, check that the case IDs, input text, model, and inference settings are unchanged. Match predictions to labels by case ID, rather than relying on file order.

Keep both evaluations in Polyaxon

  1. Preserve the original labels and save the corrected labels as a new snapshot, such as labels-v8. Dataset versioning covers saving and registering these snapshots.
  2. Create a new scoring run that reads the saved predictions and corrected labels. Log both files as inputs, plus the metrics your evaluator calculates.
  3. Compare the original and corrected evaluation runs. Keep the model and prediction inputs fixed so the comparison shows the effect of the label correction.

This helper adds the input references, label version, and calculated accuracy to your existing scoring job. Call it between tracking.init() and tracking.end() with the files actually read by the evaluator:

from polyaxon import tracking


def record_rescore(labels_path, predictions_path, label_version, accuracy):
    tracking.log_data_ref(
        name="evaluation-labels", path=labels_path, is_input=True,
    )
    tracking.log_file_ref(
        name="saved-predictions", path=predictions_path, is_input=True,
    )
    tracking.log_inputs(label_version=label_version)
    tracking.log_metrics(accuracy=accuracy)

The tracking API records these inputs and metrics with the run; its lineage view lets you inspect the file references. Reference logging identifies the files. Keep the actual snapshots and predictions in artifact storage so they remain available for rescoring.

If the original score supported a model release, revisit that decision using the corrected evaluation. Retain the old run to show what was known at the time and link the new result when recording the revised decision.