Polyaxon v3 is coming →

Evaluation articles

Browse Polyaxon articles about Evaluation. Page 1 of 2.

Separate evaluator changes from application improvements

Separate evaluator changes from application improvements

Compare old and new application outputs under both evaluator versions, inspect changed decisions, and retain the four comparisons in Polyaxon.

Sep 22, 2026

Polyaxon

LLMOpsEvaluation
Keep failed candidates visible in pipeline reports

Keep failed candidates visible in pipeline reports

Reconcile evaluation results against an expected candidate manifest, preserve failed and missing outcomes, and separate pipeline reporting from release approval.

Sep 19, 2026

Polyaxon

MLOpsOrchestration
Version the dataset behind every evaluation

Version the dataset behind every evaluation

Use Polyaxon data references, artifact logging, and registered versions to retain the dataset and split behind each evaluation.

Sep 18, 2026

Polyaxon

DataOpsMLOps
Prevent data leakage in ML and LLM evaluation datasets

Prevent data leakage in ML and LLM evaluation datasets

Choose evaluation boundaries, keep related examples together, audit duplicate overlap, and retain reproducible split manifests for ML and LLM experiments.

Sep 16, 2026

Polyaxon

DataOpsMLOps
Connect LLM production traces to evaluation runs

Connect LLM production traces to evaluation runs

Keep release, dataset, evaluator, and sampling context connected as you investigate LLM failures and compare fixes with Polyaxon.

Sep 10, 2026

Polyaxon

LLMOpsObservability
Retry failed evaluation shards with Indexed Jobs

Retry failed evaluation shards with Indexed Jobs

Give evaluation shards stable indexes and independent retry budgets with Kubernetes Indexed Jobs, while keeping incomplete results visible.

Sep 3, 2026

Polyaxon

KubernetesEvaluation
Run Promptfoo evaluations on Kubernetes with Polyaxon

Run Promptfoo evaluations on Kubernetes with Polyaxon

Package a Promptfoo suite as a Polyaxon job, export evaluation reports, track artifacts, and verify that failed checks fail the workload.

Sep 2, 2026

Polyaxon

PromptfooEvaluation
Beyond agent traces: State, artifacts, and recovery

Beyond agent traces: State, artifacts, and recovery

Connect agent traces to checkpoint state, versioned artifacts, action receipts, and evaluations so failures lead to controlled recovery and reproducible improvements.

Aug 25, 2026

Polyaxon

AI AgentsObservability
Design measurable AI agents before you build

Design measurable AI agents before you build

Define agent outcomes, prohibited actions, events, evaluators, budgets, segments, and feedback loops before implementation begins.

Aug 25, 2026

Polyaxon

AI AgentsEvaluation
Run batch LLM evaluations on Kubernetes

Run batch LLM evaluations on Kubernetes

Design reliable batch LLM evaluations with stable shards, bounded concurrency, GPU-aware placement, resumable results, and complete aggregation.

Aug 20, 2026

Polyaxon

EvaluationKubernetes
Build a shared library of agent evaluators

Build a shared library of agent evaluators

Package application-owned agent evaluators as versioned Polyaxon components with typed contracts, reusable execution profiles, and consistent result evidence.

Aug 20, 2026

Polyaxon

AI AgentsEvaluation
Continuous AI red teaming in CI/CD

Continuous AI red teaming in CI/CD

Turn AI security findings into repeatable CI checks with versioned cases, complete result manifests, explicit release gates, and retained evidence.

Aug 13, 2026

Polyaxon

Red TeamingPipelines
Build guardrails for AI-generated code

Build guardrails for AI-generated code

Control what coding agents can read, change, execute, merge, and deploy with layered checks from context selection through production.

Aug 8, 2026

Polyaxon

GuardrailsSecurity
How to evaluate LLM guardrails

How to evaluate LLM guardrails

Compare LLM guardrails using attack blocking, legitimate task success, false refusals, latency, cost, and explicit handling of errors.

Jul 30, 2026

Polyaxon

GuardrailsRed Teaming
AI red teaming metrics: measuring failures and coverage

AI red teaming metrics: measuring failures and coverage

Measure AI red teaming with explicit attack success rates, attempt budgets, coverage, false refusals, severity, and evaluator uncertainty.

Jul 23, 2026

Polyaxon

Red TeamingEvaluation
How to evaluate LLM routers for cost, quality, and latency

How to evaluate LLM routers for cost, quality, and latency

Test LLM routing policies with task-level quality, cost, latency, fallbacks, route stability, and model-aware production evidence.

Jul 21, 2026

Polyaxon

LLMOpsEvaluation
Red teaming RAG systems

Red teaming RAG systems

Test RAG systems for poisoned documents, cross-tenant retrieval, stale permissions, citation leaks, and unauthorized tool actions.

Jul 16, 2026

Polyaxon

Red TeamingRag
How to test prompt injection in LLM applications

How to test prompt injection in LLM applications

Build prompt injection tests for user input, retrieved documents, and tool responses, with checks for data access and actual side effects.

Jul 2, 2026

Polyaxon

Red TeamingEvaluation
Verified data pipelines for reliable AI agents

Verified data pipelines for reliable AI agents

Design agent data paths that preserve source identity, versions, permissions, freshness, retrieval evidence, and verification results.

Jun 26, 2026

Polyaxon

AI AgentsDataOps
Offline vs. online evaluation for generative AI

Offline vs. online evaluation for generative AI

Use offline evaluation for reproducible release decisions and online evaluation for real production behavior, then connect both in one feedback loop.

Jun 25, 2026

Polyaxon

EvaluationLLMOps
Build evaluation leaderboards with Polyaxon joins

Build evaluation leaderboards with Polyaxon joins

Use Polyaxon joins to select comparable evaluation runs and build application-owned leaderboards with explicit cohort, completeness, and artifact evidence.

Jun 18, 2026

Polyaxon

AI AgentsEvaluation
LLM-as-a-judge: Design and validate model-based evaluators

LLM-as-a-judge: Design and validate model-based evaluators

Use LLMs as scalable evaluators without treating them as ground truth: design clear rubrics, calibrate against humans, and monitor bias and drift.

Jun 18, 2026

Polyaxon

EvaluationLLMOps
Turn production traces into regression tests

Turn production traces into regression tests

Turn a failure cluster into a reviewed regression case, preserve the causal tool response, and record baseline and candidate checks with Polyaxon.

Jun 11, 2026

Polyaxon

LLMOpsObservability
How to evaluate RAG systems

How to evaluate RAG systems

Evaluate retrieval and generation separately and end to end with representative datasets, groundedness checks, retrieval metrics, and production feedback.

Jun 4, 2026

Polyaxon

RagEvaluation