Evaluation articles
Browse Polyaxon articles about Evaluation. Page 1 of 2.

Separate evaluator changes from application improvements
Compare old and new application outputs under both evaluator versions, inspect changed decisions, and retain the four comparisons in Polyaxon.
Sep 22, 2026
Polyaxon
LLMOpsEvaluation
Keep failed candidates visible in pipeline reports
Reconcile evaluation results against an expected candidate manifest, preserve failed and missing outcomes, and separate pipeline reporting from release approval.
Sep 19, 2026
Polyaxon
MLOpsOrchestration
Version the dataset behind every evaluation
Use Polyaxon data references, artifact logging, and registered versions to retain the dataset and split behind each evaluation.
Sep 18, 2026
Polyaxon
DataOpsMLOps
Prevent data leakage in ML and LLM evaluation datasets
Choose evaluation boundaries, keep related examples together, audit duplicate overlap, and retain reproducible split manifests for ML and LLM experiments.
Sep 16, 2026
Polyaxon
DataOpsMLOps
Connect LLM production traces to evaluation runs
Keep release, dataset, evaluator, and sampling context connected as you investigate LLM failures and compare fixes with Polyaxon.
Sep 10, 2026
Polyaxon
LLMOpsObservability
Retry failed evaluation shards with Indexed Jobs
Give evaluation shards stable indexes and independent retry budgets with Kubernetes Indexed Jobs, while keeping incomplete results visible.
Sep 3, 2026
Polyaxon
KubernetesEvaluation
Run Promptfoo evaluations on Kubernetes with Polyaxon
Package a Promptfoo suite as a Polyaxon job, export evaluation reports, track artifacts, and verify that failed checks fail the workload.
Sep 2, 2026
Polyaxon
PromptfooEvaluation
Beyond agent traces: State, artifacts, and recovery
Connect agent traces to checkpoint state, versioned artifacts, action receipts, and evaluations so failures lead to controlled recovery and reproducible improvements.
Aug 25, 2026
Polyaxon
AI AgentsObservability
Design measurable AI agents before you build
Define agent outcomes, prohibited actions, events, evaluators, budgets, segments, and feedback loops before implementation begins.
Aug 25, 2026
Polyaxon
AI AgentsEvaluation
Run batch LLM evaluations on Kubernetes
Design reliable batch LLM evaluations with stable shards, bounded concurrency, GPU-aware placement, resumable results, and complete aggregation.
Aug 20, 2026
Polyaxon
EvaluationKubernetes
Build a shared library of agent evaluators
Package application-owned agent evaluators as versioned Polyaxon components with typed contracts, reusable execution profiles, and consistent result evidence.
Aug 20, 2026
Polyaxon
AI AgentsEvaluation
Continuous AI red teaming in CI/CD
Turn AI security findings into repeatable CI checks with versioned cases, complete result manifests, explicit release gates, and retained evidence.
Aug 13, 2026
Polyaxon
Red TeamingPipelines
Build guardrails for AI-generated code
Control what coding agents can read, change, execute, merge, and deploy with layered checks from context selection through production.
Aug 8, 2026
Polyaxon
GuardrailsSecurity
How to evaluate LLM guardrails
Compare LLM guardrails using attack blocking, legitimate task success, false refusals, latency, cost, and explicit handling of errors.
Jul 30, 2026
Polyaxon
GuardrailsRed Teaming
AI red teaming metrics: measuring failures and coverage
Measure AI red teaming with explicit attack success rates, attempt budgets, coverage, false refusals, severity, and evaluator uncertainty.
Jul 23, 2026
Polyaxon
Red TeamingEvaluation
How to evaluate LLM routers for cost, quality, and latency
Test LLM routing policies with task-level quality, cost, latency, fallbacks, route stability, and model-aware production evidence.
Jul 21, 2026
Polyaxon
LLMOpsEvaluation
Red teaming RAG systems
Test RAG systems for poisoned documents, cross-tenant retrieval, stale permissions, citation leaks, and unauthorized tool actions.
Jul 16, 2026
Polyaxon
Red TeamingRag
How to test prompt injection in LLM applications
Build prompt injection tests for user input, retrieved documents, and tool responses, with checks for data access and actual side effects.
Jul 2, 2026
Polyaxon
Red TeamingEvaluation
Verified data pipelines for reliable AI agents
Design agent data paths that preserve source identity, versions, permissions, freshness, retrieval evidence, and verification results.
Jun 26, 2026
Polyaxon
AI AgentsDataOps
Offline vs. online evaluation for generative AI
Use offline evaluation for reproducible release decisions and online evaluation for real production behavior, then connect both in one feedback loop.
Jun 25, 2026
Polyaxon
EvaluationLLMOps
Build evaluation leaderboards with Polyaxon joins
Use Polyaxon joins to select comparable evaluation runs and build application-owned leaderboards with explicit cohort, completeness, and artifact evidence.
Jun 18, 2026
Polyaxon
AI AgentsEvaluation
LLM-as-a-judge: Design and validate model-based evaluators
Use LLMs as scalable evaluators without treating them as ground truth: design clear rubrics, calibrate against humans, and monitor bias and drift.
Jun 18, 2026
Polyaxon
EvaluationLLMOps
Turn production traces into regression tests
Turn a failure cluster into a reviewed regression case, preserve the causal tool response, and record baseline and candidate checks with Polyaxon.
Jun 11, 2026
Polyaxon
LLMOpsObservability
How to evaluate RAG systems
Evaluate retrieval and generation separately and end to end with representative datasets, groundedness checks, retrieval metrics, and production feedback.
Jun 4, 2026
Polyaxon
RagEvaluation