Evaluation articles
Browse Polyaxon articles about Evaluation. Page 2 of 2.

Schedule agent evaluations and backfill missed days
Run recurring agent evaluations and controlled historical backfills with Polyaxon schedules, date-range matrices, explicit window identity, and fresh results.
Apr 16, 2026
Polyaxon
AI AgentsEvaluation
Build an LLM pipeline to classify ML failure reports
Use Polyaxon DAGs, tracked outputs, artifacts, and run comparison to develop a reviewable classifier for ML workload failure reports.
Apr 9, 2026
Polyaxon
LLMOpsDataOps
Build and evaluate LLM function calling
Connect bounded LLM tools to Polyaxon sandboxes, validate calls in the host, retain execution receipts, and compare tool behavior in repeatable evaluation runs.
Mar 20, 2026
Polyaxon
AI AgentsTools
Build and evaluate AI agent guardrails in Polyaxon
Use Polyaxon to version guardrail experiments, compare case-level outcomes, constrain execution, and require evidence before promoting an agent release.
Feb 20, 2026
Polyaxon
AI AgentsSecurity
Cache pipeline steps without hiding LLM regressions
Cache deterministic Polyaxon preparation steps while forcing fresh LLM evaluations, using explicit dependency identities and separate cache policies.
Feb 12, 2026
Polyaxon
AI AgentsEvaluation
Evaluate zero-shot prompting for production tasks
Use Polyaxon runs, versioned components, grid searches, case artifacts, and comparison views to qualify zero-shot prompts for production ML workflows.
Jan 20, 2026
Polyaxon
PromptingLLMOps
Stop expensive agent evaluation sweeps early
Use Polyaxon failure and metric early-stopping rules to bound agent evaluation sweeps while preserving complete evidence and avoiding premature quality decisions.
Dec 18, 2025
Polyaxon
AI AgentsEvaluation
Tune RAG retrieval with Polyaxon experiment matrices
Compare RAG chunk sizes, retrieval depth, and model choices with Polyaxon experiment matrices, fixed evaluation inputs, and quality-aware run comparisons.
Nov 13, 2025
Polyaxon
AI AgentsEvaluation
Optimize LLM performance and cost with controlled experiments
Track usage and pricing assumptions in Polyaxon, compare cost against quality in run dashboards, and control resources and concurrency during LLM experiments.
Aug 14, 2025
Polyaxon
LLMOpsEvaluation