Polyaxon v3 is coming →

LLM evaluation

Evaluate retrieval, generation, and agent behavior with representative cases, calibrated judges, and offline and online feedback.

4 guides · Suggested reading order

How to evaluate RAG systems

Measure retrieval and generation separately, then assess the complete task.

Read the guide
  1. How to evaluate AI agents

    Evaluate task outcomes, tool use, trajectories, latency, and cost.

  2. LLM-as-a-judge: Design and validate model-based evaluators

    Design rubrics and validate model-based scores against human decisions.

  3. Offline vs. online evaluation for generative AI

    Connect reproducible release checks to production feedback.

Explore practical guides and resources to take the next step.