Polyaxon v3 is coming →

LLM evaluation

Evaluate retrieval, generation, and agent behavior with representative cases, calibrated judges, and offline and online feedback.

Start here

How to evaluate RAG systems

Measure retrieval and generation separately, then assess the complete task.

Read the guide

Continue learning

  1. How to evaluate AI agents

    Evaluate task outcomes, tool use, trajectories, latency, and cost.

  2. LLM-as-a-judge: Design and validate model-based evaluators

    Design rubrics and validate model-based scores against human decisions.

  3. Offline vs. online evaluation for generative AI

    Connect reproducible release checks to production feedback.

More on this topic

Compare platforms

Apply the concepts above to a documented platform decision, including where each option fits and when they can coexist.

Put it into practice