LLM evaluation
Evaluate retrieval, generation, and agent behavior with representative cases, calibrated judges, and offline and online feedback.
4 guides · Suggested reading order
Start here
How to evaluate RAG systems
Measure retrieval and generation separately, then assess the complete task.
Continue the learning path
How to evaluate AI agents
Evaluate task outcomes, tool use, trajectories, latency, and cost.
LLM-as-a-judge: Design and validate model-based evaluators
Design rubrics and validate model-based scores against human decisions.
Offline vs. online evaluation for generative AI
Connect reproducible release checks to production feedback.
Go deeper
Run batch LLM evaluations on Kubernetes
Scale evaluations with stable shards, bounded concurrency, and complete reports.
Run Promptfoo evaluations on Kubernetes with Polyaxon
Run a Promptfoo suite as a tracked job and verify its failure gate.
How to evaluate LLM guardrails
Compare protection, false refusals, legitimate task success, latency, and cost.
Red teaming RAG systems
Test poisoned context, tenant boundaries, permission changes, and cached answers.
Put the foundations into practice
Explore practical guides and resources to take the next step.