LLM evaluation
Evaluate retrieval, generation, and agent behavior with representative cases, calibrated judges, and offline and online feedback.
Start here
How to evaluate RAG systems
Measure retrieval and generation separately, then assess the complete task.
Continue learning
- Open
How to evaluate AI agents
Evaluate task outcomes, tool use, trajectories, latency, and cost.
- Open
LLM-as-a-judge: Design and validate model-based evaluators
Design rubrics and validate model-based scores against human decisions.
- Open
Offline vs. online evaluation for generative AI
Connect reproducible release checks to production feedback.
Run batch LLM evaluations on Kubernetes
Scale evaluations with stable shards, bounded concurrency, and complete reports.
Run Promptfoo evaluations on Kubernetes with Polyaxon
Run a Promptfoo suite as a tracked job and verify its failure gate.
How to evaluate LLM guardrails
Compare protection, false refusals, legitimate task success, latency, and cost.
Red teaming RAG systems
Test poisoned context, tenant boundaries, permission changes, and cached answers.
Compare platforms
Apply the concepts above to a documented platform decision, including where each option fits and when they can coexist.