Agents
Build an operating model for AI agents that connects development, tool use, evaluation, tracing, and production feedback.
Start here
The AI agent lifecycle
Explore the development, evaluation, and observation loop.
Continue learning
- Open
How to evaluate AI agents
Define success for outcomes, actions, and efficiency.
- Open
AI agent tracing: How to debug tools, loops, and handoffs
Follow model calls, tool use, loops, and handoffs.
- Open
MCP observability: Monitor tools, resources, and context
Inspect context and operations across MCP clients and servers.
Train agents with Ray and RAGEN
Run a RAGEN experiment with explicit Ray resources and persistent checkpoints.
Verified data pipelines for reliable AI agents
Preserve source versions, access decisions, and evidence across retrieval and agent actions.
Designing a control plane for AI agents
Coordinate versions, evaluation, policy, rollout, and audit evidence across the agent lifecycle.
Design measurable AI agents before you build
Define outcomes, events, evaluators, and budgets before implementation.
How to red team AI agents
Test tool permissions, memory, handoffs, retries, and execution budgets.
MCP security testing: tools, permissions, and untrusted content
Verify MCP authorization and independently observe downstream tool actions.