Section 10: Evaluation and Observability
You cannot improve what you cannot measure, and you cannot debug what you cannot trace. Agents are non-deterministic systems. The same input can produce different outputs depending on tool results, retrieved documents, conversation history, and model behavior. Traditional software testing (assert that output equals expected value) does not fully apply. This section covers how to evaluate agent quality, test systems with variable outputs, trace agent decisions for debugging, and monitor production agents continuously.
What We’ll Cover:
Why Evaluation Is Harder for Agents Than Traditional Software:
| Traditional Software | Agent Systems |
|---|---|
| Deterministic: same input produces same output | Non-deterministic: output varies across runs |
| Binary correctness: pass or fail | Spectrum of quality: correct, partially correct, wrong tone, right answer but missing context |
| Test at deploy time, stable until code changes | Model behavior can drift without any code change (provider updates, retrieval changes) |
| Failures are reproducible | Failures may not reproduce with the same input |
| Unit tests cover individual functions | Agent behavior emerges from the interaction of LLM, tools, retrieval, and prompts |
This is why evaluation for agents requires new approaches. Every technique in this section addresses one of these differences.
The Four Dimensions of Agent Quality
Every agent should be measured across four dimensions. Optimizing for one at the expense of the others produces a system that is fast but wrong, accurate but expensive, or reliable but slow.
| Dimension | What It Measures | Example Metric | Target for Support Agent |
|---|---|---|---|
| Accuracy | Did the agent give the correct answer? | % of responses graded as correct | > 90% |
| Reliability | Does it give correct answers consistently? | Variance across repeated runs on same input | < 5% variance |
| Latency | How long does each interaction take? | p50, p95, p99 response time | p50 < 3s, p95 < 5s |
| Cost | How much does each interaction cost? | Average cost per conversation | < $0.15 |