Section 10: Evaluation and Observability


SLIDE 67: Section Intro, Evaluation and Observability

Section 10: Evaluation and Observability

You cannot improve what you cannot measure, and you cannot debug what you cannot trace. Agents are non-deterministic systems. The same input can produce different outputs depending on tool results, retrieved documents, conversation history, and model behavior. Traditional software testing (assert that output equals expected value) does not fully apply. This section covers how to evaluate agent quality, test systems with variable outputs, trace agent decisions for debugging, and monitor production agents continuously.

What We’ll Cover:

  1. Measuring agent quality across four dimensions
  2. Testing non-deterministic systems
  3. Tracing agent decision paths for debugging
  4. Monitoring and alerting in production
  5. Common evaluation mistakes and how to avoid them

Why Evaluation Is Harder for Agents Than Traditional Software:

Traditional Software Agent Systems
Deterministic: same input produces same output Non-deterministic: output varies across runs
Binary correctness: pass or fail Spectrum of quality: correct, partially correct, wrong tone, right answer but missing context
Test at deploy time, stable until code changes Model behavior can drift without any code change (provider updates, retrieval changes)
Failures are reproducible Failures may not reproduce with the same input
Unit tests cover individual functions Agent behavior emerges from the interaction of LLM, tools, retrieval, and prompts

This is why evaluation for agents requires new approaches. Every technique in this section addresses one of these differences.


SLIDE 68: Measuring Agent Quality

The Four Dimensions of Agent Quality

Every agent should be measured across four dimensions. Optimizing for one at the expense of the others produces a system that is fast but wrong, accurate but expensive, or reliable but slow.

Dimension What It Measures Example Metric Target for Support Agent
Accuracy Did the agent give the correct answer? % of responses graded as correct > 90%
Reliability Does it give correct answers consistently? Variance across repeated runs on same input < 5% variance
Latency How long does each interaction take? p50, p95, p99 response time p50 < 3s, p95 < 5s
Cost How much does each interaction cost? Average cost per conversation < $0.15