C12: Agentic AI Foundations for Interviews

Section 8: Reliability and Failure Handling


SLIDE 55: Section Intro, Reliability and Failure Handling

Section 8: Reliability and Failure Handling

Agents fail. LLMs hallucinate. Tools time out. External APIs return 500 errors. Context windows overflow. Agents enter infinite reasoning loops and consume hundreds of dollars in tokens before anyone notices. The difference between a demo and a production system is not whether failures happen, but how the system behaves when they do.

A demo agent handles the happy path. A production agent handles every path. This section covers the engineering patterns that make agents reliable: retries, fallbacks, loop budgets, graceful degradation, and idempotency.

What We’ll Cover

  1. Retry strategies and fallback routing: what to do when a tool call or LLM call fails
  2. Loop budgets and termination conditions: preventing runaway agents
  3. Graceful degradation: maintaining useful behavior when components are down
  4. Idempotency: ensuring retried actions do not cause duplicate side effects

Where This Shows Up in Agent Systems

Failure Type What Goes Wrong Engineering Pattern Impact If Unhandled
Tool call fails (API timeout) Agent cannot complete the action Retry with backoff, then fallback Agent crashes or hallucinates result
LLM call fails (provider outage) Agent cannot reason at all Fallback to backup model Entire agent goes down
Agent loops (circular reasoning) Agent reasons indefinitely, burns tokens Loop budget, time limit, token limit $50+ in wasted API calls per incident
Partial system outage Knowledge base or one tool is down Graceful degradation Agent gives wrong answers from hallucination
Network retry on write action Same action executes twice Idempotency with request IDs Customer gets refunded twice

Every one of these failures has happened in production agent systems. Every one is preventable with the patterns in this section. These patterns connect directly to the monitoring and alerting infrastructure we will build in Section 10.


SLIDE 56: Retry Strategies and Fallback Routing

When Things Fail, Try Again. Then Try Something Else.

The first response to a transient failure (network timeout, rate limit, temporary server error) is a retry. Not all failures are transient, so retries need structure: how many attempts, how long to wait between them, and what to do when retries are exhausted.