Section 9: Security and Guardrails


SLIDE 61: Section Intro, Security and Guardrails

Section 9: Security and Guardrails

Agents take actions in the real world. They process refunds, look up customer data, modify account settings, and interact with production databases. A compromised or poorly guarded agent is not just a bad chatbot. It is a system that can leak sensitive data, process unauthorized transactions, or be manipulated into behaviors its designers never intended. Security is not a feature you add after launch. It is a design constraint from day one.

What We’ll Cover:

  1. Prompt injection and input validation
  2. Output filtering and content safety
  3. Authorization and action boundaries
  4. Human-in-the-loop for high-risk actions

Where Security Fits in the Agent Architecture:

Component Security Concern Defense
User input Prompt injection overrides agent instructions Input sanitization, instruction hierarchy
LLM reasoning Model follows malicious instructions from retrieved content Indirect injection detection, trust boundaries
Tool execution Agent performs unauthorized actions (excessive refunds, data access) Authorization policies, action boundaries
Agent output Response leaks system prompt, PII, or internal data Output filtering, PII redaction
Human handoff High-risk actions executed without oversight Human-in-the-loop approval gates

Every layer of the agent architecture from Section 3 has a corresponding security surface. This section teaches you to defend each one.


SLIDE 62: Prompt Injection and Input Validation

What Is Prompt Injection?

Prompt injection is an attack where a user crafts input that overrides or manipulates the agent’s system prompt. The LLM treats all text in its context window as instructions. If a user’s message contains something that looks like a system-level instruction, the model may follow it.

This is the most discussed security risk in LLM-based systems, and it is the one interviewers ask about most frequently.

Two Types of Prompt Injection: