Section 6: RAG and Knowledge for Agents
RAG (Retrieval-Augmented Generation) is how agents access knowledge beyond their training data. Instead of relying on what the LLM memorized during training, the agent retrieves relevant documents at query time and includes them in the prompt. This grounds responses in facts, reduces hallucination, and enables the agent to work with private, current, or domain-specific data that the model has never seen.
In Section 2, we identified hallucination and stale knowledge as two core LLM limitations. RAG is the primary architectural solution for both. In Section 5, we covered context engineering and token budgets. RAG is the largest consumer of that context budget. This section covers the mechanics: how to turn documents into searchable vectors, how to retrieve the right chunks, and how to decide whether RAG is even the right approach for your use case.
What We’ll Cover:
Where This Fits in the Customer Support Agent:
Our support agent needs to answer questions about return policies, product details, shipping timelines, and FAQ topics. This information lives in company documents, not in the LLM’s training data. Without RAG, the agent either refuses to answer (“I don’t have that information”) or hallucinates a policy that does not exist. With RAG, the agent retrieves the actual policy document and generates a response grounded in real, current company data.
Embedding: Text to Numbers
An embedding converts a piece of text into a numerical vector, a list of numbers that captures the meaning of the text. The key property: texts with similar meaning produce vectors that are close together in mathematical space. Texts with different meaning produce vectors that are far apart.