Section 2: LLM Foundations for Agents
Agents run on LLMs. Every agent you build, every tool call it makes, every reasoning step it takes is powered by an LLM generating tokens. If you do not understand how LLMs work, where they fail, and what knobs you can turn, you will build unreliable agents and you will not be able to debug them when they break.
This section covers LLMs at interview depth, not research depth. You need to explain transformer architecture, tokenization, and inference clearly in 2-3 minutes during an interview. You do not need to derive the attention equation or explain backpropagation.
What We’ll Cover
Where This Shows Up in Agent Systems
| Agent Component | LLM Foundation It Depends On |
|---|---|
| Core reasoning loop (Section 3) | Transformer inference, next-token prediction |
| Tool calling (Section 4) | Structured output generation, low temperature |
| RAG and retrieval (Section 6) | Context window limits, token budgets |
| Memory systems (Section 5) | Context window overflow, conversation history |
| Cost and latency (Section 11) | Token pricing, model selection, inference speed |
| Reliability (Section 8) | Hallucination, non-determinism, failure modes |
Every section in this course builds on these foundations. When something goes wrong with your agent, the root cause often traces back to one of the LLM behaviors covered here.
What a Transformer Does
A transformer takes a sequence of tokens as input and predicts the next token. It does this repeatedly, generating one token at a time, until it produces a complete response. That is the entire mechanism behind ChatGPT, Claude, Gemini, and every other LLM powering today’s agents.
The Three Concepts You Need