Section 5: Context, Memory and State
An LLM has no memory between calls. Every request starts from zero. The model does not remember the customer’s name from two messages ago, does not know they called last week, and cannot recall that they prefer email over phone. If the agent needs to remember anything, across turns within a session, across sessions over weeks, or across customers over months, you must build the memory system yourself.
This is one of the most commonly misunderstood aspects of agent design. Candidates assume the LLM “remembers” the conversation. It does not. It re-reads the entire conversation history on every single call. That re-reading has a hard limit (the context window from Section 2), and managing what fits inside that limit is a core engineering problem.
What We’ll Cover
Where This Shows Up in Agent Systems
| Memory Type | What It Stores | Persistence | Agent Example |
|---|---|---|---|
| Short-term (conversation) | Current session messages, tool results | Single session | “You said earlier you want a refund” |
| Long-term (persistent) | User profile, past interactions, preferences | Across sessions | “Welcome back. Last time we resolved a shipping issue for you” |
| Context engineering | What fits in the window right now | Per LLM call | Deciding which 5 of 20 retrieved chunks to include |
| Scratchpad (reasoning) | Intermediate thinking, eligibility checks | Single decision | Hidden reasoning about whether a refund qualifies |
This section builds directly on the context window and token concepts from Section 2. There, we calculated that the support agent gets roughly 27-89 turns depending on retrieval verbosity. Here, we solve the problem of what to do when that limit is reached, and how to use the available space wisely.
How Short-Term Memory Works
Every LLM API call is stateless. The model receives a list of messages (system prompt, user turns, assistant turns, tool results) and generates the next response. To maintain a conversation, you append each new turn to this list and send the full list on every call. The model does not “remember.” It re-reads everything, every time.