C12: Agentic AI Foundations for Interviews

Section 5: Context, Memory and State


SLIDE 31: Section Intro, Context Memory and State

Section 5: Context, Memory and State

An LLM has no memory between calls. Every request starts from zero. The model does not remember the customer’s name from two messages ago, does not know they called last week, and cannot recall that they prefer email over phone. If the agent needs to remember anything, across turns within a session, across sessions over weeks, or across customers over months, you must build the memory system yourself.

This is one of the most commonly misunderstood aspects of agent design. Candidates assume the LLM “remembers” the conversation. It does not. It re-reads the entire conversation history on every single call. That re-reading has a hard limit (the context window from Section 2), and managing what fits inside that limit is a core engineering problem.

What We’ll Cover

  1. Short-term memory: conversation history, sliding windows, and summarization
  2. Long-term memory: persistent storage, retrieval, and session continuity
  3. Context engineering: token budgets, compression, and relevance filtering
  4. Reasoning scratchpads: hidden intermediate reasoning for complex decisions

Where This Shows Up in Agent Systems

Memory Type What It Stores Persistence Agent Example
Short-term (conversation) Current session messages, tool results Single session “You said earlier you want a refund”
Long-term (persistent) User profile, past interactions, preferences Across sessions “Welcome back. Last time we resolved a shipping issue for you”
Context engineering What fits in the window right now Per LLM call Deciding which 5 of 20 retrieved chunks to include
Scratchpad (reasoning) Intermediate thinking, eligibility checks Single decision Hidden reasoning about whether a refund qualifies

This section builds directly on the context window and token concepts from Section 2. There, we calculated that the support agent gets roughly 27-89 turns depending on retrieval verbosity. Here, we solve the problem of what to do when that limit is reached, and how to use the available space wisely.


SLIDE 32: Short-Term Memory and Conversation Context

How Short-Term Memory Works

Every LLM API call is stateless. The model receives a list of messages (system prompt, user turns, assistant turns, tool results) and generates the next response. To maintain a conversation, you append each new turn to this list and send the full list on every call. The model does not “remember.” It re-reads everything, every time.