A vendor-neutral taxonomy of agent memory: short-term vs. long-term, episodic vs. semantic vs. procedural, working memory vs. retrieval memory, and the failure modes -- memory bloat, stale memory, retrieval drift -- that kill production agents.
The context window is not memory. This distinction matters more than most agent tutorials admit. The context window is what the model can attend to right now -- a finite buffer that is flushed when the session ends. Memory is what persists: summaries of past sessions, retrieved facts, learned preferences, cached tool results. Conflating the two leads to architectures that either stuff too much into context (performance degrades) or throw away everything between turns (agents that cannot learn). The first design decision in any agent memory system is clarifying which layer a piece of information belongs in.
The useful taxonomy has two axes. The first axis is duration: short-term (what survives within a session) versus long-term (what survives across sessions). The second axis is content type: episodic (what happened), semantic (what is generally true), and procedural (how to do things). Most agent systems conflate these -- they have a single "memory store" and dump everything into it. Separating them allows you to apply different retention policies, retrieval strategies, and expiration rules to each type. A recalled fact about the world (semantic) should be stored and retrieved differently from a record of a past user interaction (episodic) or a proven prompt template (procedural).
Short-term memory is your context engineering problem. What goes into the context window at the start of each turn? In what order? At what compression ratio? Most agents under-invest here. The context window is not an inbox -- you do not append everything relevant. You curate: the most recent N turns verbatim, a compressed summary of earlier turns, retrieved facts from long-term memory with source tags, and working state from mid-task tool calls. The assembly of the context window is a data pipeline that runs on every request. Build it deliberately, measure what survives in each slot, and treat the composition as a first-class engineering artifact.
Working memory is the subset of context that tracks the agent's current task state: what step it is on, what it has already tried, what it is waiting for, what constraints it cannot violate. In a ReAct loop, this is the Thought-Action-Observation chain that builds up during a multi-step task. In a graph-based framework like LangGraph, it is explicit state on each graph node. The common failure mode is losing working memory mid-task -- a context overflow partway through a long task causes the agent to forget that it already tried approach A and loop back to it. Always design for explicit working memory checkpointing, especially for tasks that run longer than a few turns.
Long-term memory is where most engineering effort belongs, and where most projects underinvest. The three content types map to different engineering patterns. Episodic memory -- logs of specific past events, user interactions, agent decisions -- is an append-only event store: write on every relevant interaction, retrieve with recency and similarity weighting. Semantic memory -- world knowledge, user preferences, domain facts the agent has accumulated -- is a key-value or vector store where you need both retrieval and invalidation, because facts go stale. Procedural memory -- learned tool use patterns, successful prompt templates for recurring task types -- is the hardest: it is effectively fine-tuning territory for most tasks, though few-shot retrieval of past successful traces approximates it without retraining.
Agent memory is related to RAG but it is not RAG. The plumbing overlaps -- both use embeddings, vector search, and chunk retrieval -- but the purpose is different. RAG retrieves from a curated external knowledge base to ground factual claims. Agent memory retrieves from the agent's own accumulated experience to maintain continuity and avoid re-deriving known conclusions. The ownership model differs too: RAG documents are curated by you; agent memories are accumulated dynamically by the agent and need different quality controls. Stale RAG documents are a curation failure you can fix by updating the source. Stale agent memories are a system design failure -- you built a memory that does not expire. Expiration and revision must be in your memory schema from day one.
Memory bloat is the first failure mode. Agents that write eagerly accumulate irrelevant, redundant, and contradictory memories. Retrieval quality degrades as the signal-to-noise ratio in the memory store falls. Apply the same discipline you apply to data warehouse design: every write should satisfy a schema, only persist what you will actually retrieve, and prune on a schedule. A memory entry that has never been retrieved in 30 days is probably not worth keeping. A memory entry that directly contradicts a newer entry should be superseded, not kept alongside it.
Stale memory and retrieval drift are the second and third failure modes, and they are harder to detect. Stale memory occurs when a fact that was true when stored is no longer true -- user preferences change, external facts change, prior conclusions become outdated. Without explicit invalidation logic, agents reason from stale premises with full confidence. Retrieval drift occurs when the embedding model or retrieval strategy changes, causing the same query to return different memories than before -- effective behavior changes silently because the retrieval layer changed. Both require the same treatment: version your memories, timestamp them, and treat memory quality as a measurable system property with its own eval harness. The teams that get agent memory right are the ones that treat it as a data product with SLAs, not as an implementation detail.
The best conceptual map of the agent memory design space. Weng's memory taxonomy -- sensory, short-term, long-term -- borrows from cognitive science and maps cleanly onto the engineering choices you actually face. Required reading before designing any memory layer.
The paper that framed the agent context window as "main memory" and external storage as "disk" -- the most productive mental model for thinking about memory hierarchy in LLM agents. Read this before designing any long-term memory system.
The most comprehensive recent survey of agent memory designs in production systems. Covers the full taxonomy from in-context to external, with analysis of retrieval strategies and freshness mechanisms.
Maps agent components -- including memory -- onto established cognitive science frameworks. Clarifies the episodic/semantic/procedural distinction with precision and gives you vocabulary for reasoning about what an agent should retain across task boundaries.
New resources and perspective on building AI-ready data systems, a few times a month. No spam.