Agent Memory
How AI agents store and retrieve memory — short-term vs long-term, semantic vs episodic vs procedural, and the storage systems behind them.
Memory is crucial for AI agents as it allows them to remember previous interactions, learn from feedback and adapt to user preferences. There are two types of memory, split into short-term memory and long-term memory.
Short-term memory (working memory): This tracks the ongoing conversation by maintaining message history within a single session, allowing the application to remember previous interactions.
Long-term memory: Split into three different types – episodic memory, procedural memory and semantic memory. This allows applications to retain information across multiple sessions or different conversations.

Semantic Memory
This type of memory involves the retention of specific facts and concepts and is often used to personalize applications by remembering facts/concepts from past interactions. These memories can be managed in different ways and there are trade-offs to both:
1. Profile
Semantic memories can be a single unified profile of well-scoped and specific information about a user, organization or any entity. This is usually a JSON document with various key-value pairs that are selected to represent the domain. You will need to constantly update this profile as new information is collected or passing in the old profile and asking the model to generate a new profile.
However, this can become error prone as the profile gets larger and may benefit from splitting the profile into multiple documents.

2. Collection
Another way of managing semantic memory is through a collection of documents that are continuously updated and extended over time. Each individual memory can be more narrowly scoped and easier to generate which leads to a lower chance of loss of information over time. Document collection leads to higher recall as a result.
However, this adds complexity to memory updating as the model must delete or update existing items in the list and some models might default to over-inserting or over-updating. Using a collection also makes it challenging to provide comprehensive context to the model. The structure of individual memories might not capture the full context/relationship between memories, causing the model to lack important contextual information which might be more readily available in a unified profile.
Episodic Memory
Episodic memory refers to recalling past events/actions. Episodic memories are usually implemented through few-shot examples to help agents learn from past sequences to perform tasks correctly. Few-shot prompting allows you to ‘program’ your LLM by updating the prompt with input-output examples to illustrate the intended behavior. The challenge lies in selecting the most relevant examples based on user input.
Procedural Memory
Procedural memory, in both humans and AI agents, involves remembering the rules used to perform tasks. In humans, procedural memory is like the internalized knowledge of how to perform tasks, such as riding a bike via basic motor skills and balance. For AI agents, procedural memory is a combination of model weights, agent code, and agent’s prompt that collectively determine the agent’s functionality.
Every memory system runs on four core operations:
- ADD: Store a completely new fact
- UPDATE: Modify an existing memory when new information complements or corrects it
- DELETE: Remove a memory when new information contradicts it
- SKIP: Do nothing when information is a repeat or irrelevant
Modern memory systems delegate these decisions to the LLM itself rather than using brittle if/else logic. The extraction phase ingests context sources (the latest exchange, a rolling summary, recent messages) and uses the LLM to extract candidate memories. The update phase compares each new fact against the most similar entries in the vector database, using conflict detection to determine whether to add, merge, update, or skip.
Writing Memories
2 primary methods for agents to write memories: ‘in the hot path’ and ‘in the background’.

In the Hot Path
Creating memories during runtime has its own advantages and disadvantages. On one hand, this allows for real time updates, making new memories immediately available for use in subsequent interactions.
However, this might increase complexity if the agent requires a new tool to decide what to commit to memory. The process of reasoning about what to save to memory can impact agent latency. Finally, because the agent has to multitask between memory creation and other tasks, this might affect the quality and quantity of memories created.
In the Background
Creating memories as a separate background task offers several advantages. It eliminates latency in primary applications and separates application logic from memory management, allowing the agent to have more focused task completion.
However, determining the frequency of updates becomes crucial as infrequent updates might cause new conversation sessions to have missing contextual information.
Storage options for memory

Most systems today use a combination of all three that work together in a single database.
1. Vector stores for semantic memory
Taking text and converting them into vector embeddings before storing in vector databases. Retrieval works through vector search, against vectors that are indexed using HNSW (Hierarchical Navigable Small World, an algorithm used for ANN). Vector search captures semantic similarity well, but misses out on structural relationships, which is where knowledge graphs come in.
2. Knowledge graphs for relationship memory
Vector search can tell you that a user mentioned coffee but is unable to tell that they prefer a specific coffeeshop, or what their order preferences are. Knowledge graphs store facts as entities and relationships, with edges capturing how they connect. Adding bi-temporal modelling (tracking both when events happened and when the system learned about them) on top of knowledge graphs allows you to see not just what you know, but also what you know at any point in time.
3. Structured databases for factual memory
Relational databases store the structured data: user profiles, access controls, session metadata, audit logs. Vectors deliver high-recall semantic candidates (what feels similar), while graphs provide the structure to trace relationships across entities and time (how things relate). Relational tables anchor both with the transactional guarantees that production systems demand.
Forgetting
Effective forgetting can be implemented with decay functions applied to vector relevance scores: by analyzing the results of vector search, old and unreferenced embeddings naturally fade from the agent’s attention, imitating biological human memory decay patterns. In a database, this is straightforward. A recency-weighted scoring function multiplies semantic similarity by an exponential decay factor based on time since last access. Memories that have not been recalled recently gradually lose relevance.
Things to take note of
1. Security & Isolation
An agent’s long-term memory should not be shared across different users but should have unique long-term memories for each user/team/company. Without strict boundaries, an attacker could feed harmful or misleading data into an agent’s memory (memory poisoning), causing the AI to make bad decisions or leak data later.
2. Multi-tenancy
Client data must be kept separate, as simple vector databases often rely on basic grouping (namespace) which falls short of compliance standards in industries like finance and healthcare.