Skip to content
BotServBotServ
Agent MemoryMemoryAI AgentShort-term MemoryLong-term MemoryVector Database

Agent Memory: How AI Agents Remember

How AI agents store information: short-term and long-term memory, implementation with vector databases and frameworks.

S

schutzgeist

17 min read
Agent Memory: How AI Agents Remember

Agent Memory: Memory for AI Agents

What This Article Covers

  • What agent memory is and why agents without memory quickly get stuck in loops
  • The different types of memory, from short-term to procedural memory
  • How short-term memory works through the context window and when summarization becomes necessary
  • How long-term memory works with vector databases and embeddings
  • How to implement memory using frameworks like LangGraph, CrewAI, and Mem0

Introduction: Agent Memory Explained

An AI agent without memory is like a colleague who forgets what they just said after every sentence. You explain a task, they start working, and two steps later they’ve forgotten what they’ve already completed. They start over, repeat themselves, or lose the thread. Agent memory, often called memory in English, solves exactly this problem.

This article is written for newcomers who want to understand how agents retain information. You don’t need prior knowledge of databases or vector math. After reading, you’ll know the main types of memory, understand how short-term and long-term memory work technically, and know which frameworks can help you implement them. If you already understand what an agent system is, this article will help you grasp the next building block.

Why Do I Need Agent Memory?

Imagine you’re building an agent that sorts your emails and drafts replies. In step 1, it reads an email from a customer named Ms. Weber. In step 2, it pulls customer data from your CRM. In step 3, it should write a reply that references both the email content and the customer data.

Without memory, the agent is blind in step 3. It’s forgotten what happened in step 1 and what it retrieved in step 2. It would have to repeat both steps or you’d need to pass all the information again. That’s inefficient and error-prone.

With short-term memory, the agent retains all intermediate results from the current task. In step 3, it still knows what happened in steps 1 and 2, and can craft the reply with that context. With long-term memory, it additionally remembers that Ms. Weber has contacted you three times with the same issue, and can address that. Memory is the difference between an agent acting blind and one acting with awareness of context.

Agent Memory at a Glance

Agent memory works similarly to human memory. Imagine you’re cooking a new recipe. Your short-term memory holds which step you’re on and which ingredients you’ve already added. It’s limited because you can’t remember every detail of a long recipe all at once. Your long-term memory, meanwhile, stores that you cooked a similar dish three months ago and the rice took too long then. Next time, you draw on that experience and cook the rice shorter.

An agent works the same way. Short-term memory stores the current context, meaning what’s happened in this session. It’s limited by the model’s context window. Long-term memory stores information across multiple sessions, such as past conversations, learned facts, or successful approaches. It typically uses a vector database where information is stored as embeddings and retrieved as needed.

Who Is This Article For?

This article is for you if you’re building or planning agents and want to understand how memory works. You don’t need to be a database expert. It helps if you have a rough idea of what a language model does and what an AI agent is. If you’re already working with planning and reflection, memory will help you improve these processes across multiple tasks.

If you only want to know whether your agent needs memory, this article is enough. If you want to implement memory, you’ll find further links and code examples in the sections below.

Key Terms Around Agent Memory

TermMeaning
Short-term memoryStores the context of the current session, such as messages already read or data retrieved
Long-term memoryStores information across multiple sessions, such as past conversations or learned facts
Working memoryThe part of short-term memory that is currently being actively processed, comparable to what you’re currently focusing on
Episodic memoryStores concrete events, such as “In the last support request, solution X worked”
Semantic memoryStores facts and knowledge, such as “Customer Y has contract Z”
Vector databaseA database that stores information as vectors and finds similar content through distance calculation
EmbeddingA numeric representation of text that makes semantic similarity measurable
Context windowThe maximum amount of text a model can process at once, measured in tokens
SummarizationCondensing older messages when the context window is full to free up space
RetrievalDeliberately accessing stored information from long-term memory

Types of Agent Memory

Agent memory breaks down into several types that complement each other. None of them suffices alone, but the combination makes an agent truly useful.

Short-term Memory

Short-term memory stores what happens in the current session. It holds the message history, meaning all messages, tool calls, and results that occur during the current task. The agent thereby knows which step is next and what’s already been completed. Short-term memory is limited by the model’s context window. If the history becomes too long, it must be trimmed or summarized.

Long-term Memory

Long-term memory stores information across multiple sessions. It typically uses a vector database where past conversations, learned facts, and outcomes are stored as embeddings. When a new task begins, the agent searches for similar entries and loads them into context. This lets it recall earlier interactions without you having to pass them manually. Long-term memory isn’t limited by the context window but by the storage capacity of the database.

Episodic Memory

Episodic memory stores concrete events, meaning what happened and when. Example: “On August 15th, the customer asked about a billing error, and solution X worked.” It’s closely tied to long-term memory but more specific. It helps the agent learn from past situations and solve similar cases faster. In practice, each completed task stores an entry in episodic memory.

Semantic Memory

Semantic memory stores facts and knowledge independent of when they were learned. Examples include “Customer Y has contract Z” or “The support team is available Monday through Friday.” These facts help the agent categorize requests correctly without needing to ask repeatedly. Semantic memory functions like a knowledge base, but the agent builds and queries it autonomously.

Procedural Memory

Procedural memory stores how things are done, including workflows and processes. For example: “To review a refund, first retrieve the order data, then check the reason for the claim, then obtain approval.” It connects closely with planning. In practice, procedural memory often lives as fixed workflows or prompts that the agent retrieves and adapts as needed.

How Short-Term Memory Works

An agent’s short-term memory is based on the language model’s context window. The context window is the maximum amount of text the model can process at once, measured in tokens. Learn more in the article Context Length. Everything the agent does during a session becomes a message in the history and gets sent to the model again at each new step.

A typical conversation history looks like this:

  1. System Prompt: The agent’s role and rules.
  2. User Message: The original task.
  3. Tool Call: The agent invokes a tool, such as a database query.
  4. Tool Result: The tool returns data.
  5. Agent Response: The agent evaluates the result and plans the next step.
  6. Another Tool Call: Next action.

Each step generates tokens. With a model that has an 8,000 token context window, the history fills up quickly, especially if tool results are lengthy. When the history exceeds the limit, several strategies exist:

  • Truncation: The oldest messages are removed. Simple, but the agent loses early steps from memory.
  • Summarization: The oldest part of the history gets condensed. Ten messages become a brief summary that preserves key points. This is the most common approach.
  • Sliding Window: Only the last N messages remain; everything else is discarded. Similar to truncation, but with a fixed boundary.

In practice, summarization is the best choice because it compresses information rather than throwing it away. Most frameworks provide built-in functions for this.

How Long-Term Memory Works

Long-term memory operates differently from short-term memory. Instead of storing information in the context window, it uses an external database, typically a vector database. Learn more in the article Vector Databases. The process has two phases: storage and retrieval.

Storage

When a conversation or task completes, the agent extracts the most important information. This can happen automatically, for instance through a prompt like “Summarize the key facts from this conversation.” The extracted information is converted into an embedding, which is a numerical representation that captures the semantic meaning of the text. This embedding is stored in the vector database alongside the original text and metadata like timestamp and user ID.

Retrieval

For a new task, the agent converts the current request into an embedding and searches the vector database for similar entries. Similarity is measured through distance calculations, such as cosine similarity. The matching entries are loaded into the context window so the agent can use them for the current task.

Here’s an example: A customer asks about their order status. The agent converts the question into an embedding, searches the vector database for similar entries, and finds a saved note: “Customer Y previously experienced a delay but has already been notified.” The agent loads this information into the context and can address it without the customer having to repeat themselves.

This principle is closely related to Local RAG. The difference lies in scope: RAG retrieves external documents, while agent memory retrieves the agent’s own past experiences.

Implementation with Frameworks

Different frameworks take different approaches to agent memory. The three most common are LangGraph, CrewAI, and Mem0.

LangGraph

LangGraph is a framework from LangChain that models agents as graphs. It offers built-in memory features through checkpointers. A checkpointer saves the agent’s state after each step, allowing the agent to pause and resume later.

from langgraph.checkpoint.memory import MemorySaver
from langgraph.prebuilt import create_react_agent

memory = MemorySaver()
agent = create_react_agent(model, tools, checkpointer=memory)

# First query
config = {"configurable": {"thread_id": "kunde-123"}}
agent.invoke(
    {"messages": [{"role": "user", "content": "Ich heisse Frau Weber."}]},
    config
)

# Second query, agent remembers
agent.invoke(
    {"messages": [{"role": "user", "content": "Wie heisse ich?"}]},
    config
)
# Response: "Frau Weber"

The thread_id groups messages from a single session. The checkpointer stores the history per thread. For long-term memory across multiple sessions, LangGraph additionally uses a vector store, for example through the BaseStore interface.

CrewAI

CrewAI is a framework for multi-agent systems. It provides two types of memory: short-term memory within a task and long-term memory across multiple tasks. Memory is activated when creating the crew.

from crewai import Crew, Agent, Task

agent = Agent(
    role="Support-Agent",
    goal="Kundenanfragen beantworten",
    backstory="Ein erfahrener Support-Mitarbeiter.",
    verbose=True
)

task = Task(
    description="Beantworte die Frage des Kunden.",
    expected_output="Eine hilfreiche Antwort.",
    agent=agent
)

crew = Crew(
    agents=[agent],
    tasks=[task],
    memory=True  # Enables short- and long-term memory
)

result = crew.kickoff()

Setting memory=True automatically activates both memory types in CrewAI. For long-term memory, it internally uses a vector database, defaulting to ChromaDB. You can also integrate your own database.

Mem0

Mem0 is a specialized memory framework that integrates into various agent systems. It focuses exclusively on memory, not planning or tool calling. Mem0 automatically extracts important facts from conversations, stores them, and retrieves them when needed.

from mem0 import Memory

m = Memory()

# Store information
m.add(
    messages=[
        {"role": "user", "content": "Ich bevorzuge Antworten per E-Mail."},
        {"role": "assistant", "content": "Verstanden, ich notiere mir das."}
    ],
    user_id="kunde-123"
)

# Retrieve information
results = m.search(
    query="Wie möchte der Kunde kontaktiert werden?",
    user_id="kunde-123"
)
# results contains the stored entry

Mem0 handles extraction, storage, and retrieval. You don’t need to worry about embeddings or vector databases. It works well if you want to add memory to an existing agent without rebuilding it.

Example: An Agent with Memory

Here’s a step-by-step walkthrough of a customer support agent that remembers past interactions.

Step 1: First interaction. Ms. Weber contacts support asking whether her order has shipped. The agent calls the order tool, sees that shipment is still pending, and responds: “Your order is scheduled to ship tomorrow.” At the same time, the agent stores this interaction in long-term memory: “Ms. Weber asked about shipment status on August 20th; shipment was still pending.”

Step 2: Second interaction, days later. Ms. Weber reaches out again, this time asking whether her delivery has arrived. The agent converts her question into an embedding and searches long-term memory. It finds the August 20th entry. It knows the inquiry concerns the same order without Ms. Weber needing to explain. It calls the order tool, sees the delivery was completed, and responds: “Your delivery was completed on August 22nd. When you last contacted us on August 20th, shipment was still pending.”

Step 3: Third interaction, new issue. Ms. Weber contacts the agent weeks later reporting a defect in the delivered product. The agent searches long-term memory and finds both previous entries. It knows when the order was placed, shipped, and delivered, and can launch the return process directly. It doesn’t need to ask Ms. Weber for her order number because it’s already in memory.

Step 4: Knowledge building. After the return is resolved, the agent stores a new entry in semantic memory: “Ms. Weber has Product X, Order Number Y, and reported a defect on September 10th.” For future inquiries about the same product, the agent can access this information instantly.

Without memory, Ms. Weber would need to repeat her order number, dates, and issue details with each contact. With memory, the agent appears to know the customer, which substantially improves satisfaction.

Common Pitfalls with Agent Memory

Memory is powerful but error-prone. Here are the most frequent issues and how to avoid them.

1. Context window overflow. As history grows, the model truncates or stops processing older messages. The agent loses track of early steps. Solution: Use summarization to compress old messages before hitting the limit. Most frameworks do this automatically.

2. Incorrect or stale memories. Long-term memory stores everything you give it. If a fact becomes outdated, say, an old address, it stays stored and gets retrieved. Solution: Validate stored entries for accuracy and allow the agent to overwrite or delete outdated information.

3. Too many irrelevant matches. When searching long-term memory, the vector database may return entries that are semantically similar but contextually irrelevant. The agent loads unnecessary information into context and gets distracted. Solution: Use filters such as user ID or date, and limit the number of retrieved entries.

4. Privacy and sensitive data. Long-term memory stores conversations and facts that may contain personal information. If this data sits unprotected in a vector database, it becomes a privacy risk. Solution: Anonymize sensitive data before storing, use local databases for confidential information, and document what gets saved.

5. High costs from continuous retrieval. Every search in long-term memory generates embeddings and token costs, especially with cloud models. If the agent retrieves on every step, costs add up. Solution: Retrieve only when necessary, such as at the start of a new task, not at every intermediate step.

6. Inconsistent storage. If the agent sometimes saves and sometimes doesn’t, memory becomes incomplete. The agent remembers some things and forgets others. Solution: Define clear rules for what and when to store, such as after every completed operation or after each user interaction.

7. Complexity in multi-agent systems. In multi-agent systems, you must decide whether each agent has its own memory or if they share one. Separate memories lead to siloed knowledge; a shared memory can become unwieldy. Solution: Establish your agents’ memory strategy early and define how they exchange information.

Hardware, Costs, and Security with Agent Memory

Hardware

Short-term memory requires no additional hardware because it runs within the model’s context window. Long-term memory, however, needs a vector database. For small projects, a local database like ChromaDB or FAISS running on a standard computer is sufficient. For larger datasets or many users, you’ll need more RAM and storage. If you compute embeddings with a local model, that requires additional compute capacity depending on model size.

Costs

Agent memory costs fall into two categories. First, token costs for loading memory contents into the context window. The more entries you retrieve, the more tokens you consume. Second, vector database costs, which cloud providers charge per stored data or per query. Local databases incur only hardware and electricity costs. A rough rule of thumb: short-term memory is cheap, long-term memory gets expensive when you store and retrieve many entries frequently.

Security

Memory stores information that may be sensitive, introducing risks. Key safeguards include:

  • Local storage: For confidential data, use a local vector database so nothing leaves your own network.
  • Anonymization: Remove personal data before storing it in long-term memory.
  • Deletion functions: Enable targeted removal of individual entries, such as on user request.
  • Access control: Restrict who can access memory, especially in multi-user systems.
  • Logging: Record every storage and retrieval operation so you can audit what was saved and retrieved.

Further Reading and Resources on Agent Memory

FAQ: Agent Memory - Common Questions

What’s the difference between short-term and long-term memory in agents?

Short-term memory stores the context of the current session within the model’s context window. It’s limited and disappears when the session ends. Long-term memory stores information in a vector database and persists across multiple sessions. An agent can recall earlier interactions even weeks later.

Does every agent need long-term memory?

No. For one-off tasks completed in a single session, short-term memory is sufficient. Long-term memory makes sense when an agent needs to act with context awareness across multiple sessions, such as with support agents or personal assistants.

What happens when the context window is full?

When the conversation history exceeds the model’s limit, old messages must be removed or condensed. The most common approach is summarization, where the oldest portion is compressed into a brief summary. This preserves the essential information without the context window overflowing.

Which vector database is best for getting started?

ChromaDB and FAISS are good choices for beginners because they run locally and are simple to set up. For larger projects, Qdrant or Weaviate are solid options. Most frameworks like LangGraph and CrewAI support multiple databases and come with default configurations.

Can an agent forget what it has stored?

Yes, if you implement deletion functions or entries become stale. Unlike short-term memory, which disappears automatically, long-term memory persists until you actively delete it. You should define rules for when old entries are removed or updated.

How do I prevent the agent from retrieving irrelevant memories?

Use filters during searches, such as by user ID, date, or topic. Also limit the number of retrieved entries to the most relevant ones. Most vector databases support metadata filtering, which lets you narrow down results precisely.

Is agent memory the same as RAG?

No, though the principles overlap. RAG is about retrieving external documents, such as from a knowledge base. Agent memory is about storing and recalling the agent’s own past experiences. Both use vector databases and embeddings, but the source of information differs.

How much does agent memory cost?

Short-term memory costs only tokens since it runs in the context window. Long-term memory incurs additional costs for the vector database and embedding computation. With local databases, you only pay for hardware and electricity. With cloud providers, you pay per storage size and query.

Can I run agent memory locally?

Yes. With a local vector database like ChromaDB or FAISS and a local embedding model, you can run the entire memory system on your own machine. This is especially important for sensitive data that shouldn’t leave your network.

Which framework is best for memory?

It depends on your use case. LangGraph offers flexible checkpointers for short-term memory and extends well. CrewAI enables memory with a single parameter and is great for quick prototypes. Mem0 specializes in memory and integrates into existing systems without major refactoring.

How should I handle personal data in memory?

Anonymize sensitive data before storing it, and use a local database for confidential information. Also implement deletion functions so users can remove their data. Log every storage and retrieval operation so you can track what was stored.

What’s the difference between episodic and semantic memory?

Episodic memory stores concrete events, such as “On August 15th, the customer asked about a billing error.” Semantic memory stores facts independent of time, such as “Customer Y holds Contract Z.” Both complement each other: episodic memory helps with similar situations, semantic memory with general knowledge about the user.

Sources and Further Reading

  • LangGraph Documentation, Memory and Checkpointer section
  • CrewAI Documentation, Memory section
  • Mem0 Project page and Documentation
  • LangChain Documentation on Memory Concepts
  • Anthropic: Building Effective Agents
  • ChromaDB and FAISS Documentation
Back to Blog
Share:

Related Posts