Guide · 6 min read
How to add memory to an AI agent
A step-by-step way to give an AI agent long-term memory: decide what to keep, extract it, store it, recall it within a token budget and keep it up to date.
Memory is a loop around the model, not a setting inside it. You can build a useful first version in a day if you build the parts in this order.
1. Decide what a memory is
Write one sentence that defines what is worth keeping for your product. For a personal assistant: “stable facts and preferences the user stated about themselves”. For a coding assistant: “conventions and decisions for this repository”. Everything later depends on this sentence, because it is the instruction your extractor follows.
Decide the scope at the same time. Does a memory belong to a user, a team, a project? The scope becomes the key every read and write is filtered by.
2. Extract
After a session, send the transcript to a model with your definition and ask for a list of short, self-contained notes. Good notes:
- make sense alone: “Maya is vegetarian”, not “she said yes to that”;
- record what was said, not what the model inferred;
- carry a date and, where it helps, the source message.
Ask for an empty list when nothing qualifies. Most conversations contain nothing worth keeping, and that is fine.
3. Store
Start with the simplest thing that fits your scale.
| Scale | Store |
|---|---|
| Under about 50 notes per scope | A text field or table; send all of it |
| Hundreds to thousands | A table plus an embedding column for search by meaning |
| Linked facts across many entities | Add a graph or structured tables beside the vectors |
4. Recall within a budget
Before each answer, search the scope’s notes with the user’s message and put the best matches in the prompt under a clear heading. Set a token budget for this section and never exceed it. A budget forces ranking, and ranking is what keeps a prompt short as the store grows.
5. Update and forget
When a new note arrives, look up the most similar existing ones. If it repeats one, skip it. If it contradicts one, replace the old note. If the user asks you to forget something, delete it everywhere, including from any summaries built on it.
6. Show it
Give users a page listing what is remembered, with edit and delete. It builds trust, and it is the fastest way to find extraction mistakes.
Then measure
Before you tune anything, write twenty questions whose answers depend on something said in an earlier session and check how many the agent gets right. See how to evaluate AI memory.