Guide · 5 min read
How to evaluate AI memory
A practical way to test whether an AI agent's memory works: five kinds of test question, what to measure on each turn, and how to build a small test set.
“It seems to remember” is not a result. Memory fails quietly: the agent still answers, just without the fact it should have had. You need tests that make the absence visible.
Five kinds of question
Write scripted conversations spread over several sessions, then ask:
- Single fact. Something stated once, several sessions ago. “What am I allergic to?”
- Updated fact. Something that changed. The user said Lisbon, later Porto. Which does the agent use?
- Combined facts. An answer needing two memories from different sessions. “Book somewhere for Ana’s birthday that I can eat at.”
- Time. “What did we decide last week?” needs dates, not only meaning.
- Nothing to recall. A question the memory cannot answer. The right behaviour is to say so, not to invent.
The fifth kind is the one most often left out and the one that catches hallucinated memories.
Measure the two halves separately
A wrong answer has two possible causes, and the fixes differ.
Recall. Was the needed note among those retrieved? If not, the problem is extraction (it was never written) or search (it was not found).
Use. The note was in the prompt, but the answer ignored or contradicted it. That is a prompt or model problem.
Log the retrieved notes for every test question so you can tell which half failed.
What to track
| Measure | What it tells you |
|---|---|
| Accuracy on memory questions | Whether memory works at all |
| Needed note retrieved (yes/no) | Extraction and search quality |
| Irrelevant notes retrieved | Noise that crowds the prompt |
| Tokens of memory per turn | Cost, and whether the budget holds |
| Added latency per turn | Whether users will feel it |
| Store size per user over time | Whether forgetting works |
Always compare with two baselines
- No memory. If the score barely changes, your questions do not depend on memory.
- Full history in the prompt. Often the accuracy ceiling for short histories, at a much higher token cost. Memory should come close to it for far fewer tokens.
Keep it honest
Do not tune on the test set until it passes and then report that number. Hold back a few conversations you never look at while tuning. And when you read published benchmark figures, check what was measured, on which data and against which baseline before you compare them with your own.