Glossary
AI agent memory glossary
The words that come up when a team starts building memory, explained in a sentence or two each.
- Agent
- A program that uses a language model to decide what to do next, call tools and work towards a goal over several steps.
- Chunk
- A piece of a longer document, cut to a size that can be embedded and retrieved on its own.
- Compaction
- Replacing older turns of a conversation with a shorter summary so that the context window does not overflow.
- Context window
- The maximum amount of text, measured in tokens, that a model can read in a single request. Everything the model knows about the current task must fit inside it.
- Decay
- Lowering the weight of a memory as it gets older or goes unused, so that fresh facts are preferred at recall.
- Embedding
- A list of numbers that represents the meaning of a piece of text. Texts with similar meanings have embeddings that are close together.
- Episodic memory
- Memory of events: what happened, when, and in what order. "Last Tuesday the import failed on row 40."
- Extraction
- The step that reads a conversation and pulls out the facts worth keeping as memories.
- Hybrid search
- Retrieval that combines search by meaning with search by keyword, which helps with names, codes and other exact terms.
- Knowledge graph
- A store of entities and the relationships between them, such as "Ana — sister of — Maya". Useful for questions that follow links.
- Long-term memory
- Facts kept in a store outside the model that persist across sessions.
- Memory scope
- What a memory belongs to: a user, a team, a project or a single task. Scope decides who can recall it.
- Procedural memory
- Memory of how to do things: rules, habits and steps an agent has learned to follow.
- Prompt caching
- A provider feature that charges less for the part of a prompt that repeats exactly from one request to the next.
- RAG
- Retrieval-augmented generation. Relevant passages are fetched from a document collection and added to the prompt before the model answers.
- Recall
- Finding the stored memories that matter for the current question and placing them in the prompt.
- Reranking
- A second pass that reorders retrieved items with a more careful model, so the best few are kept.
- Semantic memory
- Memory of facts that are true regardless of when they were learned. "Maya is vegetarian."
- Short-term memory
- The current conversation as held in the context window. Also called working memory.
- Token
- The unit a model reads and writes: a word or a piece of one. Prices and limits are counted in tokens.
- Top-k
- The number of results kept from a search. Top-5 means the five best matches.
- Vector database
- A database built to store embeddings and find the ones nearest to a query quickly.
Pack the doko. Ask it anything.
Nine memories, one question, no account. See which facts an agent would carry into its next answer.