Skip to content
    Agentic AI

    AI Agent Memory: What to Keep and What to Forget

    JK
    7 min read

    TL;DR

    1

    2

    3

    4

    5

    If you are the founder, you already know the cost of an assistant that forgets. You explain your pricing rule on Monday and explain it again on Thursday. AI agent memory is what an agent keeps between sessions and pulls back when it needs it. It is not the context window, which is only what the agent can see right now. It is not RAG either, which looks up documents on demand. You need memory when an agent has to remember facts, decisions or preferences past a single conversation. If a fresh context window does the job, skip it.

    Why AI agent memory matters more than the model

    Here is the thing nobody tells you. An AI model does not get smarter on its own. The weights are frozen, the same every session. The "self-improving agent" pitch is marketing.

    What compounds is the memory and context you give it. Structure beats raw intelligence.

    Most AI setups reset every time you open them. That is why the agent forgets your pricing rule and treats a ten-year client like a stranger. A bigger model will not fix that. Better memory will.

    The four memory types

    The types differ in what they hold and how long they last.

    TypeWhat it holdsHow long it lasts
    WorkingWhat the agent needs for the task in handUntil the session ends
    EpisodicSpecific past events, such as a refund request on 3 MarchLong term, with a date
    SemanticFacts and preferences, such as a client who prefers emailLong term, until it changes
    ProceduralLearned steps for doing a taskLong term, with versions

    You may already use procedural memory without naming it. A saved set of writing rules is the simplest example. You write the instructions once and load them into ChatGPT or Claude as a custom style. BabyLoveGrowth's guide on how to humanise AI text with instructions has a template for it.

    One more split matters. Some memory is baked into the model through training. The rest sits outside it, in a store you can read, fix and delete. Use the second kind for anything a client sees, because you can prove what the agent knows.

    For a longer primer, Oracle's beginner's guide to agent memory covers the basics.

    How memory works: capture, store, retrieve, consolidate

    Memory runs on a loop with four steps.

    1. Capture. Log the fact or the decision, along with where it came from.
    2. Store. Index it so it can be found later.
    3. Retrieve. Pull back the right memories, ranked by how useful they are now.
    4. Consolidate. Decide what to keep, what to summarise and what to bin.

    Capture is where trust starts. If you do not know where a memory came from, you cannot judge how far to trust it. Store the source and the date next to the content.

    Retrieval is more than finding the closest match. Good systems rank on a mix of signals: how recent the memory is, who it is about, and how close it is to the question. A note from yesterday about this client beats a similar note from six months ago about someone else.

    Consolidation is the hard step. Summaries save space and lose detail. Full records keep detail and grow fast. A sound rule: a memory only becomes permanent after it has proved useful more than once. One good result is not a rule.

    Which architecture fits which job

    PatternGood atWatch for
    Vector storeFinding things with a similar meaningWeak on exact relationships. Needs scoping by user
    Graph memoryQuestions about how things connectMore work to build and keep up
    Layered short and long termJobs where recency mattersRules for what moves between layers
    Zero-token pipelineHigh volume, tight budgetsLess flexible than model-written summaries

    Zero-Mem asks whether memory needs extra model calls at all. It keeps the original interaction traces as the record and organises them two ways: a graph of entities and context, and a timeline. With the same reader and context budget, it cut the time spent on memory operations by 57.6% against the fastest baseline it tested.

    AdMem uses three agents: an actor, a memory agent and a critic. Between them they write memories, score them and retrieve them. Long-term memory is then merged and pruned based on those scores. The authors report that it holds up better, and succeeds more often, on long multi-turn tasks than the baselines they compared.

    DevWiz covers more patterns in its write-up of agent memory design frameworks, including how to let old entries fade.

    No pattern wins everywhere. Pick on what your agent has to remember, not on what sounds impressive.

    A practical checklist

    Start small. Do not stand up a vector database on day one.

    PingCAP makes the case for SQLite as a local-first memory store: no server, one portable file, and extensions for full-text and vector search. It fits a single user on one machine. Move to something bigger when you serve many users at once.

    Whatever you store it in, each memory needs these fields:

    • user_id. Who the memory belongs to.
    • memory_type. Working, episodic, semantic or procedural.
    • written_at. When it was captured.
    • last_confirmed_at. When someone last checked it was still true.
    • validity_basis. Why you believe it is still true.

    Then add three lifecycle rules:

    1. Expiry. Anything that goes stale on its own gets a time limit.
    2. Decay. Memories nobody uses drop down the ranking.
    3. Invalidation. When a fact changes, such as a cancelled subscription, the old memory dies at once.

    If you want a head start, Mem0 is open source on GitHub.

    What the research measured

    The Mem0 paper tested a memory layer against feeding an agent its full history every time. It reports 91% lower p95 latency and more than 90% token cost savings. It also reports a 26% relative improvement over OpenAI's memory on its judged accuracy metric.

    P95 latency is the response time that 95 out of 100 requests beat. It tells you how slow the slow ones are, which is what your clients feel.

    To test your own system, measure four things:

    • Token cost per task, with memory and without.
    • P95 latency on any reply that depends on memory.
    • Retrieval accuracy against a small test set you labelled by hand.
    • Staleness. How often the agent serves a fact that is out of date.

    Run these on a schedule, not once. Our guide to AI agent testing covers how to build the checks into your release process.

    Where memory goes wrong

    Four mistakes to watch for.

    • Promoting an untested procedure. One good run does not prove a method is safe to repeat.
    • No scoping. Without a user and tenant on every memory, one client's details can surface in another client's reply.
    • No expiry. A stale memory is often worse than no memory.
    • Hidden memory. Show people what the agent remembers and how fresh it is. They will catch errors before you do.

    The part most teams get wrong

    Most teams treat memory as a storage problem. Pick a database, index everything, done.

    The hard part is deciding what deserves to survive. A system with perfect retrieval and no rules will serve stale or wrong facts with full confidence, because nothing tells it to stop.

    In our own stack we keep two clean stores. One holds how things connect. We call it the map. The other holds what happened. We call it the diary. Keeping them apart means "how does this work" and "what did we decide" never get mixed up. The map is a wiki the agents read before they act. Every meeting, call and work session feeds it, so the next answer starts from what we already know.

    That is what an AI Operating System does that a loose set of tools cannot. Your IP lives in the memory layer, and AI employees built with Claude Code read it before they act. So the work scales without you in every loop. It is what we build with founders in the 90-day program.

    If you are starting today, do not pick the storage engine first. Write the rules for what gets kept and what gets dropped. Then decide what infrastructure you need. Give your agents less memory than you think they need, and more rules. Our guide to building an LLM wiki for your AI stack shows how the map is built.

    James Killick

    Find out what your agents need to remember

    Memory is only worth building around the decisions that still run through you. Start by finding out which ones those are.

    Take the Founder Bottleneck Assessment. It scores five dimensions in six minutes, then names your next move.

    Sources

    Frequently Asked Questions

    JK

    James Killick

    Founder

    The AI Orchestrator. 10+ years building digital products and 200+ apps shipped, now helping $1M+ educators and consultants turn their IP into AI-powered delivery systems.

    James Killick founded and runs The AI Orchestrators.

    Ready to find out where your biggest AI opportunity is?

    Take the assessment. It takes about 5 minutes. You'll get a clear picture of how ready your business is.