All concepts

Memory Architecture

Working memory is in-window and volatile; long-term memory is written out and retrieved — the policy between them is the design.

Agentic Engineering · Intermediate · ~6 min

In plain English

Deciding what the system remembers, for how long, and who it belongs to — this conversation, this person, or the whole organization.

Why it's worth your time

Memory is what makes a product feel personal; blurred memory tiers are how it becomes a privacy incident.

If you remember three things

  • Session, user and organization memory are different things
  • Write on an explicit signal, never on every turn
  • Retrieval applies to memory too — fetch, don't paste everything

Overview

"Agent memory" is not one thing. Short-term memory is whatever sits in the context window: fast, free to read, and gone the moment the run ends. Long-term memory is written outside the model — files, key stores, a vector index, a database — and it is durable, costs a retrieval, and survives the session boundary. The architecture is the policy between them: what crosses the boundary when a session ends, how you get it back (keyword, vector, or a deterministic path), and the write rules that keep the store trustworthy — dedupe, update-over-append, source and timestamp, expiry. An unbounded, un-deduped memory doesn't make an agent smarter; it makes it a confident source of stale wrong answers.

How it works

  1. Short-term vs long-term In-window working set versus a written store. Different cost, different lifetime, different failure modes.
  2. The session boundary Preferences, corrections, commitments and entity resolutions cross it. Raw transcripts and scratch work do not.
  3. Retrieval strategies Keyword for ids and codes, vectors for paraphrase, files for deterministic and auditable lookups.
  4. When to write A correction UPDATES an existing memory rather than appending a contradiction. Dedupe, stamp, expire.
  5. The payoff A cold session retrieves what matters and starts already knowing the user's conventions.

In an interview

I separate working memory from long-term memory. Working memory is the context window: fast, free to read, volatile — it dies with the run. Long-term memory is written outside the model and retrieved, so it costs a lookup but survives. The design work is the policy between them: what crosses the session boundary — stable preferences, corrections, commitments, verified entity resolutions — and what must not, like raw transcripts or re-fetchable values that will go stale. For recall I run hybrid retrieval, because keyword nails ids while vectors handle paraphrase. And writes follow rules: a correction updates the existing memory rather than appending a contradiction, everything is deduped, stamped with source and time, and expired.

Production defaults

Tiers
session (ephemeral) · user (durable, explicit, editable) · org (shared, reviewed)
Write policy
explicit signal required. Auto-remembered noise is worse than no memory
Visibility
the user can see and delete anything stored about them
Decay
TTL or relevance decay, so old preferences don't outlive their truth

What breaks

  • One user's memory surfaced for another — Tier boundaries not enforced at the storage layer. This is an access-control bug.
  • Memory contradicts itself — No update path — you appended instead of replacing. Memory needs edits, not just inserts.

Watch it explained

Memory in AI agents — Google Cloud Tech, 4:34

Related