Working memory is in-window and volatile; long-term memory is written out and retrieved — the policy between them is the design.
Deciding what the system remembers, for how long, and who it belongs to — this conversation, this person, or the whole organization.
Memory is what makes a product feel personal; blurred memory tiers are how it becomes a privacy incident.
"Agent memory" is not one thing. Short-term memory is whatever sits in the context window: fast, free to read, and gone the moment the run ends. Long-term memory is written outside the model — files, key stores, a vector index, a database — and it is durable, costs a retrieval, and survives the session boundary. The architecture is the policy between them: what crosses the boundary when a session ends, how you get it back (keyword, vector, or a deterministic path), and the write rules that keep the store trustworthy — dedupe, update-over-append, source and timestamp, expiry. An unbounded, un-deduped memory doesn't make an agent smarter; it makes it a confident source of stale wrong answers.
I separate working memory from long-term memory. Working memory is the context window: fast, free to read, volatile — it dies with the run. Long-term memory is written outside the model and retrieved, so it costs a lookup but survives. The design work is the policy between them: what crosses the session boundary — stable preferences, corrections, commitments, verified entity resolutions — and what must not, like raw transcripts or re-fetchable values that will go stale. For recall I run hybrid retrieval, because keyword nails ids while vectors handle paraphrase. And writes follow rules: a correction updates the existing memory rather than appending a contradiction, everything is deduped, stamped with source and time, and expired.
Memory in AI agents — Google Cloud Tech, 4:34