The most useful thing published in the last week wasn't a model. It was a convergence in how practitioners describe a split many teams have been living with silently: agent memory and RAG are not the same system, and building one as a thin wrapper on the other is where production agents rot.
Three pieces landed within days of each other. MemTensor published "Agent Memory Is Not RAG: A 2026 Production Field Guide" on the Hugging Face blog, cross-posted to DEV. Supermemory put out "RAG vs Agent Memory: What Each Does and When to Combine Them." And on arXiv, "Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation" argues the retrieval step itself needs different machinery when the corpus is the agent's own history rather than a document store.
Why the distinction is architectural, not semantic
RAG has one write path and it's offline: you chunk, embed, index, and the corpus is assumed correct. Memory has a write path that runs at inference time, under uncertainty, with conflicting facts arriving from the same user across sessions. That single difference cascades.
Deduplication becomes a runtime problem. If a user says "we moved off Postgres" in week three, a document index would happily return both the old and new statement with similar cosine scores. A memory system has to decide whether the new fact supersedes, qualifies, or contradicts the old one — and that decision needs provenance, timestamps, and a policy, not a reranker.
Recall targets invert. RAG wants high recall on a fixed corpus: miss the right chunk and the answer is wrong. Memory wants aggressive forgetting. Every stale preference you retrieve is context budget spent on something that will actively mislead the model. The arXiv framing — decoupling retrieval from aggregation — matters here precisely because top-k over episodic history returns fragments that need to be composed into a current-state view, not concatenated.
Evaluation diverges completely. You can build a golden set for RAG. For memory, correctness is time-indexed: the right answer in March is wrong in September, and your eval harness has to model that or it will report 90% on a system users find broken.
What it means for what you're building
If you have one vector store serving both your document corpus and your agent's user history, this is the week to split them. They need different TTLs, different write policies, different eval loops, and probably different storage — memory's access pattern is recency- and entity-skewed in ways an ANN index over uniform chunks handles badly.
The product side is already ahead of most internal stacks. Anthropic shipped unified memory across Claude chat and Cowork in late August, on by default, with controls for sensitive topics. Cross-surface persistent memory is becoming table stakes, which means your users will expect it from your internal tools too — and will notice when yours contradicts itself.

