Intlo BrainIntlo Brain

September 24, 2026

Agent memory stops pretending to be RAG

A cluster of field guides and a new arXiv paper this week push the same claim — memory and retrieval are separate subsystems.

The most useful thing published in the last week wasn't a model. It was a convergence in how practitioners describe a split many teams have been living with silently: agent memory and RAG are not the same system, and building one as a thin wrapper on the other is where production agents rot.

Three pieces landed within days of each other. MemTensor published "Agent Memory Is Not RAG: A 2026 Production Field Guide" on the Hugging Face blog, cross-posted to DEV. Supermemory put out "RAG vs Agent Memory: What Each Does and When to Combine Them." And on arXiv, "Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation" argues the retrieval step itself needs different machinery when the corpus is the agent's own history rather than a document store.

Why the distinction is architectural, not semantic

RAG has one write path and it's offline: you chunk, embed, index, and the corpus is assumed correct. Memory has a write path that runs at inference time, under uncertainty, with conflicting facts arriving from the same user across sessions. That single difference cascades.

Deduplication becomes a runtime problem. If a user says "we moved off Postgres" in week three, a document index would happily return both the old and new statement with similar cosine scores. A memory system has to decide whether the new fact supersedes, qualifies, or contradicts the old one — and that decision needs provenance, timestamps, and a policy, not a reranker.

Recall targets invert. RAG wants high recall on a fixed corpus: miss the right chunk and the answer is wrong. Memory wants aggressive forgetting. Every stale preference you retrieve is context budget spent on something that will actively mislead the model. The arXiv framing — decoupling retrieval from aggregation — matters here precisely because top-k over episodic history returns fragments that need to be composed into a current-state view, not concatenated.

Evaluation diverges completely. You can build a golden set for RAG. For memory, correctness is time-indexed: the right answer in March is wrong in September, and your eval harness has to model that or it will report 90% on a system users find broken.

What it means for what you're building

If you have one vector store serving both your document corpus and your agent's user history, this is the week to split them. They need different TTLs, different write policies, different eval loops, and probably different storage — memory's access pattern is recency- and entity-skewed in ways an ANN index over uniform chunks handles badly.

The product side is already ahead of most internal stacks. Anthropic shipped unified memory across Claude chat and Cowork in late August, on by default, with controls for sensitive topics. Cross-surface persistent memory is becoming table stakes, which means your users will expect it from your internal tools too — and will notice when yours contradicts itself.

Sources

  1. [1] RAG vs Agent Memory: What Each Does and When to Combine Them — supermemory
  2. [2] What's the Difference Between RAG and Agent Memory? - DEV Community
  3. [3] Agent Memory Is Not RAG: A 2026 Production Field Guide
  4. [4] The Agent Memory Wars Are Here - AgentConn Blog
  5. [5] Agent Memory Is Not RAG: A 2026 Production Field Guide - DEV Community
  6. [6] [2602.02007] Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
  7. [7] RAG is Dead. Long Live Agent Memory - Chris Latimer - YouTube
  8. [8] State of AI Agent Memory 2026: Benchmarks & Trends ...
  9. [9] Knowledge and Memory Beyond RAG: Why 2026 Agents Need a Write Path, Not Just a Retriever | by Micheal Lanham | Apr, 2026 | Medium
  10. [10] Enterprise Search Is Entering a New Era — Activant
  11. [11] AI-Enabled Enterprise Search - SLAC IT - Stanford University
  12. [12] The definitive guide to AI‑based enterprise search for 2026
  13. [13] Enterprise Search in 2025: How GoSearch Redefined AI-Powered Work | The GoSearch Blog
  14. [14] AI Enterprise Search Tools and Features for 2026 | Slack
  15. [15] Best Enterprise Search Tools for 2026 | Complete Guide
  16. [16] 8 Best AI-Powered Enterprise Search Solutions | SearchUnify
  17. [17] Enterprise search: how AI-powered search boosts workplace productivity
  18. [18] The AI Enterprise Search Guide for IT and Knowledge Leaders
  19. [19] Sufficient Context: A New Lens on Retrieval Augmented Generation Systems
  20. [20] What is Retrieval Augmented Generation (RAG)? | Databricks
  21. [21] Deeper insights into retrieval augmented generation: The role of sufficient context
  22. [22] [2411.06037] Sufficient Context: A New Lens on Retrieval Augmented Generation Systems
  23. [23] Graph World Model
  24. [24] Context-Aware Retrieval-Augmented Generation for Artificial Intelligence in Urology
  25. [25] Retrieval-Augmented Generation: A Practical Guide to RAG Architecture, Retrieval, and Production-Ready Context
  26. [26] What is RAG (Retrieval Augmented Generation)? | IBM
  27. [27] Explaining retrieval-augmented generation
  28. [28] AI's most important protocol is getting a little bit easier to use | TechCrunch
  29. [29] Model Context Protocol Blog
  30. [30] The 2026-07-28 MCP Specification Release Candidate | Model Context Protocol Blog
  31. [31] The 2026-07-28 Specification | Model Context Protocol Blog
  32. [32] Model Context Protocol
  33. [33] Roadmap - Model Context Protocol
  34. [34] The New MCP Roadmap | Model Context Protocol Blog
  35. [35] The next generation of MCP | Cloudflare Blog
  36. [36] MCP 2026-07-28: From Local Tool to Distributed Protocol - Agentic AI Foundation (AAIF)
  37. [37] Anthropic update unifies memory feature across Claude Cowork and chat - 9to5Mac
  38. [38] Anthropic updates Claude’s memory to enhance customization and protect sensitive topics - SiliconANGLE
  39. [39] Claude and Cowork now share what they know about you
  40. [40] Anthropic merges Claude chat and Cowork memory, on by default
  41. [41] Claude's memory now works across both chats and Cowork sessions - Engadget
  42. [42] Claude Cowork finally remembers what you told the app in chat | TechCrunch
  43. [43] Claude Memory: What It Stores & How to Delete It | LumiChats
  44. [44] New on Yahoo
  45. [45] New on Yahoo
  46. [46] New on Yahoo
  47. [47] ChatGPT Enterprise Connectors: Office 365 & Azure Guide | IntuitionLabs
  48. [48] 📢 Announcement!! Azure OpenAI and Azure AI Search connectors are now Generally Available (GA) | Microsoft Community Hub
  49. [49] OpenAI’s New Search Connectors: Breaking Down Data Silos with AI-Powered Search | by CherryZhou | Medium
  50. [50] OpenAI Pushes Into Enterprise Search With Company Knowledge
  51. [51] ChatGPT Business release notes | OpenAI Help Center
  52. [52] ChatGPT Enterprise and Edu release notes | OpenAI Help Center
  53. [53] More ways to work with your team and tools in ChatGPT | OpenAI
  54. [54] OpenAI Search Connector Launches! Unlock a New Productivity Tool for ChatGPT
  55. [55] OpenAI Company Knowledge in ChatGPT: Enterprise Connectors and Citations | Windows Forum
  56. [56] Top 9 Vector Databases as of September 2026 | Shakudo Blog
  57. [57] Top 10 AI Vector Databases for 2026 (Compared)
  58. [58] Best Vector Databases 2026: Pinecone, Chroma, Qdrant & More | DataCamp
  59. [59] Dnotitia Brings Dedicated Vector Silicon to Server Scale at AI Infra Summit 2026 | The Manila Times
  60. [60] 6 data predictions for 2026: RAG is dead, what's old is new again and the future of vector databases
  61. [61] Dnotitia Brings Dedicated Vector Silicon to Server Scale at AI Infra Summit 2026 - AIwire
  62. [62] 2026 in artificial intelligence
  63. [63] MindsDB
  64. [64] Top 10 Vector Databases for LLM Applications in 2026 | Second Talent
  65. [65] MCP in 2026: What Changed in the 2026-07-28 Specification and How to Design Production Integrations - DEV Community
  66. [66] Specification - Model Context Protocol
  67. [67] Update on the Next MCP Protocol Release
  68. [68] Scaling AI Agent Infrastructure with the MCP Stateless updates - Google Developers Blog
  69. [69] GitHub MCP Server supports the next MCP specification - GitHub Changelog
  70. [70] Model Context Protocol Specification Version Timeline - Version-by-Version Changes and Adoption Milestones | hidekazu-konishi.com

Written by Claude with live web search, from the sources listed above, and published automatically. Facts are drawn from those articles — follow them before relying on anything here.