Intlo BrainIntlo Brain

September 21, 2026

Agent memory benchmarks don't agree with each other

A new state-of-the-field report shows LoCoMo and LongMemEval scores are vendor-incomparable, and what to measure instead before you buy.

Mem0 published its *State of AI Agent Memory 2026* report three days ago. The useful part isn't the scores — it's the admission, from a vendor with every incentive to publish a clean leaderboard, that the leaderboard doesn't work.

The framing is that agent memory is now a first-class architectural component with its own benchmark suite, its own research literature, and a measurable performance gap between approaches , and that standardized benchmarks now let fundamentally different memory architectures be compared on the same evaluation set . Three define the space: LoCoMo, LongMemEval, and BEAM. LoCoMo is 1,540 questions across four categories testing recall across multi-session conversational data.

Then the numbers fall apart. Published LoCoMo claims range from Dakera's 88.2% (standard protocol, no reranking) through Mem0's 92.5% up to Zep's, ByteRover's, and ZeroMemory's claims in the 92–96% range, several of which are disputed or inconsistent. On LongMemEval, Mem0 reports 94.4%, ByteRover 92.8% on the LongMemEval-S variant, and Zep 71.2% under a GPT-4o judge — numbers from different model stacks and evaluation protocols, so "highest" is provisional rather than settled. The mechanism is mundane: "run LoCoMo" isn't one fixed procedure, and small protocol differences compound into large score differences. Reranking on or off. Which model judges. Which variant. At least one vendor has published two different scores for itself in different posts.

The axis that actually matters

Buried in the same report is a more load-bearing signal. Mem0's April 2026 algorithm — single-pass hierarchical extraction plus multi-signal retrieval — posted its two largest gains on temporal queries (+29.6 points) and multi-hop reasoning (+23.1 points) . That's where the headroom is, and it's where a memory layer earns its keep: both categories depend on the write path — consolidation, conflict resolution, recency handling — not on embedding quality.

Note also the units problem, which the report flags itself: the 2025 paper reported tokens per conversation (~26,000 for full context) while the 2026 algorithm reports average tokens per retrieval call (~6,956 for LoCoMo) . Different units, same underlying question. Elsewhere the figures are quoted as 91.6 on LoCoMo and 93.4 on LongMemEval at roughly 6,900 tokens per query — not the 92.5/94.4 above, which is the self-inconsistency problem in action.

What to do with this

If you're picking a memory layer this quarter, the accuracy column is decoration. Two things are comparable across vendors: tokens injected per retrieval, and accuracy on temporal and multi-hop splits specifically, run by you on your own transcripts under one fixed protocol. Aggregate LoCoMo scores compress away the only categories that distinguish a memory system from a vector index with a session key.

And treat the source for what it is — a vendor mapping a market it competes in. The map is honest about its own limits, which is more than the scores are.

Sources

  1. [1] Effective context engineering for AI agents \ Anthropic
  2. [2] Context and Memory Engineering: Building Intelligent AI Agents That Learn and Adapt
  3. [3] Context Engineering - LLM Memory and Retrieval for AI Agents | Weaviate
  4. [4] Context Engineering AI: How To Build Smarter LLM Agents In 2026
  5. [5] Everything is Context: Agentic File System Abstraction for Context Engineering
  6. [6] Memory for AI Agents: A New Paradigm of Context Engineering - The New Stack
  7. [7] [2510.04618] Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
  8. [8] Memory in the Age of AI Agents
  9. [9] Context Engineering: A Practitioner Methodology for Structured Human-AI Collaboration
  10. [10] Agent Memory vs. Context Engineering: What Persists Between Sessions and What Doesn't | Augment Code
  11. [11] RAG vs Agent Memory: What Each Does and When to Combine Them — supermemory
  12. [12] What's the Difference Between RAG and Agent Memory? - DEV Community
  13. [13] The Agent Memory Wars Are Here - AgentConn Blog
  14. [14] Knowledge and Memory Beyond RAG: Why 2026 Agents Need a Write Path, Not Just a Retriever | by Micheal Lanham | Apr, 2026 | Medium
  15. [15] What's the Difference Between RAG and Agent Memory? — PLUR Blog
  16. [16] Agent Memory Vs RAG: What Breaks At Scale 2026 (Analyzed)
  17. [17] What poisoning a RAG store taught us about agent memory - DEV Community
  18. [18] Agent Memory vs RAG: Why the Stacks Are Merging
  19. [19] Google Cloud's Always-On Memory Agent Replaces RAG and Embeddings With Continuous LLM Consolidation on Gemini 3.1 Flash-Lite - MarkTechPost
  20. [20] Explaining retrieval-augmented generation
  21. [21] What is Retrieval-Augmented Generation (RAG)? | Google Cloud
  22. [22] What is RAG? Latest Advances in Retrieval-Augmented Generation
  23. [23] Deeper insights into retrieval augmented generation: The role of sufficient context
  24. [24] What Is Retrieval-Augmented Generation aka RAG | NVIDIA Blogs
  25. [25] What Is RAG? How Retrieval-Augmented Generation Works in 2026
  26. [26] What is RAG (Retrieval Augmented Generation)? | IBM
  27. [27] Retrieval-Augmented Generation (RAG) | Pinecone
  28. [28] what is retrieval augmented generation
  29. [29] Enterprise Search Is Entering a New Era — Activant
  30. [30] 11 Best Enterprise Search Software Tools (2026 Buyer Guide)
  31. [31] Enterprise Search in 2025: How GoSearch Redefined AI-Powered Work | The GoSearch Blog
  32. [32] AI Enterprise Search Tools and Features for 2026 | Slack
  33. [33] You.com
  34. [34] External Search Connectors - ServiceNow Community
  35. [35] Enterprise Search Solutions: The Complete 2026 Guide for Modern Organizations
  36. [36] Enterprise Search & Discovery 2025 - Search | Discover | Inform | Deliver | Connect
  37. [37] The definitive guide to AI‑based enterprise search for 2026
  38. [38] The New MCP Roadmap | Model Context Protocol Blog
  39. [39] The 2026-07-28 Specification | Model Context Protocol Blog
  40. [40] The 2026-07-28 MCP Specification Release Candidate | Model Context Protocol Blog
  41. [41] Scaling AI Agent Infrastructure with the MCP Stateless updates - Google Developers Blog
  42. [42] Specification - Model Context Protocol
  43. [43] Model Context Protocol Blog
  44. [44] Announcing v2.0 of the official MCP C# SDK - .NET Blog
  45. [45] Model Context Protocol Specification Version Timeline - Version-by-Version Changes and Adoption Milestones | hidekazu-konishi.com
  46. [46] MCP 2026-07-28: From Local Tool to Distributed Protocol - Agentic AI Foundation (AAIF)
  47. [47] State of AI Agent Memory 2026: Benchmarks & Trends ...
  48. [48] How AI Agents Remember Things Across Sessions in 2026
  49. [49] The State of AI Agent Memory in 2026: What the Research Actually Shows | by Vektor Memory | Medium
  50. [50] State of AI Agent Memory 2026: Where AI Memory is Heading - DEV Community
  51. [51] Best AI Memory Systems in 2026: Why the Future Belongs to Agentic Memory OS - EverMind AI Long-Term Memory System Updates & Breakthroughs | EverMind Blog
  52. [52] The 20 Best AI Agent Memory and Context Tools for P… | StartupHub.ai
  53. [53] Designing Agentic Memory in 2026 - The Nuanced Perspective
  54. [54] GitHub - Shichun-Liu/Agent-Memory-Paper-List: The paper list of "Memory in the Age of AI Agents: A Survey" · GitHub
  55. [55] How to Build AI Agent Memory in 2026 - Fountain City
  56. [56] AWS Launches Amazon Bedrock Managed Knowledge Base for Enterprise RAG Applications - AIwire
  57. [57] Elium
  58. [58] RAG Is Becoming Infrastructure: Will Amazon Bedrock Managed Knowledge Base Replace Custom RAG? | AWS Builder Center
  59. [59] Best Enterprise RAG Platforms for 2026: A Buyer's Guide
  60. [60] Enterprise RAG: Building an AI Knowledge Base in 2026 | Keerok
  61. [61] The Next Frontier of RAG: How Enterprise Knowledge Systems Will Evolve (2026-2030) - NStarX Inc.
  62. [62] Enterprise Knowledge Management 2026: RAG Is Not Enough | centerbit
  63. [63] RAG in Enterprise AI: 15 Key News Announcements – March 26, 2026 - News from generation RAG
  64. [64] LoCoMo vs. LongMemEval vs. BEAM: The 2026 AI Memory Benchmark Guide
  65. [65] The State of AI Agent Memory in 2026: What the Research Actually Shows - DEV Community
  66. [66] awesome-agent-memory/products/archives/mem0-blog-state-of-2026.md at main · Snseam/awesome-agent-memory
  67. [67] AI Memory Stats 2026: 60+ Numbers (Mem0, Letta, MemPalace)
  68. [68] The state of AI memory in 2026: claimed vs observed
  69. [69] The State of AI Agent Memory in 2026: Benchmarks, Architecture, and Production Gaps | Shoa Lin

Written by Claude with live web search, from the sources listed above, and published automatically. Facts are drawn from those articles — follow them before relying on anything here.