Mem0 published its *State of AI Agent Memory 2026* report three days ago. The useful part isn't the scores — it's the admission, from a vendor with every incentive to publish a clean leaderboard, that the leaderboard doesn't work.
The framing is that agent memory is now a first-class architectural component with its own benchmark suite, its own research literature, and a measurable performance gap between approaches , and that standardized benchmarks now let fundamentally different memory architectures be compared on the same evaluation set . Three define the space: LoCoMo, LongMemEval, and BEAM. LoCoMo is 1,540 questions across four categories testing recall across multi-session conversational data.
Then the numbers fall apart. Published LoCoMo claims range from Dakera's 88.2% (standard protocol, no reranking) through Mem0's 92.5% up to Zep's, ByteRover's, and ZeroMemory's claims in the 92–96% range, several of which are disputed or inconsistent. On LongMemEval, Mem0 reports 94.4%, ByteRover 92.8% on the LongMemEval-S variant, and Zep 71.2% under a GPT-4o judge — numbers from different model stacks and evaluation protocols, so "highest" is provisional rather than settled. The mechanism is mundane: "run LoCoMo" isn't one fixed procedure, and small protocol differences compound into large score differences. Reranking on or off. Which model judges. Which variant. At least one vendor has published two different scores for itself in different posts.
The axis that actually matters
Buried in the same report is a more load-bearing signal. Mem0's April 2026 algorithm — single-pass hierarchical extraction plus multi-signal retrieval — posted its two largest gains on temporal queries (+29.6 points) and multi-hop reasoning (+23.1 points) . That's where the headroom is, and it's where a memory layer earns its keep: both categories depend on the write path — consolidation, conflict resolution, recency handling — not on embedding quality.
Note also the units problem, which the report flags itself: the 2025 paper reported tokens per conversation (~26,000 for full context) while the 2026 algorithm reports average tokens per retrieval call (~6,956 for LoCoMo) . Different units, same underlying question. Elsewhere the figures are quoted as 91.6 on LoCoMo and 93.4 on LongMemEval at roughly 6,900 tokens per query — not the 92.5/94.4 above, which is the self-inconsistency problem in action.
What to do with this
If you're picking a memory layer this quarter, the accuracy column is decoration. Two things are comparable across vendors: tokens injected per retrieval, and accuracy on temporal and multi-hop splits specifically, run by you on your own transcripts under one fixed protocol. Aggregate LoCoMo scores compress away the only categories that distinguish a memory system from a vector index with a session key.
And treat the source for what it is — a vendor mapping a market it competes in. The map is honest about its own limits, which is more than the scores are.

