The most consequential thing in the connector layer this quarter isn't a new integration catalog. It's that the Model Context Protocol is shedding its stateful past. The 2026-07-28 specification revision — covered by The Register on July 23 under the headline that MCP "prepares to break with its stateful past," and by InfoWorld the next day framing the change as going stateless "to make scaling simpler" — reorganizes the protocol around requests that don't assume a long-lived session between client and server. MCP's roadmap post, published a few weeks ago, continues in that direction.
If you run MCP servers, the operational upside is obvious: no sticky sessions, no session affinity in your load balancer, no in-memory handle that has to survive a pod restart. You can put the thing behind a plain autoscaler and stop reasoning about connection lifetime as part of your capacity model.
The architectural cost is where teams should be paying attention. Statelessness doesn't delete state; it relocates it. Anything that used to live implicitly inside a session — pagination cursors, which documents you've already returned, tenant and permission context, accumulated retrieval history, what the agent has already learned — now has to be carried explicitly in the request or owned by your orchestrator. For a retrieval-backed system, that's not a minor refactor. It means every connector call needs auth context and dedup context passed in, and it means the memory layer is unambiguously your code, not the protocol's.
Which makes the memory retrieval problem harder to duck
That lands next to a real argument about *how* to retrieve from memory. A recent arXiv paper, "Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation," makes the case that standard RAG is poorly matched to agent memory in the first place. The reasoning is worth internalizing even if you skip the method: agent memory is a bounded, coherent interaction stream where many spans are near-duplicates of each other, not a large heterogeneous corpus. Flat top-k similarity retrieval over that stream mostly returns redundant context — five paraphrases of the same fact — while summary-centric hierarchies smooth away the small details that distinguish one candidate memory from another. The authors propose decoupling before aggregation: isolate reusable facts, updates, and distinguishing details first, then organize them for retrieval. Their system, xMemory, builds a revisable hierarchy from raw messages up through segments, memory components, and groups.
Who should care: anyone whose agent memory is currently a vector index over conversation turns with `k=10`. That design has two failure modes now visible at once — it wastes context on duplicates, and it was quietly relying on session state that the connector layer is no longer promising to keep.
Practical read for this week: audit your MCP servers for hidden session assumptions before you're forced to, and measure redundancy in your memory retrievals — dedup rate and unique-fact-per-token, not just recall@k.

