The Model Context Protocol maintainers published a new roadmap on the official MCP blog yesterday ([blog.modelcontextprotocol.io](https://blog.modelcontextprotocol.io/posts/mcp-roadmap/)). It arrives about a month after the `2026-07-28` specification shipped, which itself went through a public release candidate — a cadence worth noting on its own, because it means the connector layer under most enterprise RAG stacks now has a versioned, dated spec train rather than a moving target.
If you maintain MCP servers in front of a knowledge base, read the roadmap before you read anything else this week. But the change already in the wild is the one to plan around: Google's developer blog published guidance on scaling agent infrastructure with MCP's stateless updates roughly three weeks ago. Statelessness is not a cosmetic protocol detail. It is the difference between an MCP server you can run as one long-lived process per session and one you can put behind an ordinary load balancer with N replicas and no sticky routing.
Why this hits retrieval teams specifically
Most internal knowledge connectors were prototyped as stateful sessions — open a connection, negotiate capabilities, hold auth context and cursors in memory, stream results. That works for one developer on a laptop and falls over the moment 400 employees hit the same Confluence or Snowflake connector through an assistant. Session affinity becomes a hard dependency, deploys drop in-flight work, and horizontal scaling stops being free.
A stateless request/response shape pushes that state somewhere you control explicitly: a token, a cache, an external store. The tradeoff is real, not free. You pay in re-authentication or token validation per call, you can no longer amortize an expensive index handle or warm embedding client across a session, and cursored pagination over large document sets gets more awkward when the server is not allowed to remember where you were. Teams doing hybrid retrieval with rerankers will feel the cold-start cost most.
The mitigation is unglamorous and familiar: move the expensive state to a shared cache keyed by an opaque cursor, keep tool responses small and self-describing, and treat the MCP server as a thin retrieval facade over infrastructure that was already horizontally scalable.
Who should care, and who shouldn't yet
If you run MCP servers in production for more than a handful of internal users, this is a migration to schedule, not to admire. If you are still at the pilot stage, the practical move is narrower: pin to a dated spec version, and design new connectors so no request depends on a prior one. That constraint costs almost nothing to adopt now and is expensive to retrofit later.
Everything else in the memory-and-retrieval discourse this week — the agent-memory-versus-RAG framing, the write-path arguments — is downstream of whether your retrieval surface can actually scale. Fix the plumbing first.

