The most consequential thing in this space right now isn't a model release. It's the settling-in period after the 2026-07-28 MCP specification, published on the Model Context Protocol blog after a public release candidate earlier that month. VentureBeat called it the biggest update MCP has had. The Register, writing on 23 July, described the direction bluntly: the protocol is preparing to break with its stateful past. Cloudflare published its own "next generation of MCP" piece a few weeks back, and the official roadmap page was revised again within the last four days — meaning the surface is still moving and anything you pin today may shift.
Why statefulness was the tax
If you run an MCP server in front of a knowledge base — Confluence, Drive, a warehouse, a vector index — the session-oriented model has been the quiet source of your operational pain. A long-lived session implies the server remembers who you are, where your cursor is, and what it already returned. That is fine on a laptop and miserable in production. It forces sticky routing, so you can't load-balance freely across replicas. It makes serverless deployment awkward, because the runtime that started the session may not exist when the next call lands. It turns a restart into a correctness problem, not just a latency blip. And it makes horizontal scaling of a read-heavy retrieval connector — which should be the easiest thing in your stack to scale — needlessly hard.
A stateless direction removes that tax, but it relocates the work rather than deleting it. The implications, and these are my inference rather than quoted spec text, are worth planning around now: continuation state for large result sets has to become an explicit, serializable token you hand back to the client rather than a pointer you keep in memory; auth and tenancy have to be re-established per request instead of resolved once at handshake; and any caching you were getting implicitly from session locality now needs to be a deliberate layer — a shared cache keyed on query plus tenant, not a warm process.
What to do about it
Read the 2026-07-28 spec directly rather than trusting summaries, including this one. Then audit your servers for the specific thing that breaks: anywhere you store a cursor, a partial result, or a resolved identity in process memory. Those are your migration items. If you're on a managed path instead — AWS introduced Bedrock Managed Knowledge Base on 26 June — you're buying out of some of this, at the cost of controlling your own chunking and ranking.
Adjacent and worth a skim: the research line arguing agent memory needs more than retrieval, including "Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation" (arXiv 2602.02007). The through-line with MCP is the same question — where state lives, and who owns it.

