The most consequential thing for anyone shipping retrieval infrastructure right now isn't a model release. It's plumbing: the Model Context Protocol published a [specification release candidate dated 2026-07-28](https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/), and Google's developer relations team followed with a post on [scaling agent infrastructure with MCP's stateless updates](https://developers.googleblog.com/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates/). If your internal knowledge base is exposed to agents through an MCP server, this is the layer where your reliability problems currently live.
Why statelessness is the whole game
The original MCP shape assumed a session. A client connects, the server holds context for that connection, and the two exchange messages over a long-lived channel. That's fine for a desktop client talking to a local server over stdio. It is a poor fit for the way most teams actually deploy an enterprise connector — behind a load balancer, on autoscaled containers, with instances coming and going.
Session-bound servers force session affinity. Session affinity forces sticky routing, which breaks the moment an instance is drained during a deploy. Teams work around it with an external session store, which means every tool call now pays a round trip to Redis to rehydrate state that mostly didn't need to exist. A stateless server that treats each request as self-contained sidesteps all of it: any instance can serve any call, scale-to-zero becomes viable, and a rolling restart stops being an incident.
What this means if you own a knowledge connector
The tradeoff is that state you were implicitly keeping now has to be made explicit and carried. In a retrieval connector, that state is usually three things:
Auth and permission context. Per-user access control is the hard part of enterprise search, and if you were caching a resolved permission set per session, you now need it either recomputed per call or cached against a key you can derive from the request itself. Recomputing ACL filters on every query is expensive against systems like SharePoint or Jira; plan for a cache keyed on user plus source plus a short TTL, and accept the staleness window explicitly rather than by accident.
Pagination and cursors. Anything where the server was quietly remembering "where we were" in a result set has to move into an opaque cursor the client hands back.
Long-running work. Reindexing, deep crawls, and multi-hop retrieval don't fit a single request. These need a job handle the client can poll, not a held connection.
None of that is new engineering. It's the same discipline that moved web applications off sticky sessions two decades ago. The reason it matters now is that most internal MCP servers were written fast, as prototypes, by people solving a retrieval problem rather than a distributed systems problem — and they are quietly accumulating session state that will not survive contact with production traffic.
If you're building on MCP, read the RC before your next connector refactor, not after.

