The most consequential change for anyone running retrieval connectors right now is not a model release. It's the MCP 2026-07-28 specification, which landed after a release candidate in late July and has spent the last several weeks working its way through the implementation stack.
The headline is the removal of session state. MCP co-creator David Soria Parra described the release candidate plainly: the protocol is now stateless, with "no handshake, no session id, any request can hit any server instance," alongside extensions as first-class citizens (MCP Apps, Tasks), auth hardening, and a formal deprecation policy. The Register covered the direction in late July as MCP breaking with its stateful past.
Why this matters if you run connectors in production
The original design assumed a long-lived session between client and server: initialize, negotiate capabilities, then issue tool calls against that session. That works on a laptop talking to a local server. It fails badly the moment you put an MCP server behind a load balancer. Session affinity means sticky routing, server-side session stores, and a deployment story where restarts drop live agent work. Teams building enterprise search connectors — Confluence, Jira, S3, a Postgres index — hit this the first time they tried to scale past one instance.
Statelessness removes that constraint. Any request lands on any replica, serverless deployment becomes viable, and rolling restarts stop killing sessions. Google's developer blog published guidance on scaling agent infrastructure with the stateless updates in early August, and Microsoft published parallel material on what the change means for hosting MCP servers on App Service. The official MCP C# SDK shipped v2.0 the same day as the spec.
The tradeoff is real and lands on the client. Without a session, capability negotiation and tool discovery either repeat per request or get cached client-side with a staleness problem. Auth context travels with every call rather than being established once. For high-frequency retrieval loops, that's a per-call overhead you now have to design around. Long-running work moves to the Tasks extension rather than being implicitly held open by the connection.
State didn't disappear, it moved
Strip state out of the transport and it reappears in the model layer. Anthropic's context management work on the Claude Developer Platform points the same direction: a memory tool plus context editing, which on their internal agentic search evaluation improved performance 39% over baseline when combined, with context editing alone at 29%. In a 100-turn web search evaluation, context editing let agents finish workflows that would otherwise exhaust context while cutting token consumption 84%.
Read those together and the architecture is clearer than it was six months ago. Connectors become dumb, horizontally scalable, per-request. Continuity lives in an explicit memory store the agent reads and writes, not in a socket. If you're still holding agent state in connection lifetime, that's now a migration, not a preference.

