The thread worth following this week is not a model release. It's the Model Context Protocol's move toward stateless operation, which lands squarely on anyone running connectors into an internal knowledge base.
The sequence: InfoWorld reported on 24 July that MCP was going stateless "to make scaling simpler." The 2026-07-28 specification shipped on schedule after a release candidate the same day. Google's developer blog followed on 5 August with guidance on scaling agent infrastructure against the stateless updates, and the MCP project published a new roadmap on 22–23 August. That's a spec change with a major cloud vendor writing deployment guidance a week later — a reasonable signal that the session model is being treated as the real bottleneck, not tool schema design.
Why session affinity was the problem
The original MCP shape assumed a long-lived connection between client and server. For a local server wrapping a filesystem that's fine. For a retrieval connector — Confluence, Slack, a warehouse, a vector index — it's a load balancing problem dressed up as a protocol. Session state means sticky routing, which means you can't treat connector instances as fungible, can't scale to zero, can't roll a deploy without dropping live agent sessions, and can't easily run the same connector fleet across regions. Teams have been working around this with external session stores and affinity rules at the proxy, which is infrastructure you have to operate for no user-visible benefit.
Statelessness pushes that state up to the client or into an explicit store you control. Every request carries what the server needs. The connector becomes an ordinary horizontally-scaled HTTP service, and the ops story collapses into something your platform team already knows how to run.
What it costs you
Nothing is free here. Stateless request handling means re-establishing context per call: re-authenticating to the upstream system, re-resolving tenant and permission scope, potentially re-paying connection setup to the underlying data store. For retrieval workloads where a single agent turn fans out to six connectors, that overhead is real and it lands on latency. Expect to compensate with aggressive credential and connection pooling on the server side, and with clients that batch rather than chatter.
The second cost is that anything genuinely session-scoped — cursors, in-flight pagination, partial result sets — now needs an explicit home. If you were implicitly relying on server memory to hold a query cursor between turns, that's now your design problem.
Who should act
If you run more than a handful of connectors in production, this is a planning item for this quarter: audit which of your MCP servers hold state, and whether that state is incidental or load-bearing. If you run two connectors on a single box, ignore it for now.
Separately worth reading: The Next Platform published a piece on 22 September arguing vector search has become a data type rather than a product category, while the cost arithmetic remains unresolved — a framing that matches what most teams find when they price a dedicated index against extending the database they already run.

