The most consequential thing for anyone maintaining connectors right now is not a model release. It's the direction the Model Context Protocol has taken since the 2026-07-28 specification, which InfoWorld summarized bluntly in July: MCP is going stateless to make scaling simpler. The project has since published a new roadmap on blog.modelcontextprotocol.io, and the practical consequences land squarely on retrieval and knowledge-base teams.
Here's why it matters. The original MCP shape assumed a long-lived session between client and server: initialize, negotiate capabilities, then hold that connection while tools get called. If your server was a thin wrapper over a search index, this was tolerable overhead. If it was a knowledge connector — Confluence, a data warehouse, a document store with per-user ACLs — the session became a place where real state accumulated: auth context, cursors, cached embeddings, partial result sets. That state is exactly what makes horizontal scaling painful. You either pin clients to instances or you replicate session state across them, and both are load-balancer problems dressed up as protocol problems.
A stateless posture means each request carries what it needs. Operationally that's a big win: connectors become ordinary stateless HTTP services you can run behind a normal autoscaler, redeploy without draining sessions, and run at zero cost when idle. Serverless deployment of a knowledge connector stops being a hack.
The tradeoff is that the state doesn't disappear — it moves. Three places it lands, and each is a design decision you now own explicitly:
Auth and authorization. Per-user permission filtering on retrieval was often resolved once at session setup. Stateless means resolving identity and entitlements per request, which means you need a fast path for ACL lookup — a cached permissions index, or filters pushed into the vector/keyword query itself. If your current design fetches group membership from an upstream API at session start, you have a latency problem waiting.
Pagination and iterative retrieval. Agentic retrieval loops — search, read, refine, search again — leaned on server-side cursors. Opaque cursor tokens the client hands back are the clean answer, but they have to be signed and they have to survive an index refresh.
Memory. This is the part that shouldn't have been in your connector anyway. Conversation and long-term agent memory belongs in a dedicated store with its own write path, retention policy, and eviction logic, not smuggled into MCP session state. Mem0 published a State of AI Agent Memory report last week; the fact that memory is now a distinct infrastructure category rather than a feature of your RAG pipeline is the relevant signal.
What to do this week: audit each MCP server you run for anything held between calls. If a connector can't survive its process being killed mid-conversation, that's your migration list. Teams running one large stateful connector per tenant will feel this most.

