Intlo BrainIntlo Brain

August 25, 2026

MCP goes stateless and connector design changes

A July 28 MCP spec release candidate and Google's stateless-server guidance push knowledge connectors toward horizontally scalable, session-free design.

The most consequential thing for anyone shipping retrieval infrastructure right now isn't a model release. It's plumbing: the Model Context Protocol published a [specification release candidate dated 2026-07-28](https://blog.modelcontextprotocol.io/posts/2026-07-28-release-candidate/), and Google's developer relations team followed with a post on [scaling agent infrastructure with MCP's stateless updates](https://developers.googleblog.com/scaling-ai-agent-infrastructure-with-the-mcp-stateless-updates/). If your internal knowledge base is exposed to agents through an MCP server, this is the layer where your reliability problems currently live.

Why statelessness is the whole game

The original MCP shape assumed a session. A client connects, the server holds context for that connection, and the two exchange messages over a long-lived channel. That's fine for a desktop client talking to a local server over stdio. It is a poor fit for the way most teams actually deploy an enterprise connector — behind a load balancer, on autoscaled containers, with instances coming and going.

Session-bound servers force session affinity. Session affinity forces sticky routing, which breaks the moment an instance is drained during a deploy. Teams work around it with an external session store, which means every tool call now pays a round trip to Redis to rehydrate state that mostly didn't need to exist. A stateless server that treats each request as self-contained sidesteps all of it: any instance can serve any call, scale-to-zero becomes viable, and a rolling restart stops being an incident.

What this means if you own a knowledge connector

The tradeoff is that state you were implicitly keeping now has to be made explicit and carried. In a retrieval connector, that state is usually three things:

Auth and permission context. Per-user access control is the hard part of enterprise search, and if you were caching a resolved permission set per session, you now need it either recomputed per call or cached against a key you can derive from the request itself. Recomputing ACL filters on every query is expensive against systems like SharePoint or Jira; plan for a cache keyed on user plus source plus a short TTL, and accept the staleness window explicitly rather than by accident.

Pagination and cursors. Anything where the server was quietly remembering "where we were" in a result set has to move into an opaque cursor the client hands back.

Long-running work. Reindexing, deep crawls, and multi-hop retrieval don't fit a single request. These need a job handle the client can poll, not a held connection.

None of that is new engineering. It's the same discipline that moved web applications off sticky sessions two decades ago. The reason it matters now is that most internal MCP servers were written fast, as prototypes, by people solving a retrieval problem rather than a distributed systems problem — and they are quietly accumulating session state that will not survive contact with production traffic.

If you're building on MCP, read the RC before your next connector refactor, not after.

Sources

  1. [1] Agent Memory vs RAG: Key Differences Explained - Vectorize
  2. [2] Agent Memory Vs RAG: What Breaks At Scale 2026 (Analyzed)
  3. [3] AI Agent Memory 2026: Progress Benchmark Report Evaluations
  4. [4] RAG is Dead. Long Live Agent Memory - Chris Latimer - YouTube
  5. [5] Knowledge and Memory Beyond RAG: Why 2026 Agents Need a Write Path, Not Just a Retriever | by Micheal Lanham | Apr, 2026 | Medium
  6. [6] [2602.02007] Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
  7. [7] Machine Learning Pills
  8. [8] Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
  9. [9] Microsoft 365 connector for Claude brings enterprise search to users By Investing.com
  10. [10] Enterprise Search in 2025: How GoSearch Redefined AI-Powered Work | The GoSearch Blog
  11. [11] The AI Enterprise Search Guide for IT and Knowledge Leaders
  12. [12] The definitive guide to AI‑based enterprise search for 2025
  13. [13] Best Enterprise Search Tools for 2026 | Complete Guide
  14. [14] AI Enterprise Search Tools and Features for 2026 | Slack
  15. [15] Enterprise Search Solutions: The Complete 2026 Guide for Modern Organizations
  16. [16] LLMs with retrieval-augmented generation: Good or bad for privacy compliance? | IAPP
  17. [17] RAG is DEAD!. Retrieval-Augmented Generation ruled… | by Reliable Data Engineering | Medium
  18. [18] A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems:Progress, Gaps, and Future Directions
  19. [19] Retrieval-Augmented Generation for AI-Generated Content: A Survey | Data Science and Engineering | Springer Nature Link
  20. [20] Enhancing LLM Performance with Retrieval-Augmented Generation
  21. [21] 1 Retrieval-Augmented Generation for Large Language Models: A Survey
  22. [22] News from generation RAG - Dive deep into the transformative world of AI Retrieval Augmented Generation (RAG) technologies
  23. [23] AI Agents News — Week of August 24, 2026 (Daily Updates)
  24. [24] Daily AI Agent News - August 2026
  25. [25] GitHub - ARUNAGIRINATHAN-K/awesome-ai-agents-2026: Awesome AI Agents for 2026 · GitHub
  26. [26] AI News Today, August 21 — Top AI Stories & Live Updates | AI Weekly
  27. [27] Best AI Memory Systems in 2026: Why the Future Belongs to Agentic Memory OS - EverMind AI Long-Term Memory System Updates & Breakthroughs | EverMind Blog
  28. [28] MemoraX AI Ranks #1 on Agent Memory Leaderboard, Signaling a New Phase for Long-Term AI Memory
  29. [29] Agentic AI News — August 2026 Launches, Models & Research | Agentic.ai
  30. [30] You.com
  31. [31] Lucidworks
  32. [32] 2026 in technology and computing
  33. [33] Google AI Mode
  34. [34] Glean Technologies
  35. [35] AI Mode
  36. [36] Contextual AI
  37. [37] AI Model Releases, August 2026: What Changes for Your Visibility
  38. [38] New AI Model Releases News | August, 2026 (STARTUP EDITION)
  39. [39] Mistral Release Notes - August 2026 Latest Updates - Releasebot
  40. [40] AI Model Context Window Comparison 2026: Advertised vs. Real - elvex
  41. [41] AI Updates Today (August 2026) – Latest AI Model Releases
  42. [42] LLM Context Window Statistics (2026): Token Limit Data | BenchLM.ai
  43. [43] Major LLM Model Release Timeline [2025-2026] Claude Opus 5 / Gemini 3.6 Flash — LLM Data Hub
  44. [44] Meta Ads Updates (August 2026): What's Changing and What to Do
  45. [45] MCP in 2026: which AI agents support custom connectors (and how)
  46. [46] Scaling AI Agent Infrastructure with the MCP Stateless updates - Google Developers Blog
  47. [47] The 2026-07-28 MCP Specification Release Candidate | Model Context Protocol Blog
  48. [48] MCP Hits 10,000+ Servers as Biggest Update Ships [2026] – Tech Insider Ireland
  49. [49] How to Use Claude Connectors & MCP Servers: Complete Guide 2026 | explainx.ai Blog | explainx.ai
  50. [50] The 2026-07-28 Specification | Model Context Protocol Blog
  51. [51] Stateless MCP: What the 2026-07-28 specification changes for security | Equixly
  52. [52] MCP Goes Stateless, and Developers Ask Whether That Just Makes it an API Again - InfoQ
  53. [53] MCP Just Went Stateless — What the 2026 Spec Changes About Scaling on App Service | Microsoft Community Hub
  54. [54] The next generation of MCP | Cloudflare Blog
  55. [55] MCP Goes Stateless: What the 2026 Release Candidate ...
  56. [56] MCP 2026-07-28 spec: every breaking change, with fixes · Stacktree

Written by Claude with live web search, from the sources listed above, and published automatically. Facts are drawn from those articles — follow them before relying on anything here.