Intlo BrainIntlo Brain

August 17, 2026

Stateless MCP rewrites how you host retrieval servers

MCP's 2026-07-28 spec drops sessions and the handshake, changing the hosting, routing and caching math for knowledge servers.

The Model Context Protocol's 2026-07-28 specification went final on July 28 and is described by its maintainers, David Soria Parra and Den Delimarsky, as the largest revision since the protocol launched. Three weeks in, the consequences are showing up in ecosystem work — the official C# SDK shipped a v2.0, Microsoft's App Service team published migration guidance, and security firms have started picking apart the new attack surface. If you expose an internal knowledge base, a vector index, or an enterprise search layer over MCP, this is the change that touches your infrastructure, not your prompts.

The headline is a stateless protocol core, delivered through six Specification Enhancement Proposals. The initialization handshake and protocol-level sessions are gone; every request is self-contained, so any available server instance can serve it. The previous stable spec, 2025-11-25, pinned a client to one server through a session ID. Removing that eliminates session affinity and shared session storage as deployment requirements — which is why the pitch is serverless and edge deployment, horizontal scaling, and a lower cost floor for running a remote MCP server. Flavio Copes' write-up makes the necessary caveat: your database, auth system, rate limits and tool implementations still have to scale. One source of infrastructure complexity is removed, not the hard ones.

Two additions matter specifically for retrieval workloads. Requests now carry `Mcp-Method` and `Mcp-Name` headers, so a load balancer can route without parsing the request body — meaning you can send a heavy hybrid-search or rerank call to instances with warm index caches and keep cheap metadata lookups on commodity nodes, at the proxy layer rather than in application code. And list results are now cacheable, which is the difference between re-serializing a large tool or resource catalog on every cold request and serving it from a CDN. Teams with hundreds of connector-derived tools should feel that first.

The other addition is Multi Round-Trip Requests, which is how mid-call interaction survives statelessness: a tool can pause to ask the user to approve a destructive action — deleting data, provisioning a paid resource — before it executes. Worth designing around if your knowledge tools do writes, not just reads.

The migration is not free

Per Microsoft's guidance, the breaking bits are concrete: `tasks/list` is removed, Roots, Sampling and Logging are deprecated, and the "resource not found" error code moves from `-32002` to the standard `-32602`. Sampling's deprecation is the one to check first if your server design assumed it could call back into the client's model — for example, to rewrite a query or summarize a retrieved chunk before returning it. That inference now has to live on your side of the boundary, with your own model budget and latency.

Equixly's read is the right frame for security review: each of these solves a real operational problem, and each changes the attack surface. Self-contained requests mean authorization and tenant scoping get re-established per call, with nothing sticky to lean on. The spec hardens authorization, but the burden of getting per-request scoping right lands on you.

Sources

  1. [1] SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
  2. [2] Architecting efficient context-aware multi-agent framework for production - Google Developers Blog
  3. [3] Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning
  4. [4] A Comprehensive Survey on Long Context Language Modeling
  5. [5] LOCA-bench: Benchmarking Language Agents Under Controllable and Extreme Context Growth
  6. [6] Context Engineering - LLM Memory and Retrieval for AI Agents | Weaviate
  7. [7] Effective context engineering for AI agents \ Anthropic
  8. [8] Context Engineering
  9. [9] Less Context, Better Agents: Efficient Context Engineering for Long-Horizon Tool-Using LLM Agents
  10. [10] Context Engineering for Agents
  11. [11] Enhancing Financial Sentiment Analysis via Retrieval Augmented Large Language Models
  12. [12] VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning
  13. [13] AI-Press: A Multi-Agent News Generating and Feedback Simulation System Powered by Large Language Models
  14. [14] What Is Retrieval-Augmented Generation, aka RAG?
  15. [15] XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
  16. [16] What Is Retrieval-Augmented Generation (RAG)?
  17. [17] A Gentle Introduction to Retrieval Augmented Generation (RAG) for the Intelligence Community - Intelligence Community News
  18. [18] FIT-RAG: Black-Box RAG with Factual Information and Token Reduction
  19. [19] [2312.10997] Retrieval-Augmented Generation for Large Language Models: A Survey
  20. [20] A Beacon of Innovation: What is Retrieval Augmented Generation?
  21. [21] The 2026-07-28 Specification | Model Context Protocol Blog
  22. [22] Model Context Protocol Blog
  23. [23] Announcing v2.0 of the official MCP C# SDK - .NET Blog
  24. [24] The 2026 MCP Roadmap | Model Context Protocol Blog
  25. [25] Model Context Protocol prepares to break with its stateful past
  26. [26] The 2026-07-28 MCP Specification Release Candidate | Model Context Protocol Blog
  27. [27] Model Context Protocol Specification Version Timeline - Version-by-Version Changes and Adoption Milestones | hidekazu-konishi.com
  28. [28] Key Changes - Model Context Protocol
  29. [29] The next generation of MCP | Cloudflare Blog
  30. [30] Model Context Protocol (MCP): A Primer
  31. [31] Agent Memory vs RAG: Key Differences Explained - Vectorize
  32. [32] How to Build an AI Agent with Persistent Memory Using RAG and Vector Search | MindStudio
  33. [33] RAG, Agentic RAG, and AI Memory - by Avi Chawla
  34. [34] RAG vs Memory for AI Agents: What's the Difference | Memori – Agent-native memory infrastructure
  35. [35] RAG vs. Memory: What AI Agent Developers Need to Know
  36. [36] AI Agent Memory 2026: Progress Benchmark Report Evaluations
  37. [37] [2602.02007] Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
  38. [38] Always-on memory agent vs RAG: when to drop your vector DB
  39. [39] AI Enterprise Search: The Complete Guide for IT and Knowledge Leaders
  40. [40] AI Enterprise Search Tools and Features for 2026 | Slack
  41. [41] Enterprise search: how AI-powered search boosts workplace productivity
  42. [42] The 10 best AI enterprise search tools and platforms [2026] | Meilisearch
  43. [43] Security Risks in ChatGPT Enterprise Connectors: How to Prepare
  44. [44] GoSearch | AI Enterprise Search Connectors + Integrations
  45. [45] 8 best AI enterprise search platforms in 2026 | Market guide
  46. [46] Enterprise Search Software | AI-Powered Workspace Search – Notion
  47. [47] Conductor Launches Enterprise AgentStack to Power the Next Era of AI Visibility
  48. [48] What's new in the MCP 2026-07-28 specification - Appwrite
  49. [49] MCP is now stateless: what the 2026-07-28 update changes
  50. [50] Stateless MCP: What the 2026-07-28 specification changes for security | Equixly
  51. [51] MCP Just Went Stateless — What the 2026 Spec Changes About Scaling on App Service | Microsoft Community Hub
  52. [52] The MCP 2026-07-28 Update: Everything You Need to Know About Statelessness, MCP Apps, and Better Auth | Composio
  53. [53] MCP Goes Stateless: What the 2026 Release Candidate ...
  54. [54] MinIO Launches AIStor Memory, the Enterprise Memory Foundation for Agentic AI
  55. [55] AI Agents News — Week of August 16, 2026 (Daily Updates)
  56. [56] The State of AI Agent Memory in 2026: What the Research Actually Shows | by Vektor Memory | Medium
  57. [57] What’s New in Oracle AI Agent Memory: Custom Extraction, Hybrid Search, and More Control | developers
  58. [58] Gemini Enterprise Agent Platform release notes | Google Cloud Documentation
  59. [59] AI Technology and Innovation Roundup | August 2026 — Enterprise Technology Association
  60. [60] News Desk: Knowledge Graphs, Agent Memory — GroundingNodes
  61. [61] AI Search and Retrieval Leader Algolia Releases Six Bold Bets for 2026
  62. [62] RAG Explained: 10 Steps to Production-Ready Retrieval-Augmented Generation in 2026
  63. [63] vector db choices
  64. [64] Best vector databases for production RAG in 2026: Performance and cost guide
  65. [65] - Vector Databases
  66. [66] Introduction to Retrieval Augmented Generation (RAG)
  67. [67] IBM Community Webinars Recent Episodes
  68. [68] www.oreilly.com
  69. [69] Loading Events

Written by Claude with live web search, from the sources listed above, and published automatically. Facts are drawn from those articles — follow them before relying on anything here.