Why Your MCP Server Loses Session State Behind a Load Balancer
MCP's 2026-07-28 spec removes protocol-level sessions, allowing you to run servers behind round-robin load balancers. Here is the migration guide.
Join the DZone community and get the full member experience.
Join For FreeThe Problem
The Model Context Protocol has an initialize handshake. A client connects, the server hands back an Mcp-Session-Id header, and every request after that carries the same ID.
That's fine with one server. It falls apart with three.
A load balancer doesn't know or care about Mcp-Session-Id. It's an application-layer detail, and by default, the load balancer routes based on whatever algorithm it uses: round robin, least connections, or source IP hash. The client's second request can land on a completely different instance than the one that issued the session. That instance has never heard of the session ID, so it returns an error or, worse, silently starts a new session with no memory of the tools the client already listed or the state it had built up.
Teams hit this the same way every time. Local development works because there's one process. Staging works because there's one pod. Production breaks the first time the deployment scales beyond a single replica, and the failure looks like a flaky client rather than an infrastructure problem, so it takes a while to trace.
The standard fixes were the same ones every stateful HTTP service has used for 20 years: sticky sessions at the load balancer (route by session ID or client IP, accept the uneven load distribution), a shared session store like Redis so any instance can pick up any session, or deep packet inspection at the gateway to route on the Mcp-Session-Id header specifically. All three work. All three add an operational dependency to what should be a stateless API call.
What Changed in the Spec
The 2026-07-28 MCP specification release candidate removes protocol-level session management. The Mcp-Session-Id header is gone, and so is the session it represented. Connection metadata that used to be exchanged once at initialize time (protocol version, client info, client capabilities) now travels _meta on every request instead.
The practical effect: a remote MCP server can sit behind a plain round-robin load balancer, route traffic on the Mcp-Method header if you want method-aware routing, and let clients cache tools/list responses for as long as the server's ttlMs value says they're valid. No sticky sessions. No shared session store just to keep the protocol working.
This doesn't mean your server has to be stateless. It means the protocol stopped assuming state lives at the transport layer. If your server genuinely needs to remember something between calls (a shopping basket, an open browser tab, a half-finished multi-step operation), you handle it the way HTTP APIs have always handled it: mint an explicit handle from a tool call and have the model pass that handle back as an ordinary argument on the next call.
Migrating a Stateful Server
Say you have an MCP tool that opens a browser session and later needs to make calls to act on that same browser. Before the spec change, you'd have been tempted to key that off the transport-level session ID. After that, you do it explicitly:
from fastmcp import FastMCP
import uuid
mcp = FastMCP("browser-tools")
class BrowserHandle:
"""Minimal stand-in for a real browser automation session (e.g. Playwright)."""
def __init__(self, start_url: str):
self.start_url = start_url
def click(self, selector: str) -> None:
# Real implementation would drive an actual browser page here.
pass
# In-memory for the example; use Redis or a DB in production
_sessions: dict[str, BrowserHandle] = {}
@mcp.tool()
def open_browser(start_url: str) -> dict:
"""Open a browser and return a handle for later calls."""
handle_id = str(uuid.uuid4())
_sessions[handle_id] = BrowserHandle(start_url)
return {"browser_id": handle_id, "status": "opened"}
@mcp.tool()
def click_element(browser_id: str, selector: str) -> dict:
"""Click an element in a previously opened browser."""
session = _sessions.get(browser_id)
if session is None:
return {"error": f"No browser session for id {browser_id}. Call open_browser first."}
session.click(selector)
return {"status": "clicked", "selector": selector}
The model now carries browser_id as an ordinary tool argument, the same way it would carry a basket_id or an order_id. That handle can live in Redis with a TTL, in a database row, wherever makes sense for your durability requirements. It's no longer the protocol's job to keep it alive.
If you're running behind a load balancer today with sticky sessions configured specifically to work around the old MCP session model, this is your cue to check whether you can drop that configuration once your server and client both support the 2026-07-28 spec, or its final released version. Check the actual spec version your SDK negotiates before you rip out sticky sessions. Older clients still speaking the pre-07-28 protocol will still expect Mcp-Session-Id to work, and mixed-version fleets are exactly the kind of thing that turns a clean migration into a bad on-call week.
What to Check Before You Migrate
A few things worth confirming before you touch the load balancer config:
- Check your MCP SDK version and whether it implements the stateless spec or still assumes session pinning. Not every SDK moved at the same pace.
- Check whether any of your tools rely on server-side state that is implicitly tied to the session lifecycle (e.g., an open file handle, a database transaction, or a lock). Those need an explicit handle now, not an assumption that "the session" will still be around.
- Check your load balancer's health check and routing rules for anything referencing
Mcp-Session-Idspecifically. If your ops team added rules to route on that header, they can likely come out. - Test with a client that is legitimately routed to different instances across consecutive calls, not just a local single-instance setup. That's the scenario the old model broke on, and it's the one you want to confirm the new model handles.
References
Opinions expressed by DZone contributors are their own.
Comments