Repository navigation
Rejected streamable-HTTP requests leave live sessions behind: the session is registered before the request is validated #3228
Description
Activity
Independent confirmation on the released line, plus one behavior that I think is a separate defect from the one described here.
Reproduced on
mcp1.29.0 from PyPI (this issue is against 2.xmain), lowlevelServerwired straight toStreamableHTTPSessionManagerwith stock defaults, no FastMCP. Session counts read from_server_instances, 100 requests per probe:GET -> 406 +100 sessions HEAD -> 405 +100 sessions OPTIONS -> 405 +100 sessions GET + DELETE -> DELETE returns 200, +100 sessions total still registered in _server_instances: 400 refused GET returned 406 with Mcp-Session-Id: 4dfad7d0d20e4c8e888e10ed5ab56ca1 reusing it on a valid SSE GET -> 200Three notes.
1. It covers 405 as well, so the property is method-independent.
HEADandOPTIONSare refused with 405 from insideStreamableHTTPServerTransport.handle_request, which runs after the session is already registered. So the "server refuses the request and keeps the session anyway" behavior is not limited to 400/406/421; it includes requests whose method the transport does not implement at all.2. A successful
DELETEdoes not deregister the session either. This is the one I would flag as distinct from this issue. 100GETs, each followed by aDELETEof the returned session id, all returning 200, still leave +100 entries in_server_instances.The cause is in
run_server'sfinally:if (http_transport.mcp_session_id and http_transport.mcp_session_id in self._server_instances and not http_transport.is_terminated): del self._server_instances[http_transport.mcp_session_id]
terminate()sets_terminated = Truebefore the session task unwinds, so the cleanup is skipped for exactly the sessions that shut down cleanly. The transport's own streams are closed, so most of the memory is released, but the registry entry is permanent and the dict grows without bound.That is the well-behaved client path rather than the rejected-request path, so a fix scoped to "validate before registering" would not address it. Happy to split it into its own issue if you would rather keep this one focused.
3. Confirming your
RequireAuthMiddlewareexemption. I checked this because I had assumed the opposite. 300 unauthenticated requests behind auth return 401 and create zero sessions, with no session-id header, since the middleware wraps the ASGI app and returns before the manager runs. Your "not affected" note is correct, and it is worth keeping visible: it is the difference between this being reachable pre-auth and post-auth.Reproduction (SDK only,
uv run)# /// script # requires-python = ">=3.10" # dependencies = ["mcp==1.29.0", "httpx>=0.27", "uvicorn>=0.30", "starlette>=0.37"] # /// import contextlib, socket, threading, time import httpx, uvicorn from starlette.applications import Starlette from starlette.routing import Mount from mcp.server.lowlevel import Server from mcp.server.streamable_http_manager import StreamableHTTPSessionManager N = 100 manager = StreamableHTTPSessionManager(app=Server("repro")) # stateful, no idle timeout async def handle(scope, receive, send): await manager.handle_request(scope, receive, send) @contextlib.asynccontextmanager async def lifespan(app): async with manager.run(): yield app = Starlette(routes=[Mount("/mcp", app=handle)], lifespan=lifespan) with socket.socket() as s: s.bind(("127.0.0.1", 0)); port = s.getsockname()[1] server = uvicorn.Server(uvicorn.Config(app, host="127.0.0.1", port=port, log_level="error")) threading.Thread(target=server.run, daemon=True).start() while not server.started: time.sleep(0.05) base = f"http://127.0.0.1:{port}/mcp/" sessions = lambda: len(manager._server_instances) with httpx.Client(timeout=10.0) as c: for method in ("GET", "HEAD", "OPTIONS"): before, status = sessions(), None for _ in range(N): status = c.request(method, base).status_code print(f"{method:<8} -> {status} +{sessions() - before} sessions") before = sessions() for _ in range(N): sid = c.get(base).headers.get("mcp-session-id") if sid: d = c.request("DELETE", base, headers={"Mcp-Session-Id": sid}).status_code print(f"\nGET + DELETE -> DELETE returns {d}, +{sessions() - before} sessions") print(f"total still registered: {sessions()}")
- added a commit that references this issue
on Aug 13, 2026 Thanks — this is a genuinely useful report, and confirming it on the released line rather than
mainis the part I could not do myself.On 405. You are right that the property is method-independent, and it is a better framing than the one I opened with.
_handle_unsupported_requestruns downstream of registration exactly like the validation failures, and it echoes the session id back the same way. The fix in #3229 already covers it without a production change, because it keys off the establishing response status (>= 400) rather than an enumerated list of validation failures — so 405 was subsumed by construction. I have pushed a commit addingHEADandOPTIONSparams to that PR's test so a later refactor cannot narrow the predicate and silently reopen the method-independent half. Credited to you in the commit message.On DELETE. Confirmed, and I have split it out as #3300 as you offered. Your root-cause read is exactly right:
terminate()sets_terminated = Truebefore the streams close, sorun_server'sfinallyguard skips thedelfor precisely the sessions that shut down cleanly.One thing worth flagging about the evidence, since it took me a while to convince myself. Your probe establishes each session with
c.get(base)— a bareGETwith noAccept: text/event-stream— which the transport refuses with 406. So all 100 sessions in theGET + DELETErow were created by a request that was itself rejected, which is the defect this issue describes. The numbers are real, but they cannot separate "leaked because refused" from "leaked because DELETEd." I re-ran it establishing sessions with a realinitializereturning 200, which leaves theDELETEas the only variable, and it still leaks — so the conclusion holds, it just needed a cleaner isolation to be safe from that objection in review.The sharpest version turned out to be a contrast rather than a count. Same idle timeout, two clients:
B legacy, client DELETEs on close, idle_timeout=0.3 -> 3 of 3 still registered after 3s C legacy, client just vanishes, idle_timeout=0.3 -> 0 reaped correctlyThe polite client is punished and the rude one is cleaned up. That reads better to a maintainer than "the dict grows," and it also makes clear that
session_idle_timeoutis not a workaround.Two scope corrections I put in #3300 that are worth knowing here, because they cut against us: the leak is confined to the handshake-era path — the 2.x client default
mode="auto"negotiates2026-07-28, routes tohandle_modern_request, and never touches_server_instancesat all (measured 0 entries) — and sessions plusDELETEare markedremoved_in="2026-07-28"in this repo's own conformance table. So it is a real bug on a deprecated-but-supported path covering the entire pre-2026 installed base, which is a narrower claim than either of us started with. There is also prior art in #791, which fixed this on 1.x in May 2025 and was self-closed without review.On the auth exemption. Thanks for checking it independently — that is the kind of confirmation that is easy to skip because it only produces a negative result.
- added a commit that references this issue
on Aug 16, 2026 - addedv1Affects the v1.x maintenance lineAffects the v1.x maintenance linev2Affects the v2 line (2.x on main)Affects the v2 line (2.x on main)
on Aug 18, 2026 Confirming this with production numbers on
mcp1.27.1. In one 18.2 hour run, 109 of 224
created sessions came fromPOST /mcprequests that ended in 400. In
streamable_http_manager.py,self._server_instances[http_transport.mcp_session_id] = http_transportruns at line 254 and the task is started at line 303, but the request is
not validated untilhttp_transport.handle_request(...)at line 306 (validation lives in
streamable_http.py:487-505/_validate_request_headersat line 825). A request that
fails validation therefore leaves a fully registered session and a liverun_servertask
behind, with no client and no way to ever reach them again, since the client never learned
the session id.The ordering is unchanged in 2.0.0 (
streamable_http_manager.py:302registers, then the
request is handled), so the fix is still needed there.Suggested fix: either validate before registering, or register only after
handle_requesthas produced a non-error response, or pop the session id and terminate
the transport when the initializing request is rejected.Thanks @finedesignz — production numbers are the thing this issue was missing, and yours are worse than I expected. 109 of 224 sessions in an 18.2 hour window is roughly half of everything the manager registered being unreachable garbage.
Consolidating what's now on the thread, since it's spread across two issues:
Source Line Evidence @pete-builds released 1.29.0 400 leaked sessions across GET/HEAD/OPTIONS/DELETE probes; also found refusals hand back a usable Mcp-Session-Id@finedesignz released 1.27.1 109/224 sessions leaked in 18.2h of production traffic this issue 2.x mainordering unchanged — register at streamable_http_manager.py:302, validate insidehandle_requestat :359Three independent reproductions on three different versions, one of them from production.
@finedesignz your third suggested option — pop the session id and terminate the transport when the establishing request is rejected — is what #3229 implements. It keys off the establishing response status rather than an enumerated list of validation failures, so it also covers the 405 case @pete-builds found (
HEAD/OPTIONSare refused by_handle_unsupported_request, which runs downstream of registration in the same way). CI is green there.Two process notes rather than a nudge:
#3229 is still a draft. I opened it that way by mistake and can't flip it —
markPullRequestReadyForReviewneeds write access on the base repo, so it returnsFORBIDDENfor the PR's own author. It needs a maintainer to either mark it ready or assign me to this issue, at which pointrequire-linked-issueis satisfied too. Happy for someone else to carry the fix instead if that's easier — the diff is small and the approach is in the PR.Also worth noting the leak has no deployable mitigation on the affected path:
session_idle_timeoutdefaults toNoneand isn't exposed bystreamable_http_app(),MCPServer, or FastMCP, so a server built the documented way can't configure the reaper at all. Given that plus @finedesignz's ~49% figure, the currentP3may be worth a second look — entirely your call.- added a commit that references this issue
on Aug 21, 2026
Release line
2.x (
main), reproduced ata4f4ccd0.Bug description
In the default stateful streamable-HTTP configuration,
StreamableHTTPSessionManagermints asession id, registers a transport in
_server_instances, and starts a long-lived session taskbefore the request has been validated in any way. Every actual check lives downstream in the
transport, so a request the server itself refuses —
400,406, or421— still leaves a fullyregistered, non-terminated session behind.
Nothing reclaims it.
session_idle_timeoutdefaults toNoneand is not reachable fromstreamable_http_app()(that half is #2455), so the only reaper cannot be switched on by anoperator using the documented API.
Not filed as a security issue, and why
I originally wrote this up as a vulnerability report and then talked myself out of it, so it seems
worth stating plainly rather than leaving you to work it out.
It is unauthenticated, remotely triggerable, present on the documented default configuration, and
retains roughly 22.6 KiB per rejected request (measured two independent ways, ~2,240 req/s from a
single client). But it grants an attacker nothing they did not already have: on the same no-auth
server, 200 valid
initializerequests create 200 sessions — verified — so unbounded sessiongrowth is already reachable with entirely legitimate traffic. Closing the rejected-request path does
not change that, and I am not claiming otherwise. The real availability gap is the absence of a
session cap plus an idle reaper that is off by default and unreachable, which is #2455.
So this is a correctness / defense-in-depth defect, and a public issue is the honest channel for it.
What is worth fixing regardless of any attacker
test suite already asserts by name for the
413path (see below).421) create a session anyway, so asecurity control fires and state is created regardless.
/mcp— accumulatessessions forever on a default server.
406still returns a usableMcp-Session-Id, and a follow-up request on thatnever-initialized id is served
200by the leaked session's running server loop. An unauthenticatedcaller obtains a working session handle from a request the server rejected.
Affected
Any server built the documented way —
MCPServer(...).streamable_http_app(),Server(...).streamable_http_app(), ormcp.run("streamable-http")— i.e.stateless_http=False(default),
session_idle_timeout=None(default).Not affected:
stateless_http=True(transport terminated per request); servers behindRequireAuthMiddleware, where401is returned before the manager runs; requests carryingMCP-Protocol-Version: 2026-07-28, which route to the sessionless modern handler. Note that headeris client-controlled — omitting it selects the sessionful path — so it is not something an operator
can rely on.
Root cause
src/mcp/server/streamable_http_manager.py,_handle_stateful_request: the branch is taken purelyon
request_mcp_session_id is None. Inside it the code mintsuuid4().hex, constructs aStreamableHTTPServerTransport, inserts it intoself._server_instances, andawait self._task_group.start(run_server)— and only then callshttp_transport.handle_request(...),which is where every check lives (Host / DNS-rebinding,
Accept,Content-Type, JSON parse,JSON-RPC validation, and the "Missing session ID" check for non-
initializePOSTs).No error path calls
terminate()or pops the dict, andrun_server'sfinallyonly runs when thesession task exits — which it never does, because
serve_loopblocks on its read stream and the idlescope has no deadline when
session_idle_timeout is None.Steps to reproduce
Default configuration, zero arguments.
_server_instancesis only read, to observe the effect.Actual behaviour
Every session-less vector leaks exactly one session, each rejected with a different status:
GETno session idDELETEno session idPOSTnon-initializePOSTmalformed bodyPOSTbadAcceptGETbadHost(DNS-rebinding protection)300 rejected POSTs →
sessions 2 -> 302, exactly one per request, 302 non-terminated transportsstill registered.
The one pre-session guard that does exist works correctly: an oversized declared body is rejected
413byRequestBodyLimitMiddlewarewith no session created.Expected behaviour
A request the server refuses does not leave a registered, non-terminated session behind, and a
refused response does not hand back a usable
Mcp-Session-Id.Why this is not already covered
session_idle_timeoutnot exposed viastreamable_http_app()— covers only theunreachable-mitigation half. It does not mention that a rejected request creates a session.
accumulating, and was closed by the idle-timeout feature, which is off by default.
comment says "Reject with 405 BEFORE creating any transport or session to avoid leaking resources",
so the concern is understood — but the guard is
request.method == "GET"only. The POST and DELETEvectors above survive it unchanged.
Existing evidence this is unintended
tests/server/test_streamable_http_manager.pycontainstest_oversized_content_length_is_rejected_before_body_read_or_session_creation, assertingmanager._server_instances == {}. "Rejected without creating a session" is already treated as thedesired property — it was simply only implemented for the
413path.docs/run/legacy-clients.mdlikewise describes session creation as a consequence ofinitialize:"The moment a pre-2026 client sends
initialize, the SDK mints anMcp-Session-Id… and keeps alive record behind it."
Suggested fix
Two shapes; you are better placed to choose.
A. Only
initializemay establish a session. When there is no session id, reject anything that isnot a POST carrying an
initializerequest before creating a transport. Most faithful to thedocumented model, and subsumes #3129. Cost: the manager must peek the body, which currently lives in
the transport (
is_initialization_request).B. Discard a provisional session whose establishing request was refused. Wrap
sendto observethe response status; if
>= 400, pop the session andterminate()the transport. Smaller, needs nobody parsing. I have a working patch that closes all vectors in the table (0 leaked across 300
requests) while a legitimate
initializestill establishes a session and follow-up requests stillreturn
200.One thing to know before fixing, whichever shape you pick
tests/server/test_streamable_http_manager.pycurrently establishes sessions via requests thetransport rejects:
_open_sessionPOSTs an empty body, which is answered400/406and yet returnsa session id. Eight tests depend on it, so any correct fix fails them until that helper uses a real
initialize. The suite encodes the buggy behaviour as the way to obtain a session, which is plausiblywhy this went unnoticed.
With patch B applied and no test changes the suite is 5572 passed / 8 failed — all 8 in that one file,
all from that helper. Rewriting it is not a one-liner: a successful
initializein SSE mode keepshandle_requestopen, so the helper must read the response headers without waiting for the stream tofinish.
Disclosure
Investigated with AI assistance (Claude Code). Filed by me as the human contributor. Happy to open a
PR once you have indicated which fix shape you prefer.