Repository navigation
streamable HTTP (stateful): concurrent requests with duplicate JSON-RPC ids cross-wire responses — one request receives another's payload, the other hangs #3060
Description
Activity
I can take this one. I have a fix working locally: reject a POST whose request id is already in flight on the session (400 + JSON-RPC -32600, per the spec's requirement that ids be unique within a session), while keeping sequential id reuse after completion working since deployed clients rely on it. Includes regression tests for both cases; full test suite passes. Will open a PR shortly.
The dispatcher layer on current
mainhas the same defect, and it looks like it's already on the maintainers' radar:jsonrpc_dispatcher.pyregisters in-flight requests with a blind overwrite (self._in_flight[coerce_request_id(req.id)] = ..., ~L578), annotatedTODO(maxisbey): duplicate ids blind-overwrite (v1/TS parity); revisit rejecting with INVALID_REQUEST(added in #3046). Two session-layer consequences beyond the transport-level cross-wiring reported here:notifications/cancelledcorrelates against_in_flight(~L610), so under duplicate in-flight ids a cancellation always targets the newest request — the older one becomes uncancellable the moment the duplicate arrives.- The completion path already has an identity guard (
entry.dctx is dctx, ~L706) so the older request's cleanup won't evict the newer entry — but the older entry itself is silently evicted from the table at registration time.
Worth noting
direct_dispatcher.pyalready rejects duplicate in-flight ids for caller-supplied ids (request id ... is already in flight), so rejecting at the wire dispatcher would make the two paths consistent.@maxisbey since the TODO is yours and #3046 is fresh — would you take a PR that flips it to reject with
INVALID_REQUEST, plus regression tests for the cancellation-targeting case? Intended as complementary to the transport-level guard @Sammy-Dabbas has in flight. Happy to put it up quickly, and equally happy to leave it with you if you'd rather fold it into the cancellation work.Agreed, these compose well. The transport guard matters independently because the slot overwrite in _request_streams happens at POST handling time, before the dispatcher ever sees the message, so dispatcher-level rejection alone would not stop the stream cross-wiring. Your dispatcher fix covers the other transports plus the cancellation-targeting case. My PR is transport-level only and leaves the L575 TODO untouched. Opening it now.
Put the dispatcher-layer piece up as #3064 (green). Scope is exactly as discussed: it flips the _in_flight blind-overwrite at the TODO(maxisbey) in jsonrpc_dispatcher.py to reject duplicate in-flight ids with INVALID_REQUEST, plus a regression test for the cancellation-targeting case. It's complementary to @Sammy-Dabbas's transport-level #3063 — his guards the streamable-HTTP stream routing, mine covers the dispatcher (and the other transports); neither closes this issue alone.
Happy to close.- addedbugSomething isn't workingSomething isn't workingP2Moderate issues affecting some users, edge cases, potentially valuable featureModerate issues affecting some users, edge cases, potentially valuable featurev1Affects the v1.x maintenance lineAffects the v1.x maintenance linev2Affects the v2 line (2.x on main)Affects the v2 line (2.x on main)
on Aug 14, 2026 Thank you for this work, Sammy. If helpful, I can also reproduce this problem in my environment on
mcp 1.29.0andmcp 2.0.0using Python 3.13.14 on macOS.request A (asked for PAYLOAD-FOR-A) received: HUNG (no response) [*** WRONG ***] request B (asked for PAYLOAD-FOR-B) received: PAYLOAD-FOR-A [*** WRONG ***]I'll test your patch in my system & happy to share more details about how I'm testing if it helps. I'd love to see this fix mainlined so I can get rid of our patch eventually.
Summary
In the stateful streamable-HTTP server transport, two concurrent POSTs on the same session that carry the same JSON-RPC request id cross-wire: the later request receives the earlier request's response (entire envelope, wrong arity and all), the later request's own response is silently dropped, and the earlier request's SSE stream hangs forever (keep-alive pings defeat client read timeouts).
Duplicate in-flight ids violate the spec ("The request ID MUST NOT have been previously used by the requestor within the same session"), but a real production client — claude.ai's custom-connector MCP client sends every request with
id: 1— triggers this constantly, and the server neither rejects nor tolerates the violation: it silently mis-routes user data. We observed one user's tool response delivered to a different conversation's request in production before isolating this mechanism.Mechanism (v1.27.0; the same code is present on current
main)mcp/server/streamable_http.py:request_id = str(message.root.id)thenself._request_streams[request_id] = anyio.create_memory_object_stream(...)— the routing table is keyed by request id alone. A second concurrent POST with the same id silently overwrites the first's slot; the first POST's reader keeps a now-orphaned stream.message_router(~L997–1045): the handler's response is routed byresponse_idto whichever POST currently owns the slot — i.e. the latest arrival, regardless of which request it answers.sse_writercleanup (~L597–616): after delivering one response the slot is popped, so the second response finds no slot and is dropped (debug log:Request stream 1 not found).Result for two overlapping requests A (id=1, slow) then B (id=1): B receives A's response; B's response is dropped; A hangs indefinitely.
mcp/shared/session.py(~L375) has the same single-key assumption in_in_flight[responder.request_id], which additionally breaks cancellation targeting under duplicate ids.Reproduction
Toy server (echo tool with random 0–2s sleep):
Client: one initialized session, 12 concurrent
tools/callPOSTs, each with a unique sentinel; modesame_idusesid: 1for all (mimicking claude.ai), modeunique_idis the control:Observed results (mcp 1.27.0, Python 3.12, Linux)
same_id(12 concurrent, allid: 1)unique_idcontrolPairwise variant (A slow, B posted 300ms later, both
id: 1, 10 trials): A never receives a response; B receives A's payload in 7/10 trials (B's own in the rest). This exactly matched our production incident, including the response-arity mismatch.stateless_http=Trueis immune by construction (per-request transport): same burst is 12/12 OK.Suggested fix
At minimum, reject a POST whose request id is already in flight on the session with a JSON-RPC
-32600error instead of silently overwriting the routing slot — a protocol violation should fail loudly, not deliver one user's data to another request. (Queueing/serializing duplicate-id requests would also work and keeps the misbehaving-but-widespread client functional.)The TypeScript SDK server has the same unguarded pattern (
_requestToStreamMapping.set(message.id, streamId)); filing separately there. The client-side id-reuse is also being reported to Anthropic.