Repository navigation
streamable_http: early response.aclose() poisons keepalive connection, causes ~260ms latency on every subsequent tool call #2707
Description
Activity
- addedtriageQueued for automated analysis — bot will process and remove this labelQueued for automated analysis — bot will process and remove this label
on May 31, 2026 code paths identified but latency penalty not reproducible against the SDK's standalone test server. the reporter says the same in PR #2712: ~4 ms either way against uvicorn+sse-starlette, only ~265 ms against a FastMCP server co-resident with a desktop GUI on Windows. PR #2712 is open and proposes drain-to-EOF at all three early-close sites; it's sound HTTP/1.1 hygiene (don't release un-drained streaming bodies back to the pool) but cannot be guarded by a latency assertion in CI.
three early-close sites on main (commit 616476f) and origin/v1.x:
_handle_sse_responseat src/mcp/client/streamable_http.py:364_handle_resumption_requestat line 251_handle_reconnectionat line 421
repro (Linux, main 616476f): avg ~7 ms — no penalty observed
per-call (ms): ['7.8', '8.1', '7.8', '7.4', '7.3', '7.2', '7.5', '8.0', '7.0', '7.2', '7.0', '6.9', '7.4', '8.3', '7.8', '7.1', '7.0', '7.4', '7.1', '7.1'] avg = 7.42 ms min = 6.92 ms max = 8.30 ms NOT REPRODUCED: avg latency is lowwith PR #2712 applied: avg = 7.17 ms (statistically indistinguishable).
"""Reproducer for issue #2707: streamable_http early aclose() poisons keepalive. Spawns an MCP streamable_http server in a subprocess and measures latency of serial tools/call requests over a single ClientSession (SSE response mode). """ from __future__ import annotations import asyncio import multiprocessing import socket import sys import time import uvicorn from starlette.applications import Starlette from starlette.routing import Mount from mcp.client.session import ClientSession from mcp.client.streamable_http import streamable_http_client from mcp.server import Server, ServerRequestContext from mcp.server.streamable_http_manager import StreamableHTTPSessionManager from mcp.server.transport_security import TransportSecuritySettings from mcp.types import ( CallToolRequestParams, CallToolResult, ListToolsResult, TextContent, Tool, ) SERVER_NAME = "repro-2707" async def _list_tools(ctx, params=None) -> ListToolsResult: return ListToolsResult( tools=[ Tool( name="ping", description="Cheap tool that just returns ok", input_schema={"type": "object", "properties": {}}, ) ] ) async def _call_tool(ctx: ServerRequestContext, params: CallToolRequestParams) -> CallToolResult: return CallToolResult(content=[TextContent(type="text", text="ok")]) def _make_app() -> Starlette: server: Server = Server( SERVER_NAME, on_list_tools=_list_tools, on_call_tool=_call_tool, ) security = TransportSecuritySettings( allowed_hosts=["127.0.0.1:*", "localhost:*"], allowed_origins=["http://127.0.0.1:*", "http://localhost:*"], ) manager = StreamableHTTPSessionManager( app=server, json_response=False, security_settings=security ) return Starlette( debug=False, routes=[Mount("/mcp", app=manager.handle_request)], lifespan=lambda app: manager.run(), ) def _run_server(port: int) -> None: config = uvicorn.Config( app=_make_app(), host="127.0.0.1", port=port, log_level="warning", timeout_keep_alive=30, access_log=False, ) uvicorn.Server(config=config).run() def _wait_for_port(port: int, timeout: float = 10.0) -> None: deadline = time.time() + timeout while time.time() < deadline: with socket.socket() as s: try: s.connect(("127.0.0.1", port)); return except OSError: time.sleep(0.05) raise RuntimeError(f"server on port {port} did not come up in {timeout}s") async def _measure(url: str, n_warmup: int = 2, n_measure: int = 20) -> list[float]: async with streamable_http_client(url) as (r, w): async with ClientSession(r, w) as s: await s.initialize() for _ in range(n_warmup): await s.call_tool("ping", {}) times_ms: list[float] = [] for _ in range(n_measure): t0 = time.perf_counter() await s.call_tool("ping", {}) times_ms.append((time.perf_counter() - t0) * 1000.0) return times_ms def main() -> int: with socket.socket() as s: s.bind(("127.0.0.1", 0)); port = s.getsockname()[1] proc = multiprocessing.Process(target=_run_server, args=(port,), daemon=True) proc.start() try: _wait_for_port(port) url = f"http://127.0.0.1:{port}/mcp" times_ms = asyncio.run(_measure(url)) avg = sum(times_ms) / len(times_ms) print(f"per-call (ms): {[f'{t:.1f}' for t in times_ms]}") print(f"avg = {avg:.2f} ms min = {min(times_ms):.2f} ms max = {max(times_ms):.2f} ms") if avg > 100.0: print("REPRODUCED: avg latency is high — keepalive poisoning suspected"); return 1 print("NOT REPRODUCED: avg latency is low"); return 0 finally: proc.kill(); proc.join(timeout=2) if __name__ == "__main__": sys.exit(main())
run:
uv run python repro.pywhat would be needed to ask the reporter
a runnable repro that reliably triggers the stall — ideally the exact desktop/GUI host harness, or a minimal stand-in (e.g. running the MCP server in a thread under a different asyncio loop policy on Windows) so it can be exercised on CI. without that, the fix has to land on hygiene grounds rather than a latency regression test.
- addedbugSomething isn't workingSomething isn't workingneeds reproneeds additional information to be able to reproduce bugneeds additional information to be able to reproduce bugand removedtriageQueued for automated analysis — bot will process and remove this labelQueued for automated analysis — bot will process and remove this label
on Jun 1, 2026 Closing this as no reproduction was provided
Summary
In
mcp.client.streamable_http.StreamableHTTPTransport._handle_sse_response, the client callsawait response.aclose()immediately after receiving the first JSON-RPC response event. This early close leaves the underlying HTTP/1.1 keepalive connection in a state where the next request reusing the same connection blocks for ~260 ms before the server's response status arrives.The result is that every
session.call_tool(...)(andsend_ping,list_tools, ...) overstreamable_httppays a fixed ~260 ms penalty when calls are serial on a single connection.Removing the early
aclose()and draining the SSE stream to EOF eliminates the penalty entirely (37× speedup: 265 ms → 7 ms per call in my setup).Environment
mcp == 1.27.1mcp.server.streamable_http(also 1.27.1), localhost, SSE response modeSymptom (numbers)
Same
tools/call, same server, samehttpx.AsyncClient, all on localhost:httpx.AsyncClient.stream("POST", ...)+aiter_bytes()to EOFClientSession.call_tool(...)(current code)ClientSession.call_tool(...)(withaclose()removed)Status code arrival timing (measured with raw httpx on the same client/headers):
So the server replies in single-digit ms. The 260 ms appears only after MCP's early
aclose()on the previous request.Reproducer
Assuming any reachable streamable_http MCP server with one cheap read-only tool. Replace
URLandTOOL_NAME:On my setup this prints
avg = 267.40 ms. After the patch below it printsavg = 7.28 ms.Root cause
In
src/mcp/client/streamable_http.py,_handle_sse_response:After the response event is received, the SSE stream is force-closed before reaching EOF. The connection is then returned to the keepalive pool in a "not fully drained" state. The next POST attempting to reuse this connection blocks for ~260 ms before status arrives (likely a server-side SSE idle/reconnect window —
sse_starlette.EventSourceResponsekeeps the writer task alive after sending its only event).Confirming evidence (instrumented timings across many runs):
_handle_post_requestfrom entry to first_handle_sse_eventcall: 266 ms (always)client.stream(POST)issued on the samehttpx.AsyncClientand same event loop, outside the MCP call path: 5 msclient.stream(POST) + aiter_bytes() to EOFissued inside MCP'spost_writersubtask, immediately followed by the original_handle_post_request: bare = 266 ms, orig = 5 ms (the next call on the same connection is fast because the previous one drained to EOF)So the penalty is paid on the request following every early-aclose, not on the request that did the aclose.
Proposed fix
Drain the SSE stream to EOF instead of aborting early:
(The
last_event_id/ reconnect bookkeeping below is unaffected: we still observe every event and the loop now exits naturally on EOF.)Caveat
This relies on the server closing the SSE stream after sending the response (which
sse_starlette.EventSourceResponsedoes oncesse_writerexits viabreakon JSONRPCResponse — seemcp/server/streamable_http.py). If a server is configured to keep the stream open for multi-message responses, draining will wait for those too — which is the desired behavior. Arequest_read_timeout_seconds-aware variant could be added if needed.Happy to send a PR.