Skip to content

streamable_http: early response.aclose() poisons keepalive connection, causes ~260ms latency on every subsequent tool call #2707

Description

@whocareyw

Summary

In mcp.client.streamable_http.StreamableHTTPTransport._handle_sse_response, the client calls await response.aclose() immediately after receiving the first JSON-RPC response event. This early close leaves the underlying HTTP/1.1 keepalive connection in a state where the next request reusing the same connection blocks for ~260 ms before the server's response status arrives.

The result is that every session.call_tool(...) (and send_ping, list_tools, ...) over streamable_http pays a fixed ~260 ms penalty when calls are serial on a single connection.

Removing the early aclose() and draining the SSE stream to EOF eliminates the penalty entirely (37× speedup: 265 ms → 7 ms per call in my setup).

Environment

  • mcp == 1.27.1
  • Python 3.12.8, Windows 11
  • Server: mcp.server.streamable_http (also 1.27.1), localhost, SSE response mode
  • Transport: streamable HTTP, single long-lived client session, sequential requests

Symptom (numbers)

Same tools/call, same server, same httpx.AsyncClient, all on localhost:

Path Avg latency
Raw httpx.AsyncClient.stream("POST", ...) + aiter_bytes() to EOF ~5 ms
ClientSession.call_tool(...) (current code) ~265 ms
ClientSession.call_tool(...) (with aclose() removed) ~7 ms

Status code arrival timing (measured with raw httpx on the same client/headers):

  • Status: 1.5 ms
  • First chunk: 4.7 ms
  • EOF: 5 ms

So the server replies in single-digit ms. The 260 ms appears only after MCP's early aclose() on the previous request.

Reproducer

Assuming any reachable streamable_http MCP server with one cheap read-only tool. Replace URL and TOOL_NAME:

import asyncio, time, httpx
from mcp import ClientSession
from mcp.client.streamable_http import streamablehttp_client

URL = "http://localhost:PORT/mcp"
TOOL_NAME = "your_cheap_tool"
TOOL_ARGS = {}

async def main():
    async with streamablehttp_client(URL) as (r, w, _):
        async with ClientSession(r, w) as s:
            await s.initialize()
            # warm up
            for _ in range(2):
                await s.call_tool(TOOL_NAME, TOOL_ARGS)
            # measure
            times = []
            for _ in range(10):
                t0 = time.perf_counter()
                await s.call_tool(TOOL_NAME, TOOL_ARGS)
                times.append((time.perf_counter() - t0) * 1000)
            print(f"avg = {sum(times)/len(times):.2f} ms")

asyncio.run(main())

On my setup this prints avg = 267.40 ms. After the patch below it prints avg = 7.28 ms.

Root cause

In src/mcp/client/streamable_http.py, _handle_sse_response:

async def _handle_sse_response(self, response, ctx, is_initialization=False):
    ...
    async for sse in event_source.aiter_sse():
        ...
        is_complete = await self._handle_sse_event(...)
        if is_complete:
            await response.aclose()   # <-- offending line
            return

After the response event is received, the SSE stream is force-closed before reaching EOF. The connection is then returned to the keepalive pool in a "not fully drained" state. The next POST attempting to reuse this connection blocks for ~260 ms before status arrives (likely a server-side SSE idle/reconnect window — sse_starlette.EventSourceResponse keeps the writer task alive after sending its only event).

Confirming evidence (instrumented timings across many runs):

  • Time inside _handle_post_request from entry to first _handle_sse_event call: 266 ms (always)
  • Bare client.stream(POST) issued on the same httpx.AsyncClient and same event loop, outside the MCP call path: 5 ms
  • Bare client.stream(POST) + aiter_bytes() to EOF issued inside MCP's post_writer subtask, immediately followed by the original _handle_post_request: bare = 266 ms, orig = 5 ms (the next call on the same connection is fast because the previous one drained to EOF)

So the penalty is paid on the request following every early-aclose, not on the request that did the aclose.

Proposed fix

Drain the SSE stream to EOF instead of aborting early:

@@ async def _handle_sse_response(self, response, ctx, is_initialization=False):
-        try:
-            event_source = EventSource(response)
-            async for sse in event_source.aiter_sse():
-                ...
-                is_complete = await self._handle_sse_event(...)
-                if is_complete:
-                    await response.aclose()
-                    return
-        except Exception as e:
-            logger.debug(f"SSE stream ended: {e}")
+        try:
+            event_source = EventSource(response)
+            async for sse in event_source.aiter_sse():
+                ...
+                await self._handle_sse_event(...)
+        except Exception as e:
+            logger.debug(f"SSE stream ended: {e}")

(The last_event_id / reconnect bookkeeping below is unaffected: we still observe every event and the loop now exits naturally on EOF.)

Caveat

This relies on the server closing the SSE stream after sending the response (which sse_starlette.EventSourceResponse does once sse_writer exits via break on JSONRPCResponse — see mcp/server/streamable_http.py). If a server is configured to keep the stream open for multi-message responses, draining will wait for those too — which is the desired behavior. A request_read_timeout_seconds-aware variant could be added if needed.

Happy to send a PR.

Activity

  1. XidaoApi commented on May 29, 2026

    @XidaoApi
  2. added
    triageQueued for automated analysis — bot will process and remove this label
    on May 31, 2026
  3. mcp-claude commented on Jun 1, 2026

    @mcp-claude

    code paths identified but latency penalty not reproducible against the SDK's standalone test server. the reporter says the same in PR #2712: ~4 ms either way against uvicorn+sse-starlette, only ~265 ms against a FastMCP server co-resident with a desktop GUI on Windows. PR #2712 is open and proposes drain-to-EOF at all three early-close sites; it's sound HTTP/1.1 hygiene (don't release un-drained streaming bodies back to the pool) but cannot be guarded by a latency assertion in CI.

    three early-close sites on main (commit 616476f) and origin/v1.x:

    • _handle_sse_response at src/mcp/client/streamable_http.py:364
    • _handle_resumption_request at line 251
    • _handle_reconnection at line 421
    repro (Linux, main 616476f): avg ~7 ms — no penalty observed
    per-call (ms): ['7.8', '8.1', '7.8', '7.4', '7.3', '7.2', '7.5', '8.0', '7.0', '7.2', '7.0', '6.9', '7.4', '8.3', '7.8', '7.1', '7.0', '7.4', '7.1', '7.1']
    avg = 7.42 ms   min = 6.92 ms   max = 8.30 ms
    NOT REPRODUCED: avg latency is low
    

    with PR #2712 applied: avg = 7.17 ms (statistically indistinguishable).

    """Reproducer for issue #2707: streamable_http early aclose() poisons keepalive.
    
    Spawns an MCP streamable_http server in a subprocess and measures latency of
    serial tools/call requests over a single ClientSession (SSE response mode).
    """
    
    from __future__ import annotations
    
    import asyncio
    import multiprocessing
    import socket
    import sys
    import time
    
    import uvicorn
    from starlette.applications import Starlette
    from starlette.routing import Mount
    
    from mcp.client.session import ClientSession
    from mcp.client.streamable_http import streamable_http_client
    from mcp.server import Server, ServerRequestContext
    from mcp.server.streamable_http_manager import StreamableHTTPSessionManager
    from mcp.server.transport_security import TransportSecuritySettings
    from mcp.types import (
        CallToolRequestParams,
        CallToolResult,
        ListToolsResult,
        TextContent,
        Tool,
    )
    
    
    SERVER_NAME = "repro-2707"
    
    
    async def _list_tools(ctx, params=None) -> ListToolsResult:
        return ListToolsResult(
            tools=[
                Tool(
                    name="ping",
                    description="Cheap tool that just returns ok",
                    input_schema={"type": "object", "properties": {}},
                )
            ]
        )
    
    
    async def _call_tool(ctx: ServerRequestContext, params: CallToolRequestParams) -> CallToolResult:
        return CallToolResult(content=[TextContent(type="text", text="ok")])
    
    
    def _make_app() -> Starlette:
        server: Server = Server(
            SERVER_NAME,
            on_list_tools=_list_tools,
            on_call_tool=_call_tool,
        )
        security = TransportSecuritySettings(
            allowed_hosts=["127.0.0.1:*", "localhost:*"],
            allowed_origins=["http://127.0.0.1:*", "http://localhost:*"],
        )
        manager = StreamableHTTPSessionManager(
            app=server, json_response=False, security_settings=security
        )
        return Starlette(
            debug=False,
            routes=[Mount("/mcp", app=manager.handle_request)],
            lifespan=lambda app: manager.run(),
        )
    
    
    def _run_server(port: int) -> None:
        config = uvicorn.Config(
            app=_make_app(), host="127.0.0.1", port=port,
            log_level="warning", timeout_keep_alive=30, access_log=False,
        )
        uvicorn.Server(config=config).run()
    
    
    def _wait_for_port(port: int, timeout: float = 10.0) -> None:
        deadline = time.time() + timeout
        while time.time() < deadline:
            with socket.socket() as s:
                try:
                    s.connect(("127.0.0.1", port)); return
                except OSError:
                    time.sleep(0.05)
        raise RuntimeError(f"server on port {port} did not come up in {timeout}s")
    
    
    async def _measure(url: str, n_warmup: int = 2, n_measure: int = 20) -> list[float]:
        async with streamable_http_client(url) as (r, w):
            async with ClientSession(r, w) as s:
                await s.initialize()
                for _ in range(n_warmup):
                    await s.call_tool("ping", {})
                times_ms: list[float] = []
                for _ in range(n_measure):
                    t0 = time.perf_counter()
                    await s.call_tool("ping", {})
                    times_ms.append((time.perf_counter() - t0) * 1000.0)
                return times_ms
    
    
    def main() -> int:
        with socket.socket() as s:
            s.bind(("127.0.0.1", 0)); port = s.getsockname()[1]
        proc = multiprocessing.Process(target=_run_server, args=(port,), daemon=True)
        proc.start()
        try:
            _wait_for_port(port)
            url = f"http://127.0.0.1:{port}/mcp"
            times_ms = asyncio.run(_measure(url))
            avg = sum(times_ms) / len(times_ms)
            print(f"per-call (ms): {[f'{t:.1f}' for t in times_ms]}")
            print(f"avg = {avg:.2f} ms   min = {min(times_ms):.2f} ms   max = {max(times_ms):.2f} ms")
            if avg > 100.0:
                print("REPRODUCED: avg latency is high — keepalive poisoning suspected"); return 1
            print("NOT REPRODUCED: avg latency is low"); return 0
        finally:
            proc.kill(); proc.join(timeout=2)
    
    
    if __name__ == "__main__":
        sys.exit(main())

    run: uv run python repro.py

    what would be needed to ask the reporter

    a runnable repro that reliably triggers the stall — ideally the exact desktop/GUI host harness, or a minimal stand-in (e.g. running the MCP server in a thread under a different asyncio loop policy on Windows) so it can be exercised on CI. without that, the fix has to land on hygiene grounds rather than a latency regression test.

  4. added
    bugSomething isn't working
    needs reproneeds additional information to be able to reproduce bug
    and removed
    triageQueued for automated analysis — bot will process and remove this label
    on Jun 1, 2026
  5. norika1207-lab commented on Jun 2, 2026

    @norika1207-lab
  6. maxisbey commented on Jun 9, 2026

    @maxisbey
    Contributor

    Closing this as no reproduction was provided

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingneeds reproneeds additional information to be able to reproduce bug

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions