A server that answers the legacy initialize with HTTP 200 but reports
2026-07-28 (regardless of the version offered) makes the SDK raise
'Unsupported protocol version from the server'. auto mode then falls back
to server/discover, which such a server rejects with a non-JSON 4xx, and
the connect died with 'both Streamable HTTP and SSE transports failed
(Streamable HTTP: Server returned an error response; SSE: 405)' even
though the same initialize/tools/list POSTs succeed from curl (#113359).
When both fail for that reason, re-run initialize ourselves, adopt the
result pinned to the version we offered (keeps later requests
legacy-shaped, which the handshake just proved the server accepts), send
notifications/initialized and proceed to tools/list.
mcp >= 2.0's Streamable HTTP client folds any non-2xx whose body it cannot parse
as a JSON-RPC error into the opaque `-32603 Server returned an error response`.
Hermes printed that text verbatim in the SSE-fallback warning and in the
"both transports failed" ConnectionError, so users saw no status, no URL and
none of the server's own words (e.g. `400 {"code":-32020,"message":"Unsupported
MCP-Protocol-Version"}`) and had to reach for curl to learn what the server
actually said (#114350, #113359).
- `_make_http_rejection_recorder`: response hook on the owned SDK-httpx client
that remembers the last 4xx/5xx (status, method, URL, head of the body; SSE
bodies are never read). Sibling of the redirect-header stripper hook.
- `_describe_http_failure`: appends that detail only when the root cause is
the SDK's opaque -32603 text, so a real JSON-RPC error or an httpx status
error is never duplicated.
- `_run_http`: the fallback warning and both raise paths (both-transports
ConnectionError; the no-fallback re-raise after a proven session, strict
redirect headers or a non-rejection) carry the detail. Debug-logs the
endpoint each connect attempt uses.
- Docs: troubleshooting entry for reading the new message.
Verified live against a real Streamable HTTP server (`hermes mcp test`, temp
HERMES_HOME): base prints `Streamable HTTP: Server returned an error response`;
fixed head prints `... (HTTP 400 from POST http://127.0.0.1:PORT/mcp:
{"jsonrpc":"2.0","id":null,"error":{"code":-32020,...}})`. Control: a 400
served as application/json already surfaces the JSON-RPC message and gets no
appendix; servers negotiating `initialize` down to 2025-06-18 (fixture and a
real FastMCP on mcp 1.12.4) connect and list tools on base and fix alike — the
pinned mcp 2.0.0 stamps the negotiated version on every post-handshake request
(wire-recorded), so the sticky seed in `_run_http` is not the cause of the
reported 400 and stays as designed (#14816).
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
The per-event body cap searched each streamed chunk in isolation, so an
event terminator split across two chunks was never seen: the finished
event's bytes were charged to the next event and the cap tripped early
on ordinary TCP fragmentation. A last-boundary rfind also lumped
multiple completed events in one chunk into a single charge, and the
separator tuple missed spec-legal CR and mixed line-ending styles.
Scan the carried suffix plus each new chunk with a terminator-pair
regex, walking every boundary left to right so each completed event is
charged once. The lookahead keeps a lone CRLF line ending from parsing
as a boundary, so multi-line events cannot evade the cap.
Review finding: _exc_children returned only .exceptions for a group, so
_is_session_expired_error missed a session-expiry marker (or the
InterruptedError override) hanging off a group's __cause__/__context__
that main used to inspect. Groups now yield nested + chain like every
other node; _flatten_messages' "group str() is opaque" rule is unchanged.
The salvaged fix gave `_find_missing` and `_flatten_messages` each their own
visited-set loop, next to the one `_is_session_expired_error` already had —
three copies of the same idiom in one module. Collapse them into
`_iter_exception_nodes` (pre-order, left-to-right, each node once, bounded by
`_EXC_TRAVERSAL_MAX_NODES`) and read all three scans off that list. Acyclic
output is byte-identical: the missing-executable search keeps its depth-first
order and a message-less leaf still renders as its class name.
Tests move from the issue-numbered file into `tests/tools/test_mcp_tool_errors.py`
(mirror of the source module): a two-node cycle renders the real messages, and a
missing stdio binary wrapped deeper than the recursion limit with the chain
looping back to the top is still reported as the missing executable. Both are
red on origin/main (RecursionError).
Co-authored-by: Stephan Mongstad <stephan@users.noreply.github.com>
Port from openclaw/openclaw#123194: a hostile or misbehaving remote MCP
server could stream an unbounded HTTP catalog/tool-result body that the
MCP SDK buffers and JSON-parses before any of Hermes' post-parse limits
(resource cap, tool-result truncation) run.
New _make_mcp_body_cap_transport wraps the owned httpx AsyncClient's
transport on the Streamable HTTP (mcp >= 1.24) and SSE paths:
- finite HTTP bodies capped at 10 MiB (Content-Length rejected up front,
streamed bodies capped chunk-by-chunk);
- each SSE event capped at 10 MiB, with accounting reset at completed
event boundaries so long-lived streams/keepalives are unlimited;
- violations raise httpx.ReadError naming the byte cap, handled by the
existing transport teardown/reconnect path (#66092).
verify/cert now live on the inner AsyncHTTPTransport (client-level TLS
kwargs are inert once a custom transport is passed); the SSE
httpx_client_factory is always injected so the cap applies with default
TLS too. Legacy mcp < 1.24 path (SDK-internal client, no hook) stays
uncapped — same degradation as strict_redirect_headers.
Relocated onto the decomposed module layout and hardened:
- Trigger covers the rejection CLASS, not just literal 400: SSE-only
servers' load balancers answer the chunked Streamable HTTP initialize
POST with 400/405/406/411, and the mcp>=2.0 SDK surfaces many such
rejections as an opaque -32603 'Server returned an error response'
(error class per #104363 by @RohithPariki). Timeouts and 5xx never
trigger the fallback: they are not transport mismatches.
- Reconnect exclusion via _ever_connected instead of _ready: run()
clears _ready before re-entering the transport, so the original guard
also fired on reconnects after a proven session.
- Successful fallback latches _sse_fallback so reconnects go straight
to SSE, and logs a warning suggesting the user pin transport: sse.
- Both transports failing raises a ConnectionError naming both errors
and suggesting transport: sse / checking the URL.
- No fallback with strict_redirect_headers (SSE cannot enforce that
boundary) or when transport is explicitly configured.
- Tests trimmed to 3 invariant contracts (proven red on base): fallback
connects + latches; no fallback on reconnect/timeout/5xx; both-fail
error is actionable.
The extracted SSE path reuses _sse_transport/_serve_transport from main,
preserving the bounded handshake timeout and reconnect-retry semantics.
Fixes#53676
mcp.servers.status now rides the shared _mcp_rpc decorator (profile scope, 4064,
5024 with the real message) instead of a hand-rolled try/finally with a blanket
except. Drop the _MCPConnectErrorText str subclass and reason taxonomy: the
existing status/error fields already carry the state, and a whitelist on the RPC
keeps error text out of the wire. The Desktop connections.health contribution
contract is held back until its consumer plugin is public. Tests trimmed to the
scope invariants (per-profile runtime visibility, scoped shutdown clears only its
own status, launch runtime never leaks into another profile).
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.