fix(api_server): honour platforms.api_server.tool_progress_events for Chat Completions SSE (#12020)

Strict OpenAI clients reject the named hermes.tool.progress SSE frames. Setting
tool_progress_events: false under platforms.api_server (loaded into
PlatformConfig.extra by from_dict) now drops them; default stays on.
Reimplements the intent of #42640 against the adapter config actually read in
production. Overlaps #49069 (erikerosev).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
This commit is contained in:
kshitijk4poor
2026-09-24 17:02:26 +05:30
committed by kshitij
parent 10963689dc
commit eea4419ddc
5 changed files with 57 additions and 2 deletions

View File

@@ -1199,6 +1199,11 @@ class APIServerAdapter(OpenAICompatRoutesMixin, BasePlatformAdapter):
# @mssteuer.)
self._direct_model_requests: bool = _coerce_request_bool(
extra.get("direct_model_requests"), default=False)
# ``platforms.api_server.tool_progress_events: false`` drops the custom
# ``hermes.tool.progress`` SSE frames from Chat Completions streams for strict OpenAI
# clients that choke on named events (#12020). Default on.
self._tool_progress_events: bool = _coerce_request_bool(
extra.get("tool_progress_events"), default=True)
self._app: Optional["web.Application"] = None
self._runner: Optional["web.AppRunner"] = None
self._site: Optional["web.TCPSite"] = None

View File

@@ -882,6 +882,8 @@ class OpenAICompatRoutesMixin:
break
if isinstance(delta, tuple) and len(delta) == 2 and delta[0] == "__tool_progress__":
# Custom event: tool lifecycle for frontends without markers in history.
if not self._tool_progress_events:
continue # opted out for strict OpenAI clients (#12020)
await response.write(_sse_frame(delta[1], event="hermes.tool.progress"))
elif isinstance(delta, tuple) and len(delta) == 2 and delta[0] == "__reasoning__":
# DeepSeek-style ``delta.reasoning_content`` (#99552), the field Open WebUI,

View File

@@ -0,0 +1,48 @@
"""``platforms.api_server.tool_progress_events: false`` drops the custom
``hermes.tool.progress`` SSE frames from streaming Chat Completions (#12020)."""
import asyncio
from unittest.mock import AsyncMock, MagicMock, patch
def _stream_body(platform_cfg):
from aiohttp import web
from gateway.config import PlatformConfig
from gateway.platforms.api_server import APIServerAdapter, ThreadSafeAsyncQueue
# from_dict is the production loader: a top-level platform key lands in ``extra``.
adapter = APIServerAdapter(PlatformConfig.from_dict(platform_cfg))
written = []
async def fake_agent():
return {"final_response": "done", "completed": True}, {
"input_tokens": 1, "output_tokens": 1, "total_tokens": 2}
async def run():
stream_q = ThreadSafeAsyncQueue()
stream_q.put_nowait(("__tool_progress__", {"tool": "terminal", "toolCallId": "c1", "status": "running"}))
stream_q.put_nowait("done")
stream_q.put_nowait(None)
agent_task = asyncio.ensure_future(fake_agent())
resp = AsyncMock(spec=web.StreamResponse)
resp.write = AsyncMock(side_effect=lambda data: written.append(data))
resp.prepare = AsyncMock()
req = MagicMock()
req.headers = {}
with patch("gateway.platforms.api_server.web.StreamResponse", return_value=resp):
await adapter._write_sse_chat_completion(req, "cmpl-1", "m", 1, stream_q, agent_task)
asyncio.run(run())
return b"".join(written).decode()
def test_tool_progress_frames_emitted_by_default():
body = _stream_body({"enabled": True, "token": "k"})
assert "event: hermes.tool.progress" in body
assert '"content": "done"' in body
def test_tool_progress_events_false_suppresses_frames_but_keeps_content():
body = _stream_body({"enabled": True, "token": "k", "tool_progress_events": False})
assert "hermes.tool.progress" not in body
assert '"content": "done"' in body

View File

@@ -111,7 +111,7 @@ Uploaded files (`file` / `input_file` / `file_id`) and non-image `data:` URLs re
All SSE streams (Chat Completions, Responses, `/api/sessions/{id}/chat/stream`, `/v1/runs/{id}/events`) emit a `: keepalive` comment line whenever no event has been sent for 10 seconds, so long tool calls do not trip client idle timeouts. Standard SSE clients ignore comment lines; custom parsers must skip lines that start with `:`.
**Tool progress in streams**:
- **Chat Completions**: Hermes emits `event: hermes.tool.progress` for tool-start visibility without polluting persisted assistant text.
- **Chat Completions**: Hermes emits `event: hermes.tool.progress` for tool-start visibility without polluting persisted assistant text. Strict OpenAI clients that choke on named SSE events can turn these frames off with `gateway.platforms.api_server.tool_progress_events: false` (default `true`); content chunks are unaffected. The opt-out applies only to Chat Completions — `/v1/runs/{id}/events` always emits tool events, which is what the `tool_progress_events` feature in `/v1/capabilities` describes.
- **Responses**: Hermes emits spec-native `function_call` and `function_call_output` output items during the SSE stream, so clients can render structured tool UI in real time.
**Model reasoning** (emitted only when the model actually produces reasoning and the resolved `reasoning` config allows it; the input-side opt-out is `model_options.reasoning.enabled: false`):

View File

@@ -109,7 +109,7 @@ curl http://localhost:8642/v1/chat/completions \
**流式传输**(`"stream": true`):返回逐 token 响应块的 Server-Sent Events(SSE)。对于 **Chat Completions**,流使用标准 `chat.completion.chunk` 事件,以及 Hermes 自定义的 `hermes.tool.progress` 事件用于工具启动的 UX 展示。对于 **Responses**,流使用 OpenAI Responses 事件类型,如 `response.created`、`response.output_text.delta`、`response.output_item.added`、`response.output_item.done` 和 `response.completed`。
**流中的工具进度:**
- **Chat Completions**:Hermes 发出 `event: hermes.tool.progress` 以提供工具启动可见性,同时不污染持久化的 assistant 文本。
- **Chat Completions**:Hermes 发出 `event: hermes.tool.progress` 以提供工具启动可见性,同时不污染持久化的 assistant 文本。无法处理具名 SSE 事件的严格 OpenAI 客户端可设置 `gateway.platforms.api_server.tool_progress_events: false`(默认 `true`)关闭这些帧;内容块不受影响。该开关仅作用于 Chat Completions——`/v1/runs/{id}/events` 始终发出工具事件,`/v1/capabilities` 中的 `tool_progress_events` 功能描述的正是这一点。
- **Responses**:Hermes 在 SSE 流期间发出符合规范的 `function_call` 和 `function_call_output` 输出项,让客户端能够实时渲染结构化工具 UI。
**模型推理**(仅当模型确实产生了推理内容且解析后的 `reasoning` 配置允许时才会发出;输入侧的关闭方式是 `model_options.reasoning.enabled: false`):
- **Chat Completions**:推理增量以 `choices[0].delta.reasoning_content` 块的形式到达(DeepSeek 风格的字段,Open WebUI、opencode 和 Vercel AI SDK 会将其渲染为思考块);回答文本仍留在 `delta.content` 中。