fix(api_server): honour platforms.api_server.tool_progress_events for Chat Completions SSE (#12020)
Strict OpenAI clients reject the named hermes.tool.progress SSE frames. Setting tool_progress_events: false under platforms.api_server (loaded into PlatformConfig.extra by from_dict) now drops them; default stays on. Reimplements the intent of #42640 against the adapter config actually read in production. Overlaps #49069 (erikerosev). Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
This commit is contained in:
@@ -1199,6 +1199,11 @@ class APIServerAdapter(OpenAICompatRoutesMixin, BasePlatformAdapter):
|
||||
# @mssteuer.)
|
||||
self._direct_model_requests: bool = _coerce_request_bool(
|
||||
extra.get("direct_model_requests"), default=False)
|
||||
# ``platforms.api_server.tool_progress_events: false`` drops the custom
|
||||
# ``hermes.tool.progress`` SSE frames from Chat Completions streams for strict OpenAI
|
||||
# clients that choke on named events (#12020). Default on.
|
||||
self._tool_progress_events: bool = _coerce_request_bool(
|
||||
extra.get("tool_progress_events"), default=True)
|
||||
self._app: Optional["web.Application"] = None
|
||||
self._runner: Optional["web.AppRunner"] = None
|
||||
self._site: Optional["web.TCPSite"] = None
|
||||
|
||||
@@ -882,6 +882,8 @@ class OpenAICompatRoutesMixin:
|
||||
break
|
||||
if isinstance(delta, tuple) and len(delta) == 2 and delta[0] == "__tool_progress__":
|
||||
# Custom event: tool lifecycle for frontends without markers in history.
|
||||
if not self._tool_progress_events:
|
||||
continue # opted out for strict OpenAI clients (#12020)
|
||||
await response.write(_sse_frame(delta[1], event="hermes.tool.progress"))
|
||||
elif isinstance(delta, tuple) and len(delta) == 2 and delta[0] == "__reasoning__":
|
||||
# DeepSeek-style ``delta.reasoning_content`` (#99552), the field Open WebUI,
|
||||
|
||||
48
tests/gateway/test_chat_completions_tool_progress_optout.py
Normal file
48
tests/gateway/test_chat_completions_tool_progress_optout.py
Normal file
@@ -0,0 +1,48 @@
|
||||
"""``platforms.api_server.tool_progress_events: false`` drops the custom
|
||||
``hermes.tool.progress`` SSE frames from streaming Chat Completions (#12020)."""
|
||||
|
||||
import asyncio
|
||||
from unittest.mock import AsyncMock, MagicMock, patch
|
||||
|
||||
|
||||
def _stream_body(platform_cfg):
|
||||
from aiohttp import web
|
||||
from gateway.config import PlatformConfig
|
||||
from gateway.platforms.api_server import APIServerAdapter, ThreadSafeAsyncQueue
|
||||
|
||||
# from_dict is the production loader: a top-level platform key lands in ``extra``.
|
||||
adapter = APIServerAdapter(PlatformConfig.from_dict(platform_cfg))
|
||||
written = []
|
||||
|
||||
async def fake_agent():
|
||||
return {"final_response": "done", "completed": True}, {
|
||||
"input_tokens": 1, "output_tokens": 1, "total_tokens": 2}
|
||||
|
||||
async def run():
|
||||
stream_q = ThreadSafeAsyncQueue()
|
||||
stream_q.put_nowait(("__tool_progress__", {"tool": "terminal", "toolCallId": "c1", "status": "running"}))
|
||||
stream_q.put_nowait("done")
|
||||
stream_q.put_nowait(None)
|
||||
agent_task = asyncio.ensure_future(fake_agent())
|
||||
resp = AsyncMock(spec=web.StreamResponse)
|
||||
resp.write = AsyncMock(side_effect=lambda data: written.append(data))
|
||||
resp.prepare = AsyncMock()
|
||||
req = MagicMock()
|
||||
req.headers = {}
|
||||
with patch("gateway.platforms.api_server.web.StreamResponse", return_value=resp):
|
||||
await adapter._write_sse_chat_completion(req, "cmpl-1", "m", 1, stream_q, agent_task)
|
||||
|
||||
asyncio.run(run())
|
||||
return b"".join(written).decode()
|
||||
|
||||
|
||||
def test_tool_progress_frames_emitted_by_default():
|
||||
body = _stream_body({"enabled": True, "token": "k"})
|
||||
assert "event: hermes.tool.progress" in body
|
||||
assert '"content": "done"' in body
|
||||
|
||||
|
||||
def test_tool_progress_events_false_suppresses_frames_but_keeps_content():
|
||||
body = _stream_body({"enabled": True, "token": "k", "tool_progress_events": False})
|
||||
assert "hermes.tool.progress" not in body
|
||||
assert '"content": "done"' in body
|
||||
@@ -111,7 +111,7 @@ Uploaded files (`file` / `input_file` / `file_id`) and non-image `data:` URLs re
|
||||
All SSE streams (Chat Completions, Responses, `/api/sessions/{id}/chat/stream`, `/v1/runs/{id}/events`) emit a `: keepalive` comment line whenever no event has been sent for 10 seconds, so long tool calls do not trip client idle timeouts. Standard SSE clients ignore comment lines; custom parsers must skip lines that start with `:`.
|
||||
|
||||
**Tool progress in streams**:
|
||||
- **Chat Completions**: Hermes emits `event: hermes.tool.progress` for tool-start visibility without polluting persisted assistant text.
|
||||
- **Chat Completions**: Hermes emits `event: hermes.tool.progress` for tool-start visibility without polluting persisted assistant text. Strict OpenAI clients that choke on named SSE events can turn these frames off with `gateway.platforms.api_server.tool_progress_events: false` (default `true`); content chunks are unaffected. The opt-out applies only to Chat Completions — `/v1/runs/{id}/events` always emits tool events, which is what the `tool_progress_events` feature in `/v1/capabilities` describes.
|
||||
- **Responses**: Hermes emits spec-native `function_call` and `function_call_output` output items during the SSE stream, so clients can render structured tool UI in real time.
|
||||
|
||||
**Model reasoning** (emitted only when the model actually produces reasoning and the resolved `reasoning` config allows it; the input-side opt-out is `model_options.reasoning.enabled: false`):
|
||||
|
||||
@@ -109,7 +109,7 @@ curl http://localhost:8642/v1/chat/completions \
|
||||
**流式传输**(`"stream": true`):返回逐 token 响应块的 Server-Sent Events(SSE)。对于 **Chat Completions**,流使用标准 `chat.completion.chunk` 事件,以及 Hermes 自定义的 `hermes.tool.progress` 事件用于工具启动的 UX 展示。对于 **Responses**,流使用 OpenAI Responses 事件类型,如 `response.created`、`response.output_text.delta`、`response.output_item.added`、`response.output_item.done` 和 `response.completed`。
|
||||
|
||||
**流中的工具进度:**
|
||||
- **Chat Completions**:Hermes 发出 `event: hermes.tool.progress` 以提供工具启动可见性,同时不污染持久化的 assistant 文本。
|
||||
- **Chat Completions**:Hermes 发出 `event: hermes.tool.progress` 以提供工具启动可见性,同时不污染持久化的 assistant 文本。无法处理具名 SSE 事件的严格 OpenAI 客户端可设置 `gateway.platforms.api_server.tool_progress_events: false`(默认 `true`)关闭这些帧;内容块不受影响。该开关仅作用于 Chat Completions——`/v1/runs/{id}/events` 始终发出工具事件,`/v1/capabilities` 中的 `tool_progress_events` 功能描述的正是这一点。
|
||||
- **Responses**:Hermes 在 SSE 流期间发出符合规范的 `function_call` 和 `function_call_output` 输出项,让客户端能够实时渲染结构化工具 UI。
|
||||
**模型推理**(仅当模型确实产生了推理内容且解析后的 `reasoning` 配置允许时才会发出;输入侧的关闭方式是 `model_options.reasoning.enabled: false`):
|
||||
- **Chat Completions**:推理增量以 `choices[0].delta.reasoning_content` 块的形式到达(DeepSeek 风格的字段,Open WebUI、opencode 和 Vercel AI SDK 会将其渲染为思考块);回答文本仍留在 `delta.content` 中。
|
||||
|
||||
Reference in New Issue
Block a user