fix(gateway): bounded redacted result preview on tool.completed run events

Slims the salvaged preview helper (drop the try/except around json.dumps —
default=str cannot raise on tool results) and documents the tool.completed
SSE shape. Adds the control test that the multimodal envelope dict is still
classified as a success, so the dict passthrough only widens the failure
detection to real structured results.

The idea of carrying the tool result on tool.completed for /v1/runs
consumers was first proposed in #22362; that PR's executor half
(result=function_result) is already on main, and its wire half is landed
here in redacted, bounded form instead of the raw payload.

Salvages #111821 (@KoNit-K), part of #111815.

Co-authored-by: kidrauhl123 <105764349+kidrauhl123@users.noreply.github.com>
This commit is contained in:
teknium1
2026-09-15 11:51:53 -07:00
committed by Teknium
parent c89209d564
commit 339fa6d918
3 changed files with 21 additions and 10 deletions

View File

@@ -51,18 +51,14 @@ _TOOL_COMPLETED_PREVIEW_MAX_CHARS = 500
def _tool_completed_preview(result: Any, redact_sensitive_text: Callable[..., str]) -> str:
"""Return a bounded, secret-redacted completion summary for public run events."""
"""Bounded, secret-redacted result summary for the public run stream — redacted BEFORE
truncation so a cut never leaves a secret's prefix on the wire."""
if result is None:
return ""
if not isinstance(result, str):
try:
result = json.dumps(result, ensure_ascii=False, default=str)
except (TypeError, ValueError):
result = str(result)
preview = redact_sensitive_text(result, force=True)
if len(preview) > _TOOL_COMPLETED_PREVIEW_MAX_CHARS:
return preview[:_TOOL_COMPLETED_PREVIEW_MAX_CHARS - 3] + "..."
return preview
text = result if isinstance(result, str) else json.dumps(result, ensure_ascii=False, default=str)
preview = redact_sensitive_text(text, force=True)
limit = _TOOL_COMPLETED_PREVIEW_MAX_CHARS
return preview if len(preview) <= limit else preview[: limit - 3] + "..."
def _remember_room_retention(request: "web.Request", claims: dict[str, Any]) -> None:

View File

@@ -73,10 +73,16 @@ class TestDetectToolFailureTerminal:
assert "Terminal backend degraded:" not in suffix
def test_nonzero_dict_result_is_a_failure(self):
"""An already-parsed terminal result (a plugin tool_execution middleware may hand back the
dict instead of the JSON string) must classify like its JSON form (#111815)."""
is_failure, suffix = _detect_tool_failure("terminal", {"exit_code": 2})
assert is_failure is True
assert suffix == " [exit 2]"
def test_multimodal_envelope_dict_is_still_a_success(self):
envelope = {"_multimodal": True, "content": [{"type": "text", "text": "screenshot taken"}]}
assert _detect_tool_failure("browser_exec", envelope) == (False, "")
def test_degraded_backend_without_hint_shows_reason_alone(self):
result = json.dumps({"output": "", "exit_code": -1, "status": "degraded",
"reason": "SSH connection to bob@host timed out", "retry_hint": "",

View File

@@ -475,6 +475,15 @@ Statuses are retained briefly after terminal states (`completed`, `failed`, or `
Server-Sent Events stream of the run's tool-call progress, token deltas, and lifecycle events. Designed for dashboards and thick clients that want to attach/detach without losing state.
Tool lifecycle events carry `tool.started` (`tool`, `preview` of the arguments) and
`tool.completed` (`tool`, `duration` in seconds, `error`, and a `preview` of the result). The
`error` flag reflects the tool's own outcome — a non-zero terminal `exit_code`, a structured
`{"error": ...}` result, a denied approval — whether the result arrives as a JSON string or an
already-parsed object. The completion `preview` is the result text (structured results are
JSON-encoded), passed through forced secret redaction and then truncated to 500 characters, so a
client can tell an approval refusal (`BLOCKED: ...`) from an ordinary failure without receiving
the unbounded tool payload.
When the agent delegates work to background subagents, the stream also carries
`subagent.start` and `subagent.complete` lifecycle events, so clients can
observe delegation outcomes — including timeouts and failures — instead of the