From b2963f8034bcd2e5f6b6c32cef2a11ced187242d Mon Sep 17 00:00:00 2001 From: Teknium <127238744+teknium1@users.noreply.github.com> Date: Sun, 2 Aug 2026 15:38:52 -0700 Subject: [PATCH] docs: document agent.session_stall_timeout + compression timeout keys (re-review #7) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - website/docs/user-guide/configuration.md (en) and the zh-Hans translation gain a 'Session Stall Watchdog' section: default 300, 0=disabled, notify-only semantics (never kills the turn — contrast gateway_timeout), one notification per stall episode, and the exact stall message text so it is greppable. - cli-config.yaml.example: the two in-agent compression timeout keys (compression.context_timeout_seconds / compression.context_total_ceiling_seconds) are shown as commented lines next to session_stall_timeout's example for discoverability. --- cli-config.yaml.example | 9 +++++++++ website/docs/user-guide/configuration.md | 19 +++++++++++++++++++ .../current/user-guide/configuration.md | 19 +++++++++++++++++++ 3 files changed, 47 insertions(+) diff --git a/cli-config.yaml.example b/cli-config.yaml.example index bbffe27087..04ccfb6da4 100644 --- a/cli-config.yaml.example +++ b/cli-config.yaml.example @@ -818,6 +818,15 @@ agent: # the turn (see gateway_timeout). 0 = disable. Default 300. # session_stall_timeout: 300 + # Related in-agent compression timeouts (they live under the top-level + # compression: block, shown here for discoverability next to the stall + # watchdog they complement — a hung compression is a common stall cause): + # compression: + # context_timeout_seconds: 120 # inactivity budget for in-agent + # # compress_context (0 = disable) + # context_total_ceiling_seconds: 600 # absolute cap on the pre-commit + # # wait even while tokens stream + # Graceful drain timeout for gateway stop/restart (seconds). # Default 0 = no drain: a restart interrupts in-flight agents immediately, # cleans up, and exits. Set a positive value only if you want a grace diff --git a/website/docs/user-guide/configuration.md b/website/docs/user-guide/configuration.md index a65f8e88ed..85d321335f 100644 --- a/website/docs/user-guide/configuration.md +++ b/website/docs/user-guide/configuration.md @@ -881,6 +881,25 @@ Points at a custom OpenAI-compatible endpoint. Uses `OPENAI_API_KEY` for auth. The summary model **must** have a context window at least as large as your main agent model's. The compressor sends the full middle section of the conversation to the summary model — if that model's context window is smaller than the main model's, the summarization call will fail with a context length error. When this happens, the middle turns are **dropped without a summary**, losing conversation context silently. If you override the model, verify its context length meets or exceeds your main model's. ::: +## Session Stall Watchdog + +The gateway runs a notify-only stall watchdog (`agent.session_stall_timeout`, default `300` seconds, `0` = disabled). When a busy session has a **pending inbound follow-up** and the agent's shared activity clock has been idle for at least this long, the gateway logs a WARNING and sends the user a one-shot notification: + +``` +⚠️ Agent session appears stalled (last activity N min ago). Try /new to reset. +``` + +Semantics: + +- **Notify-only.** The watchdog never kills the turn — contrast `agent.gateway_timeout`, which cancels a run after prolonged inactivity. The stall notice just tells you the agent looks wedged so you can decide (`/new`, `/stop`, or keep waiting). +- **One notification per stall episode.** The latch clears when the pending inbound drains or activity resumes, so a session that recovers and stalls again notifies again. +- Progress comes only from the shared activity snapshot (tool calls, API stream progress, compression heartbeats). Pending inbound is a notify gate, not a progress clock. + +```yaml +agent: + session_stall_timeout: 300 # seconds; 0 disables the watchdog +``` + ## Context Engine The context engine controls how conversations are managed when approaching the model's token limit. The built-in `compressor` engine uses lossy summarization (see [Context Compression](/developer-guide/context-compression-and-caching)). Plugin engines can replace it with alternative strategies. diff --git a/website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/configuration.md b/website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/configuration.md index 5ec1239468..4f7f03909f 100644 --- a/website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/configuration.md +++ b/website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/configuration.md @@ -665,6 +665,25 @@ auxiliary: 摘要模型**必须**具有至少与您的主 agent 模型一样大的上下文窗口。压缩器将对话的完整中间部分发送给摘要模型 —— 如果该模型的上下文窗口小于主模型的,摘要调用将因上下文长度错误而失败。发生这种情况时,中间轮次将**在没有摘要的情况下被丢弃**,静默丢失对话上下文。如果您覆盖模型,请验证其上下文长度满足或超过您的主模型。 ::: +## 会话卡死监视器(Session Stall Watchdog) + +Gateway 运行一个仅通知的卡死监视器(`agent.session_stall_timeout`,默认 `300` 秒,`0` = 禁用)。当一个忙碌的会话存在**待处理的入站后续消息**,且 agent 的共享活动时钟空闲达到该时长时,gateway 会记录一条 WARNING 日志并向用户发送一次性通知: + +``` +⚠️ Agent session appears stalled (last activity N min ago). Try /new to reset. +``` + +语义: + +- **仅通知。** 监视器绝不会终止当前轮次 —— 对比 `agent.gateway_timeout`(长时间无活动后取消运行)。卡死通知只是告诉您 agent 看起来卡住了,由您决定(`/new`、`/stop` 或继续等待)。 +- **每个卡死周期只通知一次。** 待处理入站消息被消化或活动恢复时闩锁清除,因此恢复后再次卡死的会话会再次通知。 +- 进度仅来自共享活动快照(工具调用、API 流式进度、压缩心跳)。待处理入站消息是通知门槛,不是进度时钟。 + +```yaml +agent: + session_stall_timeout: 300 # 秒;0 禁用监视器 +``` + ## 上下文引擎 上下文引擎控制在接近模型 token 限制时如何管理对话。内置的 `compressor` 引擎使用有损摘要(参见[上下文压缩](/developer-guide/context-compression-and-caching))。插件引擎可以用替代策略替换它。