feat(config): wire compression.tail_mode + docs (en/zh)
#87326 shipped the lean-compaction capability on the compressor; this adds the config.yaml surface (compression.tail_mode: legacy|lean, default legacy), DEFAULT_CONFIG entry, and docs on both the dev-guide compression page and the user-guide configuration page, with zh-Hans parity.
This commit is contained in:
@@ -1987,6 +1987,11 @@ def init_agent(
|
||||
compression_enabled = str(_compression_cfg.get("enabled", True)).lower() in {"true", "1", "yes"}
|
||||
compression_target_ratio = float(_compression_cfg.get("target_ratio", 0.20))
|
||||
compression_protect_last = int(_compression_cfg.get("protect_last_n", 20))
|
||||
# Tail retention mode (compression.tail_mode). "legacy" (default) keeps
|
||||
# the 0.20*window verbatim tail; "lean" switches to the clamped
|
||||
# 2.5%/10K-25K tail with recovery-pointer machinery (#87326). Unknown
|
||||
# values fall back to legacy inside the compressor.
|
||||
compression_tail_mode = str(_compression_cfg.get("tail_mode", "legacy")).strip().lower()
|
||||
# Minimum REAL (actionable) user messages guaranteed to survive in the
|
||||
# uncompressed tail (compression.min_tail_user_messages). Default 1
|
||||
# preserves current behavior exactly — the existing single-user tail
|
||||
@@ -2617,6 +2622,7 @@ def init_agent(
|
||||
proactive_prune_min_result_chars=compression_proactive_prune_min_chars,
|
||||
proactive_prune_min_reclaim_tokens=compression_proactive_prune_min_reclaim,
|
||||
min_tail_user_messages=compression_min_tail_users,
|
||||
tail_mode=compression_tail_mode,
|
||||
)
|
||||
_bind_session_state = getattr(agent.context_compressor, "bind_session_state", None)
|
||||
if callable(_bind_session_state):
|
||||
|
||||
@@ -659,6 +659,17 @@ DEFAULT_CONFIG = {
|
||||
# threshold and this token count. Clamped to
|
||||
# the model's context length at apply-time.
|
||||
"target_ratio": 0.20, # fraction of threshold to preserve as recent tail
|
||||
"tail_mode": "legacy", # tail retention policy (#87326):
|
||||
# "legacy" — 0.20×window verbatim tail (default)
|
||||
# "lean" — clamped 2.5%-of-window tail
|
||||
# (10K floor / 25K cap) plus chunked
|
||||
# digests, a mechanical anchor index,
|
||||
# verbatim user messages, and
|
||||
# session_search recovery pointers in
|
||||
# the summary. ~3x fewer retained
|
||||
# tokens after compaction; costs a few
|
||||
# extra summarizer calls at the
|
||||
# compaction boundary.
|
||||
"protect_last_n": 20, # minimum recent messages to keep uncompressed
|
||||
"min_tail_user_messages": 1, # REAL (actionable) user messages guaranteed to
|
||||
# survive in the uncompressed tail. 1 = existing
|
||||
|
||||
@@ -86,6 +86,7 @@ compression:
|
||||
# "glm-5.2": 0.40 # longest key wins). See "Per-model threshold
|
||||
# "claude-sonnet": 0.35 # overrides" below.
|
||||
target_ratio: 0.20 # How much of threshold to keep as tail (default: 0.20)
|
||||
tail_mode: legacy # Tail retention policy: legacy | lean (default: legacy)
|
||||
protect_last_n: 20 # Minimum protected tail messages (default: 20)
|
||||
min_tail_user_messages: 1 # Real user messages guaranteed in the tail (default: 1)
|
||||
codex_gpt55_autoraise: true # gpt-5.5 on Codex OAuth: raise trigger to 85% (default: true)
|
||||
@@ -109,7 +110,8 @@ auxiliary:
|
||||
|-----------|---------|-------|-------------|
|
||||
| `threshold` | `0.50` | 0.0-1.0 | Compression triggers when prompt tokens ≥ `threshold × context_length` |
|
||||
| `model_thresholds` | `{}` | map | Per-model overrides of `threshold`. Keys are substring-matched against the model name (longest match wins). The small-context floor still applies on top (see below) |
|
||||
| `target_ratio` | `0.20` | 0.10-0.80 | Controls tail protection token budget: `threshold_tokens × target_ratio` |
|
||||
| `target_ratio` | `0.20` | 0.10-0.80 | Controls tail protection token budget: `threshold_tokens × target_ratio` (legacy mode only — `lean` uses its own clamp) |
|
||||
| `tail_mode` | `legacy` | `legacy`, `lean` | Tail retention policy. `legacy` keeps a `target_ratio`-sized verbatim tail (~100K+ tokens on big-window models). `lean` keeps a clamped tail of `2.5% × context window` (10K floor, 25K cap) and instead carries continuity in the summary: chunked identifier-preserving digests of the compacted region, a mechanically extracted anchor index (PR numbers, SHAs, paths, error strings — regex, never paraphrased), every real user message quoted verbatim (newest-first budget), and a `session_search` recovery pointer so the agent can re-access anything summarized away. Result on 500K-token real sessions: ~49K retained vs ~162K, with higher recall when paired with recovery (see `evals/compaction/results/`). Costs a few extra summarizer calls at the compaction boundary. Old tool results inside the lean tail are demoted to one-line stubs carrying a recovery pointer |
|
||||
| `protect_last_n` | `20` | ≥1 | Minimum number of recent messages always preserved |
|
||||
| `min_tail_user_messages` | `1` | ≥1 | Minimum number of REAL (actionable) user messages guaranteed to survive in the uncompressed tail. `1` = the existing single last-user anchor (behavior-preserving default). Raise to e.g. `3` to keep the last 3 real user turns verbatim even when bulky tool outputs fill the tail token budget. Blank platform echoes, compaction handoffs, and synthetic continuation rows never count toward N. The guarantee wins over the tail token budget — the tail may exceed the budget when the anchor pulls the cut back |
|
||||
| `protect_first_n` | `3` | (hardcoded) | System prompt + first exchange always preserved |
|
||||
|
||||
@@ -807,6 +807,7 @@ compression:
|
||||
threshold: 0.50 # Compress at this % of context limit
|
||||
threshold_tokens: null # Absolute token cap (optional) — takes lower of ratio vs absolute
|
||||
target_ratio: 0.20 # Fraction of threshold to preserve as recent tail
|
||||
tail_mode: legacy # Tail retention: "legacy" (0.20×window verbatim tail) or "lean" (clamped 2.5% tail, 10K-25K, with digests + anchor index + session_search recovery pointers in the summary — ~3x fewer retained tokens after compaction)
|
||||
protect_last_n: 20 # Min recent messages to keep uncompressed
|
||||
protect_first_n: 3 # Non-system head messages pinned across compactions (0 = pin nothing)
|
||||
in_place: true # Compact on the same session id (no rotation) — see below
|
||||
|
||||
@@ -80,6 +80,7 @@ compression:
|
||||
enabled: true # Enable/disable compression (default: true)
|
||||
threshold: 0.50 # Fraction of context window (default: 0.50 = 50%)
|
||||
target_ratio: 0.20 # How much of threshold to keep as tail (default: 0.20)
|
||||
tail_mode: legacy # 尾部保留策略:legacy | lean(默认 legacy)
|
||||
protect_last_n: 20 # Minimum protected tail messages (default: 20)
|
||||
|
||||
# Summarization model/provider configured under auxiliary:
|
||||
@@ -96,6 +97,7 @@ auxiliary:
|
||||
|-----------|---------|-------|-------------|
|
||||
| `threshold` | `0.50` | 0.0-1.0 | 当 prompt token 数 ≥ `threshold × context_length` 时触发压缩 |
|
||||
| `target_ratio` | `0.20` | 0.10-0.80 | 控制尾部保护 token 预算:`threshold_tokens × target_ratio` |
|
||||
| `tail_mode` | `legacy` | `legacy`、`lean` | 尾部保留策略。`legacy` 保留 `target_ratio` 大小的逐字尾部(大窗口模型约 100K+ token)。`lean` 保留截取后的尾部(窗口的 2.5%,下限 10K、上限 25K),并将连续性移入摘要:压缩区域的分块保标识符摘录、机械提取的锚点索引(PR 编号、SHA、路径、报错文本 — 正则提取,绝不改写)、逐字引用的全部真实用户消息,以及 `session_search` 恢复指引。在 500K token 的真实会话上:保留约 49K(对比 162K)。压缩边界会多出数次摘要模型调用。lean 尾部中较旧的工具输出会降级为带恢复指引的单行占位 |
|
||||
| `protect_last_n` | `20` | ≥1 | 始终保留的最近消息最小数量 |
|
||||
| `protect_first_n` | `3` | (硬编码)| 系统提示词 + 首次交互始终保留 |
|
||||
|
||||
|
||||
@@ -598,6 +598,7 @@ compression:
|
||||
enabled: true # 开启/关闭压缩
|
||||
threshold: 0.50 # 在上下文限制的此百分比时压缩
|
||||
target_ratio: 0.20 # 保留为最近尾部的阈值分数
|
||||
tail_mode: legacy # 尾部保留策略:"legacy"(0.20×窗口的逐字尾部)或 "lean"(截取 2.5% 窗口、10K-25K 上下限的精简尾部,摘要中附带分块摘录、锚点索引与 session_search 恢复指引 — 压缩后保留 token 约减少 3 倍)
|
||||
protect_last_n: 20 # 保持未压缩的最少最近消息数
|
||||
hygiene_hard_message_limit: 5000 # Gateway 安全阀 —— 见下文
|
||||
context_timeout_seconds: 120 # Agent 侧 compress_context 无进展超时(秒)—— 见下文
|
||||
|
||||
Reference in New Issue
Block a user