OpenRouter and the Nous Portal replay reasoning_details for multi-turn reasoning
continuity; every other OpenAI-compatible route either ignores the field or, when
its schema is strict (Groq, Mistral, Cerebras, opencode relays), rejects the whole
request with 400/422 once an earlier reasoning turn is in history — wedging the
session after an in-session model switch (#70233). Strip the field from the wire
copy in ChatCompletionsTransport.convert_messages (keyed on the target base_url),
mirror it in the auxiliary wire boundary and the iteration-summary path; state.db
history keeps the field so switching back to OpenRouter/Nous replays it again.
test_residency_header_from_jwt_claims asserted the shared helper directly, so
reverting the hermes_cli/codex_models.py or agent/model_metadata.py call sites
to base stayed green. Drive _fetch_models_from_api (httpx.get) and
_fetch_codex_oauth_context_lengths_with_source (requests.get) with a
compute-residency-only JWT and assert x-openai-internal-codex-residency and
ChatGPT-Account-ID reach the outgoing headers. Red with either call site
reverted to base; green at head.
Five change-detector tests collapse into one positive invariant (data
residency wins, compute residency is the fallback, the sibling /usage
header builder derives the same value) and one control (no claim or a
malformed token never carries the header). The residency value must
otherwise be probed live; see the PR body.
The Codex models catalog probe (agent/model_metadata.py), the picker catalog
fetch (hermes_cli/codex_models.py), the quota-restored probe
(hermes_cli/auth_codex.py) and the /usage dashboard call
(agent/account_usage.py) each re-decoded the OAuth JWT for
ChatGPT-Account-ID and would 401 on residency-enforced workspaces exactly
like the chat client did (#23896). They now share
agent.codex_headers.codex_account_headers, which emits the account and
x-openai-internal-codex-residency headers from one decode; the three
duplicate `_extract_chatgpt_account_id` decoders are gone.
account_usage keeps auth.json's account_id as the winner over the JWT claim
(pool-only credentials still omit it); header casing unifies on the
codex-rs canonical `ChatGPT-Account-ID` (HTTP header names are
case-insensitive on the wire; the two tests asserting the old casing follow).
Extract chatgpt_data_residency (fallback: chatgpt_compute_residency) from
the OAuth JWT and send as x-openai-internal-codex-residency header on
requests to chatgpt.com/backend-api/codex. Residency-enforced workspaces
return 401 without this header.
Fixes#23896
The source-aware classifier replaced base's '401 unauthorized' needle with a
\b401\b regex on the primary error, so any unrelated primary message
carrying a bare 401 token (e.g. 'request body exceeded limit by 401 bytes')
produced the re-login hint and retired the session. Restore the base needle
alongside 'unauthorized' (which already covers the issue's case) and drop
the regex and its 're' import; pin the negative case.
Two invariants for #75167, replacing the seven change-detector tests from
#75182: the classifier table (issue's four cases, parametrized) and the
session path (turn/start and thread/compact/start RPC errors keep the
primary error and stderr tail visible, no retire, no `codex login` hint).
The existing empty-input test moves to the keyword-only `stderr=` form.
`_classify_oauth_failure` joined the primary JSON-RPC error and the codex
stderr tail into one haystack and matched broad tokens ("unauthorized",
"401 unauthorized", "oauth"). codex writes independent ChatGPT plugin
prewarm failures ("HTTP 401 Unauthorized") to stderr while the core
JSON-RPC server keeps working, so any unrelated RPC error, timeout or
subprocess exit was rewritten into the `codex login` hint and the real
error plus stderr tail disappeared.
Classify by source: generic 401/unauthorized/oauth text is authoritative
only in the operation's own error; ambient stderr needs a strong
credential signal (invalid_grant, refresh/expired token, no auth profile).
Every call site (turn error, request timeout, dead subprocess, and the
compaction paths that share them) passes stderr by keyword.
Salvaged from #75182 by @cosin2077, hand-applied onto the refactored
session module (call sites collapsed into `_set_classified_error` /
`_request_for` / `_subprocess_died`).
Fixes#75167
A fallback_providers entry naming a `providers.<name>` block (or
`custom:<name>`) without its own api_mode was re-detected from the
resolved host: an Anthropic-Messages proxy on a plain host or a
Responses-only relay behind a generic gateway landed on chat_completions
while resolve_provider_client had already built the declared client
(#33062 bottom thread, #81932 provider-level `transport:`). The hint pass
now reads the named block's api_mode/transport as explicit, and entry-level
`transport:` is accepted as an alias of `api_mode` with the same
canonicalisation the providers block uses (`responses` -> codex_responses).
Fixes#33062Fixes#81932
The direct-call helper test stayed green when convert_tools / the auxiliary
adapter re-flattened strict to False. Drive ResponsesApiTransport.build_kwargs
and _CodexCompletionsAdapter._build_responses_kwargs instead and assert an
explicit strict: True reaches kwargs['tools'] on both routes (#105401 parity).
`agent.reasoning_effort: {enabled, effort}` now reads back as its tier name
in the TUI `config.get reasoning` result and in the setup wizard's
"currently in use" lookup instead of `str(dict)` (a disabling dict reads
as `none`). Both route the dict through `parse_reasoning_effort` so the
dict semantics live in one place. Docs note the dict form is config.yaml-only.
`agent.reasoning_effort` and `agent.reasoning_overrides` values now accept
`{enabled: true, effort: <level>}`; the level is passed through verbatim to
the wire (CustomProfile / clamp_effort already forward unknown names), so a
relay exposing `fast`/`thinking` can be asked for its real tier instead of
silently running at the default `medium`.
Bare strings stay strict: a non-ladder string is still rejected with the
existing warning, so a typo like `hgih` never reaches a request. Only the
config parser (`hermes_constants.parse_reasoning_effort`) rejected custom
names — the transport layer was already designed to pass them through.
Slim redo of the change proposed in PR #93239 (the base function had since
been compacted, so the hunk no longer applied); docs and example config
updated in the same change.
Co-authored-by: HermesDev-Bot <309177324+HermesDev-Bot@users.noreply.github.com>
_newest_reasoning_only pops older turns' codex_reasoning_items before the
converter runs, so those turns looked reasoning-free and replayed their
msg_* id with neither the reasoning item nor its rs_* id on the wire — the
exact orphan shape #97427 rejects. The trimmed row now carries a transient
codex_reasoning_trimmed marker and _replay_message_items honours it; the
single-use _turn_has_encrypted_reasoning wrapper is inlined into that
predicate. Test drives ResponsesApiTransport.build_kwargs on an Azure host.
Stateless Responses replay (store=False) strips every reasoning item's rs_*
id, but the assistant message minted in the same response kept its msg_* id.
GPT-5.6-family endpoints validate that link and reject the continuation with
HTTP 400 "Item 'msg_…' of type 'message' was provided without its required
'reasoning' item: 'rs_…'" on every post-tool turn, deterministically.
The converter now drops the message id whenever the stored turn carried
encrypted reasoning (replayed, suppressed by recovery, or dropped as a
foreign-issuer blob), keeping content/status/phase. Reasoning-free turns keep
their id for prefix-cache affinity. Done in _replay_message_items so the main
transport, preflight and the auxiliary Codex adapter all emit the same shape.
Co-authored-by: salehelsayed <saleh.fekry@gmail.com>
The expiring-singleton status test never reached the singleton resolver:
`load_pool("openai-codex")` mirrors the singleton as a `device_code` pool
entry, so with a still-valid token `pool.peek` answered and the
`get_codex_auth_status()` wiring under test was never exercised (the test
stayed green with `resolve=` reverted). The token is now already expired,
the status result is asserted to come from `hermes-auth-store`, and the
`read_only` beats `force_refresh` call is a secondary assertion on the
resolver itself.
`resolve_codex_runtime_credentials(read_only=True)` reads the store without
`_auth_store_lock`: `_save_auth_store` replaces auth.json atomically, so a
lock-free read never sees a torn file, and materialising `auth.lock` is
itself a write a diagnostic must not make (#68004). Once the pool has
mirrored the singleton a status read leaves the HERMES_HOME manifest
byte-identical.
`_codex_catalog` (the `/model` picker) now reports the stored login
read-only, matching the intent stated on `read_only`; an expired stored
token yields the hardcoded catalog until the runtime lease refreshes it.
`hermes status` / `hermes doctor` / the dashboard cards and the `/model` picker
call `get_codex_auth_status()`, whose singleton fallback ran
`resolve_codex_runtime_credentials()` with runtime defaults: a store missing
its refresh_token imported the Codex CLI's single-use token family, and an
expiring token was refreshed and written back. A diagnostic that spends or
borrows a rotating refresh token logs the other program out (#68004).
`resolve_codex_runtime_credentials(read_only=True)` reports the stored state
as-is (no CLI adoption, no refresh, no pool forced refresh, no write) and wins
over `force_refresh`; the Codex status snapshot uses it, and the xAI OAuth
snapshot passes `refresh_if_expiring=False` for the same reason. The pool side
was already an observation (`peek`, 964fbaae2c).
Superseded #68224 (@GauravPatil2515) — same mechanism, re-done on the current
`auth_codex` layout.
Co-authored-by: Gaurav Patil <gauravpatil2516@gmail.com>
The wrapped-message-only branch of _is_codex_masked_replay_rejection (SDK paths
that surface no parsed body) had no test; it joins the existing parametrize as
codex-message-only. Neutering the branch turns that id red.
The second failure mode in #51512: with no reasoning replay at all, a single
``{"role": "user", "content": "<text>"}`` item still 400s on the ChatGPT Codex
backend with ``{"detail": "Unsupported content type"}``. The classifier mapping
from the previous commit cannot recover that turn (turn_recovery's strip needs
cached codex_reasoning_items), so the wire shape has to be right up front.
``_chat_messages_to_responses_input`` now wraps string user/assistant text as
``input_text`` / ``output_text`` parts when the issuer is ``codex_backend``; every
other Responses route keeps the string shorthand it has always received. The
preflight already validates typed parts, so the real call path (build_kwargs ->
preflight_kwargs) needs no separate change.
Test drives ResponsesApiTransport.build_kwargs + preflight_kwargs, red on the
old head for the codex case. Three test_native_compaction asserts pinned the
assistant string incidentally (they check history survives, not its shape) and
now expect the typed part on the codex route.
Two Responses-API 400s meant "the replayed encrypted reasoning was rejected"
but never reached the one-shot recovery in turn_recovery (disable replay,
strip codex_reasoning_items, retry):
- OpenAI's ``thinking_signature_invalid`` code contains "thinking" and
"signature", so the Anthropic thinking-block heuristic claimed it first;
that recovery strips Anthropic fields and resends the same stale encrypted
item on every retry (#70595).
- The ChatGPT Codex backend returns a bare ``{"detail": "Unsupported content
type"}`` for the same rejection, which matched nothing and aborted the turn
as a non-retryable format_error (#51512). It joins the #92353 exact-envelope,
provider-gated codex mapping, so a no-replay 400 keeps aborting as before
(the recovery still requires cached codex_reasoning_items).
Co-authored-by: ooiuuii <al3060388206@gmail.com>
A finite `hermes chat -q` / `--oneshot` run has no later session in its HERMES_HOME
to learn for, yet it ran the full interactive self-improvement loop. Measured over
21 one-shot benchmark trajectories: 7 skills created and a bundled one patched
mid-task, 37 of ~215 tool calls on skill_view/skill_manage, skill text = 34% of all
tool-result bytes fed back into context, plus reviewer subagents spawned on the
agent's own diff (one task: 5 delegations, 62 subagent API calls, each re-paying a
cold system prompt).
Keyed on the existing HERMES_SINGLE_QUERY_SESSION marker (approval gate, delegation
dispatcher), so interactive and gateway sessions are byte-identical:
* agent/oneshot_footprint.py (new sibling): skill_manage is pruned from the tool
set; the ## Skills block keeps the index + skill_view but drops the record/patch/
offer-to-save coaching and the "load process skills for work you already know"
push (SKILLS_GUIDANCE follows because it is gated on skill_manage).
* delegation.oneshot_max_children (default 2, 0 = unlimited): total children a
one-shot run may spawn; past it delegate_task returns a tool error telling the
model to finish inline.
* requesting-code-review skill: reviewer/fixer subagents (Steps 5 and 7) are
interactive-only; one-shot applies the checklist inline.
Live: `hermes chat -q "list tools starting with skill_"` on the portal — base
"skill_manage, skill_view, skills_list", fix "skill_view, skills_list"; a real
AIAgent under a temp HERMES_HOME shows the interactive prompt unchanged (6643 chars
both) and the one-shot prompt without skill_manage / offer-to-save.
The three-child live test sampled /api/health/idle once, the instant each
child printed READY. With HERMES_DESKTOP=1 the child runs the in-process cron
ticker, and every tick holds retirement admission for its whole scan (#98745,
on purpose), so a sample inside a tick says `retirement_admission` for an idle
resident, or names the admission instead of `cron:<job>` for the busy child.
~1 in 10 runs locally, twice in a row on PR CI at 96 workers.
The contract is "stable between ticks", so poll until the verdict names the
stable state (idle for residents, the cron ledger for the busy child); the last
verdict is still asserted, so a genuinely wrong state fails. 20/20 green after.
Telemetry writes a bundled skill's usage record the moment it is seeded, so by
the curator's first pass with prune_builtins on, created_at can be months old and
every never-used built-in is marked stale at once (71 on the reporting install,
#79295). The first-sight seed only covered records that did not exist yet.
apply_automatic_transitions now re-anchors a bundled, never-used record without
first_seen_at to now (skill_usage.reanchor_clock) and defers it one pass; the
stamp makes it one-shot so the skill still ages out after a full window. A record
the bug already marked stale is reactivated in the same step (the reviewer's
migration edge on #79311), instead of waiting a week for the next pass.
Design and first implementation by @webtecnica in #79311; re-applied by hand
because that branch predates the skill_usage decomposition.
Co-authored-by: webtecnica <webtecnica@gmail.com>
Second half of #107539 on top of @liuhao1024's snapshot_paths filter: the same
TRANSIENT_DIRS set (venv, node_modules, caches, .git) now gates the whole-tree
snapshot (directory parts only, so a file named `venv` is still skill content),
and rollback carries a live nested venv back the way it already carries `.git`.
Nothing ever pruned the blob store; `hermes curator ledger --compact` now
deletes every blob no ledger entry references (98.9% of 47k blobs on the
reporting install). A malformed ledger line aborts the sweep — an unreadable
entry may still hold references.
Supersedes the curator_backup half of #100545 / #81669.
snapshot_paths() hashed every file under a skill dir with no exclusion
filter, so a stray venv/node_modules/__pycache__/.git under a skill was
copied content-addressed into ~/.hermes/.curator_backups/blobs/ on every
mutation — and nothing ever prunes that store, so the blobs grew without
bound (real install: 47k blobs / 1.3 GB in one day, 98.9% unreferenced).
Filter at capture: skip any file whose parent-chain component matches
_SNAPSHOT_EXCLUDE_DIRS (the same set proposed for the curator tarball in
transient dir is still captured — only files inside those dirs are
dropped.
The unreferenced-blob GC suggested in the issue is intentionally left
out; capture-side filtering stops the growth, and pruning existing
garbage is a separate, riskier change.
- config.yaml -> `_apply_yaml_config` seeding for discord.free_response_auto_thread
(the operator path, previously only proven by an ad-hoc probe)
- `auto_thread: false` still disables threading with the opt-in on
Adds the key to cli-config.yaml.example and the multi-profile per-key list,
records the env-over-YAML precedence, and drops the duplicated
no_thread_channels clause in the new section.
- clear DISCORD_FREE_RESPONSE_AUTO_THREAD in the adapter fixture, so a contributor
shell exporting it can no longer decide the default-path test
- no_thread_channels precedence now shows the flip inside one test rather than
passing on the default path alone
- voice-linked channels stay inline with the opt-in on: once skip_thread is cleared
that exclusion is the only thing holding them back
The quick-start table and the auto_thread section still promised inline replies
unconditionally, and configuration.md's discord block omitted the key. Adds a
discord.free_response_auto_thread section with the precedence rules (no_thread_channels
wins, voice-linked channels unaffected).
Mirrors the _discord_require_mention()/_discord_max_attachment_bytes() shape, makes the
lookup lazy (it only runs on the free-channel path now), and keeps the in-code key
manifests in _handle_message and register() listing the new key.
Free-response channels skip auto-threading by default so the bot replies
inline (lightweight chat mode). This prevented users who wanted BOTH
mention-free replies AND per-conversation threads from getting either.
Add a new opt-in `discord.free_response_auto_thread` (env:
`DISCORD_FREE_RESPONSE_AUTO_THREAD`, default false) that, when true,
re-enables auto-threading in free-response channels. Voice-linked
channels continue to skip auto-thread regardless, and the flag is
gated behind the global `DISCORD_AUTO_THREAD=true`.
Default behavior is unchanged; all 291 existing discord tests pass.
Review findings folded into the exception added by the previous commit:
- The anchored-region check now reuses `_walk_tail_budget` instead of a
second token sum with the default thought-charge rule. The walk charges
thinking only on the newest assistant turn unless the route replays stale
thinking (#73624/#84371), so the second rule could fire the exception on a
region the walk itself considers inside the ceiling.
- Two conjuncts (`last_user_idx >= head_end`, `last_user_idx < cut_idx`) are
implied by `user_anchored_cut < cut_idx`, which only changes when the
anchor found a real user turn inside the compressible region; the check
that depended on them is documented where it is read.
- The latest-assistant anchor is no longer computed and discarded under the
exception (it also logged an anchor it never applied).
- `_find_last_user_message_idx` early-exits instead of materialising every
actionable user index on the per-attempt boundary path.
- The exception logs at debug, like the sibling anchor decisions, rather
than at info from both `_compress_window` and `has_content_to_compress`.
Adds the missing guard test for the new ceiling check: a transcript that
fits the tail budget must still anchor the active request rather than
triggering the exception. Red when the ceiling conjunct is removed.
The oversized-turn exception (#80449) relaxed three tail anchors at once,
which voided two guarantees that hold on main:
- `compression.min_tail_user_messages` was skipped whenever the exception
fired, so an N-user tail could come back with no user turn at all. The
N-user anchor now always runs: the setting is a user-facing promise and
outranks the budget.
- The exception fired even when the oversized weight was the active turn's
own newest tool group. The pre-anchor cut retains that group anyway, so
the split bought no reclaim while taking the active request out of the
tail and losing the #10896 anchor. The exception now requires real turn
body (a tool-call group) between the opening request and the cut.
Both failures are bound by tests, each red when its conjunct is removed:
`TestTailTokenBudgetCeiling::test_message_floor_does_not_unboundedly_override_soft_ceiling`
and `TestMinTailUserMessages::test_n_guarantee_wins_over_tail_token_budget_and_floor`
(each passes on main), plus `test_n_user_tail_guarantee_outranks_the_split`
for the N-user guarantee on a genuinely oversized turn.
Second review round on the previous commit, both findings reproduced:
- `startswith(prefix, head_chars)` still exempted the imitation shape #83714 is
about: a replayed leaf of head + marker followed by new content was never
shrunk again (a 5,278-char leaf stayed 5,278). The guard now also requires
the marker to close the leaf, so that shape shrinks to head + marker with
true counts.
- Per-leaf savings were compared in characters, but the final re-serialise
adds separator whitespace, so compact args with many keys could come back
LONGER (measured 3,511 -> 4,110 chars) and be counted as reclaimed pressure.
The helper now returns the caller's string unless the whole rewrite is a net
reduction.
Tests cover both shapes plus a many-key compact payload.
Review findings on the previous commit, all reproduced:
- The "already marked" guard was a substring test, so a leaf that merely
contains the marker — including one a model imitated into a new call, the
#83714 failure mode itself — was exempt from shrinking forever. The marker
is always written at `head_chars`, so the guard now tests that position: a
1,550-char imitated leaf shrank to 423 again.
- The helper re-serialised even when nothing was replaced, so compact wire
JSON came back with inserted spaces and callers read it as a change
(rewriting replayed history and counting a pressure hit for zero reclaim).
Nothing replaced now returns the original string: a 547-char compact blob
is byte-identical.
The anti-imitation marker is ~220 chars, longer than the 200-char head it
follows, which made the replayed-arg rewrite unsound in two ways:
- A leaf just over `head_chars` came back LONGER: a 628-char args blob
shrank to 455 chars before this change and grew to 863 after it.
- The result was not a fixed point. `_shrink` re-ran on every later
compaction, so a leaf's marker was rewritten from the true count
("2,800 of 3,000 chars omitted") to a self-referential one
("223 of 423"), destroying the per-instance count property the marker
relies on and churning those bytes on each pass.
A leaf now keeps its head+marker replacement only when that is strictly
shorter, and an already-marked leaf is left alone. The two tests pin the
behaviour contract: never grow, and `shrink(shrink(x)) == shrink(x)`.
Root cause for #83714 (write_file/patch_tool writing literal
"...[truncated]" into files, PR #83752's guard is the safety net, not
the fix): _truncate_tool_call_args_json() in the compression pass
shrinks long string values inside a PAST assistant message's
tool_calls[].function.arguments — the exact field that represents the
model's own prior generated output, replayed back to it verbatim on
every subsequent turn. The old marker, a bare "...[truncated]" suffix,
is indistinguishable from something the model itself could have
written (it's exactly the kind of terse ellipsis abbreviation models
already produce). A model conditioned on seeing itself "get away with"
that pattern in its own history imitates it in a new tool call,
writing the literal marker instead of real content.
This is the second bug from the same root text. The first (#11762,
MiniMax 400s from unterminated JSON) was fixed by shrinking inside the
parsed structure so the JSON stays valid, but kept the same visible
marker text — fixing the syntax problem while leaving the imitation
problem untouched.
Fix: replace the marker with one deliberately NOT shaped like prose a
model would write — distinctive non-ASCII delimiters, an explicit "not
part of the original tool call" disclaimer, and a per-instance
char-count that won't match the next omission point even if copied
verbatim. The shrunk value stays a plain string (not a nested object)
so the #11762 valid-JSON/matching-shape contract is unchanged — only
the marker text changed.
Checked context_compressor.py's other "...[truncated]" call sites
(_serialize_for_summary, _compact_fallback_turn, the user-message-only
one near _ACTIVE_TASK_MAX_CHARS) — none of them write into a value
that gets replayed as the main model's own assistant/tool_calls
history, so they don't share this priming risk and were left as-is.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The per-file method() wrapper only composed the two registry decorators;
methods_free_tier/_complete/_config_set/_session_control stack them directly.
Test: drop the manual multiplex save/restore (tests/conftest.py resets the
latch per test) and the inert OP_SERVICE_ACCOUNT_TOKEN setenv.
methods_vault.py carried its own scoping decorator that returned early for the
launch profile (`_profile_home()` → None). That is correct only while the
process serves one profile: the first secondary home served flips
`set_multiplex_active(True)` and unscoped `get_secret()` reads fail closed, so
every vault.sources / vault.list call for the launch profile (the Desktop sends
no `profile` for it) raised UnscopedSecretError inside
OnePasswordLoginBackend.__init__ and the Passwords & Logins panel showed
"Could not load vault items" until `hermes gateway restart`.
Route the handlers through server.py's `_profile_scoped` via
`HandlerRegistry.profile_scoped`, the path every other methods_*.py uses: it
binds `launch_secret_scope()` for the launch profile under multiplex and also
honours a session_id-only call. Unknown `profile` names now surface as the
dispatcher's 4064 instead of a vault 5095.
Regression test: vault.sources for the launch profile with multiplex active
(red on main with the production traceback). E2E A→B→A over
`handle_request` with two on-disk homes resolves ops_LAUNCH_A / ops_WORK_B /
ops_LAUNCH_A.
* feat(mcp): add the official n8n server to the catalog
Connect to the user's instance over HTTP with browser OAuth. Keep a
separate n8n-official identifier so the retired n8n bridge is neither
relabeled nor overwritten and retains its credentials and tool filter.
Save ordinary catalog setup values in server config, retaining secret
references in .env. Pass field secrecy through the catalog API so URLs
and client IDs remain visible while credentials stay masked.
Keep the existing install-then-authorize lifecycle. Transactional setup
and cancellation changes are outside this catalog addition.
* refactor(mcp): limit n8n PR to catalog addition
Remove shared installer, storage, field-masking, and input-handler changes.
Those behaviors are being handled in a separate PR. Restore their tests
and Asana guidance to the base branch.
Keep only the official n8n manifest and setup documentation, using the
catalog's existing setup and persistence behavior.
Stop offering the third-party bridge for new catalog installs. Existing
connections keep their saved transport, credentials, and tool selection;
the runtime and configured-server controls do not require a manifest.
Update CLI examples and document that catalog reinstall is unavailable.
Adding n8n's official server remains separate work.
Three passages assumed an uncapped 1M trigger (grok 375K example, legacy tail size, the 850K -> 512K
feasibility example) and the delegation doc said children compact at the ratio only.
`preview_threshold_tokens` restated the resolve -> floor -> compute -> cap chain that `update_model`
runs; two copies of the trigger math drift the next time a step is added — the exact bug class
#83450 fixes (the guard quoting a number the compressor will not install). `_derive_trigger` is the
single pure derivation; the auxiliary-summariser ceiling stays in `_apply_threshold_tokens_cap`
because it is per-runtime, not per-model.
The startup banner names the cap only when it set the trigger; on windows where the ratio already
sits below it "(capped at 256,000)" was noise. Comments no longer repeat the default literal.