OpenRouter and the Nous Portal replay reasoning_details for multi-turn reasoning
continuity; every other OpenAI-compatible route either ignores the field or, when
its schema is strict (Groq, Mistral, Cerebras, opencode relays), rejects the whole
request with 400/422 once an earlier reasoning turn is in history — wedging the
session after an in-session model switch (#70233). Strip the field from the wire
copy in ChatCompletionsTransport.convert_messages (keyed on the target base_url),
mirror it in the auxiliary wire boundary and the iteration-summary path; state.db
history keeps the field so switching back to OpenRouter/Nous replays it again.
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.
- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".
Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
Follow-up on the salvaged #112721 commits (@fangliquanflq):
- agent/turn_context.py, tui_gateway/methods_prompt.py: add ``session_id`` to the explicit
snapshot dicts instead of switching them to ``agent._current_main_runtime()``. The titling
prologue is duck-typed (its tests drive a minimal stub) and the snapshot values feed the
``runtime_validator`` equality checks, where ``_current_main_runtime()``'s ``"" `` for a
missing attribute would no longer match the live ``None``. Same outcome — the background
request inherits the conversation's ``x-opencode-session`` — without changing what the
callers read.
- tests/agent/test_opencode_session_affinity.py: the salvaged title test passed on an
unfixed tree because the titler thread republishes the conversation contextvar
(``set_conversation_context``) and the affinity header falls back to it. Replace it with
two invariant tests (sync + async ``call_llm(main_runtime=...)`` with EVERY ambient source
unset via the ``out_of_turn`` fixture): red on origin/main, green here; both also pin
that the explicit binding does not leak past the call.
- website/docs/integrations/providers.md: name the background/out-of-turn auxiliary calls
the header now covers.
Fixes#112717
- build_api_messages reads agent._current_turn_timestamp directly: a caller that skipped the
turn prologue now raises instead of silently falling back to per-request wall time, which
would re-create the mid-turn drift the fix removes. Only production caller
(assemble_api_request) runs after _reset_per_turn_agent_state; cross-reference to the
tripwire _inflight_turn_started so the two clocks are not "unified" by mistake.
- is_dangerous_confirmation lowercases once instead of once per pattern (now on the per-request path).
- Tests: one _send(idx=) helper instead of three spellings of the builder call; the
untrustworthy-stamp contract is its own test.
Follow-up to the two cherry-picked commits from #105308 (@JoaoMarcos44), closing the
three blockers raised on that thread plus one regression the salvage found:
- Prefix-only on the send path. build_api_messages now canonicalizes only
messages[:current_turn_user_idx]; rows the current turn appended (its own tool
calls/results) pass through verbatim. Canonicalizing the live tail rewrote a block
the previous iteration had already sent whenever a tool result matched the
interrupt heuristic, which is exactly the mid-turn prefix rewrite this fix exists
to remove, and it also made the dangling-tail transform order-dependent on when
the user row was appended.
- Exact interrupt marker. is_interrupted_tool_result matched
"exit_code" + ("130" | "-1") + "interrupt" as substrings, so an ordinary
`grep KeyboardInterrupt` result next to a diff hunk header rewrote a terminal
result to an orphan notice (or dropped a read-only block). That heuristic was
tolerable at resume time only; it now runs per request. Match the executors'
bracketed markers ("[Command interrupted", "[execution interrupted") and nothing else.
- Admission-time clock. The frozen expiry clock was the input's platform-event
stamp, so a message queued 70 s before the turn ran kept a 129 s-old confirmation
live on the send path while replay expired it. _reset_per_turn_agent_state stamps
time.time() once at admission; the three other writes (bind identity, stage
message, build_api_messages write-back under suppress(Exception)) are gone.
- Fail closed on corrupt stamps. A present-but-unparseable timestamp (`"nan"`,
`"not_a_number"`) made strip_stale_dangerous_confirmations keep the confirmation
and its api_content sidecar. Coerce through hermes_cli.timefmt.coerce_epoch and
treat an unknowable age as expired; missing stamps (legacy rows) are still left
alone.
- Shape: drop the canonicalize_history_for_send alias (no consumer, never existed on
main), the `now=` kwarg (no production caller), and the getattr/hasattr rewrite of
_reset_per_turn_agent_state (only the test double needed it).
- Tests: 17 → 2 invariant tests. Real SessionDB round trip → canonicalize →
ChatCompletionsTransport bytes, equal to the send path with sidecars applied and
the durable list untouched, live tail preserved; admission-clock freeze across
iterations + corrupt-stamp fail-closed. Each is red under the matching mutation
(send path unpatched, whole-list canonicalization, per-request clock, fail-open,
loose heuristic).
The Telegram adapter asks the authorization check before dispatch, the ingress gate asks it
again, and the busy path asks a third time. Each call counted one loop-guard event, so a
Telegram bot tripped the budget after a third of the configured messages. The verdict now only
refuses a chat that is cooling down. The ingress gate counts an admitted bot message once.
`parse_turn_author` treats only booleans, integers and the strings true/1/yes as a bot flag,
and returns None for an author with neither id nor name. Names keep format characters and
non-breaking spaces so emoji sequences survive. The quiet one-shot pops HERMES_TURN_AUTHOR
before the turn so tool subprocesses do not inherit it. `max_events` must be a whole positive
number. Issue numbers move out of code comments.
on_turn_start already received the author trio. sync_turn did not, so a
provider that wanted to write the turn under its author had to stash state
between the two hooks. sync_turn now takes turn_author as a keyword-only
argument, and MemoryManager sends it only to providers whose signature
accepts it, so existing providers keep working unchanged.
build_turn_context resets the author on the agent at the start of every turn
so a cached gateway agent never carries a bot author into the next human
turn. agent/turn_author.py holds the parsing and the HERMES_TURN_AUTHOR
carrier.
MemoryProvider.identity_signature() is a new optional hook: the identity
values a provider writes under, declared by the provider itself, for the
gateway's agent cache to key on.
`on_turn_start` documents a per-turn kwargs channel — "kwargs may include:
remaining_tokens, model, platform, tool_count" — and `MemoryManager`
forwards whatever it receives. Its only caller passed nothing, so a memory
provider had no way to learn who wrote the turn it was being told about.
Providers that key durable state on identity resolve one identity when the
session is created. A shared session does not work that way: threads are
shared by default (`thread_sessions_per_user` is False), so alice, bob, and
another agent all write turns into a session whose peer is whoever spoke
first. The gateway's answer today is the `[name]` prefix it prepends to the
message text, which the model reads and a provider cannot.
`turn_author` now travels from the gateway through `run_conversation` into
`build_turn_context`, which forwards `author_id`, `author_name`, and
`author_is_bot` to every provider. It stops there — the trio never reaches
the model, and providers that ignore the kwargs are unaffected.
The bot flag is sent on every transport, not only shared sessions: a
provider deciding whether a turn may write to durable memory needs it in a
DM too.
`SessionSource.is_bot` is only as good as its producers. `build_source`
defaults it to False and 3 of 32 adapter call sites pass it, so most
platforms still report every author as human. Populating the rest is
follow-up work; nothing here depends on the flag being right yet.
Follow-up to #106312. _preflight_timeout_result carries the prior history without this turn's
user row (#7100); with a repeated prompt ("continue") the verbatim scan resolved to the
historical copy and exported it as this turn's proven boundary — the exact relabeling the export
exists to prevent. Nothing is exported for that envelope now.
The trailing `agent._persist_user_message_idx = idx` ran after finalize_turn had already flushed
the transcript, so it never influenced a persist and the next turn reset it: dead state, removed.
Hosts that settle their own transcript by index (hermes-webui) cannot prove which
row of result["messages"] is the current user turn once this loop rewrote history
(alternation repair, compaction, post-turn micro-compaction): the instance-side
_persist_user_message_idx predates those rewrites, and a text match relabels an
identical historical prompt and claims its old answer. Only the producer can
assert the coordinate against the exact list it returns.
run_conversation now wraps the turn (_run_conversation_turn) and stamps the pair
through export_current_turn_boundary on every envelope that leaves the loop
(success, partial/error, interrupt, retry-exhausted, tool-limit, preflight
timeout, codex runtime), computed on the final messages after finalize_turn and
micro-compaction. The pair is exported only when the addressed row is this turn's
user message verbatim (reanchor's last-match rule); a rewritten row exports
nothing so hosts fail closed. The final index is mirrored into
_persist_user_message_idx for the persist override.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013gp366ijf39n4UUtJhZuMh
The salvaged commit forked an explicit /refine under a new "refine_review" origin so
the memory delete gate would not treat it as unattended. But is_background_review()
is the key for every other review guard — skill_manager_guards (curator-owned-only,
read-before-write), skill_manager_tool (archive instead of rmtree), skill_ledger
actor, write_approval staging, the [auto] tag — so a /refine fork silently escaped
all of them.
Carry attendedness separately: the fork keeps origin "background_review" and sets
_review_attended; turn_context binds it beside the origin ContextVar; the memory
gate keys on the new is_unattended_review(). Also run the gate AFTER
_validate_single_op / the operations list check, as memory_tool's own docstring
requires, so an invalid replace is rejected now rather than staged and failed at
approve time.
_session_db is always a SessionDB (agent_init / delegate_tool), so the
"fail closed on a store wrapper" hasattr was defense around code that
cannot fail; the store's own guard binds the value into SQL, so the
prologue only needs the sibling idiom isinstance(_row_id, int) that
session_persistence and transcript_repair already use. The positional
hazard is explained once, on set_latest_user_api_content. The in-place
compaction test duplicated test_api_content_sidecar's
test_inplace_compaction_backfills_sidecar_into_db verbatim (its row_id
parameter was never varied); dropped, as was the positional-helper tail
of test_older_identical_row_is_untouched already covered there.
The turn-start stamp had grown its own copy of the "what does the current
user row hold" rule (persist override = clean transcript, live content =
wire bytes = sidecar when they differ) that _db_flush_row already
implements. Two copies drift; extract durable_user_row_content() in
session_persistence and call it from both.
Also: reuse _persist_lock() instead of a third open-coded lock/nullcontext
ladder; drop the hasattr guard on set_latest_user_api_content (it predates
this fix and exists on every SessionDB); cut the comment to the WHY;
trim the new test file from 18 cases to the 7 invariants (real close
flush E2E, repeated-"ok" positional protection, API-only pre-flushed
turn, normal path writes nothing, compaction keeps positional, store
guards). Still 3 red / 4 green when agent/turn_context.py is swapped
for main's copy.
_stamp_api_content_sidecar read _row_id without holding
_session_persist_lock. A close/early flush holds that lock while it
commits the row and only afterwards writes _row_id back onto the live
dict; a stamp that ran in between saw no id and skipped the backfill,
the flush finished with api_content = NULL and marked the message
persisted, and the turn-start persist skipped it — the row kept the
wrong bytes with no writer left to fix it.
Run the _row_id read and the DB backfill under the (re-entrant) lock,
re-checking _row_id after acquiring it. Race reported by @ehz0ah on
Co-authored-by: sal <141555468+salch-cred@users.noreply.github.com>
#102411; same fix shape as @salch-cred's follow-up on #103721.
The api_content sidecar ('persist what you send') preserves prompt-cache
stability across turn boundaries by persisting the exact API-bound bytes
(including memory-manager prefetch, plugin injections, and API-only notes)
and substituting them on replay.
When a user turn was already materialized in the database before the
sidecar could be composed (in-place preflight compaction or a close/early
flush racing the prologue on the CLI path), the turn-start crash persist
marker-skips that message. Previously, the backfill was gated strictly on
in-place compaction (_preflight_compressed and _last_compaction_in_place),
so racing CLI flushes left api_content = NULL in SQLite and broke prompt
caching on subsequent turns (#102194).
Positional approaches (such as #102239 and #102286) using LIMIT 1 on the
newest active user row are unsafe: repeated common inputs ('ok', 'yes',
'continue') cause the backfill to match and overwrite the PREVIOUS turn's
row with the new turn's sidecar, corrupting history and breaking cache parity.
Resolve all landing blockers and review feedback from #102411:
1. Bounded state owner (Sahilvishnaliya):
Add SessionDB.set_message_api_content(session_id, row_id, content, api_content)
to SessionMessagesMixin in hermes_state_messages.py instead of growing
hermes_state.py. Update set_latest_user_api_content docstring with durable
warning on the positional hazard.
2. API-only turns & durable content selection (ehz0ah):
When a pre-flushed clean input has an API-only difference (e.g. voice
prefix or model-switch note):
- Retain the differing API-facing bytes as api_content even when no
new memory or plugin context was injected.
- Derive the durable content guard using _override_replaces_content so
the SQL 'content IS ?' guard matches the clean override text stored
in the DB row rather than the restored wire text.
3. Turn prologue gating (_row_id) & fail-closed store duck-typing (ehz0ah):
In agent/turn_context.py::_stamp_api_content_sidecar: check _row_id on
the live user dict (stamped by _insert_message_rows and synced by
sync_flushed_message_markers). If valid (positive int, not bool), address
by exact ID. Do NOT fall back to positional matching when a row ID is
present: if an external or custom wrapper lacks set_message_api_content,
fail closed and skip rather than corrupting a neighbouring row. If absent
but in-place compacted, fall back to positional update. On normal turns,
skip the backfill entirely (single atomic INSERT).
4. Real lifecycle test coverage (salch-cred, ehz0ah):
Comprehensive tests in tests/agent/test_api_content_row_addressed_backfill.py
covering store guards, surrogate scrubbing, gate non-arming, older identical
row protection, real close-flush row_id synchronization, API-only clean
override preservation with exact wire replay, and duck-typed store fail-closed
verification when set_message_api_content is absent.
Fixes#102194.
Closes#102411.
- _transcript_row_texts re-implemented agent.message_content.flatten_message_text
and the api_content sidecar rule; the note can only land on a user row,
so the transcript scan now skips assistant/tool rows (the bulk of the bytes).
- Three sites computed "names of agent.tools"; tools.mcp_tool_agent gains
agent_tool_names() used by the switch note and conversation_loop, which
also stops importing the private _def_name across modules. The name list
is only captured when a switch was announced.
- split_runtime_boundary() is the single owner of the runtime-block
rpartition/END check for both identity_line_value and
_stored_prompt_matches_runtime.
- platform_surface_hint was a public alias of _platform_hint; the function is
now platform_hint (its docstring pointed at the pre-move module).
- consume_gateway_turn_context_notes and consume_surface_switch_note share
_pop_turn_note so the two one-shot channels have identical semantics.
- platform check hoisted above the transcript scan.
Move the six surface-switch helpers out of the conversation_loop facade
into agent/surface_switch.py (AGENTS.md: new behaviour goes in a topical
sibling), and fold the review findings on #104494:
- MoA and codex_app_server turns never stamp the api_content sidecar, so
the staged note could not be read back from the transcript and was
re-sent on every turn after a switch. Those modes now skip the note
(stored prompt still reused).
- The announced surface was parsed with split(".") — a plugin platform
with a dot in its name would never compare equal and re-stage the note
every turn. The note now closes the name with a fixed terminator.
- One identity-line parser (identity_line_value) shared by
_stored_prompt_matches_runtime and the switch detector instead of two
copies of the runtime-boundary/rpartition logic; tool names via the
existing tools.mcp_tool_agent._def_name; the transcript scan is bounded
to the last 200 rows (it ran every turn over the whole history).
- consume_surface_switch_note reduced to a plain pop; developer-guide
prompt-assembly.md updated (Platform is no longer an identity field);
17 new tests trimmed to 10 (same-shape pin/retire variants folded).
Restoring Platform as an identity field still turns 5 tests red.
`_stored_prompt_matches_runtime` treated `Platform` as a runtime-identity field, so
answering a live session from another surface — desktop -> TUI, or a resume after a
dashboard restart whose chat is a PTY TUI child — declared the stored prompt stale and
rebuilt it. The system prompt is the first thing in the request, so changing any byte of
it moves the first divergent byte to the head of a 220K-token request and the entire
conversation behind it re-prefills: a session that was hitting 240000/240287 came back
at 1536/219861.
The guard was not wrong about correctness — a desktop-built prompt on a terminal session
advertises inline widgets and a MEDIA: channel the TUI does not have — but the surface is
advisory metadata about the renderer, not a cache domain. Model/provider and cwd drift
change what the prompt should SAY; the surface changes only one paragraph.
Reuse the stored bytes across a surface switch and correct the paragraph where it costs
nothing to cache: `_stage_surface_switch_note` stages a one-shot note carrying the CURRENT
surface's guidance on the same per-turn user-message channel the gateway's must-deliver
notes use. It lands after the cached prefix and is stamped into the byte-stable
`api_content` sidecar, so later turns replay it instead of re-prefilling, and the prompt
converges at the next compaction — a boundary that already breaks the cache.
The saved tool_names prefix is not pinned across a switch: the tool registry is
process-global, so `_merge_preserving_prefix` would carry a saved-but-unloaded tool
forward (under `coding_context: focus` desktop gets a desktop_ui toolset the TUI cannot
run). On the same surface the tools freeze is untouched.
Namespace delivery markers and assign fresh keyless turn identities instead
of inferring ownership from IDs or process-local row baselines. Query only
marker existence on the canonical live compression continuation and ancestors.
Preserve raw reply IDs and exclude metadata from provider wire messages.
Expand the two existing invariants with resumed cross-chat ID collisions,
a real independent SQLite writer, reaped siblings, and archived-history
allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass.
Give dispatcher-owned workers a tool-capable reporting opportunity before the
hard iteration cap, without accepting arbitrary diffs or weakening failure
counting. Add opt-in per-turn iteration checkpoints for ordinary agents.
Persist checkpoint text with the fresh tool result, never rewrite cached rows.
Salvages the opt-in ratio and per-turn reset implementation from #104683;
credits the earlier default-off signpost proposal in #92438.
Local fixture wire A/B: Kanban ready/1 failure -> done/0; deliberately stuck
workers still reach blocked/2 after two runs. Default-off control unchanged.
Targeted and affected-directory suites queued behind campaign test lock.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: C. Michael Gibbs <252231331+MikeGibbsOnyx@users.noreply.github.com>
Track raw task identities across an agent's turns and match them against
process owner_task_id during close. Session IDs and shared terminal keys
are not process ownership, so the old bulk cleanup missed delegated work.
Preserve parent/sibling processes and consume teardown notifications.
Move task-resource cleanup into the lifecycle mixin, add real-process
isolation regressions, and document background process lifetime.
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).
The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.
- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
read the same bound value, so trigger and walk agree; the per-message memo now caches text
tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.
evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.
Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).
Now there is one authority:
- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
over threshold waits ONE request for the provider's real count instead of compressing on a guess
(first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
(note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
past the whole window, and provider-proven overflow all compress immediately; the post-compaction
latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
that scripted whole-history estimates now state the fact they relied on (provider omits usage).
Fixes#103391 (closes#103397 by construction — the baseline it repaired no longer exists).
The usage anchor (real usage.prompt_tokens + delta estimate of what was appended since)
identified the priced transcript by id() of the last message, so it was None on EVERY
gateway turn (history is re-read from the DB each turn) and in every fresh process
(--resume, desktop per-turn serve). Those are exactly the surfaces where the bytes/4
estimate then fired local compression against payloads the provider priced far under
threshold (#99421, #104462).
- agent/usage_anchor.py owns the anchor: content fingerprint instead of id(), persisted on
the session row (model_config._usage_anchor) via set_usage_anchor(), restored on the first
resumed turn while the durable transcript still matches, cleared with the row on
compaction / codex-native rewrite / session reset.
- Callers repointed from model_metadata (the compat table follows).
Design and persistence slot from #99585 by @686f6c61; re-authored against the Sep 2026
layout (the branch predates the model_metadata / agent_init split).
A native Responses compaction checkpoint is opaque ciphertext; the rough
preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against
a 204K trigger) and fires local compression on a request whose real prompt is
~116K. Arm the existing one-response real-usage latch when a replayable
checkpoint is captured (build_assistant_message) or restored into a fresh
agent (_hydrate_from_history), honor it in the post-tool gate and idle
compaction, and require non-empty encrypted_content for a checkpoint.
Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365,
bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by
patch application. Source delta is byte-identical to the PR head d6ce3e236d.
Fixes#100611
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
- turn_context_compaction: shared home for _clear_overflow_warn,
_reset_retry_state_after_compaction, _blocked_compress_reason,
_apply_grown_window, _refund_api_call (used by turn_preflight and the
turn-start preflight passes).
- turn_context: _with_persist_lock + _clear_staged_cli_input_if_persisted +
two try/except/finally wrappers -> one _persist_under_lock; MoA sanitize
model extracted to _sanitize_model_for; predicates flattened.
- turn_preflight: PreflightVerdict merged into PreflightGateVerdict (the
gate's carrier) and mutated in place; duplicate guard chain + refund blocks
collapsed onto the shared helpers.
- turn_preflight_gate: builds the carrier once, no re-unpack of the inner
verdict.
- turn_iteration_prep: step-callback tool-round scan and steer injection lifted
into helpers; dead _verdict closures on always-fallthrough paths removed.
Idle compaction (`compression.idle_compact_after_seconds`) skipped work only
when the context was at or below `threshold_tokens × summary_target_ratio`.
That is a *theoretical* post-compaction target: a real pass lands well above
it because the system prompt, the tool schemas and the protected head/tail
are an incompressible floor.
So a session that compacted to well above the target stayed above it forever,
and every later idle resume re-ran a full summarisation over a transcript that
had not grown. In the reported session the 17:01 pass reduced 64,105 -> 44,579
tokens; the 17:36 resume re-fired because 44,579 was still above the 25,502
theoretical floor, blocking the prompt for another 256 s on a ~55 tok/s local
route and reclaiming nothing. Neither pass was followed by a single API call.
`_should_idle_compact` now also honours `last_compression_rough_tokens` — what
the previous pass on this session actually produced, recorded by
`compress_context` with the same `estimate_request_tokens_rough` shape the idle
estimate uses. When it is known, the transcript must accumulate at least one
`floor_tokens` worth of new content on top of it before another pass is worth
its wall clock. A raised floor is a deferral, not an off switch.
`0` — nothing compacted yet, or the counter cleared by a rebind/recalibration
— keeps the original semantics exactly, so the first idle compaction of any
session is unaffected. The call site type-pins the read so compressor doubles
that expose a Mock there fall back to the original floor.
Fixes a root cause behind #97239. Deliberately does not touch the TUI status
surface (#97253), the digest chain / total ceiling (#93241), or turn-settle
scheduling (#96891).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HV3k1v5nx9ag5d5wXSFZ7o
The per-turn MCP refresh re-derives `agent.tools` from live availability and
publishes the result wholesale. Two kinds of bytes move as a result:
* a tool whose `check_fn` merely flapped (headless browser probe, expired
credential, docker blip) disappears from the array, and
* a late-landing MCP tool splices into sorted position, which can be index 0.
Providers that render `tools` ahead of the messages re-prefill the entire
history behind any moved byte, so either case costs a full re-prefill of the
session — the measured 2% cache hit in #100336. The caller's own comment
claimed the refresh "only ever extends a fresh request prefix"; it did not.
`refresh_agent_mcp_tools(..., preserve_prefix=True)` makes that claim true.
The live order becomes authoritative: existing tools keep their slot (fresh
schemas still land), a tool that is still registered but momentarily
unavailable is carried forward, a tool that genuinely left the registry is
still dropped, and new tools are appended at the tail. Explicit `/reload-mcp`
and the compaction boundary keep the plain rebuild.
Refs #100336