Follow-up to the salvaged #114480 (@jonameijers) and #114553 (@JoaoMarcos44).
A small summarizer can emit both the canonical "## Historical Task Snapshot"
and the legacy "## Active Task" heading in one summary (the 4B model in the
issue "duplicates the section set"). With count=1 the grounding pass replaced
only the first match and left the second as a live-looking, undisclaimed
task section. Replace the first task section with the deterministic snapshot
and drop every later one.
Tests: fold the two contributor test files into one file with two invariants
(prompt names the emitted heading; grounding collapses alias + duplicate
sections). The original #114480 test asserted the pre-fix prepend behaviour,
which #114553 makes false.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
Align the iterative update instruction with HISTORICAL_TASK_HEADING and
treat leftover ## Active Task sections as the same snapshot so grounding
replaces them instead of prepending a second live-looking task heading.
Fixes#114479
The iterative-update instruction told the summarizer to update
"## Active Task", but the template it is given emits HISTORICAL_TASK_HEADING
("## Historical Task Snapshot"). Leftover from #44454, which renamed the
heading in the template and SUMMARY_PREFIX but not in this literal.
A summarizer that follows the literal emits a section
_HISTORICAL_TASK_SECTION_RE cannot match, so
_ground_historical_task_snapshot prepends instead of replacing and the
summary carries two task sections. Only the grounded one is disclaimed by
SUMMARY_PREFIX, leaving the stale one readable as live work - the hijack
class #44454 closed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Both the auxiliary ladder (agent/auxiliary_client.py::_resolve_call_client) and main-agent
init (agent/agent_init.py::_routed_client_kwargs) told users to "Set the
<PROVIDER_ID>_API_KEY environment variable" when an explicit provider had no credentials.
Deriving the name from the id invents variables nothing reads: MINIMAX-OAUTH_API_KEY for
`minimax-oauth` (not even a valid shell name), ALIBABA_API_KEY where the registry reads
DASHSCOPE_API_KEY. agent_init already consulted PROVIDER_REGISTRY but fell back to the
invented name for OAuth providers, whose api_key_env_vars is deliberately empty.
One helper, agent/auxiliary_unavailable.py::missing_provider_credentials_message, now
builds the sentence for both surfaces from the registry: the first registered env var for
API-key providers, `hermes auth add <provider>` for OAuth providers, and only "switch
provider" for registry rows with neither (bedrock, vertex, external-process). The
compression permanent-failure classifier learns the new "no credentials were found" phrase
so an OAuth aux provider without a login still stops the retry loop.
Salvages #89517 (@liuhao1024, aux surface, earliest for #114405) and #79007 (@TUARAN,
main-init surface, #78996); supersedes #114410, #114430, #114582, #90222, #90281, #89541.
Co-authored-by: CodeMiner-掘金安东尼 <729922845@qq.com>
Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Co-authored-by: rocks737 <234251857+rocks737@users.noreply.github.com>
A turn-start preflight pass whose summary stalled had no deterministic exit: the wrapper handed the
transcript back unchanged, _fail_closed_after_preflight_timeout raised for ANY over-threshold request
(even one that fits the model window — #113646: 99K of a 120K window), the loop labelled it
compression_exhausted, the messaging gateway auto-reset the session (#114594), and the existing
deterministic escalation (DETERMINISTIC_SUMMARY_ROUTE, #112420) was gated on a PRIOR stall in the same
session — unreachable once the first stall had already wiped it. /compress rode the same wrapper, so the
suggested recovery reproduced the same loop.
- request_exceeds_model_window(agent, tokens): one predicate, two consumers.
- Fits the window: the request is sent uncompressed this turn (the cooldown-blocked path already does
exactly this every turn); the summary-failure cooldown stops the retry from repeating.
- Above the window: the stall retry ladder escalates to the deterministic fallback summary on the FIRST
stall (old tool results pruned, static handoff committed through the normal lease/fence pipeline).
compression_exhausted / auto-reset is the last resort, when even that cannot shrink the transcript.
- Per-attempt observable: one INFO line when the summary call is dispatched (model, prompt chars,
prompt build ms) so a stalled attempt is distinguishable from a slow prompt build.
- Deterministic-rung wording no longer claims "again after a stall backoff".
Live repro (real AIAgent + SessionDB + local OpenAI-compatible stand-in whose summariser never answers
within the idle budget): before — FITS(73K/200K) and OVER(73K/64K) both end failed=True,
compression_exhausted=True, main_calls=0; after — FITS completes with the request sent uncompressed
(main_calls=1), OVER commits the deterministic fallback (103->25 rows) and completes.
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).
- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
_apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
every runtime change: switch_model (outside the rollback guard), fallback activation
and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).
WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.
Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.
Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
`_image_payload` ran `json.dumps` over every image part on every request (pass runs on
each API call, before the HTTP client serializes the same bytes again). Measured per
request on the send-path clone, 20 images:
per image json.dumps len(data)
256 KB 50 ms 0.02 ms
1 MB 226 ms 0.01 ms
5 MB ~1.0 s 0.03 ms
origin/main's keep-3 walk was 0.4 ms. The image payload is ASCII base64, so the data-URL
(or Anthropic `source.data`) length is the byte count within ~17 bytes of JSON framing
per part — noise against a 24 MB budget with 8 MB of headroom.
The compressor pass and the Anthropic wire pass each carried their own copy of the
limit/batch/floor constants (under different names) and of the retire-count loop.
Retuning one alone would make the wire pass re-evict on a second frontier after the
compressor pass — the exact prefix-rewrite class #113517 fixes.
- agent/image_eviction_policy.py: stdlib-only leaf holding the constants and
outbound_image_retire_count(); the wire converter stays a leaf (no compaction import).
- evict_stale_outbound_tool_images loses `keep_newest`: no caller passed it, and its
default was the COMPACTION window constant the comments say the send path must not use.
- _retire_stale_tool_result_images (compaction) back to its main body; behaviour was
unchanged, so the rewrite was churn.
- One `_tool_result_parts` unwrap instead of three copies; one `_image_payload` walk
per message instead of two.
- Wire pass counts blocks via `_block_type` (non-dict inner blocks no longer raise).
Tests import the constants from the policy module; the two "stops at the floor" tests
assert the invariant (newest floor survives, request under the limit) instead of the
exact frontier. test_user_uploads_are_not_evicted now builds a request that actually
trips eviction — at 6 blocks it never fired and proved nothing.
Both eviction passes retire the oldest image-bearing tool results in whole
batches (8) and let the floor shelter a breach only when reserved uploads alone
exceed the ceiling. When uploads fill most of the ceiling without exceeding it,
the first batch step overshoots past the newest frames: fifteen uploads plus six
one-frame tool results retired all six (the model blind on the frame it had just
asked for) although keeping the newest three already clears the 20-block limit.
The batch step is now cut back to `total - floor` when it would cross the floor
and keeping the floor fits; a floor that does not fit (one 25-block tool_result,
byte pressure) still retires past it, and uploads alone over the ceiling still
shelter the floor. Pure-screenshot sessions are unchanged: same batch-aligned
frontier at 21/29/37 images.
Review-fix on #113519 (found while checking @ehz0ah's mixed-session boundary on
the sibling PR #113619).
Three residual witnesses from review on this branch and #113619.
W1 (@ehz0ah): reserved user uploads were counted against the BLOCK ceiling but
not against the byte budget, so five 5 MB uploads plus three 3 MB tool images
serialized to 34.0 MB against Anthropic's 32 MB Messages limit and eviction
returned pruned=0. The API answers 413 request_too_large.
W2/W3 (@andrexibiza, @teknium1): the floor applied unconditionally, so it
sheltered breaches the tool content itself caused. One tool_result carrying 25
image blocks left retire=0 because only one carrier existed against a floor of
three; four 10-block results retired one and kept 30 blocks.
Reserve upload bytes alongside blocks, and make the floor conditional on
satisfiability: it shelters a violation only when retiring every removable tool
image still cannot bring the request inside the ceiling, and never when bytes
are the breach, since the request-size limit is hard. Uploads are still never
rewritten, and the unsatisfiable case still keeps the newest frames.
W1 34.0 MB -> 25.0 MB, pruned=3
W2 25 blocks -> 0, pruned=1
W3 40 blocks -> 0, pruned=4
floor intact: 21 uploads + 5 screenshots keeps 3
Ref: https://platform.claude.com/docs/en/api/overview#request-size-limits
Two gaps in the limit-triggered eviction, both found by review on #113517.
The ceiling is a count of API image blocks, but the send path counted tool
MESSAGES. Eight tool results carrying three screenshots each are 24 blocks
against a 20-block ceiling; counting messages saw 8 and evicted nothing.
User uploads occupy the same per-request budget and were not counted at all, so
a mixed session sailed past the limit. They must be reserved against the ceiling
while never being rewritten -- the user attached them by hand.
Reserving uploads introduces a failure of its own: once uploads alone reach the
ceiling, a pure limit rule retires every screenshot and leaves the model blind
on the frames it was just asked about, which is worse than the stricter
dimension cap the limit exists to avoid. keep_newest therefore becomes a FLOOR
rather than a trigger.
Both stages now weight by image-part count and reserve non-tool image blocks.
Three regression tests cover the block accounting, the reserved-upload
accounting, and the floor; all three fail when either fix is reverted.
Review fix on top of the limit-triggered eviction. Both send-path stages ran
statelessly on a fresh clone every request, and each retired a FIXED single
batch (`min(batch, count - keep)`) once over the limit. That caps total
eviction at 8 per stage forever: once a session passes 20 + 8 (+ 8 from the
adapter) images, the outbound request grows unbounded again (24 image blocks
at 40 vision calls, 44 at 60 in the issue's own replay), so the documented
20-image threshold and the 24 MB byte budget were not enforced above the
first window. The PR's 40/60-row "3 collapses" result was this saturation,
not a held frontier under the limit.
Retire `ceil(overshoot / batch) * batch` instead, extending by whole batches
while the surviving payload still exceeds the byte budget. The frontier is
still a step function of the image count (one move per batch), and the
request now stays at or below the limit at every size:
images collapses max outbound images (PR head -> fixed)
40 3 -> 4 24 -> 20
60 3 -> 6 44 -> 20
2 MB x 30 3 -> 4 40 MB -> 22 MB
Both frontier tests now span three batch windows and assert the outbound
count never exceeds the limit; both are red on the previous head.
Both send-path image evictions retired the oldest payload as soon as more than
three images were present. Retiring an image edits a message the provider has
already cached, and Anthropic matches its prompt cache on an exact byte prefix,
so each retirement re-wrote the entire conversation. Because the window was a
count, one more message retired on every new image and nearly every turn became
a full-prefix miss.
Observed on a real 104-request session: requests reported cache_read=16,076 --
system prompt plus tool schemas and nothing else -- with cache_write up to
127,720, against a steady-state read of ~145,000. Those requests produced 62
percent of the session's entire cache-write volume.
Trigger eviction on the API's own per-request image and byte limits instead, and
retire a batch when it fires. Under the limit nothing is rewritten and the
request is append-only, so the cached prefix survives; over it, the cost is one
slower turn per batch rather than one per image. This matches the documented
behaviour of the reference Anthropic client.
Simulated across session sizes, prefix collapses fall from 3/8/18/38/58 to
1/1/1/3/3 at 5/10/20/40/60 vision calls.
Compaction keeps the count-based window: it commits its rewrite into the
canonical transcript once and has no prefix to preserve.
After a no-progress stall the ladder's first rung (60s) was shorter than the
default idle stall window (120s), so the next oversized automatic turn
re-entered the same silent summary route about a minute after burning the
whole window (#112420: "made no progress ... continuing without
compression" followed by "context compression started" 62s later).
record_timeout_failure now floors every timeout-class cooldown at the
configured idle window; a window below the rung leaves the ladder untouched.
The "Compaction rebuilt a drifted system prompt" line fired on every compact
of a long session (19x/day observed). The rebuild stays mandatory; only the
first drift per session is INFO, later ones DEBUG.
Both hunks re-applied from PR #112504 (@JoaoMarcos44); its no-LLM prune on
stall is a product decision and was not ported.
Part of #112420
Follow-up to the salvaged #112719 commit:
- `_skill_result_failure_suffix()` is shared by the `skill_manage` and
`skills_list` stubs instead of two inline copies, and also fires on a
payload that says `success: false` without an `error` string (previously
such a result still compressed into a success-shaped line — the exact
class the issue reports). The error preview is whitespace-collapsed so
the stub stays one line even when the tool error carries newlines.
- The `skills_list` comment no longer names a `query` arg the schema does
not have (`SKILLS_LIST_SCHEMA` takes `category` only).
- Tests trimmed to two invariants in the existing summarizer module
(`tests/agent/test_context_compressor.py`): skill_manage names its ops
(operations array and legacy flat shape) and keeps FAILED visible with
the error text; skills_list renders category/count and marks failure,
with skill_view as the unchanged control. The standalone 9-test file
from the contributor PR is folded into these.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
_sum_named read a top-level `name` arg that only skill_view has: skill_manage
names live in operations[] (or the legacy flat shape) and skills_list has none,
so every compressed line rendered `name=?`. The stub also ignored the result
payload, collapsing failed batches into success-shaped lines and dropping the
error text. Dedicated summarizers keep the op list and surface FAILED plus the
leading error text; _sum_named is gone.
Fixes#112710
context_breakdown._chars_to_tokens and native_compaction._approx_tokens did raw
chars//4, under-counting CJK/Cyrillic by 2-4x next to the conversation slice
that already used estimate_tokens_rough — the /context pie chart mixed two
estimators. Both now call the canonical. The four private `= 4` ratio
constants import one CHARS_PER_TOKEN from agent/model_metadata.py. Estimates
only feed UI and budgets; no prompt or message bytes change.
A bare `astra: 0.85` in compression.model_thresholds was written for the Codex
OAuth route, where Astra is capped at 272K and 50% would compact at ~136K. The
key is substring-matched on the model name alone, so it also fired on
openai/gpt-6-astra via OpenRouter and Nous, where the window is 1.1M: the user's
0.5 global threshold was silently replaced by 0.85 and the session sat at 620K
(~59%) without compacting.
Keys may now carry a provider prefix: `"openai-codex:astra": 0.85` applies only
when the session's provider is openai-codex; bare keys keep their route-agnostic
behaviour. Ranking is by model-substring length with scope as the tie-break, so
`astra-900k` still outranks `openai-codex:astra` for the 900K picker. The
provider flows through ContextCompressor (ctor + update_model), the ContextEngine
base class and the TUI hot-reload path, so a /model switch between routes
re-scopes the override.
_sum_clarify only extracted the top-level ``user_response`` key, so batch clarify
results (questions=[...] -> responses[].user_response) fell through to the generic
placeholder and the summarizer never saw the user's answer/permission decision.
Closes#106077.
Follow-up to #106317. Typing the steer row (display_kind="steer") for the renderer and the
alternation-repair guard collided with the convention that any display_kind on a user row means
scaffolding: is_user_originated_turn / _is_actionable_user_turn / split_user_originated_turn
returned False for it (tail anchoring, auto-focus, dispatcher views, resume counts) while
_is_real_user_message returned True (anchor restoration) — the two predicate families disagreed
on the same row, and list_recent_user_messages (/undo, /rewind) skipped it in SQL. A steer
carries full user authority; the steer kind is now whitelisted in all four.
Also: the pre-API drain's requeue tail reuses _requeue_pending_steer instead of a copy; the TUI
history projection compares against STEER_DISPLAY_KIND; the steer() docstring describes the row.
The previous test replaced _resolve_context_length itself with a fake that
re-did the threading, so it could not fail if the production call dropped the
kwarg; the sibling test only called get_model_context_length directly (green on
main, and its list-shaped `models` fixture is silently ignored by the config
loader). One test now patches agent.context_compressor.get_model_context_length
where production reads it and asserts the kwarg arrives; file moved next to the
other compressor tests. The defensive list copy is dropped (no caller mutates).
ContextCompressor._resolve_context_length called get_model_context_length
without custom_providers, so a per-model context_length override in
custom_providers (e.g. 256K for baseten-deepseek-flash) was skipped and the
hardcoded catalog default (128K for 'deepseek') won instead. /context then
reported the wrong window. Thread custom_providers through the compressor
constructor and into resolution.
A flat per-image constant (1500 in the trigger estimator, 1600 in the tail-budget walk) is wrong in
both directions: a screenshot costs ~1,100 tokens on one provider and 4,000+ on a local mmproj
model. In a GUI loop on a 64K window the estimate sat at ~20K while the real prompt passed 80K,
so compaction never fired and the provider rejected every request (#70328).
The provider prices every image exactly on the request that carries it, so the cost is
observable from usage alone, with no vendor formula: with a fresh usage anchor, the residual
between the next real prompt_tokens and anchor + text-only delta is the price of the N images
that delta introduced.
- agent/image_token_cost.py: calibrate_from_usage() runs in record_response_usage before the new
anchor is captured; the learned value (EMA, plausibility-banded) is kept per model@host in
~/.hermes/cache/image_token_costs.json and bound per turn through a ContextVar.
- estimate_messages_tokens_rough, _content_length_for_budget (tail walk) and gateway hygiene all
read the same bound value, so trigger and walk agree; the per-message memo now caches text
tokens and image COUNT so a recalibration re-prices cached rows.
- One flat default (1500) remains only until the first vision turn; the duplicate 1600 is gone.
evals/token_accounting/ab_image_cost_calibration.py (real AIAgent, fake provider pricing images
at 4,000, one screenshot per turn, 64K window): main learns nothing (1500) and the tail walk
under-prices its own protected tail by 56.5%; this branch learns 4,374 after one vision turn
and the walk's error is +8.5%.
Reporter and first-fix credit: @JonthanaHanh (#70328, #70463).
Codex Responses reasoning and compaction items carry ciphertext the provider prices by its own
token count, never by bytes; a single native compaction checkpoint is ~5M chars, which the
bytes/4 estimator turned into ~1.29M "tokens" against a 204K threshold (#100611). #104192 deferred
that decision for one request; this removes the mis-pricing at the source so the preflight
estimator and the tail-budget walk agree (a mismatched size class protects blob-heavy rows as
"small" and compaction re-fires). Only real usage prices these items, and the usage anchor carries
that price forward.
evals/native_compaction/ab_checkpoint_preflight.py: preflight estimate after checkpoint
1,292,413 -> 58 rough tokens; the over-threshold negative arm still compresses.
Two parallel "real usage" mechanisms fought each other: the usage anchor (real + delta) and the
compressor's rough/real projection (should_defer_preflight_to_real_usage with
last_rough_tokens_when_real_prompt_fit / _pending_request_rough_tokens / note_request_rough_estimate
baselines). The projection stored an anchored, real-scale figure as its "rough" baseline, so a
rewind that invalidated the anchor produced phantom growth and a spurious compaction (#103391).
Now there is one authority:
- Post-tool gate (turn_preflight.compress_after_tool_results): anchored figure first (the raw
last_prompt_tokens ignored the tool results just appended), then real, then rough.
- Gateway hygiene (run_turn._hmwa_hygiene_plan): real session count, else the anchor persisted on
the session row, else rough.
- Preflight / pre-API gates: an anchored figure is never deferred. A whole-context rough estimate
over threshold waits ONE request for the provider's real count instead of compressing on a guess
(first request, rewind/edit-resend, reloaded history without a persisted anchor).
- The wait is one request, never a disable: a provider that omits usage
(note_usage_less_response, #2153 class), a real reading already over threshold, a rough figure
past the whole window, and provider-proven overflow all compress immediately; the post-compaction
latch (#36718 / #104192) is unchanged.
- Projection baselines and their bookkeeping deleted (-101 LOC in context_compressor); the fixtures
that scripted whole-history estimates now state the fact they relied on (provider omits usage).
Fixes#103391 (closes#103397 by construction — the baseline it repaired no longer exists).
A native Responses compaction checkpoint is opaque ciphertext; the rough
preflight estimator counts it as text (5.17M chars -> ~1.29M tokens against
a 204K trigger) and fires local compression on a request whose real prompt is
~116K. Arm the existing one-response real-usage latch when a replayable
checkpoint is captured (build_assistant_message) or restored into a fresh
agent (_hydrate_from_history), honor it in the post-tool gate and idle
compaction, and require non-empty encrypted_content for a checkpoint.
Squash of the author's source commits from #100642 (0e3c234ea0, 771e1b3365,
bb1505a119) plus the fdf140c81d test refresh, re-based onto current main by
patch application. Source delta is byte-identical to the PR head d6ce3e236d.
Fixes#100611
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Post-review cleanup on the salvage stack:
- The re-warn-after-reset guard now drives on_session_reset() instead of
the private helper, so a site regressing to a bare rearm-zero fails it.
- One _over_threshold_warnings() helper replaces four inline caplog filters.
- Comment the redundant None check that narrows current_tokens for ty.
Follow-up to the salvaged #101894 (@jwilson411). The over-threshold
"reclamation did not run" warning is deduped on (reason, rearm mark), and
the key was only cleared when a prune committed. Every other path that
zeroes the rearm mark — compress(), on_session_reset/on_session_end,
bind_session_state, update_model — left the key in place, so a lockout
that warned at rearm=0, then a full compaction, then the same lockout
again was silent, contradicting the helper's own "warns again" contract
(and leaking the key across sessions on a rebound compressor).
- ContextCompressor._reset_proactive_prune_rearm(): one helper for the
five rearm-to-zero sites; clears the dedup key alongside the mark.
- _warn_reclamation_no_op(): dropping back under threshold releases the
key (mirrors _clear_context_overflow_warn semantics on the agent side).
- test_proactive_prune_loop_wiring: the attempts_exhausted fixture now
models the only state the real engine can produce for that branch
(should_compress() is should_compress_info()[0]) — budget spent
(max_compression_attempts=0) with the engine saying RUN, instead of
should_compress=False paired with (True, None).
- Two guards: lockout warns again after a rearm reset; dropping under
threshold releases the key. Both fail with the clears removed.
Message-only rearm could sit just above the body estimate while provider
prompt_tokens (system + tool schemas) already exceeded threshold_tokens,
so prune no-oped forever with no log. Bypass that rearm short-circuit on
the billed basis, warn once when over-threshold reclamation no-ops, and
name attempts_exhausted when should_compress_info says run but the loop
skips.
Fixes#101889
Review follow-ups on the in-flight replay:
- Run _sanitize_tool_pairs BEFORE the re-append. Its trailing-in-flight
exemption (#79278) walks back from the list end; with the replay user row
there, a genuinely pending assistant(tool_calls) looked orphaned and had its
calls stripped, so the executor's late tool result was dropped.
- When the restatement is merged onto a user-pinned summary carrier, flag the
carrier (_inflight_replay_merged). The carrier's metadata marks it synthetic,
so conversation_compression._ensure_compressed_has_user_turn inserted a second
copy of the same request; it now treats the flag as intent-present, and the
next cycle recognises the carrier as the task instead of losing it.
- Header idempotency: restate the text after the last header so a task that
survives several compactions carries one header and one copy.
- Exclude metadata-flagged scaffolding (_todo_snapshot_synthetic, recovery
nudges) from the in-flight scan via the shared _is_real_user_message.
Tests cover all four; each fails with its fix reverted.
The in-flight replay checked only compressed[-1] before choosing between
appending a user row and merging onto the summary carrier. A user-pinned
summary followed by an exempt tool_calls/tool tail therefore gained a second
visible user turn and broke the Mistral-style alternation pre-flight
(tests/agent/test_summary_role_template_alternation.py::test_zero_user_guard_still_forces_user).
Look through the exempt tail to the last template-visible role instead.
Compaction of a single-prompt (cron) session left no user message after
the handoff, so SUMMARY_PREFIX ordered the model to do nothing and the
scheduler recorded success. Re-append the in-flight task after the
summary.
Fixes#100818
(cherry picked from commit c0ea50ac02fd7553b1ef7deb0162483cb9a1b6ac)