Desktop paste/file attachments land in Hermes-managed staging dirs on the
GATEWAY (composer-pastes/ for large text pastes, attachments/ for dropped
files), but on the Remote SSH topology the workspace root (TERMINAL_CWD) is a
path on the SSH HOST - the two filesystems are fully disjoint, as the issue
thread confirms. Two gaps combined to reject every staged attachment with
"path is outside the allowed workspace":
- _resolve_path admitted only allowed_root + composer-paste roots, so a
gateway-staged attachments/ path was refused outright. Admit the
_CACHE_DIRS staging roots (attachments/, images/, cache/*, composer-pastes/)
via a helper that asks get_cache_directory_mounts - the gateway's OWN
payload is never a workspace escape, and the path-traversal and
credential-deny guards in _ensure_reference_path_allowed still run after.
Anything else outside the workspace stays blocked.
- composer-pastes/ was missing from _CACHE_DIRS, so its bytes never reached
the remote: ssh/daytona/vercel_sandbox sync via iter_sync_files ->
iter_cache_files, and to_agent_visible_cache_path only translates mounted
dirs - a paste attached on a fresh session dangled on the remote host.
Tests cover the disjoint-filesystem SSH topology end-to-end (text inlines,
binary renders the synced ~/.hermes path), the still-refused stranger path,
local-backend unchanged, and the composer-pastes mount+sync enumeration.
Consolidates PR #110387 by Finn763 (the _agent_staged_path guard widening and
the SSH-topology tests, adapted to the current _ensure_reference_path_allowed
ordering) with PR #103412 by ericmaddox (whose mapping insight is subsumed by
the _CACHE_DIRS entry, which fixes both the sync and the translation).
Co-authored-by: ericmaddox <ericmaddox@users.noreply.github.com>
Desktop rebuilds the AIAgent per turn (idle reap -> next message re-mint),
and agent_init.py gives every rebuilt agent a brand-new, empty _credits_latch
via new_credits_latch(). seed_credits_at_session_start() -> _hydrate_seed_state()
then unconditionally primes latch["seen_below_90"] on that fresh latch and
evaluates once -- correct for a genuinely new session opening mid-band, but on
a reap/resume rebuild it makes evaluate_credits_notices() see
shown_band=None vs. current_band=<the same band as before>, so it re-fires
"You've used $X of your $Y cap" on every message even though the user already
saw that exact notice moments ago on the previous incarnation of the same
session (#101578).
agent.session_id is stable across these rebuilds even though the agent object
and its latch are not, so add a small process-lifetime cache
(_seen_usage_bands, bounded to 500 entries, MRU eviction) keyed by session_id
that remembers the last usage_band actually shown. _hydrate_seed_state()
restores it into the fresh latch before evaluating, so a rebuild with unchanged
usage stays quiet, while a rebuild after a genuine crossing (recorded via the
same warm-path write in rate_limit_credits._emit_credits_notices, the single
chokepoint both the seed and live-header paths share) still fires normally.
Deliberately not persisted anywhere durable -- a real process restart is a real
"session open" and should still warn immediately, matching the existing
cold-start seed behavior for a session that opens already in a band. An agent
with no session_id (plain CLI, never rebuilt) falls back to the pre-fix
behavior unchanged.
Tests (tests/agent/test_credits_cold_start.py): a rebuild with the same
session_id and unchanged usage does not re-fire; a rebuild after a genuine
band change still fires (and clears the old key); a different session_id is
never suppressed by another session's history; an agent without a session_id
degrades to the old always-prime behavior without raising.
Fixes#101578
Co-authored-by: Edizzier <umit.ediz@hotmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Memory, context files and plugin prose precede the Model:/Provider:/Platform:
trailer and may contain lines of their own with those labels. When the live
value is empty the trailer omits the line, so a prose line stood in for it:
with the identity check now failing closed on an emptied value, that read as
a mismatch on every turn and rebuilt the prompt each time.
Route commits no longer NULL the stored prompt, so _stored_prompt_matches_runtime is the
only rebuild trigger. A model-only request that clears agent.provider skipped the compare
and replayed the previous route's prompt verbatim. The builder omits empty trailer lines,
so the rebuilt prompt matches from the next turn (no rebuild loop); prompts without
identity lines keep reusing.
The summary-interrupt path added a second copy of the string the loop's
interrupt path builds; turn_explainers matches its prefix, so both now
share interrupted_during_api_call_reason().
An InterruptedError from the now-interruptible max-iteration summary was
swallowed by handle_max_iterations' broad except and delivered as the
max_iterations_no_summary fallback with interrupted=False, so finalize_turn
cleared the pending interrupt message and CLI/gateway had nothing to requeue.
Propagate the cancellation, drop the unanswered summary nudge, and surface it
in the finalizer exactly like an interrupted loop API call: interrupted=True,
interrupted_during_api_call exit reason, INTERRUPT_WAITING_FOR_MODEL_PREFIX
text, and result["interrupt_message"] preserved.
#123987 added the "do the task first" carve-out only to PLAIN_INTRO_NOTE,
which is used only when onboarding.profile_build is "off". The default is
"ask", so a fresh default install still received profile_build_directive,
which opens with "After a one-sentence introduction ... OFFER" and lets
the intro/profile offer replace a real first-message task.
Factor the carve-out into TASK_FIRST_CLAUSE and lead both first-contact
notes with it. The consent-gated profile-build steps are unchanged.
The zero-session first-contact sidecar note (PLAIN_INTRO_NOTE /
_hmwa_first_contact_notes) unconditionally told the model to just
introduce itself, with no carve-out for the case where the user's
first-ever message IS a real task. On a fresh tenant whose voice task
was the very first message, the model followed the note verbatim and
replied with a static 'I'm Hermes. /help shows the available commands.'
- no tool call, no attempt at the task at all. This is a third shape of
the first-turn-onboarding-hijack class (turn replaced outright, not
just augmented with the known profile-build pitch).
Fix: PLAIN_INTRO_NOTE now instructs the model to do the task first
(call whatever tools it needs) and fold the one-line intro into the
close of that same reply; only a message with no real request gets the
old bare intro behavior. Also de-duplicated the literal note text in
gateway/run_turn.py, which had drifted into an inline copy instead of
importing agent.onboarding.PLAIN_INTRO_NOTE.
(cherry picked from commit fc7c3839b0b6774133d4fe8df59a9e98d23bdd8b)
'.../.ssh/link/' split to an empty basename, so the entry check degenerated
to checking the link's TARGET: get_write_denied_error(entry=True) and
is_protected_path(follow=False) let the delete remove a link inside ~/.ssh,
and _resolve_entry_for_task fell back to full resolution, so a V4A
'*** Delete File: dir/link/' deleted the file the link points to (the
original bug, trailing-slash form).
split_entry() drops trailing separators (keeping a bare '/' or drive root)
before the parent/leaf split and is used by all three entry-mode sites.
is_protected_path(follow=False) now normcases the joined entry, not only
its parent, so a case-variant spelling of the exe/venv entry still matches
on Windows.
The previous fold guarded each Delete/Move entry by running
get_write_denied_error on dirname(path). That coordinate is wrong both ways:
- runtime self-protection treats ANCESTORS of the running venv/interpreter
as protected, so a plain file directly in ~, ~/.hermes, the checkout root
or the uv python dir could no longer be deleted or moved ("'/Users/x' is a
protected system/credential file");
- credential-dir prefixes end in os.sep and match via startswith, so the
bare dir ~/.ssh never matched and a link directly inside ~/.ssh, ~/.aws,
~/.gnupg, ... was unlinked/renamed.
The existing classifier gains an entry=True mode (get_write_denied_error /
_classify_write_denial, and is_protected_path(follow=False)) that vets
realpath(parent)/basename — the entry, leaf not dereferenced — in addition
to the resolved target. A file inside a protected dir or prefix is denied;
a file merely beside the venv is allowed. delete_file and move_file make
one entry-mode call per entry instead of the duplicated path+dirname loop.
The new parametrized test covers both directions (plain Delete/Move next to
a monkeypatched runtime venv succeeds; a link in <home>/.ssh is refused with
link and target intact); all four cases fail on the previous fold head.
_INFLIGHT_REPLAY_MERGED_KEY was set at exactly one site, on a summary
carrier that has just had the replay header and a non-empty task appended
after its (last) end marker, so the content predicate
_has_merged_inflight_replay is always true wherever the flag is. Keeping
both left two sources of truth for one fact, and the flag is the one lost
on SessionDB reload -- the bug this stack fixes. Detect the merge from
content only so the in-memory and reloaded paths share one code path.
The cron reappend test now asserts the merge layout through the predicate
instead of the private dict key.
Co-authored-by: ppazosp <pablopazosp3@gmail.com>
_has_merged_inflight_replay partitioned on the FIRST end marker. A
merged-into-tail carrier can embed an older carrier (marker + replay
header + task) in its prior-context block, ahead of the new summary's own
marker, so a carrier with nothing after its real boundary was misread as
already carrying the active request. _find_inflight_user_task would stop
there and the reappend would drag the new summary body into the task
text. rpartition anchors on the real boundary; summary bodies never keep
an inner marker and the _force_user_leading layout has exactly one, so
every legitimate layout still matches.
Also compute the stripped remainder once and use removeprefix instead of
a manual length slice.
Co-authored-by: ppazosp <pablopazosp3@gmail.com>
The retain early return coerced agent._cached_system_prompt with `or ""` and
wrote it back. The only setter of _retain_seeded_system_prompt,
gateway.run._seed_hygiene_system_prompt, always stores a str ("" or the
stored prompt) in the same call, and nothing between the seed and the commit
boundary resets it on the detached hygiene / manual-compress agent (the only
None writers are invalidate_system_prompt, which this branch returns before,
and switch_model, which those agents never run). So the normalize/write-back
was dead defence; the empty-seed case still returns "".
The retain branch in _rebuild_system_prompt_at_boundary returns before
_refresh_agent_tool_definitions. That is intentional: refresh_agent_mcp_tools
-> persist_agent_tool_names would otherwise pin the live session row's
tools[] to the detached hygiene agent's memory-only toolset. Requested
inline by teknium on #122825 so a later refactor does not move the refresh
above the early return.
Gateway hygiene and the gateway /compress handler compact with a detached
AIAgent(enabled_toolsets=["memory"]) seeded with the session's stored prompt
(_seed_hygiene_system_prompt, 76a17046e2 / 678916b427). Since #98426 the
commit boundary always rebuilds the prompt, so the seed was discarded and the
reduced-toolset build was persisted over the live session's snapshot: the
skills index (## Skills (mandatory) / <available_skills>) and Skill Safety
Rule vanish, and every later fresh agent for that session restores the
degraded bytes verbatim (_stored_prompt_matches_runtime does not compare
platform).
The seed now marks the agent (_retain_seeded_system_prompt) and the
commit-boundary rebuild keeps the seeded bytes for it. A live agent's own
compaction still rebuilds, so builder updates keep reaching long sessions.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit dd5c8ef4e845b2d34e826e4d92afad73a12eefa6)
repair-prompts keys its degraded-prompt detector on this heading. A hand-copied
literal would silently drift if the heading is reworded, turning every healthy
skill_manage row into a scan finding that --apply clears. SKILLS_GUIDANCE output
is byte-identical (prompt-cache neutral).
Every context engine inherits ContextEngine.compression_count: int = 0 and
the only tagged forks in tests (the review-fork Recorder and the scope tests)
have no compressor or an int counter, so the isinstance check guarded nothing.
Also cut the _apply_fork_tag docstring to the WHY; the measurement numbers
belong in the PR body, not the source.
A substring check on the raw base URL also matches proxies such as
https://proxy.test/api.x.ai/v1, so is_slot_keyed_cache_route now compares
utils.base_url_hostname() exactly, like chat_completion_helpers and
codex_responses_adapter already do.
The aggregator Grok prefixes were spelled out in both prompt_cache_scope and
the OpenRouter profile; one GROK_AGGREGATOR_MODEL_PREFIXES constant keeps the
fork-scope check and the x-grok-conv-id header on the same model set.
Review feedback on #123554: tagging every fork made request #1 of every
xAI review and every /btw a cold full-context prefill, which is the cost
#109964 removed. A same-model fork sends the full parent snapshot, so
until its own in-place compaction rewrites the transcript its stream is
a strict prefix extension of the parent's and cannot evict the parent's
slot.
- _apply_fork_tag keeps the parent's scope until the fork's own
context_compressor.compression_count >= 1 (committed compaction; the
counter is rolled back on aborted attempts). No new config or env.
- Only background_review forks are tagged; /btw (side_question) never is.
- The tag is set inside the same-model branch, so routed forks are not
tagged.
- Note on is_slot_keyed_cache_route that only the OpenRouter profile
consumes cache_scope_id for the sticky header.
#109964 made same-model cache-parity forks (background review, /btw) inherit
the parent's resolved cache scope. That is the right call for
content-addressed caches (Anthropic, DeepSeek, Gemini) and for OpenAI's
routing-only prompt_cache_key. xAI is different: x-grok-conv-id /
prompt_cache_key select ONE server-side conversation slot, so a fork of
similar size that diverges early evicts the parent's slot and the parent's
next call reads cold.
Measured on grok-4.7 (xai-oauth Responses), ~160k context, fork of similar
size diverging right after the system prompt, parent's next call after the
fork:
shared scope (current): 1,152 / 162,239 read (2/2 runs)
derived scope (this PR): 162,176 / 162,239 read (2/2 runs)
no fork (control): 177,536 / 177,639 read (2/2 runs)
build_cache_parity_fork tags the fork (_prompt_cache_fork_tag). The
resolver derives "<scope>::<tag>" for tagged agents on slot-keyed routes
only (xai, xai-oauth, api.x.ai, OpenRouter x-ai/grok-*). The inherited
scope is kept everywhere else. The tag is applied outside the memo, so a
mid-run provider fallback re-evaluates. OpenRouter's Grok x-grok-conv-id
now honours a fork scope over the ambient affinity/conversation scope
(chat_completions threads cache_scope_id to the profile). session_id and
transcript identity are unchanged.
(cherry picked from commit 40fbb39d978648c76940cfb7835f77a42d3719b0)
At turn start the child prompt is checked against the (possibly fallback)
runtime before _restore_primary_runtime runs. A false reject only leaves the
slot None, and the normal restore re-reads the row and re-checks it, so the
order is intentional; say so, so nobody reorders it later.
The clear-then-conditionally-set sequence and its two overlapping comment
blocks said the same thing twice. One conditional assignment keeps the
behaviour identical (None when the child prompt is empty or fails the
runtime-identity check, the child prompt otherwise) and is easier to read.
_adopt_live_compression_child moves agent.session_id onto the live compression child,
but when the child's stored prompt fails _stored_prompt_matches_runtime the parent's
non-null _cached_system_prompt stayed in place. turn_context restores/rebuilds only
while that slot is None, so the child turn could send the parent session's prompt with
no validation. Clear the slot right after child confirmation, before the conditional
seed; a rejected tip then rebuilds through the canonical restore path and a matching
tip is still adopted verbatim.
The regression test now seeds a NON-NULL parent cache — the previous null-path-only
seed hid the leak — and asserts the rejected-tip agent carries session_id="child"
with the slot cleared (red on the previous head).
Review fix for #121843 (ehz0ah, agent/conversation_compression.py:1703).
(cherry picked from commit 531a5072ff4a2e74bd7003094786e32136d2545b)
_adopt_live_compression_child seeds agent._cached_system_prompt straight from the tip's row,
and turn_context only runs _restore_or_build_system_prompt while that slot is None. Until now
a route commit NULLed the row, so a tip whose model moved was never seeded; with the writers
preserving the snapshot, the tip carries the old Model:/Provider: trailer and would be served
to the new model with no check at all. Gate the seed on _stored_prompt_matches_runtime, the
same predicate the restore path applies; an unseeded slot rebuilds on the next restore.
Review fix for #121843 (found in review of the widened invariant; probe: parent ended by
compression, child /model-committed to model-b, adopting agent on model-b was seeded
"Model: model-a" on the PR head and nothing on main).
(cherry picked from commit d7143de9ce7f43201b17e8df5c0344aee777b7fc)
assemble_api_request closes with three passes: strip, tool-call
canonicalization, and _sanitize_messages_surrogates. The summary builder
now mirrors the first two but skipped the third, so a history row with a
lone surrogate (some Ollama-served models emit them) diverged the warmed
prefix and, with the SDK transform bypassed, made the ensure_ascii=False
utf-8 wire encode raise and burn the summary call's retry budget.
The sanitizer is in-place and the summary rows still share nested dicts
with the persisted transcript, so clone them with the main path's
_clone_message_for_send (cheaper than deepcopy, same copy-on-write
guarantee) before the pass. Test ported from #123004 into
TestSummaryPrefixParity.
Surrogate gap reported by @Enough1122 on #123004.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
_iteration_summary_api_messages mirrors most of the send-path passes but
skipped the closing normalization that assemble_api_request applies to
every other request (turn_request_assembly.py:180-184): string-content
strip + _canonicalize_api_tool_calls. The summary request therefore went
out with tool-call arguments in stored key order and unstripped content,
so the rendered prompt diverged from the cached prefix and providers with
prefix caching (e.g. local vLLM) re-read the entire conversation.
Apply the same two passes at the end of the summary builder, reusing
_canonicalize_api_tool_calls from conversation_loop.
Fixes#123002
(cherry picked from commit 4319c67d5b682c5845ee77bd177d9f2f6860de60)
_redecorate_prompt_cache_for_provider returned a 2-tuple when
tools_for_api was None but a 3-tuple otherwise, while its only production
caller (agent/turn_api_request.py) always unpacks three values. Passing
None raised "not enough values to unpack (expected 3, got 2)". The path is
latent today (agent.tools is always a list), but the union return type
invited exactly this mismatch.
Always return (messages, prepared, planned_tools); planned_tools falls
back to agent.tools when tools_for_api is None, as it already did
internally. Update the two tests that unpacked two values and add one
parametrized invariant over None / [] / list x cache-off / native.
Salvaged from #122441, rebuilt on current main with only the redecorate
hunks: the original commit was made on a stale conversation_loop.py and
would have reverted the continuation-overlap dedupe, the legacy
length-continuation stub, the tools-available stub wording and the Bot
Chat NULL guard.
The Codex app-server compaction path was the only caller passing
skills=False; now that it re-arms skill_view dedup like every other
rewrite boundary, the kwarg and its early return are dead. Removing them
keeps the helper's contract simple: every boundary resets both caches,
so a future caller cannot reintroduce the #101518 asymmetry by accident.
skill_view returns a stub on a repeat view of an unchanged file, pointing
the model at the earlier full result in its own context. That contract only
holds while the earlier result is still there, so a compaction boundary has
to advance the dedup generation -- which is what _reset_read_dedup_caches()
is for, and why its stub message promises "Re-issued after context
compression, this returns the full content again."
The Codex app-server path passed skills=False and reset only the file-read
half. A skill re-viewed after that compaction got a stub naming content the
compaction had just summarised away, so the model had nothing to copy from
and reconstructed it from memory instead. Observed with a delegating skill
whose child contract is a JSON schema: successive dispatches shipped
progressively degraded schemas -- a dropped allOf, then a bare
{"type": "array"} -- until the generation-time check no longer constrained
anything and the work was abandoned.
The opt-out carried the pre-refactor asymmetry forward mechanically
(0e9d46511a lifted these call sites without changing behaviour); nothing
depends on it. Drop it so both paths reset both caches. Tests cover the
helper and drive the codex path end to end.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: aurorabotticus-svg <aurorabotticus-svg@users.noreply.github.com>
(cherry picked from commit 22f36fe10ec27bf68e003058b3f1dda292e535e8)
skill_view, read_file and the computer_use screenshot dedup answer a
repeat of unchanged content with an "unchanged" stub that points at the
earlier result. Compaction and a committed proactive prune already
advance all three to a fresh generation through
_reset_read_dedup_caches, so the first repeat after the boundary serves
content again (#32106, #84857, #112763).
Two rewrites never reached that reset:
- post-turn micro-compaction splices older exchanges into the rolling
summary (_micro_compact_after_turn), taking their tool results with
them;
- a replayable native Responses compaction checkpoint makes the next
request drop every item before it (build_assistant_message).
After either, a repeat read got a stub for a body the model could no
longer see. Both now call _reset_read_dedup_caches for the turn's task
and session. Micro-compaction resets only on a real splice: no-op and
defrag passes return the input list and leave the caches alone. The
checkpoint reset no longer depends on the context engine exposing
note_native_compaction_checkpoint, and it is skipped without a task id,
where the helper would clear every task's caches.
No message, tool schema or system-prompt bytes change.
(cherry picked from commit b7c8ecbbc20cd95e70003dce8120b02dbd2ab581)
_prune_old and _is_restorable only mentioned .rollback-staging-*; say
outright that retained .rollback-unrestored-* dirs are neither pruned nor
listed as snapshots, so nobody "fixes" the prune filter to include them
and reintroduces the data loss.
_retain_staging renamed the staging dir unguarded, and it only runs after
a move has already failed. On Windows the same lock that failed the move
blocks the rename, and two failed rollbacks in the same second collide on
the .rollback-unrestored-<id> name. Either way rollback() raised an
OSError (CLI traceback) and the only copy stayed under the prunable
.rollback-staging- prefix, so the next snapshot deleted it (#122210).
Now the rename is wrapped: on failure it logs, reports the original path
and tells the user to move it before the next curator run. A same-second
collision gets a -N suffix. The helper also builds the whole
(False, msg, None) result, which collapses three copy-pasted retain+message
blocks into one-line call sites with a single message format.
Co-authored-by: liuzikaii <2319582736@qq.com>
When rollback cannot put entries back (staging, extract or carry
failure), it keeps the staging dir because it holds the only copy of
nested .git/.hub metadata that snapshots exclude. But the dir kept the
.rollback-staging- prefix, and _prune_old() deletes every dir with that
prefix as stale on the next snapshot_skills() - including the mandatory
safety snapshot of a retry rollback - destroying the copy the error
message told the user to recover from.
Rename retained staging to .rollback-unrestored-<id> via one helper used
by all three retain sites, and report that path. Crashed-rollback
staging is still cleaned up as before.
Blocker raised by Enough1122 in review of #122217.
Same race as the ACP client: _subprocess_died reads stderr_tail() the moment
is_alive() turns False, before the reader thread has the crash lines, so the user
saw 'exited unexpectedly' with no cause (CI on main: test_crash_mid_item...).
stderr_tail() now joins the reader briefly once the process has exited.
The request loop left as soon as poll() saw the exit and read stderr_tail before the
pump thread had the crash text, so a dead CLI raised TimeoutError, which the agent loop
retries on a different budget (4-5 model calls for a 2-retry config in CI). Join the
stderr pump briefly once the process has exited, and report an empty-stderr exit as
"exited early (exit code N)" instead of a timeout.
- salvage summary cap and _bound_oversized_record still composed bare
truncation idioms in model-visible text; route them through elide /
elide_middle so a copied marker is guard-visible.
- the active-task line repr()'d the elided text, escaping the marker's
apostrophe when the user text held both quote kinds and hiding it from
the guard; elide after repr instead (text within the cap stays whole).
- a leftover budget smaller than the marker produced a content-free,
over-budget marker line in _build_verbatim_user_section and the Slack
nested-attachment path; skip the item instead.
- _build_verbatim_user_section elided twice, reporting the wrong total;
one elide at min(cap, remaining).
- drop redundant len() pre-checks before elide() and name verification
stop's repeated 1200.
Co-authored-by: ahisblessed <ahisblessed@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
The guard only catches a copied elision marker because its first sentence is
byte-identical to _COMPRESSION_MARKER_TEMPLATE's; re-typing it by hand left
that invariant to a comment. Derive it the same way _COMPRESSION_MARKER_RE is
derived (verified byte-identical to the previous literal).
Also fix elide_middle(tail=0), where text[-0:] appended the whole text, and
expose ELISION_MARKER_MAX_LEN so callers with a small leftover budget can
skip an item instead of emitting a marker-only line.
Co-authored-by: ahisblessed <ahisblessed@users.noreply.github.com>
#121548's sweep missed three model-visible elisions whose wording differs
from the bare idiom the tokenizer test looked for:
- context_compressor fallback "...[previous summary snapshot truncated]"
- context_compressor fallback "...[fallback summary truncated]"
- context_engine memory-provider "...[memory provider context truncated]..."
They are the same imitable shape, so a model copying them into a durable
write would bypass the dispatch-boundary guard. Route them through
elide()/elide_middle() (caps still held; the helpers size against the
marker) and widen the no-idiom invariant to any "...[<words> truncated]"
variant so a reworded marker cannot slip back in.
The #83714 fix hardened only the tool-call args renderer; five other
renderers (plus three same-idiom siblings) still composed the bare
truncation marker, which the model imitated from replayed context into
new durable writes (#83435).
This routes all model-visible elision in agent/ through shared helpers
in agent/compression_marker.py:
- elide(text, limit): head + marker cap with accurate omitted/total counts
- elide_middle(text, head, tail): kept head/tail with the middle elided
The elision marker's first sentence is byte-identical to the args marker's,
so the existing _COMPRESSION_MARKER_RE (and the dispatch-boundary guard in
tool_dispatch_helpers) rejects a copied marker regardless of which renderer
leaked it; only the tool-call-specific second sentence is dropped so the
marker still fits small caps (the clarify summary cap is 199 chars).
Routed sites:
- context_compressor.py: static-fallback turn renderer, _sum_clarify(),
summarizer prompt builder (middle elision), active-task snapshot,
lean user-message quotes (2 sites)
- skill_preprocessing.py: inline-shell output embedded into skill bodies
- lsp/reporter.py: truncate() for <diagnostics> tool output
- verification_stop.py: verify-on-stop nudge output summary
Regression test asserts the imitable idiom appears nowhere in agent/
strings (comments excluded), so a new open-coded renderer cannot land
silently.
Fixes#121548
(cherry picked from commit 69c63d50b4ababedf4fb4ff88c08f5d1e144a79f)
The cut clamp n - max(walk_floor, 1) runs for every caller of
_find_tail_cut_by_tokens (batch compress and micro-compaction) and
keeps the newest row whatever its role. The old comment described it
only for tool rounds, so a reader could take it for tool-round-specific
logic.