The Slack Block Kit payload dump and the nested-attachment text budget
are both fed to the agent, and trajectory_compressor's summarizer input
becomes training data; all three still used the imitable bare
"... [truncated]" idiom. Route them through elide()/elide_middle() with
module-level imports and extend the no-idiom invariant to scan
plugins/platforms/slack and trajectory_compressor.py.
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
#121548's sweep missed three model-visible elisions whose wording differs
from the bare idiom the tokenizer test looked for:
- context_compressor fallback "...[previous summary snapshot truncated]"
- context_compressor fallback "...[fallback summary truncated]"
- context_engine memory-provider "...[memory provider context truncated]..."
They are the same imitable shape, so a model copying them into a durable
write would bypass the dispatch-boundary guard. Route them through
elide()/elide_middle() (caps still held; the helpers size against the
marker) and widen the no-idiom invariant to any "...[<words> truncated]"
variant so a reworded marker cannot slip back in.
The #83714 fix hardened only the tool-call args renderer; five other
renderers (plus three same-idiom siblings) still composed the bare
truncation marker, which the model imitated from replayed context into
new durable writes (#83435).
This routes all model-visible elision in agent/ through shared helpers
in agent/compression_marker.py:
- elide(text, limit): head + marker cap with accurate omitted/total counts
- elide_middle(text, head, tail): kept head/tail with the middle elided
The elision marker's first sentence is byte-identical to the args marker's,
so the existing _COMPRESSION_MARKER_RE (and the dispatch-boundary guard in
tool_dispatch_helpers) rejects a copied marker regardless of which renderer
leaked it; only the tool-call-specific second sentence is dropped so the
marker still fits small caps (the clarify summary cap is 199 chars).
Routed sites:
- context_compressor.py: static-fallback turn renderer, _sum_clarify(),
summarizer prompt builder (middle elision), active-task snapshot,
lean user-message quotes (2 sites)
- skill_preprocessing.py: inline-shell output embedded into skill bodies
- lsp/reporter.py: truncate() for <diagnostics> tool output
- verification_stop.py: verify-on-stop nudge output summary
Regression test asserts the imitable idiom appears nowhere in agent/
strings (comments excluded), so a new open-coded renderer cannot land
silently.
Fixes#121548
(cherry picked from commit 69c63d50b4ababedf4fb4ff88c08f5d1e144a79f)
test_prompt_submit_consecutive_rewinds_with_returned_survivor_row_ids[False]
flaked in CI with `assert 3 == 2`. The in-process path starts a real turn
(dispatch thread -> prompt-turn-* worker); the test cleared sess["running"]
by hand while turn 1's worker could still be alive, an overlap production's
running gate never allows. On a slow runner that worker reached
_adopt_out_of_band_turns inside rewind 2's _persist_submit_user_row window
(row 8 written, _submit_user_row popped but not yet re-staged), read
ceiling=None, and appended rewind 2's own user row to history as a
"foreign" turn. The leaked workers also produced the post-teardown
"state.db connection ... closed while a read was still in flight" warnings.
Repro: delay turn 1 0.3s before _wait_agent_for_prompt and rewind 2's
persist 3s between the DB write and staging -> base fails 12/12, fixed
passes 12/12. Fix: _join_turn_thread() follows _run_thread from the
dispatch thread to the worker and joins it after each rewind.
A dispatcher SIGKILLed between _call_spawn_fn and _set_worker_pid leaves a
live worker on a run with worker_pid NULL. release_stale_claims only extends
an expired claim for a recorded live pid, so on TTL expiry it reclaimed the
card and spawned a second worker beside the first: double billing, double
side effects, and a board showing one clean completed run (the first
worker's kanban_complete is refused as stale). Main CI hit it in
test_dispatcher_sigkill_mid_tick_never_destroys_or_duplicates_cards.
The worker now records its own pid on its run before the first model call
(adopt_worker_pid, worker_registered event, host-local claims only) and
exits without working the card when its run was already reclaimed. The
reclaim UPDATE also compares worker_pid so a registration landing between
the stale-claim SELECT and the UPDATE keeps the claim.
Repro: temporary sleep between spawn and pid record + kill 0.2 s after the
spawned event + slow first model reply -> 4/4 red on main with the CI
signature, 8/8 green here.
Fixes#121556
Points at deepinfra/deepinfra-hermes-sandbox@259a4775, which adds the
root-level __init__.py needed for the catalog's directory-install path
(the plugin's code lives under deepinfra_hermes_sandbox/ for pip installs;
this shim re-exports register()/DeepInfraProvider so `hermes plugins
install` can load it directly too). Resolves the pinned-source-validate
failure flagged in review.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
DeepInfra's deep_sands terminal backend plugin
(deepinfra/deepinfra-hermes-sandbox), pinned at the commit that added
its plugin.yaml manifest (PR deepinfra/deepinfra-hermes-sandbox#1).
Registers via TerminalEnvironmentProvider (agent/terminal_env_provider.py,
introduced in #94400); no dedicated sandbox/infra category exists yet,
so category: platform is the closest fit.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- sha 4a83520 -> 4c5f8a5: the plugin repo commit that pins the README install
snippet to rtk v0.49.0 and drops the `rtk-ai-plugin` install key (docs-only,
manifest 1.0.1), so `version` follows the manifest as the catalog contract asks.
- Disclosure now states the behaviour the reviewer's listing text describes:
every terminal tool command goes to the rtk binary on PATH and is replaced with
rtk's rewrite before approval and execution; fail-open when rtk is missing or
errors. Provenance claim kept but pointed at hooks/hermes/rtk-rewrite.
- Trailing newline added (file previously ended without one).
Both catalog gates re-run locally at the new pin: scripts/validate_plugin_catalog.py
(structural) and `hermes plugins validate --install-deps` + the JS self-updater grep
(pinned-source): all pass.
- Catalog name now matches the manifest name (loader key follows the manifest),
dropping the implied rtk-ai affiliation the disclosure disclaims.
- update.sh + AGENT-GUIDE.md pin RTK to v0.49.0 instead of releases/latest.
- sha bumped to 4a83520 (repo commit carrying the pin).
The cut clamp n - max(walk_floor, 1) runs for every caller of
_find_tail_cut_by_tokens (batch compress and micro-compaction) and
keeps the newest row whatever its role. The old comment described it
only for tool rounds, so a reader could take it for tool-round-specific
logic.
The spare covered only the pending round's tool results. The owning
assistant(tool_calls) row sits at spared.start - 1, below the lowered
prune boundary, so pass 3 (range(prune_boundary)) and pass 4's
_shrink_at still truncated its arguments: a pending 60 KB write_file
call came back as 456 chars next to its intact, unread result. The
model then reads its result beside a mangled copy of the call it just
made.
_spared_pending_tool_round now includes the owning assistant row, so
the boundary and _shrink_at skip it along with the results. That row
also counts toward the hard-share check, so the >20% escape hatch
(#61932) still applies to the whole round.
The existing pending-round test now gives the pending call >500-char
args and asserts the assistant's tool_calls survive unchanged (red
before, green after).
Co-authored-by: MongLong0214 <97578200+MongLong0214@users.noreply.github.com>
The salvage keeps the parametrized pending-round verbatim test (0/1/2
steers) and the four-image spared-round test; together they pin the
#120582 pending-round contract. The two single-prompt replay/oversized
image tests overlapped the four-image case and exceeded the stack's
two-invariant-test budget.
Co-authored-by: MongLong0214 <97578200+MongLong0214@users.noreply.github.com>
Compaction spared a pending tool round's text results but not its
images. The image-retirement pass kept only the newest three image
results across the transcript, so four parallel vision calls in one
unread round lost the oldest image before the model saw any of them.
On a single-prompt run the final media pass lost three of the four:
compaction re-appends the task as a user row after the round, and the
pending round was looked up after that row was added, so nothing was
spared.
Both image passes now skip a pending round that fits the hard share,
and the pending round is found before the task row is re-appended.
Completed rounds and a pending round over the hard share keep the
existing policy.
(cherry picked from commit eb7661f4365f009d5f9ef85e26f2f4b6ac83632e)
The pending-round check skipped only one trailing /steer row. Two can
land in the same iteration: one is appended when the tool batch ends,
and a steer sent after that is drained before the next request and
inserted right after the newest tool result. The transcript then ends
tool -> steer -> steer, so the check stopped on the second steer row,
found no pending round, and preflight compaction replaced the unread
tool output with the one-line pressure stub.
The check now skips every contiguous trailing steer row, identified the
same way as before, and still stops on any other user row. The existing
regression gains a case with two steer rows after the round.
(cherry picked from commit 455a1a3868b3843d15401781363ce796bae96b5b)
Preflight compaction can fire right after a tool round, before the model
has read its results. The protected-tail passes then treated that round
like any old output: pass 2 and the pressure pass (#61932) replaced its
results with one-line stubs, and when the newest row alone exceeded the
tail ceiling the cut landed at the end of the transcript and summarised
the round with its whole turn. The model then re-ran the call, side
effects included, or answered without the output.
_pending_tool_round finds the tool results the transcript ends with,
skipping one trailing /steer row (a steer is delivered after the newest
result, before the next API call). Both passes spare that round, and the
tail cut aligns from the row before the end so the group stays whole.
The one exception is a round that alone exceeds the tail's hard share
(20%) of the input budget, the window minus the output reservation: it
still gives way, as #61932 requires. _effective_input_window is
extracted from _compute_threshold_tokens so both use the same budget;
thresholds are unchanged.
(cherry picked from commit 9d74e22379cd7dc39636c522175772d0165db4ac)
Pin bgrablin/hermes-switchyard at signed release v0.5.1
(6a719e128315be899ff39f03e8d71bcb0fe6ad6c) for community
catalog listing. Refresh tools/hooks/middleware inventory for
adaptive reasoning, session_search rerank, and related 0.5.x
capabilities.
The LocalEnvironment payload-mode test passed on base: it exercises
base.py/local.py, which this stack does not change, so it guarded nothing
here. The Daytona/Vercel staged-stdin transport had no coverage at all.
Swap it for one Vercel test on the existing fake SDK. It checks that the
160 KiB + NUL/0xFF payload is uploaded byte-exact with mode 0o600, never
appears in run_command argv, and that the script starts with the
exec-redirect prefix. It fails on base, where the payload is dropped. Also
drop the change-detector _stdin_mode assert from the Modal test; the
assertions after it already prove the behaviour. The stack still adds two
tests (Modal + Vercel).
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
execute() can pass stdin_data="" (e.g. write_file of empty content). The
`is not None` guard then paid an upload (plus chmod on Daytona) and a
longer shell command just to feed zero bytes. Base heredoc mode skipped
empty stdin, and neither SDK exec attaches a stdin, so treating "" as no
stdin gives exactly the base command (checked: identical argv, no upload).
state["staged"] was set only after the upload returned. A kill() during
the upload therefore saw nothing to scrub, and exec_fn then hit the
cancelled gate and returned 130 without deleting. The staged file (which
can hold the sudo password line) stayed in the sandbox, whose filesystem
persists by default on both Daytona and Vercel.
Factor the scrub into a lock-held helper and call it from exec_fn's
cancelled branch as well as from cancel(). Also skip the upload entirely
when cancel already won before it started.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>