Commit Graph

44160 Commits

Author SHA1 Message Date
kshitijk4poor
f7122daaab fix(slack,trajectory): mint the compression marker for agent-facing truncations
The Slack Block Kit payload dump and the nested-attachment text budget
are both fed to the agent, and trajectory_compressor's summarizer input
becomes training data; all three still used the imitable bare
"... [truncated]" idiom. Route them through elide()/elide_middle() with
module-level imports and extend the no-idiom invariant to scan
plugins/platforms/slack and trajectory_compressor.py.

Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
2026-09-26 23:49:25 +05:30
kshitijk4poor
21edc762d5 fix(agent): route summary-snapshot and memory-context truncations through elide()
#121548's sweep missed three model-visible elisions whose wording differs
from the bare idiom the tokenizer test looked for:

- context_compressor fallback "...[previous summary snapshot truncated]"
- context_compressor fallback "...[fallback summary truncated]"
- context_engine memory-provider "...[memory provider context truncated]..."

They are the same imitable shape, so a model copying them into a durable
write would bypass the dispatch-boundary guard. Route them through
elide()/elide_middle() (caps still held; the helpers size against the
marker) and widen the no-idiom invariant to any "...[<words> truncated]"
variant so a reworded marker cannot slip back in.
2026-09-26 23:49:25 +05:30
ahisblessed
77b9ea0831 fix(agent): route every model-visible elision through the non-imitable compression marker (#121548)
The #83714 fix hardened only the tool-call args renderer; five other
renderers (plus three same-idiom siblings) still composed the bare
truncation marker, which the model imitated from replayed context into
new durable writes (#83435).

This routes all model-visible elision in agent/ through shared helpers
in agent/compression_marker.py:

- elide(text, limit): head + marker cap with accurate omitted/total counts
- elide_middle(text, head, tail): kept head/tail with the middle elided

The elision marker's first sentence is byte-identical to the args marker's,
so the existing _COMPRESSION_MARKER_RE (and the dispatch-boundary guard in
tool_dispatch_helpers) rejects a copied marker regardless of which renderer
leaked it; only the tool-call-specific second sentence is dropped so the
marker still fits small caps (the clarify summary cap is 199 chars).

Routed sites:
- context_compressor.py: static-fallback turn renderer, _sum_clarify(),
  summarizer prompt builder (middle elision), active-task snapshot,
  lean user-message quotes (2 sites)
- skill_preprocessing.py: inline-shell output embedded into skill bodies
- lsp/reporter.py: truncate() for <diagnostics> tool output
- verification_stop.py: verify-on-stop nudge output summary

Regression test asserts the imitable idiom appears nowhere in agent/
strings (comments excluded), so a new open-coded renderer cannot land
silently.

Fixes #121548

(cherry picked from commit 69c63d50b4ababedf4fb4ff88c08f5d1e144a79f)
2026-09-26 23:49:25 +05:30
teknium1
419d1d20f5 chore(kiwi): map contributor email(s) (#123255) 2026-09-26 11:19:21 -07:00
harrylabsj
b40d9c302c chore(plugin-catalog): bump kiwi to 1.2.1 (sha ac5a2dc) 2026-09-26 11:19:21 -07:00
teknium1
dcb85f54a2 chore(plan-mode): catalog review edits, map contributor email(s) (#123212) 2026-09-26 11:18:50 -07:00
100yenadmin
b0d31f037e plugin-catalog: pin plan-mode to 0.2.0 (33e59f94); declare plan_mode tool and post_tool_call hook 2026-09-26 11:18:50 -07:00
100yenadmin
12b5054d8f plugin-catalog: pin plan-mode to 0.1.7 (d7976454) 2026-09-26 11:18:50 -07:00
100yenadmin
e1b41acb5f docs(plugin-catalog): plan-mode gateway/TUI/Desktop enforcement is released in v2026.9.24 2026-09-26 11:18:50 -07:00
teknium1
8ba6d72620 chore(secret-drop): catalog review edits, map contributor email(s) (#123211) 2026-09-26 11:18:24 -07:00
Goitonthefloor
fd250cb8f8 catalog: re-pin secret-drop to 6217414 (README security-model note) 2026-09-26 11:18:24 -07:00
Goitonthefloor
cc70521661 catalog: add secret-drop plugin entry 2026-09-26 11:18:24 -07:00
teknium1
53dae5c863 chore(githermes): map contributor email(s) (#123208) 2026-09-26 11:17:59 -07:00
claudiumio
960cd341e7 fix(githermes): update catalog entry for github owner rename
claudioorjunior -> claudiumio. Pin SHA unchanged, old URLs redirect.
2026-09-26 11:17:59 -07:00
teknium1
f4ea2947ea chore(sigrex): catalog review edits, map contributor email(s) (#123198) 2026-09-26 11:17:30 -07:00
0x7s0lt1
c530521cb8 feat(plugin-catalog): add sigrex platform plugin 2026-09-26 11:17:30 -07:00
teknium1
d36ac6e61a test(tui_gateway): join the rewind turn thread instead of forcing running=False
test_prompt_submit_consecutive_rewinds_with_returned_survivor_row_ids[False]
flaked in CI with `assert 3 == 2`. The in-process path starts a real turn
(dispatch thread -> prompt-turn-* worker); the test cleared sess["running"]
by hand while turn 1's worker could still be alive, an overlap production's
running gate never allows. On a slow runner that worker reached
_adopt_out_of_band_turns inside rewind 2's _persist_submit_user_row window
(row 8 written, _submit_user_row popped but not yet re-staged), read
ceiling=None, and appended rewind 2's own user row to history as a
"foreign" turn. The leaked workers also produced the post-teardown
"state.db connection ... closed while a read was still in flight" warnings.

Repro: delay turn 1 0.3s before _wait_agent_for_prompt and rewind 2's
persist 3s between the DB write and staging -> base fails 12/12, fixed
passes 12/12. Fix: _join_turn_thread() follows _run_thread from the
dispatch thread to the worker and joins it after each rewind.
2026-09-26 11:17:12 -07:00
teknium1
63e44332f5 fix(kanban): a worker the dispatcher never recorded registers itself instead of being run twice
A dispatcher SIGKILLed between _call_spawn_fn and _set_worker_pid leaves a
live worker on a run with worker_pid NULL. release_stale_claims only extends
an expired claim for a recorded live pid, so on TTL expiry it reclaimed the
card and spawned a second worker beside the first: double billing, double
side effects, and a board showing one clean completed run (the first
worker's kanban_complete is refused as stale). Main CI hit it in
test_dispatcher_sigkill_mid_tick_never_destroys_or_duplicates_cards.

The worker now records its own pid on its run before the first model call
(adopt_worker_pid, worker_registered event, host-local claims only) and
exits without working the card when its run was already reclaimed. The
reclaim UPDATE also compares worker_pid so a registration landing between
the stale-claim SELECT and the UPDATE keeps the claim.

Repro: temporary sleep between spawn and pid record + kill 0.2 s after the
spawned event + slow first model reply -> 4/4 red on main with the CI
signature, 8/8 green here.

Fixes #121556
2026-09-26 11:16:45 -07:00
teknium1
2d5cc44fbf chore(plugin-catalog): deepinfra-sandbox — add the billable-sandbox/upload disclosure line; map contributor email 2026-09-26 11:16:19 -07:00
Milos Milutinovic
a005db38dd plugin-catalog: bump deepinfra-sandbox sha to __init__.py shim fix
Points at deepinfra/deepinfra-hermes-sandbox@259a4775, which adds the
root-level __init__.py needed for the catalog's directory-install path
(the plugin's code lives under deepinfra_hermes_sandbox/ for pip installs;
this shim re-exports register()/DeepInfraProvider so `hermes plugins
install` can load it directly too). Resolves the pinned-source-validate
failure flagged in review.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 11:16:19 -07:00
Milos Milutinovic
ddcd03f079 plugin-catalog: add deepinfra-sandbox
DeepInfra's deep_sands terminal backend plugin
(deepinfra/deepinfra-hermes-sandbox), pinned at the commit that added
its plugin.yaml manifest (PR deepinfra/deepinfra-hermes-sandbox#1).

Registers via TerminalEnvironmentProvider (agent/terminal_env_provider.py,
introduced in #94400); no dedicated sandbox/infra category exists yet,
so category: platform is the closest fit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-26 11:16:19 -07:00
teknium1
d59a71a7e0 chore(plugin-catalog): hermes-insight — add provider-key disclosure line; map contributor email 2026-09-26 11:16:06 -07:00
BrokeSkill
0f6b48a70f hermes-insight: bump pin to 1fb40f7 (review fixes: no HTTP server, no daemon, desktop half in backend-plugin/desktop) 2026-09-26 11:16:06 -07:00
BrokeSkill
39d3815954 Add hermes-insight to the plugin catalog 2026-09-26 11:16:06 -07:00
kerrz2020
b5701bdb6b catalog: rtk-rewrite — bump pin to 4c5f8a5, close the two review asks
- sha 4a83520 -> 4c5f8a5: the plugin repo commit that pins the README install
  snippet to rtk v0.49.0 and drops the `rtk-ai-plugin` install key (docs-only,
  manifest 1.0.1), so `version` follows the manifest as the catalog contract asks.
- Disclosure now states the behaviour the reviewer's listing text describes:
  every terminal tool command goes to the rtk binary on PATH and is replaced with
  rtk's rewrite before approval and execution; fail-open when rtk is missing or
  errors. Provenance claim kept but pointed at hooks/hermes/rtk-rewrite.
- Trailing newline added (file previously ended without one).

Both catalog gates re-run locally at the new pin: scripts/validate_plugin_catalog.py
(structural) and `hermes plugins validate --install-deps` + the JS self-updater grep
(pinned-source): all pass.
2026-09-26 11:15:55 -07:00
kerrz2020
5dd02ae5bf catalog: rename rtk-ai-plugin -> rtk-rewrite; pin rtk v0.49.0; bump sha
- Catalog name now matches the manifest name (loader key follows the manifest),
  dropping the implied rtk-ai affiliation the disclosure disclaims.
- update.sh + AGENT-GUIDE.md pin RTK to v0.49.0 instead of releases/latest.
- sha bumped to 4a83520 (repo commit carrying the pin).
2026-09-26 11:15:55 -07:00
kerrz2020
fbed1b1e50 pin rtk-ai-plugin to a7b420a (doc: dual-name note) 2026-09-26 11:15:55 -07:00
kerrz2020
738bdd1f07 catalog: add rtk-ai-plugin (community RTK command-rewriting plugin) 2026-09-26 11:15:55 -07:00
teknium1
0aef4b12ea chore(plugin-catalog): freight-reconciliation — add the loopback admin-page disclosure line 2026-09-26 11:15:41 -07:00
haidong
822c4b760a chore(plugin-catalog): re-pin freight-reconciliation to v0.4.1
v0.4.1 makes every confirmation kind mark its request CONFIRMED (the mapping
branch left a ghost entry on the admin page pending list).
2026-09-26 11:15:41 -07:00
haidong
90680fd14c chore(plugin-catalog): re-pin freight-reconciliation to v0.4.0
Admin-page hardening (HTML escaping + exact origin check) and English-default
UI strings / MCP tool descriptions per review feedback.
2026-09-26 11:15:41 -07:00
haidong
32502b6c7f chore(plugin-catalog): bump freight-reconciliation sha (engine pinned to immutable commit) 2026-09-26 11:15:41 -07:00
harrylabsj
14a2dd6648 feat(plugin-catalog): add freight-reconciliation community plugin 2026-09-26 11:15:41 -07:00
kshitijk4poor
3aa8a7f393 docs(compression): say the newest-row tail floor applies to every row
The cut clamp n - max(walk_floor, 1) runs for every caller of
_find_tail_cut_by_tokens (batch compress and micro-compaction) and
keeps the newest row whatever its role. The old comment described it
only for tool rounds, so a reader could take it for tool-round-specific
logic.
2026-09-26 23:45:36 +05:30
kshitijk4poor
3231237448 fix(compression): keep the pending round's tool-call args verbatim too
The spare covered only the pending round's tool results. The owning
assistant(tool_calls) row sits at spared.start - 1, below the lowered
prune boundary, so pass 3 (range(prune_boundary)) and pass 4's
_shrink_at still truncated its arguments: a pending 60 KB write_file
call came back as 456 chars next to its intact, unread result. The
model then reads its result beside a mangled copy of the call it just
made.

_spared_pending_tool_round now includes the owning assistant row, so
the boundary and _shrink_at skip it along with the results. That row
also counts toward the hard-share check, so the >20% escape hatch
(#61932) still applies to the whole round.

The existing pending-round test now gives the pending call >500-char
args and asserts the assistant's tool_calls survive unchanged (red
before, green after).

Co-authored-by: MongLong0214 <97578200+MongLong0214@users.noreply.github.com>
2026-09-26 23:45:36 +05:30
kshitijk4poor
55f77c1bb1 test(compression): trim pending-round tests to two invariants
The salvage keeps the parametrized pending-round verbatim test (0/1/2
steers) and the four-image spared-round test; together they pin the
#120582 pending-round contract. The two single-prompt replay/oversized
image tests overlapped the four-image case and exceeded the stack's
two-invariant-test budget.

Co-authored-by: MongLong0214 <97578200+MongLong0214@users.noreply.github.com>
2026-09-26 23:45:36 +05:30
MongLong0214
bb17b1f74c fix(compression): keep the pending round's images when it fits
Compaction spared a pending tool round's text results but not its
images. The image-retirement pass kept only the newest three image
results across the transcript, so four parallel vision calls in one
unread round lost the oldest image before the model saw any of them.
On a single-prompt run the final media pass lost three of the four:
compaction re-appends the task as a user row after the round, and the
pending round was looked up after that row was added, so nothing was
spared.

Both image passes now skip a pending round that fits the hard share,
and the pending round is found before the task row is re-appended.
Completed rounds and a pending round over the hard share keep the
existing policy.

(cherry picked from commit eb7661f4365f009d5f9ef85e26f2f4b6ac83632e)
2026-09-26 23:45:36 +05:30
MongLong0214
65f185c594 fix(compression): keep the pending tool round when two steers follow it
The pending-round check skipped only one trailing /steer row. Two can
land in the same iteration: one is appended when the tool batch ends,
and a steer sent after that is drained before the next request and
inserted right after the newest tool result. The transcript then ends
tool -> steer -> steer, so the check stopped on the second steer row,
found no pending round, and preflight compaction replaced the unread
tool output with the one-line pressure stub.

The check now skips every contiguous trailing steer row, identified the
same way as before, and still stops on any other user row. The existing
regression gains a case with two steer rows after the round.

(cherry picked from commit 455a1a3868b3843d15401781363ce796bae96b5b)
2026-09-26 23:45:36 +05:30
MongLong0214
18eb09f5db fix(compression): keep the unanswered tool round verbatim through mid-turn compaction
Preflight compaction can fire right after a tool round, before the model
has read its results. The protected-tail passes then treated that round
like any old output: pass 2 and the pressure pass (#61932) replaced its
results with one-line stubs, and when the newest row alone exceeded the
tail ceiling the cut landed at the end of the transcript and summarised
the round with its whole turn. The model then re-ran the call, side
effects included, or answered without the output.

_pending_tool_round finds the tool results the transcript ends with,
skipping one trailing /steer row (a steer is delivered after the newest
result, before the next API call). Both passes spare that round, and the
tail cut aligns from the row before the end so the group stays whole.
The one exception is a round that alone exceeds the tail's hard share
(20%) of the input budget, the window minus the output reservation: it
still gives way, as #61932 requires. _effective_input_window is
extracted from _compute_threshold_tokens so both use the same budget;
thresholds are unchanged.

(cherry picked from commit 9d74e22379cd7dc39636c522175772d0165db4ac)
2026-09-26 23:45:36 +05:30
kshitijk4poor
27c27c9b0b chore: map MongLong0214 for salvage of #122638 2026-09-26 23:45:36 +05:30
teknium1
5fe8392342 chore(plugin-catalog): hermes-switchyard — state the hosted skill-routing egress bound (4,000 chars/turn) in the disclosure 2026-09-26 11:15:31 -07:00
bgrablin
eb6a23667e chore(catalog): pin Hermes Switchyard v0.5.3 and disclose Jev egress 2026-09-26 11:15:31 -07:00
Brian Grablin
ed703fcf61 feat(catalog): pin Hermes Switchyard v0.5.2
Bump catalog sha/version to Switchyard release that clamps Codex/Responses/Astra
reasoning_effort none|minimal → low (HTTP 400 fix).
2026-09-26 11:15:31 -07:00
Brian Grablin
b25fd864a8 feat(catalog): add Hermes Switchyard v0.5.1
Pin bgrablin/hermes-switchyard at signed release v0.5.1
(6a719e128315be899ff39f03e8d71bcb0fe6ad6c) for community
catalog listing. Refresh tools/hooks/middleware inventory for
adaptive reasoning, session_search rerank, and related 0.5.x
capabilities.
2026-09-26 11:15:31 -07:00
Apostol Apostolov
209f505cbf feat(plugin-catalog): update sidebar-manager to SDK-backed pin 2026-09-26 11:15:17 -07:00
Apostol Apostolov
dd16078697 feat(plugin-catalog): update better-session-appearance to SDK-backed pin 2026-09-26 11:15:07 -07:00
kshitijk4poor
95fc717460 chore(environments): drop uuid imports left unused by the shared staged-stdin path 2026-09-26 23:44:55 +05:30
kshitijk4poor
7d29687ae6 test(environments): cover Vercel stdin staging instead of the local contract
The LocalEnvironment payload-mode test passed on base: it exercises
base.py/local.py, which this stack does not change, so it guarded nothing
here. The Daytona/Vercel staged-stdin transport had no coverage at all.

Swap it for one Vercel test on the existing fake SDK. It checks that the
160 KiB + NUL/0xFF payload is uploaded byte-exact with mode 0o600, never
appears in run_command argv, and that the script starts with the
exec-redirect prefix. It fails on base, where the payload is dropped. Also
drop the change-detector _stdin_mode assert from the Modal test; the
assertions after it already prove the behaviour. The stack still adds two
tests (Modal + Vercel).

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 23:44:55 +05:30
kshitijk4poor
8f5b5cd88b perf(environments): skip stdin staging for empty payloads
execute() can pass stdin_data="" (e.g. write_file of empty content). The
`is not None` guard then paid an upload (plus chmod on Daytona) and a
longer shell command just to feed zero bytes. Base heredoc mode skipped
empty stdin, and neither SDK exec attaches a stdin, so treating "" as no
stdin gives exactly the base command (checked: identical argv, no upload).
2026-09-26 23:44:55 +05:30
kshitijk4poor
e4f51c654e fix(environments): scrub staged stdin when cancel lands during the upload
state["staged"] was set only after the upload returned. A kill() during
the upload therefore saw nothing to scrub, and exec_fn then hit the
cancelled gate and returned 130 without deleting. The staged file (which
can hold the sudo password line) stayed in the sandbox, whose filesystem
persists by default on both Daytona and Vercel.

Factor the scrub into a lock-held helper and call it from exec_fn's
cancelled branch as well as from cancel(). Also skip the upload entirely
when cancel already won before it started.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 23:44:55 +05:30