Commit Graph

3443 Commits

Author SHA1 Message Date
Royalaid
9e47ac17bd fix(search): serialize filename walks by root 2026-09-03 04:56:59 +05:30
Royalaid
af98771b45 fix(search): normalize dash-prefixed find roots 2026-09-03 04:56:59 +05:30
Royalaid
3c864fe331 fix(search): terminate rg root options 2026-09-03 04:56:59 +05:30
Royalaid
092b355f90 fix(search): scope macOS globs and mark bounded totals 2026-09-03 04:56:59 +05:30
Royalaid
4c189cfd41 fix(search): accept bounded rg SIGPIPE and full SemVer 2026-09-03 04:56:59 +05:30
Royalaid
6d9022a55e fix(search): harden bounded fallback handling 2026-09-03 04:56:59 +05:30
Royalaid
e11d91339d fix(search): fail closed and preserve explicit roots 2026-09-03 04:56:59 +05:30
Royalaid
06086b95e1 fix(search): unify fallback scans and reject partial errors 2026-09-03 04:56:59 +05:30
Royalaid
3880e4af62 fix(search): close ordering and resolver gaps 2026-09-03 04:56:59 +05:30
Royalaid
413fd2a1fc feat(search): add fast discovery file ordering 2026-09-03 04:56:59 +05:30
Royalaid
aa446092d0 fix(file-ops): normalize Windows find fallback paths 2026-09-03 04:56:59 +05:30
John Paul Soliva
1a451bfaeb fix(mcp): pick the Windows launcher, not the sh script, from an npx cache
Review follow-up, and a real platform bug. On Windows npm's `.bin` holds three
siblings per binary — the extensionless sh script, `<name>.cmd` and
`<name>.ps1`. Spawning the sh one from a Windows process fails, and
`os.access(X_OK)` there is effectively an existence check, so the previous
candidate test could not tell them apart: it would have picked a file that
cannot run and broken server startup that works today via npx.

Select by extension instead (`.cmd`, then `.exe`), the same precedence
hermes_constants._candidate_node_command_names already uses for npm/npx/node,
and fall back to npx when no launcher is present.

The platform branch moved into `_npx_bin_candidates(..., windows=...)` so it is
testable by injection. Patching `os.name` instead took pytest's own traceback
formatting down with an INTERNALERROR — a test that cannot report its own
failure is worse than no test.

Also from review: an `npx pkg -y` shape (flag AFTER the spec) now falls back to
npx, since those args are forwarded verbatim and would hand the server a flag
npx would have consumed. And the structural OSV-ordering guard reports a rename
explicitly instead of raising a bare ValueError that reads like a broken test.
2026-09-03 04:40:44 +05:30
John Paul Soliva
11ed840431 perf(mcp): skip npx's resident parent when the package is already cached
`npx` resolves a package and then FORKS: it stays alive as the real MCP
server's parent for the whole process lifetime while doing no work. Measured
on a 4-agent host, that is ~48 MB of private memory per stdio MCP server —
and it buys nothing here, because Hermes already wraps the child in its own
parent-death watchdog, so npx's supervision is a second parent nobody reads.
The process tree for one server is watchdog -> `npm exec <pkg>` -> node.

When the package is already in npx's cache, spawn its binary directly and
drop the middle process. A cache miss changes nothing: the command stays
`npx`, so a cold machine still installs on first run.

The swap deliberately happens AFTER the OSV malware preflight.
`_infer_ecosystem` keys off the command basename being `npx`/`uvx`/`pipx`, so
rewriting first makes `check_package_for_malware` return None and silently
turns the gate into a no-op. Two tests pin that ordering — one behavioural,
one structural — because a future edit that moved the swap earlier would
disable the malware check without failing any test of either piece alone.
This is also why doing it in code beats the config workaround of pointing
`command` straight at the cached path: that loses the preflight.

Conservative by construction — falls back to npx for a version-pinned spec
(npx owns that resolution), an ambiguous or absent `bin` map, a missing or
non-executable binary, an unreadable cache entry, or no cache at all.

Measured on a Mac mini running four agent gateways plus a dashboard, ten
stdio MCP servers total: resident `npm exec` parents 5 -> 0, tracked
footprint 1531 MB -> 1324 MB across 30 -> 25 processes, free memory
464 MB -> 697 MB on a host that was actively swapping. Every MCP server kept
working; tool calls verified against a live Linear server afterwards.
2026-09-03 04:40:44 +05:30
RelaxJonh
6fdd93ab45 fix(tools): release per-task file registry state (#86514)
Clear file-operation trackers when a terminal task ends and remove per-path
locks after the last holder or waiter exits. This prevents long-lived gateway
processes from retaining file state indefinitely while preserving sibling-write
tracking and path-level serialization.

Fixes #86514
2026-09-03 04:34:45 +05:30
Jason Wu
98a6e4a90e fix(session-search): bound recent-session browsing
Preselect indexed recent candidates before rich hydration, cap and deduplicate compression lineage traversal, and interrupt sustained SQLite work through a cooperative progress deadline. Fail closed when the bounded browse API is unavailable and cover legacy-schema reconciliation plus malformed lineage cases.
2026-09-03 04:23:35 +05:30
kshitijk4poor
d99eed7d83 perf(cron): bound the lifecycle guard's whole-walk scan work
The referenced-script walk in cron/lifecycle_guard.py capped each file
(1 MiB) and the recursion depth (8) but not the walk: a command referencing
hundreds of scripts, or one enormous shlex token, held the GIL for minutes
on every gateway terminal call (#78398).

Add a per-walk _LifecycleScanBudget (bytes, lines, longest line, unique
paths, remote reads) charged BEFORE any text reaches shlex, and cap each
referenced read at the remaining byte budget so an oversized file is never
read whole. Exhaustion fails closed (the existing contract for one oversized
file) and is logged at WARNING so operators can tell it from a genuine
lifecycle block. Limits are sized so real wrapper graphs never hit them:
a 200-script benign graph is allowed and a restart hidden behind it is
still caught.

tools/terminal_tool.py gates its optional launchctl pre-scan (which also
tokenizes) on the same budget; the full guard still runs afterwards.

Redesigned from #83821 by @Riccardo-Vecchi, which introduced the budget
idea but blocked benign wide graphs (64-path cap) and bundled a suffix
classification change that is left out here.

Refs #78398
2026-09-03 03:33:24 +05:30
kshitijk4poor
62f2c82f2b fix(terminal): &> after a backgrounded compound is a redirect, not a terminator
Follow-up on the #42278 salvage (#98222): `a && b & &>/dev/null c` is valid
bash but rewrote to `a && { b & } &>/dev/null c`, which binds the redirect to
the brace group and orphans `c`. Treat a suffix starting with `&>` as command
text and insert the `;` separator. Adds bash -n coverage for that shape, for
the `;;` case-arm terminator (must NOT gain a separator), a multi-line mixed
form, and separator idempotence.
2026-09-03 03:32:32 +05:30
entropy-0x
77dd1c6534 fix(tools): restore separator after backgrounded compound rewrite
`_rewrite_compound_background` rewrites `A && B &` into `A && { B & }` to
avoid the subshell-wait process leak. When another statement follows the
backgrounded compound on the same line (`A && B & C`), the trailing `&`
was the only separator between the compound and `C`. The rewrite consumed
that `&` into the brace group and produced `A && { B & } C`. A brace group
must be terminated by `;`, `&`, `|`, a newline, or `)`/`}` before the next
command, so the result is a bash syntax error and the entire command fails
to run — neither `A`, `B`, nor `C` execute. The rewrite is applied by
default to every foreground command, so a valid command is silently
mangled into one that errors out.

This restores a `;` separator after the closing `}` when the suffix
resumes with command text. Only spaces and tabs are stripped before the
check; a newline already terminates the group, and an existing separator
(`;`, `&`, `|`, `)`, `}`) is left untouched, so previously-correct rewrites
are unchanged.

## What does this PR do?

Fixes a rewrite in `_rewrite_compound_background` that turned a valid
foreground command of the form `A && B & C` into the bash syntax error
`A && { B & } C`, causing the whole command to fail. A `;` is now inserted
after the brace group whenever a further statement follows on the same
line, while leaving commands that already end in a separator/newline
untouched.

## Related Issue

N/A

## Type of Change

- [x] 🐛 Bug fix (non-breaking change that fixes an issue)
- [ ] ✨ New feature (non-breaking change that adds functionality)
- [ ] 🔒 Security fix
- [ ] 📝 Documentation update
- [ ] ✅ Tests (adding or improving test coverage)
- [ ] ♻️ Refactor (no behavior change)
- [ ] 🎯 New skill (bundled or hub)

## Changes Made

- `tools/terminal_tool.py`: in `_rewrite_compound_background`, insert a `;`
  after the rewritten `{ ... & }` brace group when the trailing suffix
  begins with command text rather than a separator/terminator.
- `tests/tools/test_terminal_compound_background.py`: add
  `TestTrailingStatementSeparator` for the string-level rewrites and
  `TestRewriteIsValidBash`, which runs the rewriter output through
  `bash -n` for parse validity plus one end-to-end execution check.

## How to Test

1. Run `pytest tests/tools/test_terminal_compound_background.py -q`.
2. Before the fix, `_rewrite_compound_background("echo hi && sleep 5 & echo done")`
   returns `echo hi && { sleep 5 & } echo done`; `bash -n -c` on that string
   exits 2 with `syntax error near unexpected token 'echo'`.
3. After the fix it returns `echo hi && { sleep 5 & } ; echo done`, which
   parses and runs, and the trailing statement executes.

## Checklist

### Code

- [x] I've read the Contributing Guide
- [x] My commit messages follow Conventional Commits
- [x] I searched for existing PRs to make sure this isn't a duplicate
- [x] My PR contains only changes related to this fix
- [x] I've run `pytest tests/tools/test_terminal_compound_background.py -q` (50 passed)
- [x] I've added tests for my changes
- [x] I've tested on my platform: macOS 15.5

### Documentation & Housekeeping

- [x] I've updated relevant documentation — N/A
- [x] I've updated `cli-config.yaml.example` if I added/changed config keys — N/A
- [x] I've updated `CONTRIBUTING.md` or `AGENTS.md` — N/A
- [x] I've considered cross-platform impact: the rewrite is
  platform-independent; the `bash -n` test is skipped when bash is absent
- [x] I've updated tool descriptions/schemas — N/A
2026-09-03 03:32:32 +05:30
Teknium
66414dedd3 Port from cline/cline#13525: bound giant single-line matches in content search
A search_files hit inside a serialized dump (multi-MB single-line JSON,
minified bundle) made rg/grep emit the entire matched line into stdout:
head -n counts lines, so a 40MB match line crossed the exec transport
untruncated and was buffered whole into Python before the per-match
[:500] clamp ran. Measured on main: 42MB transport payload and ~180MB
peak Python allocation for a single match.

Fix at the engine layer, all three pipelines:
- rg: --max-columns 2000 --max-columns-preview (preview keeps the match
  visible instead of omitting it)
- grep fallback + darwin pruned-grep fallback: | cut -c1-2000
- files_only/count modes skipped (lines are paths/counts, never giant)

2000 cols exceeds the existing 500-char content clamp, so no
previously-visible content changes.

Adapted from cline/cline#13525 (search_codebase RangeError OOM crash on
giant single-line files) — hermes buffers in Python rather than a JS
string, so the failure mode is memory/transport blowup rather than an
uncaughtException, but the class is identical.
2026-09-03 03:18:44 +05:30
Hermes
cee12cc34d Port from MoonshotAI/kimi-code#3234 + #3227: MCP structuredContent dedup + dropped-block notices
content and structuredContent are now alternatives — never both forwarded
to the model. Spec-following servers render their data into content (the
verbatim dual-emit SHOULD or a faithful reorganisation), so forwarding
both sent the same information twice per tool call. structuredContent
fills in only when content blocks render effectively empty (whitespace-only
text counts as empty), preserving structuredContent-only servers. _meta
passthrough is unchanged.

Also surfaces unsupported content-block drops to the MODEL as
'[MCP content dropped: unsupported block (...)]' notices carrying
type/mime/uri/size handles (kimi-code#3227) instead of a log-only warning;
drop notices do not count as usable content for the arbitration.

Live E2E: real stdio MCP server (mcp 2.x structured_output tool) through
register_mcp_servers + _make_tool_handler — origin/main forwarded the
payload twice, branch forwards content only; text-only path unchanged.
2026-09-03 03:18:19 +05:30
kshitijk4poor
3bf4fa45b8 refactor(browser): use lru_cache for the lightpanda flag probe
/simplify-code pass on the salvage stack. The hand-rolled module-global
probe cache ignored its binary argument (stale verdict survives a binary
swap) and logged per launch; switch to the repo's established
functools.lru_cache capability-probe pattern (cua_backend, browser_tool),
probe stdout+stderr, share the --http-cache-dir literal via a module
constant, tighten the probe timeout to 3s, and cover the probe's
except branch in tests. Mutation-checked: gate tests fail with the
gate removed and with the probe hardcoded True.
2026-09-03 03:18:08 +05:30
kshitijk4poor
eac466d7bd fix(browser): gate lightpanda --http-cache-dir on binary support
Salvage follow-up for PR #100269. The flag landed upstream in 0.3.x;
binaries before it (e.g. 0.2.8, verified locally) fatally reject the
flag with 'unknown argument', breaking every Browser Use launch.
Probe 'lightpanda help' once per process and omit the flag when the
binary predates it. Also soften the unverified concurrency claim in
the _http_cache_dir docstring and tell docs readers to stop sessions
before deleting the live sqlite cache.
2026-09-03 03:18:08 +05:30
Adrià Arrufat
6973ce3ac2 feat(browser): give Lightpanda an on-disk HTTP cache
Lightpanda's HTTP cache is opt-in (`--http-cache-dir`, off by default), and
the launcher never passed it, so every Browser Use navigation re-fetched
every asset.

Point all Hermes-spawned instances at one shared cache under
$HERMES_HOME/cache/browser-use/lightpanda/http-cache. Sharing it across
sessions keeps assets warm through session churn; Lightpanda stores it in
sqlite (WAL), so a write that loses a race degrades to a cache miss rather
than a failed load, and --http-cache-entry-limit (default 1000) bounds the
directory without Hermes managing eviction.

Measured over 25 navigations across 5 sites, median warm navigation drops
from 0.40s to 0.18s on news.ycombinator.com and 0.14s to 0.10s on
wikipedia; total navigation time 7.0s -> 5.9s.

The flag has existed since Lightpanda 0.3.x (April 2026), so this needs no
minimum-version bump.
2026-09-03 03:18:08 +05:30
kshitijk4poor
87597d30c8 perf(mcp): apply the -t spawn filter before the mcp SDK import
With servers configured but none selected by -t/--toolsets, discovery still
paid the ~260ms mcp SDK import before discovering it had nothing to spawn.
Filter first; the test now asserts the SDK probe is not reached.
2026-09-03 03:16:04 +05:30
Michel Alexander
e73257ccc1 feat(cli): filter MCP server spawning by -t/--toolsets flag + fix orphan subprocess leak
v5 — addresses GPT-5.5 fusion-judge SHIP_WITH_FIXES on v4.
One-line cosmetic fix: SIGTERM handler now uses cleanup-scope local
alias `_os_local.kill(_os_local.getpid(), signum)` instead of bare
`os.kill(os.getpid(), signum)`, completing the self-containment intent.
Behavior unchanged (verified: SIGTERM still exits 143, SIGINT still 130).

═══ Code review history ═══

- v1: qwen3-max single review → SHIP_WITH_FIXES
- v2: Fusion (Codex+Sonnet+Gemini, GPT-5.5 judge) → NEEDS_REWORK (7 fixes)
- v3: Fusion → SHIP_WITH_FIXES (3 required + 4 hardening fixes)
- v4: Fusion → SHIP_WITH_FIXES (1 cosmetic — use _os_local in SIGTERM)
- v5 (this commit): cosmetic fix applied; behavior verified unchanged

═══ Final fix tracking ═══

| v2 required fix                       | Final |
|---------------------------------------|-------|
| 1. Fail-open import behavior          | FIXED |
| 2. SIGTERM exit code 143              | FIXED |
| 3. Unmatched -t warning logic         | FIXED |
| 4. _stdio_pids robustness             | FIXED |
| 5. time/os imports verified           | FIXED |
| 6. Poll-loop logger.debug             | FIXED |
| 7. Idempotency guard                  | FIXED |
| 8. SIGTERM uses self-contained _os    | FIXED (v5) |

═══ Verified locally (final) ═══

  hermes -z "ACK" -t web        : 4.8s wall (was 65s — 93% reduction)
  hermes -z "ACK" -t slack      : 5.5s wall (was 67s — 92% reduction)
  hermes -z "ACK"  (no -t)      : 75s     (unchanged — backwards-compat)
  SIGINT mid-flight             : exit 130 ✓
  SIGTERM mid-flight            : exit 143 ✓
  Orphan accumulation           : 0 new PPID=1 across 10+ runs

Related issues:
- Fixes the MCP subprocess component of #18438 (gateway memory leak)
- Supersedes the startup-drag portion of #18523 (closed unmerged)
- Extends toolset-gating pattern from #18166 and #5788 (memory) to MCPs

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-03 03:16:04 +05:30
Enzo Adami
02a6e95228 fix(files): allow one full read after compaction 2026-09-03 03:07:37 +05:30
webtecnica
4a0aae8aae fix(agent): preserve tool-output dedup state across context compaction (#84857) 2026-09-03 03:07:37 +05:30
kshitijk4poor
8cab422ab0 perf(file-ops): count the native read's tail with memchr; reuse the local-env gate
Two /simplify-code findings on the salvaged #95160:

- The native read walked every line of the whole file with per-line
  Python bookkeeping even after the requested window had passed, because
  total_lines and the trailing byte are still needed. Once lineno is past
  end_line, switch to chunk.count(b'\\n') for the remainder -- same result,
  C speed. Page 1 of a 3M-line (123 MB) file: 448 ms -> 49 ms. Parity
  harness native vs shell: 20 shapes (1 MiB boundary, no trailing newline,
  past-EOF offsets, long lines) identical.

- _native_read_enabled re-implemented the LocalEnvironment check that
  _lsp_local_only() already does in this class (same env-None / import
  policy) and memoised it on self although self.env is bound once in
  __init__ and never rebound. Call the existing helper, drop the cache.
2026-09-03 02:27:52 +05:30
Alex Jia
e1c0c4bb6e chore(file-ops): log compound probe fallbacks at debug level
(cherry picked from commit 345f8994f3d422eb39f1807006ca41ce9086c65a)
2026-09-03 02:27:52 +05:30
Alex Jia
d3275acf80 perf(file-ops): native read_file fast path on local POSIX environments
(cherry picked from commit 04b27898724d5ca756bf1e7c61578bd28986364a)
2026-09-03 02:27:52 +05:30
Alex Jia
9fbb3b716a perf(file-ops): merge write_file's pre-write probes into one shell call
(cherry picked from commit fe68f00b22dcfe4706e38a3ca4f310b7501ba131)
2026-09-03 02:27:52 +05:30
Alex Jia
6a2b3f1ebf perf(file-ops): collapse read_file's shell probes into one compound command
(cherry picked from commit 4e6c0beff77e24c46641d09ce77258c819ea9098)
2026-09-03 02:27:52 +05:30
kshitij
04eaaab410 Merge pull request #101622 from kshitijk4poor/fix/92164-utf8-seam
fix(process): keep the sandbox log poller from splitting a UTF-8 character across polls
2026-09-03 02:26:12 +05:30
kshitijk4poor
126668c72d fix(mcp): wait() the released supervisor so it does not linger as a zombie
Phase-2c finding on #101598: the empty-set release closed the pipe and
dropped the Popen without reaping it, so an idle gateway that once
connected a stdio server held one zombie until the next Popen in the
process. wait(timeout=5) after close -- the supervisor exits on EOF
immediately with nothing registered (E2E: ps shows no entry at all).

Also: the replay test asserted against the last line only, which could
not catch the regression it names; assert against the full stream and a
post-unregister replay. Module docstring now states the stdlib-only /
no tools/ import constraint and why _reap's sweep is a deliberate copy.
2026-09-03 02:24:06 +05:30
kshitijk4poor
8f5749db4b fix(mcp): respawn the death supervisor on any lifecycle event while groups remain
Addresses @andrexibiza's Blocker 1 on #93517 (reproduced): after a
broken-pipe write dropped the supervisor while groups were still
registered, the no-spawn fast path was keyed on the incoming verb
(`unregister` + no proc => return), so a clean teardown of one server
left the survivors recorded in _supervised_pgids but unsupervised until
the next register. Key the fast path on the supervised set being empty
instead; any later call respawns and replays the survivors.

Regression test: two live groups, write fails, unregister one ->
replacement receives `register` for the survivors. Mutation-checked
against the old verb-keyed guard.
2026-09-03 02:24:06 +05:30
kshitijk4poor
4a2b23d77e fix(mcp): release the death supervisor once nothing is left to reap
Follow-up to the salvaged #93517. Two gaps found in review:

- After the last unregister the supervisor stayed resident for the life of
  the process (~15 MB + a pipe) in any gateway/cron that ever connected a
  stdio server; main's per-server watchdog exited with its server. Close
  our write end when the supervised set empties: EOF with nothing
  registered makes the supervisor exit without reaping, and the next
  register already respawns and replays.
- A child that raced and exited before os.getpgid dropped its group from
  coverage entirely. The SDK spawns stdio servers as session leaders
  (pgid == pid), so fall back to the pid instead; the prune forgets the
  group once nothing in it is alive.

Tests: fake-supervisor release/re-spawn sequence, and a real-process EOF
release (supervisor exits 0, unregistered child untouched). Mutation
checked: disabling the release branch fails both.
2026-09-03 02:24:06 +05:30
John Paul Soliva
b5b2ab1bb0 fix(mcp): forget supervised groups once nothing is left alive
Review raised a real gap: we reap by pgid, so a registration is only as
meaningful as the group's identity. A group we deliberately keep registered --
an orphan teardown failed to kill, such as the `node` mcp-remote leaves behind
-- can later exit on its own, after which the kernel may hand that pgid to an
unrelated process owned by the same user. An ungraceful Hermes death while the
registration is stale would then signal a stranger. `_is_safe_target` cannot
catch it, because the value is stale rather than invalid.

Prune registrations whose group has no members left on every registration
change, and tell the supervisor to forget them. Signal 0 is a pure existence
question -- it cannot terminate anything -- so this is cheap and safe to run on
the hot path. An ambiguous answer (EPERM: exists but not ours) keeps the
registration, since dropping real coverage is the more expensive mistake.

This narrows the window rather than closing it: a group can still die and its
pgid be recycled between two probes. Closing it completely means proving group
identity at reap time, e.g. stamping children with a boot-unique env marker and
checking a member still carries it. That was judged not worth putting a `ps`
parse into the one process whose job is to stay simple enough to always work,
so the residual is now documented in the module docstring instead, along with
the note that Hermes's existing killpg-based orphan cleanup already carries the
same exposure (upstream #88350).

Also records why supervisor recovery is deliberately two-step: a failed write
drops the handle, and the next call rebuilds coverage from `_supervised_pgids`,
which -- not the pipe -- is the record of what needs reaping.

The protocol tests register synthetic pgids that were never real process
groups, so they now state that precondition through an `all_groups_alive`
fixture instead of depending on pid-space luck.

29 tests pass. Verified the change adds no failures: same selection, my changes
stashed vs applied, 42 failed either way (pre-existing in a locally rebuilt
venv) with passed going 617 -> 619 for the two new tests. Footgun linter clean.

(cherry picked from commit 3a7d219620a8cb57a7fe2c2346ce3f055320e743)
2026-09-03 02:24:06 +05:30
John Paul Soliva
2d783a15eb perf(mcp): one parent-death supervisor per process, not one per stdio server
Every stdio MCP server was wrapped in its own CPython watchdog that polled
getppid() every 2s to notice an ungraceful Hermes exit (kill -9, OOM, crash,
force-quit), since macOS has no PR_SET_PDEATHSIG. That is a whole interpreter
per server for a job that does nothing until the moment Hermes dies: 10.1 MB
physical footprint each, measured on macOS/arm64.

Replace the fleet of pollers with a single supervisor per Hermes process
holding the read end of a pipe only Hermes writes to. Death detection becomes
EOF on that pipe -- exact and instant, rather than up to a poll interval late.
Hermes sends `register <pgid>` / `unregister <pgid>` as servers come and go;
on EOF the supervisor killpg's whatever is still registered, which is exactly
the set whose teardown never ran. A clean shutdown unregisters as it goes, so
EOF then finds nothing to kill.

Servers are now spawned unwrapped. The MCP SDK already starts each stdio child
in its own session, so the pgid recorded for killpg is the server's own group
and the existing cleanup paths reach it unchanged. That also deletes the
signal-forwarding layer the wrapper needed: wrapping had put the real server in
a different session from the pgid being tracked, so a graceful killpg would
have hit only the wrapper.

Measured on a 5-gateway host: 10 watchdogs (~98 MB) -> 5 supervisors (~49 MB).
One supervisor costs about what one watchdog did (9.9 vs 10.1 MB), so the win
is (servers_per_process - 1) x ~10 MB, and a process with no stdio servers now
spawns nothing at all.

The supervisor reads length-capped lines rather than iterating the stream: a
writer that never sends a newline would otherwise grow it without bound, which
it must not be vulnerable to when it is the last defense against leaked
servers. Found by feeding it /dev/zero, where it reached 15 GB.

Verified beyond unit coverage: a real stdio MCP server connects and its tools
are discovered on the unwrapped path; with a live Hermes holding a real
connection, kill -9 reaped the server, its grandchild (in the server's group,
the mcp-remote `node` case), and the supervisor exited on its own. The reap
tests were sabotage-checked in both directions -- a no-op reaper fails all
three, while the test pinning that a cleanly unregistered server survives keeps
passing -- and each wiring half fails independently when removed.

(cherry picked from commit a252d4ce7ff1722f687635fdbf0cff79f538c3f1)
2026-09-03 02:24:06 +05:30
kshitijk4poor
75bcd86687 fix(process): keep the sandbox log poller from splitting a UTF-8 character across polls
Follow-up to #92164: the delta window now ends on a character boundary
(up to 3 trailing continuation bytes are held for the next poll), so
multibyte output no longer decodes to U+FFFD at the seam. Verified on
bash, dash and busybox sh; exhaustive-prefix regression test added.
2026-09-03 02:20:46 +05:30
kshitijk4poor
09fa753366 refactor(terminal): share the force_remove-aware env teardown between cleanup_vm and the probe
Extract the signature-checking cleanup dispatch that cleanup_vm already had
into tools.terminal_tool._cleanup_env and call it from the prompt-time
backend probe instead of carrying a second copy. Comment on the probe
trimmed to the non-obvious part (why ssh is skipped).
2026-09-03 02:07:42 +05:30
kshitijk4poor
c81cec5797 refactor(tools): collapse the OSV disk-cache load except ladder
Same retry semantics (only a non-FileNotFoundError OSError leaves the latch
unset), one latch assignment instead of five; docstring now says the call
runs per get/put and does work once.
2026-09-03 02:05:42 +05:30
kshitijk4poor
cf71a60c65 refactor(tools): write the OSV disk cache through utils.atomic_write_text
Follow-up to the cherry-picked change: reuse the repo's shared atomic writer
(adds the fsync the hand-rolled mkstemp+os.replace skipped, keeps 0600 on
create) and document the clean-verdict staleness trade-off the cross-process
cache introduces.
2026-09-03 02:05:42 +05:30
Christopher
c7a87bb110 fix(tools): retry OSV disk-cache load after transient I/O
A busy or briefly unreadable cache file must not disable disk loads for
the rest of the process; only missing or permanently malformed files
should mark the cache as loaded.
2026-09-03 02:05:42 +05:30
Christopher
8196d409a0 perf(tools): persist OSV malware-check verdict cache to disk
The OSV preflight is a synchronous network POST to api.osv.dev (up to 10s timeout) on every MCP stdio server start. Tools like `hermes mcp test` and MCP reconnect ladders spawn the same package repeatedly, but the previous in-process cache was empty after every process restart, so each run re-queried OSV and added 5.91x variance to the connection-time span.

Persist the malware-check verdict cache to `<hermes_home>/cache/osv_check.json`. Cache expiry is stored as an absolute wall-clock timestamp, so it survives restarts and monotonic-clock skew. Loading only adds missing keys so an in-memory overwrite (e.g. a test forcing expiry) is not silently reversed by the disk copy. Writes are atomic (temp file + rename) and happen under the existing cache lock.

Fixes the Hermes-owned OSV preflight variance component of #68416. Server-side `initialize` time is outside Hermes' control.

- Adds `hermes_constants.get_hermes_home()` lazy import to keep `tools/osv_check.py` import-safe (stdlib + typing only at module scope).

- Switches cache timestamps from `time.monotonic()` to `time.time()` for persistence compatibility.

- Updates `tests/tools/test_osv_check.py` fixture to isolate disk cache per test via `HERMES_HOME` + `tmp_path`, and adds regression tests for persistence, reload, and disk format.
2026-09-03 02:05:42 +05:30
Adolanium
8b681f70ea perf(process): poll sandbox job logs for new bytes only
The background-process poller for non-local backends ran `cat` on the
whole log file every two seconds, then threw away everything except the
part it had not seen yet. The offset it needed was already tracked one
line below, so the full read was pure waste.

Cost of one poll grew with the total output so far, which makes the cost
of a run grow with the square of its length. A job writing 10 MB over an
hour moved about 9 GB across the docker or SSH channel to deliver 10 MB
of output.

The poller now asks the shell for the file size and the bytes after the
offset in one command. Reading the size first and cutting the tail at
that same size keeps the two in step, so a file that grows mid-command
never sends a byte twice. A file that shrank was rotated or truncated,
so the offset drops back to 0 and the buffer is dropped.

The output buffer is now appended to rather than replaced, matching the
local reader loops, and the offset is counted in bytes because the shell
counts bytes.
2026-09-03 02:00:35 +05:30
kshitijk4poor
ff7233b815 fix(credential_files): apply the same exclusions to the symlink-safe mount copy
_safe_skills_path() is the sibling of iter_skills_files(): when a symlink in
skills/ forces a sanitized copy for mount-based backends (Docker/Singularity),
it rglob-copied the whole tree — .hub, .curator_backups, node_modules and all.
Prune EXCLUDED_SKILL_DIRS before descending, same rule as the sync generator,
so the mounted copy never carries (or walks) the bookkeeping trees either.
2026-09-03 01:33:18 +05:30
kshitijk4poor
1d06ef3a5d refactor(credential_files): prune excluded dirs before descending in the sync walk
Replaces the three hand-copied rglob loops + post-hoc parts check with one
os.walk generator that drops EXCLUDED_SKILL_DIRS from dirnames before
recursing. Same file set as the cherry-picked fix (the test binds it), but
the walk no longer stats every file under .hub/.curator_backups/node_modules
on each 5s FileSyncManager tick.

Bench (synthetic skills tree: 20 skills + 400 .hub files + 5x8MB curator
tarballs + 50 archived files): iter_skills_files() 35ms -> 2.4ms.
2026-09-03 01:33:18 +05:30
Carry00
edac49e473 fix(skills): stop syncing bookkeeping dirs to sandboxes
iter_skills_files() walked the skills tree with a bare rglob("*"), so the
.hub download cache, .archive, curator backups, and any node_modules/.git
under a skill package were uploaded to the sandbox on every sync. The
sandbox never reads them: skill content is resolved host-side.

EXCLUDED_SKILL_DIRS is already the canonical exclusion set, honoured by
discovery and backup. Apply it to the sync path too, across all three
roots iter_skills_files() walks (local, external, project-local), and add
.curator_backups to the set.

Measured on a local install: 900 files / 67.3 MB -> 771 files / 8.4 MB.

This is not just wasted bandwidth on the SSH backend, where the oversized
payload can exceed the 120s _ssh_bulk_upload deadline and surface as the
agent hanging on every tool call.

The filter intentionally does not reuse is_excluded_skill_path(), which
also prunes references/, templates/, assets/ and scripts/ -- those hold
support files and bundled scripts the sandbox does read and execute.
2026-09-03 01:33:18 +05:30
Teknium
f6234d00c5 fix(security): close GitSpawn RCE class — malicious repo .git/config no longer executes on context gathering (GHSA-7x36-8jrh-v4pw)
Hermes gathers workspace context by running git against the session
directory automatically — the coding-workspace snapshot, gateway
project-tree build, /diff, @diff|@staged context refs, goal-gate
fingerprint, and -w startup worktree add — before any prompt, tool call,
approval, or trust gate. Those probes ran the system git without
stripping the repository's own config, so a repo delivered as files with
its .git directory intact (a shared zip, sync folder, or USB stick;
git clone never transfers .git/config) could set an execution-sink git
setting and get arbitrary host code execution as the user with nothing
on screen.

- core.fsmonitor / core.hooksPath / pager / editor / credential helper:
  neutralized by routing every automatic probe through
  noninteractive_git_env(), which pins those keys to inert values via
  GIT_CONFIG_* and ignores global/system config. bounded_git_probe (the
  reported sink, coding_context._git + tui_gateway.git_probe) now defaults
  to that env; worktree-add, working_diff, web_git, context_references,
  goals, and subagent_worktree route through it too.
- Attribute-scoped [diff "x"] command=/textconv= drivers: the attacker
  names the driver in .gitattributes, so GIT_CONFIG_KEY overrides can't
  enumerate them. Added harden_git_argv(), which inserts
  --no-ext-diff --no-textconv on diff-rendering subcommands (diff/show/
  log/blame) only — status et al reject the flags. Both flags required
  (verified empirically; each alone leaves the other live).

Builds on the noninteractive_git_env config-scrubbing from the
gemini-cli #28792 port. Real-git E2E regression suite arms a malicious
repo and asserts every automatic path neutralizes fsmonitor, hooks,
external-diff, and textconv; a baseline test proves the repo is armed.
2026-09-02 10:33:43 -07:00
Teknium
7840a0e2d9 feat: delegation batch tags read "set N" instead of a hex id slice
Interleaved subagent fan-outs were tagged with the first 4 hex chars of the
delegation id ([b2ac 3/9]), which is attributable but unreadable. Batches are
now numbered in order of appearance per process: [set 1 · 3/9], [set 2 · 1/7].
Desktop /agents already labels groups "Delegation N", so its duplicate hex
badge is dropped.
2026-09-02 10:12:54 -07:00