Commit Graph

4626 Commits

Author SHA1 Message Date
kshitijk4poor
8ad7ae06e7 test(mcp-oauth): bind the 400-recovery path to issuer binding, guard live-TTL on an empty context
The 400-recovery reload installs a disk pair and only then re-runs issuer
binding on it. When that pair was minted by a different issuer the enforcer
strips its refresh token, and the reload must report "no recovery" so the
session is cleared like any other dead grant. No test pinned that verdict:
a reload that ignored the install result would keep a stripped pair in the
context and return True. The new invariant drives a real 400 through
_handle_refresh_response against a foreign-bound disk pair and asserts the
result is False, the context is cleared, and the foreign refresh token does
not survive on disk.

_hermes_live_ttl read expires_in off current_tokens without checking for
None; getattr(None, ...) happened to yield the default and report "live".
Both current callers install a pair first, but the helper is now explicit
that an empty context is never live, so a future call site cannot adopt
nothing.
2026-09-15 10:50:13 +05:30
kshitijk4poor
f4bf786c57 fix(mcp-oauth): return a verdict from disk-pair install, count a missing expiry as live
Two defects in the peer-adoption path of the refresh fence:

1. `_hermes_install_disk_pair` raised `_RefreshCompletedByPeer` when issuer
   binding stripped the candidate's refresh token. That is the right outcome
   for `_refresh_token` (restart the flow so the SDK lands in 401 -> full
   auth), but `_hermes_reload_tokens_after_refresh_failure` shares the helper
   and must instead treat the candidate as rejected: restore the previous
   pair and return False so the caller clears state and prompts. The helper
   now returns whether a refresh token survived binding and each caller
   decides; the one-shot adopt wrapper is inlined into `_refresh_token`.

2. `_hermes_live_ttl` treated `expires_in is None` as expired. RFC 6749 makes
   `expires_in` optional, the SDK's `is_token_valid()` is True with no expiry,
   and `_rebase_expires_in` preserves None on read, so a peer's rotated pair
   without an expiry was never adopted and we POSTed its refresh token
   anyway, burning a generation on single-use providers. None now counts as
   live; the try/except around a pydantic `int | None` field is dropped.

One new test drives the real auth flow against a peer pair with no
`expires_in` and asserts the pair is adopted with zero POSTs.
2026-09-15 10:50:13 +05:30
kshitijk4poor
2e89c5da48 fix(mcp-oauth): keep the refresh-fence sidecar when removing token state
flock is bound to an inode, not a path. Unlinking `<srv>.json.refresh.lock`
from `remove()` while a peer still holds the fence lets the next acquirer
open and lock a brand-new inode, so two processes hold "the" fence at once
and the single-use refresh token can be consumed twice. On Windows the unlink
of a locked file raises PermissionError straight out of `remove()`/`restore()`.

A 0-byte 0600 sidecar in a 0700 directory is harmless, so leave it in place.
It stays out of `_state_paths()` so `restore(only_if_absent=True)` still keys
off real token state only.
2026-09-15 10:50:13 +05:30
kshitijk4poor
60262f71bd refactor(mcp-oauth): split the refresh fence into acquire/release functions, drop the lock sidecar on remove
The SDK drives a refresh as a generator (request yielded from
_refresh_token, response consumed in _handle_refresh_response), so the
fence was a hand-driven @asynccontextmanager: __aenter__ in one method,
__aexit__ in another, generator object stashed on the provider. A plain
`acquire_refresh_fence(path, timeout) -> fd` / `release_refresh_fence(fd)`
pair says what actually happens and leaves nothing half-entered to leak.
The descriptor is opened with os.open at 0600 and closed on every
acquisition failure.

The `_refresh_token` release-on-exception stays: `_refresh_token` is also
reachable outside `async_auth_flow` (tests call it directly), and the
wrapper's finally only covers the generator-driven path.

HermesTokenStorage.remove() now unlinks the `.refresh.lock` sibling so
logout leaves no stray file. It is deliberately NOT added to
_state_paths(): snapshot()/restore(only_if_absent=True) treat any
existing state path as "newer state exists", and a lingering lock file
would silently veto a rollback.

tests/tools/test_mcp_oauth.py: drop the unused `Path` import (`time` is
still used by the socket poll helper).
2026-09-15 10:50:13 +05:30
kshitijk4poor
f07ea70de7 fix(mcp-oauth): only treat contention errnos as "a peer holds the fence"
The poll loop swallowed every OSError from the lock syscall as contention,
so a filesystem that cannot take advisory locks at all (ENOLCK on some
network mounts, EMFILE, ...) stalled for the full 60 s deadline and then
blamed a peer. Only EWOULDBLOCK/EAGAIN/EACCES/EDEADLK mean "held by
someone else"; anything else now raises RefreshFenceTimeout immediately
with the real errno. Still fails closed -- the refresh is never POSTed
without ownership -- but the failure is diagnosable and instant.

The errno set mirrors cron.scheduler._is_lock_contention_errno; it is
duplicated rather than imported because importing the scheduler pulls in
the whole cron module graph for a four-value tuple.
2026-09-15 10:50:13 +05:30
kshitijk4poor
8933d355a9 fix(mcp-oauth): share one rotated-candidate rule between adopt and reload, re-bind issuer on every disk pair
Both fence paths that pull a peer's pair off disk now go through
_hermes_rotated_candidate (different, non-empty refresh token + non-empty
access token) and _hermes_install_disk_pair, which runs
enforce_refresh_token_issuer on the installed pair. Before, the adopt path
skipped the issuer check entirely, so a pair minted by a different issuer
could be POSTed straight to the new one.

The adopt path installs the candidate even when its access token has
already expired: the POST we are about to build needs the new refresh
token, and skipping the POST (_RefreshCompletedByPeer) is only correct
when the peer's access token is live with a positive TTL, mirroring the
reload path's clamp-to-zero guard. When the issuer enforcer strips the
refresh token there is nothing to refresh with, so the flow restarts into
401 -> full authorization instead of failing with OAuthTokenError.

The reload helper drops its outer except-Exception: get_tokens already
returns None for absent or corrupt files, so the blanket catch only hid
programming errors. Comment updated: this path exists for writers outside
the fence (interactive login, pre-fence Hermes), not for a fenced peer.
2026-09-15 10:50:13 +05:30
kshitijk4poor
76e2e8fb64 refactor(mcp-oauth): drop the per-access token-store lock now that the fence owns the refresh
`_token_store_lock` serialized a single get_tokens()/set_tokens() call and
then released. Its two justifications no longer hold:

- torn reads: `_write_json` goes through `atomic_json_write` (write to a
  sibling, rename), so a reader can never observe a half-written token
  file, locked or not;
- the read-modify-write of a single-use refresh token: a lock released
  between the read and the POST cannot close that race. `_refresh_fence`
  now spans read -> POST -> persist, and every store access on the refresh
  path (adopt-from-disk read, post-failure reload, `_store_tokens` write)
  runs inside it.

The remaining unfenced accesses are the cold `_initialize` read and the
authorization-code exchange's full overwrite -- neither is a
read-modify-write, so neither needs mutual exclusion. Keeping a second,
fail-open lock layer only adds a 10 s stall on a stale lock file with no
correctness gain. Tests that exercised the removed lock go with it.
2026-09-15 10:50:13 +05:30
kshitijk4poor
1a1345aba4 fix(mcp-oauth): make the refresh fence async and skip the POST after adopting a peer's rotation
The fence is entered from the SDK's coroutine-driven auth flow, so its
acquire loop spun on time.sleep(0.05) and blocked the whole event loop
for up to 60 s while a peer finished its network round trip. The fence
is now an async context manager that polls a non-blocking flock with
asyncio.sleep, and it creates the token directory itself (the parent may
not exist yet on a first refresh; secure_parent_dir only chmods).

The in-process RLock layer is dropped: an advisory lock on a fresh
descriptor already excludes sibling tasks and threads of the same
process, and a thread RLock is reentrant across asyncio tasks on one
thread, so it excluded nothing there anyway.

After acquiring the fence the provider re-reads the store; when a peer
already rotated the pair and the adopted access token is valid, it
raises _RefreshCompletedByPeer instead of building the refresh request.
The auth-flow wrapper restarts the SDK flow so the original request goes
out with the winner's access token. Previously the loser adopted the
new pair but still presented its stale refresh token, burning a
generation on every single-use provider.

Design lifted from #71715.

Co-authored-by: Kevin Yin <182213728+yinkev@users.noreply.github.com>
2026-09-15 10:50:13 +05:30
anhtahaylove
a0810c9cc9 fix(mcp-oauth): fence one refresh generation across the consuming POST
The token-store lock is per-operation: get_tokens() and set_tokens() each
take it and release it. With a provider that issues single-use refresh
tokens, two processes can therefore both read R1, both POST it, and the
loser gets invalid_grant on a session that was healthy:

    A: get_tokens() -> R1   (lock taken and RELEASED)
    B: get_tokens() -> R1   (lock taken and RELEASED)
    A: POST R1              -> 200, receives R2
    B: POST R1              -> 400, credential already burned

Add _refresh_fence(), held across read -> POST -> persist so exactly one
process consumes a refresh generation. It fails CLOSED: unlike the
token-store lock it raises RefreshFenceTimeout instead of degrading to
unlocked, because proceeding without ownership is the race itself. It
locks a .refresh.lock sibling rather than the token file, since
flock/msvcrt locks are per-descriptor and nesting one path would
self-deadlock on Windows and silently no-op on POSIX.

The provider takes the fence before its final read, re-reads under it so
a peer rotation is adopted instead of overwritten, and releases in
_handle_refresh_response. async_auth_flow also releases on abandonment:
a cancelled generator never reaches the handler, which would strand the
fence and turn the race into a deadlock. That wrapper delegates send and
throw manually -- async generators have no yield from, and `async for`
would feed the SDK response to the inner generator as None.

tests/tools/test_mcp_oauth_refresh_fence.py is an acceptance test, not a
unit test: two real OS processes refresh against a single-use-token
authorization server that audits every redemption. It asserts R1 is
presented exactly once, neither process clears the session, and disk
converges on the newest token. Verified to FAIL without the fence
(audit=[rt-1, rt-1]); a threading-only lock cannot catch this.

(cherry picked from commit c1508a1d47e2cbd94e05fa507684f3716a1e988c)
2026-09-15 10:50:13 +05:30
anhtahaylove
47ee79a647 fix(mcp-oauth): serialize token-store access across processes
The desktop app spawns 'serve' while the scheduled task runs 'gateway run',
so two backends routinely share one HERMES_HOME. With a provider that issues
single-use refresh tokens, both could POST the same token and the loser's
refresh was rejected, clearing an otherwise healthy session.

Guard the token file's read-modify-write with a bounded advisory file lock
(fcntl on Unix, msvcrt on Windows, in-process only where neither exists),
mirroring how cron/jobs.py guards jobs.json. Acquisition is non-blocking with
a 10s ceiling: a briefly-contended refresh beats a permanently stuck client.

(cherry picked from commit 16f0d811de446a66ed5fd061fd7594fca3d230da)
2026-09-15 10:50:13 +05:30
kshitijk4poor
73f808e47f fix(plugins): only mark a timed-out hook worker abandoned while it still holds its token
The timeout branch of _run_hook_callback_bounded unconditionally added
gate_key to _hook_abandoned. A worker that finishes between done.wait()
returning False and the caller taking the lock has already popped its
token via _release_token, so nothing would ever clear that entry: the
callback stayed blocked for every later call id until reload with no
thread behind it. Guard the insert on the worker still being registered.

The new test makes the race deterministic by swapping the module's
threading.Event for one whose wait() lets the worker finish and then
reports a timeout, and asserts a fresh call id still runs.

Also pass tool_call_id inline from terminal_tool_result instead of the
conditional dict plumbing: an empty id is already treated as "no
identity" by _hook_call_identity and unknown fields are withheld from
narrow-signature callbacks (same shape as _fire_approval_hook). Update
the stale "(hook_name, id(cb))" comment above _hook_running_callbacks.
2026-09-15 10:48:49 +05:30
kshitijk4poor
4bd38ec9fd fix(plugins): give output-transform hooks a call identity for the callback gate
Hook callbacks are gated per (hook, callback, call identity). Two hooks
on the tool-loop path fired without any identity, so concurrent terminal
calls in one turn, or overlapping turns, still collapsed onto a single
gate key and the second invocation was skipped as if a callback had hung.

transform_terminal_output now forwards the tool_call_id bound in the
approval context around dispatch (only when set); transform_llm_output
forwards the turn_id already in scope. Payloads are additive: the
dispatcher withholds unknown fields from narrow-signature callbacks.
2026-09-15 10:48:49 +05:30
teknium1
ab0d4735a7 fix(skills-guard): socat only flags a reverse shell when an address spec follows
`\bsocat\b` under IGNORECASE matched "SOCAT", the Surface Ocean CO2 Atlas,
in every oceanography skill of a 2,110-file research bundle (17 critical
findings in one file), burying the bundle's real issues under noise. A real
socat relay always names an address type (TCP:/UDP:/OPENSSL:/EXEC:/SYSTEM:/
PTY:/UNIX-…:), so the pattern now requires one on the same line. `nc -l` /
`ncat -l` are unchanged. Scanner version bumped to v5 so cached verdicts
re-scan.
2026-09-14 21:55:24 -07:00
teknium1
f5a457ad5b fix(tools): one-shot linger waits for a completion that is mid-publish
`ProcessRegistry._move_to_finished` pops the session out of `_running`, then
saves the receipt, releases handles and writes the checkpoint, and only THEN
enqueues the completion and sets `_completion_event`. A quiet one-shot parent
whose turn ends inside that window called `wait_for_pending_completions`,
found nothing in `_running`, drained an empty queue and exited without the
follow-up turn. That is the CI flake in
tests/tools/test_completed_process_results.py::test_headless_terminal_result_survives_cli_exit
(`follow_ups == []`), which also hit unrelated branches.

Consider `_finished` sessions whose event is not yet set as pending too.

Repro: a 6s sleep before the enqueue plus a 2s delay before the parent's first
wait fails the E2E 2/2 on main and passes 2/2 with this change.
2026-09-14 20:44:56 -07:00
teknium1
a4d474777f fix(computer-use): doctor diagnoses a denied cua-driver spawn instead of crashing
On Windows the Hermes venv interpreter cannot CreateProcess a binary under
C:\Program Files\WindowsApps (WinError 5) even though the shell resolves it,
so `hermes computer-use doctor` died with a raw PermissionError traceback
from _open_mcp. Catch the spawn OSError and print what failed, why the tool
may still work (PATH resolves another copy), and the fix (reinstall outside
WindowsApps or HERMES_CUA_DRIVER_CMD), exit 2.
2026-09-14 17:34:31 -07:00
Jeff J Hunter
65625dfe87 fix(computer_use): attach element_token when the driver schema accepts it
cua-driver 0.21 refuses a bare element_index:

    click: bare element_index is not accepted; pass element_token,
    or snapshot_id together with element_index

_maybe_attach_element_token gated solely on the trycua/cua#1961 capability
vocabulary. 0.21 stopped publishing per-tool capability sets — every tool
reports an empty set — while still accepting element_token in its input
schema. The gate therefore fails closed on 0.21.x and we send the bare
index, so the driver rejects the call.

The effect is total: every element-targeted click is refused, leaving
agents with only blind pixel coordinates. Observed against cua-driver
0.21.0 on X11, where a capture returned 762 elements with 762 tokens
cached and every subsequent click still failed with snapshot_id_required.

Check the live input schema first — supports_input_property() already
exists for exactly this, and its docstring notes it "deliberately inspects
tools/list rather than ... requiring a capability token the driver never
shipped". The capability check is retained as a fallback so drivers that
did ship the vocabulary are unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-14 17:34:31 -07:00
teknium1
5d8390d1a4 fix(cron): retain Bot Chat output while a CLI owner is open
Keep never-started output behind unsupported owners and drain in admission
order after release. Persist claims before execution and never replay uncertain
started turns. Existing supported-owner receipts keep their authority.

Credits 686f6c61's residual queue proposal in #100319. This is a scoped
implementation, not general retry of failed CLI subprocesses.

Native Electron before/after: CLI-owned target previously returned
SESSION_NOT_OWNED and remained empty after release/tick; now its queued
output and reply appear once in the target Bot Chat. Nested quiet CLI
message_agent delivery to a named Desktop owner also passes on base.
2026-09-14 17:29:32 -07:00
liuhao1024
37243bd668 fix(tools): resolve hermes CLI beside the interpreter in bot_mode_dm deliveries
Bot-to-bot message_agent delivery builds both transport argvs (local
teammate chat and peer dm) with a bare "hermes" as argv[0]. Since #96631
the delivery runner spawns under terminal_tool's isolated host-local
environment, which does not inherit the gateway's PATH — so on
docker/service installs (venv at /opt/hermes/.venv) every delivery exits
with FileNotFoundError: 'hermes'.

Resolve the CLI with bot_relay._hermes_cli() (#93590) — the venv sibling
of this interpreter, then shutil.which, then the bare name — at both
argv construction sites. The turn-lock matcher in _delivery_lock()
already matches argv[0] by basename, so absolute paths lock exactly as
before.

Fixes #100662
2026-09-14 17:04:56 -07:00
teknium1
1a990f3062 fix(skills): keep scanning link-shaped arguments inside fenced code blocks
Masking every balanced [x](dest) in a .md file also blanked
`cp [k](../../../.ssh/id_rsa) /tmp` inside a ```sh fence, which main scored
caution and the branch let through as safe. A link inside a code fence is a
command argument, not a hyperlink: toggle masking off between fence markers.
2026-09-14 16:28:49 -07:00
Puvaan Raaj
b259f2401d fix(skills): ignore traversal in markdown links
(cherry picked from commit 92b48941b707df4fd3c13029b157a3f28188b4c4)
2026-09-14 16:28:49 -07:00
teknium1
de16ce9d0c fix(threat-scanner): gate ssh_access on every mutating verb, not just copy verbs
Review of the write-verb gate found sed -i, chmod, truncate, curl -o, wget -O,
git clone and a scripted open(...) against ~/.ssh all slipping to no finding,
where the bare path regex on main caught them. Add those verbs and the open(
shape to the gate; read-only mentions stay clear.
2026-09-14 16:16:08 -07:00
teknium1
b1733fd085 fix: keep the ssh_access id, word-bound the verb gate, collapse tests to two invariants
Salvage trim of #89249:
- keep pattern id `ssh_access` (test_memory_tool and callers key on it; renaming buys nothing)
- word-bound the verb alternation: the unanchored form matched `add` inside "address" and
  `dd` inside "middle", so read-only prose still fired; `\b` closes that leak
- fold the second "bare leading redirect" regex into the same alternation (`>>?` branch)
- tests: one parametrized "write shapes still fire" (echo, cat, cp, tee, mv, install,
  printf, dd, scp, rsync, ln, leading redirect, option clusters) and one "read-only
  mention does not fire"; the trade-off/change-detector tests are dropped
- test_memory_tool: the persistence fixture used a read-only mention; use a write shape
2026-09-14 16:16:08 -07:00
Martin Mogis
fe68349cd6 fix(threat-scanner): close reviewer-noted ssh_access_write bypass shapes
Extend the write-verb alternation with mv/install/printf/dd/scp/rsync/ln,
allow short option clusters between verb and path, and add a bare-leading-
redirect branch (a > ~/.ssh/... line carries no verb word at all). Prose
that names a write primitive before an SSH path still fails closed — a
false positive costs a review, a false negative is a backdoor.

(cherry picked from commit 4865b96ca3c8661c0ea6f5c479c198ccf68f86c4)
2026-09-14 16:16:08 -07:00
Martin Mogis
34304b722d fix(threat-scanner): require a write verb before SSH paths in the ssh_access strict pattern
The bare \$HOME/.ssh|~/.ssh regex fired on ANY mention of SSH paths in
scanned content, so operational documentation (VPS recovery notes, SSH
configuration write-ups stored in memory or skills) was blocked as a
persistence threat. Require an echo/cat/cp/tee/append/add/write/>>
verb before the path so only backdoor-insertion shapes match.

(cherry picked from commit 8c9b4c8972ac849e4c82fa20d355733ca3e5a231)
2026-09-14 16:16:08 -07:00
unsupportedpastels
4f00c456f0 fix(skills-guard): stop flagging sudo.request/sudo.respond event names as sudo usage
`sudo.request` and `sudo.respond` are gateway wire events: the masked
sudo-password prompt the terminal tool raises, which every client surface
(desktop, TUI, any plugin that relays secure prompts to another device)
has to name to forward it. The `sudo_usage` rule matched the bare word,
so any plugin listing those events scored `high` and every install of it
landed on `caution`, which Hermes Desktop cannot confirm past.

A dotted event name is never a shell `sudo` invocation. Exclude exactly
those two names with a negative lookahead; a real `sudo cmd` still fires.

(cherry picked from commit a7a3a31126de8057c2dbb9b7048724aa9d55a7fd)
2026-09-14 16:14:23 -07:00
chenxue
63395922bf fix(skills-guard): stop shell_rc_mod matching attribute access
The shell-startup-file pattern is
`\.(bashrc|zshrc|profile|bash_profile|bash_login|zprofile|zlogin)\b`.
Six of those names are distinctive enough that seeing them after a dot
means the file. `profile` is not: it is also how every language spells
attribute access, so `self.profile`, `user.profile`, and
`request.profile` each score a medium persistence finding.

The cost is signal, not blocking -- `_determine_verdict` treats
medium/low alone as informational -- but a plugin that happens to name a
field `profile` buries the findings a reviewer has to read. A model-
provider plugin whose tests exercise a `profile` object contributed 36
of 38 findings in its scan report, all of them this pattern.

Split `profile` into its own entry anchored on a non-identifier
character before the dot. Real references keep matching in the forms they
actually take (`~/.profile`, `"$HOME/.profile"`, `./.profile`,
bare `.profile`); attribute reads no longer do. The other six names are
untouched. Both entries keep the `shell_rc_mod` id, and scan_file
deduplicates on (pattern_id, line), so a line holding both still yields
one finding.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 705d0bddc61a1c55be82956d343b229544350eae)
2026-09-14 16:14:23 -07:00
teknium1
818d781c0b test(tools): trim the salvaged #110632 / #110307 tests to two invariants each
The file-safety salvage carried five tests and the execute_code salvage four;
fold them into the two behaviour contracts per fix: (a) the named-profile scope
exempts the root's direct files and the write lands with no prompt; a child env
sees the ACTIVE profile's home per turn; (b) negatives hold: a checkout's
.hermes/config.yaml stays gated fail-closed, a lookalike profiles/ tree exempts
nothing; a dedicated process with no override is untouched. Both red on base.
Also compress the _hermes_exempt_homes docstring (WHY only; cite #110630).
2026-09-14 16:13:51 -07:00
teknium1
09a7c297ec fix(tools): uncached check_fn probes classify UnscopedSecretError from the live scope
Under `gateway.multiplex_profiles: true` every gateway start logged a WARNING +
traceback from `tools.registry` for `_check_vault_available`. The browser vault
gate is a `no_cache_check_fn`, so `_check_fn_cached` short-circuited into
`_run_check_fn_uncached(fn)` without consulting the cache scope and the default
`unresolved_scope=False` hint reported an EXPECTED boot-time fail-closed read as
"scope was resolved" — the #100697 fix only reached the cached branch.

Derive the verdict from `current_secret_scope()` at the catch site instead of a
branch-derived hint: no scope installed → DEBUG without traceback (expected,
re-probes on the first scoped turn); scope installed but the read still failed
closed → WARNING + traceback (the probe dropped the scope on a bare thread/hop,
a spawn-site bug that must stay loud). Every uncached probe that reads a
credential is covered, vault included.

Fixes #110635. Credit @KoNit-K (#110638) for the report-side diagnosis; that PR
disabled the legacy-cloud probe for the vault gate only, which would have hidden
the vault from legacy Browser Use cloud users and left the class open for every
other uncached probe.
2026-09-14 16:13:51 -07:00
Kevin Rajan
7ef0b98272 fix(tools): honor multiplexed per-turn HERMES_HOME in execute_code child env
Under a multiplexed Desktop/Dashboard connection, one server process serves
multiple profiles, binding a context-local HERMES_HOME override per turn.
_build_child_env scrubbed the server process's os.environ - which carries the
machine-default HERMES_HOME - so skill scripts run via execute_code silently
read/wrote the wrong profile's directory.

Rewrite HERMES_HOME in the child env from the active override when one is
bound (the same per-turn rewrite apply_subprocess_home_env already does for
HOME). With no override (dedicated per-profile process) the inherited value
is left untouched.

Fixes #110303.
2026-09-14 16:13:51 -07:00
moxian
ab78e71a9e fix(file-safety): exempt the Hermes ROOT, not just the profile home
Under a named profile (``hermes -p <name>``, i.e. ``HERMES_HOME=<root>/profiles/<name>``)
the protected-instruction gate exempted only the profile dir, so the ROOT's DIRECT
files (LEDGER.md / MEMORY.md / USER.md / SOUL.md / AGENTS.md / DECISIONS.md ...) fell
through to the ``.hermes`` component rule and were treated as project-local
``<repo>/.hermes/config.yaml``. That gate always-asks and fails closed without a human
channel, so every write there was refused headless — it blocked #54 (LEDGER.md edit
could not land). Subdirectories (scripts/, cron/, data/) were unaffected, which is why
#55 sailed through and #54 did not.

``_hermes_exempt_homes()`` now returns the active profile home plus the Hermes ROOT
when that home really is a named profile (``named_profile_home`` validates: ``.hermes``
name, home markers, tombstone, or the resolved default root), so a coincidental
``profiles/`` directory never exempts its parent. ``_get_real_hermes_home()`` keeps its
per-profile semantics — the #107327 multiplex regression test pins that — and the
exemption consumes the new helper instead.

Refs: https://gitlab.yx.netease.com/agent-projects/platforms/agent-mentor/-/issues/60
2026-09-14 16:13:51 -07:00
teknium1
36c7f6c89d refactor(skills): trim the denylist demotion to one owner regex and two invariants
Fold the cherry-picked mechanism into the main file's compact style: one
denylist-owner regex replaces the assignment-target + name regex pair and the
_in_denylist_construct helper; the comment-prefix table loses its dead .css
entry; the denylist demotion lands on high (a confirmable caution) instead of
medium so the finding still gates the install and shell_rc_mod, already
medium, leaves the set. The verb guard widens to stem-prefix matching and
copyfile/copy2/sendfile per the review thread on #92632, closing the
Path.read_text()/shutil.copyfile shape. Tests trimmed to two invariants: a
skill's own denylist is caution (force-overridable), a real write to
authorized_keys stays dangerous.
2026-09-14 16:13:35 -07:00
Jack Lau
52bec9d77e fix(skills): stop scoring a skill's own denylist as an access
`_scan_file` matches every threat pattern against every line with no notion
of what the line is. The path-token patterns — `authorized_keys`, `~/.aws`,
`~/.ssh` — therefore cannot tell `cat ~/.ssh/authorized_keys` from a skill
that spells the path in order to REFUSE to read it. One `critical` becomes
`dangerous` in `_determine_verdict`, and on a community source that blocks the
install with no `--force`, so the skill that bothered to skip credential
files is the one that gets quarantined (#92478).

Demote, do not drop, following the precedent `allowed_tools_field` already
sets in this file: the finding keeps its file, line and matched text so an
auditor still sees the token; it just stops deciding the verdict alone.

Two contexts qualify, and only for the eight path-reference pattern ids:

- a whole-line comment, in a language that HAS comments, drops to `low`.
  Markdown is deliberately excluded: `#` opens a heading there, and Markdown
  prose is the prompt-injection surface itself.
- a line inside a construct NAMED as a denylist (`SKIP_PATTERNS`, `DENY_*`,
  `EXCLUDE_*`), carrying no verb that could touch the path, drops to `medium`.

The issue also suggested demoting any quoted token on a verb-free line. That
is wider than it looks — a fragment in quotes can be interpolated into a
command a line later — so the name test is the primary rule and the verb test
only guards it, because the construct's name is attacker-chosen.

The denylist check is statement-aware rather than per-line: the reported
match sat on a continuation line of a multi-line regex whose name is four
lines up, which a per-line test reads as anonymous.

8 tests. Reverting the demotion fails the two behaviour tests and leaves the
six contract tests green; dropping the verb guard, the Markdown exclusion, or
the statement-awareness each fails exactly its own test.

(cherry picked from commit f50acd34f62375926bd5843125e106f7e7eb6ee8)
2026-09-14 16:13:35 -07:00
teknium1
bb1d255a77 chore(scanners): bump skills-guard to v4 and plugin-guard to v3
Five scanner rule changes land together (prose/comment demotion, own-denylist
demotion, shell_rc/sudo token fixes, Markdown link masking, ssh write-verb
gate). Cached verdicts keyed on the old versions would keep previously
blocked skills and plugins blocked; one bump re-scans them.
2026-09-14 16:12:07 -07:00
teknium1
05e74268ea test: trim prose/comment guard coverage to invariants; drop the unused CSS comment prefix
Fold the three PRs' overlapping tests into two invariants per behaviour:
- comment/changelog prose demoted to caution and confirmable; trailing comments,
  runtime code and agent-facing docs still dangerous (#111193)
- Markdown plan/design prose demoted to caution; the same content in runtime
  code still dangerous (#103364)
- skills_guard: context_exfil needs a transfer directive; rm -rf under temp
  roots is not destructive_root_rm

The #111199 regression test is kept (behaviour is the same); its mechanism
(re-read the file per finding, cap every .md) was not carried since the
match-based cap already covers it without touching agent-facing docs.
'//' is not a CSS comment marker, so .css is dropped from the prefix table.
2026-09-14 16:12:07 -07:00
webtecnica
fd0de74bd9 fix(plugins): cut plugin-guard false positives on prose and agent-config-file refs (#103364)
(cherry picked from commit 4ca17d1de68c25e56fabd0a72dbca6c90b877bba)
2026-09-14 16:12:07 -07:00
Konstantin Khlopkov
a8b7af8586 fix(plugins): stop the install scanner from scoring hardening comments and changelogs as un-overridable criticals
A whole-line code comment or a changelog entry describing the threat a defense
rejects ("# a symlink could point at /etc/passwd, so ...") scored full severity,
driving the verdict to dangerous and making security-hardened community plugins
un-installable: --force does not override dangerous, so the only way through was
deleting the documentation of what was defended against. Whole-line comments and
CHANGELOG entries now cap one severity step lower (critical->high), keeping the
finding visible and the verdict at caution: blocked by default, --force
overridable. Trailing comments (executable code on the line), runtime code and
agent-facing docs keep full severity.

Fixes #111193

(cherry picked from commit 2a79f0a67e70eddcd387e759b1570d4feaee1e97)
2026-09-14 16:12:07 -07:00
Siddharth Balyan
cf35e7351e fix(connectors): manage_connections is absent for accounts the portal has not enabled (#111238)
A signed-in, paid Nous account that the portal had not enabled for
connectors got `manage_connections` in its schema and a raw "tool gateway
request failed with status 404" back from every call. The gateway answers
404 for any such account by design, and Hermes gated the tool on paid access
or a free tool pool, which says nothing about that.

The gate now reads the portal's own answer: a `managed_tools` token claim,
plus the existing free-tier leg. The gate is also the tool's check_fn, so a
session without the claim never sees the tool and the model has no 404 to
narrate. A token without the claim reads as not enabled.
2026-09-14 22:01:32 +00:00
Adolanium
30d78cd85e fix(plugin-guard): reconcile current scanner rules and test install confirmation 2026-09-14 23:42:17 +03:00
Siddharth Balyan
ee2f5629b8 Desktop connect runs on the connection operation: one card, no link to the model, no renderer polling (NS-868) (#110574)
* refactor(connectors): cut comments that restate the code

Connector modules (tools/connectors, tui_gateway connector RPCs, desktop
connector card/store) keep only comments that carry a non-derivable why or
a cross-module contract. No behaviour change.

* feat(connectors): managed connect runs on the connection operation

Managed `connect` / `reconnect` mint one ConnectionOperation for every target and, on a
desktop session, block the tool turn until the operation settles; the result is per-target
outcomes and never carries a connect link. Off the desktop the result carries the links and
returns at once (PR3 delivers them as their own message).

Why: the previous leg handed the model a URL and a `wait` verb, and the renderer ran its own
2s poller on top of the backend's 5s one; both walked the whole gateway catalog at two vendor
calls per page to read one row (~3 Composio calls/s per pending target). A hidden composer
message started the model's `wait` on the user's behalf. None of it was observable from the
operation the MCP leg already used.

What the operation looks like now:
- `contract.py`: TargetState / Actor / SettleReason enums and the `(kind, from) -> {to: actor}`
  transition table. `operation.transition()` enforces it; a card cannot claim a managed
  target `connected`, only the backend watcher can.
- `live.py`: one open operation per session, found by `op_id`. `connectors.operation.status`
  reads it, `connection.respond` drives it, `pending_connection` on resume replays it.
- `run.py`: the one lifecycle for both target kinds (prepare -> card -> wake/observe loop ->
  settle -> result). The managed `observe` hook polls the gateway list once per tick for the
  whole operation; the exact-status route replaces that call when the gateway ships it.
- `connection.update` is emitted on every transition and on settlement; registered in the
  shared event contract with the operation vocabulary typed on the TS side.
- `wait`, `_rendered_links`, `_seen_instructions`, the just-minted bounce and `_clamp_timeout`
  are deleted. `force` on `reconnect` always reinitiates; plain `reconnect` repairs only what
  the gateway reports disconnected.
- `connections.wait_timeout_seconds` is removed from config defaults, the example and the
  docs. The deadline is `OPERATION_DEADLINE_SECONDS = 300` in `operation.py`; the key was
  added on this unmerged train so no migration is needed.
- Wire model: `statusReason` parsed on connection results; the seven-state `connectionStatus`
  is typed on list items and an unknown value fails validation; `CONNECTION_REQUIRED` carries
  `connect_card_available` instead of the link when the session platform is `desktop`.

Session platform, not callback presence, decides whether a card exists: the GUI bridge
attaches callbacks to every backend session, terminal TUI included.

* feat(desktop): connector card subscribes to the connection operation

The card renders from the backend's operation instead of driving its own: `connector-flow.ts`
(the renderer's 2s `connectors.list` poller, its 120s client deadline and `keepWaiting`) is
deleted, and both hidden composer submits in `connector-tool.tsx` go with it. The model is
never nudged into a `wait`; the tool call is blocked on the backend until the operation
settles.

- `connection-request.ts` is the operation store: keyed by `op_id`, one entry per session,
  `applyOperationStatus` / `applyConnectionUpdate` as pure reducers, `respond` leaves the
  entry in place (the backend answers with `connection.update`), `ConnectionTargetOutcome`
  is a discriminated union the backend's transition table accepts.
- `input-requests.ts` applies `connection.update`; `connection.expire` and the resume
  snapshot correlate by `op_id` (a snapshot has no `request_id`).
- `ConnectorOffer` renders one `ConnectorCard` per target from a single
  `Record<ConnectionTargetState, phase>` table; Connect opens the stored link, Try again on
  failed / expired reissues through `connectors.connect` on the open operation, Not now is a
  per-target `skipped`, Continue settles. A settled operation renders `ConnectorSummary` rows
  with no live control.
- `tool-render-class.ts`: `manage_connections` renders the card regardless of
  `HERMES_GUEST_ONBOARDING`; the flag still gates the onboarding flow, not the card. The
  backend gate already decided admission; a card only exists because the tool was admitted.
- `mcp-setup-tool.tsx` speaks the same outcome vocabulary (connected / skipped / failed).
- `ConnectorRow.connectionStatus` is the seven-state literal union, not `string | null`.
- The guided-onboarding poller (`first-build-connectors.ts`) keeps its own row/phase types
  and compiles unchanged; PR3 moves it onto the operation.

anti-slop: no net-new findings (17 touched files vs 11d1a12472).

* fix(connectors): the card never parks the tool thread; every update carries the snapshot

Found by the pre-PR adversarial review and a real-path E2E test (both left in the tree).

- The desktop `connection_callback` was still `_block("connection.request", ...)`, which parked
  the tool thread on a private request-id Event until a `_respond` that no longer exists for
  this event. `connection.respond` settled the operation but the tool waited its full deadline
  before the watcher loop even started. The callback now only emits the card; the operation's
  own wake loop is the wait. The MCP leg's blocking bridge goes with it: the card answers
  through `connection.respond` like every other card.
- `connection.request` and every `connection.update` frame carry the full target snapshot
  (state, link, detail). The initial mint happened before the card existed, so the renderer
  never saw the links and Connect stayed disabled; a Continue settlement stamped
  `not_connected` on the backend while the card still showed `initiated`. The store now
  overlays the snapshot; no state is reconstructed from deltas.
- The `connection.update` emitter is a class-level `on_change` slot on the operation, set
  once by `register()` (a second `register()` no longer stacks wrappers); session lookup takes
  `_sessions_lock`; a re-minted link on an `initiated` target goes through `refresh_link()`
  and emits, instead of a bare attribute write.
- `session.interrupt` is checked before the first observe, so an interrupted call settles
  `interrupt`, not `all_resolved`.
- A gateway list reporting `expired` for an initiated target is recorded with actor `clock`
  (the contract's owner of that edge); it raised `IllegalTransition` before.
- Dead `keepWaiting` i18n keys from the deleted renderer poller removed.

tests/tui_gateway/test_connector_operation_e2e.py runs the desktop lifecycle through the real
tool, registry, gateway RPC handlers and callback bridge with only the HTTP client faked.

* docs(connectors): prompts and docs describe the operation, not the deleted wait verb

The onboarding prompts told the model to call action="wait" with timeout_seconds and to
expect a hidden [setup]/[connectors] note; both are gone. tool-search.md and
toolsets-reference.md said the model gets a connect link on the desktop. tui_gateway/AGENTS.md
gains the connection-operation row of the surface table.

* fix(connectors): the panel re-mints only a dead link

Try again on a failed or expired target mints a fresh link on the open operation. A waiting
target keeps the link it was minted with; the card reopens it and connectors.connect refuses
to spend a second mint (LINK_STILL_VALID). The unused refresh_link() goes. The package
docstring names the new siblings; the nine-name public surface is unchanged.

* test(connectors): the local-batch test answers the operation the way the card does

The callback stopped returning an answer in f782b26d98 (the card answers through
connection.respond); this test still returned one and waited out the 300s deadline in CI.

* ci: retrigger

* fix(connectors): the desktop card appears outside guided onboarding

Live on a signed-in macOS desktop, the two-app connect never showed a card. Three
defects, each hidden by a test that bound state the running app never binds.

The backend read the surface from HERMES_SESSION_PLATFORM only. The desktop and TUI
gateway bind it as HERMES_SESSION_SOURCE (_set_session_context), so session_platform()
was "" and managed connects took the off-desktop branch: links in the model's message,
no operation. session_platform() now reads platform, then source. The E2E test binds
through server._set_session_context instead of set_session_vars(platform="desktop").

The renderer routed manage_connections to the card only under isOnboardingEnabled(),
the HERMES_GUEST_ONBOARDING launch flag, in message-parts.tsx and the run splitter in
fallback.tsx. tool-render-class.ts had already dropped that gate in this PR; the two
routers had not. Both now route on the tool name alone.

ConnectorTool resolved the session owner by the runtime id. Owner routes, hints and
session rows are keyed by the stored id, so in registry topology the owner never
resolved and the card rendered null while the tool blocked. It now resolves by the
stored id, matching the PR1.5 card and every other owner lookup.

message-parts-connectors.test.tsx mounts the real Fallback router with the onboarding
flag off and distinct runtime/stored ids; red before each renderer fix, green after.

* style(connectors): shorter comments, no module mock in the card router test

The router test mocked isOnboardingEnabled to false; jsdom has no preload bridge, so the
real function already returns false. Comments that restated the code are cut to one line.

anti-slop: no net-new findings (25 touched files)

* fix(connectors): Connect on a waiting row opens the stored link

ConnectorCard derived the button's loading state from the phase label, so a managed row that
read "Finish connecting in your browser" (every row, since links are minted up front) had a
disabled Connect button. Nothing on the desktop could open the sign-in link; every managed
connect ended skipped, not_connected, or at the deadline.

The card now takes `busy` for "the action itself is running" and keeps `phase` as a label.
The MCP card passes its in-flight flag; the connector card passes the re-mint wait. Red before:
the Connect button on an initiated row rendered disabled and a click opened nothing.

* fix(connectors): a settled card stays dead; the card binds to its tool call only

A second connect for the same apps revived the finished card on the old tool row. The
connection.request payload carried no id, so the renderer fell back to matching rows by
connector names, and any row with those names qualified, settled or not.

The operation now records the model's tool_call_id and sends it in connection.request and in
the resume snapshot. The card binds to the tool row with that id and to nothing else; the
name-match fallback is deleted. A payload without the id is rejected by the store.

`reason` is removed from the tool: it was the only text the card ever showed from the model
and its absence forked a second tool part, since `reason` doubled as the row-correlation key
in tool-parts.ts. The card never needed it.

`connection.expire` is deleted from the contract and from _EXPIRING_REQUESTS: the card is
raised with _emit, not _block, so nothing has emitted it since the operation lifecycle landed.

Sid's rule of record: a resolved card is fully dead; no path brings it back.

* fix(connectors): the watch loop settles once, on time, and never raises into the result

Three findings from the live review, one loop.

Continue racing a finished sign-in: the loop ran the gateway read, then settled. A read that
returned `connected` for an already-settled or failed target raised IllegalTransition out of
the tool and the model got a generic error instead of the per-app outcomes. The read now skips
targets that are not live (pending, initiated) and skips a settled operation; the loop checks
`settled` after every read.

Settle reason as row text: `settle()` wrote `continue`/`deadline` into each unresolved target's
`detail`, and the card printed it in red. The reason stays on the operation only.

Stop and the deadline waited for the next tick: `/stop` sets a per-thread flag with no wake
hook, so the sleep is sliced at 250 ms and the flag and clock are read each slice. The clock is
also checked before each read, not only after.

Tests: a failed mint that later reads connected settles cleanly; Continue during a read keeps
the settled result; no reason in detail; an interrupt settles within the same second.

* fix(connectors): MCP setup off the desktop returns unavailable instead of blocking

run_mcp_operation treated a non-None connection_callback as "a card exists". Every tui_gateway
session has that callback, the Ink TUI included, so an MCP install from the terminal UI blocked
until the 300 s deadline while the docs promised `unavailable` with the terminal commands.

The MCP path now reads the session surface the same way the managed path does; the callback is
never the predicate. Test binds the surface to `tui` with the callback attached.

* fix(connectors): a failed Try again shows the failure, not the old dead link

The panel's re-mint ignored the gateway's per-app status and moved the row to `initiated` with
whatever link came back, `None` included, so a mint that failed again rendered as waiting on the
link that had already died.

One reader of a mint response now serves both the first mint and Try again
(`managed.mint`, with the actor as a parameter). A repeated failure keeps the row `failed`,
drops the link, and carries the vendor's new text through `operation.refresh`, which emits a
frame without a state change so the card redraws.

* fix(connectors): a forced reconnect waits for the new sign-in before it reports connected

`reconnect` with `force: true` is the account switch. The vendor keeps the old account active
while the new link waits, so the first list read after the mint said `connected` and the
operation settled at once: the new link was dropped and the model was told the switch was done.

A forced target is marked awaiting_new_attempt after the mint. The watcher ignores its row until
the list shows the new attempt (`connectionStatus: initiated`) once, then trusts `connected`.

* fix(connectors): the operation registers under the gateway session key

The tool registered the operation under the agent's session_id; every RPC (connection.respond,
connectors.operation.status, the panel's connectors.connect) and the update emitter looked it up
by the gateway's session key. Those agree until compaction rotates the agent id mid-turn; then
the card's clicks find nothing, no update reaches it, and the tool waits out the deadline.

The registration key is now the bound HERMES_SESSION_KEY, with the agent id as the fallback for
callers with no gateway (unit tests, a bare CLI). The E2E passes a rotated agent id and drives
the card by the gateway key.

* fix(connectors): the forced-reconnect gate reads any non-active row; a failed re-mint of an expired row is failed

Three follow-ups from the verification of the fix pass.

The awaiting_new_attempt gate cleared only on the literal `connectionStatus: initiated`. The
field is optional on the wire and `initializing`, `failed`, `expired` are valid values, so a
forced reconnect could wait the full 300 s and swallow a failed new attempt. The gate now holds
only while the row still reads as the old account (`connected` or `active`) and releases on
anything else.

Try again on an `expired` row whose re-mint fails raised IllegalTransition (no expired → failed
edge). The re-mint steps through `initiated` as the user's attempt, then `failed`, then drops the
dead link.

`detail` never carries a state name any more: `failed` as detail rendered as the row label and
made agent/display.py tag the settled result as a tool error. Only vendor text goes there.

`connection.expire` removed from the renderer's unscoped-stream set; nothing emits it.
2026-09-15 00:41:14 +05:30
Siddharth Balyan
d105376b21 Connector code lives in one package, tools/connectors/ (move only; NS-868 prep) (#110368)
* refactor(tools): discovery also scans tools/<pkg>/tool.py

A tool family that is a whole package had no way to register: discovery
globbed tools/*.py only and derived the module name from the filename.
Now the candidate list is tools/*.py plus tools/*/tool.py, merged and
sorted once so import order does not depend on depth (register() lets a
same-name duplicate overwrite silently), and the module name comes from
the path relative to tools/. Only tool.py is scanned inside a package,
so its siblings are libraries by construction. A package without an
__init__.py is skipped with a warning rather than registering from a
checkout and vanishing from the installed wheel.

The AST prefilter and the (mtime, size) disk cache are per absolute path
and work unchanged. The two hand-rolled tools/*.py enumerators in tests
now use the same candidate helper.

* refactor(connectors): one package for the connector domain, tools/connectors/

The connector code was spread across six flat files and a root-level
module that was a sibling of model_tools.py only by address:

  tools/connections_tool.py            -> tools/connectors/tool.py     (schema, register, dispatcher)
                                          tools/connectors/managed.py  (the managed leg, split out)
  tools/connections_tool_mcp.py        -> tools/connectors/mcp.py      (validation split out ->)
                                          tools/connectors/targets.py  (normalize_targets, validate_action)
  tools/connections_tool_operation.py  -> tools/connectors/operation.py
  tools/connector_search.py            -> tools/connectors/search.py
  model_tools_connectors.py            -> tools/connectors/dispatch.py
  tools/tool_gateway/                  -> tools/connectors/gateway/

Move only; every function body is unchanged. tools/connectors/__init__.py
is the door: nine names, the whole cross-package surface. model_tools and
tool_search deep-import a few helpers past it on purpose and the docstring
says so. The two split files make the import graph one-directional
(tool -> mcp -> targets, tool -> managed) where the old layout had
connections_tool importing validation out of the MCP file.

One behaviour-neutral seam change: the _connectors_available try/except
wrapper is gone. connectors_available() already fails closed, and both
the registry handler and the inline executor now read it as a module
attribute (gateway.config.connectors_available), so tests patch it in
one place instead of two. _default_client lives in managed.py, the only
module that calls it.

tools/managed_tool_gateway.py and tools/managed_gateway_auth.py stay:
they are gateway identity shared by tts, transcription, image and modal.

Test files follow their modules. No docs referenced the old paths; no
compat pointer is added (in-tree moves get none).

* ci: retrigger (zero-job dispatch on 142466de6b)
2026-09-15 00:41:13 +05:30
Siddharth Balyan
e0ef0eb9c3 manage_connections covers local MCP servers; setup_mcp leaves the schema (NS-867, PR1) (#109517)
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema

One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.

MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.

Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.

`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.

Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.

`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.

The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.

* wip(desktop): connection.request store, resume restore, card routing for MCP targets

Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.

* fix(config): hermes update turns on the connections toolset for saved toolset lists

`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.

Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.

`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.

* refactor: anti-slop pass on the desktop slice; shorten added comments

Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.

* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed

The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.

session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.

lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.

vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.

* chore: drop __pycache__ files swept in by an over-broad git add

* fix(desktop): correlate the connection.request row with the model's tool call by reason

The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.

* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds

* fix(connections): settle reason derives from target state, never from the renderer

A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.

* fix(desktop): a pending connection card re-arms on resume and activate

The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.

* style: literal wording in added comments, docstrings and docs

* fix: shared gateway-event contract and config-schema category for the connection events

connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.

* style: import order (perfectionist) in the desktop and shared files this PR touches

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-15 00:41:13 +05:30
teknium1
e383c28d2f fix(approval): judge shell quoting on the raw command, not the escape-stripped one
The hardline floor tracked quote state on text that
_normalize_command_for_detection had already rewritten (`\"` -> `"`).
That flips quote parity and broke both ways:

- false positive: a shell-valid `grep -o "[^\"]*" f` lexed as an
  unterminated quote and hit the unconditional "malformed executable
  payload" block. 118 of the 125 hardline blocks in one week of real
  agent use on this install were this shape, every one a benign grep,
  and the block cannot be bypassed by --yolo or approvals.mode=off.
- bypass: `cat "f\"n.txt"; rm -rf --no-preserve-root /` put the `; rm`
  start "inside" a phantom quote, so no command start was marked and the
  floor let the root wipe through with approved=True.

Fix: the malformed-quoting verdict reads the raw command (only quoted
newlines masked, which keeps quoting intact), and
_command_detection_variants adds a variant whose command starts were
marked on the raw command before normalization. The marker is " \n" so a
preceding literal backslash cannot eat it as a line continuation.
_iter_shell_command_starts no longer treats the `{` of `${...}` as a
brace-group opener, so the new variant does not split `${IFS}` and defeat
the IFS collapse.

Direction from #85922 by @Soju06, re-implemented onto the decomposed
tools/approval_detection.py; the parameter-expansion scanner from that PR
is replaced by the one-character `${` check above.

Co-authored-by: Soju06 <qlskssk@gmail.com>
2026-09-14 09:19:25 -07:00
kshitijk4poor
d328adf611 fix(bot-mode): strip canonical session env names, keep HERMES_SESSION_* knobs
The HERMES_SESSION_* prefix match also stripped non-identity knobs
(HERMES_SESSION_STALL_TIMEOUT tunes gateway stall watchers). The strip now
uses gateway.session_context's session env names (_VAR_MAP) — synced with
the session binding surface as vars are added — imported unguarded like
tools/approval_context already does on a hotter path; the except fallback
was an unreachable branch whose 3-name tuple silently narrowed the strip.
Test pins STALL_TIMEOUT as the kept negative case.
2026-09-14 21:18:47 +05:30
xxxigm
9a6c7a9314 fix(bot-mode): resume nested one-shot DMs on the recipient's session
Quiet chat -Q inherited the dispatcher's HERMES_SESSION_KEY, so a nested
message_agent notify was addressed to the grandparent and B never woke.
Bind this session's key, strip inherited session identity from the delivery
child env, and continue owned notify completions in-process before stdout.
2026-09-14 21:18:47 +05:30
teknium1
939a2f64b4 fix: long-lived processes stop minting duplicate state.db writer handles
Gateway, dashboard, ACP server and the CLI already hold one registry-shared
SessionDB per state.db path, yet several in-process call paths still opened
a bare SessionDB() beside it. Each one is a full writer: schema init, write
lock, token-writer thread and a close-time WAL checkpoint. On a dashboard
serving overlapping requests that stacked up to the "5 live SessionDB
handles" precursor within seconds; per #110544 they are now harmless to each
other's WAL generation, but the leak itself remained.

Pure readers attach read_only=True (no writer connection, no write lock):
  - plugins/hermes-achievements/dashboard/plugin_api.py::scan_sessions
    (dashboard, per background scan and per /rescan; highest-frequency site)
  - hermes_cli/console_engine.py::_session_db (dashboard console; list,
    stats and export are reads; rename/optimize opt in to a writer)
  - tools/process_registry_results.py::_owns_result (gateway, per retained
    result load)
  - hermes_cli/main.py::_session_db (last-session / title / cwd lookups)
  - hermes_cli/terminal_breadcrumbs.py, hermes_cli/status.py,
    hermes_cli/main_tui_launch.py (lookup one-shots)

Writers share the process's registry handle (hermes_state_registry.acquire;
close()/release_or_close release one refcount):
  - acp_adapter/session.py::SessionManager._get_db — the AIAgent it builds
    acquires the same path, so the ACP server held two writers per process
  - hermes_cli/kanban_db_dispatch.py::_retag_legacy_worker_sessions
    (gateway dispatcher tick)
  - hermes_cli/main.py::_create_titled_session, hermes_cli/oneshot.py,
    hermes_cli/foreign_sessions.py — the CLI acquires the same handle a
    moment later

The "N live SessionDB handles" warning now counts only writable members:
read-only attaches are the sanctioned per-request shape for dashboard
routers and CLI lookups, and counting them turned a healthy topology into
an operator alarm (#100896 field reports of restart loops keyed on it).

Live repro (one registry writer + 4 overlapping dashboard/gateway paths in
one process): before 5 writable opens, 5 live handles, warning fired;
after 1 writable open (the registry handle), 0 from the request paths,
no warning.

Refs #100896 #103339
2026-09-14 08:10:36 -07:00
teknium1
d28938d3da fix(computer-use): screenshot dedup forgets its last frame at a compaction boundary
The "screen unchanged" result points the model at its previous capture. After
context compression that capture may be summarized away, so the note would refer
to pixels no longer in context. Mirror read_file's reset_file_dedup: the
compaction boundary (both the summary path and the codex app-server path) now
clears the session's screenshot digest, and the first capture afterwards delivers
the image again even when the screen is byte-identical.
2026-09-14 07:21:18 -07:00
teknium1
31964ff4c6 fix(computer-use): dedup keyed by the scoped session, cleared on release
Key the screenshot-dedup state by the same profile-scoped session id the
backend cache uses, so two multiplexed profiles sharing a session id (or a
DISPLAY) never dedup against each other's frames, and forget the state in
release_computer_use_session so a re-created session's first capture
always delivers pixels. Reword the unchanged note to cover the aux-vision
path (where the prior result was an analysis, not an image). Tests: the
dispatch path (explicit capture + capture_after) honours the streak cap;
release forgets state.
2026-09-14 07:21:18 -07:00
Teknium
682b973b32 feat(computer-use): stop resending unchanged screenshots
Port from openclaw/openclaw#129924: a capture whose pixels are
byte-identical to the previous capture of the same target in the same
session returns its full text metadata (element index included) plus an
explicit 'screen unchanged' note instead of the multimodal image block.

Adapted for Hermes: openclaw gates dedup on per-frame context-epoch
tracking; Hermes bounds staleness with a consecutive-omission streak cap
(2) so full pixels are re-delivered before compaction could evict the
referenced image. Dedup state is per-session (no cross-session leaks),
append-only (no history rewrites — prompt cache prefixes untouched),
and skipped entirely when no session_id is present.
2026-09-14 07:21:18 -07:00
teknium1
ebe8cda8ea feat(tui_gateway): real JSON-RPC server→client requests replace the *.request / *.respond notification pair
The backend never sent a JSON-RPC request; when it needed an answer from the
renderer it hand-correlated a `*.request` notification with a later `*.respond`
method through four module-level dicts, a timeout thread and 13 derived
`*.expire` names, plus a separate reconnect snapshot per prompt kind. That is a
second request/response layer built on a protocol that already has one.

`tui_gateway/server_requests.py` sends `{id: "srq-…", method, params}` and
blocks on the response frame with that id (string ids never collide with the
clients' integer ids). One `request.cancel {id, method, reason}` notification
withdraws a request on timeout / interrupt / session close. `open_requests` on
`session.resume` / `session.activate` / `session.events.since` re-delivers
unanswered requests after a reconnect; the shared TypeScript channel does that
itself before the caller sees the result. Batch clarify keeps its per-question
locks as a normal `clarify.lock` RPC (the last lock resolves the request).
Approvals stay queue-backed (`tools.approval` owns the timeout, `/approve all`,
coalescing): the request resolves the queue entry and the entry's own
resolution withdraws the request through `register_gateway_settle`.

Deleted: `_block`, `_respond`, `_pending`, `_answers`,
`_pending_prompt_payloads`, `_batch_clarify`, `_EXPIRING_REQUESTS`, the
`*.respond` methods, every `*.request` / `*.expire` event, `pending_clarify`.
Compute-host (turn isolation) mirrors the child's open request and relays the
response frame / lock to it. Desktop, TUI and shared clients register
`onRequest` handlers where they used to switch on `*.request` events; answers
are response frames over the socket the request arrived on, so #91684's
owner-routing class cannot recur for prompts.
2026-09-14 06:02:05 -07:00
teknium1
60559d4e0e fix(delegation): late-attached children take the parent's soft/hard stop kind; dedupe the fallback replay
_attach_child now mirrors a pending parent stop with the same split
interrupt() uses for its own fan-out (hard -> hard_interrupt, soft ->
interrupt), so a redirect is not turned into a cancel on a child that was
attached late. _restore_parent_cancellation collapses to re-attaching the
rejected unit's children: the replay is the attach step's job now.

Test fixture: _Batch gained origin_session_history_delivery on main after
the salvaged PR was written.

Co-authored-by: illidan <noequal666@gmail.com>
2026-09-13 21:31:57 -07:00