Commit Graph

362 Commits

Author SHA1 Message Date
teknium1
e234408607 fix(tools): tear down a child in the caller's own process group by PID, never killpg it
A spawner that skipped setsid (the Darwin gateway's posix_spawn shim) leaves
tool children in the gateway's process group, and _kill_process_group_posix
then killpg()s the gateway itself — launchd logs 'Killed: 9' and KeepAlive
respawns it. Branch on pgid == os.getpgrp() as a plain if/else (not a raised
PermissionError routed into the EPERM handler) and signal the wrapper plus
its snapshotted descendants by PID; the EPERM fallback shares that helper.

Fixes #107029
2026-09-20 16:32:19 -07:00
teknium1
d4b772cbda fix(tools): search_files keeps its matches when killpg is refused
On macOS a search that reaches its `limit` takes the early-stop branch that
TERMs rg's process group; rg has often already exited, and macOS answers
killpg on a zombie-only group with EPERM instead of ESRCH. The
PermissionError escaped _kill_process_group_posix, unwound to search_tool
and the drained matches were replaced by
`{"error": "[Errno 1] Operation not permitted"}` (#116855, same symptom
as #104696). The helper now treats EPERM like "nothing left to signal"
and falls back to killing the known PIDs so a live child cannot escape.
It also never killpg's the caller's own group (#107029): a child that
shares our pgid is torn down by PID.

Diagnosis credit: #116949 (@liuhao1024), redone slim.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 16:32:19 -07:00
chugiacan
0abc3b044b fix(agent,cli): skip WSL/system bash.exe in Windows bash discovery (#116818)
_windows_bash_candidates in tools/environments/local.py appended
shutil.which("bash") to the candidate list without filtering — when Git Bash
is installed but PATH resolves `bash` to WSL's stub (C:\Windows\System32\bash.exe)
or WindowsApps bash.exe, that stub enters the candidate list.

Fix: skip paths whose normalized form contains `system32` or `windowsapps`,
logging at debug level. Git Bash (already searched explicitly in the roots list)
remains the first usable candidate.

Fixes #116818
2026-09-20 15:50:26 -07:00
teknium1
940c610994 fix(security): keep _HERMES_PROVIDER_ENV_BLOCKLIST importable from tools.environments.local
Switching local.py to the folded _is_provider_env_blocklisted helper dropped the
blocklist name from its import list, which removed the re-export that
tests/providers/test_auth_registry_import_order.py (and any out-of-tree caller)
reaches through tools.environments.local. Re-import it alongside the helper so
the module keeps exposing the set; behaviour is unchanged.
2026-09-20 12:53:47 -07:00
beardthelion
b534f4b8c8 fix(security): match credential env names case-insensitively
The provider-credential blocklist and every adjacent env-name check ran
case-sensitive membership tests while the Windows environment block
resolves names case-insensitively. A skill or terminal.env_passthrough
entry registering openai_api_key was accepted, then resolved to the real
OPENAI_API_KEY by os.getenv() and forwarded into SSH/Docker exec envs,
the same GHSA-rhgp-j443-p4rf tunnel the blocklist closes. The same gap
let variant-cased credentials survive the source-side strip
(_filter_secret_env, _scrub_credentials), evade the docker_forward_env
and docker_extra_args egress collision guards, and ride
strip_launch_profile_env residue into a routed sibling profile's child.

Add _is_provider_env_blocklisted (exact + folded membership) and apply
it at every layer: registration refusal, both scrub paths (Tier-1 and
plugin strip sets fold too; inherit_credentials still inherits), the
remote exec-env builder, and the docker egress collision checks. Fold
_matches_terminal_first_party_prefix symmetrically so a lowercase-stored
buzz_private_key keeps its terminal carve-out, and fold the launch-
residue pop (selection folds too so a lowercase path in .env stays
global). _HERMES_FORCE_* opt-in transport and docker_env literal
container names stay case-sensitive on purpose.
2026-09-20 12:53:47 -07:00
beardthelion
e885d9cdc2 fix(docker): drop -f from orphan reaper rm so running containers fail safe
reap_orphan_containers snapshots exited containers, then removes each
candidate with docker rm -f. The exited-only filter is only as fresh as
the docker ps snapshot: a sibling process can legitimately restart an
exited container between the snapshot and the rm (the reuse path calls
docker start on matching exited containers), and FinishedAt still
reports the previous exit, so the age check passes. The -f then kills
the sibling's live container and its in-container state.

A plain docker rm fails atomically in the daemon on a running container,
which is exactly the skip semantic the exited-only guard intends. Every
intended target of this sweep is already exited, so -f bought nothing.

The new test drives the real reaper with a simulated restart: rm argv
must not contain -f, and a daemon refusal on a running container is
skipped while the rest of the sweep continues.
2026-09-20 00:05:30 -07:00
teknium1
16ae9cfccb refactor: cache probe verdicts on exit-0 / capability rejection only, trim tests
Drop the nondeterminate-stderr fragment list: the rule is now simply "exit 0
caches True; a nonzero exit whose stderr names the probed capability
(cgroup / storage) caches False; everything else is uncached and retried on
the next spawn". A curated list of pull/daemon phrases would need to track
every docker CLI wording; the capability keyword is the property the cache
exists to record. Tests are trimmed to the invariants: the production
DockerEnvironment path recovers on the next spawn, a definitive cgroup
rejection is still cached, and the storage-opt probe caches only definitive
answers.
2026-09-20 00:04:54 -07:00
beardthelion
504e79d0d5 fix(docker): stop caching transient capability-probe failures as unsupported
_cgroup_limits_available latched _cgroup_limits_ok=False on ANY probe
failure — a TimeoutExpired while docker run auto-pulls an uncached image,
a daemon cold-start, a manifest/pull error, or a missing docker binary —
permanently stripping --cpus/--memory/--pids-limit from every later
container in the process. The sibling _storage_opt_supported latched the
same way, and additionally read a failed `docker info` (returncode never
checked) as "not overlay2", disabling disk quota for the process.

Both probes now cache only definitive answers: exit 0 -> True; a daemon
rejection naming the probed capability -> False. Any other failure
(missing docker, pull/manifest/daemon errors, timeouts) degrades that
spawn and is retried on the next. A shared nondeterminate-stderr list is
checked before the capability keyword so an image named e.g.
"cgroup-tools" whose pull fails is not misread as a cgroup rejection.
2026-09-20 00:04:54 -07:00
teknium1
9182ad0860 test: trim egress shorthand tests to two invariants, ignore positionals in the flag scan
Fold the salvaged parametrized positive/negative tables and the redundant
env-file/real-chain e2e tests into one parametrized invariant (every pflag
spelling of -e collides, legitimate args do not) plus one production-entry
test through DockerEnvironment. Args that do not start with "-" are
positionals under pflag and are never env flags, so the shorthand scanner
skips them instead of reading a container name like "de" as "-d -e".
2026-09-20 00:04:19 -07:00
beardthelion
10c30c8264 fix(docker): scan docker_extra_args env flags under pflag shorthand semantics
The egress collision scan only recognized -e/--env/--env-file as standalone
tokens, so the combined-shorthand spellings docker accepts, -eNAME=v joined
and -iteNAME=v chained after boolean shorthands, injected critical env names
past the guard. Parse each arg the way pflag does (boolean ditPq shorthands
may precede the terminal value shorthand) and stop at the -- terminator.

Fixes #115887
2026-09-20 00:04:19 -07:00
beardthelion
578c7aeae4 fix(docker): guard every egress-written env name in collision checks
_critical_egress_env_names seeded the forward_env/extra_args collision
set by scanning env_overrides for *_API_KEY/*_TOKEN suffixes. Mapped
tokens land under arbitrary real_env_name and alias_env_names entries
from mappings.json, so a non-suffixed credential name (a custom mapping,
an alias, AWS_SECRET_ACCESS_KEY) was injectable via docker_forward_env
or -e in docker_extra_args without tripping the enforced refusal,
placing the real credential inside the sandbox.

Every name the egress layer writes is egress-owned, so the critical set
is now env_overrides itself plus the proxy-control vars and NODE_OPTIONS:
complete by construction, no load_mappings dependency, and correct as
providers and aliases are added. This matches what
check_docker_env_collisions already protects via load_mappings.

Tests cover both surfaces: docker_extra_args -e and docker_forward_env
entries naming a non-suffixed mapped credential now refuse under
enforcement.
2026-09-20 00:04:19 -07:00
teknium1
4586cde64d chore: mark the deliberate /tmp literals and shrink the lint baseline to one code block
The seventeen remaining literals are container-side paths, AF_UNIX socket-path-limit
candidates on darwin, detection needles, guard regexes and guidance text that tells the
model to avoid /tmp. Each carries an inline `no-tmp: ok — <why>` so the reason lives
next to the line; the baseline keeps only a fenced tree listing where a marker would render.
2026-09-19 10:44:26 -07:00
teknium1
d51f9c8464 refactor: real-home and terminal temp-dir fallbacks stop naming /tmp
get_real_home() fell back to a literal /tmp when no OS home could be found, and
LocalEnvironment.get_temp_dir() probed /tmp by hand before consulting
tempfile.gettempdir(), whose own candidate walk already covers the system temp
dir (and honours the scratch TMPDIR Hermes now exports). Both defer to
tempfile.gettempdir(); a relative gettempdir() result is made absolute instead of
being swapped for /tmp.
2026-09-19 10:44:26 -07:00
teknium1
2dcebe6471 feat: Hermes-owned scratch dir replaces the system temp dir for every process and child
Hermes and everything it launches (browser profiles, PTY probes, skill scripts,
tempfile defaults in delegated code) wrote to the system temp dir, which is a
RAM-backed tmpfs on most Linux hosts and containers and fills under agent load.

- hermes_constants.get_scratch_dir(): HERMES_HOME/cache/scratch (0700), entries
  older than 72h pruned once per process / once per hour across processes.
- apply_scratch_tmp_env(env) / export_scratch_tmp_env(): TMPDIR/TMP/TEMP point at
  the scratch dir when the user or OS has not set them; a value Hermes itself
  exported (== HERMES_SCRATCH_DIR) is re-derived for a re-homed process or a
  child served under another profile, so profiles never share scratch.
- hermes_bootstrap runs the export on import (every entry point); hermes_cli.main
  re-runs it after --profile resolution; the subprocess HOME contract
  (apply_subprocess_home_env) and the routed-home rewrites in code_execution_env
  and served_profile_child_env apply it to child envs.
- The runtime-environment prompt block names the scratch dir so the model stops
  reaching for the system temp dir by reflex; hermes doctor reports the dir, its
  size and whether a user TMPDIR overrides it.
2026-09-19 10:44:26 -07:00
teknium1
1bc953d342 docs(terminal): name the stdin=DEVNULL atom of the Git Bash probe (#73403)
Review follow-up: the comment in _bash_starts now says why stdin=DEVNULL matters
beyond the bounded cleanup — cygwin init's handle_to_fn/NtQueryObject stalls on a
pipe file object that has a read pending (the ACP host's stdin), which is why the
probe blew its timeout inside ACP hosts — and why stderr stays captured (the
Mandatory-ASLR remediation keys off bash's dofork:/child_copy: text).
2026-09-19 02:36:56 -07:00
teknium1
edae76fecd fix(terminal): bound the Git Bash startup probe's timeout cleanup (#73403)
_bash_starts used subprocess.run(timeout=15). On Windows run()'s post-timeout
cleanup is an unbounded communicate(): the probe's MSYS children (true/cat)
outlive the killed bash holding the captured pipe write ends, so the reader
join never returned and the ACP host wedged for minutes at the first terminal
tool (faulthandler dumps in the thread sit in _bash_starts → _communicate).

Route the probe through hermes_cli._subprocess_compat.bounded_probe_run — own
Popen (stdin=DEVNULL, hidden window), communicate(timeout), process-tree kill
and a 1 s bounded drain — so a probe that cannot finish fails fast and
_find_bash falls through to its last-resort candidate. A timeout is recorded as
the probe detail so the ASLR diagnostics still see it.

Supersedes #69083 (@fangliquanflq), whose bounded_captured_run design landed on
main as bounded_probe_run for the git probes (#68997); this is the surviving
_bash_starts call-site conversion from that PR.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-19 02:36:56 -07:00
teknium1
f0032daf94 fix(terminal): restore the unconditional stdout.close() after the Windows drain
Reverts c26e9ccb712. The close guard was written against the old blocking
os.read drain, where close() serialized on the CRT per-fd lock held by the
reader (#67362). _drain_fd_windows only calls os.read after PeekNamedPipe
reports bytes (and caps the read at that count), so it never blocks there and
close() cannot contend after join(2). The guard's only reachable branch — a
grandchild still streaming past the join — leaked the fd, the daemon thread
and collector appends for the grandchild's lifetime on Windows while POSIX
closed; closing unconditionally makes the drain thread exit on its next read
like the POSIX path. #67362 is closed by the PeekNamedPipe drain itself.

Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
2026-09-19 02:35:45 -07:00
teknium1
428777a669 fix(terminal): skip stdout.close() while a Windows drain thread is still reading (#67362)
On Windows close() serializes on the CRT per-fd lock held by a drain thread
blocked in os.read, so the natural-exit close blocked until the grandchild
exited — after the poll loop, where no timeout could see it. With the
PeekNamedPipe drain the thread normally exits before the join deadline; when it
has not (loaded host), leave the fd to the daemon thread / Popen.__del__.

Ported from #67373 (@PRATHAMESH75) and #67448 (@webtecnica).

Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
2026-09-19 02:35:45 -07:00
joaomarcos
e7a82eaded fix(terminal): bound the Windows stdout drain with PeekNamedPipe (#105865)
select() cannot poll pipe fds on Windows, so _drain_stdout used a bare
blocking os.read loop that only returned at true EOF. Any grandchild that
inherited the pipe's write handle (backgrounded `cmd &`, MSYS helpers)
kept the terminal/file tools hung for its whole lifetime — up to the 420s
tool timeout during init_session on real Windows hosts.

_drain_fd_windows polls PeekNamedPipe and mirrors _drain_fd_select's
stop-event + ~300ms idle-after-exit bound, so the drain returns shortly
after bash exits regardless of who still holds the write end.

Salvaged from #105981 by @JoaoMarcos44 (tests trimmed to two windows_only
invariants that run on the windows-latest job).
2026-09-19 02:35:45 -07:00
teknium1
3ed40556ce fix(profiles): a child spawned for another profile no longer inherits the spawner's authorization gates
A `hermes -p B` child built from a process that loaded profile A's env (a gateway, the
dashboard, the post-update fleet restart) started with A's `DISCORD_ALLOWED_CHANNELS`,
`TELEGRAM_GROUP_ALLOWED_CHATS`, `GATEWAY_ALLOW_ALL_USERS`... and enforced them as its
own: gates are not credentials (no secret scrub sees them), a unit-file `Environment=`
or operator export is in no dotenv (no name-list strip sees them), and B's own `.env`
rarely defines the key (its dotenv load never overwrites the inherited value). Observed
as profile B's gateway rejecting every message in B's own channel after a per-profile
restart issued from A (#113270).

- `local_env_policy.is_profile_gate_env` / `strip_profile_gate_env`: gates matched by
  shape (`_ALLOWED_`, `_ALLOW_ALL_`, `_ALLOW_FROM`, `_ALLOW_BOTS`, `_IGNORED_CHANNELS`,
  ...), never `HERMES_*`, so a gate added to any adapter is covered without a second edit.
- `strip_launch_profile_env` drops them on its existing routed-home branch: the seam
  `served_profile_child_env`, the kanban dispatcher, cron workers and the dashboard
  action env already funnel through (the dashboard site now calls it for its target).
  Same-home children keep an operator export.
- `update_restart_recovery._child_environment(profile)`: the one site that bypassed
  every helper (bare `os.environ.copy()` relaunching EVERY profile) strips gates when
  the profile is not the one the updater runs as; the module stays stdlib-only at
  import time.

Live: fresh-process `hermes_cli.update_restart_recovery --stdin` with three gates in
the updater env — base hands all three to profile b's relaunch, fixed hands none and
keeps them for the launch profile.

Refs #113270; supersedes #113308 (@yashraj4, static key list + always-strip; this keeps
same-profile children intact and covers the per-adapter gate set).
2026-09-18 15:11:47 -07:00
teknium1
547fff7500 fix(tools): scope-only passthrough overlay raises instead of silently dropping the declared secret
_scrubbed_env wrapped the scoped_passthrough_additions overlay in try/except Exception with a
debug log; a scope or config failure there would silently drop the declared secret again — the
exact failure mode #114209 complains about. The import cannot fail in-tree and _scrub_child_env
already calls it unguarded, so the guard is removed and one test pins that both local surfaces
propagate the error.
2026-09-18 10:04:57 -07:00
teknium1
e1c0896518 fix(cron): no_agent script env comes from the factory's own snapshot, not a raw copy at the spawn site
tests/agent/test_subprocess_env_guard.py flagged cron/scheduler_script.py:362 as a new raw
os.environ.copy() spawn-env site. build_subprocess_env gains strip_launch_profile=True so the
launch profile's .env residue is still dropped from the base before the secret scrub (same
order and semantics as before: strip first, then scope-overlay the owning profile's declared
names), and the spawn site no longer snapshots the environ itself.
2026-09-18 10:04:57 -07:00
teknium1
802a9975d2 fix(cron): no_agent scripts get the owning profile's declared secret, never the launch profile's
A no_agent cron script owned by a profile served by a multi-profile gateway (or the
Desktop/dashboard backend) could not receive its own credential: the served profile's
.env never enters the process env (load_hermes_dotenv skips the process-global load
for a routed home), and the child-env sanitizer only resolved a terminal.env_passthrough
name through the profile secret scope when the launch environment already carried that
name. So the documented declaration mechanism (SECURITY.md 2.3) could not supply a
value that exists only in the served profile's scope, while the launch profile's own
.env credentials rode into the served profile's script child unstripped.

- tools/env_passthrough.scoped_passthrough_additions: declared names the bound scope
  holds but the env being filtered lacks; reads the scope alone (no os.environ, no
  other profile), empty without a scope so single-profile spawns are byte-identical.
- _scrubbed_env (terminal, cron scripts, bg processes, search workers) and
  execute_code's _scrub_child_env overlay those names after filtering.
- cron _run_job_script and terminal _make_run_env strip the launch profile's .env
  residue from the base (strip_launch_profile_env, a no-op for the launch profile's
  own jobs) before the filter, so the scope overlay lands after the strip.

Why not the 1Password-only shape of #114218: the gap is the declaration mechanism, not
one vendor's token, and an unconditional token pop broke the single-profile .env flow.
Docs: cron no_agent credential section, secrets child-process section, security
passthrough note.

Fixes #114209
Supersedes #114218
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-18 10:04:57 -07:00
kokhlo
b911b9794e fix(tools): sanitize cwd overrides written into live container envs
A cwd override registered by a desktop/TUI surface is a raw host path. On
container backends it was written verbatim into the cached live env's cwd,
so every file-tools command wrapper did `builtin cd -- <host path> || exit 126`
and all writes outside the mapped workspace failed with an unrelated `cd:`
error while terminal commands kept working (their per-command resolver
already sanitizes).

The live-env write now mirrors the creation-path guards: a host/relative cwd
is remapped to /workspace when it is the directory mounted there (docker cwd
passthrough), otherwise not applied to the live env at all. The session cwd
record keeps the raw path and non-container backends apply overrides
verbatim, so ACP project-root switching is unchanged.

Fixes #113894

(cherry picked from commit d712c6b47d28f213773f40ebac81dd9a45c6e94f)
2026-09-18 09:54:45 -07:00
teknium1
775aedabce fix(file-sync): sync-back cap is a config.yaml key, not an env var
Non-secret settings live in config.yaml: the extraction cap override becomes
terminal.sync_back_max_bytes in DEFAULT_CONFIG (read via load_config() like the
rest of the terminal section) and the HERMES_SYNC_BACK_MAX_BYTES env var goes
away; docs and the cap tests follow. psutil is a pinned core dependency, so
_temp_entry_owner_alive imports it unconditionally instead of guarding an
ImportError that cannot happen.
2026-09-18 09:30:16 -07:00
teknium1
d007afe130 fix(file-sync): bound sync-back leaks by owner PID, make the cap overridable, cover every tar backend
Follow-up to the salvaged #114440 commits:

- Stale sweep: a dead owner's entry goes at once, but a live (or recycled)
  PID no longer exempts an entry from the 30 min age cutoff — a recycled PID
  must not pin a multi-GB leak forever. Liveness via psutil.pid_exists (the
  footgun scanner rejects os.kill(pid, 0): it terminates on Windows), so the
  os.name guard and the two extra helpers collapse into one.
- HERMES_SYNC_BACK_MAX_BYTES overrides the 2 GiB extraction cap for trees that
  legitimately exceed it (the reporter's 4.26 GB tree was downloaded and
  discarded on every attempt); the skip warning names the override.
- Vercel sandbox bulk download excludes *.sock like SSH/Modal/Daytona.
- Tests trimmed to two invariants per file; docs mention the cap override,
  the socket skip and the per-PID temp naming.
2026-09-18 09:30:16 -07:00
liuhao1024
e928655b85 fix(tools): reclaim sync-back temps by owner PID and harden socket tolerance
Review follow-ups on #114440:

- Drop the HERMES_SYNC_BACK_MAX_BYTES env override: a malformed value
  crashed the whole backend at import time, and the rubric keeps env
  vars for secrets only. The 2 GiB cap stays hardcoded (raising it can
  be a separate change).
- Embed the owning PID in sync-back temp names and reclaim dead-owner
  entries immediately, mirroring daytona's PID-suffixed remote temp.
  Age no longer decides liveness (a staging dir's mtime does not move
  while content streams into subdirectories, and the download bound is
  SSH/Modal-only), and concurrent gateway processes under different
  HERMES_HOMEs never sweep each other's transfers. Legacy (pre-PID)
  names and Windows hosts keep the 30-minute cutoff.
- Anchor the rc=2 socket tolerance to lines ending in ': socket
  ignored' and require a non-empty diagnostic, so whitespace-only
  stderr or a filename merely containing the marker still fails.
- Add --exclude='*.sock' to the Modal and Daytona bulk-download tars
  (same failure class as the SSH backend).
2026-09-18 09:30:16 -07:00
liuhao1024
2f62b1f4de fix(tools): stop socket-ignoring tars from failing sync-back and leaking GBs
Every SSH sync-back on a host with a live gateway.sock failed: GNU tar
cannot archive a Unix socket, prints "socket ignored", and exits 2 —
so the download raised, retried three times, and (when the process was
hard-killed mid-download, e.g. by the gateway's tool-subprocess kill)
left its multi-GB temp tar behind. With the stale-temp sweep window at
6 h, a crash loop accumulated 45 GB in /tmp and filled the rootfs
(#114437).

- Exclude *.sock from the remote tar up front; tolerate a rc=2 whose
  stderr is only "socket ignored" lines (defensive for tars without
  --exclude support or unmatched socket names) — anything else still
  fails the transfer.
- Shrink the stale sweep cutoff from 6 h to 30 min; the download is
  bounded by a 120 s subprocess timeout, so a live transfer is minutes
  old at most and the cutoff only has to sit above that.
- Make the 2 GiB extraction cap overridable via HERMES_SYNC_BACK_MAX_BYTES
  for synced trees that legitimately exceed it.

Fixes #114437
2026-09-18 09:30:16 -07:00
kshitijk4poor
e4739c6d2d fix(cli): doctor and terminal setup resolve the container runtime too
`hermes doctor` and `hermes setup terminal` only ever looked for a `docker` binary
(`_safe_which("docker")` / `shutil.which("docker")`) and probed `docker version` by
literal name, so a podman-only machine was told Docker was missing while the docker
terminal backend was already running containers through Podman — the same mis-report
the dashboard probe had.

Both now resolve the CLI through `find_docker()` (HERMES_DOCKER_BINARY → docker →
podman → macOS Docker Desktop paths) and name the runtime they actually found, via
`docker_runtime_name()` and `docker_runtime_start_hint()` next to `find_docker()` —
shared with the dashboard probe instead of a per-module copy. The probe's version call
goes through the backend's `run_capture()`, and a Podman row no longer advises starting
a daemon: Podman is daemonless, so it points at `podman machine start`.

Tests that pinned the old resolution (`_safe_which` / the global `shutil.which`) now pin
`find_docker()`, so they no longer depend on whether the host has Docker Desktop at a
known macOS path.
2026-09-18 21:47:13 +05:30
kshitijk4poor
27442b5799 test(tools): prove the snapshot dump drops every bridged var and the delegation marker
Replace the string-presence snippet test and the file-content e2e with two
behaviour contracts: the real bash dump must drop every gateway-bridged name
(gateway.session_context._VAR_MAP) plus HERMES_DELEGATED_CHILD_CONTEXT while
keeping ordinary exports — this is the test that would have caught the
HERMES_CRON_SESSION drift (in the regex contract, missing from the unset list)
that was live on main — and a real LocalEnvironment run where a delegated
child still sees its marker and the next ordinary command does not. Add the
marker to the regex contract and cut the WHAT comment down to the WHY.

Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
2026-09-17 23:47:17 +05:30
liuhao1024
c44b423e52 fix(tools): stop the terminal snapshot from persisting delegation markers
_export_dump_excluding_session_vars unsets the session bridged vars,
the attribution markers, and HERMES_UI_SESSION_ID before export -p, but
not HERMES_DELEGATED_CHILD_CONTEXT or HERMES_CRON_SESSION. Both are
scope-limited: scrub_kanban_env() injects the delegated-child marker
into the SUBPROCESS env while a delegate_task child is live, and the
cron session bridge tracks the owning cron scope. A snapshot captured
in that window persists them — the marker outlives the child whose
ContextVar was correctly reset on exit, and every later `source` of the
snapshot re-asserts it. In the parent session this locks out all kanban
mutations ("delegate_task child contexts cannot mutate Kanban tasks or
boards") even though no child is running (#90782).

Add both names to the unset list in the dump snippet. Regression tests
assert the snippet unsets them and that an end-to-end command run with
the marker set leaves no trace of it in the snapshot file.
2026-09-17 23:47:17 +05:30
Xuxyyy
84893f8342 fix(tools): honor false interrupt debug values 2026-09-16 17:21:14 -07:00
teknium1
3fe8e5e443 fix(multiplex): routed children never inherit launch-only credentials, with or without the multiplex flag
Two authority gaps in served_profile_child_env (#111617 review, andrexibiza P1 #1/#2,
kvnloo finding 1):

- The base was hermes_subprocess_env(inherit_credentials=True) = the launch environ's
  provider credentials; strip_launch_profile_env only knows names with .env/source
  provenance, so a key systemd/Compose/the shell injected into the launch process
  survived into profile B's child whenever B did not define the same name. Now a ROUTED
  target scrubs every Tier-1/Tier-2 credential from the base regardless of provenance
  before B's own scope is overlaid (the child boundary gets get_secret's contract: a
  scoped miss is no credential, never ambient fallback). The launch profile's own child
  keeps its env. bot_relay's base=os.environ goes through the same scrub.
- strip_launch_profile_env / the scrub keyed on is_multiplex_active(); the Desktop and
  dashboard backends serve ?profile=B by installing the HERMES_HOME override without
  that flag, so B's slash worker / helper children kept A's .env and settings. The
  authority test is now "is the target a routed home" (target != process home).
- _build_browser_env resolved the passthrough keys via get_secret, which falls through
  to os.environ on a scoped miss while multiplexing is inactive: a routed B with no
  Firecrawl key got A's. Under serves_routed_profile() the bound scope is the only source.
- served_profile_child_env(inherit_credentials=True) with no target and no scope bound
  under multiplex minted with the launch credentials (key_cmd TTL refresh on a worker
  thread); it now raises UnscopedSecretError like get_secret.

tests/tui_gateway/test_served_profile_child_env_authority.py: ambient-only A key + B
missing it (mux on), flag-off routed B (helper child + browser), real child observation.
3/3 red on base.
2026-09-16 00:35:00 -07:00
teknium1
be67e1d31e fix(tools): tolerate an untraversable HOME when probing ~/.local/bin
CI runs the suite as an unprivileged user with HOME=/root in one fixture;
Path.is_dir() raised PermissionError from _user_local_bin_entries and the
run-env builder crashed. An unreadable home has no usable ~/.local/bin, so
treat the OSError as absent.
2026-09-15 18:49:29 -07:00
teknium1
43e7e830fd fix(tools): fold ~/.local/bin into the POSIX PATH completion siblings, tests + docs
Slim follow-up to the salvaged #111790: the helper becomes a list-returning
sibling of _managed_runtime_path_entries (same shape, same "only when it
exists" convention) and loses the Windows check the caller already performs.

Why here and not in the Electron remote spawn: propagating the login-shell PATH
that locateHermes discovered into `exec env HERMES_DESKTOP=1 … hermes serve`
would fix only the Desktop SSH surface; the terminal environment's PATH
completion is the seam every thin-PATH launcher (SSH, systemd, launchd, cron)
already goes through, so the class closes once. Windows twin out of scope.

Tests move to the mirror dir tests/tools/environments/ with an absent-dir
control; FAQ documents the terminal PATH composition.

Fixes #111778
2026-09-15 18:49:29 -07:00
KoNit-K
c68e306ea4 fix(tools): include user local bin in POSIX PATH 2026-09-15 18:49:29 -07:00
teknium1
bae9f8ab85 fix(file-sync): sweep stale sync-back dirs too and tighten the stale window to 6 h
The salvaged commit only swept `hermes-sync-back-*.tar`. The extraction
staging dir (`tempfile.TemporaryDirectory(prefix="hermes-sync-back-")`) is
leaked by the same hard kill, so the sweep now reclaims both shapes and
returns the count. The window drops from 24 h to 6 h (the reporter's
value): archives appear every few minutes on a busy gateway and a live
transfer is never hours old. Tests trimmed to two invariants: the sweep
touches only stale prefixed entries, and a real sync_back reclaims a
leaked archive while writing its own tar under the identifiable prefix.
2026-09-15 05:35:34 -07:00
KoNit-K
6a03d5a94c fix(file-sync): clean stale sync-back archives 2026-09-15 05:35:34 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1
d1794d5539 fix(multiplex): children spawned for a served profile start from that profile's env
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:

- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
  base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
  BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.

tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.

Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
2026-09-15 03:47:15 -07:00
kshitijk4poor
dcdbcb8a2b fix(env-loader): split source_supplied_names() out of secret_source_names()
Widening secret_source_names() to include skipped_existing names silently changed
tools/mcp_tool_config.py::_build_safe_env, an untouched consumer that forwards every
returned name into MCP stdio child envs. That consumer wants only names a source
actually APPLIED (pre-stack semantics), so secret_source_names() goes back to
tuple(_SECRET_SOURCES).

The routed-child scrub in strip_launch_profile_env is the one site that must also see
names a source supplied but lost to a pre-existing process value, so it reads the new
source_supplied_names() accessor instead. tools/mcp_tool_config.py is byte-identical to
origin/main.
2026-09-15 11:03:39 +05:30
kshitijk4poor
f367ebeb5d refactor(cron): strip external-source residue in strip_launch_profile_env itself
_run_job_script popped the launch profile's secret-source names in an
inline loop right above strip_launch_profile_env, so only the no_agent
child got that protection; the four other callers of the same helper
(the external cron worker, scheduler_delivery, kanban dispatch and the
byterover plugin) still handed a served profile the launch vault or
1Password names. Fold the names into the helper's residue set, which is
already gated on multiplex and on the target not being the launch
profile, and keeps administrator-managed keys.
2026-09-15 11:03:39 +05:30
John Paul Soliva
62b4488cb5 fix(cron): routed fires are multiplexed at the worker handoff; managed keys keep policy precedence
Review findings on f5f88d5058. Three are defects the previous round introduced.

Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.

Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.

Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.

Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.

Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.

(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
2026-09-15 11:03:39 +05:30
John Paul Soliva
9d7de6c140 fix(cron): close three launch-residue leaks into a routed no_agent child
Review findings on d8c467f223, each reproduced through its production path.

Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.

Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.

Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).

Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.

(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
2026-09-15 11:03:39 +05:30
teknium1
1512edfb85 refactor(tools): terminal, execute_code, MCP and the bounded collector truncate through one head/tail helper
Four copies of the 40/60 head/tail algorithm with a near-identical notice
(terminal_tool_result, mcp_tool_content, code_execution_tool,
environments/base_output) collapse into tools/tool_output_truncate.py, so the
ratio and the `... [<LABEL> TRUNCATED - N <unit> omitted out of T total] ...`
marker are defined once. execute_code keeps byte mode + spill path and only
shares the notice/split. Visible change: the terminal notice now uses
thousands separators like the other three (`9,000 chars` not `9000 chars`).

kanban_specify._truncate: comment claimed escape stripping the body never did;
comment now says what the plain clamp is for.
2026-09-13 05:09:43 -07:00
Teknium
a49a9d79b3 Port from lobehub/lobehub#19329: surface environment recreation in terminal tool results
When a persistent Docker container is removed out-of-band or a Vercel
sandbox hits a terminal state, the backend silently recreates it and
retries. The model then keeps assuming background processes and
non-persisted files from earlier commands still exist.

Backends now call _mark_recreated() after a successful recovery;
BaseEnvironment.execute() folds the one-shot flag into the result as
environment_recreated, and finalize_foreground_result() attaches a
model-facing warning field explaining what may have been lost.

Ported from lobehub/lobehub#19329 (sandbox recreation surfacing),
adapted to hermes environment backends and tool-result JSON.
2026-09-12 20:55:04 -07:00
Teknium
284d220ba4 fix(multiplex): cron, kanban, /loop and completion paths for a served profile match its standalone gateway
Under gateway.multiplex_profiles a secondary profile X is ticked, dispatched
and notified from the default profile's process, where os.environ holds the
DEFAULT profile's .env and X's values live only in the per-turn secret scope /
HERMES_HOME override. Every remaining read that skipped that scope made X
behave differently from `hermes -p X gateway run`:

- cron: HERMES_CRON_TIMEOUT, HERMES_MODEL (job/preflight fallback),
  HERMES_CRON_MAX_PARALLEL, inflight allowance, prefill file and the script
  timeout were bare os.getenv → the default profile's values; a job without a
  model silently ran on the default's HERMES_MODEL instead of refusing.
  cron/env_settings.py::cron_env_setting reads the scope (fire) or the ticked
  home's .env (tick thread), plain environ when multiplexing is off.
- child env: the restart-safe cron worker, the Bot Chat delivery child and the
  kanban worker inherited the launch profile's non-credential .env settings
  and bridged TERMINAL_* policy (TERMINAL_ENV=docker, default's image,
  HERMES_MODEL) — X's worker ran in the default's docker image on the
  default's model. tools/environments/local.py::strip_launch_profile_env drops
  them when the child targets another served profile.
- kanban: the worker --toolsets pin was silently dropped for every served
  assignee (toolset probes call get_secret without a scope → swallowed
  UnscopedSecretError); notifier pings, artifact uploads and the wake text ran
  under the default's media policy / display language (only wake() was scoped).
- /loop: _post_turn_loop_completion hopped to the executor without contextvars,
  writing the completed tick into the DEFAULT profile's state.db and leaving
  X's row awaiting_response forever; the --until judge ran with the default's
  aux credentials.
- background processes: a secondary's processes.json (scope-relative since
  adf23550f5) was never read at startup; its processes were not re-adopted
  and notify_on_complete notices were lost. Startup recovers every served
  home under its scope; recovery adopts each session once.
- completion delivery: background_process_notifications was evaluated once per
  drain for the ambient profile (default's mode for everyone; X's `off`
  dropped a sibling's `all` event), recovered watchers used the default's
  mode, HERMES_BACKGROUND_NOTIFICATIONS was read raw from environ;
  _deliver_platform_notice used the default's GatewayConfig so a secondary's
  notice_delivery: private went public.

Not changed: gateway/run.py and tools/async_delegation.py (PR #106742
rewrites both). Known residue left for the env-bridge lane:
HERMES_SESSION_STALL_TIMEOUT is bridged once from the launch config.
scheduler bug, not a parity gap; unchanged here.
2026-09-11 19:58:07 -07:00
Teknium
adf23550f5 fix(tools): profile-scoped checkpoint/snapshot paths, tool caches, TZ and schema paths under multiplex
Under `gateway.multiplex_profiles` one gateway process serves every profile
under ~/.hermes/profiles/NAME/; each routed turn runs with a context-local
HERMES_HOME override while `os.environ` still holds the DEFAULT profile's
values. Anything evaluated once at import, or memoised in a single unkeyed
module slot, therefore freezes the LAUNCH profile's value and leaks it into
every other profile's turns. This lands the tools-side half of that class:

- tools/process_registry.py, tools/environments/{modal,singularity}.py:
  `_checkpoint_path()` / `_snapshot_store()` resolve `get_hermes_home()` at
  call time (same seam as `tools/skills_tool._skills_dir`, so the existing
  `monkeypatch.setattr(CHECKPOINT_PATH)` test sites keep working). Completes
  the checkpoint_manager / sticker_cache half cherry-picked from #56315.
- plugins/platforms/feishu/feishu_comment_rules.py: `_MtimeCache` is now
  path-keyed (accepts a Path or a zero-arg resolver, one (mtime, data) slot
  per resolved path) with `invalidate()`; `_rules_file()` / `_pairing_file()`
  resolve the routed profile's files. Proposed in #63962.
- tools/tool_output_limits.py, tools/browser_tool.py, tools/browser_camofox.py:
  the process-lifetime config caches are dicts keyed by `hermes_home_key()`;
  the `_X_resolved` flags and the lifecycle reset keep their shape.
  tools/file_tools.py drops its private `file_read_max_chars` memo and reads
  the already mtime+path-cached `load_config_readonly()`.
- hermes_time.py: `get_timezone_name()`; when `is_multiplex_active()` the
  env `HERMES_TIMEZONE` (bridged from the default profile's config at gateway
  startup) is ignored in favour of the routed profile's config.yaml. Both
  sandbox TZ sites (code_execution_env/_tool) now use it.
- tools/cronjob_tools.py, tools/tts_tool.py, tools/skill_manager_tool.py:
  the static schema text is profile-neutral and `dynamic_schema_overrides=`
  rebuilds the `display_hermes_home()` / create-dir hint per
  `get_definitions()`, so a routed profile's model sees its own paths.

Refs #95685.

Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
(cherry picked from commit 6d3fc6b07b3155c6196b1fd61a829283f1d7855c)
2026-09-11 15:44:00 -07:00
kshitijk4poor
4fc64b6730 refactor(terminal): probe NOPASSWD only when a prompt would fire; drop the host-only probe
With BaseEnvironment supplying a backend-scoped probe to every production
caller, the module-level host-only `_sudo_nopasswd_works` (and its
TERMINAL_ENV gate) had no callers left; the `or` fallback and the outer
try/except around a callback that already fails closed were dead too.

Move the probe under `should_prompt_for_sudo`: headless callers (gateway,
cron, delegated children) reach `(command, None)` whether or not the probe
runs, so the extra backend round trip — an ssh exec on the SSH backend —
was pure waste on every headless sudo command.

test_subagent_sudo_prompt no longer needs to patch the host probe out:
bare `_transform_sudo_command(cmd)` calls have no probe by construction.
2026-09-11 11:19:53 +05:30
fangliquanflq
39abca492d fix(terminal): gate sudo probes by backend cancellation safety
Opt in only Local, Docker, SSH and Singularity: a timed-out `sudo -n true`
probe on those backends kills one process, while SDK adapters (Modal,
Daytona, Vercel) cancel by terminating the whole sandbox.

Net of the original PR's commits f7db10ef + 9965b1b9 (the intermediate
_ThreadedProcessHandle special-case was superseded by this gate).
2026-09-11 11:19:53 +05:30