Commit Graph

407 Commits

Author SHA1 Message Date
ethernet
6716b72ef2 refactor: separate checkpoint store maintenance from snapshots
Move retention, orphan pruning, status and clear operations into a topical sibling and route callers and tests directly to their defining module. Keep the shared store paths and git execution in checkpoint_manager; shorten the local child-env WHAT docstring.
2026-09-24 01:55:51 -04:00
ethernet
5b9b997c42 fix(tools): read snapshot store without check-then-read race 2026-09-24 01:39:08 -04:00
ethernet
038d7f797f fix(pm): avoid duplicate Daytona install and decode Store output explicitly 2026-09-24 01:36:14 -04:00
ethernet
d75d3fe15c Merge origin/main into ethie/pm-clean
Conflict resolutions and semantic fixups:

- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
  fence missed, deregister from _live_foreground in a finally) wrapped around
  pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
  the old early-recovery block stays gone; main's interrupted-pull restore
  (auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
  pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
  and is cleared once git is done. The marker's target is the ref git actually
  moves to (a release tag, not always origin/<branch>), since the restore
  compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
  had dropped and main's auto-merged restore needs (NameError on the first
  launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
  settings.about block; trim it to `updates` as pm-clean's type and the other
  overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
  names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
  key-leak switch leg needs the SDK, and the api_server two-tenant test needs
  aiohttp, both PM runtime extras the test env does not carry.
2026-09-23 21:55:59 -04:00
teknium1
3a001a4d7f fix(tui-gateway): the hard-exit kill owns every foreground spawn and never blocks
Re-review findings on the foreground-process registry:

- Spawn vs exit races. A child spawned but not yet registered when the hard-exit kill ran, or a
  command launched after the kill took its snapshot, survived under init. The hard-exit kill now
  raises a one-way exit fence and waits (bounded) for spawns already past it to register; a spawn
  refuses once the fence is up, and one that registers after a timed-out wait kills itself.
- The immediate kill fell back to proc.kill() for every handle. On Modal/Daytona/Vercel that is a
  blocking SDK cancel (an 8s cancel made _hard_exit take 8s). Popen handles are still killed
  inline (killpg, never blocks); every other handle's kill runs on a daemon thread under one
  shared 0.5s deadline.
- The registry lock was a plain Lock: a signal landing on a thread that held it deadlocked the
  exit. It is now an RLock (via a Condition) and the hard-exit path only takes it with a timeout,
  falling back to a lock-free copy.
- The kanban worker's SIGALRM deadman os._exit()ed without the kill; it now goes through it.
2026-09-23 17:09:54 -07:00
teknium1
17a137d8a3 fix(tui-gateway): a SIGTERM-ignoring command no longer survives the gateway's SIGTERM exit
The SIGTERM handler arms a 1s os._exit timer, then runs _shutdown_sessions: a flush of up to
5s, then _stop_turns_before_exit, whose kill was the graceful TERM, wait 1s, KILL. A command
that ignores SIGTERM was still alive when the timer fired, and os._exit left it reparented to
init (live: `trap '' TERM; sleep 3600` survived a SIGTERM to `python -m tui_gateway.entry`).

- kill_live_foreground_processes(now=True): SIGKILL each in-flight foreground tree at once,
  no TERM grace, no wait (BaseEnvironment._force_kill_process; LocalEnvironment kills the
  recorded process group, never our own).
- The grace timer's exit (entry._hard_exit) runs it before os._exit.
- _stop_turns_before_exit SIGKILLs whatever is still alive halfway through its settle budget
  (it ignored the interrupt's TERM), so the tool call still ends with a result the teardown
  persists instead of a dangling tool_call in state.db.
- The other hard exits that skip cleanup do the same before os._exit: the serve parent-death
  watchdog, the CLI exit watchdog, the kanban worker's SIGTERM path, and the messaging
  gateway's shutdown and loop-liveness watchdogs.
- Deflake test_shutdown_mid_tool_kills_the_command_and_keeps_its_result: the 0.5s settle
  budget was too tight under -n 40 (1 red in 9 runs); the join returns when the turn ends.
2026-09-23 17:09:54 -07:00
teknium1
537ce77f52 fix(tui-gateway): exiting mid-tool no longer orphans the foreground command's process tree
Foreground terminal commands run in their own session (start_new_session) so an
interrupt can kill the whole tree, which also puts them outside the host's process
group. When the tui_gateway left mid-command (client closed stdin, or SIGTERM) nothing
killed them: _shutdown_sessions closed the agents, the SIGTERM path hard-exits after a
1s grace, and the `bash -c` + child tree survived, reparented to init.

- tools/environments/base.py: execute() records every in-flight foreground command;
  kill_live_foreground_processes() kills their trees through the backend's own
  _kill_process (the same kill an interrupt uses).
- cleanup_all_environments() (the exit funnel of the CLI, one-shot, messaging gateway
  and terminal_tool's atexit, so `hermes serve` too) kills them first.
- tui_gateway _shutdown_sessions (EOF atexit + SIGTERM handler) and the serve
  SIGTERM/SIGINT exit-flush handler interrupt running turns, wait up to 0.5s for them
  to settle so the tool call ends with a result the final persist records (no dangling
  tool_call in state.db), then kill any foreground command still alive.
- ComputeHost.close() kills them too: every caller os._exit()s right after.

Covers the case of b9dac83d366c (Desktop quit: serve SIGTERM handler) on every host.
2026-09-23 17:09:54 -07:00
ethernet
16652eea18 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	gateway/config.py
#	gateway/config_loader.py
#	gateway/readiness.py
#	hermes_cli/managed_scope.py
#	hermes_cli/plugin_python_deps.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/update_cmd_maint.py
#	plugin-catalog/hindsight.yaml
#	plugins/plugin_loader.py
#	providers/__init__.py
#	scripts/run_tests.sh
#	tests/gateway/test_control_socket_windows_live.py
#	tests/gateway/test_gateway_streaming_nested_config.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_plan_reconciliation_windows_live.py
#	tests/hermes_cli/test_update_apply_shallow_count.py
#	tests/hermes_cli/test_update_concurrent_quarantine.py
#	tests/hermes_cli/test_update_shim_self_lock.py
#	tests/hermes_cli/test_verify_console_scripts.py
#	tests/tools/test_lazy_deps.py
#	tests/tui_gateway/test_subprocess_encoding.py
#	tools/lazy_deps.py
2026-09-23 15:26:34 -04:00
tancou
c7c9c18ccf fix(profiles): pin the launch home so a mirrored HERMES_HOME cannot flip routed-profile decisions
Symptom: a host that serves several profiles from one process and mirrors
the active turn's profile into `os.environ["HERMES_HOME"]` for legacy
readers (Hermes WebUI does this on every chat turn, next to the
context-local override) makes every launch-home decision see the served
profile as the launch profile. Two profiles that both configure `atlassian`
with different credentials share whichever MCP connection came first: a
READ_ONLY_MODE=false profile ends up calling a read-only server
(nesquena/hermes-webui#7721). The same misjudgement leaves the launch
residue in the served profile's child env, seeds the launch profile's
bridged allow-all grant into the served profile's secret scope, and lets
the served profile's `terminal.*` config bridge into the shared process env.

Cause: four launch-home checks compare the task's override with
`get_process_hermes_home()`, which reads `HERMES_HOME` live:
`agent.secret_scope.serves_routed_profile` (keys the MCP ledger via
`_mcp_registry_scope`, #108352 / #111481, and the check_fn cache, #111151),
`agent.secret_scope._is_process_home`, `tools.environments.local._is_routed_home`
and `hermes_cli.env_loader._process_hermes_home`. Under the mirror the two
sides are equal for every turn.

Change: `hermes_constants.pin_process_hermes_home(path | None)` lets the
host record the home it serves as its own; `get_routing_process_hermes_home()`
returns the pin when set, else `get_process_hermes_home()`; the four checks
compare against it. The pin is deliberately NOT folded into
`get_process_hermes_home()`: `get_hermes_home()` falls back to it for tasks
carrying no override (MCP loop, spawners), and the host's mirror exists
precisely so those readers see the served profile. Only "is this task
routed / is this the launch home" changes. Unpinned, behaviour is
byte-for-byte the old one; hosts that never mutate `HERMES_HOME` need not
call it. `activate_multi_profile_hosting()` is not the seam for this: it
flips `get_secret` fail-closed process-wide and freezes the launch env,
which an embedding host cannot adopt as a bug fix.

Tests (2 invariants, parametrized over the four checks plus the MCP ledger
key; red on main, green here): pinned + mirrored env -> the served home is
routed and the launch home is not, the MCP key is `(home_key, name)`,
`get_process_hermes_home()` still follows the env var; never pinned or
pinned-then-cleared -> old semantics, including "a mirrored env var IS the
launch home". `tests/conftest.py` resets the pin per test so the
module-global cannot leak between files.

Live repro (WebUI + a stdio FastMCP server named `atlassian` in two
profiles, one gated by READ_ONLY_MODE): base -> one ledger key
`'atlassian'`, the write profile lists only the read-only tools; fixed ->
`(<read_home_key>, 'atlassian')` and `(<write_home_key>, 'atlassian')`,
each profile lists its own tools.

Docs: `gateway/AGENTS.md` § Profile scope (one launch-home identity) and the
isolation table in `website/docs/user-guide/multi-profile-gateways.md`.
Also maps the author e-mail under contributors/emails/ (attribution check).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 08:09:47 -07:00
ethernet
f4a38b1ee8 Merge origin/main into ethie/pm-clean 2026-09-23 10:13:26 -04:00
teknium1
b88b3de6e3 fix(profile-scope): stamp profile_home only for FOREIGN homes; cache the gate-prefix scan
The profile_home stamp exists to let serves_routed_profile() see a foreign-home scope without the
HERMES_HOME override. Stamping the LAUNCH home from the own-home binders (launch_profile_runtime_scope,
_worker_profile_scope's launch arm, model_switch's launch arm) added nothing for production but, under
a test conftest where server._hermes_home differs from HERMES_HOME, flipped the launch profile to
'routed' and charged every turn a routed tool/plugin discovery. Own-home binders now leave the stamp
None; the foreign-home arms keep it.

is_profile_gate_env resolved the platform prefix set (Platform enum + bundled plugin scan + registry)
once per env key; strip_profile_gate_env walks the whole child env, so a spawn paid the scan
hundreds of times. The static half is lru_cached, the dynamic registry union is taken once per strip.
2026-09-23 06:46:02 -07:00
teknium1
ebb9bc5699 fix(env-policy): a profile gate is a platform prefix plus a gate suffix, not a bare name shape
is_profile_gate_env matched any name containing _ALLOWED_ / _ALLOW_ALL_ / ... so an operator's own
DEMO_ALLOWED_SENDER (script data, not a Hermes authorization gate) was deleted from every routed
child env built with strip_launch_profile=True, breaking no_agent cron scripts under the host
gateway. The owner prefix now has to be a platform: built-in Platform values, bundled platform
plugins (directory names and manifest aliases), runtime-registered plugin adapters, plus GATEWAY_
(pairing) and QQ_ (qqbot). Every gate the adapters read still strips; operator variables survive.

Fixes #119539
2026-09-23 06:46:02 -07:00
ethernet
c9bd7459c2 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	scripts/run_tests_parallel.py
#	tui_gateway/model_switch.py
2026-09-22 00:28:30 -04:00
teknium1
10e7de79a9 fix(terminal): cache cleanup keeps the max_age_hours kwarg the housekeeping loop calls
7fe77f29a7 renamed cleanup_terminal_temp_cache's kwarg to max_idle_hours;
gateway/run.py::_run_media_cache_cleanup calls every registered cleanup
as fn(max_age_hours=24), so the terminal sweep raised TypeError inside
housekeeping on every tick. The kwarg is the shared cleanup_*_cache
signature (bot_relay notes the same parity); the idle semantics stay.
A test binds the loop's call shape against every registered cleanup.
2026-09-21 18:54:23 -07:00
teknium1
7fe77f29a7 fix(scratch): reap orphan processes and worktree registrations on prune; cache/terminal on the same 24h-idle rule
Deleting an idle scratch entry left two things behind. Headless browsers
started by a lane's e2e run kept running for days with a "(deleted)" cwd
(about 20 on one host; the process registry never knew them because they
were grandchildren of a shell that had exited). Repos whose linked
worktree lived in the entry kept a dangling registration until someone
ran `git worktree prune` by hand (10 in one repo).

The prune now lives in hermes_constants_scratch (hermes_constants keeps the
entry point). Before an idle entry is removed, same-user processes whose
cwd is inside it, or inside any scratch path that no longer exists, are
TERMed then KILLed; the deleted-cwd sweep runs on every pass so orphans
from earlier deletions are caught too. `.git` files found in the entry
name their repo, which gets `git worktree prune` after the rmtree.

cache/terminal used its own 72h fixed-age sweep; it now shares the 24h
idle rule and the subtree check. `hermes doctor` warns about cache-root
directories over 1 GiB that neither pruner covers, since finished campaign
trees parked there sat for weeks (95 GB on one host). System-prompt
scratch line updated to match.
2026-09-21 18:37:21 -07:00
ethernet
3917f11d79 fix(pm): skip System32/WindowsApps bash.exe, prefer Program Files Git
pm.shell.bash() returned shutil.which("bash") first on Windows. On most
machines that is C:\Windows\System32\bash.exe, the WSL launcher stub
(exits 1, "no installed distributions"), so a source install without the
staged git package lost its shell — the #116818 regression main fixed in
0abc3b04. The System32 guard survived only in local._windows_bash_candidates,
which had no callers after _find_bash collapsed onto pm.shell.

Move the ladder into pm.shell as a pure function (Program Files Git first,
then PATH minus the System32/WindowsApps stubs) so it is testable on any
host, and delete the dead copy in tools.environments.local.
2026-09-21 18:36:17 -04:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
teknium1
e234408607 fix(tools): tear down a child in the caller's own process group by PID, never killpg it
A spawner that skipped setsid (the Darwin gateway's posix_spawn shim) leaves
tool children in the gateway's process group, and _kill_process_group_posix
then killpg()s the gateway itself — launchd logs 'Killed: 9' and KeepAlive
respawns it. Branch on pgid == os.getpgrp() as a plain if/else (not a raised
PermissionError routed into the EPERM handler) and signal the wrapper plus
its snapshotted descendants by PID; the EPERM fallback shares that helper.

Fixes #107029
2026-09-20 16:32:19 -07:00
teknium1
d4b772cbda fix(tools): search_files keeps its matches when killpg is refused
On macOS a search that reaches its `limit` takes the early-stop branch that
TERMs rg's process group; rg has often already exited, and macOS answers
killpg on a zombie-only group with EPERM instead of ESRCH. The
PermissionError escaped _kill_process_group_posix, unwound to search_tool
and the drained matches were replaced by
`{"error": "[Errno 1] Operation not permitted"}` (#116855, same symptom
as #104696). The helper now treats EPERM like "nothing left to signal"
and falls back to killing the known PIDs so a live child cannot escape.
It also never killpg's the caller's own group (#107029): a child that
shares our pgid is torn down by PID.

Diagnosis credit: #116949 (@liuhao1024), redone slim.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 16:32:19 -07:00
chugiacan
0abc3b044b fix(agent,cli): skip WSL/system bash.exe in Windows bash discovery (#116818)
_windows_bash_candidates in tools/environments/local.py appended
shutil.which("bash") to the candidate list without filtering — when Git Bash
is installed but PATH resolves `bash` to WSL's stub (C:\Windows\System32\bash.exe)
or WindowsApps bash.exe, that stub enters the candidate list.

Fix: skip paths whose normalized form contains `system32` or `windowsapps`,
logging at debug level. Git Bash (already searched explicitly in the roots list)
remains the first usable candidate.

Fixes #116818
2026-09-20 15:50:26 -07:00
teknium1
940c610994 fix(security): keep _HERMES_PROVIDER_ENV_BLOCKLIST importable from tools.environments.local
Switching local.py to the folded _is_provider_env_blocklisted helper dropped the
blocklist name from its import list, which removed the re-export that
tests/providers/test_auth_registry_import_order.py (and any out-of-tree caller)
reaches through tools.environments.local. Re-import it alongside the helper so
the module keeps exposing the set; behaviour is unchanged.
2026-09-20 12:53:47 -07:00
beardthelion
b534f4b8c8 fix(security): match credential env names case-insensitively
The provider-credential blocklist and every adjacent env-name check ran
case-sensitive membership tests while the Windows environment block
resolves names case-insensitively. A skill or terminal.env_passthrough
entry registering openai_api_key was accepted, then resolved to the real
OPENAI_API_KEY by os.getenv() and forwarded into SSH/Docker exec envs,
the same GHSA-rhgp-j443-p4rf tunnel the blocklist closes. The same gap
let variant-cased credentials survive the source-side strip
(_filter_secret_env, _scrub_credentials), evade the docker_forward_env
and docker_extra_args egress collision guards, and ride
strip_launch_profile_env residue into a routed sibling profile's child.

Add _is_provider_env_blocklisted (exact + folded membership) and apply
it at every layer: registration refusal, both scrub paths (Tier-1 and
plugin strip sets fold too; inherit_credentials still inherits), the
remote exec-env builder, and the docker egress collision checks. Fold
_matches_terminal_first_party_prefix symmetrically so a lowercase-stored
buzz_private_key keeps its terminal carve-out, and fold the launch-
residue pop (selection folds too so a lowercase path in .env stays
global). _HERMES_FORCE_* opt-in transport and docker_env literal
container names stay case-sensitive on purpose.
2026-09-20 12:53:47 -07:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00
beardthelion
e885d9cdc2 fix(docker): drop -f from orphan reaper rm so running containers fail safe
reap_orphan_containers snapshots exited containers, then removes each
candidate with docker rm -f. The exited-only filter is only as fresh as
the docker ps snapshot: a sibling process can legitimately restart an
exited container between the snapshot and the rm (the reuse path calls
docker start on matching exited containers), and FinishedAt still
reports the previous exit, so the age check passes. The -f then kills
the sibling's live container and its in-container state.

A plain docker rm fails atomically in the daemon on a running container,
which is exactly the skip semantic the exited-only guard intends. Every
intended target of this sweep is already exited, so -f bought nothing.

The new test drives the real reaper with a simulated restart: rm argv
must not contain -f, and a daemon refusal on a running container is
skipped while the rest of the sweep continues.
2026-09-20 00:05:30 -07:00
teknium1
16ae9cfccb refactor: cache probe verdicts on exit-0 / capability rejection only, trim tests
Drop the nondeterminate-stderr fragment list: the rule is now simply "exit 0
caches True; a nonzero exit whose stderr names the probed capability
(cgroup / storage) caches False; everything else is uncached and retried on
the next spawn". A curated list of pull/daemon phrases would need to track
every docker CLI wording; the capability keyword is the property the cache
exists to record. Tests are trimmed to the invariants: the production
DockerEnvironment path recovers on the next spawn, a definitive cgroup
rejection is still cached, and the storage-opt probe caches only definitive
answers.
2026-09-20 00:04:54 -07:00
beardthelion
504e79d0d5 fix(docker): stop caching transient capability-probe failures as unsupported
_cgroup_limits_available latched _cgroup_limits_ok=False on ANY probe
failure — a TimeoutExpired while docker run auto-pulls an uncached image,
a daemon cold-start, a manifest/pull error, or a missing docker binary —
permanently stripping --cpus/--memory/--pids-limit from every later
container in the process. The sibling _storage_opt_supported latched the
same way, and additionally read a failed `docker info` (returncode never
checked) as "not overlay2", disabling disk quota for the process.

Both probes now cache only definitive answers: exit 0 -> True; a daemon
rejection naming the probed capability -> False. Any other failure
(missing docker, pull/manifest/daemon errors, timeouts) degrades that
spawn and is retried on the next. A shared nondeterminate-stderr list is
checked before the capability keyword so an image named e.g.
"cgroup-tools" whose pull fails is not misread as a cgroup rejection.
2026-09-20 00:04:54 -07:00
teknium1
9182ad0860 test: trim egress shorthand tests to two invariants, ignore positionals in the flag scan
Fold the salvaged parametrized positive/negative tables and the redundant
env-file/real-chain e2e tests into one parametrized invariant (every pflag
spelling of -e collides, legitimate args do not) plus one production-entry
test through DockerEnvironment. Args that do not start with "-" are
positionals under pflag and are never env flags, so the shorthand scanner
skips them instead of reading a container name like "de" as "-d -e".
2026-09-20 00:04:19 -07:00
beardthelion
10c30c8264 fix(docker): scan docker_extra_args env flags under pflag shorthand semantics
The egress collision scan only recognized -e/--env/--env-file as standalone
tokens, so the combined-shorthand spellings docker accepts, -eNAME=v joined
and -iteNAME=v chained after boolean shorthands, injected critical env names
past the guard. Parse each arg the way pflag does (boolean ditPq shorthands
may precede the terminal value shorthand) and stop at the -- terminator.

Fixes #115887
2026-09-20 00:04:19 -07:00
beardthelion
578c7aeae4 fix(docker): guard every egress-written env name in collision checks
_critical_egress_env_names seeded the forward_env/extra_args collision
set by scanning env_overrides for *_API_KEY/*_TOKEN suffixes. Mapped
tokens land under arbitrary real_env_name and alias_env_names entries
from mappings.json, so a non-suffixed credential name (a custom mapping,
an alias, AWS_SECRET_ACCESS_KEY) was injectable via docker_forward_env
or -e in docker_extra_args without tripping the enforced refusal,
placing the real credential inside the sandbox.

Every name the egress layer writes is egress-owned, so the critical set
is now env_overrides itself plus the proxy-control vars and NODE_OPTIONS:
complete by construction, no load_mappings dependency, and correct as
providers and aliases are added. This matches what
check_docker_env_collisions already protects via load_mappings.

Tests cover both surfaces: docker_extra_args -e and docker_forward_env
entries naming a non-suffixed mapped credential now refuse under
enforcement.
2026-09-20 00:04:19 -07:00
ethernet
e1576d06a6 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
2026-09-19 22:57:07 -04:00
teknium1
4586cde64d chore: mark the deliberate /tmp literals and shrink the lint baseline to one code block
The seventeen remaining literals are container-side paths, AF_UNIX socket-path-limit
candidates on darwin, detection needles, guard regexes and guidance text that tells the
model to avoid /tmp. Each carries an inline `no-tmp: ok — <why>` so the reason lives
next to the line; the baseline keeps only a fenced tree listing where a marker would render.
2026-09-19 10:44:26 -07:00
teknium1
d51f9c8464 refactor: real-home and terminal temp-dir fallbacks stop naming /tmp
get_real_home() fell back to a literal /tmp when no OS home could be found, and
LocalEnvironment.get_temp_dir() probed /tmp by hand before consulting
tempfile.gettempdir(), whose own candidate walk already covers the system temp
dir (and honours the scratch TMPDIR Hermes now exports). Both defer to
tempfile.gettempdir(); a relative gettempdir() result is made absolute instead of
being swapped for /tmp.
2026-09-19 10:44:26 -07:00
teknium1
2dcebe6471 feat: Hermes-owned scratch dir replaces the system temp dir for every process and child
Hermes and everything it launches (browser profiles, PTY probes, skill scripts,
tempfile defaults in delegated code) wrote to the system temp dir, which is a
RAM-backed tmpfs on most Linux hosts and containers and fills under agent load.

- hermes_constants.get_scratch_dir(): HERMES_HOME/cache/scratch (0700), entries
  older than 72h pruned once per process / once per hour across processes.
- apply_scratch_tmp_env(env) / export_scratch_tmp_env(): TMPDIR/TMP/TEMP point at
  the scratch dir when the user or OS has not set them; a value Hermes itself
  exported (== HERMES_SCRATCH_DIR) is re-derived for a re-homed process or a
  child served under another profile, so profiles never share scratch.
- hermes_bootstrap runs the export on import (every entry point); hermes_cli.main
  re-runs it after --profile resolution; the subprocess HOME contract
  (apply_subprocess_home_env) and the routed-home rewrites in code_execution_env
  and served_profile_child_env apply it to child envs.
- The runtime-environment prompt block names the scratch dir so the model stops
  reaching for the system temp dir by reflex; hermes doctor reports the dir, its
  size and whether a user TMPDIR overrides it.
2026-09-19 10:44:26 -07:00
teknium1
1bc953d342 docs(terminal): name the stdin=DEVNULL atom of the Git Bash probe (#73403)
Review follow-up: the comment in _bash_starts now says why stdin=DEVNULL matters
beyond the bounded cleanup — cygwin init's handle_to_fn/NtQueryObject stalls on a
pipe file object that has a read pending (the ACP host's stdin), which is why the
probe blew its timeout inside ACP hosts — and why stderr stays captured (the
Mandatory-ASLR remediation keys off bash's dofork:/child_copy: text).
2026-09-19 02:36:56 -07:00
teknium1
edae76fecd fix(terminal): bound the Git Bash startup probe's timeout cleanup (#73403)
_bash_starts used subprocess.run(timeout=15). On Windows run()'s post-timeout
cleanup is an unbounded communicate(): the probe's MSYS children (true/cat)
outlive the killed bash holding the captured pipe write ends, so the reader
join never returned and the ACP host wedged for minutes at the first terminal
tool (faulthandler dumps in the thread sit in _bash_starts → _communicate).

Route the probe through hermes_cli._subprocess_compat.bounded_probe_run — own
Popen (stdin=DEVNULL, hidden window), communicate(timeout), process-tree kill
and a 1 s bounded drain — so a probe that cannot finish fails fast and
_find_bash falls through to its last-resort candidate. A timeout is recorded as
the probe detail so the ASLR diagnostics still see it.

Supersedes #69083 (@fangliquanflq), whose bounded_captured_run design landed on
main as bounded_probe_run for the git probes (#68997); this is the surviving
_bash_starts call-site conversion from that PR.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-19 02:36:56 -07:00
teknium1
f0032daf94 fix(terminal): restore the unconditional stdout.close() after the Windows drain
Reverts c26e9ccb712. The close guard was written against the old blocking
os.read drain, where close() serialized on the CRT per-fd lock held by the
reader (#67362). _drain_fd_windows only calls os.read after PeekNamedPipe
reports bytes (and caps the read at that count), so it never blocks there and
close() cannot contend after join(2). The guard's only reachable branch — a
grandchild still streaming past the join — leaked the fd, the daemon thread
and collector appends for the grandchild's lifetime on Windows while POSIX
closed; closing unconditionally makes the drain thread exit on its next read
like the POSIX path. #67362 is closed by the PeekNamedPipe drain itself.

Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
2026-09-19 02:35:45 -07:00
teknium1
428777a669 fix(terminal): skip stdout.close() while a Windows drain thread is still reading (#67362)
On Windows close() serializes on the CRT per-fd lock held by a drain thread
blocked in os.read, so the natural-exit close blocked until the grandchild
exited — after the poll loop, where no timeout could see it. With the
PeekNamedPipe drain the thread normally exits before the join deadline; when it
has not (loaded host), leave the fd to the daemon thread / Popen.__del__.

Ported from #67373 (@PRATHAMESH75) and #67448 (@webtecnica).

Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
2026-09-19 02:35:45 -07:00
joaomarcos
e7a82eaded fix(terminal): bound the Windows stdout drain with PeekNamedPipe (#105865)
select() cannot poll pipe fds on Windows, so _drain_stdout used a bare
blocking os.read loop that only returned at true EOF. Any grandchild that
inherited the pipe's write handle (backgrounded `cmd &`, MSYS helpers)
kept the terminal/file tools hung for its whole lifetime — up to the 420s
tool timeout during init_session on real Windows hosts.

_drain_fd_windows polls PeekNamedPipe and mirrors _drain_fd_select's
stop-event + ~300ms idle-after-exit bound, so the drain returns shortly
after bash exits regardless of who still holds the write end.

Salvaged from #105981 by @JoaoMarcos44 (tests trimmed to two windows_only
invariants that run on the windows-latest job).
2026-09-19 02:35:45 -07:00
ethernet
bbec973514 refactor(pm): pm owns the dependency-environment layout and interpreter paths
hermes_cli.runtime_paths (venv generations, selection, activation) moves to
pm.environments, and gains venv_bin_dir / venv_python / project_python. Every
in-tree caller asks pm for an interpreter now; pm no longer reaches back into
hermes_cli for its own environment layout (pm.packages, pm.extras, pm.ensure,
pm.paths imported hermes_cli.runtime_paths). The three open-coded
"Scripts/python.exe or bin/python" ladders in pm collapse onto venv_python.

hermes_constants.venv_python_path / venv_bin_dir and hermes_cli.runtime_paths
stay as frozen-updater-surface shims only (tests/compat/old_updater_surface.json).

To keep the boot path light, pm/__init__ resolves its facade lazily (PEP 562)
and pm.registry loads the built-in package definitions on first read instead of
at import: `import hermes_bootstrap` now loads pm + pm.environments only (25ms,
was 37ms with the eager facade dragging in the downloader). The stripped-payload
fixtures that ship only pre-import files keep working for the same reason.

Also restores two frozen-surface re-exports the F401 sweep dropped
(banner._github_compare_behind, cua_backend.resolve_cua_driver_cmd).
2026-09-18 20:02:36 -04:00
ethernet
ee2b165e78 chore: drop merge-resolution notes, dead imports and a dead probe module
- 15 `MERGE-CHECK:` conflict-resolution comments removed from prod code (two were
  TODOs already done: the utf-8-sig sessions.json read lives in session_persistence,
  the pm-aware cron script helpers in scheduler_script).
- 49 imports the branch left unused (ruff F401, none present at the merge base,
  none inside PLUGIN-COMPAT blocks). update_cmd's frozen-surface re-exports are
  trimmed to the names tests/compat/old_updater_surface.json actually lists under
  hermes_cli.update_cmd; the rest resolve through hermes_cli.main.__getattr__.
- tools/environments/local_gitbash_probe.py: nothing imported it once _find_bash
  delegated to pm.shell().
- Three try/except wrappers around calls that cannot raise (install_truststore,
  get_hermes_home, and a duplicated except clause in supermemory).
2026-09-18 19:31:50 -04:00
ethernet
82a5affdd3 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/backup.py
#	tests/hermes_cli/test_gateway_restart_loop.py
#	website/docs/developer-guide/web-search-provider-plugin.md
#	website/docs/getting-started/installation.md
#	website/docs/getting-started/updating.md
#	website/docs/index.mdx
#	website/docs/reference/cli-commands.md
#	website/docs/user-guide/docker.md
#	website/docs/user-guide/windows-wsl-quickstart.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/developer-guide/plugins/index.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/developer-guide/web-search-provider-plugin.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/index.mdx
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/reference/cli-commands.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/reference/environment-variables.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/docker.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/features/plugins.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/security.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/windows-wsl-quickstart.md
2026-09-18 18:41:16 -04:00
teknium1
3ed40556ce fix(profiles): a child spawned for another profile no longer inherits the spawner's authorization gates
A `hermes -p B` child built from a process that loaded profile A's env (a gateway, the
dashboard, the post-update fleet restart) started with A's `DISCORD_ALLOWED_CHANNELS`,
`TELEGRAM_GROUP_ALLOWED_CHATS`, `GATEWAY_ALLOW_ALL_USERS`... and enforced them as its
own: gates are not credentials (no secret scrub sees them), a unit-file `Environment=`
or operator export is in no dotenv (no name-list strip sees them), and B's own `.env`
rarely defines the key (its dotenv load never overwrites the inherited value). Observed
as profile B's gateway rejecting every message in B's own channel after a per-profile
restart issued from A (#113270).

- `local_env_policy.is_profile_gate_env` / `strip_profile_gate_env`: gates matched by
  shape (`_ALLOWED_`, `_ALLOW_ALL_`, `_ALLOW_FROM`, `_ALLOW_BOTS`, `_IGNORED_CHANNELS`,
  ...), never `HERMES_*`, so a gate added to any adapter is covered without a second edit.
- `strip_launch_profile_env` drops them on its existing routed-home branch: the seam
  `served_profile_child_env`, the kanban dispatcher, cron workers and the dashboard
  action env already funnel through (the dashboard site now calls it for its target).
  Same-home children keep an operator export.
- `update_restart_recovery._child_environment(profile)`: the one site that bypassed
  every helper (bare `os.environ.copy()` relaunching EVERY profile) strips gates when
  the profile is not the one the updater runs as; the module stays stdlib-only at
  import time.

Live: fresh-process `hermes_cli.update_restart_recovery --stdin` with three gates in
the updater env — base hands all three to profile b's relaunch, fixed hands none and
keeps them for the launch profile.

Refs #113270; supersedes #113308 (@yashraj4, static key list + always-strip; this keeps
same-profile children intact and covers the per-adapter gate set).
2026-09-18 15:11:47 -07:00
ethernet
a6ae6ace51 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.github/workflows/js-tests.yml
#	agent/model_metadata.py
#	apps/desktop/electron/main.ts
#	apps/desktop/scripts/bundle-electron-main.mjs
#	apps/desktop/src/app/settings/about-settings.tsx
#	apps/desktop/src/app/settings/gateway-settings.test.tsx
#	apps/desktop/src/app/settings/gateway-settings.tsx
#	apps/desktop/src/app/updates-overlay.tsx
#	gateway/shutdown_flush.py
#	hermes_bootstrap.py
#	hermes_cli/local_runtime/binaries.py
#	hermes_cli/main.py
#	hermes_cli/managed_uv.py
#	hermes_cli/update_cmd.py
#	hermes_cli/update_cmd_deps.py
#	hermes_cli/update_cmd_fleet.py
#	hermes_cli/update_cmd_maint.py
#	hermes_cli/update_receipt.py
#	hermes_cli/update_serve_obligations.py
#	hermes_constants.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_managed_uv.py
#	tests/hermes_cli/test_pending_supervisor_recovery.py
#	tests/hermes_cli/test_startup_fast_guards.py
#	tests/hermes_cli/test_update_desktop_stale_warning.py
#	tests/hermes_cli/test_update_fleet_restart_pending.py
#	tests/hermes_state/test_hermes_state.py
#	tests/tools/test_tirith_security.py
#	tools/bot_relay.py
#	tools/checkpoint_manager.py
#	tools/write_approval.py
#	website/docs/getting-started/updating.md
#	website/docs/reference/environment-variables.md
2026-09-18 17:26:10 -04:00
teknium1
547fff7500 fix(tools): scope-only passthrough overlay raises instead of silently dropping the declared secret
_scrubbed_env wrapped the scoped_passthrough_additions overlay in try/except Exception with a
debug log; a scope or config failure there would silently drop the declared secret again — the
exact failure mode #114209 complains about. The import cannot fail in-tree and _scrub_child_env
already calls it unguarded, so the guard is removed and one test pins that both local surfaces
propagate the error.
2026-09-18 10:04:57 -07:00
teknium1
e1c0896518 fix(cron): no_agent script env comes from the factory's own snapshot, not a raw copy at the spawn site
tests/agent/test_subprocess_env_guard.py flagged cron/scheduler_script.py:362 as a new raw
os.environ.copy() spawn-env site. build_subprocess_env gains strip_launch_profile=True so the
launch profile's .env residue is still dropped from the base before the secret scrub (same
order and semantics as before: strip first, then scope-overlay the owning profile's declared
names), and the spawn site no longer snapshots the environ itself.
2026-09-18 10:04:57 -07:00
teknium1
802a9975d2 fix(cron): no_agent scripts get the owning profile's declared secret, never the launch profile's
A no_agent cron script owned by a profile served by a multi-profile gateway (or the
Desktop/dashboard backend) could not receive its own credential: the served profile's
.env never enters the process env (load_hermes_dotenv skips the process-global load
for a routed home), and the child-env sanitizer only resolved a terminal.env_passthrough
name through the profile secret scope when the launch environment already carried that
name. So the documented declaration mechanism (SECURITY.md 2.3) could not supply a
value that exists only in the served profile's scope, while the launch profile's own
.env credentials rode into the served profile's script child unstripped.

- tools/env_passthrough.scoped_passthrough_additions: declared names the bound scope
  holds but the env being filtered lacks; reads the scope alone (no os.environ, no
  other profile), empty without a scope so single-profile spawns are byte-identical.
- _scrubbed_env (terminal, cron scripts, bg processes, search workers) and
  execute_code's _scrub_child_env overlay those names after filtering.
- cron _run_job_script and terminal _make_run_env strip the launch profile's .env
  residue from the base (strip_launch_profile_env, a no-op for the launch profile's
  own jobs) before the filter, so the scope overlay lands after the strip.

Why not the 1Password-only shape of #114218: the gap is the declaration mechanism, not
one vendor's token, and an unconditional token pop broke the single-profile .env flow.
Docs: cron no_agent credential section, secrets child-process section, security
passthrough note.

Fixes #114209
Supersedes #114218
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-18 10:04:57 -07:00
kokhlo
b911b9794e fix(tools): sanitize cwd overrides written into live container envs
A cwd override registered by a desktop/TUI surface is a raw host path. On
container backends it was written verbatim into the cached live env's cwd,
so every file-tools command wrapper did `builtin cd -- <host path> || exit 126`
and all writes outside the mapped workspace failed with an unrelated `cd:`
error while terminal commands kept working (their per-command resolver
already sanitizes).

The live-env write now mirrors the creation-path guards: a host/relative cwd
is remapped to /workspace when it is the directory mounted there (docker cwd
passthrough), otherwise not applied to the live env at all. The session cwd
record keeps the raw path and non-container backends apply overrides
verbatim, so ACP project-root switching is unchanged.

Fixes #113894

(cherry picked from commit d712c6b47d28f213773f40ebac81dd9a45c6e94f)
2026-09-18 09:54:45 -07:00
teknium1
775aedabce fix(file-sync): sync-back cap is a config.yaml key, not an env var
Non-secret settings live in config.yaml: the extraction cap override becomes
terminal.sync_back_max_bytes in DEFAULT_CONFIG (read via load_config() like the
rest of the terminal section) and the HERMES_SYNC_BACK_MAX_BYTES env var goes
away; docs and the cap tests follow. psutil is a pinned core dependency, so
_temp_entry_owner_alive imports it unconditionally instead of guarding an
ImportError that cannot happen.
2026-09-18 09:30:16 -07:00
teknium1
d007afe130 fix(file-sync): bound sync-back leaks by owner PID, make the cap overridable, cover every tar backend
Follow-up to the salvaged #114440 commits:

- Stale sweep: a dead owner's entry goes at once, but a live (or recycled)
  PID no longer exempts an entry from the 30 min age cutoff — a recycled PID
  must not pin a multi-GB leak forever. Liveness via psutil.pid_exists (the
  footgun scanner rejects os.kill(pid, 0): it terminates on Windows), so the
  os.name guard and the two extra helpers collapse into one.
- HERMES_SYNC_BACK_MAX_BYTES overrides the 2 GiB extraction cap for trees that
  legitimately exceed it (the reporter's 4.26 GB tree was downloaded and
  discarded on every attempt); the skip warning names the override.
- Vercel sandbox bulk download excludes *.sock like SSH/Modal/Daytona.
- Tests trimmed to two invariants per file; docs mention the cap override,
  the socket skip and the per-PID temp naming.
2026-09-18 09:30:16 -07:00
liuhao1024
e928655b85 fix(tools): reclaim sync-back temps by owner PID and harden socket tolerance
Review follow-ups on #114440:

- Drop the HERMES_SYNC_BACK_MAX_BYTES env override: a malformed value
  crashed the whole backend at import time, and the rubric keeps env
  vars for secrets only. The 2 GiB cap stays hardcoded (raising it can
  be a separate change).
- Embed the owning PID in sync-back temp names and reclaim dead-owner
  entries immediately, mirroring daytona's PID-suffixed remote temp.
  Age no longer decides liveness (a staging dir's mtime does not move
  while content streams into subdirectories, and the download bound is
  SSH/Modal-only), and concurrent gateway processes under different
  HERMES_HOMEs never sweep each other's transfers. Legacy (pre-PID)
  names and Windows hosts keep the 30-minute cutoff.
- Anchor the rc=2 socket tolerance to lines ending in ': socket
  ignored' and require a non-empty diagnostic, so whitespace-only
  stderr or a filename merely containing the marker still fails.
- Add --exclude='*.sock' to the Modal and Daytona bulk-download tars
  (same failure class as the SSH backend).
2026-09-18 09:30:16 -07:00