446 Commits

Author SHA1 Message Date
John Paul Soliva
136135e653 fix(cron): overlay the routed no_agent scope inside build_subprocess_env
Main moved the no_agent script env to the factory (build_subprocess_env(strip_launch_profile=True),
e1c0896518) so the spawn site takes no raw os.environ snapshot. The restored routed-fire overlay
(installed scope, then managed keys, both before the scrub) now lives behind that same flag instead
of a dict(os.environ) copy at the spawn site: strip -> overlay -> sanitize, unchanged order, and a
no-op outside multiplex semantics, so a single-profile child stays byte-identical. The flag's only
caller is the cron script runner.

(cherry picked from commit 800c37e7323443c4d4719c0d6a43e028018286e3)
2026-09-29 22:57:48 -04:00
kshitijk4poor
adf3948d76 fix(env-loader): split source_supplied_names() out of secret_source_names()
Widening secret_source_names() to include skipped_existing names silently changed
tools/mcp_tool_config.py::_build_safe_env, an untouched consumer that forwards every
returned name into MCP stdio child envs. That consumer wants only names a source
actually APPLIED (pre-stack semantics), so secret_source_names() goes back to
tuple(_SECRET_SOURCES).

The routed-child scrub in strip_launch_profile_env is the one site that must also see
names a source supplied but lost to a pre-existing process value, so it reads the new
source_supplied_names() accessor instead. tools/mcp_tool_config.py is byte-identical to
origin/main.

(cherry picked from commit dcdbcb8a2b)
(cherry picked from commit b1b07b03440e1e35a2a2c8e33738618399c8b900)
2026-09-29 22:57:48 -04:00
kshitijk4poor
03d3c52855 refactor(cron): strip external-source residue in strip_launch_profile_env itself
_run_job_script popped the launch profile's secret-source names in an
inline loop right above strip_launch_profile_env, so only the no_agent
child got that protection; the four other callers of the same helper
(the external cron worker, scheduler_delivery, kanban dispatch and the
byterover plugin) still handed a served profile the launch vault or
1Password names. Fold the names into the helper's residue set, which is
already gated on multiplex and on the target not being the launch
profile, and keeps administrator-managed keys.

(cherry picked from commit f367ebeb5d)
(cherry picked from commit a5cba0db7743b6a927df56bc09df5be6bb45210f)
2026-09-29 22:57:48 -04:00
John Paul Soliva
65bd3eb336 fix(cron): routed fires are multiplexed at the worker handoff; managed keys keep policy precedence
Review findings on f5f88d5058. Three are defects the previous round introduced.

Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.

Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.

Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.

Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.

Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.

(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
(cherry picked from commit 62b4488cb5)
(cherry picked from commit d3319102a7e4dcce35ec2676fe22aed41613e893)
2026-09-29 22:57:48 -04:00
John Paul Soliva
66e3090a4f fix(cron): close three launch-residue leaks into a routed no_agent child
Review findings on d8c467f223, each reproduced through its production path.

Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.

Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.

Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).

Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.

(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
(cherry picked from commit 9d7de6c140)
(cherry picked from commit b0beb731b18edd884e36a30ab91c91687bc7a965)
2026-09-29 22:57:48 -04:00
liuhao1024
89dd61a281 fix(terminal): route GUI launches to the running Bot Desktop screen
A terminal command that starts a GUI app (flatpak run, google-chrome,
xdg-open) inherited the gateway's own DISPLAY, so while a Bot Screen was
running the window opened on the user's seat and xrandr/xdotool probed
the user's display — the exact babysitting loop of #125830. Merge the
launcher's published DISPLAY/XAUTHORITY/DBUS into _make_run_env while
the Bot Desktop is up (same routing the browser tool and cua backend
already apply), via published_env() so a plain command does not stamp
screen activity and idle_stop_minutes stays effective. An inline
DISPLAY=:N prefix still wins; with the desktop down the seat env passes
through untouched.
2026-09-29 18:49:11 -05:00
fangliquanflq
e1ab557716 fix(desktop): secure SSH workspace browsing
Route remote workspace filesystem operations through the extracted files router without starting agent-file synchronization.

Resolve SSH download targets before applying sensitive-path policy and stream authorized files without buffering or preview-size limits.

Tests: python -m pytest tests/hermes_cli/test_web_server_fs.py tests/hermes_cli/test_ssh_workspace_fs.py tests/tools/test_ssh_environment.py -q.

Original implementation authored by fangliquanflq.
2026-09-29 15:28:19 -05:00
kshitijk4poor
a0e20b60d9 fix(child-env): a partial plugin scan keeps the denials it already knew
andrexibiza (review on 60bdd5fbe3): skipping an unreadable plugin dir,
or an unreadable or unparsable flat manifest, replaced the home's
cached secret set with the partial result. A secret declared a moment
ago could then reach children while its manifest was unreadable. And
because the partial result was cached under the same file signature,
recovery could keep hitting it.

platform_manifest_secret_scan() now reports whether the scan was
complete. A partial scan unions in the names this home already had,
and is cached with no stamp, so the next spawn rescans and a recovery
is seen at once. A deleted plugin still releases its names, because
that scan is complete. An unreadable plugins/platforms/ manifest still
raises.
2026-09-29 02:06:58 +05:30
kshitijk4poor
9e26d34475 fix(child-env): manifest scan skips only what plugin discovery skips
Per-commit review follow-ups:

- Dot directories are scanned again, matching plugins_discovery. A
  platform plugin that discovery loads from plugins/.x now also has
  its secrets declared.
- Only PermissionError means a plugin directory cannot be searched. A
  symlink loop is treated as missing, so plugin.yml is still checked,
  as discovery does. A plugin.yaml that is not a regular file is
  ignored rather than failing every strict read.
- Dropped a redundant union: the home set is already stripped in
  Tier 1.
- Tests now pin the dunder skip, the unlistable platforms root, and a
  stamp change on an edit that keeps the mtime.
2026-09-29 02:06:58 +05:30
kshitijk4poor
b1206115a2 perf(child-env): one scandir manifest stamp per scrub, file-signature keyed
Per-commit review follow-ups on the per-home plugin declarations:

- The manifest stamp walked each plugin directory with pathlib
  is_dir/exists/stat on every spawn, and _scrub_credentials took it
  twice. It is now one scandir per root and one stat per candidate,
  taken once per scrub.
- The stamp uses utils.file_signature, so a manifest replaced with its
  mtime preserved (cp -p, rsync -t) still invalidates the cache.
- An unsearchable plugin directory, including an unlistable flat
  plugins/ root, is reported once as a directory entry. The stamp walk
  stays quiet; the warning comes from the manifest read, which runs
  only on a cache miss.
- The dot/dunder skip now matches plugins_discovery.scan_directory, and
  the manifest source argument is a Literal.
- Tests: the unsearchable-dir case uses a non-dunder dir, so both
  guards are pinned. New tests cover the bundled read failing closed,
  manifest deletion releasing its names, and a symlinked home alias
  sharing one cache entry.
2026-09-29 02:06:58 +05:30
kshitijk4poor
8ffcd7488f fix(child-env): openviking drops the profile overlay's bot tokens; unprovable manifests don't block spawns
Phase 2 review follow-ups:

- OpenViking overlaid the bound profile's whole .env after the scrub,
  so that profile's bot, dashboard and relay tokens reached the server
  again. The overlaid env is now scrubbed a second time for Tier 1.
  Provider keys still pass, and they still come from the bound profile.
- The strict per-home manifest read raised on any unreadable entry
  under <home>/plugins, which failed every child spawn for that
  profile. It is now strict only where the manifest is known to be a
  platform's: bundled or plugins/platforms/. Dunder and dot
  directories, unsearchable plugin directories and unreadable flat
  plugins/* manifests are skipped with a warning, because they can't
  load as plugins either.
- The per-home cache is keyed on hermes_home_key(), so a symlinked
  alias no longer gets a second entry.
- The modal snapshot test swapped hermes_cli.config for a stub, which
  the policy import can no longer use. HERMES_HOME already points the
  real module at the test home, so the stub is gone.
- The env_passthrough docstring now describes declared names rather
  than the dropped prefix rule.
2026-09-29 02:06:58 +05:30
kshitijk4poor
1ede0b805b refactor(child-env): one secret-suffix rule, one env-table walk, one bundled manifest read
/simplify-code follow-ups on the declared-secret policy:

- The policy and the manifest classifier each defined which env-name
  suffixes make a platform variable a secret, and they disagreed on
  _JSON. PLATFORM_SECRET_ENV_SUFFIXES in hermes_cli/config.py is now
  the only definition.
- The gateway env-override table was walked by hand a fourth time,
  through private gateway.config_env names. The walk is now
  profile_channels.config_env_table_keys(), shared with
  declared_channel_env_keys.
- The bundled platform manifests were parsed twice at import (about
  16 ms). The config injection reads them once, strictly, and keeps
  their secret names. The policy re-reads only if that read hit an
  I/O error.
- The per-home cache was keyed on directory mtimes, so an in-place
  plugin.yaml edit went unseen until restart. It is now keyed on
  every manifest file's mtime.
- The core-declared snapshot is a module-level constant rather than
  a global set inside the injector. The two manifest-source booleans
  are now one source argument. The Tier 1 set is upper-cased once
  at import, not on every spawn.
2026-09-29 02:06:58 +05:30
kshitijk4poor
3a6406137d fix(child-env): per-profile plugin secrets, fail-closed declaration scans, dashboard auth in Tier 1
Review follow-ups on the declared-secret policy:

- A profile's user-installed platform plugins declared secrets into
  one process-wide set, read from the launch home at import. Under
  multiplex that stripped profile B's same-named user variable and
  missed profile A's own declarations. Bundled manifests stay
  process-wide; user manifests are now read per bound home, cached
  until that home's plugin dirs change, and are Tier 1 for that
  profile only.
- The declaration scans no longer fail open. A registry error is no
  longer swallowed into an empty set, non-string required_env entries
  are skipped rather than breaking the set, and a manifest that
  cannot be read raises instead of vanishing from the policy. A
  malformed manifest still declares nothing.
- A plugin manifest can no longer reclassify a core-declared name
  such as OPENAI_API_KEY; the same rule the config form applies.
- The dashboard basic-auth password and signing secret, the OIDC
  client secret and the drain bearer move to Tier 1. Credentialed
  CLIs (claude, codex) no longer receive them.
- openviking-server starts from served_profile_child_env, so it gets
  the bound profile's provider keys, never the launch profile's.
  With no bound profile under multiplex the start is refused.
- NOUS_API_KEY and QWEN_API_KEY are listed statically. Discovering
  provider plugins while the policy module imported re-mirrored them
  over a plugin's own auth registry entry
  (tests/providers/test_auth_registry_import_order.py).
2026-09-29 02:06:58 +05:30
kshitijk4poor
6c2a53367f fix(terminal-env): block Hermes' own dashboard, anon and Meet secrets
Review of the adapter-secret salvage found Hermes-owned secrets that no
declared source covers, so they reached terminal and execute_code
children and skill passthrough accepted them: the dashboard basic-auth
password and signing secret (enough to forge dashboard sessions; the
session token beside them is already Tier 1), the dashboard drain and
OIDC client secrets, HERMES_ANON_API_SECRET (a provider-category entry,
which the OPTIONAL_ENV_VARS loop never blocks) and the Google Meet
realtime key. Add them to the static blocklist.
2026-09-29 02:06:58 +05:30
kshitijk4poor
ee942a7bbf fix(terminal-env): adapter secrets are the declared ones, not any platform-named variable
The shape rule (a platform prefix plus _TOKEN/_SECRET/_PASSWORD/_KEY) also
matched variables Hermes never reads: the platform list holds plain words
(LOCAL, GATEWAY, WEBHOOK, SLACK), so a user's SLACK_USER_TOKEN,
LOCAL_LLM_API_KEY or GATEWAY_API_KEY vanished from the terminal and
terminal.env_passthrough could not bring them back. The prefix census it
leaned on also failed open (an unreadable plugins dir cached an empty set).

Adapter secrets now come from what adapters declare: password entries of
the messaging OPTIONAL_ENV_VARS (built-ins plus every platform plugin
manifest) and secret-named keys of the gateway env-override table, both
Tier 1 and refused by passthrough; plus the secret-named required_env of
adapters registered in the current profile scope, read per spawn without
loading deferred adapters, Tier 2 only because required_env is an
unchecked setup list. Secrets nothing declared get a manifest entry
(TELEGRAM_WEBHOOK_SECRET, PHOTON_SIDECAR_TOKEN, A2A_PUSH_SECRET,
TEAMS_GRAPH_ACCESS_TOKEN, TEAMS_INCOMING_WEBHOOK_URL) or join the
policy's read-in-code list (QQ_STT_API_KEY, the two MSGRAPH names).
2026-09-29 02:06:58 +05:30
John Paul Soliva
dca8684c59 fix(terminal-env): adapter secrets are Tier 1, owners are read per call
An inheriting child (claude/codex/gemini) kept TELEGRAM_WEBHOOK_SECRET, WHATSAPP_CLOUD_ACCESS_TOKEN and the Graph secrets because the shape rule sat only in the provider tier; it now strips with the bot tokens. The shape rule reads the per-call adapter census, so a plugin adapter registered after import, in the bound profile only, is covered. The OAuth provider scan takes bundled profiles only, so the process-wide blocklist no longer freezes the home bound at import.
2026-09-29 02:06:58 +05:30
John Paul Soliva
66019e5597 fix(terminal-env): strip adapter secrets and OAuth-profile keys from child envs
The child-env blocklist is derived from the provider registry and
OPTIONAL_ENV_VARS, but many gateway adapters read their secrets straight
from the environment without listing them there: WHATSAPP_CLOUD_ACCESS_TOKEN,
WHATSAPP_CLOUD_APP_SECRET, WEIXIN_TOKEN, YUANBAO_APP_SECRET, FEISHU_ENCRYPT_KEY,
TELEGRAM_WEBHOOK_SECRET, PHOTON_SIDECAR_TOKEN and others. They reached
terminal, background/PTY, cron-script and hermes_subprocess_env children,
and a skill could register them as env passthrough, while the documented
bot tokens next to them were stripped. OAuth provider profiles (nous,
qwen-oauth) also accept a pasted key (NOUS_API_KEY, QWEN_API_KEY), but the
registry mirror copies env_vars only for api_key profiles, so those passed
through too.

Match adapter secrets by shape, the way authorization gates already are: a
built-in or bundled adapter prefix plus a _TOKEN/_SECRET/_PASSWORD/_KEY
suffix, so a new adapter secret is covered without a second edit. Add every
provider profile's env_vars regardless of auth_type, and list the two
Microsoft Graph secrets no adapter prefix owns.
2026-09-29 02:06:58 +05:30
teknium1
8d8836ddb1 feat(telemetry): harness-accuracy shared metrics for the agent loop
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):

- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
  matcher reports the strategy that landed (or no_match / ambiguous) into a
  context-local probe opened only around the patch/write_file handlers, so we
  learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
  the guardrail already warns/blocks/halts and at turn end for the iteration
  budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
  next_outcome. One row per failed tool call, resolved against the model's
  next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
  a table lookup on the first program word, never the text. Hermes' own
  deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
  backend result, so a command's own `exit 124` reads nonzero, not timeout.
  Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
  truncation only from the structured finish reason; empty/reasoning_only
  from the normalized message; issue=none per response as the denominator
  (model_route counts attempts and files empty replies as failures).

Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
2026-09-28 12:43:03 -07:00
Teknium
7c799a6565 feat(vercel): fresh sandboxes use a managed image (universal:latest); runtime presets deprecated, migration 49 (#126741)
* feat(vercel): start fresh sandboxes from a managed image instead of the deprecated runtime

Vercel deprecated Sandbox runtimes (node24/node22/python3.13) in Aug 2026 in favour of
images, and rejects runtime+image together and runtime with a snapshot source. New
terminal.vercel_image (default vercel/sandbox/universal:latest, Node 24 + Python 3.14)
picks the image for fresh sandboxes; a pinned terminal.vercel_runtime still works, wins
over the image and logs a deprecation warning; snapshot restores send neither.

Setup wizard prompts for the image, dashboard exposes both keys, status/config show the
effective choice, TERMINAL_VERCEL_IMAGE bridges config to the tool like its siblings.

* feat(config): migration 49 drops the seeded node24 Vercel runtime pin

Every pre-49 config.yaml carries terminal.vercel_runtime: node24 (the template default) and
the setup wizard mirrored it into .env as TERMINAL_VERCEL_RUNTIME. Both are the default
copied, not a choice, so the migration drops them and fresh sandboxes follow vercel_image;
node22 / python3.13 pins are the user's and survive. Persisted sandboxes are unaffected:
a snapshot restore never sends a runtime or an image.
2026-09-28 12:33:01 -07:00
Teknium
3f871425af fix(modal): persistent sandbox snapshots no longer expire after 30 days (modal 1.5.5, ttl=None) (#126740)
* fix(modal): keep persistent-sandbox snapshots past the SDK's 30-day TTL

modal>=1.5 gives Sandbox.snapshot_filesystem() a default ttl of 30 days, so an idle
persistent Modal sandbox silently lost its filesystem and restarted from the base image.
Pass ttl=None (retain until deleted) and bump the modal extra from 1.3.4 (no ttl
parameter; legacy RPC) to 1.5.5 so the kwarg exists on every install.

* fix(modal): drop the dead modal.Mount credential-mount block

modal.Mount left the public API in modal 1.0, so _modal.Mount.from_local_file raised
AttributeError into the surrounding except on every sandbox start and the block never
mounted anything. The FileSyncManager created right after already uploads the same
credential, skills and cache files (iter_sync_files), so delete the duplicate; the test
fake stops exporting a Mount the real SDK does not have.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-28 12:31:03 -07:00
Teknium
a5bd246865 Old pre-decomposition import paths are gone: plugin compat layer removed on schedule (#126164)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule

The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:

- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
  names, restored dead definitions) and the three re-export stub modules
  (gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
  the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
  `plugins.allow_deprecated_imports` escape hatch

An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.

hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).

In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.

* chore: retrigger CI (zero-job startup_failure phantom)

* test: drop resolution allowlist rows for the two deleted which() sites

hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
2026-09-28 10:21:41 -07:00
teknium1
4cb2c8ae76 fix(review): a named profile's restart --all never runs the root gateway in-process
MAJOR: `hermes -p X gateway restart --all` with no installed service re-entered
`run_gateway(replace=True)` IN the CLI process. `_home_env` swapped only HERMES_HOME
while os.environ already held X's dotenv (loaded at CLI startup) and gateway/run.py's
root .env load does not clear inherited keys, so the default gateway came up with X's
TELEGRAM_BOT_TOKEN / OPENAI_API_KEY. And any child spawned inside the swap (Windows
`start --all`, the launchd/detached fallback, migrate's secondary spawns) read the
swapped home as the launch profile, so `served_profile_child_env` neither stripped nor
scrubbed X's env.

- `_home_env` pins the process's launch home for the swap (`pin_process_hermes_home`),
  so routed-home decisions inside it keep the real identity; `strip_launch_profile_env`
  reads the residue list from that pinned home too.
- `_restart_all_as_host` spawns the root detached (`_spawn_detached_gateway`, scrubbed
  base env + the root's own secret scope, `gateway run --replace`) when the CLI runs
  under a named profile; the default profile keeps its foreground run. The Desktop
  update hand-off's `gateway start --all` under a named active profile is still served.
- test: `test_named_profile_restart_all_with_host_down_restarts_the_default_root` now
  asserts the child env carries none of the named profile's keys and `run_gateway` is
  never entered (red on the PR head: "the root ran IN the named profile's process").

MINORS:
- (a) `gateway_migrate._service_op('enable')` raises on a non-zero `systemctl enable`;
  a fresh apply treats it as a preflight refusal before the flag/manifest write, a
  resume reports it and converges. New `test_migrate_refuses_when_the_survivor_cannot_
  be_enabled` (red on head: exit 0 with the ⚠ line).
- (b) `systemd_install` already-current branch warns when `systemctl enable` fails.
- (c) `remove_system_systemd_unit` reports a failed `systemctl stop` instead of ✓.
2026-09-28 04:04:23 -07:00
Teknium
399d956903 feat(bot_desktop): Bot Screen, computer_use and the browser run inside the terminal backend (#121169)
* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends

The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.

docker/sandbox-desktop.Dockerfile is that base plus:
  - the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
    zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
  - the exact package set the Hermes -desktop image installs (TigerVNC,
    Xfce components, dbus, xauth, fonts)
  - Playwright's headed Chromium (same build as the -desktop image)
  - cua-driver 0.28.2 from its pinned release tarball

No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.

docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.

* feat(docker): sandbox desktop base on python3.13-nodejs26

Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.

* feat(docker): bake agent-browser into hermes-sandbox:desktop

The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.

* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend

A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.

Now the desktop lives where the terminal lives:

- tools/environments/streams.py: one primitive per spawn-per-call backend, the
  local argv prefix that runs its remainder inside the sandbox with stdio open
  (`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
  vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
  the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
  prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
  gateway). `auto` follows the backend; a sandbox that cannot host a screen
  REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
  lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
  there; the host driver's runtime contract is irrelevant then; check_fn is
  true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
  with the daemon, socket dir and profile inside the sandbox; screenshots are
  fetched back so MEDIA: paths keep working; recycle closes the sandbox
  daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
  in a sandbox, a unix socket otherwise.

Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.

* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only

DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).

* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read

Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.

* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order

Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host

sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.

* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox

Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).

* chore: retrigger CI (zero-job dispatch failure)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)

* chore(config): template stamps v47 and shows the new default sandbox image

The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.

* feat(sandbox): the default image change is a decision, not a surprise

A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.

Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.

Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.

Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.

* fix(config): both plain defaults that preceded the desktop sandbox image are template copies

main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.

* ci(docker): build the sandbox image on release/dispatch, not every main push

Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.

* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta

The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.

* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)

* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)

* fix(bot_desktop): placement is the authority; sandbox screen survives restarts

Review findings on the sandbox-hosted Bot Screen, each reproduced live first.

Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.

Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.

SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.

CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.

pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.

Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.

Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.

Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.

* docs(bot-screen): no literal tmp path in the profile-location note

* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)

* fix(bot_desktop): adopting a screen the sandbox kept records the marker

Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.

Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).

* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)

* docs(bot-screen): what the Apptainer path inherits from the image and what it does not

* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)

* chore: retrigger CI (zero-job dispatch failure, auto-heal)
2026-09-28 03:34:07 -07:00
brooklyn!
1a72042115 fix(docker): bind a Windows workspace when /workspace is already claimed
A volume that already owns /workspace skipped the configured working
directory, so tools treated that host path as unmounted. Bind it at a
second mount, or point tools at the volume that already has it, for any
drive path.
2026-09-26 18:27:10 -05:00
beardthelion
5b8fd7fc32 fix(code-execution): lock down remote kernel/RPC dirs, keep RPC token out of argv
On shared remote backends the execute_code channel created kernel and
sandbox dirs under shared temp at the process umask (775 group-writable
under umask 002), wrote request/result files group-readable, and carried
HERMES_RPC_TOKEN on remote command lines where co-tenant users read argv
via ps for the whole run. A co-tenant could read tool arguments and
results, and on group-writable dirs forge RPC requests dispatched under
the user's approval context.

- All remote dirs are created owner-only (umask 077 + chmod 700, checked
  fail-closed) and every Hermes file write is mode 600.
- The token travels in a sourced env file inside a subshell so the vars
  never enter the backend's session-snapshot dump, and ships via stdin on
  pipe-capable backends so it never enters argv at all.
- The RPC poll loop rejects non-int seq requests before dispatch instead
  of replaying them every cycle.
- tool_result_storage gets the same owner-only treatment for archived
  tool output.

(cherry picked from commit aef21731d7fb8a4e0a6ada4ff9889264df4a8893)
2026-09-27 00:56:31 +05:30
kshitijk4poor
95fc717460 chore(environments): drop uuid imports left unused by the shared staged-stdin path 2026-09-26 23:44:55 +05:30
kshitijk4poor
8f5b5cd88b perf(environments): skip stdin staging for empty payloads
execute() can pass stdin_data="" (e.g. write_file of empty content). The
`is not None` guard then paid an upload (plus chmod on Daytona) and a
longer shell command just to feed zero bytes. Base heredoc mode skipped
empty stdin, and neither SDK exec attaches a stdin, so treating "" as no
stdin gives exactly the base command (checked: identical argv, no upload).
2026-09-26 23:44:55 +05:30
kshitijk4poor
e4f51c654e fix(environments): scrub staged stdin when cancel lands during the upload
state["staged"] was set only after the upload returned. A kill() during
the upload therefore saw nothing to scrub, and exec_fn then hit the
cancelled gate and returned 130 without deleting. The staged file (which
can hold the sudo password line) stayed in the sandbox, whose filesystem
persists by default on both Daytona and Vercel.

Factor the scrub into a lock-held helper and call it from exec_fn's
cancelled branch as well as from cancel(). Also skip the upload entirely
when cancel already won before it started.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 23:44:55 +05:30
kshitijk4poor
bc815821f2 refactor(environments): share the staged-stdin path and redirect prefix
Daytona and Vercel built the staged stdin path and the
`exec 0< f || exit $?; rm -f -- f || exit $?` prefix byte for byte the
same. That prefix is the security handoff (the shell takes ownership of the
payload and unlinks it before the user command runs), so it should live in
one place: two small BaseEnvironment helpers next to _embed_stdin_heredoc.
Upload and cancel lifecycle stay per-backend.

Also document the "payload" _stdin_mode; the old comment still claimed
Modal/Daytona use heredoc, which no built-in backend does now.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 23:44:55 +05:30
kshitijk4poor
9c17837d41 fix(daytona): upload stdin bytes directly and call delete_file with its real signature
The pinned SDK (daytona 0.155.0) upload_file() accepts bytes, so the host
NamedTemporaryFile round-trip was redundant and briefly wrote the merged
stdin (which can start with the sudo password line) to the host disk.

delete_file() is delete_file(path, recursive=False); the extra
request_timeout kwarg raised TypeError inside contextlib.suppress, so the
pre-dispatch cancel scrub silently never ran and the staged payload stayed
in the sandbox /tmp.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 23:44:55 +05:30
kshitijk4poor
9c2b2c5aed fix(daytona): drop redundant staged-stdin deletes and restart path
The script already rm-s the staged file before running the user command,
so the success-path delete_file was a wasted round-trip, and post-cancel
deletes ran against a stopped sandbox. Only delete on kill() when the file
was uploaded but exec was never dispatched (checked under the env lock,
bounded request_timeout); exec_fn skips dispatch once cancelled.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 23:44:55 +05:30
kshitijk4poor
1a50db7e67 fix(vercel): scrub staged stdin only on a pre-dispatch cancel
#122218 stopped the sandbox on ANY exec exception (killing background
processes on a transient SDK error, even for stdin-less commands) and
re-stopped/rm-ed on every cancel, failing the existing cancel test.

Track staged/dispatched under the env lock: kill() overwrites the staged
file only if it was uploaded but never dispatched, then stops once as
before. After dispatch the user shell opens and unlinks it itself, so no
extra command runs. exec_fn skips dispatch if cancel already won.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 23:44:55 +05:30
JoaoMarcos44
2a0f0f4a66 fix(environments): preserve byte-exact SDK stdin
(cherry picked from commit 8dfa82862c5a86736b57756532b3b265e953cfe5)
2026-09-26 23:44:55 +05:30
JoaoMarcos44
79b3369333 fix(environments): keep SDK stdin out of command argv
(cherry picked from commit 8a60cc2dcfda619483de4ae9c13ec37affdbb528)
2026-09-26 23:44:55 +05:30
kshitijk4poor
641c49f841 fix(tools): resolve ssh paths against the remote home, keep guards intact
Follow-up to the ssh path fix salvaged from #121686.

- Read the ssh anchor raw (session record, override, TERMINAL_CWD): the
  shared workspace-root helper expands ~ on the Hermes host, so
  TERMINAL_CWD='~/proj' still resolved into the container home.
- Resolve ~ to the remote home the SSH environment detects at connect,
  bringing the environment up through the file tools' own creator
  (_get_file_ops, same cwd and cache) when none is live. A failed bring-up
  is remembered per container for 30s so one call's several resolutions
  don't each retry; a live environment is always used first. SSHEnvironment
  now records whether the home was detected, and a guessed /home/<user>
  (echo $HOME failed) is not used. Results are
  absolute and stable from the first call (read tracking and staleness
  checks key on them), and '..' normalizes to the real target: relative
  traversal like ../../../etc/x from ~ was refused on main and slipped past
  the sensitive-path guard on the PR head.
- If the remote home cannot be detected, an ssh ~-path that climbs above ~
  cannot be classified; the write guard refuses it.
- ~user passes through for the remote shell instead of becoming ~/~user.
- coerce_ssh_remote_cwd maps paths under the host subprocess home onto ~/,
  except when that home is the OS user's real home.
- The outside-workspace warning compares in the remote namespace (it fired
  on every correct relative write when the anchor was ~).
- The backend type is looked up once per resolution again (the PR head did three per local path).
2026-09-26 06:20:21 +05:30
John Paul Soliva
66186a8c04 fix(environments): deliver heredoc stdin byte-exact
A heredoc body always ends in a newline, so on the heredoc-stdin backends
(Modal, Daytona, Vercel) every stdin payload arrived with one extra byte.
write_file and patch verify an exact sha256 of the written file, so even
with the command grouped every write failed verification and left
content + "\n" on disk.

Feed the heredoc through a process substitution that re-emits the body
minus that last character. The heredoc stays outside <( ) because bash
3.2 mis-parses heredoc bodies inside it, and the reader tolerates read's
EOF status under an inherited set -e.

The tool-result spill no longer needs its +1 byte allowance for heredoc
mode, so the size check is exact on every backend again.
2026-09-25 14:14:06 -04:00
josh snider
a781ff257a fix(environments): route heredoc stdin to compound commands 2026-09-25 14:14:06 -04:00
brooklyn!
1ac24fa209 fix(gateway): do not let a launching profile own the host gateway
A profile-scoped parent donated its environ to the host multiplexer, so a
named launcher was treated as the primary adapter owner and its platform
token became the primary claim. Spawn the host with served_profile_child_env
for the default root, mark multiplex active before that primary load, and
name the env-derived side in a duplicate-credential refusal.
2026-09-25 12:01:44 -05:00
ethernet
49a49e8635 fix(environments): log activity-callback failures at debug
The heartbeat that reports long-running command progress swallowed any
activity-callback exception with a bare pass, so a broken callback left
no trace. Log it at debug with exc_info like the sibling best-effort
handlers in the same module; the command still never fails because of
a progress hiccup.
2026-09-24 11:50:30 -04:00
ethernet
6716b72ef2 refactor: separate checkpoint store maintenance from snapshots
Move retention, orphan pruning, status and clear operations into a topical sibling and route callers and tests directly to their defining module. Keep the shared store paths and git execution in checkpoint_manager; shorten the local child-env WHAT docstring.
2026-09-24 01:55:51 -04:00
ethernet
5b9b997c42 fix(tools): read snapshot store without check-then-read race 2026-09-24 01:39:08 -04:00
ethernet
038d7f797f fix(pm): avoid duplicate Daytona install and decode Store output explicitly 2026-09-24 01:36:14 -04:00
ethernet
d75d3fe15c Merge origin/main into ethie/pm-clean
Conflict resolutions and semantic fixups:

- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
  fence missed, deregister from _live_foreground in a finally) wrapped around
  pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
  the old early-recovery block stays gone; main's interrupted-pull restore
  (auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
  pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
  and is cleared once git is done. The marker's target is the ref git actually
  moves to (a release tag, not always origin/<branch>), since the restore
  compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
  had dropped and main's auto-merged restore needs (NameError on the first
  launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
  settings.about block; trim it to `updates` as pm-clean's type and the other
  overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
  names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
  key-leak switch leg needs the SDK, and the api_server two-tenant test needs
  aiohttp, both PM runtime extras the test env does not carry.
2026-09-23 21:55:59 -04:00
teknium1
3a001a4d7f fix(tui-gateway): the hard-exit kill owns every foreground spawn and never blocks
Re-review findings on the foreground-process registry:

- Spawn vs exit races. A child spawned but not yet registered when the hard-exit kill ran, or a
  command launched after the kill took its snapshot, survived under init. The hard-exit kill now
  raises a one-way exit fence and waits (bounded) for spawns already past it to register; a spawn
  refuses once the fence is up, and one that registers after a timed-out wait kills itself.
- The immediate kill fell back to proc.kill() for every handle. On Modal/Daytona/Vercel that is a
  blocking SDK cancel (an 8s cancel made _hard_exit take 8s). Popen handles are still killed
  inline (killpg, never blocks); every other handle's kill runs on a daemon thread under one
  shared 0.5s deadline.
- The registry lock was a plain Lock: a signal landing on a thread that held it deadlocked the
  exit. It is now an RLock (via a Condition) and the hard-exit path only takes it with a timeout,
  falling back to a lock-free copy.
- The kanban worker's SIGALRM deadman os._exit()ed without the kill; it now goes through it.
2026-09-23 17:09:54 -07:00
teknium1
17a137d8a3 fix(tui-gateway): a SIGTERM-ignoring command no longer survives the gateway's SIGTERM exit
The SIGTERM handler arms a 1s os._exit timer, then runs _shutdown_sessions: a flush of up to
5s, then _stop_turns_before_exit, whose kill was the graceful TERM, wait 1s, KILL. A command
that ignores SIGTERM was still alive when the timer fired, and os._exit left it reparented to
init (live: `trap '' TERM; sleep 3600` survived a SIGTERM to `python -m tui_gateway.entry`).

- kill_live_foreground_processes(now=True): SIGKILL each in-flight foreground tree at once,
  no TERM grace, no wait (BaseEnvironment._force_kill_process; LocalEnvironment kills the
  recorded process group, never our own).
- The grace timer's exit (entry._hard_exit) runs it before os._exit.
- _stop_turns_before_exit SIGKILLs whatever is still alive halfway through its settle budget
  (it ignored the interrupt's TERM), so the tool call still ends with a result the teardown
  persists instead of a dangling tool_call in state.db.
- The other hard exits that skip cleanup do the same before os._exit: the serve parent-death
  watchdog, the CLI exit watchdog, the kanban worker's SIGTERM path, and the messaging
  gateway's shutdown and loop-liveness watchdogs.
- Deflake test_shutdown_mid_tool_kills_the_command_and_keeps_its_result: the 0.5s settle
  budget was too tight under -n 40 (1 red in 9 runs); the join returns when the turn ends.
2026-09-23 17:09:54 -07:00
teknium1
537ce77f52 fix(tui-gateway): exiting mid-tool no longer orphans the foreground command's process tree
Foreground terminal commands run in their own session (start_new_session) so an
interrupt can kill the whole tree, which also puts them outside the host's process
group. When the tui_gateway left mid-command (client closed stdin, or SIGTERM) nothing
killed them: _shutdown_sessions closed the agents, the SIGTERM path hard-exits after a
1s grace, and the `bash -c` + child tree survived, reparented to init.

- tools/environments/base.py: execute() records every in-flight foreground command;
  kill_live_foreground_processes() kills their trees through the backend's own
  _kill_process (the same kill an interrupt uses).
- cleanup_all_environments() (the exit funnel of the CLI, one-shot, messaging gateway
  and terminal_tool's atexit, so `hermes serve` too) kills them first.
- tui_gateway _shutdown_sessions (EOF atexit + SIGTERM handler) and the serve
  SIGTERM/SIGINT exit-flush handler interrupt running turns, wait up to 0.5s for them
  to settle so the tool call ends with a result the final persist records (no dangling
  tool_call in state.db), then kill any foreground command still alive.
- ComputeHost.close() kills them too: every caller os._exit()s right after.

Covers the case of b9dac83d366c (Desktop quit: serve SIGTERM handler) on every host.
2026-09-23 17:09:54 -07:00
ethernet
16652eea18 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	gateway/config.py
#	gateway/config_loader.py
#	gateway/readiness.py
#	hermes_cli/managed_scope.py
#	hermes_cli/plugin_python_deps.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/update_cmd_maint.py
#	plugin-catalog/hindsight.yaml
#	plugins/plugin_loader.py
#	providers/__init__.py
#	scripts/run_tests.sh
#	tests/gateway/test_control_socket_windows_live.py
#	tests/gateway/test_gateway_streaming_nested_config.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_plan_reconciliation_windows_live.py
#	tests/hermes_cli/test_update_apply_shallow_count.py
#	tests/hermes_cli/test_update_concurrent_quarantine.py
#	tests/hermes_cli/test_update_shim_self_lock.py
#	tests/hermes_cli/test_verify_console_scripts.py
#	tests/tools/test_lazy_deps.py
#	tests/tui_gateway/test_subprocess_encoding.py
#	tools/lazy_deps.py
2026-09-23 15:26:34 -04:00
tancou
c7c9c18ccf fix(profiles): pin the launch home so a mirrored HERMES_HOME cannot flip routed-profile decisions
Symptom: a host that serves several profiles from one process and mirrors
the active turn's profile into `os.environ["HERMES_HOME"]` for legacy
readers (Hermes WebUI does this on every chat turn, next to the
context-local override) makes every launch-home decision see the served
profile as the launch profile. Two profiles that both configure `atlassian`
with different credentials share whichever MCP connection came first: a
READ_ONLY_MODE=false profile ends up calling a read-only server
(nesquena/hermes-webui#7721). The same misjudgement leaves the launch
residue in the served profile's child env, seeds the launch profile's
bridged allow-all grant into the served profile's secret scope, and lets
the served profile's `terminal.*` config bridge into the shared process env.

Cause: four launch-home checks compare the task's override with
`get_process_hermes_home()`, which reads `HERMES_HOME` live:
`agent.secret_scope.serves_routed_profile` (keys the MCP ledger via
`_mcp_registry_scope`, #108352 / #111481, and the check_fn cache, #111151),
`agent.secret_scope._is_process_home`, `tools.environments.local._is_routed_home`
and `hermes_cli.env_loader._process_hermes_home`. Under the mirror the two
sides are equal for every turn.

Change: `hermes_constants.pin_process_hermes_home(path | None)` lets the
host record the home it serves as its own; `get_routing_process_hermes_home()`
returns the pin when set, else `get_process_hermes_home()`; the four checks
compare against it. The pin is deliberately NOT folded into
`get_process_hermes_home()`: `get_hermes_home()` falls back to it for tasks
carrying no override (MCP loop, spawners), and the host's mirror exists
precisely so those readers see the served profile. Only "is this task
routed / is this the launch home" changes. Unpinned, behaviour is
byte-for-byte the old one; hosts that never mutate `HERMES_HOME` need not
call it. `activate_multi_profile_hosting()` is not the seam for this: it
flips `get_secret` fail-closed process-wide and freezes the launch env,
which an embedding host cannot adopt as a bug fix.

Tests (2 invariants, parametrized over the four checks plus the MCP ledger
key; red on main, green here): pinned + mirrored env -> the served home is
routed and the launch home is not, the MCP key is `(home_key, name)`,
`get_process_hermes_home()` still follows the env var; never pinned or
pinned-then-cleared -> old semantics, including "a mirrored env var IS the
launch home". `tests/conftest.py` resets the pin per test so the
module-global cannot leak between files.

Live repro (WebUI + a stdio FastMCP server named `atlassian` in two
profiles, one gated by READ_ONLY_MODE): base -> one ledger key
`'atlassian'`, the write profile lists only the read-only tools; fixed ->
`(<read_home_key>, 'atlassian')` and `(<write_home_key>, 'atlassian')`,
each profile lists its own tools.

Docs: `gateway/AGENTS.md` § Profile scope (one launch-home identity) and the
isolation table in `website/docs/user-guide/multi-profile-gateways.md`.
Also maps the author e-mail under contributors/emails/ (attribution check).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 08:09:47 -07:00
ethernet
f4a38b1ee8 Merge origin/main into ethie/pm-clean 2026-09-23 10:13:26 -04:00
teknium1
b88b3de6e3 fix(profile-scope): stamp profile_home only for FOREIGN homes; cache the gate-prefix scan
The profile_home stamp exists to let serves_routed_profile() see a foreign-home scope without the
HERMES_HOME override. Stamping the LAUNCH home from the own-home binders (launch_profile_runtime_scope,
_worker_profile_scope's launch arm, model_switch's launch arm) added nothing for production but, under
a test conftest where server._hermes_home differs from HERMES_HOME, flipped the launch profile to
'routed' and charged every turn a routed tool/plugin discovery. Own-home binders now leave the stamp
None; the foreign-home arms keep it.

is_profile_gate_env resolved the platform prefix set (Platform enum + bundled plugin scan + registry)
once per env key; strip_profile_gate_env walks the whole child env, so a spawn paid the scan
hundreds of times. The static half is lru_cached, the dynamic registry union is taken once per strip.
2026-09-23 06:46:02 -07:00