_run_foreground recorded outcome=timeout whenever a backend exception's
text contained "timeout" (an SSH connect timeout, a sandbox API timeout).
That command never reached an exit status, and the doc says timeout comes
only from Hermes' own deadline flag (hermes_timed_out), which
terminal_outcome() already reads. Owner ruling: never from exception text.
The exception path now records nothing; the model-facing result (exit 124)
is unchanged.
Probe (code-read finding; the new test drives the real terminal_tool with a
backend raising "... Connection timeout"):
before: terminal.outcome {backend: local, command_kind: shell_builtin, outcome: timeout} x1
after: no terminal.outcome row; tool result still exit_code 124
test_backend_exception_mentioning_timeout_is_not_a_terminal_timeout is RED on
base, green here.
- hermes.tool_unavailable.count: model called a shipped built-in (BUILTIN_TOOL_NAMES) not enabled in
the session; tool name + catalog provider/model. Unknown names stay in v4 unknown_tool quality.
- hermes.provider_setup.count: started/completed/abandoned/failed per provider and surface
(cli_setup, cli_model, tui, desktop, dashboard) with a closed failure_class. Pending marker per flow;
a dead or stale marker is reported abandoned at the next start (v4 process-marker pattern).
- hermes.feature_adoption.count: once per feature per install at first real use, derived from existing
counters in the subscriber (metric -> feature table) plus Bot Mode / Projects hooks; bucketed by the
owning profile's install age.
- hermes.feature_disabled.count: turning off a default-on toolset/skill/plugin/platform/setting (and the
re-enable back to default), diffed at the config write chokepoints; once per (kind,name,event)/day.
Surface comes from the user entry point; setup/migrations record nothing.
Five opt-in counters that tune the agent loop itself, recorded through
record_process_mark (disabled config => zero rows, zero files):
- hermes.file_edit.count: tool, mode, outcome, match_strategy. The fuzzy
matcher reports the strategy that landed (or no_match / ambiguous) into a
context-local probe opened only around the patch/write_file handlers, so we
learn which strategies earn their keep without counting internal callers.
- hermes.loop_guard.count: provider, model, signal, detector. Hooked where
the guardrail already warns/blocks/halts and at turn end for the iteration
budget; latched once per turn per (signal, detector).
- hermes.tool_recovery.count: provider, model, tool, next_tool,
next_outcome. One row per failed tool call, resolved against the model's
next round (same tool first) or the turn end (no_tool_call / gave_up).
- hermes.terminal.outcome.count: backend, command_kind, outcome. The kind is
a table lookup on the first program word, never the text. Hermes' own
deadline/interrupt now set hermes_timed_out / hermes_interrupted on the
backend result, so a command's own `exit 124` reads nonzero, not timeout.
Disjoint dims from hermes.execution_backend.count (served vs not).
- hermes.model_reply_issue.count: provider, model, issue. Refusal and
truncation only from the structured finish reason; empty/reasoning_only
from the normalized message; issue=none per response as the denominator
(model_route counts attempts and files empty replies as failures).
Background review / curator loops and _host_local / unmetered terminal
calls never count. Helpers live in shared_metrics_harness.py; call sites
are one or two lines each.
- memory tool: one row per operation (a batch counts each op); gate/validation
refusals are "rejected", store-level non-application "failed". Memory provider
tools are counted in MemoryManager.handle_tool_call with the op read from the
action arg or tool-name verb.
- curator: run_curator_review records one row per pass (dry run = skipped) with
the before/after diff bucketed; a scheduled pass whose claim is held by another
process counts as skipped. The home is captured before the review thread starts.
- delegate_task: _run_batch opens the call, each joined unit folds its results in,
and the last unit emits the call's single row, so group-split background calls
are not counted per unit.
- terminal / execute_code / browser: counted per call that reached the backend;
the backend is resolved from the owning profile at call time (terminal plan
env_type, code local/remote, browser CDP > Camofox > cloud provider > engine, or
"extension" when the extension controller served the call). Guard refusals and
Hermes' own _host_local commands are not counted.
Tests read rows back from the real store; each call-site group is red with its
source reverted.
Every path that adds an MCP server now emits one
hermes.extension.install.count event through record_extension_install:
- catalog installs (hermes mcp install, the picker, the dashboard route and
its background CLI action) via mcp_catalog.install_entry
- catalog installs from connector cards and the agent's catalog tool via
_CatalogBackend.install / start_install_oauth (OAuth: success at commit,
failed when the flow cannot start; an abandoned browser step is a cancel)
- custom servers from `hermes mcp add`, POST /api/mcp/servers and the
mcp.add RPC (source url for http, local for stdio, name None)
A reinstall or overwrite of an already-configured server is not counted,
so the metric measures new installs rather than config churn. Catalog
names pass through raw; the contract reports non-catalog names as custom.
* feat(vercel): start fresh sandboxes from a managed image instead of the deprecated runtime
Vercel deprecated Sandbox runtimes (node24/node22/python3.13) in Aug 2026 in favour of
images, and rejects runtime+image together and runtime with a snapshot source. New
terminal.vercel_image (default vercel/sandbox/universal:latest, Node 24 + Python 3.14)
picks the image for fresh sandboxes; a pinned terminal.vercel_runtime still works, wins
over the image and logs a deprecation warning; snapshot restores send neither.
Setup wizard prompts for the image, dashboard exposes both keys, status/config show the
effective choice, TERMINAL_VERCEL_IMAGE bridges config to the tool like its siblings.
* feat(config): migration 49 drops the seeded node24 Vercel runtime pin
Every pre-49 config.yaml carries terminal.vercel_runtime: node24 (the template default) and
the setup wizard mirrored it into .env as TERMINAL_VERCEL_RUNTIME. Both are the default
copied, not a choice, so the migration drops them and fresh sandboxes follow vercel_image;
node22 / python3.13 pins are the user's and survive. Persisted sandboxes are unaffected:
a snapshot restore never sends a runtime or an image.
* fix(modal): keep persistent-sandbox snapshots past the SDK's 30-day TTL
modal>=1.5 gives Sandbox.snapshot_filesystem() a default ttl of 30 days, so an idle
persistent Modal sandbox silently lost its filesystem and restarted from the base image.
Pass ttl=None (retain until deleted) and bump the modal extra from 1.3.4 (no ttl
parameter; legacy RPC) to 1.5.5 so the kwarg exists on every install.
* fix(modal): drop the dead modal.Mount credential-mount block
modal.Mount left the public API in modal 1.0, so _modal.Mount.from_local_file raised
AttributeError into the surrounding except on every sandbox start and the block never
mounted anything. The FileSyncManager created right after already uploads the same
credential, skills and cache files (iter_sync_files), so delete the duplicate; the test
fake stops exporting a Mount the real SDK does not have.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
The warm_command/release_command hook thread was a bare threading.Thread, so under
the multiplexer resolve_passthrough_value() saw no secret scope, raised
UnscopedSecretError, and the best-effort hook swallowed it at debug level: every
command provider with a non-empty env_passthrough silently never warmed or released.
Bind the thread with ctx_bound (copy_context().run), the same seam the keep-warm
timer already uses. Regression test: the hook thread resolves the bound profile's
passthrough value.
Docs (review minor): the multiplex isolation table now names per-profile slash-command
gating and the fail-closed empty-admin policy for a served profile with no cached config.
run_command_provider forwarded declared env_passthrough keys from os.environ,
which under the multiplexer is the LAUNCH profile's .env. A served profile's
command TTS/STT subprocess (voice-note transcription, speech, warm/release
hooks) got the launch profile's credential and never its own. Resolve each key
with resolve_passthrough_value, as the terminal and code-execution spawns do.
(cherry picked from commit 22a299d5f562e1d174fa5d37590cd46361a1383a)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
A script that calls hermes_tools.write_file/patch and never prints the
returned {"error": ...} got status=success with empty output, so the
model believed the write happened. In a cross-agent bench this cost 3 of
18 Opus runs 4-6 extra turns (write_file refused by the read-before-write
guard, then re-probing and re-emitting whole files). The result now
carries tool_errors for that cell, and the in-script write_file doc says
existing files must be read first.
`hermes logs --since` and `--level` passed every line that had no leading
timestamp. A traceback's frames are written without one, so an old error
printed its frames without the header that was filtered out, and
`--component` dropped the frames of a matching record. `hermes logs` now
reads each unstamped line as part of the record above it: the line gets
that record's verdict for every filter, in the tail read and in `-f`.
Lines before the first stamp in the read window have an unknown time and
level, so they are dropped when `--since` or `--level` is set.
mcp-stderr.log had no parseable stamp at all: the banner started with
`=====` and server output was copied raw, so `hermes logs mcp --since`
printed the whole file. The stderr tee already reads each server's
stderr in a thread, so it now writes one line at a time, each prefixed
with the asctime-shaped local stamp the Python logs use. The banner
starts with the same stamp. The stamp comes from new public
`timestamp()`/`stamp_line()` in hermes_cli/stderr_timestamp.py, which
stays stdlib-only. The desktop MCP log view accepts both banner shapes.
A test now requires a real writer's sample line for every LOG_FILES
entry to parse with `_parse_line_timestamp`. The docs no longer say the
`timezone` key changes log timestamps; log lines use the machine's
local time.
Trim the salvaged detector/delivery pair to the salvage bar and fix the
delivery half. The PR resolved the refresh selection as
platform_toolsets.<surface> (desktop/tui), a key no config path writes,
so _get_platform_tools fell back to the constructed hermes-<surface>
composite that resolves to 0 tools (a cold resume went 30 -> 1 tool).
Delivery now goes through the builder new desktop/TUI sessions use
(tui_gateway.server._load_enabled_toolsets(platform) +
_load_disabled_toolsets) as an explicit refresh_agent_mcp_tools
override, so the refreshed set equals a fresh session's on that surface;
only desktop/tui agents are long-lived, every other surface builds a
fresh agent per process/turn and is left alone. Drops the
toolsets.py/tui_gateway policy relocation and the raw-value fallback in
the fingerprint; keeps the detector (platform_toolsets +
agent.disabled_toolsets replace the dead tools.enabled_toolsets key)
and the tools re-pin after the refresh.
Tests: 2 invariants in tests/agent/test_bot_chat_toolset_refresh.py
(desktop -> builder consulted, disabled override passed, re-pinned;
cli -> untouched) and the existing fingerprint axis test now edits
the key `hermes tools enable/disable` writes.
The capability epoch watched tools.enabled_toolsets, a key no surface
writes; real hermes tools enable/disable edits (platform_toolsets.* plus
agent.disabled_toolsets) never flipped it. Even when stale, the refresh
rebuilt only the prompt and never tools[]. Watch the real keys and
rebuild + re-pin the tool snapshot on Bot Chat capability refresh.
Closes#124211
Review finding (major): ProcessRegistry.restore_completions() is once-per-process but read
async_delegation._db_path() (ContextVar-aware) under whatever scope the FIRST consumer ran in.
In the TUI gateway both first consumers (the session notification poller and the prompt_turn
drain) run inside _session_profile_runtime_scope(session), so under multi-profile
`hermes serve` the first session's profile ledger was replayed and the LAUNCH profile's
undelivered completions were never replayed for the life of the process (origin/main's
import-time replay always covered the launch profile).
Fix: restore_completions() clears the hermes-home override for the duration of the replay
(set_hermes_home_override(None) + reset), so the once-per-process replay always reads the
launch ledger whoever gets there first; the caller's scope is restored afterwards. Smaller
than a per-home restored-set plus a tui_gateway boot hook: it restores exactly main's
invariant with no new boot seam, and secondaries stay where they were (gateway
_restore_secondary_completion_ledgers). Docstring states the invariant.
Tests:
- tests/tools/test_process_registry_lazy_restore.py::test_first_drain_under_secondary_scope_replays_the_launch_ledger
(red on the PR head: replayed profiles/b/state.db; green now)
- tests/gateway/test_multiplex_unserved_shared_ingress.py::test_boot_replays_the_launch_ledger_before_secondaries_and_watchers
(minor: pins the gateway boot hook — launch-scope replay in _start_secondary_profiles, i.e.
before the secondary bind and before _async_delegation_watcher, the only production consumer
that reads completion_queue.get_nowait() without drain_notifications)
WHAT
- tools/process_registry.py: ProcessRegistry.__init__ no longer calls
restore_undelivered_completions(); a new once-per-process
restore_completions() does, invoked by the first consumer:
drain_notifications() (CLI process_loop / TUI prompt turn), the gateway
startup (run_startup._start_secondary_profiles, right before the secondary
ledgers are replayed) and the TUI session notification poller.
- tools/async_delegation.py: restore_undelivered_completions() returns 0 when
<HERMES_HOME>/state.db does not exist (a replay must not create or migrate
the ledger); _connect() creates the parent via mkdir_under_hermes_home()
instead of a bare mkdir so a late writer cannot resurrect a deleted
(tombstoned) or missing named profile.
- gateway/run_notifications.py: docstring no longer claims the import restores
the launch ledger.
WHY
`process_registry = ProcessRegistry()` runs at module import, and model_tools
imports it transitively (tools.registry -> close_terminal_tool), so
`python -c 'import model_tools'` (hermes doctor, any tool-registry consumer)
under a fresh/typo HERMES_HOME created the full profile skeleton + state.db,
recreated a tombstoned `profiles/.deleted/<name>` profile, and ran
reconcile_state_schema() against an existing store — bypassing the
assert_named_profile_home_live / mkdir_under_hermes_home guards (#97128,
#112592). Restoring on first consume keeps the replay for every process that
actually drains the queue while an import performs no state.db I/O.
Salvaged from #123348 (@webtecnica: lazy restore + missing-ledger guard) and
#123298 (@Wenfengcheng: mkdir_under_hermes_home in _connect), trimmed to the
minimal shape (no completion_queue property, no read-only probe).
Fixes#123265
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
(cherry picked from commit b77c4ff49062d2936271ded80821b48b0c11d368)
MAJOR — `_builtin_gateway_liveness` trusted a bare-epoch ticker heartbeat for
~200 s after its writer died, for every profile, and the multiplexer pid gate
had been dropped: a killed serve/Desktop ticker read as alive and
`hermes -p X cron run` queued a primary-routed run "for the gateway's next
tick" with no ticker present.
- `cron/jobs.py::record_ticker_heartbeat` stamps `<epoch> <pid>`;
`get_ticker_heartbeat_age` parses the first field (legacy bare stamps still
yield an age); new `ticker_heartbeat_writer_alive` requires the stamped pid
to be alive and treats a bare stamp as NOT proof by itself.
- The heartbeat-only rung is now `fresh AND (served-by-multiplexer OR writer
alive)`: the multiplexer record proves the host process for a served named
profile (and covers a stale-code multiplexer still writing bare stamps), so
only the in-process serve/Desktop ticker relies on the heartbeat, and then
only with a live writer. `cron status`'s in-process-ticker rung applies the
same rule.
MINOR (a) — `_hand_off_primary_routed_run` gated on `is True`; an unknown
(None) probe returns an error naming the uncertainty instead of "queued …
runs and delivers it".
MINOR (b) — `_run_claimed_job` wraps a shared-bot satellite's resolved map in
`SharedRouteAdapters(primary, _primary_profile_routes_for_current_home())`,
the same grant the ticker's `tick_adapters_for` makes, instead of handing it
the full primary adapter map.
Tests (each red on the previous head a18977090be):
- test_cron_satellite_diagnostics.py::test_in_process_ticker_heartbeat_counts_only_while_its_writer_lives
- test_cronjob_run_primary_routed.py::test_routed_run_without_a_serving_gateway_fails_before_the_turn[None]
- test_cronjob_run_immediate.py::test_execute_job_now_grants_a_shared_bot_satellite_only_its_routed_targets
A multiplexed satellite profile with no platforms.<p> credential of its own
posts through the primary's bot via a root gateway.profile_routes entry.
The delivery preflight lets such a job through (#97476) on the assumption
that the primary gateway's live adapters send it, which holds on a
scheduler tick but not for a manual run: `hermes -p <profile> cron run`
executed the whole agent turn in the CLI process, which has no sender for
the route, then failed delivery with "platform 'telegram' not
configured/enabled" and overwrote last_status with delivery_failed.
cronjob(action='run') now checks, before the in-process claim and before
the background dispatch, whether any delivery platform of a runnable job
is reachable only through the primary route (routed to this profile, no
connected credential here). If so:
- a gateway that serves the profile is live (or liveness is unknown):
queue the run with trigger_job so that gateway's ticker runs and
delivers it; job status is untouched and `cron run` prints "It will run
on the next scheduler tick";
- no gateway serves it: fail fast before the agent turn, like the
relay-fronted forward does when its api_server is unreachable.
Runs inside the gateway process (its live adapter delivers, #89302),
paused jobs (trigger_job would resume them), local delivery, and profiles
with their own credential keep the in-process path unchanged.
Fixes#120330
(cherry picked from commit 5a0db1c4e37de865f23034937b93d38fa87414f4)
Review feedback: the previous try/except restored runner.adapters on any
resolution error, which is the cross-profile misdelivery this fix
prevents. Let the error propagate to _run_claimed_job's handler, which
marks the run failed and surfaces the message. Runners without
_adapters_for_profile (shims, tests) still keep runner.adapters.
Adds a regression test where resolution raises: run not fired, marked
failed, error surfaced.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 183547ea6b3798874dc8a98f943ef98ee710911b)
A manual `cronjob(action="run")` fired from a secondary profile's agent
resolved delivery through `runner.adapters`, which is the default
profile's map. Under `gateway.multiplex_profiles` HERMES_HOME is
overridden to the owner profile mid-turn, so the result left through the
default profile's bot (a Telegram DM arrived from the wrong bot while the
run was correctly recorded on the secondary profile's cron).
Resolve the owner profile from HERMES_HOME and use
`runner._adapters_for_profile()`, the same fail-closed path notifications
and goal loops already use. Runners without that method (tests, older
shims) keep the previous behaviour.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 350cfc1436504b92d5f9e924f43f81b232beefd7)
_install_owner_secret_scope / _owner_secret_scope rebuilt every owner's mapping with
build_profile_secret_scope (.env + external sources only). For the LAUNCH profile under
multiplexing the caller's bound mapping is launch_secret_scope's (frozen launch env under its
files), so a credential injected only by systemd Environment= / `op run` / Compose vanished on
the rebuild, the remote header stayed the literal ${VAR} and the new fail-closed check parked a
server that worked on main. Route the rebuild through _owner_secret_mapping: launch_secret_scope
for the process home (same rule kanban_db_dispatch applies), build_profile_secret_scope for a
served profile.
Also: retry a not-fully-hydrated home's secret sources at most once per 30 s per home instead of
on every connect/reconnect (each retry is a helper subprocess); document the remote url/headers
${VAR} fail-closed error and the scoped A2A/Buzz gates in the MCP config reference and the
multiplexing guide.
Under multiplex a served secondary profile's remote MCP server whose config has a
`${VAR}` Authorization header was brought up at gateway boot with the placeholder
unresolved (HTTP 401) and retried that same rendering forever.
Mechanism: `_owner_scope_home()` trusted any caller-bound secret scope, and the
gateway's boot-time `_profile_runtime_scope` binding is a SNAPSHOT taken before the
profile's external secret source (`secrets.command`) may have answered. The run
task copies that context, and `_refresh_remote_config` re-read config.yaml on every
probe but rendered it under the same frozen mapping, so a parked server sent the
literal `Bearer ${VAR}` every ~5 min for the life of the process.
- `_owner_scope_home()` now returns the owner home (the scope's stamped home, else
the registry scope) even when a scope is bound: MCP rebuilds the owner's scope
fresh, which retries hydration (cached once it succeeds). Single-profile
processes (scope key None) are unchanged.
- `MCPServerTask.run()` binds that fresh owner scope around the rebuild-time config
refresh, so a reconnect heals once the source hydrates.
- `_require_rendered_remote` fails closed: a remote `url`/`headers` still carrying
a `${VAR}` after rendering raises naming the variable instead of sending it.
- Salvaged from #119097 (@JoaoMarcos44): the connect-time re-render, which also
covers a lazy server whose config was rendered at boot; its mechanism alone was a
no-op for the reported topology because the owner-scope install was skipped
whenever a (frozen) scope was already bound.
Row L1/L3, topology T2. Tests: A→B→A under set_multiplex_active(True) with two
temp homes, red on origin/main (literal placeholder sent), green on head.
Co-authored-by: JoaoMarcos44 <JoaoMarcos44@users.noreply.github.com>
Independent-review minors on #126157 (#123989 class):
- session.create is a plain @method, so tui_gateway/session_workdir.py::_register_session_cwd
wrote the cwd record under the RAW session key while the scoped turn read
profile:<p>:<key> and missed it until the first `cd`. The writer now binds the
session's own profile_home around register_task_env_overrides.
- tools/terminal_tool.py::_task_env_overrides stayed raw-keyed, so two profiles
registering the same task id (docker_image / cwd) still collided. Writer, clear
and both readers (_has_isolation_overrides, resolve_task_overrides) now use
_qualify_task_key; the isolation-override branch of _resolve_container_task_id
returns the qualified key so the env cache and the override record agree.
- gateway/platforms/api_server.py::_derive_chat_session_id prefixed the seed for
the NAMED LAUNCH profile addressed via /p/<launch>/, so prefixed and
un-prefixed requests on one profile derived different ids. The prefix is now
skipped when the routed name is the process's launch profile
(_names_launch_profile vs get_routing_process_hermes_home()).
Tests (red on 3720519199c, green here):
tests/tui_gateway/test_session_cwd_profile_key.py::test_secondary_profile_session_cwd_is_found_inside_its_scope
tests/gateway/test_api_server.py::TestDeriveChatSessionId::test_launch_profile_prefix_keeps_the_unprefixed_id
Fold the two salvaged fixes to the class and the salvage bar:
- tools/terminal_tool.py: `_qualify_task_key` reuses `_routed_home_task_key`
instead of a duplicate qualifier; branch 2 (session-isolated sandboxes) AND
branch 3's `session:<key>` (SSH-style backends) carry the routed profile, so a
session id two profiles share (header-less API fingerprint, a DM chat id served
by two bots) never resolves to one `_active_environments` slot. Cwd records
keep the symmetric qualification. No routed home → raw key, unchanged.
- gateway/platforms/api_server.py: `_derive_chat_session_id` namespaces only
named routed profiles, so default/standalone ids are byte-identical and live
Open WebUI conversations survive the upgrade.
- Tests trimmed to two invariants: A→B→A across two temp profile homes under
multiplex (distinct sandbox keys, alias follows the qualified parent, cwd
isolated per profile), and derived-id namespacing without moving default.
Both fail on 5912ed81ed, pass here.
A multiplexed host serves every profile in one process, and header-less API-server
clients derive their session id from a fingerprint of the system prompt + first
user message. Under per-session isolation (docker + container_persistent: false,
plus plugin backends declaring session_isolated_when_nonpersistent) branch 2 of
_resolve_container_task_id keyed the raw id, so two profiles whose conversations
open with the same text shared one _active_environments slot: the second profile's
turn attached to the first profile's sandbox, with the first profile's image,
mounts and exported environment. Verified live against real docker (both
directions on unmodified main; isolated after the fix).
Routed scopes (HERMES_HOME override: multiplex turns named AND default,
per-profile TUI/desktop RPC, cron ticks) now qualify the key with the profile
name or home realpath; CLI and single-profile gateways keep the historical raw
key. The session-cwd record/get/clear trio gets the same qualification so the
fingerprint collision cannot carry one profile's cd state into another's.
Regression tests in TestRoutedScopeQualification fail on pre-fix code (4/6) and
pass with the fix; persistent-Docker behavior is byte-identical (labels stay
profile:<name> / default).
MAJOR: `hermes -p X gateway restart --all` with no installed service re-entered
`run_gateway(replace=True)` IN the CLI process. `_home_env` swapped only HERMES_HOME
while os.environ already held X's dotenv (loaded at CLI startup) and gateway/run.py's
root .env load does not clear inherited keys, so the default gateway came up with X's
TELEGRAM_BOT_TOKEN / OPENAI_API_KEY. And any child spawned inside the swap (Windows
`start --all`, the launchd/detached fallback, migrate's secondary spawns) read the
swapped home as the launch profile, so `served_profile_child_env` neither stripped nor
scrubbed X's env.
- `_home_env` pins the process's launch home for the swap (`pin_process_hermes_home`),
so routed-home decisions inside it keep the real identity; `strip_launch_profile_env`
reads the residue list from that pinned home too.
- `_restart_all_as_host` spawns the root detached (`_spawn_detached_gateway`, scrubbed
base env + the root's own secret scope, `gateway run --replace`) when the CLI runs
under a named profile; the default profile keeps its foreground run. The Desktop
update hand-off's `gateway start --all` under a named active profile is still served.
- test: `test_named_profile_restart_all_with_host_down_restarts_the_default_root` now
asserts the child env carries none of the named profile's keys and `run_gateway` is
never entered (red on the PR head: "the root ran IN the named profile's process").
MINORS:
- (a) `gateway_migrate._service_op('enable')` raises on a non-zero `systemctl enable`;
a fresh apply treats it as a preflight refusal before the flag/manifest write, a
resume reports it and converges. New `test_migrate_refuses_when_the_survivor_cannot_
be_enabled` (red on head: exit 0 with the ⚠ line).
- (b) `systemd_install` already-current branch warns when `systemctl enable` fails.
- (c) `remove_system_systemd_unit` reports a failed `systemctl stop` instead of ✓.
* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends
The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.
docker/sandbox-desktop.Dockerfile is that base plus:
- the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
- the exact package set the Hermes -desktop image installs (TigerVNC,
Xfce components, dbus, xauth, fonts)
- Playwright's headed Chromium (same build as the -desktop image)
- cua-driver 0.28.2 from its pinned release tarball
No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.
docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.
* feat(docker): sandbox desktop base on python3.13-nodejs26
Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.
* feat(docker): bake agent-browser into hermes-sandbox:desktop
The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.
* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend
A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.
Now the desktop lives where the terminal lives:
- tools/environments/streams.py: one primitive per spawn-per-call backend, the
local argv prefix that runs its remainder inside the sandbox with stdio open
(`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
gateway). `auto` follows the backend; a sandbox that cannot host a screen
REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
there; the host driver's runtime contract is irrelevant then; check_fn is
true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
with the daemon, socket dir and profile inside the sandbox; screenshots are
fetched back so MEDIA: paths keep working; recycle closes the sandbox
daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
in a sandbox, a unix socket otherwise.
Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.
* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only
DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).
* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read
Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.
* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order
Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host
sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.
* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox
Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).
* chore: retrigger CI (zero-job dispatch failure)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore(config): template stamps v47 and shows the new default sandbox image
The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.
* feat(sandbox): the default image change is a decision, not a surprise
A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.
Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.
Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.
Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.
* fix(config): both plain defaults that preceded the desktop sandbox image are template copies
main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.
* ci(docker): build the sandbox image on release/dispatch, not every main push
Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.
* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta
The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.
* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)
* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)
* fix(bot_desktop): placement is the authority; sandbox screen survives restarts
Review findings on the sandbox-hosted Bot Screen, each reproduced live first.
Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.
Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.
SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.
CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.
pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.
Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.
Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.
Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.
* docs(bot-screen): no literal tmp path in the profile-location note
* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)
* fix(bot_desktop): adopting a screen the sandbox kept records the marker
Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.
Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).
* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)
* docs(bot-screen): what the Apptainer path inherits from the image and what it does not
* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
The default browser_exec tool ran the `browser-use` CLI from a PM side
environment (browser-use==0.13.10 in <home>/environments/browser-use),
provisioned by the installers and `hermes update`. Sealed Desktop payloads
skip that step, so the Desktop app never had it and silently fell back to
the built-in tools; the side env was also per-profile and 225 MB.
The CLI's execution path is only `browser_harness.run.main()`; the
browser-use agent framework (anthropic/openai/google-api pins, 93 MB of
googleapiclient) is never imported. browser-harness itself is 2.6 MB of
pure Python whose pins (Pillow 12.3.0, websockets 15.0.1) already match
Hermes's own, so it becomes a core dependency and runs on sys.executable:
- pyproject/uv.lock: browser-harness==0.1.13 (+ cdp-use, fetch-use).
- _find_cli() returns [sys.executable, -m, browser_harness.run]; the child
env points PYTHONPATH at the harness site dir (the Desktop store
interpreter boots without a venv and the harness daemon re-runs
sys.executable), replacing whatever the agent inherited.
- The side-env provisioning (install_cli, the update/installer step) goes.
Post-write .js/.ts lint on the local backend runs PM's node/npx instead of
the terminal shell's PATH (user dirs first). pm.activate always puts the
store dirs in front, also on drift, activating only packages whose whole
chain is installed. tui_gateway puts the store dirs ahead of ~/.local/bin.
Bare npx/npm/node/uv/uvx MCP commands resolve to PM's copies; the system
fallback tables are gone and an absolute command: stays the user's.
Hermes runs stdio MCP servers on its own packaged Node only. A server whose
native addon was compiled by the user's Node (typically through the ~/.npm/_npx
cache the user's npx shares with ours) dies at startup with NODE_MODULE_VERSION /
ERR_DLOPEN_FAILED; the SDK only saw "Connection closed", Hermes retried three
times and parked the server, and the cause lived only in logs/mcp-stderr.log.
The stdio child's stderr now goes through a pipe that still copies every byte
into the shared log but keeps the last 16 KB readable. When the session fails
and that tail shows an addon load failure, the error becomes a
NodeAbiMismatchError naming the addon, both ABI versions and the remedy with
Hermes's own paths: delete the npx cache entry (Hermes's npx reinstalls it) or
`PATH=<managed node dir>:"$PATH" <managed npm> rebuild <pkg> --prefix <root>`.
It is classified permanent, so the server parks at once and self-probes back
after the rebuild. The message reaches every MCP status surface through
_format_connect_error / str(exc): startup banner, `hermes mcp test`, the TUI
and Desktop probes and the dashboard.
The hosts E2E cell that asserted the user's Node wins now asserts the ruling:
the managed Node runs the server and the ABI failure surfaces the remedy.
Refs #124264
Add an optional per-job `interpreter` field so a cron Python `script` /
`monitor_script` can run under a user-managed venv instead of Hermes' own
Python, letting scripts import packages the Hermes runtime does not carry
(#8714). Nothing is installed, frozen, or restored automatically.
- cron/jobs.py: persist + normalize the field (absent => record unchanged;
empty string clears it on update).
- cron/scheduler_script.py: _resolve_cron_interpreter() validates the path
at run time (absolute/~ required, regular file, executable on POSIX);
_script_argv runs [interpreter, script] and skips the managed-store
bootstrap/PYTHONPATH overlays, which exist for Hermes' own venv.
Threaded through _run_job_script, the claim-heartbeat wrapper, the
pre-run prompt path and monitor scripts.
- hermes_cli: --interpreter on `cron create` / `cron edit`; shown in
details and `cron list`.
- tools/cronjob_tools.py: programmatic/CLI lane only, like model and
reasoning_effort — absent from the model-facing schema.
Shell scripts (.sh/.bash) still always run under bash. Revives #8741.
Ported onto current main from #70500 (the scheduler moved to
cron/scheduler_script.py and the CLI/tool became table-driven since the
PR's base).
Co-authored-by: MestreY0d4-Uninter <241404605+MestreY0d4-Uninter@users.noreply.github.com>
This reverts commit 79dbb1450e (#124792).
_resolve_stdio_command passes _prepend_path the directory the command
resolved into, not the managed Node directory. After 79dbb145 that
directory moved to PATH[0] even when it was already on PATH, so:
- the reported layout (system Node with npx first, managed dir behind)
was unchanged: which() finds the system npx, whose dir is already first;
- a command found later on PATH now shadowed every earlier entry for the
child's other bare lookups. With the store-first PATH pm.activate()
gives the Hermes process, a brew-installed MCP command hands its
children brew's node/python3/git instead of the pinned store copies.
Restore the prepend-only-when-absent behaviour.
is_todo_tool_name returns False for non-string names (a malformed list/dict
name used to raise TypeError where the old check returned False), and the
kept regression test imports tui_gateway.server at module level so it no
longer depends on another test importing it first. Docstrings updated.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
is_todo_tool_call lived in agent/tool_executor.py and went through
canonical_tool_name, which imports model_tools. TUI resume calls it from
_todo_state_from_history on the RPC path, so the first resume in a
gateway loaded ~405 modules (2-3s) synchronously. tui_gateway/server.py
and run_agent.py also imported agent.tool_executor at module level,
adding ~142 modules to every TUI/desktop launch and breaking run_agent's
lazy-forward rule.
The predicate now lives in tools/todo_tool.py, which both startup paths
already load. It matches TODO_TOOL_NAMES ({TODO_SCHEMA name} + the legacy
aliases) and imports the bridge parser only when a tool_call entry's
args mention "todo". model_tools._LEGACY_TOOL_ALIASES derives its todo
entry from TODO_LEGACY_ALIASES, so there is one source of truth ("todo"
is the only alias mapping to todo_list). The live tool.complete path in
tool_progress uses is_todo_tool_name and the hand-kept _TODO_TOOL_NAMES
tuple is gone. The server.py noqa import is replaced by a function-local
import next to MAX_TODO_RESULT_CHARS, so a pruned name can't be swallowed
by the broad except. run_agent imports lazily. The dead TypeError arm is
dropped, and field reads use message_sanitization._tc_field.
agent/tool_executor.py is back to its pre-stack state.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Since 4317ed0e71 the heartbeat schema allows 0 (disabled) and clamps
positive values to 60, but the validation error still said "min 60",
steering models away from the valid 0. Flagged on #119202.
The subagent timeout diagnostic and the session_search eval harness still
appended a bare "...[truncated]" marker, the imitable wording #121548
replaced everywhere else. Route both through agent.compression_marker.elide
so every elision in the tree mints the same counted, guard-matched marker.
Salvaged from #122392 (only the two call-site hunks; base already ships
the elide helpers the PR re-defined). Refs #121572.
Fresh fix for #119975 (PR #119993 was deleted; nothing to salvage).
Three paths collapsed into app_not_running with the sentence '<slug> is
not running. Start <slug>': (1) a server_json probe with the app running
AND its endpoint PRESENT — the exact case from the report, where the only
thing missing is Hermes' own MCP connection; (2) every non-server_json,
non-interactive_session liveness kind; (3) static and unregistered
liveness, which cannot observe the app at all yet still claimed it was
stopped. Telling the user to start an app that IS running is the wrong
instruction.
Add a hermes_not_connected LivenessState ('<app>'s MCP connection is
missing. Reconnect <app> in Hermes, then try again.'), map the
running+endpoint-present branch and the static/unknown fallback to it, and
wire it through the TUI gateway contract (PluginServerState) and the
Desktop Plugins tab (AgentPluginServerState, SERVER_TONE, serverStates
i18n in en/de/es/fr).
Also stop composing the sentence from the declaration's slug: describe()
takes an optional display_name, and _plugin_server_rows passes the
curated catalog title (fallback: the manifest name) so the Plugins tab
reads 'NVIDIA App' instead of a raw server slug.
Fixes#119975
An OpenAI-SDK-shaped transcription response can be a structured object whose
``text`` is None and whose ``error`` carries the provider's failure. Every
caller fell back to ``str(transcription)``, so the object repr —
``Transcription(text=None, logprobs=None, usage=None, error='Transcription failed')``
— was logged as a successful transcript and returned in the
``{"success": true, "transcript": ...}`` envelope. Desktop conversation mode
then injected that repr as the user's message instead of the audio.
``_extract_transcript_text`` now raises ``STTResponseError`` (a ``ValueError``)
for any structured response — SDK object or JSON dict — with a missing or
non-string ``text``: the provider's ``error`` when there is one, else
"Transcription response contained no text". ``_with_openai_client`` and
``_cloud_failure`` surface that message verbatim, so the openai, groq,
deepinfra, mistral, xAI and ElevenLabs paths all return their existing failure
envelope instead. ``_transcribe_groq`` uses the shared normalizer rather than
its own ``str(transcription)``. Plain strings, objects/dicts with a string
``text`` (including ``""``, so silence stays non-fatal) and unknown scalars are
unchanged; only the repr fallback for structured responses is gone. No desktop
change is required.
Fixes#78098
An MCP tool result reaches the model as the handler envelope
`{"result": <text>, ...}` (tools/mcp_tool_handlers.py::_render_call_tool_result) so
structured metadata survives inline delivery. When that string crossed the persistence
threshold it was written to $HERMES_HOME/cache/spillover verbatim, so a ~200 KB document
landed on ONE line with every newline escaped (`\n`), making the read_file offset/limit
pagination the <persisted-output> block recommends unusable.
maybe_persist_tool_result now unwraps that envelope before persisting: the spill file and
the preview carry the model-facing text with real newlines. The envelope is recognized by
SHAPE (a JSON object whose keys are a subset of {"result", "structuredContent", "_meta"}
with a non-empty string "result") rather than by tool name, so opaque JSON from any other
tool is still persisted verbatim -- and the aggregate path is covered too, since
enforce_turn_budget persists under __budget_enforcement__ where a `mcp__` prefix test would
miss exactly the results it has to fix. Sibling members (structuredContent/_meta) are
appended after the text in a delimited metadata block instead of being dropped: they are
the payloads _render_call_tool_result keeps for the model on purpose (#115430), and the
spill file is the only copy left once the envelope is replaced by the preview.
Fixes#90426
`_prepend_path` inserted the resolved command's directory only when it was
absent from the child's PATH. The Hermes installer appends its managed Node
dir to the user PATH, so for anyone with a system Node (<22.12) earlier on
PATH the check no-oped and the managed dir stayed behind it. npm lifecycle
children (`node install.js`) then resolved the older system Node and failed
with ERR_REQUIRE_ESM even though Hermes had provisioned a compatible runtime.
Strip every existing case/trailing-separator variant of the directory first,
then prepend it, so the canonical entry is the one that wins and PATH does
not grow duplicates.
Fixes#82309
A `terminal(background=true, heartbeat=N)` tick queued a notification every N seconds
whether or not the process had printed anything, and every queued event costs the owning
session a full model turn. On Desktop and the TUI that turn painted the wake as a user
bubble ("[Background process ... heartbeat #9 ... (no new output since the last
heartbeat)]") followed by the model's "Still running normally." — over and over, for a
process whose row on the status stack already said it was running — and while the wake
held the session's turn, the user's own prompt sat queued behind it.
- `ProcessRegistry._emit_heartbeat` skips a tick with no new output. The sequence counts
delivered beats only; the "(no new output)" placeholder in the formatter is gone.
- TUI/Desktop type heartbeat rows `display_kind: hidden` (the kind both clients and the
transcript preview already honour); the CLI paints a one-line receipt and persists the
row hidden, so reopening the session in Desktop shows only the agent's reply.
- Desktop hydration drops heartbeat rows persisted by older backends the same way.
- `display.background_process_notifications: off` is honored by the TUI/Desktop poller and
the CLI drain, not just the messaging gateway. `off` mutes process-driven wakes only:
a finished `delegate_task(background=true)` still lands.
Supersedes #123123 (cherry-picked; scoped so `off` keeps subagent results) and #119202
(cherry-picked; `heartbeat: 0` is schema-valid so models that materialize every field
stop tripping the foreground guard).
Windows has no POSIX parent-death supervisor/killpg safety net, so an
ungraceful exit of the hermes process left every stdio MCP child tree
(npx.cmd -> node.exe) running as orphans with ParentId=null, piling up
across session restarts.
- _run_stdio now attaches the process to a KILL_ON_JOB_CLOSE job object
before spawning stdio children (self-guarded no-op off Windows), so the
whole child tree dies with the parent at the kernel level.
- Windows reaps kill the process tree (direct child + descendants) in the
lifecycle orphan sweep and the spawn-ledger startup sweep, where there
is no pgid to group-kill.
Fixes#61059