- memory tool: one row per operation (a batch counts each op); gate/validation
refusals are "rejected", store-level non-application "failed". Memory provider
tools are counted in MemoryManager.handle_tool_call with the op read from the
action arg or tool-name verb.
- curator: run_curator_review records one row per pass (dry run = skipped) with
the before/after diff bucketed; a scheduled pass whose claim is held by another
process counts as skipped. The home is captured before the review thread starts.
- delegate_task: _run_batch opens the call, each joined unit folds its results in,
and the last unit emits the call's single row, so group-split background calls
are not counted per unit.
- terminal / execute_code / browser: counted per call that reached the backend;
the backend is resolved from the owning profile at call time (terminal plan
env_type, code local/remote, browser CDP > Camofox > cloud provider > engine, or
"extension" when the extension controller served the call). Guard refusals and
Hermes' own _host_local commands are not counted.
Tests read rows back from the real store; each call-site group is red with its
source reverted.
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends
The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.
docker/sandbox-desktop.Dockerfile is that base plus:
- the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
- the exact package set the Hermes -desktop image installs (TigerVNC,
Xfce components, dbus, xauth, fonts)
- Playwright's headed Chromium (same build as the -desktop image)
- cua-driver 0.28.2 from its pinned release tarball
No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.
docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.
* feat(docker): sandbox desktop base on python3.13-nodejs26
Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.
* feat(docker): bake agent-browser into hermes-sandbox:desktop
The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.
* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend
A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.
Now the desktop lives where the terminal lives:
- tools/environments/streams.py: one primitive per spawn-per-call backend, the
local argv prefix that runs its remainder inside the sandbox with stdio open
(`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
gateway). `auto` follows the backend; a sandbox that cannot host a screen
REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
there; the host driver's runtime contract is irrelevant then; check_fn is
true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
with the daemon, socket dir and profile inside the sandbox; screenshots are
fetched back so MEDIA: paths keep working; recycle closes the sandbox
daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
in a sandbox, a unix socket otherwise.
Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.
* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only
DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).
* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read
Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.
* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order
Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host
sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.
* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox
Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).
* chore: retrigger CI (zero-job dispatch failure)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore(config): template stamps v47 and shows the new default sandbox image
The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.
* feat(sandbox): the default image change is a decision, not a surprise
A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.
Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.
Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.
Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.
* fix(config): both plain defaults that preceded the desktop sandbox image are template copies
main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.
* ci(docker): build the sandbox image on release/dispatch, not every main push
Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.
* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta
The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.
* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)
* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)
* fix(bot_desktop): placement is the authority; sandbox screen survives restarts
Review findings on the sandbox-hosted Bot Screen, each reproduced live first.
Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.
Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.
SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.
CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.
pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.
Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.
Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.
Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.
* docs(bot-screen): no literal tmp path in the profile-location note
* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)
* fix(bot_desktop): adopting a screen the sandbox kept records the marker
Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.
Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).
* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)
* docs(bot-screen): what the Apptainer path inherits from the image and what it does not
* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
- pm.environment: the `--no-install-package <name>` root read parses [project].name without
tomllib on the pre-3.11 bootstrap python (Docker arm64 stage_runtime). Both parsers proven
to agree on the real pyproject.
- windows-build-deps.ps1 / run_tests.sh: with DISTUTILS_USE_SDK, setuptools takes link.exe
from PATH; under a bash-hosted step Git for Windows' coreutils `link` shadowed MSVC's.
The MSVC linker directory now leads PATH (cl.exe was already found — this was the next
failure in the ruamel-yaml-clib build).
- tests/install/e2e-assets/source-build-env.ps1: clear the identity variables through the
env: drive — [Environment]::SetEnvironmentVariable(..., 'Process') on .NET/Unix does not
reach spawned children, so the stamp child still saw GITHUB_SHA. Red→green under nix pwsh.
- old-updater surface: warm_agent_browser_npx_cache is a permanent def in both facade and
sibling (the frozen surface names both); its compat pointer is retired, and the shims test's
__module__ check holds.
- tests re-seamed / de-faked: import guard tests use the probe_root fixture (CI has no
editable finder), the takeover child tree gets a hermes_constants stub, posix.sh hand-off
test pins HERMES_HOME (our script honours the ambient one), the two Windows-layout
PYTHONPATH tests are platforms("windows") (they fake `Lib/site-packages` on Linux; pm's
site_packages() is host-correct), setup-pin test picks Git bash over System32's WSL stub,
expose_cli's Windows test asserts the installer-convention convergence (the branch retired
"windows-installer-owned"), plugin-manifest satisfied-dep fixture uses a core dep (pyyaml is
gone), housekeeping test yaml imports go through hermes_yaml.
- apps/desktop main.ts: three imports restored in round 4 whose users main removed.
Merge fallout (my resolution errors, all caught by CI):
- hermes_cli/backup.py + gateway.py: `theirs` on those hunks re-imported clusters HEAD had
already moved to backup_restore.py / kept in the facade. backup.py loses the 349-line
duplicate (main's #110179 fix is ported into backup_restore._import_db_member); the
systemd service-unit cluster returns to gateway.py (PM's _prepare_service_launcher /
_pm_managed_node_dirs / _systemd_command have no home in main's extraction) with main's
utf-8-sig read. gateway_service_unit.py is dropped.
- gateway/run.py: main's plugin-update chore is not profile-scoped (the housekeeping
ordering test pins the scope/drain sequence).
- pyproject + 30 test files: `import yaml` -> `import hermes_yaml as yaml` (pm-clean has no
pyyaml); gateway/config._bundled_platform_manifest_name reads through hermes_yaml.
- tests re-seamed onto pm-clean's shape: residency admission (installed_engine),
supervisor child env (binary is a constructor argument), update import guard
(update_cmd_deps is gone; our probe already scrubs PYTHONPATH — both #115032 invariants
pass), shallow-count git responses (stash path asks `status --porcelain -z`); dropped
tests for retired code (_run_node_bootstrap/_ensure_tui_node, Windows resume demotion).
- tests/tools/test_local_env_blocklist.py: restore the two helpers the suite-reduction
commit dropped and the blocklist import.
Real fixes:
- pm: classify_uv_failure/ResolutionConflict move beside the uv runner (pm.environment,
stdlib-only). pm.workspace imports tomllib at module level and cannot load on the 3.10
bootstrap python that streams uv output in the Docker arm64 image.
- tools/browser_tool.warm_agent_browser_npx_cache: back as a permanent definition — it is on
the frozen old-updater surface, and the revert-scheduled compat pointer does not count.
- hermes_cli/memory_setup: the dashboard's pip row uses pm.environments.
running_from_selected_environment for installed vs restart_required.
- scripts/windows-build-deps.ps1: export DISTUTILS_USE_SDK/MSSdk so setuptools trusts the
primed MSVC environment instead of asking vswhere (`env -i` test runner on win32-arm64
compiling ruamel-yaml-clib); run_tests.sh forwards them.
- tests/pm/test_windows_build_deps.py: start the protocol test from a parent env without the
toolchain variables the runner job already exports.
- tests/conftest.py scrubs HERMES_BUNDLED_PLUGINS (Nix-wrapped hermes on the dev host);
tests/home_io_guard.py treats sys.path site-packages under the real home as the
interpreter's installation (PM-activated developer shell).
- tests-js: four `curly` lint errors from main's new scripts.
Production:
- agent/bedrock_adapter.py, agent/vertex_adapter.py: pm.ensure_import ran at
module import. In any process that imports these modules without a committed
PM selection (CI's build_environment test venv, a fresh checkout) that sync
rebuilt the dependency environment mid-process and replaced sys.path with a
generation missing the caller's own packages (anthropic, aiohttp vanished).
The extra is now ensured at first client build / credential request.
- plugins/platforms/matrix/adapter.py: a complete install needs no
ensure_and_bind round trip; only a partial one syncs.
- tools/browser_tool.py: drop the facade's duplicate warm_agent_browser_npx_cache
shim; the compat pointer already resolves to browser_tool_install.
Test harness:
- tests/home_io_guard.py: PATH-entry probes (shutil.which) and the running
interpreter's own installation (stdlib reads, realpath ancestry, fixture
symlinks into it) are not Hermes state; a patched Path.expanduser must not
crash the guard. run_tests.sh no longer filters PATH — the guard owns it.
- tests/tui_gateway/conftest.py: import hermes_bootstrap before any file opens
a MagicMock hermes_constants window (6 files exited the process at boot).
- tests/hermes_cli/conftest.py probe_root: scratch checkouts the import guard
probes need hermes_bootstrap.py (the launcher imports it).
- tests/pm/_fixtures.py stage_host_python: a copied relocatable python needs
its stdlib beside it (No module named 'encodings' on CI).
- tests/install/e2e-assets/smoke-env.mjs: dependency-free env shaping so the
source-build-env probe runs under bare node (main deleted the Playwright
entry it was imported through).
- adapt main's new tests to branch seams (model_metadata_http, launch
completion tail, CI toolchain exports uv after python, source_launch
hermes_cli stub, systemd_notify single marker).
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
Chrome puts its SingletonSocket under $TMPDIR and AF_UNIX paths cap at 104/108 bytes, so a
deep HERMES_HOME (profile homes, test homes) made the new scratch TMPDIR kill Chrome at
startup ("Socket path too long") — two browser test files went red on the branch and green
on main. hermes_constants.socket_safe_tmpdir() keeps the scratch root when it fits and falls
back to the OS root for sockets only; the browser env and the code kernel RPC socket use it.
check_no_tmp_literals walked git-ignored runner artifacts (test_durations.json) and failed on
whatever the last test run wrote; it now skips `git ls-files --others --ignored` paths.
run_tests_parallel created its per-file temp roots in the system temp dir and exported no
TMPDIR, so a full-suite run wrote gigabytes of fixtures to tmpfs (3,225 leftover roots, 9.9 GB,
were sitting in /tmp on the dev box). Roots now live under HERMES_HOME/cache/scratch/pytest and
the test process inherits TMPDIR=<root>, so the existing cleanup removes every temp file.
The seventeen remaining literals are container-side paths, AF_UNIX socket-path-limit
candidates on darwin, detection needles, guard regexes and guidance text that tells the
model to avoid /tmp. Each carries an inline `no-tmp: ok — <why>` so the reason lives
next to the line; the baseline keeps only a fenced tree listing where a marker would render.
_build_browser_env() is shared by the npx cache warmer (hermes update / doctor
--fix), the lazy Chromium auto-installer, the Lightpanda engine's env and
browser_use_cli. Hooking bot_desktop.auto_start there made every one of those
block up to 15 s spawning Xvnc+Xfce with browser.headed on, against
desktop_env()'s own "never starts anything" contract.
The hook now sits where a headed Chromium is actually spawned for a tool
action, mirroring computer_use dispatch: _spawn_and_collect (the first
agent-browser command forks the daemon; Lightpanda engine excluded), the Chrome
fallback from Lightpanda, and the real-profile Chrome launch. The regression
test asserts both halves: the env builder never starts the screen, the headed
Chromium spawn does, a headless or Lightpanda spawn does not.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
Conflicts (all keep-both): tools/browser_tool.py takes main's scope-bound
passthrough + loopback NO_PROXY and still routes the env through the Bot
Desktop's desktop_env; i18n gains main's sudoDesc/sudoCommandUnavailable
beside our sudoInstallDesc; test_profiles keeps both sides' new tests.
A declared `security.fake_ip_ranges` block is trusted like `allow_private_urls`
for whatever it names, and `_resolve_fake_ip_ranges` accepted the entries
verbatim. Declaring 10.0.0.0/8, 127.0.0.0/8, 100.64.0.0/10, fc00::/7, or a
catch-all 0.0.0.0/0 / ::/0 therefore made loopback, RFC 1918, CGNAT and ULA
hosts dialable at pre-flight and at TCP connect, while the docs promised those
classes "stay blocked". Only the link-local/metadata floor actually held.
A local proxy owns none of those classes, so an entry overlapping one of them
(or the unspecified address, which is how the catch-alls are caught) is now
dropped with a warning instead of widening the guard; the other declared
entries keep working. The docs sentence now says so.
The browser hybrid-routing oracle `_url_is_private` classified the declared
sentinel as private via `ipaddress.is_private` without consulting the
declaration, so with a cloud browser provider and
`browser.auto_local_for_private_urls` (default on) every URL on a fake-ip
host was routed to the local sidecar. The sentinel is not private for routing
either: the name is public and the cloud browser resolves it itself.
Two authority gaps in served_profile_child_env (#111617 review, andrexibiza P1 #1/#2,
kvnloo finding 1):
- The base was hermes_subprocess_env(inherit_credentials=True) = the launch environ's
provider credentials; strip_launch_profile_env only knows names with .env/source
provenance, so a key systemd/Compose/the shell injected into the launch process
survived into profile B's child whenever B did not define the same name. Now a ROUTED
target scrubs every Tier-1/Tier-2 credential from the base regardless of provenance
before B's own scope is overlaid (the child boundary gets get_secret's contract: a
scoped miss is no credential, never ambient fallback). The launch profile's own child
keeps its env. bot_relay's base=os.environ goes through the same scrub.
- strip_launch_profile_env / the scrub keyed on is_multiplex_active(); the Desktop and
dashboard backends serve ?profile=B by installing the HERMES_HOME override without
that flag, so B's slash worker / helper children kept A's .env and settings. The
authority test is now "is the target a routed home" (target != process home).
- _build_browser_env resolved the passthrough keys via get_secret, which falls through
to os.environ on a scoped miss while multiplexing is inactive: a routed B with no
Firecrawl key got A's. Under serves_routed_profile() the bound scope is the only source.
- served_profile_child_env(inherit_credentials=True) with no target and no scope bound
under multiplex minted with the launch credentials (key_cmd TTL refresh on a worker
thread); it now raises UnscopedSecretError like get_secret.
tests/tui_gateway/test_served_profile_child_env_authority.py: ambient-only A key + B
missing it (mux on), flag-off routed B (helper child + browser), real child observation.
3/3 red on base.
Move the loopback NO_PROXY merge from browser_tool into agent/proxy_bypass.py (the
module that already owns NO_PROXY semantics) and reuse no_proxy_entries() so comma-
and whitespace-separated operator values are both preserved. Add
loopback_connect_kwargs() and pass proxy=None on the two in-process websockets
dials to loopback CDP endpoints (browser_cdp_tool._cdp_call, BrowserSupervisor._run):
those never see the child env, so the env merge alone left them routed through a
macOS system proxy. Remote CDP URLs keep the default proxy behaviour.
Tests trimmed to two invariants: the built child env appends loopback to an
operator NO_PROXY in both casings, and only loopback URLs get proxy=None.
Sibling helper in tools/browser_use_cli (#110570) is redundant once the shared
env carries the entries.
websockets>=14 defaults to proxy=True and resolves proxies via
urllib.request.getproxies(), which reads the macOS/Windows system
proxy config even with no *_proxy env vars set. Local CDP endpoints
(ws://127.0.0.1:<port>/devtools/...) were therefore dialed through
the system proxy and the handshake failed with "did not receive a
valid HTTP response" (#110565).
Append 127.0.0.1/localhost/::1 to NO_PROXY/no_proxy (both casings)
in _build_browser_env so every browser subprocess (Browser Use CLI,
agent-browser, Chromium, Lightpanda) bypasses proxies for loopback.
Operator-provided NO_PROXY entries are preserved.
Fixes#110565
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:
- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.
tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.
Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
`browser_console(expression=...)` answers over the CDP supervisor's persistent
WebSocket and returns BEFORE `_run_browser_command`, which is where the Bot
Desktop lease fence lived. With a human holding the lease every other browser
command returned `human_has_control` while the one command that evaluates
arbitrary JS still read the page the human was typing into.
The fence is now ONE helper, `browser_tool_session.run_fenced(session_info, fn)`
(admit -> run -> epoch check), used by both the subprocess path and the eval
fast path, so a future third path cannot fork the policy again.
Test: tests/tools/test_bot_desktop_browser_fence.py — with a human lease and a
fake supervisor returning a value, browser_console must return
human_has_control and never evaluate the expression (red on bc36ddb5f9).
A bot running on a headless Linux gateway now gets its own desktop (TigerVNC
Xvnc + a minimal Xfce session, one per profile) that Hermes Desktop streams
live. The user can watch the bot work, take over to type a login / 2FA code /
CAPTCHA, and hand control back; the bot refuses every computer_use action
(screenshots included) while a human holds the screen, then resumes with the
session cookies the human just created.
Why this shape:
- The screen lives on the machine Hermes runs on, not in a vendor cloud browser,
so it works for any app the bot drives and keeps the session on the user's host.
- Xfce components are launched individually (xfwm4, xfce4-panel, xfdesktop,
xfsettingsd) under a private dbus session rather than xfce4-session/the
metapackage: no screensaver, power manager or polkit agent to lock or prompt
a headless desktop.
- Transport is raw RFB over a WebSocket beside /api/ws, authenticated with a
one-shot ticket minted through the already-authenticated RPC channel; noVNC
runs in the Electron renderer. The bridge parses the RFB client stream and
drops keyboard/pointer/clipboard (incl. QEMU Extended KeyEvent, which noVNC
switches to once Xvnc advertises it) from any viewer that does not hold the
lease, so viewOnly is enforced server-side, not by the client.
- One lease per profile (agent | human viewer) is the single truth for the RFB
bridge, the computer_use tool and the Desktop UI; taking control evicts other
viewers' input with close code 4000 control-taken.
- Auto-start happens only at the computer_use tool boundary (headless host,
packages present, bot_desktop.auto_start=true); env builders stay pure so
status probes and tests never spawn X servers. tests/tools/conftest.py pins
the binaries to "missing" for the same reason the browser-use fixture does.
Surfaces: Desktop (Bots → right-click → Open Screen; Take over / Hand back),
CLI (`hermes computer-use screen status|start|stop|install-deps`), tool
(`computer_use` actions request_handoff / wait_for_human), gateway RPCs
(display.status/start/stop/observe/lease.acquire/lease.release + display.lease
event), docs page user-guide/features/bot-screen.
Competing installers and checkout-local venv assumptions bypassed PM
selection, install consent, and generation lifetimes. Route consumers
through PM and installation-bound launchers. Refresh source launchers
before obsolete Python entries can be collected.
Remove Node, browser, and CUA acquisition engines, obsolete venv-holder
handling, detached sync, and unused PM APIs. Keep historical updater
exports inert and preserve external tool ownership and native integration.
Share product freshness and prepared inputs across builders. Align plugin
admission, Docker provisioning, setup instructions, and behavioral tests.
Verified targeted Python and JavaScript tests, desktop and web typechecks,
scoped lint, real product builds, and the Docker frontend smoke test.
The missed post-setup test cleanup is included and verified.
Native Windows/macOS execution, full Rust compilation, and the complete
repository suite remain unverified. Historical compatibility requirements
were preserved and extended, not fully rescanned.
Old updaters keep running after the checkout changes. Returning None
from their removed uv helpers enables a pip fallback against the new tree.
Keep the historical imports as inert shims and stop dependency entrypoints
with a relaunch message instead. Do not call PM or write recovery markers
from that mixed-version process.
Self-managed source launches use PM's successful input stamp to decide
when dependencies need a sync. Restart on the managed interpreter before
activating the selected generation. Preserve launcher forms and options,
and do not sync while a live updater owns the installation.
Targeted runtime batch: 225 passed, 8 platform skips. Real PM worker tests
build and publish disposable dependency generations, retain prior state on
failure, and exercise fresh-process relaunch before dependency activation.
The final launch guard test also passes. Full suite and native Windows
execution were not run locally.
Under `gateway.multiplex_profiles` one gateway process serves every profile
under ~/.hermes/profiles/NAME/; each routed turn runs with a context-local
HERMES_HOME override while `os.environ` still holds the DEFAULT profile's
values. Anything evaluated once at import, or memoised in a single unkeyed
module slot, therefore freezes the LAUNCH profile's value and leaks it into
every other profile's turns. This lands the tools-side half of that class:
- tools/process_registry.py, tools/environments/{modal,singularity}.py:
`_checkpoint_path()` / `_snapshot_store()` resolve `get_hermes_home()` at
call time (same seam as `tools/skills_tool._skills_dir`, so the existing
`monkeypatch.setattr(CHECKPOINT_PATH)` test sites keep working). Completes
the checkpoint_manager / sticker_cache half cherry-picked from #56315.
- plugins/platforms/feishu/feishu_comment_rules.py: `_MtimeCache` is now
path-keyed (accepts a Path or a zero-arg resolver, one (mtime, data) slot
per resolved path) with `invalidate()`; `_rules_file()` / `_pairing_file()`
resolve the routed profile's files. Proposed in #63962.
- tools/tool_output_limits.py, tools/browser_tool.py, tools/browser_camofox.py:
the process-lifetime config caches are dicts keyed by `hermes_home_key()`;
the `_X_resolved` flags and the lifecycle reset keep their shape.
tools/file_tools.py drops its private `file_read_max_chars` memo and reads
the already mtime+path-cached `load_config_readonly()`.
- hermes_time.py: `get_timezone_name()`; when `is_multiplex_active()` the
env `HERMES_TIMEZONE` (bridged from the default profile's config at gateway
startup) is ignored in favour of the routed profile's config.yaml. Both
sandbox TZ sites (code_execution_env/_tool) now use it.
- tools/cronjob_tools.py, tools/tts_tool.py, tools/skill_manager_tool.py:
the static schema text is profile-neutral and `dynamic_schema_overrides=`
rebuilds the `display_hermes_home()` / create-dir hint per
`get_definitions()`, so a routed profile's model sees its own paths.
Refs #95685.
Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
(cherry picked from commit 6d3fc6b07b3155c6196b1fd61a829283f1d7855c)
browser_navigate / browser_exec / web_extract refused any URL whose query carried a
credential-NAMED parameter (token, signature, access_token, ...) when the backend was
a cloud provider. That is exactly the shape of magic links, OAuth callbacks and signed
CDN assets, so on Browserbase/Browser Use the agent could not finish a sign-in flow or
open an X video asset ("Blocked: URL contains a credential-like query parameter").
The floor protected nothing: the cloud browser already sees every cookie and typed
password of the session, and with the credential vault it receives the real password at
fill time. Hermes' own secrets leaking into a URL stay blocked by the value-shaped
_PREFIX_RE check (_secret_url_error), which is backend-independent. IMDS and
private-address floors are unchanged.
33 read sites across 25 files used BOM-intolerant encoding='utf-8'.
Windows tooling BOMs files it touches; json.load on a BOM'd file fails
with 'Expecting value'. Reads now use utf-8-sig (writes unchanged).
check-windows-footguns.py --all: 33 → 0.
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.
Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.
Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).
hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.