500 Commits

Author SHA1 Message Date
MongLong0214
64804eea40 fix(gateway): stop the orphan reaper claiming another home's gateway
When the Desktop server starts, its lifespan reaps unsupervised gateway
orphans. The scan behind it (_scan_gateway_pids) counted a gateway as
"ours" whenever its argv named no profile or home, and the exclusion
spared only the service manager's own pid. So a Desktop server on a temp
or custom home, which a test suite starts, SIGTERMed the native home's
launchd-managed gateway.

A gateway whose argv names no home got that home from its environment
(launchd plist, systemd unit, a test's child), so argv proves nothing
about it. The scan now reads the home from argv first, then from the
process's readable environment, replaying the profile resolution the
target ran at startup (dashboard_procs._hermes_home_for_pid). When the
environment cannot be read, only the native default home keeps the old
bare-argv claim, so `hermes gateway stop` on the native home still
reaches its gateway; a custom or temp home skips it.
2026-09-27 11:48:22 -07:00
kshitijk4poor
09472a30a7 fix(gateway): trim reap-grace docs and dead test state (#122533)
Review cleanups on the orphan-reap startup grace:
- Drop the docstring claim that the Desktop boot sweep is "the one caller
  that races a launch": web_server._spawn_gateway_restart also reaps
  (grace-less) before its coalesce check, so the claim was wrong.
- Cut the 7-line lifespan comment to one line; the full reasoning lives in
  the _reap_unsupervised_gateway_orphans docstring, so the two can't drift.
- Drop the never-asserted seen["extra_exclude"] and the constant-only
  `_REAP_MIN_AGE_SECONDS > 0` assert; a grace-less revert already fails the
  `seen["min_age_s"] == _REAP_MIN_AGE_SECONDS` check.

Co-authored-by: Halldrix <halldrix@users.noreply.github.com>
2026-09-27 18:28:26 +05:30
kshitijk4poor
b7c5207b9e refactor(gateway): fold the reap-grace age probe into the filter
The standalone _gateway_process_age_s wrapper only re-wrapped
dashboard_procs._process_age_seconds in a try/except. Inline it as a local
fail-closed predicate (same shape as dashboard_procs._is_stale_orphan) so
the grace lives entirely inside the one reaper that uses it; an
undeterminable age still never widens the reap.

Co-authored-by: Halldrix <halldrix@users.noreply.github.com>
2026-09-27 18:28:26 +05:30
Halldrix
bbaf54584e fix(gateway): reuse the shared process-age probe in the orphan reap grace (#122533)
(cherry picked from commit 1661c66c0a9015ad085edb739b580e445725f804)
2026-09-27 18:28:26 +05:30
Halldrix
33a7ce0fa6 fix(gateway): spare a still-booting gateway from the Desktop orphan reap (#122533)
(cherry picked from commit b616766e30c830cf07096f02cedb1ca5098d3433)
2026-09-27 18:28:26 +05:30
teknium1
5a0225dfff fix(gateway): restart watcher names its checkout instead of inheriting the cwd
The detached watcher runs as `<python> -c <program>`, so `hermes_cli` resolved only
because update_completion happened to spawn it with cwd=<checkout>. The program now puts
the checkout on sys.path itself, and the bare-Python test runs it from an unrelated cwd
(red without the sys.path line).

Also drops the two unused re-export aliases in gateway.status (`_posix_is_zombie`,
`_pid_exists_win32_ctypes`): nothing imports them and neither is in the old-updater
compat surface.
2026-09-27 03:37:09 -07:00
teknium1
029445545c fix(gateway): keep the update restart watcher stdlib-only so it survives the bare store Python
After the package-manager handoff, hermes update finishes on the bare store
interpreter and spawns the detached restart watcher as sys.executable -c.
The watcher imported gateway.status (utils -> hermes_yaml -> ruamel) and
hermes_cli.config, died with ModuleNotFoundError before relaunching, and
left every manually started gateway down after the update (gated on #124649).

Move the stdlib liveness probe (zombie-aware POSIX kill(0), Windows
OpenProcess) into hermes_cli._subprocess_compat, have gateway.status's
fallback delegate to it, and import only stdlib-backed modules in the
watcher.

Fixes #124649.
2026-09-27 03:37:09 -07:00
kshitijk4poor
8afaab3703 fix(gateway): restart wait survives a non-finite drain; tighten its tests
Follow-up to the salvaged restart-wait commits:

- A drain or cron timeout of .inf now means "wait indefinitely" instead of
  an OverflowError from the integer stop envelope, which crashed
  `hermes gateway restart` and made `hermes update` silently fall back to
  its 45s floor. The fleet "draining (up to Ns)" lines and the drain
  progress report format the budget instead of int()-ing it, so an
  unbounded wait no longer crashes them either (it did on main too).
- cron_drain_timeout is required: a 0.0 default meant "cron opted out",
  the under-budget this fix exists to remove.
- Docstrings describe what the budget actually covers (PID exit, not
  replacement startup).
- Tests assert the observer outlasts after-turn + the supervisor stop
  envelope and that configured cron reaches the CLI wait, instead of
  re-deriving the formula; the negative wording assertion on the pending
  footer is dropped (change-detector).
2026-09-26 20:07:47 +05:30
Brian Le
5aad192995 fix(gateway): include cron drain in restart exit wait
(cherry picked from commit aefbb0741ed3d8491ede6e2d26689022c69a85bc)
2026-09-26 20:07:47 +05:30
Austin Pickett
fae9e5677a fix(update): stop reading gateway identity off the Windows restart watcher's argv (#107002) (#121635)
* fix(update): stop reading gateway identity off the restart watcher's argv (#107002)

The detached restart watcher is spawned as
`python -c <watcher source> <old_pid> <python> -m hermes_cli.main gateway run`.
Its trailing argv is the command it will spawn LATER, but the canonical matchers
read identity straight off the joined command line, so the watcher itself was
classified as a live `gateway run` process — the documented "never infer process
identity from argv substrings" bug class, on the exact surface `hermes update`
uses to verify a post-update relaunch.

Also budget the post-relaunch liveness poll against the watcher's own deadline:
the watcher respawns the gateway only after the PID it was handed exits, so a
30 s window can expire before the relaunch it verifies was scheduled to start.

* test(windows-live): run the gateway-ancestor harness parent from a script file

A `python -c <src>` parent is an interpreter running inline source and carries
no readable Hermes identity, so it is no longer a gateway to any classifier —
the harness's own comment already said a realistic gateway argv is not a -c blob.

* test(windows-live): share one sleeper SCRIPT across the live process-topology fixtures

Four live Windows E2E files stood processes up as `python -c "sleep" <hermes argv tail>`.
That shape no longer carries a readable Hermes identity, so the fixtures stopped standing
in for the gateways they simulate. One shared sleeper script replaces the -c spelling.

* fix(gateway): drop the duplicated _INLINE_SOURCE_FLAG_RE definition

The constant was emitted twice around command_line_runs_inline_source. Same
pattern both times, so behaviour is unchanged — but one definition is enough.

* test(stderr-timestamp): run the gateway-lookalike children from script files

Both lookalikes stood a gateway child up as `python -c <src> <gateway tail>`.
That shape no longer carries a readable Hermes identity (#107002), so the
wrapper correctly stopped treating them as gateway spawns and the tests failed.
A `-c` tail is data for a program the inline source may spawn LATER, never the
child's own identity — the real wrapper child is `python -m hermes_cli.main
gateway run`, which has no `-c`. Running the stand-ins from a real script file
restores what the tests mean to assert without depending on the misread.

* test(windows-live): restore the tempfile import dropped with the local sleeper helper

* test(windows-live): wait on the sleeper SCRIPT name, not its source text

The live fixtures proved argv visibility by waiting for `time.sleep(120)` in the
spawned process's command line. That string only ever appeared there because the
sleeper was spelled `python -c "import time; time.sleep(120)"`; now that it runs
from a file the source is in the file, so the probe timed out ("sleeper argv
never visible") even though the argv was perfectly visible.

Wait on the script name instead, exported as SLEEPER_MARKER next to the script
so the probe and the spelling cannot drift apart again.

* fix(tests,gateway): keep the live-system guard blocking -c-wrapped gateway spawns

The #107002 identity fix made _gateway_command_subcommand return None for
'python -c <src> … -m hermes_cli.main gateway run'. tests/_fixtures/live_system_guard.py
shares that matcher, so the autouse guard stopped blocking the detached restart
watcher: real gateways leaked out of the e2e run and squatted the webhook port.

Add gateway.status.gateway_spawn_intent_subcommand — the spawn-intent mirror of the
identity matcher, peeling the inline-source wrapper token-wise and re-running the same
canonical matcher on each suffix (still no substring matching) — and point the guard at
it. Read-only subcommands stay spawnable.

* fix(gateway): make the inline-source option walk value-aware so -X utf8 -c is not read as a gateway

The walk that decides whether a command line is an interpreter running inline
source (`python -c <src> ...`) treated every token starting with `-` as a flag
and the first non-flag token as the end of the option block. CPython options
that take a SEPARATE operand (-X/-W/-Q, --check-hash-based-pycs, --jit) break
that model: the operand was mistaken for the end of the block, so the walk
never reached the -c behind it and the watcher was read as a live gateway
again -- exactly the #107002 misclassification, one shape further out.

- Reuse the canonical operand sets from hermes_state_holders rather than
  hand-rolling a second copy (AGENTS.md: parser-derived flag sets).
- Walk case-preserving tokens: operand-taking -Q/-W/-X must not be conflated
  with operand-less -q/-b, so callers no longer lowercase before the walk.
- Handle clustered short options precisely (-uc is inline source, -Xc is -X c).
- Replace the ad-hoc _INLINE_SOURCE_FLAG_RE rescan in
  gateway_spawn_intent_subcommand with the index the same walk returns; the
  regex could not find spellings the walk accepts and would raise
  StopIteration.

Reported by an automated review on PR #121635 and reproduced here.

Refs #107002

---------

Co-authored-by: Austin Pickett <austinpickett@users.noreply.github.com>
2026-09-25 14:15:22 -04:00
Hermes Agent
d0cb567273 fix(gateway): take host identity from the live host record on replay
The settled-flag fix covers a process that multiplexes itself. The update
and fleet processes replay a FOREIGN gateway's captured argv with no
settled flag of their own, so a selector-less argv fell back to the
ambient HERMES_HOME comparison — the exact coordinate the review rejects
(#93943): a host launched from a named profile was replayed as that
profile, donating the named credentials to the respawned host.

Both restart edges now consult, in order: this process's settled
multiplex verdict, then the live host gateway's published rendezvous
record (its SETTLED served set, proven live), and only then the
compatibility default-root comparison. The raw config re-read stays last
so no settled identity exists => unchanged compatibility behavior.

Regressions: selector-less replay from a named home with a live host
record is host; without one it stays profile-scoped; the restart watcher
env takes the default root and drops the named token when only the host
record proves hostness.
2026-09-25 12:01:44 -05:00
Hermes Agent
8bde3a72e8 fix(gateway): preserve settled host identity across restart 2026-09-25 12:01:44 -05:00
brooklyn!
1ac24fa209 fix(gateway): do not let a launching profile own the host gateway
A profile-scoped parent donated its environ to the host multiplexer, so a
named launcher was treated as the primary adapter owner and its platform
token became the primary claim. Spawn the host with served_profile_child_env
for the default root, mark multiplex active before that primary load, and
name the env-derived side in a duplicate-credential refusal.
2026-09-25 12:01:44 -05:00
ethernet
746d861504 fix(install): don't ask the gateway install questions twice
The setup stage installs the gateway service through
ensure_gateway_service. On Windows that asks the start-now, Scheduled
Task and UAC questions. The gateway stage then ran `hermes gateway
install`, which asked them all again.

`gateway install --if-missing` does nothing when a service is already
installed. Both installers' gateway stages use it, so they ask only
when setup did not install the service.
2026-09-24 23:49:18 -04:00
ethernet
16652eea18 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	gateway/config.py
#	gateway/config_loader.py
#	gateway/readiness.py
#	hermes_cli/managed_scope.py
#	hermes_cli/plugin_python_deps.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/update_cmd_maint.py
#	plugin-catalog/hindsight.yaml
#	plugins/plugin_loader.py
#	providers/__init__.py
#	scripts/run_tests.sh
#	tests/gateway/test_control_socket_windows_live.py
#	tests/gateway/test_gateway_streaming_nested_config.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_plan_reconciliation_windows_live.py
#	tests/hermes_cli/test_update_apply_shallow_count.py
#	tests/hermes_cli/test_update_concurrent_quarantine.py
#	tests/hermes_cli/test_update_shim_self_lock.py
#	tests/hermes_cli/test_verify_console_scripts.py
#	tests/tools/test_lazy_deps.py
#	tests/tui_gateway/test_subprocess_encoding.py
#	tools/lazy_deps.py
2026-09-23 15:26:34 -04:00
Hermes Agent
b70ec27fa7 fix(gateway): multiplex convergence works inside an s6 container; a standalone-on-a-guard host says so loudly
Inside the official image every named profile has an s6 slot that the container's boot
registers DOWN (multiplex-only). `implicit_multiplex_blocker` still called
`_host_supports_migration`, whose s6 branch refused unconditionally ("Restart the container"),
so a Hermes Cloud host with the key UNSET booted standalone after an in-place update and every
other profile's bot went silent, while an explicit `true` bypassed the guard and worked. The
guard was vetoing a second gateway that could not exist.

- Only a named slot that is UP is a blocker; registered-down slots never veto the default.
  `hermes gateway migrate --multiplex` (and the `hermes update` hook) now fold an UP slot
  in-process: `s6-svc -d` + `down` file, root slot (re)started, no container restart. Boot
  and migration share ONE fold rule (`gateway_multiplex_s6.fold_named_slot_intent`).
- A multi-profile host that resolves standalone on a guard prints a boxed warning at gateway
  start, in the `hermes update` summary, in `hermes gateway status`, and the dashboard
  `/api/status` carries `multiplex_standalone_reason` with a banner. Single-profile installs
  are not warned.
- The resolved unset→on default is written as `gateway.multiplex_profiles: true` into the
  default profile's config.yaml (comment-preserving writer, once, never on a guard refusal).
2026-09-23 10:58:15 -07:00
Victor Kyriazakos
4c342c05de feat(gateway): stop, start and restart one profile under the host multiplexer
Under the host multiplexer every installed named profile came online because
its directory existed, and the only way to stop one profile's bots was to stop
the host, which stopped everyone's. Both are fleet-operator blockers.

`hermes -p X gateway stop` on a profile served by the host now parks it: it
writes `profiles/X/gateway.parked` first (so the 30s reconcile cannot re-add X
between the verb and the marker) and sends the new `unserve-profile` control
verb, which tears down X's adapters, reconnects and cron inside X's own scope
(the teardown `_unserve_profile` already used for deleted profiles). `start`
removes the marker and sends `serve-profile`, which runs the same add-path the
reconcile loop uses. `restart` cycles both without parking. The default
profile keeps today's whole-host meaning.

`profiles_to_serve()` skips parked profiles, so adapters, cron, ingress
membership and the served record follow from one chokepoint; roster callers
that mean "every installed profile" (plugin deps, Windows update, dashboard
listing and topology, migration inventory) pass `include_parked=True`.
Provisioning may pre-create the marker: an installed profile stays offline
until an operator starts it. Host boot logs one INFO per parked profile;
`gateway status` shows `parked (hermes -p X gateway start)`.

Tests: profiles_to_serve contract, control verbs through the runner
(round-trip and every refusal), CLI marker-before-socket ordering, reconcile
honouring the marker both ways, migration inventory retaining parked
profiles, two-home E2E through the real loaders.
2026-09-23 08:25:28 -07:00
ethernet
063560696e Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	tests/gateway/test_launchd_exit_timeout_drain_cap.py
2026-09-23 10:52:20 -04:00
John Paul Soliva
693e4aa12c fix(gateway): launchd install honours --no-start-now and the wizard's "start now" answer
On macOS, `hermes gateway install --no-start-now` started the gateway anyway.
`_cmd_install` forwarded the start flags to the systemd and Windows backends but
called `launchd_install(force)` alone, and `launchd_install` always ran
`launchctl bootstrap`. The plist sets RunAtLoad, so bootstrapping it starts the
gateway immediately, and the command still printed "Service installed and
loaded!". The setup wizard had the same gap: answering No to "Start the gateway
now?" and Yes to login auto-start still called `launchd_install(force=False)` and
started the gateway on the spot.

launchd_install now takes start_now. When it is False and launchd is not already
running the gateway, the install writes the plist and does not load it. It also
boots out any idle registration left from before, such as a job parked after a
clean exit, because `hermes gateway start` would kickstart that registration's
old definition instead of loading the new plist. The outdated-plist repair takes
the same path, since its bootout/bootstrap reload would start a stopped gateway.
A gateway that launchd already runs is reloaded as before, not stopped. With
the plist in ~/Library/LaunchAgents, the gateway starts at the next login or on
`hermes gateway start`. `_cmd_install` and the wizard now pass the answer
through.

Measured with launchctl recorded rather than run: `install --no-start-now` went
from 1 bootstrap to 0, the wizard's No went from 1 bootstrap to 0, and the
repair of an outdated plist with a stopped gateway went from a reload to a
rewrite only. The --no-start-on-login half on launchd is #91549 and is not
touched here.
2026-09-23 07:48:32 -07:00
ethernet
f4a38b1ee8 Merge origin/main into ethie/pm-clean 2026-09-23 10:13:26 -04:00
ethernet
cbb9660c2c Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/plugin_python_deps.py
2026-09-22 23:08:39 -04:00
teknium1
9fe737aef2 gateway: frame gateway.standalone as a TEMPORARY compatibility shim, not a kept topology
Multiplex-only is the direction. This key exists so fleets that lost per-profile gateways in
the switch keep working while the remaining gaps (per-profile stop/restart, WhatsApp
bridge/relay on secondaries, dashboard scoping) are closed, and it goes away once they are.
Every surface that names the key now says so through one shared notice
(STANDALONE_DEPRECATION_NOTICE): the install/start refusal hint, per-profile and default
`gateway status`, the migrate plan, and the host boot log (WARNING, not INFO). The user
guide carries a deprecation admonition where the key is introduced and no longer describes
it as opting out "for good".

Also: the Windows cold-start test fixture stubbed profiles_to_serve(multiplex) without the
new include_standalone kwarg (the PR's one red on CI).
2026-09-22 19:55:02 -07:00
Victor Kyriazakos
0238c9d740 feat(gateway): gateway.standalone opts a named profile out of the host multiplexer
A named profile that authors `gateway.standalone: true` in its own config.yaml
runs its own gateway again, the pre-multiplex topology, while the default
gateway keeps serving every other profile. Topology becomes something the
operator authors per profile instead of something the box infers from boot
state, which is what a fleet running per-profile gateways lost when
`gateway.multiplex_profiles: false` was retired.

Changed
- hermes_cli/profiles.py: `profile_is_standalone(home)` reads the profile's
  own config.yaml (memo by file signature, tolerant of malformed yaml, always
  False for the default profile with one warning). `profiles_to_serve()`
  excludes standalone profiles; roster callers that mean "every installed
  profile" (plugin deps, Windows update, launch policy, dashboard listing and
  topology) pass `include_standalone=True`.
- gateway/host_attach.py: `standalone_attach_decision` starts a standalone
  profile's gateway beside the host multiplexer once every live gateway
  confirms it does not serve that profile; refuses with a rescan message while
  one still does. Used by the initial attach check and the lock-losing race.
- hermes_cli/gateway_multiplex_mode.py: a standalone launcher never becomes
  the host multiplexer (`STANDALONE_PROFILE_REASON`), including callers that
  supply an explicit GatewayConfig.
- hermes_cli/gateway.py, web_server_gateway.py: `hermes -p X gateway
  install/start/run` proceeds without --force for a standalone profile; the
  refusal text for other profiles points at the opt-out; status shows
  "standalone (gateway.standalone: true)" and the default lists skipped
  profiles.
- gateway/run_profile_reconcile.py: the host does not re-adopt a profile whose
  own gateway is live (removing the key while it runs no longer double-binds).
- hermes_cli/gateway_migrate.py: standalone profiles are neither blocker nor
  fold target; the plan lists them as "standalone by config".
- gateway/run.py: one INFO line per standalone profile at host boot.

Tests: two-home E2E through real loaders and resolve_multiplex_mode, decide()
with fake host records for both arms, lock-losing branch, reconcile guard,
migrate plan, refusal predicate both ways, topology, memo and malformed-yaml
contracts. All red on base.
2026-09-22 19:55:02 -07:00
ethernet
c9bd7459c2 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	scripts/run_tests_parallel.py
#	tui_gateway/model_switch.py
2026-09-22 00:28:30 -04:00
teknium1
ed69aa9e6e refactor(doctor,gateway): share the migrate preflight's duplicate-credential check with doctor and status
Under multiplex-only the duplicated platform token is exactly what the migration
preflight refuses to fold on, so doctor and `gateway status` reuse that check
(`gateway_migrate.duplicate_credential_findings`) instead of a second scanner
in profile_channels. All three surfaces print one shared line naming the
profiles, the env key NAME (TELEGRAM_BOT_TOKEN, never a value or fingerprint),
why one token can serve only one gateway, and the remedy (own token / remove the
key from the non-owner / profile_routes) ending in `hermes gateway migrate
--multiplex`. Doctor folds the finding into its existing Profiles section.

Tests trimmed to two invariants: identical wording across doctor, status and
the preflight with the secret absent from output; distinct tokens and a shared
non-singleton key (model API key) produce no finding.
2026-09-21 18:53:36 -07:00
KoNit-K
54817f463b fix(gateway): detect cross-profile credential collisions 2026-09-21 18:53:36 -07:00
ethernet
4efb36f81e merge: integrate upstream desktop features through PM preparation
Merge origin/main at 8e806ae1b2. Keep native helper compilation in
prepareDesktopNativeDependencies and keep bundling/beforePack consume-only.
Bind helper sources and headers into preparation identities and cache keys;
copy admitted executable resources beside node_modules and preserve signing
semantics in product freshness checks.

Verified desktop typecheck, focused native/packaging/UI and gateway/cache
tests, and the real Linux preparation/copy/Xvfb execution path. Incoming
upstream anti-slop findings remain unchanged; no baseline was raised.
2026-09-21 15:17:44 -04:00
teknium1
d27ba15669 fix(gateway): a named profile cannot install a new standalone gateway without --force
`hermes -p <name> gateway install|start` (and `run`, `restart`, the setup flows
through `ensure_gateway_service`) refused only while a live multiplexer already
served the profile. On a host with no multiplexer running -- or one that had not
rescanned yet -- the same command wrote a brand-new per-profile unit: a fleet
member the multiplex-only topology (#118273) no longer supports.

`_named_profile_refused_under_multiplexer` now refuses every
`<root>/profiles/<name>` home that has no service of its own: the served form
keeps naming the owner PID and `gateway restart`; the new form points at
`hermes gateway install` (default profile) and `hermes gateway migrate
--multiplex`. `--force` stays the single escape (UNIX-user / out-of-tree
HERMES_HOME fleets), and a `--force`-installed service stays startable without
it so its supervisor (ExecStart carries no --force) keeps relaunching it. The
dashboard `/api/gateway/start` twin (`multiplexed_profile_refusal`) applies the
same rule and text; `stop` is unchanged.

Tests: red on base for the no-multiplexer refusal (exit 78, no unit written,
A->B->A over two profile homes, dashboard shape); controls for the default's own
install, the served-by-PID form, and the --force path; the #111958 setup test
now asserts the unserved named profile gets no unit either.

Completes #118273 (item B). Part of #109417
2026-09-21 11:40:48 -07:00
ethernet
4b9e6839d0 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-21 13:19:07 -04:00
teknium1
ed703493da gateway: converge on multiplex-only — real migrate, no opt-out, no rollback, host lock refuses
One gateway per host is the only supported topology (#100896). This is the convergence
path that makes it true on an upgraded machine.

`gateway migrate --multiplex` is now defined by TOPOLOGY, not by the config flag: a host is
converged when no secondary profile owns a gateway process or a supervisor unit. A
half-migrated host (flag flipped, a unit left behind, a crash between the two) therefore
converges on the next run instead of reporting "already multiplexed" — that flag-only test
made the re-run a no-op on exactly the host that needed it. The manifest is the resume
record: present means UNFINISHED, so a confirmed convergence clears it and any manifest
found on disk is resumed from rather than refused.

Windows Scheduled Tasks (and the Startup-folder fallback `gateway install` writes when it
cannot register a task) are now detected and removed like any other unit; Windows used to
be a flat refusal with hand-migration instructions. s6 stays refused, because there the
per-profile gateways are slots the container's own boot registers — and that boot now
registers named slots DOWN unconditionally, which is the s6 leg of the convergence. It
used to read `gateway.multiplex_profiles`, so the UNSET default (on) booted the slots
anyway and the image shipped the opt-out topology by accident.

A running gateway is never SIGTERMed silently: the plan names every pid it will signal and
prints before anything is signalled, with a dry run that changes nothing.

`--standalone` is gone. Reinstalling per-profile services is not a supported target, so
there is no rollback command; the machinery survives only as the compensator inside a
single failed apply, because the one outcome worse than a per-profile fleet is a profile
with no gateway at all.

`gateway.multiplex_profiles: false` is retired as a topology opt-out: it still parses and
still carries the runtime mode every scoped code path reads, but it can no longer pin a
second gateway process — it resolves like an unset key, warns, and points at the migration.
The unset path is not optimistic (it refuses to multiplex while a real blocker holds), so a
host that genuinely cannot fold still comes up standalone and says why. The eager Desktop
activation stops reading it too: a stale `false` there made a multi-home host serve a
second profile with the LAUNCH profile's credentials.

The host gateway lock flips from observe-only to a refusal. `host_attach` already
attaches/rescans/refuses before anything binds, but it reads a RECORD published a moment
after the owner starts, so two gateways launched together can both see no owner. The lock
is the only atomic arbiter of that race. The refusal names the owner, prints the migrate
command, and exits 75 (EX_TEMPFAIL) so a supervisor RETRIES — never the parking 78 — and on
the retry the record exists and the attach path resolves it properly. `--force`, `--replace`
and an unopenable lock dir are not refusals.
2026-09-21 10:11:19 -07:00
ethernet
f622ea76b3 merge: integrate origin/main while preserving PM ownership 2026-09-21 13:09:28 -04:00
teknium1
bb359ec5c1 fix(gateway): log the converge hint on a standalone owner; put a refusal in the profile's own log
Two gaps left by the standalone-owner START (field report on #118097):

* decide() now logs ONE WARNING when it starts beside another profile's
  standalone gateway, naming `hermes gateway migrate --multiplex`. The
  legacy per-profile topology stays live, but the fleet should be able to
  find out it is still on it from its own logs.

* A REFUSE on `gateway run` prints to stdout, which under launchd is the
  unit's stdout file; the wrapper then maps 78 to 0 and the unit is
  parked with nothing in that profile's gateway/errors log. Log the
  verdict and the remedy at WARNING before exiting. The exit code is
  unchanged: a multiplexing owner that excludes the profile is a
  config-derived, permanent refusal, and 75 would make launchd relaunch
  it every ThrottleInterval forever (#89477) — the smaller correct
  change is to stop the refusal being silent, not to make it retry.
2026-09-21 08:38:15 -07:00
ethernet
9d036f8ac9 merge origin/main (15 commits) into ethie/pm-clean
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
2026-09-21 08:16:57 -04:00
teknium1
ea0c2b820b fix(host-topology): a createTime-null record is not a live gateway, and the restart we print must run
Review findings on this PR.

- `_from_host_record` accepted a record with `createTime: null`: `_same_incarnation` treats a
  missing create_time as "matches", so `liveness_is_proven()` blessed whatever process happens to
  own that PID today. A record pointing at an unrelated `sleep 60` made `cron status` print
  "the host gateway (PID ...) serving profiles default, served". Require a recorded create_time.
- `cron status`'s host-record rung printed `hermes gateway restart`, which exits 78 for exactly
  its audience (a served NAMED profile); use the `--profile default` form both rungs now share.
- doctor's WORST state (supervision slots exist, nothing owns the gateway role) emitted a warning
  but appended no issue, while the strictly less severe LEGACY case did.
- Drop three defensive wrappers in `host_topology`: a try/except around an always-importable
  `normalize_profile_name` that silently degraded to a DIFFERENT normalisation, a blanket except
  per rung, and a swallowed `get_active_profile_name`.
- `get_systemd_unit_path`, `_legacy_unit_search_paths`, the doctor host-unit probe and the
  profile-delete service cleanup hardcoded `~/.config/systemd/user`, so on a host that moves
  $XDG_CONFIG_HOME the unit probe false-negatived -- the very bug this PR fixes survived there.
  One `user_systemd_unit_dir()` helper honours the spec default.

Tests: the two "Scheduler host: the host gateway" assertions were prefixes of BOTH rungs, so the
rungs were indistinguishable and the host-record rung was never exercised (which is how the
restart-command defect shipped green). Both now assert the full line, plus a served-host-record
case. Rendezvous isolation uses `tmp_path` instead of a leaked `tempfile.mkdtemp()` and also
covers `test_gateway_multiplex_served_record.py`, which read the real host record; the
`tests/conftest.py` hook in #118097 supersedes all of them once it lands.
2026-09-21 05:06:45 -07:00
teknium1
a10620a669 fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.

- ATTACH now requires a LIVE `identify` answer. The claim-time record is
  published with NO served set (the runner settles multiplex a moment later),
  and `host_gateway()` reports `served_known=False` when nothing answers. An
  owner whose served set is unknown yields a TRANSIENT refusal, never an
  attach: previously `default`'s claim published "default,other" before its
  socket bound, `other`'s systemd unit read that as "I am served", exited 78,
  and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
  instead of forcing `multiplex=True`, so a standalone gateway stops claiming
  the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
  `--force` into `_host_attach_or_none`. Both previously exited in the guard
  before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
  config refusal every supervisor parks on; "someone else serves me right now"
  is a runtime observation that ends when that process does. No unit files
  change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
  finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
  would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
  and re-enters with `replace=True`, so it can no longer attach to the corpse
  it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
  `st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
  it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
  generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
  `gateway status`/doctor across N profiles pays one probe, not N.

Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
2026-09-21 05:02:29 -07:00
teknium1
a20a88398f gateway: lifecycle verbs mean "the one host multiplexer"
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:

- `gateway run` for a profile the host gateway already serves ATTACHES: print
  its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
  to re-scan `profiles/` (control socket) and attach once the answer includes
  it. Refuse only when the host gateway cannot be made to serve it. Under a
  service supervisor the attach exits 78 instead of 0 so a redundant unit is
  parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
  landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
  the owner's control socket, so it no longer depends on the claim ordering in
  start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
  process: they restart the host multiplexer and preserve its served set. A
  secondary still running its own gateway is reported with the
  `gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
  argv: a host singleton runs bare/default argv and can never prove it serves
  profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
  multiplexer is whichever profile launched the one host process.

Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
2026-09-21 05:02:29 -07:00
ethernet
9ac6af3809 merge origin/main (28 commits) into ethie/pm-clean
main extracted the launchd backend and the setup wizard out of hermes_cli/gateway.py. The
eight launchd functions pm-clean had changed are ported into gateway_launchd.py in its
_gw() style: runtime_command/installation_command for the gateway argv (no VIRTUAL_ENV in
the plist, XML-escaped args), _prepare_service_launcher before every plist write,
utf-8-sig plist reads, and launchd_restart's refresh-first + bounded bootstrap revival.
The systemd service-unit cluster stays in the facade (no gateway_service_unit sibling).
backup.py takes main's browser_profiles backup-only exclusion on our profile_root_entry
shape. main.ts takes main's attach-first backend block with our explicit types.
2026-09-21 05:07:21 -04:00
teknium1
4a641d7e92 refactor(gateway): move the platform setup wizard out of hermes_cli/gateway.py
The `hermes gateway setup` wizard - the _PLATFORMS registry, platform status table,
per-platform setup prompts (standard/weixin/qqbot/signal), the service offer and the
wizard loop (33 names) - moves to hermes_cli/gateway_setup_wizard.py.

Same late-binding shape as gateway_launchd.py: bodies read facade helpers via _gw();
table values that bind at import (print_success/print_warning in _WEIXIN_DM_POLICIES)
import the neutral hermes_cli.setup module directly.

Part of #116913
2026-09-21 01:24:28 -07:00
teknium1
ae85aaa366 refactor(gateway): move the launchd backend out of hermes_cli/gateway.py
hermes_cli/gateway.py (6,511 lines) keeps the facade; the launchd (macOS LaunchAgent)
cluster - plist generation/refresh, launchctl bootstrap/kickstart, start/stop/restart/status
and the detached-process degrade path - moves to hermes_cli/gateway_launchd.py (41 names).

Moved bodies reach facade helpers through _gw() (late binding on hermes_cli.gateway) so
monkeypatches on the facade keep intercepting; the launchd-domain cache stays a facade
global and is rebound through the facade (a moved global would have rebound the sibling
module and left the facade reader None forever).

Part of #116913
2026-09-21 01:24:28 -07:00
ethernet
f3b1399211 fix: round-5 CI backlog after the 779-commit main merge
Merge fallout (my resolution errors, all caught by CI):
- hermes_cli/backup.py + gateway.py: `theirs` on those hunks re-imported clusters HEAD had
  already moved to backup_restore.py / kept in the facade. backup.py loses the 349-line
  duplicate (main's #110179 fix is ported into backup_restore._import_db_member); the
  systemd service-unit cluster returns to gateway.py (PM's _prepare_service_launcher /
  _pm_managed_node_dirs / _systemd_command have no home in main's extraction) with main's
  utf-8-sig read. gateway_service_unit.py is dropped.
- gateway/run.py: main's plugin-update chore is not profile-scoped (the housekeeping
  ordering test pins the scope/drain sequence).
- pyproject + 30 test files: `import yaml` -> `import hermes_yaml as yaml` (pm-clean has no
  pyyaml); gateway/config._bundled_platform_manifest_name reads through hermes_yaml.
- tests re-seamed onto pm-clean's shape: residency admission (installed_engine),
  supervisor child env (binary is a constructor argument), update import guard
  (update_cmd_deps is gone; our probe already scrubs PYTHONPATH — both #115032 invariants
  pass), shallow-count git responses (stash path asks `status --porcelain -z`); dropped
  tests for retired code (_run_node_bootstrap/_ensure_tui_node, Windows resume demotion).
- tests/tools/test_local_env_blocklist.py: restore the two helpers the suite-reduction
  commit dropped and the blocklist import.

Real fixes:
- pm: classify_uv_failure/ResolutionConflict move beside the uv runner (pm.environment,
  stdlib-only). pm.workspace imports tomllib at module level and cannot load on the 3.10
  bootstrap python that streams uv output in the Docker arm64 image.
- tools/browser_tool.warm_agent_browser_npx_cache: back as a permanent definition — it is on
  the frozen old-updater surface, and the revert-scheduled compat pointer does not count.
- hermes_cli/memory_setup: the dashboard's pip row uses pm.environments.
  running_from_selected_environment for installed vs restart_required.
- scripts/windows-build-deps.ps1: export DISTUTILS_USE_SDK/MSSdk so setuptools trusts the
  primed MSVC environment instead of asking vswhere (`env -i` test runner on win32-arm64
  compiling ruamel-yaml-clib); run_tests.sh forwards them.
- tests/pm/test_windows_build_deps.py: start the protocol test from a parent env without the
  toolchain variables the runner job already exports.
- tests/conftest.py scrubs HERMES_BUNDLED_PLUGINS (Nix-wrapped hermes on the dev host);
  tests/home_io_guard.py treats sys.path site-packages under the real home as the
  interpreter's installation (PM-activated developer shell).
- tests-js: four `curly` lint errors from main's new scripts.
2026-09-21 02:47:48 -04:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
teknium1
e7e1928174 fix(update): the invoking profile's launchd restart is claimed only on a pid that changed
`hermes update` on macOS printed "✓ Restarted ai.hermes.gateway" for the
profile running the update whenever launchd reported *any* pid for the label —
including the pre-update process the restart was supposed to replace.
f29ee96dd3 closed this for SIBLING labels only
(`_wait_for_launchd_service_pid(label, old_pid=old_pid, ...)`); the invoking
profile kept `wait_for_launchd_gateway_supervision` →
`_launchctl_label_supervising_process`, a predicate that is true for "launchctl
list exits 0 and some positive pid" and never compares against the pid seen
before the restart.

- `_launchctl_supervised_pid()` exposes the pid `_launchctl_label_supervising_process`
  already parsed; the boolean is now a thin wrapper over it.
- `wait_for_launchd_gateway_supervision(..., old_pid=...)` accepts a fresh pid
  only, mirroring the sibling loop's contract. `old_pid=None` keeps the old
  "any supervised pid" meaning for callers with no pre-restart observation.
- `_restart_launchd_gateway_after_update()` snapshots the pid before
  `launchd_restart()` and passes it in; the failure line now says launchd is
  not supervising a NEW process. The snapshot is verification-only, so the
  no-`launchctl list`-gating invariant of #74973 is untouched.

Also from round-2 review of this PR:
- the supervised-serve discharge test parametrized 5 supervisors, 3 of which
  the inventory writer can never put on a serve/dashboard row; cut to the two
  it can emit and documented the parity-only members.
- `_marker_only_restart_obsolete` now names who still owns a discharged
  launchd-supervised serve (update-time `report_unaccounted_runtimes`, exit 1).
2026-09-20 20:13:57 -07:00
teknium1
70fceb80ab refactor(gateway): move the systemd service-unit cluster out of hermes_cli/gateway.py
Mechanical extraction of the service-definition cluster (generate_systemd_unit,
systemd_unit_is_current, refresh_systemd_unit_if_needed and their
_systemd_*/_service_*/_ld_library_path/_temp_home helpers) into
hermes_cli/gateway_service_unit.py. Behaviour-neutral: moved bodies read facade
helpers (get_python_path, _profile_arg, _run_systemctl, load_gateway_config, ...)
and each other through _gw() — a late import of hermes_cli.gateway — so every
seam tests and callers patch on the facade keeps intercepting the moved code;
the facade keeps one import block re-binding the names.

hermes_cli/gateway.py: 6808 -> 6496 lines.

Part of #116913
2026-09-20 18:33:07 -07:00
joaomarcos
7dcb667716 fix(windows): ask gateway setup install questions once and stop after UAC hand-off
The setup wizard asked start-now/start-on-login, then called the Windows
installer without forwarding the answers, so the installer asked the same
two questions again. After a UAC hand-off (nothing registered yet in the
parent), the wizard then started the service, which re-entered the install
flow and re-offered the UAC prompt while the elevated child was still
waiting on consent.

Forward both answers into the Windows installer and let the installer own
the start decision: it starts the gateway itself on both the Scheduled
Task and Startup-folder paths, and reports a UAC hand-off as False so the
wizard parent stops instead of starting an unregistered service. Fixes #116550.
2026-09-20 10:58:14 -07:00
teknium1
9cb5c8ae09 fix: HERMES_HOME argv values containing a space still identify their own gateway (review follow-up)
hermes_home_assignments parses a space-joined command line, so an unquoted
`HERMES_HOME=C:\Users\John Doe\.hermes` token was cut to `c:/users/john` and the
default-profile predicate (assignments non-empty, own home absent) judged the
gateway NOT to belong to its own home -- a regression from the substring test
this PR replaced. Add command_line_names_hermes_home: the token-bounded parse
first, then a token-bounded literal match of the whole home, used by both
gateway.status and the hermes_cli PID scan so the two predicates stay mirrors.
2026-09-20 10:27:31 -07:00
liuhao1024
deda2c9dc3 fix(gateway): accept trailing-separator HERMES_HOME spellings, bound the assignment name
Review follow-up (#115039): a supervisor that exports HERMES_HOME with a
trailing separator (systemd Environment=, sh -c wrappers) put a spelling on
the command line the token-bounded comparison rejected, so gateway status
reported that gateway not-running for its own home. Strip one trailing
separator on both the extracted values and the profile home, keeping the
longer-sibling rejection (ops/2 is still a different home).

Also bound the assignment NAME: FOO=hermes_home=/x embeds the name inside
another token and is not a HERMES_HOME assignment (the value side already
was token-bounded; the substring test this replaces matched it).

Drive _scan_gateway_pids in a test so the mirrored predicates in
gateway.status and hermes_cli.gateway cannot drift apart again.
2026-09-20 10:27:31 -07:00
liuhao1024
5cca9b250c fix(gateway): profile identity predicate must not accept a longer sibling HERMES_HOME
The named-profile branch of _command_line_belongs_to_profile tested
`hermes_home={home}` as a substring of the command line, so a process
declaring HERMES_HOME=/root/profiles/ops2 satisfied the ops profile's
predicate (the -p side already uses token equality for exactly this
reason). The default-home branch and the CLI mirror
hermes_cli.gateway._matches_current_profile had the same hole.

Extract every HERMES_HOME=<value> assignment token-bounded with quotes
stripped (hermes_home_assignments) and compare full values in all three
places. Quoted assignments (ps/wmic re-quoting paths with spaces) now
match too.
2026-09-20 10:27:31 -07:00
fangliquan
b2269603a9 fix(gateway): skip unreadable Windows SCM services 2026-09-20 10:22:27 -07:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00
Kyzcreig
aa0289f307 fix(gateway): cap signal-driven stop drain to launchd's live ExitTimeOut
launchd's per-user (gui) domain clamps ExitTimeOut at 60s regardless of the
plist value (measured on macOS 26.6.1: 215->60, 90->60, 60->60, 30->30). A
gateway configured with restart_drain_timeout above that drains past the
budget and is SIGKILLed mid-teardown (adapters still connected, SQLite
checkpoint/close never runs). That unclean exit is one half of the
recurring state.db "structural corruption" class.

- gateway/restart.py: parse_launchd_exit_timeout / read_launchd_exit_timeout_s
  (launchctl print gui/<uid>/<label>, bounded, fail-open None when not
  launchd-owned or launchctl fails), resolve_launchd_capped_drain
  (min(drain, exit_timeout - 10s cleanup reserve); never extends a drain),
  effective_stop_drain_timeout (duck-typed for shutdown-path test doubles).
- gateway/run.py: read the live budget once at boot and WARN when the
  configured drain exceeds it; the SIGTERM/SIGINT handler marks the stop as
  signal-driven.
- gateway/run_shutdown.py: _stop_impl caps the drain, the watchdog leash and
  the cron drain ceiling to the launchd budget for those stops only. In-band
  SIGUSR1 restarts and --replace takeovers are not launchd-timed and keep the
  configured drain. systemd/s6/foreground: unchanged (no budget -> no cap).
- hermes_cli/gateway.py: plist template ExitTimeOut 25 -> 60 so LaunchAgents
  ask for the full clamp instead of a quarter of it.
2026-09-20 18:49:23 +05:30