One statement of the ownership rule (module doc) instead of three; the
lock-refusal docstring keeps only why --replace is not offered. The
record-less-holder note names the real cause (a holder this process cannot
interrogate), and the start-beside INFO line and the lock refusal take the
migrate command from MIGRATE_COMMAND / _migrate_command() like their
siblings.
decide() runs twice per start (the CLI guard in hermes_cli.gateway and again
from start_gateway -> _host_attach_or_none), and the lock-losing path in
_claim_host_gateway_role re-derives the same fact via _owner_is_standalone(),
so every boot of a generated unit beside another profile's standalone gateway
logged the "starting beside it" WARNING three times. decide() is a verdict
function; reporting belongs at the action site. Demote host_attach's copy to
INFO and keep run.py's lock-claim WARNING, which fires exactly once on both the
--replace and plain paths and carries the `gateway migrate --multiplex` hint.
The lifecycle test still asserts the converge hint is logged by decide(); it
now captures at INFO.
decide() grants REPLACE_HOST for an owner whose served set is unknown, but
_replace_target_belongs_to_other_profile can then only prove ownership from
THIS home's gateway.pid record (owner.serves() is False when served_known is
False, so the cross-home served-set shortcut never fires). "Whichever home
launched it" is therefore true only for the serves-this-profile leg; a
served-unknown owner launched from another home is refused (exit 1, one
supervisor retry). Say so at the host_attach module doc, the decide()
REPLACE_HOST comment and the _host_attach_or_none comment, and qualify the
`--replace` remedy in _unknown_served_message the same way.
_refuse_message is reached for an owner known NOT to serve this profile, so
its "take the host over: --replace" line pointed at a path that lands on the
same refusal (decide() only returns REPLACE_HOST for an owner that serves us
or has an unknown served set). Drop it in favour of a note, and interpolate
MIGRATE_COMMAND like the sibling refusal in run.py instead of hard-coding the
command. Text/comment changes only; no behaviour change.
The host_attach module doc, the REPLACE_HOST consumer comment in
_host_attach_or_none and the _refuse_second_host_gateway refusal still
promised the old semantics: --replace takes the host over from whichever
home launched it. Since decide() only returns REPLACE_HOST for an owner
that serves this profile (or has not published its served set), and the
lock claim no longer treats --replace as --force, that advice is untrue
on exactly the path where it was printed.
_refuse_second_host_gateway is reached after losing the host lock to a
live process -- including one with no readable record (publish_record
failed after the claim, or a different HOST_PROTOCOL_VERSION during a
rolling upgrade). Offering --replace there sends the operator around a
loop that ends in the same exit 75. The message now says to stop the
other gateway or use --force, and states that --replace does not skip
the lock check. Text/comment changes only; no behaviour change.
gateway.standalone wins over the parked marker: profile_lifecycle() returns
False for an opted-out profile, so `-p X gateway stop|start` keeps addressing
X's own gateway process and never writes gateway.parked, even while a stale
host record still lists X. Every installed-roster caller now threads both
kwargs (`include_standalone=True, include_parked=True`): the host-attach peer
walk and the standalone boot notice were reading the roster without parked
profiles.
Dashboard twin (#119886 class): `/api/gateway/stop?profile=X` on a served
profile no longer answers 409 — the spawned `hermes -p X gateway stop` parks
it; `/api/gateway/start` on a parked profile is allowed while a host
multiplexer is live (the child unparks it) and still refused when nothing can
serve it. Tests trimmed to the invariants: the marker-appearing case was a
subset of the boot-and-reconcile test.
host_attach.decide() returned REPLACE_HOST for any live host owner whenever
--replace was passed, before checking whether that owner serves this profile.
On a one-process-per-profile fleet the owner is another profile's standalone
gateway: _replace_target_belongs_to_other_profile correctly refuses to signal
it (fail closed), the gateway exits, and the supervisor restarts it into the
same refusal. hermes gateway install generates --replace for every unit, so
every profile but the one holding the host lock respawn-storms.
The non-replace path already handles this owner by starting beside it; only
--replace skipped that branch. --replace now targets the owner only when it
serves this profile, or when its served set is not known yet (the boot race,
where replacing keeps --replace's authority and the ownership guard still
decides). An owner known not to serve us takes the non-replace path.
Reproduced on a live macOS launchd fleet after updating to 0.21.4: the first
profile to restart claimed the host lock, and the default profile's unit then
looped on "Refusing --replace: PID <other profile's gateway> cannot be proven
to belong to this profile's gateway" until the respawn-storm breaker engaged.
A named profile that authors `gateway.standalone: true` in its own config.yaml
runs its own gateway again, the pre-multiplex topology, while the default
gateway keeps serving every other profile. Topology becomes something the
operator authors per profile instead of something the box infers from boot
state, which is what a fleet running per-profile gateways lost when
`gateway.multiplex_profiles: false` was retired.
Changed
- hermes_cli/profiles.py: `profile_is_standalone(home)` reads the profile's
own config.yaml (memo by file signature, tolerant of malformed yaml, always
False for the default profile with one warning). `profiles_to_serve()`
excludes standalone profiles; roster callers that mean "every installed
profile" (plugin deps, Windows update, launch policy, dashboard listing and
topology) pass `include_standalone=True`.
- gateway/host_attach.py: `standalone_attach_decision` starts a standalone
profile's gateway beside the host multiplexer once every live gateway
confirms it does not serve that profile; refuses with a rescan message while
one still does. Used by the initial attach check and the lock-losing race.
- hermes_cli/gateway_multiplex_mode.py: a standalone launcher never becomes
the host multiplexer (`STANDALONE_PROFILE_REASON`), including callers that
supply an explicit GatewayConfig.
- hermes_cli/gateway.py, web_server_gateway.py: `hermes -p X gateway
install/start/run` proceeds without --force for a standalone profile; the
refusal text for other profiles points at the opt-out; status shows
"standalone (gateway.standalone: true)" and the default lists skipped
profiles.
- gateway/run_profile_reconcile.py: the host does not re-adopt a profile whose
own gateway is live (removing the key while it runs no longer double-binds).
- hermes_cli/gateway_migrate.py: standalone profiles are neither blocker nor
fold target; the plan lists them as "standalone by config".
- gateway/run.py: one INFO line per standalone profile at host boot.
Tests: two-home E2E through real loaders and resolve_multiplex_mode, decide()
with fake host records for both arms, lock-losing branch, reconcile guard,
migrate plan, refusal predicate both ways, topology, memo and malformed-yaml
contracts. All red on base.
Two gaps left by the standalone-owner START (field report on #118097):
* decide() now logs ONE WARNING when it starts beside another profile's
standalone gateway, naming `hermes gateway migrate --multiplex`. The
legacy per-profile topology stays live, but the fleet should be able to
find out it is still on it from its own logs.
* A REFUSE on `gateway run` prints to stdout, which under launchd is the
unit's stdout file; the wrapper then maps 78 to 0 and the unit is
parked with nothing in that profile's gateway/errors log. Log the
verdict and the remedy at WARNING before exiting. The exit code is
unchanged: a multiplexing owner that excludes the profile is a
config-derived, permanent refusal, and 75 would make launchd relaunch
it every ThrottleInterval forever (#89477) — the smaller correct
change is to stop the refusal being silent, not to make it retry.
Review follow-ups on the standalone-owner fix:
- _request_serve_profile built the HostGateway twice three lines apart; the standalone
branch also copied the probe's stale profiles instead of the rescan answer's served set.
One construction from the answer, standalone= from its multiplex key.
- START no longer carries an owner nobody reads.
- Module doc said 'exactly one gateway per host' / 'Four outcomes' / 'never a second gateway',
all contradicted by the START-beside-standalone outcome; scoped to 'beside a multiplexer'
and pointed at #109417 for the forced migration that removes the standalone branch.
`gateway/host_attach.py::decide()` treated a host owner that answers `multiplex: False`
to the rescan — another profile's standalone gateway — the same as a multiplexer whose
roster excludes us, and issued the permanent REFUSE. On the supervised path that is
exit 78, which `hermes_cli/stderr_timestamp.py` maps to 0 so launchd's
`KeepAlive.SuccessfulExit=false` parks the unit. On a one-process-per-profile fleet
(the topology `multi-profile-gateways.md` documents and `multiplex_profiles: false`
promises to keep) every launchd gateway except the first to claim the host lock was
parked at boot, silently; which ones survived was a boot race against the owner's
record landing.
`request_serve_profile` now returns the owner flagged `standalone` instead of `None`
for a `multiplex: False` answer, and `decide()` turns that into START — this profile
runs its own gateway beside the owner, as before #118097; the host-lock claim still
logs the two-gateway topology and `gateway migrate --multiplex` stays the converge
path. REFUSE is unchanged for a multiplexing owner that excludes the profile.
A/B against a REAL owner process (real host lock, record and control socket answering
the rescan): base → `refuse`, exit 78 → launchd 0 (parked); fix → `start`. Negative
control (owner answers `multiplex: True`, roster excludes the profile): `refuse` on
both.
Field report: debug share 5024e996 (11 launchd profile gateways, 0.21.3 ea0c2b82 —
only 2 of 11 came back after `hermes update`), discussed on #118097.
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.
- ATTACH now requires a LIVE `identify` answer. The claim-time record is
published with NO served set (the runner settles multiplex a moment later),
and `host_gateway()` reports `served_known=False` when nothing answers. An
owner whose served set is unknown yields a TRANSIENT refusal, never an
attach: previously `default`'s claim published "default,other" before its
socket bound, `other`'s systemd unit read that as "I am served", exited 78,
and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
instead of forcing `multiplex=True`, so a standalone gateway stops claiming
the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
`--force` into `_host_attach_or_none`. Both previously exited in the guard
before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
config refusal every supervisor parks on; "someone else serves me right now"
is a runtime observation that ends when that process does. No unit files
change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
and re-enters with `replace=True`, so it can no longer attach to the corpse
it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
`st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
`gateway status`/doctor across N profiles pays one probe, not N.
Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:
- `gateway run` for a profile the host gateway already serves ATTACHES: print
its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
to re-scan `profiles/` (control socket) and attach once the answer includes
it. Refuse only when the host gateway cannot be made to serve it. Under a
service supervisor the attach exits 78 instead of 0 so a redundant unit is
parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
the owner's control socket, so it no longer depends on the claim ordering in
start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
process: they restart the host multiplexer and preserve its served set. A
secondary still running its own gateway is reported with the
`gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
argv: a host singleton runs bare/default argv and can never prove it serves
profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
multiplexer is whichever profile launched the one host process.
Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.