docs: multi-profile gateway pages match the multiplex-only topology

Every line still promising that `gateway.multiplex_profiles: false` keeps
per-profile gateways, or documenting the deleted `migrate --standalone`
rollback, now describes the shipped behaviour: one gateway per host serves
every profile; `false` no-ops with a warning; `hermes update` folds a fleet
unless a real boundary (different UNIX user, HERMES_HOME outside profiles/)
holds; a blocked fleet keeps `--force` per profile; re-running
`migrate --multiplex` converges a half-migrated host; a manifest on disk is the
resume record, never a rollback. Adds the new "No new per-profile gateways"
section with the exact refusal `hermes -p <name> gateway install` prints.

Files: website/docs/user-guide/multi-profile-gateways.md,
website/docs/reference/cli-commands.md (migrate row),
website/docs/developer-guide/multiplexing-gateway.md (eager activation no
longer honours `false`), hermes_cli/AGENTS.md (migrate + refusal seams).

Completes #118273 (docs item). Part of #109417
This commit is contained in:
teknium1
2026-09-21 11:17:32 -07:00
committed by Teknium
parent d27ba15669
commit 91d1339a8d
4 changed files with 89 additions and 49 deletions

View File

@@ -216,7 +216,11 @@ no preflight blocker, migratable host → `True`; else `False` + a logged reason
through. CLI/dashboard readers use `default_gateway_multiplexes` (live `served_profiles` record, then
the explicit flag) — never the merged default, which would guess a verdict only the gateway makes.
Migration from per-profile gateways: `hermes_cli/gateway_migrate.py` (`hermes gateway migrate
--multiplex|--standalone`, table-driven `_PREFLIGHT_CHECKS`, manifest `<default>/gateway_migration.json`);
--multiplex`, the only mode — `--standalone` is deleted and a per-profile fleet is not a supported
target; table-driven `_PREFLIGHT_CHECKS`; manifest `<default>/gateway_migration.json` = UNFINISHED,
a re-run resumes from it; a named profile's `gateway install|start|run` refuse without `--force` via
`gateway.py::_named_profile_refused_under_multiplexer`, dashboard twin
`web_server_gateway.py::multiplexed_profile_refusal`);
`update_cmd_fleet._verify_fleet_after_update` calls `maybe_auto_migrate_after_update` on the success
path only; `gateway_migrate_guards.py` holds the auto-path-only refusals (table `_AUTO_MIGRATION_GUARDS`:
other service domain / UNIX user / HERMES_HOME outside `profiles/` — notices for the explicit command,

View File

@@ -66,8 +66,9 @@ is documented as a known limitation at the end of this document.
only source for launch keys with no `.env` to rebuild from (systemd
`Environment=`, `op run`, Compose) — a key injected or rotated after the
freeze is invisible for the process lifetime. A genuinely single-profile
host never activates, and neither does one that pinned
`gateway.multiplex_profiles: false`; an unreadable `profiles/` directory
host never activates; `gateway.multiplex_profiles: false` is retired and
deliberately NOT consulted here (honouring it would serve a second profile
with the launch profile's credentials); an unreadable `profiles/` directory
fails closed (it activates) and logs a WARNING.
- With the guard armed the **launch profile is a tenant too**: a body with no
routed profile binds `launch_profile_scope_if_multiplexed()` rather than

View File

@@ -327,7 +327,7 @@ Subcommands:
| `install` | Install as a systemd (Linux) or launchd (macOS) background service. |
| `uninstall` | Remove the installed service. |
| `setup` | Interactive messaging-platform setup. |
| `migrate` | Move per-profile standalone gateways onto one multiplexed default gateway (`--multiplex`, the default) or roll back from the recorded manifest (`--standalone`). Runs a preflight (duplicate bot tokens, secondary port-binders without a `/p/<profile>/` ingress) and changes nothing when blocked. Flags: `--dry-run`, `-y`/`--yes`. See [Migrating from per-profile gateways](../user-guide/multi-profile-gateways.md#migrating-from-per-profile-gateways). |
| `migrate` | Fold per-profile standalone gateways onto the one host gateway (`--multiplex`, the only mode — `hermes update` runs it automatically unless a real boundary blocks it). Re-running it converges a half-migrated host; a manifest on disk is the resume record, never a rollback (there is no `--standalone`). Runs a preflight (duplicate bot tokens, secondary port-binders without a `/p/<profile>/` ingress) and changes nothing when blocked. Flags: `--dry-run`, `-y`/`--yes`. See [Migrating from per-profile gateways](../user-guide/multi-profile-gateways.md#migrating-from-per-profile-gateways). |
| `migrate-legacy` | Remove legacy `hermes.service` units left over from pre-rename installs. Profile units (`hermes-gateway-<profile>.service`) and unrelated services are never touched. Flags: `--dry-run`, `-y`/`--yes`. |
| `enroll` | Experimental: enroll this gateway with a relay connector and save relay credentials for connector-backed platforms. See [Hermes Relay](../user-guide/messaging/relay.md). |

View File

@@ -100,14 +100,17 @@ s6 container or Windows Scheduled Tasks). Otherwise it comes up exactly as
before — serving the default profile only — and logs the blocker plus the
`hermes gateway migrate --multiplex` one-liner. Nothing is changed on disk.
An **explicit** value is never second-guessed:
An **explicit** `true` is never second-guessed:
- `gateway.multiplex_profiles: true` (what the migration writes) multiplexes
regardless of the preflight — you, or the migration, made the call.
- `gateway.multiplex_profiles: false` (what `--standalone` restores) keeps
per-profile gateways for good. When it's off, nothing on this page changes —
every behavior below is inert.
- `GATEWAY_MULTIPLEX_PROFILES` in the process environment overrides both.
- `gateway.multiplex_profiles: false` is **retired**. It used to keep
per-profile gateways for good; now it resolves exactly like an unset key and
the gateway logs a warning pointing at `hermes gateway migrate --multiplex`.
One gateway per host serves every profile — the only way to run a separate
per-profile gateway is `--force` (below).
- `GATEWAY_MULTIPLEX_PROFILES` in the process environment overrides the
unset-key decision the same way an explicit `true` does.
Other processes (`hermes -p <name> gateway start`, the dashboard, `hermes gateway
migrate`) never guess how an unset flag was settled: they read the running
@@ -121,19 +124,22 @@ only when no gateway runs.
- Many low-traffic profiles that don't each justify a full process.
- You want a single thing to start, monitor, and restart.
Stick with one-process-per-profile when you want hard process-level isolation
between profiles (separate memory footprints, independent crash domains, the
ability to restart one profile without touching the others).
One-process-per-profile is no longer a supported topology to *choose*: a named
profile's `gateway install` / `gateway start` refuses without `--force` (see
[No new per-profile gateways](#no-new-per-profile-gateways)). It survives only
where a real boundary blocks the fold — a fleet split across UNIX users, or a
`HERMES_HOME` outside `<default home>/profiles/` — and there every profile keeps
`--force` as its stated path.
### Pinning the flag
With the flag unset, the default gateway decides at each boot (above). To pin
it, set it on the profile whose gateway runs as the host process (usually the
**default profile**) and restart its gateway — `true` forces multiplexing even where the boot preflight would have
held back, `false` opts out durably:
**default profile**) and restart its gateway — `true` forces multiplexing even
where the boot preflight would have held back (`false` is retired and ignored):
```bash
hermes config set gateway.multiplex_profiles true # or false
hermes config set gateway.multiplex_profiles true
hermes gateway restart
```
@@ -154,10 +160,43 @@ keys** — credentials are never shared across profiles.
You do **not** run `hermes gateway start` for the secondary profiles — the
default gateway serves them. See the contract changes below.
### No new per-profile gateways
Because one host gateway serves every profile, a named profile never gets a
gateway of its own. `hermes -p coder gateway install` (or `start`, `run`, and
the service step of `hermes -p coder setup`) refuses with exit 78 whether or not
a host gateway is running right now:
```
❌ Profile 'coder' does not get a gateway of its own.
Exactly one gateway per host is the inbound process for every
profile. Starting a separate gateway for this profile would
double-bind its platforms (two pollers on one bot token, port
conflicts).
Install or start the host gateway from the default profile; it serves this one too:
hermes gateway install
Or fold an existing per-profile fleet onto one host gateway:
hermes gateway migrate --multiplex
A separate per-profile gateway (for a fleet split across UNIX users or a
HERMES_HOME outside profiles/) needs --force: hermes -p coder gateway install --force
```
When the host gateway is already running and serves the profile, the first
line reads `The host gateway already serves profile 'coder'.` with the owner's
PID and served set, and the pointer is `hermes -p default gateway restart`.
The dashboard's **Start** button for a named profile returns the same refusal.
`--force` is the one escape: it installs a real per-profile service, and that
service (its `ExecStart` carries no `--force`) keeps starting normally afterwards.
### What changes when multiplexing is on
Enabling the flag changes how a few things behave. All of these revert the
moment the flag is off.
Multiplexing changes how a few things behave. None of these apply to a profile
that runs a separate `--force` gateway on a blocked host.
#### 1. Secondary profiles must not start their own gateway
@@ -890,18 +929,22 @@ grep -H 'TELEGRAM_BOT_TOKEN\|DISCORD_BOT_TOKEN' \
## Migrating from per-profile gateways
If your profiles each run their own gateway today (one systemd unit or launchd
agent per profile), the default gateway's boot preflight keeps it standalone
(the unset default never double-binds a running fleet). Fold them into a single
multiplexed default gateway with one command — and roll back with another.
Standalone per-profile gateways remain fully supported; this is an optional
migration, not a removal.
agent per profile, from a release before multiplex-only), the default gateway's
boot preflight keeps it standalone until they are folded (the unset default
never double-binds a running fleet). `hermes update` folds them for you unless a
real boundary blocks it (below); the same fold is one command, and re-running it
on a half-migrated host (flag on, a unit left behind, a crash between the two)
finishes the job instead of reporting "already multiplexed":
```bash
hermes gateway migrate --multiplex --dry-run # print the plan and any blockers; changes nothing
hermes gateway migrate --multiplex # apply (asks for confirmation on a TTY; -y skips)
hermes gateway migrate --standalone # roll back to per-profile gateways
```
There is no `--standalone` reverse command: a per-profile fleet is not a
supported target. A blocked fleet keeps running as it is, and each profile
keeps `hermes -p <name> gateway install --force` as its stated path.
### What `hermes update` does
After a successful update, when the install has two or more profiles, at least
@@ -935,7 +978,8 @@ A standalone secondary behind any of these boundaries stops the automatic path:
In that case `hermes update` prints the boundary it found plus
`hermes gateway migrate --multiplex`, and changes nothing — no unit is removed
and `gateway.multiplex_profiles` stays off. Collapsing such a fleet replaces a
and the per-profile gateways keep running (`--force` remains their stated
path). Collapsing such a fleet replaces a
kernel-enforced boundary (file ownership, `User=`) with in-process isolation,
which is an operator's decision. The explicit command still makes it: the same
findings appear as **notices** in `hermes gateway migrate --multiplex --dry-run`
@@ -962,10 +1006,8 @@ behaviour described above.
The explicit command is different: `hermes gateway migrate --multiplex` with
two or more profiles and **no** standalone secondary gateway still applies the
one remaining step — it sets `gateway.multiplex_profiles: true`, (re)starts the
default gateway and writes the same rollback manifest (with an empty
`secondaries` list), so `--standalone` undoes it. You asked for multiplex; you
get multiplex.
one remaining step — it sets `gateway.multiplex_profiles: true` and (re)starts
the default gateway. You asked for multiplex; you get multiplex.
:::tip Clones do not carry channels
`hermes profile create --clone` leaves the source's bot tokens and allowlists
@@ -992,7 +1034,7 @@ Older clones that still carry them are flagged by `hermes profile list`.
| Blocker | Why | Fix |
|---|---|---|
| Two profiles configure the same platform credential (e.g. the same `TELEGRAM_BOT_TOKEN`) | Under one process a bot token can only be polled once; the multiplexer would park the duplicate and that profile's bot would go silent | Remove the token from the second profile, or keep it in `default` and route that profile's chats with [`profile_routes`](#routing-shared-bot-chats-to-profiles-profile_routes) |
| A secondary profile enables a port-binding platform that has **no** `/p/<profile>/` ingress on the default listener | The multiplexer skips that whole profile (see [rule 2](#2-http-inbound-platforms-are-reached-via-a-pprofile-url-prefix)) | Disable the platform in that profile (`platforms.<name>.enabled: false`), or keep the profile on a standalone gateway with `hermes -p <name> gateway start --force` |
| A secondary profile enables a port-binding platform that has **no** `/p/<profile>/` ingress on the default listener | The multiplexer skips that whole profile (see [rule 2](#2-http-inbound-platforms-are-reached-via-a-pprofile-url-prefix)) | Disable the platform in that profile (`platforms.<name>.enabled: false`), or keep the profile on a standalone gateway with `hermes -p <name> gateway install --force` (a served profile's `install`/`start` refuse without it) |
The credential check reuses the gateway's own conflict detection, so its verdict
matches what the multiplexer does at startup. Which port-binding platforms have
@@ -1022,18 +1064,9 @@ above). `hermes profile create` confirms this when the live multiplexer picked t
profile up; it prints the `hermes gateway restart` reminder only when it could not
reach the multiplexer (for example, a gateway started from an older build).
### Rollback
### Failure handling and resuming
```bash
hermes gateway migrate --standalone
```
reads `gateway_migration.json`, sets `gateway.multiplex_profiles` back to its
previous value, restarts the default gateway, and reinstalls/starts every
recorded per-profile service (a system unit comes back with the `User=` it had).
The manifest is removed once everything is back.
The forward migration is transactional in the same way. Failures it can see
The migration is transactional. Failures it can see
coming from the plan (a system unit that would have to run as root without a
recorded `User=`, a config file it cannot rewrite) are refused before any
per-profile gateway is stopped. Anything that fails after the manifest is
@@ -1043,14 +1076,14 @@ profile is left without a gateway. Should the process die anywhere in that
window, the next `hermes gateway migrate --multiplex` sees the flag on, the
manifest, and no live multiplexer serving the migrated profiles (an installed
but stopped default unit does not count) and resumes from the manifest instead
of reporting "already multiplexed".
If no manifest exists (you enabled multiplexing by hand), leave multiplex mode
with `hermes config set gateway.multiplex_profiles false && hermes gateway restart`
and reinstall the per-profile services you want.
of reporting "already multiplexed". A manifest on disk always means
*unfinished*: it is the resume record, not a rollback command — there is no
`--standalone` reverse, and the compensator above only ever runs inside a
single failed apply so that no profile is left without a gateway.
Not covered automatically: s6-supervised containers (set the flag on the
default profile and restart the container) and Windows Scheduled Tasks (set the
flag, stop the per-profile tasks, `hermes gateway restart`). The dashboard's
Not covered automatically: s6-supervised containers — they converge on the next
container start (the per-profile slots are registered down and the root gateway
multiplexes). Windows Scheduled Tasks are folded by the command. The dashboard's
System page offers the same migration as a button when the preflight finds an
eligible install.
@@ -1065,9 +1098,11 @@ hermes-gateways restart
```
Running gateways are restarted by the update itself; on an install that still
runs one gateway per profile, the update then offers the
runs one gateway per profile, the update then runs the
[migration to a single multiplexed gateway](#migrating-from-per-profile-gateways)
— automatically when nothing blocks it, otherwise as a warning with the fixes.
— automatically when nothing blocks it, otherwise as a warning naming the
boundary (different UNIX user, `HERMES_HOME` outside `profiles/`) and the
one-liner to run yourself.
User-modified skills are never overwritten.