Commit Graph

14 Commits

Author SHA1 Message Date
teknium1
3ed40556ce fix(profiles): a child spawned for another profile no longer inherits the spawner's authorization gates
A `hermes -p B` child built from a process that loaded profile A's env (a gateway, the
dashboard, the post-update fleet restart) started with A's `DISCORD_ALLOWED_CHANNELS`,
`TELEGRAM_GROUP_ALLOWED_CHATS`, `GATEWAY_ALLOW_ALL_USERS`... and enforced them as its
own: gates are not credentials (no secret scrub sees them), a unit-file `Environment=`
or operator export is in no dotenv (no name-list strip sees them), and B's own `.env`
rarely defines the key (its dotenv load never overwrites the inherited value). Observed
as profile B's gateway rejecting every message in B's own channel after a per-profile
restart issued from A (#113270).

- `local_env_policy.is_profile_gate_env` / `strip_profile_gate_env`: gates matched by
  shape (`_ALLOWED_`, `_ALLOW_ALL_`, `_ALLOW_FROM`, `_ALLOW_BOTS`, `_IGNORED_CHANNELS`,
  ...), never `HERMES_*`, so a gate added to any adapter is covered without a second edit.
- `strip_launch_profile_env` drops them on its existing routed-home branch: the seam
  `served_profile_child_env`, the kanban dispatcher, cron workers and the dashboard
  action env already funnel through (the dashboard site now calls it for its target).
  Same-home children keep an operator export.
- `update_restart_recovery._child_environment(profile)`: the one site that bypassed
  every helper (bare `os.environ.copy()` relaunching EVERY profile) strips gates when
  the profile is not the one the updater runs as; the module stays stdlib-only at
  import time.

Live: fresh-process `hermes_cli.update_restart_recovery --stdin` with three gates in
the updater env — base hands all three to profile b's relaunch, fixed hands none and
keeps them for the launch profile.

Refs #113270; supersedes #113308 (@yashraj4, static key list + always-strip; this keeps
same-profile children intact and covers the per-adapter gate set).
2026-09-18 15:11:47 -07:00
teknium1
08bb2273e4 fix: make the allow_all_users env bridge own what it writes and re-derive it on restart
The bridge wrote GATEWAY_ALLOW_ALL_USERS into os.environ only when unset and
never cleared it. In-process restart paths (gateway restart watcher, dashboard
profile actions) copy os.environ into the child, so a config.yaml grant became
a sticky env var: flipping allow_all_users to false and restarting left the
gateway OPEN. The bridge now tracks its own write (module flag), overwrites or
clears it on reload, exports only a truthy grant (presence-based readers such
as the Telegram intake prefilter treated "false" as configured auth), and the
two restart env builders drop the bridge-owned value so the child re-derives
the posture from its own config.yaml. Under multiplex_profiles the DEFAULT
profile's events are authorized inside its secret scope, where gate readers
never fall to os.environ; the bridged grant is now seeded into that profile's
scope mapping only (secondaries never inherit it).

Review finding: bridged GATEWAY_ALLOW_ALL_USERS survives restart and overrides a flipped config.yaml; inert for the default profile under multiplex; "false" exported as configured auth.
2026-09-15 04:37:41 -07:00
teknium1
b2577df807 fix(dashboard): accept command-scoped NOPASSWD sudo for system gateway actions
The dashboard's system-scope elevation gate hard-failed on a refused
`sudo -n true`, but the `hermes update` fleet restart it claims to
mirror treats that blanket probe as inconclusive and falls back to
probing the targeted command. A host with a sudoers entry scoped to
the hermes command (the hardened shape) therefore updated fine from
the CLI while the dashboard reported "passwordless sudo is
unavailable".

Factor the fleet's two-step probe into update_cmd_fleet
._sudo_noninteractive_ok and call it from both sites: the fleet keeps
its `reset-failed <unit>` fallback, the dashboard falls back to a
non-destructive `sudo -n -l -- <exact argv>` check before spawning.

Review finding: dashboard sudo gate diverged from the fleet posture it
borrowed (_needs_sudo only) and rejected targeted NOPASSWD sudoers.
2026-09-15 04:08:00 -07:00
teknium1
05fb879609 fix(dashboard): share the fleet's sudo posture; trim to two invariant tests
The root/sudo decision now reuses `update_cmd_fleet._needs_sudo` (the helper `hermes update`'s
own fleet restart already uses for `sudo -n systemctl --no-ask-password`) instead of a second
euid check. Tests reduced to one parametrized argv invariant (system-scope lifecycle verbs get
`sudo -n`; status and both-units-installed never do) plus the no-passwordless-sudo request
failure. Dashboard docs note the passwordless-sudo requirement on system-scope installs.
2026-09-15 04:08:00 -07:00
Baris Sencan
eaf700c67e fix(dashboard): elevate system-scope gateway lifecycle actions (#110820)
The Restart Gateway button spawned `hermes gateway restart` as the dashboard's
own user. On a systemd *system* install the CLI refuses every lifecycle verb
below root (`_require_root_for_system_service`), so the button could never work:
the refusal landed in `~/.hermes/logs/gateway-restart.log` while the endpoint
reported a started action. `start`/`stop` shared the defect.

`_spawn_hermes_action` now prefixes `sudo -n` for gateway restart/start/stop when
the action resolves to the system unit. Scope comes from the CLI's own picker
(`_select_systemd_scope`) evaluated against the profile the action addresses, so
a host with a user unit installed keeps the unprivileged path and root never
shells out to sudo. With no passwordless path the request fails with an
actionable message instead of reporting a started action whose child refuses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 04:08:00 -07:00
Teknium
b0c383cdf7 fix(desktop): served profiles show running and route lifecycle to the multiplexer from a pooled local backend
Electron sends a local sub-profile's REST to its pooled `hermes --profile X serve` without
?profile=; inside that process the unscoped branches never reached the multiplexer rung, so a
profile served by the default multiplexer read as 'Messaging gateway stopped' on the system and
messaging pages, start/stop spawned a child that exited 78 while the UI reported success, and
restart ran `gateway restart` under X's HOME (same exit 78). Remote-backend topology was already
correct because its requests carry ?profile=.

Unscoped liveness/status/messaging now take the multiplexer rung for the process's own home;
lifecycle verbs resolve the own profile, refuse start/stop with 409 and restart the multiplexer via
-p default; Electron routes POST /api/gateway/{restart,start,stop} through the primary with
?profile= so the action lives on the backend the status poll asks and outside the pooled
backend's shutdown SIGTERM.
2026-09-12 06:13:44 -07:00
Teknium
df95f378a6 feat(dashboard): "Migrate to a single multiplexed gateway" on the System page
GET /api/gateway/migrate/plan returns the CLI plan JSON; POST
/api/gateway/migrate spawns `hermes gateway migrate --multiplex --yes`
detached (action log gateway-migrate.log). The Gateway card shows the
button only for a multi-profile install that is not yet multiplexed, and
disables it while listing the blockers.
2026-09-12 01:49:28 -07:00
BowmanStephen
4d47891b4a fix(dashboard): isolate the environment of named-profile actions
_spawn_hermes_action copied the dashboard's os.environ verbatim into every
detached `hermes ...` action. The dashboard runs inside the gateway and has
loaded its own profile's .env into the process environment, so
`hermes -p <other> gateway restart` (and every other dashboard-driven
profile action) started with the DEFAULT profile's platform credentials
and ports already present. load_hermes_dotenv does not override keys that
are already set, so the named profile's own .env could not displace them:
an A2A-only profile ended up claiming the default Discord bot token and
binding the default API server / BlueBubbles ports.

For actions carrying a profile selector (`-p X`, `--profile X`,
`--profile=X`), build the child env from the standard scrubbed subprocess
environment, drop _PROFILE_MANAGED_ENV_KEYS plus every key defined by the
dashboard/default profile's .env and its hydrated secret sources, and pin
HERMES_HOME to the target profile so the child's normal startup loads that
profile's .env. Only the leading selector is inspected: argv after the
subcommand may legitimately contain -p for a nested process. Actions
without a selector keep the historical environment byte-for-byte.

Test: the new case asserts the leaked keys are gone, benign keys survive,
and (by running the real dotenv loader in a fresh interpreter with the
captured env) that the target profile's values load without reviving any
default-profile value.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 19:50:46 -07:00
Teknium
446f6f79a8 fix(multiplex): dashboard and gateway stop see a served profile's gateway as the multiplexer
A profile served by the default multiplexer owns no gateway.pid / gateway_state.json,
so every surface that reads per-profile identity files called it stopped while the CLI
status surfaces (hermes -p X status / gateway status / cron status) said "running via the
default-profile multiplexer":

- `/api/status?profile=X` and `/api/messaging/platforms?profile=X` reported
  gateway_running=false / state=None / "gateway_stopped" in the same body that listed X
  under gateways[].served_profiles. The shared ladder `resolve_gateway_liveness` gains a
  fourth rung for a named profile_dir: the live default multiplexer that records X in
  served_profiles IS X's gateway (pid = multiplexer pid, runtime = its record, X's
  `<X>:<platform>` entries re-keyed to the standalone shape).
- `POST /api/gateway/stop?profile=X` spawned `hermes -p X gateway stop`, which printed
  "No gateway running for this profile" (exit 0) into the action log while the UI flipped
  to stopped and the multiplexer kept serving X; `/api/gateway/restart?profile=X` spawned
  a `-p X gateway restart` that only exits 78. start/stop now answer 409 with the
  multiplexer explanation (one helper shared with the existing start refusal) and restart
  targets the multiplexer, the process that actually serves X. A `--force`-started
  separate gateway for X (own pid file) keeps normal per-profile management.
- CLI `hermes -p X gateway stop` refuses with exit 78 like run/start/install/restart when
  X has no gateway of its own, instead of a contradictory exit-0 "not running".

Docs: multi-profile-gateways.md §1 and §5 describe stop + the dashboard behaviour.
2026-09-11 19:38:36 -07:00
Teknium
5f1feb5344 simplify(compat): web_server — drop 221 re-exports (config/status/shutil/run_in_threadpool, lifecycle, 13 web_server_<concern> blocks, 47 route-handler legacy re-exports); web_deps.late()/LateState() take an owning-module arg; concern modules import each other directly (62 lazy sites) 2026-09-03 14:21:32 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
71a0a35e88 refactor(web_server helpers): second pass — helper dedupe in windows_ssh_runtime, webhook via atomic_json_write, table packing
- windows_ssh_runtime: _open_existing shared by _read_shared/remove_artifact; _win32 builds
  its namespace via importlib; reparse/path checks folded
- webhook: _save_subscriptions -> utils.atomic_json_write(mode=0o600) (same fchmod-before-
  rename + post-replace chmod semantics); base URL builder tightened
- web_server_messaging: override table tuples single-line, catalog entry builder flattened,
  WhatsApp payload from a field tuple; channel keys as one comprehension
- web_server_oauth: poller bodies read sess fields inline; status dicts packed
- web_server_gateway: health URL normalisation via one regex; Popen detach kwargs inline;
  topology cache getter collapsed
- web_server_cron/xai_retirement/win_pty_bridge/worktree_gc: small collapses

Verification: routes identical, --help byte-identical (webhook/worktree + subcommands),
golden corpus identical, 189 test files / 2925 passed / 0 failed.
2026-09-02 23:08:44 -07:00
Teknium
f0a451f1da refactor(web_server helpers): compact gateway/messaging/cron/oauth helpers and CLI utils (-21% LOC, zero behavior change)
- web_server_cron: one _cron_store_scope ctx manager replaces 3 copies of the
  set_hermes_home_override + use_cron_store + reset pattern
- web_server_oauth: _token_status helper for the 2 credential-status dict builders,
  logged-out sentinel, catalog table packed one card per 2-3 lines (key order kept)
- web_server_messaging: telegram request collapsed to shared detail strings, catalog
  loop unified, env-prefix aliases lifted to a module table, catalog tuples single-line
- web_server_gateway: mode ladder -> dict lookup, owned-platforms comprehension,
  action log table derived; docstrings compacted (every WHY kept)
- windows_ssh_runtime: _read_shared/_write_new/_check/_sid_str helpers, dispatch()
  if-chain -> _OPERATIONS table (arity-checked), _win32 returns a namespace
- worktree_gc/webhook/xai_retirement/write_approval_commands/win_pty_bridge/
  worktree_cmd: dead _require_webhook_enabled/_repo_root inlined, _run helper,
  duplicated except branches merged, docstrings compacted

Parity: route table identical, webhook/worktree --help byte-identical, golden
corpus of 27 pure-function outputs identical vs base 113f04616b.
2026-09-02 21:54:36 -07:00
Teknium
3a056dd466 refactor(web_server): extract gateway topology/actions/process helpers to web_server_gateway; restore late-bound config re-exports 2026-09-02 15:59:08 -07:00