Commit Graph

39743 Commits

Author SHA1 Message Date
teknium1
5bb314fa01 docs: link the Skills and Plugins hubs from the doc sidebar
The hubs are navbar-only. On a phone Docusaurus opens the drawer on the
doc sidebar and puts the navbar behind 'Back to main menu', so mobile
readers cannot find Skills or Plugins at all. Two top-level sidebar
links make them one tap away on every doc page (desktop sidebar too).
2026-09-21 07:32:44 -07:00
Austin Pickett
782886f87d fix(desktop): bound the Chromium log while the shell is running, not only at launch
The startup reclaim caps the log a previous run left behind, but a shell
that stays up for days writing Chromium ERRORs is unbounded until it
restarts — the case @ehz0ah flagged in review.

Chromium owns that descriptor in append mode for the life of the
process, so rotation is the wrong primitive: renaming leaves the writer
on the renamed inode and the cap silently stops applying. Poll the live
size and truncate in place instead; an O_APPEND writer resumes at offset
0, so a noisy process stays bounded at ~cap plus one interval's output.
The timer is unref'd and every tick is best-effort.

Co-authored-by: ehz0ah <ehz0ah@users.noreply.github.com>

Refs #100573
2026-09-21 10:23:04 -04:00
Austin Pickett
9a21794450 fix(desktop): optional Linux crash diagnostics must not be fatal or unbounded
Two follow-ups to #117851, both raised in review by @ehz0ah:

1. The diagnostics block created HERMES_HOME/logs with an unguarded
   module-level mkdirSync, before app readiness. A read-only or invalid
   logs path terminated the Linux desktop at startup — optional
   diagnostics killing the app they exist to diagnose. Wiring now runs
   through enableLinuxCrashDiagnostics(), where every step is
   best-effort: no writable logs dir degrades to "no Chromium log" (and
   keeps the crash reporter, which writes elsewhere), and a Crashpad
   handler that refuses to start is not a startup failure either.

2. Electron opens an explicit --log-file with APPEND_TO_OLD_LOG_FILE, so
   desktop-chromium.log accumulated ERROR/FATAL output across launches
   with no bound. desktop.log is capped at 10 MiB x 3 backups precisely
   because it once reached ~326 GB and exhausted the disk; the new file
   now goes through the same planner before Chromium appends to it.

Co-authored-by: ehz0ah <ehz0ah@users.noreply.github.com>

Refs #100573
2026-09-21 10:23:04 -04:00
Austin Pickett
5c57bfdcaa refactor(desktop): a shared, path-parameterized log rotation planner
desktop.log's cap (10 MiB x 3 backups, plus the discard ceiling that
reclaims a boot-loop log outright instead of renaming it to .1) lived
inline in main.ts, keyed to one hardcoded path, and main.ts cannot be
imported by a test.

Move the planner to its own module, parameterized by base path, so the
Chromium diagnostic log added for #100573 gets the same bound and the
behaviour is provable without booting Electron.

Refs #100573
2026-09-21 10:23:04 -04:00
teknium1
3be255eca6 test(update): the host obligation is host-scoped state, so each test must discharge it
Rebasing onto main turned 21 tests red in four files whenever the operator supplies
HERMES_GATEWAY_LOCK_DIR. Not a behaviour collision: the record this PR introduces lives in
the per-OS-USER host state dir, and the root conftest pins that dir per test only when the
caller supplied nothing (#118097 deliberately keeps the documented override working). With
one supplied, every test in a file shares the dir, so a case that arms the obligation makes
the next case read a restart it never owed.

Clearing the record around each hermes_cli test — rather than re-pinning the dir — fixes the
leak without touching #118097's resolution rule or any expectation.
2026-09-21 07:14:37 -07:00
teknium1
953b6f6f08 fix(update): never disarm the host restart obligation, and never discharge unknown terms
Review findings on the host-scoped update→restart obligation.

- update_cmd_fleet: an unwritable host state dir (read-only HERMES_GATEWAY_LOCK_DIR,
  container UID that does not own $HOME) made write_host_obligation return False and the
  caller ignored it, so an interrupted update left ZERO obligation — stale code, no
  warning, no catch-up restart. The return is now propagated: the legacy per-home marker
  (still read by every reader) carries it, and a host that can write neither says so.
- update_cmd_fleet::_obligation_fields: a PRESENT but unparseable/foreign-versioned host
  record no longer falls through to the legacy marker; terms nobody can read cannot be
  discharged by another record's terms.
- update_cmd_fleet::_restart_identity_sha: zip/pip/Docker installs resolve no checkout
  SHA, so the restart-once stamp was "" and could never match — every profile's update
  re-killed the one shared multiplexer. Falls back to the record's expected_sha, then the
  receipt's post-update identity.
- update_host_obligation: any main_pid probe error is unproven identity (keep its own
  restart), never an aborted restart pass.
- run_notifications: the online notice dedupes per home CHAT, so two served profiles
  sharing one chat get one message (accounting stays per profile); transport resolution is
  isolated per profile, so one broken adapter no longer starves the rest of the fan-out.
- run_adapters / run_profile_reconcile: _profile_configs is pruned with the served set, so
  a failed or removed profile no longer owes a notice nothing can deliver and
  .restart_pending.json is unlinked.

Tests cover the new format's own hazards: unwritable record dir, foreign-version record
with a legacy marker present, non-git install, shared home chat, broken adapter, pruned
config, plus a parity test for the duplicated host-state-dir resolver.
2026-09-21 07:14:37 -07:00
teknium1
061195fac1 FLEET: one host-scoped update-restart obligation, one restart per host
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.

- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
  gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
  unit->live-MainPID collapse rule. The legacy per-home marker stays readable
  and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
  idempotent per host (a completed restart onto the checkout SHA is never
  repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
  restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
  profile's home channels; the marker survives until each was reached.
2026-09-21 07:14:37 -07:00
kshitijk4poor
8f7b3b48fc fix(desktop): close the Linux notification bus on will-quit
The session-bus socket was created lazily and never closed. main.ts
already tears down its pooled keep-alive sockets in will-quit so nothing
holds the event loop open or leaks an fd past app teardown; the new
long-lived socket joins that policy. Disposal goes through the existing
close path, so live notifications are failed the same way a daemon crash
fails them.
2026-09-21 19:27:10 +05:30
kshitijk4poor
a327f1ec14 test(desktop): cover both notification transports from any host
The transport was picked from process.platform, so the ipc test mocked
the Linux module and the only vitest lane (ubuntu) never exercised the
Electron Notification branch that Windows and macOS run, while the Linux
tests skipped everywhere else. Make the platform an injectable host
parameter: the ipc test drives the Electron branch as darwin, the Linux
tests pass linux explicitly and lose their skipIf.
2026-09-21 19:27:10 +05:30
kshitijk4poor
c93b157249 fix(desktop): no notification cooldown after a benign daemon swap on Linux
Every failure in show() armed the global 10 s cooldown, including the
ones caused by the notification daemon being replaced mid-call: the
owner-changed fences, and the bus's NameHasNoOwner error for a call
already addressed to the vanished unique name. Linux has no Electron
fallback, so every notification in that window was silently dropped
exactly when the new daemon was healthy.

Classify by generation rather than by error: a failure against an owner
that is still current is the daemon's fault and cools down; one against
an owner that has since been replaced does not. The fixture now answers
calls to a stale unique name with NameHasNoOwner like the real bus, and
the race test asserts the next notification reaches the new owner
immediately.
2026-09-21 19:27:10 +05:30
kshitijk4poor
f09af639ac refactor(desktop): drop unreachable fences in the Linux notification transport
`finished` cannot be observed inside show(): close() aborts the signal
first, and dbus-native settles an aborted invoke synchronously, so the
awaiting code throws instead of resuming past the check. `connection !==
state` is implied by the generation check (disconnect bumps the
generation before clearing the connection), `connection !== target.state`
in close() is implied by `delivered` (disconnect fails every live entry,
which clears it), and receive() can never run after release() deleted its
map entry.

A malformed Notify id now gets its own error instead of being reported as
an owner change.
2026-09-21 19:27:10 +05:30
BearHuddleston
ab4455d965 fix(desktop): release naturally closed Linux notifications
(cherry picked from commit 5b0a1e3ab1982c2db8dbbceadda1499d72827631)
2026-09-21 19:27:10 +05:30
BearHuddleston
326ce363a3 fix(desktop): prevent Linux native notification freezes
(cherry picked from commit 3dd43eb59480e279835ce0f6409567ee864d47fc)
2026-09-21 19:27:10 +05:30
Austin Pickett
274bc7b8f6 docs(desktop): where the Linux Chromium log and minidumps live 2026-09-21 09:00:05 -04:00
Austin Pickett
00e08cec30 fix(desktop): a Linux SIGTRAP leaves its FATAL line and a minidump behind
Every Chromium CHECK/LOG(FATAL) traps at the same instruction (the
ImmediateCrash tail of logging::LogMessage::HandleFatal), so the cores
collected for #100573 all share one address and none say which check
fired. The message goes to stderr, which the .desktop entry and the
Omarchy wrapper both discard, and the app never started Crashpad.

On Linux, route Chromium's log (ERROR and above) to
HERMES_HOME/logs/desktop-chromium.log and keep local minidumps; nothing
is uploaded. Child processes inherit the switches, so zygote/GPU checks
land in the same file.

Refs #100573
2026-09-21 09:00:05 -04:00
teknium1
ea0c2b820b fix(host-topology): a createTime-null record is not a live gateway, and the restart we print must run
Review findings on this PR.

- `_from_host_record` accepted a record with `createTime: null`: `_same_incarnation` treats a
  missing create_time as "matches", so `liveness_is_proven()` blessed whatever process happens to
  own that PID today. A record pointing at an unrelated `sleep 60` made `cron status` print
  "the host gateway (PID ...) serving profiles default, served". Require a recorded create_time.
- `cron status`'s host-record rung printed `hermes gateway restart`, which exits 78 for exactly
  its audience (a served NAMED profile); use the `--profile default` form both rungs now share.
- doctor's WORST state (supervision slots exist, nothing owns the gateway role) emitted a warning
  but appended no issue, while the strictly less severe LEGACY case did.
- Drop three defensive wrappers in `host_topology`: a try/except around an always-importable
  `normalize_profile_name` that silently degraded to a DIFFERENT normalisation, a blanket except
  per rung, and a swallowed `get_active_profile_name`.
- `get_systemd_unit_path`, `_legacy_unit_search_paths`, the doctor host-unit probe and the
  profile-delete service cleanup hardcoded `~/.config/systemd/user`, so on a host that moves
  $XDG_CONFIG_HOME the unit probe false-negatived -- the very bug this PR fixes survived there.
  One `user_systemd_unit_dir()` helper honours the spec default.

Tests: the two "Scheduler host: the host gateway" assertions were prefixes of BOTH rungs, so the
rungs were indistinguishable and the host-record rung was never exercised (which is how the
restart-command defect shipped green). Both now assert the full line, plus a served-host-record
case. Rendezvous isolation uses `tmp_path` instead of a leaked `tempfile.mkdtemp()` and also
covers `test_gateway_multiplex_served_record.py`, which read the real host record; the
`tests/conftest.py` hook in #118097 supersedes all of them once it lands.
2026-09-21 05:06:45 -07:00
teknium1
f66089174c fix(host-topology): only a PROVEN-live host record names the host gateway
``host_rendezvous``'s contract is that a record whose (pid, createTime) cannot
be POSITIVELY matched is a candidate, never an owner. Accepting an unprovable
record turned a leftover record into a permanent "gateway running" answer on
every reporting surface. Require ``liveness_is_proven`` before treating the
record as the host gateway, with an invariant test.

Two existing tests that simulate "no gateway anywhere" now point
HERMES_GATEWAY_LOCK_DIR at an empty dir, so an unrelated host record on the
machine cannot make the profile look served.
2026-09-21 05:06:45 -07:00
teknium1
a1398a6f67 fix: doctor, cron status, claw and gateway status report the real host topology
Multiplex-only (Teknium ruling) runs exactly ONE `hermes gateway run` per host,
multiplexing every profile. Four reporting surfaces still asked a per-PROFILE
process question ("does MY profile own a gateway process?"), so a SERVED profile
answered "no" and the output lied:

* `hermes -p served cron status` printed "Gateway is not running" and told the
  user to run `hermes gateway install` / `gateway run` for that profile — i.e.
  to start a SECOND host process, which the architecture forbids.
* `hermes doctor` reported "Per-profile gateways: up/total" from s6 slots and
  checked systemd linger against the CURRENT profile's unit, so a doctor run
  under a served profile skipped the check entirely.
* `hermes claw`'s destructive-action warning was gated on `get_running_pid()`,
  which returns None for a served profile — the token-conflict warning was
  SILENTLY SKIPPED (a real safety hole).
* `gateway.status.multiplexer_liveness_for_profile()` returned None for the
  default home by construction, so `default` could never be reported as SERVED.
* `hermes doctor`'s state.db holder + WAL wording implied one gateway per
  profile.

New `gateway/host_topology.py` resolves "which single process owns the gateway
role on this host, and which profiles does it serve" once, from the host
rendezvous record (`gateway/host_rendezvous.py`), falling back to the default
home's recorded `served_profiles` for a gateway that predates the record.
`default` is just another served profile there.

Root cause: every surface derived gateway identity from per-profile artifacts
(argv `-p <name>`, `gateway.pid`, the active profile's systemd unit, s6 slot
counts) instead of the host record that actually names the owner.
2026-09-21 05:06:45 -07:00
teknium1
a10620a669 fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.

- ATTACH now requires a LIVE `identify` answer. The claim-time record is
  published with NO served set (the runner settles multiplex a moment later),
  and `host_gateway()` reports `served_known=False` when nothing answers. An
  owner whose served set is unknown yields a TRANSIENT refusal, never an
  attach: previously `default`'s claim published "default,other" before its
  socket bound, `other`'s systemd unit read that as "I am served", exited 78,
  and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
  instead of forcing `multiplex=True`, so a standalone gateway stops claiming
  the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
  `--force` into `_host_attach_or_none`. Both previously exited in the guard
  before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
  config refusal every supervisor parks on; "someone else serves me right now"
  is a runtime observation that ends when that process does. No unit files
  change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
  finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
  would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
  and re-enters with `replace=True`, so it can no longer attach to the corpse
  it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
  `st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
  it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
  generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
  `gateway status`/doctor across N profiles pays one probe, not N.

Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
2026-09-21 05:02:29 -07:00
teknium1
a20a88398f gateway: lifecycle verbs mean "the one host multiplexer"
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:

- `gateway run` for a profile the host gateway already serves ATTACHES: print
  its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
  to re-scan `profiles/` (control socket) and attach once the answer includes
  it. Refuse only when the host gateway cannot be made to serve it. Under a
  service supervisor the attach exits 78 instead of 0 so a redundant unit is
  parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
  landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
  the owner's control socket, so it no longer depends on the claim ordering in
  start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
  process: they restart the host multiplexer and preserve its served set. A
  secondary still running its own gateway is reported with the
  `gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
  argv: a host singleton runs bare/default argv and can never prove it serves
  profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
  multiplexer is whichever profile launched the one host process.

Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
2026-09-21 05:02:29 -07:00
kshitijk4poor
7b660e66ee fix(homeassistant): detect the supervised launch through is_supervised_gateway_launch
gateway.restart already answers "was this gateway launched by a generated
service" (HERMES_SUPERVISED_CHILD or launchd's XPC_SERVICE_NAME, so a plist
that predates the marker still counts); the hint reuses it instead of a
second env read. The negative branch (plain unreachable host without a
supervisor) is now an unmarked test so the Linux lane keeps covering the
function; the macOS-only test holds the positive branch.
2026-09-21 17:16:23 +05:30
kshitijk4poor
eda2432cce fix(gateway): scope the Local Network hint and keep the osascript wrapper out of gateway process scans
Review of the first cut:

- The Home Assistant errno hint matched a bare 65 and two error strings on
  every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
  errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
  genuinely unreachable HA host was told macOS was blocking launchd. Gate on
  darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
  generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
  `gateway run`, so `hermes gateway stop` would have signalled osascript
  alongside the gateway. The canonical matcher now bails on an exact
  argv[0] basename of osascript (the gateway is its child and is matched on
  its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
  carries (osascript's own output) instead of implying it is redundant.
2026-09-21 17:16:23 +05:30
kshitijk4poor
5d63d3e5c5 fix(gateway): launchd gateway reaches the local network on macOS (#71206, #57812)
macOS Local Network Privacy attributes a socket to the executable launchd
spawned for the job. A bare venv Python has no application identity and is
not platform-entitled, so every LAN connect from the supervised gateway
failed with errno 65 (No route to host) while the same URL worked from
Terminal — and nehelper never showed a prompt that could grant it, so a
headless Mac had no way out.

Run the job through /usr/bin/osascript: `do shell script "exec …"` spawns
its child as osascript-responsible — an Apple platform binary — and the
child is exempt. Verified live on macOS 26.3.1 with a fresh ad-hoc binary
under the gui launchd domain: bare → 65, `/bin/sh -c exec` → 65,
`/usr/bin/time` → 65, osascript → reachable; an ad-hoc-signed helper .app
with NSLocalNetworkUsageDescription (the #115196 approach) stays denied with
no prompt even after lsregister, matching the #57812 dead-end table.

`do shell script` buffers the child's output until exit, so the shell
command redirects stdout/stderr to the log files the plist already routes;
`exec` keeps the gateway in the job's process group, so `launchctl bootout`
still delivers SIGTERM to it (live: stop → no orphan, restart → exit 75 →
KeepAlive respawn, stop → parked). The existing plist-staleness refresh
picks the new definition up on `hermes gateway install`/`start`.

Mechanism proposed by @leewaiho in #57812.

Co-authored-by: leewaiho <18321182+leewaiho@users.noreply.github.com>
2026-09-21 17:16:23 +05:30
xxxigm
61938453c0 test(homeassistant): errno 65 under launchd names the Local Network remedy (#71206) 2026-09-21 17:16:23 +05:30
xxxigm
1d871110e1 fix(homeassistant): name the macOS Local Network block behind errno 65 under launchd (#71206)
A launchd-run gateway that macOS Local Network Privacy denies sees every LAN
connect fail with EHOSTUNREACH while the same URL works from Terminal.
Annotate the Home Assistant connect/reconnect log lines with the cause and
the remedy so the failure is actionable instead of a bare "No route to host".

Salvaged from #115196 (remedy text points at the regenerated launchd job).
2026-09-21 17:16:23 +05:30
kshitijk4poor
6668ac14a0 fix(desktop): count the answer row in a page-boundary fold
A hydrated page can start mid-turn: the first row is the tool-only
assistant, the second its tool result, and the answer arrives third with
no active bubble to append to. The bubble that fold creates stands for
all three backend rows, but only the two pending rows were counted
(serverRowSpan: 2), so the release rewound the older-page offset short
of what it gave up and 'Show earlier' skipped rows the reader could no
longer reach.

Count the current answer row too, and pin it with a regression whose
page starts on the tool-only assistant/result pair.

Found in review of 72a5aa3fd2 by @ehz0ah.
2026-09-21 17:14:45 +05:30
kshitijk4poor
3f985c84cd fix(desktop): count released transcript rows in backend rows
The hydration fold merges a turn's tool rows into the bubble they belong
to, so a ChatMessage is not one backend row (a ten-tool-call turn is one
message over eleven rows). The older-page offset is counted in backend
rows, so rewinding it by the store's message count skipped history: the
next 'Show earlier' page started past rows the reader could never reach.

Carry the count out of the fold as ChatMessage.serverRowSpan and rewind
by it, which makes the rewind exact instead of approximate. Drop the
'has a rowId' release guard for the same reason it was wrong before: a
folded bubble can have no single durable id while every row behind it is
persisted and re-fetchable; what must not be released is a row still in
flight (pending).
2026-09-21 17:14:45 +05:30
kshitijk4poor
7c799c26ce fix(desktop): rewind the older-page offset in the backend's own units
The backend pages display history (include_compacted reads group by display
order; inactive rows are not counted), so a locally counted row total is not
guaranteed to be the offset's unit. Decrement the offset the backend reported by
the rows the store gave up, clamped at zero: a mismatch can only overlap a page
already in memory (which the merge dedupes), where an over-counted absolute
offset would skip rows the reader could then never reach.
2026-09-21 17:14:45 +05:30
kshitijk4poor
71a0ce913c refactor(desktop): share the tail-entry lookup, and plan nothing when nothing is released
Three copies of the tail-entry resolution (record the page, rewind, read the
state) now go through one resolveTailEntry; the retention planner reports
released:false without walking the weights or counting persisted rows, and the
hook checks for a fetch route before planning instead of discarding the plan
afterwards.
2026-09-21 17:14:45 +05:30
kshitijk4poor
14b4327445 fix(desktop): the store releases paged-through history instead of holding it for the window's lifetime (#77311)
The transcript window bounds what reaches assistant-ui, but the session store
underneath keeps every message it ever materialized: the tail hydration, every
older page "Show earlier" fetched, and every turn the window streamed. Renderer
retention therefore grows with session content and never comes back down — the
remaining half of #77311 after #117681.

`boundRetainedTranscript` keeps the live window plus one window page of slack
(so the first "Show earlier" still pages through memory) and releases the rows
older than that, which are off-screen and already persisted. Two invariants
keep the release safe: nothing is released without a durable `rowId` (a row the
backend cannot serve again would be content lost), and a branch group is never
split at the cut, for the same reason the window cut does not.

`rewindTranscriptTail` points the session's tail bookkeeping at the retained
prefix, so the released rows stay reachable: the existing older-page backfill
fetches them back on the next "Show earlier", and the rewind is refused outright
when the session has no recorded page route to fetch from.

`useTranscriptRetention` runs on the window's own cut moves — streaming grows
the tail, "Show earlier" prepends, a re-cut moves the anchor — so the weight walk
never happens per token. It computes the cut inside the store updater, from the
array it is about to replace: a view snapshot can be a flush behind the store,
and trimming a stale copy back into it would drop the newest rows of a live turn.
2026-09-21 17:14:45 +05:30
kshitijk4poor
db1f3f4564 fix(computer-use): doctor names the stale TCC row on the health_report path too
The first cut appended the recovery only in `_tcc_row`, which is built by the
0.10 fallback probes. Drivers that serve health_report (0.22+, the ones a
stale row actually bites) render their own tcc_* rows untouched, so `doctor`
never showed it. Apply the hint at the report seam like the display-count
guard so both paths get it.

Also: one CUA_DRIVER_BUNDLE_ID (permissions.py is the leaf; the daemon
imports it), the CLI derives the field set from the same table, and the
unverified "--upgrade repairs stale rows" claim is dropped from hint + docs
(the user hitting this is already on a current driver).
2026-09-21 16:21:00 +05:30
kshitijk4poor
42ae8b4410 fix(computer-use): name the stale TCC row when CuaDriver shows ON but is denied
`hermes computer-use permissions status` and the doctor's fallback `tcc_*`
rows told a user whose System Settings toggle already showed CuaDriver ON to
"grant it in System Settings" — the one step that cannot help. macOS keys the
TCC row to the app's code-signing requirement; a row written for an earlier
CuaDriver build stops matching after a driver update and flipping the toggle
does not rewrite it (trycua/cua#3170, repaired by the cua-driver >= 0.22
installer on update). The daemon then reports Accessibility / Screen Recording
false while the pane shows ON, and every click silently no-ops (#99732).

`permissions.stale_tcc_grant_hint(*missing)` renders the reset for exactly
the missing services (`tccutil reset Accessibility|ScreenCapture
com.trycua.driver`, then `hermes computer-use permissions grant`); both
surfaces append it only when a grant is actually reported False. Docs: the
computer-use page still said bounded/unrestricted daemons run "under the
Hermes host identity" — they launch through CuaDriver.app since #95381 — and
gains the stale-row troubleshooting entry.
2026-09-21 16:21:00 +05:30
Siddharth Balyan
080907e3b7 chore(auth): drop the unused tool:invoke OAuth scope from the step-up request (#118036)
Request only inference and billing access when no prior Nous scope exists.
2026-09-21 15:55:09 +05:30
teknium1
bfcd209906 docs(multiplexing): record the config escape hatch and the .env-over-frozen-env precedence
Activation is not purely stricter for the LAUNCH tenant: inside the launch scope its own
.env wins over the frozen boot env, so a key present in both reads differently before and
after activation. Also documents the explicit multiplex_profiles: false escape hatch and
the fail-closed unresolved-profile case.
2026-09-21 02:58:40 -07:00
teknium1
af380f66ae fix(mcp): divide the shutdown budget across profile passes and stop poisoning the teardown context
Each per-profile pass could spend 15s against a 5s caller budget, so on a host with
several served profiles the trailing wildcard pass — the only one that stops the shared
MCP loop and its connections — might never run. shutdown_mcp_servers takes a timeout and
the caller gives each pass budget/(N+1).

The worker also ran under copy_context(), carrying a caller's HERMES_HOME override into
a pass that documents it has none: from inside profiles/poison's scope the wildcard pass
resolved live_home to that profile. It now starts from a fresh Context().

Both call sites replace `with suppress(Exception)` with a logged WARNING: the suppress
swallowed a real TypeError from this signature change, which surfaced only as a missing
teardown in an unrelated test.
2026-09-21 02:58:40 -07:00
teknium1
bc0a42fd96 fix(kanban): scrub the dispatcher's credentials from another profile's worker on any host
_default_spawn gated the scrub on is_multiplex_active(), so on a single-profile host —
the common Kanban deployment — build_subprocess_env never reached _sanitize_subprocess_env
and profile B's worker env was byte-identical to the dispatcher's, OPENAI_API_KEY and
systemd-injected tokens included. The authority test is "is this worker acting for a
ROUTED home", exactly as served_profile_child_env decides it, not the gateway-wide flag.

Drops install_profile_terminal_scope from the bind_home=False branch: that path has no
terminal-scope consumer and it cost a config.yaml parse per spawn. The test now asserts
the worker ENV rather than that a scope object was bound.
2026-09-21 02:58:40 -07:00
teknium1
d80aaf1c07 fix(serve): honour multiplex_profiles: false, freeze the launch env last, be loud on failure
Three defects in the eager activation gate:

- profiles_to_serve() takes the multiplex flag as an ARGUMENT and never reads config, so
  any host with >=2 profile dirs armed the fail-closed guard at boot — including hosts
  that deliberately pinned gateway.multiplex_profiles: false. The gate now reads
  GATEWAY_MULTIPLEX_PROFILES then config.yaml and skips activation when it is explicitly
  false; the lazy backstop is unchanged.
- The snapshot was taken at the TOP of start_server, before keepalive / auth gate /
  uvicorn build. It is the ONLY source for launch keys with no .env to rebuild from
  (systemd Environment=, op run, Compose), so a credential injected or rotated by a later
  boot step was invisible for the process lifetime. Activation moves to the last boot step.
- A failing probe returned False silently: the guard never armed and the caller's
  diagnostic was dead code. It now logs at WARNING with exc_info and fails CLOSED — an
  unreadable profiles/ cannot prove the host is single-profile.

Also: a profile dir holding only an EMPTY .env is what a crashed `hermes profile create`
leaves behind, and counting it flipped a single-profile host fail-closed at its next
boot. The gate counts named_profile_has_servable_identity instead.
2026-09-21 02:58:40 -07:00
teknium1
cbeedffc90 fix(tui_gateway): never run an external secret command inside the exit-flush budget
_session_profile_runtime_scope hydrates a profile's EXTERNAL secret sources at entry —
it shells out to the operator's secret command (op run / bws / sh -c) with a 30s CLI
budget while holding the process-global source-cache lock. The exit-flush worker has a
5s TOTAL budget, so with a slow source it flushed 0 transcripts where base flushed 1,
and the loss was silent: both callers discard the return value. The periodic reaper
tick paid the same cost per session per tick.

Persisting a transcript needs no external credential: pass hydrate_secrets=False (the
flag this branch already uses in run_profile_reconcile), and WARN when the exit flush
persists fewer transcripts than it found.

Docstrings corrected while here: the transcript ROWS were never at risk (a profile
session owns a SessionDB whose db_path is frozen at construction); the real defect is
the config / token-ledger / provider resolution that _persist_session does at call time.
2026-09-21 02:58:40 -07:00
teknium1
4aa9baf139 fix(multiplex): a profile whose home does not resolve must not borrow the launch profile's secrets
_profile_home_or_none answered None both for "this body is the launch profile's own
work" and for "this NAMED profile's home no longer resolves" (deleted or renamed
mid-run, invalid name, transient OSError), and the handler factories captured that
answer once. With the launch-profile fallback added earlier on this branch, an inbound
message on secondary B's own bot was then handled with the LAUNCH profile's .env and
frozen env — a silent cross-tenant credential fallback where origin/main raised.

_routed_profile_home now returns an UNRESOLVED_PROFILE_HOME sentinel and logs a
WARNING; that case binds nothing, so an unscoped get_secret raises UnscopedSecretError
under multiplexing. None keeps exactly one meaning: launch-owned.
2026-09-21 02:58:40 -07:00
teknium1
c718267ef3 fix(multiplex): close the profile-scope holes the scope machinery misses
One host process serves every profile, so every execution point must bind the
profile it is acting FOR. These six ran unscoped (or bound only part of a
scope) and resolved get_hermes_home()/credentials against the LAUNCH profile:

- tui_gateway/session_reaper: the idle-reaper and exit-flush transcript writes
  now enter the SESSION's profile scope, the same chokepoint _finalize_session
  already binds. A served profile's transcript was landing in the launch home.
- gateway/run: MCP shutdown tears down per served profile inside that
  profile's scope (mirrors startup discovery and the reconcile chore), with a
  trailing wildcard pass under the launch profile's own scope.
- gateway/run_profile_reconcile: _unserve_profile's adapter teardown, agent
  eviction and state/memory handle release now run inside the deleted
  profile's scope.
- gateway/run_adapters + run_goals + run_notifications: a body with no routed
  profile no longer means "no scope". launch_profile_scope_if_multiplexed()
  binds the launch profile once the process multiplexes; before activation it
  is still literally a nullcontext, so single-profile hosts are unchanged.
- hermes_cli/kanban_db_dispatch: one _worker_profile_scope helper binds the
  assignee's secret AND terminal scope for toolset resolution and the spawn-env
  build, unconditionally instead of only under multiplex.
- hermes_cli/web_server: `hermes serve` activates multi-profile hosting at boot
  when the host has more than one servable profile home, instead of lazily on
  the first ?profile= request after earlier work already ran unscoped.

Secret scope is never widened: a non-launch home resolves from its own .env and
sources only; the launch home keeps its existing env-over-.env precedence.
2026-09-21 02:58:40 -07:00
teknium1
8b71dc31e4 fix(state): VACUUM admission counts live holders only; argv proves the home, not the install
Independent review found the previous commit could make the unbounded-growth
symptom it fixes PERMANENT, and that its argv narrowing re-opened #92401 inside
a single install.

- other_generations_for_path() counts only LIVE generations. A retired one is
  already write-fenced (StateDbReplacedError, close-time checkpoint disabled)
  and leaves the registry only when its holder releases -- which a gateway
  handle does not do before shutdown. One inode replacement (repair swap,
  backup restore, snapshot) therefore skipped auto-VACUUM for that path for the
  whole process lifetime.
- _argv_scoped_to_other_home is ranked evidence now. argv[0] is the SHARED
  install binary for every profile on a host, so it is neutral, never proof of a
  hold; the process's own --hermes-home / HERMES_HOME= / --profile / -p
  selection decides which home it serves, and a token naming ANOTHER profile's
  store is tested before any own-prefix check. argv[0] also stops dismissing a
  holder of a store whose home is not part of an install layout (a custom
  HERMES_HOME is served BY the binary under ~/.hermes).
- install_root comes from hermes_constants.named_profile_home, not
  basename(parent) == "profiles": an arbitrary <X>/profiles/<n>/ tree no longer
  promotes all of <X> to "ours".
- The launch profile keeps its gateway.sessions_dir override when its own store
  is pruned; every other served profile prunes under its own <home>/sessions.
  Pruning under the wrong dir orphaned transcripts forever.
- glob.escape on the request_dump_<id>_* sweep (pre-existing).

Tests: the retired-generation test asserted the starvation mechanism; it is
replaced by the invariant (a live sibling defers VACUUM) plus a red-on-base test
that a retired, write-fenced generation does not. New argv cases cover the
shared binary with -p other, another profile's store token, and a non-Hermes
<X>/profiles/ tree; the housekeeping fixture now asserts the unpinned store
still resolves inside the sandbox before yielding.
2026-09-21 02:56:31 -07:00
teknium1
a2c0af430c fix(state): per-profile store housekeeping under one multiplexed process
Three silent data-correctness bugs on a host where ONE process serves every
profile:

- The gateway constructor ran auto-archive and auto-prune/VACUUM once, on a
  handle pinned to the construction-time launch home, so a served secondary
  profile's state.db was never pruned or vacuumed by anybody. Both now run per
  SERVED profile, inside that profile's runtime scope, against its own
  SessionDB, its own `sessions:` config and its own transcript dir.
- The dashboard/`hermes serve` auto-archive sweep archived an arbitrary
  profile's store but read the config through the PROCESS HERMES_HOME, so one
  profile's sessions.auto_archive_days governed every other profile's
  retention. It now loads the config of the home whose store it sweeps.
- Auto-VACUUM admission read "no FOREIGN holder" as "the store is quiet", but
  the /proc scan skips our own pid by design. VACUUM plus its TRUNCATE
  checkpoint therefore retired a WAL generation another live SessionDB in THIS
  process still held. The path-keyed registry now answers the in-process half.

Also: argv naming only the install root no longer dismisses a multiplexer as
"another instance" when the store is a profile store beneath that root — under
one-process-per-host that argv is exactly what the holder looks like.
2026-09-21 02:56:31 -07:00
teknium1
9ecd22e5da fix(cron): release a profile's in-flight claim under the key it registered with
Under one multiplexing ticker, every non-launch profile's in-flight claim
leaked on every run. The cron home scope is a ContextVar: the claim is taken on
the ticker thread inside `_profile_cron_scope`, but the pool worker's `finally`
sits outside `ctx.run`, where the worker thread resolves the LAUNCH home — so
the discard missed the real key. The job then skipped a fire window until the
force-release backstop swept it, and the shutdown drain plus
`hermes.cron.jobs.running` saw phantom work.

`_submit_with_guard` now captures the registering home and passes it to
`release_running_job(job_id, home=...)` on every release path.

Also in this pass:

- Stand down for a profile that runs its OWN gateway (`_cron_profile_gate`,
  the gate the serve/Desktop ticker already passes). On a host pinned to
  per-profile gateways both processes raced that profile's tick lock, and when
  the launch gateway won, delivery went through SharedRouteAdapters/fail-closed
  instead of the profile's live adapters. The gate compares the liveness PID
  against `os.getpid()`: this process holds the launch `gateway.pid` and
  publishes every served profile in `served_profiles`, so a bare liveness answer
  would have stood cron down host-wide.
- Job-liveness consumers ask `is_job_running(job_id, home=...)` instead of the
  host-wide bare-id union, which let profile A's running `daily-brief` report
  profile B's idle one as running and keep B's stale one-shot alive.
- `register_ticked_homes` reaps the parallel pools of homes that leave the
  ticked set; pools lived until `atexit`, so every home ever ticked kept a
  ThreadPoolExecutor and its worker threads.
- `mark_running_jobs_interrupted` reads the real home Path from
  `_inflight_home_path` instead of rebuilding it from the normcased key half.
- The ownership test imports `cron.scheduler_ownership` per test, so the file
  fails behaviourally on base instead of as a collection error.
2026-09-21 02:52:15 -07:00
teknium1
cb647a018f fix(cron): one host ticker owns every profile's cron, per profile
The cron ticker multiplexes N profiles from one process while its ownership
predicates and its in-flight bookkeeping still assumed one profile per process.

- `_should_yield_tick_to_fresh_gateway` asked a process-global boolean
  (`owns_gateway_runtime_lock`) and a launch-home lock probe, so one answer
  covered every profile ticked. It now asks `scheduler_ownership`:
  `owns_cron_tick_for(home)` (this process is the host gateway AND ticks that
  home) and `live_gateway_ticking(home)` (another live host gateway whose
  published served set covers that home).
- In-flight state (`_running_job_ids`, `_running_since`, `_running_futures`,
  `_running_allowance_s`, `_running_worker_pids`, `_running_fire_owners`,
  `_restart_safe_waiter_job_ids`, `_interrupted_job_ids`) is keyed by
  `_inflight_key(job_id)` = `(home key, job id)`; two profiles carrying a
  `daily-brief` no longer read as one job. The public accessors still report
  the host-wide union of bare job ids for the shutdown drain.
- The parallel worker pool is keyed by home: `cron.max_parallel_jobs` is a
  per-profile key, and the single global pool was sized by whichever profile
  ticked first and torn down by the next one.
- `gateway/run.py` no longer gates the cron tick set on
  `gateway.multiplex_profiles`: that flag gates adapters, and with it off every
  non-launch profile's jobs sat in a store no ticker visited.
2026-09-21 02:52:15 -07:00
teknium1
5289a0058d docs: unescape a quote in the HostLockOutcome docstring 2026-09-21 02:51:58 -07:00
teknium1
f9b8cf7c42 fix(gateway): a failed host attach never exits 0, and SIGTERM clears the record
Reviewer findings on the host rendezvous record + serve attach.

- Attach now PROVES the owner before exiting 0: a bounded TCP connect to the
  recorded endpoint plus a token-authenticated GET /api/host/identity that must
  answer as the recorded pid+role. A supervisor or `hermes update` relaunch
  landing in the old process's graceful-shutdown window (socket closed, atexit
  not yet run) previously exited 0 with NOTHING listening, so the service
  reported success for a dead backend. Anything short of a proven owner falls
  through to the bind.
- Unprovable liveness (no psutil, an unexpected psutil error) is a CANDIDATE,
  not an owner: it goes to the same probe instead of exiting 0, which is what
  turned a record for a long-dead pid into a permanent silent outage.
- An explicitly typed --port/--host the owner cannot serve is a non-zero refusal
  naming the owner, never a silent loopback redirect; and a `hermes dashboard`
  user is never routed to a headless `serve` backend (servesSpa in the identity
  answer).
- SIGTERM — the NORMAL stop (systemd stop, docker stop, the update relaunch) —
  now clears the record and the 0600 token and releases the host lock. atexit
  does not run on it: uvicorn's capture_signals re-raises into the default
  disposition, so a live session token outlived its process indefinitely. The
  handler only prepends cleanup and hands off to the previous handler, leaving
  the shutdown sequence unchanged.
- Host lock claim is tri-state (acquired / held-by-other / could-not-open) and
  logs the OSError: an unwritable lock dir used to be reported as "another
  gateway owns this host", sending operators hunting a process that never
  existed. Its handle cache is keyed by (role, resolved lock path), so a changed
  lock dir can no longer make owns_host_lock() lie.
- Windows: the token is written through the SSH runtime's protected
  owner+SYSTEM DACL writer (os.open(0o600) sets no ACLs there, and os.replace
  fails against an open reader); read_token's docstring no longer claims the
  mode bits prove same-OS-user.
- gateway/status: a RELATIVE $XDG_STATE_HOME is ignored per the XDG spec — it
  made the host lock dir CWD-relative, so two serves started from different
  directories shared no singleton.

Live A/B (real processes): kill -TERM of a fully started `hermes serve` left
host-serve.json + host-serve.token on disk on base, removes both on head (exit
status -15 unchanged). A record with a LIVE pid and a closed port made base exit
0 with no bind; head binds. `serve --port 8899` against an owner on another port
exits 1 naming the owner.
2026-09-21 02:51:58 -07:00
teknium1
837b905168 test: host singleton invariants — one host lock, attach, stale record ignored
Red on base: `gateway.host_rendezvous` does not exist there, and the real-process
A/B shows base binding two ports for two profiles where head binds one.

- two HERMES_HOMEs each take their OWN per-home gateway lock (unchanged) but
  only the first takes the host lock;
- a live record ends a second `hermes serve` at exit 0 before the bind;
- a record whose creation time does not match its live PID is a recycled PID
  and is ignored; `--isolated` never attaches.
2026-09-21 02:51:58 -07:00
teknium1
9eb90b0a03 feat(gateway): host-wide singleton lock + rendezvous record
Multiplex-only (Teknium ruling): exactly ONE `hermes serve` and ONE
`hermes gateway run` per host, each multiplexing every profile. The existing
gateway lock/PID files are anchored to HERMES_HOME, so N profiles were N
independently acquirable locks and `hermes serve` had no singleton at all.

- `gateway/host_rendezvous.py`: a host lock + a published record (pid,
  creation time, port, protocolVersion, tokenFingerprint, served-profile set)
  under the one cross-profile lock root already in the tree. A record whose PID
  is dead, or alive with a different creation time (PID reuse), is stale and is
  never attached to. Removed on clean exit.
- Gateway: takes the host lock ALONGSIDE the per-home lock and publishes the
  record. Observe-only — a second gateway still starts, and logs the owner.
- Serve: publishes the record plus a 0600 token file on bind, and a second
  `hermes serve` for ANY profile discovers it, prints the live pid/port and
  exits 0 instead of binding a second port. The bare-TCP `_dashboard_listening`
  probe (which proved only that *something* answers) is replaced by
  record-based discovery with creation-time proof; the unified re-exec stays as
  the no-record fallback. `--isolated` and HERMES_DESKTOP=1 opt out.

`spawn-ledger.json` keeps being written unchanged, so the merged Desktop attach
ladder (#117983) keeps working.
2026-09-21 02:51:58 -07:00
teknium1
3917c7dd6a fix: RoomLink catalog names the served profile and reads its policy from that profile
`catalog_mapping` took an optional `target_profile` and fell back to the launch
process's `HERMES_PROFILE` (then "default"), and `execution_policy_mapping`
loaded whatever gateway home was active while labelling the result with the
requested `target_profile`. On a multiplexed / Desktop-spawned backend a remote
Bot invited for profile B could receive a catalog signed for the launch
profile, or B's name over A's approvals/turn-limit/toolsets.

- `target_profile` is a required keyword on `catalog_mapping`; an empty or
  mismatched profile fails loudly naming the field, no env fallback.
- `execution_policy_mapping(config=None)` resolves the SERVED profile's own
  config: a no-op when it is the active home (callers that already scope are
  unchanged), else `_profile_runtime_scope(get_profile_dir(profile))`; a profile
  that does not exist is refused instead of silently reading the launch config.
- `issue_room_grant`'s default digest inherits the served-profile resolution.
- Tests: 27 call sites now name the profile; two invariants red on base.
- Docs: multi-profile resolution table row for the RoomLink catalog/policy.

Fixes #116900
2026-09-21 02:08:18 -07:00
teknium1
8863b36fd6 fix(cron): stale-code yield reads as an outage in cron status; hermes update restarts proven-stale gateways
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:

* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
  loop) and the success marker; with no success marker on disk it printed
  "✓ Gateway is running — cron jobs will fire automatically", and with a stale one
  it pointed at "Check the gateway log" instead of naming the cause. The persisted
  `CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
  and reported as "Gateway is running STALE code — fires NOTHING", with both
  revisions and the restart command. A fresh heartbeat with a recorded error and no
  success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
  left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
  to the existing drain-first `request_restart` path (SIGUSR1) via
  `hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
  code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
  contract the restart phase already uses for unmapped manual gateways. The drain
  budget computation is shared (`_gateway_drain_budget`).

The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).

Fixes #117275
2026-09-21 01:57:31 -07:00