source ./activate and . .\activate.ps1 always run setup, and setup always
runs `python -m pm.cli install`. An explicit install re-hashes every
published tool entry so it can repair a corrupted one. On this machine
that hash reads about 550 MiB and takes 6.0 of the 6.6 seconds, even
when nothing changed.
Shell activation now passes --trust-recorded and trusts the digest the
install recorded, the same check startup already uses. A missing tool
is still installed and a stale venv is still rebuilt. `hermes pm
install` and `hermes update` keep the byte check, and the flag refuses
names, --extra, and --target so it cannot narrow an install someone
asked for by name.
Measured on this machine, already up to date: the activation path drops
from 6.6 s to 0.9 s. A bare `python -m pm.cli install` stays at 7.2 s.
The remaining 0.9 s is process startup and imports, not hashing.
Making install-stamp.json the single version identity removed
hermes_cli.__version__, but shipped updaters import that name after the
checkout swap (tests/compat/old_updater_surface.json), so an old updater
pulling this tree would die on ImportError mid-update.
Restore it as a compat stub that reads the stamp's baseVersion through
steward.read_install_stamp, the one stamp locator, so Nix's
HERMES_INSTALL_ROOT is honoured. It never runs git (it is evaluated on
every package import) and keeps the pre-stamp 0.0.0 placeholder when a
checkout has no stamp. In-tree code keeps using get_version_info().
A desktop hand-off on Windows stalled at "Another Hermes update is already
running (process <the hand-off itself>)". The hand-off PowerShell claims the
marker and runs the old `hermes update`, whose _old_updater shim spawns
_update_takeover.py as python -I -S -B. HERMES_UPDATE_HANDOFF_PID names the
old updater, not the marker owner, so adoption relies on the ancestry walk;
without psutil (-S) that walk falls back to /proc or ps, and Windows has
neither, so the grandchild refused its own orchestrator and exited 2.
Read the parent pid from a Toolhelp32 snapshot through stdlib ctypes, and
treat a parent created after its child as a recycled pid, as psutil does.
Reproduced and verified on a real Windows venv: the grandchild now adopts
its ancestor's marker and still refuses an unrelated live holder.
The regression tests were unmarked, and the Windows lane imports only files
carrying a platforms marker, so the Windows half never ran: mark both
stdlib-walk tests for every lane.
Also print a line when the old updater hands off to the new (PM) updater,
so logs show where the switch happened.
cp -c clones file by file: 21.5s for a real 148k-entry HERMES_HOME. APFS
clones a whole directory tree in a single clonefileat(2), which a stock Mac
can call through /usr/bin/perl: 5.9s for the same home, identical in name,
type, size, mode, mtime and link target. The kernel stamps cloned
directories with the current time, so a second pass restores each
directory's atime/mtime. On any failure the entry is removed and cloned
with cp -c instead, with a warning the smoke test asserts is absent.
New `-Phase verify-stamp` in the windows install/update driver, wired into
install-e2e-windows-run.yml AFTER the update phase — including the new
runtime's launch and smoke checks — because the bootstrap marker can
complete on a later run than the install itself. Known-failure legs skip
it: they legitimately never reach the new HEAD.
The phase runs scripts/verify-bootstrap-version-stamp.py against the
installed checkout with --expect-commit from the staged serve state, so
a real windows-latest run now asserts both the bootstrap-complete
receipt and the checkout's install-stamp.json tell the truth about the
final HEAD (native coverage the Linux-only suite can't provide).
Validation: windows-e2e.ps1 parses clean under real pwsh; workflow YAML
parses; tests/scripts/test_powershell_script_syntax.py green.
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
pre tarred the entire HERMES_HOME and post extracted it over an emptied
home: two full copies of a tree that is mostly node_modules and venvs.
pre now takes a `hermes backup` of the user's data, then clones every
top-level entry of HERMES_HOME except cache/ (and the Electron userData)
into the backup dir: copy-on-write where the filesystem can (clonefile on
APFS, reflink on btrfs/xfs), a plain copy elsewhere. post moves the live
entries aside and the clones in, by rename only, so the home directory
itself never moves (it may be a mountpoint or symlink target) and cache/
stays where it is. It then deletes the moved-aside post-update trees;
hermes-backup.zip is never deleted.
Windows deletes a VSS snapshot once the old copies of rewritten blocks
exceed the volume's shadow-storage cap, which defaults to ~1% of the disk
(17.8 GB on a 1.8 TB drive). A long rehearsal with other disk activity can
blow that and cost the tester the rollback.
pre now raises the cap to 128 GB (a ceiling, not a reservation) right after
taking the snapshot, recording the exact original in shadowstorage.txt; post
puts it back, also on the snapshot-gone path. An already-larger cap is left
alone.
pre tarred the entire HERMES_HOME, which is slow on a real home (hundreds of
thousands of files) and silently left out files other processes held open:
on a live 33GB home the tar came out at 4.7GB with no error.
pre now takes a `hermes backup` of the user's data, then a Volume Shadow
Copy snapshot of the volume(s) holding HERMES_HOME and the Electron userData.
The snapshot is copy-on-write: 2s, nothing copied, locked files included.
post checks every snapshot still exists before touching anything, then
robocopy /MIR's both trees back from it (only changed files are copied,
files the update added are deleted; HERMES_HOME\cache is left alone), and
deletes the snapshot. If Windows dropped the snapshot, post changes nothing
and points at the hermes backup zip instead.
pre and post now need an elevated PowerShell.
Under multiplexing `get_provider_env` resolved a scoped miss (`get_env_value` -> None) by falling
through to a bare `os.getenv`, which is the LAUNCH profile's `.env`: a routed profile with no
EXA_API_KEY/PARALLEL_API_KEY of its own searched on another profile's key, and the (key, client)
cache in `cached_sdk_client` faithfully built it a client on that borrowed key. Independent-review
major on #120092 (pre-existing on main, but this PR's claim covered only the both-have-keys case).
The bare `os.getenv` rung now runs only when no profile secret scope is bound (stripped installs,
plain CLI, systemd-injected env keep working); with a scope bound a miss is "" and the provider
raises its usual missing-key error. Isolation is between profiles; no environ fallthrough on a
scoped miss.
Tests: the A->B->A row gains a no-key profile (refused, launch key never used) plus a
single-profile control; the openrouter video sibling gets the same absence row (already correct
on this PR's head, red on main).
Sibling sites of #119986's class: the openai and meta-ai image backends
resolved their API key through get_secret but the base URL through
os.environ, so on a multiplexed gateway a routed profile's key was sent to
the launch profile's endpoint. Both fields now come from the same scope
(get_secret_str, like the DeepInfra video backend after #119986).
The salvaged #120037 test keyed on mem0 rendering "unavailable" under
multiplex, but mem0's status is "unavailable" whenever mem0ai is not
installed (dependencies_installed=False short-circuits before any secret
read), so the test was red on head in a venv without the extra. Replace
it with a provider whose is_available() IS a scoped get_secret read
against the launch home's .env: red on base (UnscopedSecretError
swallowed -> available=False), green with the launch scope bound.
GET /api/cron/delivery-targets ran cron_delivery_targets() outside any
profile secret scope. Once the dashboard/desktop `serve` backend hosts a
second profile home (it flips multiplex active on the first `?profile=`
request), agent.secret_scope.get_secret fails closed, so every poll raised
UnscopedSecretError — caught, logged as an error, and the response silently
lost every configured platform, leaving only the implicit `local` entry.
Run the read inside _config_profile_scope, matching the sibling cron routes,
and thread an optional `profile` query param through this route and
GET /api/cron/blueprints, which shares the same call. Add a regression test
that reproduces the fail-closed read (red on base, green with the scope).
Regression test for the multiplex fail-closed path: with
set_multiplex_active(True), _merged_plugins_hub's unscoped
_discover_memory_provider_statuses call makes mem0's schema probe raise
UnscopedSecretError (swallowed by probe_availability) and the provider
renders 'unavailable'. The launch scope binding fixes it. Proven red on
upstream/main, green with the fix.
_merged_plugins_hub builds the provider picker through
_discover_memory_provider_statuses, which reads credentials via get_secret
(mem0's get_config_schema/is_available). On a multiplexed host an unscoped
read raises UnscopedSecretError; probe_availability swallows it and every
memory provider renders 'unavailable' in the dashboard with no user-visible
error. Wrap the discovery in launch_profile_scope_if_multiplexed (no-op on
single-profile hosts) so the dashboard's own profile resolves its secrets.
The OpenRouter video backend read OPENROUTER_API_KEY and OPENROUTER_BASE_URL
straight from os.environ. That broke two setups:
- A key added with `hermes auth add openrouter` (API key or OAuth) lives in
the credential pool, not the environment. Chat and image_gen/openrouter find
it through resolve_runtime_provider. video_gen reported OpenRouter
unavailable, and generate() returned missing_credentials.
- On a multiplexed gateway, os.environ holds the launch profile's .env. A
routed profile's video jobs were submitted, polled and downloaded with the
launch profile's key and billed to that account. A profile whose key lived
only in its own .env could not use the backend at all.
The backend now resolves (api_key, base_url) with
resolve_runtime_provider(requested="openrouter"), the same call
image_gen/openrouter makes. generate() resolves once and passes the pair to
submit, poll and download. With a round-robin pool, resolving per request
would poll with a different account's key than the one that created the job.
OpenAICompatibleVideoGenProvider, which the DeepInfra video backend uses, had
the same raw reads of <NAME>_API_KEY and <NAME>_BASE_URL. Both now go through
get_secret_str, as image_gen/deepinfra already does.
cached_sdk_client returned the client cached on tools.web_tools before it
read the key, so the Exa, Parallel and AsyncParallel clients kept the key
they were first built with for the life of the process. A key fixed in .env
and applied with /reload still sent the old key (401s until a restart), and
on a gateway serving multiplexed profiles every profile's Exa and Parallel
calls went out on whichever profile's key built the client first, billed to
that account. A key removed from the environment also kept being used.
Resolve the key on every call and reuse the cached client only when it was
built with that key; the slot now holds (key, client) as one value so two
builds racing under different keys cannot record one key beside the other
key's client. Firecrawl already compares its credential before reusing its
client; this brings the two SDK-backed providers in line.
The independent review caught a half-finished rewrite: the head parametrized
test_capped_drain_fits_inside_launchd_budget_minus_reserve over (platform, label, expected) but
its body ignored the params (three identical runs of the pure cap arithmetic), while _read_budget
— the intended parametrized body, referencing free names `platform`/`expected` — was never called.
Restore the pure arithmetic test unparametrized (as on main) and put the reader rows back in
test_launchd_reader_yields_a_budget_only_for_hermes_jobs, now driven from the process environment
end to end (environ -> launchd_service_label -> read_launchd_exit_timeout_s -> capped drain) with
`platform` as data instead of macos_only/linux_only host markers: the pre-PR direct-label rows
(ai.hermes budget, app-coalition label ignored, linux leak ignored) plus the wrapper-forwarded
HERMES_LAUNCHD_LABEL grandchild rows this PR adds (budget on darwin, app label still not ours,
ignored on linux) and the unlabelled grandchild. Sabotage (drop LAUNCHD_LABEL_ENV from
launchd_job_label): darwin-grandchild and the every-reader supervisor test go red; the direct/app/
linux rows are controls the base already satisfied and stay green.
Salvage trim of #119924: three tests became two (mail the gateway admits or
answers reaches it; mail it would ignore, or that forges From:, is dropped).
The bare-entry rows are drop cases and live in that table.
The salvaged #119607 taught only the drain cap (launchd_service_label) to read
HERMES_LAUNCHD_LABEL. The gateway grandchild under the generated plist also
reads XPC_SERVICE_NAME=0 in is_gateway_supervisor_process (exit-75 restart
route) and control_socket._detect_supervisor (identify payload), so one
process was "launchd" to the drain cap and "manual" to the restart route.
One seam: gateway.restart.launchd_job_label applies the ai.hermes predicate to
XPC_SERVICE_NAME then HERMES_LAUNCHD_LABEL; the drain cap, the restart route,
the control socket and the wrapper's export all call it. launchd_service_label
and read_launchd_exit_timeout_s take `platform` as data so the mapping is
tested on Linux without patching sys.platform (AGENTS.md: don't fake the host
OS); the salvaged tests are trimmed to two invariants each and lose their
monkeypatch of sys.platform.
Not live-run: this host is Linux, launchd is code-path + wrapper-subprocess
proof only.
Under the generated launchd plist the stderr-timestamp wrapper is the job
process and the gateway is a grandchild: launchd stamps XPC_SERVICE_NAME
only on the wrapper, so the gateway reads "0" there, launchd_service_label()
returns None, and the ExitTimeOut drain cap (aa0289f307) silently never
applies — the exact SIGKILL-mid-teardown shape it was added to prevent.
The wrapper now re-exports its ai.hermes.* job label to the child as
HERMES_LAUNCHD_LABEL, and launchd_service_label() falls back to it when
XPC_SERVICE_NAME carries no usable label. Foreground/unsupervised starts
(no label to export) and app-coalition labels keep failing open exactly as
before.
Fixes#119598
EmailAdapter._sender_accepted runs before any MessageEvent exists and
read only EMAIL_ALLOWED_USERS. Unset, it dropped every sender unless
allow-all was on; set, it dropped everyone not listed. The gateway's
own handling therefore never ran for email:
platforms.email.unauthorized_dm_behavior "pair" (the setup wizard's
"Use DM pairing") and "decline" sent nothing, and a sender admitted by
GATEWAY_ALLOWED_USERS or an approved pairing was dropped. bb304b4914
turned the empty-allowlist branch into drop-all after #50568 had made
"pair" email's explicit opt-in.
The gate now keeps a sender listed by address in EMAIL_ALLOWED_USERS
or GATEWAY_ALLOWED_USERS, a sender the registered gateway
authorization check admits (that is the only reader of the pairing
store), and, under an explicit pair or decline, an unknown sender the
gateway will answer. The default "ignore" still drops unknown senders
before a MessageEvent exists, so the mail-loop guard from fd9c32c0f2
holds.
Three guards keep the wider gate from widening access, and close two
forged-From: paths main already had:
- A sender admitted only so the gateway can answer it (pair or
decline) must authenticate its From:, open access or not: the
pairing code or refusal is mailed back to that address. A granted
sender still needs it short of open access, since a pairing grant
keys on From: just as the allowlist does. Open access follows the
gateway's own order: EMAIL_ALLOW_ALL_USERS wins over a list, while
GATEWAY_ALLOW_ALL_USERS beside a list admits nobody extra, so it no
longer exempts a listed address from From: authentication either
(on main a forged From: of a listed address got through there).
- Open access comes from the gateway's own verdict when a check is
registered. GATEWAY_ALLOW_ALL_USERS beside a GATEWAY_ALLOWED_USERS
list grants a stranger nothing there, so the env flag alone no
longer exempts one from From: authentication (that path mailed a
pairing code to a forged From: on main too).
- A sender whose local part alone matches an allowlist entry is
dropped. The gateway's check also matches an address by its bare
local part (#119446), so without this, GATEWAY_ALLOWED_USERS=alice
(a chat username) would admit or pair alice@<any domain>. The lists
are parsed as the gateway parses them, JSON list literals included,
or '["alice"]' would slip past this guard.
_allowlist_in_effect only served the old condition and is removed.
The scope tests now assert the same scoped reads through
_sender_accepted, with GATEWAY_ALLOWED_USERS covered as well.
Measured end to end with the real GatewayRunner callback wired
(adapter -> gateway ingress):
- pair, decline, GATEWAY_ALLOWED_USERS and an approved pairing each
went from 0 events reaching the gateway to 1. pair mails a pairing
code, decline mails one refusal.
- An unauthenticated From: in pair mode, for a paired address or for a
GATEWAY_ALLOWED_USERS address still reaches nothing.
- A bare GATEWAY_ALLOWED_USERS=stranger entry lets nothing from
stranger@<domain> through, under ignore or pair. Without the
local-part guard that mail reached the gateway in both.
- The same holds for a JSON-literal list, and a pair-mode stranger
with a forged From: under allow-all beside an EMAIL_ or
GATEWAY_ALLOWED_USERS list reaches nothing.
- The default still drops.
The analytics routes answer a corrupt state.db with a 503 whose detail is an
object, {error: "state_db_corrupt", message: "state.db corrupt — run `hermes
doctor` ...", path}. extractDetail only accepted a string or a validation list
under detail/message/error, so it returned null and apiErrorFromResponse fell
back to humanizeStatus(503): "The Hermes service is not ready yet. Try again in
a moment." The Models page (every visit) and the Analytics page (with token
analytics on) told the operator to wait, which never helps, and dropped the
one instruction that does.
extractDetail now reads the string `message` of an object value, so any
structured detail surfaces its sentence, not its machine code. String and
validation-list details are unchanged, and an object without a message still
falls back to the status sentence.
gateway_state.json keeps platform entries across restarts, so a platform
removed since (here Feishu, last written by a gateway process from a week
earlier) stays "connected" on disk. The topology already filters its
platform map by writer identity (_owned_profile_platforms), but ports were
computed from the raw map with only dead states dropped, so gateways[].ports
advertised a port nothing listens on. Ports now come from the owned entries.
On macOS, `hermes gateway install --no-start-now` started the gateway anyway.
`_cmd_install` forwarded the start flags to the systemd and Windows backends but
called `launchd_install(force)` alone, and `launchd_install` always ran
`launchctl bootstrap`. The plist sets RunAtLoad, so bootstrapping it starts the
gateway immediately, and the command still printed "Service installed and
loaded!". The setup wizard had the same gap: answering No to "Start the gateway
now?" and Yes to login auto-start still called `launchd_install(force=False)` and
started the gateway on the spot.
launchd_install now takes start_now. When it is False and launchd is not already
running the gateway, the install writes the plist and does not load it. It also
boots out any idle registration left from before, such as a job parked after a
clean exit, because `hermes gateway start` would kickstart that registration's
old definition instead of loading the new plist. The outdated-plist repair takes
the same path, since its bootout/bootstrap reload would start a stopped gateway.
A gateway that launchd already runs is reloaded as before, not stopped. With
the plist in ~/Library/LaunchAgents, the gateway starts at the next login or on
`hermes gateway start`. `_cmd_install` and the wizard now pass the answer
through.
Measured with launchctl recorded rather than run: `install --no-start-now` went
from 1 bootstrap to 0, the wizard's No went from 1 bootstrap to 0, and the
repair of an outdated plist with a stopped gateway went from a reload to a
rewrite only. The --no-start-on-login half on launchd is #91549 and is not
touched here.
* refactor(fallback): share the pinned-owner chain rule
delegate_task's _resolve_child_fallback_chain decides which fallback chain
a child may walk: a pinned child never borrows the parent chain, an explicit
[] disables fallback, a declared list is the child's own. Cron needs the
same rule for pinned jobs (#100437), so the body moves to
hermes_cli.fallback_config.scoped_fallback_chain and the delegation helper
becomes a thin caller. Behaviour is unchanged; the delegation matrix test
still pins every cell.
* fix(cron): a pinned job never falls back to the global chain
A job with its own provider, model or base_url is an explicit operator pin
(since 0469740ab3 unpinned jobs store none of these). It still walked the
global fallback_providers chain in two places, so a pinned job could run
on a different provider and model than the one chosen:
- _resolve_job_runtime walked the chain on an AuthError or transient
network failure while resolving the pinned primary;
- _resolve_cron_agent_setup handed the global chain to every cron agent as
fallback_model, so the conversation loop's provider ladder could swap a
pinned job mid-run.
Both now read _job_fallback_chain(job, cfg), which returns no chain for a
pinned job through the same scoped_fallback_chain rule delegate_task uses
for pinned children. The pre-dispatch key check reads it too: the global
chain used to skip that check for every job, so a pinned job with a
missing key now blocks before the agent is built instead of failing in the
resolver. The transient-failure notice for a pinned job says it does not
fall back and names --unpin, instead of "No backup provider succeeded".
Unpinned jobs (including legacy *_snapshot records) and same-provider
credential-pool rotation are unchanged. The two scheduler tests that
asserted atomic provider+model fallback swaps used pinned jobs; they now
use unpinned jobs and keep the same assertions.
No per-job fallback_providers list: jobs have no generic override field
(create_job/update_job, the cronjob tool schema and the CLI enumerate each
field), so an opt-in chain would be a new surface on all of them. The
escape hatch is to leave the job unpinned and pick its model with
cron.model / cron.model_provider.
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
* docs(cron): pinned jobs do not use fallback_providers
cron.md "Provider recovery" and the pre-dispatch key check, the cron rows
and section in fallback-providers.md, and the developer notes in
cron-internals.md / provider-runtime.md said every cron job inherits the
global chain. State the new rule, the compatibility note for users who
relied on a pinned job landing on the chain, and the unpinned + cron.model
alternative.
---------
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
_for_each_served_profile ran a housekeeping body once per served profile
with no boundary between them, and _housekeeping_chore only catches at the
tick level. One profile's unreadable store or broken .env therefore ended
the loop, and every profile after it lost its state.db archive/prune,
curator, skill-sync and MCP reconcile pass on every tick.
That state is reachable: _init_session_db tolerates a failed launch store
and keeps running, and hermes serve defers each served profile's
auto-archive to this loop (#117746), so a satellite behind a broken launch
store had no sweeper at all.
Each profile now gets its own try/except, logged at debug like
_housekeeping_chore.
Co-authored-by: Baris Sencan <b.sencan@equalsmoney.com>
* fix(state): publish structural state.db corruption as one profile-level state
A structurally corrupt state.db showed up differently on every surface: the
sidebar endpoint returned 200 with empty slices plus an errors row, /api/sessions
returned 500, /api/status said components.storage ok and readiness was green.
None of them said the store was damaged, so Desktop rendered it as deleted
history (#72046).
hermes_state_health is now the single latch, keyed by resolved state.db path:
- SessionDB._halt_db_corrupt, the SessionDB read helpers, the web profile
reader and the readiness probe publish into it, only for structural
corruption (not FTS-scoped damage, not the malformed-schema case the web
open path heals).
- gateway.readiness reports it (state_db degraded/corrupt, and session_store
unavailable/corrupt even when the handle cache says ok), which also feeds
/api/status components.storage (now with reason: corrupt).
- /api/sessions, /api/profiles/sessions and /api/profiles/sessions/sidebar
carry storage: {profile: "corrupt"}; /api/sessions returns 503
state_db_corrupt instead of 500.
- A peer SessionDB handle in the same process refuses writes on a latched
path with the existing StateDbCorruptError, so gateway/agent transcript
diversion and classify_persistence_error keep working unchanged.
The latch never clears on its own and resets on restart, the recovery boundary
StateDbCorruptError already documents.
Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>
* fix(desktop): say the session store is damaged instead of an empty sidebar
The sidebar reads the list endpoints' new storage map into
$corruptSessionStores and renders a persistent destructive Alert above the
session list naming the affected profile(s). The copy says missing chats were
not deleted and points at the non-destructive path (quit Hermes, then
`hermes sessions recover --source <state.db> --inspect-only` or restore a
snapshot) plus the recovery guide; it does not recommend `sessions repair`
for structural damage.
Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>
---------
Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>
Review follow-up on the host-rendezvous isolation:
* Excluding Desktop-owned children from ROLE_SERVE broke #119644's terminal
path: `hermes plugins install` found "the running Desktop backend" only via
read_record(ROLE_SERVE), so on a Desktop-only box open chats no longer lit
up. A Desktop child now publishes ROLE_DESKTOP_SERVE (own lock, record and
0600 token). The attach/refuse ladder (_host_backend_attachment) still reads
ROLE_SERVE only, so a supervised public dashboard is never blocked by it;
notify_serve_backend prefers the host owner and falls back to the Desktop
record. /api/host/identity reports the role actually published.
* is_desktop_owned_backend() missed Desktop's SSH spawn — `env HERMES_DESKTOP=1
hermes serve --isolated --ssh-session-token-file F` carries NO token env var
(remote-lifecycle tests assert the var name never appears) — so the SSH child
still claimed ROLE_SERVE on the remote host and its MCP discovery flipped to
deferred. The predicate now accepts the token-FILE argv shape; the argv half
lives in _startup_fast.is_desktop_ssh_backend_argv (stdlib-only, importable
before main.py's import wall) and replaces the two duplicate substring checks
in main.py and dashboard_procs.py.
* web_server.py's remaining three bare HERMES_DESKTOP reads (cron ticker,
managed-gateway teardown, orphan serve reap) route through the predicate;
the tests that modelled a Desktop child with the bare flag set the spawn
token, as a real pool child does.
Widen the two salvaged fixes to the whole class and make the refusal
supervisor-safe:
- `process_identity.is_desktop_owned_backend()` is the single discriminator
(HERMES_DESKTOP=1 AND the per-spawn HERMES_DASHBOARD_SESSION_TOKEN). The
attach bypass, the named-profile reroute, the env sanitizer, the MCP
discovery timing and the host-rendezvous publish all key on it now; a shell
that merely inherited the flag from Desktop is treated as a normal launch
(#119210). The publish skip from #119832 keyed on the bare flag, which
would have hidden a supervised service launched from a Desktop shell.
- `_attach_to_host_backend` refuses with exit 78 (EX_CONFIG) instead of 1.
Exit 1 under `Restart=always` was an infinite restart loop with nothing
listening on the ingress port (2035 restarts, #119824); 78 is the
deliberate-refusal code `RestartPreventExitStatus=78` parks on, the same
contract the gateway unit already uses.
- Docs: the systemd example carries RestartPreventExitStatus=78 and the
user-side remedy (stop the owner, or `--isolated`).
- Tests trimmed to ≤2 invariants per fix; the control case in the desktop
tests sets the token so it models a real pool child.
When submit cannot prove the pane's runtime owns the selected stored session
(reverse binding lost to eviction, reconnect, or a compression rotation the
selection never followed), it resumes the stored session and continues on
the runtime id that session.resume returns. That path pinned only
activeSessionIdRef. ChatView renders the $sessionStates slice named by
$activeSessionId, so the optimistic prompt, the reply, and every later turn
landed in a slice the pane never painted, while the legacy $messages mirror
(which the ref drives) looked correct. The chat stayed frozen until a relaunch.
Rebind the atom together with the ref, as the session-not-found recovery in
the same file already does, and carry the transcript the pane was showing
into the resumed runtime's empty slice (session.resume omits messages) when
both name the same conversation by lineage. A pane runtime that belongs to a
different stored session is never carried over.
Refs #71733
Refs #117867
archive_and_compact takes the newest tail_count durable rows as the carried tail's superseded
originals (active=0, compacted=0: hidden from display and session_search). The CLI and the gateway
persist a turn's user row only after the turn-start preflight, so at that compaction the row rides
in the carried tail with no durable original of its own. The positional rewind then reached one
row past the tail and flagged the newest summarized message as a superseded duplicate. Each
turn-start compaction hid one more (A5, then A16 in a two-compaction run), on the default config.
The TUI/Desktop are immune because they persist the user row at submit.
tail_count now leaves out this turn's rows that never reached state.db: no persisted marker and
no _row_id, counted within the carried tail window only. That includes the user row and any
unflushed scaffolding. It does so only while the turn holds the session turn lease. Between turns
(a manual /compress, the gateway's pre-turn hygiene) the turn anchor is left over from the last
turn, and a gateway transcript reload is durable but unmarked. Counting those rows would leave
carried originals at compacted=1, the duplicate recall #86366 fixed.