Commit Graph

40736 Commits

Author SHA1 Message Date
kshitijk4poor
8646ddc19a chore: map contributor emails for the perf salvage stack 2026-09-23 21:26:49 +05:30
hermes-seaeye[bot]
6cd727bbfa fmt(js): npm run fix on merge (#120400)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-23 15:53:24 +00:00
hermes-seaeye[bot]
8191aa0660 fmt(js): npm run fix on merge (#120392)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-23 15:46:42 +00:00
doresa0
03544a73be fix(desktop): keep reasoning settings label on one line
Prevent short CJK labels like 推理 from wrapping awkwardly next to the
effort select in Model settings.

Closes #83118
2026-09-23 10:38:40 -05:00
teknium1
38c9611791 test(gateway): served-profile stop parks under the host (rebase over #120143)
test_pooled_served_profile_backend_unscoped pinned the pre-parking contract (stop on a served
profile refuses). With per-profile parking, stop on a served profile is the parking verb and
returns no refusal; start while unparked still refuses. Same one-line contract change already
made in test_gateway_multiplex_served_record.py.
2026-09-23 08:25:28 -07:00
teknium1
9ab2ad3d3f docs(gateway): parked vs standalone precedence; gap (1) of the standalone shim is closed by parking 2026-09-23 08:25:28 -07:00
teknium1
f6ce02eafa fix(gateway): parking composes with gateway.standalone and the dashboard Stop/Start twin
gateway.standalone wins over the parked marker: profile_lifecycle() returns
False for an opted-out profile, so `-p X gateway stop|start` keeps addressing
X's own gateway process and never writes gateway.parked, even while a stale
host record still lists X. Every installed-roster caller now threads both
kwargs (`include_standalone=True, include_parked=True`): the host-attach peer
walk and the standalone boot notice were reading the roster without parked
profiles.

Dashboard twin (#119886 class): `/api/gateway/stop?profile=X` on a served
profile no longer answers 409 — the spawned `hermes -p X gateway stop` parks
it; `/api/gateway/start` on a parked profile is allowed while a host
multiplexer is live (the child unparks it) and still refused when nothing can
serve it. Tests trimmed to the invariants: the marker-appearing case was a
subset of the boot-and-reconcile test.
2026-09-23 08:25:28 -07:00
Victor Kyriazakos
85d28d343a docs(gateway): stopping one profile without stopping the host
The stop, start and restart verbs on a profile served by the host multiplexer,
the gateway.parked marker (provisioning may pre-create it), reconcile timing,
and what parked means in gateway status.
2026-09-23 08:25:28 -07:00
Victor Kyriazakos
4c342c05de feat(gateway): stop, start and restart one profile under the host multiplexer
Under the host multiplexer every installed named profile came online because
its directory existed, and the only way to stop one profile's bots was to stop
the host, which stopped everyone's. Both are fleet-operator blockers.

`hermes -p X gateway stop` on a profile served by the host now parks it: it
writes `profiles/X/gateway.parked` first (so the 30s reconcile cannot re-add X
between the verb and the marker) and sends the new `unserve-profile` control
verb, which tears down X's adapters, reconnects and cron inside X's own scope
(the teardown `_unserve_profile` already used for deleted profiles). `start`
removes the marker and sends `serve-profile`, which runs the same add-path the
reconcile loop uses. `restart` cycles both without parking. The default
profile keeps today's whole-host meaning.

`profiles_to_serve()` skips parked profiles, so adapters, cron, ingress
membership and the served record follow from one chokepoint; roster callers
that mean "every installed profile" (plugin deps, Windows update, dashboard
listing and topology, migration inventory) pass `include_parked=True`.
Provisioning may pre-create the marker: an installed profile stays offline
until an operator starts it. Host boot logs one INFO per parked profile;
`gateway status` shows `parked (hermes -p X gateway start)`.

Tests: profiles_to_serve contract, control verbs through the runner
(round-trip and every refusal), CLI marker-before-socket ordering, reconcile
honouring the marker both ways, migration inventory retaining parked
profiles, two-home E2E through the real loaders.
2026-09-23 08:25:28 -07:00
teknium1
bd970b0588 fix(profiles): route-only launch pin; profile delete names its own unit under multiplex
Builds on tancou's #119129 (cherry-picked above): the pin now lives in
get_routing_process_hermes_home() and only the four routed-profile DECISIONS read it.
get_process_hermes_home()/get_hermes_home() keep following HERMES_HOME, so an env-only
home switch in a multiplexed process resolves as before.

- set_multiplex_active(True) pins the launch home only when no host pin exists, and
  set_multiplex_active(False) releases only the pin it created itself. A transient toggle
  (gateway_migrate._multiplex_read_mode, cron external-worker restore) no longer drops an
  embedding host's explicit pin_process_hermes_home(launch).
- profiles._cleanup_gateway_service binds set_hermes_home_override(profile_dir) beside the
  env write. Under the previous head, DELETE /api/profiles/<x> from a multi-profile dashboard
  resolved get_service_name() against the pinned launch home -> bare `hermes-gateway`, and
  disabled/stopped/unlinked the HOST multiplexer's unit. Same path serves rename_profile.

Tests (red on the previous head): explicit pin survives True->False; env readers follow the
env while pinned; two-home delete removes hermes-gateway-victim and leaves hermes-gateway.
2026-09-23 08:09:47 -07:00
tancou
c7c9c18ccf fix(profiles): pin the launch home so a mirrored HERMES_HOME cannot flip routed-profile decisions
Symptom: a host that serves several profiles from one process and mirrors
the active turn's profile into `os.environ["HERMES_HOME"]` for legacy
readers (Hermes WebUI does this on every chat turn, next to the
context-local override) makes every launch-home decision see the served
profile as the launch profile. Two profiles that both configure `atlassian`
with different credentials share whichever MCP connection came first: a
READ_ONLY_MODE=false profile ends up calling a read-only server
(nesquena/hermes-webui#7721). The same misjudgement leaves the launch
residue in the served profile's child env, seeds the launch profile's
bridged allow-all grant into the served profile's secret scope, and lets
the served profile's `terminal.*` config bridge into the shared process env.

Cause: four launch-home checks compare the task's override with
`get_process_hermes_home()`, which reads `HERMES_HOME` live:
`agent.secret_scope.serves_routed_profile` (keys the MCP ledger via
`_mcp_registry_scope`, #108352 / #111481, and the check_fn cache, #111151),
`agent.secret_scope._is_process_home`, `tools.environments.local._is_routed_home`
and `hermes_cli.env_loader._process_hermes_home`. Under the mirror the two
sides are equal for every turn.

Change: `hermes_constants.pin_process_hermes_home(path | None)` lets the
host record the home it serves as its own; `get_routing_process_hermes_home()`
returns the pin when set, else `get_process_hermes_home()`; the four checks
compare against it. The pin is deliberately NOT folded into
`get_process_hermes_home()`: `get_hermes_home()` falls back to it for tasks
carrying no override (MCP loop, spawners), and the host's mirror exists
precisely so those readers see the served profile. Only "is this task
routed / is this the launch home" changes. Unpinned, behaviour is
byte-for-byte the old one; hosts that never mutate `HERMES_HOME` need not
call it. `activate_multi_profile_hosting()` is not the seam for this: it
flips `get_secret` fail-closed process-wide and freezes the launch env,
which an embedding host cannot adopt as a bug fix.

Tests (2 invariants, parametrized over the four checks plus the MCP ledger
key; red on main, green here): pinned + mirrored env -> the served home is
routed and the launch home is not, the MCP key is `(home_key, name)`,
`get_process_hermes_home()` still follows the env var; never pinned or
pinned-then-cleared -> old semantics, including "a mirrored env var IS the
launch home". `tests/conftest.py` resets the pin per test so the
module-global cannot leak between files.

Live repro (WebUI + a stdio FastMCP server named `atlassian` in two
profiles, one gated by READ_ONLY_MODE): base -> one ledger key
`'atlassian'`, the write profile lists only the read-only tools; fixed ->
`(<read_home_key>, 'atlassian')` and `(<write_home_key>, 'atlassian')`,
each profile lists its own tools.

Docs: `gateway/AGENTS.md` § Profile scope (one launch-home identity) and the
isolation table in `website/docs/user-guide/multi-profile-gateways.md`.
Also maps the author e-mail under contributors/emails/ (attribution check).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-23 08:09:47 -07:00
teknium1
11d5c710fd test(i18n): a served profile's turn binds the home override, not a mutated HERMES_HOME
`test_language_is_per_profile_under_multiplex` switched profiles by rewriting
`os.environ["HERMES_HOME"]` after `set_multiplex_active(True)`. That is the T1
standalone contract (environ IS the profile); a multiplexed turn binds
`set_hermes_home_override`, and since the launch home is pinned at activation
(#119242) a later env mutation is deliberately ignored. The two assertions are
unchanged; only the profile-switch mechanism now matches production.
2026-09-23 08:09:47 -07:00
teknium1
64c1cba883 chore: map contributor email for salvaged kanban commit 2026-09-23 08:09:47 -07:00
John Paul Soliva
3e61247c55 fix(api-server): /v1/responses state belongs to the routed profile
A multiplexed API server mirrors every route under /p/<profile>/ and
authenticates each mirror with that profile's key, but it kept one
ResponseStore at the home it was constructed in. Conversation names are
client-chosen strings ("main", "my-project"), so another profile's key
could post `conversation: <name>`, receive that profile's transcript,
instructions and session id as its agent's context, become the
conversation's tip (the owner's next turn replayed the intruder's
messages), and GET or DELETE the owner's responses by id.

The adapter now resolves the store from the request's profile home, as
the SessionDB cache already does: the construction home keeps
self._response_store (and its response_store.db), every other routed
home gets its own <home>/response_store.db, opened on first use and
closed on disconnect. The stream state captures its store when the
request starts, so a snapshot written after the scope ends (disconnect)
still lands in the right one.

Rows a secondary profile wrote into the shared store before this change
stay there, visible to the construction home's profile only; they are
not migrated.
2026-09-23 08:09:47 -07:00
teknium1
c18b6904ad test(logging): keep two invariant cases for routed-profile redaction
The launch-opt-out case pins the routing handler's home binding; the .env case pins the
listener-thread .env read in _redact_enabled and also exercises the binding. The config.yaml
case duplicated the binding proof.

Salvaged from #120019
2026-09-23 08:09:47 -07:00
John Paul Soliva
d9199d8e8d fix(logging): a routed profile's own log files follow its own redaction policy
Under multiplex (and the Desktop/dashboard backend) every record is
formatted on the log QueueListener thread, after the profile scope that
produced it is gone. _ProfileRoutingFileHandler still writes a routed
profile's records to <profile>/logs/*.log, but RedactingFormatter ran
there with no home override, so _redact_enabled() returned the LAUNCH
profile's import-time snapshot and the vault scrub keyed on the launch
home. A launch profile with security.redact_secrets: false (or
HERMES_REDACT_SECRETS=false) therefore wrote every routed profile's
credentials raw into that profile's agent.log/errors.log/gateway.log,
and a routed profile's own opt-out was ignored.

The routing handler now binds the record's stamped home while a record
for another profile is formatted. Launch-profile records are untouched.
With no live secret scope (that thread has none), _redact_enabled reads
the profile's own .env for HERMES_REDACT_SECRETS, as its scope would;
without it the first call there cached a config-only answer for the
process.
2026-09-23 08:09:47 -07:00
John Paul Soliva
2e2fd69687 fix(gateway): a served profile's reconnect replays the owed restart notice and resumes deferred sessions
A multiplexed gateway defers two things for a platform that is offline at boot: the
planned-restart online notice (the marker keeps the target owed "for its reconnect", #112109)
and restart-interrupted sessions (left resume_pending "for the reconnect watcher"). The owed
set spans every served profile's home channel, and a served profile's session never falls
back to the default bot, so both wait on that profile's own reconnect.

Only the primary reconnect (_install_reconnected_adapter) acted on either.
_run_secondary_profile_reconnect published the adapter, redelivered failed obligations and
returned: the notice never went out and the marker outlived the outage, and the deferred
sessions stayed stranded until the freshness window aged them out.

The secondary reconnect now does what the primary does, through one shared helper for the
notice replay (still lock-serialized, delivered targets recorded, so a concurrent primary
replay cannot double-send).
2026-09-23 08:09:47 -07:00
teknium1
8932450779 fix(profiles): freeze the launch home once the process serves several profiles
Every "does this task serve a ROUTED home" decision (serves_routed_profile,
_is_process_home, _is_routed_home, env_loader._process_hermes_home) compares the
home override with get_process_hermes_home(), which read os.environ["HERMES_HOME"]
live. A host that mirrors the served profile into that env var per turn (hermes-webui)
made every served profile look like the launch one: MCP registry scope None, bare
cross-profile connection names, launch residue kept in served child envs, the launch
GATEWAY_ALLOW_ALL_USERS grant seeded into the served scope.

set_multiplex_active(True) now pins the launch home (hermes_constants.
pin_process_hermes_home; first pin wins, an embedding host may pin explicitly) and
get_process_hermes_home() returns the frozen value while multiplex is active.
Standalone hermes -p x gateway run (multiplex inactive) keeps following the env.
No os.environ fallthrough is added anywhere.

Closes #119242
2026-09-23 08:09:47 -07:00
teknium1
6dc6fb593a test(cron): own-profile bot-chat turn from a multiplexed tick runs in the ticking home
Regression guard for #119858. The fix itself is already on main: 3b0fe0cc2b pinned
the child HERMES_HOME to the override-aware source home and 786c0e3f9d moved the
spawn onto served_profile_child_env(target_home=home, inherit_credentials=True).
No existing test asserted the OWN-profile (no -p) leg A->B->A under multiplex; these
two do (red at 786c0e3f9dc~1, green on main).
2026-09-23 08:09:47 -07:00
teknium1
9267706d6e fix(kanban): board identity follows the bound profile override, not the launch env
One resolver, hermes_cli.profiles.current_profile_name(): the HERMES_HOME override
(a multiplexed cron tick or routed gateway turn) names the profile first; the
dispatcher's HERMES_PROFILE pin is consulted only when no override is bound; the
process home last. #112888 added the HERMES_HOME-derived fallback but kept the env
pin FIRST, so under a multiplexer whose launch process carries a HERMES_PROFILE the
served profile's writes were still re-labelled with the host's name. The same
resolver replaces the per-module copies in hermes_cli/kanban.py, kanban_specify.py,
cron/lifecycle_guard.py and the kanban notify-target default.

Tests trimmed to two invariants (A->B->A under the override; control pin + generic
absence).

Closes #119859
Supersedes #112888
2026-09-23 08:09:47 -07:00
max
253a956144 fix(kanban): persist active-profile identity in comment author / created_by when HERMES_PROFILE unpinned
Board records (comment author, task creator) were written as the generic
"worker" whenever the dispatcher did not pin HERMES_PROFILE, even though
the active Hermes profile was resolvable. Add _persisted_identity():
environment (HERMES_PROFILE_NAME / HERMES_PROFILE) first, else the active
profile derived from HERMES_HOME via hermes_cli.profiles, else "worker".
Wire it into _handle_comment author, _handle_create created_by, and the
own-comment skip filter in inject_new_comments_from_env so a worker's own
notes never re-enter its live turn as fake operator steering.

Identity is never taken from tool args: board records are injected into
future workers' prompts, so a caller-supplied author override could forge
an authoritative-looking directive (see #19713).

Tests: regression tests for env present, env absent with active profile,
and no profile; injection echo guard without env profile.
2026-09-23 08:09:47 -07:00
John Paul Soliva
4d3815d7c2 fix(gateway): keep a served profile's delivery-ledger rows in the launch state.db
A multiplexed gateway connects each served profile's adapter inside
_profile_runtime_scope(<profile home>). The receive loop an adapter starts
while connecting inherits that home override, so every final reply the bot
sends is recorded from it. The ledger resolved its path through
get_hermes_home(), which follows the override, and the rows landed in
profiles/<name>/state.db. The boot sweep (sweep_recoverable) and the boot
flood-timer arming (pending_retries) run in the launch context and open the
launch state.db, so they never saw those rows. A served bot's reply cut off by
a crash or SIGKILL between finalize and platform ACK was never redelivered, a
flood-refused reply that spanned a restart was never retried, and
resume_pending was not cleared for a session whose answer sat in the ledger.

The ledger is meant to be one shared store: the boot sweep already scopes rows
by (platform, adapter_profile), and the profile purge terminalizes rows in the
shared store. _db_path now resolves from get_process_hermes_home(), as the
gateway's other process-level files do (gateway.status). It deliberately skips
the get_hermes_home() fallback that lifecycle_ledger uses when HERMES_HOME is
unset: a default gateway started in the foreground has no HERMES_HOME, and that
fallback would follow the override again.

Rows an earlier build already wrote to a profile's state.db stay where they are.
2026-09-23 08:09:47 -07:00
teknium1
6b6c7f4a99 fix(web): a routed profile without an Exa/Parallel key is refused, not served on the launch key
Under multiplexing `get_provider_env` resolved a scoped miss (`get_env_value` -> None) by falling
through to a bare `os.getenv`, which is the LAUNCH profile's `.env`: a routed profile with no
EXA_API_KEY/PARALLEL_API_KEY of its own searched on another profile's key, and the (key, client)
cache in `cached_sdk_client` faithfully built it a client on that borrowed key. Independent-review
major on #120092 (pre-existing on main, but this PR's claim covered only the both-have-keys case).

The bare `os.getenv` rung now runs only when no profile secret scope is bound (stripped installs,
plain CLI, systemd-injected env keep working); with a scope bound a miss is "" and the provider
raises its usual missing-key error. Isolation is between profiles; no environ fallthrough on a
scoped miss.

Tests: the A->B->A row gains a no-key profile (refused, launch key never used) plus a
single-profile control; the openrouter video sibling gets the same absence row (already correct
on this PR's head, red on main).
2026-09-23 07:49:35 -07:00
teknium1
e6f47e6474 chore(contributors): map rmohid@rmohid.dev -> @Rmohid 2026-09-23 07:49:35 -07:00
teknium1
6bbbc65b45 fix(image_gen): resolve OpenAI and Meta image base URLs through the secret scope
Sibling sites of #119986's class: the openai and meta-ai image backends
resolved their API key through get_secret but the base URL through
os.environ, so on a multiplexed gateway a routed profile's key was sent to
the launch profile's endpoint. Both fields now come from the same scope
(get_secret_str, like the DeepInfra video backend after #119986).
2026-09-23 07:49:35 -07:00
teknium1
f686368b2f test(dashboard): prove the plugins-hub launch scope with a credentialed provider
The salvaged #120037 test keyed on mem0 rendering "unavailable" under
multiplex, but mem0's status is "unavailable" whenever mem0ai is not
installed (dependencies_installed=False short-circuits before any secret
read), so the test was red on head in a venv without the extra. Replace
it with a provider whose is_available() IS a scoped get_secret read
against the launch home's .env: red on base (UnscopedSecretError
swallowed -> available=False), green with the launch scope bound.
2026-09-23 07:49:35 -07:00
rmohid
ce4e90ec4c fix(cron): bind profile secret scope in the delivery-targets route
GET /api/cron/delivery-targets ran cron_delivery_targets() outside any
profile secret scope. Once the dashboard/desktop `serve` backend hosts a
second profile home (it flips multiplex active on the first `?profile=`
request), agent.secret_scope.get_secret fails closed, so every poll raised
UnscopedSecretError — caught, logged as an error, and the response silently
lost every configured platform, leaving only the implicit `local` entry.

Run the read inside _config_profile_scope, matching the sibling cron routes,
and thread an optional `profile` query param through this route and
GET /api/cron/blueprints, which shares the same call. Add a regression test
that reproduces the fail-closed read (red on base, green with the scope).
2026-09-23 07:49:35 -07:00
wangtao
2cad86502d fix(memory): recover catalog provider once per profile 2026-09-23 07:49:35 -07:00
wangtao
20552e1f5e fix(memory): resolve OAuth flow in requested profile scope 2026-09-23 07:49:35 -07:00
sambouli
7ec5c412e9 test(dashboard): plugins hub keeps memory providers available under multiplex
Regression test for the multiplex fail-closed path: with
set_multiplex_active(True), _merged_plugins_hub's unscoped
_discover_memory_provider_statuses call makes mem0's schema probe raise
UnscopedSecretError (swallowed by probe_availability) and the provider
renders 'unavailable'. The launch scope binding fixes it. Proven red on
upstream/main, green with the fix.
2026-09-23 07:49:35 -07:00
sambouli
09f76522dd fix(dashboard): bind the launch secret scope for the plugins-hub memory options
_merged_plugins_hub builds the provider picker through
_discover_memory_provider_statuses, which reads credentials via get_secret
(mem0's get_config_schema/is_available). On a multiplexed host an unscoped
read raises UnscopedSecretError; probe_availability swallows it and every
memory provider renders 'unavailable' in the dashboard with no user-visible
error. Wrap the discovery in launch_profile_scope_if_multiplexed (no-op on
single-profile hosts) so the dashboard's own profile resolves its secrets.
2026-09-23 07:49:35 -07:00
John Paul Soliva
bb3011a202 fix(video_gen): resolve OpenRouter and DeepInfra credentials per profile, not from os.environ
The OpenRouter video backend read OPENROUTER_API_KEY and OPENROUTER_BASE_URL
straight from os.environ. That broke two setups:

- A key added with `hermes auth add openrouter` (API key or OAuth) lives in
  the credential pool, not the environment. Chat and image_gen/openrouter find
  it through resolve_runtime_provider. video_gen reported OpenRouter
  unavailable, and generate() returned missing_credentials.
- On a multiplexed gateway, os.environ holds the launch profile's .env. A
  routed profile's video jobs were submitted, polled and downloaded with the
  launch profile's key and billed to that account. A profile whose key lived
  only in its own .env could not use the backend at all.

The backend now resolves (api_key, base_url) with
resolve_runtime_provider(requested="openrouter"), the same call
image_gen/openrouter makes. generate() resolves once and passes the pair to
submit, poll and download. With a round-robin pool, resolving per request
would poll with a different account's key than the one that created the job.

OpenAICompatibleVideoGenProvider, which the DeepInfra video backend uses, had
the same raw reads of <NAME>_API_KEY and <NAME>_BASE_URL. Both now go through
get_secret_str, as image_gen/deepinfra already does.
2026-09-23 07:49:35 -07:00
John Paul Soliva
707852d92f fix(web): rebuild the Exa and Parallel clients when their API key changes
cached_sdk_client returned the client cached on tools.web_tools before it
read the key, so the Exa, Parallel and AsyncParallel clients kept the key
they were first built with for the life of the process. A key fixed in .env
and applied with /reload still sent the old key (401s until a restart), and
on a gateway serving multiplexed profiles every profile's Exa and Parallel
calls went out on whichever profile's key built the client first, billed to
that account. A key removed from the environment also kept being used.

Resolve the key on every call and reuse the cached client only when it was
built with that key; the slot now holds (key, client) as one value so two
builds racing under different keys cannot record one key beside the other
key's client. Firecrawl already compares its credential before reusing its
client; this brings the two SDK-backed providers in line.
2026-09-23 07:49:35 -07:00
teknium1
38d86b0749 test(gateway): fold the launchd drain-cap reader rows into one parametrized end-to-end test
The independent review caught a half-finished rewrite: the head parametrized
test_capped_drain_fits_inside_launchd_budget_minus_reserve over (platform, label, expected) but
its body ignored the params (three identical runs of the pure cap arithmetic), while _read_budget
— the intended parametrized body, referencing free names `platform`/`expected` — was never called.

Restore the pure arithmetic test unparametrized (as on main) and put the reader rows back in
test_launchd_reader_yields_a_budget_only_for_hermes_jobs, now driven from the process environment
end to end (environ -> launchd_service_label -> read_launchd_exit_timeout_s -> capped drain) with
`platform` as data instead of macos_only/linux_only host markers: the pre-PR direct-label rows
(ai.hermes budget, app-coalition label ignored, linux leak ignored) plus the wrapper-forwarded
HERMES_LAUNCHD_LABEL grandchild rows this PR adds (budget on darwin, app label still not ours,
ignored on linux) and the unlabelled grandchild. Sabotage (drop LAUNCHD_LABEL_ENV from
launchd_job_label): darwin-grandchild and the every-reader supervisor test go red; the direct/app/
linux rows are controls the base already satisfied and stay green.
2026-09-23 07:48:32 -07:00
teknium1
0e973b5453 test(email): fold the bare-allowlist-entry cases into the drop invariant
Salvage trim of #119924: three tests became two (mail the gateway admits or
answers reaches it; mail it would ignore, or that forges From:, is dropped).
The bare-entry rows are drop cases and live in that table.
2026-09-23 07:48:32 -07:00
teknium1
67f385f2d8 fix(gateway): every launchd-identity reader accepts the wrapper-forwarded label
The salvaged #119607 taught only the drain cap (launchd_service_label) to read
HERMES_LAUNCHD_LABEL. The gateway grandchild under the generated plist also
reads XPC_SERVICE_NAME=0 in is_gateway_supervisor_process (exit-75 restart
route) and control_socket._detect_supervisor (identify payload), so one
process was "launchd" to the drain cap and "manual" to the restart route.

One seam: gateway.restart.launchd_job_label applies the ai.hermes predicate to
XPC_SERVICE_NAME then HERMES_LAUNCHD_LABEL; the drain cap, the restart route,
the control socket and the wrapper's export all call it. launchd_service_label
and read_launchd_exit_timeout_s take `platform` as data so the mapping is
tested on Linux without patching sys.platform (AGENTS.md: don't fake the host
OS); the salvaged tests are trimmed to two invariants each and lose their
monkeypatch of sys.platform.

Not live-run: this host is Linux, launchd is code-path + wrapper-subprocess
proof only.
2026-09-23 07:48:32 -07:00
liuhao1024
fa34336c92 fix(gateway): resolve the launchd drain cap's job label through the wrapper
Under the generated launchd plist the stderr-timestamp wrapper is the job
process and the gateway is a grandchild: launchd stamps XPC_SERVICE_NAME
only on the wrapper, so the gateway reads "0" there, launchd_service_label()
returns None, and the ExitTimeOut drain cap (aa0289f307) silently never
applies — the exact SIGKILL-mid-teardown shape it was added to prevent.

The wrapper now re-exports its ai.hermes.* job label to the child as
HERMES_LAUNCHD_LABEL, and launchd_service_label() falls back to it when
XPC_SERVICE_NAME carries no usable label. Foreground/unsupervised starts
(no label to export) and app-coalition labels keep failing open exactly as
before.

Fixes #119598
2026-09-23 07:48:32 -07:00
John Paul Soliva
b5a300fe34 fix(email): pairing, decline and gateway grants reach the gateway instead of dying in the adapter pre-gate
EmailAdapter._sender_accepted runs before any MessageEvent exists and
read only EMAIL_ALLOWED_USERS. Unset, it dropped every sender unless
allow-all was on; set, it dropped everyone not listed. The gateway's
own handling therefore never ran for email:
platforms.email.unauthorized_dm_behavior "pair" (the setup wizard's
"Use DM pairing") and "decline" sent nothing, and a sender admitted by
GATEWAY_ALLOWED_USERS or an approved pairing was dropped. bb304b4914
turned the empty-allowlist branch into drop-all after #50568 had made
"pair" email's explicit opt-in.

The gate now keeps a sender listed by address in EMAIL_ALLOWED_USERS
or GATEWAY_ALLOWED_USERS, a sender the registered gateway
authorization check admits (that is the only reader of the pairing
store), and, under an explicit pair or decline, an unknown sender the
gateway will answer. The default "ignore" still drops unknown senders
before a MessageEvent exists, so the mail-loop guard from fd9c32c0f2
holds.

Three guards keep the wider gate from widening access, and close two
forged-From: paths main already had:
- A sender admitted only so the gateway can answer it (pair or
  decline) must authenticate its From:, open access or not: the
  pairing code or refusal is mailed back to that address. A granted
  sender still needs it short of open access, since a pairing grant
  keys on From: just as the allowlist does. Open access follows the
  gateway's own order: EMAIL_ALLOW_ALL_USERS wins over a list, while
  GATEWAY_ALLOW_ALL_USERS beside a list admits nobody extra, so it no
  longer exempts a listed address from From: authentication either
  (on main a forged From: of a listed address got through there).
- Open access comes from the gateway's own verdict when a check is
  registered. GATEWAY_ALLOW_ALL_USERS beside a GATEWAY_ALLOWED_USERS
  list grants a stranger nothing there, so the env flag alone no
  longer exempts one from From: authentication (that path mailed a
  pairing code to a forged From: on main too).
- A sender whose local part alone matches an allowlist entry is
  dropped. The gateway's check also matches an address by its bare
  local part (#119446), so without this, GATEWAY_ALLOWED_USERS=alice
  (a chat username) would admit or pair alice@<any domain>. The lists
  are parsed as the gateway parses them, JSON list literals included,
  or '["alice"]' would slip past this guard.

_allowlist_in_effect only served the old condition and is removed.
The scope tests now assert the same scoped reads through
_sender_accepted, with GATEWAY_ALLOWED_USERS covered as well.

Measured end to end with the real GatewayRunner callback wired
(adapter -> gateway ingress):
- pair, decline, GATEWAY_ALLOWED_USERS and an approved pairing each
  went from 0 events reaching the gateway to 1. pair mails a pairing
  code, decline mails one refusal.
- An unauthenticated From: in pair mode, for a paired address or for a
  GATEWAY_ALLOWED_USERS address still reaches nothing.
- A bare GATEWAY_ALLOWED_USERS=stranger entry lets nothing from
  stranger@<domain> through, under ignore or pair. Without the
  local-part guard that mail reached the gateway in both.
- The same holds for a JSON-literal list, and a pair-mode stranger
  with a forged From: under allow-all beside an EMAIL_ or
  GATEWAY_ALLOWED_USERS list reaches nothing.
- The default still drops.
2026-09-23 07:48:32 -07:00
John Paul Soliva
92eaf21eb5 fix(web): show a structured error detail's message instead of the status sentence
The analytics routes answer a corrupt state.db with a 503 whose detail is an
object, {error: "state_db_corrupt", message: "state.db corrupt — run `hermes
doctor` ...", path}. extractDetail only accepted a string or a validation list
under detail/message/error, so it returned null and apiErrorFromResponse fell
back to humanizeStatus(503): "The Hermes service is not ready yet. Try again in
a moment." The Models page (every visit) and the Analytics page (with token
analytics on) told the operator to wait, which never helps, and dropped the
one instruction that does.

extractDetail now reads the string `message` of an object value, so any
structured detail surfaces its sentence, not its machine code. String and
validation-list details are unchanged, and an object without a message still
falls back to the status sentence.
2026-09-23 07:48:32 -07:00
John Paul Soliva
4935536e2a fix(dashboard): /api/status reports only the ports the live gateway binds
gateway_state.json keeps platform entries across restarts, so a platform
removed since (here Feishu, last written by a gateway process from a week
earlier) stays "connected" on disk. The topology already filters its
platform map by writer identity (_owned_profile_platforms), but ports were
computed from the raw map with only dead states dropped, so gateways[].ports
advertised a port nothing listens on. Ports now come from the owned entries.
2026-09-23 07:48:32 -07:00
John Paul Soliva
693e4aa12c fix(gateway): launchd install honours --no-start-now and the wizard's "start now" answer
On macOS, `hermes gateway install --no-start-now` started the gateway anyway.
`_cmd_install` forwarded the start flags to the systemd and Windows backends but
called `launchd_install(force)` alone, and `launchd_install` always ran
`launchctl bootstrap`. The plist sets RunAtLoad, so bootstrapping it starts the
gateway immediately, and the command still printed "Service installed and
loaded!". The setup wizard had the same gap: answering No to "Start the gateway
now?" and Yes to login auto-start still called `launchd_install(force=False)` and
started the gateway on the spot.

launchd_install now takes start_now. When it is False and launchd is not already
running the gateway, the install writes the plist and does not load it. It also
boots out any idle registration left from before, such as a job parked after a
clean exit, because `hermes gateway start` would kickstart that registration's
old definition instead of loading the new plist. The outdated-plist repair takes
the same path, since its bootout/bootstrap reload would start a stopped gateway.
A gateway that launchd already runs is reloaded as before, not stopped. With
the plist in ~/Library/LaunchAgents, the gateway starts at the next login or on
`hermes gateway start`. `_cmd_install` and the wizard now pass the answer
through.

Measured with launchctl recorded rather than run: `install --no-start-now` went
from 1 bootstrap to 0, the wizard's No went from 1 bootstrap to 0, and the
repair of an outdated plist with a stopped gateway went from a reload to a
rewrite only. The --no-start-on-login half on launchd is #91549 and is not
touched here.
2026-09-23 07:48:32 -07:00
Austin Pickett
e7bff4b6d8 fix(cron): a pinned job never falls back to the global fallback chain (#120312)
* refactor(fallback): share the pinned-owner chain rule

delegate_task's _resolve_child_fallback_chain decides which fallback chain
a child may walk: a pinned child never borrows the parent chain, an explicit
[] disables fallback, a declared list is the child's own. Cron needs the
same rule for pinned jobs (#100437), so the body moves to
hermes_cli.fallback_config.scoped_fallback_chain and the delegation helper
becomes a thin caller. Behaviour is unchanged; the delegation matrix test
still pins every cell.

* fix(cron): a pinned job never falls back to the global chain

A job with its own provider, model or base_url is an explicit operator pin
(since 0469740ab3 unpinned jobs store none of these). It still walked the
global fallback_providers chain in two places, so a pinned job could run
on a different provider and model than the one chosen:

- _resolve_job_runtime walked the chain on an AuthError or transient
  network failure while resolving the pinned primary;
- _resolve_cron_agent_setup handed the global chain to every cron agent as
  fallback_model, so the conversation loop's provider ladder could swap a
  pinned job mid-run.

Both now read _job_fallback_chain(job, cfg), which returns no chain for a
pinned job through the same scoped_fallback_chain rule delegate_task uses
for pinned children. The pre-dispatch key check reads it too: the global
chain used to skip that check for every job, so a pinned job with a
missing key now blocks before the agent is built instead of failing in the
resolver. The transient-failure notice for a pinned job says it does not
fall back and names --unpin, instead of "No backup provider succeeded".

Unpinned jobs (including legacy *_snapshot records) and same-provider
credential-pool rotation are unchanged. The two scheduler tests that
asserted atomic provider+model fallback swaps used pinned jobs; they now
use unpinned jobs and keep the same assertions.

No per-job fallback_providers list: jobs have no generic override field
(create_job/update_job, the cronjob tool schema and the CLI enumerate each
field), so an opt-in chain would be a new surface on all of them. The
escape hatch is to leave the job unpinned and pick its model with
cron.model / cron.model_provider.

Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>

* docs(cron): pinned jobs do not use fallback_providers

cron.md "Provider recovery" and the pre-dispatch key check, the cron rows
and section in fallback-providers.md, and the developer notes in
cron-internals.md / provider-runtime.md said every cron job inherits the
global chain. State the new rule, the compatibility note for users who
relied on a pinned job landing on the chain, and the unpinned + cron.model
alternative.

---------

Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
2026-09-23 10:42:18 -04:00
Austin Pickett
d2ef7db751 fix(gateway): one served profile's failing chore no longer skips the rest (#120268)
_for_each_served_profile ran a housekeeping body once per served profile
with no boundary between them, and _housekeeping_chore only catches at the
tick level. One profile's unreadable store or broken .env therefore ended
the loop, and every profile after it lost its state.db archive/prune,
curator, skill-sync and MCP reconcile pass on every tick.

That state is reachable: _init_session_db tolerates a failed launch store
and keeps running, and hermes serve defers each served profile's
auto-archive to this loop (#117746), so a satellite behind a broken launch
store had no sweeper at all.

Each profile now gets its own try/except, logged at debug like
_housekeeping_chore.

Co-authored-by: Baris Sencan <b.sencan@equalsmoney.com>
2026-09-23 10:19:39 -04:00
Austin Pickett
bd945ec384 fix(state): surface corrupt state.db as one degraded session-storage state (#120274)
* fix(state): publish structural state.db corruption as one profile-level state

A structurally corrupt state.db showed up differently on every surface: the
sidebar endpoint returned 200 with empty slices plus an errors row, /api/sessions
returned 500, /api/status said components.storage ok and readiness was green.
None of them said the store was damaged, so Desktop rendered it as deleted
history (#72046).

hermes_state_health is now the single latch, keyed by resolved state.db path:

- SessionDB._halt_db_corrupt, the SessionDB read helpers, the web profile
  reader and the readiness probe publish into it, only for structural
  corruption (not FTS-scoped damage, not the malformed-schema case the web
  open path heals).
- gateway.readiness reports it (state_db degraded/corrupt, and session_store
  unavailable/corrupt even when the handle cache says ok), which also feeds
  /api/status components.storage (now with reason: corrupt).
- /api/sessions, /api/profiles/sessions and /api/profiles/sessions/sidebar
  carry storage: {profile: "corrupt"}; /api/sessions returns 503
  state_db_corrupt instead of 500.
- A peer SessionDB handle in the same process refuses writes on a latched
  path with the existing StateDbCorruptError, so gateway/agent transcript
  diversion and classify_persistence_error keep working unchanged.

The latch never clears on its own and resets on restart, the recovery boundary
StateDbCorruptError already documents.

Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>

* fix(desktop): say the session store is damaged instead of an empty sidebar

The sidebar reads the list endpoints' new storage map into
$corruptSessionStores and renders a persistent destructive Alert above the
session list naming the affected profile(s). The copy says missing chats were
not deleted and points at the non-destructive path (quit Hermes, then
`hermes sessions recover --source <state.db> --inspect-only` or restore a
snapshot) plus the recovery guide; it does not recommend `sessions repair`
for structural damage.

Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>

---------

Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>
2026-09-23 10:17:46 -04:00
Xipong
5c6b56a866 fix(desktop): keep offset-drifted backfill rows in order (#119606)
Co-authored-by: Xipong <217837358+Xipong@users.noreply.github.com>
2026-09-23 14:12:13 +00:00
teknium1
746f7ea21b fix(dashboard): Desktop children publish their own host role; SSH spawn shape is Desktop-owned
Review follow-up on the host-rendezvous isolation:

* Excluding Desktop-owned children from ROLE_SERVE broke #119644's terminal
  path: `hermes plugins install` found "the running Desktop backend" only via
  read_record(ROLE_SERVE), so on a Desktop-only box open chats no longer lit
  up. A Desktop child now publishes ROLE_DESKTOP_SERVE (own lock, record and
  0600 token). The attach/refuse ladder (_host_backend_attachment) still reads
  ROLE_SERVE only, so a supervised public dashboard is never blocked by it;
  notify_serve_backend prefers the host owner and falls back to the Desktop
  record. /api/host/identity reports the role actually published.

* is_desktop_owned_backend() missed Desktop's SSH spawn — `env HERMES_DESKTOP=1
  hermes serve --isolated --ssh-session-token-file F` carries NO token env var
  (remote-lifecycle tests assert the var name never appears) — so the SSH child
  still claimed ROLE_SERVE on the remote host and its MCP discovery flipped to
  deferred. The predicate now accepts the token-FILE argv shape; the argv half
  lives in _startup_fast.is_desktop_ssh_backend_argv (stdlib-only, importable
  before main.py's import wall) and replaces the two duplicate substring checks
  in main.py and dashboard_procs.py.

* web_server.py's remaining three bare HERMES_DESKTOP reads (cron ticker,
  managed-gateway teardown, orphan serve reap) route through the predicate;
  the tests that modelled a Desktop child with the bare flag set the spawn
  token, as a real pool child does.
2026-09-23 07:09:45 -07:00
teknium1
9ca3b54e91 fix(dashboard): one desktop-owned-child predicate, exit 78 on host-owner refusal
Widen the two salvaged fixes to the whole class and make the refusal
supervisor-safe:

- `process_identity.is_desktop_owned_backend()` is the single discriminator
  (HERMES_DESKTOP=1 AND the per-spawn HERMES_DASHBOARD_SESSION_TOKEN). The
  attach bypass, the named-profile reroute, the env sanitizer, the MCP
  discovery timing and the host-rendezvous publish all key on it now; a shell
  that merely inherited the flag from Desktop is treated as a normal launch
  (#119210). The publish skip from #119832 keyed on the bare flag, which
  would have hidden a supervised service launched from a Desktop shell.
- `_attach_to_host_backend` refuses with exit 78 (EX_CONFIG) instead of 1.
  Exit 1 under `Restart=always` was an infinite restart loop with nothing
  listening on the ingress port (2035 restarts, #119824); 78 is the
  deliberate-refusal code `RestartPreventExitStatus=78` parks on, the same
  contract the gateway unit already uses.
- Docs: the systemd example carries RestartPreventExitStatus=78 and the
  user-side remedy (stop the owner, or `--isolated`).
- Tests trimmed to ≤2 invariants per fix; the control case in the desktop
  tests sets the token so it models a real pool child.
2026-09-23 07:09:45 -07:00
KoNit-K
269a982b64 fix(dashboard): require Desktop spawn credential for host bypass 2026-09-23 07:09:45 -07:00
KoNit-K
95d1faf291 fix(web): isolate desktop host rendezvous 2026-09-23 07:09:45 -07:00
Austin Pickett
8fb0fc6ae6 fix(desktop): a submit that resumes the selected session moves the chat to the resumed runtime (#120273)
When submit cannot prove the pane's runtime owns the selected stored session
(reverse binding lost to eviction, reconnect, or a compression rotation the
selection never followed), it resumes the stored session and continues on
the runtime id that session.resume returns. That path pinned only
activeSessionIdRef. ChatView renders the $sessionStates slice named by
$activeSessionId, so the optimistic prompt, the reply, and every later turn
landed in a slice the pane never painted, while the legacy $messages mirror
(which the ref drives) looked correct. The chat stayed frozen until a relaunch.

Rebind the atom together with the ref, as the session-not-found recovery in
the same file already does, and carry the transcript the pane was showing
into the resumed runtime's empty slice (session.resume omits messages) when
both name the same conversation by lineage. A pane runtime that belongs to a
different stored session is never carried over.

Refs #71733
Refs #117867
2026-09-23 10:06:51 -04:00