Inside the official image every named profile has an s6 slot that the container's boot
registers DOWN (multiplex-only). `implicit_multiplex_blocker` still called
`_host_supports_migration`, whose s6 branch refused unconditionally ("Restart the container"),
so a Hermes Cloud host with the key UNSET booted standalone after an in-place update and every
other profile's bot went silent, while an explicit `true` bypassed the guard and worked. The
guard was vetoing a second gateway that could not exist.
- Only a named slot that is UP is a blocker; registered-down slots never veto the default.
`hermes gateway migrate --multiplex` (and the `hermes update` hook) now fold an UP slot
in-process: `s6-svc -d` + `down` file, root slot (re)started, no container restart. Boot
and migration share ONE fold rule (`gateway_multiplex_s6.fold_named_slot_intent`).
- A multi-profile host that resolves standalone on a guard prints a boxed warning at gateway
start, in the `hermes update` summary, in `hermes gateway status`, and the dashboard
`/api/status` carries `multiplex_standalone_reason` with a banner. Single-profile installs
are not warned.
- The resolved unset→on default is written as `gateway.multiplex_profiles: true` into the
default profile's config.yaml (comment-preserving writer, once, never on a guard refusal).
The fast-mode docs list Claude Opus 4.8, Opus 5 and Opus 5.5. The old
"opus-5" substring matched any future Opus 5.x, and the API answers
`speed` on a model it doesn't list with an error.
One exact list in agent/model_metadata.py now backs both the wire gate
(anthropic_adapter) and the /fast toggle (hermes_cli/models.py). It
accepts vendor-prefixed, dotted and dated ids.
Hermes's 14-day `[tool.uv] exclude-newer` quarantine applies to Hermes's own
dependencies only (uv lock/sync, `hermes update`, LAZY_DEPS extras via
`ensure()`). A plugin's declared `python_dependencies` install under the
PLUGIN's policy: `install_specs(policy="plugin")` runs uv with `--no-config`
from any cwd, still inside the core constraints file.
Reverses item 3 of #118841, which ran the uv tier with cwd=<checkout> for
every install so the quarantine reached plugin deps from any cwd. That made
catalog re-pins floored on a <14-day release uninstallable (#120076:
"only hindsight-client<=0.9.2 is available"; #114530 held on the same gate).
Maintainer ruling (Teknium): "plugins dont have to abide by our 14 day rule
btw. They can have their own security policy on that. Only hermes'
dependencies themselves have to. We should recommend that they do this for
their plugins and we should give guidance to plugin devs that they should
though."
- tools/lazy_deps.py: INSTALL_POLICIES ("core" | "plugin"); `_uv_policy_args`
replaces `_uv_policy_cwd`; `_venv_pip_install(policy=)` defaults to core
(ensure/LAZY_DEPS), `install_specs(policy=)` defaults to plugin.
- hermes_cli/plugin_python_deps.py: `resolve()` passes policy="plugin".
- Docs: developer guide "Dependency security policy" section, catalog README
admission rule 9, AGENTS.md pinning policy — plugin authors are responsible
for their deps and strongly recommended to pin upper bounds, floor on the
oldest API-compatible version and run their own release quarantine
(`uv --exclude-newer` in their CI); operators can set UV_EXCLUDE_NEWER.
- Tests: the #118841 cwd test is replaced by two invariants — a plugin install
carries `--no-config` and no checkout cwd (red on base), a core lazy install
keeps the checkout cwd and no `--no-config`.
The classic CLI dock above the status bar painted subagents and background
processes only. A standing /goal (and whether it is active, parked or
paused) was visible only as a "⊙ goal N/M" status-bar segment, and prompts
waiting in /queue could not be seen at all while the agent worked.
The dock now paints the goal's status line on its top row and a Queue block
(count, first three previews, +N more) on its bottom rows, closest to the
input they will be sent from; F7's summary line carries "goal active|parked|
paused" and "N queued". Goal state comes only from the GoalManager the CLI
already holds for the current session, so the 1 Hz dock refresh never opens
state.db and never shows a previous session's goal after /new.
/queue now dispatches inline while the agent runs, honouring its
busy_policy="dispatch": typed mid-run it was queued as the raw command, so
"/queue X" only re-enqueued X after the turn and /queue list|rm|edit could
not inspect the queue until it had drained.
The "display index not backfilled" probe was spelled twice in
hermes_state_messages.py, and the copy in the delete fence tested only
display_order while _ensure_display_order tests display_order OR
display_identity. Hoist one _DISPLAY_INDEX_MISSING_SQL beside
_DISPLAY_ACTIVE_CLAUSE and use it at both sites. The fence now refuses
whenever the read path would have backfilled instead of projecting; on
every reachable delete the two probes agree (the read path backfills both
halves before the export snapshot exists), so this only tightens fail-closed.
hermes_state_timeline keeps its own probe: it carries a role slot.
_cmd_export re-derived "which formats are human-read transcripts" as
`format == "html" or only`; use SAVE_TRANSCRIPT_FORMATS instead. Equivalent
on every reachable path: md/qmd without --only never reach _collect_sessions
(they route to _export_markdown), and --only forces the transcript view.
Bind include_compacted once in a local `_one` instead of threading it.
The IOERR test now injects on "WITH page AS", the display CTE's own opener,
so reshaping the SQL cannot let _ensure_display_order's SELECT probe absorb
the failure and make the case pass vacuously. Comments that narrated
removed guards now state the current reason.
--delete-after-verified had three guards after the merge: #119935's pre-delete
message re-count, #120065's caller-side `previous != snapshot` compare between
exported items, and #120065's expected_display_messages compare inside
delete_session's BEGIN IMMEDIATE. Only the last one is race-free: it compares
the exact display rows against the store at DELETE time, so a same-count
rewrite, an append after the re-count, or inter-item drift is refused there.
The two caller-side checks run outside the transaction and can only ever
duplicate a verdict the fence already gives, so drop them and merge the
snapshots straight into expected_messages.
Verified with the S1 probe (--race): compacted_fires false, same_count_fires
false with the caller-side guards removed.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
In-place compaction is the default. It soft-archives every earlier row of
a session under the same id (active = 0, compacted = 1), and Desktop and
the dashboard still show those turns. The md/qmd export read the session
through export_session -> get_messages with the default live-only clause,
so it wrote only the compaction summary and the carried tail.
verify_export_file then compared the file with that same dict, and
delete_session removed every row of the session, including the archived
turns that never reached the file.
The md/qmd export now reads the display history (include_compacted), for
a single session and for --lineage logical. export_session and
export_session_lineage take include_compacted, off by default:
import_sessions inserts every message as live context, so the JSON export
and stranded-session adoption keep reading live rows only.
The other transcripts people read had the same hole without the delete:
`/save md` and `/save html` (CLI and gateway), `sessions export --format
html` (one session or all of them) and `--only user-prompts` each held 1
of 6 answers on a six-turn session after one compaction. They now read
the display history too (SAVE_TRANSCRIPT_FORMATS; export_all gains
include_compacted and reads per session then, since the display read
dedupes per session). `/save json`, JSONL and the dashboard's JSON export
stay live-only for the import reason above.
Before deleting, the verify step also re-counts the store's display rows
for every session the file covers and refuses on a mismatch. A message
that lands while the files are written, or a later export change that
reads a narrower view, now refuses the delete instead of being removed
unseen. Like the adoption retire loop, the re-count runs just before
delete_session, not inside its transaction. Rewind rows (undone turns,
the superseded originals of a carried tail) are still deleted without
being exported, as `hermes sessions delete` does: they are not part of
the history the session shows.
Measured through the real CLI on a session with 6 turns and one default
in-place compaction (15 rows, 13 shown): before, 3 messages were exported
and all 15 rows deleted, with answers 1-5 missing from the file; after,
13 messages are exported in display order, then deleted.
(cherry picked from commit adeaff1e33f1ae2b8a066ef2374bb29ade50b3b8)
_disk_serve_tier owns the ollama clamp now, and get_model_capabilities /
get_model_info already treat config=None as load-it-yourself.
Six config=config pass-through ternaries in agent/models_dev.py whose callees
have no test double (_configured_catalog_provider, _models_dev_id,
_get_provider_models) collapse to a single call; the _cfg_get and
_load_model_overrides ternaries stay because tests replace those functions.
The parallel prefetch classified any cache entry past
_PROVIDER_MODELS_CACHE_TTL as needing a fetch and blocked the picker on
the thread pool until every one returned. But cached_provider_model_ids()
has two non-blocking tiers, not one: past the TTL and inside
_PROVIDER_MODELS_STALE_SERVE_MAX it still returns the cached list
immediately and revalidates off-thread. Prefetching those slugs traded a
non-blocking serial call for a blocking parallel one.
Since _PROVIDER_MODELS_STALE_SERVE_MAX is far longer than
_PROVIDER_MODELS_CACHE_TTL, every picker open more than a TTL after the
previous one paid for it: locally, 26 providers with a cache 60s past TTL
took 3.1-4.3s, bounded only by the slowest provider's round-trip. Gating
on usability instead of freshness brings that to 0.45s.
The gate now keys on the row the serial call actually reads via
_normalized_cache_slug: a bare "ollama" stays its own cache key rather
than folding into "custom", because the local native catalog has its own
TTL and its own empty-is-authoritative rule. And an empty catalog counts
as servable only for ollama inside _OLLAMA_LOCAL_MODELS_CACHE_TTL —
cached_provider_model_ids returns it with no round-trip, so prefetching
it is redundant work the picker waits on. Past that TTL an empty row gets
no stale-serve window and really does block, so it stays in the prefetch.
Entries the serial path genuinely cannot serve — missing, fingerprint
mismatched, or past the stale-serve window — still prefetch in parallel,
so a cold cache is unaffected. refresh=True already skips the prefetch
entirely, so explicit refresh still forces every provider.
python -m pytest tests/hermes_cli/test_model_cache_parallel_prefetch.py -q
19 passed
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
(cherry picked from commit fb8954a37816002bc5db9a592ebc2d8f1af87673)
Widen #117979 to the sibling routes in the same router: /api/files/read
(inline read_bytes+b64encode), /api/fs/read-text (bounded prefix read) and
/api/fs/read-data-url (_fs_read_bytes+b64encode) all still ran their file
I/O on the loop thread. Reuse the same helper / asyncio.to_thread pattern so
a slow disk or large bounded payload no longer stalls every other request.
Same class as the config/manifest loader swap earlier in this stack. The
/ready probe re-read and pure-Python-parsed config.yaml on every poll, inside
the aiohttp handler (110 KB seeded config: 2332 -> 487 ms per probe on a
loaded runner). It now uses utils.load_yaml_file_readonly: the C loader, and
repeat probes reuse the parse until the file signature changes. Parse errors
are not cached, so an edit that breaks or fixes the file shows on the next
probe. The bundled-platform manifest reader was the one startup manifest
reader left on yaml.safe_load. managed_scope imports fast_safe_load at module
level next to file_signature instead of lazily.
_count_skills() captured `now` BEFORE walking the skills dir and stored it as the
cache timestamp, so a scan longer than the TTL published an already-expired entry
(the next caller re-walked immediately). Concurrent direct callers (profile info,
dump, non-lazy list) could also each run the same cold walk.
Stamp the cache entry after the walk, and take a per-skills-dir scan lock with a
re-check under it so overlapping callers for the same profile share one walk while
different profiles' background scans still run in parallel.
Partial salvage of #107137: kept the stamp-after-walk + locked double-check hunk,
replaced the single global _SKILL_COUNT_SCAN_LOCK with a per-key lock so N profiles'
scans don't serialize, and trimmed the threading.Event test file to one invariant
test (published timestamp >= scan end).
(cherry picked from commit 30aba1f737)
(cherry picked from commit c1ec369828)
utils.fast_safe_load already exists, is pinned by tests/test_fast_safe_load.py, and
its comment names exactly these payers: 'startup parses config.yaml and every plugin
manifest, so the slow path cost ~0.9 s of cold start'. The migration was started —
hermes_cli/config.py uses it eight times, hermes_cli/main.py and hermes_cli/plugins.py
too — but the file-level loaders it was written for were never converted.
The cost is config SIZE, and the size is the installer's doing: it seeds config.yaml
by copying cli-config.yaml.example, 120,897 bytes of mostly comments. Nothing caches
load_gateway_config() and it has 238 production call sites.
Profiled before assuming a cause — reader.forward 34 ms, scanner.scan_to_next_token
28 ms, reader.peek 13 ms: the pure-Python PyYAML scanner, nothing else.
Measured A/B on one realistic pass (gateway config + every bundled plugin
description), medians of five runs, __pycache__ cleared between arms:
48.8 ms -> 2.46 ms. Per path: load_gateway_config 48.3 -> 1.84 ms, managed
config.yaml 44.4 -> 1.03 ms, 105 plugin.yaml manifests 59.3 -> 5.9 ms.
Same parse, same restricted tag set, same result — only the loader changes. Drops the
three 'import yaml' statements the swap orphaned.
(cherry picked from commit a4efd6a4506043159310eba02c32b673c455b88b)
gateway.standalone wins over the parked marker: profile_lifecycle() returns
False for an opted-out profile, so `-p X gateway stop|start` keeps addressing
X's own gateway process and never writes gateway.parked, even while a stale
host record still lists X. Every installed-roster caller now threads both
kwargs (`include_standalone=True, include_parked=True`): the host-attach peer
walk and the standalone boot notice were reading the roster without parked
profiles.
Dashboard twin (#119886 class): `/api/gateway/stop?profile=X` on a served
profile no longer answers 409 — the spawned `hermes -p X gateway stop` parks
it; `/api/gateway/start` on a parked profile is allowed while a host
multiplexer is live (the child unparks it) and still refused when nothing can
serve it. Tests trimmed to the invariants: the marker-appearing case was a
subset of the boot-and-reconcile test.
Under the host multiplexer every installed named profile came online because
its directory existed, and the only way to stop one profile's bots was to stop
the host, which stopped everyone's. Both are fleet-operator blockers.
`hermes -p X gateway stop` on a profile served by the host now parks it: it
writes `profiles/X/gateway.parked` first (so the 30s reconcile cannot re-add X
between the verb and the marker) and sends the new `unserve-profile` control
verb, which tears down X's adapters, reconnects and cron inside X's own scope
(the teardown `_unserve_profile` already used for deleted profiles). `start`
removes the marker and sends `serve-profile`, which runs the same add-path the
reconcile loop uses. `restart` cycles both without parking. The default
profile keeps today's whole-host meaning.
`profiles_to_serve()` skips parked profiles, so adapters, cron, ingress
membership and the served record follow from one chokepoint; roster callers
that mean "every installed profile" (plugin deps, Windows update, dashboard
listing and topology, migration inventory) pass `include_parked=True`.
Provisioning may pre-create the marker: an installed profile stays offline
until an operator starts it. Host boot logs one INFO per parked profile;
`gateway status` shows `parked (hermes -p X gateway start)`.
Tests: profiles_to_serve contract, control verbs through the runner
(round-trip and every refusal), CLI marker-before-socket ordering, reconcile
honouring the marker both ways, migration inventory retaining parked
profiles, two-home E2E through the real loaders.
Builds on tancou's #119129 (cherry-picked above): the pin now lives in
get_routing_process_hermes_home() and only the four routed-profile DECISIONS read it.
get_process_hermes_home()/get_hermes_home() keep following HERMES_HOME, so an env-only
home switch in a multiplexed process resolves as before.
- set_multiplex_active(True) pins the launch home only when no host pin exists, and
set_multiplex_active(False) releases only the pin it created itself. A transient toggle
(gateway_migrate._multiplex_read_mode, cron external-worker restore) no longer drops an
embedding host's explicit pin_process_hermes_home(launch).
- profiles._cleanup_gateway_service binds set_hermes_home_override(profile_dir) beside the
env write. Under the previous head, DELETE /api/profiles/<x> from a multi-profile dashboard
resolved get_service_name() against the pinned launch home -> bare `hermes-gateway`, and
disabled/stopped/unlinked the HOST multiplexer's unit. Same path serves rename_profile.
Tests (red on the previous head): explicit pin survives True->False; env readers follow the
env while pinned; two-home delete removes hermes-gateway-victim and leaves hermes-gateway.
Symptom: a host that serves several profiles from one process and mirrors
the active turn's profile into `os.environ["HERMES_HOME"]` for legacy
readers (Hermes WebUI does this on every chat turn, next to the
context-local override) makes every launch-home decision see the served
profile as the launch profile. Two profiles that both configure `atlassian`
with different credentials share whichever MCP connection came first: a
READ_ONLY_MODE=false profile ends up calling a read-only server
(nesquena/hermes-webui#7721). The same misjudgement leaves the launch
residue in the served profile's child env, seeds the launch profile's
bridged allow-all grant into the served profile's secret scope, and lets
the served profile's `terminal.*` config bridge into the shared process env.
Cause: four launch-home checks compare the task's override with
`get_process_hermes_home()`, which reads `HERMES_HOME` live:
`agent.secret_scope.serves_routed_profile` (keys the MCP ledger via
`_mcp_registry_scope`, #108352 / #111481, and the check_fn cache, #111151),
`agent.secret_scope._is_process_home`, `tools.environments.local._is_routed_home`
and `hermes_cli.env_loader._process_hermes_home`. Under the mirror the two
sides are equal for every turn.
Change: `hermes_constants.pin_process_hermes_home(path | None)` lets the
host record the home it serves as its own; `get_routing_process_hermes_home()`
returns the pin when set, else `get_process_hermes_home()`; the four checks
compare against it. The pin is deliberately NOT folded into
`get_process_hermes_home()`: `get_hermes_home()` falls back to it for tasks
carrying no override (MCP loop, spawners), and the host's mirror exists
precisely so those readers see the served profile. Only "is this task
routed / is this the launch home" changes. Unpinned, behaviour is
byte-for-byte the old one; hosts that never mutate `HERMES_HOME` need not
call it. `activate_multi_profile_hosting()` is not the seam for this: it
flips `get_secret` fail-closed process-wide and freezes the launch env,
which an embedding host cannot adopt as a bug fix.
Tests (2 invariants, parametrized over the four checks plus the MCP ledger
key; red on main, green here): pinned + mirrored env -> the served home is
routed and the launch home is not, the MCP key is `(home_key, name)`,
`get_process_hermes_home()` still follows the env var; never pinned or
pinned-then-cleared -> old semantics, including "a mirrored env var IS the
launch home". `tests/conftest.py` resets the pin per test so the
module-global cannot leak between files.
Live repro (WebUI + a stdio FastMCP server named `atlassian` in two
profiles, one gated by READ_ONLY_MODE): base -> one ledger key
`'atlassian'`, the write profile lists only the read-only tools; fixed ->
`(<read_home_key>, 'atlassian')` and `(<write_home_key>, 'atlassian')`,
each profile lists its own tools.
Docs: `gateway/AGENTS.md` § Profile scope (one launch-home identity) and the
isolation table in `website/docs/user-guide/multi-profile-gateways.md`.
Also maps the author e-mail under contributors/emails/ (attribution check).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One resolver, hermes_cli.profiles.current_profile_name(): the HERMES_HOME override
(a multiplexed cron tick or routed gateway turn) names the profile first; the
dispatcher's HERMES_PROFILE pin is consulted only when no override is bound; the
process home last. #112888 added the HERMES_HOME-derived fallback but kept the env
pin FIRST, so under a multiplexer whose launch process carries a HERMES_PROFILE the
served profile's writes were still re-labelled with the host's name. The same
resolver replaces the per-module copies in hermes_cli/kanban.py, kanban_specify.py,
cron/lifecycle_guard.py and the kanban notify-target default.
Tests trimmed to two invariants (A->B->A under the override; control pin + generic
absence).
Closes#119859
Supersedes #112888
GET /api/cron/delivery-targets ran cron_delivery_targets() outside any
profile secret scope. Once the dashboard/desktop `serve` backend hosts a
second profile home (it flips multiplex active on the first `?profile=`
request), agent.secret_scope.get_secret fails closed, so every poll raised
UnscopedSecretError — caught, logged as an error, and the response silently
lost every configured platform, leaving only the implicit `local` entry.
Run the read inside _config_profile_scope, matching the sibling cron routes,
and thread an optional `profile` query param through this route and
GET /api/cron/blueprints, which shares the same call. Add a regression test
that reproduces the fail-closed read (red on base, green with the scope).
_merged_plugins_hub builds the provider picker through
_discover_memory_provider_statuses, which reads credentials via get_secret
(mem0's get_config_schema/is_available). On a multiplexed host an unscoped
read raises UnscopedSecretError; probe_availability swallows it and every
memory provider renders 'unavailable' in the dashboard with no user-visible
error. Wrap the discovery in launch_profile_scope_if_multiplexed (no-op on
single-profile hosts) so the dashboard's own profile resolves its secrets.
The salvaged #119607 taught only the drain cap (launchd_service_label) to read
HERMES_LAUNCHD_LABEL. The gateway grandchild under the generated plist also
reads XPC_SERVICE_NAME=0 in is_gateway_supervisor_process (exit-75 restart
route) and control_socket._detect_supervisor (identify payload), so one
process was "launchd" to the drain cap and "manual" to the restart route.
One seam: gateway.restart.launchd_job_label applies the ai.hermes predicate to
XPC_SERVICE_NAME then HERMES_LAUNCHD_LABEL; the drain cap, the restart route,
the control socket and the wrapper's export all call it. launchd_service_label
and read_launchd_exit_timeout_s take `platform` as data so the mapping is
tested on Linux without patching sys.platform (AGENTS.md: don't fake the host
OS); the salvaged tests are trimmed to two invariants each and lose their
monkeypatch of sys.platform.
Not live-run: this host is Linux, launchd is code-path + wrapper-subprocess
proof only.
Under the generated launchd plist the stderr-timestamp wrapper is the job
process and the gateway is a grandchild: launchd stamps XPC_SERVICE_NAME
only on the wrapper, so the gateway reads "0" there, launchd_service_label()
returns None, and the ExitTimeOut drain cap (aa0289f307) silently never
applies — the exact SIGKILL-mid-teardown shape it was added to prevent.
The wrapper now re-exports its ai.hermes.* job label to the child as
HERMES_LAUNCHD_LABEL, and launchd_service_label() falls back to it when
XPC_SERVICE_NAME carries no usable label. Foreground/unsupervised starts
(no label to export) and app-coalition labels keep failing open exactly as
before.
Fixes#119598
gateway_state.json keeps platform entries across restarts, so a platform
removed since (here Feishu, last written by a gateway process from a week
earlier) stays "connected" on disk. The topology already filters its
platform map by writer identity (_owned_profile_platforms), but ports were
computed from the raw map with only dead states dropped, so gateways[].ports
advertised a port nothing listens on. Ports now come from the owned entries.
On macOS, `hermes gateway install --no-start-now` started the gateway anyway.
`_cmd_install` forwarded the start flags to the systemd and Windows backends but
called `launchd_install(force)` alone, and `launchd_install` always ran
`launchctl bootstrap`. The plist sets RunAtLoad, so bootstrapping it starts the
gateway immediately, and the command still printed "Service installed and
loaded!". The setup wizard had the same gap: answering No to "Start the gateway
now?" and Yes to login auto-start still called `launchd_install(force=False)` and
started the gateway on the spot.
launchd_install now takes start_now. When it is False and launchd is not already
running the gateway, the install writes the plist and does not load it. It also
boots out any idle registration left from before, such as a job parked after a
clean exit, because `hermes gateway start` would kickstart that registration's
old definition instead of loading the new plist. The outdated-plist repair takes
the same path, since its bootout/bootstrap reload would start a stopped gateway.
A gateway that launchd already runs is reloaded as before, not stopped. With
the plist in ~/Library/LaunchAgents, the gateway starts at the next login or on
`hermes gateway start`. `_cmd_install` and the wizard now pass the answer
through.
Measured with launchctl recorded rather than run: `install --no-start-now` went
from 1 bootstrap to 0, the wizard's No went from 1 bootstrap to 0, and the
repair of an outdated plist with a stopped gateway went from a reload to a
rewrite only. The --no-start-on-login half on launchd is #91549 and is not
touched here.
* refactor(fallback): share the pinned-owner chain rule
delegate_task's _resolve_child_fallback_chain decides which fallback chain
a child may walk: a pinned child never borrows the parent chain, an explicit
[] disables fallback, a declared list is the child's own. Cron needs the
same rule for pinned jobs (#100437), so the body moves to
hermes_cli.fallback_config.scoped_fallback_chain and the delegation helper
becomes a thin caller. Behaviour is unchanged; the delegation matrix test
still pins every cell.
* fix(cron): a pinned job never falls back to the global chain
A job with its own provider, model or base_url is an explicit operator pin
(since 0469740ab3 unpinned jobs store none of these). It still walked the
global fallback_providers chain in two places, so a pinned job could run
on a different provider and model than the one chosen:
- _resolve_job_runtime walked the chain on an AuthError or transient
network failure while resolving the pinned primary;
- _resolve_cron_agent_setup handed the global chain to every cron agent as
fallback_model, so the conversation loop's provider ladder could swap a
pinned job mid-run.
Both now read _job_fallback_chain(job, cfg), which returns no chain for a
pinned job through the same scoped_fallback_chain rule delegate_task uses
for pinned children. The pre-dispatch key check reads it too: the global
chain used to skip that check for every job, so a pinned job with a
missing key now blocks before the agent is built instead of failing in the
resolver. The transient-failure notice for a pinned job says it does not
fall back and names --unpin, instead of "No backup provider succeeded".
Unpinned jobs (including legacy *_snapshot records) and same-provider
credential-pool rotation are unchanged. The two scheduler tests that
asserted atomic provider+model fallback swaps used pinned jobs; they now
use unpinned jobs and keep the same assertions.
No per-job fallback_providers list: jobs have no generic override field
(create_job/update_job, the cronjob tool schema and the CLI enumerate each
field), so an opt-in chain would be a new surface on all of them. The
escape hatch is to leave the job unpinned and pick its model with
cron.model / cron.model_provider.
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
* docs(cron): pinned jobs do not use fallback_providers
cron.md "Provider recovery" and the pre-dispatch key check, the cron rows
and section in fallback-providers.md, and the developer notes in
cron-internals.md / provider-runtime.md said every cron job inherits the
global chain. State the new rule, the compatibility note for users who
relied on a pinned job landing on the chain, and the unpinned + cron.model
alternative.
---------
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
* fix(state): publish structural state.db corruption as one profile-level state
A structurally corrupt state.db showed up differently on every surface: the
sidebar endpoint returned 200 with empty slices plus an errors row, /api/sessions
returned 500, /api/status said components.storage ok and readiness was green.
None of them said the store was damaged, so Desktop rendered it as deleted
history (#72046).
hermes_state_health is now the single latch, keyed by resolved state.db path:
- SessionDB._halt_db_corrupt, the SessionDB read helpers, the web profile
reader and the readiness probe publish into it, only for structural
corruption (not FTS-scoped damage, not the malformed-schema case the web
open path heals).
- gateway.readiness reports it (state_db degraded/corrupt, and session_store
unavailable/corrupt even when the handle cache says ok), which also feeds
/api/status components.storage (now with reason: corrupt).
- /api/sessions, /api/profiles/sessions and /api/profiles/sessions/sidebar
carry storage: {profile: "corrupt"}; /api/sessions returns 503
state_db_corrupt instead of 500.
- A peer SessionDB handle in the same process refuses writes on a latched
path with the existing StateDbCorruptError, so gateway/agent transcript
diversion and classify_persistence_error keep working unchanged.
The latch never clears on its own and resets on restart, the recovery boundary
StateDbCorruptError already documents.
Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>
* fix(desktop): say the session store is damaged instead of an empty sidebar
The sidebar reads the list endpoints' new storage map into
$corruptSessionStores and renders a persistent destructive Alert above the
session list naming the affected profile(s). The copy says missing chats were
not deleted and points at the non-destructive path (quit Hermes, then
`hermes sessions recover --source <state.db> --inspect-only` or restore a
snapshot) plus the recovery guide; it does not recommend `sessions repair`
for structural damage.
Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>
---------
Co-authored-by: konsisumer <konsisumer@users.noreply.github.com>
Review follow-up on the host-rendezvous isolation:
* Excluding Desktop-owned children from ROLE_SERVE broke #119644's terminal
path: `hermes plugins install` found "the running Desktop backend" only via
read_record(ROLE_SERVE), so on a Desktop-only box open chats no longer lit
up. A Desktop child now publishes ROLE_DESKTOP_SERVE (own lock, record and
0600 token). The attach/refuse ladder (_host_backend_attachment) still reads
ROLE_SERVE only, so a supervised public dashboard is never blocked by it;
notify_serve_backend prefers the host owner and falls back to the Desktop
record. /api/host/identity reports the role actually published.
* is_desktop_owned_backend() missed Desktop's SSH spawn — `env HERMES_DESKTOP=1
hermes serve --isolated --ssh-session-token-file F` carries NO token env var
(remote-lifecycle tests assert the var name never appears) — so the SSH child
still claimed ROLE_SERVE on the remote host and its MCP discovery flipped to
deferred. The predicate now accepts the token-FILE argv shape; the argv half
lives in _startup_fast.is_desktop_ssh_backend_argv (stdlib-only, importable
before main.py's import wall) and replaces the two duplicate substring checks
in main.py and dashboard_procs.py.
* web_server.py's remaining three bare HERMES_DESKTOP reads (cron ticker,
managed-gateway teardown, orphan serve reap) route through the predicate;
the tests that modelled a Desktop child with the bare flag set the spawn
token, as a real pool child does.
Widen the two salvaged fixes to the whole class and make the refusal
supervisor-safe:
- `process_identity.is_desktop_owned_backend()` is the single discriminator
(HERMES_DESKTOP=1 AND the per-spawn HERMES_DASHBOARD_SESSION_TOKEN). The
attach bypass, the named-profile reroute, the env sanitizer, the MCP
discovery timing and the host-rendezvous publish all key on it now; a shell
that merely inherited the flag from Desktop is treated as a normal launch
(#119210). The publish skip from #119832 keyed on the bare flag, which
would have hidden a supervised service launched from a Desktop shell.
- `_attach_to_host_backend` refuses with exit 78 (EX_CONFIG) instead of 1.
Exit 1 under `Restart=always` was an infinite restart loop with nothing
listening on the ingress port (2035 restarts, #119824); 78 is the
deliberate-refusal code `RestartPreventExitStatus=78` parks on, the same
contract the gateway unit already uses.
- Docs: the systemd example carries RestartPreventExitStatus=78 and the
user-side remedy (stop the owner, or `--isolated`).
- Tests trimmed to ≤2 invariants per fix; the control case in the desktop
tests sets the token so it models a real pool child.
`hermes sessions repair-profiles --apply` copied a stranded session into
its owning profile's store verbatim, title included. Titles are unique
per store only (idx_sessions_title_unique), so when the target profile
already held an unrelated session with the same title, which is common
for generic auto-titles, the insert raised IntegrityError.
_MoveBatch.run imported every row before deleting any and had no
per-row handling. The exception escaped before the delete phase and
before the result was memoised. Rows copied earlier in the batch were
left in both stores. Every later finding re-ran the whole batch and hit
the same error. Every row in the batch was reported failed on every
run, the command exited 1 forever, and a non-colliding row ended up
duplicated across two profiles.
- import_moved_session gives a colliding title the moved row's id tail,
the same convention import_foreign_history uses, capped at
MAX_TITLE_LENGTH. The resident row keeps its name, so resolving it by
title is unchanged.
- _MoveBatch.run handles each row separately. A failed import or delete
is that row's failure alone, reported through apply(). The rest of
the batch still moves, and the batch runs once. The failed row's
lineage waits with it: its descendants are not imported without it,
and the parent it still points at in the source is not deleted. The
next run moves the lineage whole.
Measured on main: with two stranded rows and one title collision, both
rows failed on every run and one was left in both stores. With the fix,
one run moves both, and a second run finds nothing.
Gateway liveness reports a profile the multiplexer serves as running on
the multiplexer's PID (446f6f79a8). The dashboard read that answer as
"this profile runs its own gateway" in two places. multiplexed_profile_refusal
returned None for every served profile, so while a multiplexer was live
(the only time it matters) start/stop were never refused and a restart
was never rewritten to the multiplexer. The /api/status topology listed
one phantom gateway per served profile beside the host.
_has_own_gateway() answers the actual question: a live gateway whose PID
is not the multiplexer serving the profile. The existing lifecycle test
stubbed _check_gateway_running to False, which is what hid it; it now
runs on the real liveness.
`_gateway_subcommand` rewrites a restart of a profile the multiplexer
serves into a restart of the multiplexer. From a dashboard whose own home
is the default one it emitted a bare `gateway restart`, assuming bare
means the root. It does not: the child's profile pre-parse does not trust
a root HERMES_HOME as-is and re-reads the sticky active_profile (#22502),
so the restart went to that profile, which the multiplexer serves, and
the multiplexer was never restarted. The child now always gets
`-p default`, as the pooled-backend path already did.
`_marker_only_restart_obsolete` bailed out with "a newer pull moved HEAD" whenever
`checkout_sha != expected_sha`, and otherwise held the fleet to `expected_sha` exactly.
A cherry-picked hotfix on top of the pulled SHA moves HEAD without any pull arming a
fresh obligation, so the recorded one could never discharge: every CLI start warned
and every no-op `hermes update` exited 1 through the armed veto in
`_pending_fleet_restart_needed`, while the gateway verifiably served HEAD.
The bail-out now fires only when HEAD no longer CONTAINS `expected_sha`
(`checkout_contains`, `git merge-base --is-ancestor`, fail-closed on any probe
failure), and the live fleet is held to the checkout SHA — the code it actually runs.
A gateway on genuinely stale code still keeps the obligation armed.
Closes#119367
`_critical_module_import_failures` spawned `<venv python> -c ...` with
cwd=PROJECT_ROOT; `-c` puts the cwd at sys.path[0], so every first-party import
resolved from the tree and the probe was green while the installed editable
finder's MAPPING lacked `hermes_platform` — the exact state that crash-loops a
gateway started from `/` (#119466, maintainer triage gap 1).
When the venv carries an `__editable__*hermes_agent*finder.py` the probe now
runs under `-P` (Python >= 3.11, matches requires-python), so it imports through
the same finder the gateway will. A dev checkout with no editable install keeps
the cwd path and its advisory verdict.
Live: scratch venv, `uv pip install -e .` at 9863e315 (pre-hermes_platform),
checkout a27b1305: base probe NONE (green) while `cd / && python -c 'import
hermes_constants'` dies; head probe reports ModuleNotFoundError
hermes_platform; after a plain `uv pip install -e .` the probe is green again.
Part of #119466
`hermes update --yes` from a cron job inside the gateway takes the #100179
ancestor branch (`_drain_or_signal_gateway_for_update` -> self-restart request,
fire-and-forget) — the gateway can only restart after the updater exits. The
post-restart fleet matrix then saw that same gateway on the pre-update code_sha,
printed STALE + "Update not complete" and exited 1 on every nightly run (#119597).
The restart phase now records the pids that accepted a self-restart
(`_GatewayRestartOutcome.self_restart_pending_pids`, threaded from the systemd,
manual and launchd paths) and `collect_fleet_versions(self_restart_pending=)`
turns exactly those rows into `restart_pending` — rendered as "restart pending
(deferred until this process exits)" and excluded from the stale_or_down
verdict. Every other stale/down gateway keeps its verdict; the same pid without
the recorded acceptance is still STALE. The receipt keeps the row's old sha, so
the next `hermes update` / CLI startup hint still verifies the restart landed
via the live fleet (`_receipt_reports_stale_runtime` -> `_live_fleet_covers_receipt`).
Closes#119597
A stale editable finder cannot import a new top-level name unless the
checkout is the process cwd. python -m only adds that cwd, so a hand-off
spawned from elsewhere died before the map check could refresh the finder.
Review fix for #119525. The stale-map detection stays; the forced
`--reinstall` goes, for three reasons verified live:
- uv (0.12.13, the version the updater ships) reinstalls an editable
source tree on every `uv pip install -e .` (`~ hermes-agent==0.21.4`
even with no change). Adding a root module and running the plain
install rewrote the finder MAPPING without touching packaging files,
so once the skip is refused the existing install already refreshes it.
- `--reinstall` reinstalls EVERY package in the venv, not just
hermes-agent; the narrow spelling would be `--reinstall-package
hermes-agent`, and neither is needed.
- `--reinstall` is not a pip option. On the pip fallback path (no uv:
Termux without uv, some site-packages installs) a stale map turned the
update into a hard `no such option` failure, worse than the skip.
Also removes the second `_editable_finder_mapping_current` call in
`_sync_python_dependencies_after_pull` (the predicate already ran it)
and the flag's test.
The packaging-file skip treated an unchanged pyproject as proof the
setuptools finder could still import the checkout. That finder freezes
top-level names at install time, so a later package such as
hermes_platform stays missing and the gateway crash-loops after a
successful update.
The profile_home stamp exists to let serves_routed_profile() see a foreign-home scope without the
HERMES_HOME override. Stamping the LAUNCH home from the own-home binders (launch_profile_runtime_scope,
_worker_profile_scope's launch arm, model_switch's launch arm) added nothing for production but, under
a test conftest where server._hermes_home differs from HERMES_HOME, flipped the launch profile to
'routed' and charged every turn a routed tool/plugin discovery. Own-home binders now leave the stamp
None; the foreign-home arms keep it.
is_profile_gate_env resolved the platform prefix set (Platform enum + bundled plugin scan + registry)
once per env key; strip_profile_gate_env walks the whole child env, so a spawn paid the scan
hundreds of times. The static half is lru_cached, the dynamic registry union is taken once per strip.
Eight secret readers wrapped the scoped get_secret() call in a broad
except Exception / contextlib.suppress and fell through to os.environ.
Under multiplex that env holds the default profile's value, so a bound
scope whose resolution fails silently borrowed another profile's
credential: the pairing allowlist reader could then persist the foreign
list into the served profile's .env, and the proxy key, tool gateway
token, OpenRouter and aux provider keys, ElevenLabs key, and Slack token
probe had the same shape.
Keep the deliberate UnscopedSecretError -> os.environ fallback (the
unscoped default-profile path legitimately reads its own env) and let
every other scoped-read failure propagate; the two availability probes
fail closed instead. config._scoped_environ_get now propagates as its
docstring already claimed.
Several profile-scope binders called set_hermes_home_override and
set_secret_scope before the try/finally that releases them. A raise in
scope setup (a corrupt or removed profile home) propagated with the
foreign override still bound to the caller's context, silently re-homing
every later read — and in _reregister_orphaned_adopters it also skipped
every remaining adopter. The set calls now run inside the try with
None-guarded resets; the same shape is fixed in the routed-turn scope,
the cron external worker (which also leaked the multiplex flag), the
kanban worker scope, the MCP OAuth paths, the launch-profile policy,
and model_switch, which releases partially-bound scopes on raise.