Follow-up to the cherry-picked fix from #112724 (@poijygfdyy):
- tools/checkpoint_profile_migration.py -> tools/checkpoint_manager_profile_rename.py, the
repo's `<stem>_<topic>.py` sibling convention for code that extends checkpoint_manager.
- Replace the fail-closed target-collision check plus temp-file/rollback choreography with an
idempotent rekey: every step overwrites and the old project metadata is removed last, so a
mid-way failure is repaired by `hermes profile migrate-identity` redoing the same writes.
A genuine collision cannot occur — `profiles/<new>` must not exist for the rename to run.
192 -> 98 lines.
- Metadata/ledger writes go through the same idiom as checkpoint_manager itself
(`_register_project` plain write, `_save_ledger`), dropping the private temp-file helpers.
- Keep the git-present precondition as a single early check: without git the ref cannot move
and rekeying only the metadata would orphan the history.
- Test: `create_profile` now seeds `workspace/`, so the fixture uses `project/`; add the control
assertion that a workdir outside the profile dir keeps its history unchanged.
- Docs: the profile rename / migrate-identity reference notes that checkpoint history is preserved.
Fixes#112973
Same behaviour as the cherry-picked fix, one branch instead of a nested
if/else: a not-running gateway is "stopped" when the operator's durable
``desired_state`` says so or when the retained state is not one of the two
terminal states; ``exit_reason`` is dropped only on "stopped" so a live
``startup_failed`` (desired_state=running) keeps its diagnostic.
Live probe (real /api/status, temp HERMES_HOME): stopped default profile
with retained port-conflict failure -> unscoped ``gateway_state='stopped'``,
``gateway_exit_reason=None``; control ``?profile=worker`` (desired running)
-> still ``startup_failed`` + 'telegram: token rejected'.
A delete can end with the profile directory already removed and its durable
session/routing identity settlement still pending. delete_profile raised a bare
RuntimeError for that state, so a caller wanting to distinguish the completed
filesystem half from a real failure had only the message to go on — and the
dashboard's DELETE /api/profiles/{name} folded it into its generic 500, making a
client read the completed delete as "delete failed" (its retry then 404'd).
The partial settlement is now the typed ProfileIdentitySettlementPending — a
RuntimeError subclass carrying profile / path / retry_command:
- the CLI's delete handler keeps its non-zero exit unchanged: the type subclasses
RuntimeError, so its existing catch tuple covers it;
- DELETE /api/profiles/{name} catches exactly that type and answers
200 {"ok": true, "path": ..., "identity_settled": false, "settlement_pending":
true, "retry_command": "hermes profile purge-identity <name>"};
- a genuine filesystem failure stays a plain RuntimeError and keeps the 500 — the
API does not catch RuntimeError broadly.
Tests: the live-multiplexer delete test now pins the typed payload (profile,
retry_command, the removed path, RuntimeError subclassing); a new endpoint test
drives the real delete chain — settle-pending answers the partial success with the
directory really gone, and an rmtree failure still answers 500. Red on base (the
endpoint answered 500 for the completed delete; the type is absent) and green with
the fix: 347 passed / 0 failed across the 10 affected files.
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:
- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
`_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
(`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
`_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
`create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
leaves of a profile's conversation record is `delete_profile`'s business — it removes the
profile's own home, `state.db` included.
Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).
Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.
Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
`_finish_dashboard_update_cleanup` runs pulled `dashboard_procs` inside the pre-pull
interpreter. A stale-symbol failure there (#112604: `AttributeError: module
'hermes_cli.main_dashboard' has no attribute '_loaded_launchd_backend_jobs'`, reproduced on
the maintainer's own box) propagated out of `_verify_fleet_after_update`, skipping the
fleet version matrix, plan-vs-execution reconciliation and the inner receipt finalize; the
command boundary then stamped the receipt `failed` with the traceback as stop_reason even
though the code update and gateway restart had succeeded.
The cleanup is now isolated like every sibling post-update step: the exception is
printed with a manual-restart hint, logged at WARNING, and recorded as a failed
`dashboard_cleanup` step on the receipt. A dashboard/serve left on pre-update code is still
escalated by the survivor probe → reconciliation (exit 1). Covers the git and ZIP paths.
Refs #112604 (residual class after #112753). Test is red on origin/main.
`hermes update` writes its receipt from the PRE-pull interpreter after the post-pull
module purge. `_receipt_dir()` resolved the home through `hermes_cli.config`, so the
write re-executed the pulled `config.py` against whatever was still cached — on a pull
that added a symbol to a root-level module (`utils.file_signature`, `base_url_origin`)
that import raised and the whole receipt was dropped. The failure was logged at DEBUG,
which the updater's INFO log discards, so an activation run left no receipt while the
following no-op run wrote one, and `latest.json` kept pointing at an older run.
- `_receipt_dir()` uses `hermes_constants.get_hermes_home` (purge-protected, stdlib-only).
- A failed write prints `⚠ Update receipt not written: <exc>` and logs at WARNING.
Refs #112465, #112558 (Finding A). Both new tests are red on origin/main.
* fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping
systemd already stops restarting on exit 78; launchd KeepAlive=true
respawned the same token/port collision every 30s. Map 78 to a clean
stop under SuccessfulExit=false so the job stays down until the holder
releases the lock.
* test(launchd): pin EX_CONFIG park under SuccessfulExit KeepAlive
The plist must not use unconditional KeepAlive, and the stderr wrapper
must turn gateway exit 78 into a clean stop without swallowing exit 75
or a non-gateway child's 78.
MintFailure.as_payload() reported ceil(not_before - monotonic()). With
not_before = now + 60, the subtraction is 60.000000000000455 for some values
of now, and the ceil made it 61: users saw "retry in 61s" and CI failed
test_the_background_loop_retries_a_transient_failure_until_it_settles
(slept == [15, 61]) whenever the runner uptime landed on such a value —
twice in a row on #113055. remaining() now rounds to the millisecond before
the ceil. One test pins a concrete `now` that produced the dust.
Union Alpha is OpenRouter's new $0 stealth model (262,144 ctx, tools +
tool_choice supported, text+image in). It carries no ":free" suffix, so it
needs an explicit "free, stealth model" description and an _OPENROUTER_ONLY
entry to keep it out of the derived Nous Portal list, plus a "union-alpha"
DEFAULT_CONTEXT_LENGTHS key so the offline resolver returns 262144 instead
of the 256K catch-all. Docs manifest regenerated in the same commit.
Live gate: max_tokens=64 tool-calling completion on origin/main key returned
200, echoed model stealth/union-alpha, usage.cost=0, finish_reason=tool_calls.
_live_writer_holds_db in hermes_state_repair already binds the repair connector
(the sibling _write_health_reason uses it); the held-branch detail now says
'or state.db cannot be inspected' so the console line matches the comment above
it. Test setup for a 51 MB WAL is one helper shared by all three tests.
`hermes doctor` (without --fix) warned "WAL file is large — run 'hermes doctor
--fix' to checkpoint" without checking whether Desktop or the gateway held
the database; the holder scan only ran inside --fix. A large WAL is normal for
a live writer, and that nudge is how users in the #110054 threads became the
second writer that the deleted-WAL guard then fired on.
The holder scan now runs on the warn path too: held (or unprovable) reports
the size as normal while Desktop/gateway run and orders "stop" before any
--fix; no holder keeps the checkpoint suggestion, also stop-first.
Refs #110054 (item 3 of the proposed fix), #110073.
A profile served by the live default multiplexer runs no gateway of its
own, but a gateway_state.json left behind by a pre-multiplex or standalone
run is a non-None stale record, so the multiplexer fallback in
_platform_payloads (guarded by 'if runtime is None') never ran: the bare
platform-key lookup found nothing and the Messaging card fell through the
liveness ladder to pending_restart — a permanent "Restart needed" for a
platform that was connected and working (#112765).
Probe the multiplexer unconditionally: a served profile's own file always
describes a dead process, so the live multiplexer's record wins whether
the leftover exists or not.
The boot guard logged its refusal once and nothing else knew. A user whose default says multiplex
but whose gateway serves one profile would read `hermes gateway status` and see a healthy gateway.
The verdict now lands in gateway_state.json (multiplex_standalone_reason) and status prints it with
both remedies; a boot that does multiplex clears it.
DEFAULT_CONFIG now ships gateway.multiplex_profiles: true. GatewayConfig keeps an UNSET flag as
None so the boot can tell "the operator chose" from "the default applies"; every reader tests
truthiness, so an undecided flag never multiplexes by accident.
hermes_cli/gateway_multiplex_mode.py settles the unset default once per boot (called from
load_gateway_config_for_runner and `gateway run --config`): the same preflight `hermes gateway
migrate --multiplex` runs — default profile, >= 2 profiles, no secondary running its own gateway
(live pid or installed unit), no duplicate-credential / port-binder blocker, migratable host. A
refusal is a logged warning naming the blocker and the migrate one-liner; the gateway comes up
standalone exactly as before. Explicit values (config.yaml, GATEWAY_MULTIPLEX_PROFILES) pass through
verbatim; `--standalone` already pins false.
Other processes stop guessing the verdict from the merged default: named_profile_served_by_running
_multiplexer, the enroll warning, the dashboard listener guard, the cron-fire port resolver, container
boot and the migration plan (_read_multiplex_flag) read the live gateway's served_profiles record
first and the EXPLICIT flag second — so a per-profile fleet with the flag unset still reads as
"not yet multiplexed" and the fold proceeds.
Docs: multi-profile-gateways.md, multiplexing-gateway.md, hermes_cli/AGENTS.md.
Hoist the "enter the launch profile's scope under multiplex, hop to the executor with
copy_context" into GatewayStartupMixin._run_boot_probe_in_launch_scope so the warm-up has one
call site instead of two executor branches, and the sibling boot probes can share it.
_scoped_operator_override now takes the candidate names: _nous_portal_env_override used to call
it twice (HERMES_PORTAL_BASE_URL, then the NOUS_PORTAL_BASE_URL alias), so one lost-scope event
logged two identical WARNINGs; one try around the list logs once and names both.
Tests: drop the dead UnscopedSecretError arm in the probe (the helper swallows it), fold the
no-scope-left-behind assertions into the multiplex test, keep the environ-semantics contract for
single-profile. Docstrings keep the WHY; the incident chain lives in the PR.
Scope entry that cannot be built falls back to the unscoped executor call (the
load_gateway_config_for_runner shape) instead of aborting startup. The no-scope-left-behind
assertions run in the warm-up's own task — asyncio.run copies the context, so asserting from the
caller proved nothing.
_warm_turn_prerequisites ran _warm_turn_machinery_sync on a bare executor thread with no
profile scope. get_tool_definitions runs every check_fn; the vision probe resolves live Nous
runtime credentials (resolve_vision_provider_client -> _resolve_nous_runtime_api ->
resolve_nous_runtime_credentials -> _nous_effective_routing). Under GATEWAY_MULTIPLEX_PROFILES
get_secret fails closed there, so HERMES_PORTAL_BASE_URL / NOUS_INFERENCE_BASE_URL read as
absent, the Portal allowlist heals to production, and a staging deployment's refresh token is
POSTed to portal.nousresearch.com: invalid_grant, "Nous OAuth state quarantined", ~10 s after
every boot and before any inbound turn. Reproduced live on hermes-agent-stg-gg-probe-test-0062
with the override present in the profile .env (four NAS re-seeds, four boot-time deaths); the
unscoped call stack was captured on the box.
The warm-up now enters the launch profile's _profile_runtime_scope and carries the context
into the executor thread with copy_context().run — the same shape _discover_gateway_mcp_tools
uses (#95518). Single-profile gateways are unchanged.
_scoped_operator_override logs which override was unreadable and why; the downstream
"ignoring invalid portal_base_url" warning never named the caller that lost its scope.
Tests: scoped warm-up resolves the .env override on the executor thread (red on main), no
scope leaks past the warm-up, single-profile keeps environ semantics.
`_purge_stale_hermes_modules` reads the cache through the `_m().sys.modules`
seam (tests patch `hermes_cli.main.sys`), but `_evict_module` popped from the
bare `sys.modules` and looked the parent up there — mutating a different dict
than the one the caller iterates. Pass the table in.
Also retarget the issue citations: #111689 is the launchd bug whose fix
ADDED `_loaded_launchd_backend_jobs`; the crash this fixes is #112604.
`_evict_module` compared the package attribute's `__name__` to the evicted
name. Every module bound under that attribute has that name, so the guard
also deleted a same-named module something else had already rebound on the
package — the newer object, which must stay. Compare the attribute against
the module object popped from `sys.modules` instead; only the exact stale
object is unbound. Read through `vars()` so a lazy package `__getattr__`
(`providers`) is never triggered mid-purge.
Follow-up to djsanchezsuarez's #112598 (cherry-picked as the previous
commit). Supersedes the reload-list carriers #112614 / #112630 for #112604.
`hermes update` runs in the PRE-pull process. `_purge_stale_hermes_modules()`
evicted cached Hermes submodules from `sys.modules`, but `hermes_cli.main`
imports `main_dashboard` at CLI start, so the module also lives as an ATTRIBUTE
of the (purge-protected) `hermes_cli` package. `from hermes_cli import
main_dashboard` is satisfied by that attribute (`_handle_fromlist`), so the
post-update dashboard cleanup kept resolving the PRE-pull module and died on a
symbol the pull had just added:
File "hermes_cli/dashboard_procs.py", line 366, in _kill_stale_dashboard_processes
launchd_jobs = _dash._loaded_launchd_backend_jobs() if sys.platform != "win32" else []
AttributeError: module 'hermes_cli.main_dashboard' has no attribute
'_loaded_launchd_backend_jobs'
It lands in `_finish_dashboard_update_cleanup`, i.e. right after the update has
already reported the dashboard restart, so the code update itself succeeds but
the run aborts non-zero and the remaining cleanup/verification never happens.
The symbol came from the launchd-supervision series (#111690, merged as
#112240 for #111689); this crash is a separate stale-module handoff.
`_evict_module()` now also removes the parent package's attribute when it is
the module being evicted, so the cleanup's call-time import re-reads the pulled
source. Both regression tests reproduce the reported traceback at the pre-fix
checkout and pass with the fix.
#111617 review (andrexibiza P1 #3/#4, kvnloo nit):
- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
a reboot, and that counter does not, so an unrelated process on a later boot with the
same PID and the same tick value passed _start_times_agree(). The fingerprint is now
"<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
the witness the drain marker already uses); both halves must match. Integer values on
rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
row and falls back to bare PID existence - a new spawn silently recreated the #89614/
#99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
and must not outlive its pid.
Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.
Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
Three edges of the fail-closed multi-profile host (#111620 review, andrexibiza P1 + P2,
kvnloo finding 1):
- `send` under a routed profile's scope `update()`d the installed scope from raw `.env`,
reversing build_profile_secret_scope's precedence (user .env, then external secret
sources) for the rest of the request; a stale user value beat the secret-manager one.
The installed scope is authoritative as-is; only the config.yaml setdefault bridge runs.
- The launch profile's body was scoped only when `is_multiplex_active()` was already true
at entry, while get_secret consults that global on every read. A launch RPC / dashboard
request entering single-profile and resuming after a concurrent first `?profile=B`
activation raised UnscopedSecretError mid-request. The launch profile's secret scope
(its .env + external sources over the launch env: live while single-profile, the frozen
snapshot once multiplexing is active) is now bound for every launch-profile body, so the
credential source is fixed at entry. The terminal policy overlay stays multiplex-only
(standalone terminal execution keeps its os.environ bridge). _publish_env_value mirrors a
same-request .env write into that scope AND os.environ for the launch profile, only into
the scope for a routed one (serves_routed_profile).
- _release_profile_runtime_scope_tokens reset terminal → secret → home in sequence under
one outer suppress; a failing terminal reset left the previous profile's secrets and
HERMES_HOME installed for the next body in that context. Each reset is now independent;
the first failure is re-raised after every scope is released.
tests/tui_gateway/test_multi_profile_hosting_transitions.py: manager-vs-dotenv precedence
through _load_hermes_env, TUI-RPC and dashboard barrier tests (launch enters single-profile,
B activates on another thread, launch resumes and still resolves its injected credential,
never B's), forced terminal-reset failure still releases secret + home. 4/4 red on base.
Three P1 findings from the review of #111062 (fixes#110850 remainder):
- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
installed-but-dead unit; `systemd_install` writes the unit before the start that
can still be killed, so an apply interrupted at default/start (or mid-restart of
a stopped default) was reported as "already multiplexed" and never recovered.
`interrupted` is now derived from the manifest vs the LIVE default's recorded
served_profiles: flag on + manifest + not every migrated profile served = resume.
- the compensation `try` began after `_remove_secondary_gateways` and the flag
write, so a later secondary's stop, a unit's daemon-reload or the config write
failing left the first secondary removed with no rollback. The flag write,
removals and default bring-up now all sit inside one boundary that rolls back
through the manifest; `_preflight_apply` refuses failures knowable from the plan
(root system unit without a recorded User=, unresolvable recorded user, an
unwritable config.yaml) before any working gateway is stopped.
- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
`default_uid is None`; a default system unit whose User= this host cannot
resolve is the same boundary from the other side and now blocks the update
hook (same known uid still folds, different known uid still refuses).
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.
The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.
Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.
Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.
- `hermes profile migrate-identity <old> <new>`: retries the migration —
delegates to the gateway control verb while a multiplexer is live, performs the
durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
naming the offending database on a collision, a lock, or a partial failure. Only
the name format and the existence of the new profile are checked; the old profile
directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
`click`: a failed second database raised NameError instead of printing its warning.
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.
The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:
- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
tables, matching the agent:<name>: namespace by exact prefix (substr, not
LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
the profile inside routing/origin JSON, and REFUSING on a target collision
(routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
(keys + origin.profile) then persist — the half a DB write cannot reach.
Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
params only to handlers that declare them, bare handlers unchanged) so a
live gateway rekeys its in-memory copy AND both durable stores (routing home
+ the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
does NOT fall back to a racing CLI-side write: it prints a warning telling
the operator to restart the gateway and retry. With no live gateway it
performs the durable rewrite itself (safe: nothing else holds the store
open).
Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.
Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.
Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
Production code must not depend on website/docs being present; the
"Hermes does not read this name" note now keys only on the code
registry (OPTIONAL_ENV_VARS / _EXTRA_ENV_KEYS / setup-hidden suffixes).
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.
- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
into config.yaml (--force included); the env writer's denylist
(HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
bridging the value into os.environ; a name neither registered nor in the
environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
and lowercase bare keys are unchanged.
Fixes#111848 (first half landed in #112250).
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.
After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.
Fixes#111804 (its discovery half landed in #112293).
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.
- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
approval prompt could not be delivered or was not answered (<cause>)" with
outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.
Fixes#22992
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.
`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.
`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.
This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.
Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.
Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
get_mcp_status reports status='disabled' (never 'lazy') for a disabled
server and derives the 'disabled' flag from that same status, so the
extra 'and not srv.get("disabled")' check could never change the branch.
A `lazy: true` MCP server registers its tools from the schema cache and
spawns on first use. Three consumers still equated "alive" with a live
session, so a healthy all-lazy startup was reported as a total failure:
- `get_mcp_status()` fell through to `status: configured, tools: 0` for a
lazily registered server. It now reports `lazy` with the cached tool
count (`connected: False`); an in-flight or failed first-use connect
still outranks it because the error is the actionable part.
- `discover_mcp_tools()`'s summary counted every name absent from
`_servers` as failed, logging `MCP: 0 tool(s) from 0 server(s) (2
failed)` right after registering every cached tool, and re-announced
the same "failure" on every repeat discovery. Lazy servers are now
reported as `(N lazy, not spawned yet)` and an already-lazy server is
not re-announced.
- `hermes_cli/mcp_startup.py` judged a discovery run by `connected` at
two sites, so every startup logged `Background MCP discovery completed
with zero connected servers` and every later call re-spawned the
discovery thread as a retry. One predicate,
`_discovery_registered_servers`, treats a lazy registration as a
usable outcome at both sites.
- `hermes_cli/banner.py` rendered the unknown `lazy` status through the
red "could not connect" line; it now shows the cached tool count with
`(lazy, starts on first use)`.
Ported from #100648 (core hunks only; the toolsets-filter predicate
branch, the Ink TUI component extraction and 13 tests were not ported).
Fixes#111717