Commit Graph

7827 Commits

Author SHA1 Message Date
teknium1
c055eee1fe refactor(checkpoints): slim the profile-rename rekey and name it as a checkpoint_manager sibling
Follow-up to the cherry-picked fix from #112724 (@poijygfdyy):

- tools/checkpoint_profile_migration.py -> tools/checkpoint_manager_profile_rename.py, the
  repo's `<stem>_<topic>.py` sibling convention for code that extends checkpoint_manager.
- Replace the fail-closed target-collision check plus temp-file/rollback choreography with an
  idempotent rekey: every step overwrites and the old project metadata is removed last, so a
  mid-way failure is repaired by `hermes profile migrate-identity` redoing the same writes.
  A genuine collision cannot occur — `profiles/<new>` must not exist for the rename to run.
  192 -> 98 lines.
- Metadata/ledger writes go through the same idiom as checkpoint_manager itself
  (`_register_project` plain write, `_save_ledger`), dropping the private temp-file helpers.
- Keep the git-present precondition as a single early check: without git the ref cannot move
  and rekeying only the metadata would orphan the history.
- Test: `create_profile` now seeds `workspace/`, so the fixture uses `project/`; add the control
  assertion that a workdir outside the profile dir keeps its history unchanged.
- Docs: the profile rename / migrate-identity reference notes that checkpoint history is preserved.

Fixes #112973
2026-09-16 14:34:22 -07:00
NanPan
8a2c9edac6 fix(profiles): preserve checkpoint history across rename 2026-09-16 14:34:22 -07:00
teknium1
059c0269e1 refactor(status): collapse the stopped-intent ladder into one condition
Same behaviour as the cherry-picked fix, one branch instead of a nested
if/else: a not-running gateway is "stopped" when the operator's durable
``desired_state`` says so or when the retained state is not one of the two
terminal states; ``exit_reason`` is dropped only on "stopped" so a live
``startup_failed`` (desired_state=running) keeps its diagnostic.

Live probe (real /api/status, temp HERMES_HOME): stopped default profile
with retained port-conflict failure -> unscoped ``gateway_state='stopped'``,
``gateway_exit_reason=None``; control ``?profile=worker`` (desired running)
-> still ``startup_failed`` + 'telegram: token rejected'.
2026-09-16 14:34:22 -07:00
KoNit-K
aa711728e8 fix(status): honor stopped gateway intent 2026-09-16 14:34:22 -07:00
KoNit-K
43e28551bb fix(profiles): release routed log locks before deletion 2026-09-16 14:34:22 -07:00
xielevi
aa90c01e9a fix(profiles): report a settle-pending delete as a typed partial success
A delete can end with the profile directory already removed and its durable
session/routing identity settlement still pending. delete_profile raised a bare
RuntimeError for that state, so a caller wanting to distinguish the completed
filesystem half from a real failure had only the message to go on — and the
dashboard's DELETE /api/profiles/{name} folded it into its generic 500, making a
client read the completed delete as "delete failed" (its retry then 404'd).

The partial settlement is now the typed ProfileIdentitySettlementPending — a
RuntimeError subclass carrying profile / path / retry_command:

- the CLI's delete handler keeps its non-zero exit unchanged: the type subclasses
  RuntimeError, so its existing catch tuple covers it;
- DELETE /api/profiles/{name} catches exactly that type and answers
  200 {"ok": true, "path": ..., "identity_settled": false, "settlement_pending":
  true, "retry_command": "hermes profile purge-identity <name>"};
- a genuine filesystem failure stays a plain RuntimeError and keeps the 500 — the
  API does not catch RuntimeError broadly.

Tests: the live-multiplexer delete test now pins the typed payload (profile,
retry_command, the removed path, RuntimeError subclassing); a new endpoint test
drives the real delete chain — settle-pending answers the partial success with the
directory really gone, and an rmtree failure still answers 500. Red on base (the
endpoint answered 500 for the completed delete; the type is absent) and green with
the fix: 347 passed / 0 failed across the 10 affected files.
2026-09-16 14:21:14 -07:00
xielevi
a41552fad4 fix(profiles): purge a deleted profile's session/routing identity on delete
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:

- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
  `_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
  also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
  namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
  (`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
  in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
  writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
  `_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
  has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
  partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
  profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
  `create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
  delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
  leaves of a profile's conversation record is `delete_profile`'s business — it removes the
  profile's own home, `state.db` included.

Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
2026-09-16 14:21:14 -07:00
teknium1
034313e7cd feat: plugin catalog entries carry an optional version label and card image
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).

Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.

Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
2026-09-16 14:18:39 -07:00
brooklyn!
d5b3cad7b9 fix(mcp): isolate unsupported inferred app labels 2026-09-16 15:54:47 -05:00
brooklyn!
74ca4f28ba fix(onboarding): consume future catalog entries without custom metadata 2026-09-16 15:54:47 -05:00
brooklyn!
a38d544f1b feat(mcp): expose optional backend-local app discovery 2026-09-16 15:16:09 -05:00
Yags
61154b6f68 fix(models): route Union Alpha through messages 2026-09-16 11:41:18 -07:00
teknium1
5f071fb3ad fix(update): a failed dashboard cleanup no longer aborts fleet verification and the receipt
`_finish_dashboard_update_cleanup` runs pulled `dashboard_procs` inside the pre-pull
interpreter. A stale-symbol failure there (#112604: `AttributeError: module
'hermes_cli.main_dashboard' has no attribute '_loaded_launchd_backend_jobs'`, reproduced on
the maintainer's own box) propagated out of `_verify_fleet_after_update`, skipping the
fleet version matrix, plan-vs-execution reconciliation and the inner receipt finalize; the
command boundary then stamped the receipt `failed` with the traceback as stop_reason even
though the code update and gateway restart had succeeded.

The cleanup is now isolated like every sibling post-update step: the exception is
printed with a manual-restart hint, logged at WARNING, and recorded as a failed
`dashboard_cleanup` step on the receipt. A dashboard/serve left on pre-update code is still
escalated by the survivor probe → reconciliation (exit 1). Covers the git and ZIP paths.

Refs #112604 (residual class after #112753). Test is red on origin/main.
2026-09-16 11:06:52 -07:00
teknium1
1259b140c9 fix(update): the receipt survives a mixed sys.modules graph and a lost write is visible
`hermes update` writes its receipt from the PRE-pull interpreter after the post-pull
module purge. `_receipt_dir()` resolved the home through `hermes_cli.config`, so the
write re-executed the pulled `config.py` against whatever was still cached — on a pull
that added a symbol to a root-level module (`utils.file_signature`, `base_url_origin`)
that import raised and the whole receipt was dropped. The failure was logged at DEBUG,
which the updater's INFO log discards, so an activation run left no receipt while the
following no-op run wrote one, and `latest.json` kept pointing at an older run.

- `_receipt_dir()` uses `hermes_constants.get_hermes_home` (purge-protected, stdlib-only).
- A failed write prints `⚠ Update receipt not written: <exc>` and logs at WARNING.

Refs #112465, #112558 (Finding A). Both new tests are red on origin/main.
2026-09-16 11:06:52 -07:00
xxxigm
2eb5395d3f fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping (#112079)
* fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping

systemd already stops restarting on exit 78; launchd KeepAlive=true
respawned the same token/port collision every 30s. Map 78 to a clean
stop under SuccessfulExit=false so the job stays down until the holder
releases the lock.

* test(launchd): pin EX_CONFIG park under SuccessfulExit KeepAlive

The plist must not use unconditional KeepAlive, and the stderr wrapper
must turn gateway exit 78 into a clean stop without swallowing exit 75
or a non-gateway child's 78.
2026-09-16 17:58:25 +00:00
kshitijk4poor
6ff1a6c8a2 fix(anon): retry_after is the whole wait, not the wait plus float dust
MintFailure.as_payload() reported ceil(not_before - monotonic()). With
not_before = now + 60, the subtraction is 60.000000000000455 for some values
of now, and the ceil made it 61: users saw "retry in 61s" and CI failed
test_the_background_loop_retries_a_transient_failure_until_it_settles
(slept == [15, 61]) whenever the runner uptime landed on such a value —
twice in a row on #113055. remaining() now rounds to the millisecond before
the ceil. One test pins a concrete `now` that produced the dust.
2026-09-16 23:22:49 +05:30
teknium1
d0aecdf352 feat(models): add stealth/union-alpha to the OpenRouter free-tier catalog
Union Alpha is OpenRouter's new $0 stealth model (262,144 ctx, tools +
tool_choice supported, text+image in). It carries no ":free" suffix, so it
needs an explicit "free, stealth model" description and an _OPENROUTER_ONLY
entry to keep it out of the derived Nous Portal list, plus a "union-alpha"
DEFAULT_CONTEXT_LENGTHS key so the offline resolver returns 262144 instead
of the 256K catch-all. Docs manifest regenerated in the same commit.

Live gate: max_tokens=64 tool-calling completion on origin/main key returned
200, echoed model stealth/union-alpha, usage.cost=0, finish_reason=tool_calls.
2026-09-16 10:52:32 -07:00
kshitijk4poor
acc78451ca fix(doctor): the stop-first issue line says run, not re-run — it is now also emitted before any --fix 2026-09-16 23:19:45 +05:30
kshitijk4poor
7af776db62 refactor(doctor): reuse the repair holder-check wrapper and keep the honest disjunction in the warning
_live_writer_holds_db in hermes_state_repair already binds the repair connector
(the sibling _write_health_reason uses it); the held-branch detail now says
'or state.db cannot be inspected' so the console line matches the comment above
it. Test setup for a 51 MB WAL is one helper shared by all three tests.
2026-09-16 23:19:45 +05:30
kshitijk4poor
84a5f1cd1e fix(doctor): a large WAL under a live writer is reported as normal, never as "run --fix"
`hermes doctor` (without --fix) warned "WAL file is large — run 'hermes doctor
--fix' to checkpoint" without checking whether Desktop or the gateway held
the database; the holder scan only ran inside --fix. A large WAL is normal for
a live writer, and that nudge is how users in the #110054 threads became the
second writer that the deleted-WAL guard then fired on.

The holder scan now runs on the warn path too: held (or unprovable) reports
the size as normal while Desktop/gateway run and orders "stop" before any
--fix; no holder keeps the checkpoint suggestion, also stop-first.

Refs #110054 (item 3 of the proposed fix), #110073.
2026-09-16 23:19:45 +05:30
Konstantin Khlopkov
d4ff9e852e fix(desktop): resolve multiplexed secondary-profile platform state past a stale own record 2026-09-16 10:27:34 -07:00
liuhao1024
1db1303964 fix(messaging): ignore a served profile's stale own runtime record
A profile served by the live default multiplexer runs no gateway of its
own, but a gateway_state.json left behind by a pre-multiplex or standalone
run is a non-None stale record, so the multiplexer fallback in
_platform_payloads (guarded by 'if runtime is None') never ran: the bare
platform-key lookup found nothing and the Messaging card fell through the
liveness ladder to pending_restart — a permanent "Restart needed" for a
platform that was connected and working (#112765).

Probe the multiplexer unconditionally: a served profile's own file always
describes a dead process, so the live multiplexer's record wins whether
the leftover exists or not.
2026-09-16 10:27:34 -07:00
teknium1
0a1716cb27 gateway status shows why an unset-default gateway stayed standalone
The boot guard logged its refusal once and nothing else knew. A user whose default says multiplex
but whose gateway serves one profile would read `hermes gateway status` and see a healthy gateway.
The verdict now lands in gateway_state.json (multiplex_standalone_reason) and status prints it with
both remedies; a boot that does multiplex clears it.
2026-09-16 08:57:18 -07:00
teknium1
a10bbf95bb feat(gateway): gateway.multiplex_profiles defaults to on, gated by a boot-time serve guard
DEFAULT_CONFIG now ships gateway.multiplex_profiles: true. GatewayConfig keeps an UNSET flag as
None so the boot can tell "the operator chose" from "the default applies"; every reader tests
truthiness, so an undecided flag never multiplexes by accident.

hermes_cli/gateway_multiplex_mode.py settles the unset default once per boot (called from
load_gateway_config_for_runner and `gateway run --config`): the same preflight `hermes gateway
migrate --multiplex` runs — default profile, >= 2 profiles, no secondary running its own gateway
(live pid or installed unit), no duplicate-credential / port-binder blocker, migratable host. A
refusal is a logged warning naming the blocker and the migrate one-liner; the gateway comes up
standalone exactly as before. Explicit values (config.yaml, GATEWAY_MULTIPLEX_PROFILES) pass through
verbatim; `--standalone` already pins false.

Other processes stop guessing the verdict from the merged default: named_profile_served_by_running
_multiplexer, the enroll warning, the dashboard listener guard, the cron-fire port resolver, container
boot and the migration plan (_read_multiplex_flag) read the live gateway's served_profiles record
first and the EXPLICIT flag second — so a per-profile fleet with the flag unset still reads as
"not yet multiplexed" and the fold proceeds.

Docs: multi-profile-gateways.md, multiplexing-gateway.md, hermes_cli/AGENTS.md.
2026-09-16 08:57:18 -07:00
Konstantin Khlopkov
948e970661 fix(vault): default detected password managers to enabled (opt-out contract) (#109558) 2026-09-16 14:26:13 +00:00
brooklyn!
7e69873ca6 perf(sessions): index prompts without hydrating full transcripts 2026-09-16 06:39:15 -05:00
kshitijk4poor
c71dbf7078 refactor(gateway): one scoped boot-probe hop; one lost-scope warning per resolve
Hoist the "enter the launch profile's scope under multiplex, hop to the executor with
copy_context" into GatewayStartupMixin._run_boot_probe_in_launch_scope so the warm-up has one
call site instead of two executor branches, and the sibling boot probes can share it.

_scoped_operator_override now takes the candidate names: _nous_portal_env_override used to call
it twice (HERMES_PORTAL_BASE_URL, then the NOUS_PORTAL_BASE_URL alias), so one lost-scope event
logged two identical WARNINGs; one try around the list logs once and names both.

Tests: drop the dead UnscopedSecretError arm in the probe (the helper swallows it), fold the
no-scope-left-behind assertions into the multiplex test, keep the environ-semantics contract for
single-profile. Docstrings keep the WHY; the incident chain lives in the PR.

Scope entry that cannot be built falls back to the unscoped executor call (the
load_gateway_config_for_runner shape) instead of aborting startup. The no-scope-left-behind
assertions run in the warm-up's own task — asyncio.run copies the context, so asserting from the
caller proved nothing.
2026-09-16 16:16:24 +05:30
Ben Barclay
a3c4fc84e9 fix(gateway): run the boot warm-up inside the launch profile's scope under multiplex
_warm_turn_prerequisites ran _warm_turn_machinery_sync on a bare executor thread with no
profile scope. get_tool_definitions runs every check_fn; the vision probe resolves live Nous
runtime credentials (resolve_vision_provider_client -> _resolve_nous_runtime_api ->
resolve_nous_runtime_credentials -> _nous_effective_routing). Under GATEWAY_MULTIPLEX_PROFILES
get_secret fails closed there, so HERMES_PORTAL_BASE_URL / NOUS_INFERENCE_BASE_URL read as
absent, the Portal allowlist heals to production, and a staging deployment's refresh token is
POSTed to portal.nousresearch.com: invalid_grant, "Nous OAuth state quarantined", ~10 s after
every boot and before any inbound turn. Reproduced live on hermes-agent-stg-gg-probe-test-0062
with the override present in the profile .env (four NAS re-seeds, four boot-time deaths); the
unscoped call stack was captured on the box.

The warm-up now enters the launch profile's _profile_runtime_scope and carries the context
into the executor thread with copy_context().run — the same shape _discover_gateway_mcp_tools
uses (#95518). Single-profile gateways are unchanged.

_scoped_operator_override logs which override was unreadable and why; the downstream
"ignoring invalid portal_base_url" warning never named the caller that lost its scope.

Tests: scoped warm-up resolves the .env override on the executor thread (red on main), no
scope leaks past the warm-up, single-profile keeps environ semantics.
2026-09-16 16:16:24 +05:30
kshitij
b1d38c2e5a refactor(update): evict through the caller's module table
`_purge_stale_hermes_modules` reads the cache through the `_m().sys.modules`
seam (tests patch `hermes_cli.main.sys`), but `_evict_module` popped from the
bare `sys.modules` and looked the parent up there — mutating a different dict
than the one the caller iterates. Pass the table in.

Also retarget the issue citations: #111689 is the launchd bug whose fix
ADDED `_loaded_launchd_backend_jobs`; the crash this fixes is #112604.
2026-09-16 16:04:27 +05:30
kshitij
d2f1587aa2 fix(update): unbind the purged submodule by object identity, not name
`_evict_module` compared the package attribute's `__name__` to the evicted
name. Every module bound under that attribute has that name, so the guard
also deleted a same-named module something else had already rebound on the
package — the newer object, which must stay. Compare the attribute against
the module object popped from `sys.modules` instead; only the exact stale
object is unbound. Read through `vars()` so a lazy package `__getattr__`
(`providers`) is never triggered mid-purge.

Follow-up to djsanchezsuarez's #112598 (cherry-picked as the previous
commit). Supersedes the reload-list carriers #112614 / #112630 for #112604.
2026-09-16 16:04:27 +05:30
djsanchezsuarez
5a942f95fc fix(update): drop the stale submodule attribute in the post-pull purge
`hermes update` runs in the PRE-pull process. `_purge_stale_hermes_modules()`
evicted cached Hermes submodules from `sys.modules`, but `hermes_cli.main`
imports `main_dashboard` at CLI start, so the module also lives as an ATTRIBUTE
of the (purge-protected) `hermes_cli` package. `from hermes_cli import
main_dashboard` is satisfied by that attribute (`_handle_fromlist`), so the
post-update dashboard cleanup kept resolving the PRE-pull module and died on a
symbol the pull had just added:

    File "hermes_cli/dashboard_procs.py", line 366, in _kill_stale_dashboard_processes
        launchd_jobs = _dash._loaded_launchd_backend_jobs() if sys.platform != "win32" else []
    AttributeError: module 'hermes_cli.main_dashboard' has no attribute
    '_loaded_launchd_backend_jobs'

It lands in `_finish_dashboard_update_cleanup`, i.e. right after the update has
already reported the dashboard restart, so the code update itself succeeds but
the run aborts non-zero and the remaining cleanup/verification never happens.
The symbol came from the launchd-supervision series (#111690, merged as
#112240 for #111689); this crash is a separate stale-module handoff.

`_evict_module()` now also removes the parent package's attribute when it is
the module being evicted, so the cleanup's call-time import re-reads the pulled
source. Both regression tests reproduce the reported traceback at the pre-fix
checkout and pass with the fix.
2026-09-16 16:04:27 +05:30
teknium1
db54f5448d fix(kanban): worker fingerprint carries a boot witness; an uncaptured fingerprint never authorizes a signal
#111617 review (andrexibiza P1 #3/#4, kvnloo nit):

- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
  is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
  a reboot, and that counter does not, so an unrelated process on a later boot with the
  same PID and the same tick value passed _start_times_agree(). The fingerprint is now
  "<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
  the witness the drain marker already uses); both halves must match. Integer values on
  rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
  row and falls back to bare PID existence - a new spawn silently recreated the #89614/
  #99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
  claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
  by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
  once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
  the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
  and must not outlive its pid.

Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.

Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
2026-09-16 00:35:00 -07:00
teknium1
8609743389 fix(serve): launch-profile scope decided at entry; send keeps scope authority; per-reset release
Three edges of the fail-closed multi-profile host (#111620 review, andrexibiza P1 + P2,
kvnloo finding 1):

- `send` under a routed profile's scope `update()`d the installed scope from raw `.env`,
  reversing build_profile_secret_scope's precedence (user .env, then external secret
  sources) for the rest of the request; a stale user value beat the secret-manager one.
  The installed scope is authoritative as-is; only the config.yaml setdefault bridge runs.
- The launch profile's body was scoped only when `is_multiplex_active()` was already true
  at entry, while get_secret consults that global on every read. A launch RPC / dashboard
  request entering single-profile and resuming after a concurrent first `?profile=B`
  activation raised UnscopedSecretError mid-request. The launch profile's secret scope
  (its .env + external sources over the launch env: live while single-profile, the frozen
  snapshot once multiplexing is active) is now bound for every launch-profile body, so the
  credential source is fixed at entry. The terminal policy overlay stays multiplex-only
  (standalone terminal execution keeps its os.environ bridge). _publish_env_value mirrors a
  same-request .env write into that scope AND os.environ for the launch profile, only into
  the scope for a routed one (serves_routed_profile).
- _release_profile_runtime_scope_tokens reset terminal → secret → home in sequence under
  one outer suppress; a failing terminal reset left the previous profile's secrets and
  HERMES_HOME installed for the next body in that context. Each reset is now independent;
  the first failure is re-raised after every scope is released.

tests/tui_gateway/test_multi_profile_hosting_transitions.py: manager-vs-dotenv precedence
through _load_hermes_env, TUI-RPC and dashboard barrier tests (launch enters single-profile,
B activates on another thread, launch resumes and still resolves its injected credential,
never B's), forced terminal-reset failure still releases secret + home. 4/4 red on base.
2026-09-16 00:35:00 -07:00
teknium1
7dde7a2424 fix(gateway): migrate --multiplex resumes from live state, compensates the whole destructive phase, and refuses an unknown default principal
Three P1 findings from the review of #111062 (fixes #110850 remainder):

- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
  installed-but-dead unit; `systemd_install` writes the unit before the start that
  can still be killed, so an apply interrupted at default/start (or mid-restart of
  a stopped default) was reported as "already multiplexed" and never recovered.
  `interrupted` is now derived from the manifest vs the LIVE default's recorded
  served_profiles: flag on + manifest + not every migrated profile served = resume.

- the compensation `try` began after `_remove_secondary_gateways` and the flag
  write, so a later secondary's stop, a unit's daemon-reload or the config write
  failing left the first secondary removed with no rollback. The flag write,
  removals and default bring-up now all sit inside one boundary that rolls back
  through the manifest; `_preflight_apply` refuses failures knowable from the plan
  (root system unit without a recorded User=, unresolvable recorded user, an
  unwritable config.yaml) before any working gateway is stopped.

- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
  `default_uid is None`; a default system unit whose User= this host cannot
  resolve is the same boundary from the other side and now blocks the update
  hook (same known uid still folds, different known uid still refuses).
2026-09-16 00:33:03 -07:00
Sora-bluesky
9085ef967c fix(profiles): sweep the remaining pre-write mkdirs under the deleted-profile guard
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.

The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.

Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
2026-09-16 00:32:15 -07:00
teknium1
5d97d5ed6d refactor(profiles): move rename identity migration off the facades; trim tests to invariants
- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
  migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
  now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
  profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
  gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
  hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
  tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.

Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
2026-09-16 00:32:15 -07:00
KoNit-K
91a38622db fix: guard atomic writers after profile deletion 2026-09-16 00:32:15 -07:00
xielevi
81140e4546 fix(profiles): ship a retry path for the rename identity migration
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.

- `hermes profile migrate-identity <old> <new>`: retries the migration —
  delegates to the gateway control verb while a multiplexer is live, performs the
  durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
  naming the offending database on a collision, a lock, or a partial failure. Only
  the name format and the existence of the new profile are checked; the old profile
  directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
  answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
  command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
  `click`: a failed second database raised NameError instead of printing its warning.
2026-09-16 00:32:15 -07:00
xielevi
4ba717df12 fix(profiles): migrate session/routing identity on profile rename
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.

The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:

- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
  tables, matching the agent:<name>: namespace by exact prefix (substr, not
  LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
  the profile inside routing/origin JSON, and REFUSING on a target collision
  (routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
  (keys + origin.profile) then persist — the half a DB write cannot reach.
  Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
  params only to handlers that declare them, bare handlers unchanged) so a
  live gateway rekeys its in-memory copy AND both durable stores (routing home
  + the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
  does NOT fall back to a racing CLI-side write: it prints a warning telling
  the operator to restart the gateway and retry. With no live gateway it
  performs the durable rewrite itself (safe: nothing else holds the store
  open).

Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.

Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
2026-09-16 00:32:15 -07:00
chelsealong
3b9c1118cd fix(kanban): surface the worker's own last output on a dead-worker reap
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.

Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
2026-09-15 22:47:03 -07:00
teknium1
08f36192b5 fix(config): drop the "Hermes does not read this" note on config set
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
2026-09-15 21:47:48 -07:00
teknium1
fd74bf6f58 fix(config): drop the runtime read of the docs env registry
Production code must not depend on website/docs being present; the
"Hermes does not read this name" note now keys only on the code
registry (OPTIONAL_ENV_VARS / _EXTRA_ENV_KEYS / setup-hidden suffixes).
2026-09-15 21:47:48 -07:00
teknium1
f8e8cacf35 fix(config): route every UPPER_SNAKE key from hermes config set to .env by shape
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.

- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
  into config.yaml (--force included); the env writer's denylist
  (HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
  bridging the value into os.environ; a name neither registered nor in the
  environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
  and lowercase bare keys are unchanged.

Fixes #111848 (first half landed in #112250).
2026-09-15 21:47:48 -07:00
teknium1
db25a7852e fix(plugins): install refuses to ship an unreadable plugin tree (#111804)
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.

After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.

Fixes #111804 (its discovery half landed in #112293).
2026-09-15 21:47:18 -07:00
teknium1
2dfb795cb7 fix(approval): undelivered or unanswered CLI approval prompts are not user denials
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.

- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
  sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
  approval prompt could not be delivered or was not answered (<cause>)" with
  outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
  reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
  schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.

Fixes #22992
2026-09-15 21:46:37 -07:00
teknium1
2246c245f5 refactor(update): single defer flag for the deferred catch-up; trim tests; document the flag
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
2026-09-15 19:28:39 -07:00
Turgut Kural
7b27ea3639 fix(cli): allow hermes update without gateway restart for cron (rebased on upstream/main)
(cherry picked from commit e70f78e54a96f2e8037f7e385fc31563bbeaf392)
(cherry picked from commit 225f56ab29e91977f1fef2e743c50dd6e2e89b60)
2026-09-15 19:28:39 -07:00
dmelkk-secondbrain
be9d4369a7 fix(cli): one-shot -q runs report their outcome in the exit code
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.

`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.

`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.

This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.

Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.

Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
2026-09-15 19:28:32 -07:00
teknium1
fff10484d6 fix: drop the dead disabled guard on the lazy MCP banner line
get_mcp_status reports status='disabled' (never 'lazy') for a disabled
server and derives the 'disabled' flag from that same status, so the
extra 'and not srv.get("disabled")' check could never change the branch.
2026-09-15 19:06:54 -07:00
John Paul Soliva
a17d0409be fix(mcp): report lazily registered servers as lazy, not configured or failed
A `lazy: true` MCP server registers its tools from the schema cache and
spawns on first use. Three consumers still equated "alive" with a live
session, so a healthy all-lazy startup was reported as a total failure:

- `get_mcp_status()` fell through to `status: configured, tools: 0` for a
  lazily registered server. It now reports `lazy` with the cached tool
  count (`connected: False`); an in-flight or failed first-use connect
  still outranks it because the error is the actionable part.
- `discover_mcp_tools()`'s summary counted every name absent from
  `_servers` as failed, logging `MCP: 0 tool(s) from 0 server(s) (2
  failed)` right after registering every cached tool, and re-announced
  the same "failure" on every repeat discovery. Lazy servers are now
  reported as `(N lazy, not spawned yet)` and an already-lazy server is
  not re-announced.
- `hermes_cli/mcp_startup.py` judged a discovery run by `connected` at
  two sites, so every startup logged `Background MCP discovery completed
  with zero connected servers` and every later call re-spawned the
  discovery thread as a retry. One predicate,
  `_discovery_registered_servers`, treats a lazy registration as a
  usable outcome at both sites.
- `hermes_cli/banner.py` rendered the unknown `lazy` status through the
  red "could not connect" line; it now shows the cached tool count with
  `(lazy, starts on first use)`.

Ported from #100648 (core hunks only; the toolsets-filter predicate
branch, the Ink TUI component extraction and 13 tests were not ported).

Fixes #111717
2026-09-15 19:06:54 -07:00