Commit Graph

4406 Commits

Author SHA1 Message Date
teknium1
c055eee1fe refactor(checkpoints): slim the profile-rename rekey and name it as a checkpoint_manager sibling
Follow-up to the cherry-picked fix from #112724 (@poijygfdyy):

- tools/checkpoint_profile_migration.py -> tools/checkpoint_manager_profile_rename.py, the
  repo's `<stem>_<topic>.py` sibling convention for code that extends checkpoint_manager.
- Replace the fail-closed target-collision check plus temp-file/rollback choreography with an
  idempotent rekey: every step overwrites and the old project metadata is removed last, so a
  mid-way failure is repaired by `hermes profile migrate-identity` redoing the same writes.
  A genuine collision cannot occur — `profiles/<new>` must not exist for the rename to run.
  192 -> 98 lines.
- Metadata/ledger writes go through the same idiom as checkpoint_manager itself
  (`_register_project` plain write, `_save_ledger`), dropping the private temp-file helpers.
- Keep the git-present precondition as a single early check: without git the ref cannot move
  and rekeying only the metadata would orphan the history.
- Test: `create_profile` now seeds `workspace/`, so the fixture uses `project/`; add the control
  assertion that a workdir outside the profile dir keeps its history unchanged.
- Docs: the profile rename / migrate-identity reference notes that checkpoint history is preserved.

Fixes #112973
2026-09-16 14:34:22 -07:00
NanPan
8a2c9edac6 fix(profiles): preserve checkpoint history across rename 2026-09-16 14:34:22 -07:00
KoNit-K
aa711728e8 fix(status): honor stopped gateway intent 2026-09-16 14:34:22 -07:00
teknium1
3946949646 test(profiles): prove delete_profile releases routed log handlers end to end
Replace the mocked delete_profile test (it patched release_profile_log_handlers
and only checked the call happened) with a real-path invariant: route a record
to a fresh profile — explicitly, as the Desktop cron startup does, and through
the setup_logging(hermes_home=<profile>) adoption path agent_init takes — then
delete_profile(name, yes=True) and assert no router still holds a handler or an
open stream for that home. Both parametrizations fail on origin/main (the
router's _profile_handlers keep two open streams into the removed directory)
and pass with the release in place.

A windows_only companion asserts the directory is actually removed: on win32
those streams are ConcurrentRotatingFileHandler .__agent.lock / .__errors.lock
handles, the WinError 32 the reporter hit. The Linux test covers the same seam
because rmtree of open files succeeds on POSIX and would hide the leak.

Shape follows #112543's fail-before tests by the reporter.

Refs #112538

Co-authored-by: Sora-bluesky <sora.bluesky.dev@gmail.com>
2026-09-16 14:34:22 -07:00
KoNit-K
43e28551bb fix(profiles): release routed log locks before deletion 2026-09-16 14:34:22 -07:00
teknium1
ba8b458c5b test: trim purge-identity tests to behaviour invariants
Drop the three cases that pin trivia the invariant tests already cover
(idempotence is exercised by the routing-store test; the empty-name /
no-store guard and the default-profile refusal are one-line argument
checks). Salvage bar: invariants only.
2026-09-16 14:21:14 -07:00
xielevi
aa90c01e9a fix(profiles): report a settle-pending delete as a typed partial success
A delete can end with the profile directory already removed and its durable
session/routing identity settlement still pending. delete_profile raised a bare
RuntimeError for that state, so a caller wanting to distinguish the completed
filesystem half from a real failure had only the message to go on — and the
dashboard's DELETE /api/profiles/{name} folded it into its generic 500, making a
client read the completed delete as "delete failed" (its retry then 404'd).

The partial settlement is now the typed ProfileIdentitySettlementPending — a
RuntimeError subclass carrying profile / path / retry_command:

- the CLI's delete handler keeps its non-zero exit unchanged: the type subclasses
  RuntimeError, so its existing catch tuple covers it;
- DELETE /api/profiles/{name} catches exactly that type and answers
  200 {"ok": true, "path": ..., "identity_settled": false, "settlement_pending":
  true, "retry_command": "hermes profile purge-identity <name>"};
- a genuine filesystem failure stays a plain RuntimeError and keeps the 500 — the
  API does not catch RuntimeError broadly.

Tests: the live-multiplexer delete test now pins the typed payload (profile,
retry_command, the removed path, RuntimeError subclassing); a new endpoint test
drives the real delete chain — settle-pending answers the partial success with the
directory really gone, and an rmtree failure still answers 500. Red on base (the
endpoint answered 500 for the completed delete; the type is absent) and green with
the fix: 347 passed / 0 failed across the 10 affected files.
2026-09-16 14:21:14 -07:00
xielevi
a41552fad4 fix(profiles): purge a deleted profile's session/routing identity on delete
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:

- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
  `_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
  also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
  namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
  (`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
  in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
  writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
  `_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
  has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
  partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
  profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
  `create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
  delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
  leaves of a profile's conversation record is `delete_profile`'s business — it removes the
  profile's own home, `state.db` included.

Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
2026-09-16 14:21:14 -07:00
teknium1
034313e7cd feat: plugin catalog entries carry an optional version label and card image
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).

Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.

Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
2026-09-16 14:18:39 -07:00
brooklyn!
d5b3cad7b9 fix(mcp): isolate unsupported inferred app labels 2026-09-16 15:54:47 -05:00
brooklyn!
74ca4f28ba fix(onboarding): consume future catalog entries without custom metadata 2026-09-16 15:54:47 -05:00
brooklyn!
a38d544f1b feat(mcp): expose optional backend-local app discovery 2026-09-16 15:16:09 -05:00
Yags
61154b6f68 fix(models): route Union Alpha through messages 2026-09-16 11:41:18 -07:00
teknium1
47685348ea test(update): receipt fixture sets HERMES_HOME, the seam _receipt_dir now reads
The fixture patched hermes_cli.config.get_hermes_home, which the receipt writer no longer
imports; locally run_tests.sh's temp HERMES_HOME hid the mismatch, CI counted files in the
wrong directory.
2026-09-16 11:06:52 -07:00
teknium1
5f071fb3ad fix(update): a failed dashboard cleanup no longer aborts fleet verification and the receipt
`_finish_dashboard_update_cleanup` runs pulled `dashboard_procs` inside the pre-pull
interpreter. A stale-symbol failure there (#112604: `AttributeError: module
'hermes_cli.main_dashboard' has no attribute '_loaded_launchd_backend_jobs'`, reproduced on
the maintainer's own box) propagated out of `_verify_fleet_after_update`, skipping the
fleet version matrix, plan-vs-execution reconciliation and the inner receipt finalize; the
command boundary then stamped the receipt `failed` with the traceback as stop_reason even
though the code update and gateway restart had succeeded.

The cleanup is now isolated like every sibling post-update step: the exception is
printed with a manual-restart hint, logged at WARNING, and recorded as a failed
`dashboard_cleanup` step on the receipt. A dashboard/serve left on pre-update code is still
escalated by the survivor probe → reconciliation (exit 1). Covers the git and ZIP paths.

Refs #112604 (residual class after #112753). Test is red on origin/main.
2026-09-16 11:06:52 -07:00
teknium1
1259b140c9 fix(update): the receipt survives a mixed sys.modules graph and a lost write is visible
`hermes update` writes its receipt from the PRE-pull interpreter after the post-pull
module purge. `_receipt_dir()` resolved the home through `hermes_cli.config`, so the
write re-executed the pulled `config.py` against whatever was still cached — on a pull
that added a symbol to a root-level module (`utils.file_signature`, `base_url_origin`)
that import raised and the whole receipt was dropped. The failure was logged at DEBUG,
which the updater's INFO log discards, so an activation run left no receipt while the
following no-op run wrote one, and `latest.json` kept pointing at an older run.

- `_receipt_dir()` uses `hermes_constants.get_hermes_home` (purge-protected, stdlib-only).
- A failed write prints `⚠ Update receipt not written: <exc>` and logs at WARNING.

Refs #112465, #112558 (Finding A). Both new tests are red on origin/main.
2026-09-16 11:06:52 -07:00
xxxigm
2eb5395d3f fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping (#112079)
* fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping

systemd already stops restarting on exit 78; launchd KeepAlive=true
respawned the same token/port collision every 30s. Map 78 to a clean
stop under SuccessfulExit=false so the job stays down until the holder
releases the lock.

* test(launchd): pin EX_CONFIG park under SuccessfulExit KeepAlive

The plist must not use unconditional KeepAlive, and the stderr wrapper
must turn gateway exit 78 into a clean stop without swallowing exit 75
or a non-gateway child's 78.
2026-09-16 17:58:25 +00:00
kshitijk4poor
09828fbcfc test(anon): pin the float-dust now as the literal the docstring quotes 2026-09-16 23:22:49 +05:30
kshitijk4poor
6ff1a6c8a2 fix(anon): retry_after is the whole wait, not the wait plus float dust
MintFailure.as_payload() reported ceil(not_before - monotonic()). With
not_before = now + 60, the subtraction is 60.000000000000455 for some values
of now, and the ceil made it 61: users saw "retry in 61s" and CI failed
test_the_background_loop_retries_a_transient_failure_until_it_settles
(slept == [15, 61]) whenever the runner uptime landed on such a value —
twice in a row on #113055. remaining() now rounds to the millisecond before
the ceil. One test pins a concrete `now` that produced the dust.
2026-09-16 23:22:49 +05:30
kshitijk4poor
7af776db62 refactor(doctor): reuse the repair holder-check wrapper and keep the honest disjunction in the warning
_live_writer_holds_db in hermes_state_repair already binds the repair connector
(the sibling _write_health_reason uses it); the held-branch detail now says
'or state.db cannot be inspected' so the console line matches the comment above
it. Test setup for a 51 MB WAL is one helper shared by all three tests.
2026-09-16 23:19:45 +05:30
kshitijk4poor
84a5f1cd1e fix(doctor): a large WAL under a live writer is reported as normal, never as "run --fix"
`hermes doctor` (without --fix) warned "WAL file is large — run 'hermes doctor
--fix' to checkpoint" without checking whether Desktop or the gateway held
the database; the holder scan only ran inside --fix. A large WAL is normal for
a live writer, and that nudge is how users in the #110054 threads became the
second writer that the deleted-WAL guard then fired on.

The holder scan now runs on the warn path too: held (or unprovable) reports
the size as normal while Desktop/gateway run and orders "stop" before any
--fix; no holder keeps the checkpoint suggestion, also stop-first.

Refs #110054 (item 3 of the proposed fix), #110073.
2026-09-16 23:19:45 +05:30
teknium1
a09195b1e8 fix(messaging): dual-live and cross-profile invariants for the served-profile card
Widen the #112765 regression (from #112784, @liuhao1024) to the two cases the three
community fixes split on:

- a sibling served profile with no ``<profile>:telegram`` entry in the shared record
  must not inherit another profile's connected verdict (from #112786, @kvnloo);
- the transitional dual-live case — a profile running its own standalone gateway while
  the live multiplexer still lists it in ``served_profiles`` — keeps the own record, the
  same rung order ``resolve_gateway_liveness`` uses, so liveness and platform state never
  come from two different gateways (from #112790, @kokhlo). Red on the "multiplexer always
  wins" shape, green on the precedence-mirroring one.

Also: every ``write_text`` in the file carries ``encoding="utf-8"`` (windows-footgun gate).
2026-09-16 10:27:34 -07:00
liuhao1024
1db1303964 fix(messaging): ignore a served profile's stale own runtime record
A profile served by the live default multiplexer runs no gateway of its
own, but a gateway_state.json left behind by a pre-multiplex or standalone
run is a non-None stale record, so the multiplexer fallback in
_platform_payloads (guarded by 'if runtime is None') never ran: the bare
platform-key lookup found nothing and the Messaging card fell through the
liveness ladder to pending_restart — a permanent "Restart needed" for a
platform that was connected and working (#112765).

Probe the multiplexer unconditionally: a served profile's own file always
describes a dead process, so the live multiplexer's record wins whether
the leftover exists or not.
2026-09-16 10:27:34 -07:00
teknium1
0a1716cb27 gateway status shows why an unset-default gateway stayed standalone
The boot guard logged its refusal once and nothing else knew. A user whose default says multiplex
but whose gateway serves one profile would read `hermes gateway status` and see a healthy gateway.
The verdict now lands in gateway_state.json (multiplex_standalone_reason) and status prints it with
both remedies; a boot that does multiplex clears it.
2026-09-16 08:57:18 -07:00
teknium1
a10bbf95bb feat(gateway): gateway.multiplex_profiles defaults to on, gated by a boot-time serve guard
DEFAULT_CONFIG now ships gateway.multiplex_profiles: true. GatewayConfig keeps an UNSET flag as
None so the boot can tell "the operator chose" from "the default applies"; every reader tests
truthiness, so an undecided flag never multiplexes by accident.

hermes_cli/gateway_multiplex_mode.py settles the unset default once per boot (called from
load_gateway_config_for_runner and `gateway run --config`): the same preflight `hermes gateway
migrate --multiplex` runs — default profile, >= 2 profiles, no secondary running its own gateway
(live pid or installed unit), no duplicate-credential / port-binder blocker, migratable host. A
refusal is a logged warning naming the blocker and the migrate one-liner; the gateway comes up
standalone exactly as before. Explicit values (config.yaml, GATEWAY_MULTIPLEX_PROFILES) pass through
verbatim; `--standalone` already pins false.

Other processes stop guessing the verdict from the merged default: named_profile_served_by_running
_multiplexer, the enroll warning, the dashboard listener guard, the cron-fire port resolver, container
boot and the migration plan (_read_multiplex_flag) read the live gateway's served_profiles record
first and the EXPLICIT flag second — so a per-profile fleet with the flag unset still reads as
"not yet multiplexed" and the fold proceeds.

Docs: multi-profile-gateways.md, multiplexing-gateway.md, hermes_cli/AGENTS.md.
2026-09-16 08:57:18 -07:00
brooklyn!
7e69873ca6 perf(sessions): index prompts without hydrating full transcripts 2026-09-16 06:39:15 -05:00
kshitij
16bf9a1c49 test(update): one stale-main_dashboard helper; fixture restores bindings only
- `_install_stale_main_dashboard(**attrs)` replaces the two hand-rolled
  copies of the stale-module install (built on the file's `_fake_module`).
- The per-test `finally` teardown duplicated the autouse fixture, which now
  restores the package namespace as well — dropped.
- The fixture snapshots/restores only module-valued package attributes (the
  one thing `_evict_module` touches), so any other package global a test
  leaks is still visible as pollution; the dead
  `sys.modules.get(name) is not mod` guard (always False after the
  preceding restore loop) is gone.
- End-to-end test: the process table is stubbed one layer below the module
  boundary (`dashboard_procs._iter_process_table`), which only the pulled
  scanner crosses — hermetic on the green leg, still dying on the real
  `dashboard_procs.py:366` line at main; a read of that table proves which
  module the cleanup resolved.
- Cite #112604 (the crash), not #111689 (the feature that added the symbol).
2026-09-16 16:04:27 +05:30
djsanchezsuarez
5a942f95fc fix(update): drop the stale submodule attribute in the post-pull purge
`hermes update` runs in the PRE-pull process. `_purge_stale_hermes_modules()`
evicted cached Hermes submodules from `sys.modules`, but `hermes_cli.main`
imports `main_dashboard` at CLI start, so the module also lives as an ATTRIBUTE
of the (purge-protected) `hermes_cli` package. `from hermes_cli import
main_dashboard` is satisfied by that attribute (`_handle_fromlist`), so the
post-update dashboard cleanup kept resolving the PRE-pull module and died on a
symbol the pull had just added:

    File "hermes_cli/dashboard_procs.py", line 366, in _kill_stale_dashboard_processes
        launchd_jobs = _dash._loaded_launchd_backend_jobs() if sys.platform != "win32" else []
    AttributeError: module 'hermes_cli.main_dashboard' has no attribute
    '_loaded_launchd_backend_jobs'

It lands in `_finish_dashboard_update_cleanup`, i.e. right after the update has
already reported the dashboard restart, so the code update itself succeeds but
the run aborts non-zero and the remaining cleanup/verification never happens.
The symbol came from the launchd-supervision series (#111690, merged as
#112240 for #111689); this crash is a separate stale-module handoff.

`_evict_module()` now also removes the parent package's attribute when it is
the module being evicted, so the cleanup's call-time import re-reads the pulled
source. Both regression tests reproduce the reported traceback at the pre-fix
checkout and pass with the fix.
2026-09-16 16:04:27 +05:30
teknium1
287c56e95a test(kanban): archive-terminates-worker fixture records a verified spawn fingerprint
_set_worker_pid on a fake PID (54321) now persists the 'unverified' marker, under which
the archive path correctly refuses to signal; the test pins the verified-spawn behaviour,
so it stubs the fingerprint capture.
2026-09-16 00:35:00 -07:00
teknium1
db54f5448d fix(kanban): worker fingerprint carries a boot witness; an uncaptured fingerprint never authorizes a signal
#111617 review (andrexibiza P1 #3/#4, kvnloo nit):

- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
  is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
  a reboot, and that counter does not, so an unrelated process on a later boot with the
  same PID and the same tick value passed _start_times_agree(). The fingerprint is now
  "<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
  the witness the drain marker already uses); both halves must match. Integer values on
  rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
  row and falls back to bare PID existence - a new spawn silently recreated the #89614/
  #99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
  claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
  by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
  once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
  the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
  and must not outlive its pid.

Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.

Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
2026-09-16 00:35:00 -07:00
teknium1
7dde7a2424 fix(gateway): migrate --multiplex resumes from live state, compensates the whole destructive phase, and refuses an unknown default principal
Three P1 findings from the review of #111062 (fixes #110850 remainder):

- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
  installed-but-dead unit; `systemd_install` writes the unit before the start that
  can still be killed, so an apply interrupted at default/start (or mid-restart of
  a stopped default) was reported as "already multiplexed" and never recovered.
  `interrupted` is now derived from the manifest vs the LIVE default's recorded
  served_profiles: flag on + manifest + not every migrated profile served = resume.

- the compensation `try` began after `_remove_secondary_gateways` and the flag
  write, so a later secondary's stop, a unit's daemon-reload or the config write
  failing left the first secondary removed with no rollback. The flag write,
  removals and default bring-up now all sit inside one boundary that rolls back
  through the manifest; `_preflight_apply` refuses failures knowable from the plan
  (root system unit without a recorded User=, unresolvable recorded user, an
  unwritable config.yaml) before any working gateway is stopped.

- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
  `default_uid is None`; a default system unit whose User= this host cannot
  resolve is the same boundary from the other side and now blocks the update
  hook (same known uid still folds, different known uid still refuses).
2026-09-16 00:33:03 -07:00
teknium1
5d97d5ed6d refactor(profiles): move rename identity migration off the facades; trim tests to invariants
- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
  migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
  now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
  profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
  gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
  hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
  tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.

Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
2026-09-16 00:32:15 -07:00
KoNit-K
91a38622db fix: guard atomic writers after profile deletion 2026-09-16 00:32:15 -07:00
xielevi
81140e4546 fix(profiles): ship a retry path for the rename identity migration
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.

- `hermes profile migrate-identity <old> <new>`: retries the migration —
  delegates to the gateway control verb while a multiplexer is live, performs the
  durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
  naming the offending database on a collision, a lock, or a partial failure. Only
  the name format and the existence of the new profile are checked; the old profile
  directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
  answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
  command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
  `click`: a failed second database raised NameError instead of printing its warning.
2026-09-16 00:32:15 -07:00
xielevi
4ba717df12 fix(profiles): migrate session/routing identity on profile rename
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.

The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:

- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
  tables, matching the agent:<name>: namespace by exact prefix (substr, not
  LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
  the profile inside routing/origin JSON, and REFUSING on a target collision
  (routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
  (keys + origin.profile) then persist — the half a DB write cannot reach.
  Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
  params only to handlers that declare them, bare handlers unchanged) so a
  live gateway rekeys its in-memory copy AND both durable stores (routing home
  + the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
  does NOT fall back to a racing CLI-side write: it prints a warning telling
  the operator to restart the gateway and retry. With no live gateway it
  performs the durable rewrite itself (safe: nothing else holds the store
  open).

Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.

Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
2026-09-16 00:32:15 -07:00
teknium1
2cfb655d52 test(kanban): connect via kanban_db_connect, not the facade compat pointer 2026-09-15 22:47:03 -07:00
chelsealong
3b9c1118cd fix(kanban): surface the worker's own last output on a dead-worker reap
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.

Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
2026-09-15 22:47:03 -07:00
teknium1
3abeca16e6 fix(kanban): 5xx and timeouts requeue the worker instead of spending its retry budget
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
2026-09-15 22:05:38 -07:00
teknium1
f0fd0650b5 test(update): autostash suite keeps the launchd restart scope off the host (#111866)
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
2026-09-15 21:49:33 -07:00
teknium1
08f36192b5 fix(config): drop the "Hermes does not read this" note on config set
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
2026-09-15 21:47:48 -07:00
teknium1
f8e8cacf35 fix(config): route every UPPER_SNAKE key from hermes config set to .env by shape
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.

- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
  into config.yaml (--force included); the env writer's denylist
  (HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
  bridging the value into os.environ; a name neither registered nor in the
  environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
  and lowercase bare keys are unchanged.

Fixes #111848 (first half landed in #112250).
2026-09-15 21:47:48 -07:00
teknium1
db25a7852e fix(plugins): install refuses to ship an unreadable plugin tree (#111804)
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.

After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.

Fixes #111804 (its discovery half landed in #112293).
2026-09-15 21:47:18 -07:00
teknium1
23863ccbaf fix(cli): one-shot chat -q exits non-zero on failure; 75 covers upstream 429 and overload
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
2026-09-15 21:46:41 -07:00
teknium1
2dfb795cb7 fix(approval): undelivered or unanswered CLI approval prompts are not user denials
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.

- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
  sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
  approval prompt could not be delivered or was not answered (<cause>)" with
  outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
  reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
  schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.

Fixes #22992
2026-09-15 21:46:37 -07:00
teknium1
2246c245f5 refactor(update): single defer flag for the deferred catch-up; trim tests; document the flag
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
2026-09-15 19:28:39 -07:00
Turgut Kural
7b27ea3639 fix(cli): allow hermes update without gateway restart for cron (rebased on upstream/main)
(cherry picked from commit e70f78e54a96f2e8037f7e385fc31563bbeaf392)
(cherry picked from commit 225f56ab29e91977f1fef2e743c50dd6e2e89b60)
2026-09-15 19:28:39 -07:00
teknium1
647263cca9 test(cli): trim exit-contract tests to the two invariants; fix stale budget comment 2026-09-15 19:28:32 -07:00
dmelkk-secondbrain
be9d4369a7 fix(cli): one-shot -q runs report their outcome in the exit code
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.

`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.

`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.

This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.

Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.

Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
2026-09-15 19:28:32 -07:00
teknium1
abdb402701 fix(mcp): carry the lazy status across the TUI wire, tests and docs
Follow-up to the ported status fix:

- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
  closed wire enum; `mcp.servers.status` would raise `ContractViolation`
  on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
  contract files.
- `ui-tui` session panel: an unknown status fell through to the red
  `failed` branch; render `lazy` with its cached tool count (inline
  branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
  yields `status: lazy` with the cached tool count and a summary without
  `failed` (eager control stays `configured`, live control stays
  `connected`); a lazy-only run neither warns nor re-arms the startup
  retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
  `cli-config.yaml.example`, the MCP config reference and the MCP guide.
2026-09-15 19:06:54 -07:00
teknium1
a2837ec088 docs(docker): overriding entrypoint: drops the zombie reaper — document init: true; trim the PID-1 warning
Follow-up to the cherry-picked #111584 (@chelsealong):

- website/docs/user-guide/docker.md: new warning block next to the existing
  "do not override the entrypoint" note explaining WHY (with `/init` gone the
  hermes process is PID 1 and nothing reaps orphaned browser/MCP/shell
  children), the Compose `init: true` / `docker run --init` remedy, and that
  supervision is still lost on that path; plus a Troubleshooting entry for
  `<defunct>` processes under PID 1.
- hermes_cli/main.py: `_warn_if_unsupervised_pid1` keeps the `os.getpid() == 1`
  check and drops the `platform.system()` gate and the blanket
  `try/except Exception: pass` — a user process is never PID 1 on any host OS
  (PID 1 is init/launchd; Windows PIDs are multiples of 4), and nothing in the
  check can raise.
- tests trimmed to two invariants (warns at pid 1 / silent otherwise).

Not done, on purpose: a `prctl(PR_SET_CHILD_SUBREAPER)` + SIGCHLD reaper in
main-wrapper/hermes. As PID 1 hermes already receives the orphans; what is
missing is a `waitpid(-1)` loop, and a process-wide one races
`subprocess.Popen` for exit statuses. The maintainer decides whether that
runtime change is wanted; docs + the startup warning cover the reported
deployment.
2026-09-15 19:04:00 -07:00