Follow-up to the cherry-picked fix from #112724 (@poijygfdyy):
- tools/checkpoint_profile_migration.py -> tools/checkpoint_manager_profile_rename.py, the
repo's `<stem>_<topic>.py` sibling convention for code that extends checkpoint_manager.
- Replace the fail-closed target-collision check plus temp-file/rollback choreography with an
idempotent rekey: every step overwrites and the old project metadata is removed last, so a
mid-way failure is repaired by `hermes profile migrate-identity` redoing the same writes.
A genuine collision cannot occur — `profiles/<new>` must not exist for the rename to run.
192 -> 98 lines.
- Metadata/ledger writes go through the same idiom as checkpoint_manager itself
(`_register_project` plain write, `_save_ledger`), dropping the private temp-file helpers.
- Keep the git-present precondition as a single early check: without git the ref cannot move
and rekeying only the metadata would orphan the history.
- Test: `create_profile` now seeds `workspace/`, so the fixture uses `project/`; add the control
assertion that a workdir outside the profile dir keeps its history unchanged.
- Docs: the profile rename / migrate-identity reference notes that checkpoint history is preserved.
Fixes#112973
Replace the mocked delete_profile test (it patched release_profile_log_handlers
and only checked the call happened) with a real-path invariant: route a record
to a fresh profile — explicitly, as the Desktop cron startup does, and through
the setup_logging(hermes_home=<profile>) adoption path agent_init takes — then
delete_profile(name, yes=True) and assert no router still holds a handler or an
open stream for that home. Both parametrizations fail on origin/main (the
router's _profile_handlers keep two open streams into the removed directory)
and pass with the release in place.
A windows_only companion asserts the directory is actually removed: on win32
those streams are ConcurrentRotatingFileHandler .__agent.lock / .__errors.lock
handles, the WinError 32 the reporter hit. The Linux test covers the same seam
because rmtree of open files succeeds on POSIX and would hide the leak.
Shape follows #112543's fail-before tests by the reporter.
Refs #112538
Co-authored-by: Sora-bluesky <sora.bluesky.dev@gmail.com>
Drop the three cases that pin trivia the invariant tests already cover
(idempotence is exercised by the routing-store test; the empty-name /
no-store guard and the default-profile refusal are one-line argument
checks). Salvage bar: invariants only.
A delete can end with the profile directory already removed and its durable
session/routing identity settlement still pending. delete_profile raised a bare
RuntimeError for that state, so a caller wanting to distinguish the completed
filesystem half from a real failure had only the message to go on — and the
dashboard's DELETE /api/profiles/{name} folded it into its generic 500, making a
client read the completed delete as "delete failed" (its retry then 404'd).
The partial settlement is now the typed ProfileIdentitySettlementPending — a
RuntimeError subclass carrying profile / path / retry_command:
- the CLI's delete handler keeps its non-zero exit unchanged: the type subclasses
RuntimeError, so its existing catch tuple covers it;
- DELETE /api/profiles/{name} catches exactly that type and answers
200 {"ok": true, "path": ..., "identity_settled": false, "settlement_pending":
true, "retry_command": "hermes profile purge-identity <name>"};
- a genuine filesystem failure stays a plain RuntimeError and keeps the 500 — the
API does not catch RuntimeError broadly.
Tests: the live-multiplexer delete test now pins the typed payload (profile,
retry_command, the removed path, RuntimeError subclassing); a new endpoint test
drives the real delete chain — settle-pending answers the partial success with the
directory really gone, and an rmtree failure still answers 500. Red on base (the
endpoint answered 500 for the completed delete; the type is absent) and green with
the fix: 347 passed / 0 failed across the 10 affected files.
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:
- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
`_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
(`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
`_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
`create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
leaves of a profile's conversation record is `delete_profile`'s business — it removes the
profile's own home, `state.db` included.
Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).
Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.
Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
The fixture patched hermes_cli.config.get_hermes_home, which the receipt writer no longer
imports; locally run_tests.sh's temp HERMES_HOME hid the mismatch, CI counted files in the
wrong directory.
`_finish_dashboard_update_cleanup` runs pulled `dashboard_procs` inside the pre-pull
interpreter. A stale-symbol failure there (#112604: `AttributeError: module
'hermes_cli.main_dashboard' has no attribute '_loaded_launchd_backend_jobs'`, reproduced on
the maintainer's own box) propagated out of `_verify_fleet_after_update`, skipping the
fleet version matrix, plan-vs-execution reconciliation and the inner receipt finalize; the
command boundary then stamped the receipt `failed` with the traceback as stop_reason even
though the code update and gateway restart had succeeded.
The cleanup is now isolated like every sibling post-update step: the exception is
printed with a manual-restart hint, logged at WARNING, and recorded as a failed
`dashboard_cleanup` step on the receipt. A dashboard/serve left on pre-update code is still
escalated by the survivor probe → reconciliation (exit 1). Covers the git and ZIP paths.
Refs #112604 (residual class after #112753). Test is red on origin/main.
`hermes update` writes its receipt from the PRE-pull interpreter after the post-pull
module purge. `_receipt_dir()` resolved the home through `hermes_cli.config`, so the
write re-executed the pulled `config.py` against whatever was still cached — on a pull
that added a symbol to a root-level module (`utils.file_signature`, `base_url_origin`)
that import raised and the whole receipt was dropped. The failure was logged at DEBUG,
which the updater's INFO log discards, so an activation run left no receipt while the
following no-op run wrote one, and `latest.json` kept pointing at an older run.
- `_receipt_dir()` uses `hermes_constants.get_hermes_home` (purge-protected, stdlib-only).
- A failed write prints `⚠ Update receipt not written: <exc>` and logs at WARNING.
Refs #112465, #112558 (Finding A). Both new tests are red on origin/main.
* fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping
systemd already stops restarting on exit 78; launchd KeepAlive=true
respawned the same token/port collision every 30s. Map 78 to a clean
stop under SuccessfulExit=false so the job stays down until the holder
releases the lock.
* test(launchd): pin EX_CONFIG park under SuccessfulExit KeepAlive
The plist must not use unconditional KeepAlive, and the stderr wrapper
must turn gateway exit 78 into a clean stop without swallowing exit 75
or a non-gateway child's 78.
MintFailure.as_payload() reported ceil(not_before - monotonic()). With
not_before = now + 60, the subtraction is 60.000000000000455 for some values
of now, and the ceil made it 61: users saw "retry in 61s" and CI failed
test_the_background_loop_retries_a_transient_failure_until_it_settles
(slept == [15, 61]) whenever the runner uptime landed on such a value —
twice in a row on #113055. remaining() now rounds to the millisecond before
the ceil. One test pins a concrete `now` that produced the dust.
_live_writer_holds_db in hermes_state_repair already binds the repair connector
(the sibling _write_health_reason uses it); the held-branch detail now says
'or state.db cannot be inspected' so the console line matches the comment above
it. Test setup for a 51 MB WAL is one helper shared by all three tests.
`hermes doctor` (without --fix) warned "WAL file is large — run 'hermes doctor
--fix' to checkpoint" without checking whether Desktop or the gateway held
the database; the holder scan only ran inside --fix. A large WAL is normal for
a live writer, and that nudge is how users in the #110054 threads became the
second writer that the deleted-WAL guard then fired on.
The holder scan now runs on the warn path too: held (or unprovable) reports
the size as normal while Desktop/gateway run and orders "stop" before any
--fix; no holder keeps the checkpoint suggestion, also stop-first.
Refs #110054 (item 3 of the proposed fix), #110073.
Widen the #112765 regression (from #112784, @liuhao1024) to the two cases the three
community fixes split on:
- a sibling served profile with no ``<profile>:telegram`` entry in the shared record
must not inherit another profile's connected verdict (from #112786, @kvnloo);
- the transitional dual-live case — a profile running its own standalone gateway while
the live multiplexer still lists it in ``served_profiles`` — keeps the own record, the
same rung order ``resolve_gateway_liveness`` uses, so liveness and platform state never
come from two different gateways (from #112790, @kokhlo). Red on the "multiplexer always
wins" shape, green on the precedence-mirroring one.
Also: every ``write_text`` in the file carries ``encoding="utf-8"`` (windows-footgun gate).
A profile served by the live default multiplexer runs no gateway of its
own, but a gateway_state.json left behind by a pre-multiplex or standalone
run is a non-None stale record, so the multiplexer fallback in
_platform_payloads (guarded by 'if runtime is None') never ran: the bare
platform-key lookup found nothing and the Messaging card fell through the
liveness ladder to pending_restart — a permanent "Restart needed" for a
platform that was connected and working (#112765).
Probe the multiplexer unconditionally: a served profile's own file always
describes a dead process, so the live multiplexer's record wins whether
the leftover exists or not.
The boot guard logged its refusal once and nothing else knew. A user whose default says multiplex
but whose gateway serves one profile would read `hermes gateway status` and see a healthy gateway.
The verdict now lands in gateway_state.json (multiplex_standalone_reason) and status prints it with
both remedies; a boot that does multiplex clears it.
DEFAULT_CONFIG now ships gateway.multiplex_profiles: true. GatewayConfig keeps an UNSET flag as
None so the boot can tell "the operator chose" from "the default applies"; every reader tests
truthiness, so an undecided flag never multiplexes by accident.
hermes_cli/gateway_multiplex_mode.py settles the unset default once per boot (called from
load_gateway_config_for_runner and `gateway run --config`): the same preflight `hermes gateway
migrate --multiplex` runs — default profile, >= 2 profiles, no secondary running its own gateway
(live pid or installed unit), no duplicate-credential / port-binder blocker, migratable host. A
refusal is a logged warning naming the blocker and the migrate one-liner; the gateway comes up
standalone exactly as before. Explicit values (config.yaml, GATEWAY_MULTIPLEX_PROFILES) pass through
verbatim; `--standalone` already pins false.
Other processes stop guessing the verdict from the merged default: named_profile_served_by_running
_multiplexer, the enroll warning, the dashboard listener guard, the cron-fire port resolver, container
boot and the migration plan (_read_multiplex_flag) read the live gateway's served_profiles record
first and the EXPLICIT flag second — so a per-profile fleet with the flag unset still reads as
"not yet multiplexed" and the fold proceeds.
Docs: multi-profile-gateways.md, multiplexing-gateway.md, hermes_cli/AGENTS.md.
- `_install_stale_main_dashboard(**attrs)` replaces the two hand-rolled
copies of the stale-module install (built on the file's `_fake_module`).
- The per-test `finally` teardown duplicated the autouse fixture, which now
restores the package namespace as well — dropped.
- The fixture snapshots/restores only module-valued package attributes (the
one thing `_evict_module` touches), so any other package global a test
leaks is still visible as pollution; the dead
`sys.modules.get(name) is not mod` guard (always False after the
preceding restore loop) is gone.
- End-to-end test: the process table is stubbed one layer below the module
boundary (`dashboard_procs._iter_process_table`), which only the pulled
scanner crosses — hermetic on the green leg, still dying on the real
`dashboard_procs.py:366` line at main; a read of that table proves which
module the cleanup resolved.
- Cite #112604 (the crash), not #111689 (the feature that added the symbol).
`hermes update` runs in the PRE-pull process. `_purge_stale_hermes_modules()`
evicted cached Hermes submodules from `sys.modules`, but `hermes_cli.main`
imports `main_dashboard` at CLI start, so the module also lives as an ATTRIBUTE
of the (purge-protected) `hermes_cli` package. `from hermes_cli import
main_dashboard` is satisfied by that attribute (`_handle_fromlist`), so the
post-update dashboard cleanup kept resolving the PRE-pull module and died on a
symbol the pull had just added:
File "hermes_cli/dashboard_procs.py", line 366, in _kill_stale_dashboard_processes
launchd_jobs = _dash._loaded_launchd_backend_jobs() if sys.platform != "win32" else []
AttributeError: module 'hermes_cli.main_dashboard' has no attribute
'_loaded_launchd_backend_jobs'
It lands in `_finish_dashboard_update_cleanup`, i.e. right after the update has
already reported the dashboard restart, so the code update itself succeeds but
the run aborts non-zero and the remaining cleanup/verification never happens.
The symbol came from the launchd-supervision series (#111690, merged as
#112240 for #111689); this crash is a separate stale-module handoff.
`_evict_module()` now also removes the parent package's attribute when it is
the module being evicted, so the cleanup's call-time import re-reads the pulled
source. Both regression tests reproduce the reported traceback at the pre-fix
checkout and pass with the fix.
_set_worker_pid on a fake PID (54321) now persists the 'unverified' marker, under which
the archive path correctly refuses to signal; the test pins the verified-spawn behaviour,
so it stubs the fingerprint capture.
#111617 review (andrexibiza P1 #3/#4, kvnloo nit):
- worker_started_at persisted only gateway.status.get_process_start_time(): on Linux that
is /proc/<pid>/stat field 22, clock ticks since THIS boot. The threat is a row surviving
a reboot, and that counter does not, so an unrelated process on a later boot with the
same PID and the same tick value passed _start_times_agree(). The fingerprint is now
"<gateway.drain_control.current_instantiation_epoch()>|<start>" (boot_id + PID-1 start,
the witness the drain marker already uses); both halves must match. Integer values on
rows written before this change keep the start-time-only comparison.
- A failed capture persisted NULL, which _pid_recycled treats as the legacy pre-fingerprint
row and falls back to bare PID existence - a new spawn silently recreated the #89614/
#99558 kill authority. A failed capture now persists UNVERIFIED_WORKER_FINGERPRINT: the
claim is held while the PID is live (never released beside it, never SIGTERM/SIGKILLed
by timeout, stale-claim, manual reclaim, archive or the terminal reaper) and reclaimed
once it is gone. NULL stays legacy-only.
- Every tasks UPDATE that nulls worker_pid nulls worker_started_at too (archive_task and
the reclaim/timeout/reopen paths): the fingerprint is part of the kill-authority tuple
and must not outlive its pid.
Live (real sleeper child): reboot-shaped row (same pid, same tick, other boot id) ->
reclaimed to ready, child untouched; matching fingerprint -> SIGTERM delivered, exit -15.
tests/hermes_cli/test_kanban_worker_pid_fingerprint.py: +2 hostile tests, both red on base.
Not changed: the check-then-act window between _pid_recycled and kill (kvnloo P2) is
real but needs pidfd_open/pidfd_send_signal (Linux 5.3+) to close atomically; left as
the documented residual of "never kills a DETECTED recycled PID".
Three P1 findings from the review of #111062 (fixes#110850 remainder):
- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
installed-but-dead unit; `systemd_install` writes the unit before the start that
can still be killed, so an apply interrupted at default/start (or mid-restart of
a stopped default) was reported as "already multiplexed" and never recovered.
`interrupted` is now derived from the manifest vs the LIVE default's recorded
served_profiles: flag on + manifest + not every migrated profile served = resume.
- the compensation `try` began after `_remove_secondary_gateways` and the flag
write, so a later secondary's stop, a unit's daemon-reload or the config write
failing left the first secondary removed with no rollback. The flag write,
removals and default bring-up now all sit inside one boundary that rolls back
through the manifest; `_preflight_apply` refuses failures knowable from the plan
(root system unit without a recorded User=, unresolvable recorded user, an
unwritable config.yaml) before any working gateway is stopped.
- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
`default_uid is None`; a default system unit whose User= this host cannot
resolve is the same boundary from the other side and now blocks the update
hook (same known uid still folds, different known uid still refuses).
- hermes_cli/profiles.py was already past the 2,000-line gate; the rename identity
migration (migrate_profile_identity, _migrate_profile_identity, control-answer helpers)
now lives in hermes_cli/profile_identity.py, imported late from rename_profile and the
profile subcommand.
- The migrate-profile-identity control-verb handler moves out of gateway/run.py into
gateway/run_profile_reconcile.py (migrate_profile_identity_verb), beside the other
hot-serve control-verb logic.
- tests/utils/ is not a mirrored source dir: the atomic-writer tests move to
tests/test_utils_atomic_writers_deleted_profile.py (root-module test placement).
- Drop duplicate no-op/idempotent/raw-answer tests so each fix carries invariant tests only.
Behaviour unchanged; authored commits from @xielevi, @KoNit-K and @kokhlo are preserved.
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.
- `hermes profile migrate-identity <old> <new>`: retries the migration —
delegates to the gateway control verb while a multiplexer is live, performs the
durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
naming the offending database on a collision, a lock, or a partial failure. Only
the name format and the existence of the new profile are checked; the old profile
directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
`click`: a failed second database raised NameError instead of printing its warning.
Renaming a profile moved profiles/<old>/ to profiles/<new>/, so the row DATA
travelled with the directory, but the profile name is also baked into
keys/values the move left untouched: session keys (agent:<old>:* namespace),
sessions.profile_name (fail-closed owner ladder / Desktop sidebar scope /
@session: deep links), sessions.origin_json.profile,
gateway_heartbeats.profile, delivery_obligations (session_key +
adapter_profile), telegram_dm_topic_* profile_name bindings, and the
gateway_routing index. Left stale, every inbound event on a chat keyed to the
old name resolved to a profile that no longer exists — flooding errors.log
with "Profile <old> does not exist ... falling back to global HERMES_HOME"
every few seconds — and renamed sessions dropped out of the sidebar / broke
their deep links.
The routing index is held in memory by a live multiplexer and written back
periodically, so a CLI-side DB rewrite alone is clobbered. Fix in layers:
- SessionDB.rekey_profile_state: atomic durable rewrite of the state.db
tables, matching the agent:<name>: namespace by exact prefix (substr, not
LIKE — '_' is a legal profile-name character and a LIKE wildcard), rewriting
the profile inside routing/origin JSON, and REFUSING on a target collision
(routing rows or telegram bindings) instead of silently merging.
- SessionStore.rekey_profile_routing: rekey the in-memory routing index
(keys + origin.profile) then persist — the half a DB write cannot reach.
Raises on a target-key collision before mutating.
- Control verb migrate-profile-identity (params-carrying; the socket passes
params only to handlers that declare them, bare handlers unchanged) so a
live gateway rekeys its in-memory copy AND both durable stores (routing home
+ the renamed profile's own state.db).
- rename_profile calls the verb when a multiplexer is live and, if it fails,
does NOT fall back to a racing CLI-side write: it prints a warning telling
the operator to restart the gateway and retry. With no live gateway it
performs the durable rewrite itself (safe: nothing else holds the store
open).
Checkpoints keyed by the profile's workdir path are a known related gap,
tracked separately, not addressed here.
Tests: rekey_profile_state (all tables, routing/origin JSON, collisions,
idempotent, no-op), rekey_profile_routing (namespace + origin, no-op, no
overwrite), control verb param passing, and rename end-to-end for both the
live-gateway (delegates, refuses unsafe fallback) and no-gateway (durable
rewrite) paths.
A `chat -q` worker's stdout/stderr go to its per-task log, so when it exits
without a terminal board call the reason is usually right there: the model's
explanation of why it could not comply (#88603) or the rendered provider error
(#46593). The reap discarded it and stamped a canned "protocol violation" /
"pid N exited with code C" on every retry. `_worker_final_output` reads the log
tail (trimming the CLI exit summary, Rich panel chrome and the session_id
trailer) and folds it into `last_failure_error` and the reap event payload as
`worker_output`, for clean exits AND crashes; `board` is threaded from the
dispatching tick so non-default boards find their own log directory.
Ported from #88815 (chelsealong) onto the decomposed dispatcher; widened to the
crash branch. Earliest attempt at the symptom: #46985 (joyjit).
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
The autouse fixture neutralised gateway discovery and the systemd branch
but not the launchd one. On a macOS host `_restart_macos_launchd_gateways`
derives its labels from the profile layout, so a default profile alone
hands it `ai.hermes.gateway`, the label never "comes back", and nine
unrelated update tests exit 1 with "Update incomplete". No OS is faked:
the seam is stubbed the same way test_update_fleet_restart_pending does.
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.
- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
into config.yaml (--force included); the env writer's denylist
(HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
bridging the value into os.environ; a name neither registered nor in the
environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
and lowercase bare keys are unchanged.
Fixes#111848 (first half landed in #112250).
A clone can land unreadable (Windows ACL inheritance -> WinError 5, a
mode-000 file). Discovery now skips such a dir instead of aborting
(#112293), but the install that produced it still exited 0, so the user
got a plugin that silently never loads.
After the clone and before anything moves into place, walk the staged
tree and open every file / list every dir. On failure repair u+rX where
the OS honours mode bits; if still unreadable raise
PluginOperationError naming the file and the fix (icacls / chmod). The
staging dir is cleaned up, nothing is installed, exit is non-zero.
Fixes#111804 (its discovery half landed in #112293).
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.
- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
approval prompt could not be delivered or was not answered (<cause>)" with
outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.
Fixes#22992
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.
`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.
`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.
This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.
Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.
Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
Follow-up to the ported status fix:
- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
closed wire enum; `mcp.servers.status` would raise `ContractViolation`
on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
contract files.
- `ui-tui` session panel: an unknown status fell through to the red
`failed` branch; render `lazy` with its cached tool count (inline
branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
yields `status: lazy` with the cached tool count and a summary without
`failed` (eager control stays `configured`, live control stays
`connected`); a lazy-only run neither warns nor re-arms the startup
retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
`cli-config.yaml.example`, the MCP config reference and the MCP guide.
Follow-up to the cherry-picked #111584 (@chelsealong):
- website/docs/user-guide/docker.md: new warning block next to the existing
"do not override the entrypoint" note explaining WHY (with `/init` gone the
hermes process is PID 1 and nothing reaps orphaned browser/MCP/shell
children), the Compose `init: true` / `docker run --init` remedy, and that
supervision is still lost on that path; plus a Troubleshooting entry for
`<defunct>` processes under PID 1.
- hermes_cli/main.py: `_warn_if_unsupervised_pid1` keeps the `os.getpid() == 1`
check and drops the `platform.system()` gate and the blanket
`try/except Exception: pass` — a user process is never PID 1 on any host OS
(PID 1 is init/launchd; Windows PIDs are multiples of 4), and nothing in the
check can raise.
- tests trimmed to two invariants (warns at pid 1 / silent otherwise).
Not done, on purpose: a `prctl(PR_SET_CHILD_SUBREAPER)` + SIGCHLD reaper in
main-wrapper/hermes. As PID 1 hermes already receives the orphans; what is
missing is a `waitpid(-1)` loop, and a process-wide one races
`subprocess.Popen` for exit statuses. The maintainer decides whether that
runtime change is wanted; docs + the startup warning cover the reported
deployment.