Follow-up to the cherry-picked fix from #112724 (@poijygfdyy):
- tools/checkpoint_profile_migration.py -> tools/checkpoint_manager_profile_rename.py, the
repo's `<stem>_<topic>.py` sibling convention for code that extends checkpoint_manager.
- Replace the fail-closed target-collision check plus temp-file/rollback choreography with an
idempotent rekey: every step overwrites and the old project metadata is removed last, so a
mid-way failure is repaired by `hermes profile migrate-identity` redoing the same writes.
A genuine collision cannot occur — `profiles/<new>` must not exist for the rename to run.
192 -> 98 lines.
- Metadata/ledger writes go through the same idiom as checkpoint_manager itself
(`_register_project` plain write, `_save_ledger`), dropping the private temp-file helpers.
- Keep the git-present precondition as a single early check: without git the ref cannot move
and rekeying only the metadata would orphan the history.
- Test: `create_profile` now seeds `workspace/`, so the fixture uses `project/`; add the control
assertion that a workdir outside the profile dir keeps its history unchanged.
- Docs: the profile rename / migrate-identity reference notes that checkpoint history is preserved.
Fixes#112973
`hermes profile delete` removes the profile directory and tears its runtime down, but the name is
also baked into durable identity the delete path never touches — `agent:<name>:*` routing keys,
`gateway_heartbeats.profile` and `delivery_obligations`. An inbound event on a chat keyed to the
dead name then enters the routing index, resolves a profile whose directory is gone, and logs
`Profile '<name>' does not exist` on every event for the life of the store (the #111926 flood,
reached from a *deleted* rather than a renamed profile). The delete side is now symmetric with the
rename rekey (`rekey_profile_state` / `rekey_profile_routing` / `migrate-profile-identity`), with
the same ownership rule:
- `SessionDB.purge_profile_state(name)` — the mirror of `rekey_profile_state`, in one
`_execute_write` transaction. Routing keys, heartbeat rows and the telegram topic rows the rekey
also owns are hard-deleted (a binding is matched by `profile_name` OR its `session_key`
namespace, because the rename rewrites both); `delivery_obligations` rows are terminalized
(`state='abandoned'`) rather than dropped, so pending delivery state is not lost silently.
- `SessionStore.purge_profile_routing(name)` — the mirror of `rekey_profile_routing`: drops the
in-memory entries and persists the drop. Mandatory, not belt-and-braces — the owning process
writes its in-memory copy back, so a durable delete made elsewhere is undone by its next save.
- A delete-only control verb `purge-profile-identity`, deliberately NOT inside
`_unserve_profile()`: that hook also unserves a rename's old name, whose identity the rekey still
has to migrate. `hermes profile delete` requires the owner's `{"ok": true}` answer and reports a
partial settlement (naming the retry) instead of a clean success.
- The retry is the new `hermes profile purge-identity <name>`. It refuses a name that is a live
profile again: the purge keys off the name alone, so `delete foo` (settlement pending) →
`create foo` → `purge-identity foo` would otherwise delete the NEW incarnation's identity. The
delete path tombstones the directory before it purges, so the guard never blocks the delete.
- `sessions` rows are not deleted by the purge: it settles identity, not history. What a delete
leaves of a profile's conversation record is `delete_profile`'s business — it removes the
profile's own home, `state.db` included.
Tests (`scripts/run_tests.sh`, red on base → green): `tests/hermes_state/test_purge_profile_state.py`,
`tests/gateway/test_purge_profile_routing.py`, `tests/gateway/test_profile_identity_purge.py`,
`tests/hermes_cli/test_profile_identity_purge_cmd.py` and `TestDeleteProfile` in
`tests/hermes_cli/test_profiles.py` — 95 passed, 0 failed across those five files.
repo social banners are 2:1, so that shape fills the slot without a crop;
the recommended size is in the README/docs. Desktop AgentPluginRow gains
catalog_version to match the generated contract.
The 40-hex sha stays the release, but nobody reads one. Entries may now add
`version: "1.4.0"` (free-form, <=32 chars, never parsed) and `image:` (an https
URL on raw.githubusercontent.com / github.com / *.githubusercontent.com).
Why GitHub-only: the Desktop catalog browser deliberately never fetches from
third-party hosts, and a raw URL pinned to the entry commit is as immutable as
the sha it decorates.
Readers updated together: PluginCatalogEntry + entry_from_mapping (drop with a
warning, entry survives), validate_plugin_catalog.py (admission error), the
site extractor (drop, never fatal), the /docs/plugins card (banner + version
pill + "1.4.0 @ abcd1234" pin), the CLI table/info (pin_label), the TUI-gateway
plugin row (catalog_version -> Desktop "Update to 1.4.0"), and the Desktop
catalog detail header (image).
* fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping
systemd already stops restarting on exit 78; launchd KeepAlive=true
respawned the same token/port collision every 30s. Map 78 to a clean
stop under SuccessfulExit=false so the job stays down until the holder
releases the lock.
* test(launchd): pin EX_CONFIG park under SuccessfulExit KeepAlive
The plist must not use unconditional KeepAlive, and the stderr wrapper
must turn gateway exit 78 into a clean stop without swallowing exit 75
or a non-gateway child's 78.
Union Alpha is OpenRouter's new $0 stealth model (262,144 ctx, tools +
tool_choice supported, text+image in). It carries no ":free" suffix, so it
needs an explicit "free, stealth model" description and an _OPENROUTER_ONLY
entry to keep it out of the derived Nous Portal list, plus a "union-alpha"
DEFAULT_CONTEXT_LENGTHS key so the offline resolver returns 262144 instead
of the 256K catch-all. Docs manifest regenerated in the same commit.
Live gate: max_tokens=64 tool-calling completion on origin/main key returned
200, echoed model stealth/union-alpha, usage.cost=0, finish_reason=tool_calls.
DEFAULT_CONFIG now ships gateway.multiplex_profiles: true. GatewayConfig keeps an UNSET flag as
None so the boot can tell "the operator chose" from "the default applies"; every reader tests
truthiness, so an undecided flag never multiplexes by accident.
hermes_cli/gateway_multiplex_mode.py settles the unset default once per boot (called from
load_gateway_config_for_runner and `gateway run --config`): the same preflight `hermes gateway
migrate --multiplex` runs — default profile, >= 2 profiles, no secondary running its own gateway
(live pid or installed unit), no duplicate-credential / port-binder blocker, migratable host. A
refusal is a logged warning naming the blocker and the migrate one-liner; the gateway comes up
standalone exactly as before. Explicit values (config.yaml, GATEWAY_MULTIPLEX_PROFILES) pass through
verbatim; `--standalone` already pins false.
Other processes stop guessing the verdict from the merged default: named_profile_served_by_running
_multiplexer, the enroll warning, the dashboard listener guard, the cron-fire port resolver, container
boot and the migration plan (_read_multiplex_flag) read the live gateway's served_profiles record
first and the EXPLICIT flag second — so a per-profile fleet with the flag unset still reads as
"not yet multiplexed" and the fold proceeds.
Docs: multi-profile-gateways.md, multiplexing-gateway.md, hermes_cli/AGENTS.md.
The desktop guide and the wake-word page told users to click an ear in the
composer row; it now fans out of the microphone on hover. The folded voice
menu, not the ear, carries the silent-mic hint.
kvnloo (#111620 review) asked for the operator-visible statement that once a serve /
dashboard process hosts a second profile, the launch profile's env-only credentials are
frozen at activation and a rotation in the process env needs a restart. Also states the
routed-child rule with and without the multiplex flag (#111617). Adds encoding='utf-8'
to the probe child in the child-env authority test (windows-footguns lint).
Three P1 findings from the review of #111062 (fixes#110850 remainder):
- interrupted-detection keyed off `plan.default.has_gateway`, which is true for an
installed-but-dead unit; `systemd_install` writes the unit before the start that
can still be killed, so an apply interrupted at default/start (or mid-restart of
a stopped default) was reported as "already multiplexed" and never recovered.
`interrupted` is now derived from the manifest vs the LIVE default's recorded
served_profiles: flag on + manifest + not every migrated profile served = resume.
- the compensation `try` began after `_remove_secondary_gateways` and the flag
write, so a later secondary's stop, a unit's daemon-reload or the config write
failing left the first secondary removed with no rollback. The flag write,
removals and default bring-up now all sit inside one boundary that rolls back
through the manifest; `_preflight_apply` refuses failures knowable from the plan
(root system unit without a recorded User=, unresolvable recorded user, an
unwritable config.yaml) before any working gateway is stopped.
- `_guard_unix_user` blocked an unknown SECONDARY principal but accepted
`default_uid is None`; a default system unit whose User= this host cannot
resolve is the same boundary from the other side and now blocks the update
hook (same known uid still folds, different known uid still refuses).
A rename under a live multiplexer that could not reach the control verb warned and
stopped there, leaving the operator with no way to finish: the rename cannot be
repeated (profiles/<old> is gone) and the CLI deliberately never rewrites the
routing DB a live gateway holds in memory.
- `hermes profile migrate-identity <old> <new>`: retries the migration —
delegates to the gateway control verb while a multiplexer is live, performs the
durable rewrite of both state DBs when none is. Idempotent, and exits non-zero
naming the offending database on a collision, a lock, or a partial failure. Only
the name format and the existence of the new profile are checked; the old profile
directory is expected to be gone.
- An older gateway that does not implement the verb is reported as such (`identify`
answers while the migrate verb does not), not as "no gateway".
- `_migrate_profile_identity` returns an explicit success/failure result so the
command can set its exit code; the rename warning now names the exact invocation.
- A failed control answer keeps the raw payload when it carries no reason field.
- The offline failure branch called `click.echo` in a module that never imports
`click`: a failed second database raised NameError instead of printing its warning.
Auto-generated display titles (LLM or derived) were passed to Honcho's
resolve_session_name() as authoritative, so a titled per-repo,
per-directory, or global session silently remapped onto a second Honcho
session named after the generated title. Only explicit /title commands
(user provenance) should act as an intentional session-name override.
Thread session_title_source from the session DB through
agent_init into the Honcho provider, and skip title-based remapping
when the source is 'derived' or 'llm'. Missing provenance keeps the
legacy explicit-title behavior for callers that predate source
threading. Gateway per-chat keys and per-session identity safeguards
are unchanged.
Adds regressions for titled per-repo, per-directory, and global
sessions at both the resolver and provider level.
Fixes#24740
(cherry picked from commit e7ba26ee15821baa382a397ce9ce9cd57a260188)
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
The code registry cannot tell a plugin-only name from one the gateway
reads straight off os.environ (TELEGRAM_GROUP_ALLOWED_USERS), so the
note was false for real settings. Every UPPER_SNAKE name simply lands
in .env; the docs say so.
`hermes config set TELEGRAM_GROUP_ALLOWED_USERS ...` (and ~290 other documented
variables Hermes reads straight from os.getenv without registering them in
OPTIONAL_ENV_VARS) still landed as a config.yaml top-level scalar with a notice,
while the setup flows write .env and one-shot CLI readers never bridge YAML
scalars — two writers, two readers. #112250 routed the registered names; this
closes the class with a shape rule: any bare ^[A-Z][A-Z0-9_]*$ key is an
environment setting.
- set: writes .env, drops a stale config.yaml copy, never writes UPPER_SNAKE
into config.yaml (--force included); the env writer's denylist
(HERMES_YOLO_MODE, PATH, ...) now refuses cleanly instead of the YAML detour
bridging the value into os.environ; a name neither registered nor in the
environment-variables reference gets a one-line note but is still saved.
- get: .env first; a leftover top-level config.yaml copy is reported as stale.
- unset: removes the .env entry and the stale copy.
- Registered names, credentials (credential lifecycle + masking), dotted paths
and lowercase bare keys are unchanged.
Fixes#111848 (first half landed in #112250).
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.
- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
approval prompt could not be delivered or was not answered (<cause>)" with
outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.
Fixes#22992
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
The 1Password page describes the minimal allowlisted child environment;
list the config-location variables so a user moving the op config dir
(unwritable ~/.config in containers) knows the setting is honoured.
Follow-up to the ported status fix:
- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
closed wire enum; `mcp.servers.status` would raise `ContractViolation`
on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
contract files.
- `ui-tui` session panel: an unknown status fell through to the red
`failed` branch; render `lazy` with its cached tool count (inline
branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
yields `status: lazy` with the cached tool count and a summary without
`failed` (eager control stays `configured`, live control stays
`connected`); a lazy-only run neither warns nor re-arms the startup
retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
`cli-config.yaml.example`, the MCP config reference and the MCP guide.
Document that an untrusted-server write-capable MCP tool now parks a run in
waiting_for_approval and is resolved through POST /v1/runs/{id}/approval,
the same bridge dangerous-command approvals already use.
Part of #111526
Follow-up to the salvaged #111796 commit:
- NO_PROXY matching goes through `agent.proxy_bypass.should_bypass_proxy` (the one
matcher the LLM transport and the gateway adapters already use), so CIDR ranges and
`*.host` patterns bypass the proxy for MCP servers exactly as they do for the model
endpoint. The stdlib `proxy_bypass` stays for the OS bypass list (Windows
ProxyOverride / macOS exceptions). Live probe: NO_PROXY=10.255.255.0/24 still routed
the MCP request through the proxy before this commit, direct after.
- Drop the try/except around `getproxies()` / `proxy_bypass()`: the stdlib guards its
own registry/sysconf reads and httpx calls the same functions unguarded.
- Trim the six contributor tests to two invariants (mount + NO_PROXY incl. CIDR; both
client builders carry mounts next to the body-cap transport). Fixture uses the stdlib
`getproxies_environment` / `proxy_bypass_environment` instead of a hand-rolled copy and
skips when the mcp SDK is absent.
- Docs: one sentence on the MCP page about proxy resolution for HTTP/SSE servers.
- contributors/emails mapping for the PR author.
Follow-up to the cherry-picked #111584 (@chelsealong):
- website/docs/user-guide/docker.md: new warning block next to the existing
"do not override the entrypoint" note explaining WHY (with `/init` gone the
hermes process is PID 1 and nothing reaps orphaned browser/MCP/shell
children), the Compose `init: true` / `docker run --init` remedy, and that
supervision is still lost on that path; plus a Troubleshooting entry for
`<defunct>` processes under PID 1.
- hermes_cli/main.py: `_warn_if_unsupervised_pid1` keeps the `os.getpid() == 1`
check and drops the `platform.system()` gate and the blanket
`try/except Exception: pass` — a user process is never PID 1 on any host OS
(PID 1 is init/launchd; Windows PIDs are multiples of 4), and nothing in the
check can raise.
- tests trimmed to two invariants (warns at pid 1 / silent otherwise).
Not done, on purpose: a `prctl(PR_SET_CHILD_SUBREAPER)` + SIGCHLD reaper in
main-wrapper/hermes. As PID 1 hermes already receives the orphans; what is
missing is a `waitpid(-1)` loop, and a process-wide one races
`subprocess.Popen` for exit statuses. The maintainer decides whether that
runtime change is wanted; docs + the startup warning cover the reported
deployment.
`hermes mcp login <server> --flow device` took `authorization_servers[0]`
from the protected-resource metadata and failed when that entry was a
browser-only or issuer-inconsistent server, even though a later entry was
the issuer-bound device_code server meant for headless clients (Higgsfield
advertises exactly this shape: a PKCE server first, the device server second).
Discovery now tries each advertised server in order and binds to the first
whose metadata issuer matches its advertised URL and that offers device
authorization. Issuer validation (RFC 8414 / SEP-2468) is unchanged per
server; a single-server resource raises exactly the error it raised before,
and a multi-server resource with no usable entry reports every attempt.
The browser path (`tools/mcp_oauth_manager.py` pre-flight) is deliberately
left on the SDK's own first-entry selection: the SDK's 401-branch discovery
re-selects `authorization_servers[0]` itself, so a divergent pre-flight pick
would only desynchronise the cached metadata from what the SDK authorizes against.
Follow-up to the cherry-picked watchdog from #112052 (@KoNit-K), finishing the
class the reporter of #112049 laid out:
- `_websocket_loop` runs the read loop and the discovery sweep as sibling
tasks and ends the connection when EITHER finishes. The discovery sweep
re-raises `ConnectionClosed` instead of logging it and retrying next tick:
a send that sees the socket closed is proof the read the loop is parked on
will never return. That is exactly the traceback the reporter watched for
22-86 h while inbound stayed silent.
- Health is invalidated while reconnecting: the first disconnect publishes
`retrying` (`_mark_degraded`) and a successful re-subscribe publishes
`connected` again. Before, `connect()` wrote "connected" once and nothing
ever changed it, so `/health/detailed` claimed delivery during the silence.
- The teardown awaits both tasks with `gather(return_exceptions=True)` instead
of a bare `except (CancelledError, Exception): pass`, which could swallow a
`disconnect()` cancellation landing mid-teardown.
- Slims the salvaged read-loop hunk: the extra "receive task remained parked"
warning and the in-loop `_mark_degraded()` are dropped; the reconnect log
line and the loop-level health flip cover both.
Docs: the Buzz page still described inbound as poll-only and the WebSocket
transport as a future optimization; it now describes the watchdog and the
`retrying` health state.
Under multiplex a secondary profile reads FEISHU_GROUP_POLICY from its own
secret scope only (deliberate 0.21.3 isolation), so a profile whose .env
carries no FEISHU_* policy keys falls back to `allowlist` with an empty
FEISHU_ALLOWED_USERS and every human group message is rejected while DMs
keep working. That deny was logged only at DEBUG, making it look like the
events never arrived (#111420).
Keep the scoped read as is — no environ fallthrough. Instead, the first
group drop caused by the untouched allowlist default logs once at WARNING
naming the chat and the keys to set (FEISHU_GROUP_POLICY /
FEISHU_ALLOWED_USERS in the profile's own .env, or group_rules in its
config.yaml). Operator-configured denies (populated allowlist, per-chat
rule, non-allowlist policy) and later drops stay at DEBUG. The predicate
lives in the topical sibling feishu_admission_diagnostics.py.
Docs: the Group Message Policy section now states the per-profile read and
where to put the keys under a multiplexed gateway.
Co-authored-by: bear0328 <bear0328@users.noreply.github.com>
Co-authored-by: NanPan <111261006+poijygfdyy@users.noreply.github.com>
`terminal(background=true, notify_on_complete=true)` appended its watcher descriptor to
`process_registry.pending_watchers`, which only the post-turn hooks drain. A process that
finished while the turn that launched it was still running (an agent sleep-polling for
hours) had no watcher task at all: the completion_queue entry sat inert, nothing was
injected, and the chat stayed mute until that turn ended (#112033).
- `_register_completion_watcher` arms the watcher on the live gateway loop at registration
(`GatewayRunner.arm_process_watcher`, via the existing `_gateway_runner_ref` /
`_gateway_loop` seam that send_message and cron already use); `pending_watchers` stays
the fallback while the gateway is not serving (checkpoint recovery at startup, shutdown).
- The agent-notify branch of `_run_process_watcher` keeps its design (the agent's next turn
is the user-facing report) but, when the launching turn is still active at process exit,
the injection only queues a follow-up — so the concise receipt is sent to the chat right
away instead of never. The busy check is taken before injection because the injected turn
itself installs the adapter's session guard.
Live probe (real process, real GatewayRunner loop, fake telegram adapter, busy session):
before — pending_watchers=1 after exit, 0 watcher tasks, 0 injections, 0 receipts;
after — pending_watchers=0, watcher task armed at launch, 1 injection, 1 concise receipt.
Control (idle session): 1 injection, 0 receipts, unchanged.
Slimmer redo of #112038 by @KoNit-K: same two gaps closed, without a second scheduler
registry / loop attribute on ProcessRegistry and GatewayRunner.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Users pointed at the Desktop app as a workaround-breaker ("don't keep the shared session open"); state
the ownership rule so the behaviour is discoverable: a chat-registered heartbeat or /loop is fired by the
gateway and replies into the chat even while a TUI / Desktop viewer has the same session open.
The previous commit claims every outbox before delivering and runs
deliveries per target profile, but `drainBusy` still spanned the delivery
phase: a `bot_relay.outbox.pending` push that arrived while one lane ran a
long turn (up to RELAY_DELIVER_TIMEOUT_MS) only set `drainRerun`, and the new
envelope was claimed after that turn — the reporter's step 3 (bot C mails D
while A→B runs) still ended in `queued_expired`, because the gateway checks
the TTL at the claim.
Scope `drainBusy` to the claim phase and make the delivery lanes module
state: `relayLanes` maps `target_connection::target_profile` to the tail of
that target's in-flight deliveries, so a later drain appends to the running
lane (same target stays ordered, one turn at a time) or starts a new one
(other targets run now). Lane entries drop once idle; stopBotRelay clears
them so a restart begins fresh.
Tests: keep the contributor's red-on-base test (claims every outbox first,
delivers to different targets concurrently) and replace the ordering-only
test — green on base — with one that pins the mid-delivery claim plus the
same-target ordering (red on base AND on the previous commit alone).
Docs: bot-mode.md states the delivery concurrency contract.
Part of #111587 (with the previous commit: Fixes#111587)