`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).
The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.
New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
BasePlatformAdapter._get_human_delay read HERMES_HUMAN_DELAY_MODE/_MIN_MS/_MAX_MS from the
process environment at every send, so under multiplexing the launch profile's pacing applied
to every served profile, and the documented `human_delay:` config section (mode/min_ms/max_ms,
already in DEFAULT_CONFIG) was never consulted. The runner now resolves `human_delay` per
profile through the same seam as the busy-text timings (`_human_delay_from_config`,
snapshotted in `_snapshot_profile_busy_modes`, installed by `_wire_adapter_handlers`) and the
adapter only consumes the installed range. Invalid `custom` bounds (non-integer, negative,
inverted) warn naming the key and fall back to the natural range.
Fixes#116895
The line did int(os.getenv("HERMES_MAX_ITERATIONS", "500")): agent.max_turns: unlimited
made int() raise inside suppress() so no line was logged at all, and the invented 500
default never matched what the turn loop enforces. Resolve through resolve_turn_limit
like _current_max_iterations does. The "bypass" half of #116888 is already config-first:
_bridge_max_turns_to_env writes agent.max_turns over the env slot unconditionally (PR
#18413) and routed profiles read their own config (819517fbac).
Fixes#116888
BasePlatformAdapter read HERMES_GATEWAY_BUSY_TEXT_{MODE,DEBOUNCE_SECONDS,HARD_CAP_SECONDS}
at construction, freezing the launch profile values into every profile adapter under
multiplexing. The mode was already re-synced per profile by the runner; the two timing
knobs were not. They are now display.busy_text_debounce_seconds /
display.busy_text_hard_cap_seconds, snapshotted per profile next to the busy modes and
installed by _wire_adapter_handlers. Invalid values warn naming the key and fall back.
Fixes#116893
Under gateway multiplexing, _run_agent_display_settings read
HERMES_TOOL_PROGRESS_MODE via a raw os.getenv, which answers whichever
profile's .env loaded last into the shared process os.environ. A turn
served for profile B could observe profile A's tool-progress setting.
Route the read through agent.secret_scope.get_secret, the fail-closed,
per-profile scope every other profile-varying setting in this call
path already uses (see agent/i18n.py's HERMES_LANGUAGE read).
Review follow-up for #117505: delivery_ledger, api_server_runs and
async_delegation reconciled a live owner to dead/unknown on the same 1 s
same-host fingerprint drift that cron/executions had just stopped doing.
One comparator (gateway.status.start_time_fingerprints_match) replaces the
private cron tolerance and the three exact-equality sites.
Operators expect `systemctl reload hermes-gateway` to be an in-process
config reload; the unit maps it to SIGUSR1, i.e. drain-and-relaunch. The
SIGUSR1 handler now says so on INFO when the signal lands, and the handler
factory mirrors _start_gateway_make_shutdown_signal_handler so the real
start_gateway wiring is testable.
CreateProcess resolves a bare "bash" to System32 WSL launcher before PATH,
and shutil.which("bash") inherits PATH order (#115124), so node bootstrap,
the TUI node probe and webhook filter scripts ran the wrong interpreter on
Windows. All four sites now use tools.environments.local._find_bash (Git Bash
first, probed). rc!=0 with no output at all is now a WARNING in webhook
filters and an explicit [inline-shell exit N with no output] marker in skills.
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
The base media dispatch passes is_voice= to send_voice; line, mattermost and
weixin had explicit signatures without it, so a non-image MEDIA attachment
routed as audio raised TypeError and was dropped — the same class as the
Matrix report (#102221, #116776). Adds a repo-wide signature invariant test.
Pause the chat in _typing_paused right before the first delivery attempt so a
final send that never returns (platform accepted, HTTP ack lost) cannot keep
_keep_typing refreshing an idle agent; _stop_typing_refresh still discards the
pause in the turn finally. Adds an entry-point test that drives
_process_message_background with a stalled _send_final_text.
Salvaged from PR #117340 (author email in that commit was not a valid address;
re-authored to the GitHub noreply form for the attribution gate).
A restart-safe cron worker runs outside the gateway, queues its final send in
cron/deliveries.db and polls that row once a second. The gateway drained the
queue only as a housekeeping chore at the 60 s tick, so a reminder whose turn
ended in 8 s reached the user a minute later (the issue's 60.86 s silence).
The housekeeping thread now watches the queue file (and its WAL) between
ticks and drains as soon as a stamp moved; the tick's drain stays as the
fallback. The watch never resets itself after draining: the drain's own status
write costs one empty pass, a worker that enqueued during the drain still
wakes it, and an empty pass writes nothing, so the stamp settles.
Fixes#117307
`_read_process_cmdline` tried `/proc`, then `ps`, then psutil. macOS and BSD have no `/proc`, so
every call there forked `ps -p <pid> -o command=` while psutil — a pinned core dependency, which
pyproject calls "the canonical answer" for PID questions — sat unused behind it, returning the same
string in-process.
That call is the gateway-identity check guarding against PID reuse, reached for every profile with
a live gateway, and `list_profiles()` runs it per profile. `list_profiles` is the shared body of
`GET /api/profiles` and `profiles.list`, which the Bots roster polls every 5s per connection, so
the forks repeat forever on an idle machine.
Measured on macOS: one call 4.17ms via `ps` against 0.022ms via psutil. End to end against a home
whose profiles each hold a live gateway lock, A/B in one process state, two rounds: 4 gateways
23.9ms -> 4.0ms, 8 gateways 49.0ms -> 9.3ms.
`ps` stays as the fallback rather than being replaced: on macOS psutil raises AccessDenied for a
process owned by another user, which `ps` still reports — real when a dashboard probes root-owned
LaunchDaemon gateways. That path is unchanged and simply pays a cheap failed lookup first. For a
readable process both sources return byte-identical strings, so the command-line matchers see no
difference, and Linux still answers from /proc without reaching either.
Regressions: psutil answers without any fork (a raising `subprocess.run` proves it); the `ps`
fallback still answers when psutil refuses, as it does across users; a dead PID reports nothing
from either; and both sources agree on a live gateway-shaped process. Restoring the old order
fails the first.
Fixes#117270
`_adapter_for_subscription` fail-closed on ANY connected secondary adapter of
the pinned profile: a profile that ran Signal/Home Assistant bots but held no
Telegram token (removed on purpose to avoid a duplicate-credential collision
with the shared bot) could never receive kanban notifications in a Telegram
group that `gateway.profile_routes` pins to it — the claim rewound every tick
and the docs' route-only promise could not be met because the profile is never
route-only.
Only an adapter for the subscription's OWN platform is a credential boundary
(`_authorization_adapter` already answered for it). Adapters on other
platforms no longer gate delivery; the exact-route match still authorizes the
primary solely for a chat `profile_routes` pins to that served profile — the
same authority the primary already exercises for that chat's inbound turns.
The "other-platform adapters but none for X" warning added for this branch is
superseded by delivery; the stamped-with-the-wrong-profile warning stays.
Fixes#115460
_execute_run_via_live_owner's _finish lacked the _shutdown_interrupted_run_ids
guard that the native _execute_run._finish has. On gateway shutdown
_mark_shutdown_interrupted_runs published "interrupted", then the task cancel
landed in `except CancelledError` and _finish("cancelled") overwrote it, so
`peer status` showed a shutdown-killed run as "cancelled" with no error text.
Mirror the native guard: force status "interrupted" with the shutdown error.
Local message_agent, the Desktop relay, `hermes peer dm` and `hermes peer run`
all hand a turn to the live Bot Chat owner's mailbox and then need the same
thing: poll the receipt until the owner settles it, the budget lapses, or the
caller wants out. Each lane carried (or lacked) its own copy of that loop,
which is how the class drifted — the relay answered with a receipt sentence
(#115316), the two peer lanes never asked the owner at all (#114959, #115174).
`tools/bot_live_delivery.py::await_delivery` (+ `await_delivery_async` for the
aiohttp handlers, which must not park a worker thread for 300 s) is now the one
loop; `_wait_live_dm`, `bot_relay.deliver`, `_answer_through_live_bot_chat`
and `_execute_run_via_live_owner` call it. `should_stop` carries the peer-run
`/stop` check the runs lane needs.
On Discord, Telegram, Slack and Matrix a plain `/branch` used to rebind the
CURRENT chat/thread's session key to the clone, ending the original session
on that surface. The user could not keep the original path live while
exploring an alternate one — the opposite of what a branch is for.
Now the handler opens a sibling thread through the adapter's existing
`create_handoff_thread` BEFORE cloning (a failed create never orphans a
branch row), binds the thread's own session key to the clone with the
thread's routing columns written at create time, and leaves the origin key
untouched. `/branch --here` keeps the legacy in-place switch; platforms
without threads, DMs, unknown Discord parents and adapters that cannot open
a thread fall back to in-place with a one-line note. The CLI strips the
flag through the same parser so `--here` never becomes a session title.
Destination source shapes mirror each adapter's inbound key (Discord keys
threads on their own id; Telegram/Slack/Matrix on the parent chat), the
same rules the CLI->platform handoff uses.
Live repro (real gateway + real Slack adapter against a stand-in Slack
Socket Mode/Web API): base ends the origin session and rebinds its key;
fixed posts the thread seed, replies "this chat stays on it", the origin
thread keeps its session and the follow-up typed in the new thread lands on
the branch (parent_session_id = origin).
Design and first implementation by Angello Picasso (#66014, #66024);
this is a slim port onto the split slash_commands_* layout.
Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
_pid_exists() ran psutil.Process(pid).status() (~7 ms) unconditionally before the
Windows branch — once per registry entry inside the unfair exclusive session file
lock, starving every 2 Hz poller at >=7 leases (#115578). Windows has no zombie
state, so the probe is pure cost there; pid_exists()/ctypes decide instead.
active_session_registry_snapshot() read AND pruned under _FileLock. It now snapshots
raw entries under the lock, probes liveness after release, and re-locks only to drop
the lease ids proven dead, so a lease acquired in between survives.
Tests trimmed from the contributor branch to two invariants + a POSIX control; the
Windows half is proven on the real runner via the wine2e red/green receipt.
Salvages #115591
The count was incremented on the submitting thread and released only in the
worker's finally; when loop.run_in_executor itself raised (default executor
already shut down during quiesce -> RuntimeError) no worker ever ran and
_API_WORKER_LIVE stayed elevated for the process lifetime, making the shutdown
close gate skip the SessionDB close forever. _submit_api_worker now owns both
sides of the submission.
Handler-side _inflight_agent_runs drops in the handler finally on
cancellation while the run_in_executor thread lives on, letting the
SessionDB close gate observe zero live runs under a live writer.
Count the worker lifetime itself (increment before run_in_executor,
decrement in the worker finally) and gate
_stop_quiesce_and_close_session_dbs on it alongside ctx.api_live.
Fixes#116535
With busy_input_mode: interrupt, a message B arriving while A's turn runs
redirects the turn: the answer now answers B, but the final send was still
bracketed against the event that OPENED the turn, so platforms with reply /
quote semantics (Feishu, Moti, Telegram DMs) quoted A under an answer to B
(#115001, Repro A). The reply anchor, the ledger identity and the
TurnContext anchor (read by the queued-first-response lane) were all bound
at turn start and nothing moved them on a successful redirect.
Both redirect entry points (interrupt-mode busy path and the PRIORITY path
in run_inbound) now go through _redirect_active_turn, which on success
rebinds the running turn's opening event (reply_anchor_override +
ledger_message_id) and its TurnContext to the redirecting message. The
running turn's event/ctx live on TurnState (set at claim / promotion,
cleared with the slot), and the rebind is skipped when the slot is already
owned by a newer agent, so a displaced turn's redirect can never re-anchor
a later turn.
Live repro (real GatewayRunner._handle_message with a fake platform adapter
and an agent that honours redirect): before, final reply_to='A' for
ANSWER-TO<What day is tomorrow?>; after, reply_to='B'. The busy ack was
already anchored to B on both sides.
The queue half of #115001 (both Moti quote blocks on B) is not touched: the
gateway already hands each lane a different anchor and needs the adapter
boundary datapoint the thread asks for.
Supersedes #115065 (@Finn763): same reply_anchor_override seam, without the
generation-tuple state and the queued-chain re-anchoring.
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
`_wedged_agent_count` only ever looked at chat agents, so a cron run that would
never finish (a no-agent job whose delivery hung on a dead transport) was
structurally un-skippable: `hermes update` sat in "draining" for the full
`agent.restart_after_turn_timeout` printing "0 wedged and excluded" while the
script had finished 8 seconds in.
Cron has no per-turn activity clock, but the scheduler already defines when an
in-flight claim can no longer be making progress: `sweep_stale_inflight`'s
`max(2 * interval, cron.inflight_max_minutes)` allowance. It cannot release a
claim whose worker thread is still alive, so expose that judgement as
`cron.scheduler.get_wedged_job_ids()` and let the drain count those runs as
wedged (restart is their remedy), the way it already treats idle chat turns.
`_describe_active_work` marks the cron unit `wedged` so the status line names it.
Fixes#115469 (Defect B; Defect A is the bounded standalone send this branch
stacks on).
PRAGMA quick_check(1) is one monolithic SQLite call and a lease is clamped to
900 s, so on a large store (37 GB / 4159 s and 5.4 GB / 197 s reported) the
startup watchdog fired exit 75 before the check could finish and the supervisor
restarted the gateway from zero forever. Take a phase-owned lease
("state_db_unclean_integrity_check") at entry and renew it from a SQLite
progress handler for as long as the PRAGMA is advancing; the handler always
returns 0 so it can never abort the verdict, and the ok/absent/first-complaint
contract is unchanged.
Salvaged from #115557 with the defensive try/except padding removed:
report_startup_progress is documented never-raising and a no-op when unarmed.
Refs #115542.
Co-authored-by: teknium1 <127238744+teknium1@users.noreply.github.com>
An interim edit skipped because the chat's shared send+edit slot was busy
returned a plain SendResult(success=True), so the stream consumer recorded
the never-shown text as _last_sent_text and reset _flood_strikes. A later
turn-final flood then saw _visible_prefix() == final text and either marked
the turn delivered or entered fallback with an empty continuation — the user
never saw the tail. The adapter now flags the skip in
raw_response={"skipped": True} and _edit_existing leaves the visible prefix
and flood state untouched, so the next tick retries and a flood fallback
re-sends exactly the unseen tail.
`_on_edit_failure` treated a flood-refused edit as a generic failure: a strike plus
`min(interval * 2, 10)`. Starting from the 0.8s default that is 1.6s then 3.2s, so all
three strikes (and three more refused requests, each extending the ban) were spent in
about five seconds of a penalty Telegram had already told us is 9s or longer, and
edits were then abandoned for the rest of the turn.
- The interim interval now becomes `max(doubling, retry_after)` (capped at 30s; interim
edits are skipped, not slept, so a long wait only costs a stale preview).
- `_should_edit`'s `buffer_threshold` clause no longer overrides an active flood backoff:
once the reply passed 24 codepoints every 50ms tick re-edited regardless of the
interval, which made both the legacy doubling and any server wait dead letters.
Slim redo of the retry_after half of #105340 (analysis by @AlexxRussell on #116312);
the pause/join-budget machinery there is not needed once the interval itself is honoured.
When editing content grows past a streaming edit tick, the fallback prefix
can end mid-word. Continuing from that cut drops the broken word's tail and
reads as stutter. Back the cut up to the last space or newline so the tail
of the split word is re-sent; a boundary-less prefix (one long token) keeps
the original cut instead of re-sending the whole reply.
- request_restart() opens the same drain window as stop() (new turns refused,
in-flight work awaited), so pollers of GET /v1/runs/{id} need the
shutdown_requested_at marker from that moment as well; the marker is idempotent
so stop() re-marking after the restart wait is a no-op.
- Collapse the three adapter-level tests into one invariant that drives the real
GET /v1/runs/{id} route: live run keeps status=running but carries the marker
(durably), a status set after the boundary inherits it, a terminal run never does.
- Shutdown-path fakes gain _mark_api_runs_shutdown_requested (the real mixin method
where the fake already borrows _api_server_hook, a 0-stub where it has no adapter).
- Document the field on the runs API page.
Refs #115133.
`hermes config set KEY '["-100","-200"]'` used to write the literal as one quoted
YAML string; the writer now emits a real list (acaac9a18, #88163) and the Telegram
gate decodes the legacy string shape (122ad719, #110213). The Discord, WhatsApp and
DingTalk gate parsers still comma-split that string into `{'["-100"', '"-200"]'}`,
so a config written before the writer fix silently locks every allowlisted chat,
channel or user out — with no warning.
Route every remaining comma-split gate through the shared
`gateway/platforms/_shared.py::decode_json_list_literal`:
- Discord `_gate_csv_set` (allowed/ignored/no-thread channels, allowed users/roles),
`_discord_free_response_channels` and `_missed_message_backfill_channels` now share
the one parser instead of three hand-rolled splits.
- WhatsApp `_coerce_allow_list` (allow_from, group_allow_from, free_response_chats).
- DingTalk `_csv_set` (allowed_users, allowed_chats, free_response_chats).
Plain CSV strings, YAML lists and malformed JSON keep their previous meaning.
POST /api/sessions/{id}/chat/stream is the SSE sibling of the chat route and
went _prepare_session_chat -> _run_agent with no owner check, so it stayed a
second writer into a live-owned canonical Bot Chat (#114959, "Remaining
siblings"). It now admits through the same _admit_to_live_bot_chat door; the
owner's settled receipt is streamed as the run's single assistant.completed
event, a receipt still open at the budget as run.queued (the 202 shape), a
failed one as an error event carrying the reason. The receipt wait is shared
with the JSON route (_await_live_bot_chat_receipt) and sends SSE keepalives
while it waits.
Docs: the peer dm paragraph now carries both the read-timeout wording
(#116885) and the open-chat wording in one paragraph so either landing order
resolves to this text.
`hermes peer run` posts POST /v1/runs with the peer's canonical Bot Chat as
session_id. Like the /chat transport before #114959, the run executed here
while a Desktop session held that chat's lease — a second writer the open chat
never showed, with the two transcripts interleaved in state.db.
The admission that /api/sessions/{id}/chat now performs moves onto the adapter
as one helper both peer transports call, so the two lanes cannot drift. When
the selected session is the live-held canonical Bot Chat, /v1/runs admits the
message to the owner's mailbox and drives the run from the owner's receipt
instead of an executor: `settled` completes it with the reply, a failed
receipt fails it with the owner's classified reason, and the run retires the
way an executor-backed one does. `peer run` keeps its run_id and `peer status`
keeps working; the status carries the delivery_id.
/stop cannot reach the owner's turn — the mailbox has no recall once a record
is claimed — so a stop ends this run as cancelled while the chat finishes on
its own; the stop handler already reports a run without an in-process agent as
not interruptible here.
`hermes peer dm` posts POST /api/sessions/{id}/chat. When the target is the
profile's canonical Bot Chat and a Desktop holds it live, that Desktop session
owns the chat's single-writer lease, and every other writer is refused
SESSION_NOT_OWNED — per-session exclusivity is correctness, enforced
unconditionally (hermes_cli/active_sessions.py). The API server neither takes
that lease nor checks it, so the turn ran beside the owner: the open chat never
showed the message or the reply, the live session's context never learned of
them, and two writers appended to one transcript in state.db.
Hand the message to the owner's mailbox instead, as local DMs
(tools/bot_mode_dm.py) and relayed DMs (tui_gateway/methods_bot_relay.py,
budget so the peer still gets the reply on the same call. A turn still running
at that deadline answers 202 with the delivery id, and `peer dm` reports the
message as queued in that chat instead of printing "(no reply)".
Only the canonical Bot Chat's own compression lineage is handed off: a peer turn
into any other session, or into a Bot Chat nobody holds, runs here as before.
Tests: the four-row table is the whole discriminator (owner answers, owner still
running, another session, nobody holding the chat) and the client row pins that a
queued answer reads as delivered.
filter_media_delivery_paths kept only the survivors, so a MEDIA path that did
not exist on the host (or was denied by the delivery policy) vanished with a
host-side warning while the caller got success:true / exit 0 and booked a
delivery that never happened. _validated_delivery_path and
filter_media_delivery_paths take an optional `dropped` list that collects
{path, reason}; send_message_tool passes it and, when anything was dropped,
returns success:false + partial_success:true + media_dropped + an error line,
which hermes send already turns into a non-zero exit. Slim redo of #115913
(@fangliquanflq): same payload shape, out-parameter instead of a wrapper pair.
Co-authored-by: fangliquan <fangliquan@qq.com>
HomeChannel normalization covers env and YAML homes, but cron
`deliver: discord:<link>` targets and thread metadata reach the adapter as
written and still died in int(). Share one helper
(gateway.config.discord_channel_id_from_link) between HomeChannel and the
resolver every outbound Discord target passes through; message links and
non-link strings keep their existing error path. Tests trimmed to two
invariants: config load (env + YAML) and the resolver entry point.
A bundled platform whose plugin.yaml `name:` differs from its directory
(dir `a2a`, name `a2a-platform`) could be enabled under the manifest name
— `plugins enable` writes that key and `gate_manifest` matches it — but
`Platform._missing_` only knew directory names, so `platforms.a2a-platform:`
in config.yaml raised inside `GatewayConfig.from_dict` and was silently
dropped: the gateway booted without the platform and its port never bound.
`_scan_bundled_plugin_platforms` now also collects each manifest `name:`
as an alias mapped to its directory, and `_missing_` resolves such an
alias to the directory-name member, so every consumer (registry lookup,
adapter creation, config round-trips) sees the canonical value and the
identity invariant holds for both spellings. Aliases never shadow a
directory name.
Fixes#116180
hermes_home_assignments parses a space-joined command line, so an unquoted
`HERMES_HOME=C:\Users\John Doe\.hermes` token was cut to `c:/users/john` and the
default-profile predicate (assignments non-empty, own home absent) judged the
gateway NOT to belong to its own home -- a regression from the substring test
this PR replaced. Add command_line_names_hermes_home: the token-bounded parse
first, then a token-bounded literal match of the whole home, used by both
gateway.status and the hermes_cli PID scan so the two predicates stay mirrors.
Review follow-up (#115039): a supervisor that exports HERMES_HOME with a
trailing separator (systemd Environment=, sh -c wrappers) put a spelling on
the command line the token-bounded comparison rejected, so gateway status
reported that gateway not-running for its own home. Strip one trailing
separator on both the extracted values and the profile home, keeping the
longer-sibling rejection (ops/2 is still a different home).
Also bound the assignment NAME: FOO=hermes_home=/x embeds the name inside
another token and is not a HERMES_HOME assignment (the value side already
was token-bounded; the substring test this replaces matched it).
Drive _scan_gateway_pids in a test so the mirrored predicates in
gateway.status and hermes_cli.gateway cannot drift apart again.
The named-profile branch of _command_line_belongs_to_profile tested
`hermes_home={home}` as a substring of the command line, so a process
declaring HERMES_HOME=/root/profiles/ops2 satisfied the ops profile's
predicate (the -p side already uses token equality for exactly this
reason). The default-home branch and the CLI mirror
hermes_cli.gateway._matches_current_profile had the same hole.
Extract every HERMES_HOME=<value> assignment token-bounded with quotes
stripped (hermes_home_assignments) and compare full values in all three
places. Quoted assignments (ps/wmic re-quoting paths with spaces) now
match too.
`hermes peer dm` posts to POST /api/sessions/{id}/chat, the third Bot-DM
transport. The local (`tools.bot_mode_dm`) and relayed
(`tui_gateway.methods_bot_relay`) lanes both re-run a transiently failed turn
once under the shared policy (`tools.bot_failure_reasons.retry_action`) and
resume the row the failed attempt left as the transcript's unanswered tail;
this lane ran the turn once and handed the provider's 429 paragraph to the
sender as the reply (#115325).
The policy is asked about a result dict now, not two streams: `result_retry_action`
joins `error` + `failure_reason` (the turn loop's own typed verdict) so one
classifier serves every lane, and the server-error rule accepts the providers'
`server_error` / `overloaded_error` spellings — the codes the in-process lanes
key on instead of a status number.
The resume half is the CLI lane's rule extracted to `agent.session_persistence.
adopt_unanswered_turn`, which `quiet_single_query` (env-gated dispatcher re-run)
and the API lane (in-process re-run, on the agent it just built) now share.
The regression drives the real route and the real `_run_agent` over a real
store: a 429 re-runs the same DM once with the persisted row adopted as this
turn's user message (so no second copy), a 401 still reports one attempt.
(cherry picked from commit 8fe6d46ada5b8064bc7132ee956fda998244c10a)
_prepare_runtime_status_update deep-copied the module snapshot into the new
payload, deep-copied that again as previous_payload, and deep-copied the
result back into the module snapshot; submit() then copies once more. The
module snapshot is only ever reassigned (never mutated in place), and the
transition emitter only reads previous_payload, so hand out the old snapshot
as previous_payload and store the new payload directly.
Both callers of _get_runtime_status_writer() already hold
_runtime_status_state_lock (an RLock); the flush paths read the module
attribute directly and never initialise. A second lock plus double-checked
init protected nothing, so the getter serialises on the state lock instead.
The 18-keyword signature was copy-pasted onto write_runtime_status and
publish_runtime_status and forwarded one by one. Keep the explicit
keyword-only signature on _prepare_runtime_status_update (typo safety at the
boundary) and let the two public wrappers forward **fields. No positional
callers exist (all call sites use keywords).
flush_runtime_status_async re-implemented _RuntimeStatusWriter.flush() with a
hand-rolled relay thread parked on the writer Condition plus a future and
call_soon_threadsafe/shield/wait_for. flush() already returns the same
True/False outcome at its deadline, so run it under asyncio.to_thread.
With the relay gone, wait() loops on settled() instead of re-inlining the
same predicate.