- Dialog 2 now shows the read-then-answer step for one benign prompt and
warns against blind timed Enter, and names --permission-mode acceptEdits
as the narrower opt-in (idea from PR #113456).
- Quick Reference row and pitfall #2 carry the opt-in wording instead of
teaching Down+Enter as the expected move.
- Regenerated the claude-code docs page.
_interrupt_and_clear_session popped the adapter's single pending slot and dropped whatever it
held ("consume and discard", 59575d6a91). That was right when the slot only ever carried the
user's stale follow-up text, but internal wakes (async-delegation completion notices,
kanban/cron notify+wake) now park in the same slot and are claim-settled the moment the adapter
admits them, so the pop lost them for good: the drain that runs after the command found an
empty slot and the session idled until the next user message (#114456, ~6 min stall in the
reported session).
Now the human follow-up is still discarded, but an internal wake stays parked (promoted out of
the overflow FIFO when a discarded human head occupied the slot) and the post-command drain
starts it immediately. Applies to every caller of the helper — /stop (busy fast path, handler,
pending sentinel, thread sibling), /new and /reset — since a wake that arrives a second after
/new runs against the fresh session anyway; whether a completion pinned to the closed session
may run stays with _resolve_async_delegation_session (fail-closed).
Salvaged from #114538 (@whyyagswhy, earliest filer): the pop-and-re-park mechanism is theirs;
widened here to the /new and /reset callers and the overflow promotion, and the invariant tests
rewritten against the real adapter drain. #114540 (@JoaoMarcos44) reached the same fix
independently; its stop-reason taxonomy and overflow analysis informed the class coverage.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
The reporter's case in #113372 is "not a crash": the process stays alive,
`gateway_state.json` keeps saying `running`, and housekeeping, cron and the
kanban dispatcher are frozen. Nothing wrote `updated_at` periodically, so the
file was a stored status, not a heartbeat, and both `hermes gateway status`
and `/api/status` rendered a wedged gateway as healthy (the existing stale arm
only fires when the recorded PID is gone).
Make the housekeeping tick re-stamp `gateway_state.json` first thing every
tick (60 s), so `updated_at` is a heartbeat that stops when the thread — or a
chore blocked on the loop — wedges. Readers then warn on `running`/`starting`
+ stale stamp + live PID: `hermes gateway status` prints
`⚠ Gateway heartbeat stale: housekeeping has not refreshed gateway_state.json
for N s (event loop or housekeeping wedged; pid X alive)`, `/api/status`
carries `gateway_heartbeat_stale_s` (null when healthy) and the sidebar strip
shows "Heartbeat stale". Liveness (`gateway_running`, busy/drainable) still
keys off the PID, never the stamp. Draining is excluded: shutdown drains the
housekeeping thread before the process exits, so its stamp legitimately ages.
Invariant tests: the CLI line names the age and the live PID and is silent on
a fresh stamp; /api/status sets/clears `gateway_heartbeat_stale_s` the same way.
`/api/status` mapped a not-running gateway's retained record through
`retained_gateway_state`, which only kept `startup_failed`; a watchdog-stamped
`degraded` + `exit_reason` of a dead PID became a bare `stopped` with the
reason nulled, so the sidebar strip and System page read "Stopped" while
`hermes gateway status` said "exited degraded: event loop stopped dispatching".
Retain `degraded` under the same rule as `startup_failed` (only while
`desired_state` still wants the gateway running, and only for the watchdog
reasons in the new `gateway.status.WATCHDOG_EXIT_REASONS`); the resolver
already keeps `gateway_exit_reason` for any non-`stopped` verdict. The sidebar
strip gains the `degraded` label (warning while live, destructive when the
process is gone) and the System page describes a dead `degraded` record as a
watchdog exit. `/api/messaging/platforms` keeps yielding `gateway_stopped` for
it — the channels really are down.
Invariant tests: retained_gateway_state keeps/drops the verdict by
desired_state and exit_reason; /api/status carries degraded + exit_reason for
the dead PID and stopped/null after `hermes gateway stop`.
The salvaged commit (#113513) stamps gateway_state.json `degraded` + exit_reason
inside `_mark_exited_quietly`, so both out-of-loop watchdogs (loop liveness, shutdown)
publish before os._exit. This follow-up closes the class:
- Order: lifecycle ledger first, runtime status LAST, immediately before os._exit —
the runtime record is the one housekeeping refreshes, so nothing may overwrite the
degraded stamp after it lands (ordering delta from #113386 by @KoNit-K).
- `restart_requested` is asserted only for the supervisor-restart exit code (75); the
shutdown watchdog's exit code 1 no longer clobbers a recorded restart intent to False.
- `hermes gateway status` (`_runtime_health_lines`) had no `degraded` arm, so the new
stamp would still have been invisible on the surface operators use; it now renders
"⚠ Gateway exited degraded: <operator wording for the watchdog reason>". The
startup-time `degraded` (retryable platforms queued, exit_reason None) stays silent.
- Docs: the built-in loop liveness watchdog and what it writes/exits with.
Why: housekeeping, cron and the embedded kanban dispatcher share one asyncio loop, and
liveness was a stored status field written by that loop — a dead loop left the file at
`running` (#113372). The watchdog thread already detected the stall and exited; it just
never told the file or the status command.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
- test_shell_snippets_paste_safe parametrizes over every SKILL.md under
skills/ and optional-skills/ and fails on a backslash followed by a
comment inside a fenced bash block. A fixed file list cannot catch new
instances (the notion one was missed by two earlier PRs).
- test_os_bound_skill_is_not_offered_on_the_wrong_platform drives the real
loader (agent.skill_utils.skill_matches_platform) for heartmula on
win32 and tensorrt-llm on darwin.
- Regenerated the four touched skill docs pages.
Two rejections from the tool_call bridge fuelled identical-retry loops on
small models because they stated a constraint without a correction:
* A multi-entry batch naming a local tool got "Local tools require one entry
per tool_call; mixed and multi-local batches are not supported." A 9B model
re-sent the same two-entry array until identical_call_streak_halt. The
error now restates the valid shape with the caller's OWN first entry
("you sent 2. Retry with only: {"calls":[<first entry>]} then issue the
remaining 1 call(s) as separate tool_call invocations"), which is the
hand-written correction that unstuck the reporter's session. The refusal
itself stays: b4d04eb8fd keeps local entries out of batches deliberately
(a batch bypasses per-tool admission, path-overlap serialization and the
per-server MCP parallel opt-in). Shared helper local_batch_error() feeds
both resolve_underlying_call and the connector dispatcher's guard.
* "'<name>' is not a deferrable tool. If it appears in the model-facing
tools list already, call it directly" fired for two opposite mistakes.
For a directly-listed tool the advice is right; for an unknown name —
typically a deferred MCP tool cited by its bare suffix
('mempalace_search' for 'mcp__mempalace__mempalace_search') — "call it
directly" is the opposite of what the model must do. not_deferrable_error()
now distinguishes them, suggests the registered full name when one ends
with '__<name>', and points at tool_search otherwise. tool_describe's
wrong-door error uses the same helper.
Also ports one parametrize row from #114488 (a stringified single object
without a name flows into the legacy single-shape tolerance) and adds two
invariant tests; docs updated.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
The salvaged mark-then-interrupt path covers every run whose agent exists.
A run whose executor task was created but has not taken its first tick has
no agent to interrupt: _execute_run would re-persist `running` over the
shutdown mark and start a full turn that outlives the gateway (only forced
to `interrupted` when it finally returned). Check the shutdown set before
the `running` transition and finish immediately.
Docs: `interrupted` joins the terminal status vocabulary for /v1/runs.
Co-authored-by: Mark Rockminster <markrockminster@M1-U-48-128.local>
The tracked-item delete path in quick() called shutil.rmtree() on any tracked
directory without consulting the protection list, which only the empty-dir
sweep used. guess_category() files every path under cache/ as "temp", so a
terminal command that merely mentioned $HERMES_HOME/cache tracked the
directory itself, and 7 days later quick() removed cache/ wholesale — taking
cache/terminal (terminal snapshots) with it and breaking every later command
with a mktemp "No such file or directory".
Fix: _is_protected_dir() — a tracked DIRECTORY that is HERMES_HOME itself or
sits under an _EMPTY_DIR_PROTECTED_TOP_LEVEL tree is never tracked by
guess_category(), never listed by dry_run(), and skipped (logged SKIPPED,
entry dropped) by quick(). Files under cache/ still age out as before.
kanban/ (task attachments and workspaces have their own lifecycle) is added to
both _NEVER_TRACK_TOP_LEVEL and _EMPTY_DIR_PROTECTED_TOP_LEVEL; stale pre-fix
"test" entries under it are dropped by the existing re-validation instead of
deleted. The kanban row was first proposed in #80842 (@nicha16).
Fixes#114552
Follow-up to the two salvaged commits. The nudge text asserted the card was
"still `running`" and offered only kanban_complete / kanban_block, so even
when it fired legitimately it steered a review-bound card toward a false
completion. It now states what the transcript actually shows (no terminal
board call yet) and lists kanban_complete / kanban_request_review /
kanban_block, plus the reviewer exits.
Every remaining copy of the "terminal tools" knowledge is brought in line:
turn_stop_gates docstring + diagnostic status, goals.py finalize comment,
kanban_db_dispatch grace/concurrency comments, the orchestrator-only
refusal in kanban_tools, and the kanban / tutorial / codex-runtime docs.
Fixes#114598
Salvage follow-up to the cherry-picked #114529 (@Roblmvp):
- `hermes kanban block --kind dependency` now says the card was blocked as
needs_input because no parent is open, and the triage verdict keys off the
landed block_kind (a re-kinded dependency block is a human question).
- kanban_block tool: report the landed kind + requested_kind + a note telling
the worker why it did not park in todo (dropped the redundant rekind_reason
echo — it lives in the event payload).
- Tests trimmed to two invariants in test_kanban_block_kinds.py (terminal
parents -> blocked/needs_input -> triage at BLOCK_RECURRENCE_LIMIT; open
parent -> todo with no recurrence across a dispatch tick) plus one tool
surface test.
- Docs: kanban_block row, `blocked`/`dependency_wait` event rows.
- contributors/emails mapping for the salvaged author.
Fixes#114627
_hmwa_heal_telegram_topic_binding awaits get_telegram_topic_binding and get_compression_tip between reading the route and calling switch_session, the same shape as the async-delegation re-pin fixed in this branch. Pass the snapshot session id as expected_session_id so the compare-and-swap refuses to overwrite a concurrent /new or /resume. The developer-guide Operations table now documents the CAS kwarg.
Follow-up to the salvaged #114614 pick:
- `_codex_login_post`: the `for` loop ended in an unreachable `raise … # pragma: no cover`
(needed only to satisfy the return type). A `while True` with the terminal condition
folded into the except branch has no dead path and the same three-attempt bound.
- `_codex_poll_authorization_code`: the pick duplicated the `except KeyboardInterrupt`
handler; the second copy was unreachable.
- Tests: six change-detector tests collapsed into two invariants (poll survives blips but
never retries a non-transport exception; one-shot POST retries once and keeps the typed
AuthError + TLS hint + cause chain at the cap). The existing
test_codex_device_login_ssl_hint.py still pins the poll's terminal hint path.
- Docs: providers.md notes that a single dropped connection during device login is no
longer fatal.
The timeout text tells the agent not to retry, and users read that as a
session-wide lock with no way back (#114204). Document the supported
path: a new user message raises a fresh tool call and a fresh approval
card; timeouts do not count toward the denial breaker.
Fixes#114204
Document satellite heartbeat output, late and missed-fire findings, and doctor among the last_fire_error surfaces. Retain the R2-2 ruling and explain the watchdog consequence.
Reported-by: kshitijk4poor
Remove the keyless free tier from the provider table (providers.md,
fallback-providers.md), the .env.example section, and the 'three OpenCode
providers' wording now that only Zen and Go ship built-in.
The thread-sibling tier short-circuited the chat-scope tier, so when a per-user
thread sibling was live the handler interrupted ONLY that run and replied
"Stopped" while a same-thread run under a differently shaped key — the
channel-keyed #286 shape this fallback exists for — kept going. The chat tier is
a superset of the sibling tier (a sibling needs the caller's own thread slot,
which satisfies the chat predicate), so it is the set acted on; the reason still
resolves to `stop_command_thread_sibling` when the chat tier found nothing beyond
siblings, so hook consumers keep that label.
Both tiers also share ONE scan now (they each called `_same_chat_runs`, building
the running-agent snapshot twice), and the matcher is a single text pass:
`_same_chat_key_slots` takes the caller's key prefix instead of re-deriving the
namespace per key, and the whole-slot rule lives once in `_strip_slot` (the tail
check uses it too). Concrete annotations on the three methods (`List[...]` was
missing from the typing import), and the prefix construction reads as a named
namespace rather than a nested f-string.
WhatsApp DMs: `build_session_key` canonicalises the DM chat id, so the fallback
matched nothing there — it now canonicalises the same way, which makes the
chat-scope fallback reach a run keyed from any JID/LID alias.
Docs: `sessions.md` described the tiers sequentially and over-claimed the
widening's reach — an in-thread stop reaches that thread plus a thread-less
room-wide run, never another thread or a peer's per-sender top-level run.
`hooks.md` pointed at `gateway/run.py` for `_interrupt_and_clear_session`, which
lives in `gateway/run_agent_cache.py`.
Guards: the partial-stop case (per-user sibling + same-thread channel run, both
must be interrupted), the lone-sibling reason label, and the WhatsApp canonical
id. Each is mutation-checked individually: restoring the short-circuit, forcing
the reason to chat_scope, or dropping the canonicalisation each fails exactly
its own test. 21 passed in the two /stop test files.
Two defects found by the pre-arm review gate on the matcher above.
- The scope-slot branch ran on every platform, but only Slack ever emits that
slot (build_session_key appends scope_id for Slack alone). A non-Slack DM keys
chat_id as the USER id, so a per-sender group key `group:<chat>:<user>` parsed
as "<user>'s chat": a Telegram user's /stop in her DM interrupted her group
run in a different chat. The branch is now Slack-only.
- The thread boundary only fired for `thread`-typed keys, so a top-level channel
turn — which keeps chat_type "channel" with the relay-stamped reply-thread ts —
stayed reachable from a /stop in a DIFFERENT reply thread of the same channel.
The boundary now keys on the caller's thread alone: a stop sent from inside a
thread reaches only runs whose own first trailing slot is that thread, or that
carry no trailing slot at all (the slotless rolling-DM shape).
Guards: one test per defect (Telegram DM vs group run ending in the same user id;
channel-keyed run belonging to another reply thread), and the workspace/profile
negative now keeps a same-chat run live so a matcher that returns nothing can no
longer satisfy it. Each guard was mutation-checked individually.
- Negatives strengthened: the isolation tests now keep a same-chat run AND a
foreign run live and assert exactly which one was interrupted (the previous
version also passed on base and would have passed for any early return);
added the chat-id prefix case ("C90") and the workspace-scope + profile case.
- New: a different thread of the same channel is never reached; a pending
sentinel is never claimed as stopped; a peer's per-sender group run IS
reached (documents the intentional widening of #113846).
- Assertions now read the real reply (`t("gateway.stop.*")`) and the invalidation
reason instead of an English substring, so a locale or copy change can't
silently satisfy them.
- Docs: `sessions.md` no longer claims interrupt handling is strictly per-user;
`hooks.md` lists the new `stop_command_chat_scope` reason.
Slack seals a native stream server-side after a few minutes (live-observed
at ~5m20s on three independent long turns, 2026-09-15/16; the lifetime is
not documented). The next chat.appendStream on the card fails with
message_not_in_streaming_state. The adapter returned a bare failure, the
TurnRunner latched native_failed, and the rest of the turn rendered as an
edited text bullet list. Long autonomous turns lost the card UX exactly
when it mattered.
On message_not_in_streaming_state from appendStream, drop the dead
stream_ts and chat.startStream a fresh plan-mode card in the same thread,
then append the current frame there. Every frame already carries the full
visible task projection, so no task state is lost. One reopen per update;
a second rejection surfaces as a real failure. The sealed card is a plain
message now, so no stopStream is sent to it; the turn-final stop targets
the reopened card. Error matching reads SlackApiError.response["error"],
never the message text.
Tests assert the wire sequence (start, append, rejected append, start,
append on the new ts), the reopened frame's task states, the cache pointing
at the new card, and the stop targeting it; plus the one-reopen bound.
Mutation: forcing the expiry branch off turns both tests red.
- `is_diagnostic_notice()` replaces three drifting copies: the gateway muted every
`credits.*` notice, the TUI and CLI only `warn`/`error`, so `credits.restored` was hidden
on Telegram and shown in the TUI for the same config. Every credit-service notice is an
automatic diagnostic (a "restored" line after a hidden depletion notice is orphan noise).
- `effective_user_config()` is the single fail-open effective-config read; the two extra
`deepcopy`s per foreground turn go (the loader already returns a fresh copy and the
snapshot is read-only).
- `diagnostic_metadata(event)` replaces the repeated
`{"notification_category": "diagnostic"} if event.internal and ... else {}` literal in
gateway/run_turn.py; the gateway-side imports of the resolver are module-level (no cycle:
it imports only gateway.display_config).
- `display.suppress_warning_notifications` is listed with its sibling display keys in
cli-config.yaml.example and the configuration reference; the messaging guide states that a
muted diagnostic wake still runs (and bills) its agent turn.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
Discord paces a bot's sends at roughly one per second, so chunk 3 of a long
handoff lands after the 2s window anchored on the tag and was still dropped —
the symptom the continuation window exists to fix. Each admitted continuation
now re-arms the window; the gateway bot loop guard bounds a bot that never
stops. The flush-delay override moves onto the BasePlatformAdapter seam
(_text_batch_delay_for) that main relocated the batcher to.
Tests trimmed to the invariants: reply-ping-only bot message rejected by
default (and admitted with the explicit opt-out), and a 3-chunk tagged
handoff paced past the original window arrives as one batched event; the
defensive getattr fallbacks that only served test construction are gone.
* refactor(desktop): extract titleBarOverlayOptions into a tested helper
getTitleBarOverlayOptions in main.ts branched inline on mac/windows/wsl.
Move the decision into titlebar-overlay-width.ts so the WSLg → false rule
(renderer paints its own controls there) is unit-tested next to the width
reservation it pairs with. main.ts keeps a thin caller.
Co-authored-by: null-runner <nicholas.mariani@hotmail.it>
* fix(desktop): renderer-drawn window controls on WSLg
Under WSLg the frameless window (titleBarStyle: 'hidden') had no
minimize/maximize/close. getTitleBarOverlayOptions returned false on the
premise that the RDP host paints replacement controls; it does not for a
frameless window. Electron's native overlay is not the fix either: its
cluster's hit-region drifts from the rendered buttons under the RAIL
compositor.
The renderer now paints Windows-style min/max/close (wslg-window-controls)
routed over a hermes:window-control IPC channel (window-controls.ts →
preload → main). getWindowState reports customWindowControls/isMaximized so
the renderer knows when to mount them and which glyph to show. The cluster
is pinned to TITLEBAR_HEIGHT px (the contrib shell zeroes --titlebar-height
for content subtrees) with 46px native-width caption buttons.
Co-authored-by: Austin Pickett <pickett.austin@gmail.com>
* fix(desktop): wire WSLg window-control IPC and snap maximize onto the work area
Register hermes:window-control in main, report windowControlState from
getWindowState, and re-send window state on maximize/unmaximize so the
renderer's maximize/restore glyph tracks the window.
maximizedBoundsCorrection snaps a WSLg-maximized frameless window back onto
the display work area (RAIL can settle it offset — reported on WSLg 1.0.65).
It is a no-op wherever native maximize already fills the work area, so
healthy compositors are never fought and setBounds cannot loop.
Co-authored-by: null-runner <nicholas.mariani@hotmail.it>
* fix(desktop): call WSLg window-control bridge without the click event
contextBridge structured-clones every argument, and a React SyntheticEvent
is not cloneable, so onClick={controls.minimize} threw "An object could
not be cloned" before ipcRenderer.send ran. The buttons rendered but did
nothing. Wrap the handlers so the bridge is called with no arguments, and
pin that in the test.
* fix(desktop): remove recursive WSLg maximize bounds correction
* fix(desktop): select native Wayland before Electron initialization on WSLg
* fix(desktop): preserve ready pipe and debugger across WSLg launch
* fix(desktop): keep WSLg caption controls scoped to every window
* style(desktop): order window chrome event import
---------
Co-authored-by: null-runner <nicholas.mariani@hotmail.it>
The updating guide describes the rename-over-previous-build step; mention that a
scanner briefly holding release/win-unpacked is ridden out with short retries
(#112544) so the paragraph matches the promotion behaviour.
Two open atoms of #83390 (DeepSeek "This response_format type is unavailable now"):
* `_call_fallback_candidate_sync/_async` only special-cased auth errors, so when the primary
aux provider failed (timeout, rate limit, payment) and the fallback landed on a provider that
rejects `json_schema`, the 400 re-raised and the whole task died — the primary-path rung from
#89589 never applied there. Both fallback paths now retry once without `response_format`.
* Every structured aux call (titles, kanban decomposer, goal judge, plugin structured calls)
paid a guaranteed-fail request on providers that lack `json_schema` before the retry. A
provider profile can now declare `unsupported_response_formats` (DeepSeek: json_schema, per
https://api-docs.deepseek.com/guides/json_mode) and the recovery ladder remembers any route
that rejected a type once (host:port scoped), so `_build_call_kwargs` — shared by the primary
and fallback paths — omits the field before the first request. Dropping rather than
downgrading to json_object matches the end state the retry already produced; json_object
needs a JSON-mentioning prompt and some relays return empty content under it.
New logic lives in agent/auxiliary_structured_output.py; the facade only gains the fallback rung
next to the predicate it uses. tests/agent/conftest.py resets the process-level memo per test.
Fixes#83390, #105191. Closes duplicates #84976, #88830, #102849, #113064.
Co-authored-by: Legion-is-life <Legion-is-life@users.noreply.github.com>
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).
- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
inherited UI transport label without an explicit --source) resolves to `oneshot`; the
platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
(HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
`sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
`--source` flag reference (explicit flag always stored as given).
The late-failure watch treated every non-"sent" outcome as definitive: a relay
lost-ack (raw_response.ambiguous=True, the card may well have posted) arriving
after SEND_ACK_WINDOW tore down the registration and returned the delivery
notice, so the user's button tap on a card that WAS rendered found no pending
entry and was lost. That breaks the invariant _abort_for_outcome (and main)
keeps for an immediate ambiguous outcome. _on_card_done now returns early on
"ambiguous" and lets the bounded wait's own timeout cover a card that truly
never arrived.
Also while here:
- a late DECLINE releases with UNDELIVERED_DECLINED, matching the immediate
path, instead of the generic notice (_release carries the outcome);
- when the text fallback cannot even be scheduled (fallback() -> None) the
card's "failed" verdict stands instead of re-classifying None, which logged a
misleading "no scheduling future (loop unavailable)";
- test_card_failing_after_the_ack_window_... now asserts the text prompt was
sent exactly once and the wait released within ack window + 2 s, and the
helper answers only after observing the prompt — with fallback/late-watch
disabled it stayed green before (the helper answered on a 10 s deadline);
- telegram.md: clarify_timeout default is 3600 s (resolve_clarify_timeout,
configuration.md), not 600.
Part of #112684
A native clarify card on a messaging platform (Telegram inline keyboard, Slack blocks) could
fail without the user ever seeing a question, and the agent then waited out the full
clarify_timeout and reported "[user did not respond within Nm]" (#112684):
- the platform rejects the card at once -> the wait aborted with a delivery sentinel but the
question was never re-asked;
- the send outruns the 15 s acknowledgement window and only then fails (pool/connect timeouts,
a stale-thread retry) -> the runner never looked at the send future again and blocked for
the whole timeout for a card that never posted;
- no status adapter at all -> _ask_clarify_question returned ('', False), so a batch reported
timed_out=True with an empty notice.
gateway/run_turn_runner_clarify_delivery.py (new topical sibling; the send-disposition helpers
move out of the gateway/run.py facade) now retries every definitive card failure once through the
adapter's plain-text send_clarify (numbered list + text capture) - except a connector egress
DECLINE, where re-sending the text is the exfiltration the guard exists to stop - and watches a
possibly-delivered send so a late failure releases the waiter with
"[clarify prompt could not be delivered]". The no-surface case reports
"[clarify prompt could not be delivered: no chat surface]" (same prefix every consumer already
treats as a non-answer).
Live probe (real TurnRunner + real clarify_gateway, Telegram-shaped adapter, timeout 30 s):
card fails after 16 s -> before 45.2 s and "[user did not respond within 0m]", after 16.7 s and
the typed answer to the text prompt; card rejected at once -> before sentinel with no text prompt,
after the numbered prompt is sent and "2" resolves to the choice.
Why: an integration that answered server→client requests but never sent
client.capabilities now has every clarify/approval/sudo/… request refused at
once with no grace path. Marking a connection as answering after its first
response frame cannot help — the frame is never sent to it — so the break is
documented instead.
Item 2 of #112548: a Desktop/dashboard build that predates server→client
requests has no response path, so every clarify/approval/sudo/secret/vault/
connection/bridge request sat for the full deadline (clarify: 300s). Only the
tour probed. Clients now advertise once per connection
(`client.capabilities {server_requests: true}`, sent by the shared TypeScript
channel on `gateway.ready`); `send()` / `send_async()` return the
error-response shape (None) at once when every WebSocket peer of the session
is a build that never advertised. Sessions with no client attached still wait
so the reconnect replay (`open_requests`) keeps working; the stdio TUI ships
with the backend and is not gated. The advertisement is dropped on disconnect.
Reviewer minors from #113227:
- tools/approval_gateway_wait.py: the verdict is the choice committed under
the approval lock while leaving the queue, so an /approve that lands after
the deadline check but before the entry is dropped is an answer, not a
timeout (the client was already acked "ok").
- tests/tui_gateway/test_protocol.py: the error-fails-fast test that only
restated pre-existing behaviour is replaced by the two capability
invariants (never advertised → fails fast; advertised → frame written,
waits, forgotten on disconnect).
- server_requests.send try/finally around event.wait already landed on main
(4371ed34a9); nothing to change.
Docs: programmatic-integration.md (advertise once per connection; method
list), tui_gateway/AGENTS.md; contracts regenerated.
`hermes dashboard --stop` / `hermes update` signal only the backend PIDs. When the
lifespan teardown wedges (stop_hosted_room_service, PTY close_all never reached) the
10s grace loses, SIGKILL lands mid-teardown, and the hosted ui-tui / tui_gateway.entry
child is reparented to init still holding the state.db-wal inode; the next start
refuses with DeletedWalGenerationError (#112631, residual of #111912). No finite
root grace covers an unbounded teardown.
_kill_pids_posix now snapshots the dashboard-owned descendant tree BEFORE the kill
(the PPID link is gone once the root dies), and after the root phase SIGTERM→SIGKILLs
the descendants that are still alive, waiting for the tree to be gone before
returning. Descendants are re-checked against their snapshotted start-time
fingerprint (the same PID-reuse guard _kill_pids_windows uses) — no `ps -o lstart`
per system PID.
Detached session leaders without a controlling terminal are pruned from the sweep
with their subtrees: those are the messaging-gateway bots and profile actions the
dashboard launched with start_new_session from /api/gateway/*, which belong to the
user, not the dashboard. A hosted TUI is a session leader too (pty.fork) but owns
the pts whose master the dashboard held, so its tty column is set and it is swept.
The Desktop boot reaper (_reap_orphaned_desktop_local_serves) SIGKILLs the same
class of backend after 1.5s and had the same hole; it now SIGKILLs the surviving
snapshotted descendants (no second grace: the boot path runs under a 10s probe).
Live: pyteman hermes-111912 wedged leg on origin/main
`child_orphan_alive=True deleted_sidecar_holders=2 guard=FATAL DeletedWalGenerationError`,
on this head `child_orphan_alive=False deleted_sidecar_holders=0 guard=clean`; the
start_new_session control sibling survives on both.
The warn-once reporter added for hooks (and for PluginManager.invoke_middleware)
left the execution chain out: hermes_cli/middleware.py::_run_execution_chain — the
tool_execution / llm_execution frames that run once per tool or LLM call — still
logged "Middleware '%s' callback %s raised" at WARNING on every invocation. A
mis-declared callback (e.g. a signature naming tool_data) therefore flooded the
log exactly like the hook case #111922 reported: 5 calls -> 5 WARNING lines.
Route the frame's except through the manager's _report_hook_failure with the
"Middleware" surface label, so the first failure warns (listing the fields the
middleware does provide) and identical repeats go to DEBUG; the set is already
forgotten on plugin reload. The frame's skip-and-continue semantics are unchanged.
Part of #111922
User-visible knobs need a home: the Vision feature page explains why native embeds ride
the session and what each key does (subagent-only default cap, clamp range), the
configuration reference points at it next to auxiliary.vision so the two sections are
not confused, cli-config.yaml.example carries the commented block, and tools/AGENTS.md
names vision_tools_history_budget.py as the single owner of embed-cost policy.
A terminally rejected refresh token (invalid_grant / invalid_token /
refresh_token_reused) is the moment a login is lost. Three gaps remained
after #113197 (which added the WARNING for Codex/xAI/Nous):
- Anthropic: the token endpoint's HTTPError carried no classifiable code,
so a dead Anthropic / Claude Code grant fell through to a transient
'exhausted' bench at DEBUG and was replayed every hour. The endpoint
error is now a structured AnthropicOAuthError (status + OAuth error
code); _recover_failed_refresh logs one WARNING with the repair command
and marks the row DEAD. A dead grant is not replayed at the fallback
endpoint. The auxiliary Claude Code refresher (sibling path) warns the
same way. Claude Code's own credentials file is never touched.
- Codex/xAI/Nous: _quarantine_sources drops only singleton-seeded rows, so
an independent `hermes auth add` (manual:*) login survived unmarked and
re-fired the WARNING on every later refresh attempt. The survivor is now
marked DEAD (leaves rotation until a write-side re-auth clears it).
- Hints on live Hermes paths (anthropic 401 troubleshooting block, the
no-credentials error) recommended an external CLI's login command,
which cannot repair Hermes' own login; they now say
`hermes auth add anthropic` / `hermes auth list anthropic`.
Docs: credential-pools.md documents the dead-login behaviour.
Part of #113023.
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).
WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.
Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.
Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
Closes the remaining atoms of #112600.
A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
its siblings still resolved credentials against config's `default`: the CLI
auth-fallback rung, `--resume` credential re-resolution, the gateway
provider-override helper (channel overrides, persisted /model switches,
API-server provider refresh), the gateway fallback chain, the TUI /model
switch-from runtime and ACP agent construction. With a `*-free` default the
OpenCode free-tier rung fired first and a Go-only model was built against
the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
optional `target_model` and the two test stubs of it accept the kwarg.
B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
provider matched by opencode_provider_family, including custom providers
merely named after a family (`opencode-go-bridge`, #85589) whose relay the
user declared explicitly in `providers:`. The family heal now applies to the
built-in canonical providers only; custom prefix-named providers keep their
per-model api_mode routing and /v1 handling. Documented in the providers
guide.
C) Same function: the official-host check uses parsed.hostname (a port no
longer defeats the heal) and only the path is edited, so query/fragment
round-trip instead of being dropped.
Fixes#112600
Follow-up to the cherry-picked #106018 (@liuhao1024). The alias-table fix resolved the
keyless lane, but the reporter's second variant (provider: ollama + bare base_url +
explicit api_key → 404) still failed through the real config path:
_resolve_task_provider_model collapses "base_url + api_key" lanes to "custom" before
resolve_provider_client runs, so the custom branch never saw the ollama alias and
skipped the /v1 tail. Keep the local-server alias identity on both base_url paths
(config lane and explicit kwargs via _preserve_provider_with_base_url) so the tail
applies whenever the user wrote provider: ollama/vllm/llamacpp.
Shape: inline the 5-line _bare_host_base_url helper (urlparse never raises on str;
no new facade helper), trim the five contributor tests to two invariants that drive
the reporter's lane through _resolve_task_provider_model, and document the alias
group under the auxiliary provider list.
Live: stand-in mirroring Ollama's /v1 OpenAI-compatible surface — before: RuntimeError
"no API key was found" (keyless) / 404 on /chat/completions (explicit key); after: both
POST /v1/chat/completions and return OK, sync and async; provider: custom + /v1 unchanged.
The dispatcher forced `-Q` on goal_mode cards because cli.py only ran the
kanban judge loop in the fully-quiet one-shot branch. `-Q` strips every tool
callback, so the dashboard's Worker log (and `hermes kanban log`) stayed
blank for the whole run while non-goal cards logged normally; users read
that as "the worker is doing nothing" (Discord report, Sep 2026).
Run the judge loop on the `-q` path too, driving follow-up turns through
cli.chat so each turn's tool activity lands on stdout (= the worker log),
and print each judge verdict there as well. The dispatcher spawns goal_mode
and one-shot workers with the identical argv; the mode travels only in
HERMES_KANBAN_GOAL_MODE. The `-Q` hook stays for manual quiet runs.
Live A/B (real `hermes kanban dispatch` + spawned worker against a scripted
loopback provider that answers a terminal tool call then text, judge always
"continue", --goal-max-turns 2): base log = 94 bytes, 0 tool-feed lines,
0 verdict lines; fix log = 1910 bytes, 4 tool-feed lines, 3 verdict lines;
both arms end blocked "exhausted 2/2 turns" (loop behaviour unchanged).
The external cron worker is spawned as `sys.executable -m cron.scheduler`. Its
entry module is `cron.scheduler`, not `hermes_cli.main`, so it never runs the
bootstrap that puts the checkout on the gateway's sys.path; it only imported
`cron` at all through the implicit `-m` cwd entry (cwd was already the repo
root). That implicit path breaks on real hosts: a venv whose editable install
maps a moved or deleted checkout (the finder in this repo's own venv points at a
worktree that no longer exists), or a host that sets PYTHONSAFEPATH so `-m`
ignores cwd. The worker then dies with "No module named 'cron'" before its
ownership ack and every fire records
"cron external worker exited before ownership acknowledgement (exit 1)" (#112729).
The shared subprocess sanitizer strips Hermes-owned PYTHONPATH entries because
user children must not see our tree; this child IS Hermes, so after the env is
built the worker gets an explicit PYTHONPATH: the checkout `cron/scheduler.py`
lives in first, then whatever PYTHONPATH the gateway itself was started with.
The cwd stays the same checkout. Kanban workers spawn `-m hermes_cli.main` and
already get the bootstrap, so this is the only affected entry point.
Live probe (worktree python from cwd=/tmp with PYTHONSAFEPATH=1, real
_launch_external_cron_worker + real Popen): before, the worker stderr tail is
"Error while finding module specification for 'cron.scheduler'
(ModuleNotFoundError: No module named 'cron')"; after, the worker imports
cron.scheduler, loads the payload and reaches the durable-ownership check.
The two review minors on the first pass (worker unlinks its own stderr capture;
comments no longer claim DEVNULL) already landed in f857a4ed99 and are covered
by that PR's tests.