Commit Graph

2964 Commits

Author SHA1 Message Date
teknium1
c55b5bb6ce docs(claude-code): show how to answer a single prompt; opt-in wording in the flag table
- Dialog 2 now shows the read-then-answer step for one benign prompt and
  warns against blind timed Enter, and names --permission-mode acceptEdits
  as the narrower opt-in (idea from PR #113456).
- Quick Reference row and pitfall #2 carry the opt-in wording instead of
  teaching Down+Enter as the expected move.
- Regenerated the claude-code docs page.
2026-09-18 09:27:35 -07:00
Yagna Vudathu
87cd8a3c84 fix(gateway): /stop, /new and /reset keep a parked internal wake instead of discarding it
_interrupt_and_clear_session popped the adapter's single pending slot and dropped whatever it
held ("consume and discard", 59575d6a91). That was right when the slot only ever carried the
user's stale follow-up text, but internal wakes (async-delegation completion notices,
kanban/cron notify+wake) now park in the same slot and are claim-settled the moment the adapter
admits them, so the pop lost them for good: the drain that runs after the command found an
empty slot and the session idled until the next user message (#114456, ~6 min stall in the
reported session).

Now the human follow-up is still discarded, but an internal wake stays parked (promoted out of
the overflow FIFO when a discarded human head occupied the slot) and the post-command drain
starts it immediately. Applies to every caller of the helper — /stop (busy fast path, handler,
pending sentinel, thread sibling), /new and /reset — since a wake that arrives a second after
/new runs against the fresh session anyway; whether a completion pinned to the closed session
may run stays with _resolve_async_delegation_session (fail-closed).

Salvaged from #114538 (@whyyagswhy, earliest filer): the pop-and-re-park mechanism is theirs;
widened here to the /new and /reset callers and the overflow promotion, and the invariant tests
rewritten against the real adapter drain. #114540 (@JoaoMarcos44) reached the same fix
independently; its stop-reason taxonomy and overflow analysis informed the class coverage.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-18 09:27:02 -07:00
teknium1
a3cce7b974 fix(gateway): warn when a live gateway's heartbeat goes stale instead of printing running
The reporter's case in #113372 is "not a crash": the process stays alive,
`gateway_state.json` keeps saying `running`, and housekeeping, cron and the
kanban dispatcher are frozen. Nothing wrote `updated_at` periodically, so the
file was a stored status, not a heartbeat, and both `hermes gateway status`
and `/api/status` rendered a wedged gateway as healthy (the existing stale arm
only fires when the recorded PID is gone).

Make the housekeeping tick re-stamp `gateway_state.json` first thing every
tick (60 s), so `updated_at` is a heartbeat that stops when the thread — or a
chore blocked on the loop — wedges. Readers then warn on `running`/`starting`
+ stale stamp + live PID: `hermes gateway status` prints
`⚠ Gateway heartbeat stale: housekeeping has not refreshed gateway_state.json
for N s (event loop or housekeeping wedged; pid X alive)`, `/api/status`
carries `gateway_heartbeat_stale_s` (null when healthy) and the sidebar strip
shows "Heartbeat stale". Liveness (`gateway_running`, busy/drainable) still
keys off the PID, never the stamp. Draining is excluded: shutdown drains the
housekeeping thread before the process exits, so its stamp legitimately ages.

Invariant tests: the CLI line names the age and the live PID and is silent on
a fresh stamp; /api/status sets/clears `gateway_heartbeat_stale_s` the same way.
2026-09-18 09:25:59 -07:00
teknium1
993d1b9891 fix(status): dashboard keeps a watchdog-exited gateway degraded with its reason
`/api/status` mapped a not-running gateway's retained record through
`retained_gateway_state`, which only kept `startup_failed`; a watchdog-stamped
`degraded` + `exit_reason` of a dead PID became a bare `stopped` with the
reason nulled, so the sidebar strip and System page read "Stopped" while
`hermes gateway status` said "exited degraded: event loop stopped dispatching".

Retain `degraded` under the same rule as `startup_failed` (only while
`desired_state` still wants the gateway running, and only for the watchdog
reasons in the new `gateway.status.WATCHDOG_EXIT_REASONS`); the resolver
already keeps `gateway_exit_reason` for any non-`stopped` verdict. The sidebar
strip gains the `degraded` label (warning while live, destructive when the
process is gone) and the System page describes a dead `degraded` record as a
watchdog exit. `/api/messaging/platforms` keeps yielding `gateway_stopped` for
it — the channels really are down.

Invariant tests: retained_gateway_state keeps/drops the verdict by
desired_state and exit_reason; /api/status carries degraded + exit_reason for
the dead PID and stopped/null after `hermes gateway stop`.
2026-09-18 09:25:59 -07:00
teknium1
8bfc66b252 fix(cron): render watchdog-degraded exits in hermes gateway status; degraded is the last writer
The salvaged commit (#113513) stamps gateway_state.json `degraded` + exit_reason
inside `_mark_exited_quietly`, so both out-of-loop watchdogs (loop liveness, shutdown)
publish before os._exit. This follow-up closes the class:

- Order: lifecycle ledger first, runtime status LAST, immediately before os._exit —
  the runtime record is the one housekeeping refreshes, so nothing may overwrite the
  degraded stamp after it lands (ordering delta from #113386 by @KoNit-K).
- `restart_requested` is asserted only for the supervisor-restart exit code (75); the
  shutdown watchdog's exit code 1 no longer clobbers a recorded restart intent to False.
- `hermes gateway status` (`_runtime_health_lines`) had no `degraded` arm, so the new
  stamp would still have been invisible on the surface operators use; it now renders
  "⚠ Gateway exited degraded: <operator wording for the watchdog reason>". The
  startup-time `degraded` (retryable platforms queued, exit_reason None) stays silent.
- Docs: the built-in loop liveness watchdog and what it writes/exits with.

Why: housekeeping, cron and the embedded kanban dispatcher share one asyncio loop, and
liveness was a stored status field written by that loop — a dead loop left the file at
`running` (#113372). The watchdog thread already detected the stall and exited; it just
never told the file or the status command.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:25:59 -07:00
teknium1
f52bf846cd test(skills): guard paste-safe shell snippets and OS-bound platform gates
- test_shell_snippets_paste_safe parametrizes over every SKILL.md under
  skills/ and optional-skills/ and fails on a backslash followed by a
  comment inside a fenced bash block. A fixed file list cannot catch new
  instances (the notion one was missed by two earlier PRs).
- test_os_bound_skill_is_not_offered_on_the_wrong_platform drives the real
  loader (agent.skill_utils.skill_matches_platform) for heartmula on
  win32 and tensorrt-llm on darwin.
- Regenerated the four touched skill docs pages.
2026-09-18 09:25:26 -07:00
teknium1
13fc2e1044 fix(tools): tool_call rejections restate the single-entry shape and stop misdirecting deferred names
Two rejections from the tool_call bridge fuelled identical-retry loops on
small models because they stated a constraint without a correction:

* A multi-entry batch naming a local tool got "Local tools require one entry
  per tool_call; mixed and multi-local batches are not supported." A 9B model
  re-sent the same two-entry array until identical_call_streak_halt. The
  error now restates the valid shape with the caller's OWN first entry
  ("you sent 2. Retry with only: {"calls":[<first entry>]} then issue the
  remaining 1 call(s) as separate tool_call invocations"), which is the
  hand-written correction that unstuck the reporter's session. The refusal
  itself stays: b4d04eb8fd keeps local entries out of batches deliberately
  (a batch bypasses per-tool admission, path-overlap serialization and the
  per-server MCP parallel opt-in). Shared helper local_batch_error() feeds
  both resolve_underlying_call and the connector dispatcher's guard.

* "'<name>' is not a deferrable tool. If it appears in the model-facing
  tools list already, call it directly" fired for two opposite mistakes.
  For a directly-listed tool the advice is right; for an unknown name —
  typically a deferred MCP tool cited by its bare suffix
  ('mempalace_search' for 'mcp__mempalace__mempalace_search') — "call it
  directly" is the opposite of what the model must do. not_deferrable_error()
  now distinguishes them, suggests the registered full name when one ends
  with '__<name>', and points at tool_search otherwise. tool_describe's
  wrong-door error uses the same helper.

Also ports one parametrize row from #114488 (a stringified single object
without a name flows into the legacy single-shape tolerance) and adds two
invariant tests; docs updated.

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:24:54 -07:00
teknium1
5811653914 fix(api): a run admitted but not yet started settles interrupted at shutdown
The salvaged mark-then-interrupt path covers every run whose agent exists.
A run whose executor task was created but has not taken its first tick has
no agent to interrupt: _execute_run would re-persist `running` over the
shutdown mark and start a full turn that outlives the gateway (only forced
to `interrupted` when it finally returned). Check the shutdown set before
the `running` transition and finish immediately.

Docs: `interrupted` joins the terminal status vocabulary for /v1/runs.

Co-authored-by: Mark Rockminster <markrockminster@M1-U-48-128.local>
2026-09-18 09:24:21 -07:00
teknium1
fdca601381 docs: notebook output truncation names the original notebook path 2026-09-18 09:23:48 -07:00
teknium1
5cee761d54 docs: phonetic guides are excluded from XLSX/DOCX extraction 2026-09-18 09:22:45 -07:00
teknium1
33f3d96b99 fix(disk-cleanup): never rmtree protected top-level dirs; never track kanban/
The tracked-item delete path in quick() called shutil.rmtree() on any tracked
directory without consulting the protection list, which only the empty-dir
sweep used. guess_category() files every path under cache/ as "temp", so a
terminal command that merely mentioned $HERMES_HOME/cache tracked the
directory itself, and 7 days later quick() removed cache/ wholesale — taking
cache/terminal (terminal snapshots) with it and breaking every later command
with a mktemp "No such file or directory".

Fix: _is_protected_dir() — a tracked DIRECTORY that is HERMES_HOME itself or
sits under an _EMPTY_DIR_PROTECTED_TOP_LEVEL tree is never tracked by
guess_category(), never listed by dry_run(), and skipped (logged SKIPPED,
entry dropped) by quick(). Files under cache/ still age out as before.

kanban/ (task attachments and workspaces have their own lifecycle) is added to
both _NEVER_TRACK_TOP_LEVEL and _EMPTY_DIR_PROTECTED_TOP_LEVEL; stale pre-fix
"test" entries under it are dropped by the existing re-validation instead of
deleted. The kanban row was first proposed in #80842 (@nicha16).

Fixes #114552
2026-09-18 09:22:13 -07:00
teknium1
41fe679d96 fix(kanban): stop-guard nudge names every worker exit; sync terminal-tool copies
Follow-up to the two salvaged commits. The nudge text asserted the card was
"still `running`" and offered only kanban_complete / kanban_block, so even
when it fired legitimately it steered a review-bound card toward a false
completion. It now states what the transcript actually shows (no terminal
board call yet) and lists kanban_complete / kanban_request_review /
kanban_block, plus the reviewer exits.

Every remaining copy of the "terminal tools" knowledge is brought in line:
turn_stop_gates docstring + diagnostic status, goals.py finalize comment,
kanban_db_dispatch grace/concurrency comments, the orchestrator-only
refusal in kanban_tools, and the kanban / tutorial / codex-runtime docs.

Fixes #114598
2026-09-18 09:21:42 -07:00
teknium1
b7e6acf468 fix(kanban): report the re-kinded dependency block on every surface, trim tests
Salvage follow-up to the cherry-picked #114529 (@Roblmvp):

- `hermes kanban block --kind dependency` now says the card was blocked as
  needs_input because no parent is open, and the triage verdict keys off the
  landed block_kind (a re-kinded dependency block is a human question).
- kanban_block tool: report the landed kind + requested_kind + a note telling
  the worker why it did not park in todo (dropped the redundant rekind_reason
  echo — it lives in the event payload).
- Tests trimmed to two invariants in test_kanban_block_kinds.py (terminal
  parents -> blocked/needs_input -> triage at BLOCK_RECURRENCE_LIMIT; open
  parent -> todo with no recurrence across a dispatch tick) plus one tool
  surface test.
- Docs: kanban_block row, `blocked`/`dependency_wait` event rows.
- contributors/emails mapping for the salvaged author.

Fixes #114627
2026-09-18 09:21:11 -07:00
teknium1
3e408dcc2a fix(gateway): Telegram topic-binding heal cannot re-pin a route /new moved during its lookups
_hmwa_heal_telegram_topic_binding awaits get_telegram_topic_binding and get_compression_tip between reading the route and calling switch_session, the same shape as the async-delegation re-pin fixed in this branch. Pass the snapshot session id as expected_session_id so the compare-and-swap refuses to overwrite a concurrent /new or /resume. The developer-guide Operations table now documents the CAS kwarg.
2026-09-18 09:20:39 -07:00
teknium1
383f8c0248 docs(messaging): queue mode gives each media follow-up its own turn; photo bursts merge 2026-09-18 09:19:35 -07:00
teknium1
227384e332 fix(auth): drop dead code from the Codex login retry; trim to two invariant tests
Follow-up to the salvaged #114614 pick:

- `_codex_login_post`: the `for` loop ended in an unreachable `raise … # pragma: no cover`
  (needed only to satisfy the return type). A `while True` with the terminal condition
  folded into the except branch has no dead path and the same three-attempt bound.
- `_codex_poll_authorization_code`: the pick duplicated the `except KeyboardInterrupt`
  handler; the second copy was unreachable.
- Tests: six change-detector tests collapsed into two invariants (poll survives blips but
  never retries a non-transport exception; one-shot POST retries once and keeps the typed
  AuthError + TLS hint + cause chain at the cap). The existing
  test_codex_device_login_ssl_hint.py still pins the poll's terminal hint path.
- Docs: providers.md notes that a single dropped connection during device login is no
  longer fatal.
2026-09-18 09:18:00 -07:00
teknium1
9244ec1d77 docs(security): explain how to re-request an approval after a prompt times out
The timeout text tells the agent not to retry, and users read that as a
session-wide lock with no way back (#114204). Document the supported
path: a new user message raises a fresh tool call and a fresh approval
card; timeouts do not count toward the denial breaker.

Fixes #114204
2026-09-18 09:16:25 -07:00
yoniebans
248d20b127 docs(cron): R3-W1+W2 sync diagnostics and doctor exit contract
Document satellite heartbeat output, late and missed-fire findings, and doctor among the last_fire_error surfaces. Retain the R2-2 ruling and explain the watchdog consequence.

Reported-by: kshitijk4poor
2026-09-18 17:10:12 +02:00
yoniebans
23e43a6459 docs(update): explain manual backend reminders and owned recovery 2026-09-18 17:09:28 +02:00
Ritesh Patel
9927199998 docs: drop opencode-free references (provider removed)
Remove the keyless free tier from the provider table (providers.md,
fallback-providers.md), the .env.example section, and the 'three OpenCode
providers' wording now that only Zen and Go ship built-in.
2026-09-18 15:40:37 +05:30
kshitijk4poor
0fb56906fc fix(gateway): a /stop in a thread stops every run of that thread
The thread-sibling tier short-circuited the chat-scope tier, so when a per-user
thread sibling was live the handler interrupted ONLY that run and replied
"Stopped" while a same-thread run under a differently shaped key — the
channel-keyed #286 shape this fallback exists for — kept going. The chat tier is
a superset of the sibling tier (a sibling needs the caller's own thread slot,
which satisfies the chat predicate), so it is the set acted on; the reason still
resolves to `stop_command_thread_sibling` when the chat tier found nothing beyond
siblings, so hook consumers keep that label.

Both tiers also share ONE scan now (they each called `_same_chat_runs`, building
the running-agent snapshot twice), and the matcher is a single text pass:
`_same_chat_key_slots` takes the caller's key prefix instead of re-deriving the
namespace per key, and the whole-slot rule lives once in `_strip_slot` (the tail
check uses it too). Concrete annotations on the three methods (`List[...]` was
missing from the typing import), and the prefix construction reads as a named
namespace rather than a nested f-string.

WhatsApp DMs: `build_session_key` canonicalises the DM chat id, so the fallback
matched nothing there — it now canonicalises the same way, which makes the
chat-scope fallback reach a run keyed from any JID/LID alias.

Docs: `sessions.md` described the tiers sequentially and over-claimed the
widening's reach — an in-thread stop reaches that thread plus a thread-less
room-wide run, never another thread or a peer's per-sender top-level run.
`hooks.md` pointed at `gateway/run.py` for `_interrupt_and_clear_session`, which
lives in `gateway/run_agent_cache.py`.

Guards: the partial-stop case (per-user sibling + same-thread channel run, both
must be interrupted), the lone-sibling reason label, and the WhatsApp canonical
id. Each is mutation-checked individually: restoring the short-circuit, forcing
the reason to chat_scope, or dropping the canonicalisation each fails exactly
its own test. 21 passed in the two /stop test files.
2026-09-18 15:32:17 +05:30
kshitijk4poor
04c0bec905 fix(gateway): chat-scope match must not alias a chat_id or a reply thread
Two defects found by the pre-arm review gate on the matcher above.

- The scope-slot branch ran on every platform, but only Slack ever emits that
  slot (build_session_key appends scope_id for Slack alone). A non-Slack DM keys
  chat_id as the USER id, so a per-sender group key `group:<chat>:<user>` parsed
  as "<user>'s chat": a Telegram user's /stop in her DM interrupted her group
  run in a different chat. The branch is now Slack-only.
- The thread boundary only fired for `thread`-typed keys, so a top-level channel
  turn — which keeps chat_type "channel" with the relay-stamped reply-thread ts —
  stayed reachable from a /stop in a DIFFERENT reply thread of the same channel.
  The boundary now keys on the caller's thread alone: a stop sent from inside a
  thread reaches only runs whose own first trailing slot is that thread, or that
  carry no trailing slot at all (the slotless rolling-DM shape).

Guards: one test per defect (Telegram DM vs group run ending in the same user id;
channel-keyed run belonging to another reply thread), and the workspace/profile
negative now keeps a same-chat run live so a matcher that returns nothing can no
longer satisfy it. Each guard was mutation-checked individually.
2026-09-18 12:18:09 +05:30
kshitijk4poor
b3b2e508b7 test(gateway): pin the /stop chat-scope contract and its bounds
- Negatives strengthened: the isolation tests now keep a same-chat run AND a
  foreign run live and assert exactly which one was interrupted (the previous
  version also passed on base and would have passed for any early return);
  added the chat-id prefix case ("C90") and the workspace-scope + profile case.
- New: a different thread of the same channel is never reached; a pending
  sentinel is never claimed as stopped; a peer's per-sender group run IS
  reached (documents the intentional widening of #113846).
- Assertions now read the real reply (`t("gateway.stop.*")`) and the invalidation
  reason instead of an English substring, so a locale or copy change can't
  silently satisfy them.
- Docs: `sessions.md` no longer claims interrupt handling is strictly per-user;
  `hooks.md` lists the new `stop_command_chat_scope` reason.
2026-09-18 12:18:09 +05:30
teknium1
d177b119e9 docs: setup UX a standalone memory provider keeps; catalog migration note for users 2026-09-17 20:35:22 -07:00
Victor Kyriazakos
5648f81431 fix(slack): reopen a native task card when Slack seals the stream mid-turn
Slack seals a native stream server-side after a few minutes (live-observed
at ~5m20s on three independent long turns, 2026-09-15/16; the lifetime is
not documented). The next chat.appendStream on the card fails with
message_not_in_streaming_state. The adapter returned a bare failure, the
TurnRunner latched native_failed, and the rest of the turn rendered as an
edited text bullet list. Long autonomous turns lost the card UX exactly
when it mattered.

On message_not_in_streaming_state from appendStream, drop the dead
stream_ts and chat.startStream a fresh plan-mode card in the same thread,
then append the current frame there. Every frame already carries the full
visible task projection, so no task state is lost. One reopen per update;
a second rejection surfaces as a real failure. The sealed card is a plain
message now, so no stopStream is sent to it; the turn-final stop targets
the reopened card. Error matching reads SlackApiError.response["error"],
never the message text.

Tests assert the wire sequence (start, append, rejected append, start,
append on the new ts), the reopened frame's task states, the cache pointing
at the new card, and the stop targeting it; plus the one-reopen bound.
Mutation: forcing the expiry branch off turns both tests red.
2026-09-17 18:52:15 -07:00
Victor Kyriazakos
541b20290c fix(bedrock): restore Grok context with provider-confirmed cache provenance 2026-09-17 18:36:10 -07:00
kshitijk4poor
94a029091d refactor(notifications): one notice classification, one config read, one diagnostic-metadata helper
- `is_diagnostic_notice()` replaces three drifting copies: the gateway muted every
  `credits.*` notice, the TUI and CLI only `warn`/`error`, so `credits.restored` was hidden
  on Telegram and shown in the TUI for the same config. Every credit-service notice is an
  automatic diagnostic (a "restored" line after a hidden depletion notice is orphan noise).
- `effective_user_config()` is the single fail-open effective-config read; the two extra
  `deepcopy`s per foreground turn go (the loader already returns a fresh copy and the
  snapshot is read-only).
- `diagnostic_metadata(event)` replaces the repeated
  `{"notification_category": "diagnostic"} if event.internal and ... else {}` literal in
  gateway/run_turn.py; the gateway-side imports of the resolver are module-level (no cycle:
  it imports only gateway.display_config).
- `display.suppress_warning_notifications` is listed with its sibling display keys in
  cli-config.yaml.example and the configuration reference; the messaging guide states that a
  muted diagnostic wake still runs (and bills) its agent turn.
2026-09-18 01:43:35 +05:30
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
teknium1
a566d20d22 docs(discord): drop a stray conflict marker left by the #114356 salvage resolution 2026-09-17 12:08:01 -07:00
teknium1
b675e6dee7 fix(discord): re-arm the bot handoff window per chunk; trim salvage to invariants
Discord paces a bot's sends at roughly one per second, so chunk 3 of a long
handoff lands after the 2s window anchored on the tag and was still dropped —
the symptom the continuation window exists to fix. Each admitted continuation
now re-arms the window; the gateway bot loop guard bounds a bot that never
stops. The flush-delay override moves onto the BasePlatformAdapter seam
(_text_batch_delay_for) that main relocated the batcher to.

Tests trimmed to the invariants: reply-ping-only bot message rejected by
default (and admitted with the explicit opt-out), and a 3-chunk tagged
handoff paced past the original window arrives as one batched event; the
defensive getattr fallbacks that only served test construction are gone.
2026-09-17 12:07:09 -07:00
simpolism
2c5504bf26 docs(discord): clarify deployed bot handoff behavior and verify ingress 2026-09-17 12:07:09 -07:00
simpolism
5d220ed693 fix(discord): batch bot tag continuations 2026-09-17 12:07:09 -07:00
simpolism
f708640561 fix(discord): require hard mentions from bots 2026-09-17 12:07:09 -07:00
Austin Pickett
8ffc2f0369 fix(desktop): working WSLg window controls and native Wayland launch (#113247)
* refactor(desktop): extract titleBarOverlayOptions into a tested helper

getTitleBarOverlayOptions in main.ts branched inline on mac/windows/wsl.
Move the decision into titlebar-overlay-width.ts so the WSLg → false rule
(renderer paints its own controls there) is unit-tested next to the width
reservation it pairs with. main.ts keeps a thin caller.

Co-authored-by: null-runner <nicholas.mariani@hotmail.it>

* fix(desktop): renderer-drawn window controls on WSLg

Under WSLg the frameless window (titleBarStyle: 'hidden') had no
minimize/maximize/close. getTitleBarOverlayOptions returned false on the
premise that the RDP host paints replacement controls; it does not for a
frameless window. Electron's native overlay is not the fix either: its
cluster's hit-region drifts from the rendered buttons under the RAIL
compositor.

The renderer now paints Windows-style min/max/close (wslg-window-controls)
routed over a hermes:window-control IPC channel (window-controls.ts →
preload → main). getWindowState reports customWindowControls/isMaximized so
the renderer knows when to mount them and which glyph to show. The cluster
is pinned to TITLEBAR_HEIGHT px (the contrib shell zeroes --titlebar-height
for content subtrees) with 46px native-width caption buttons.

Co-authored-by: Austin Pickett <pickett.austin@gmail.com>

* fix(desktop): wire WSLg window-control IPC and snap maximize onto the work area

Register hermes:window-control in main, report windowControlState from
getWindowState, and re-send window state on maximize/unmaximize so the
renderer's maximize/restore glyph tracks the window.

maximizedBoundsCorrection snaps a WSLg-maximized frameless window back onto
the display work area (RAIL can settle it offset — reported on WSLg 1.0.65).
It is a no-op wherever native maximize already fills the work area, so
healthy compositors are never fought and setBounds cannot loop.

Co-authored-by: null-runner <nicholas.mariani@hotmail.it>

* fix(desktop): call WSLg window-control bridge without the click event

contextBridge structured-clones every argument, and a React SyntheticEvent
is not cloneable, so onClick={controls.minimize} threw "An object could
not be cloned" before ipcRenderer.send ran. The buttons rendered but did
nothing. Wrap the handlers so the bridge is called with no arguments, and
pin that in the test.

* fix(desktop): remove recursive WSLg maximize bounds correction

* fix(desktop): select native Wayland before Electron initialization on WSLg

* fix(desktop): preserve ready pipe and debugger across WSLg launch

* fix(desktop): keep WSLg caption controls scoped to every window

* style(desktop): order window chrome event import

---------

Co-authored-by: null-runner <nicholas.mariani@hotmail.it>
2026-09-17 14:51:18 -04:00
teknium1
64ea66b03d docs: note the Windows file-lock retry in the desktop stage-and-swap description
The updating guide describes the rename-over-previous-build step; mention that a
scanner briefly holding release/win-unpacked is ridden out with short retries
(#112544) so the paragraph matches the promotion behaviour.
2026-09-17 09:19:19 -07:00
teknium1
df1074b4e5 fix(aux): structured-output rejection no longer kills fallback candidates or costs a doomed first request
Two open atoms of #83390 (DeepSeek "This response_format type is unavailable now"):

* `_call_fallback_candidate_sync/_async` only special-cased auth errors, so when the primary
  aux provider failed (timeout, rate limit, payment) and the fallback landed on a provider that
  rejects `json_schema`, the 400 re-raised and the whole task died — the primary-path rung from
  #89589 never applied there. Both fallback paths now retry once without `response_format`.
* Every structured aux call (titles, kanban decomposer, goal judge, plugin structured calls)
  paid a guaranteed-fail request on providers that lack `json_schema` before the retry. A
  provider profile can now declare `unsupported_response_formats` (DeepSeek: json_schema, per
  https://api-docs.deepseek.com/guides/json_mode) and the recovery ladder remembers any route
  that rejected a type once (host:port scoped), so `_build_call_kwargs` — shared by the primary
  and fallback paths — omits the field before the first request. Dropping rather than
  downgrading to json_object matches the end state the retry already produced; json_object
  needs a JSON-mentioning prompt and some relays return empty content under it.

New logic lives in agent/auxiliary_structured_output.py; the facade only gains the fallback rung
next to the predicate it uses. tests/agent/conftest.py resets the process-level memo per test.

Fixes #83390, #105191. Closes duplicates #84976, #88830, #102849, #113064.
Co-authored-by: Legion-is-life <Legion-is-life@users.noreply.github.com>
2026-09-17 09:11:38 -07:00
teknium1
ae1b5d79b2 fix(sessions): one-shot runs get a distinct oneshot source that pickers hide
`hermes chat -q`/`--oneshot`/`-Q` and `hermes -z` (both set HERMES_SINGLE_QUERY_SESSION=1)
persisted their session as `cli` — and, before the first pass, as the inherited
`tui`/`desktop` transport label — so finite automation runs sat in the TUI, Desktop and
dashboard session pickers next to real conversations (#112550).

- run_agent._session_source_for_agent: a single-query run whose source is empty (or an
  inherited UI transport label without an explicit --source) resolves to `oneshot`; the
  platform gate keeps delegate children (`subagent`) untouched; an explicit `--source`
  (HERMES_SESSION_SOURCE_EXPLICIT=1 from main.py) still wins.
- hermes_state_sessions.INTERNAL_LISTING_SOURCES = (kanban, tool, oneshot) replaces the
  three copied `["kanban", "tool"]` literals (tui_gateway session.list, console
  `sessions list`/`stats`, in-chat /sessions), and the Desktop project tree / sidebar
  recents and the dashboard automation set exclude `oneshot` too.
- `hermes -c` / `--resume latest` still chain on the previous one-shot (PR #105957's
  documented flow): the CLI MRU lookup matches the cli family {cli, oneshot} and
  search_sessions accepts several sources; one-shots keep stamping their launch cwd so the
  workspace-scoped lookup keeps working.
- Compression child: the rotated child is published with the PARENT ROW's persisted source
  instead of bare agent.platform, so a `--source tool` / `oneshot` / inherited `kanban`
  session does not degrade to a picker-visible `cli` row after compaction.
- Docs: sessions source table (+ oneshot/kanban/tool rows, compression note) and the
  `--source` flag reference (explicit flag always stored as given).
2026-09-17 09:06:59 -07:00
teknium1
9790d730c0 fix: a clarify card that resolves AMBIGUOUS after the ack window stays armed for the late tap
The late-failure watch treated every non-"sent" outcome as definitive: a relay
lost-ack (raw_response.ambiguous=True, the card may well have posted) arriving
after SEND_ACK_WINDOW tore down the registration and returned the delivery
notice, so the user's button tap on a card that WAS rendered found no pending
entry and was lost. That breaks the invariant _abort_for_outcome (and main)
keeps for an immediate ambiguous outcome. _on_card_done now returns early on
"ambiguous" and lets the bounded wait's own timeout cover a card that truly
never arrived.

Also while here:
- a late DECLINE releases with UNDELIVERED_DECLINED, matching the immediate
  path, instead of the generic notice (_release carries the outcome);
- when the text fallback cannot even be scheduled (fallback() -> None) the
  card's "failed" verdict stands instead of re-classifying None, which logged a
  misleading "no scheduling future (loop unavailable)";
- test_card_failing_after_the_ack_window_... now asserts the text prompt was
  sent exactly once and the wait released within ack window + 2 s, and the
  helper answers only after observing the prompt — with fallback/late-watch
  disabled it stayed green before (the helper answered on a 10 s deadline);
- telegram.md: clarify_timeout default is 3600 s (resolve_clarify_timeout,
  configuration.md), not 600.

Part of #112684
2026-09-17 09:06:17 -07:00
teknium1
98b4efbf48 fix: clarify cards that cannot render are re-asked as plain text, never reported as user inactivity
A native clarify card on a messaging platform (Telegram inline keyboard, Slack blocks) could
fail without the user ever seeing a question, and the agent then waited out the full
clarify_timeout and reported "[user did not respond within Nm]" (#112684):

- the platform rejects the card at once -> the wait aborted with a delivery sentinel but the
  question was never re-asked;
- the send outruns the 15 s acknowledgement window and only then fails (pool/connect timeouts,
  a stale-thread retry) -> the runner never looked at the send future again and blocked for
  the whole timeout for a card that never posted;
- no status adapter at all -> _ask_clarify_question returned ('', False), so a batch reported
  timed_out=True with an empty notice.

gateway/run_turn_runner_clarify_delivery.py (new topical sibling; the send-disposition helpers
move out of the gateway/run.py facade) now retries every definitive card failure once through the
adapter's plain-text send_clarify (numbered list + text capture) - except a connector egress
DECLINE, where re-sending the text is the exfiltration the guard exists to stop - and watches a
possibly-delivered send so a late failure releases the waiter with
"[clarify prompt could not be delivered]". The no-surface case reports
"[clarify prompt could not be delivered: no chat surface]" (same prefix every consumer already
treats as a non-answer).

Live probe (real TurnRunner + real clarify_gateway, Telegram-shaped adapter, timeout 30 s):
card fails after 16 s -> before 45.2 s and "[user did not respond within 0m]", after 16.7 s and
the typed answer to the text prompt; card rejected at once -> before sentinel with no text prompt,
after the numbered prompt is sent and "2" resolves to the choice.
2026-09-17 09:06:17 -07:00
teknium1
091b8c4f53 docs: call out the client.capabilities fail-closed gate as breaking for third-party WS clients
Why: an integration that answered server→client requests but never sent
client.capabilities now has every clarify/approval/sudo/… request refused at
once with no grace path. Marking a connection as answering after its first
response frame cannot help — the frame is never sent to it — so the break is
documented instead.
2026-09-17 09:04:38 -07:00
teknium1
f9d178f78e fix(tui_gateway): old app builds no longer stall the agent on clarify/approval; late approval choices count
Item 2 of #112548: a Desktop/dashboard build that predates server→client
requests has no response path, so every clarify/approval/sudo/secret/vault/
connection/bridge request sat for the full deadline (clarify: 300s). Only the
tour probed. Clients now advertise once per connection
(`client.capabilities {server_requests: true}`, sent by the shared TypeScript
channel on `gateway.ready`); `send()` / `send_async()` return the
error-response shape (None) at once when every WebSocket peer of the session
is a build that never advertised. Sessions with no client attached still wait
so the reconnect replay (`open_requests`) keeps working; the stdio TUI ships
with the backend and is not gated. The advertisement is dropped on disconnect.

Reviewer minors from #113227:
- tools/approval_gateway_wait.py: the verdict is the choice committed under
  the approval lock while leaving the queue, so an /approve that lands after
  the deadline check but before the entry is dropped is an answer, not a
  timeout (the client was already acked "ok").
- tests/tui_gateway/test_protocol.py: the error-fails-fast test that only
  restated pre-existing behaviour is replaced by the two capability
  invariants (never advertised → fails fast; advertised → frame written,
  waits, forgotten on disconnect).
- server_requests.send try/finally around event.wait already landed on main
  (4371ed34a9); nothing to change.

Docs: programmatic-integration.md (advertise once per connection; method
list), tui_gateway/AGENTS.md; contracts regenerated.
2026-09-17 09:04:38 -07:00
teknium1
f3bdcd0877 fix: dashboard stop sweeps the wedged descendants that outlive the SIGKILLed backend
`hermes dashboard --stop` / `hermes update` signal only the backend PIDs. When the
lifespan teardown wedges (stop_hosted_room_service, PTY close_all never reached) the
10s grace loses, SIGKILL lands mid-teardown, and the hosted ui-tui / tui_gateway.entry
child is reparented to init still holding the state.db-wal inode; the next start
refuses with DeletedWalGenerationError (#112631, residual of #111912). No finite
root grace covers an unbounded teardown.

_kill_pids_posix now snapshots the dashboard-owned descendant tree BEFORE the kill
(the PPID link is gone once the root dies), and after the root phase SIGTERM→SIGKILLs
the descendants that are still alive, waiting for the tree to be gone before
returning. Descendants are re-checked against their snapshotted start-time
fingerprint (the same PID-reuse guard _kill_pids_windows uses) — no `ps -o lstart`
per system PID.

Detached session leaders without a controlling terminal are pruned from the sweep
with their subtrees: those are the messaging-gateway bots and profile actions the
dashboard launched with start_new_session from /api/gateway/*, which belong to the
user, not the dashboard. A hosted TUI is a session leader too (pty.fork) but owns
the pts whose master the dashboard held, so its tty column is set and it is swept.

The Desktop boot reaper (_reap_orphaned_desktop_local_serves) SIGKILLs the same
class of backend after 1.5s and had the same hole; it now SIGKILLs the surviving
snapshotted descendants (no second grace: the boot path runs under a 10s probe).

Live: pyteman hermes-111912 wedged leg on origin/main
`child_orphan_alive=True deleted_sidecar_holders=2 guard=FATAL DeletedWalGenerationError`,
on this head `child_orphan_alive=False deleted_sidecar_holders=0 guard=clean`; the
start_new_session control sibling survives on both.
2026-09-17 09:04:11 -07:00
teknium1
0f98d5a1d9 fix(plugins): execution-chain middleware failures are reported once, not per call
The warn-once reporter added for hooks (and for PluginManager.invoke_middleware)
left the execution chain out: hermes_cli/middleware.py::_run_execution_chain — the
tool_execution / llm_execution frames that run once per tool or LLM call — still
logged "Middleware '%s' callback %s raised" at WARNING on every invocation. A
mis-declared callback (e.g. a signature naming tool_data) therefore flooded the
log exactly like the hook case #111922 reported: 5 calls -> 5 WARNING lines.

Route the frame's except through the manager's _report_hook_failure with the
"Middleware" surface label, so the first failure warns (listing the fields the
middleware does provide) and identical repeats go to DEBUG; the set is already
forgotten on plugin reload. The frame's skip-and-continue semantics are unchanged.

Part of #111922
2026-09-17 09:03:29 -07:00
teknium1
ed7eb5685a docs(vision): document vision.embed_target_bytes and vision.max_calls_per_image
User-visible knobs need a home: the Vision feature page explains why native embeds ride
the session and what each key does (subagent-only default cap, clamp range), the
configuration reference points at it next to auxiliary.vision so the two sections are
not confused, cli-config.yaml.example carries the commented block, and tools/AGENTS.md
names vision_tools_history_budget.py as the single owner of embed-cost policy.
2026-09-17 09:02:55 -07:00
teknium1
d6add16059 fix(auth): dead OAuth logins are reported once and leave rotation; hints point at hermes auth
A terminally rejected refresh token (invalid_grant / invalid_token /
refresh_token_reused) is the moment a login is lost. Three gaps remained
after #113197 (which added the WARNING for Codex/xAI/Nous):

- Anthropic: the token endpoint's HTTPError carried no classifiable code,
  so a dead Anthropic / Claude Code grant fell through to a transient
  'exhausted' bench at DEBUG and was replayed every hour. The endpoint
  error is now a structured AnthropicOAuthError (status + OAuth error
  code); _recover_failed_refresh logs one WARNING with the repair command
  and marks the row DEAD. A dead grant is not replayed at the fallback
  endpoint. The auxiliary Claude Code refresher (sibling path) warns the
  same way. Claude Code's own credentials file is never touched.
- Codex/xAI/Nous: _quarantine_sources drops only singleton-seeded rows, so
  an independent `hermes auth add` (manual:*) login survived unmarked and
  re-fired the WARNING on every later refresh attempt. The survivor is now
  marked DEAD (leaves rotation until a write-side re-auth clears it).
- Hints on live Hermes paths (anthropic 401 troubleshooting block, the
  no-credentials error) recommended an external CLI's login command,
  which cannot repair Hermes' own login; they now say
  `hermes auth add anthropic` / `hermes auth list anthropic`.

Docs: credential-pools.md documents the dead-login behaviour.
Part of #113023.
2026-09-17 09:02:26 -07:00
teknium1
07c92d675a fix(compression): repeated summary stall escalates to the deterministic fallback summary
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).

WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.

Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.

Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
2026-09-17 08:59:50 -07:00
teknium1
931b5ff9e7 fix(opencode): every credential-resolution surface keys off the model it will send; family heal only for built-in providers
Closes the remaining atoms of #112600.

A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
   its siblings still resolved credentials against config's `default`: the CLI
   auth-fallback rung, `--resume` credential re-resolution, the gateway
   provider-override helper (channel overrides, persisted /model switches,
   API-server provider refresh), the gateway fallback chain, the TUI /model
   switch-from runtime and ACP agent construction. With a `*-free` default the
   OpenCode free-tier rung fired first and a Go-only model was built against
   the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
   the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
   optional `target_model` and the two test stubs of it accept the kwarg.

B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
   provider matched by opencode_provider_family, including custom providers
   merely named after a family (`opencode-go-bridge`, #85589) whose relay the
   user declared explicitly in `providers:`. The family heal now applies to the
   built-in canonical providers only; custom prefix-named providers keep their
   per-model api_mode routing and /v1 handling. Documented in the providers
   guide.

C) Same function: the official-host check uses parsed.hostname (a port no
   longer defeats the heal) and only the path is edited, so query/fragment
   round-trip instead of being dropped.

Fixes #112600
2026-09-17 08:59:22 -07:00
teknium1
200f8a6b4e fix(aux): keep the /v1 tail on the config path and trim the salvage to two invariants
Follow-up to the cherry-picked #106018 (@liuhao1024). The alias-table fix resolved the
keyless lane, but the reporter's second variant (provider: ollama + bare base_url +
explicit api_key → 404) still failed through the real config path:
_resolve_task_provider_model collapses "base_url + api_key" lanes to "custom" before
resolve_provider_client runs, so the custom branch never saw the ollama alias and
skipped the /v1 tail. Keep the local-server alias identity on both base_url paths
(config lane and explicit kwargs via _preserve_provider_with_base_url) so the tail
applies whenever the user wrote provider: ollama/vllm/llamacpp.

Shape: inline the 5-line _bare_host_base_url helper (urlparse never raises on str;
no new facade helper), trim the five contributor tests to two invariants that drive
the reporter's lane through _resolve_task_provider_model, and document the alias
group under the auxiliary provider list.

Live: stand-in mirroring Ollama's /v1 OpenAI-compatible surface — before: RuntimeError
"no API key was found" (keyless) / 404 on /chat/completions (explicit key); after: both
POST /v1/chat/completions and return OK, sync and async; provider: custom + /v1 unchanged.
2026-09-17 08:56:27 -07:00
teknium1
3187d68b49 fix(kanban): goal_mode workers keep the live tool feed in their worker log
The dispatcher forced `-Q` on goal_mode cards because cli.py only ran the
kanban judge loop in the fully-quiet one-shot branch. `-Q` strips every tool
callback, so the dashboard's Worker log (and `hermes kanban log`) stayed
blank for the whole run while non-goal cards logged normally; users read
that as "the worker is doing nothing" (Discord report, Sep 2026).

Run the judge loop on the `-q` path too, driving follow-up turns through
cli.chat so each turn's tool activity lands on stdout (= the worker log),
and print each judge verdict there as well. The dispatcher spawns goal_mode
and one-shot workers with the identical argv; the mode travels only in
HERMES_KANBAN_GOAL_MODE. The `-Q` hook stays for manual quiet runs.

Live A/B (real `hermes kanban dispatch` + spawned worker against a scripted
loopback provider that answers a terminal tool call then text, judge always
"continue", --goal-max-turns 2): base log = 94 bytes, 0 tool-feed lines,
0 verdict lines; fix log = 1910 bytes, 4 tool-feed lines, 3 verdict lines;
both arms end blocked "exhausted 2/2 turns" (loop behaviour unchanged).
2026-09-17 08:55:01 -07:00
teknium1
9269b19e4f fix(cron): restart-safe worker imports the gateway's own checkout instead of relying on cwd
The external cron worker is spawned as `sys.executable -m cron.scheduler`. Its
entry module is `cron.scheduler`, not `hermes_cli.main`, so it never runs the
bootstrap that puts the checkout on the gateway's sys.path; it only imported
`cron` at all through the implicit `-m` cwd entry (cwd was already the repo
root). That implicit path breaks on real hosts: a venv whose editable install
maps a moved or deleted checkout (the finder in this repo's own venv points at a
worktree that no longer exists), or a host that sets PYTHONSAFEPATH so `-m`
ignores cwd. The worker then dies with "No module named 'cron'" before its
ownership ack and every fire records
"cron external worker exited before ownership acknowledgement (exit 1)" (#112729).

The shared subprocess sanitizer strips Hermes-owned PYTHONPATH entries because
user children must not see our tree; this child IS Hermes, so after the env is
built the worker gets an explicit PYTHONPATH: the checkout `cron/scheduler.py`
lives in first, then whatever PYTHONPATH the gateway itself was started with.
The cwd stays the same checkout. Kanban workers spawn `-m hermes_cli.main` and
already get the bootstrap, so this is the only affected entry point.

Live probe (worktree python from cwd=/tmp with PYTHONSAFEPATH=1, real
_launch_external_cron_worker + real Popen): before, the worker stderr tail is
"Error while finding module specification for 'cron.scheduler'
(ModuleNotFoundError: No module named 'cron')"; after, the worker imports
cron.scheduler, loads the payload and reaches the durable-ownership check.

The two review minors on the first pass (worker unlinks its own stderr capture;
comments no longer claim DEVNULL) already landed in f857a4ed99 and are covered
by that PR's tests.
2026-09-17 08:54:39 -07:00