1334 Commits

Author SHA1 Message Date
ethernet
df1b647b42 fix(security): don't warn while tirith's first download runs
The first CLI launch after a PM install starts tirith's download in the
background, so ensure_installed() returns None and the CLI printed
"tirith security scanner enabled but not available". The scanner was on
its way, not missing; the same warning also fired when lazy installs are
disabled by the operator's own policy.

missing_is_expected() tells those by-design states apart. The CLI logs
them and keeps the visible warning for a missing explicit tirith_path or
a finished download that failed.
2026-09-24 14:23:47 -04:00
ethernet
cb7b18431a Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/src/app/settings/connections-registry.tsx
#	scripts/install.ps1
#	scripts/install.sh
#	tests/hermes_cli/test_update_autostash.py
2026-09-24 03:48:31 -04:00
teknium1
0004e3a8df refactor(cli): move the output-history row helpers out of the cli facade (#95375)
_output_history_lines / _output_history_rows are new in this PR; they belong with the
rest of the history rendering in hermes_cli/cli_render.py, not appended to cli.py.
cli.py imports them for _replay_output_history; the terminal mixin imports them from
cli_render directly.
2026-09-24 00:17:13 -07:00
teknium1
e389aa4061 fix(cli): close resize-drag row loss and modal refill dups (#95375)
- Gate every paint and chrome render on the live width: a paint or render that
  would land after the width changed but before its recovery now waits for it
  (checked when it actually runs on the loop, not when it was scheduled).
- Recovery that sees the width move again keeps holding; the next recovery runs
  right after the next width change once the 0.85s hold cap passed.
- Track the narrowest width of a drag: narrowing and widening back still refills.
- Chrome floor: after a refill the chrome is drawn down to the bottom row (what
  CPR would tell prompt_toolkit), so the next count is exact when it shrank.
- Rows output scrolled while the width changed under it are repainted by the
  next refill (twice rather than truncated for good).
- A resize whose viewport holds the whole history erases from its oldest row,
  keeping the unrecorded startup banner.
- _terminal_reflows: TERM decides before inherited env (st/urxvt from tmux,
  kitty or vscode); rxvt added; vte-256color no longer matches 'vt'.
2026-09-24 00:17:13 -07:00
teknium1
68b0f85d46 fix(cli): resize keeps each transcript row once on tmux and never loses rows elsewhere (#95375)
- Resize no longer erases the viewport and replays it: prompt_toolkit's erase is aimed
  at the chrome's top as the terminal re-wrapped it, the transcript left in place.
- Unknown terminals (TERM=xterm-256color over ssh): the rows the aim would take on a
  non-reflowing terminal are replayed, so an erase never takes more than it puts back.
  st/screen/mosh/vt/cons TERMs count as non-reflowing; LC_TERMINAL=iTerm2 as reflowing.
- Debounce 0.3 s: tmux re-wraps a pane at once but signals it at most every 250 ms.
- No app.invalidate() paints while a resize recovery is pending.
- Ctrl+L/focus refill paints synchronously (tmux copies the screen on ED from home) and
  erases every viewport row (CUP rows are 1-based).
- Bare print() in the chat/stream mixins goes through _cprint; held paints are released
  at app exit; narrower excepts.
- tmux-backed E2E cell pins all of it.
2026-09-24 00:17:13 -07:00
teknium1
38f97640c0 fix(cli): size the resize replay for how the terminal re-wraps
Live tmux probes showed the painted-width budget from the previous
commit duplicating rows on reflowing terminals: a shrink re-wraps the
rows the replay counted at their old width (Rich pads lines to the full
width, so every reply line grows) and pushes the extra into scrollback.
Terminals that keep rows in place (xterm) need the opposite count, and
the app cannot see which kind it runs in.

- A shrink replays only on a reflowing terminal (_terminal_reflows: TMUX
  yes; STY, XTERM_VERSION, TERM=linux no; unknown counts as reflowing),
  counting rows at the new width and the old chrome as re-wrapped to it.
  Elsewhere the transcript stays as the terminal truncated it and
  prompt_toolkit's own erase covers the chrome: nothing is lost.
- Ctrl+L counts rows at their painted width only where rows stay put.
- _cprint output is held from the resize signal until its recovery ran:
  a paint in between erased the re-wrapped chrome from prompt_toolkit's
  stale cursor, leaving two rows of it in the transcript that the
  replay could not account for (two duplicated lines per shrink).
2026-09-24 00:17:13 -07:00
teknium1
d4865368bc fix(cli): a resize mid-stream repaints each transcript line once
The viewport-sized replay (#95375) still duplicated rows when the width
changed while a reply streamed: the review's e2e cell showed the warmup
reply twice after two widens.

- A widen leaves the transcript alone. It wraps nothing into extra rows,
  so prompt_toolkit's own stale-cursor erase still covers the chrome, and
  no replay size is right for both terminals that keep rows in place
  (xterm) and reflowing ones. The 3J rebuild opt-in still rebuilds.
- Output is recorded in _OUTPUT_HISTORY when it is painted, not when a
  worker requests it: a replay no longer prints lines still queued for
  the app loop (they printed again right after).
- Each recorded line carries the terminal width it was painted at, and
  the replay budget counts it at that width: a line painted after the
  terminal shrank but before the debounced handler ran wraps into more
  rows than the old width says, which pushed the budget into scrollback.
- The room above the chrome is the renderer's last paint (it can be
  taller than the preferred height), not a re-fit of the future chrome.
- A line taller than the room keeps its bottom rows (_ansi_drop_cells)
  instead of vanishing.
- The two blank separator lines of every turn go through _cprint, so
  the replay budget sees them (print() through patch_stdout bypassed it).

Tests: widen skips the replay; worker prints enter the history when
painted, tagged with the width; the tail fit counts painted widths and
keeps a tall line's bottom; the width-change test asserts the fit.
2026-09-24 00:17:13 -07:00
teknium1
f748481247 fix(cli): refill only the viewport on redraw and erase it without CSI 2J
The bounded replay still stacked a copy per resize/Ctrl+L: counting logical
lines against the terminal height ignores soft-wrapped rows and the prompt
chrome, and CSI 2J makes scroll-on-clear terminals (tmux, VTE) copy the
whole visible screen into scrollback before the replay repaints it.

- _clear_prompt_toolkit_screen erases the viewport row by row (EL) when
  scrollback is kept and returns the room above the chrome: terminal rows
  minus the layout's preferred height, at the current width. The 3J
  rebuild path (display.cli_rebuild_scrollback_on_redraw) keeps 2J+3J and
  replays the whole history, which is what rebuilds scrollback.
- _replay_output_history(fit) keeps the newest lines whose wrapped height
  fits that room (_output_tail_fitting in cli_render).
- the resize path budgets against the full chrome: the status bar/rules
  hidden while the reflow settles come back and pushed the top replayed
  rows into scrollback.

Tests: the two tests that pinned CSI 2J on the keep-scrollback path now
assert the row erase; replay stubs take the fit argument; the salvaged
visible-height tests collapse into one wrapped-height invariant.

Fixes #95375
2026-09-24 00:17:13 -07:00
Jackal991
e7bcdcc0c9 fix(cli): bound persistent_output replay to visible terminal height
Replay re-printed the full 200-line _OUTPUT_HISTORY buffer on every
resize/redraw, but the pre-replay clear only wipes the visible screen
(CSI 2J), so each replay appended a duplicate copy of the history to
scrollback. Bound the replay to the visible row count so a redraw
restores the visible transcript without stacking duplicate blocks.

Refs #95375
2026-09-24 00:17:13 -07:00
ethernet
44f99caacd fix(entrypoints): tolerate only an absent bootstrap, not a broken one
Every entry point wrapped `import hermes_bootstrap` in
`except ModuleNotFoundError: pass` for a partial update that left the
bootstrap unregistered. It also swallowed a module the bootstrap itself
failed to import, and since the bootstrap now owns PM activation that
silently ran the tree on stale dependencies: exactly how a pre-PM
editable venv hid its unreachable `pm` until it crashed on ruamel.

Re-raise unless the missing module is hermes_bootstrap, at all six entry
points. The stale "only Windows UTF-8 stdio suffers" comments go with it.
2026-09-23 14:15:29 -04:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
ethernet
3b181dd848 merge origin/main into ethie/pm-clean 2026-09-21 00:38:40 -04:00
teknium1
2ca53fc386 refactor(cli): move cli.py module-level helper clusters into topical siblings; cli.py lands under 2,000 lines (#116911)
Mechanical, behaviour-neutral extraction along the existing hermes_cli/cli_*.py
pattern. Six clusters of module-level helpers leave cli.py:

- cli_config_load.py     prefill messages, reasoning/service-tier parsing, terminal
                         env mirroring, CLI defaults + YAML merge, logging bootstrap
- cli_render.py          reasoning-tag stripping, ANSI/skin colours, light-mode
                         detection, markdown/final rendering, output-history
                         recording, _cprint, ChatConsole, compact banner, panel wrap
- cli_terminal_input.py  file drops/attachments, bracketed-paste patch, extended
                         Enter keys, CPR guards, TUI input height, query images
- cli_shutdown.py        process session-id sync, exit watchdog, cleanup steps,
                         session-finalize notifications, one-shot finalize
- cli_single_query.py    kanban goal loops, exit-code mapping, quiet -q runner,
                         image routing, signal handlers, single-query mode
- cli_auto_maintenance.py state-db / checkpoint startup maintenance

Every moved body is AST-identical to the base copy modulo two mechanical
edits: cli-level names are late-bound with a call-time `from cli import ...`
(the cli_init_mixin pattern) and mutable cli state is read through `_cli().NAME`
(the gateway_service_unit `_gw()` pattern), so every monkeypatch seam on the
`cli` facade still intercepts moved code and no read snapshots stale state.
Mutable module state and every `global`-writing function stay in cli.py; the
facade re-exports the moved names in one import block.

test_bracketed_paste_timeout AST-loads the paste helper from its new home.
cli.py 3959 -> 1820 lines.
2026-09-20 20:42:17 -07:00
teknium1
aaafa80577 refactor(cli): move HermesCLI init phases and TUI run-loop phases into mixin siblings (#116911)
Mechanical, behaviour-neutral extraction along the existing cli_*_mixin.py
pattern: the ten _init_* constructor phases become CLIInitMixin
(hermes_cli/cli_init_mixin.py) and the fourteen _tui_* run-loop phases
(input dispatch, after-turn, startup banner/prewarm/maintenance, application
build, signal handlers, shutdown) become CLITuiRuntimeMixin
(hermes_cli/cli_tui_runtime_mixin.py). __init__ and run() stay in cli.py as
the orchestrators. Every moved body is AST-identical to the base copy; cli.py
module names are resolved lazily through cli so monkeypatch seams survive.
cli.py 4835 -> 3960 lines.
2026-09-20 18:43:11 -07:00
teknium1
265e68d784 fix(auth): one primary_failure_wording() helper labels quota vs auth failure at all three fallback surfaces
- hermes_cli.auth.primary_failure_wording(exc) -> (log, user) phrase; reused by
  cli_agent_setup_mixin._resolve_fallback_runtime, runtime_provider's fallback
  logger and the TUI gateway/Desktop _resolve_runtime_with_fallback (#117482
  sibling: 'Primary auth failed' for a 429 on the gateway surface).
- Drop the dead credentials_rate_limited kwarg at the post-turn exit site (the
  flag is only True when _ensure_runtime_credentials returned False).
- Fold the new tests: 2 parametrized production-path tests + 1 gateway test.
2026-09-20 16:23:22 -07:00
686f6c61
89096ba56c fix(cli): treat credential-resolution 429 as quota, not auth
A Codex quota wall at startup was labeled "Primary auth failed" and
exited 1, so Kanban counted a failure and operators went looking for
a bad token. Use is_rate_limited_auth_error() for the fallback notice
and EX_TEMPFAIL (75) when HERMES_KANBAN_TASK is set.
2026-09-20 16:23:22 -07:00
John Paul Soliva
76a486a1eb fix(bot-relay): a relayed turn is booked from its turn report at the cap, not killed with its handoff
The relay's delivery child shared one 600s deadline between the target's turn
and the one-shot exit linger, whose own budget is the same 600s — so a turn that
answered in seconds and then handed off to a teammate was killed mid-linger,
reported to the sender as delivery_timeout (auto-retried: the turn ran twice),
its reply lost, and the handoff delivery the linger protected destroyed.

The -Q child's turn report (#113608) now carries the answer the run will print,
rewritten when a follow-up turn displaces it, and the poll loop that books a
child from that report moves next to the contract as
quiet_single_query.run_reported_turn. The cron lane keeps its policy (book after
a 2s exit grace); the relay waits for exit under the cap as before — a teammate's
reply during the linger may still become the printed answer — and at the cap
books a reported child from its latest report and leaves it to finish. Only a
turn that never ends is a timeout. The report is 0600 from creation now that it
carries the answer.

Fixes #114980
2026-09-20 11:55:07 -07:00
teknium1
0469740ab3 feat(cron): jobs follow the main agent model at fire time; pinned locks it on request
An unpinned cron job used to snapshot the global provider/model at creation and treat that
snapshot as its effective pin (#44585), so `hermes model` / `/model` never moved the fleet and
`hermes cron resnap` existed to catch jobs up. New ruling: jobs run on whatever the main agent
model is when they fire. Resolution is per-job pin > cron.model / cron.model_provider (the cron
fleet default) > model.default.

`pinned` replaces the implicit snapshot with an explicit lock: create/update with pinned=true
writes the CURRENT main provider+model onto the job as an ordinary per-job pin; pinned=false
releases both. The cronjob tool exposes it (schema: only when the user asks; it can only lock
the main model, never point spend at a different one) and reports `pinned` per job; the CLI
gets `--pin` / `--unpin`. Legacy records that still carry *_snapshot keys follow the main model.

Removed with the snapshot: `hermes cron resnap`, the tool's resnap action + `all` param, the
"N unpinned jobs keep running on ..." notice in `hermes model` / `hermes config set` / the
dashboard model assignment, and the Desktop cron-model-impact card (setMainModelAssignment
keeps the expensive-model confirm flow in store/model-assignment.ts).

Live A/B (real store + run_job against a temp HERMES_HOME): main-model X -> Y, unpinned job
fires on X before, Y after; pinned job stays on X; unpin -> Y; legacy snapshot record -> Y.
2026-09-20 09:20:51 -07:00
ethernet
e1576d06a6 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
2026-09-19 22:57:07 -04:00
ethernet
d29da5fe20 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the branch: PM owns dependency preparation, the
Windows shim re-exec/hand-off path stays retired (main's shim-parent wait,
gateway-resume env token and update_cmd_deps tests dropped), docs describe
the PM update flow. The docker workflow parks install-stamp.json around the
toolchain step instead of deleting it so tests/docker can compare provenance.
2026-09-19 13:45:07 -04:00
teknium1
64c703ee95 fix(cli): a WAF-blocked Kanban worker exits KANBAN_TERMINAL_PROVIDER_EXIT_CODE, not 1
upstream_blocked was in neither provider-reason set, so a worker hitting a User-Agent
firewall exited 1 and the dispatcher re-spawned it into the same wall until
kanban.failure_limit was spent (on base the same 403 classified auth and parked the card
after one spawn). A header change heals it, a retry never does: it is terminal.
2026-09-19 09:57:21 -07:00
teknium1
2a0a6c5c08 fix(curator): keep the CLI prompt off the curator's critical path; stop snapshotting the ledger and .archive
A due weekly curator pass held the classic CLI between the banner and the
input box for 5m54s (355s recorded in .curator_state). Two causes:

- cli.py::_tui_startup_background_maintenance ran the deterministic curator
  pass and the skill-sync pulls synchronously on the main thread, despite its
  name; _tui_build_layout / app.run() waited on them. It now spawns a daemon
  thread (the curator's LLM pass already ran on one).

- Every skill archive records a "complete package" ledger capture, which
  gunzips the newest curator snapshot in full to fill missing files. The
  snapshot rolled in .curator_ledger.jsonl (652 MB) and .archive/ (95 MB) on
  top of ~40 MB of live skills, so each of 57 archives inflated 820 MB to
  recover two files (~9 s apiece under load). Neither belongs in a snapshot:
  the ledger is the append-only audit log and .archive/ the recoverable store,
  and rolling either back to an older copy loses entries/skills. Both join
  _EXCLUDE_TOP_LEVEL, which snapshot_skills and rollback already share.

Live A/B on a copy of a 1.9 GB skills tree with the same 46 stale skills
restored: see PR body.
2026-09-19 03:15:10 -07:00
teknium1
d6f1de3f74 fix(lsp): release a worktree's language servers on removal; reap deleted roots
`LSPService` kept one client per `(server_id, workspace_root)` for the life
of the process.  In a long-running gateway that outlives its coding sessions
the client for a removed worktree stayed registered with its stdio pipes
held open — tsserver heaps of several GiB pointed at trees that no longer
existed (#102345).  The idle reaper (d7578018c5) does not cover this: a
client whose root vanished is not idle from the server's point of view.

- `LSPService.release_workspace(path)`: detaches every client whose folders
  live under `path` under `_state_lock`, waits for in-flight spawns so a
  concurrent `_get_or_spawn` cannot reinsert a client after release, prunes
  the delta baselines and broken-set entries beneath the path, and shuts the
  clients down on the service loop.  Multi-root servers (pyright) only drop
  the folder (`LSPClient.remove_workspace_folder`) so siblings keep their
  shared process.  Idempotent; best-effort; returns the count.
- The reaper sweep now also uses that primitive for clients whose every
  workspace folder no longer exists (externally deleted roots).
- `agent.lsp.release_workspace()` reaches every started service without
  creating one; `hermes_cli.worktree_ops.release_lsp_clients()` is called
  from both worktree-removal paths (`cli._cleanup_worktree`, kanban
  `_cleanup_worktree_workspace`) BEFORE `git worktree remove`.

Salvages the direction of #102381 (@Sahilvishnaliya, release_workspace +
cli hook) and the deleted-root eligibility of #95047 (@israellot); both
matched on `key[1]`, which is `""` for multi-root servers and would have
reaped pyright on every sweep — the primitive here matches on
`client.workspace_folders` instead.

Co-authored-by: Sahil Vishnalya <222165401+Sahilvishnaliya@users.noreply.github.com>
Co-authored-by: Israel Lot <840042+israellot@users.noreply.github.com>
2026-09-19 01:29:03 -07:00
ethernet
873e6e21df Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-19 01:43:50 -04:00
ethernet
d70feca03d Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/update_cmd.py
#	tests/hermes_cli/test_cmd_update.py
2026-09-19 01:14:17 -04:00
kshitijk4poor
cf13448580 refactor(agent): name the tool-call XML namespace prefix once
Gate fold: the optional-prefix regex was spelled three times in the agent
stripper; it is one constant now (_NS_PREFIX). The cli.py docstring pointed
at run_agent._strip_think_blocks, which moved to
agent.agent_runtime_helpers.strip_think_blocks; stray blank lines dropped
from the test file.
2026-09-19 10:41:08 +05:30
DoGMaTiiC
60735cbf81 fix(agent): strip namespace-prefixed text-channel tool-call XML from visible text
muse-spark (opencode-go, Responses wire) occasionally serializes its next native
tool call onto the text channel as <atem:function_calls>…</atem:function_calls>
instead of a function_call item. With real tool work earlier in the turn and
finish_reason=stop, the literal XML is then delivered as the turn's final answer
(observed on a production cron worker: the turn ended on the full XML block; the
task itself had executed).

Teach the tool-call XML patterns an optional namespace prefix (?:[\w.-]+:)? on
the open/close tags and the unterminated-opener tail, in every copy of the
stripper: agent/agent_runtime_helpers.py (_TOOL_CALL_BLOCK_PATTERNS,
_STRAY_TOOL_CALL_CLOSER_PATTERN, _UNTERMINATED_TOOL_CALL_PATTERN) and the
display-side cli.py::_strip_reasoning_tags, which the test suite requires to stay
in sync. Covered positions: closed block, stray closer, unterminated opener
(stream cut mid-serialization).

With the block stripped to nothing, the existing empty-response recovery path
(agent/turn_final_response.py) re-prompts the model instead of shipping XML as
the answer.

Refs #103483 (the native-call XML leaking as text is listed there as not covered yet).

Tests: tests/agent/test_strip_reasoning_tags_cli.py — full block and
cut-mid-serialization tail, each asserted against both strippers. Proven red on
base: 2 failed / 4 passed without the pattern change; 6 passed with it.
2026-09-19 10:41:08 +05:30
teknium1
c20432fffa fix(kanban): worker exits EX_CONFIG on a terminal provider error
A dispatcher-spawned worker whose turn failed on a provider error that no
retry can heal (credential rejected: auth / auth_permanent, model_not_found,
ssl_cert_verification) now exits KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78,
BSD EX_CONFIG) instead of a plain 1, so the dispatcher can tell "the
provider will reject every further spawn" from "this attempt crashed".
The classification is the agent's own FailoverReason verdict already on
the turn result (failure_reason) — no error-text regex. billing stays in
the transient set: credit comes back, a revoked key does not.

Shared by both one-shot paths (-q and -Q) and gated on HERMES_KANBAN_TASK,
so a person's `hermes chat -q` keeps exiting 1 on the same error.

Part of #114587
2026-09-18 20:34:16 -07:00
ethernet
82a5affdd3 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/backup.py
#	tests/hermes_cli/test_gateway_restart_loop.py
#	website/docs/developer-guide/web-search-provider-plugin.md
#	website/docs/getting-started/installation.md
#	website/docs/getting-started/updating.md
#	website/docs/index.mdx
#	website/docs/reference/cli-commands.md
#	website/docs/user-guide/docker.md
#	website/docs/user-guide/windows-wsl-quickstart.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/developer-guide/plugins/index.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/developer-guide/web-search-provider-plugin.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/index.mdx
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/reference/cli-commands.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/reference/environment-variables.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/docker.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/features/plugins.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/security.md
#	website/i18n/zh-Hans/docusaurus-plugin-content-docs/current/user-guide/windows-wsl-quickstart.md
2026-09-18 18:41:16 -04:00
ethernet
a6ae6ace51 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.github/workflows/js-tests.yml
#	agent/model_metadata.py
#	apps/desktop/electron/main.ts
#	apps/desktop/scripts/bundle-electron-main.mjs
#	apps/desktop/src/app/settings/about-settings.tsx
#	apps/desktop/src/app/settings/gateway-settings.test.tsx
#	apps/desktop/src/app/settings/gateway-settings.tsx
#	apps/desktop/src/app/updates-overlay.tsx
#	gateway/shutdown_flush.py
#	hermes_bootstrap.py
#	hermes_cli/local_runtime/binaries.py
#	hermes_cli/main.py
#	hermes_cli/managed_uv.py
#	hermes_cli/update_cmd.py
#	hermes_cli/update_cmd_deps.py
#	hermes_cli/update_cmd_fleet.py
#	hermes_cli/update_cmd_maint.py
#	hermes_cli/update_receipt.py
#	hermes_cli/update_serve_obligations.py
#	hermes_constants.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_managed_uv.py
#	tests/hermes_cli/test_pending_supervisor_recovery.py
#	tests/hermes_cli/test_startup_fast_guards.py
#	tests/hermes_cli/test_update_desktop_stale_warning.py
#	tests/hermes_cli/test_update_fleet_restart_pending.py
#	tests/hermes_state/test_hermes_state.py
#	tests/tools/test_tirith_security.py
#	tools/bot_relay.py
#	tools/checkpoint_manager.py
#	tools/write_approval.py
#	website/docs/getting-started/updating.md
#	website/docs/reference/environment-variables.md
2026-09-18 17:26:10 -04:00
teknium1
edd4ec3124 fix(kanban): book a dead worker the same way whichever process notices it
The dead-worker sweep learned a worker's exit status only from
``_recent_worker_exits``, which ``reap_worker_zombies`` fills via
``os.waitpid`` — so only the process that spawned the worker ever knows
how it exited. With ``kanban.dispatch_in_gateway: false`` every
``hermes kanban dispatch`` tick is a fresh process, the registry is
empty, and a worker that exited rc=0 without a terminal board call was
booked as a bare ``crashed`` / ``pid N not alive``: no
``protocol_violation`` marker, no corrective error text for the retry
worker, no violation streak — and a rc=75 quota wall was counted as a
failure instead of a neutral ``rate_limited`` requeue.

A Kanban worker (``HERMES_KANBAN_TASK`` set) now writes
``[kanban-worker-exit] rc=<code>`` as the last line of its own log on
every one-shot exit path; when the registry has no entry for a dead PID
the sweep reads that trailer and books the exit through the same
code -> kind mapping. A worker killed before its exit epilogue leaves no
trailer and stays a plain crash. ``_worker_final_output`` strips the
trailer so it never leaks into the board diagnostic.

Direction (a durable, process-independent witness in the worker log)
from PR #113638; its predicate keyed on the ``Resume this session with:``
summary, which the CLI prints before ``sys.exit`` for rc 0, 1, 75 and 130
alike and so would have booked quota walls and failed turns as protocol
violations.

Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
2026-09-18 12:48:16 -07:00
teknium1
4ce832eeb9 fix(cron): bot-chat delivery cap bounds the bot's turn, not its exit linger
A cron job delivering to Bot Chat booked "timed out after 600s" for turns
that finished in seconds: cron/scheduler_delivery.py::_deliver_to_bot_chat
waited for the `hermes chat -Q` child to EXIT, but when the bot's turn had
messaged a teammate (message_agent -> notify_on_complete runner) the child
then runs the one-shot exit linger, bounded by
terminal.oneshot_completion_wait_seconds (default 600) — the same default as
cron.bot_chat_delivery_timeout_seconds. The cap counts from the claim, the
linger only starts when the turn ends, so the cap expired first on every
such delivery, booked a completed turn as a timeout, held the job's fire
fence for the full cap, and killed the child mid-linger — tearing down the
reply the linger exists to protect (#90879).

Ordering the two bounds cannot fix it (the turn has no duration bound; the
linger has its own contract), so the cap stops competing with the linger:

- hermes_cli/quiet_single_query.py: the -Q child accepts a per-process
  report path (HERMES_QUIET_TURN_REPORT_FILE), popped before the turn like
  HERMES_TURN_AUTHOR so nothing the turn spawns inherits it; cli.py writes
  {pid, exit_code, error} there the moment the turn ends, BEFORE the linger.
- cron: _run_bot_chat_turn polls the child and that report under the cap.
  Report present -> the delivery is booked from it (real exit code and
  stream tails when the child exits within a short grace) and the still-
  lingering child is left running, drained and reaped by a daemon thread.
  No report by the cap -> the turn never ended: killed and booked as a
  timeout, exactly as before. The linger itself is untouched.

Live repro (real _deliver_to_bot_chat, real `hermes chat -Q` against a
loopback provider whose turn spawns `sleep 90` with notify_on_complete,
cap 45s): base books the timeout at 45.7s for a turn that ended at +5s and
kills the child; fixed head books success at 11.2s, the child lingers, the
teammate follow-up turn runs at +97s and the child exits on its own.

Supersedes #113649's cap = delivery + linger (the thread shows a headroom
only moves the race and lengthens the fence hold); analysis credit to the
reporter and the thread's independent verification.

Fixes #113608

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:29:13 -07:00
teknium1
e63da95318 fix(cli): auth.json-only login with a benched credential is explained, not sent to the wizard
Two gaps from review of #113720's fix:

1. A profile logged in via auth.json (active_provider: nous) with no
   model.provider in config.yaml resolves as "auto". The ladder's OAuth rung
   swallowed the AuthError for "auto" and fell through to the keyless OpenRouter
   fallback, so the startup probe returned (False, None) and the first-run
   wizard ran anyway. The ladder now catches the AuthError in _ladder_rungs,
   still falls through for "auto", but stamps the swallowed error on a keyless
   fallback as `auth_error`; _probe_runtime_credentials returns it so the notice
   names the real failure.

2. The gate itself lived only in cli.py::_tui_print_startup and was untested at
   the seam (reverting cli.py left the suite green). It is now one mixin method,
   _maybe_offer_first_run_setup (tty check → probe → explain → offer), called
   from _tui_print_startup, and both tests drive that method with stdin.isatty
   patched True and _offer_first_run_setup asserting it is not called.

Tests: the benched-credential test now covers the gate and the cooldown headline
wording; the blank-install control is folded into the new auth.json-only test.
2026-09-18 10:24:15 -07:00
teknium1
d98b1480b9 fix(cli): a benched or signed-out credential prints its reason instead of the first-run provider wizard
`_runtime_credentials_ready()` collapsed "nothing is configured" and "the
configured credential is unusable right now" into one False, so a profile
whose only Nous credential was benched by a failed refresh (or quarantined
DEAD) was told "No inference provider is configured yet" and offered the
provider picker — re-running setup on an OAuth provider with single-use
refresh tokens can rotate the grant away from the session that was working
(#113720).

The readiness probe now returns the resolver's exception too
(`_probe_runtime_credentials`). At CLI startup only the resolver's
`no_provider_configured` verdict reaches the wizard; any other failure is
explained by `_explain_unusable_credentials`: `format_auth_error(exc)`
plus, from the provider's pool, "cooling down ... re-enters rotation in
about Nm" or "sign-in was lost (<reason>); run `hermes auth add <provider>`".
An actually-empty profile still gets the wizard.

Fixes #113720
Supersedes #113732 (@whyyagswhy): its `_inventory_other_providers()` gate
returns False for a Nous-configured profile, so the reported Nous case would
still have reached the wizard, and it printed no reason.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
2026-09-18 10:24:15 -07:00
izumi0uu
9104b25cb9 fix(clarify): CLI honours agent.clarify_timeout instead of a hidden 120 s cap
Remove the legacy clarify.timeout=120 CLI default so resolve_clarify_timeout
falls through to agent.clarify_timeout (3600 s default, <= 0 unlimited), the
single documented source of truth every surface shares.

The classic CLI built CLI_CONFIG from cli._cli_config_defaults(), which
seeded ``clarify: {timeout: 120}``; the resolver prefers an explicit legacy
key, so every CLI clarify modal (single question and batch panel) auto-
proceeded after 120 s even with ``agent.clarify_timeout: 900`` in the user's
config, while the TUI/Desktop and gateway paths honoured 900.

Salvage note: the one-line removal is the earliest filer's diff (#72714);
#87463, #96223, #97925 and #113879 carry the same hunk. The test drives the
CLI defaults + file merge through the resolver (a resolver-only test is
always green on main). #113879's exempt-set hunk is dropped: clarify is a
``_NEVER_PARALLEL_TOOLS`` entry and runs inline before any deadline is
armed (61645cde82), which tests/agent/test_sequential_tool_timeout.py pins.

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 10:16:14 -07:00
KoNit-K
0781f4458f fix(cli): re-resolve reasoning config after fallback 2026-09-18 09:35:49 -07:00
Hukla
a1b1c0e328 fix(cli): preserve startup alias base URL (#103938) 2026-09-17 20:29:30 +00:00
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
ethernet
b4a294fff9 Merge origin/main; keep PM as plugin dependency owner
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
2026-09-17 13:52:05 -04:00
teknium1
3187d68b49 fix(kanban): goal_mode workers keep the live tool feed in their worker log
The dispatcher forced `-Q` on goal_mode cards because cli.py only ran the
kanban judge loop in the fully-quiet one-shot branch. `-Q` strips every tool
callback, so the dashboard's Worker log (and `hermes kanban log`) stayed
blank for the whole run while non-goal cards logged normally; users read
that as "the worker is doing nothing" (Discord report, Sep 2026).

Run the judge loop on the `-q` path too, driving follow-up turns through
cli.chat so each turn's tool activity lands on stdout (= the worker log),
and print each judge verdict there as well. The dispatcher spawns goal_mode
and one-shot workers with the identical argv; the mode travels only in
HERMES_KANBAN_GOAL_MODE. The `-Q` hook stays for manual quiet runs.

Live A/B (real `hermes kanban dispatch` + spawned worker against a scripted
loopback provider that answers a terminal tool call then text, judge always
"continue", --goal-max-turns 2): base log = 94 bytes, 0 tool-feed lines,
0 verdict lines; fix log = 1910 bytes, 4 tool-feed lines, 3 verdict lines;
both arms end blocked "exhausted 2/2 turns" (loop behaviour unchanged).
2026-09-17 08:55:01 -07:00
teknium1
05e3b864a9 fix(bot-mode): the DM delivery retry reads the stream the CLI writes and resumes the persisted row
Both bot-to-bot DM retry gates (tools/bot_mode_dm.py::_run_local_turn and
tui_gateway/methods_bot_relay.py bot_relay.deliver) classified a failed turn from
`(proc.stderr or proc.stdout)`. A failed `hermes … -Q` turn prints the provider
prose on stdout (it is the turn's final_response) and `session_id: …` on stderr
on every run, so the classifier only ever saw the banner: every 429/5xx/context
overflow classified `unknown`, the policy-gated re-run never fired, and the relay
sender always got `reason: unknown`. Classify from both streams in the order the
CLI writes them (tools.bot_failure_reasons.turn_failure_text) at both sites.

Classifying alone double-delivers: the failed attempt's turn-start persist already
wrote the DM as the Bot Chat's unanswered tail row, and the re-run (a fresh
process replaying the same session + payload) appended a second identical copy.
The re-run's child env now carries HERMES_RESUME_UNANSWERED_TURN=1
(tools.bot_relay.retry_turn_env); the quiet one-shot consumes it before the turn
(hermes_cli.quiet_single_query.adopt_unanswered_turn) and re-stages the
identical unanswered tail row as the pending CLI user dict, stamped durable, so
_stage_turn_user_message reuses it and the flush writes no second row. Opt-in
by design: an identical tail alone cannot tell a re-run from a person's
deliberate re-send.

Slimmer redo of #105529 by @jonpol01 (same diagnosis and shape; the marker is
consumed at the CLI seam that already reads the dispatcher's env instead of
inside agent/turn_context.py).

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-16 17:19:58 -07:00
Sora-bluesky
9085ef967c fix(profiles): sweep the remaining pre-write mkdirs under the deleted-profile guard
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.

The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.

Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
2026-09-16 00:32:15 -07:00
teknium1
3abeca16e6 fix(kanban): 5xx and timeouts requeue the worker instead of spending its retry budget
`server_error` and `timeout` join the transient-provider set that makes a Kanban
worker exit 75 (EX_TEMPFAIL). A provider outage or a hung connection says
nothing about the task, so the dispatcher requeues without a failure tick
rather than counting toward the circuit breaker (#91206 proposed the same set).
2026-09-15 22:05:38 -07:00
teknium1
23863ccbaf fix(cli): one-shot chat -q exits non-zero on failure; 75 covers upstream 429 and overload
The non-quiet one-shot path exited 0 unless a Kanban worker was running, so
scripts could not tell a failed `hermes chat -q` from a good one and an
incomplete turn (partial, iteration budget) still read as success (#111770).
Both one-shot paths now share one contract: 0 completed, 1 failed / partial /
incomplete / never ran, 130 interrupted. The Kanban EX_TEMPFAIL sentinel also
fires for `upstream_rate_limit` (aggregator's upstream 429) and `overloaded`
(503/529): neither says anything about the task, so the dispatcher should
requeue without a failure tick rather than count it toward the breaker.
2026-09-15 21:46:41 -07:00
teknium1
647263cca9 test(cli): trim exit-contract tests to the two invariants; fix stale budget comment 2026-09-15 19:28:32 -07:00
dmelkk-secondbrain
be9d4369a7 fix(cli): one-shot -q runs report their outcome in the exit code
The Kanban dispatcher spawns workers as `hermes ... chat -q <prompt>`
(`kanban_db.py::_default_spawn`). That path ran the turn and fell through
to an implicit 0 whatever happened — success, failure, or a provider
quota wall.

`detect_crashed_workers` reads rc=0 with the task still `running` as a
protocol violation, and protocol violations trip the breaker at
`failure_limit=1`, so a single HTTP 429 blocked the card permanently and
every card queued behind it stayed in `todo` forever waiting on a parent
that could never reach `done`.

`KANBAN_RATE_LIMIT_EXIT_CODE` (EX_TEMPFAIL) exists precisely to prevent
this: `_classify_worker_exit` maps it to a `rate_limited` kind and the
task is released back to `ready` without counting a failure. The consumer
end was complete and tested. The producer end was wired into the `-Q`
path only — the one the dispatcher does not use.

This extracts that mapping into `_single_query_exit_code()` and applies it
on both one-shot paths. `chat()` returns the rendered response string, so
the non-quiet path could not see the outcome; `_chat_settle_turn` now
records the raw turn result for it to read.

Scope is deliberately narrow. The non-quiet path only exits non-zero when
`HERMES_KANBAN_TASK` is set, so interactive runs and ordinary `hermes chat
-q` invocations still exit 0 exactly as before. For a dispatcher-spawned
worker the full contract now applies: 0 on success, 1 on failure, and the
sentinel on a rate-limit/billing wall.

Tests cover the path that was missed rather than the one that already
worked: 16 of the 17 new assertions fail on the parent commit, and the
key regression fails as `assert None == 75` — the exact rc=0 fall-through
— rather than on a missing symbol. The seventeenth asserts that a human's
one-shot run keeps exiting 0, and passes both before and after.
2026-09-15 19:28:32 -07:00
teknium1
804707bea6 fix: checkpoint store gc never runs inside a tool call or gateway startup
Symptom: `hermes update` sat for ~40s after "Refreshing cua-driver" and ended
with "Fleet version check returned no rows" (exit 1); the restarted gateway
took 26s to reach "Starting Hermes Gateway" instead of the usual 3s. The
gateway constructor was running `maybe_auto_prune_checkpoints` synchronously,
before the control socket, adapters and the code_sha stamp, and on a 1.2 GB
store its `git gc --prune=now` (a full repack) takes 20-28s — twice, because
the size-cap shrink gc'd again even when it could drop nothing.

The same defect sat on the tool-call path: `CheckpointManager._take` ran
`_enforce_size_cap`, whose `_shrink_store_to_cap` returned True without
dropping anything and triggered a 20-28s gc on the first file-mutating tool
call of every turn once the store was over the cap. That loop also re-measured
a pack size that cannot move without a gc, so a single over-cap checkpoint
dropped 20 rounds of history and flattened every project to one snapshot.

- `_take` never gcs: `_prune` and `_enforce_size_cap` rewrite refs (cheap),
  drop at most one snapshot round, and mark the store `.gc-pending`.
- `prune_checkpoints` gcs only when a ref moved (project deleted, or the
  pending marker), and its cap loop is drop -> gc -> re-measure.
- `maybe_auto_prune_checkpoints` claims the interval marker before the run
  so a failing prune costs one day, not a gc per housekeeping tick.
- `auto_prune_from_config` is the one config-driven entry point; the gateway
  calls it from the housekeeping tick (last chore), the CLI from a daemon
  thread. Nothing on either startup path waits for git.

Live A/B on a copy of a real 1.2 GB / 224-ref store: checkpoint 20.5s ->
1.2-1.6s (0 inline gc); the single repack (19.6s) now runs in the prune.
2026-09-15 10:57:16 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1
e43f2f6816 fix(cli): retire the 0.0 monotonic sentinel in the remaining input-mode throttles
Review follow-up on #91651: _recover_terminal_input_modes and the termios
drift check used the same `now - 0.0 < interval` idiom as the repaint
throttles. Their windows (0.5s / 1.0s) are unreachable in practice, but
converting them to the None sentinel retires the bug class instead of the
instance, so nobody "simplifies" a None back to 0.0 later. Also adds the
missing regression test for _invalidate, the throttle the PR title is about.
2026-09-15 04:08:53 -07:00
Teknium
eab377a407 fix(cli): first repaint no longer swallowed when monotonic clock is small
time.monotonic() counts from an arbitrary epoch (boot on Linux). Two CLI
repaint throttles used 0.0 as the never-fired sentinel, so on a freshly
booted VM (CI runners, containers) now - 0.0 < min_interval suppressed
the FIRST repaint ever requested:

- _schedule_focus_regain_redraw: min_interval=60 suppressed the first
  focus-regain redraw whenever uptime < 60s — the exact failure in CI
  run 32494557030 (test_focus_regain_redraw_is_rate_limited, both
  attempts red on a fresh runner, green everywhere else).
- _invalidate: same 0.0 sentinel; a first spinner/stream repaint inside
  the first 250ms of uptime was droppable the same way.

Both now use None as the never-fired sentinel. Regression test pins
monotonic()=3.0 with min_interval=60 and asserts the first redraw fires.
2026-09-15 04:08:53 -07:00