70 nested closures (every prompt_toolkit key handler and handler factory, the
clarify/sudo/secret/approval/model-picker/command-palette/stash display
renderers, hint/placeholder/spinner callables, process_loop, spinner_loop,
_signal_handler, wake-startup) become HermesCLI._tui_* methods; run() only
binds them (kb.add(...)(self._tui_x)). The run() locals they closed over are
published on self at their original binding sites (_tui_multiline_shortcuts,
the four paste list-cells; process_loop reads self._app). Then four zero-
back-ref prologue regions become _tui_print_startup, _tui_init_run_state,
_tui_build_key_bindings -> kb and _tui_build_layout(kb) -> (layout, style).
All moved bodies verified AST-identical modulo cli_ref->self and the free-var
rewrites; unresolved-free-name check on every new method is empty. Source-
inspection tests repointed: getsource(run) -> _tui_process_loop, AST lookup of
handle_enter -> _tui_handle_enter.
process_command resolves the canonical command (hermes_cli/commands.py registry)
to a handler via _slash_handler(): a 44-entry _SLASH_DISPATCH table for the
irregular cases (no-arg handlers, differently named methods, /exit and /update
return-value adapters) and the _handle_<name>_command(cmd) naming convention for
the other 47. Inline branch bodies became small _cmd_<name> methods; the
else-fallthrough (quick_commands -> plugins -> bundles -> skills -> prefix
expansion -> unknown) became _process_unregistered_slash. Pre-dispatch side
effects (pre_command hook, pending-resume disarm) and the False-exits-REPL
contract are unchanged. Also unifies the three random-tip blocks into
_print_random_tip and the repeated arg parsing into _slash_args; removes two
zero-reference helpers (_run_curses_picker, _try_launch_chrome_debug).
Parity guard: tests/cli/test_slash_dispatch_table.py asserts every command of
the old chain resolves and that no other registry command silently gained a
handler.
GatewayRunner.__init__ snapshotted _ephemeral_system_prompt once from the
launch profile's config and _get_system_prompt_for_channel returned that
string for every source, so under multiplex a routed profile's
display.personality / agent.system_prompt never injected (#89161), and
/personality from any chat rewrote the one process-global attribute for
everyone.
Drop the snapshot: _get_system_prompt_for_channel now calls
_load_ephemeral_system_prompt() (env var, then
resolve_ephemeral_system_prompt_from_config(_load_gateway_runtime_config()))
on each call. Its caller run_sync already runs inside
_profile_runtime_scope, so the routed profile's config.yaml is what gets
read; single-profile hot-edits of the personality also take effect on the
next turn instead of requiring a restart. /personality only persists via
persist_personality() (get_hermes_home()/config.yaml = the routed profile)
and no longer touches in-memory state.
Fixes#89161
Co-authored-by: worlldz <101180447+worlldz@users.noreply.github.com>
Adds two bounded fast modes on top of the static /fast toggle, default OFF:
- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
window; requests inside it carry the provider fast param, later tool-loop
requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
user/assistant/tool history).
agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.
resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.
Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.
Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes#64785, #74730.
Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
The speed=fast allowlist still gates on Opus 4.6, but the fast-mode
matrix has changed twice since it was written (verified against the
live docs, platform.claude.com/docs/en/build-with-claude/fast-mode):
- Opus 4.8 and Opus 5 SUPPORT fast mode (research preview, Claude API
only — not Bedrock/Vertex/Foundry).
- Opus 4.6 LOST fast mode on 2026-06-29. The parameter does not error:
requests silently run at standard speed and bill standard rates
(usage.speed: 'standard'). Today's allowlist therefore shows 4.6
users a fast toggle that does nothing, while denying it to the two
models that actually support it.
- Opus 4.7 never had it and hard-400s (unchanged).
- Dedicated '…-fast' model ids (OpenRouter's claude-opus-4.8-fast)
select fast inference via the model field and are explicitly
excluded from the param gate.
Both gates move in lock-step as before: the adapter param gate
(agent.anthropic_adapter._supports_fast_mode) and the CLI toggle gate
(hermes_cli.models._is_anthropic_fast_model). Docstrings now record
the history in both directions so the next matrix change has context.
## How to test
scripts/run_tests.sh tests/agent/test_anthropic_adapter.py tests/cli/test_fast_command.py -- -q
113 tests pass. The updated predicate/matrix tests fail against the
previous allowlist (verified by stashing the source changes). Tested
on Linux (aarch64).
Desktop's cold resume (defer_history + omit_messages, transcript paged over
REST) only ever holds the live tip segment in memory, but session.resume
bounded it against the FULL compression lineage (sessions.max_resume_messages,
default 20000). A Bot Chat with 85 compaction segments / ~29k lineage rows
behind a ~700-row tip was refused at 20001, sent zero model prompts, and sat on
"Waking up default…" forever — the healthiest possible session shape, rejected
by a guard sized for in-memory materialization.
- hermes_state: one `_resume_lineage_ids` definition shared by the resume
readers (get_resume_conversations, get_ancestor_display_prefix) and the
guard (assert_resume_safe / get_resume_message_count). Guard grows
`tip_only=` and names the scope it counted; the branch-aware lineage the
readers already used is now what the guard counts too (a /branch copy was
being counted against its parent's rows).
- tui_gateway session.resume: deferred, omit_messages and lazy resumes are
bounded by the tip; only the full in-memory lineage resume keeps the
lineage-wide bound. Deferred hydration falls back to tip-only history when
the lineage exceeds the limit instead of loading the rows the guard refused.
- CLI mid-setup tip-only path routes through the same guard instead of
borrowing assert_export_safe.
- docs: sessions.max_resume_messages / max_export_messages documented with the
per-surface scope.
Live repro (real SessionDB fixture, 85 segments / 29,226 lineage rows / 666 tip
rows, real tui_gateway.server.handle_request): before — deferred resume ->
4130; after — ok, hydrated history=666 prefix=0; the non-deferred full resume
still returns 4130 on the same fixture.
hermes chat -q wired the interactive prompt_toolkit clarify callback
unconditionally, but a -q turn never builds the prompt_toolkit
application — the modal can never be painted or answered, so the turn
polls its response queue until agent.clarify_timeout expires (default
3600 s, 0 = unlimited). The gateway, cron jobs, the kanban dispatcher
and inter-agent wakeups all deliver work as -q turns. Route the
single-query case to a headless callback at the agent-construction site
that already knows _single_query_mode, mirroring _oneshot_clarify_callback
on the -z path (#94943; third member of the family after #86909 and
#88013).
Follow-up to the salvaged #88097: the same normalization covers the
Shift+letter class reported in #92343 (xterm modifyOtherKeys and both
kitty CSI-u codepoint forms), plus a guard that plain ASCII typing
never triggers the ESC-prefix predicate.
Two growth leaks closed:
1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
trees, ~18GB on the reporting box): managed installs fetch with a
single-branch refspec, so pushed PR branches never get refs/remotes/*
entries and read as 'unpushed' forever. When a clean tree's branch head
EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
is redundant: reap the TREE, keep the BRANCH ref (shielded from the
orphaned-branch pass). Anything diverged/unverifiable stays preserved.
Applied to both the startup pruner and hermes worktree prune/list.
2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
gateway-driven boxes accumulated trees for days. The scheduler tick now
dispatches the same conservative pruner on a daemon thread, throttled
to once per 6h, against the install checkout + job-workdir repos that
have a .worktrees/ dir.
Follow-up to the cherry-picked #41909/#92696 + #39760 + #97970 cluster:
- single field-key namespace (display.status_bar.fields) instead of the
second tui_statusbar_fields list; cache_hit/latency/tps/stash/battery/
title join the existing key set
- cache-hit % prefers the baseline-delta regime (resets on model switch
and compression) and hides on zero cache reads instead of alarming 0%
- latency/tps segments added to the styled fragment renderer too
- docs updated in website/docs/user-guide/configuration.md
- 7 new tests: rolling latency/t/s, NaN/negative guard, field filtering,
baseline resets, title badge gating
Add a ◎ XX% indicator to the CLI status bar showing the prompt cache
hit rate (cache_read / prompt_tokens). This helps users monitor how
effectively their provider's prompt caching is working.
Features:
- Color-coded: green (≥70%), yellow (40-70%), red (<40%)
- Adaptive precision: integer on narrow terminals, one decimal on wide
- Only shown when cache data is available (provider supports it)
- Compatible with OpenAI, Anthropic, DeepSeek, xiaomi, and other
providers that return prompt_tokens_details.cached_tokens
Tests: 6 new test cases, 43/43 passing
Allow users to control which fields appear in the interactive CLI status
bar via display.status_bar.fields in config.yaml.
Available fields: model, context_pct, context_detail, compressions,
bg_tasks, bg_processes, duration, prompt_elapsed, yolo, total_tokens.
When the list is empty (default), all fields are shown as before.
The field order is fixed (model always first); the config controls
visibility only. Narrow terminals (<76 cols) automatically drop
context_detail regardless of config.
total_tokens is opt-in only (not shown by default) to avoid width
overflow in the prompt_toolkit fragment renderer.
Closes#41909
Live-reproduced on main: /handoff poll-waited a flat 60s for a TERMINAL
state, but the gateway's dispatch is a full synthetic agent turn (whole
transcript replay + delivery) that routinely exceeds 60s on long sessions.
The CLI then printed "Timed out waiting for the gateway. Is `hermes
gateway` running?" (false diagnosis), called fail_handoff() on the RUNNING
row (stomping the gateway's claim), and promised "Your CLI session is
intact" after switch_session had already re-pointed the session. The
watcher later overwrote failed -> completed: split-brain.
- hermes_state.fail_handoff gains only_states CAS; waiters can only fail
rows still pending. Owner (gateway watcher) keeps the unconditional form.
- CLI wait loop is two-phase: 60s for the CLAIM (pending) — a timeout
there really does mean no gateway — then up to 15 min for the claimed
dispatch with 30s heartbeats; a running row is never failed by the CLI.
- Desktop handoff.fail RPC now CAS-fails pending rows only; a running row
returns {failed: false, state: running} instead of stomping the claim.
Repro (real _handoff_watcher, real state.db, CLI as separate process,
75s dispatch): before — CLI timeout @60s + false message + row stomped;
after — pending->running@5s->completed@80s, clean CLI exit.
74a95a3ddf promoted /bg and /btw to independent canonical commands and
retired /background entirely (hermes_cli/commands.py's COMMAND_REGISTRY
has no "background" command or alias). Three places still taught the
old name:
- skills/autonomous-ai-agents/hermes-agent/references/slash-commands.md:
the bundled hermes-agent skill's own slash-command reference — the
skill's SKILL.md explicitly routes the model here for in-session
command questions, so a model following it would emit the dead
`/background <prompt>` and never learn /btw exists.
- ui-tui/README.md: listed /btw as an alias of /background, which is
simply wrong post-split (both are independent, alias-free commands).
- tests/cli/test_cli_background_status_indicator.py: docstring/comments
described the ▶ indicator by the retired command name.
No behavior change; corrects documentation only.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
Reuse kitty Unicode placeholders plus after_render write_raw so
prompt_toolkit's screen-diff can host the same crisp sprite as the TUI.
Re-queue the transmit after Ctrl+L / resize so the image comes back.
Co-authored-by: Sam Foreman <saforem2@gmail.com>
The protected agent-instruction gate grants one operation and persists
nothing, but only the TUI/desktop and Runs transports were taught that.
The prompt_toolkit panel, the input() fallback, and the ACP editor menu
still rendered "Allow for session", so a user editing SOUL.md tapped it,
got re-prompted on the next write, and read the gate as broken.
Thread allow_session through prompt_dangerous_approval so a caller that
re-asks every time collapses every surface to once/deny, and cover the
producer-to-transport contract end to end.
text.splitlines() returns [] for empty strings. Accessing msg_lines[0]
then raises IndexError, making session resume crash when the session
contains a message with empty or whitespace-only text (e.g. reasoning-only
turns, tool-only assistant messages).
Guard with `or [""]` in all three branches (user, assistant_last,
regular assistant) so an empty message renders as a blank line.
Fixes#59265
Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
tests/tools/test_website_policy.py and tests/cli/test_surrogate_sanitization.py
repeatedly failed under the parallel CI runner (process-teardown timeouts /
async-timeout flakes) across multiple unrelated PRs this cycle, while passing
locally. Removed per maintainer direction to stop the flake taxing every PR.
Builds on @fattchris resolve_turn_limit salvage (#67696): flips the default
from a numeric cap to unlimited across all construction paths (CLI, agent_init,
run_agent subagents), adds inf/infinity/null to the unlimited spellings, and
sets DEFAULT_CONFIG agent.max_turns to null. The turn cap caused more problems
than it solved (silent mid-task truncation).
Two gaps behind the recurring 'hermes -w timed out after 30 seconds':
1. Rebase-merge leak: git cherry only catches patch-identical commits.
Salvage flows routinely change the diff (conflict resolution, follow-up
commits), so 12 of 22 'unpushed' trees on the incident box had MERGED
PRs and were preserved forever. The pruner now falls back to
'gh pr list --head <branch> --state merged' — authoritative, memoized
on (branch, head_sha) with True-only caching, fail-safe to preserve.
2. Creation timeout 30s -> 120s: the ~10k-file checkout measured 113s at
near-zero CPU under multi-agent disk contention vs 1.2s idle. 30s
killed legitimate creates and threw away completed work.
Follow-up to the salvaged #89676 + #90291 lock-bit fixes: extract a shared _lock_variants() helper and cover the sites both PRs missed - install_shift_enter_alias / install_ctrl_enter_alias / install_cmd_backspace_alias CSI-u spellings, legacy CSI-letter and CSI-tilde navigation twins derived from the existing table for ALL modifiers 1-16 (not just plain/shift), plain F1-F4 SS3 fallback, unmodified CSI-u keys (Tab/Enter/Space/Backspace), and kitty PUA functional keys (keypad, F13-F24, Ignore range) under lock bits. 8 new tests.
kitty and ghostty OR the CapsLock (64) / NumLock (128) state into the
CSI-u modifier parameter. With NumLock on, Ctrl+C arrives as
ESC[99;133u (5 + 128) instead of ESC[99;5u; the alias table had no
entry for it, so every key combo leaked as literal text like
[127;133u (#89651). Install every CSI-u alias with the lock-bit
variants (+64/+128/+192); the xterm modifyOtherKeys encoding never
carries lock bits, so the ESC[27;N;CP~ form is left untouched. The
Esc-key registration now covers modifier 1 as well (1+128=129 is a
lone Esc with NumLock on).
Widen the cli.py Ghostty exception to the sibling sites the review found: the Ink TUI pushes CSI >1u at raw-mode entry (App.tsx), on alt-screen exit, and on the extended-keys re-assert path (ink.tsx) for every EXTENDED_KEYS_TERMINALS entry including ghostty - same Alt-stripping bug. New skipKittyKeyboardProtocol() helper in terminal.ts gates the ENABLE push at all 3 sites; the DISABLE (pop) stays unconditional since popping an empty stack is a spec no-op. Also fix the cli.py comment citing the modifyOtherKeys encoding where the kitty CSI-u form (ESC[127;3u) is what the broken path expected, dedupe the quadruplicated Ghostty comment, and update the stale 'mirroring the Ink TUI' docstring. 7 new vitest cases.
Ghostty's Kitty disambiguate-mode implementation strips the Alt modifier
from the Backspace key — Option+Backspace arrives as bare \x7f instead of
the expected \x1b[27;3;127~, breaking backward-kill-word. This was a
regression introduced when PR #87630 re-added the CSI >1u Kitty protocol
push for all allowlisted terminals including Ghostty.
Under modifyOtherKeys mode (CSI >4;2m), Ghostty correctly sends
\x1b[27;3;127~ for Option+Backspace, which the alias table in
pt_input_extras already maps to (Escape, ControlH) = backward-kill-word.
Fix: for Ghostty only, push just modifyOtherKeys and skip the Kitty
protocol push. All other terminals (iTerm2, WezTerm, kitty, tmux, VS Code)
still get the full dual-protocol push.
Ghostty upstream tracking: discussion #9560, issue #9895 (cmd+backspace
variant of the same root cause).
Shift-Tab walks backwards through the questions, with wrap, the same
way Tab walks forward. A locked answer now renders on its own indented
line in a distinct color under its question, instead of an arrow
suffix on the status line, so the current answers stay readable while
the cursor moves.
A re-visited question restores its earlier state: a choice answer puts
the cursor back on that choice, a typed answer highlights the Other
row and shows the typed text next to it. Enter on an answered Other
switches to freetext with the composer prefilled with the earlier
text, so the user edits instead of retyping. Answer metadata records
how each answer was produced to drive the restore.
The clarify callback accepts a questions list and renders a batch
panel. The batch panel shows all questions as a status list with one
expanded active question. Enter locks the active answer and moves to
the next unanswered question. Tab cycles questions for any-order
answering. Locked answers stay editable until the batch completes. A
timeout returns the locked partial answers with a timed_out flag. The
single-question panel is unchanged.
Bot Mode's bot-to-bot send (`hermes -p <bot> chat --in ~ -c "Bot Chat"
--create-if-missing -Q -q "..."`) runs one turn and exits. When the turn's
in-loop transcript flush failed transiently (state.db write-lock contention
with a multiplex gateway), the one-shot path had no end-of-run durable
retry: the reply reached stdout and agent.log while the resumed titled
session's stored history never changed (#88583). The interactive CLI is
immune — it retries the flush on the next persist point and finalizes the
row on quit — but every one-shot exit path lacked both.
Fix the whole class with cli._flush_one_shot_session_store():
- final _persist_session retry at one-shot exit (idempotent — per-message
persisted-marker stamps mean already-written turns are not re-written)
- drain queued async token-accounting deltas
- end_session(..., "cli_close") so resumed/created titled session rows no
longer dangle open forever after one-shot runs
Wired into _finalize_single_query (quiet -Q -q AND human -q paths, ahead
of memory-provider shutdown so nothing later can lose the turn) and into
the kanban SIGTERM handler before os._exit(0), which skips atexit and the
SessionDB token-drain hook entirely (same gap class as PR #50881).
Handed-off sessions (#88234) and persistence-isolated forks
(_persist_disabled) are skipped.
Fixes#88583🤖 Generated with Hermes Agent
CI git consolidates during incremental pack creation differently per
build (4 packs from 6 attempts on ubuntu-latest, 6 locally, 3 on the
previous run) — even pack-objects counts drift with auto-maintenance.
The fixture now only guarantees strictly-more-packs-than-threshold and
the test asserts consolidation strictly decreases the count.
Incremental 'git repack' consolidates small packs on newer git builds
(CI produced 3 packs from 6 commits), making the sprawl fixture count
nondeterministic. pack-objects with an explicit sha per commit creates
exactly one pack each on every git version.
Two gaps from the Aug 2026 'hermes -w timed out after 30s' incident:
1. Atomic failure cleanup: a timed-out/failed `git worktree add` left a
partially-materialized directory plus a LOCKED admin entry under
.git/worktrees/ (lock pid = the live hermes process that timed out),
which the startup pruner's dead-pid unlock never reaps — retries of
the same name fail forever. _cleanup_failed_worktree_add sweeps dir,
admin entry, and orphaned branch on every failure path (timeout,
nonzero exit, remote-base retry).
2. Pack maintenance: nothing consolidated the object store; on a
multi-agent box packs sprawl (39 packs / 638MB at the incident) and
every object lookup scans all pack indexes until worktree creation
blows its timeout. _maintain_pack_health repacks (niced, background,
fail-soft) when *.pack count reaches 15, wired into the existing
startup maintenance thread on both the CLI (-w) and TUI paths.
gc --auto doesn't cover this: its threshold is 50 packs.
Both sabotage-verified; full repack on the incident box: 39 packs ->
2, 638MB -> 287MB, worktree add 30s-timeout -> 0.5s.
/simplify-code review found _notify_single_query_session_finalize was
missing the _handed_off_session_ids guard that _should_emit_cleanup_session_finalize
and _emit_interrupted_session_end already had. One-shot CLI queries that
somehow handed off would still finalize the session via this path.
Added guard + test.
Two data-loss bugs reported by users:
1. /handoff CLI→gateway race (#88234): After /handoff completed, CLI
cleanup called finalize_session on the session the gateway just
reopened. This set end_reason on a row the gateway was actively
writing to, causing the handoff leg to vanish from session history
and breaking session_search recall. Fix: add _handed_off_session_ids
module-level set (mirrors _single_query_finalize_attempted_session_ids
pattern). _handle_handoff_command registers the session_id on
completion; _should_emit_cleanup_session_finalize and
_emit_interrupted_session_end check it before firing.
2. state.db corruption silent failure (#88235): When SessionDB init
failed at gateway startup, the error stayed in logs — messages
flowed but nothing was persisted, with no user-visible indication.
Fix: store _session_db_init_error on GatewayRunner, broadcast a
recovery-guidance message to all home channels via
_send_session_db_warning_notifications() after the gateway connects.
Also improved the 'corrupt' persistence cause wording in
_format_turn_completion_explanation to include the full recovery
path (hermes doctor --fix, sqlite3 .recover, backups).
Tests: 6 new tests for handoff cleanup race, 3 for corruption wording.
All existing CLI/turn-completion tests pass.