Commit Graph

1201 Commits

Author SHA1 Message Date
Teknium
73475d62f6 refactor(cli): main() 393 -> 226 via _build_cli_from_args / _start_worktree_setup / _run_legacy_gateway
The worktree block keeps its 'global _active_worktree' writer inside cli.py (the
helper is module-level in the same file), so the cleanup path's reads are unchanged;
main() no longer declares the global it never wrote directly.
2026-09-02 16:34:20 -07:00
Teknium
84d04094e3 refactor(cli): compact seven long narrative comment blocks (WHY kept, prose trimmed) 2026-09-02 16:19:03 -07:00
Teknium
7e250ce03e refactor(cli): no function over 300 LOC — split __init__/main/run/load_cli_config + two TUI builders
__init__ 659 -> 39 (_init_display_options / _init_model_routing / _init_runtime_state);
main 651 -> 403 (_install_single_query_signal_handlers, _collect_kanban_task_images,
_route_single_query_images, _run_quiet_single_query as module-level helpers);
run 384 -> 293 (_tui_shutdown); load_cli_config 366 -> 115 (_cli_config_defaults,
_mirror_config_to_env); cli_tui_mixin: _tui_build_layout 352 -> 196 (_tui_build_input_area,
_tui_set_base_style), _tui_handle_enter 333 -> 192 (_tui_enter_overlay -> bool).
All regions are zero-back-ref lifts, ast.dump-identical bodies, unresolved-free-names
check clean. Two source-inspection tests repointed at the helper that now owns the code.
2026-09-02 16:11:02 -07:00
Teknium
7f9f274cde refactor(cli): chat() 855 -> ~240 LOC via _ChatTurn dataclass + 5 phase methods
Per-turn locals (result, TTS pipeline state, voice prefix) move onto a
_ChatTurn dataclass shared with the agent worker thread (same-object
visibility, so the abandoned-thread late write still lands). Four zero-
back-ref regions and the run_agent closure body become
_chat_setup_turn_audio / _chat_run_agent / _chat_monitor_agent_thread /
_chat_settle_turn / _chat_render_turn. Bodies are ast.dump-identical after
the rename; each helper passed the unresolved-free-names check.
2026-09-02 15:53:22 -07:00
Teknium
eb74a00c71 refactor(cli): split HermesCLI into 10 cohesive mixins (cli.py 22,284 -> 9,150)
326 methods lifted by AST (bodies identical; ast.dump-verified) into
hermes_cli/cli_{tui,status_bar,voice,model_switch,session,stream,modal,
terminal,info,loops}_mixin.py. cli.py-internal symbols resolve via lazy
'from cli import ...' inside each method (no import cycle; patch('cli.X')
keeps working). The three 'global' writers (_skill_commands, _cli_wake_owner)
now write the cli module attribute explicitly so the origin's readers still
see them. Dropped imports left unused in cli.py; kept display_hermes_home /
build_welcome_banner as re-exports (mixins + tests resolve them via cli).
Repointed two AST change-detector tests to cli_tui_mixin.py; one test
fixture now keeps 'cli' in sys.modules across its patch.dict scope.
2026-09-02 15:42:24 -07:00
Teknium
3e38ef2126 refactor(cli): drop unused browser_connect imports; unify Alt+Enter/Ctrl+J newline handlers; collapse 8 trivial if/else forms
Also restores the reload-mcp auto-reload rationale in _confirm_and_reload_mcp's docstring.
2026-09-02 13:29:43 -07:00
Teknium
57b2dfdf0e refactor(cli): decompose HermesCLI.run() (3433 -> 384 LOC) into _tui_* methods
70 nested closures (every prompt_toolkit key handler and handler factory, the
clarify/sudo/secret/approval/model-picker/command-palette/stash display
renderers, hint/placeholder/spinner callables, process_loop, spinner_loop,
_signal_handler, wake-startup) become HermesCLI._tui_* methods; run() only
binds them (kb.add(...)(self._tui_x)). The run() locals they closed over are
published on self at their original binding sites (_tui_multiline_shortcuts,
the four paste list-cells; process_loop reads self._app). Then four zero-
back-ref prologue regions become _tui_print_startup, _tui_init_run_state,
_tui_build_key_bindings -> kb and _tui_build_layout(kb) -> (layout, style).

All moved bodies verified AST-identical modulo cli_ref->self and the free-var
rewrites; unresolved-free-name check on every new method is empty. Source-
inspection tests repointed: getsource(run) -> _tui_process_loop, AST lookup of
handle_enter -> _tui_handle_enter.
2026-09-02 13:29:43 -07:00
Teknium
0227cf7d9c refactor(cli): unify the 3 copies of the TUI panel helpers into module-level functions
_panel_box_width/_wrap_panel_text/_append_panel_line/_append_blank_panel_line
were defined identically inside run(), _get_approval_display_fragments and
_get_slash_confirm_display_fragments (the latter two with whitespace-preserving
wrap and different width defaults). One module-level set; the two fragment
renderers alias _wrap_panel_text_keep_ws and pass their width defaults explicitly.
2026-09-02 13:29:42 -07:00
Teknium
9d745e3162 refactor(cli): replace the 93-branch slash-command if/elif chain with convention-based dispatch
process_command resolves the canonical command (hermes_cli/commands.py registry)
to a handler via _slash_handler(): a 44-entry _SLASH_DISPATCH table for the
irregular cases (no-arg handlers, differently named methods, /exit and /update
return-value adapters) and the _handle_<name>_command(cmd) naming convention for
the other 47. Inline branch bodies became small _cmd_<name> methods; the
else-fallthrough (quick_commands -> plugins -> bundles -> skills -> prefix
expansion -> unknown) became _process_unregistered_slash. Pre-dispatch side
effects (pre_command hook, pending-resume disarm) and the False-exits-REPL
contract are unchanged. Also unifies the three random-tip blocks into
_print_random_tip and the repeated arg parsing into _slash_args; removes two
zero-reference helpers (_run_curses_picker, _try_launch_chrome_debug).

Parity guard: tests/cli/test_slash_dispatch_table.py asserts every command of
the old chain resolves and that no other registry command silently gained a
handler.
2026-09-02 13:29:41 -07:00
Teknium
f6234d00c5 fix(security): close GitSpawn RCE class — malicious repo .git/config no longer executes on context gathering (GHSA-7x36-8jrh-v4pw)
Hermes gathers workspace context by running git against the session
directory automatically — the coding-workspace snapshot, gateway
project-tree build, /diff, @diff|@staged context refs, goal-gate
fingerprint, and -w startup worktree add — before any prompt, tool call,
approval, or trust gate. Those probes ran the system git without
stripping the repository's own config, so a repo delivered as files with
its .git directory intact (a shared zip, sync folder, or USB stick;
git clone never transfers .git/config) could set an execution-sink git
setting and get arbitrary host code execution as the user with nothing
on screen.

- core.fsmonitor / core.hooksPath / pager / editor / credential helper:
  neutralized by routing every automatic probe through
  noninteractive_git_env(), which pins those keys to inert values via
  GIT_CONFIG_* and ignores global/system config. bounded_git_probe (the
  reported sink, coding_context._git + tui_gateway.git_probe) now defaults
  to that env; worktree-add, working_diff, web_git, context_references,
  goals, and subagent_worktree route through it too.
- Attribute-scoped [diff "x"] command=/textconv= drivers: the attacker
  names the driver in .gitattributes, so GIT_CONFIG_KEY overrides can't
  enumerate them. Added harden_git_argv(), which inserts
  --no-ext-diff --no-textconv on diff-rendering subcommands (diff/show/
  log/blame) only — status et al reject the flags. Both flags required
  (verified empirically; each alone leaves the other live).

Builds on the noninteractive_git_env config-scrubbing from the
gemini-cli #28792 port. Real-git E2E regression suite arms a malicious
repo and asserts every automatic path neutralizes fsmonitor, hooks,
external-diff, and textconv; a baseline test proves the repo is armed.
2026-09-02 10:33:43 -07:00
Teknium
8e4366d358 fix(tools): freeze tools[] across agent-cache eviction; make /reload-mcp the re-probe hatch
Policy: availability-gated tools (check_fn probes — Docker, HASS_TOKEN,
OAuth…) are frozen for the life of a session. tools[] only changes on
/new, /reload-mcp, or compaction. Two doors remained after #100638:

* Gateway agent-cache eviction (LRU/idle sweep/cross-process invalidation)
  rebuilds a fresh AIAgent for the SAME session and agent_init re-derives
  agent.tools from live probes with no predecessor to preserve. Persist
  the session's resolved tool-name order in a new `sessions.tool_names`
  JSON column (declarative reconciliation, SCHEMA_VERSION 28), written
  alongside the system prompt and re-pinned on every published refresh
  (so /reload-mcp and compaction naturally reset it; /new mints a new
  row). On restore-for-existing-session the fresh definitions are folded
  onto the saved order via the SAME `_merge_preserving_prefix` helper —
  a probe-flipped tool is carried forward from the registry schema, a
  deregistered one dropped, new tools appended at the tail.

* /reload-mcp (CLI, gateway, TUI RPC) now also calls
  `reprobe_tool_availability()` — drops the check_fn verdict cache and the
  get_tool_definitions memo — so a user can consciously pick up a
  credential/daemon that appeared mid-session. Docs updated.
2026-09-02 07:22:59 -07:00
Teknium
632078bca7 feat(cli): OSC 9 + Warp OSC 777 notifications ride on the bell flags
Extend _ring_bell() so display.bell_on_prompt / bell_on_complete also
emit terminal-native desktop notifications from the same six call sites
(clarify, clarify batch, approval incl. computer_use, sudo password,
secret capture, turn complete). No new config keys.

- OSC 9 (ESC ] 9 ; body BEL): Ghostty / iTerm2 / Kitty / WezTerm raise an
  OS notification; unknown terminals drop it. Body is "Hermes: <context>"
  with C0 controls and DEL stripped. Written to /dev/tty (prompt_toolkit's
  stdout wrapper can buffer/strip raw escapes) with a sys.stdout fallback.
- Warp OSC 777 warp://cli-agent (agent "hermes", event permission_request
  / stop, compact JSON mirroring build-payload.sh). Gated on
  TERM_PROGRAM=WarpTerminal + WARP_CLI_AGENT_PROTOCOL_VERSION + the
  should-use-structured.sh broken-build floor (stable/preview builds at or
  before v0.2026.03.25.08.24.*_05 rejected). Never raises.

Salvages #58957 and #100805.

Co-authored-by: glitchbunny0 <glitchbunny0@proton.me>
Co-authored-by: harsha-usethread <harsha@usethread.io>
2026-09-02 06:17:10 -07:00
Teknium
552159d222 feat(cli,tui): collapse bell_on_clarify/approval into display.bell_on_prompt
One key covers every blocking prompt modal: clarify (single + batch),
dangerous-command approval (incl. computer_use), sudo password, and
secret capture. CLI gets a _ring_bell() helper shared with
bell_on_complete; TUI rings on clarify/approval/sudo/secret .request
events (isTTY-gated). 'hermes config' Bell summary shows both flags.
2026-09-02 05:34:35 -07:00
Turgut Kural
3082a34669 feat(cli,tui): add display.bell_on_approval + fix eslint error
- display.bell_on_approval (default false): same BEL mechanism as
  bell_on_complete, rings when a dangerous-command approval prompt
  opens (_approval_callback / approval.request event). Complements
  bell_on_clarify from the previous commit.
- fix(ui-tui): eslint curly error in useConfigSync.applyDisplay
  (if without braces) that failed the CI JS & TS checks job.
2026-09-02 05:34:35 -07:00
Turgut Kural
ef6d3367a6 feat(cli,tui): add display.bell_on_clarify — terminal bell on clarify prompts
Same BEL mechanism as display.bell_on_complete (\a / \x07), gated by
display.bell_on_clarify (default false). CLI rings in _clarify_callback
and _clarify_callback_batch before _paint_now(); TUI rings on
clarify.request when bellOnClarify && stdout.isTTY. Docs in
cli-config.yaml.example and website/docs/user-guide/configuration.md.
2026-09-02 05:34:35 -07:00
Teknium
aaa34b0e08 fix(desktop): model picker no longer hardcodes --global; one persist policy for every surface (#90235)
Symptom: picking a model in the Desktop composer for the primary chat
silently rewrote config.yaml (model.default + model.provider) as the
profile default, ignoring model.persist_switch_by_default. A throwaway
pick that resolved to e.g. openai-api (no key) left the profile with an
unusable default on the next launch (#90235).

Root cause: 7d96537bc8 (#86414) made use-model-controls.ts send --global
for every primary-tile pick so a fresh profile would get a persisted
provider instead of falling through to a leftover OPENAI_API_KEY env var.
That put a persistence policy in the client, contradicting the
server-side rule /model uses (resolve_persist_behavior).

Fix:
- resolve_persist_behavior gains one rule, ahead of the --provider
  session-only rule: when neither model.default nor model.provider is
  configured yet, persist. This preserves #86414's first-pick motivation
  for CLI, gateway and Desktop alike. With a default configured, a plain
  pick is session-only unless --global / persist_switch_by_default.
- Desktop primary-tile picks send no scope flag and let the gateway decide.
  Secondary tiles and MoA presets still send --session.
- /model help text in cli.py said "(persists)"; it now matches reality and
  lists --global.
- Docs: desktop.md picker note + slash-commands /model row.

Tests: test_first_pick_persists_then_session_only (fails on main), and the
existing use-model-controls vitest updated to assert the flag-less request.
2026-09-02 05:33:33 -07:00
Teknium
c7e2e0b779 feat(fast): bounded /fast auto|cold windows behind one route-aware gate
Adds two bounded fast modes on top of the static /fast toggle, default OFF:

- `auto`: every user turn opens a `agent.fast_auto_seconds` (default 60s)
  window; requests inside it carry the provider fast param, later tool-loop
  requests fall back to standard pricing.
- `cold`: the same window, but only on the first turn of a session (no prior
  user/assistant/tool history).

agent/fast_mode.py holds the whole policy: `begin_turn()` at the
run_conversation ingress arms `agent._fast_until`; `effective_request_overrides()`
is consumed in the ONE place request_overrides feed the transports
(build_api_kwargs), so the fast param is a per-request kwarg only. System
prompt, tools and messages are untouched — the prompt cache is preserved.

resolve_fast_mode_overrides() is now the single gate for static and bounded
modes and accepts provider/base_url: OpenRouter, Nous, Copilot, Azure,
Bedrock and custom base_urls never receive service_tier/speed (#34308's
route gating). Both existing callers (CLI turn route, gateway turn route)
and the TUI config.set path pass the route.

Surfaces: config `agent.service_tier: auto|cold` + `agent.fast_auto_seconds`,
`/fast auto|cold` in CLI, gateway (picker gains both entries), TUI/desktop
config.set; status shows the mode; web dashboard select lists the real
values. Docs: configuration.md Fast Mode section with mode table + cost note,
slash-commands, cli-config.yaml.example, locale strings for the two picker
entries.

Salvages #89991 (bounded fast modes) and #34308 (route gating).
Fixes #64785, #74730.

Co-authored-by: Eva <239388517+100yenadmin@users.noreply.github.com>
Co-authored-by: kbaicai <kbaicai@qq.com>
2026-09-02 05:33:13 -07:00
Teknium
4e3feb8bbb feat(tts): speech toggles warm up and unload local TTS engines (#100881)
Desktop "Read replies aloud" / voice conversation, TUI and CLI /voice tts
now hold a lease on the TTS engine. Acquiring pre-loads the configured
provider (piper/kittentts model into the same LRU slot synthesis reads;
lazily-installed cloud SDKs), so the first spoken reply no longer pays the
model load as dead air. Releasing the last lease across surfaces unloads
resident local models.

- tools/tts_tool.py: warm_tts_provider / release_tts_provider /
  acquire_tts_lease / release_tts_lease over a _LOCAL_TTS_MODEL_CACHES
  registry; piper/kittentts loaders extracted so warm-up and synthesis
  share one resolution path.
- web_server: POST /api/audio/tts-lease (profile-scoped, off-loop,
  failures reported in body never as HTTP errors).
- tui_gateway voice.toggle + cli.py /voice tts|on|off wire the lease.
- desktop: lib/tts-lease.ts (dedupe, per-lease serialization, latest
  intent wins) driven from useComposerVoice; setTtsLease API client.
- docs: features/tts.md section.

Live (real piper, isolated HERMES_HOME): first synthesis 988ms cold →
92ms after the toggle warmed the engine; release drops the model.
2026-09-01 21:43:59 -07:00
Teknium
c0495c6bce fix(cli): context meter no longer sawtooths on reasoning models — show durable transcript, not last-request replay
On reasoning models a long tool loop replays the current turn's thinking +
scaffolding on every request, so the LAST request's prompt_tokens can exceed
the durable transcript by hundreds of K — all of which evaporates at the turn
boundary. The status bar and /context breakdown rendered that raw figure, so
users watched 'context' jump (e.g.) 850K -> 600K across a turn boundary and
read it as a broken compaction.

- conversation_loop: capture a turn-base usage anchor from the turn's FIRST
  provider response (api_call_count == 1), where replay is minimal.
- anchored_context_tokens: new charge_stale_thinking kwarg forwarded to the
  delta estimate (stale reasoning excluded on all but the newest assistant
  message).
- cli status snapshot + context_breakdown: prefer the turn-base anchored
  figure; fall back to last-response anchor / raw last_prompt_tokens.
- All _usage_anchor invalidation sites also clear _turn_base_usage_anchor.

Display-only: compression trigger math keeps using real last-request usage
(the inflated request is what actually risks the window mid-loop).
2026-09-01 15:34:03 -07:00
Teknium
33797073bb fix: harden startup route salvage — aggregator-slug guard, alias credential ownership, oneshot dedup
Follow-ups on top of #87210 (@liuhao1024) and #87246 (@JoaoMarcos44):
- resolve_startup_model_route: aggregator-native slugs stay on the current
  routing aggregator (bare vendor slugs resolve WITHIN the aggregator first);
  URL-bearing aliases resolve via direct_alias_runtime_request so a foreign
  provider label never carries the vendor token to the alias host (#28660);
  route carries the alias's own api_key.
- cli.py: pass current_provider; explicit --api-key wins over alias key.
- Drop #87246's oneshot double-handling (main's oneshot alias+detection path
  already covers it once #87210's detection fix is in) and the PR-body SVG.
- Rewrote/extended startup-route tests for the hardened semantics.
2026-09-01 12:07:35 -07:00
joaomarcos
4c870951e2 fix(cli): resolve startup model routes before provider defaults
Resolve configured aliases and provider/model inputs before HermesCLI attaches the configured default provider. Keep aggregator namespaces intact, cover oneshot startup, and document the supported CLI forms.
2026-09-01 12:07:35 -07:00
Shannon Sands
f5bb1e144d fix(gateway): address startup-watchdog review findings (OOF-298, PR #89750)
Independent review of the initial startup-liveness watchdog surfaced two
P1s and three P2s. All are addressed here.

P1 — legitimate slow startups (large state.db schema migrations inside
SessionDB.__init__, which run synchronously before the loop starts) could
exceed the fixed 300s deadline and restart-loop. The watchdog now checks
process CPU time (time.process_time(), process-wide) when the deadline
expires: continuous CPU consumption means a live migration, so the deadline
is extended (with a warning log per extension). The OOF-298 deadlock class
parks every thread in futex waits and accrues ~zero CPU, so it still fires
on schedule. Documented limitation: a spinning busy-wait deadlock reads as
progress and won't fire — the observed incident class is parked threads.

P1 — import-time deadlocks were outside coverage. The implementation moved
to a stdlib-only top-level module (hermes_startup_watchdog), and
hermes_cli/main.py arms it via an argv fast-path ("gateway" + "run" in
argv) BEFORE the heavy module-level import graph. gateway/startup_watchdog
remains as a re-export shim so the intuitive import path keeps working for
the disarm site, tests, and REPL use. Import-lightness is a correctness
property, tested via AST inspection: at fire time the wedged main thread
may hold the import lock, so the fire path performs no imports on its own
thread — the lifecycle-ledger write runs on a bounded-join helper thread
and os._exit happens regardless.

P2 — disarm/fire race: the handle now has an explicit state machine
(armed → disarmed | firing) guarded by a lock; whichever transition takes
the lock first wins, so a disarm landing after deadline expiry but before
the fire transition is honored. Regression test forces the exact
interleaving by blocking inside the CPU probe.

P2 — uncovered entry points: cli.py --gateway and scripts/hermes-gateway
run_gateway() now arm the watchdog before importing the gateway graph.
hermes_cli/gateway.py run_gateway() keeps an idempotent backstop arm for
programmatic callers.

P2 — respawn-storm backoff interaction: the storm breaker's intentional
backoff sleep (up to minutes, ~zero CPU — indistinguishable from a parked
deadlock) now calls kick_startup_watchdog(extra_s=backoff) so the deadline
is pushed past the sleep instead of firing mid-backoff.

Also: the faulthandler stack dump is now additionally written to
logs/gateway-startup-watchdog.log (stderr may be absent on detached/
windowless runs); the disarm site in gateway/run.py moved inside the
loop-confirmed branch (if the loop is NOT live, the milestone was not
reached and the watchdog must stay armed); hermes_startup_watchdog added
to pyproject py-modules so sealed venvs ship it; SERVICE_RESTART_EXIT_CODE
is duplicated in the stdlib-only module with a parity test against
gateway.restart.

Tests: 38 in tests/gateway/test_startup_watchdog.py (contracts incl.
stdlib-only AST check and shim re-export identity, config resolution,
arm/disarm/kick, CPU-progress extension vs no-progress fire, probe-failure
fails toward firing, disarm-vs-fire race, dump record + file stacks,
lifecycle ledger, custom exit code).
2026-08-31 14:01:39 -07:00
Futahua
a5f0fbb262 fix(sessions): per-session exclusivity is correctness, not a capacity policy
Cherry-picked from PR #94595 (author: Futahua) onto current main, with the
maintainer-review revision points folded in during the rebase:

- the lease engages UNCONDITIONALLY: try_acquire_active_session no longer
  returns a disabled no-op lease when max_concurrent_sessions is unset;
  the concurrency cap stays an orthogonal, optional policy checked second
- ownership uncertainty fails CLOSED (SESSION_COORDINATION_UNAVAILABLE)
  instead of degrading to an untracked go-ahead: a corrupt/unreadable
  registry must not be collapsed into 'no owner exists' (review blocker 2)
- the ownership admission sits at the _run_prompt_submit chokepoint that
  EVERY fresh turn source crosses, and crash auto-continue acquires (or
  bails) BEFORE emitting message.start — closing the #94778 bypass where
  backend B's auto-continue ran a duplicate turn while backend A was live
  (review blocker 1)
- the TUI gateway claim helper fails closed on claim exceptions for every
  surface, not just desktop
- CLI and messaging-gateway call sites pass live_session_id metadata so
  the (pid, live id) re-entrancy identity protects them from self-fencing
  on a leaked lease

Co-authored-by: teknium1 <teknium1@users.noreply.github.com>
2026-08-31 12:36:33 -07:00
HexLab98
fa3471d462 fix(cli): print partial-update hint when chat startup hits a first-party ImportError (#96900)
HermesCLI construction imports helpers from hermes_cli.config before the agent-setup mixin can run, so a mixed-version tree crashed with a raw traceback. Catch that ImportError on the chat entry path and tell the user to run hermes update.
2026-09-01 00:26:55 +05:30
webtecnica
839de43d52 fix(cli): stop raw CSI bytes from Shift+Space leaking into buffer (#88071) 2026-08-31 10:08:48 -07:00
Teknium
3a351a9665 feat(worktree): pushed open-PR lanes reclaim their disk; cron tick prunes worktrees
Two growth leaks closed:

1. Pushed-branch tier (the dominant survivor class — 24 of 33 preserved
   trees, ~18GB on the reporting box): managed installs fetch with a
   single-branch refspec, so pushed PR branches never get refs/remotes/*
   entries and read as 'unpushed' forever. When a clean tree's branch head
   EXACTLY matches origin (one lazy git ls-remote per sweep), the checkout
   is redundant: reap the TREE, keep the BRANCH ref (shielded from the
   orphaned-branch pass). Anything diverged/unverifiable stays preserved.
   Applied to both the startup pruner and hermes worktree prune/list.

2. Cron-tick maintenance: the pruner only ran on hermes -w launches, so
   gateway-driven boxes accumulated trees for days. The scheduler tick now
   dispatches the same conservative pruner on a daemon thread, throttled
   to once per 6h, against the install checkout + job-workdir repos that
   have a .worktrees/ dir.
2026-08-31 07:27:50 -07:00
Stephen Chin
48a4201f40 fix(compaction): preserve switch compatibility fixtures
Keep model-switch callers compatible with result objects created before runtime_capabilities was added, and do not roll back minimal agents that lack optional LM Studio helpers. Preserve rollback for real helper failures.
2026-08-30 05:16:10 -07:00
Stephen Chin
08c7879ca1 fix(compaction): preserve native capability across runtime switches
Stage destination native-compaction capabilities until the complete runtime and context setup succeeds, and restore them with primary and fallback runtimes. Keep native compaction default-deny across live switches and session reconstruction.\n\nVerification: uv run --with pytest --with pyyaml python -m pytest tests/run_agent/test_switch_model_context.py tests/run_agent/test_native_compaction.py tests/run_agent/test_native_compaction_switch_capabilities.py tests/run_agent/test_switch_model_rollback.py tests/run_agent/test_fallback_reasoning_override.py tests/run_agent/test_primary_runtime_restore.py tests/run_agent/test_provider_fallback.py -q -o 'addopts='; uv run --with ruff ruff check <touched files>; git diff --check
2026-08-30 05:16:10 -07:00
webtecnica
5c6e5e7ea3 feat(cli): add /plan command (#67264)
Generate a structured execution plan without executing tools.
Uses _pending_agent_seed injection (same pattern as /moa).
2026-08-29 19:14:15 -07:00
Teknium
9e017428ba fix: unify status-bar field keys, docs, and tests for salvaged cluster
Follow-up to the cherry-picked #41909/#92696 + #39760 + #97970 cluster:
- single field-key namespace (display.status_bar.fields) instead of the
  second tui_statusbar_fields list; cache_hit/latency/tps/stash/battery/
  title join the existing key set
- cache-hit % prefers the baseline-delta regime (resets on model switch
  and compression) and hides on zero cache reads instead of alarming 0%
- latency/tps segments added to the styled fragment renderer too
- docs updated in website/docs/user-guide/configuration.md
- 7 new tests: rolling latency/t/s, NaN/negative guard, field filtering,
  baseline resets, title badge gating
2026-08-29 18:34:51 -07:00
Turgut Kural
3548fc809b feat(cli): tui status bar per-field toggle + cache/latency/tps
- Add rolling status bar metrics:
  - cache hit ratio (◈) delta since model/compression reset
    (hit = cache_read / prompt, verified against live logs)
  - avg latency (◷) and throughput (↑ t/s) over last 10 API calls
    (deque in agent, displayed in wide bar only)
- Add display.tui_statusbar_fields list to filter segments:
  model, ctx, ctx_bar, cache_hit, latency, tps, compressions,
  bg_tasks, bg_processes, bg_subagents, goal, duration, prompt,
  idle, focus, yolo, stash, battery, title
  Missing/null -> all enabled (backward compat). Unknown keys ignored.
  Title gated via right-align; stash/battery also gated.

- Wide bar (≥76 cols) respects fields, narrow/medium filtered,
  overflow trim preserved. Battery also respects display.battery.

No private data; mock data in tests.

Test: pytest tests/cli/test_cli_status_bar.py etc. 68 passed,
check-windows-footguns clean.
2026-08-29 18:34:51 -07:00
Cheri Wen
4bc7e624d6 feat(cli): show prompt cache hit rate in status bar
Add a ◎ XX% indicator to the CLI status bar showing the prompt cache
hit rate (cache_read / prompt_tokens). This helps users monitor how
effectively their provider's prompt caching is working.

Features:
- Color-coded: green (≥70%), yellow (40-70%), red (<40%)
- Adaptive precision: integer on narrow terminals, one decimal on wide
- Only shown when cache data is available (provider supports it)
- Compatible with OpenAI, Anthropic, DeepSeek, xiaomi, and other
  providers that return prompt_tokens_details.cached_tokens

Tests: 6 new test cases, 43/43 passing
2026-08-29 18:34:51 -07:00
liuhao1024
fb786d2f5b feat(cli): add display.status_bar.fields config for customizing status bar
Allow users to control which fields appear in the interactive CLI status
bar via display.status_bar.fields in config.yaml.

Available fields: model, context_pct, context_detail, compressions,
bg_tasks, bg_processes, duration, prompt_elapsed, yolo, total_tokens.

When the list is empty (default), all fields are shown as before.
The field order is fixed (model always first); the config controls
visibility only. Narrow terminals (<76 cols) automatically drop
context_detail regardless of config.

total_tokens is opt-in only (not shown by default) to avoid width
overflow in the prompt_toolkit fragment renderer.

Closes #41909
2026-08-29 18:34:51 -07:00
Teknium
74a95a3ddf feat: /btw now answers side questions with conversation context; /background renamed to /bg
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.

/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.

Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
2026-08-29 07:25:17 -07:00
Brooklyn Nicholson
9f90cd438c fix(cli): keep pet kitty flush off the import-time prompt_toolkit path
ColorDepth is lazy so MagicMock prompt_toolkit stubs can still import
cli. The redraw/resize kitty re-queue no-ops when the pet pane was
never initialized.
2026-08-28 23:38:59 -05:00
Brooklyn Nicholson
fac3c62334 feat(cli): render Ghostty-level pets in the interactive pane
Reuse kitty Unicode placeholders plus after_render write_raw so
prompt_toolkit's screen-diff can host the same crisp sprite as the TUI.
Re-queue the transmit after Ctrl+L / resize so the image comes back.

Co-authored-by: Sam Foreman <saforem2@gmail.com>
2026-08-28 23:38:59 -05:00
Teknium
a5c7eed5f3 feat(cli/tui): -q now seeds a live interactive session; prompts submit literally
On a real TTY, `hermes chat -q "…"` (and `--tui -q`) now starts a normal
interactive session with the prompt submitted literally as the first turn —
no slash-command routing, no '!' shell dispatch, no $(...) interpolation,
no file-drop rewriting — matching how other coding agents handle seeded
launches (Omarchy prompted agent terminals, basecamp/omarchy#8705).

Legacy answer-and-exit is preserved everywhere automation depends on it:
- new `hermes chat --oneshot` flag (distinct dest from top-level -z)
- -Q/--quiet machine-readable contract
- any non-TTY stdio (kanban workers, cron, pipes, A2A)
- top-level `hermes -z` unchanged

CLI: seeded prompt rides a _SeededQueryMessage sentinel through
process_loop, which skips the slash/!/file-drop dispatchers for that one
message. TUI: STARTUP_QUERY submits via a new literal path (submitLiteral)
that bypasses dispatchSubmission and the input.detect_drop rewrite.
2026-08-28 05:15:06 -07:00
fangliquanflq
f5200a4c10 fix(config): declare shared Docker container key 2026-08-25 04:00:27 -07:00
fangliquanflq
7a67bd07a7 feat(docker): support shared container identities 2026-08-25 04:00:27 -07:00
Teknium
a75ea37dc5 feat: browser snapshots drop LLM summarization — truncate-and-store like web_extract; auxiliary.web_extract slot removed
web_extract stopped using an auxiliary LLM long ago (deterministic
truncate-and-store), but browser snapshots still routed oversized
accessibility trees through the auxiliary web_extract model, keeping a
dead-looking aux slot alive across every config/picker surface.

- tools/browser_tool.py: remove _extract_relevant_content and
  _get_extraction_model; oversized snapshots always truncate at line
  boundaries, store the full tree to cache/web, and append a read_file
  pointer (element refs beyond the cut live in the file)
- tools/browser_camofox.py: same — no LLM path
- Remove auxiliary.web_extract slot: config_defaults (removal note, same
  pattern as session_search/PR #27590), cli.py defaults + env bridge,
  gateway/run.py bridged keys, hermes config display, hermes model picker,
  dashboard REST slots, desktop + web AUX_TASKS, i18n labels (en/zh/
  zh-hant/ja/ar)
- Docs: env-vars, configuration, fallback-providers, browser + zh-Hans
  mirrors (web-search zh-Hans was stale on the old LLM pipeline — synced
  to truncate-and-store truth)
- Tests updated: aux bridge uses approval slot, browser tests assert the
  LLM path is gone and stored files are secret-redacted
2026-08-24 20:11:18 -07:00
Teknium
1a95d0d58e Merge branch 'pr-81234' into salv/81234-retry-carrier 2026-08-24 03:15:07 -07:00
Teknium
12395e57b4 feat: /review command — independent reviewer subagent on every surface
/review takes the last 10 chat messages plus optional instructions,
spawns a full-privilege background subagent (the async delegation
rail) that investigates the referenced work (PR, code, docs), and its
complete review re-enters the spawning session as a normal
async-delegation completion the primary agent can act on.

- agent/review_engine.py: shared engine (snapshot, briefing,
  auxiliary.review credential resolution, dispatch, note formatting)
- tools/delegate_tool.py: internal credentials_cfg per-call override
  (never model-facing) resolved through the same credential system as
  delegation.provider pins
- auxiliary.review config block (provider/model/base_url/api_key/
  api_mode); provider auto + empty model = inherit the main model
- Surfaces: CLI process_command, gateway run.py dispatch +
  slash_commands handler (binds the approval session key so the
  completion routes back), TUI/Desktop live dispatch in
  tui_gateway/server.py, CommandDef registry (+Slack /hermes-only cap)
- Docs: delegation.md section + slash-commands.md (both tables)
- Tests: 15 engine tests (sabotage-verified: credentials_cfg and
  dispatch tests fail without the fix), 4 gateway handler tests
  through the real async rail
2026-08-23 17:38:38 -07:00
Teknium
637716755c fix(cli): -Q stdout carries only the final response — no tool diffs, spinner lines, or reasoning
Widens the cherry-picked reasoning-callback fix to the whole leak class
(#93220):

- quiet branch also neutralizes tool_progress_callback,
  tool_start_callback, tool_complete_callback (inline diff rendering via
  render_edit_diff_with_delta was gated by NEITHER quiet_mode nor
  tool_progress_mode) and syncs agent.tool_progress_mode='off'.
- _should_emit_quiet_tool_messages() returns False under
  suppress_status_output: with callbacks neutralized, the quiet-mode
  KawaiiSpinner fallback printed '[tool]'/'[done]' lines into captured
  stdout. Also covers oneshot.py and background-review forks, which set
  the same flag and expect strict silence.

E2E (isolated HERMES_HOME, live model, write_file turn): base leaks
'┊ review diff' + full SVG source into stdout; head emits exactly the
final response. Regression tests pin the quiet-branch statements and the
gate (sabotage-verified).

Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
2026-08-23 17:01:31 -07:00
Chris Cowan
8e2e3202da fix: suppress reasoning display in quiet single-query mode
The -Q quiet single-query path suppresses stream_delta_callback and
tool_gen_callback to keep stdout machine-readable, but missed
reasoning_callback. When display.show_reasoning is on (the default),
the reasoning box leaks into stdout before the final response,
corrupting output for automation wrappers and third-party integrations
using --source tool.

Before:
  hermes chat -Q --source tool -q "Reply with exactly: PING_OK"
  ┌─ Reasoning ──────────────────────┐
  The user wants me to reply...
  PING_OK

After:
  PING_OK
2026-08-23 17:01:31 -07:00
Brooklyn Nicholson
165d1849e2 fix(approval): stop the CLI and ACP offering a scope the protected gate discards
The protected agent-instruction gate grants one operation and persists
nothing, but only the TUI/desktop and Runs transports were taught that.
The prompt_toolkit panel, the input() fallback, and the ACP editor menu
still rendered "Allow for session", so a user editing SOUL.md tapped it,
got re-prompted on the next write, and read the gate as broken.

Thread allow_session through prompt_dangerous_approval so a caller that
re-asks every time collapses every surface to once/deny, and cover the
producer-to-transport contract end to end.
2026-08-23 17:45:47 -05:00
Teknium
9e18197745 fix(cli): one-shot runs linger for notify_on_complete background processes so Bot Mode replies survive parent exit
A Bot Mode agent invoked by a handoff runs as a short-lived
`hermes -p <bot> chat -Q --query-file ...` process. When it dispatches
its reply via message_agent / bot_relay — spawned as
terminal(background=true, notify_on_complete=true) per the Bot Chat
protocol — the one-shot parent exits as soon as the turn ends. The
reply child writes to a stdout pipe owned by the dying parent and is
destroyed a few seconds later, so the handoff reply is silently lost
while the sender waits for a notification that can never come (#90879).

Fix (class-wide, not DM-specific): before the one-shot exit paths tear
down, the parent now lingers — bounded by the new
terminal.oneshot_completion_wait_seconds config (default 600s, 0
disables) — for every tracked background process spawned with
notify_on_complete=true. Plain background processes (servers, daemons,
watch-pattern monitors) carry no completion contract and are never
waited on.

- tools/process_registry.py: ProcessRegistry.wait_for_pending_completions()
  — bounded, interrupt-safe wait over pending notify_on_complete
  sessions; reconciles orphaned-pipe exits (#17327) each pass so a
  wedged reader cannot burn the full bound; KeyboardInterrupt aborts
  the linger without skipping the caller's durable teardown.
- cli.py: _finalize_single_query() lingers first, before the durable
  session flush / cleanup (covers -q and -Q, i.e. the DM recipient
  shape and bot_relay waiter spawns from one-shot agents).
- hermes_cli/oneshot.py: same linger before agent.close() (which
  kill_all()s the task's processes) on the -z path.
- hermes_cli/config_defaults.py: terminal.oneshot_completion_wait_seconds.

Tests: tests/tools/test_oneshot_completion_linger.py — unit coverage of
the wait semantics (no-op, completion, timeout, task filter, disable,
config fallback, reconcile path), exit-path ordering contracts, and a
real-process E2E: a short-lived python parent spawns a delivery child
through the real ProcessRegistry, lingers, exits, and the delivery
completes; sabotaging the linger makes the same E2E reproduce the
destroyed-delivery symptom.

Fixes #90879
2026-08-23 03:56:37 -07:00
poisdahl
fd41164861 fix(history): keep carrier rewinds race-safe after refresh 2026-08-22 17:30:35 +02:00
poisdahl
abf87e7248 Merge current main into composite-carrier fix 2026-08-21 15:56:45 +02:00
Teknium
a83c3915a3 fix(opencode): family-wide provider predicate + reserved tool-name aliases for custom opencode-* providers
Builds on @Lesnak1's #85619 (issue #85589):

- New opencode_provider_family() single-owner predicate in
  hermes_cli/models.py — resolves built-in AND custom family providers
  (opencode-go-bridge, OpenCode-Zen-Custom, ...) case-insensitively.
  Migrated all 8 inlined family checks (models.py x3, runtime_provider.py
  x4 from the salvaged commits) plus 4 sibling sites the PR missed:
  cli.py api_mode sync, agent_runtime_helpers.py double-/v1 guard,
  model_normalize.py flat-namespace strip, model_switch.py base_url
  normalization.
- Responses transport: alias OpenCode-reserved function names
  (web_search, search_files -> hermes_*) on the wire and map them back on
  dispatch — same pattern as the xAI web_search collision fix. Matches
  family providers and any base_url on opencode.ai. Fixes the HTTP 400
  'custom function name X is reserved' half of #85589.
- Tests: custom-provider routing assertions + 5 new transport alias tests.
2026-08-20 20:21:12 -07:00
Chris Fontes
5046282867 feat(config): resolve_turn_limit — first-class 'none'/'unlimited' for agent.max_turns
Previously agent.max_turns only accepted positive integers. Setting it to
'none', 'unlimited', or 0 — all natural ways to say 'no limit' — either
crashed int() or was silently skipped by `or` checks, falling back to 90.

This adds resolve_turn_limit() in hermes_cli/config.py as the single
normalization point. It accepts:

  - int/float → int(raw) (floats truncated)
  - numeric string ('120') → int(raw)
  - 'none'/'unlimited'/'infinite'/'∞'/'-1'/'0' (case-insensitive,
    whitespace-tolerant) → sys.maxsize sentinel
  - YAML None/null → default (90)
  - bool/list/dict/garbage → default (with debug log)

All config-reading sites (cli.py, gateway/run.py, cron/scheduler.py) now
call this instead of bare int(), so agent.max_turns: none in config.yaml
becomes a first-class supported spelling of 'unlimited'.

The sentinel (sys.maxsize) survives the str()→int() round-trip through
the HERMES_MAX_ITERATIONS env-var bridge in gateway/run.py and works in
every <, >=, remaining = max - used comparison without requiring call
sites to learn about a special value.

Includes 38 tests covering the full spelling table, the str→int env-var
round-trip, and sentinel properties.
2026-08-20 04:50:39 -07:00