Commit Graph

7 Commits

Author SHA1 Message Date
teknium1
2b3e1c3995 feat(metrics): startup latency and install version-lag/hardware fields
Two product questions shared metrics could not answer: how long each surface
takes from launch to usable (so startup regressions show per release), and how
current, on which channel and on what class of machine installs run.

hermes.startup.latency {surface, latency_bucket} records once per process start:
cli (process creation -> first rendered prompt, or -q dispatch; Kanban workers
excluded), gateway_boot (-> GatewayRunner.start done), serve_boot (-> hermes
serve listening), and tui / desktop_attach reported by the clients through the
new shared_metrics.startup_latency RPC. The clients declare their surface
because a Desktop on a URL/cloud backend has no HERMES_DESKTOP there; env
detection is the fallback for older clients. In-process surfaces measure from
psutil's process create time, the earliest timestamp available, and hand the
runtime start to a daemon thread under the caller's context so no event loop
waits on it. Everything goes through _emit, so disabled profiles record nothing.

The install snapshot gains release_channel, version_age_bucket, behind_bucket,
ram_bucket, gpu_class and local_model_provider_used. All are read offline:
the installed commit's own date, the channel record / packaged channel / checkout
branch (never the remote URL or branch name), and the update check's existing
cache for this exact revision (never a network call). Rows counted before these
fields existed stay valid as a legacy field set.
2026-09-28 12:43:03 -07:00
funky-xamarin
25b6a9f010 fix(kanban): preserve credential startup failure exit codes 2026-09-28 03:37:09 -07:00
teknium1
63e44332f5 fix(kanban): a worker the dispatcher never recorded registers itself instead of being run twice
A dispatcher SIGKILLed between _call_spawn_fn and _set_worker_pid leaves a
live worker on a run with worker_pid NULL. release_stale_claims only extends
an expired claim for a recorded live pid, so on TTL expiry it reclaimed the
card and spawned a second worker beside the first: double billing, double
side effects, and a board showing one clean completed run (the first
worker's kanban_complete is refused as stale). Main CI hit it in
test_dispatcher_sigkill_mid_tick_never_destroys_or_duplicates_cards.

The worker now records its own pid on its run before the first model call
(adopt_worker_pid, worker_registered event, host-local claims only) and
exits without working the card when its run was already reclaimed. The
reclaim UPDATE also compares worker_pid so a registration landing between
the stale-claim SELECT and the UPDATE keeps the claim.

Repro: temporary sleep between spawn and pid record + kill 0.2 s after the
spawned event + slow first model reply -> 4/4 red on main with the CI
signature, 8/8 green here.

Fixes #121556
2026-09-26 11:16:45 -07:00
teknium1
3a001a4d7f fix(tui-gateway): the hard-exit kill owns every foreground spawn and never blocks
Re-review findings on the foreground-process registry:

- Spawn vs exit races. A child spawned but not yet registered when the hard-exit kill ran, or a
  command launched after the kill took its snapshot, survived under init. The hard-exit kill now
  raises a one-way exit fence and waits (bounded) for spawns already past it to register; a spawn
  refuses once the fence is up, and one that registers after a timed-out wait kills itself.
- The immediate kill fell back to proc.kill() for every handle. On Modal/Daytona/Vercel that is a
  blocking SDK cancel (an 8s cancel made _hard_exit take 8s). Popen handles are still killed
  inline (killpg, never blocks); every other handle's kill runs on a daemon thread under one
  shared 0.5s deadline.
- The registry lock was a plain Lock: a signal landing on a thread that held it deadlocked the
  exit. It is now an RLock (via a Condition) and the hard-exit path only takes it with a timeout,
  falling back to a lock-free copy.
- The kanban worker's SIGALRM deadman os._exit()ed without the kill; it now goes through it.
2026-09-23 17:09:54 -07:00
teknium1
17a137d8a3 fix(tui-gateway): a SIGTERM-ignoring command no longer survives the gateway's SIGTERM exit
The SIGTERM handler arms a 1s os._exit timer, then runs _shutdown_sessions: a flush of up to
5s, then _stop_turns_before_exit, whose kill was the graceful TERM, wait 1s, KILL. A command
that ignores SIGTERM was still alive when the timer fired, and os._exit left it reparented to
init (live: `trap '' TERM; sleep 3600` survived a SIGTERM to `python -m tui_gateway.entry`).

- kill_live_foreground_processes(now=True): SIGKILL each in-flight foreground tree at once,
  no TERM grace, no wait (BaseEnvironment._force_kill_process; LocalEnvironment kills the
  recorded process group, never our own).
- The grace timer's exit (entry._hard_exit) runs it before os._exit.
- _stop_turns_before_exit SIGKILLs whatever is still alive halfway through its settle budget
  (it ignored the interrupt's TERM), so the tool call still ends with a result the teardown
  persists instead of a dangling tool_call in state.db.
- The other hard exits that skip cleanup do the same before os._exit: the serve parent-death
  watchdog, the CLI exit watchdog, the kanban worker's SIGTERM path, and the messaging
  gateway's shutdown and loop-liveness watchdogs.
- Deflake test_shutdown_mid_tool_kills_the_command_and_keeps_its_result: the 0.5s settle
  budget was too tight under -n 40 (1 red in 9 runs); the join returns when the turn ends.
2026-09-23 17:09:54 -07:00
teknium1
069d6a1f9e fix: give plugins the CLI reference in hermes chat -q/-Q too
PluginManager._cli_ref was set only by the interactive run loop
(_tui_init_run_state), so a one-shot query never bound it and
PluginContext.dispatch_tool injected no parent_agent into plugin tools. Bind it
once in _run_single_query_mode before the turn; the quiet, stream_json and
kanban-goal branches all pass through that line, and the interactive-seed
branch calls cli.run(), which already binds it.

Fixes #67597
credit: @webtecnica #67610 (slim redo at the current seam; #67611 and #67633 are later duplicates)
2026-09-21 23:38:16 -07:00
teknium1
2ca53fc386 refactor(cli): move cli.py module-level helper clusters into topical siblings; cli.py lands under 2,000 lines (#116911)
Mechanical, behaviour-neutral extraction along the existing hermes_cli/cli_*.py
pattern. Six clusters of module-level helpers leave cli.py:

- cli_config_load.py     prefill messages, reasoning/service-tier parsing, terminal
                         env mirroring, CLI defaults + YAML merge, logging bootstrap
- cli_render.py          reasoning-tag stripping, ANSI/skin colours, light-mode
                         detection, markdown/final rendering, output-history
                         recording, _cprint, ChatConsole, compact banner, panel wrap
- cli_terminal_input.py  file drops/attachments, bracketed-paste patch, extended
                         Enter keys, CPR guards, TUI input height, query images
- cli_shutdown.py        process session-id sync, exit watchdog, cleanup steps,
                         session-finalize notifications, one-shot finalize
- cli_single_query.py    kanban goal loops, exit-code mapping, quiet -q runner,
                         image routing, signal handlers, single-query mode
- cli_auto_maintenance.py state-db / checkpoint startup maintenance

Every moved body is AST-identical to the base copy modulo two mechanical
edits: cli-level names are late-bound with a call-time `from cli import ...`
(the cli_init_mixin pattern) and mutable cli state is read through `_cli().NAME`
(the gateway_service_unit `_gw()` pattern), so every monkeypatch seam on the
`cli` facade still intercepts moved code and no read snapshots stale state.
Mutable module state and every `global`-writing function stay in cli.py; the
facade re-exports the moved names in one import block.

test_bracketed_paste_timeout AST-loads the paste helper from its new home.
cli.py 3959 -> 1820 lines.
2026-09-20 20:42:17 -07:00