Two product questions shared metrics could not answer: how long each surface
takes from launch to usable (so startup regressions show per release), and how
current, on which channel and on what class of machine installs run.
hermes.startup.latency {surface, latency_bucket} records once per process start:
cli (process creation -> first rendered prompt, or -q dispatch; Kanban workers
excluded), gateway_boot (-> GatewayRunner.start done), serve_boot (-> hermes
serve listening), and tui / desktop_attach reported by the clients through the
new shared_metrics.startup_latency RPC. The clients declare their surface
because a Desktop on a URL/cloud backend has no HERMES_DESKTOP there; env
detection is the fallback for older clients. In-process surfaces measure from
psutil's process create time, the earliest timestamp available, and hand the
runtime start to a daemon thread under the caller's context so no event loop
waits on it. Everything goes through _emit, so disabled profiles record nothing.
The install snapshot gains release_channel, version_age_bucket, behind_bucket,
ram_bucket, gpu_class and local_model_provider_used. All are read offline:
the installed commit's own date, the channel record / packaged channel / checkout
branch (never the remote URL or branch name), and the update check's existing
cache for this exact revision (never a network call). Rows counted before these
fields existed stay valid as a legacy field set.
A dispatcher SIGKILLed between _call_spawn_fn and _set_worker_pid leaves a
live worker on a run with worker_pid NULL. release_stale_claims only extends
an expired claim for a recorded live pid, so on TTL expiry it reclaimed the
card and spawned a second worker beside the first: double billing, double
side effects, and a board showing one clean completed run (the first
worker's kanban_complete is refused as stale). Main CI hit it in
test_dispatcher_sigkill_mid_tick_never_destroys_or_duplicates_cards.
The worker now records its own pid on its run before the first model call
(adopt_worker_pid, worker_registered event, host-local claims only) and
exits without working the card when its run was already reclaimed. The
reclaim UPDATE also compares worker_pid so a registration landing between
the stale-claim SELECT and the UPDATE keeps the claim.
Repro: temporary sleep between spawn and pid record + kill 0.2 s after the
spawned event + slow first model reply -> 4/4 red on main with the CI
signature, 8/8 green here.
Fixes#121556
Re-review findings on the foreground-process registry:
- Spawn vs exit races. A child spawned but not yet registered when the hard-exit kill ran, or a
command launched after the kill took its snapshot, survived under init. The hard-exit kill now
raises a one-way exit fence and waits (bounded) for spawns already past it to register; a spawn
refuses once the fence is up, and one that registers after a timed-out wait kills itself.
- The immediate kill fell back to proc.kill() for every handle. On Modal/Daytona/Vercel that is a
blocking SDK cancel (an 8s cancel made _hard_exit take 8s). Popen handles are still killed
inline (killpg, never blocks); every other handle's kill runs on a daemon thread under one
shared 0.5s deadline.
- The registry lock was a plain Lock: a signal landing on a thread that held it deadlocked the
exit. It is now an RLock (via a Condition) and the hard-exit path only takes it with a timeout,
falling back to a lock-free copy.
- The kanban worker's SIGALRM deadman os._exit()ed without the kill; it now goes through it.
The SIGTERM handler arms a 1s os._exit timer, then runs _shutdown_sessions: a flush of up to
5s, then _stop_turns_before_exit, whose kill was the graceful TERM, wait 1s, KILL. A command
that ignores SIGTERM was still alive when the timer fired, and os._exit left it reparented to
init (live: `trap '' TERM; sleep 3600` survived a SIGTERM to `python -m tui_gateway.entry`).
- kill_live_foreground_processes(now=True): SIGKILL each in-flight foreground tree at once,
no TERM grace, no wait (BaseEnvironment._force_kill_process; LocalEnvironment kills the
recorded process group, never our own).
- The grace timer's exit (entry._hard_exit) runs it before os._exit.
- _stop_turns_before_exit SIGKILLs whatever is still alive halfway through its settle budget
(it ignored the interrupt's TERM), so the tool call still ends with a result the teardown
persists instead of a dangling tool_call in state.db.
- The other hard exits that skip cleanup do the same before os._exit: the serve parent-death
watchdog, the CLI exit watchdog, the kanban worker's SIGTERM path, and the messaging
gateway's shutdown and loop-liveness watchdogs.
- Deflake test_shutdown_mid_tool_kills_the_command_and_keeps_its_result: the 0.5s settle
budget was too tight under -n 40 (1 red in 9 runs); the join returns when the turn ends.
PluginManager._cli_ref was set only by the interactive run loop
(_tui_init_run_state), so a one-shot query never bound it and
PluginContext.dispatch_tool injected no parent_agent into plugin tools. Bind it
once in _run_single_query_mode before the turn; the quiet, stream_json and
kanban-goal branches all pass through that line, and the interactive-seed
branch calls cli.run(), which already binds it.
Fixes#67597
credit: @webtecnica #67610 (slim redo at the current seam; #67611 and #67633 are later duplicates)
Mechanical, behaviour-neutral extraction along the existing hermes_cli/cli_*.py
pattern. Six clusters of module-level helpers leave cli.py:
- cli_config_load.py prefill messages, reasoning/service-tier parsing, terminal
env mirroring, CLI defaults + YAML merge, logging bootstrap
- cli_render.py reasoning-tag stripping, ANSI/skin colours, light-mode
detection, markdown/final rendering, output-history
recording, _cprint, ChatConsole, compact banner, panel wrap
- cli_terminal_input.py file drops/attachments, bracketed-paste patch, extended
Enter keys, CPR guards, TUI input height, query images
- cli_shutdown.py process session-id sync, exit watchdog, cleanup steps,
session-finalize notifications, one-shot finalize
- cli_single_query.py kanban goal loops, exit-code mapping, quiet -q runner,
image routing, signal handlers, single-query mode
- cli_auto_maintenance.py state-db / checkpoint startup maintenance
Every moved body is AST-identical to the base copy modulo two mechanical
edits: cli-level names are late-bound with a call-time `from cli import ...`
(the cli_init_mixin pattern) and mutable cli state is read through `_cli().NAME`
(the gateway_service_unit `_gw()` pattern), so every monkeypatch seam on the
`cli` facade still intercepts moved code and no read snapshots stale state.
Mutable module state and every `global`-writing function stay in cli.py; the
facade re-exports the moved names in one import block.
test_bracketed_paste_timeout AST-loads the paste helper from its new home.
cli.py 3959 -> 1820 lines.