Review follow-ups on the stale-code tick yield gate:
- `_gate_mocks` patches `_current_gateway_code_sha` with `raising=False` so the
base tree (where the symbol does not exist yet) fails only on the two new
invariants instead of erroring the whole module out with AttributeError.
- Note at the pid comparison that `get_running_pid(pid_path=None)` can fall back
to `get_runtime_status_running_pid()`, which derives the pid from the same
record — the equality is tautological in that branch and the real proof there
is `runtime_status_is_stale` + `code_sha`.
- Docstring says what the record actually is: `gateway_state.json` is per-HOME
and last-writer-wins, the pid equality binds it to the lock holder, and the
`--replace` takeover window fails open by design.
Review follow-up for #117505: delivery_ledger, api_server_runs and
async_delegation reconciled a live owner to dead/unknown on the same 1 s
same-host fingerprint drift that cron/executions had just stopped doing.
One comparator (gateway.status.start_time_fingerprints_match) replaces the
private cron tolerance and the three exact-equality sites.
_owner_is_live() required the recomputed start-time fingerprint to be exactly
equal to the one recorded at claim time. On hosts where same-host readings
drift by ~1s (macOS kern.boottime adjustment), every long-running execution
was reconciled to unknown while its owner process was still alive and healthy.
A mismatch is now treated as inconclusive within a small tolerance (both
platform fingerprint scales are x100, so 200 means 2s), and an unreadable
reading is fail-safe, matching the function's existing posture that inability
to prove death must not rewrite durable state.
`deliver: bot-chat:<other profile>` spawned the destination's agent turn with the sending
gateway's whole environment: its `.env` settings, bridged `TERMINAL_*` policy, platform
authorization gates and provider credentials. The lane called
`strip_launch_profile_env(env)` with no target, so the strip resolved against the ambient home
override — which is the SENDER's home, never the destination's. On an ordinary root-profile
gateway (`hermes gateway run`, no `-p`) `_is_routed_home` is then false and the strip is a
complete no-op, including the #113270 gate strip that lives after its early return.
This is the only cron child built for a profile other than the one whose tick spawned it; the
worker lane (`scheduler.py`) targets its own home, so its no-target call is correct. Build this
one through `served_profile_child_env(target_home=home, inherit_credentials=True)` — the helper
`kanban_db_dispatch` and `web_server_gateway` already use for cross-profile spawns: it strips the
launch residue against the real target, scrubs credentials the launch process was given by
systemd/Compose/the shell (which no name-based strip can see), points TMPDIR at the destination's
scratch, and overlays the destination's own secrets, as a standalone `hermes -p <profile>` has.
A failure to build that environment (an unreadable target home under per-user 0700, a broken
secret source) is reported as a refusal string like every other failure in this lane rather than
raised: `_deliver_result`'s fan-out does not catch, unlike the deferred drain.
Regressions drive the real `_deliver_to_bot_chat`; removing the fix fails the two leak witnesses
(`HERMES_MODEL leaked from the launch profile`, and the firing profile's gate reaching another
profile's turn under an active override) and leaves the four guard tests green.
Fixes#117220
(cherry picked from commit 6cc81d7ddad8c9793f21fdda7c2e260c3ee1ca44)
The sibling module already reaches the scheduler via the late-bound _sched
reference, so the public alias was a shim (shape gate: no re-export for
internal moves). Tests kept: answer survives a huge prompt, silent run falls
through to an older archive, headingless (script-mode) archive injects whole.
scheduler_prompt reached into the private scheduler helper; a rename would
have broken it silently. Add a public module-level alias and call it
late-bound instead (review feedback).
Follow-up to the review on #117302:
- _archive_answer: rpartition — the assembled prompt half can itself carry
a literal "## Response" heading (a skill documenting its response format,
an injected previous answer quoting it), so the writer's boundary is the
LAST occurrence; an early split re-injects the prompt noise this
extraction exists to drop.
- silence check: reuse _sched._is_cron_silence_response instead of the
literal [SILENT] comparison, so the bracketless SILENT / NO_REPLY forms
the delivery lane itself suppresses (#46917, #51438) also fall through
to an older archive.
- _SELF_CONTEXT_INTRO: "most recent non-silent output" — a silent latest
run now deliberately resolves to an older archive.
context_from head-truncation amputated the ## Response section of long
archives (skill-bearing prompts routinely exceed the 8000-char budget),
so self-continuity silently became a no-op: the job was handed its own
old prompt plus a truncation marker. Reduce each agent-run archive to
its answer (text after ## Response), fall through to an older archive
on [SILENT]/blank responses, inject nothing when no archive has a
usable answer, and clip oversized answers head+tail with an explicit
chars-omitted marker since conclusions sit at the end. Script-mode
archives without the heading keep injecting whole.
Fixes#117290
A restart-safe cron worker runs outside the gateway, queues its final send in
cron/deliveries.db and polls that row once a second. The gateway drained the
queue only as a housekeeping chore at the 60 s tick, so a reminder whose turn
ended in 8 s reached the user a minute later (the issue's 60.86 s silence).
The housekeeping thread now watches the queue file (and its WAL) between
ticks and drains as soon as a stamp moved; the tick's drain stays as the
fallback. The watch never resets itself after draining: the drain's own status
write costs one empty pass, a worker that enqueued during the drain still
wakes it, and an empty pass writes nothing, so the stamp settles.
Fixes#117307
get_wedged_job_ids called cron.jobs.get_job (a full jobs.json load) per in-flight
job, and the restart drain invokes it on the event loop every 0.1 s for the whole
wait: 3 running jobs x 600 s = ~18,000 parses. Resolve each run's allowance once
(cached per job_id while it is registered, dropped on release) from a single
load_jobs() call.
`_wedged_agent_count` only ever looked at chat agents, so a cron run that would
never finish (a no-agent job whose delivery hung on a dead transport) was
structurally un-skippable: `hermes update` sat in "draining" for the full
`agent.restart_after_turn_timeout` printing "0 wedged and excluded" while the
script had finished 8 seconds in.
Cron has no per-turn activity clock, but the scheduler already defines when an
in-flight claim can no longer be making progress: `sweep_stale_inflight`'s
`max(2 * interval, cron.inflight_max_minutes)` allowance. It cannot release a
claim whose worker thread is still alive, so expose that judgement as
`cron.scheduler.get_wedged_job_ids()` and let the drain count those runs as
wedged (restart is their remedy), the way it already treats idle chat turns.
`_describe_active_work` marks the cron unit `wedged` so the status line names it.
Fixes#115469 (Defect B; Defect A is the bounded standalone send this branch
stacks on).
The relay's delivery child shared one 600s deadline between the target's turn
and the one-shot exit linger, whose own budget is the same 600s — so a turn that
answered in seconds and then handed off to a teammate was killed mid-linger,
reported to the sender as delivery_timeout (auto-retried: the turn ran twice),
its reply lost, and the handoff delivery the linger protected destroyed.
The -Q child's turn report (#113608) now carries the answer the run will print,
rewritten when a follow-up turn displaces it, and the poll loop that books a
child from that report moves next to the contract as
quiet_single_query.run_reported_turn. The cron lane keeps its policy (book after
a 2s exit grace); the relay waits for exit under the cap as before — a teammate's
reply during the linger may still become the printed answer — and at the cap
books a reported child from its latest report and leaves it to finish. Only a
turn that never ends is a timeout. The report is 0600 from creation now that it
carries the answer.
Fixes#114980
The cron bot-chat delivery child inherited the scheduler's cwd; when that
directory had been removed (a kanban worker whose scratch workspace was reaped
by completion cleanup) the child died in `hermes_cli/_startup_fast.py::
ensure_project_root_on_path` — a relative `sys.path` entry goes through
`os.getcwd()` inside `realpath`, which raises FileNotFoundError — before it
could parse argv, and the finished job was booked `delivery_failed`.
- `_run_bot_chat_turn` pins the child's `cwd` to the target home the lane has
already verified exists (`env["HERMES_HOME"]`).
- `ensure_project_root_on_path` resolves entries through a `realpath` that
tolerates a gone cwd, so any `hermes` invocation from a dead directory still
starts (the second half of #102941).
Salvaged from #102967 (@686f6c61), resolved onto the report-driven Popen lane
(#113608); tests drive the CLI entry point and the delivery spawn seam from a
deleted cwd.
An unpinned cron job used to snapshot the global provider/model at creation and treat that
snapshot as its effective pin (#44585), so `hermes model` / `/model` never moved the fleet and
`hermes cron resnap` existed to catch jobs up. New ruling: jobs run on whatever the main agent
model is when they fire. Resolution is per-job pin > cron.model / cron.model_provider (the cron
fleet default) > model.default.
`pinned` replaces the implicit snapshot with an explicit lock: create/update with pinned=true
writes the CURRENT main provider+model onto the job as an ordinary per-job pin; pinned=false
releases both. The cronjob tool exposes it (schema: only when the user asks; it can only lock
the main model, never point spend at a different one) and reports `pinned` per job; the CLI
gets `--pin` / `--unpin`. Legacy records that still carry *_snapshot keys follow the main model.
Removed with the snapshot: `hermes cron resnap`, the tool's resnap action + `all` param, the
"N unpinned jobs keep running on ..." notice in `hermes model` / `hermes config set` / the
dashboard model assignment, and the Desktop cron-model-impact card (setMainModelAssignment
keeps the expensive-model confirm flow in store/model-assignment.ts).
Live A/B (real store + run_job against a temp HERMES_HOME): main-model X -> Y, unpinned job
fires on X before, Y after; pinned job stays on X; unpin -> Y; legacy snapshot record -> Y.
The bound reads the script timeout, which falls through to load_config() when
HERMES_CRON_SCRIPT_TIMEOUT is unset. Computing it eagerly on every reap put a
config load back on the idle gateway tick (CI: tests/cron/test_idle_tick_config_skip.py).
Resolve it on the first live-owned claimed/running row instead; an idle tick with no
such rows stays config-free.
The carrier used a hardcoded 7200 s wall-clock ceiling. HERMES_CRON_TIMEOUT is
an *inactivity* limit (cron/jobs.py, _run_agent_with_watchdog), so a healthy
agent job in the external-worker topology can legitimately run past any fixed
wall clock; marking its row `unknown` releases the gateway's running guard and
the next tick double-dispatches while the first worker is still running (its
later finish_execution is fenced out). The constant also ignored
HERMES_CRON_TIMEOUT=0 (unlimited) and HERMES_CRON_SCRIPT_TIMEOUT > 7200.
Mirror cron/jobs.py::_oneshot_run_claim_ttl_seconds instead:
stale_after = max(3 × HERMES_CRON_TIMEOUT, script timeout, 7200), reusing the
existing accessors (_cron_inactivity_seconds, _get_script_timeout). When the
inactivity timeout is 0/unlimited or not a finite positive number there is no
bound to derive, so live owners are skipped entirely (fail closed — today's
behaviour). Drop the `except (ValueError, TypeError): pass` swallow (rows are
always aware ISO from hermes_time.now), fix the stray whitespace in the
SELECT, record a distinct `error` reason for the wedged-owner case, and keep
the comment honest: the deadlocked worker process is not terminated, and
in-process rows (process_id == _PROCESS_ID) remain out of scope.
`_last_dead_owner_reap_at` was a process-global float, so under
`gateway.multiplex_profiles` the first profile ticked in a cycle consumed
the 300 s reap window for every other profile: a stale claim in profile B
was not noticed until profile A's window elapsed. Key the throttle by the
resolved profile home so each profile gets its own dead-owner scan.
Carried from #115696 (whyyagswhy, first submitter for #115692); only the
throttle hunk and `test_reap_throttle_is_profile_scoped` are taken — the
warn-only stuck-claim helpers are superseded by the live-owner stale-claim
guard in this stack.
A claimed/running execution whose owner process is alive but permanently
deadlocked (e.g. futex_wait after a route/proxy flip) passes _owner_is_live()
and is never reclaimed — the execution row stays 'running' forever and the
job rejects every subsequent fire with 'Job is already running'.
recover_interrupted_executions() now also checks the wall clock: a
claimed/running row older than STALE_RUNNING_CLAIM_TIMEOUT_SECONDS (7200 s,
covering the 3600 s script timeout + 600 s inactivity watchdog + margin)
is marked unknown even though the owner PID is alive.
Tests: stale claim recovered, recent claim NOT recovered.
The recursion guard for the timeout marker matched 'DELIVERY DEGRADED' in the
payload, so a job whose own output contained that phrase would silently lose its
notice. The deferred record now carries an explicit degraded flag and the guard
reads it. The marker's id is a fresh sha256 digest of '<key>:degraded' instead of
'<key>-degraded': deferred ids double as live-owner delivery ids, which
tools.bot_live_delivery._delivery_id requires to be 32-64 hex characters.
test_bot_chat_pending documents the deliberate extra turn: a timeout drains one
short marker turn, whose own timeout queues nothing more.
The 10-minute per-key complaint cache (_fire_fence_complaints /
_fence_timeout_log) only rate-limited the log line. A contender that fails
closed on the fire fence every tick is a real contention class (cross-process
`hermes cron run`, a second gateway, or the forced-release sweep evicting the
running-set entry mid-delivery); demoting repeats to DEBUG hides it for ten
minutes instead of fixing it, and the helper was a module-level appendage on
the jobs facade. cron/jobs.py is back to main; the test that poked the private
cache goes with it.
A bot-chat delivery child that times out is killed mid-turn. The killed
child cannot complete, but its turn may have persisted the query into the
session before the kill — so replaying the full payload risks a duplicate,
while staying silent loses the alert entirely (the message was only saved
to job output; the Bot Chat consumer never saw it).
- Timeout on the CLI fallback lane queues a short degraded-delivery marker
through the existing deferred lane: it references `hermes cron runs`
and carries a 280-char excerpt instead of the full payload, so a turn
that persisted the query before the kill cannot duplicate content.
Recursion-guarded: a marker that itself times out is never re-marked.
- `_run_bot_chat_turn`: a turn report landing in the kill window books
the delivery instead of raising TimeoutExpired (delivered turns were
misreported as timeouts).
- `jobs.py`: fire-fence timeout logs ONE ERROR per episode (600s
interval; repeats log DEBUG, a successful acquire clears the episode).
One delivery turn holding the fence for minutes previously produced an
8-ERROR storm plus restore warnings from every failed tick.
The bot-chat delivery child is guaranteed UTF-8 on Windows (hermes_cli
reconfigures its streams via hermes_bootstrap even under
PYTHONIOENCODING=cp1252), but the gateway parent is not started in UTF-8
mode — its env overlay sets only PYTHONIOENCODING, which rebinds the
parent's own streams but does not change the locale codec text=True picks.
The delivery pipes were therefore decoded with the ANSI code page (cp1252):
an accented reply came back mojibake'd, and bytes undefined in cp1252
killed the reader thread, losing the captured reply and the failure
diagnostic while the delivery still booked as delivered (#115894).
Pin encoding='utf-8' on the win32 branch only — on POSIX the child keeps
the locale codec, so the locale default stays correct there (#66566) —
mirroring the win32 branch of the no_agent script lane (#45099).
A stray non-UTF-8 byte in a cron child's stdout/stderr (e.g. a grandchild
sharing the pipe interleaving a partial multi-byte write) raised
UnicodeDecodeError in the decode step and failed a run that in fact
completed: the no_agent script lane died inside communicate() with the
output discarded, and the bot-chat delivery lane lost both the captured
reply and the failure diagnostic when the drain thread died (#105582).
Decode lossily on the POSIX side of both lanes: keep the platform-default
(locale) encoding — gating encoding= to win32 was deliberate (#66566) —
but relax errors= to 'replace', matching the lossy UTF-8 decode the win32
branch of the script lane already applies (#45099).
Wrapping the coroutine in `asyncio.wait_for` outside `asyncio.run` left the
running-loop fallback (which closes the unstarted coroutine and retries in a
thread) with a never-awaited wait_for wrapper and an unbounded inner send.
Awaiting `wait_for` inside `_send` bounds every runner of the coroutine.
The invariant tests now drive the production entry point `_deliver_result`
with no live adapters (the standalone lane) instead of `_standalone_send`
directly, and the user-visible knob `cron.standalone_send_timeout_seconds`
is documented (#115469).
The standalone fallback ran _send_to_platform under a bare asyncio.run,
while the gateway-loop dispatch inside it awaits its future with a
deliberate no-timeout shield. The shield's comment assumes an outer
_run_async bound that never exists on this lane, so a mid-reconnect
transport pinned the run (and the restart drain behind it) indefinitely
(#115469). Bound the wait with asyncio.wait_for — default 60s, matching
the live lane's future.result(timeout=60); configurable via
cron.standalone_send_timeout_seconds.
_run_external_worker_payload wrote the acknowledgement JSON directly
into ack_path after creating it with O_CREAT|O_EXCL. That made an
empty 0-byte file visible to the scheduler's exists()-then-read
polling loop (every 50ms) before the JSON body was flushed, so the
poller would occasionally read an empty file and crash with
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
(observed 3x on 2026-09-09 in production logs).
Fix: write the ack payload to a per-pid temp file in the same
directory, then os.replace() it into place. Readers now only ever
see ack_path fully absent or fully written.
_live_send_text rebuilt DeliveryRouter over the plain target_adapters dict
and re-resolved, discarding the SharedRouteAdapters-authorized transport from
_resolve_target_transport. Under satellite config (own platforms.<p> block
disabled) the second resolution yields None and delivery falls through to the
credentialless standalone lane and fails.
DeliveryRouter._deliver_to_platform takes an optional transport; the live
lane passes t.transport. All router pipeline behavior (oversize cap, relay
home stamping, thread routing) is unchanged; other callers still resolve.
(cherry picked from commit 9a2a67647cd823fe9120891d863a6fc6981fec71)
The blocked_config reason for a missing provider credential now carries
"[profile '<name>', HERMES_HOME <path>]" — the home the scheduler actually
read auth.json/.env from — under the ticker's profile scope, so a
multiplexed satellite profile reports its own home, not the gateway's
launch home.
Why: #116213 reports an openai-codex cron job blocked with "No Codex
credentials stored" while an interactive session under "the same"
HERMES_HOME resolves the credential. A 5-shape x 5-scope live matrix
(singleton, expired+refreshable, pool-only, ~/.codex only, none; root,
named profile, root-only auth, multiplex default/named) on origin/main
and on the reporter's build 345cd2b0 shows interactive and cron
preflight agree in every cell — both call the same
resolve_runtime_provider ladder and read the same store. The remaining
explanation is a scheduler process reading a different home than the
shell (Docker HOME vs HERMES_HOME, a service unit without the shell's
env, a satellite profile), which the bare verdict could not reveal.
Naming the store the verdict judged makes that mismatch visible in the
one alert the user receives.
Part of #116213
update_job recomputes next_run_at from the new schedule but left quota_hold_until behind, so
the marker kept shielding a record that was no longer parked where it said. Clear it with the
schedule edit; the next fire re-parks (with a fresh notice) if the window is still closed.
cron.md claimed every 429 with a retry-after hint holds the job; only the provider-resolve
usage probe (a rate-limited AuthError) is classified. Say so.
A quota-exhausted subscription provider answers every request with a 429 and an explicit
`retry after <N>s` (Codex: ~33 h). The scheduler walked the fallback chain, failed, and
re-fired the job on its normal cadence into the same closed window — one guaranteed failure
and one delivered alert per tick until the window reopened (#89376: ~460 failed runs and
alerts across four profiles in one exhaustion).
Why: the retry-after was available at the failure site (a rate-limited AuthError in the
RuntimeError's cause chain) but nothing consumed it, and the stale-error re-arm (#62002)
would have pulled any manually parked next_run_at back after one cadence anyway.
What: cron/quota_hold.py mirrors cron/unreachable_retry.py in the opposite direction —
run_job flags the failure with the provider's remaining seconds (read only from a
rate-limited AuthError, never from arbitrary text), the one failure alert says the job is
held, and mark_job_run parks next_run_at at the first scheduled occurrence after the window
while stamping quota_hold_until; _job_is_stale_error_recurring treats an active hold as
deliberate. Any run that reaches the model clears it. Preflight no longer mislabels a
rate-limited AuthError as "provider credential missing" (blocked_config) so the hold applies
with or without a fallback chain.
Persisted marker shape (quota_hold_until) and the hold direction follow PR #89395.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Fixes#89376
- turn_recovery: the upstream_blocked hint no longer advertises model.default_headers,
which _apply_user_default_headers skips on anthropic_messages/bedrock_converse.
- cron: an upstream_blocked run says to set a User-Agent via extra_headers or pin
another provider, instead of 'run it again' — a retry never heals a firewall block.
- agent_runtime_helpers: in anthropic_messages the request dump read agent.client
(None) and printed 'Authorization: Bearer None' with a /chat/completions URL; it now
reads the live _anthropic_client key and uses /messages.
A cron job created from an api_server turn captured origin.platform="api_server".
That adapter's send() is a stub (request/response only, supports_async_delivery=False),
so every fire ran fine (last_status=ok) and then recorded
"API server uses HTTP request/response, not send()" — nothing reached the creator and
nothing warned them at creation time (#69304).
Producer: _origin_from_env() stamps no origin when the session cannot receive async
delivery, so the job takes the existing no-origin path — home-channel fallback at fire
time and the creation-time "local-only" notice (its wording now names stateless HTTP API
sessions alongside CLI/TUI). Fire time: _resolve_origin() treats an already-stamped
api_server origin as missing so existing jobs.json entries fall back too.
Co-authored-by: Chen Jin <243284244@qq.com>
The same credential-resolution fallback the messaging gateway now announces also runs in
tui_gateway/server.py (_resolve_runtime_with_fallback, used_fallback=True only logged 'Primary
auth failed, falling back to ...') and in cron/scheduler.py (_resolve_job_runtime, 'fallback
resolved to ...'), with no user-visible notice on either surface.
- hermes_cli/fallback_config.pre_agent_fallback_notice: shared text builder (gateway/run.py now
calls it) so the three pre-agent paths cannot drift in wording.
- TUI/Desktop: _resolve_agent_model_runtime tags the fallback runtime; _make_agent pops the tag
and sets agent._pending_fallback_notice, emitted once via status_callback on the first reply.
- cron: _resolve_job_runtime tags the runtime; the setup pops it and run_job prepends the notice
to a deliverable final_response ([SILENT] and the [CRON_FAILURE] first-line marker keep their
contracts) — a cron agent has no status rail, so the report is the only place it can surface.
Tests drive the production entry points (_make_agent, run_job) with the resolver and AIAgent
stubbed; each goes red with its attach/prepend removed.
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).
Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.
`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.
No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.
Part of #114587
Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
SupervisedTickerThread (#111010) treated any ended thread as dead. An external provider's
start() (Chronos) arms remote one-shots and returns by design, so every housekeeping tick
logged "Cron ticker thread died without a stop request; restarting" at ERROR and re-ran
start() - a fresh recover_interrupted + NAS list/arm reconcile once a minute on every hosted
instance. Track whether the target escaped with an exception and respawn only then; the
built-in ticker returns normally only on stop_event.
Review follow-up on #113667:
- _pattern_reaches_host_interpreter: when the `-f` head token is not an
interpreter image, require the pattern to plausibly reach the gateway
cmdline (hermes_cli / hermes+gateway tokens, mirroring Branch D). On
the previous head `pkill -f 'hermes-polis/run.sh'` and
`pkill -f my_hermes_bot.py` were hard-blocked although both are allowed
on main and cannot match the gateway; both spellings are now negative
controls in test_kill_forms_that_do_not_reach_the_gateway.
- lifecycle_ledger: the unclean-exit warning enumerated only OS causes
(SIGKILL / OOM / VM death); an agent- or descendant-issued kill of the
host interpreter leaves identical evidence and is now named in the
cause family. One invariant test.
- test file: encoding="utf-8" on the bare write_text calls the footgun
scanner flags.
An agent-issued `taskkill /F /IM python.exe` (or `pkill -9 python3`, `killall python`,
`Stop-Process -Name python`, `taskkill /FI "IMAGENAME eq python.exe"`, `pgrep python | xargs
kill`) from inside the supervised gateway killed the gateway: every branch of
_GATEWAY_LIFECYCLE_PATTERN was anchored on a hermes/gateway token, and the supervised gateway is
literally a `python` process. Branch E is token-aware (not a line regex) so option values are
never read as targets, `-f` cmdline patterns are judged as patterns (`pkill -f 'python
my_script.py'` passes, `pkill -f 'python -m hermes_cli.main'` does not), and other image names
(`taskkill /F /IM agent-browser.exe`) stay killable. Numeric-PID kills stay out of scope: the
explicit PID / `proc_*` id is the ownership-scoped route the terminal rejection now names.
The guard never ran on the Windows Scheduled-Task topology either: the launcher exports only the
generalized HERMES_SUPERVISED_CHILD marker, which gateway/restart.py never read.
is_supervised_gateway_launch() reads it and gates the self-kill guards;
is_gateway_supervisor_process() deliberately keeps ignoring it because it also selects the
exit-75 restart route, which the task has no restart policy to honour (#113670).
Supersedes the narrow `/IM python.exe` regex from #113671 (kept for authorship); the Windows
spellings from #94379 (`hermes.exe gateway restart`, `taskkill`/`Stop-Process` on hermes-gateway
tokens) ride along.
Fixes#113667
The guard that stops a supervised gateway from restarting, stopping or
uninstalling itself only knew POSIX spellings of those commands, so the
Windows spellings of the very same operations walked straight past it.
Branch A anchored on the bare CLI name, but on Windows the CLI is spelled
with an executable suffix, and npm-style shims install `hermes.cmd` and
`hermes.ps1` beside the `hermes.exe`. Every one of these was allowed
while the unsuffixed form was blocked:
hermes.exe gateway restart -> ALLOWED (now blocked)
hermes.cmd gateway stop -> ALLOWED (now blocked)
hermes.bat gateway restart -> ALLOWED (now blocked)
hermes.ps1 gateway uninstall -> ALLOWED (now blocked)
C:\tools\hermes.exe gateway stop-> ALLOWED (now blocked)
Branch D had the same gap for process termination: `\bp?kill\b` cannot
reach inside `taskkill` (there is no word boundary between the two `k`s),
so `taskkill /F /IM hermes-gateway.exe` and `Stop-Process -Name
hermes-gateway` were allowed while `pkill -f hermes gateway` was blocked.
This is reachable. The guard is gated on `_is_supervised_gateway_process()`
— PID-file ownership — not on platform, and its callers (terminal_tool,
code_execution_tool, approval, cron) all run on Windows.
Service-control spellings (`net stop`, `sc stop`, `Stop-Service`) are
deliberately left out: they presuppose a service install this guard has
no evidence of, and guessing at one risks blocking unrelated services.
`C:/tools/hermes.exe` (forward slashes) also stays allowed — the `/` in
the Branch A lookbehind is the deliberate #77173 path false-positive
fix, and narrowing it is a separate decision.
19 new cases pin the caught spellings, the quoted/wrapped forms that
reach the same place through the tokenizing rescan, and that the suffix
does not widen the match — `hermes.exe gateway start` stays benign,
`my-hermes.exe` is still a different binary. 13 of the 19 fail without
this change.
(cherry picked from commit e81c42e4482cb72a93f6468d7571f98aa5fcb848)