Files
hermes-agent/cron
Teknium 0fc2a10d82 fix(cron): stop one-shot CLI cron run from orphaning the job; reap dead-owner claims on tick
`hermes cron run <job_id>` from a one-shot CLI invocation could
background-dispatch the run onto a daemon thread of the calling process
(when the CLI inherited a gateway/desktop session env and resolved a
session key). The CLI printed "Triggered job: ..." and exited instantly,
killing the runner mid-LLM-call: the async delegation died with
state='unknown' and the job's row in cron/executions.db stayed
status='claimed' forever, blocking every subsequent run of that job.

Two-part fix:

1. hermes_cli/cron.py: `_job_action("run", ...)` declares the delivery
   channel stateless (scoped ContextVar set/reset around the call) before
   invoking the cron API, so `async_delivery_supported()` gates off
   `_try_dispatch_background_run` and the run executes synchronously to
   completion in the CLI process — the same behavior `hermes -z` already
   gets via declare_stateless_channel().

2. cron/scheduler.py: tick() now periodically invokes
   recover_interrupted_executions() (previously only run at scheduler
   startup), so execution rows whose exact owner process is provably dead
   (pid + process start time check in _owner_is_live) are reaped to
   'unknown' by the long-lived gateway ticker without a restart.
   Throttled to once per 300s so idle 60s ticks don't pay a ledger
   connection every cycle.

Tests: tests/cron/test_dead_owner_claim_reclaim.py covers the dead-owner
reap (real dead pid via a finished subprocess), live-owner rows surviving
the reap, throttle behavior, reap-failure isolation, the CLI stateless
gate (including restoration after the call), and the end-to-end refusal
of background dispatch under a stateless channel.

Fixes #86721
2026-08-15 02:44:14 -07:00
..