Files
hermes-agent/gateway/platforms
JulianCruzet 95fd5d6816 fix(cron): honor ESTOP in NAS fire webhook and misfire backstop
`agent/estop.py:1-9` documents that while the sentinel exists "the cron
scheduler ... skips work." The built-in ticker honors this
(`cron/scheduler.py:3749-3754`), but the managed-cron paths do not:

- Door 2 (NAS fire webhook `_handle_cron_fire`,
  `gateway/platforms/api_server.py:3480-3557`): no ESTOP check between
  the JWT/drain guards and `provider.claim_fire`. Added a
  `check_paused("cron-webhook")` guard inside the reservation block,
  returning 503 + Retry-After so the NAS retries after `hermes resume`
  rather than silently dropping the run.

- Door 3 (misfire backstop `fire_overdue_jobs`,
  `cron/scheduler_provider.py:253-344`): no ESTOP check at the top of
  the function. Added an early-return `check_paused("cron-misfire")`
  guard. Self-healing — the next sweep after `hermes resume` catches
  everything up via the existing claim_fire path.

Both guards use `suppress(ImportError)` matching the ticker idiom so a
broken estop module fails open rather than killing cron. Distinct
component names keep the existing log-once mechanism independent per
surface.

Manual runs (`hermes cron run`, dashboard Trigger) deliberately
unchanged: operator override is arguably a feature, and PR #105144
already rewrites that path.

Tests (3 new, 49 pre-existing in affected files all pass):
- tests/cron/test_misfire_catchup.py::test_estop_engaged_skips_backstop
- tests/cron/test_misfire_catchup.py::test_estop_release_restores_backstop
- tests/gateway/test_api_server_jobs.py::test_fire_webhook_returns_503_when_estop_engaged
2026-09-15 03:33:13 -07:00
..