Files
hermes-agent/cron
teknium1 454f9c1421 fix(cron): remind once per cooldown after the alert-once gate; migrate alerted_at
Builds on the previous commit (alerted treated like closed): a permanently
silent `alerted` incident would hide a job that stays broken for days, so the
gate now withholds only inside `cron.failure_repeat_alert_hours` (default 6,
0 = re-alert on every failing run) and lets exactly one reminder through, which
re-stamps the window.

- cron/incidents.py: `alerted_at` column (added in place to existing ledgers via
  add_column_if_missing); set_incident_state(..., "alerted") stamps it every time
  so a reminder restarts the cooldown; resolved->detected re-open clears it so
  the same error after a green run alerts immediately.
- cron/scheduler.py: `_repeat_alert_withheld` reads the stamp; a missing or
  unparseable stamp (pre-migration row) delivers rather than swallowing the
  alert; `closed` still wins; the unreadable-ledger fail-open is unchanged. The
  crash path (`_deliver_crash_failure`) shares the gate through
  `_upsert_incident_for_failure`.
- hermes_cli/config_defaults.py + website/docs cron page: document the key.
- tests: fold the contributor's two tests and the old "unacked failures keep
  alerting per run" change-detector into two invariants (unit gate incl. 0 /
  legacy row / closed; end-to-end alert once -> reminder once -> recovery re-arms).

Fixes #113665
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 09:40:00 -07:00
..