* feat(cron): durable failure incidents with signature dedup and ack
Introduce a durable cron incident store (cron_incidents in the shared
cron/executions.db) that groups "same job + same error signature" across
runs, so a known recurring failure stops re-pinging the operator every run
once it has been acknowledged.
- cron/incidents.py: lazily-created incident table (detected -> alerted ->
reviewed -> closed lifecycle; closed is per-signature terminal), sha256
signature dedup over job_id + normalized error, redacted/truncated error
storage, failure-type classification, and ack/list/get/count helpers.
- cron/scheduler.py: record an incident on the failure delivery path and
suppress the per-run failure ping when the exact signature is acked (both
the normal failure path and the processing-raised retry path). Best-effort:
an incident-store error never breaks the cron run or delivery. Streak nudge,
alert-once markers, and delivery-error behavior are untouched.
- hermes_cli: add `hermes cron incidents [--state ...]` and
`hermes cron incidents ack <id>`.
- tests/cron/test_cron_incidents.py: dedup, lifecycle, redaction,
classification, lazy-schema, scheduler gating, and CLI coverage.
Non-goals deferred to later slices: Discord buttons/review view, HMAC action
tokens, owner-agent review launch, approval-gated fixes, incident playbooks.
* refactor(cron): tighten incident lifecycle, wire alerted state and suppressed_acked outcome
Follow-ups on top of the salvaged #94692:
- Drop the dead 'reviewed' state and the SQLite CHECK (state validity
lives in INCIDENT_STATES so future slices can add states without a
table rebuild); lifecycle is detected -> alerted -> closed.
- Actually mark incidents 'alerted' after a failure ping reaches
delivery, on both the normal and exception delivery paths.
- Record ack-suppressed runs with a distinct 'suppressed_acked'
delivery outcome (registered in cron_health monitoring) instead of
the ambiguous generic 'suppressed'.
- Drift-skip alerts explicitly bypass the ack gate (they carry the
remediation command and alert once via drift_alerted already).
- Docs: failure-incidents section in the cron guide.
- Tests for the alerted transition + never-resurrect-closed.
---------
Co-authored-by: Laura López Real <113060513+laulopezreal@users.noreply.github.com>