diff --git a/cron/AGENTS.md b/cron/AGENTS.md index 15701943f0..5e354f5abe 100644 --- a/cron/AGENTS.md +++ b/cron/AGENTS.md @@ -17,6 +17,11 @@ loaded), multi-platform delivery. Hardening invariants — each guards a real failure; don't weaken without answering for it: - **3-minute hard interrupt** on cron sessions: runaway loops cannot monopolise the scheduler. - Catch-up window = half the period, clamped to 120s–2h; 120s grace for missed one-shots. +- Every recurring occurrence is accounted for: `tick()` advances `next_run_at` BEFORE dispatch + (at-most-once across a mid-run crash) and stamps `pending_slot` in the same save; a scan that + finds the stamp with a dead owner restores the instant ONCE (`cron/occurrences.py`), the + executions ledger's `scheduled_instant` blocks a second fire, `cron.catch_up_missed: false` + skips past-grace misses with a logged reason. Never drop a slot silently (#107485). - File lock `~/.hermes/cron/.tick.lock` prevents duplicate ticks across processes. - Cron sessions pass `skip_memory=True`; memory providers intentionally do not run during cron. - Cron execution has its own session. Eligible continuable deliveries may mirror or seed the diff --git a/website/docs/developer-guide/cron-internals.md b/website/docs/developer-guide/cron-internals.md index 29b2b42e6c..a8bab0a297 100644 --- a/website/docs/developer-guide/cron-internals.md +++ b/website/docs/developer-guide/cron-internals.md @@ -114,6 +114,47 @@ tick() 6. Release scheduler lock ``` +### Missed-occurrence contract (restart gaps) + +Recurring jobs are **at-most-once per occurrence, and every occurrence is +accounted for**: it either runs (one execution row carrying its +`scheduled_instant`), or its skip is logged with a reason. An occurrence is never +dropped silently. The mechanics, in the order the due scan applies them +(`cron/jobs.py::_evaluate_due_job`): + +1. **Pre-dispatch advance is provisional.** `tick()` advances `next_run_at` past + the due occurrence *before* dispatch so a crash mid-run cannot re-fire it on + every restart. Because that leaves a window — advanced, but no fire claim yet + (interpreter finalizing, executor refusing work, `SIGKILL`) — the due scan + stamps `pending_slot = {scheduled_at, at, by}` on the record in the same + save. `claim_job_for_fire` (the point after which side effects may exist) + and `mark_job_run` clear it; an explicit `schedule` / `next_run_at` / + `enabled` / `state` rewrite (edit, pause, resume, run-now) drops it. +2. **Restore once.** A later scan that finds a `pending_slot` whose owner is + provably gone (this process and the job is not in its running set; another + process past the 300 s fire-claim lease or with a dead pid) puts + `scheduled_at` back as `next_run_at`, drops the stamp, and logs a WARNING + (`cron/occurrences.py::unclaimed_pending_slot`). This happens at most once + per lost occurrence — the restored instant then meets the ordinary rules + below like any other overdue slot, so there is never a replay of N slots. +3. **Already fired → never twice.** `completed_occurrence()` consults the + executions ledger for a `completed` row with that exact `scheduled_instant` + before anything is due; a slot that ran before the restart advances without + firing. `failed` / `unknown` rows do not count as completion. +4. **Late within grace → fire late.** Grace = half the period clamped to + `[120 s, 2 h]` (`_compute_grace_seconds`); the dispatch is stamped + `last_dispatch.kind = late`. +5. **Past grace → collapse the backlog, fire once** (`kind = catch_up`), or skip + with a logged reason when the operator set `cron.catch_up_missed: false` + (planned downtime). One-shots past their 120 s grace are retired with a + diagnostic, never resurrected. +6. **Paused / disabled / terminal jobs never catch up**; the due scan drops them + before any of the above, and pause/resume clears any pending slot. + +The same store fields drive every topology: a standalone `hermes -p X gateway +run` and a profile served by the default multiplexer (`_start_multiplex` ticks +each home under `_profile_cron_scope`) evaluate the identical record. + ### Gateway Integration In gateway mode, the cron **trigger** (the part that decides *when* a due job diff --git a/website/docs/user-guide/features/cron.md b/website/docs/user-guide/features/cron.md index 0b353d5359..4ad4bc2827 100644 --- a/website/docs/user-guide/features/cron.md +++ b/website/docs/user-guide/features/cron.md @@ -882,8 +882,14 @@ The stamp always reflects **current** auto-fire health: it is overwritten by new ### Local missed-run policy -The local ticker normally runs each overdue recurring job **once**, not once per -missed slot. To avoid catch-up load after a planned gateway stop, set: +If the gateway was down (or restarting) when a recurring job's scheduled time +passed, the job **catches up once** when the scheduler is back: a slot missed +inside a restart gap fires exactly one time, a slot that already ran before the +restart is never run again, and a long outage collapses into a single run rather +than one run per missed slot. Paused jobs never catch up. Each catch-up shows in +`hermes cron list` as `⚠ late` / `⚠ catch-up after missed fire`. + +To avoid that catch-up load after a planned gateway stop, set: ```yaml cron: