docs(cron): state the missed-occurrence contract for restart gaps
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch advance is provisional, restore once, never twice, grace, opt-out, paused never catches up, same on standalone and multiplexed); the user guide describes the catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md lists it as a hardening invariant.
This commit is contained in:
@@ -17,6 +17,11 @@ loaded), multi-platform delivery.
|
||||
Hardening invariants — each guards a real failure; don't weaken without answering for it:
|
||||
- **3-minute hard interrupt** on cron sessions: runaway loops cannot monopolise the scheduler.
|
||||
- Catch-up window = half the period, clamped to 120s–2h; 120s grace for missed one-shots.
|
||||
- Every recurring occurrence is accounted for: `tick()` advances `next_run_at` BEFORE dispatch
|
||||
(at-most-once across a mid-run crash) and stamps `pending_slot` in the same save; a scan that
|
||||
finds the stamp with a dead owner restores the instant ONCE (`cron/occurrences.py`), the
|
||||
executions ledger's `scheduled_instant` blocks a second fire, `cron.catch_up_missed: false`
|
||||
skips past-grace misses with a logged reason. Never drop a slot silently (#107485).
|
||||
- File lock `~/.hermes/cron/.tick.lock` prevents duplicate ticks across processes.
|
||||
- Cron sessions pass `skip_memory=True`; memory providers intentionally do not run during cron.
|
||||
- Cron execution has its own session. Eligible continuable deliveries may mirror or seed the
|
||||
|
||||
@@ -114,6 +114,47 @@ tick()
|
||||
6. Release scheduler lock
|
||||
```
|
||||
|
||||
### Missed-occurrence contract (restart gaps)
|
||||
|
||||
Recurring jobs are **at-most-once per occurrence, and every occurrence is
|
||||
accounted for**: it either runs (one execution row carrying its
|
||||
`scheduled_instant`), or its skip is logged with a reason. An occurrence is never
|
||||
dropped silently. The mechanics, in the order the due scan applies them
|
||||
(`cron/jobs.py::_evaluate_due_job`):
|
||||
|
||||
1. **Pre-dispatch advance is provisional.** `tick()` advances `next_run_at` past
|
||||
the due occurrence *before* dispatch so a crash mid-run cannot re-fire it on
|
||||
every restart. Because that leaves a window — advanced, but no fire claim yet
|
||||
(interpreter finalizing, executor refusing work, `SIGKILL`) — the due scan
|
||||
stamps `pending_slot = {scheduled_at, at, by}` on the record in the same
|
||||
save. `claim_job_for_fire` (the point after which side effects may exist)
|
||||
and `mark_job_run` clear it; an explicit `schedule` / `next_run_at` /
|
||||
`enabled` / `state` rewrite (edit, pause, resume, run-now) drops it.
|
||||
2. **Restore once.** A later scan that finds a `pending_slot` whose owner is
|
||||
provably gone (this process and the job is not in its running set; another
|
||||
process past the 300 s fire-claim lease or with a dead pid) puts
|
||||
`scheduled_at` back as `next_run_at`, drops the stamp, and logs a WARNING
|
||||
(`cron/occurrences.py::unclaimed_pending_slot`). This happens at most once
|
||||
per lost occurrence — the restored instant then meets the ordinary rules
|
||||
below like any other overdue slot, so there is never a replay of N slots.
|
||||
3. **Already fired → never twice.** `completed_occurrence()` consults the
|
||||
executions ledger for a `completed` row with that exact `scheduled_instant`
|
||||
before anything is due; a slot that ran before the restart advances without
|
||||
firing. `failed` / `unknown` rows do not count as completion.
|
||||
4. **Late within grace → fire late.** Grace = half the period clamped to
|
||||
`[120 s, 2 h]` (`_compute_grace_seconds`); the dispatch is stamped
|
||||
`last_dispatch.kind = late`.
|
||||
5. **Past grace → collapse the backlog, fire once** (`kind = catch_up`), or skip
|
||||
with a logged reason when the operator set `cron.catch_up_missed: false`
|
||||
(planned downtime). One-shots past their 120 s grace are retired with a
|
||||
diagnostic, never resurrected.
|
||||
6. **Paused / disabled / terminal jobs never catch up**; the due scan drops them
|
||||
before any of the above, and pause/resume clears any pending slot.
|
||||
|
||||
The same store fields drive every topology: a standalone `hermes -p X gateway
|
||||
run` and a profile served by the default multiplexer (`_start_multiplex` ticks
|
||||
each home under `_profile_cron_scope`) evaluate the identical record.
|
||||
|
||||
### Gateway Integration
|
||||
|
||||
In gateway mode, the cron **trigger** (the part that decides *when* a due job
|
||||
|
||||
@@ -882,8 +882,14 @@ The stamp always reflects **current** auto-fire health: it is overwritten by new
|
||||
|
||||
### Local missed-run policy
|
||||
|
||||
The local ticker normally runs each overdue recurring job **once**, not once per
|
||||
missed slot. To avoid catch-up load after a planned gateway stop, set:
|
||||
If the gateway was down (or restarting) when a recurring job's scheduled time
|
||||
passed, the job **catches up once** when the scheduler is back: a slot missed
|
||||
inside a restart gap fires exactly one time, a slot that already ran before the
|
||||
restart is never run again, and a long outage collapses into a single run rather
|
||||
than one run per missed slot. Paused jobs never catch up. Each catch-up shows in
|
||||
`hermes cron list` as `⚠ late` / `⚠ catch-up after missed fire`.
|
||||
|
||||
To avoid that catch-up load after a planned gateway stop, set:
|
||||
|
||||
```yaml
|
||||
cron:
|
||||
|
||||
Reference in New Issue
Block a user