fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost that occurrence silently: resume_job recomputed next_run_at from now, so the elapsed slot was neither fired nor recorded — no execution row, no incident, no log line, and last_dispatch stayed on the previous run (the reporter's daily job showed next_run_at jumping two cadences with nothing in between). resume_job now leaves a past stored next_run_at in place for cron/interval jobs and logs that it did. The first tick after resume then applies the existing occurrence policy to that instant — late fire within grace, one collapsed catch-up run past grace, or the loud "missed its scheduled time" skip when cron.catch_up_missed is false — so the slot is accounted for the same way a restart-gap slot is (#107485 contract: every recurring occurrence runs once or its skip is logged). One-shots, future instants and jobs created --paused (next_run_at null) still recompute from now. Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of main's self-removal/fire-claim-skew code and issue-numbered test were dropped).
This commit is contained in:
24
cron/jobs.py
24
cron/jobs.py
@@ -2158,11 +2158,31 @@ def pause_job(job_id: str, reason: Optional[str] = None) -> Optional[Dict[str, A
|
||||
|
||||
|
||||
def resume_job(job_id: str) -> Optional[Dict[str, Any]]:
|
||||
"""Resume a paused job and compute the next future run from now. Accepts a job ID or name."""
|
||||
"""Resume a paused job. Accepts a job ID or name.
|
||||
|
||||
A recurring job paused across one of its slots must not lose that slot silently: the stored
|
||||
``next_run_at`` (already past) survives resume as the due instant, and the ordinary late /
|
||||
catch-up / ``cron.catch_up_missed`` policy in the due scan decides what happens to it — one
|
||||
fire or a logged skip, never a silent re-anchor past it (#113603). One-shots and future
|
||||
instants recompute from now as before.
|
||||
"""
|
||||
job = resolve_job_ref(job_id)
|
||||
if not job:
|
||||
return None
|
||||
next_run_at = compute_next_run(job["schedule"])
|
||||
stored_next = job.get("next_run_at")
|
||||
stored_dt = _parse_aware(stored_next) if stored_next else None
|
||||
if (
|
||||
job["schedule"].get("kind") in {"cron", "interval"}
|
||||
and stored_dt is not None
|
||||
and stored_dt <= _hermes_now()
|
||||
):
|
||||
next_run_at = stored_next
|
||||
logger.info(
|
||||
"Job '%s' resumed with occurrence %s that elapsed while paused kept due; the next "
|
||||
"tick fires it (late/catch-up) or logs the skip.",
|
||||
job.get("name", job["id"]), stored_next)
|
||||
else:
|
||||
next_run_at = compute_next_run(job["schedule"])
|
||||
if next_run_at is None and job["schedule"].get("kind") == "once":
|
||||
run_at = job["schedule"].get("run_at", "unknown")
|
||||
raise ValueError(
|
||||
|
||||
@@ -155,8 +155,13 @@ dropped silently. The mechanics, in the order the due scan applies them
|
||||
with a logged reason when the operator set `cron.catch_up_missed: false`
|
||||
(planned downtime). One-shots past their 120 s grace are retired with a
|
||||
diagnostic, never resurrected.
|
||||
6. **Paused / disabled / terminal jobs never catch up**; the due scan drops them
|
||||
before any of the above, and pause/resume clears any pending slot.
|
||||
6. **Paused / disabled / terminal jobs never fire**; the due scan drops them
|
||||
before any of the above, and pause/resume clears any pending slot. A
|
||||
recurring occurrence that came due *while paused* is not lost, though:
|
||||
`resume_job` keeps a past stored `next_run_at` as the due instant instead of
|
||||
re-anchoring from now (and logs that it did), so the first tick after
|
||||
resume applies rules 3–5 to it — one late/catch-up run, or a logged skip.
|
||||
One-shots and future instants recompute from now on resume.
|
||||
|
||||
The same store fields drive every topology: a standalone `hermes -p X gateway
|
||||
run` and a profile served by the default multiplexer (`_start_multiplex` ticks
|
||||
|
||||
@@ -694,7 +694,7 @@ hermes cron <list|create|edit|pause|resume|run|remove|status|runs|incidents|doct
|
||||
| `create` / `add` | Create a scheduled job from a prompt, optionally attaching one or more skills via repeated `--skill`. Supports a per-job reasoning pin via `--reasoning-effort <none\|minimal\|low\|medium\|high\|xhigh\|max\|ultra>`. |
|
||||
| `edit` | Update a job's schedule, prompt, name, delivery, repeat count, or attached skills. Supports `--clear-skills`, `--add-skill`, and `--remove-skill`, plus `--reasoning-effort` (empty string clears the pin). |
|
||||
| `pause` | Pause a job without deleting it. |
|
||||
| `resume` | Resume a paused job and compute its next future run. |
|
||||
| `resume` | Resume a paused job. A recurring slot that came due while paused stays due (one catch-up run or a logged skip on the next tick); otherwise the next future run is computed. |
|
||||
| `run` | Trigger a job on the next scheduler tick. |
|
||||
| `remove` | Delete a scheduled job. |
|
||||
| `status` | Check whether the cron scheduler is running. |
|
||||
|
||||
@@ -264,7 +264,7 @@ hermes cron tick
|
||||
What they do:
|
||||
|
||||
- `pause` — keep the job but stop scheduling it
|
||||
- `resume` — re-enable the job and compute the next future run
|
||||
- `resume` — re-enable the job. A recurring job whose slot came due while it was paused keeps that slot due, so the next tick fires one catch-up run (or logs the skip when `cron.catch_up_missed: false`) instead of silently jumping to the next occurrence; otherwise the next future run is computed
|
||||
- `run` — trigger the job on the next scheduler tick
|
||||
- `remove` — delete it entirely
|
||||
- `edit` — modify schedule, prompt, delivery, etc.
|
||||
|
||||
Reference in New Issue
Block a user