fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)

A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
This commit is contained in:
lEWFkRAD
2026-09-18 00:43:07 -07:00
committed by Teknium
parent 9ee7da83fa
commit cb8d652b49
4 changed files with 31 additions and 6 deletions

View File

@@ -2158,11 +2158,31 @@ def pause_job(job_id: str, reason: Optional[str] = None) -> Optional[Dict[str, A
def resume_job(job_id: str) -> Optional[Dict[str, Any]]:
"""Resume a paused job and compute the next future run from now. Accepts a job ID or name."""
"""Resume a paused job. Accepts a job ID or name.
A recurring job paused across one of its slots must not lose that slot silently: the stored
``next_run_at`` (already past) survives resume as the due instant, and the ordinary late /
catch-up / ``cron.catch_up_missed`` policy in the due scan decides what happens to it — one
fire or a logged skip, never a silent re-anchor past it (#113603). One-shots and future
instants recompute from now as before.
"""
job = resolve_job_ref(job_id)
if not job:
return None
next_run_at = compute_next_run(job["schedule"])
stored_next = job.get("next_run_at")
stored_dt = _parse_aware(stored_next) if stored_next else None
if (
job["schedule"].get("kind") in {"cron", "interval"}
and stored_dt is not None
and stored_dt <= _hermes_now()
):
next_run_at = stored_next
logger.info(
"Job '%s' resumed with occurrence %s that elapsed while paused kept due; the next "
"tick fires it (late/catch-up) or logs the skip.",
job.get("name", job["id"]), stored_next)
else:
next_run_at = compute_next_run(job["schedule"])
if next_run_at is None and job["schedule"].get("kind") == "once":
run_at = job["schedule"].get("run_at", "unknown")
raise ValueError(

View File

@@ -155,8 +155,13 @@ dropped silently. The mechanics, in the order the due scan applies them
with a logged reason when the operator set `cron.catch_up_missed: false`
(planned downtime). One-shots past their 120 s grace are retired with a
diagnostic, never resurrected.
6. **Paused / disabled / terminal jobs never catch up**; the due scan drops them
before any of the above, and pause/resume clears any pending slot.
6. **Paused / disabled / terminal jobs never fire**; the due scan drops them
before any of the above, and pause/resume clears any pending slot. A
recurring occurrence that came due *while paused* is not lost, though:
`resume_job` keeps a past stored `next_run_at` as the due instant instead of
re-anchoring from now (and logs that it did), so the first tick after
resume applies rules 3–5 to it — one late/catch-up run, or a logged skip.
One-shots and future instants recompute from now on resume.
The same store fields drive every topology: a standalone `hermes -p X gateway
run` and a profile served by the default multiplexer (`_start_multiplex` ticks

View File

@@ -694,7 +694,7 @@ hermes cron <list|create|edit|pause|resume|run|remove|status|runs|incidents|doct
| `create` / `add` | Create a scheduled job from a prompt, optionally attaching one or more skills via repeated `--skill`. Supports a per-job reasoning pin via `--reasoning-effort <none\|minimal\|low\|medium\|high\|xhigh\|max\|ultra>`. |
| `edit` | Update a job's schedule, prompt, name, delivery, repeat count, or attached skills. Supports `--clear-skills`, `--add-skill`, and `--remove-skill`, plus `--reasoning-effort` (empty string clears the pin). |
| `pause` | Pause a job without deleting it. |
| `resume` | Resume a paused job and compute its next future run. |
| `resume` | Resume a paused job. A recurring slot that came due while paused stays due (one catch-up run or a logged skip on the next tick); otherwise the next future run is computed. |
| `run` | Trigger a job on the next scheduler tick. |
| `remove` | Delete a scheduled job. |
| `status` | Check whether the cron scheduler is running. |

View File

@@ -264,7 +264,7 @@ hermes cron tick
What they do:
- `pause` — keep the job but stop scheduling it
- `resume` — re-enable the job and compute the next future run
- `resume` — re-enable the job. A recurring job whose slot came due while it was paused keeps that slot due, so the next tick fires one catch-up run (or logs the skip when `cron.catch_up_missed: false`) instead of silently jumping to the next occurrence; otherwise the next future run is computed
- `run` — trigger the job on the next scheduler tick
- `remove` — delete it entirely
- `edit` — modify schedule, prompt, delivery, etc.