Files
hermes-agent/tests/cron
pierrenode ba4c2d5253 fix(cron): make a recurring job stuck in state=error recoverable again
is_terminal_job() treats state=error identically to state=completed at
every one of its 6 call sites (all added together in c3a63a16f1, "refuse
to run terminal jobs"). That conflates two very different situations:

* state=completed: a one-shot that genuinely has no more occurrences,
  ever. Correctly terminal.
* state=error: set ONLY on a cron/interval job when compute_next_run()
  fails to produce a next occurrence (e.g. the croniter package is
  missing at runtime). _mark_job_run_locked's own comment is explicit:
  "Recurring jobs must NEVER be silently disabled" (issue #16265) — the
  job is left enabled=True specifically so it keeps being a live,
  recoverable job once the underlying issue resolves.

Because is_terminal_job() lumps both together, a recurring job that ever
reaches state=error is wedged forever, with every recovery path refusing
it:

* _get_due_jobs_locked()'s own next_run_at self-heal (a few lines below
  its own is_terminal_job() check) never runs, because the check itself
  skips the job first.
* resume_job() -> update_job() raises "Cannot activate terminal cron job
  ... use cron resume --run-now or --at."
* rearm_oneshot() (the suggested alternative in that exact error message)
  itself raises "Cannot re-arm recurring jobs: re-arm is one-shot-only."
* advance_next_runs() and _claim_job_for_fire_locked() — the pre-advance
  and claim steps the scheduler's own dispatch loop calls immediately
  after get_due_jobs() for anything that DOES make it into the due list —
  both also refuse the job, so even a manually-recovered next_run_at
  would fail to actually fire.
* pause_job() (itself just an update_job() call) can't even pause a
  broken recurring job through the normal path.

The only way out was deleting the job and recreating it.

Fix: _is_recoverable_error_job() identifies this specific case (state ==
"error" and schedule kind in {"cron", "interval"} — the only shape
state=error ever takes) and is excluded from the is_terminal_job() gate
at update_job() (both checks), advance_next_runs(),
_claim_job_for_fire_locked(), and _get_due_jobs_locked(). trigger_job()
is left untouched: its own error message already points users at "cron
resume", which this fix makes work correctly.

Empirically verified end-to-end against the real module before writing
the fix: create a recurring job, force state=error via
_mark_job_run_locked() with compute_next_run() mocked to return None
(the exact croniter-missing scenario), then confirm resume_job() raises
ValueError, rearm_oneshot() raises ValueError, and get_due_jobs() never
recovers next_run_at. Verified after the fix: all three succeed/recover,
and claim_job_for_fire()/advance_next_runs() correctly stop refusing the
job while still correctly refusing a genuinely state=completed one-shot
through every one of those same paths.

New regression tests (tests/cron/test_terminal_job_rearm.py,
TestRecurringJobStuckInErrorStateIsRecoverable, 6 tests) cover the
due-scan self-heal, resume_job, claim_job_for_fire, advance_next_runs,
and pause_job recovery paths, plus a control confirming a genuinely
completed one-shot stays blocked on every one of the same paths.

Mutation-verified: reverting the fix reproduces exactly 5 failures (all
but the completed-oneshot control, which was never broken).
2026-08-28 02:18:25 +05:30
..