Files
hermes-agent/cron/scheduler_thread.py
teknium1 45b42202fa fix(cron): a raising profile gate ticks nothing; housekeeping respawns a dead ticker (#111010)
Follow-up on the salvaged #111034:

- `_start_multiplex` published the enumerated home list before the gate had
  filtered it, so a raising `profile_gate` (the Desktop stand-down probe from
  #100489) kept the thread alive but ticked every profile UNGATED — racing the
  gateway that owns them for the same cron store. The list is now assigned only
  after gating; a gate failure yields zero ticks for that cycle.
- `cron/scheduler_thread.py::SupervisedTickerThread` wraps the gateway ticker
  thread; `_start_gateway_housekeeping` gets a per-tick "Cron ticker supervisor"
  chore that respawns a ticker that ended without a stop request and logs the
  outage at ERROR. Every guard inside `start()` keeps the loop alive, but nothing
  outside it could notice a thread that had already ended.
- Tests trimmed to the invariants, proven red on origin/main: a REAL corrupt
  `executions.db` (no patched recover) no longer kills the ticker; a raising gate
  keeps the thread alive with zero ticks; housekeeping restarts a dead ticker and
  leaves a stopped one alone.

Root cause: the unguarded pre-loop `recover_interrupted_executions()` +
`record_ticker_heartbeat()` were added by d9dd05b69d (#61791, "truthful
execution ledger", 2026-07-09). The reporter's build (e440bf35) also carried
#107485's `completed_occurrence()` in the due scan, which opens the same ledger
on every tick — the first traceback in their errors.log is that in-loop hit
(caught); the restart then hit the SAME corrupt ledger from the pre-loop
recovery scan, which nothing caught: thread dead, no heartbeat, no error marker.
2026-09-14 16:14:33 -07:00

54 lines
2.2 KiB
Python

"""Supervised daemon thread for the in-process cron ticker.
The gateway (and the Desktop backend) run ``InProcessCronScheduler.start`` on a bare daemon thread.
Every guard inside ``start`` keeps the loop alive on a per-tick failure, but nothing outside it could
notice a thread that had already ended — the gateway kept serving while ``ticker_heartbeat`` froze and
no job fired again until a restart (#111010). The supervisor is the missing outer layer: the
housekeeping loop asks it once a cycle to respawn a ticker that died while shutdown was not requested.
"""
from __future__ import annotations
import logging
import threading
from typing import Any, Callable, Mapping
logger = logging.getLogger(__name__)
class SupervisedTickerThread:
"""``threading.Thread``-shaped handle whose ``restart_if_dead`` respawns a dead ticker."""
def __init__(self, target: Callable[..., Any], *, args: tuple = (), kwargs: Mapping[str, Any] | None = None,
stop_event: threading.Event, name: str = "cron-scheduler") -> None:
self._target, self._args, self._kwargs = target, args, dict(kwargs or {})
self._stop_event, self._name = stop_event, name
self._thread = self._spawn()
self.restarts = 0
def _spawn(self) -> threading.Thread:
return threading.Thread(target=self._target, args=self._args, kwargs=self._kwargs, daemon=True, name=self._name)
def start(self) -> None:
self._thread.start()
def is_alive(self) -> bool:
return self._thread.is_alive()
def join(self, timeout: float | None = None) -> None:
self._thread.join(timeout)
def restart_if_dead(self) -> bool:
"""Respawn the ticker when it ended without ``stop_event``; True when a restart happened."""
if self._stop_event.is_set() or self._thread.is_alive():
return False
self.restarts += 1
logger.error(
"Cron ticker thread %r died without a stop request; restarting (restart #%d). Scheduled jobs "
"did not fire while it was down — check errors.log for the escaping exception.",
self._name, self.restarts,
)
self._thread = self._spawn()
self._thread.start()
return True