fix(gateway): restart wait survives a non-finite drain; tighten its tests

Follow-up to the salvaged restart-wait commits:

- A drain or cron timeout of .inf now means "wait indefinitely" instead of
  an OverflowError from the integer stop envelope, which crashed
  `hermes gateway restart` and made `hermes update` silently fall back to
  its 45s floor. The fleet "draining (up to Ns)" lines and the drain
  progress report format the budget instead of int()-ing it, so an
  unbounded wait no longer crashes them either (it did on main too).
- cron_drain_timeout is required: a 0.0 default meant "cron opted out",
  the under-budget this fix exists to remove.
- Docstrings describe what the budget actually covers (PID exit, not
  replacement startup).
- Tests assert the observer outlasts after-turn + the supervisor stop
  envelope and that configured cron reaches the CLI wait, instead of
  re-deriving the formula; the negative wording assertion on the pending
  footer is dropped (change-detector).
This commit is contained in:
kshitijk4poor
2026-09-26 19:36:05 +05:30
committed by kshitij
parent 643b2091f0
commit 8afaab3703
6 changed files with 30 additions and 34 deletions

View File

@@ -356,11 +356,15 @@ def resolve_systemd_timeout_stop_sec(
def resolve_restart_exit_wait_budget(
drain_timeout: float, after_turn_timeout: float, cron_drain_timeout: float = 0.0, *, headroom: float = 15.0,
drain_timeout: float, after_turn_timeout: float, cron_drain_timeout: float, *, headroom: float = 15.0,
) -> float:
"""Observer budget for in-band deferral, the full stop envelope, and replacement startup.
"""Seconds a CLI waits for the gateway PID to exit after SIGUSR1: the in-band after-turn
deferral, then the longest supervisor stop envelope — the one systemd's ``TimeoutStopSec`` is
sized from (cron drain + cleanup reserve, supervisor headroom, floor; #94759) — then
observer ``headroom``.
The stop envelope includes cron cleanup reserve, supervisor headroom and the floor;
deferral precedes it, while observer headroom follows it. Cron zero opts out.
A non-finite drain means "wait indefinitely", not a crash in the integer envelope.
"""
if not all(math.isfinite(_seconds(value)) for value in (drain_timeout, cron_drain_timeout)):
return math.inf
return _seconds(after_turn_timeout) + resolve_systemd_timeout_stop_sec(drain_timeout, cron_drain_timeout) + _seconds(headroom)