A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus session (containers, minimal LXCs, supervisors without linger) fails EVERY scheduled job at dispatch: restart_safe_gateway_child_argv() raises, run_one_job() records a failure, and the only symptom is silently skipped executions (a missed nightly backup, dead watchdogs, no alert). Cron now degrades to a direct external subprocess with a once-per-process warning instead of raising, unless cron.require_restart_safe_scope=true (config.yaml, default false) restores fail-closed. Degraded jobs keep process separation and the full #101940 ownership handoff - only cgroup isolation is lost, so a mid-job gateway restart kills the worker and the execution ledger records exactly that. The dispatch is a GatewayChildDispatch NamedTuple (in_process / scoped / degraded) so the degraded case can never collapse into the "not managed, stay in-process" sentinel - the failure mode that would recreate the restart-interruption edge #101940 closed. Kanban stays fail-closed (require_restart_safe_scope=True at its call sites): its workers are long-lived agentic runs, so the degrade policy is limited to bounded cron jobs in this PR. Addresses the #102431 review: the env-var flag became a config key per AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed instead of updating its tests to a degraded contract, main's enable-linger remedy message is preserved, and the degrade warning fires once per process.
3 lines
62 B
Plaintext
3 lines
62 B
Plaintext
kiwipaulrob
|
|
# PR #102431 rework (cron scope graceful degrade)
|