Session-finalize hooks ran synchronously on the gateway event loop from
three call sites (shutdown drain, session-expiry watcher, /new reset).
A plugin hook doing heavy blocking work froze the whole loop: adapter
heartbeats stopped, the drain machinery could not run, and systemd
eventually SIGKILLed the process mid-export. Observed live on a
multi-day 4.7G session where the nemo_relay observability plugin
serialized a full-session ATIF trace inside on_session_finalize.
Changes:
- gateway/run.py: new GatewayRunner._finalize_session_off_loop()
dispatches hermes_cli.lifecycle.finalize_session via the gateway
executor under asyncio.wait_for (10s budget), mirroring
_cleanup_agent_resources_off_loop (#53175). Shutdown finalize and
the session-expiry watcher now use it.
- gateway/slash_commands.py: /new reset path uses the same helper.
- plugins/observability/nemo_relay: ATIF export is now bounded
(HERMES_NEMO_RELAY_ATIF_EXPORT_TIMEOUT_S, default 30s) and skipped
entirely for sessions whose Relay scope operations already errored
(their exporter state is unreliable and the export can be
pathologically slow).
- tests/gateway/test_finalize_session_off_loop.py: regression tests
proving the loop stays live under a wedged hook and the budget is
enforced.