Files
hermes-agent/plugins/observability
Teknium fe0a56ed16 fix(nemo_relay): bound plugin Relay marks so a wedged native pipeline cannot stall the agent
The plugin's _Runtime.run_in_session wrapper serves every mark/event it
emits (turn start/end, approvals, subagent marks) and runs synchronously
on the agent's conversation thread. It passed no timeout, so the host's
run_in_session default (timeout=None) made each mark an UNBOUNDED native
call. With a wedged native Relay pipeline the agent blocked between API
calls with zero activity ticks — observed live 2026-08-15: two cron jobs
died at the 600s inactivity kill and a gateway chat session at 1800s,
all with last_activity="API call #N completed".

The core's scope push/pop/flush/close sites were bounded with
_SCOPE_OP_TIMEOUT after the 2026-08-10 delegation stall; the plugin's
event marks were the missed sibling class.

Changes:
- plugins/observability/nemo_relay: the wrapper always passes
  timeout=relay_runtime._SCOPE_OP_TIMEOUT (10s) to the host. A breach
  costs one telemetry span, never the agent; it also sets scope_errored
  (so close_session skips the ATIF export for the wedged session) and
  warns once so the sick pipeline is visible.
- tests/plugins/test_nemo_relay_bounded_marks.py: proves the budget
  reaches the host (fails on the pre-fix code — sabotage-verified),
  a TimeoutError flags the session and disables its export, and the
  generic error path keeps its scope_errored contract.
2026-08-15 14:31:32 -07:00
..