The plugin's _Runtime.run_in_session wrapper serves every mark/event it
emits (turn start/end, approvals, subagent marks) and runs synchronously
on the agent's conversation thread. It passed no timeout, so the host's
run_in_session default (timeout=None) made each mark an UNBOUNDED native
call. With a wedged native Relay pipeline the agent blocked between API
calls with zero activity ticks — observed live 2026-08-15: two cron jobs
died at the 600s inactivity kill and a gateway chat session at 1800s,
all with last_activity="API call #N completed".
The core's scope push/pop/flush/close sites were bounded with
_SCOPE_OP_TIMEOUT after the 2026-08-10 delegation stall; the plugin's
event marks were the missed sibling class.
Changes:
- plugins/observability/nemo_relay: the wrapper always passes
timeout=relay_runtime._SCOPE_OP_TIMEOUT (10s) to the host. A breach
costs one telemetry span, never the agent; it also sets scope_errored
(so close_session skips the ATIF export for the wedged session) and
warns once so the sick pipeline is visible.
- tests/plugins/test_nemo_relay_bounded_marks.py: proves the budget
reaches the host (fails on the pre-fix code — sabotage-verified),
a TimeoutError flags the session and disables its export, and the
generic error path keeps its scope_errored contract.
Session-finalize hooks ran synchronously on the gateway event loop from
three call sites (shutdown drain, session-expiry watcher, /new reset).
A plugin hook doing heavy blocking work froze the whole loop: adapter
heartbeats stopped, the drain machinery could not run, and systemd
eventually SIGKILLed the process mid-export. Observed live on a
multi-day 4.7G session where the nemo_relay observability plugin
serialized a full-session ATIF trace inside on_session_finalize.
Changes:
- gateway/run.py: new GatewayRunner._finalize_session_off_loop()
dispatches hermes_cli.lifecycle.finalize_session via the gateway
executor under asyncio.wait_for (10s budget), mirroring
_cleanup_agent_resources_off_loop (#53175). Shutdown finalize and
the session-expiry watcher now use it.
- gateway/slash_commands.py: /new reset path uses the same helper.
- plugins/observability/nemo_relay: ATIF export is now bounded
(HERMES_NEMO_RELAY_ATIF_EXPORT_TIMEOUT_S, default 30s) and skipped
entirely for sessions whose Relay scope operations already errored
(their exporter state is unreliable and the export can be
pathologically slow).
- tests/gateway/test_finalize_session_off_loop.py: regression tests
proving the loop stays live under a wedged hook and the budget is
enforced.
Adds an observability/nemo_relay section to the built-in plugins page
(the plugin had no section despite appearing in the shipped table) with
the gateway.telemetry.session_segments keys, defaults-off contract, and
segment metadata; mirrors a summary in the plugin README.
Follow-up fixes from /hermes-pr-review + /simplify-code on PR #83437:
1. Replace _redact_secrets with agent.redact.redact_sensitive_text(force=True)
— the plugin's 11-pattern list was a strict subset of the 50+ patterns in
agent/redact.py. Secrets like Stripe keys, Google API keys, GitLab tokens,
HuggingFace tokens, DB connection strings, and Telegram bot tokens would
all leak through the plugin's list but are caught by the existing redactor.
Added pk-lf- (Langfuse public key) to _PREFIX_PATTERNS in agent/redact.py.
2. Remove dead 'not isinstance(client, object)' check in on_session_finalize —
always False for any Python value.
3. Fix MoAClient.last_reference_metrics() to call the public
self.chat.completions.last_reference_metrics() instead of reaching into
the private _last_reference_metrics attribute via getattr.
4. Deduplicate _coerce_request_messages call in on_pre_llm_request — pass
pre_coerced=input_messages to _messages_for_langfuse_input to avoid
double-coercion + double _capture_content serialization per API request.
5. Add HERMES_LANGFUSE_CAPTURE to OPTIONAL_ENV_VARS in hermes_cli/config.py
for consistency with the other HERMES_LANGFUSE_* env vars.
6. Fix test_sanitized_mode_redacts_secrets test data — the old samples
('sk-abc...1234', 'sk-ant...1234', 'Authorization: Bearer ***') were too
short to match the regex thresholds and never actually tested redaction.
Updated to realistic-length secrets and changed assertions to check that
the output differs from input (redact_sensitive_text masks rather than
inserting the literal string 'REDACTED').
Salvaged from PR #83437 by @erosika, with adopted fixes from @bgodlin (#81054),
@aldoeliacim (#82332), @nftpoetrist (#42326), @rodboev (#39653), @FnExpress
(#64292, supersedes #32175 by @db-aeon), @Per0-1 (#61166), @NaMinhyeok (#64797),
and @liuhao1024 (#43130).
Widens the bundled Langfuse plugin from 6 to 11 hooks and fixes two
attribution bugs. Also adopts shutdown/atexit lifecycle fixes and composes
8 prior community PRs with interaction-fix follow-ups.
Model attribution: on_pre_llm_request and on_post_llm_call now prefer the
wire value (request body model, response model) over the agent attribute,
which goes stale after /model switch or provider fallback.
Cost total: both cost paths now send a summed total alongside the per-type
breakdown, since Langfuse does not derive calculatedTotalCost from
cost_details keys. Subscription-included routes send no cost keys at all.
New coverage: api_request_error closes failed generations with ERROR level;
on_session_finalize/on_session_end close dangling traces for tool-only and
interrupted turns; subagent_start/subagent_stop trace delegated children as
spans; MoA advisor fan-out emits one generation per advisor priced at the
advisor's own model.
Capture modes: HERMES_LANGFUSE_CAPTURE=metadata|sanitized|full (default
sanitized). Sanitized mode redacts secret patterns before truncation.
Adopted lifecycle fixes: shutdown client at session finalize when
reason=shutdown (not on session rotation); atexit finalizer ends open root
spans for short-lived processes; root context manager exited to prevent
interpreter-teardown TypeError; TOCTOU on _get_langfuse() fixed with lock;
reasoning_content surfaced in traces; system prompt included in generation
input for Anthropic/Codex/Bedrock; SDK v3 update_trace replaces set_trace_io.
Closes#29482, #43129, #72661.
Supersedes #81054, #82332, #42326, #39653, #64292, #32175, #61166, #64797, #43130.
Partially addresses #67544 (capture modes + secret redaction; user_id remains open).
Scope events export when their OWNING scope closes. Turn scopes close
every turn; session scopes close only at session end. Marks were attached
to the session handle, so a long-lived conversation — a Slack thread open
all day, the normal enterprise case — emitted no approval or turn marks
for hours, and none at all if the process died first. Audit dashboards
showed an empty approval table while approvals were demonstrably firing;
the operator had to end the session to see anything.
Attach marks to the live turn handle when one exists for the mark's
session (active_turn already validates live/same-profile/same-session/
unreleased), falling back to the session handle otherwise — correct for
session-level events like session.end and for marks emitted outside a
turn. Parentage semantics are unchanged: the turn is a child of the
session, so the session tree is identical, only export cadence changes
from per-session to per-turn.
Scoping the trace key by turn_id (the prior commit) fixed cross-turn
collisions but introduced a slow leak: _finish_trace only pops a key when a
turn ends cleanly (final response has content and no tool calls), so any
turn that is interrupted, ends on a tool call, or has empty final content
now leaves its uniquely-keyed entry in _TRACE_STATE forever. Previously the
constant per-session key was overwritten by the next turn, capping growth at
~1 entry per session.
Add an LRU cap (_MAX_TRACE_STATE) enforced by _evict_stale_locked, called
under _STATE_LOCK immediately before each insert. It evicts the
least-recently-updated entries (using the previously-dead last_updated_at
field) and ends their root span so nothing dangles. Regression test drives
50 non-finalizing turns against a cap of 8 and asserts the dict stays bounded
with the most-recent turns surviving.
The turn- and api-scoped branches each repeated the same
task/session/thread fallback ladder with only the infix differing. Extract
the shared prefix into _scope_prefix so a future scope dimension touches one
ladder instead of three. The legacy branch still returns a bare task_id (not
the task: prefix) for backward compatibility, so it stays separate.
Output key strings are unchanged; a new test pins them across every
task/session/turn/api combination since the keys are matched across hooks
and any drift would silently break trace finalization.
The Langfuse SDK treats `data:*;base64,...` strings as media and tries to
decode them. `_truncate_text` was slicing those strings mid-payload, producing
invalid base64 and noisy "Error parsing base64 data URI" logs. Observability
only needs the metadata, not raw image/audio bytes, so redact the whole data
URI (type, media_type, length) before it reaches the SDK.
Salvaged the Langfuse fix from #39682 onto current main as a standalone,
single-concern change (the dashboard `dist/**` and plugin-discovery parts of
that PR already landed separately on main).
Co-authored-by: foras910521-lab <foras910521-lab@users.noreply.github.com>
Based on #42658 by @mnajafian-nv.
Preserves the real downstream provider/tool exception when NeMo Relay's
managed adaptive execution wraps a failing callback as an internal runtime
error. Without this, the original exception (and its retry-classification
signal, e.g. status_code) is lost behind Relay's wrapper.
Salvage changes on top of the original PR:
- Tolerant Relay-wrapper match: _is_relay_wrapped_callback_error now uses
str.startswith on the "internal error: <cls>: <msg>" prefix instead of
exact equality, so a future Relay version appending a traceback/suffix
doesn't silently defeat the unwrap. On a total format change it returns
False and falls back to the pre-fix behavior (surfacing Relay's error)
rather than masking it.
- Deduplicated the LLM and tool execute paths into a shared
_run_managed_with_downstream_preservation helper, removing ~20 lines of
copy-pasted nonlocal/try-except scaffolding that could drift out of sync.
- Added a real-middleware regression guard
(test_nemo_relay_downstream_unwrap_matches_real_middleware_wrapper_shape)
that drives hermes_cli.middleware._run_execution_chain and asserts the
plugin's _original_downstream_error unwraps the actual private
_DownstreamExecutionError wrapper. The original synthetic tests modeled the
wrapper with a local class, so a rename or shape change in core middleware
would not have been caught; this test fails loudly if that contract drifts.
Co-authored-by: mnajafian-nv <mnajafian@nvidia.com>