Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.
- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
post_api_request payload shape plus `aux_task`, `api_request_id`
(`aux-...`, shared by every attempt of one logical call), `retry_count`,
`streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
turn is in flight; fail-open (a raising/hung subscriber is logged and
the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
attempt shares (_relay_sync_completion / _relay_async_completion /
_relay_sync_stream) run under the hook pair — retries and fallbacks
included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
aux_task and no api_request events; raising subscriber never breaks
the call). First is red on origin/main.
Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.
Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Follow-up trim of the #110265 salvage. `ainvoke_hook` logged raising callbacks
with a bare warning; route them through `_report_hook_failure` (warn-once per
distinct failure, #111922) and, for `_HOOK_TIMEOUT_FAIL_CLOSED_HOOKS`, append
the same named block directive the sync path emits (#109624), so the async twin
cannot drift into a fail-open policy path. Tests trimmed to the salvage bar: the
in-loop await is proven once through the real `_handle_message` path
(`test_async_hook_callback_is_awaited_on_the_gateway_loop`); the manager-level
duplicate is dropped and the narrowing test also pins failure isolation. Docs:
`pre_gateway_dispatch` callbacks may be `async def` and stay unbounded.
Credit order for the three PRs fixing this gap: #102485 (dmspark, earliest,
pre-decomposition `gateway/run.py`), #110253 (KoNit-K, bounded the hook —
rejected by design: neither fail mode is acceptable for a policy gate), #110265
(twidtwid, reporter; cherry-picked because it matches the ainvoke_hook shape,
keeps the hook unbounded, and adapts the existing sync test seams honestly).
Part of #110241
Supersedes #102485
Supersedes #110253
Co-authored-by: David Marcus <dmspark@users.noreply.github.com>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
`GatewayInboundMixin._hm_pre_gateway_dispatch_hook` was a plain `def` calling
the sync `hermes_cli.lifecycle.invoke_hook` from the async `_hm_admit_event`,
so an `async def pre_gateway_dispatch` callback was resolved through
`resolve_plugin_command_result` on a helper thread with its own loop: the
gateway loop blocked for the callback's whole duration and any loop-bound await
(an `asyncio.Event` set by a loop task, a loop-bound aiohttp session,
`asyncio.to_thread`) could never complete, failing at 30s.
Add `PluginManager.ainvoke_hook` (+ `hermes_cli.plugins.ainvoke_hook` /
`hermes_cli.lifecycle.ainvoke_hook`): same payload narrowing (shared
`_hook_callback_kwargs`), observer + isolation semantics and result contract as
`invoke_hook`, but awaitable results are awaited on the caller's loop.
`pre_gateway_dispatch` stays intentionally unbounded. The inbound hook becomes
`async def` and `_hm_admit_event` awaits it; the sync `invoke_hook` is untouched
for every other caller. Existing tests that stubbed the hook synchronously are
adapted to the async seam.
Fixes#110241
Salvages #110265
(cherry picked from commit 22bb10d305c7992f43bbfdf1481ad0ef0015b2b7)
`_run_hook_callback_bounded` treated any live abandoned worker for a callback
(`bool(self._hook_abandoned.get(suppression_key))`) as "still running", so one
never-returning `pre_tool_call` callback made every later tool call fail closed
with the timeout message until the process restarted. The timeout path
self-heals through the 60s suppression window; the abandoned path never did.
Policy now: while the suppression window is open the callback is skipped as
before. After it expires a fresh call id may start a new worker even though the
abandoned one is still alive — capped at `_HOOK_MAX_ABANDONED_WORKERS` (3) live
abandoned workers per callback so a hung plugin cannot leak a thread per call
(the #98382 constraint). At the cap the callback keeps being skipped (fail-closed
for pre_tool_call) with a WARNING naming the callback and its module, until one
of its workers finishes and frees a slot.
Test changes: `test_hung_worker_blocks_new_call_identity_after_suppression`
encoded the removed behaviour (exactly one worker, forever); it becomes
`test_hung_worker_caps_new_call_identities_after_suppression`, which pins the
same invariant it was protecting — bounded leak, never one per call — at the new
bound and checks the warning. `test_hung_worker_does_not_fail_closed_forever`
is the #105223 regression (red on base: call-c returned the block directive).
Redone slim against the per-call-id gate that landed in #111177; #105241
targeted the pre-#111177 shape and needed plugins_ledger/__init__ changes for a
one-retry-then-quarantine policy. Its analysis and shape informed this fix.
Fixes#105223
Supersedes #105241
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Follow-up trim of the #109632 salvage: the block directive a raising
`pre_tool_call` guard produces reused the timeout wording, so an operator could
not tell a crashing guard from a slow one from the tool result. Build it from
the callback name and `TypeError: ...` (error text truncated like
`_report_hook_failure`). The two salvaged tests collapse into one parametrized
invariant (caller-thread and bounded-worker path) that also pins the message
shape and that a sibling callback's result still flows.
Part of #109624
Salvages #109632
A `pre_tool_call` guard that raised failed OPEN (only `_report_hook_failure`,
results stayed empty, the tool ran) while one that timed out failed CLOSED with
a block directive. For a veto hook a crashing guard is the control quietly
disappearing; make both failure modes of `_HOOK_TIMEOUT_FAIL_CLOSED_HOOKS`
consistent: an exception appends a block directive in addition to the warning.
Fixes#109624
Salvages #109632
(cherry picked from commit 6ccf5d8a4082fc6b9af80270aae9f13b0320cfb8)
Under a multiplexed gateway / serve backend the launch profile's own turns bound
secrets and terminal policy but NO HERMES_HOME override (launch_profile_runtime_scope,
_profile_runtime_scope_tokens(None)). Once multiplexing is active an unset override
is the fail-closed "unbound context" signal (serves_routed_profile, per-home slots,
third-party runtime bindings such as OMH's pre_tool_call gate), so every plugin hook
a launch-profile turn fired looked unscoped and a fail-closed plugin vetoed every
tool call ("OMH runtime binding unavailable", #118538). Routed turns already bind
theirs via gateway/run.py::_profile_runtime_scope; the launch profile is a tenant
like any other, so its scope now binds the launch home too.
Sibling gap in the two hook kinds delivered off-turn by long-lived worker threads:
agent.plugin_stream_hooks (on_stream_*/on_interim_message) and the plugin event bus
(plugins_dispatch._dispatch_event) ran callbacks in the worker's empty context, so
even a routed turn's observers saw no scope. Both now capture the emitter's
contextvars at enqueue and run each delivery under a copy.
Single-profile hosts are unchanged (no override until activation).
Closes#118538
The hook loop now treats SystemExit like any other callback failure; the
three sibling dispatch loops in this module (event subscribers,
invoke_middleware, system prompt section content) had the same
`except Exception` gap and let a plugin dependency's sys.exit() abandon
the loop and kill the turn. KeyboardInterrupt still propagates on all
of them.
Two per-call WARNING surfaces were still outside the warn-once reporter:
- agent/plugin_stream_hooks.py::_worker — the on_stream_start/on_stream_delta/
on_stream_end consumer, which fires once per streaming delta (far more often than
per tool call). A mis-declared callback (signature naming tool_data) logged
"Hook ... raised" at WARNING on every delta: 20 deltas -> 20 WARNING lines.
- hermes_cli/plugins_dispatch.py::_deliver_event — plugin event subscribers that
raise identically were warned on every emit.
Both now go through PluginManager._report_hook_failure (keyed by module/qualname,
cleared on unload): the first failure warns and names the fields the hook/event
provides, identical repeats are DEBUG. Skip-and-continue semantics are unchanged.
Part of #111922
Review follow-up on the warn-once hook reporter. Middleware (agent_tool_execution
etc.) runs once per tool call exactly like a hook, so a mis-declared middleware
callback flooded WARNING identically; route its except through the same
_report_hook_failure helper with a "Middleware" surface label.
The reported-failure set was keyed on (hook, id(cb), repr(exc)) and never
cleared: a hook whose message embeds tool args grew it by one entry per call,
and id() recycling across plugin reloads could swallow a reloaded callback's
first failure. Key on (hook, module, qualname, exc type, str(exc)[:200]) and
clear the set in _reset_after_unload_all next to _hook_timeout_suppressed_until.
Part of #111922
A plugin callback whose signature names a parameter the hook never sends
(on_pre_tool(tool_data) where core provides tool_name/args) raises the same
TypeError on every tool call; core logged a WARNING each time — ~1700 lines an
hour in the report, burying the freeze signature it was trying to find
(finding (c) of #111922). The mis-declared signature is the plugin's bug; the
per-call repeat is ours.
invoke_hook now reports one WARNING per distinct (hook, callback, error) and
names the fields the hook actually provides so the author can fix the
signature; identical repeats are logged at DEBUG. A callback failing in a new
way still warns.
Part of #111922
The timeout branch of _run_hook_callback_bounded unconditionally added
gate_key to _hook_abandoned. A worker that finishes between done.wait()
returning False and the caller taking the lock has already popped its
token via _release_token, so nothing would ever clear that entry: the
callback stayed blocked for every later call id until reload with no
thread behind it. Guard the insert on the worker still being registered.
The new test makes the race deterministic by swapping the module's
threading.Event for one whose wait() lets the worker finish and then
reports a timeout, and asserts a fresh call id still runs.
Also pass tool_call_id inline from terminal_tool_result instead of the
conditional dict plumbing: an empty id is already treated as "no
identity" by _hook_call_identity and unknown fields are withheld from
narrow-signature callbacks (same shape as _fire_approval_hook). Update
the stale "(hook_name, id(cb))" comment above _hook_running_callbacks.
Gating hook callbacks by call identity lets two concurrent calls of the
same tool both run their hooks, but it also let a fresh tool_call_id pass
the gate once the 60 s suppression window lapsed even though the previous
worker for that callback never returned. A hung plugin then leaked one
daemon thread per minute for the life of the process; on the old
coarse-keyed gate it leaked exactly one.
Track abandoned-but-running workers per callback: the timeout branch
records the gate key, the worker's own release discards it, and the gate
treats any non-empty abandoned set as "still running" for that callback.
Healthy callbacks keep distinct-id concurrency; hung ones are back to
at most one outstanding worker.
Concurrent invocations of the same tool in one session collapsed into a single
busy key (hook_name, id(cb)): the second invocation was reported as 'still
running' and dropped. For pre_tool_call a drop is a fail-closed block, so the
gate silenced itself on an ordinary, healthy callback.
Measured on a busy profile: 3574 skip lines and 0 timeout lines in one hour —
every skip was the 'while still running' branch, i.e. pure key collision, not
slowness.
The gate now keys on the call identity that is already in the payload
(tool_call_id, else turn_id, else none — the last case behaves exactly as
before). Suppression stays keyed coarsely on (hook_name, id(cb)): a hung
callback is a fact about the callback, so its back-off must not be diluted
per call.
Refs #98382. Independent of #107894 (that one releases the slot on timeout;
this one stops healthy concurrency from colliding).
(cherry picked from commit 53b3dacd008418fcdf5fa6dfcadde575a35a776e)
Slash-command handlers gained loop-safe awaiting in ca9a61ae38, but
`PluginManager.invoke_hook` still called `async def` hook callbacks directly:
the coroutine object was appended to the results (so `pre_llm_call` context
injection silently did nothing) and Python warned "coroutine was never
awaited". `_invoke_hook_callback` now routes every return through
`resolve_plugin_command_result`, which covers both the direct and the
timeout-bounded paths and is safe under the gateway's running loop.
Fixes#12449 (remaining hook half). Salvage of #63240 by @Bartok9, applied
one layer down so the bounded-worker path is covered too.
Co-authored-by: Bartok9 <Bartok9@users.noreply.github.com>
The identity-guarded token pop appeared twice (the runner's finally and the
new worker-start except); a future edit to one copy would silently reintroduce
the sticky-token class. Hoist into a local closure called from both sites, and
extend the _run_hook_callback_bounded docstring with the new skip reason.
Behavior-preserving follow-up on #104651.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.