Closes#33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.
Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.
Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.
A multiplexed gateway answered "which bot received this / who may admit it /
where does it run" in three places (`_transport_owner`,
`_authorization_home_for_source`, `_resolve_profile_home_for_source` +
`_session_key_profile`) that agreed only because they read the same fallback
chain. `gateway/session_identity.py` answers them once: `resolve_identity()`
folds `_admit_primary_source` + `_stamp_routed_profile` + the transport-owner
lookup and pins a frozen `RoutingIdentity` (transport_profile, runtime_profile,
authorization_home, runtime_home, weak transport ref) on the source as a
wire-invisible attribute, like `_transport_adapter_ref`. Under multiplexing a
route to an unserved profile raises `IdentityUnresolved` instead of a
`None`-means-default return; `"default"` is spelled out inside the object.
Additive: the existing helpers become thin readers of the identity when it is
present and keep their fallback chain when it is not, `source.profile` stays
the serialized runtime profile (None on the wire ⇔ default) and every
historical `agent:main` key is byte-identical. `replace_source()` copies a
source without losing its provenance (run_topics used to hand-copy the
transport ref).
Phase 1 of #88715; the gateway rows of #90142 / #93943.
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:
- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
already did this in `_handle_reset_command`, the second request is
idempotent) and the idle `_handle_stop_command` tail, which replied
"No active task to stop." while a background child was running; it now
stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.
The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.
An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.
Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.
Part of #114456
NAS maps the agent-cron bearer to a provisioned instance via an `agent:*` client or the
`hermes-cli-vps` bootstrap session (hermes-portal `server/agent-cron/instance-auth.ts`). A
container whose auth.json holds a plain `hermes-cli` user login is refused with 403
invalid_client on every arm, re-arm and list, for the life of that credential - and the
re-login users try first (`hermes auth logout nous` + device code) replaces the bootstrap
session, making it permanent. Chronos previously logged one bare warning per job and left
the jobs with no trigger at all: they only ran through the misfire sweep, minutes late.
NasCronClientError now carries the HTTP status and the OAuth `error` code. On an identity
rejection the provider logs ONE warning that names the real remedy (restore the hosted
credential from the Nous Portal; re-login cannot fix it), stops calling NAS, and starts the
built-in ticker with the gateway's own adapters/loop so scheduled jobs keep firing on time.
Transient 5xx/transport failures keep retrying on the next reconcile.
Supersedes #97566 (@wesleysimplicio), whose runtime-credential swap resolves to the same
bearer on main (`agent_key` is the access token) and so could not change the 403.
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
Operators enabling compression.checkpoint_required need to know up front that
the contract is opt-in for third-party archiving providers; otherwise the flag
blocks every compression attempt and the reason is only visible after the fact.
Docs hunk of #106882 (the code change was consolidated with #106879).
The honest "unavailable (remote backend)" state (previous commit) tells the
user the copy will never happen; the tooltip now also says what does work
against a remote backend — Install from Git with the Desktop target checked
clones the desktop half onto this machine (the install modal already takes
that branch for connection.mode === 'remote'). Docs note the new state next to
the existing remote-backend paragraph.
The contributor test covers the queue fall-through; the Desktop's normal case is
a redirect-capable AIAgent, where busy_input_mode=interrupt turned the edit into
a mid-turn redirect and left the un-edited transcript in place. Pin that branch
too, and document in the rewind section that a truncating submit refuses with
4009 while a turn runs so hosts interrupt + retry (the Desktop already does).
`444c75c10a` rendered `titleBar.center` in the workspace panel header on
full pages so the kanban board switcher would sit beside the page title
instead of colliding with the sidebar tab strip. That made the area migrate
between two React subtrees on chat <-> page navigation: the old instance's
effect cleanup ran after the new instance's setup and wiped every global
side effect a third-party plugin had just re-created (#114290).
Keep both intents: `titleBar.center` is a permanent titlebar slot again (the
salvaged commit), and page-owned controls move to a dedicated
`WORKSPACE_PAGE_HEADER_AREA` (`workspace.pageHeader`) that the workspace
pane projects into its vetoed tab row while `$workspaceIsPage` holds —
exactly the placement `444c75c10a` introduced, just not through the
plugin-facing titlebar area. Kanban's board switcher contributes there.
- SDK exports `WORKSPACE_PAGE_HEADER_AREA`; `TITLEBAR_AREAS` documents the
permanent-mount contract.
- Test: a `titleBar.center` component's effect runs setup once and cleanup
never across a chat -> /skills -> chat round trip (red on base).
- Docs: plugin SDK guide covers the lifecycle contract and the page-header
area.
Follow-up to the cherry-picked reconnect cap (#114178, @KoNit-K); supersedes the
earlier #44460 (@plcunha), whose bounded-reconnect design this lands in slimmer form:
- Factor the two give-up branches into ``CDPSupervisor._reconnect_budget_spent``,
which logs the single final warning AND drops the supervisor from
``SUPERVISOR_REGISTRY`` when it is still the registered one. Consumers
(``browser_cdp`` frame routing, the eval fast path) then see "no supervisor
attached" instead of a dead entry, and the next browser call for the task
starts a fresh one via ``get_or_start``.
- Move the unit tests to ``tests/tools/test_browser_supervisor_reconnect.py``
(two invariants: bounded + unregistered after attach; first-dial failure
still fatal for ``start()``) and restore ``test_browser_supervisor.py`` to
main — that file is the opt-in real-Chrome E2E suite and its module-level
skip marker did not need to be rewritten to host unit tests.
- Docs: lifecycle section of the developer guide describes the bounded
reconnect.
Live: fake CDP endpoint killed after attach — base logs connect-failed
warnings indefinitely (7 in 45 s, thread alive, still registered); fixed head
logs 4 + one "stopped after 5 failed reconnect attempts" line, thread exits,
registry entry gone. Control: endpoint back within 2.5 s re-attaches and the
budget resets.
Co-authored-by: plcunha <jvsantos.cunha@gmail.com>
Long-running cron jobs (>60 s, i.e. past one heartbeat interval) intermittently
ended with last_status=error / "Interrupted by shutdown before terminal
completion." while their output was complete. The fire-claim heartbeat thread
took ONE sample that read the claim as not ours, latched lost_ownership, and
_FireOwnership.lost() then trusted that latch without asking the store again —
so a run whose claim still validated (the very condition under which
_record_fire_ownership_lost writes that message) was recorded as interrupted,
its output never saved. The reporter's own error string proves the miss was
transient: it is only ever written when the owner re-validates True.
Why this shape:
- _FireOwnership.lost(): an explicit transport cancel stays terminal; a latched
heartbeat miss is re-checked against the store, and a claim that still
validates keeps the run's real outcome. The owner-fenced mark_job_run and
fire_claim_fence remain the authority, so a genuinely re-owned claim still
yields (no delivery, no terminal write over the new owner, ledger discard).
An unreachable store after a latched miss stays fail-closed.
- _heartbeat_loop: a miss is re-sampled once after
_FIRE_CLAIM_MISS_CONFIRM_SECONDS before it latches. Latching cancels the live
agent run and ends the lease refresh, so a single sample must not do that; a
genuinely re-owned claim misses twice and latches ~1 s later.
The issue's other hypothesis (worker process identity under a systemd-run
--scope worker) is falsified by the code: the owner token is stored in
fire_claim.by at claim time and compared to the stored value, never recomputed
from _machine_id(), and the heartbeat thread runs in a copied context so it
reads the same store.
Slimmer redo of #113364 (same direction: revalidate before trusting the latch)
without its production-dead isinstance(_CombinedCancelEvent) branch.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
The `[Triggering message id: …]` note that `_prepend_inbound_reply_context`
adds for Discord turns is a model instruction (which id to pass to the
discord tools), not something the user wrote. `_hmwa_apply_message_timestamp`
derived `persist_user_message` from the already-wrapped text, so every
Discord-origin user row stored the note as the first line of `content` —
the desktop transcript (and FTS, memory providers, exports) showed the
envelope instead of the message.
- `run_inbound.py`: the note is now the OUTERMOST prefix (after the reply
pointer) and rendered by one function, `discord_triggering_note`;
`strip_discord_triggering_note` peels exactly that prefix for THIS event
off the persisted text, so the `[Replying to: "…"]` pointer survives.
- `run_turn.py::_hmwa_apply_message_timestamp`: persist the stripped text.
The wrapper keeps riding `message_text`; when the durable row differs, the
live bytes land in the replay-only `api_content` sidecar by the existing
persist-override contract (same as timestamps and per-turn sidecar notes).
- `run_turn.py::_run_agent_queued_followup`: the in-band queued follow-up
runs the same inbound prep but passed no persist override; it now carries
the authored text too.
Live: real inbound prep → persist seam → SessionStore.append_to_transcript on
a temp state.db. Before: content='[Triggering message id: `1550…` — use as
`message_id` …]\n\nCreate a project plan for Q4'. After: content='Create a
project plan for Q4', api_content=<wrapped bytes>; reply-pointer and
no-message_id (desktop relay) controls unchanged.
Slimmer redo of #71309 by @JonthanaHanh (regex strip on a `gateway/run.py`
site that no longer exists; tests exercised a copied regex); #71619
(@calvinnwq, +1303/-40 over 10 files) is over the salvage bar.
Fixes#71304Fixes#114719
Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).
- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
_apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
every runtime change: switch_model (outside the rollback guard), fallback activation
and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
`tools/tts_tool.py::_xai_requirements` (the provider check text_to_speech_tool
dispatch consults) still resolved xAI credentials OAuth-first, while both TTS
synthesis paths now pass `prefer_api_key=True`. Truthiness is unchanged, but a
configured key no longer routes the availability probe through the OAuth pool
(pool select / refresh) for a bearer the metered `/v1/tts` endpoint rejects.
Docs: the streaming-TTS provider table now states the ordering.
Part of the #113727 salvage of #113728 (@beardthelion); sibling site the PR
missed.
Co-authored-by: beardthelion <beardthelion@users.noreply.github.com>
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).
resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.
Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
_interrupt_and_clear_session popped the adapter's single pending slot and dropped whatever it
held ("consume and discard", 59575d6a91). That was right when the slot only ever carried the
user's stale follow-up text, but internal wakes (async-delegation completion notices,
kanban/cron notify+wake) now park in the same slot and are claim-settled the moment the adapter
admits them, so the pop lost them for good: the drain that runs after the command found an
empty slot and the session idled until the next user message (#114456, ~6 min stall in the
reported session).
Now the human follow-up is still discarded, but an internal wake stays parked (promoted out of
the overflow FIFO when a discarded human head occupied the slot) and the post-command drain
starts it immediately. Applies to every caller of the helper — /stop (busy fast path, handler,
pending sentinel, thread sibling), /new and /reset — since a wake that arrives a second after
/new runs against the fresh session anyway; whether a completion pinned to the closed session
may run stays with _resolve_async_delegation_session (fail-closed).
Salvaged from #114538 (@whyyagswhy, earliest filer): the pop-and-re-park mechanism is theirs;
widened here to the /new and /reset callers and the overflow promotion, and the invariant tests
rewritten against the real adapter drain. #114540 (@JoaoMarcos44) reached the same fix
independently; its stop-reason taxonomy and overflow analysis informed the class coverage.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
The salvaged mark-then-interrupt path covers every run whose agent exists.
A run whose executor task was created but has not taken its first tick has
no agent to interrupt: _execute_run would re-persist `running` over the
shutdown mark and start a full turn that outlives the gateway (only forced
to `interrupted` when it finally returned). Check the shutdown set before
the `running` transition and finish immediately.
Docs: `interrupted` joins the terminal status vocabulary for /v1/runs.
Co-authored-by: Mark Rockminster <markrockminster@M1-U-48-128.local>
_hmwa_heal_telegram_topic_binding awaits get_telegram_topic_binding and get_compression_tip between reading the route and calling switch_session, the same shape as the async-delegation re-pin fixed in this branch. Pass the snapshot session id as expected_session_id so the compare-and-swap refuses to overwrite a concurrent /new or /resume. The developer-guide Operations table now documents the CAS kwarg.
Two open atoms of #83390 (DeepSeek "This response_format type is unavailable now"):
* `_call_fallback_candidate_sync/_async` only special-cased auth errors, so when the primary
aux provider failed (timeout, rate limit, payment) and the fallback landed on a provider that
rejects `json_schema`, the 400 re-raised and the whole task died — the primary-path rung from
#89589 never applied there. Both fallback paths now retry once without `response_format`.
* Every structured aux call (titles, kanban decomposer, goal judge, plugin structured calls)
paid a guaranteed-fail request on providers that lack `json_schema` before the retry. A
provider profile can now declare `unsupported_response_formats` (DeepSeek: json_schema, per
https://api-docs.deepseek.com/guides/json_mode) and the recovery ladder remembers any route
that rejected a type once (host:port scoped), so `_build_call_kwargs` — shared by the primary
and fallback paths — omits the field before the first request. Dropping rather than
downgrading to json_object matches the end state the retry already produced; json_object
needs a JSON-mentioning prompt and some relays return empty content under it.
New logic lives in agent/auxiliary_structured_output.py; the facade only gains the fallback rung
next to the predicate it uses. tests/agent/conftest.py resets the process-level memo per test.
Fixes#83390, #105191. Closes duplicates #84976, #88830, #102849, #113064.
Co-authored-by: Legion-is-life <Legion-is-life@users.noreply.github.com>
Why: an integration that answered server→client requests but never sent
client.capabilities now has every clarify/approval/sudo/… request refused at
once with no grace path. Marking a connection as answering after its first
response frame cannot help — the frame is never sent to it — so the break is
documented instead.
Item 2 of #112548: a Desktop/dashboard build that predates server→client
requests has no response path, so every clarify/approval/sudo/secret/vault/
connection/bridge request sat for the full deadline (clarify: 300s). Only the
tour probed. Clients now advertise once per connection
(`client.capabilities {server_requests: true}`, sent by the shared TypeScript
channel on `gateway.ready`); `send()` / `send_async()` return the
error-response shape (None) at once when every WebSocket peer of the session
is a build that never advertised. Sessions with no client attached still wait
so the reconnect replay (`open_requests`) keeps working; the stdio TUI ships
with the backend and is not gated. The advertisement is dropped on disconnect.
Reviewer minors from #113227:
- tools/approval_gateway_wait.py: the verdict is the choice committed under
the approval lock while leaving the queue, so an /approve that lands after
the deadline check but before the entry is dropped is an answer, not a
timeout (the client was already acked "ok").
- tests/tui_gateway/test_protocol.py: the error-fails-fast test that only
restated pre-existing behaviour is replaced by the two capability
invariants (never advertised → fails fast; advertised → frame written,
waits, forgotten on disconnect).
- server_requests.send try/finally around event.wait already landed on main
(4371ed34a9); nothing to change.
Docs: programmatic-integration.md (advertise once per connection; method
list), tui_gateway/AGENTS.md; contracts regenerated.
The warn-once reporter added for hooks (and for PluginManager.invoke_middleware)
left the execution chain out: hermes_cli/middleware.py::_run_execution_chain — the
tool_execution / llm_execution frames that run once per tool or LLM call — still
logged "Middleware '%s' callback %s raised" at WARNING on every invocation. A
mis-declared callback (e.g. a signature naming tool_data) therefore flooded the
log exactly like the hook case #111922 reported: 5 calls -> 5 WARNING lines.
Route the frame's except through the manager's _report_hook_failure with the
"Middleware" surface label, so the first failure warns (listing the fields the
middleware does provide) and identical repeats go to DEBUG; the set is already
forgotten on plugin reload. The frame's skip-and-continue semantics are unchanged.
Part of #111922
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).
WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.
Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.
Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
A directory plugin's pyproject.toml [project].dependencies (or manifest
python_dependencies) are now installed into the Hermes venv on install/enable/
update, resolved under a constraints file built from Hermes' own pinned
dependencies together with every enabled peer's declarations. A candidate that
cannot resolve is refused before its tree is moved into place; nothing is
installed and no other plugin is touched. --no-deps opts a single install out.
hermes update rebuilds the venv from Hermes' lock and strips everything else, so
its git, zip and both repair paths now re-apply the union of every profile's
enabled plugins. When the union no longer resolves, each plugin is resolved on
its own so only culprits are dropped (never an alphabetical neighbour); a
plugin-vs-plugin conflict peels non-memory plugins first because a Hermes that
boots without memory reads as data loss. Dropped plugins are disabled through
the real config writer with a loud message naming the fix.
python_runtime: external lets sidecar-venv plugins (Mnemosyne's shape) opt out
of the union; hermes-agent self-dependencies and direct-URL requirements are
never installed (the latter are surfaced for the user to install by hand).
A cron agent runs as platform "cron", so its system prompt only ever got
PLATFORM_HINTS["cron"]. The job's final response lands on the `deliver`
channel (Slack, Telegram, ...), but the model was never told what renders
there: no MEDIA: attachment instructions, no "tables don't render", and a
user's `platform_hints.slack.append` never reached scheduled jobs at all
(community report: Slack-delivered cron jobs stopped sending attachments).
platform_hint() now appends "Delivery destination (<channel>): <hint>" for
cron agents, where <hint> is the channel's built-in/plugin text resolved
through the same `platform_hints.<channel>` override path a live session
uses. The channel comes from the HERMES_CRON_AUTO_DELIVER_PLATFORM
ContextVar the scheduler publishes before the agent runs (the seam
send_message already routes by). Non-cron platforms are untouched.
Desktop's appendLiveSessionProjection built the `user-inflight-<sid>` row from
`inflight.user` text alone, so a background-process completion (or async
delegation result) that was still running at reconnect showed up as a user
bubble even though history would render the same turn as a `process_complete`
timeline marker once it landed. The gateway now ships `display_kind` /
`display_metadata` on `inflight` (#112162); route a typed row through the same
`toChatMessages` projection history uses, keeping the stable id so the
activate/resume reconcile keeps deduping it. `hidden` prompts yield no row.
Genuine user input carries no display fields and renders unchanged, so a user
quoting the marker text stays a user bubble.
Docs: describe the `inflight` display projection in the programmatic
integration guide.
Part of #112144 (with the salvaged #112162 by @KoNit-K).
Every hard interrupt reached `begin_iteration` through the same `_interrupt_requested`
flag, so the turn loop booked all of them as `interrupted_by_user` — a cron run killed
by the scheduler's inactivity watchdog, a turn aborted by the liveness watchdog, a lost
session turn lease and a gateway inactivity timeout all read as a human pressing stop,
and the investigation went to the wrong subsystem (#112647).
`interrupt()` already records a trusted category per interrupt (`_tool_interrupt_reason`,
fed by the `tool_reason` every producer can pass through `request_hard_interrupt`). The
exit reason is now derived from it: the three categories `interrupt()` itself mints for
human stops keep `interrupted_by_user` / `interrupted_during_api_call`; any other
category names its producer — `interrupted_by_system(cron_inactivity_watchdog)`,
`interrupted_during_api_call(turn_liveness_watchdog)`. The system producers that passed
no `tool_reason` (cron inactivity watchdog, turn liveness watchdog, lease loss, gateway
inactivity timeout) now name themselves, so the model-visible tool-cancellation text
says the same thing. `_publish_interrupt_state` logs ONE line naming the source so the
turn record and the log agree.
No change to WHEN anything interrupts. `interrupted_during_api_call` moves to the
prefix-matched explanation table so the parameterised form keeps its user copy.
Slim redo of #112652 by @KoNit-K, which added a parallel `issuer` attribute and
keyword; this reuses the existing `tool_reason` plumbing instead.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The cherry-picked fix scoped host.onEvent disposers into the runtime
loader's per-plugin disposer list. This finishes the class:
- apps/desktop/src/contrib/plugins.ts: the bundled loader's activate()
wraps register() in the same trackGatewayEventDisposers scope, so a
Capabilities > Plugins disable/re-enable cycle no longer strands a
bundled plugin's gateway listeners either (same accumulate-per-toggle
mechanism as the disk hot reload).
- apps/desktop/src/contrib/plugin.ts: PluginContext.onEvent — a tracked
door for subscriptions made AFTER register() returns (timers, socket
callbacks), where the register-time scope cannot see them and a bare
host.onEvent still needs hand-wiring to ctx.onDispose.
- apps/desktop/src/sdk/index.ts + website desktop-plugin-sdk doc: state
the contract plugin authors can now rely on.
- Tests: one ctx.onEvent invariant (red on base: TypeError) in
plugin.test.ts; the contributor's events.test.ts control case dropped
(green on base, no invariant) to keep the fix at two tests.
Why: unloadRuntimePlugin / deactivate only run ctx-tracked disposers;
anything a plugin subscribes outside track() survives every reload, and
one relay marker then opens N sessions (#112366).
Fixes#112366
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
A clean ``stop`` with empty content and reasoning text still promotes the
reasoning to ``final_response`` (the vLLM nemotron parser files the whole
answer as reasoning; re-running the empty-response ladder re-billed the prompt
8x, #109205). What changes is what the assistant ROW carries: ``content`` stays
empty, the text lives in ``reasoning``/``reasoning_content``, and the promoted
text rides the ``api_content`` sidecar. ``build_api_messages`` substitutes the
sidecar on the wire, so the next request replays the answer byte-identically
(same bytes as before this change) and role alternation is intact, while
state.db, ``session.history`` and every other history surface no longer show
chain-of-thought as a reply indistinguishable from a real one.
The "Reasoning-only clean stop" log moves INFO -> WARNING and names model,
provider, api_calls and tool_turns: a model that ends every turn this way is
stalled (planning monologue, zero tool calls) while the turn reported
"complete".
The thinking-prefill interim row (``_thinking_prefill``) is ephemeral
scaffolding already skipped by the flush and popped before the final answer, so
it never persisted a content==reasoning row; no change there.
Live probe (real AIAgent, real SessionDB, temp HERMES_HOME, DeepSeek-shaped
reasoning_content only): before, the assistant row had content == reasoning
== reasoning_content with api_content NULL and an INFO log; after, content is
empty, api_content carries the text, the replayed second turn sends the same
assistant content on the wire, the WARNING carries the route, and a normal
"Hi there." reply persists unchanged.
Part of #111761
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
A SOUL.md in HERMES_HOME that *documents* the canonical injection phrase as
security guidance ("...content telling you to ignore previous instructions...")
was replaced wholesale by a [BLOCKED: SOUL.md ...] marker, so the agent ran
with no identity/constitution and only a one-line agent.log warning said so.
SOUL.md is the user's own file: agent writes to it always go through the
protected-instruction approval gate (tools/file_tools_write_guards.py) and no
repository checkout can plant it, so it sits in the same trust class as
config.yaml — not a cloned repo's AGENTS.md. `_scan_context_content` gains a
`user_authored` mode that still scans, logs the matched pattern(s) at WARNING
and loads the content; only `load_soul_md` uses it. Project-dir context files
(.hermes.md, AGENTS.md and the other project-dir files), subdirectory hints,
memory and tool-result scanning are unchanged and keep blocking. No threat
pattern is narrowed: the reported-speech form cannot be separated from real
attacks by regex ("I want you to ignore all previous instructions" is a
canonical payload).
The /context manifest reports such a file as `flagged` (loaded, ⚠ "review the
file") so the warning is visible on CLI, TUI, gateway and Desktop rather than
buried in the log.
Fixes#112570
Extract the contributor's staging + rename sequence into
`publishDesktopTree` and make `installDesktopPluginFromGit` use it too:
its `copyDesktopTree` had the same mkdir + cp shape, so a copy that failed
half-way left an empty `<root>/<name>/` and every later install attempt
was refused with "already exists. Enable force reinstall to replace it".
One helper, both copy sites, same guarantee: a published folder is always
complete (marker included) or absent.
Docs: describe the staged copy and the marker-less/plugin.js rule in the
desktop plugin SDK guide.
* fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping
systemd already stops restarting on exit 78; launchd KeepAlive=true
respawned the same token/port collision every 30s. Map 78 to a clean
stop under SuccessfulExit=false so the job stays down until the holder
releases the lock.
* test(launchd): pin EX_CONFIG park under SuccessfulExit KeepAlive
The plist must not use unconditional KeepAlive, and the stderr wrapper
must turn gateway exit 78 into a clean stop without swallowing exit 75
or a non-gateway child's 78.
DEFAULT_CONFIG now ships gateway.multiplex_profiles: true. GatewayConfig keeps an UNSET flag as
None so the boot can tell "the operator chose" from "the default applies"; every reader tests
truthiness, so an undecided flag never multiplexes by accident.
hermes_cli/gateway_multiplex_mode.py settles the unset default once per boot (called from
load_gateway_config_for_runner and `gateway run --config`): the same preflight `hermes gateway
migrate --multiplex` runs — default profile, >= 2 profiles, no secondary running its own gateway
(live pid or installed unit), no duplicate-credential / port-binder blocker, migratable host. A
refusal is a logged warning naming the blocker and the migrate one-liner; the gateway comes up
standalone exactly as before. Explicit values (config.yaml, GATEWAY_MULTIPLEX_PROFILES) pass through
verbatim; `--standalone` already pins false.
Other processes stop guessing the verdict from the merged default: named_profile_served_by_running
_multiplexer, the enroll warning, the dashboard listener guard, the cron-fire port resolver, container
boot and the migration plan (_read_multiplex_flag) read the live gateway's served_profiles record
first and the EXPLICIT flag second — so a per-profile fleet with the flag unset still reads as
"not yet multiplexed" and the fold proceeds.
Docs: multi-profile-gateways.md, multiplexing-gateway.md, hermes_cli/AGENTS.md.
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.
- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
approval prompt could not be delivered or was not answered (<cause>)" with
outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.
Fixes#22992
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).
The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:
- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
category — no string matching) for the interrupted state and marks a
notifier-unregister wake (event set, result None) as "the turn ended before
the prompt was answered". Both the direct and the coalesced-follower wait
return `cancelled=<cause>`; the post_approval_response hook fires
choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
"BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
with outcome="cancelled" and its own user_summary; an explicit /deny is
untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
tool_reason ("parent delegation ended"; the late-child mirror forwards the
parent's own category) so a child's pending approval can tell teardown from a
user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
instruction-file gate and MCP elicitation consume the same key instead of
reporting "denied by the user" / "decline".
Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.
The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.
`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.
CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.
Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>