Follow-up to the cherry-picked reconnect cap (#114178, @KoNit-K); supersedes the
earlier #44460 (@plcunha), whose bounded-reconnect design this lands in slimmer form:
- Factor the two give-up branches into ``CDPSupervisor._reconnect_budget_spent``,
which logs the single final warning AND drops the supervisor from
``SUPERVISOR_REGISTRY`` when it is still the registered one. Consumers
(``browser_cdp`` frame routing, the eval fast path) then see "no supervisor
attached" instead of a dead entry, and the next browser call for the task
starts a fresh one via ``get_or_start``.
- Move the unit tests to ``tests/tools/test_browser_supervisor_reconnect.py``
(two invariants: bounded + unregistered after attach; first-dial failure
still fatal for ``start()``) and restore ``test_browser_supervisor.py`` to
main — that file is the opt-in real-Chrome E2E suite and its module-level
skip marker did not need to be rewritten to host unit tests.
- Docs: lifecycle section of the developer guide describes the bounded
reconnect.
Live: fake CDP endpoint killed after attach — base logs connect-failed
warnings indefinitely (7 in 45 s, thread alive, still registered); fixed head
logs 4 + one "stopped after 5 failed reconnect attempts" line, thread exits,
registry entry gone. Control: endpoint back within 2.5 s re-attaches and the
budget resets.
Co-authored-by: plcunha <jvsantos.cunha@gmail.com>
Long-running cron jobs (>60 s, i.e. past one heartbeat interval) intermittently
ended with last_status=error / "Interrupted by shutdown before terminal
completion." while their output was complete. The fire-claim heartbeat thread
took ONE sample that read the claim as not ours, latched lost_ownership, and
_FireOwnership.lost() then trusted that latch without asking the store again —
so a run whose claim still validated (the very condition under which
_record_fire_ownership_lost writes that message) was recorded as interrupted,
its output never saved. The reporter's own error string proves the miss was
transient: it is only ever written when the owner re-validates True.
Why this shape:
- _FireOwnership.lost(): an explicit transport cancel stays terminal; a latched
heartbeat miss is re-checked against the store, and a claim that still
validates keeps the run's real outcome. The owner-fenced mark_job_run and
fire_claim_fence remain the authority, so a genuinely re-owned claim still
yields (no delivery, no terminal write over the new owner, ledger discard).
An unreachable store after a latched miss stays fail-closed.
- _heartbeat_loop: a miss is re-sampled once after
_FIRE_CLAIM_MISS_CONFIRM_SECONDS before it latches. Latching cancels the live
agent run and ends the lease refresh, so a single sample must not do that; a
genuinely re-owned claim misses twice and latches ~1 s later.
The issue's other hypothesis (worker process identity under a systemd-run
--scope worker) is falsified by the code: the owner token is stored in
fire_claim.by at claim time and compared to the stored value, never recomputed
from _machine_id(), and the heartbeat thread runs in a copied context so it
reads the same store.
Slimmer redo of #113364 (same direction: revalidate before trusting the latch)
without its production-dead isinstance(_CombinedCancelEvent) branch.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
The `[Triggering message id: …]` note that `_prepend_inbound_reply_context`
adds for Discord turns is a model instruction (which id to pass to the
discord tools), not something the user wrote. `_hmwa_apply_message_timestamp`
derived `persist_user_message` from the already-wrapped text, so every
Discord-origin user row stored the note as the first line of `content` —
the desktop transcript (and FTS, memory providers, exports) showed the
envelope instead of the message.
- `run_inbound.py`: the note is now the OUTERMOST prefix (after the reply
pointer) and rendered by one function, `discord_triggering_note`;
`strip_discord_triggering_note` peels exactly that prefix for THIS event
off the persisted text, so the `[Replying to: "…"]` pointer survives.
- `run_turn.py::_hmwa_apply_message_timestamp`: persist the stripped text.
The wrapper keeps riding `message_text`; when the durable row differs, the
live bytes land in the replay-only `api_content` sidecar by the existing
persist-override contract (same as timestamps and per-turn sidecar notes).
- `run_turn.py::_run_agent_queued_followup`: the in-band queued follow-up
runs the same inbound prep but passed no persist override; it now carries
the authored text too.
Live: real inbound prep → persist seam → SessionStore.append_to_transcript on
a temp state.db. Before: content='[Triggering message id: `1550…` — use as
`message_id` …]\n\nCreate a project plan for Q4'. After: content='Create a
project plan for Q4', api_content=<wrapped bytes>; reply-pointer and
no-message_id (desktop relay) controls unchanged.
Slimmer redo of #71309 by @JonthanaHanh (regex strip on a `gateway/run.py`
site that no longer exists; tests exercised a copied regex); #71619
(@calvinnwq, +1303/-40 over 10 files) is over the salvage bar.
Fixes#71304Fixes#114719
Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).
- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
_apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
every runtime change: switch_model (outside the rollback guard), fallback activation
and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
`tools/tts_tool.py::_xai_requirements` (the provider check text_to_speech_tool
dispatch consults) still resolved xAI credentials OAuth-first, while both TTS
synthesis paths now pass `prefer_api_key=True`. Truthiness is unchanged, but a
configured key no longer routes the availability probe through the OAuth pool
(pool select / refresh) for a bearer the metered `/v1/tts` endpoint rejects.
Docs: the streaming-TTS provider table now states the ordering.
Part of the #113727 salvage of #113728 (@beardthelion); sibling site the PR
missed.
Co-authored-by: beardthelion <beardthelion@users.noreply.github.com>
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).
resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.
Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
_interrupt_and_clear_session popped the adapter's single pending slot and dropped whatever it
held ("consume and discard", 59575d6a91). That was right when the slot only ever carried the
user's stale follow-up text, but internal wakes (async-delegation completion notices,
kanban/cron notify+wake) now park in the same slot and are claim-settled the moment the adapter
admits them, so the pop lost them for good: the drain that runs after the command found an
empty slot and the session idled until the next user message (#114456, ~6 min stall in the
reported session).
Now the human follow-up is still discarded, but an internal wake stays parked (promoted out of
the overflow FIFO when a discarded human head occupied the slot) and the post-command drain
starts it immediately. Applies to every caller of the helper — /stop (busy fast path, handler,
pending sentinel, thread sibling), /new and /reset — since a wake that arrives a second after
/new runs against the fresh session anyway; whether a completion pinned to the closed session
may run stays with _resolve_async_delegation_session (fail-closed).
Salvaged from #114538 (@whyyagswhy, earliest filer): the pop-and-re-park mechanism is theirs;
widened here to the /new and /reset callers and the overflow promotion, and the invariant tests
rewritten against the real adapter drain. #114540 (@JoaoMarcos44) reached the same fix
independently; its stop-reason taxonomy and overflow analysis informed the class coverage.
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
The salvaged mark-then-interrupt path covers every run whose agent exists.
A run whose executor task was created but has not taken its first tick has
no agent to interrupt: _execute_run would re-persist `running` over the
shutdown mark and start a full turn that outlives the gateway (only forced
to `interrupted` when it finally returned). Check the shutdown set before
the `running` transition and finish immediately.
Docs: `interrupted` joins the terminal status vocabulary for /v1/runs.
Co-authored-by: Mark Rockminster <markrockminster@M1-U-48-128.local>
_hmwa_heal_telegram_topic_binding awaits get_telegram_topic_binding and get_compression_tip between reading the route and calling switch_session, the same shape as the async-delegation re-pin fixed in this branch. Pass the snapshot session id as expected_session_id so the compare-and-swap refuses to overwrite a concurrent /new or /resume. The developer-guide Operations table now documents the CAS kwarg.
Two open atoms of #83390 (DeepSeek "This response_format type is unavailable now"):
* `_call_fallback_candidate_sync/_async` only special-cased auth errors, so when the primary
aux provider failed (timeout, rate limit, payment) and the fallback landed on a provider that
rejects `json_schema`, the 400 re-raised and the whole task died — the primary-path rung from
#89589 never applied there. Both fallback paths now retry once without `response_format`.
* Every structured aux call (titles, kanban decomposer, goal judge, plugin structured calls)
paid a guaranteed-fail request on providers that lack `json_schema` before the retry. A
provider profile can now declare `unsupported_response_formats` (DeepSeek: json_schema, per
https://api-docs.deepseek.com/guides/json_mode) and the recovery ladder remembers any route
that rejected a type once (host:port scoped), so `_build_call_kwargs` — shared by the primary
and fallback paths — omits the field before the first request. Dropping rather than
downgrading to json_object matches the end state the retry already produced; json_object
needs a JSON-mentioning prompt and some relays return empty content under it.
New logic lives in agent/auxiliary_structured_output.py; the facade only gains the fallback rung
next to the predicate it uses. tests/agent/conftest.py resets the process-level memo per test.
Fixes#83390, #105191. Closes duplicates #84976, #88830, #102849, #113064.
Co-authored-by: Legion-is-life <Legion-is-life@users.noreply.github.com>
Why: an integration that answered server→client requests but never sent
client.capabilities now has every clarify/approval/sudo/… request refused at
once with no grace path. Marking a connection as answering after its first
response frame cannot help — the frame is never sent to it — so the break is
documented instead.
Item 2 of #112548: a Desktop/dashboard build that predates server→client
requests has no response path, so every clarify/approval/sudo/secret/vault/
connection/bridge request sat for the full deadline (clarify: 300s). Only the
tour probed. Clients now advertise once per connection
(`client.capabilities {server_requests: true}`, sent by the shared TypeScript
channel on `gateway.ready`); `send()` / `send_async()` return the
error-response shape (None) at once when every WebSocket peer of the session
is a build that never advertised. Sessions with no client attached still wait
so the reconnect replay (`open_requests`) keeps working; the stdio TUI ships
with the backend and is not gated. The advertisement is dropped on disconnect.
Reviewer minors from #113227:
- tools/approval_gateway_wait.py: the verdict is the choice committed under
the approval lock while leaving the queue, so an /approve that lands after
the deadline check but before the entry is dropped is an answer, not a
timeout (the client was already acked "ok").
- tests/tui_gateway/test_protocol.py: the error-fails-fast test that only
restated pre-existing behaviour is replaced by the two capability
invariants (never advertised → fails fast; advertised → frame written,
waits, forgotten on disconnect).
- server_requests.send try/finally around event.wait already landed on main
(4371ed34a9); nothing to change.
Docs: programmatic-integration.md (advertise once per connection; method
list), tui_gateway/AGENTS.md; contracts regenerated.
The warn-once reporter added for hooks (and for PluginManager.invoke_middleware)
left the execution chain out: hermes_cli/middleware.py::_run_execution_chain — the
tool_execution / llm_execution frames that run once per tool or LLM call — still
logged "Middleware '%s' callback %s raised" at WARNING on every invocation. A
mis-declared callback (e.g. a signature naming tool_data) therefore flooded the
log exactly like the hook case #111922 reported: 5 calls -> 5 WARNING lines.
Route the frame's except through the manager's _report_hook_failure with the
"Middleware" surface label, so the first failure warns (listing the fields the
middleware does provide) and identical repeats go to DEBUG; the set is already
forgotten on plugin reload. The frame's skip-and-continue semantics are unchanged.
Part of #111922
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).
WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.
Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.
Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
A directory plugin's pyproject.toml [project].dependencies (or manifest
python_dependencies) are now installed into the Hermes venv on install/enable/
update, resolved under a constraints file built from Hermes' own pinned
dependencies together with every enabled peer's declarations. A candidate that
cannot resolve is refused before its tree is moved into place; nothing is
installed and no other plugin is touched. --no-deps opts a single install out.
hermes update rebuilds the venv from Hermes' lock and strips everything else, so
its git, zip and both repair paths now re-apply the union of every profile's
enabled plugins. When the union no longer resolves, each plugin is resolved on
its own so only culprits are dropped (never an alphabetical neighbour); a
plugin-vs-plugin conflict peels non-memory plugins first because a Hermes that
boots without memory reads as data loss. Dropped plugins are disabled through
the real config writer with a loud message naming the fix.
python_runtime: external lets sidecar-venv plugins (Mnemosyne's shape) opt out
of the union; hermes-agent self-dependencies and direct-URL requirements are
never installed (the latter are surfaced for the user to install by hand).
A cron agent runs as platform "cron", so its system prompt only ever got
PLATFORM_HINTS["cron"]. The job's final response lands on the `deliver`
channel (Slack, Telegram, ...), but the model was never told what renders
there: no MEDIA: attachment instructions, no "tables don't render", and a
user's `platform_hints.slack.append` never reached scheduled jobs at all
(community report: Slack-delivered cron jobs stopped sending attachments).
platform_hint() now appends "Delivery destination (<channel>): <hint>" for
cron agents, where <hint> is the channel's built-in/plugin text resolved
through the same `platform_hints.<channel>` override path a live session
uses. The channel comes from the HERMES_CRON_AUTO_DELIVER_PLATFORM
ContextVar the scheduler publishes before the agent runs (the seam
send_message already routes by). Non-cron platforms are untouched.
Desktop's appendLiveSessionProjection built the `user-inflight-<sid>` row from
`inflight.user` text alone, so a background-process completion (or async
delegation result) that was still running at reconnect showed up as a user
bubble even though history would render the same turn as a `process_complete`
timeline marker once it landed. The gateway now ships `display_kind` /
`display_metadata` on `inflight` (#112162); route a typed row through the same
`toChatMessages` projection history uses, keeping the stable id so the
activate/resume reconcile keeps deduping it. `hidden` prompts yield no row.
Genuine user input carries no display fields and renders unchanged, so a user
quoting the marker text stays a user bubble.
Docs: describe the `inflight` display projection in the programmatic
integration guide.
Part of #112144 (with the salvaged #112162 by @KoNit-K).
Every hard interrupt reached `begin_iteration` through the same `_interrupt_requested`
flag, so the turn loop booked all of them as `interrupted_by_user` — a cron run killed
by the scheduler's inactivity watchdog, a turn aborted by the liveness watchdog, a lost
session turn lease and a gateway inactivity timeout all read as a human pressing stop,
and the investigation went to the wrong subsystem (#112647).
`interrupt()` already records a trusted category per interrupt (`_tool_interrupt_reason`,
fed by the `tool_reason` every producer can pass through `request_hard_interrupt`). The
exit reason is now derived from it: the three categories `interrupt()` itself mints for
human stops keep `interrupted_by_user` / `interrupted_during_api_call`; any other
category names its producer — `interrupted_by_system(cron_inactivity_watchdog)`,
`interrupted_during_api_call(turn_liveness_watchdog)`. The system producers that passed
no `tool_reason` (cron inactivity watchdog, turn liveness watchdog, lease loss, gateway
inactivity timeout) now name themselves, so the model-visible tool-cancellation text
says the same thing. `_publish_interrupt_state` logs ONE line naming the source so the
turn record and the log agree.
No change to WHEN anything interrupts. `interrupted_during_api_call` moves to the
prefix-matched explanation table so the parameterised form keeps its user copy.
Slim redo of #112652 by @KoNit-K, which added a parallel `issuer` attribute and
keyword; this reuses the existing `tool_reason` plumbing instead.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The cherry-picked fix scoped host.onEvent disposers into the runtime
loader's per-plugin disposer list. This finishes the class:
- apps/desktop/src/contrib/plugins.ts: the bundled loader's activate()
wraps register() in the same trackGatewayEventDisposers scope, so a
Capabilities > Plugins disable/re-enable cycle no longer strands a
bundled plugin's gateway listeners either (same accumulate-per-toggle
mechanism as the disk hot reload).
- apps/desktop/src/contrib/plugin.ts: PluginContext.onEvent — a tracked
door for subscriptions made AFTER register() returns (timers, socket
callbacks), where the register-time scope cannot see them and a bare
host.onEvent still needs hand-wiring to ctx.onDispose.
- apps/desktop/src/sdk/index.ts + website desktop-plugin-sdk doc: state
the contract plugin authors can now rely on.
- Tests: one ctx.onEvent invariant (red on base: TypeError) in
plugin.test.ts; the contributor's events.test.ts control case dropped
(green on base, no invariant) to keep the fix at two tests.
Why: unloadRuntimePlugin / deactivate only run ctx-tracked disposers;
anything a plugin subscribes outside track() survives every reload, and
one relay marker then opens N sessions (#112366).
Fixes#112366
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
A clean ``stop`` with empty content and reasoning text still promotes the
reasoning to ``final_response`` (the vLLM nemotron parser files the whole
answer as reasoning; re-running the empty-response ladder re-billed the prompt
8x, #109205). What changes is what the assistant ROW carries: ``content`` stays
empty, the text lives in ``reasoning``/``reasoning_content``, and the promoted
text rides the ``api_content`` sidecar. ``build_api_messages`` substitutes the
sidecar on the wire, so the next request replays the answer byte-identically
(same bytes as before this change) and role alternation is intact, while
state.db, ``session.history`` and every other history surface no longer show
chain-of-thought as a reply indistinguishable from a real one.
The "Reasoning-only clean stop" log moves INFO -> WARNING and names model,
provider, api_calls and tool_turns: a model that ends every turn this way is
stalled (planning monologue, zero tool calls) while the turn reported
"complete".
The thinking-prefill interim row (``_thinking_prefill``) is ephemeral
scaffolding already skipped by the flush and popped before the final answer, so
it never persisted a content==reasoning row; no change there.
Live probe (real AIAgent, real SessionDB, temp HERMES_HOME, DeepSeek-shaped
reasoning_content only): before, the assistant row had content == reasoning
== reasoning_content with api_content NULL and an INFO log; after, content is
empty, api_content carries the text, the replayed second turn sends the same
assistant content on the wire, the WARNING carries the route, and a normal
"Hi there." reply persists unchanged.
Part of #111761
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
A SOUL.md in HERMES_HOME that *documents* the canonical injection phrase as
security guidance ("...content telling you to ignore previous instructions...")
was replaced wholesale by a [BLOCKED: SOUL.md ...] marker, so the agent ran
with no identity/constitution and only a one-line agent.log warning said so.
SOUL.md is the user's own file: agent writes to it always go through the
protected-instruction approval gate (tools/file_tools_write_guards.py) and no
repository checkout can plant it, so it sits in the same trust class as
config.yaml — not a cloned repo's AGENTS.md. `_scan_context_content` gains a
`user_authored` mode that still scans, logs the matched pattern(s) at WARNING
and loads the content; only `load_soul_md` uses it. Project-dir context files
(.hermes.md, AGENTS.md and the other project-dir files), subdirectory hints,
memory and tool-result scanning are unchanged and keep blocking. No threat
pattern is narrowed: the reported-speech form cannot be separated from real
attacks by regex ("I want you to ignore all previous instructions" is a
canonical payload).
The /context manifest reports such a file as `flagged` (loaded, ⚠ "review the
file") so the warning is visible on CLI, TUI, gateway and Desktop rather than
buried in the log.
Fixes#112570
Extract the contributor's staging + rename sequence into
`publishDesktopTree` and make `installDesktopPluginFromGit` use it too:
its `copyDesktopTree` had the same mkdir + cp shape, so a copy that failed
half-way left an empty `<root>/<name>/` and every later install attempt
was refused with "already exists. Enable force reinstall to replace it".
One helper, both copy sites, same guarantee: a published folder is always
complete (marker included) or absent.
Docs: describe the staged copy and the marker-less/plugin.js rule in the
desktop plugin SDK guide.
* fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping
systemd already stops restarting on exit 78; launchd KeepAlive=true
respawned the same token/port collision every 30s. Map 78 to a clean
stop under SuccessfulExit=false so the job stays down until the holder
releases the lock.
* test(launchd): pin EX_CONFIG park under SuccessfulExit KeepAlive
The plist must not use unconditional KeepAlive, and the stderr wrapper
must turn gateway exit 78 into a clean stop without swallowing exit 75
or a non-gateway child's 78.
DEFAULT_CONFIG now ships gateway.multiplex_profiles: true. GatewayConfig keeps an UNSET flag as
None so the boot can tell "the operator chose" from "the default applies"; every reader tests
truthiness, so an undecided flag never multiplexes by accident.
hermes_cli/gateway_multiplex_mode.py settles the unset default once per boot (called from
load_gateway_config_for_runner and `gateway run --config`): the same preflight `hermes gateway
migrate --multiplex` runs — default profile, >= 2 profiles, no secondary running its own gateway
(live pid or installed unit), no duplicate-credential / port-binder blocker, migratable host. A
refusal is a logged warning naming the blocker and the migrate one-liner; the gateway comes up
standalone exactly as before. Explicit values (config.yaml, GATEWAY_MULTIPLEX_PROFILES) pass through
verbatim; `--standalone` already pins false.
Other processes stop guessing the verdict from the merged default: named_profile_served_by_running
_multiplexer, the enroll warning, the dashboard listener guard, the cron-fire port resolver, container
boot and the migration plan (_read_multiplex_flag) read the live gateway's served_profiles record
first and the EXPLICIT flag second — so a per-profile fleet with the flag unset still reads as
"not yet multiplexed" and the fold proceeds.
Docs: multi-profile-gateways.md, multiplexing-gateway.md, hermes_cli/AGENTS.md.
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.
- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
approval prompt could not be delivered or was not answered (<cause>)" with
outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.
Fixes#22992
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).
The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:
- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
category — no string matching) for the interrupted state and marks a
notifier-unregister wake (event set, result None) as "the turn ended before
the prompt was answered". Both the direct and the coalesced-follower wait
return `cancelled=<cause>`; the post_approval_response hook fires
choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
"BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
with outcome="cancelled" and its own user_summary; an explicit /deny is
untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
tool_reason ("parent delegation ended"; the late-child mirror forwards the
parent's own category) so a child's pending approval can tell teardown from a
user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
instruction-file gate and MCP elicitation consume the same key instead of
reporting "denied by the user" / "decline".
Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.
The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.
`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.
CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.
Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.
Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).
Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The budget in _sweep_agent_cache_under_pressure comes from the gateway's own
cgroup memory.high/memory.max, but read_anon_rss_mb() read /proc/self/status
RssAnon: only the main process. Every child in the unit (execute_code kernels,
terminal commands) is charged against the same limit, so a kernel at 4.7 GiB
pushed the unit to MemoryHigh while the valve saw <1.6 GiB and never fired;
systemd's stop then SIGKILLed the gateway mid-flush (the #80764 signature).
Under a capped cgroup v2, read own memory.stat `anon` (the same scope as the
budget); uncapped, or when the file is unreadable, keep the self reading.
_finite_limit is factored out of _cgroup_limit_bytes so the cap check and
the budget agree on what "unlimited" means.
Fixes#110549
Corrections on top of #110617: hermes_state._ensure_test_isolation raises RuntimeError (the
draft said 'warning'); the older Database Location section still told readers the default
path is ~/.hermes/state.db, which the new section says never to hard-code. Mirror the new
section into the zh-Hans page.
`gateway.allow_all_users: true` (and the top-level spelling) in config.yaml
was a silent no-op: `_TOPLEVEL_BRIDGE` never forwarded it, GatewayConfig has
no field, and every allow-all reader (authz mixin default-deny branch,
startup access check, own-policy adapters, Discord/Matrix/Email plugin
gates) consults the GATEWAY_ALLOW_ALL_USERS env var only.
Bridge the YAML key into that env var in `bridge_core_env_settings`, the
one seam every reader already shares, instead of threading a new config
attribute through ten readers:
- first-writer-wins: an explicit env var beats YAML (matches every other
{PLATFORM}_* gate);
- skipped inside a multiplexed secondary profile's scope (#80099 class):
the secondary's config.yaml must not become the default profile's policy,
and the isolation test now asserts GATEWAY_ALLOW_ALL_USERS stays unset;
- a startup warning names config.yaml as the grant source, because the key
was inert until now and a forgotten `true` flips the posture to open.
Tests: both spellings authorize a stranger through `_is_user_authorized`;
`false`, absent key, and env=false over YAML=true all stay denied.
Docs: security guide, env-var reference, gateway internals.
Fixes#110690