Commit Graph

379 Commits

Author SHA1 Message Date
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
Konstantin Khlopkov
ece262150a docs(memory-provider): state that no bundled provider advertises checkpoint API v2
Operators enabling compression.checkpoint_required need to know up front that
the contract is opt-in for third-party archiving providers; otherwise the flag
blocks every compression attempt and the reason is only visible after the fact.

Docs hunk of #106882 (the code change was consolidated with #106879).
2026-09-18 13:58:52 -07:00
kshitijk4poor
9e6f02538e docs(memory-provider): agent_context carries cron/subagent, not always primary 2026-09-19 00:55:05 +05:30
teknium1
aee7de4db5 fix(desktop): point the remote-backend desktop-half tooltip at the install path
The honest "unavailable (remote backend)" state (previous commit) tells the
user the copy will never happen; the tooltip now also says what does work
against a remote backend — Install from Git with the Desktop target checked
clones the desktop half onto this machine (the install modal already takes
that branch for connection.mode === 'remote'). Docs note the new state next to
the existing remote-backend paragraph.
2026-09-18 10:53:41 -07:00
teknium1
66ba4e5114 test(tui_gateway): pin that a live-turn rewind is not redirected; document the 4009 contract
The contributor test covers the queue fall-through; the Desktop's normal case is
a redirect-capable AIAgent, where busy_input_mode=interrupt turned the edit into
a mid-turn redirect and left the un-edited transcript in place. Pin that branch
too, and document in the rewind section that a truncating submit refuses with
4009 while a turn runs so hosts interrupt + retry (the Desktop already does).
2026-09-18 10:52:27 -07:00
teknium1
14d862384b fix(desktop): page-owned header controls get their own area; titleBar.center stays permanent
`444c75c10a` rendered `titleBar.center` in the workspace panel header on
full pages so the kanban board switcher would sit beside the page title
instead of colliding with the sidebar tab strip. That made the area migrate
between two React subtrees on chat <-> page navigation: the old instance's
effect cleanup ran after the new instance's setup and wiped every global
side effect a third-party plugin had just re-created (#114290).

Keep both intents: `titleBar.center` is a permanent titlebar slot again (the
salvaged commit), and page-owned controls move to a dedicated
`WORKSPACE_PAGE_HEADER_AREA` (`workspace.pageHeader`) that the workspace
pane projects into its vetoed tab row while `$workspaceIsPage` holds —
exactly the placement `444c75c10a` introduced, just not through the
plugin-facing titlebar area. Kanban's board switcher contributes there.

- SDK exports `WORKSPACE_PAGE_HEADER_AREA`; `TITLEBAR_AREAS` documents the
  permanent-mount contract.
- Test: a `titleBar.center` component's effect runs setup once and cleanup
  never across a chat -> /skills -> chat round trip (red on base).
- Docs: plugin SDK guide covers the lifecycle contract and the page-header
  area.
2026-09-18 10:45:44 -07:00
teknium1
b941f41ead docs(acp): every tool call reaches a terminal status (event bridge) 2026-09-18 10:33:54 -07:00
teknium1
d9ca304047 docs(acp): note the failed-turn transcript boundary in the ACP lifecycle 2026-09-18 10:33:21 -07:00
teknium1
aefa503479 fix(browser): gave-up supervisor unregisters itself; tests trimmed to two invariants
Follow-up to the cherry-picked reconnect cap (#114178, @KoNit-K); supersedes the
earlier #44460 (@plcunha), whose bounded-reconnect design this lands in slimmer form:

- Factor the two give-up branches into ``CDPSupervisor._reconnect_budget_spent``,
  which logs the single final warning AND drops the supervisor from
  ``SUPERVISOR_REGISTRY`` when it is still the registered one. Consumers
  (``browser_cdp`` frame routing, the eval fast path) then see "no supervisor
  attached" instead of a dead entry, and the next browser call for the task
  starts a fresh one via ``get_or_start``.
- Move the unit tests to ``tests/tools/test_browser_supervisor_reconnect.py``
  (two invariants: bounded + unregistered after attach; first-dial failure
  still fatal for ``start()``) and restore ``test_browser_supervisor.py`` to
  main — that file is the opt-in real-Chrome E2E suite and its module-level
  skip marker did not need to be rewritten to host unit tests.
- Docs: lifecycle section of the developer guide describes the bounded
  reconnect.

Live: fake CDP endpoint killed after attach — base logs connect-failed
warnings indefinitely (7 in 45 s, thread alive, still registered); fixed head
logs 4 + one "stopped after 5 failed reconnect attempts" line, thread exits,
registry entry gone. Control: endpoint back within 2.5 s re-attaches and the
budget resets.

Co-authored-by: plcunha <jvsantos.cunha@gmail.com>
2026-09-18 10:28:40 -07:00
teknium1
f39ae9000c docs(sessions): say what user_id holds for desktop and dashboard sessions 2026-09-18 10:16:46 -07:00
teknium1
eafed27cf0 fix(cron): a completed run keeps its result when a fire-claim heartbeat sample misses
Long-running cron jobs (>60 s, i.e. past one heartbeat interval) intermittently
ended with last_status=error / "Interrupted by shutdown before terminal
completion." while their output was complete. The fire-claim heartbeat thread
took ONE sample that read the claim as not ours, latched lost_ownership, and
_FireOwnership.lost() then trusted that latch without asking the store again —
so a run whose claim still validated (the very condition under which
_record_fire_ownership_lost writes that message) was recorded as interrupted,
its output never saved. The reporter's own error string proves the miss was
transient: it is only ever written when the owner re-validates True.

Why this shape:
- _FireOwnership.lost(): an explicit transport cancel stays terminal; a latched
  heartbeat miss is re-checked against the store, and a claim that still
  validates keeps the run's real outcome. The owner-fenced mark_job_run and
  fire_claim_fence remain the authority, so a genuinely re-owned claim still
  yields (no delivery, no terminal write over the new owner, ledger discard).
  An unreachable store after a latched miss stays fail-closed.
- _heartbeat_loop: a miss is re-sampled once after
  _FIRE_CLAIM_MISS_CONFIRM_SECONDS before it latches. Latching cancels the live
  agent run and ends the lease refresh, so a single sample must not do that; a
  genuinely re-owned claim misses twice and latches ~1 s later.

The issue's other hypothesis (worker process identity under a systemd-run
--scope worker) is falsified by the code: the owner token is stored in
fire_claim.by at claim time and compared to the stored value, never recomputed
from _machine_id(), and the heartbeat thread runs in a copied context so it
reads the same store.

Slimmer redo of #113364 (same direction: revalidate before trusting the latch)
without its production-dead isinstance(_CombinedCancelEvent) branch.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:57:02 -07:00
teknium1
c61bbc2461 fix(gateway): persist the authored text, not the Discord triggering note, as the user row
The `[Triggering message id: …]` note that `_prepend_inbound_reply_context`
adds for Discord turns is a model instruction (which id to pass to the
discord tools), not something the user wrote. `_hmwa_apply_message_timestamp`
derived `persist_user_message` from the already-wrapped text, so every
Discord-origin user row stored the note as the first line of `content` —
the desktop transcript (and FTS, memory providers, exports) showed the
envelope instead of the message.

- `run_inbound.py`: the note is now the OUTERMOST prefix (after the reply
  pointer) and rendered by one function, `discord_triggering_note`;
  `strip_discord_triggering_note` peels exactly that prefix for THIS event
  off the persisted text, so the `[Replying to: "…"]` pointer survives.
- `run_turn.py::_hmwa_apply_message_timestamp`: persist the stripped text.
  The wrapper keeps riding `message_text`; when the durable row differs, the
  live bytes land in the replay-only `api_content` sidecar by the existing
  persist-override contract (same as timestamps and per-turn sidecar notes).
- `run_turn.py::_run_agent_queued_followup`: the in-band queued follow-up
  runs the same inbound prep but passed no persist override; it now carries
  the authored text too.

Live: real inbound prep → persist seam → SessionStore.append_to_transcript on
a temp state.db. Before: content='[Triggering message id: `1550…` — use as
`message_id` …]\n\nCreate a project plan for Q4'. After: content='Create a
project plan for Q4', api_content=<wrapped bytes>; reply-pointer and
no-message_id (desktop relay) controls unchanged.

Slimmer redo of #71309 by @JonthanaHanh (regex strip on a `gateway/run.py`
site that no longer exists; tests exercised a copied regex); #71619
(@calvinnwq, +1303/-40 over 10 files) is over the salvage bar.

Fixes #71304
Fixes #114719
Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>
2026-09-18 09:53:09 -07:00
teknium1
9f5b7ea02e fix(compression): keep the aux ceiling and re-probe feasibility on every main-runtime change
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).

- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
  _apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
  only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
  every runtime change: switch_model (outside the rollback guard), fallback activation
  and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
  fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
  rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:49:14 -07:00
teknium1
d7865939af fix(tts): xAI TTS availability probe prefers XAI_API_KEY like the synthesis paths
`tools/tts_tool.py::_xai_requirements` (the provider check text_to_speech_tool
dispatch consults) still resolved xAI credentials OAuth-first, while both TTS
synthesis paths now pass `prefer_api_key=True`. Truthiness is unchanged, but a
configured key no longer routes the availability probe through the OAuth pool
(pool select / refresh) for a bearer the metered `/v1/tts` endpoint rejects.

Docs: the streaming-TTS provider table now states the ordering.

Part of the #113727 salvage of #113728 (@beardthelion); sibling site the PR
missed.

Co-authored-by: beardthelion <beardthelion@users.noreply.github.com>
2026-09-18 09:47:06 -07:00
teknium1
e81d4b8d96 docs(relay): scale-to-zero obligation 7 — relay-fronted only, re-checked at suspend time 2026-09-18 09:34:42 -07:00
lEWFkRAD
cb8d652b49 fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
2026-09-18 09:32:34 -07:00
Yagna Vudathu
87cd8a3c84 fix(gateway): /stop, /new and /reset keep a parked internal wake instead of discarding it
_interrupt_and_clear_session popped the adapter's single pending slot and dropped whatever it
held ("consume and discard", 59575d6a91). That was right when the slot only ever carried the
user's stale follow-up text, but internal wakes (async-delegation completion notices,
kanban/cron notify+wake) now park in the same slot and are claim-settled the moment the adapter
admits them, so the pop lost them for good: the drain that runs after the command found an
empty slot and the session idled until the next user message (#114456, ~6 min stall in the
reported session).

Now the human follow-up is still discarded, but an internal wake stays parked (promoted out of
the overflow FIFO when a discarded human head occupied the slot) and the post-command drain
starts it immediately. Applies to every caller of the helper — /stop (busy fast path, handler,
pending sentinel, thread sibling), /new and /reset — since a wake that arrives a second after
/new runs against the fresh session anyway; whether a completion pinned to the closed session
may run stays with _resolve_async_delegation_session (fail-closed).

Salvaged from #114538 (@whyyagswhy, earliest filer): the pop-and-re-park mechanism is theirs;
widened here to the /new and /reset callers and the overflow promotion, and the invariant tests
rewritten against the real adapter drain. #114540 (@JoaoMarcos44) reached the same fix
independently; its stop-reason taxonomy and overflow analysis informed the class coverage.

Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-18 09:27:02 -07:00
teknium1
5811653914 fix(api): a run admitted but not yet started settles interrupted at shutdown
The salvaged mark-then-interrupt path covers every run whose agent exists.
A run whose executor task was created but has not taken its first tick has
no agent to interrupt: _execute_run would re-persist `running` over the
shutdown mark and start a full turn that outlives the gateway (only forced
to `interrupted` when it finally returned). Check the shutdown set before
the `running` transition and finish immediately.

Docs: `interrupted` joins the terminal status vocabulary for /v1/runs.

Co-authored-by: Mark Rockminster <markrockminster@M1-U-48-128.local>
2026-09-18 09:24:21 -07:00
teknium1
3e408dcc2a fix(gateway): Telegram topic-binding heal cannot re-pin a route /new moved during its lookups
_hmwa_heal_telegram_topic_binding awaits get_telegram_topic_binding and get_compression_tip between reading the route and calling switch_session, the same shape as the async-delegation re-pin fixed in this branch. Pass the snapshot session id as expected_session_id so the compare-and-swap refuses to overwrite a concurrent /new or /resume. The developer-guide Operations table now documents the CAS kwarg.
2026-09-18 09:20:39 -07:00
teknium1
d177b119e9 docs: setup UX a standalone memory provider keeps; catalog migration note for users 2026-09-17 20:35:22 -07:00
Victor Kyriazakos
541b20290c fix(bedrock): restore Grok context with provider-confirmed cache provenance 2026-09-17 18:36:10 -07:00
teknium1
df1074b4e5 fix(aux): structured-output rejection no longer kills fallback candidates or costs a doomed first request
Two open atoms of #83390 (DeepSeek "This response_format type is unavailable now"):

* `_call_fallback_candidate_sync/_async` only special-cased auth errors, so when the primary
  aux provider failed (timeout, rate limit, payment) and the fallback landed on a provider that
  rejects `json_schema`, the 400 re-raised and the whole task died — the primary-path rung from
  #89589 never applied there. Both fallback paths now retry once without `response_format`.
* Every structured aux call (titles, kanban decomposer, goal judge, plugin structured calls)
  paid a guaranteed-fail request on providers that lack `json_schema` before the retry. A
  provider profile can now declare `unsupported_response_formats` (DeepSeek: json_schema, per
  https://api-docs.deepseek.com/guides/json_mode) and the recovery ladder remembers any route
  that rejected a type once (host:port scoped), so `_build_call_kwargs` — shared by the primary
  and fallback paths — omits the field before the first request. Dropping rather than
  downgrading to json_object matches the end state the retry already produced; json_object
  needs a JSON-mentioning prompt and some relays return empty content under it.

New logic lives in agent/auxiliary_structured_output.py; the facade only gains the fallback rung
next to the predicate it uses. tests/agent/conftest.py resets the process-level memo per test.

Fixes #83390, #105191. Closes duplicates #84976, #88830, #102849, #113064.
Co-authored-by: Legion-is-life <Legion-is-life@users.noreply.github.com>
2026-09-17 09:11:38 -07:00
teknium1
091b8c4f53 docs: call out the client.capabilities fail-closed gate as breaking for third-party WS clients
Why: an integration that answered server→client requests but never sent
client.capabilities now has every clarify/approval/sudo/… request refused at
once with no grace path. Marking a connection as answering after its first
response frame cannot help — the frame is never sent to it — so the break is
documented instead.
2026-09-17 09:04:38 -07:00
teknium1
f9d178f78e fix(tui_gateway): old app builds no longer stall the agent on clarify/approval; late approval choices count
Item 2 of #112548: a Desktop/dashboard build that predates server→client
requests has no response path, so every clarify/approval/sudo/secret/vault/
connection/bridge request sat for the full deadline (clarify: 300s). Only the
tour probed. Clients now advertise once per connection
(`client.capabilities {server_requests: true}`, sent by the shared TypeScript
channel on `gateway.ready`); `send()` / `send_async()` return the
error-response shape (None) at once when every WebSocket peer of the session
is a build that never advertised. Sessions with no client attached still wait
so the reconnect replay (`open_requests`) keeps working; the stdio TUI ships
with the backend and is not gated. The advertisement is dropped on disconnect.

Reviewer minors from #113227:
- tools/approval_gateway_wait.py: the verdict is the choice committed under
  the approval lock while leaving the queue, so an /approve that lands after
  the deadline check but before the entry is dropped is an answer, not a
  timeout (the client was already acked "ok").
- tests/tui_gateway/test_protocol.py: the error-fails-fast test that only
  restated pre-existing behaviour is replaced by the two capability
  invariants (never advertised → fails fast; advertised → frame written,
  waits, forgotten on disconnect).
- server_requests.send try/finally around event.wait already landed on main
  (4371ed34a9); nothing to change.

Docs: programmatic-integration.md (advertise once per connection; method
list), tui_gateway/AGENTS.md; contracts regenerated.
2026-09-17 09:04:38 -07:00
teknium1
0f98d5a1d9 fix(plugins): execution-chain middleware failures are reported once, not per call
The warn-once reporter added for hooks (and for PluginManager.invoke_middleware)
left the execution chain out: hermes_cli/middleware.py::_run_execution_chain — the
tool_execution / llm_execution frames that run once per tool or LLM call — still
logged "Middleware '%s' callback %s raised" at WARNING on every invocation. A
mis-declared callback (e.g. a signature naming tool_data) therefore flooded the
log exactly like the hook case #111922 reported: 5 calls -> 5 WARNING lines.

Route the frame's except through the manager's _report_hook_failure with the
"Middleware" surface label, so the first failure warns (listing the fields the
middleware does provide) and identical repeats go to DEBUG; the set is already
forgotten on plugin reload. The frame's skip-and-continue semantics are unchanged.

Part of #111922
2026-09-17 09:03:29 -07:00
teknium1
07c92d675a fix(compression): repeated summary stall escalates to the deterministic fallback summary
A first stalled summary stream keeps today's behaviour: the transcript is left
alone, the stall-class cooldown (floored at the idle window) is armed and the
LLM route retries after it lapses. When the route stalls AGAIN while a
stall-class failure is still on the ladder, the stall retry ladder now ends
with a deterministic rung: the worker is re-run with the summary LLM skipped
(DETERMINISTIC_SUMMARY_ROUTE pin, consumed in _summarize_window) and
compress() commits its static fallback summary through the ordinary
lease/fence/watermark pipeline — the same degrade a failed summary call gets
(abort_on_summary_failure still aborts).

WHY: after "made no progress … continuing without compression" the context
stays oversized, so the next turn after the cooldown re-enters the same silent
stream and burns another full idle window; the reporter saw this every ~2 min
for hours (#112420). A route that has proven unhealthy twice must degrade once
instead of looping. The prune-on-stall hunk from #112504 was declined because
it committed outside the lease/fence; this rung reuses the same-turn fallback
worker (bypass_cooldown) and the commit path the fallback_chain retry already
uses, so no new commit surface is introduced.

Also: a pinned fallback_chain route whose summary call FAILS still commits the
static fallback summary (default abort_on_summary_failure=false); the host log
said "recovered on fallback_chain[0]" for that. It now logs "committed a
deterministic fallback summary on …" at WARNING (#112387 review caveat), keyed
on the post-commit fallback_compression_streak bump.

Docs: developer-guide failure-cooldown section + agent/AGENTS.md.
2026-09-17 08:59:50 -07:00
teknium1
96e8a23222 feat(plugins): install declared Python dependencies and re-apply them after hermes update
A directory plugin's pyproject.toml [project].dependencies (or manifest
python_dependencies) are now installed into the Hermes venv on install/enable/
update, resolved under a constraints file built from Hermes' own pinned
dependencies together with every enabled peer's declarations. A candidate that
cannot resolve is refused before its tree is moved into place; nothing is
installed and no other plugin is touched. --no-deps opts a single install out.

hermes update rebuilds the venv from Hermes' lock and strips everything else, so
its git, zip and both repair paths now re-apply the union of every profile's
enabled plugins. When the union no longer resolves, each plugin is resolved on
its own so only culprits are dropped (never an alphabetical neighbour); a
plugin-vs-plugin conflict peels non-memory plugins first because a Hermes that
boots without memory reads as data loss. Dropped plugins are disabled through
the real config writer with a loud message naming the fix.

python_runtime: external lets sidecar-venv plugins (Mnemosyne's shape) opt out
of the union; hermes-agent self-dependencies and direct-URL requirements are
never installed (the latter are surfaced for the user to install by hand).
2026-09-17 00:11:27 -07:00
teknium1
3276af0957 fix: cron agents carry their delivery channel's platform hint
A cron agent runs as platform "cron", so its system prompt only ever got
PLATFORM_HINTS["cron"]. The job's final response lands on the `deliver`
channel (Slack, Telegram, ...), but the model was never told what renders
there: no MEDIA: attachment instructions, no "tables don't render", and a
user's `platform_hints.slack.append` never reached scheduled jobs at all
(community report: Slack-delivered cron jobs stopped sending attachments).

platform_hint() now appends "Delivery destination (<channel>): <hint>" for
cron agents, where <hint> is the channel's built-in/plugin text resolved
through the same `platform_hints.<channel>` override path a live session
uses. The channel comes from the HERMES_CRON_AUTO_DELIVER_PLATFORM
ContextVar the scheduler publishes before the agent runs (the seam
send_message already routes by). Non-cron platforms are untouched.
2026-09-16 21:39:13 -07:00
KoNit-K
434c575543 fix(tui_gateway): make server request settlement atomic 2026-09-16 17:54:42 -07:00
teknium1
b75396a707 fix(desktop): render a typed synthetic in-flight prompt like its persisted row on reconnect
Desktop's appendLiveSessionProjection built the `user-inflight-<sid>` row from
`inflight.user` text alone, so a background-process completion (or async
delegation result) that was still running at reconnect showed up as a user
bubble even though history would render the same turn as a `process_complete`
timeline marker once it landed. The gateway now ships `display_kind` /
`display_metadata` on `inflight` (#112162); route a typed row through the same
`toChatMessages` projection history uses, keeping the stable id so the
activate/resume reconcile keeps deduping it. `hidden` prompts yield no row.
Genuine user input carries no display fields and renders unchanged, so a user
quoting the marker text stays a user bubble.

Docs: describe the `inflight` display projection in the programmatic
integration guide.

Part of #112144 (with the salvaged #112162 by @KoNit-K).
2026-09-16 17:54:17 -07:00
teknium1
6a99766424 fix(agent): system watchdog interrupts are attributed to their issuer, not the user
Every hard interrupt reached `begin_iteration` through the same `_interrupt_requested`
flag, so the turn loop booked all of them as `interrupted_by_user` — a cron run killed
by the scheduler's inactivity watchdog, a turn aborted by the liveness watchdog, a lost
session turn lease and a gateway inactivity timeout all read as a human pressing stop,
and the investigation went to the wrong subsystem (#112647).

`interrupt()` already records a trusted category per interrupt (`_tool_interrupt_reason`,
fed by the `tool_reason` every producer can pass through `request_hard_interrupt`). The
exit reason is now derived from it: the three categories `interrupt()` itself mints for
human stops keep `interrupted_by_user` / `interrupted_during_api_call`; any other
category names its producer — `interrupted_by_system(cron_inactivity_watchdog)`,
`interrupted_during_api_call(turn_liveness_watchdog)`. The system producers that passed
no `tool_reason` (cron inactivity watchdog, turn liveness watchdog, lease loss, gateway
inactivity timeout) now name themselves, so the model-visible tool-cancellation text
says the same thing. `_publish_interrupt_state` logs ONE line naming the source so the
turn record and the log agree.

No change to WHEN anything interrupts. `interrupted_during_api_call` moves to the
prefix-matched explanation table so the parameterised form keeps its user copy.

Slim redo of #112652 by @KoNit-K, which added a parallel `issuer` attribute and
keyword; this reuses the existing `tool_reason` plumbing instead.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 17:48:17 -07:00
teknium1
ca396e0f7f fix(desktop): track plugin gateway listeners in the bundled loader and ctx.onEvent
The cherry-picked fix scoped host.onEvent disposers into the runtime
loader's per-plugin disposer list. This finishes the class:

- apps/desktop/src/contrib/plugins.ts: the bundled loader's activate()
  wraps register() in the same trackGatewayEventDisposers scope, so a
  Capabilities > Plugins disable/re-enable cycle no longer strands a
  bundled plugin's gateway listeners either (same accumulate-per-toggle
  mechanism as the disk hot reload).
- apps/desktop/src/contrib/plugin.ts: PluginContext.onEvent — a tracked
  door for subscriptions made AFTER register() returns (timers, socket
  callbacks), where the register-time scope cannot see them and a bare
  host.onEvent still needs hand-wiring to ctx.onDispose.
- apps/desktop/src/sdk/index.ts + website desktop-plugin-sdk doc: state
  the contract plugin authors can now rely on.
- Tests: one ctx.onEvent invariant (red on base: TypeError) in
  plugin.test.ts; the contributor's events.test.ts control case dropped
  (green on base, no invariant) to keep the fix at two tests.

Why: unloadRuntimePlugin / deactivate only run ctx-tracked disposers;
anything a plugin subscribes outside track() survives every reload, and
one relay marker then opens N sessions (#112366).

Fixes #112366

Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-16 17:46:28 -07:00
teknium1
eb11e7eac9 fix(agent): reasoning promoted on a reasoning-only clean stop is never persisted as an ordinary reply
A clean ``stop`` with empty content and reasoning text still promotes the
reasoning to ``final_response`` (the vLLM nemotron parser files the whole
answer as reasoning; re-running the empty-response ladder re-billed the prompt
8x, #109205). What changes is what the assistant ROW carries: ``content`` stays
empty, the text lives in ``reasoning``/``reasoning_content``, and the promoted
text rides the ``api_content`` sidecar. ``build_api_messages`` substitutes the
sidecar on the wire, so the next request replays the answer byte-identically
(same bytes as before this change) and role alternation is intact, while
state.db, ``session.history`` and every other history surface no longer show
chain-of-thought as a reply indistinguishable from a real one.

The "Reasoning-only clean stop" log moves INFO -> WARNING and names model,
provider, api_calls and tool_turns: a model that ends every turn this way is
stalled (planning monologue, zero tool calls) while the turn reported
"complete".

The thinking-prefill interim row (``_thinking_prefill``) is ephemeral
scaffolding already skipped by the flush and popped before the final answer, so
it never persisted a content==reasoning row; no change there.

Live probe (real AIAgent, real SessionDB, temp HERMES_HOME, DeepSeek-shaped
reasoning_content only): before, the assistant row had content == reasoning
== reasoning_content with api_content NULL and an INFO log; after, content is
empty, api_content carries the text, the replayed second turn sends the same
assistant content on the wire, the WARNING carries the route, and a normal
"Hi there." reply persists unchanged.

Part of #111761

Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-16 17:21:57 -07:00
teknium1
7bb4336811 fix(prompt_builder): load the user's own SOUL.md on a scanner hit instead of blocking it
A SOUL.md in HERMES_HOME that *documents* the canonical injection phrase as
security guidance ("...content telling you to ignore previous instructions...")
was replaced wholesale by a [BLOCKED: SOUL.md ...] marker, so the agent ran
with no identity/constitution and only a one-line agent.log warning said so.

SOUL.md is the user's own file: agent writes to it always go through the
protected-instruction approval gate (tools/file_tools_write_guards.py) and no
repository checkout can plant it, so it sits in the same trust class as
config.yaml — not a cloned repo's AGENTS.md. `_scan_context_content` gains a
`user_authored` mode that still scans, logs the matched pattern(s) at WARNING
and loads the content; only `load_soul_md` uses it. Project-dir context files
(.hermes.md, AGENTS.md and the other project-dir files), subdirectory hints,
memory and tool-result scanning are unchanged and keep blocking. No threat
pattern is narrowed: the reported-speech form cannot be separated from real
attacks by regex ("I want you to ignore all previous instructions" is a
canonical payload).

The /context manifest reports such a file as `flagged` (loaded, ⚠ "review the
file") so the warning is visible on CLI, TUI, gateway and Desktop rather than
buried in the log.

Fixes #112570
2026-09-16 17:17:47 -07:00
teknium1
8170c16cca fix(desktop): route every desktop-plugin tree copy through staged publication
Extract the contributor's staging + rename sequence into
`publishDesktopTree` and make `installDesktopPluginFromGit` use it too:
its `copyDesktopTree` had the same mkdir + cp shape, so a copy that failed
half-way left an empty `<root>/<name>/` and every later install attempt
was refused with "already exists. Enable force reinstall to replace it".
One helper, both copy sites, same guarantee: a published folder is always
complete (marker included) or absent.

Docs: describe the staged copy and the marker-less/plugin.js rule in the
desktop plugin SDK guide.
2026-09-16 17:17:22 -07:00
teknium1
8188373b72 docs: cooldown floor and same-turn fallback retry in the compression failure-cooldown section 2026-09-16 17:12:33 -07:00
brooklyn!
74ca4f28ba fix(onboarding): consume future catalog entries without custom metadata 2026-09-16 15:54:47 -05:00
brooklyn!
4716ec0ba4 feat(desktop): suggest varied first tasks from real integrations 2026-09-16 15:16:09 -05:00
xxxigm
2eb5395d3f fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping (#112079)
* fix(launchd): park EX_CONFIG token conflicts instead of KeepAlive-looping

systemd already stops restarting on exit 78; launchd KeepAlive=true
respawned the same token/port collision every 30s. Map 78 to a clean
stop under SuccessfulExit=false so the job stays down until the holder
releases the lock.

* test(launchd): pin EX_CONFIG park under SuccessfulExit KeepAlive

The plist must not use unconditional KeepAlive, and the stderr wrapper
must turn gateway exit 78 into a clean stop without swallowing exit 75
or a non-gateway child's 78.
2026-09-16 17:58:25 +00:00
teknium1
a10bbf95bb feat(gateway): gateway.multiplex_profiles defaults to on, gated by a boot-time serve guard
DEFAULT_CONFIG now ships gateway.multiplex_profiles: true. GatewayConfig keeps an UNSET flag as
None so the boot can tell "the operator chose" from "the default applies"; every reader tests
truthiness, so an undecided flag never multiplexes by accident.

hermes_cli/gateway_multiplex_mode.py settles the unset default once per boot (called from
load_gateway_config_for_runner and `gateway run --config`): the same preflight `hermes gateway
migrate --multiplex` runs — default profile, >= 2 profiles, no secondary running its own gateway
(live pid or installed unit), no duplicate-credential / port-binder blocker, migratable host. A
refusal is a logged warning naming the blocker and the migrate one-liner; the gateway comes up
standalone exactly as before. Explicit values (config.yaml, GATEWAY_MULTIPLEX_PROFILES) pass through
verbatim; `--standalone` already pins false.

Other processes stop guessing the verdict from the merged default: named_profile_served_by_running
_multiplexer, the enroll warning, the dashboard listener guard, the cron-fire port resolver, container
boot and the migration plan (_read_multiplex_flag) read the live gateway's served_profiles record
first and the EXPLICIT flag second — so a per-profile fleet with the flag unset still reads as
"not yet multiplexed" and the fold proceeds.

Docs: multi-profile-gateways.md, multiplexing-gateway.md, hermes_cli/AGENTS.md.
2026-09-16 08:57:18 -07:00
brooklyn!
64a9b43261 docs: run worktree desktops side by side without port eviction 2026-09-16 01:54:36 -05:00
outpoints
86b3695248 fix(tui): synchronize runtime cwd after workspace moves 2026-09-15 22:30:11 -07:00
outpoints
9e6c79f5be docs(honcho): document workspace routing and provider context 2026-09-15 22:30:11 -07:00
teknium1
2dfb795cb7 fix(approval): undelivered or unanswered CLI approval prompts are not user denials
When the CLI approval callback raises, when no callback is registered on the
thread while prompt_toolkit owns the terminal, or when the input() read is
interrupted, prompt_dangerous_approval returned "deny" and the command gate
rendered "BLOCKED: User denied this command" — attributing a refusal to a
user who was never asked (#22992). #112308 fixed the gateway half of the
class (withdrawn prompts -> outcome "cancelled" with a cause); this closes
the CLI residual on the same shape.

- tools/approval_prompt.py: those three paths return an Unanswered("cancelled")
  sentinel carrying the cause; MCP elicitation consent maps it to "cancel".
- tools/approval.py: the CLI gate renders "BLOCKED: <noun> was not approved: the
  approval prompt could not be delivered or was not answered (<cause>)" with
  outcome "cancelled" — still fail-closed, "Silence is not consent".
- tools/file_tools_write_guards.py: the protected-instruction write gate
  reports the undelivered prompt instead of "was denied by the user".
- Shared metrics: "cancelled" is a counted approval outcome (contract + v2
  schema) instead of falling into "unknown".
- Docs: hook `choice="cancelled"` now covers the CLI causes.

Fixes #22992
2026-09-15 21:46:37 -07:00
teknium1
6332216384 fix(approval): withdrawn gateway approval prompts no longer read as a user deny
When a gateway approval wait ends without anyone answering — the parent's
delegate_task finishing and tearing the child down, a /stop, or the turn's
notifier being unregistered at turn end — the tool result said
"BLOCKED: Command denied by user" (outcome="denied", user_summary "You denied
this command"). The user never saw or answered the prompt, so the parent agent
went on reasoning about a refusal that never happened (#112026, #22992).

The action stays fail-closed (the command does not run, the model still gets
the NOT-consented stop text), but the attribution is now truthful:

- tools/approval_gateway_wait.py: `_cancel_cause()` reads the existing
  per-thread interrupt-cause channel (`get_interrupt_reason()`, a trusted fixed
  category — no string matching) for the interrupted state and marks a
  notifier-unregister wake (event set, result None) as "the turn ended before
  the prompt was answered". Both the direct and the coalesced-follower wait
  return `cancelled=<cause>`; the post_approval_response hook fires
  choice="cancelled" instead of "deny"/"timeout".
- tools/approval.py: a cancelled decision renders
  "BLOCKED: Command approval was withdrawn before the user answered (<cause>)."
  with outcome="cancelled" and its own user_summary; an explicit /deny is
  untouched.
- tools/delegate_tool_child_run.py: `_signal_child_stop` publishes a fixed
  tool_reason ("parent delegation ended"; the late-child mirror forwards the
  parent's own category) so a child's pending approval can tell teardown from a
  user /stop — previously it rode the default "explicit stop requested".
- tools/file_tools_write_guards.py / tools/approval_prompt.py: the protected
  instruction-file gate and MCP elicitation consume the same key instead of
  reporting "denied by the user" / "decline".

Co-authored-by: zccyman <16263913+zccyman@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:44:46 -07:00
teknium1
0a6c7b7fa3 fix(agent,gateway): report interrupted and unfinished turns truthfully
An interrupted turn left `finalize_turn` with a diagnostic `final_response`
("Operation interrupted: waiting for model response") and no failure, so the
result said `completed=True` — the only producer that did; `turn_recovery`
and `codex_runtime` already return `completed=False` for an interrupt and the
gateway stream gate documents that contract. `completed` now also requires
`not interrupted`.

The API server then hard-coded the terminal status: the session chat stream
emitted `assistant.completed {completed: true, interrupted: false}` and
`run.completed` for every turn that did not raise, and `/v1/runs` booked any
non-`failed` result as `completed` — including an interrupt that did not come
through `/stop` and a turn that ran out of iteration budget. Automation that
reads the run status or the terminal event saw unfinished work as delivered,
and `partial: true` could sit next to `completed: true` in one payload.

`api_server_runs.terminal_run_status()` is now the single mapping for both
surfaces: interrupted -> `cancelled`, failed/partial/`completed=False` ->
`failed` (with `turn_exit_reason` and the fallback text as `output`),
otherwise `completed`; the terminal event is always `run.<status>` and a
late `pending_steer` rides on every terminal status instead of only on
`completed`.

CLI exit codes (`-q` quiet mode, `-z` one-shot) are deliberately unchanged
here: `hermes -z` returning 0 whenever text was produced was a stated design
choice (093f567f0d) and scripts depend on it, so that flip needs a
maintainer decision.

Fixes the gateway/producer half of #111770; slimmer redo of #111785 by
@KoNit-K (same mapping idea, one helper instead of three ladders).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-15 18:21:44 -07:00
teknium1
3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
Robin Fernandes
1034215ae8 docs(free-tier): drop the rehearsal page and its references
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes
d89cacc25f chore(free-tier): keep the rehearsal server out of the repo; the docs page explains the stand-in instead
The fault-injecting server served one-off manual rehearsal only and would
drift silently from the real services; the doc now says how to point the
desktop at any local stand-in (the three env overrides) and what such a
stand-in has to speak. The dev-only HERMES_EXTRA_WELCOME_HOSTS override stays,
pinned by a test in test_anon_failure_modes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30
Robin Fernandes
51e39af967 feat(free-tier): ruled behaviour for every welcome-api failure, with friendly copy and a fault-injecting rehearsal server
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.

Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
  temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
  into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
  honours the server's wait, climbs a short ladder when the service is
  unreachable, never retries terminal codes, and yields to the user's own
  retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
  background loop retries transient failures and re-announces setup.ready.
  setup.status and free_tier.status expose the block; free_tier.provision is
  the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
  the route); model_not_free moves onto the gateway's alternate once;
  anon_on_paid_host re-reads the route once; a long rate_limited refusal
  trips the cross-session guard; a locked account is retired but never
  replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
  retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
  service is off" (what is unavailable is using Hermes without signing in,
  and signing in is free), no jargon, spoken waits.

Desktop
- A setup-failure notice above the provider picker: one sentence per code,
  a retry when the backend says one can work, the sign-in pointer only when
  the account service answered at all. The overlay re-checks readiness on
  setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.

Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
  real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
  (dev-only, env-only) lets the route rules treat it as the welcome host.
  Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-15 20:44:42 +05:30