* feat(platform): resolver core with locate/inspect/probe tiers and ordered candidates
Every resource lookup needs one result shape and one cost contract. `locate` reads
metadata only, `inspect` may open files and call OS APIs in-process, `probe` is fresh
and the only tier that may spawn or connect. `Resolution.candidates` keeps probe order
so fan-out consumers can try every present binary.
Linear NS-921.
* feat(platform): AppResolver over AppDef with plist, PE, registry, and server.json sources
Desktop apps need presence, version, and liveness as separate observations. The runtime
file's bearer token is parsed, used for one request, and discarded inside the probe;
no public type carries it. Endpoints are accepted only when loopback with a numeric port.
* refactor(copilot): gh candidates through locate_command and the Homebrew table
First consumer of the resolver. The gh token probe still tries every present binary in
order; the allowlist loses its two copilot_auth rows.
* feat(platform): availability() over an application declaration
locate() + inspect() only, never probes; the fail-closed _version in
app.py treats a vendor's plist/PE/registry entry as untrusted input.
Salvaged from PR #118122; reads any object with requires_app,
min_version, app_for(os) — nothing here imports the MCP catalog.
* feat(platform): application declarations parsed into AppDef per OS
The parser slice of PR #118122's catalog manifest, re-homed as a
catalog-free module: whoever owns an MCP server declares the app it
fronts per OS and what it needs, and registers it here. Stdlib +
hermes_platform.resolver only. register/lookup/clear are the one seam
the MCP check_fn and the skill gate both read.
* feat(mcp): check_fn honours a registered application declaration
_make_check_fn ANDs the declared app's availability into the
connection-alive check; with nothing registered for the server the
behaviour is the pre-PR3 connection check. Provenance is explicit
registration, not endpoint matching. Returns a plain bool: the registry
caches bool(fn()).
* feat(skills): requires_apps gate through registered declarations
Offer-time filter beside environments:; names resolve through
hermes_platform.declaration, an unknown name hides the skill (fail
closed). The disk snapshot carries requires_apps and the fast path
re-evaluates it (snapshot version bumped to 3): app presence is a host
fact that changes without SKILL.md changing.
* docs: application declarations page
The plugin-facing schema reference: app: and requires: blocks,
availability() states, and the two gates that read the registry.
Registered under Extending > Plugins in the docs sidebar.
* test(platform): declaration parser, availability, gates
The PR3 app-block tests re-homed off the catalog: fixtures are dicts
passed to parse_declaration, the check_fn gate keys on explicit
registration (not endpoint matching), and the import-hygiene probe now
covers hermes_platform.declaration and resolver.availability.
`snapshot_paths` / `_store_blob` write blobs OUTSIDE `_ledger_lock`, seconds
before the row that references them is appended. Since `_maintain_size` now
runs `gc_blobs()` on any append that trims, a sweep in another process during
that window deleted the in-flight capture: its row then landed with sha256
values `read_blob` cannot resolve and `rollback_entry` failed.
`_gc_blobs_locked` now skips any unreferenced blob file (incl. `.tmp-*`) whose
st_mtime is newer than `_BLOB_GC_GRACE_SECS = 3600`; no new config key.
Docs (curator.md) and the `gc_blobs` docstring say so.
PROOF: probes/s5_final_inflight_blob.py before -> `P1 blob survives P2's
sweep: False ... read_blob: None`; after -> `True ... read_blob: b'P1
in-flight skill file'`. test_concurrent_appends_never_lose_a_middle_row gained
a fresh + 2h-aged unreferenced blob pair: red with the grace set to 0 (fresh
blob deleted), green at 3600 (aged deleted, fresh kept). ruff clean,
check-windows-footguns --all clean, run_tests.sh ledger + AX files 29 passed.
_trim_oldest never parses lines: it drops the OLDEST ones whatever they
are, so 'malformed lines always survive' was false. Reword docstring,
curator.md and the kept test's docstring to the true guarantee: lines in
the retained tail are never rewritten or parsed, so a malformed line there
survives verbatim. No behaviour change.
PROOF: probes/s5_trim_probe.py (cap 4096, malformed first line, 4 fat
appends) -> old malformed survived: False, contradicting the old wording;
the kept test asserts a LAST-line malformed row, which is what the new
wording promises.
A plugin manifest's `config_schema` now reaches the Desktop: `plugins.manage list`
returns each plugin's schema with the current `plugins.entries.<id>.settings`
values (`settings_schema`), and a new `settings` action writes edits through
`hermes_cli.plugins_state.save_plugin_setting` — the writer extracted from
`PluginContext.set_config`, so the plugin, the CLI and the Desktop share one
config path, one lock and the same managed-install / managed-key refusals.
The Plugins tab grows a gear per plugin with a schema; the inline form is
table-driven (`FIELD_CONTROLS` / `INITIAL_TEXT` / `COERCE` keyed on the wire
field type) for string / number / boolean / enum / json / secret. Secrets are
declared with `type: secret`: the row carries only the `.env` name and a
presence flag, the client writes the value through the existing `PUT /api/env`
credential route, and the RPC refuses secret keys so nothing lands in
config.yaml.
Contracts regenerated; docs gain a "Settings form in the Desktop" section.
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours
The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.
- `tools/connectors/portal/`: a client for the portal's tool-list route and a
JSON cache under the Hermes home, one file per portal origin and connector.
An entry is fresh for 24 hours. After that the read revalidates with the
stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
`tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
live in `tui_gateway/methods_connectors_account.py`.
The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.
* feat(connectors): catalog, accounts and member tool rules by RPC
The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.
- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
first. The body is a union on `mode`, so a reader can name who turned a
tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
connector, or one connector on or off), with the revision the user saw. A
stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
page can show one card per app.
* feat(connectors): connect an app without a chat session
Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.
- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
`connectors.operation.wake` and `connection.respond` take `owner`, a union
on `type`: `session` (today's behaviour and authorization) or `account`
(routed by `profile`, authorized by the live transport like `mcp.*`).
`session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
thread, under the profile's scope, so the watcher reads the account and
settles the operation. A second connect for an app that is already
connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
address, so its updates go out on the session-less broadcast path.
* feat(mcp-catalog): eighteen more bundled entries name their hosted connector
A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.
* refactor(connectors): the account handlers share one gate, one params model and one write table
The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.
The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.
The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.
An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.
Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.
* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"
The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.
Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.
* fix(connectors): a connect from the page returns to the app after sign-in
The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.
Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.
* test(connectors): defer the new connector RPC coverage
The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.
Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.
Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.
* fix(cli): the connection panel hands the tool thread back at once
The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.
For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.
The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.
Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.
* fix(connectors): "run it again" lives in the library, so the classic CLI can use it
Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.
`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.
Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.
* feat(connectors): the account list and disconnect go through the portal
`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.
The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.
`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.
* fix(connectors): the account RPCs answer what the portal really sends
Checked against the portal source and against the staging and production
services.
- Errors are read from the upstream error code, not the HTTP status. A rule
write answered 409 for a stale revision and for a user with no organisation;
both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
portal's own result for this user, with its stamp and without provider or
subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.
Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.
* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link
Found by two adversarial reviews of the RPC layer and its types.
- `connectors.connect` from a chat session with no open operation is refused
(`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
registry with no card: it made a link nobody watched, returned a reply
without the required `settled` field, and named an operation that was never
registered. There is one way into an operation: the agent's call, or the
account owner's `connectors.connect`. "Run it again" inside an open
operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
profiles with the same key no longer cross-deliver a sign-in link. The event
payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
`connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
The last one is display data: the gateway enforces the rules, the backend
only passes the list on. The phantom `name` and `description` are gone, and
the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
produces it. The contract generator now fails when a contract enum and its
domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
`optional-mcps/` are untouched by this PR again.
anti-slop: no net-new findings (15 touched files).
* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect
The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.
- The agent turn now declares how a link can reach the user
(`tools/connectors/turn.py`): CARD when the agent was built with a
connection callback, SIDE for a subagent or a background turn, LINK for a
headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
tool batch in the agent loop and read by the connector dispatch path, which
never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
tool on any path that derives the tool list, and a connector call on an
unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
does, so no `connection.update` is emitted for an operation no client asked
for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
`respondToConnectionRequest` share one guard and send nothing for a settled
or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
repeated it to the user.
Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.
* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients
A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.
- `tools/tool_labels.py` is the one place that turns a bridged call into a
label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
→ "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
keeps its own emoji, verb and primary-argument preview. A batch gets exactly
one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
failure text on the row of the call that failed. With friendly labels off
it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
carry a typed `labels` field. It does not depend on the classic CLI's
display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
labels, one row per call. The labels reach the row under a key no tool
argument can use. The connect card it drew under a failed tool result is
gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
`manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
"Reading tool details · N tools".
Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.
* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly
- A failed hosted search or describe used to return nothing, by design, so the
model saw only local tools and told the user that a connected app was
missing. The local results are unchanged; when the hosted leg failed, the
`tool_search` and `tool_describe` results carry
`connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
and one hint line. A rejected token is `sign_in_expired`; an entitlement
refusal or a shut gate adds nothing. `tool_describe` no longer lists those
names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
is a hosted connector account; `mcp: true` only when the user asks for an MCP
server, a local server or an install, or when the name exists only in the
catalog; connect and reconnect are hosted verbs, install, enable and
authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
gateway does not know the connector (confirmed on that failure path) and the
name is a catalog entry does the target fail with "X is a local MCP server.
Call manage_connections with action install ...". It is a per-target
outcome: other targets of the same call keep their links and their card. A
vendor failure on a name both sides know stays an ordinary failed row. The
MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
USER asks for that app again; the description and the settled-result notes
say so. A builder saw the model refuse a direct user request.
Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.
* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled
Reproduced on the real Ink TUI with the rig, then fixed:
- The keyboard was dead during the sign-in wait: the card kept a `submitting`
flag that the normal OAuth path never cleared, and Esc went through the same
guard. The in-flight state now belongs to the answered row and clears when
that row moves, when any later frame of the operation arrives, or after
five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
(the input handler had no branch for this overlay); Shift+arrows scroll the
transcript and the card ignores them; arrow keys no longer move the text
cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
operation stayed in the store, and a resume dropped the pending card. The
flag survives idle, a resume shows the pending card again, a session switch
clears it.
- States with no branch: `not_connected` and a row with no link fell into the
credential form; `expired` vanished with no note. The title and the row text
now name the action (connect, reconnect, install, enable, authorize); a
failed or expired row with no fields offers Try again / Skip; a failed row
WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
per app states the outcome. A settled or dismissed operation id is
remembered, so no replay or resume can reopen its card. Esc in the last
"Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
the card in one sentence.
Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.
* chore(connectors): remove the comments and docstrings this branch added
Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.
Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.
* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures
Found by the end-to-end runs on the pushed head.
- Desktop: after a window reload, Continue on the restored card sent nothing.
The answer looked up the backend that holds the session with the runtime
session id, the lookup wants the stored id, and a failed lookup returned
silently. When the lookup gives no owner the answer now goes out on the
window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
a refused scope, a non-member and a missing organisation alike: the handler
runs with the gateway's globals and did not import the reason enum, so its
own error mapping raised. `connectors.accounts.remove` caught auth failures
in its generic branch. `org_required` was mapped on `policy.set` only. All
six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
`ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.
- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
post_api_request payload shape plus `aux_task`, `api_request_id`
(`aux-...`, shared by every attempt of one logical call), `retry_count`,
`streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
turn is in flight; fail-open (a raising/hung subscriber is logged and
the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
attempt shares (_relay_sync_completion / _relay_async_completion /
_relay_sync_stream) run under the hook pair — retries and fallbacks
included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
aux_task and no api_request events; raising subscriber never breaks
the call). First is red on origin/main.
Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.
Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
A plugin whose import or register() never returns (an infinite loop, a blocking
network call) held PluginManager.discover_and_load() forever, and with it every
synchronous caller: `hermes chat`, gateway startup, ACP session/new (#108139).
Each plugin's import + register() now runs under `plugins.load_timeout_seconds`
(default 10, 0 disables, max 600) on a daemon worker. On overrun the plugin is
recorded as failed with "load timed out after Ns" (same channel as every other
load failure: startup WARNING, `/plugins`, `list_plugins()`), its pre-hang
registrations are disposed, and discovery continues with the next plugin. The
abandoned worker's later `ctx.register_*`/`subscribe`/`on_unload` calls are
refused with a WARNING (the context is marked abandoned), so a late registration
can never land in a registry the failure path already swept. Abandoned loaders
are capped per process (8); past the cap further loads are refused with a named
reason rather than run inline, which would recreate the hang (#98382 shape).
Because the worker cannot own the caller's RLocks: the deferred-platform eager
fallback now runs outside the replacement transaction, discovery re-entered from
a loader worker returns on the already-set discovered flag instead of blocking on
the sweep's lock, and such a worker never joins the background discovery thread
that is waiting on it.
website/docs/user-guide/features/kanban.md:987-988 listed the two gc
retention flags without saying what the edge values do; add "(negative N
is rejected; 0 disables that sweep)" so users don't have to read the code.
hermes_cli/kanban_db.py:4279,4293 — gc_events/gc_worker_logs accept
older_than_seconds=0 as "everything older than now" while _cmd_gc maps
days=0 to "disabled" before calling them. The docstrings did not state
that split, so a library caller could assume 0 is a no-op. One line each.
Gate findings: D.2ab.md:27, D.2c.md:32.
Since the first-strike escalation, a closed transport (socket_closed /
client_closed) forces the reconnect on the first unhealthy sample and the
failure threshold only applies to soft signals (ack staleness, latency,
event silence). The user guide (website/docs/user-guide/messaging/discord.md:89)
and the env-var reference (website/docs/reference/environment-variables.md:780)
still described the threshold as gating every unhealthy sample, so an
operator reading `1/2` followed by a forced reconnect would think the
knob was ignored. One sentence each, citing #118487.
Annotated-tag pins (F8): a catalog `sha` recorded as `git rev-parse <tag>` names
the TAG object, while HEAD can only ever be the commit it points at. The scan
trust check compared HEAD against the unpeeled sha (so every tag-pinned entry
lost the reviewed-pin bypass and prompted on caution findings) and the sidecar
recorded the peeled commit, so `update_available` was true forever and every
`update` re-installed. The installer now peels the pin (`<sha>^{commit}`) for
trust, records `pin` on the catalog block only when the checkout satisfies it
(empty for an off-pin `--ref` install), and every at-pin check goes through
`at_catalog_pin(sidecar, entry_sha)` (repin, dashboard payload, TUI rows).
Re-pin consent (F10): `hermes plugins update` on a catalog install replaced the
tree without asking, even when the new pin declared new tools, hooks, Python
dependencies, host capabilities or a Desktop half. `repin_catalog_plugin` now
diffs the installed manifest against the staged clone BEFORE anything moves
(`_install_plugin_core(before_swap=...)`) and, on a widening:
- CLI: prints the delta and asks y/N (non-interactive → not applied, fail
closed); after a changed re-pin it runs the same `_run_capability_consent`
grant path as the git-pull `update`.
- `plugins.manage update` RPC and the dashboard REST route answer
`{ok: false, consent_required: true, delta, delta_lines}` with nothing
changed; a retry with `accept_capabilities: true` applies it. Desktop shows
the delta in its confirm dialog; the web dashboard uses `window.confirm`.
- Gateway contract regenerated (`accept_capabilities` param; `consent_required`,
`delta`, `delta_lines`, `error` result fields).
Catalog audit findings F8 and F10 (low severity, no issue filed).
The cached live catalog (`HERMES_HOME/cache/plugin-catalog.json`) won over the
in-tree copy unconditionally: right after `hermes update` bumped an in-tree pin
a fresh (<6 h) cache still installed the previous sha, and offline a cache of
ANY age (90 days in the audit probe) outranked the checkout's catalog.
- For an entry both sources carry at different pins the NEWER catalog wins:
the checkout's last `plugin-catalog/` commit time vs the doc's
`generated_at`; when neither resolves (release install, no timestamp) the
entries' `version` labels break the tie, else live wins as before.
- A cache older than LIVE_CATALOG_MAX_STALE_SECONDS (24 h) stops supplying
pins (in-tree takes over) but its removals still count — a kill-list entry
never expires.
- The cache is written via temp file + rename: a concurrent reader (gateway,
TUI, a second CLI) can no longer see a half-written document, which read as
a fetch failure and started a 60 s failure window in that process.
Catalog audit findings F7 and F11 (low severity, no issue filed).
Builds on #72026 (@PRATHAMESH75): list content carries the turn's memory-prefetch /
pre_llm_call context as a durable text part appended once in the prologue, in every
api mode (MoA and codex_app_server included), so the request, the persisted row,
compaction and a later resume all see the message the model saw.
Persistence gap from the #72026 review: in-place preflight compaction (and a
close/early flush that races the prologue) writes the current user row BEFORE the
part exists and the crash persist identity-skips that dict, so a resumed session
replayed the turn without the context. The list branch now pushes the appended part
into that row via set_user_message_content under the same _row_id-under-lock
protocol as the string sidecar backfill, keeping the writer's shape (compaction: raw
parts; flush: text projection).
Titling moves before the injection step so a list turn's title is derived from the
user's ask, not the injected tail. Tests trimmed to one invariant per layer: hook
edit reaches the wire on a list turn and replays after reload; in-place compaction
+ reload keeps the part (red without the backfill); memory query flattens parts.
Four error-isolation holes in the disk plugin door
(apps/desktop/src/contrib/runtime-loader.ts), found by a static+live audit of
the loader; none had an issue filed.
- A plugin whose module evaluation never settles (top-level `await` on a dead
host) hung `import()` forever and, through the scan's sequential loop and
its re-entrancy guard, froze every later plugin and all future scans until
restart. `import()` now races a 10 s deadline; the plugin errors on its own
row ("import timed out") and the scan continues.
- Timers and DOM listeners a plugin took out with bare globals survived
disable and every hot-reload. `ctx.setTimeout` / `ctx.setInterval` /
`ctx.addEventListener` are tracked with the plugin and torn down on
unload; the SDK doc says bare globals are not.
- Two folders exporting one plugin id silently last-wins: the second
disposed the first's registrations and each hot-reload flipped ownership.
The first (folder-name sorted, so deterministic) owns the id; the later
file errors on its own row ("duplicate id, already loaded from <path>").
- A save that no longer loads (syntax error, timeout, duplicate) left the old
incarnation's contributions and activate handle live beside the error
row, so the Plugins tab showed a broken file as "loaded" and could
re-enable stale code. The previous incarnation is unloaded and dropped.
Tests: one invariant per fix in runtime-loader.test.ts, all red on base
(the hang case red by timing out).
`~/.hermes/hooks/` auto-loads every valid `HOOK.yaml` + `handler.py` at
gateway startup with no `plugins.enabled` gate. That is the documented
contract since 3988c3c245 ("Implicit (dir trust)" in the comparison
table), but the plugins page's "disabled by default" promise read as if
it covered gateway hooks too (#37963). Maintainer ruling: keep implicit
dir trust, fix the docs.
- hooks.md: new "Trust model" section stating exactly what loads, when,
how, and that placing the files is the opt-in; comparison-table cell
links to it and the plugin-hooks consent cell now says
`plugins.enabled`.
- plugins.md: note scoping `plugins.enabled` away from gateway hooks.
- developer-guide/plugins: one sentence at the gateway-hook recipe.
- security.md: "Trusted-by-placement extension points" section
cross-linked from hooks.md.
skill_view loads SKILL.md whole and the content then rides in context for
every later call of the session, so body size is paid per turn. The only
size signal was the 100k hard cap in skill_manage, and agent-authored skills
grew by small patches until they sat right under it (31 of 432 local skills
over 40k chars, 7 at 100-115k; skill_view results averaging 32k chars).
- skill_linter: advisory `oversized-body` past _BODY_SOFT_BUDGET_CHARS (24k,
~3x the ~200-line standard; bundled skills average ~20k) naming the size,
the token estimate and the references/ split.
- skill_manage patch: attach lint findings the write INTRODUCED (diff of
rules before/after), so the crossing patch reports it once and a clean
patch on an already-large skill stays quiet. Create keeps reporting all.
- curator prompt: a body over the budget is itself a consolidation target.
- docs: skills.md linter paragraph.
A plugin's transform_llm_output replacement reached final_response only.
agent/turn_finalizer.py::finalize_turn fired the hook after the assistant
row had already been persisted — and the row is first written even earlier,
in agent/turn_final_response.py::finish_text_response's durable flush — so
messages[-1], the SQLite/JSON session, /resume and the next turn's replay all
kept the raw model text while the user had seen the rewritten one.
Writing the transformed text back after that flush cannot work: SQLite treats
a non-blank assistant row as settled (resolve_and_repair_transcript_batch
adopts the stored content instead of overwriting it), so the only correct seam
is BEFORE the row is first persisted. apply_llm_output_transform (new, in
turn_finalizer) fires the hook once per turn_id and records the outcome;
finish_text_response calls it ahead of append+flush and writes the result into
the row (api_content for the promoted-reasoning sidecar), finalize_turn's
_persist_step calls it ahead of the recovery-path tail close, and
_apply_output_hooks reads the recorded outcome (firing only when no earlier
seam saw a response) before post_llm_call. Only the current turn's not-yet-
written text changes — earlier turns and the system prompt are untouched.
post_llm_call is unchanged: it is an observer whose return is ignored by
contract, so there is nothing of it to persist (#14913/#44253's premise).
Fixes#44239
Slim redo of #44244 (AIalliAI, earliest; same sync-then-persist idea, moved to
the pre-flush seam) — also supersedes #65921 (SingleVirgin, sibling fix).
Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
(cherry picked from commit 5fe02a0aeecef422a4ffb4ff4385015f6a1528a0)
A plugin registered on transform_terminal_output only ever saw foreground
`terminal` output: tools/terminal_tool_result.py::_apply_output_transform_hook
runs from finalize_foreground_result and nowhere else. Background output
reached the model through a different seam — process_manage poll/wait/log/kill
results, `list` previews and the completion/heartbeat/watch notifications all
pass through tools/process_registry.py::_redact_process_result — which redacted
but never transformed, so a fleet redaction or summarising plugin silently did
nothing for backgrounded commands.
Apply the same hook helper at that shared seam (one new
transform_process_output wrapper) and at the two gateway agent-notify sites
that read session.output_buffer directly. The order matches the foreground
path and teknium1's review note on #71401: hook first, redaction after, so a
replacement the plugin returns is still masked. returncode is None while the
process runs; env_type is not recorded per process and is passed empty.
Not changed: the spawn acknowledgement ("Background process started") carries
no command output, and the persistent local shell already goes through
_run_foreground and was transformed — the issue's reading of that branch was
wrong; the real gap was the process_manage/notification seam.
Fixes#70760
Slim redo of #71401 (Christopher-Schulze): same seam and ordering, without
the ANSI-stripping relocation and render helper.
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
(cherry picked from commit 84661de52f78ccb84054e1ded92b07166e40546d)
Follow-up trim of the #110265 salvage. `ainvoke_hook` logged raising callbacks
with a bare warning; route them through `_report_hook_failure` (warn-once per
distinct failure, #111922) and, for `_HOOK_TIMEOUT_FAIL_CLOSED_HOOKS`, append
the same named block directive the sync path emits (#109624), so the async twin
cannot drift into a fail-open policy path. Tests trimmed to the salvage bar: the
in-loop await is proven once through the real `_handle_message` path
(`test_async_hook_callback_is_awaited_on_the_gateway_loop`); the manager-level
duplicate is dropped and the narrowing test also pins failure isolation. Docs:
`pre_gateway_dispatch` callbacks may be `async def` and stay unbounded.
Credit order for the three PRs fixing this gap: #102485 (dmspark, earliest,
pre-decomposition `gateway/run.py`), #110253 (KoNit-K, bounded the hook —
rejected by design: neither fail mode is acceptable for a policy gate), #110265
(twidtwid, reporter; cherry-picked because it matches the ainvoke_hook shape,
keeps the hook unbounded, and adapts the existing sync test seams honestly).
Part of #110241
Supersedes #102485
Supersedes #110253
Co-authored-by: David Marcus <dmspark@users.noreply.github.com>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
`agent/shell_hooks.py::_parse_pre_tool_call` translated only the block and
modify dialects, so a shell hook printing the documented
`{"action": "approve", ...}` parsed to None and the tool ran with no approval
prompt — silently, with exit 0, valid JSON and `hermes hooks doctor` green.
The Python-plugin side already accepts approve and routes it through
`_resolve_block_from_details` → `request_tool_approval`; the shell parser now
yields the same `{"action": "approve", "message"?, "rule_key"?}` shape (optional
fields kept only as non-empty stripped strings), so `hermes hooks test` prints
it under `parsed:` and the dispatcher escalates it. the `decision` dialect's
`{"decision": "approve"}` means auto-ALLOW, not "ask a human", so it is
deliberately not mapped; that dialect has no top-level ask dialect to mirror.
Slim redo with credit: #92562 (earliest) bundled a larger policy-authority
rework; #110325 carried the same parser change plus an unrelated rule_key
default change and 10+ tests.
Fixes#92553
Supersedes #92562
Supersedes #110325
Co-authored-by: fangliquanflq <fangliquan@qq.com>
`_get_pre_tool_call_directive_details` returned the first valid block-or-approve
in registration order, so a plugin registered earlier that returned `approve`
hid a later security plugin's `block`; under `approvals.mode: off` an approve
means no prompt at all, so the veto was dropped silently. Precedence is now
`block` > `approve` > none: a valid block still returns immediately (modify
directives seen before it stay attached, as before), a valid approve is held
back until the whole result list has been scanned for a veto, and among approves
the first valid one (with its rule_key) still wins. Modify accumulation is
unchanged and now also keeps modify directives that follow the winning approve,
since the scan no longer stops there. Docstring and hooks.md no longer describe
"first valid directive wins".
Slim redo of #68644 (earliest) and #87449 against the modify-aware shape of the
function on main; both PRs predate it and could not be cherry-picked.
Fixes#87420
Supersedes #68644
Supersedes #87449
Co-authored-by: synscott <1563043+synscott@users.noreply.github.com>
Co-authored-by: Jack Lau <72348727+jackulau@users.noreply.github.com>
`_run_hook_callback_bounded` treated any live abandoned worker for a callback
(`bool(self._hook_abandoned.get(suppression_key))`) as "still running", so one
never-returning `pre_tool_call` callback made every later tool call fail closed
with the timeout message until the process restarted. The timeout path
self-heals through the 60s suppression window; the abandoned path never did.
Policy now: while the suppression window is open the callback is skipped as
before. After it expires a fresh call id may start a new worker even though the
abandoned one is still alive — capped at `_HOOK_MAX_ABANDONED_WORKERS` (3) live
abandoned workers per callback so a hung plugin cannot leak a thread per call
(the #98382 constraint). At the cap the callback keeps being skipped (fail-closed
for pre_tool_call) with a WARNING naming the callback and its module, until one
of its workers finishes and frees a slot.
Test changes: `test_hung_worker_blocks_new_call_identity_after_suppression`
encoded the removed behaviour (exactly one worker, forever); it becomes
`test_hung_worker_caps_new_call_identities_after_suppression`, which pins the
same invariant it was protecting — bounded leak, never one per call — at the new
bound and checks the warning. `test_hung_worker_does_not_fail_closed_forever`
is the #105223 regression (red on base: call-c returned the block directive).
Redone slim against the per-call-id gate that landed in #111177; #105241
targeted the pre-#111177 shape and needed plugins_ledger/__init__ changes for a
one-retry-then-quarantine policy. Its analysis and shape informed this fix.
Fixes#105223
Supersedes #105241
Co-authored-by: fangliquanflq <fangliquan@qq.com>
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.
Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.
Fixes#61839
credit: @giggling-ginger #61995
A general-plugin context engine is one shared instance; agent init copied it per agent
with copy.deepcopy() only. Engines that hold a SQLite connection or lock (hermes-lcm)
already expose clone_for_agent() for exactly this, but it was never called, so every
init logged "could not be safely copied … falling back to built-in compressor" and the
engine was unusable through the plugin system.
ContextEngine grows clone_for_agent() (default: deepcopy, the previous behaviour) and
_select_context_engine calls it; the failure message now names the hook to override.
Docs: context-engine-plugin.md documents the per-agent clone contract.
Test change (existing on main): tests/agent/test_context_engine.py::
test_agent_init_source_deepcopies_singleton_not_aliases was a source-reading pin on the
literal `copy.deepcopy(_candidate)` line, which this fix intentionally replaces. It is
superseded by tests/agent/test_plugin_context_engine_clone.py, which drives the real
_select_context_engine seam and asserts the invariant it guarded (child update_model()
never mutates the shared singleton) plus the new clone_for_agent() path.
Fixes#99640
credit: @stephenschoettler #62374
credit: @686f6c61 #99677
ToolRegistry.dispatch spread every injected keyword (task_id, session_id,
user_task, parent_agent, ...) into the handler, so a third-party plugin tool
written as `def handle(args)` raised TypeError on every call. The plugin
contract (plugins/AGENTS.md) says optional kwargs are signature-inspected, not
forwarded unconditionally; hooks already do this in plugins_dispatch. Handlers
taking **kwargs still receive the full payload. Docs updated: the narrow
signature is supported, **kwargs opts into the whole context.
Fixes#68318
credit: @vveerrgg #22146 (slim redo; #68636 is the later duplicate)
A list/mapping slot holding ONE quoted string (`plugins:\n enabled: '["a","b"]'`,
`model_catalog:\n excluded_providers: '["openai-api"]'`) is skipped by every
isinstance-gated reader (`plugins_cmd._config_name_set` -> set(), `plugins._names`,
`inventory.excluded_providers` -> []) while `config get` echoes it back, so user
plugins silently unmount and provider exclusions silently lapse. Neither `hermes
doctor` nor the startup `print_config_warnings` banner said a word.
Reader/doctor half: one schema-aware pass in `validate_config_structure`
(`_validate_quoted_containers`) walks `DEFAULT_CONFIG` (sections included) plus
`_KNOWN_CONTAINER_TYPES` and warns when the user value is a string that parses to
a list/mapping. The message names the key, the quoted value and the remedy
(`hermes config set <key> '<literal>'`). Finding only: the file is never
rewritten. String-typed keys (`approvals.mode: "[off]"`), the `model: <name>`
shorthand and the `parse_config_string_list`-read slots
(`agent.disabled_toolsets`, `skills.disabled`) are not flagged. Feeds both the
doctor "Config Structure" section and the startup banner.
Writer half (audit of every path that can put a string in a container slot):
- `hermes config set` (`hermes_cli/config.py::set_config_value`): already parses
bracket/brace literals (#88163) and refuses wrong-shaped values
(`_refuse_container_type_mismatch`), BUT the guard only knew slots present in
DEFAULT_CONFIG or `_KNOWN_CONTAINER_TYPES`. `plugins.enabled`,
`plugins.disabled` and `model_catalog.excluded_providers` are deliberately
absent from DEFAULT_CONFIG, so `config set plugins.enabled foo` /
`plugins.enabled a,b` / `model_catalog.excluded_providers openai-api` stored a
plain string (live repro on base). Added the three keys to
`_KNOWN_CONTAINER_TYPES`: those writes are now refused with the literal hint.
- `cli.py::save_config_value` + callers (cli_*_mixin, gateway/slash_commands,
gateway/run_busy): every caller passes a bool/enum string for scalar keys; no
list-slot caller. Not reachable.
- `hermes_cli/plugins_cmd.py::_save_plugin_sets` / `_write_config_value` and
`plugins_cmd_catalog` (via `_save_enabled_set`): write `sorted(set)` — real
lists. Not reachable.
- Dashboard `PUT /api/config` (`web_routers/config_env.py::update_config` ->
`_denormalize_config_from_web` -> `save_config`): schema-driven form;
`web/src/components/AutoField.tsx` splits list-typed fields into a real array
before the PUT. Not reachable.
- tui_gateway `config.set`: fixed `_CONFIG_SETTERS` table of scalar keys only
(out of this lane's files anyway). Not reachable.
Conclusion: the quoted shapes on real machines are leftovers of pre-#88163
`config set` runs plus the DEFAULT_CONFIG-absent slots fixed here.
Live repro (fake HOME/HERMES_HOME, both quoted shapes in config.yaml):
before: validate_config_structure() -> []; startup banner silent; doctor has
no Config Structure section; _config_name_set('plugins','enabled') -> set()
after: two warnings naming plugins.enabled / model_catalog.excluded_providers
with `hermes config set ... '["a","b"]'`; doctor prints them under
Config Structure; running the remedy yields real lists,
validate -> [], _config_name_set -> {'a','b'}
writer: `config set plugins.enabled a,b` wrote `enabled: a,b` on base; now
refused ("must be a list ... nothing was written").
Fixes#83308Fixes#105706
credit: @fangliquanflq #105725 (schema-aware detection in validate_config_structure; slim redo)
credit: @Luna161 #83313 (first report + per-key warning in plugins_cmd)
Catalog trust bugs from the 2026-09-21 plugin audit (lane 3, F1-F6, F9):
- F1 (high): a URL-installed repo shipping its own .hermes-catalog.json rendered
as catalog:official everywhere and marked the real entry "installed". Provenance
now lives on the installer-owned .install-metadata.json record (catalog block
written by _install_plugin_core, sha = checked-out commit); read_catalog_sidecar
never reads the tree. Pre-fix installs are adopted once when the installer
record agrees (pinned at the sidecar sha, cloned from the entry's repo).
- F2: removed.yaml bypassed by git@/ssh:///http:///www. spellings. _normalize_repo
canonicalises to host/owner/repo (scheme, user, www., .git, slashes dropped).
- F6: removed.yaml consulted for INSTALLED plugins too: git-pull update, enable and
gate_manifest (load) refuse recalled plugins, offline (in-tree + cached live
list). --allow-removed is recorded on the install record and exempts it.
- F5: install NAME --ref X recorded the catalog pin, so list/TUI/update claimed
the reviewed pin while HEAD differed. Recorded sha is the checked-out one.
- F3: re-pin replaced the whole tree, losing the installer-created config.yaml,
data and user patches with no warning. Untracked/ignored files are carried
into the new tree; edits to tracked files are copied to
<HERMES_HOME>/plugins-backup/<name>-<sha8>/ with a warning.
- F4: a manifest rename between pins left the OLD dir installed and enabled.
The stale dir is removed and the enabled flag follows the new name.
- F9: dashboard payload `removed` list now includes live removals.
repin_catalog_plugin returns RepinResult(sha, changed, installed_name, warnings);
CLI, dashboard and TUI callers surface the warnings and the new name.
The catalog page claimed admission's lint means a marketplace install 'cannot quietly rewire the app'; the lint is a handful of regexes and plugin.js runs in the app realm with the full window.hermesDesktop bridge. The user guide, catalog trust model and SDK security section now describe the real model: human review of a pinned SHA plus two tripwires (lint + loader import allowlist), no isolation.
Symptom: a model-provider plugin installed with `hermes plugins install`
(e.g. claude-subscription-directsdk) worked from a terminal but the Desktop
app failed every session build with "Unknown provider '<name>'".
Cause: `providers/__init__.py` scanned `$HERMES_HOME/plugins` exactly once
per process, under whichever profile home happened to be bound at the first
lookup, and the registry was global. The Desktop backend and the multiplex
gateway serve several profiles from one process, so any profile other than
the first-discovered one never saw its own plugins, and a plugin installed
while the process ran was invisible until a restart.
Change: bundled, pip and legacy providers stay process-wide; `$HERMES_HOME`
plugins load into a per-home layer keyed by `hermes_home_key()` at lookup
time (`get_provider_profile`, `list_providers`, `provider_source`). The
layer rescans when the plugin directories' mtimes change, so a fresh install
is found on the next lookup. The registration target is a ContextVar so two
turn threads scanning two homes cannot cross-register, and no lock is held
across plugin imports (a lock there could deadlock against a thread
mid-`import hermes_cli.auth`). Per-home module names let two profiles carry
the same plugin. `plugin_dev` reuses the module-name helper.
Live repro (tui_gateway `session.create` on a secondary profile whose config
selects a plugin installed only there): base -> agent_error "Unknown
provider 'fakeprov-b'"; fixed -> provider resolves and the build proceeds to
the plugin's own runtime check.
Deleting an idle scratch entry left two things behind. Headless browsers
started by a lane's e2e run kept running for days with a "(deleted)" cwd
(about 20 on one host; the process registry never knew them because they
were grandchildren of a shell that had exited). Repos whose linked
worktree lived in the entry kept a dangling registration until someone
ran `git worktree prune` by hand (10 in one repo).
The prune now lives in hermes_constants_scratch (hermes_constants keeps the
entry point). Before an idle entry is removed, same-user processes whose
cwd is inside it, or inside any scratch path that no longer exists, are
TERMed then KILLed; the deleted-cwd sweep runs on every pass so orphans
from earlier deletions are caught too. `.git` files found in the entry
name their repo, which gets `git worktree prune` after the rmtree.
cache/terminal used its own 72h fixed-age sweep; it now shares the 24h
idle rule and the subtree check. `hermes doctor` warns about cache-root
directories over 1 GiB that neither pruner covers, since finished campaign
trees parked there sat for weeks (95 GB on one host). System-prompt
scratch line updated to match.
The scratch pruner deleted top-level entries whose own mtime was older than
72h. A directory's mtime only moves when a direct child is added or removed,
so a lane writing deep inside its tree looked untouched and could lose a live
worktree at the deadline, while finished trees (7 GB clones with their own
venv per campaign lane) sat for three days: 791 entries / 57 GB after four
days of campaigns on one host.
Retention is now idle-based: an entry stays while anything anywhere in its
subtree was written in the last 24h and goes a day after the last write. The
walk short-circuits at the first recent mtime, so live trees cost one stat
and only a truly idle tree pays for a full walk, once, right before deletion.
Symlinks are not followed so a link into the repo cannot keep an entry alive.
The config comment and both docs pages named six of the eight sources in
MACHINE_PACED_SOURCES; "tool" and "batch" were missing, so an operator
reading the docs could not predict the tier those sessions get.
The 1h Anthropic cache tier writes at 2x base (5m: 1.25x) and only pays off when
turns are more than five minutes apart. That is exactly the shape of an interactive
session a person parks and resumes, and exactly not the shape of a subagent, cron
run, one-shot or webhook that calls every few seconds and is gone. A single global
`cache_ttl` cannot be right for both, so operators leave it on 5m and pay a full
context re-write every time they come back to a CLI session after a coffee.
Measured on one install (2 days of per-call API logs, Claude via the Nous route):
63% of interactive cache-write tokens were cold re-writes after a 5-60 minute idle
gap; 1h would cut interactive write cost ~42% while costing ~49% more on subagents
and ~23% more on cron. `auto` resolves once per session from the session source
(`_session_source_for_agent`): 1h for cli/tui/desktop/messaging platforms, 5m for
subagent, cron, oneshot, webhook, kanban, api. Auxiliary/stub calls keep 5m; the
delegate_tool child clamp (#104168) still applies. Default stays "5m".
`terminal(background=true, heartbeat=N)` emits a `heartbeat` event every N seconds
(floor 60) carrying only the output produced since the previous one, plus the usual
completion notice. The agent stays current on a long bounded job (merge train, full
test suite, deploy) and reacts to a failure within N seconds instead of at exit.
Why: `notify=true` fires once at exit, so multi-hour merge trains ran silently — the
parent's prompt cache went cold and every "what's up" cost 5–10 rediscovery calls;
30 days of batch sessions show 4,930 hand status polls (14% of all tool calls).
`notify=[pattern]` cannot serve this: its lifetime cap (8) exists precisely to stop
periodic output from flooding the agent.
Mechanism: one daemon timer thread for every heartbeat session (reader threads block
on the pipe and cannot keep time); delta output via a total-ingested counter over the
rolling buffer; heartbeats only where the completion notice can be delivered (same
async-support/subagent gates), never after exit. `heartbeat_seconds` is checkpointed.
Rendered on CLI, gateway and TUI through the existing process-notification paths.
Every line still promising that `gateway.multiplex_profiles: false` keeps
per-profile gateways, or documenting the deleted `migrate --standalone`
rollback, now describes the shipped behaviour: one gateway per host serves
every profile; `false` no-ops with a warning; `hermes update` folds a fleet
unless a real boundary (different UNIX user, HERMES_HOME outside profiles/)
holds; a blocked fleet keeps `--force` per profile; re-running
`migrate --multiplex` converges a half-migrated host; a manifest on disk is the
resume record, never a rollback. Adds the new "No new per-profile gateways"
section with the exact refusal `hermes -p <name> gateway install` prints.
Files: website/docs/user-guide/multi-profile-gateways.md,
website/docs/reference/cli-commands.md (migrate row),
website/docs/developer-guide/multiplexing-gateway.md (eager activation no
longer honours `false`), hermes_cli/AGENTS.md (migrate + refusal seams).
Completes #118273 (docs item). Part of #109417
f50d4b3423 (a docs revert) restored a route-style /user-guide/profiles link in
nous-portal.md (en + zh-Hans); website/scripts/check_doc_links.py rejects those, so
the Docs Site check has been red on main since. Rewritten with the script's --fix.
Reverts 93889b770d ("named profiles no longer inherit the root
profile's auth.json"). After the Desktop update every bot profile that had
relied on the root OpenAI Codex login failed with "No Codex credentials
stored. Run `hermes -p <bot> auth add openai-codex --type oauth`", and users
had to re-run the device-code flow once per bot (5-6 times in the field
report). Sharing one grant across profiles is the intended design: OAuth
refresh tokens are single-use, so ONE grant lives at the root, profiles
resolve it read-only, and a refresh under a profile writes the rotated chain
back to root (Codex / xAI write-through, borrowed-row pool bookkeeping,
forked-grant heal) — never a per-profile copy.
Restored: `_global_auth_file_path` / `_load_global_auth_store` fallback in
`_load_provider_state*` / `read_credential_pool` / `_provider_state_transaction`,
Codex + xAI root write-through, `credential_pool` borrowed-root persistence,
`heal_forked_single_use_oauth_grants`, `share_auth` on profile creation
(Desktop create dialog checkbox), and the docs. `profile_credential_audit.py`
(the `hermes update` "profiles without a provider" notice) is removed with it.
Kept from after #111724: `_save_codex_tokens(set_active=...)` for image gen,
the plugin-auth `status` dispatch and the external-login notice in
`hermes auth list`, and the registry-derived env-var hint in agent_init.
`gateway/host_attach.py::decide()` treated a host owner that answers `multiplex: False`
to the rescan — another profile's standalone gateway — the same as a multiplexer whose
roster excludes us, and issued the permanent REFUSE. On the supervised path that is
exit 78, which `hermes_cli/stderr_timestamp.py` maps to 0 so launchd's
`KeepAlive.SuccessfulExit=false` parks the unit. On a one-process-per-profile fleet
(the topology `multi-profile-gateways.md` documents and `multiplex_profiles: false`
promises to keep) every launchd gateway except the first to claim the host lock was
parked at boot, silently; which ones survived was a boot race against the owner's
record landing.
`request_serve_profile` now returns the owner flagged `standalone` instead of `None`
for a `multiplex: False` answer, and `decide()` turns that into START — this profile
runs its own gateway beside the owner, as before #118097; the host-lock claim still
logs the two-gateway topology and `gateway migrate --multiplex` stays the converge
path. REFUSE is unchanged for a multiplexing owner that excludes the profile.
A/B against a REAL owner process (real host lock, record and control socket answering
the rescan): base → `refuse`, exit 78 → launchd 0 (parked); fix → `start`. Negative
control (owner answers `multiplex: True`, roster excludes the profile): `refuse` on
both.
Field report: debug share 5024e996 (11 launchd profile gateways, 0.21.3 ea0c2b82 —
only 2 of 11 came back after `hermes update`), discussed on #118097.
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.
- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
unit->live-MainPID collapse rule. The legacy per-home marker stays readable
and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
idempotent per host (a completed restart onto the checkout SHA is never
repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
profile's home channels; the marker survives until each was reached.
Multiplex-only (Teknium ruling) runs exactly ONE `hermes gateway run` per host,
multiplexing every profile. Four reporting surfaces still asked a per-PROFILE
process question ("does MY profile own a gateway process?"), so a SERVED profile
answered "no" and the output lied:
* `hermes -p served cron status` printed "Gateway is not running" and told the
user to run `hermes gateway install` / `gateway run` for that profile — i.e.
to start a SECOND host process, which the architecture forbids.
* `hermes doctor` reported "Per-profile gateways: up/total" from s6 slots and
checked systemd linger against the CURRENT profile's unit, so a doctor run
under a served profile skipped the check entirely.
* `hermes claw`'s destructive-action warning was gated on `get_running_pid()`,
which returns None for a served profile — the token-conflict warning was
SILENTLY SKIPPED (a real safety hole).
* `gateway.status.multiplexer_liveness_for_profile()` returned None for the
default home by construction, so `default` could never be reported as SERVED.
* `hermes doctor`'s state.db holder + WAL wording implied one gateway per
profile.
New `gateway/host_topology.py` resolves "which single process owns the gateway
role on this host, and which profiles does it serve" once, from the host
rendezvous record (`gateway/host_rendezvous.py`), falling back to the default
home's recorded `served_profiles` for a gateway that predates the record.
`default` is just another served profile there.
Root cause: every surface derived gateway identity from per-profile artifacts
(argv `-p <name>`, `gateway.pid`, the active profile's systemd unit, s6 slot
counts) instead of the host record that actually names the owner.
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.
- ATTACH now requires a LIVE `identify` answer. The claim-time record is
published with NO served set (the runner settles multiplex a moment later),
and `host_gateway()` reports `served_known=False` when nothing answers. An
owner whose served set is unknown yields a TRANSIENT refusal, never an
attach: previously `default`'s claim published "default,other" before its
socket bound, `other`'s systemd unit read that as "I am served", exited 78,
and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
instead of forcing `multiplex=True`, so a standalone gateway stops claiming
the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
`--force` into `_host_attach_or_none`. Both previously exited in the guard
before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
config refusal every supervisor parks on; "someone else serves me right now"
is a runtime observation that ends when that process does. No unit files
change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
and re-enters with `replace=True`, so it can no longer attach to the corpse
it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
`st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
`gateway status`/doctor across N profiles pays one probe, not N.
Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.