Commit Graph

5912 Commits

Author SHA1 Message Date
ethernet
920a57e193 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-22 12:11:43 -04:00
ethernet
74ae184245 test(desktop): isolate expected payload fact failures
The macOS digest-order test emitted the strict PM failure from its expected
negative case into the shared Vitest log. That traceback looked like a second
packaging failure even though the test passed.

Run corrupt and missing fact checks in a captured child process. Keep the PM
reader strict, verify both errors, and retain the real builder ordering proof
for valid payload facts.
2026-09-22 12:01:37 -04:00
Siddharth Balyan
d1ce67c67a fix(desktop): the connector dialog's verb and switch sit in its header (#119128)
* fix(desktop): the connector dialog's verb and switch sit in its header

The primary verb (Install, Authenticate, Connect, Reconnect) and the
server or app switch move to the header's right cluster, so a dialog
that only carried a button in its lead band loses the band and gets
squarer. Stop waiting stays next to its sentence in the lead.

An app whose sign-in expired or broke keeps the whole-app switch, so it
can be turned off; before, the switch showed only for a live connection
and an expired account could neither be switched off nor removed. When
Nous refuses to remove a sign-in, the confirm dialog now says so and
points at the switch. The tool-rule switches stay hidden until the
connection is live, not merely present. The usage footer is one line.

* fix(desktop): a squarer connector dialog; connector reads retry while the gateway opens

The dialog is 32rem wide instead of 42rem. A read that fails with
"Hermes gateway is not connected" retries with backoff up to five
times; before, a policy read that raced the socket at startup stayed
failed and the tool rules read as unchangeable until Retry.

* feat(desktop): one picker says where a paired app runs

A catalog app that Nous also runs showed two rows under "How Hermes
reaches <App>", a Connect button on one and a switch on the other, and
nothing said which one was in use. The section is now "Where <App>
runs" with one segmented control, Managed or On this device, and only
the chosen way's own control under it. Picking a segment makes that way
the one in use: it turns the other way off, and turns a ready way on.
The tool list follows the picker; the dialog's own hosted/local tabs
are gone. Paired cards always open the managed dialog, which now also
carries the local server's usage line and Advanced section.

* fix(desktop): one control on the paired app's device row; the search clear button sits at the row's edge

Under the picker, the On this device row carried a switch next to its
button. The picker already decides whether the server runs, so the row
now has its state sentence and one button: Install, Turn back on when
the server is off, Authenticate when it needs a sign-in, nothing when
it runs. The unused switch row component is gone.

The Connectors search input takes the row's width, so the clear button
sits at the right edge instead of hugging the typed text.
2026-09-22 20:40:00 +05:30
ethernet
8da6c5f386 Merge branch 'ethie/release-machinery' into ethie/pm-clean
# Conflicts:
#	tests/ci/test_desktop_store_eligibility.py
2026-09-22 11:05:14 -04:00
ethernet
e0fb555beb Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-22 10:58:18 -04:00
ethernet
093c45325c fix(pm): keep platform detection out of bootstrap imports 2026-09-22 10:58:05 -04:00
kshitijk4poor
fde4997f58 test(desktop): split the scoped ws-url cases and explain the recovery route
Follow-up to the peer-window routing fix. One it.each in gateway-ws-url.test.ts
walked three scenarios in sequence while mutating its fixture (mockClear,
deleting the *For bridge) and branching on authMode in the body, so a failure
in the registry-scoped leg hid the missing-bridge leg. Each rung of the
adapter is now its own test with its own fixture. The spliced comment in
use-gateway-request.ts pointed at a "below" that lives in requestGateway, and
the e2e allowlist gained a WHY for inheriting the temp-dir variables.
2026-09-22 19:46:42 +05:30
BearHuddleston
84f41249aa fix(desktop): preserve remote profile routing in peer windows
(cherry picked from commit b0a34ec6163dc3c6c902abfd8ca6c05658b60c0e)
2026-09-22 19:46:42 +05:30
ethernet
207f8fedfd Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/local_runtime/binaries.py
#	hermes_cli/plugins_cmd.py
#	hermes_constants.py
#	tests/test_hermes_constants.py
#	tests/tools/test_clipboard.py
#	tests/tools/test_voice_wsl_pipewire.py
#	tools/computer_use/cua_backend.py
#	tools/voice_mode.py
2026-09-22 09:55:47 -04:00
kshitijk4poor
e2f8a0731b test(desktop): one discriminating guard test per empty-page site
The "records no signature" test could not fail: the second page differed
from the first, so it landed whether or not the ignored read left a
signature. Re-feed the SAME empty page once the runtime is genuinely empty
instead — that publish is deduped away only if the signature was wrongly
recorded (mutation-checked). Fold the plain "keeps the transcript" case into
it, keep the foreign-runtime carve-out, drop the two green-on-main boundary
tests, and add one guard each for the tile refresh and the post-turn hydrate.
2026-09-22 18:35:52 +05:30
kshitijk4poor
fbb0181224 fix(desktop): a kept truncated tail entry still counts as a use
Returning early on the empty page skipped setTranscriptTailEntry, and with
it the MRU bump an identical re-record gets — a run of empty pages let the
constantly re-read active session drift toward the eviction end of the
256-entry tail map. Re-record the kept entry instead: the identical-entry
path writes nothing and bumps the key.
2026-09-22 18:35:52 +05:30
kshitijk4poor
1993202ffb refactor(desktop): read the post-await session state once per refresh
Both refresh paths read $sessionStates twice after the await (once for the
stale-read check, once for the empty-page guard); hoist it into one local
that both consume. Keep the WHY of the guard on the predicate and leave the
call site with the one fact unique to it (bail before the signature write).
Type the page as readonly.
2026-09-22 18:35:52 +05:30
kshitijk4poor
2d46d124c7 fix(desktop): tile refresh and post-turn hydrate ignore an empty page too
The same zero-row page that blanked the active pane also blanked a
populated session tile (reconcileTileTranscripts published [] over it) and
was accepted by the post-turn fallback as the turn's answer. Apply the one
predicate at both sibling sites: the tile keeps its rows and no signature is
recorded; the hydrate falls through to its next attempt instead of
publishing.
2026-09-22 18:35:52 +05:30
Calvin Ng
c785dd8001 fix(desktop): keep a populated active transcript through an empty refresh page
A `sessions.changed` refresh can read zero rows while the backend respawns or
its state.db read races the event. reconcileActiveTranscript accepted that page
as the whole transcript, which blanked the view, tripped the routed
loading branch and re-ran the composer lifecycle (flash + caret reset).

- use-background-sync: ignore an empty page when the runtime for the SAME
  stored session already holds messages; leave the accepted-signature map
  untouched so the next usable page still lands. Same rule as the
  warm-activation guard (f0748b451c).
- transcript-tail: an empty page no longer downgrades a possiblyTruncated
  entry, so "Show earlier" stays armed through the transient read.

Companion to #118855 (unanchored stale tail) and #117898 (warm-gate hold);
this covers the third seam neither touches. Refs #118850, #117867, #98005.

(cherry picked from commit 070343591027017e95fdece1aa42e85d723da4b7)
2026-09-22 18:35:52 +05:30
Gille
2c65d5aee6 fix(desktop): clarify plugins folder label (#118848) 2026-09-22 08:34:36 -04:00
hermes-seaeye[bot]
171073f528 fmt(js): npm run fix on merge (#119114)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-22 12:31:58 +00:00
Siddharth Balyan
571ceec006 Desktop plugin card shows each declared server's state and the sentence that explains it (NS-941) (#119051)
* feat: expose plugin server states

* feat: reduce plugin server snapshots

* feat: render plugin server health
2026-09-22 17:54:42 +05:30
hermes-seaeye[bot]
f14f86dd5c fmt(js): npm run fix on merge (#119091)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-22 12:04:08 +00:00
Siddharth Balyan
f21e583211 chore(desktop): remove the old MCP tab (#119083)
The Connectors page replaced the MCP tab in #119074, but `McpTab` stayed
in the tree because the desktop SDK exported it and the bundled Bots
plugin rendered it in a bot's profile editor. The SDK now exports
`ConnectorsTab` in its place (same `gateway` and `profile` props) and the
Bots editor embeds that; older shells that lack the export still get the
plugin's own checkbox list, as before.

Deleted with the tab: its avatar component and the 45 `settings.mcp` and
`skills.tabMcp` strings nothing reads any more, in every locale.
2026-09-22 17:28:06 +05:30
hermes-seaeye[bot]
725cba2f80 fmt(js): npm run fix on merge (#119079)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-22 11:51:53 +00:00
Siddharth Balyan
dc50403a81 feat(desktop): the Connectors page replaces the MCP tab (#119074)
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours

The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.

- `tools/connectors/portal/`: a client for the portal's tool-list route and a
  JSON cache under the Hermes home, one file per portal origin and connector.
  An entry is fresh for 24 hours. After that the read revalidates with the
  stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
  serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
  chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
  `tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
  live in `tui_gateway/methods_connectors_account.py`.

The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.

* feat(connectors): catalog, accounts and member tool rules by RPC

The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.

- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
  the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
  first. The body is a union on `mode`, so a reader can name who turned a
  tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
  connector, or one connector on or off), with the revision the user saw. A
  stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
  write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
  page can show one card per app.

* feat(connectors): connect an app without a chat session

Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.

- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
  `connectors.operation.wake` and `connection.respond` take `owner`, a union
  on `type`: `session` (today's behaviour and authorization) or `account`
  (routed by `profile`, authorized by the live transport like `mcp.*`).
  `session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
  thread, under the profile's scope, so the watcher reads the account and
  settles the operation. A second connect for an app that is already
  connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
  address, so its updates go out on the session-less broadcast path.

* feat(mcp-catalog): eighteen more bundled entries name their hosted connector

A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.

* refactor(connectors): the account handlers share one gate, one params model and one write table

The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.

The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.

The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.

An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.

Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.

* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"

The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.

Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.

* fix(connectors): a connect from the page returns to the app after sign-in

The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.

Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.

* test(connectors): defer the new connector RPC coverage

The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.

Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.

Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.

* fix(cli): the connection panel hands the tool thread back at once

The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.

For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.

The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.

Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.

* fix(connectors): "run it again" lives in the library, so the classic CLI can use it

Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.

`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.

Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.

* feat(connectors): the account list and disconnect go through the portal

`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.

The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.

`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.

* fix(connectors): the account RPCs answer what the portal really sends

Checked against the portal source and against the staging and production
services.

- Errors are read from the upstream error code, not the HTTP status. A rule
  write answered 409 for a stale revision and for a user with no organisation;
  both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
  403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
  sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
  the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
  portal's own result for this user, with its stamp and without provider or
  subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
  and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
  answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
  invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
  so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.

Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.

* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link

Found by two adversarial reviews of the RPC layer and its types.

- `connectors.connect` from a chat session with no open operation is refused
  (`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
  registry with no card: it made a link nobody watched, returned a reply
  without the required `settled` field, and named an operation that was never
  registered. There is one way into an operation: the agent's call, or the
  account owner's `connectors.connect`. "Run it again" inside an open
  operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
  profiles with the same key no longer cross-deliver a sign-in link. The event
  payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
  OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
  `connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
  The last one is display data: the gateway enforces the rules, the backend
  only passes the list on. The phantom `name` and `description` are gone, and
  the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
  produces it. The contract generator now fails when a contract enum and its
  domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
  as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
  instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
  on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
  `optional-mcps/` are untouched by this PR again.

anti-slop: no net-new findings (15 touched files).

* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect

The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.

- The agent turn now declares how a link can reach the user
  (`tools/connectors/turn.py`): CARD when the agent was built with a
  connection callback, SIDE for a subagent or a background turn, LINK for a
  headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
  tool batch in the agent loop and read by the connector dispatch path, which
  never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
  link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
  CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
  tool on any path that derives the tool list, and a connector call on an
  unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
  does, so no `connection.update` is emitted for an operation no client asked
  for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
  `respondToConnectionRequest` share one guard and send nothing for a settled
  or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
  repeated it to the user.

Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.

* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients

A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.

- `tools/tool_labels.py` is the one place that turns a bridged call into a
  label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
  → "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
  keeps its own emoji, verb and primary-argument preview. A batch gets exactly
  one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
  failure text on the row of the call that failed. With friendly labels off
  it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
  carry a typed `labels` field. It does not depend on the classic CLI's
  display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
  labels, one row per call. The labels reach the row under a key no tool
  argument can use. The connect card it drew under a failed tool result is
  gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
  `manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
  "Reading tool details · N tools".

Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.

* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly

- A failed hosted search or describe used to return nothing, by design, so the
  model saw only local tools and told the user that a connected app was
  missing. The local results are unchanged; when the hosted leg failed, the
  `tool_search` and `tool_describe` results carry
  `connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
  and one hint line. A rejected token is `sign_in_expired`; an entitlement
  refusal or a shut gate adds nothing. `tool_describe` no longer lists those
  names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
  is a hosted connector account; `mcp: true` only when the user asks for an MCP
  server, a local server or an install, or when the name exists only in the
  catalog; connect and reconnect are hosted verbs, install, enable and
  authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
  gateway does not know the connector (confirmed on that failure path) and the
  name is a catalog entry does the target fail with "X is a local MCP server.
  Call manage_connections with action install ...". It is a per-target
  outcome: other targets of the same call keep their links and their card. A
  vendor failure on a name both sides know stays an ordinary failed row. The
  MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
  USER asks for that app again; the description and the settled-result notes
  say so. A builder saw the model refuse a direct user request.

Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.

* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled

Reproduced on the real Ink TUI with the rig, then fixed:

- The keyboard was dead during the sign-in wait: the card kept a `submitting`
  flag that the normal OAuth path never cleared, and Esc went through the same
  guard. The in-flight state now belongs to the answered row and clears when
  that row moves, when any later frame of the operation arrives, or after
  five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
  (the input handler had no branch for this overlay); Shift+arrows scroll the
  transcript and the card ignores them; arrow keys no longer move the text
  cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
  operation stayed in the store, and a resume dropped the pending card. The
  flag survives idle, a resume shows the pending card again, a session switch
  clears it.
- States with no branch: `not_connected` and a row with no link fell into the
  credential form; `expired` vanished with no note. The title and the row text
  now name the action (connect, reconnect, install, enable, authorize); a
  failed or expired row with no fields offers Try again / Skip; a failed row
  WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
  per app states the outcome. A settled or dismissed operation id is
  remembered, so no replay or resume can reopen its card. Esc in the last
  "Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
  the card in one sentence.

Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.

* chore(connectors): remove the comments and docstrings this branch added

Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.

Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.

* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures

Found by the end-to-end runs on the pushed head.

- Desktop: after a window reload, Continue on the restored card sent nothing.
  The answer looked up the backend that holds the session with the runtime
  session id, the lookup wants the stored id, and a failed lookup returned
  silently. When the lookup gives no owner the answer now goes out on the
  window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
  a refused scope, a non-member and a missing organisation alike: the handler
  runs with the gateway's globals and did not import the reason enum, so its
  own error mapping raised. `connectors.accounts.remove` caught auth failures
  in its generic branch. `org_required` was mapped on `policy.set` only. All
  six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
  `ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.

* wip(desktop): port the Connectors tab files and wiring onto the #115191 head

* wip(desktop): Connectors tab on the #115191 contract, catalog arm removed, audit defects fixed

* wip(desktop): Connectors tab passes the anti-slop ratchet; dormant two-ways code and the Available collapse removed

* wip(mcp): every server row says whether config or a plugin provides it; writes refuse plugin rows

* wip(desktop): Connectors tab, the owner's first live round (custom MCP form, kind words, compact dialog)

* wip(desktop): the connector dialog fits its content

* wip(desktop): catalog MCPs show on the Connectors tab until the catalog dies; connector_slug pairs a manifest with its managed app; the closed-gate state

* wip(desktop): connectors cache v3, the seed shape gained connector_slug

* wip(desktop): the owner's answers on the connectors page

A plugin-provided server now shows its tool list: the dialog probes it
through the existing read-only test endpoint, shows the tools without
switches (the plugin owns them), and shows the probe's error with a
Retry when the server cannot start. Its card is named after the server
key in the plugin's mcp.json, not the namespaced runtime key.

The paste box no longer parses `--header` on a `hermes mcp add` line;
the CLI has no such flag.

The rule write sends the member layer's revision only. The portal
always returns a member layer (baseline revision when no row exists)
and compares the write against that row, so the effective revision was
never the right guess. Verified live on staging: two writes in a row,
both accepted, policy restored.

The page cache keeps every read for signed-in accounts too and only
clears itself when the account is signed out. The storage version moves
to v4 so old blobs are ignored.

* chore(desktop): strip the prose comments the connectors page branch added

Comments and docstrings this branch added relative to main are gone;
tool directives, SAFETY lines and the contract docstrings that feed the
generated OpenRPC stay. Guards: Python AST and TypeScript printer output
are identical before and after; ruff, tsc, eslint, the ratchet and the
generated contracts are unchanged.
2026-09-22 17:16:00 +05:30
kshitijk4poor
e175ee3a4d test(desktop): trim kanban scope tests to their invariants
Store test header points at the store/connections.ts comment instead of
restating it; drop the two assertions implied by the final-tag check; drawer
test invalidates via taskKey() instead of a raw prefix literal. Revert the
comment-style churn on the unrelated $activeConnectionProfile subscription.
2026-09-22 17:00:02 +05:30
kshitijk4poor
944e017cbe fix(desktop): kanban queries fetch only while keyed to the routed connection
On a connection switch the request tag moves (applyActive →
setApiRequestConnection) before the descriptor publishes, and React re-keys
the kanban observers later still. The app-wide invalidations that fire in that
window (the connection twin, the profile twin) refetched observers still
sitting on the OUTGOING scope's key with the INCOMING tag — writing the new
gateway's boards/tasks under the old connection's cache key, which then
painted on the way back. Three fetches per query per switch, one of them
poison.

Install `enabled: routedToScope` as the `['kanban']` query default in bindApi:
a kanban query only fetches while its key's scope segment matches where a
request is routed right now (`host.activeConnectionId()`). Outgoing observers
are skipped; the incoming keys are already a cache miss and fetch once on
re-render. The drawer's two `enabled: !!id` sites compose it.

Also: the connection listener uses nanostores' previous value instead of a
mutable tracker; 'local' literal hoisted to LOCAL_SCOPE; duplicated rationale
comments cut.

Test: an observer on boardsKey('local') with the route moved to spark is not
refetched by invalidateQueries(); back on local it is. Red without the
default (`[null, 'spark']`), green with it.
2026-09-22 17:00:02 +05:30
kshitijk4poor
49f47ef64c refactor(desktop): connection-scope invalidation is a listen on $activeConnectionId
$activeConnectionId is a computed of a primitive: nanostores' set() drops
equal values and listen() does not fire on attach, so the undefined
sentinel, the change guard and the comment explaining them are already
provided by the store. One listener replaces the block riding the
$activeConnectionProfile subscription; both connection-scope tests hold.
2026-09-22 17:00:02 +05:30
kshitijk4poor
e4a1969fd6 fix(desktop): kanban keeps the bare slug key for local and dials once per switch
boardSlug.<scope> for every connection abandoned the slug existing users
had persisted under the bare key and broke the lib/connection-scoped.ts
contract (the local connection keeps the bare key, byte-identical storage
for single-backend users). Local now reads and writes `boardSlug`; remote
connections get `boardSlug.<id>`.

On a connection change whose persisted slug differed, hydrating $boardSlug
already reopened the socket through the $boardSlug listener, and the
listener then dialed a second time. Reopen explicitly only when the slug is
unchanged (the backend behind it still changed), so each switch dials once.

Test covers both plus the render-time key following the connection without
a manual rerender.
2026-09-22 17:00:02 +05:30
kshitijk4poor
a625189410 fix(desktop): kanban query keys read the connection scope reactively
The list keys (BOARDS_KEY, PROFILES_KEY, PROJECTS_KEY, ORCHESTRATION_KEY)
were module-level consts, so their scope segment was evaluated once at
import and stayed 'local' for the life of the app — the boards list, the
query behind the reported symptom, never got the cache miss the scoping
promised. The per-slug builders read host.state.connectionId.get() during
render with no subscription, so when the slug was the same on both
gateways (the default '' is) the observer kept the old-scoped key, the
twin's refetch stored the new gateway's data under the old scope, and the
next re-render flipped the key into a second fetch.

Every builder now takes the scope explicitly: rendering components get it
from useKanbanScope() (useValue on the SDK atom, so a switch re-renders
them and their observers move to the new key), non-rendering actions
(mutation settles, socket frames) from kanbanConnectionScope() at call
time. boardKeyPrefix(scope) replaces the raw ['kanban', 'board'] arrays
that leaked the key shape into six call sites.
2026-09-22 17:00:02 +05:30
Ben Barclay
8027c370a1 fix(desktop): kanban (and every connection-scoped query) follows the gateway switch
A connection switch left the kanban pane painting the previous
gateway's boards. Three holes, one root cause — the switch commit
runs its cache wipe BEFORE the activation publishes the new request
scope, and nothing re-invalidates after the tags move:

- store/connections: the CONNECTION twin of profile.ts's
  $activeGatewayProfile subscription. beginGatewaySwitch →
  invalidateProfileScopedQueries fires inside beforeActivate, so its
  refetches ride the OUTGOING backend; when the connection id moves
  (profile unchanged, e.g. default→default) nothing re-invalidates.
  Ride the existing $activeConnectionProfile.subscribe (no new
  computed listener) and invalidate on the actual id change; an
  undefined sentinel suppresses the first fire because the id is
  legitimately null on the local primary.

- plugins/kanban/api: every query key embeds the active connection
  scope (kanbanConnectionScope — the hermes-bots roster pattern), so
  a switch is a clean cache miss and one gateway's boards can never
  paint under another connection's route. The persisted board slug
  becomes per-connection storage (boardSlug.<id>): a slug that pins a
  localhost board 404s on the next gateway and froze the pane on the
  last successful (localhost) data.

- plugins/kanban/api: the events socket re-opens on a connection
  change — pluginSocket resolves the backend only at connect time, so
  a socket left open across a switch streamed the old gateway's
  events against the new connection's cache forever.

Regression tests in src/store/kanban-connection-scope.test.ts drive
the REAL selectConnection chain: the post-switch refetch must be
tagged for the NEW gateway, and a no-change registry republication
must not refetch. Mutation-checked: both fail without the twin.

(cherry picked from commit 3ea28be090c3ee885b296694c74be40f84081d2c)
2026-09-22 17:00:02 +05:30
ethernet
c13287c915 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/electron/main.ts
#	hermes_cli/backup.py
#	hermes_cli/config.py
#	hermes_cli/plugin_catalog.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/plugins_cmd_catalog.py
#	hermes_cli/plugins_discovery.py
#	hermes_cli/profiles.py
#	hermes_cli/update_cmd_deps.py
#	pyproject.toml
#	tests/gateway/test_dm_topics.py
#	tests/hermes_cli/test_config.py
#	tests/hermes_cli/test_plugins_cmd.py
#	tests/hermes_cli/test_update_autostash.py
#	tests/tools/test_lazy_deps.py
#	tools/lazy_deps.py
#	tools/skill_ledger.py
#	utils.py
#	website/docs/user-guide/security.md
2026-09-22 05:16:50 -04:00
teknium1
5c0e73eff1 feat(desktop): render plugin-declared settings in the Plugins tab (#46600, #87934)
A plugin manifest's `config_schema` now reaches the Desktop: `plugins.manage list`
returns each plugin's schema with the current `plugins.entries.<id>.settings`
values (`settings_schema`), and a new `settings` action writes edits through
`hermes_cli.plugins_state.save_plugin_setting` — the writer extracted from
`PluginContext.set_config`, so the plugin, the CLI and the Desktop share one
config path, one lock and the same managed-install / managed-key refusals.

The Plugins tab grows a gear per plugin with a schema; the inline form is
table-driven (`FIELD_CONTROLS` / `INITIAL_TEXT` / `COERCE` keyed on the wire
field type) for string / number / boolean / enum / json / secret. Secrets are
declared with `type: secret`: the row carries only the `.env` name and a
presence flag, the client writes the value through the existing `PUT /api/env`
credential route, and the RPC refuses secret keys so nothing lands in
config.yaml.

Contracts regenerated; docs gain a "Settings form in the Desktop" section.
2026-09-22 01:48:18 -07:00
hermes-seaeye[bot]
30de0e01e6 fmt(js): npm run fix on merge (#118939)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-22 08:34:14 +00:00
Siddharth Balyan
70f5dc5f46 feat(connectors): the backend API for the desktop Connectors page; connect an app without a chat session (#115191)
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours

The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.

- `tools/connectors/portal/`: a client for the portal's tool-list route and a
  JSON cache under the Hermes home, one file per portal origin and connector.
  An entry is fresh for 24 hours. After that the read revalidates with the
  stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
  serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
  chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
  `tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
  live in `tui_gateway/methods_connectors_account.py`.

The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.

* feat(connectors): catalog, accounts and member tool rules by RPC

The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.

- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
  the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
  first. The body is a union on `mode`, so a reader can name who turned a
  tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
  connector, or one connector on or off), with the revision the user saw. A
  stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
  write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
  page can show one card per app.

* feat(connectors): connect an app without a chat session

Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.

- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
  `connectors.operation.wake` and `connection.respond` take `owner`, a union
  on `type`: `session` (today's behaviour and authorization) or `account`
  (routed by `profile`, authorized by the live transport like `mcp.*`).
  `session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
  thread, under the profile's scope, so the watcher reads the account and
  settles the operation. A second connect for an app that is already
  connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
  address, so its updates go out on the session-less broadcast path.

* feat(mcp-catalog): eighteen more bundled entries name their hosted connector

A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.

* refactor(connectors): the account handlers share one gate, one params model and one write table

The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.

The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.

The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.

An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.

Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.

* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"

The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.

Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.

* fix(connectors): a connect from the page returns to the app after sign-in

The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.

Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.

* test(connectors): defer the new connector RPC coverage

The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.

Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.

Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.

* fix(cli): the connection panel hands the tool thread back at once

The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.

For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.

The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.

Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.

* fix(connectors): "run it again" lives in the library, so the classic CLI can use it

Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.

`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.

Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.

* feat(connectors): the account list and disconnect go through the portal

`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.

The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.

`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.

* fix(connectors): the account RPCs answer what the portal really sends

Checked against the portal source and against the staging and production
services.

- Errors are read from the upstream error code, not the HTTP status. A rule
  write answered 409 for a stale revision and for a user with no organisation;
  both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
  403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
  sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
  the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
  portal's own result for this user, with its stamp and without provider or
  subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
  and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
  answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
  invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
  so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.

Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.

* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link

Found by two adversarial reviews of the RPC layer and its types.

- `connectors.connect` from a chat session with no open operation is refused
  (`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
  registry with no card: it made a link nobody watched, returned a reply
  without the required `settled` field, and named an operation that was never
  registered. There is one way into an operation: the agent's call, or the
  account owner's `connectors.connect`. "Run it again" inside an open
  operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
  profiles with the same key no longer cross-deliver a sign-in link. The event
  payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
  OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
  `connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
  The last one is display data: the gateway enforces the rules, the backend
  only passes the list on. The phantom `name` and `description` are gone, and
  the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
  produces it. The contract generator now fails when a contract enum and its
  domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
  as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
  instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
  on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
  `optional-mcps/` are untouched by this PR again.

anti-slop: no net-new findings (15 touched files).

* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect

The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.

- The agent turn now declares how a link can reach the user
  (`tools/connectors/turn.py`): CARD when the agent was built with a
  connection callback, SIDE for a subagent or a background turn, LINK for a
  headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
  tool batch in the agent loop and read by the connector dispatch path, which
  never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
  link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
  CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
  tool on any path that derives the tool list, and a connector call on an
  unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
  does, so no `connection.update` is emitted for an operation no client asked
  for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
  `respondToConnectionRequest` share one guard and send nothing for a settled
  or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
  repeated it to the user.

Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.

* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients

A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.

- `tools/tool_labels.py` is the one place that turns a bridged call into a
  label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
  → "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
  keeps its own emoji, verb and primary-argument preview. A batch gets exactly
  one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
  failure text on the row of the call that failed. With friendly labels off
  it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
  carry a typed `labels` field. It does not depend on the classic CLI's
  display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
  labels, one row per call. The labels reach the row under a key no tool
  argument can use. The connect card it drew under a failed tool result is
  gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
  `manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
  "Reading tool details · N tools".

Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.

* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly

- A failed hosted search or describe used to return nothing, by design, so the
  model saw only local tools and told the user that a connected app was
  missing. The local results are unchanged; when the hosted leg failed, the
  `tool_search` and `tool_describe` results carry
  `connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
  and one hint line. A rejected token is `sign_in_expired`; an entitlement
  refusal or a shut gate adds nothing. `tool_describe` no longer lists those
  names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
  is a hosted connector account; `mcp: true` only when the user asks for an MCP
  server, a local server or an install, or when the name exists only in the
  catalog; connect and reconnect are hosted verbs, install, enable and
  authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
  gateway does not know the connector (confirmed on that failure path) and the
  name is a catalog entry does the target fail with "X is a local MCP server.
  Call manage_connections with action install ...". It is a per-target
  outcome: other targets of the same call keep their links and their card. A
  vendor failure on a name both sides know stays an ordinary failed row. The
  MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
  USER asks for that app again; the description and the settled-result notes
  say so. A builder saw the model refuse a direct user request.

Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.

* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled

Reproduced on the real Ink TUI with the rig, then fixed:

- The keyboard was dead during the sign-in wait: the card kept a `submitting`
  flag that the normal OAuth path never cleared, and Esc went through the same
  guard. The in-flight state now belongs to the answered row and clears when
  that row moves, when any later frame of the operation arrives, or after
  five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
  (the input handler had no branch for this overlay); Shift+arrows scroll the
  transcript and the card ignores them; arrow keys no longer move the text
  cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
  operation stayed in the store, and a resume dropped the pending card. The
  flag survives idle, a resume shows the pending card again, a session switch
  clears it.
- States with no branch: `not_connected` and a row with no link fell into the
  credential form; `expired` vanished with no note. The title and the row text
  now name the action (connect, reconnect, install, enable, authorize); a
  failed or expired row with no fields offers Try again / Skip; a failed row
  WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
  per app states the outcome. A settled or dismissed operation id is
  remembered, so no replay or resume can reopen its card. Esc in the last
  "Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
  the card in one sentence.

Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.

* chore(connectors): remove the comments and docstrings this branch added

Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.

Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.

* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures

Found by the end-to-end runs on the pushed head.

- Desktop: after a window reload, Continue on the restored card sent nothing.
  The answer looked up the backend that holds the session with the runtime
  session id, the lookup wants the stored id, and a failed lookup returned
  silently. When the lookup gives no owner the answer now goes out on the
  window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
  a refused scope, a non-member and a missing organisation alike: the handler
  runs with the gateway's globals and did not import the reason enum, so its
  own error mapping raised. `connectors.accounts.remove` caught auth failures
  in its generic branch. `org_required` was mapped on `policy.set` only. All
  six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
  `ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
2026-09-22 13:57:51 +05:30
kshitijk4poor
88b9d35ee0 fix(desktop): log the latched-failure refusal once in the supervisor, not per request
799a4cc7f2 put a rememberLog before each of the three latched-failure
throws in runHermesStart (main.ts:13257/13262/13270). That path is not
rare while a failure is latched: ensureBackend() calls startHermes()
for every proxied primary API request, and the renderer's boot retries
(BOOT_RETRY_MAX_ATTEMPTS=5) hit it too. Each hit pushed an identical
timestamped line into the 300-line hermesLog ring and the disk flush
buffer, evicting the lines that explain the original failure -- the
same eviction 3f9380daba was fixing.

Remove the per-request lines (runHermesStart's short-circuit is as
cheap and silent as on base, WHY comments kept) and log the refusal
once at the only new consumer that needs the reason: the terminal-latch
early return in runPrimaryRecoverySpawn, which runs once per refused
respawn. Both sites now go through latchedBootFailure() (ordered
bootstrapFailure -> backendStartFailure -> remoteReauthFailure, same
order as before) so the trio can no longer drift between them.

Proof: tsc (tsconfig.electron.json) and eslint clean; vitest 25/25
(backend-start-failure + backend-exit-recovery). Per-request path: the
only rememberLog left in runHermesStart before the E2E block is the
pre-existing non-primary-instance line (awk over the function body);
`grep -c 'latched; refusing restart'` = 0. Helper probe:
(null,null,null)->null, (B,S,R)->B, (null,S,R)->S, (null,null,R)->R.
2026-09-22 13:36:49 +05:30
kshitijk4poor
774d851756 refactor(desktop): decide the supervisor-respawn latch inside shouldLatchBackendStartFailure
main.ts:13720 hand-rolled `!supervisorRecovery &&` in front of
shouldLatchBackendStartFailure(), so the pure predicate whose sole job
is "should this boot failure latch?" (backend-start-failure.ts:1-41)
gave the wrong answer for a supervisor-owned respawn, and the new rule
had no test outside the untestable 500-line runHermesStart.

Thread `supervisorRecovery?: boolean` through BackendStartFailureContext
(optional, so the existing call sites and tests are unchanged), make
the body `!attemptedRemote && !supervisorRecovery`, move the WHY
("supervisor respawn already has its own bounded crash-loop budget")
onto the predicate docblock, and add the invariant test
`{ attemptedRemote: false, supervisorRecovery: true } -> false`.

Proof: with the predicate hand-edited back to `!context.attemptedRemote`
the new test fails (1 failed | 18 passed); at head 19/19. tsc
(tsconfig.electron.json) and eslint clean. Pure boolean AND, so the
call-site behaviour is identical.
2026-09-22 13:36:49 +05:30
kshitijk4poor
1272ec3e66 test(desktop): cover retryAfterFailedStart on an unclaimed recovery latch
The `!claimed` guard in retryAfterFailedStart (backend-exit-recovery.ts:81)
is what stops a pre-ready failure that the supervisor never claimed (a
user-driven start, or a start after the previous claim was released by
reset()) from re-arming recovery and double-spending the 3/120s window. It
was exercised only by the live probe, not by a test: the existing retry
tests all start from a claimed latch.

Add the negative case: retry on a fresh latch and on a reset() latch both
return false, do not flip isCrashLooping, and leave the remaining budget
intact (the next two claims are granted, the fourth is the real
exhaustion).

RED proof: dropping `!claimed ||` from the guard fails this test at
"fresh latch has no claim to re-arm" (1 failed / 5 passed); restored ->
6/6 green.

Gate 2c suggestion: apps/desktop/electron/backend-exit-recovery.test.ts.
2026-09-22 13:36:49 +05:30
kshitijk4poor
65f623b10d docs(desktop): state the releaseStart/.catch ordering contract in startHermes
retryAfterFailedStart's correctness (main.ts:13186) depends on
`start.then(releaseStart, releaseStart)` (main.ts:13148) being registered on
the same promise startHermes returns, before runPrimaryRecoverySpawn attaches
its `.catch`. Same-promise reactions run in registration order, so
primaryStartsInFlight is already 0 and primaryRecoveryState() reports
hasPendingStart:false when the retry is evaluated (verified by microtask
probe in gate 2ab). Nothing in the code said so: a refactor that returns
`start.then(...)` or wraps `start` in a new promise would silently make every
pre-ready retry hasPendingStart:true -> refused -> reportPrimaryRecoveryCrashLoop
false -> respawn failure dies with no retry and no UI, the exact silent state
this salvage fixes.

Comment-only; no behaviour change.

Gate 2ab warning (med): main.ts:13147 vs 13176.
2026-09-22 13:36:49 +05:30
kshitijk4poor
0b283ec9c9 fix(desktop): log only the first line of re-logged boot/respawn errors
The `[boot] ... latched; refusing restart` lines (main.ts:13248/13253/13261)
and `[supervisor] backend respawn failed` (main.ts:13177) re-logged the full
error message. The pre-ready-exit error is built as `...Log: <path>\n` +
recentHermesLog() (20 prior ring lines), and rememberLog() splits on
newlines and pushes every line into the 300-line hermesLog ring and the
desktop log file. Each refused getConnection() (the renderer's
ensureGatewayOpen retries on every failed request) therefore appended ~21
lines, so ~14 refusals evicted the whole ring — including the original
failure evidence these lines exist to preserve. The respawn line is now
fired up to 3x per storm as well.

Reuse the existing module-scope `firstLine` helper (main.ts:3283) so each
re-log costs one ring line; the full message still reaches the renderer
via the thrown error / sendBackendExit unchanged.

Gate 2c warning: apps/desktop/electron/main.ts:13177,13248,13253,13261.
2026-09-22 13:36:49 +05:30
kshitijk4poor
465bab3758 fix(desktop): log the latched-failure short-circuits in runHermesStart
Once bootstrapFailure, backendStartFailure or remoteReauthFailure is
latched, every later startHermes() -- including the overlay's "Reconnect
now" -- re-throws the latched error without writing a line, so the
report log shows the original exit and then nothing, and the user's
retry looks like a dead button (#118680). Log one [boot] line at each
short-circuit so the trace shows the latch is what refused the restart.
2026-09-22 13:36:49 +05:30
kshitijk4poor
7f36291720 refactor(desktop): keep the #112344 comment on scheduleUnexpectedPrimaryRecovery
The latch-helper extraction inserted primaryRecoveryState() and friends
between the supervisor rationale comment and the function it describes,
so the "a ready primary child died" paragraph read as documentation for
a plain state getter. Move the comment back onto the entry point it
explains. Comment placement only; no behaviour change.
2026-09-22 13:36:49 +05:30
JoaoMarcos44
d905a80ac2 fix(desktop): preserve terminal recovery latches
(cherry picked from commit 3b4eef59dde4f0e8a0860a02146d9a00dc8cd38a)
2026-09-22 13:36:49 +05:30
JoaoMarcos44
71f9e94d81 chore(desktop): keep recovery tests formatted
(cherry picked from commit 3ab0117ccf3aefa16cd8d9b230097a71a4f8868b)
2026-09-22 13:36:49 +05:30
JoaoMarcos44
ff802df650 fix(desktop): retry failed supervisor respawns
(cherry picked from commit c60e707f6c82418a94cf7d5be1fa7562e3052c6a)
2026-09-22 13:36:49 +05:30
JoaoMarcos44
1751ccb628 test(desktop): cover pre-ready recovery retries
(cherry picked from commit 4eb9a7c2a9af21cf6dc60aed39831b277d23b3c2)
2026-09-22 13:36:49 +05:30
JoaoMarcos44
86d18d6f15 fix(desktop): re-arm failed pre-ready recovery
(cherry picked from commit 4b2012342a05f9799938f8c58b004cc08e705d0b)
2026-09-22 13:36:49 +05:30
teknium1
db8dcbe94c fix: catalog re-pins ask before widening a plugin; annotated-tag pins keep reviewed trust
Annotated-tag pins (F8): a catalog `sha` recorded as `git rev-parse <tag>` names
the TAG object, while HEAD can only ever be the commit it points at. The scan
trust check compared HEAD against the unpeeled sha (so every tag-pinned entry
lost the reviewed-pin bypass and prompted on caution findings) and the sidecar
recorded the peeled commit, so `update_available` was true forever and every
`update` re-installed. The installer now peels the pin (`<sha>^{commit}`) for
trust, records `pin` on the catalog block only when the checkout satisfies it
(empty for an off-pin `--ref` install), and every at-pin check goes through
`at_catalog_pin(sidecar, entry_sha)` (repin, dashboard payload, TUI rows).

Re-pin consent (F10): `hermes plugins update` on a catalog install replaced the
tree without asking, even when the new pin declared new tools, hooks, Python
dependencies, host capabilities or a Desktop half. `repin_catalog_plugin` now
diffs the installed manifest against the staged clone BEFORE anything moves
(`_install_plugin_core(before_swap=...)`) and, on a widening:
- CLI: prints the delta and asks y/N (non-interactive → not applied, fail
  closed); after a changed re-pin it runs the same `_run_capability_consent`
  grant path as the git-pull `update`.
- `plugins.manage update` RPC and the dashboard REST route answer
  `{ok: false, consent_required: true, delta, delta_lines}` with nothing
  changed; a retry with `accept_capabilities: true` applies it. Desktop shows
  the delta in its confirm dialog; the web dashboard uses `window.confirm`.
- Gateway contract regenerated (`accept_capabilities` param; `consent_required`,
  `delta`, `delta_lines`, `error` result fields).

Catalog audit findings F8 and F10 (low severity, no issue filed).
2026-09-22 01:00:09 -07:00
hermes-seaeye[bot]
8e221d12a8 fmt(js): npm run fix on merge (#118910)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-22 07:52:40 +00:00
teknium1
e1a6679399 fix(desktop): runtime plugin loader survives hung imports, leaked timers, duplicate ids and broken reloads
Four error-isolation holes in the disk plugin door
(apps/desktop/src/contrib/runtime-loader.ts), found by a static+live audit of
the loader; none had an issue filed.

- A plugin whose module evaluation never settles (top-level `await` on a dead
  host) hung `import()` forever and, through the scan's sequential loop and
  its re-entrancy guard, froze every later plugin and all future scans until
  restart. `import()` now races a 10 s deadline; the plugin errors on its own
  row ("import timed out") and the scan continues.
- Timers and DOM listeners a plugin took out with bare globals survived
  disable and every hot-reload. `ctx.setTimeout` / `ctx.setInterval` /
  `ctx.addEventListener` are tracked with the plugin and torn down on
  unload; the SDK doc says bare globals are not.
- Two folders exporting one plugin id silently last-wins: the second
  disposed the first's registrations and each hot-reload flipped ownership.
  The first (folder-name sorted, so deterministic) owns the id; the later
  file errors on its own row ("duplicate id, already loaded from <path>").
- A save that no longer loads (syntax error, timeout, duplicate) left the old
  incarnation's contributions and activate handle live beside the error
  row, so the Plugins tab showed a broken file as "loaded" and could
  re-enable stale code. The previous incarnation is unloaded and dropped.

Tests: one invariant per fix in runtime-loader.test.ts, all red on base
(the hang case red by timing out).
2026-09-22 00:46:24 -07:00
hermes-seaeye[bot]
969872ebaa fmt(js): npm run fix on merge (#118897)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-22 07:30:02 +00:00
kshitijk4poor
d8552004ca test(desktop): partially mock session-states in the status-drawer test
The background-sync hook now hydrates todos on refresh, which pulls
`@/store/todos` into this test's import graph; its module-level computed reads
`$sessionStates`, so a mock that only defines `knownOwnerForSession` fails at
load. Spread the real module and override the one function the test cares about.
2026-09-22 12:54:27 +05:30
kshitijk4poor
a835405065 fix(desktop): bound the sealed-reply settle and count hidden prompts as turns
Review of the follow-ups on the final head:

- The no-delta settle looked for ANY earlier interim row with equal text; a
  tool-round interim two turns back with the same prose would have absorbed the
  completion. Only the assistant row directly before the boundary qualifies.
- Hidden prompts (bots mode, slash commands) start turns and the completion
  reducer already treats them as boundaries, so the journal's "last committed
  turn" starts at the last user row, hidden or not.
- Pins that durable identity retires a journal row regardless of turn position
  (the branch the restored fixtures no longer exercised).
- `settleAt` replaces four copies of the settle-in-place map; the echo fold
  uses a type guard instead of casts; `identityCovers` names the rule.
2026-09-22 12:54:27 +05:30
kshitijk4poor
5aff954e15 refactor(desktop): dedupe the live-turn reconciler seams
- `reconcilePersistedLiveTurn` called `reconcileAuthoritativeChatMessages`
  back through a callback that was the caller itself; the projection-less
  branch is now `reconcileDurableHistory` in utils and imported directly.
- The activation path retried the live-turn reconcile inside its fallback
  with the same rows; `null` does not depend on `previous`, so the retry was
  dead work. The fallback now runs the durable path only.
- `isLiveTailReplyId` (spoken-reply) replaces the third and fourth copies of
  the `assistant-stream-`/`inflight-assistant-` prefix predicate; `normalizeWs`
  is exported from chat-messages instead of re-declared in coverage.ts and
  the journal; `isCommittedRow` names the "not a projection, not a recovery"
  predicate the journal repeated three times.
- Loop-invariant `lastPreviousUser` hoisted; `renderedText(raw)` computed once
  per row instead of per part; `snapshotIntervals` skips the code-point split
  when there are no corrections; `persistInFlightTurnState` checks for a
  recovered row before building the recoverable tail on every idle commit.
2026-09-22 12:54:27 +05:30