Preview-pane guest pages (Streamlit's traceback "Ask Google" / "Ask …"
buttons, plain `<a target="_blank">` anchors) could not open anything: the
`<webview>` has no `allowpopups`, so Chromium drops the popup before any
handler runs. Opening from `setWindowOpenHandler` is banned
(GHSA-9f4c-93c8-jc8g, window-open-policy.ts), so this adds an explicit
click bridge instead:
- main.ts installs a guest preload through `will-attach-webview`, keyed on
the `persist:hermes-preview` partition only.
- The preload forwards a clicked `_blank` anchor's resolved href to the
host renderer via `ipcRenderer.sendToHost`; it opens nothing itself.
- PreviewPane admits the URL and routes it through the existing audited
`hermes:openExternal` IPC.
Salvaged from PR #112959 (squash of its two commits).
Fixes#112941
Same stale invariant the ui-tui test had: since the shared channel sends one
client.capabilities frame on every gateway.ready, "no frames" is no longer
what this case guards. What it guards is that an older backend without
heartbeat gets no gateway.ping and stays open, so assert the wire holds
exactly the advertisement.
The lib test sits beside markdown-preprocess.directives.test.ts as
markdown-preprocess.reasoning.test.ts, and the frame-fidelity test joins
its markdown-text.*.test.tsx siblings as markdown-text.reasoning.test.tsx,
so a glance at the directory shows what each file exercises. Also trim the
REASONING_TAGS comment: how the two tag lists converged is git history, not
something a reader of the constant needs.
OPEN_REASONING_BLOCK_RE needs the full `<tag>`, so a frame ending in `<thin`
at a block boundary painted `<thin` as prose for one frame and then erased it
— the paint/un-paint class of #62774, one frame long. The fidelity test carved
an escape hatch around exactly those frames.
Add a third pass that drops a trailing partial open tag at a block boundary,
restricted to prefixes of the known reasoning tag names (built from
REASONING_TAGS) so `<div` at a line start still renders. This is the desktop
counterpart of agent/think_scrubber.py `_hold_partial`/`_max_partial_suffix`,
cited rather than ported. The fidelity test's escape hatch is deleted so the
monotonic-prefix invariant holds on every frame; the unit test gains the
`<thin` positive and `<div` negative cases.
stripReasoningBlocks runs on the accumulated message text every streaming
flush (markdown-text.tsx). Its replacer sliced the whole text twice per
closed block to test whether a space is needed at the seam — O(n) copies plus
a forward scan per block per flush. Read the single char on each side of the
match instead.
The seam test also had a second defect: with two adjacent closed blocks the
second block's `before` ended in the first block's `>`, so a second space was
emitted (`no Hermes`). Matching a run of adjacent blocks as one match makes
the two-edge check see the real prose on both sides. Covered by an extra
expect in the existing seam test.
WHAT: one REASONING_TAGS alternation feeds both regexes and now also covers
`thought` and `reasoning_scratchpad` (agent/think_scrubber.py THINK_TAG_NAMES);
the desktop-only `scratchpad`/`analysis` stay. Behaviour change beyond the
contributor's: closed or unterminated `<thought>`/`<reasoning_scratchpad>`
blocks are now stripped on the desktop too.
Doc comments trimmed to the WHY (block-boundary rule, why an unterminated block
is held back). Tests cut to the ≤2-invariant shape: two unit cases (seam,
unterminated-vs-mention) and one DOM frame-by-frame case; the accents/plain-prose
DOM cases were protection for unrelated code and are dropped.
preprocessMarkdown runs on the accumulated text on every streaming flush, so
the plain REASONING_BLOCK_RE replace damaged the visible answer in two ways a
settled message never shows:
- a model that inlines its chain of thought in the answer channel streams the
open tag long before the close one, and the regex needs both, so the reasoning
rendered as chat prose until the close tag arrived — and when it did the whole
span vanished in one frame, taking text the reader had already read with it
(#62774: the mid-reply "truncation" on long answers).
- the regex ate the whitespace on its seam, so a block sitting between two words
fused them: `no` + `Hermes` rendered as `noHermes`.
An unterminated block now strips at a block boundary — the same rule
agent/think_scrubber.py draws, so a real reasoning preamble (always its own
block) disappears while prose that merely mentions `<thinking>` mid-sentence
survives — and a removal between two prose fragments leaves one space instead of
deleting the seam.
Tests: markdown-reasoning-stream.test.ts (module level: seam, unterminated
blocks, per-delta leak + visibility monotonicity) and
streaming-text-fidelity.test.tsx (real surface: accents/emoji and plain prose
streamed in 1- and 3-character deltas, plus the chain of thought never painting
a frame).
With the paste-time interception gone nothing calls
`git.review.fetchPrComment`: remove the Electron handler, its preload entry,
the `parsePrCommentUrl`/`reviewFetchPrComment` helpers (a second copy of the
renderer's regex, free to drift), the remote-gateway stub and the
`HermesPrComment` typing. Reads only; no user-visible behaviour.
Separate commit so it can be dropped independently of the paste fix.
Cherry-picked hunks from 74b31c5b79d92 (#112480); the `reviewFetchPrComment`
export added on main since is removed too.
The composer's paste handler special-cased exactly one URL shape: a GitHub
PR-comment deep link (`…/pull/<n>#issuecomment-<id>` or `#discussion_r<id>`)
called `onAttachPrCommentUrl`, ran `preventDefault()`, and returned. The
clipboard payload never reached the editor — pasting into an empty composer
left it empty — and the URL existed only as an attachment pill above the
composer. Removing that pill lost the URL for good: nothing reinserted the
text and the path recorded no undo point.
Every URL now takes the ordinary link path. The renderer-side `review`
attachment machinery that only the interception fed goes with it:
`attachPrCommentUrl`, the `onAttachPrCommentUrl` prop/wiring/adapter entry,
the `review` attachment kind, `reviewCommentBlock` and `PR_COMMENT_URL_RE`
(verified by grep: no remaining consumer). The Electron `fetchPrComment` IPC
is removed in the next commit so it can be dropped independently.
Test: paste-url-is-text.test.tsx mounts the real ChatBar, fires a paste on
the contentEditable and pins the contract — the payload reaches the editor,
nothing is attached, the intercepting hook is never consulted.
Slim redo of the Desktop half of #112480 (cherry-picked from 74b31c5b79d92,
resolved onto the FloatingComposerSurface wrapper on main).
Review: the dialog test asserted a styling detail (`whitespace-pre-line`)
on top of the behavioural invariant; the textContent check already proves
the guard's paragraphs survive, so the class match only couples the test
to CSS. Also add the one missing sentence to the Desktop user guide so
"Switch anyway" / "Keep current model" is discoverable from the docs.
When the gateway answers `confirm_required` (large cached context,
expensive model, data-training tier), the Desktop asked for confirmation in
a warning toast with a single "Confirm" action and an ✕ — no button meant
"no", Enter/Esc did nothing, the four-item toast stack could evict the
pending question, and the guard's `\n`/`\n\n` paragraphs collapsed into one
run-on line. Answering after the session had moved on was a silent no-op.
`surfaceModelSwitchConfirm` (THE shared applier for both surfaces — the
composer picker via `config.set` and the Bots editor via
`profiles.configure`) now asks through `confirm()` from `@/store/confirm`,
i.e. the shell-mounted ConfirmDialog that apps/desktop/DESIGN.md names as
the only way to ask "are you sure": destructive "Switch anyway" vs "Keep
current model", Enter confirms, Esc/backdrop/✕ decline, the dialog owns
focus. Declining is free — nothing was applied before the answer. A stale
confirmed answer now toasts "Selection changed — the model switch was not
applied" instead of doing nothing. ConfirmDialog renders its description
`whitespace-pre-line` so backend-composed paragraphs survive. The applier
resolves `true`/`false` instead of returning a notification id; callers
fire-and-forget it. New i18n keys land in all six locales.
Slim redo of #112463 by @DavidMetcalfe (design: route through confirm(),
labels, stale notice). #112461 by @KoNit-K was the earliest filer (Cancel
action on the toast) and is superseded by the dialog.
Fixes#112458
Co-authored-by: DavidMetcalfe <80915+DavidMetcalfe@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
toolCallOwnerMessageId walked the whole transcript newest-first for any part
carrying the tool_call_id. Tool call ids are not unique across turns:
llama.cpp emits one constant id for every call and Hermes' own deterministic
ids repeat (#76632), so the next turn's tool.start/tool.complete was routed
into the previous turn's bubble — the new call was drawn over the finished
row (args and result overwritten) and never appeared in its own turn. That
is a regression against main, which only ever upserted into the live bubble.
Only an unresolved part (no `result` key: still running, or sealed by an
interim/settle/mid-turn boundary without its completion) can own an event.
A part that already carries its completion is a finished call from an
earlier turn; the lookup skips it and the event seeds a fresh part in the
live bubble as before. This is the same "completed" predicate
sealOpenToolParts uses, so a completion that arrives with an error and no
result still counts as resolved. The late-completion re-attach for #113035
is unchanged because sealed-but-unresolved parts remain eligible.
Tests: late-tool-complete.test.tsx and late-tool-result.test.tsx (four cases
for one fix) fold into late-tool-events.test.tsx with the two load-bearing
invariants (late completion re-attaches to the sealed bubble; a running event
on a sealed row does not seed a duplicate) plus one new invariant for the
reused id, which fails on the previous head with one row instead of two.
Salvage follow-up to the cherry-picked #113056: the completion-phase lookup
attached a late `tool.complete` to the bubble that owns the call, but a
running-phase event (`tool.start` replay / progress) for a call whose part
was already sealed still seeded a second live row with its own timer under
the user's mid-turn message — the "two rows for one call" the reporter
photographed in #113035. Both phases now resolve their target by tool id.
Why the bookkeeping change: when the target is NOT the live stream, the
event is a patch to history. The sealed bubble keeps its own pending bit
(a running event must not re-arm a settled/interim bubble), and streamId /
sawAssistantPayload / awaitingResponse are left alone so a late event from
the previous phase cannot drop the wait state of the turn now in flight.
`upsertToolPart`: a running event landing on a part sealed without a result
drops the seal, so the row ticks again instead of reading "Result
unavailable" next to a ticking duplicate.
Tests: contributor suite trimmed to two invariants (owner attach across an
interim boundary, no-owner completion still seeds); new suite covers the
post-settle re-attach and the mid-turn user insert with a running replay.
Interim commentary and turn settles seal the streaming bubble and drop the
stream id while a long-running tool is still executing. When the completion
finally arrives, mutateStream had no stream target left, so it seeded a
fresh tail bubble holding a duplicate tool row while the sealed part kept
rendering "Result unavailable" forever (#113035).
A completion now resolves its target by the call's stable id anywhere in
the transcript (toolCallOwnerMessageId) instead of by stream position, so
upsertToolPart attaches the result to the part that already owns the call.
The stream bookkeeping is left untouched: a sealed bubble stays sealed and
a live stream keeps its own id, so later deltas still start the next
bubble. Completions without an owner (reconnect gap, never-seen start)
keep seeding a bubble as before.
A handoff directive whose brief read as markdown (*by week*, a_b c_d, ~/x)
split the paragraph into element children, so the card never claimed it and
the raw ::onboarding{task="…"} line painted as the user's own message.
Directive lines are now backslash-shielded in preprocessMarkdown, next to
the inline-code and math shields, and reach the renderer as one text node.
The guide's reasoning is it reading its own runbook ("Now step 4: offer the
tour with ::ask"); shown under the greeting it breaks the conversation the
guide is trying to have. Reasoning disclosures stay hidden on the guide
thread only.
Same class as the plugin-route bug: KeybindSettings subscribed to
useContributions(KEYBINDS_AREA) but discarded the snapshot and called the
impure allKeybindActions() in render, so the React Compiler memoized the
action list without the subscription as a visible input. A plugin keybind
registered after the tab mounted never appeared, despite the comment
promising "appear/disappear live".
allKeybindActions()/contributedKeybinds() take an optional contribution
snapshot (default: registry read, so the store/imperative callers are
unchanged) and the settings tab passes its subscription through. One test,
red on origin/main.
For people who run profiles as bots, the colored profile strip at the sidebar
foot duplicates the sessions list (community request). Add a persisted
`hermes.desktop.profileRailVisible` preference (on by default) toggled from the
Sessions view menu ("Profile rail"), the shell right-click menu, ⌘K
("Toggle profile rail") and an unbound `view.toggleProfileRail` keybind.
While the rail is hidden the statusbar grows a `ProfileSwitcher` dropdown
beside the gateway switcher ("This device ⌄ · Profiles ⌄") offering the same
choices the rail does: this gateway's profiles, All profiles, every other
gateway's agents in fleet mode, New / Import / Manage. It also answers the
`profile.create` hotkey the rail used to own, so no door is lost.
A welcome-tier 403 classifies as auth_permanent, so the desktop's error
surface mapped it to "Your Nous Portal sign-in expired" with a Nous Portal
re-login button — the chat sentence never reached the user. Terminal results
on the free route now carry a structured free_tier block (kind + the chat
sentence); agent/error_surface.py turns it into a free_tier_<kind> code on
the provider layer with the sentence as `message`. The desktop gives those
codes their own titles, shows the backend sentence as the body, and offers
"Sign in with a Nous account" (the free-tier dialog) instead of the OAuth
re-login, with Retry only where a later send can succeed.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The free tier depends on the account service (NAS) and the welcome inference
host, and Hermes had no honest answer for most of the ways either can refuse
or fail: the NAS codes it matched were never sent, the tier-dark 403 carried
no message to match, a single boot-time blip disabled minting for the whole
process, and a structured rate-limit refusal never reached the cross-session
guard, so the "sign in for a bigger allowance" prompt was dead code.
Backend
- anon_auth: classify what NAS actually sends (404 not_found, 503
temporarily_disabled, 429 + Retry-After, 428 pow_*, 403 account_locked)
into one ANON_* code each, carrying retry_after / retryable on AuthError.
- Replace the process-lifetime mint memo with a per-profile cooldown that
honours the server's wait, climbs a short ladder when the service is
unreachable, never retries terminal codes, and yields to the user's own
retry (force=True).
- Bootstrap record carries error_code / retryable / retry_after; a bounded
background loop retries transient failures and re-announces setup.ready.
setup.status and free_tier.status expose the block; free_tier.provision is
the forced retry.
- Inference: a generic 403 from a welcome host is the tier refusing (keyed on
the route); model_not_free moves onto the gateway's alternate once;
anon_on_paid_host re-reads the route once; a long rate_limited refusal
trips the cross-session guard; a locked account is retired but never
replaced; terminal copy on the free route is one plain sentence.
- Sign-in: Failed keeps the service's code and wait; account_busy is
retryable; the OAuth poll reports retryable / retry_after.
- All user-facing copy rewritten for first-time users: never "the free
service is off" (what is unavailable is using Hermes without signing in,
and signing in is free), no jargon, spoken waits.
Desktop
- A setup-failure notice above the provider picker: one sentence per code,
a retry when the backend says one can work, the sign-in pointer only when
the account service answered at all. The overlay re-checks readiness on
setup.ready so a background success dismisses it.
- Sign-in dialog gains busy / unreachable / unavailable screens.
Rehearsal
- scripts/free_tier_fault_server.py stands in for both services with the
real wire contract and a CORS-open scenario switch; HERMES_EXTRA_WELCOME_HOSTS
(dev-only, env-only) lets the route rules treat it as the welcome host.
Walkthrough in website/docs/developer-guide/free-tier-fault-rehearsal.md.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
sortByProfileOrder moves to a pure lib module with a key selector so
buildRestGroups sorts its named squares directly, replacing the collator
sort that the component then re-sorted with a different comparator. The
fleet rail test no longer needs importOriginal (and three store mocks) to
reach the helper.
The Electron listener now always emits iss (null when the server sent none)
and McpOauthCallbackParams is extra="forbid", so a new Desktop against a
backend without this change would fail every remote MCP OAuth login with a
4000 - including providers that never send iss. Send the key only when set.
Also drop the deliver_callback_flow test the RPC test subsumes.
* refactor(connectors): cut comments that restate the code
Connector modules (tools/connectors, tui_gateway connector RPCs, desktop
connector card/store) keep only comments that carry a non-derivable why or
a cross-module contract. No behaviour change.
* feat(connectors): managed connect runs on the connection operation
Managed `connect` / `reconnect` mint one ConnectionOperation for every target and, on a
desktop session, block the tool turn until the operation settles; the result is per-target
outcomes and never carries a connect link. Off the desktop the result carries the links and
returns at once (PR3 delivers them as their own message).
Why: the previous leg handed the model a URL and a `wait` verb, and the renderer ran its own
2s poller on top of the backend's 5s one; both walked the whole gateway catalog at two vendor
calls per page to read one row (~3 Composio calls/s per pending target). A hidden composer
message started the model's `wait` on the user's behalf. None of it was observable from the
operation the MCP leg already used.
What the operation looks like now:
- `contract.py`: TargetState / Actor / SettleReason enums and the `(kind, from) -> {to: actor}`
transition table. `operation.transition()` enforces it; a card cannot claim a managed
target `connected`, only the backend watcher can.
- `live.py`: one open operation per session, found by `op_id`. `connectors.operation.status`
reads it, `connection.respond` drives it, `pending_connection` on resume replays it.
- `run.py`: the one lifecycle for both target kinds (prepare -> card -> wake/observe loop ->
settle -> result). The managed `observe` hook polls the gateway list once per tick for the
whole operation; the exact-status route replaces that call when the gateway ships it.
- `connection.update` is emitted on every transition and on settlement; registered in the
shared event contract with the operation vocabulary typed on the TS side.
- `wait`, `_rendered_links`, `_seen_instructions`, the just-minted bounce and `_clamp_timeout`
are deleted. `force` on `reconnect` always reinitiates; plain `reconnect` repairs only what
the gateway reports disconnected.
- `connections.wait_timeout_seconds` is removed from config defaults, the example and the
docs. The deadline is `OPERATION_DEADLINE_SECONDS = 300` in `operation.py`; the key was
added on this unmerged train so no migration is needed.
- Wire model: `statusReason` parsed on connection results; the seven-state `connectionStatus`
is typed on list items and an unknown value fails validation; `CONNECTION_REQUIRED` carries
`connect_card_available` instead of the link when the session platform is `desktop`.
Session platform, not callback presence, decides whether a card exists: the GUI bridge
attaches callbacks to every backend session, terminal TUI included.
* feat(desktop): connector card subscribes to the connection operation
The card renders from the backend's operation instead of driving its own: `connector-flow.ts`
(the renderer's 2s `connectors.list` poller, its 120s client deadline and `keepWaiting`) is
deleted, and both hidden composer submits in `connector-tool.tsx` go with it. The model is
never nudged into a `wait`; the tool call is blocked on the backend until the operation
settles.
- `connection-request.ts` is the operation store: keyed by `op_id`, one entry per session,
`applyOperationStatus` / `applyConnectionUpdate` as pure reducers, `respond` leaves the
entry in place (the backend answers with `connection.update`), `ConnectionTargetOutcome`
is a discriminated union the backend's transition table accepts.
- `input-requests.ts` applies `connection.update`; `connection.expire` and the resume
snapshot correlate by `op_id` (a snapshot has no `request_id`).
- `ConnectorOffer` renders one `ConnectorCard` per target from a single
`Record<ConnectionTargetState, phase>` table; Connect opens the stored link, Try again on
failed / expired reissues through `connectors.connect` on the open operation, Not now is a
per-target `skipped`, Continue settles. A settled operation renders `ConnectorSummary` rows
with no live control.
- `tool-render-class.ts`: `manage_connections` renders the card regardless of
`HERMES_GUEST_ONBOARDING`; the flag still gates the onboarding flow, not the card. The
backend gate already decided admission; a card only exists because the tool was admitted.
- `mcp-setup-tool.tsx` speaks the same outcome vocabulary (connected / skipped / failed).
- `ConnectorRow.connectionStatus` is the seven-state literal union, not `string | null`.
- The guided-onboarding poller (`first-build-connectors.ts`) keeps its own row/phase types
and compiles unchanged; PR3 moves it onto the operation.
anti-slop: no net-new findings (17 touched files vs 11d1a12472).
* fix(connectors): the card never parks the tool thread; every update carries the snapshot
Found by the pre-PR adversarial review and a real-path E2E test (both left in the tree).
- The desktop `connection_callback` was still `_block("connection.request", ...)`, which parked
the tool thread on a private request-id Event until a `_respond` that no longer exists for
this event. `connection.respond` settled the operation but the tool waited its full deadline
before the watcher loop even started. The callback now only emits the card; the operation's
own wake loop is the wait. The MCP leg's blocking bridge goes with it: the card answers
through `connection.respond` like every other card.
- `connection.request` and every `connection.update` frame carry the full target snapshot
(state, link, detail). The initial mint happened before the card existed, so the renderer
never saw the links and Connect stayed disabled; a Continue settlement stamped
`not_connected` on the backend while the card still showed `initiated`. The store now
overlays the snapshot; no state is reconstructed from deltas.
- The `connection.update` emitter is a class-level `on_change` slot on the operation, set
once by `register()` (a second `register()` no longer stacks wrappers); session lookup takes
`_sessions_lock`; a re-minted link on an `initiated` target goes through `refresh_link()`
and emits, instead of a bare attribute write.
- `session.interrupt` is checked before the first observe, so an interrupted call settles
`interrupt`, not `all_resolved`.
- A gateway list reporting `expired` for an initiated target is recorded with actor `clock`
(the contract's owner of that edge); it raised `IllegalTransition` before.
- Dead `keepWaiting` i18n keys from the deleted renderer poller removed.
tests/tui_gateway/test_connector_operation_e2e.py runs the desktop lifecycle through the real
tool, registry, gateway RPC handlers and callback bridge with only the HTTP client faked.
* docs(connectors): prompts and docs describe the operation, not the deleted wait verb
The onboarding prompts told the model to call action="wait" with timeout_seconds and to
expect a hidden [setup]/[connectors] note; both are gone. tool-search.md and
toolsets-reference.md said the model gets a connect link on the desktop. tui_gateway/AGENTS.md
gains the connection-operation row of the surface table.
* fix(connectors): the panel re-mints only a dead link
Try again on a failed or expired target mints a fresh link on the open operation. A waiting
target keeps the link it was minted with; the card reopens it and connectors.connect refuses
to spend a second mint (LINK_STILL_VALID). The unused refresh_link() goes. The package
docstring names the new siblings; the nine-name public surface is unchanged.
* test(connectors): the local-batch test answers the operation the way the card does
The callback stopped returning an answer in f782b26d98 (the card answers through
connection.respond); this test still returned one and waited out the 300s deadline in CI.
* ci: retrigger
* fix(connectors): the desktop card appears outside guided onboarding
Live on a signed-in macOS desktop, the two-app connect never showed a card. Three
defects, each hidden by a test that bound state the running app never binds.
The backend read the surface from HERMES_SESSION_PLATFORM only. The desktop and TUI
gateway bind it as HERMES_SESSION_SOURCE (_set_session_context), so session_platform()
was "" and managed connects took the off-desktop branch: links in the model's message,
no operation. session_platform() now reads platform, then source. The E2E test binds
through server._set_session_context instead of set_session_vars(platform="desktop").
The renderer routed manage_connections to the card only under isOnboardingEnabled(),
the HERMES_GUEST_ONBOARDING launch flag, in message-parts.tsx and the run splitter in
fallback.tsx. tool-render-class.ts had already dropped that gate in this PR; the two
routers had not. Both now route on the tool name alone.
ConnectorTool resolved the session owner by the runtime id. Owner routes, hints and
session rows are keyed by the stored id, so in registry topology the owner never
resolved and the card rendered null while the tool blocked. It now resolves by the
stored id, matching the PR1.5 card and every other owner lookup.
message-parts-connectors.test.tsx mounts the real Fallback router with the onboarding
flag off and distinct runtime/stored ids; red before each renderer fix, green after.
* style(connectors): shorter comments, no module mock in the card router test
The router test mocked isOnboardingEnabled to false; jsdom has no preload bridge, so the
real function already returns false. Comments that restated the code are cut to one line.
anti-slop: no net-new findings (25 touched files)
* fix(connectors): Connect on a waiting row opens the stored link
ConnectorCard derived the button's loading state from the phase label, so a managed row that
read "Finish connecting in your browser" (every row, since links are minted up front) had a
disabled Connect button. Nothing on the desktop could open the sign-in link; every managed
connect ended skipped, not_connected, or at the deadline.
The card now takes `busy` for "the action itself is running" and keeps `phase` as a label.
The MCP card passes its in-flight flag; the connector card passes the re-mint wait. Red before:
the Connect button on an initiated row rendered disabled and a click opened nothing.
* fix(connectors): a settled card stays dead; the card binds to its tool call only
A second connect for the same apps revived the finished card on the old tool row. The
connection.request payload carried no id, so the renderer fell back to matching rows by
connector names, and any row with those names qualified, settled or not.
The operation now records the model's tool_call_id and sends it in connection.request and in
the resume snapshot. The card binds to the tool row with that id and to nothing else; the
name-match fallback is deleted. A payload without the id is rejected by the store.
`reason` is removed from the tool: it was the only text the card ever showed from the model
and its absence forked a second tool part, since `reason` doubled as the row-correlation key
in tool-parts.ts. The card never needed it.
`connection.expire` is deleted from the contract and from _EXPIRING_REQUESTS: the card is
raised with _emit, not _block, so nothing has emitted it since the operation lifecycle landed.
Sid's rule of record: a resolved card is fully dead; no path brings it back.
* fix(connectors): the watch loop settles once, on time, and never raises into the result
Three findings from the live review, one loop.
Continue racing a finished sign-in: the loop ran the gateway read, then settled. A read that
returned `connected` for an already-settled or failed target raised IllegalTransition out of
the tool and the model got a generic error instead of the per-app outcomes. The read now skips
targets that are not live (pending, initiated) and skips a settled operation; the loop checks
`settled` after every read.
Settle reason as row text: `settle()` wrote `continue`/`deadline` into each unresolved target's
`detail`, and the card printed it in red. The reason stays on the operation only.
Stop and the deadline waited for the next tick: `/stop` sets a per-thread flag with no wake
hook, so the sleep is sliced at 250 ms and the flag and clock are read each slice. The clock is
also checked before each read, not only after.
Tests: a failed mint that later reads connected settles cleanly; Continue during a read keeps
the settled result; no reason in detail; an interrupt settles within the same second.
* fix(connectors): MCP setup off the desktop returns unavailable instead of blocking
run_mcp_operation treated a non-None connection_callback as "a card exists". Every tui_gateway
session has that callback, the Ink TUI included, so an MCP install from the terminal UI blocked
until the 300 s deadline while the docs promised `unavailable` with the terminal commands.
The MCP path now reads the session surface the same way the managed path does; the callback is
never the predicate. Test binds the surface to `tui` with the callback attached.
* fix(connectors): a failed Try again shows the failure, not the old dead link
The panel's re-mint ignored the gateway's per-app status and moved the row to `initiated` with
whatever link came back, `None` included, so a mint that failed again rendered as waiting on the
link that had already died.
One reader of a mint response now serves both the first mint and Try again
(`managed.mint`, with the actor as a parameter). A repeated failure keeps the row `failed`,
drops the link, and carries the vendor's new text through `operation.refresh`, which emits a
frame without a state change so the card redraws.
* fix(connectors): a forced reconnect waits for the new sign-in before it reports connected
`reconnect` with `force: true` is the account switch. The vendor keeps the old account active
while the new link waits, so the first list read after the mint said `connected` and the
operation settled at once: the new link was dropped and the model was told the switch was done.
A forced target is marked awaiting_new_attempt after the mint. The watcher ignores its row until
the list shows the new attempt (`connectionStatus: initiated`) once, then trusts `connected`.
* fix(connectors): the operation registers under the gateway session key
The tool registered the operation under the agent's session_id; every RPC (connection.respond,
connectors.operation.status, the panel's connectors.connect) and the update emitter looked it up
by the gateway's session key. Those agree until compaction rotates the agent id mid-turn; then
the card's clicks find nothing, no update reaches it, and the tool waits out the deadline.
The registration key is now the bound HERMES_SESSION_KEY, with the agent id as the fallback for
callers with no gateway (unit tests, a bare CLI). The E2E passes a rotated agent id and drives
the card by the gateway key.
* fix(connectors): the forced-reconnect gate reads any non-active row; a failed re-mint of an expired row is failed
Three follow-ups from the verification of the fix pass.
The awaiting_new_attempt gate cleared only on the literal `connectionStatus: initiated`. The
field is optional on the wire and `initializing`, `failed`, `expired` are valid values, so a
forced reconnect could wait the full 300 s and swallow a failed new attempt. The gate now holds
only while the row still reads as the old account (`connected` or `active`) and releases on
anything else.
Try again on an `expired` row whose re-mint fails raised IllegalTransition (no expired → failed
edge). The re-mint steps through `initiated` as the user's attempt, then `failed`, then drops the
dead link.
`detail` never carries a state name any more: `failed` as detail rendered as the row label and
made agent/display.py tag the settled result as a tool error. Only vendor text goes there.
`connection.expire` removed from the renderer's unscoped-stream set; nothing emits it.
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema
One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.
MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.
Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.
`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.
Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.
`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.
The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.
* wip(desktop): connection.request store, resume restore, card routing for MCP targets
Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.
* fix(config): hermes update turns on the connections toolset for saved toolset lists
`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.
Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.
`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.
* refactor: anti-slop pass on the desktop slice; shorten added comments
Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.
* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed
The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.
session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.
lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.
vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.
* chore: drop __pycache__ files swept in by an over-broad git add
* fix(desktop): correlate the connection.request row with the model's tool call by reason
The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.
* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds
* fix(connections): settle reason derives from target state, never from the renderer
A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.
* fix(desktop): a pending connection card re-arms on resume and activate
The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.
* style: literal wording in added comments, docstrings and docs
* fix: shared gateway-event contract and config-schema category for the connection events
connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.
* style: import order (perfectionist) in the desktop and shared files this PR touches
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
After #108525 the turn-settle seal stopped writing `result: {}` onto abandoned
tool calls and left them result-less with `completedAt`. Two clarify consumers
still keyed "pending" on `result === undefined` alone:
- `findPendingClarifyLocation` adopted the sealed call as the "sole open
clarify" fallback for an uncorrelated request, so a new clarify raised after
the user stopped an earlier one re-armed the OLD row (`streamId` = old turn)
and the new question never got a card.
- `ClarifyTool` reads the session's live clarify request, so once the new
request arrived the sealed row painted the NEW question too: two live cards
for one question, the stopped one with no request id behind it.
Sealed calls now only re-arm on a genuine correlation (request id / question
match, which is the resume case; the seal comes off so the row renders live)
and never as the uncorrelated fallback; in the timeline a settled-without-result
clarify renders as the generic history row, as every other tool already does.
Live Desktop A/B (real Electron, loopback provider, stop Q1 then ask for Q2):
before 2 live clarify cards both showing Q2; after 1 live card (Q2), Q1 shown
as "Result unavailable", answer reaches the provider.
apps/shared/src/gateway-events.ts is now a thin layer over
gateway-contract.generated.ts (client-local synthetic events + the
GatewayEvent envelope); gateway-events.json, its two rendezvous tests and
the duplicated BillingBlock / SessionInfo / ProjectInfo hand copies are
gone. Desktop, TUI, web and shared typecheck against the generated
RpcMethods / ServerRequestMap / BackendGatewayEventMap.
What tsc found once the types were honest: three phantom fields the
backend never sent (tool.start.todos, error.reason,
voice.transcript.voice_stopped) - the TUI todo tests were driving the
list through the phantom and are retargeted to tool.complete, where the
wire actually carries it; nullable fields (`None` on the wire) were typed
as plain optionals in eight places and now coerce at the boundary;
SessionResumeResult had a stale generic.
Contract fixes from the consumer pass: TranscriptMessage is the gateway
projection (text/row_id/context/args), not the stored row; SkinPayload
matches HermesSkin (empty-string defaults, never null); SessionLiveInfo
model/tools/skills are required (always emitted); BillingBlock.billing_url
is required-nullable (dataclass asdict).
tui_gateway/AGENTS.md documents the declare -> regenerate -> tsc loop.
The old suites asserted the deleted wire (`*.request` events, `*.respond`
RPCs, `pending_clarify` snapshots, `_pending`/`_answers` teardowns). Each
test keeps its invariant against the new shape: a seeded live request's
`respond` spy receives the answer object, `hasOpenServerRequest` flips, the
`approval.respond` RPC fallback is asserted ONLY for queue entries restored
without a socket, and Bot Mode rooms answer via `request.answer` /
`clarify.lock`. The group-turns test that polled forever for a
`clarify.respond` that no longer exists (20-minute hang) now completes.
Three consumers still read `result === undefined` as an open call or an error without text after the result and display-hint split.
The tool row lost the event's error explanation when the result carried none: `toolErrorText` now reads `toolResultMetadata.error` / `.message` when `isError` is set, and `toolStatus` runs the error path for sealed error rows so an envelope-only read miss stays on the notice tier. The turn-activity signature and wait narration count a sealed call as settled. Onboarding's start-with-connections offer withdraws once a connection wait is sealed.
Four tests fail on 235ec0f and pass here.