The 'execProbe keeps the parent event loop available' case asserts nothing
about timing; the 1s spawn budget only exists so a wedged child cannot stall
the run. A cold Windows CI runner can take longer than 1s just to start the
node child, which failed the test for reasons unrelated to what it guards.
5s keeps the safety bound without the flake.
runPoolBackendStart reused assertPoolEntryStillOwned at every cancellation
checkpoint, and that helper releases the local backend slot before it throws.
That is right before spawn (no child, nothing to wait for) but wrong at the
post-spawn checkpoints (after the port announcement, waitForHermes, token
adoption, WS probe): the lease was handed back while the superseded child was
still running, so a successor could spawn into an occupied slot, and the
caller's teardownFailedLocalBackend -> releaseLocalBackendSlotAfterExit then
found nothing left to release and became a no-op. This breaks the
pool-spawn-coordinator invariant that a lease is held until the child exits
or the start fails.
Add a `releaseSlot` option to assertPoolEntryStillOwned and pass
`{ releaseSlot: false }` at every site where `entry.process` is set. The
pre-spawn sites keep releasing. The post-spawn release now happens only via
teardownFailedLocalBackend (after the child has provably exited) or the
child's own exit handler.
No vitest case: the helper and its callers live in main.ts alongside the
module-level backendPool / localBackendLifecycle state and are not importable
from a unit test without extracting them; the lease-after-exit ordering
itself is already pinned by pool-spawn-coordinator.test.ts ('a rejected wait
keeps the slot occupied').
The serve-support resolver caches the probe outcome per resolved runtime
for the process lifetime. That is right for a genuine "unknown
subcommand" exit, but a probe that died by timeout says nothing about the
runtime — only that this machine was slow right then (cold AV scan on
Windows, first Python import after boot). Caching that as `false` routed
a modern runtime through the legacy `dashboard` form until the app was
relaunched.
Export isTimeoutError from backend-probes and evict the cache entry on a
timeout so the next check re-probes; non-timeout failures stay cached.
One vitest case covers timeout → re-probe; the existing case still pins
non-timeout failure → cached.
Also replace the last synchronous read on the discovery path
(dashboard.py fast path) with fs.promises.readFile, matching the stack's
goal of keeping runtime discovery off the main event loop.
isActiveRuntimeUsable relied on async-return flattening for the trailing
canImportHermesCli promise: correct today, but appending any further `&&`
operand would have made the expression truthy regardless of the probe.
Await the probe explicitly so the intent survives future edits, and mark
unwrapWindowsVenvHermesCommand async like its siblings since it always
returns a promise.
connectRemote used two inline isCurrentAttempt/throw pairs with the same
message that backendConnectionState.assertCurrentAttempt already emits,
and the third guard in the same function already uses the helper. Use it
in all three places so the superseded-attempt message has one source.
The fast-path/probe/cache explanation for serve support lived above the
resolver call site in main.ts while the logic lives in
backend-serve-support.ts; move it next to the code and leave a pointer.
The redirect out of the archive used a literal '/app.asar/' replace, so the
staged get-windows specifier built with path.join on a Windows packaged build
(...\resources\app.asar\dist\...) was never rewritten and the helper spawn
would still land inside the archive. Mirror the app.asar(?=$|[\\/]) regex
main.ts already uses for the same job, and cover the backslash path in the
test. The dev-tree no-op case folds into the segment-boundary test so the
describe block stays at three cases.
Closes#88468 (with the preceding commit from PR #89633).
`read_window_below` answers "could not enumerate windows on this system" on
every packaged macOS build. get-windows does its real work by exec'ing a Swift
helper whose path it derives from its own module URL, and execFile cannot run a
file that lives inside app.asar: the kernel sees the archive as a file, so the
spawn fails ENOTDIR. Electron does rewrite asar paths for child_process, and
only once that module has been pulled through the CJS loader, which this ESM
main process never does (every import here is `from 'node:child_process'`).
Both specifiers now resolve outside the archive. The staged copy added in
1c75e05982 is redirected, and so is the node_modules fallback, which needs
`import.meta.resolve` first to have a path to redirect. electron-builder
unpacks the whole get-windows directory, so the redirected path exists on the
real filesystem and the derived helper path is executable.
`/app.asar/` only ever appears as a complete path segment in a packaged build,
so outside one the replace is a no-op. Tests cover the packaged redirect, the
dev-tree passthrough, and an `app.asar-tools` lookalike that must be left
alone.
Verified on a packaged build, macOS 26.6.2 arm64, ad-hoc signed: the same
machine that returns the enumeration failure on unpatched main answers with the
window underneath and its title.
(cherry picked from commit dce34f626a0796fcd945aa3885384a1ec86c819f)
Dropping a link out of a browser onto the Desktop composer toasted
"Drop files — Could not attach <title>.url" and attached nothing. A browser
link drag carries `text/uri-list` plus, on Windows, a virtual `<title>.url`
shortcut File (`.webloc` on macOS) that has no on-disk path. The drop
pipeline only understood Files and in-app paths: the path-less stub went to
the upload branch, `attachContextFilePath('')` returned false, and the user
had to copy/paste the URL instead.
`extractDroppedFiles` now reads `text/uri-list`, drops the path-less
shortcut stub when a link is present, and emits `{ url }` entries;
`droppedFileInlineRef` turns them into the same `@url:` chip the "+ → Add
URL" dialog and paste-linkify produce, so every drop surface (composer
form, text box, conversation area, edit composer) gets it for free.
`dragHasAttachments` accepts `text/uri-list` so the form-level enter/over
handlers claim the drag at all. A path-less *image* dragged off a web page
keeps its bytes and wins over the link to its own src.
Attribute drive-level errors to the thread being drained, not the send
that created the queue. Reinsert repeat member failures in recency order
so the collapsed activity row cannot show an older sibling failure.
Cover both invariants and the repeated-refusal sequence in native Desktop.
Thanks to @kvnloo for identifying both review findings.
Keep failed-member exclusion across the room queue, rather than resetting
it per pending thread. A new user action after failure still permits a new
attempt. Share the drain activity epoch so a skipped queued thread cannot
hide the preceding member failure.
Proven red in real Electron: hold transport refusal, enqueue same-thread
and cross-thread sends, then release; old head submits three times, fixed
head once. Strengthen follow-up evidence with distinct provider replies,
exact public log order/count, and per-input inference counts.
Serialize room drives through their actual member completion, freeze input
watermarks by retained entry identity, and share the same completion path
with handoff continuations. Stop discards queued work without releasing an
active owner early; rename follows the existing room binding.
Observe stranded replies for the hard-cap duration plus grace after the
foreground wait, and retain unresolved failures in collapsed Activity.
Never automatically retry an ambiguous failed submit within the same drive.
Adapted from the queue and boundary approach in #92041 by @enwaiax and
harvest-budget approach in #107193 by @Finn763; #106502 by @wadib identified
failed-submit watermark consumption. The implementation retains current
numeric watermark storage, room lifecycle bindings and serial round limits.
Related: #92003, #105247, #100026
Keep the salvaged botHandle normalization, but remove the unconditional
hermes alias on remote default profiles: the parser's last-wins map
otherwise retargets a local @hermes handoff by roster order.
Consolidate regression coverage into two invariants for persisted primary
handles, both handoff directions and three-source qualified identity.
Capture real Electron screenshots, durable logs and source receipts;
exercise the reverse live handoff too. Document the repair and bump the
bundled Desktop patch version.
Real-Electron Playwright spec for #100406: a two-member room (primary
profile + code-farmer). The user addresses only @code-farmer; its
scripted reply @mentions hermes; the assertion is a `default`-authored
"B" entry in the persisted room log. On origin/main the room settles
after Code Farmer's line and the spec fails at that assertion; with the
mention-alias fix it passes.
The mock inference server gains a per-speaker script for group rooms:
`E2E_SAY(<handle>)[<line>]` tokens in the user's send answer the member
whose turn prompt opens with `You are @<handle>`; unscripted members
reply "(pass)". `{at}` stands for `@` so the script itself never
mentions anyone and round one only drives the member the user tagged.
Group mention parse used member.handle before botHandle, so a persisted
or union-stamped handle of "default" never mapped to the user-facing
@hermes alias. Bot-to-bot handoff toward the primary profile then
settled with no continuation, while the reverse direction still worked.
Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
Settings → Appearance gains a Chat Font row next to Terminal Font. The value
lands in config.yaml as desktop.font_family and is layered in front of the
active theme's fontSans when the theme paints --dt-font-sans, so an empty value
is exactly the theme and a chosen family keeps the theme's CJK/emoji fallbacks.
Why: #72485 asked for OpenDyslexic in the app; #76395 only made the terminal
pane configurable, and chat/chrome typography had no user-facing knob at all
(theme presets set fontSans, imported VS Code themes carry no font opinion).
Refs #72485, #37566
The card offers; the transcript explains. One shell, one heading, one row per app:
a 16 px state mark (hollow circle, spinning arc, check), the brand chip, the name, one
cue on the row whose link is open ("Waiting for your browser…"), and one verb in a
fixed lane (Connect / Try again / Install). A resolved row shows no control.
Continue sits below the shell and is the one exit; the per-row Not now is gone, so
every open row settles as "Not connected".
No reason or error text lives in the card. A sign-in failure is a background event:
the row's verb becomes Try again and the model's reply after Continue explains. A
click that changed nothing (a refused re-mint) is the one toast: "Could not start
authorization for <app>." A successful Try again opens the fresh link at once.
The MCP setup card is the same component with Install / Enable / Authorize as the
verb: every target gets a row (a two-server call no longer strands the second one),
Continue below, no catalog source line. The settled card keeps its scaffold rows and
says three words: Connected, Skipped, Not connected.
Design of record: Paper 01M29RZY5NDXT7CN7MBWQT0HRW, page B-0, artboard ZV6-0
("Variation 2b"). Presentation only; no wire change.
Tests are behaviour contracts, run red first: one verb per row and no reason; the
mark carries the state; Try again mints on the open operation and opens the link; a
refused Try again toasts and keeps the verb; Continue settles; settled rows carry no
detail text; the MCP card lists every target live and settled.
* refactor(connectors): cut comments that restate the code
Connector modules (tools/connectors, tui_gateway connector RPCs, desktop
connector card/store) keep only comments that carry a non-derivable why or
a cross-module contract. No behaviour change.
* feat(connectors): managed connect runs on the connection operation
Managed `connect` / `reconnect` mint one ConnectionOperation for every target and, on a
desktop session, block the tool turn until the operation settles; the result is per-target
outcomes and never carries a connect link. Off the desktop the result carries the links and
returns at once (PR3 delivers them as their own message).
Why: the previous leg handed the model a URL and a `wait` verb, and the renderer ran its own
2s poller on top of the backend's 5s one; both walked the whole gateway catalog at two vendor
calls per page to read one row (~3 Composio calls/s per pending target). A hidden composer
message started the model's `wait` on the user's behalf. None of it was observable from the
operation the MCP leg already used.
What the operation looks like now:
- `contract.py`: TargetState / Actor / SettleReason enums and the `(kind, from) -> {to: actor}`
transition table. `operation.transition()` enforces it; a card cannot claim a managed
target `connected`, only the backend watcher can.
- `live.py`: one open operation per session, found by `op_id`. `connectors.operation.status`
reads it, `connection.respond` drives it, `pending_connection` on resume replays it.
- `run.py`: the one lifecycle for both target kinds (prepare -> card -> wake/observe loop ->
settle -> result). The managed `observe` hook polls the gateway list once per tick for the
whole operation; the exact-status route replaces that call when the gateway ships it.
- `connection.update` is emitted on every transition and on settlement; registered in the
shared event contract with the operation vocabulary typed on the TS side.
- `wait`, `_rendered_links`, `_seen_instructions`, the just-minted bounce and `_clamp_timeout`
are deleted. `force` on `reconnect` always reinitiates; plain `reconnect` repairs only what
the gateway reports disconnected.
- `connections.wait_timeout_seconds` is removed from config defaults, the example and the
docs. The deadline is `OPERATION_DEADLINE_SECONDS = 300` in `operation.py`; the key was
added on this unmerged train so no migration is needed.
- Wire model: `statusReason` parsed on connection results; the seven-state `connectionStatus`
is typed on list items and an unknown value fails validation; `CONNECTION_REQUIRED` carries
`connect_card_available` instead of the link when the session platform is `desktop`.
Session platform, not callback presence, decides whether a card exists: the GUI bridge
attaches callbacks to every backend session, terminal TUI included.
* feat(desktop): connector card subscribes to the connection operation
The card renders from the backend's operation instead of driving its own: `connector-flow.ts`
(the renderer's 2s `connectors.list` poller, its 120s client deadline and `keepWaiting`) is
deleted, and both hidden composer submits in `connector-tool.tsx` go with it. The model is
never nudged into a `wait`; the tool call is blocked on the backend until the operation
settles.
- `connection-request.ts` is the operation store: keyed by `op_id`, one entry per session,
`applyOperationStatus` / `applyConnectionUpdate` as pure reducers, `respond` leaves the
entry in place (the backend answers with `connection.update`), `ConnectionTargetOutcome`
is a discriminated union the backend's transition table accepts.
- `input-requests.ts` applies `connection.update`; `connection.expire` and the resume
snapshot correlate by `op_id` (a snapshot has no `request_id`).
- `ConnectorOffer` renders one `ConnectorCard` per target from a single
`Record<ConnectionTargetState, phase>` table; Connect opens the stored link, Try again on
failed / expired reissues through `connectors.connect` on the open operation, Not now is a
per-target `skipped`, Continue settles. A settled operation renders `ConnectorSummary` rows
with no live control.
- `tool-render-class.ts`: `manage_connections` renders the card regardless of
`HERMES_GUEST_ONBOARDING`; the flag still gates the onboarding flow, not the card. The
backend gate already decided admission; a card only exists because the tool was admitted.
- `mcp-setup-tool.tsx` speaks the same outcome vocabulary (connected / skipped / failed).
- `ConnectorRow.connectionStatus` is the seven-state literal union, not `string | null`.
- The guided-onboarding poller (`first-build-connectors.ts`) keeps its own row/phase types
and compiles unchanged; PR3 moves it onto the operation.
anti-slop: no net-new findings (17 touched files vs 11d1a12472).
* fix(connectors): the card never parks the tool thread; every update carries the snapshot
Found by the pre-PR adversarial review and a real-path E2E test (both left in the tree).
- The desktop `connection_callback` was still `_block("connection.request", ...)`, which parked
the tool thread on a private request-id Event until a `_respond` that no longer exists for
this event. `connection.respond` settled the operation but the tool waited its full deadline
before the watcher loop even started. The callback now only emits the card; the operation's
own wake loop is the wait. The MCP leg's blocking bridge goes with it: the card answers
through `connection.respond` like every other card.
- `connection.request` and every `connection.update` frame carry the full target snapshot
(state, link, detail). The initial mint happened before the card existed, so the renderer
never saw the links and Connect stayed disabled; a Continue settlement stamped
`not_connected` on the backend while the card still showed `initiated`. The store now
overlays the snapshot; no state is reconstructed from deltas.
- The `connection.update` emitter is a class-level `on_change` slot on the operation, set
once by `register()` (a second `register()` no longer stacks wrappers); session lookup takes
`_sessions_lock`; a re-minted link on an `initiated` target goes through `refresh_link()`
and emits, instead of a bare attribute write.
- `session.interrupt` is checked before the first observe, so an interrupted call settles
`interrupt`, not `all_resolved`.
- A gateway list reporting `expired` for an initiated target is recorded with actor `clock`
(the contract's owner of that edge); it raised `IllegalTransition` before.
- Dead `keepWaiting` i18n keys from the deleted renderer poller removed.
tests/tui_gateway/test_connector_operation_e2e.py runs the desktop lifecycle through the real
tool, registry, gateway RPC handlers and callback bridge with only the HTTP client faked.
* docs(connectors): prompts and docs describe the operation, not the deleted wait verb
The onboarding prompts told the model to call action="wait" with timeout_seconds and to
expect a hidden [setup]/[connectors] note; both are gone. tool-search.md and
toolsets-reference.md said the model gets a connect link on the desktop. tui_gateway/AGENTS.md
gains the connection-operation row of the surface table.
* fix(connectors): the panel re-mints only a dead link
Try again on a failed or expired target mints a fresh link on the open operation. A waiting
target keeps the link it was minted with; the card reopens it and connectors.connect refuses
to spend a second mint (LINK_STILL_VALID). The unused refresh_link() goes. The package
docstring names the new siblings; the nine-name public surface is unchanged.
* test(connectors): the local-batch test answers the operation the way the card does
The callback stopped returning an answer in f782b26d98 (the card answers through
connection.respond); this test still returned one and waited out the 300s deadline in CI.
* ci: retrigger
* fix(connectors): the desktop card appears outside guided onboarding
Live on a signed-in macOS desktop, the two-app connect never showed a card. Three
defects, each hidden by a test that bound state the running app never binds.
The backend read the surface from HERMES_SESSION_PLATFORM only. The desktop and TUI
gateway bind it as HERMES_SESSION_SOURCE (_set_session_context), so session_platform()
was "" and managed connects took the off-desktop branch: links in the model's message,
no operation. session_platform() now reads platform, then source. The E2E test binds
through server._set_session_context instead of set_session_vars(platform="desktop").
The renderer routed manage_connections to the card only under isOnboardingEnabled(),
the HERMES_GUEST_ONBOARDING launch flag, in message-parts.tsx and the run splitter in
fallback.tsx. tool-render-class.ts had already dropped that gate in this PR; the two
routers had not. Both now route on the tool name alone.
ConnectorTool resolved the session owner by the runtime id. Owner routes, hints and
session rows are keyed by the stored id, so in registry topology the owner never
resolved and the card rendered null while the tool blocked. It now resolves by the
stored id, matching the PR1.5 card and every other owner lookup.
message-parts-connectors.test.tsx mounts the real Fallback router with the onboarding
flag off and distinct runtime/stored ids; red before each renderer fix, green after.
* style(connectors): shorter comments, no module mock in the card router test
The router test mocked isOnboardingEnabled to false; jsdom has no preload bridge, so the
real function already returns false. Comments that restated the code are cut to one line.
anti-slop: no net-new findings (25 touched files)
* fix(connectors): Connect on a waiting row opens the stored link
ConnectorCard derived the button's loading state from the phase label, so a managed row that
read "Finish connecting in your browser" (every row, since links are minted up front) had a
disabled Connect button. Nothing on the desktop could open the sign-in link; every managed
connect ended skipped, not_connected, or at the deadline.
The card now takes `busy` for "the action itself is running" and keeps `phase` as a label.
The MCP card passes its in-flight flag; the connector card passes the re-mint wait. Red before:
the Connect button on an initiated row rendered disabled and a click opened nothing.
* fix(connectors): a settled card stays dead; the card binds to its tool call only
A second connect for the same apps revived the finished card on the old tool row. The
connection.request payload carried no id, so the renderer fell back to matching rows by
connector names, and any row with those names qualified, settled or not.
The operation now records the model's tool_call_id and sends it in connection.request and in
the resume snapshot. The card binds to the tool row with that id and to nothing else; the
name-match fallback is deleted. A payload without the id is rejected by the store.
`reason` is removed from the tool: it was the only text the card ever showed from the model
and its absence forked a second tool part, since `reason` doubled as the row-correlation key
in tool-parts.ts. The card never needed it.
`connection.expire` is deleted from the contract and from _EXPIRING_REQUESTS: the card is
raised with _emit, not _block, so nothing has emitted it since the operation lifecycle landed.
Sid's rule of record: a resolved card is fully dead; no path brings it back.
* fix(connectors): the watch loop settles once, on time, and never raises into the result
Three findings from the live review, one loop.
Continue racing a finished sign-in: the loop ran the gateway read, then settled. A read that
returned `connected` for an already-settled or failed target raised IllegalTransition out of
the tool and the model got a generic error instead of the per-app outcomes. The read now skips
targets that are not live (pending, initiated) and skips a settled operation; the loop checks
`settled` after every read.
Settle reason as row text: `settle()` wrote `continue`/`deadline` into each unresolved target's
`detail`, and the card printed it in red. The reason stays on the operation only.
Stop and the deadline waited for the next tick: `/stop` sets a per-thread flag with no wake
hook, so the sleep is sliced at 250 ms and the flag and clock are read each slice. The clock is
also checked before each read, not only after.
Tests: a failed mint that later reads connected settles cleanly; Continue during a read keeps
the settled result; no reason in detail; an interrupt settles within the same second.
* fix(connectors): MCP setup off the desktop returns unavailable instead of blocking
run_mcp_operation treated a non-None connection_callback as "a card exists". Every tui_gateway
session has that callback, the Ink TUI included, so an MCP install from the terminal UI blocked
until the 300 s deadline while the docs promised `unavailable` with the terminal commands.
The MCP path now reads the session surface the same way the managed path does; the callback is
never the predicate. Test binds the surface to `tui` with the callback attached.
* fix(connectors): a failed Try again shows the failure, not the old dead link
The panel's re-mint ignored the gateway's per-app status and moved the row to `initiated` with
whatever link came back, `None` included, so a mint that failed again rendered as waiting on the
link that had already died.
One reader of a mint response now serves both the first mint and Try again
(`managed.mint`, with the actor as a parameter). A repeated failure keeps the row `failed`,
drops the link, and carries the vendor's new text through `operation.refresh`, which emits a
frame without a state change so the card redraws.
* fix(connectors): a forced reconnect waits for the new sign-in before it reports connected
`reconnect` with `force: true` is the account switch. The vendor keeps the old account active
while the new link waits, so the first list read after the mint said `connected` and the
operation settled at once: the new link was dropped and the model was told the switch was done.
A forced target is marked awaiting_new_attempt after the mint. The watcher ignores its row until
the list shows the new attempt (`connectionStatus: initiated`) once, then trusts `connected`.
* fix(connectors): the operation registers under the gateway session key
The tool registered the operation under the agent's session_id; every RPC (connection.respond,
connectors.operation.status, the panel's connectors.connect) and the update emitter looked it up
by the gateway's session key. Those agree until compaction rotates the agent id mid-turn; then
the card's clicks find nothing, no update reaches it, and the tool waits out the deadline.
The registration key is now the bound HERMES_SESSION_KEY, with the agent id as the fallback for
callers with no gateway (unit tests, a bare CLI). The E2E passes a rotated agent id and drives
the card by the gateway key.
* fix(connectors): the forced-reconnect gate reads any non-active row; a failed re-mint of an expired row is failed
Three follow-ups from the verification of the fix pass.
The awaiting_new_attempt gate cleared only on the literal `connectionStatus: initiated`. The
field is optional on the wire and `initializing`, `failed`, `expired` are valid values, so a
forced reconnect could wait the full 300 s and swallow a failed new attempt. The gate now holds
only while the row still reads as the old account (`connected` or `active`) and releases on
anything else.
Try again on an `expired` row whose re-mint fails raised IllegalTransition (no expired → failed
edge). The re-mint steps through `initiated` as the user's attempt, then `failed`, then drops the
dead link.
`detail` never carries a state name any more: `failed` as detail rendered as the row label and
made agent/display.py tag the settled result as a tool error. Only vendor text goes there.
`connection.expire` removed from the renderer's unscoped-stream set; nothing emits it.
* feat(connections): manage_connections covers local MCP servers; setup_mcp leaves the schema
One model tool now connects the user to apps of both kinds. A target
`{"name": "linear", "mcp": true}` is a locally configured MCP server;
`install` / `enable` / `authorize` are its verbs. Bare strings and
`{"name": ...}` stay managed connectors and that leg is unchanged.
MCP targets run through one backend-owned connection operation
(tools/connections_tool_operation.py): created with a server-side
deadline from the new config key `connections.wait_timeout_seconds`
(default 120, floor 5, no ceiling), per-target state, and exactly-once
settlement (all resolved / Continue / deadline / interrupt). Unresolved
targets freeze as `not_connected` with the settle reason.
Why the fold works now: the approval card is reached through
`agent.connection_callback` via the agent-level inline executor table,
which is the only path that carries a GUI callback. Registry dispatch
(every non-GUI surface) settles MCP targets as `unavailable` with the
`hermes mcp install / login` hint; managed targets in the same call
are unaffected.
`setup_mcp` is removed from every advertised toolset and from the
deferral list; an inline-table shim keeps calls from conversations
opened before this change dispatching (prompt-cache protection).
`_LEGACY_TOOL_ALIASES` is not the mechanism: inline tools bypass it.
Gateway: `mcp.setup.request/respond` are replaced by
`connection.request/respond/expire` (no wire compat; desktop ships
with this). The bridge waits exactly the operation's deadline. The
`session.resume` snapshot gains `pending_connection` so a reopened
window restores the card with the original deadline.
`manage_connections` joins `_SEQUENTIAL_DEADLINE_EXEMPT_TOOLS`: the
operation owns its wait; the 420s guard must not report `tool_timeout`
while the card is live.
The portal `check_fn` on the tool is dropped in favour of a
handler-level gate on the managed leg, so signed-out sessions can still
approve local MCPs.
* wip(desktop): connection.request store, resume restore, card routing for MCP targets
Renderer half of the setup_mcp fold, first slice: connection-request store
(mirrors clarify), connection.request/expire handling, pending_connection
resume restore, mcpTargets() + isCardTool(name, args) so MCP-target
manage_connections calls classify as cards. Not yet: the card component
rewrite (mcp-setup-tool.tsx), mcp-directory.ts removal, vitest, docs.
Does not typecheck until the card rewrite lands.
* fix(config): hermes update turns on the connections toolset for saved toolset lists
`hermes tools` writes an explicit `platform_toolsets.<platform>` list, and the
resolver reads absence from that list as "unchecked". The `connections`
toolset (#106842) shipped after most users last saved, so `manage_connections`
is stripped from the schema on every install that ever opened the picker.
The Nous entitlement gate never runs; the agent reports the tool as missing.
Migration 44 -> 45 (renumbered when folded into #109517; main was already at 44) appends `connections` to each explicit per-platform list
that lacks it and records the offer in `known_builtin_toolsets` where that
record exists, so a later uncheck reads as a decline. It skips: platforms
whose record already holds `connections` (the user saw the checkbox and left
it off), bare composite lists ([hermes-cli]) that already inherit it, platforms
where the toolset is not allowed, and any config whose `agent.disabled_toolsets`
names `connections` (Blank Slate, `hermes tools --disable`), because the
resolver subtracts that list last and the enable would never take effect.
The explicit-list test is the resolver's own: any configurable or plugin key.
`hermes update` runs migrations post-pull for the active profile and every
sibling, so one update is enough. Fresh installs and composite users were
never affected.
* refactor: anti-slop pass on the desktop slice; shorten added comments
Parse connection.request at the boundary with a typed wire interface instead of
unknown + typeof; mcpTargets reuses connectorText; comments cut to one or two
lines. slop-ratchet: no net-new findings in 13 touched files.
* feat(desktop): the MCP approval card answers manage_connections; MCP Directory removed
The existing card (mcp-setup-tool.tsx) now reads the connection-request store,
renders for manage_connections calls with mcp:true targets, answers through
connection.respond with a per-target outcome, and no longer calls reload.mcp
after Install; the new server's tools arrive on the between-turns refresh.
A settled operation renders the first target's frozen state.
session.resume restores a pending card with its original deadline on both the
activate and cold-resume paths.
lib/mcp-directory.ts is deleted along with its two fallback branches
(suggestion provider, card install). The catalog was already primary in both;
a catalog miss now yields no suggestion / a notInCatalog error. The GitHub
never-suggest test is rewritten on catalog-shaped data.
vitest: connection-request store (6), suggestion provider, clarify restore.
slop-ratchet: no net-new findings in 19 touched files.
* chore: drop __pycache__ files swept in by an over-broad git add
* fix(desktop): correlate the connection.request row with the model's tool call by reason
The synthetic row from connection.request and the tool.start row carried
different ids and no shared match value (op_id is not in the model's args),
so the card mounted twice. reason is the arg both sides carry.
* docs: manage_connections covers local MCP servers; connections.wait_timeout_seconds
* fix(connections): settle reason derives from target state, never from the renderer
A card that answers one of two targets and claims all_resolved must settle as
continue with the other target not_connected; found live with a two-target call.
* fix(desktop): a pending connection card re-arms on resume and activate
The store entry was restored but the transcript row was not, so navigating
away and back (or reloading) lost the card while the backend kept waiting.
restorePendingClarifyToolCall's core is generalized to any blocking tool
name and both resume paths project the connection row through it.
Verified live: card restored after navigate-away and after a full renderer
reload, deadline_at unchanged, approve settles connected.
* style: literal wording in added comments, docstrings and docs
* fix: shared gateway-event contract and config-schema category for the connection events
connection.request/expire replace mcp.setup.* in apps/shared gateway-events
(json list, BACKEND_EVENT_NAMES, GatewayEventMap) so the renderer's event
union includes them and the tui_gateway contract test passes. The new
`connections` config section folds into the agent tab like the other
single-field sections.
* style: import order (perfectionist) in the desktop and shared files this PR touches
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
HERMES_SKIP_INTRO=1 turned off the intro film but also silently killed the
entire guided first launch: the guide can only queue on the film's
completion edge (queueGuideAfterIntro requires hasSeenIntroReveal), so with
the film disabled nothing ever fired and the app booted straight into the
normal shell with the free-tier flag on.
The film-to-guide seam now fires without the film: beginOnboardingFlowWithout
Intro records the film as watched and queues the guide directly, so
HERMES_SKIP_INTRO skips exactly the film. A relaunch adopts the persisted
guide (loadGate sees cinematic + seen), and a later launch without the flag
cannot replay the film over the completed flow.
Verified: tsc -p tsconfig.json and tsconfig.electron.json clean; vitest ui
(onboarding-gate, onboarding, onboarding-never-forces-sign-in: 30 passed
incl. 2 new invariant tests) and electron (guest-onboarding-flag,
preload-flags: 6 passed) green.
After #108525 the turn-settle seal stopped writing `result: {}` onto abandoned
tool calls and left them result-less with `completedAt`. Two clarify consumers
still keyed "pending" on `result === undefined` alone:
- `findPendingClarifyLocation` adopted the sealed call as the "sole open
clarify" fallback for an uncorrelated request, so a new clarify raised after
the user stopped an earlier one re-armed the OLD row (`streamId` = old turn)
and the new question never got a card.
- `ClarifyTool` reads the session's live clarify request, so once the new
request arrived the sealed row painted the NEW question too: two live cards
for one question, the stopped one with no request id behind it.
Sealed calls now only re-arm on a genuine correlation (request id / question
match, which is the resume case; the seal comes off so the row renders live)
and never as the uncorrelated fallback; in the timeline a settled-without-result
clarify renders as the generic history row, as every other tool already does.
Live Desktop A/B (real Electron, loopback provider, stop Q1 then ask for Q2):
before 2 live clarify cards both showing Q2; after 1 live card (Q2), Q1 shown
as "Result unavailable", answer reaches the provider.
Every caller tears the SSH connection down right after `await
cancelAndWait(scope)`, so resolving on this call's own barrier alone let a
connection-apply teardown overlap the pool stop's still-running afterStop
teardown for the same key. Wait for the composed barrier instead; chain the
barriers (drain promises never reject, so this equals allSettled). The test
now waits a macrotask and asserts that neither the apply nor the new
bootstrap runs before the first teardown completes.
A pool stop blocked in teardownSshConnection() holds drains[scope]; a
concurrent connection apply calling cancelAndWait() for the same scope
replaced that barrier with its own and, having nothing to drain, cleared the
map entry as soon as it finished. start() then saw no drain and began a new
bootstrap while the first SSH teardown was still running.
cancelAndWait() now composes with the drain already in flight and the entry
is cleared only when the composed barrier settles, so start() waits for
every active teardown. Regression test reproduces the pool-stop / apply
race; it fails on the previous coordinator.
Reported by ehz0ah on #110025 (follow-up to #106935).
A dead tunnel was redialled every 2 s until the scope was torn down;
delays now double from the base to a 30 s cap and reset on open. The
generation counter duplicated what the entry map and entry.socket
already say (stop() removes the entry, connect() replaces the socket),
so abandonIfStale reads those instead. afterStop's unused entry
parameter is dropped.
afterStop awaited cancelAndWait(teardownSshConnection) with no catch; the
idle reaper calls stopPoolBackend un-awaited, so a teardown rejection
became an unhandled rejection on the main process. Log it through the SSH
log and let the fence release.
Four registry tests overlapped: sibling independence is a two-line
assertion inside hold-until-stop, and the reject-missing-target case is
the setup half of the empty-scope primary case. Folding them keeps every
invariant covered (hold, no-reconnect-after-stop, sibling isolation,
''-is-a-real-key, baseUrl/token required) with fewer fixtures to keep in
sync. The pool-stop teardown fence is already covered by
pool-stop.test.ts::afterStop and ssh-bootstrap-coordinator.test.ts.
Electron bundles a global WebSocket in the main process; a missing
constructor is a build/environment error, not a runtime state to route
around. The two silent `typeof WebSocketImpl !== 'function'` returns in
start()/connect() would turn that build error into an armed-looking
registry that never opens a socket, letting web_server_idle_exit retire
an owned SSH-isolated sibling with no log line — exactly the bug this
module exists to prevent. Let `new WebSocketImpl(url)` throw into the
existing catch, which logs and schedules a reconnect.
Review banned source-reading main.ts greps and asked the remaining registry
cases down to ~4: hold-until-stop with no reconnect, empty-string primary,
reject missing target, and sibling isolation. Prettier the helper.
Co-authored-by: Cursor <cursoragent@cursor.com>
Process-less SSH entries finish child exit immediately. Keep inFlight
and the bootstrap drain up until keepalive teardown completes, and
stop reading main.ts from tests.
Co-authored-by: Cursor <cursoragent@cursor.com>
Desktop keeps live sockets on the profile-less backend while the --profile
sibling only sees short RPC sockets; idle-exit then retires the sibling and
the app restarts broadly. Arm a main-process /api/ws keep-alive for every
published sshConnections scope (including '') until teardown, without
treating spawn artifacts as liveness (#101626).
Co-authored-by: Cursor <cursoragent@cursor.com>
A renderer built after d9834a3e86 listens for srq- request frames; a v6 backend
still emits <kind>.request notifications, so every approval/clarify card would
silently never render. The skew toast now points the user at the backend update.
tests/contracts -> tests/tui_gateway/contracts (tree-layout rule: tests mirror a source
package). test_rpc_params_cannot_spoof_runtime_artifacts: forged owner_transport /
owner_session_record / owner_token keys are now refused at the wire (4000 + key path)
instead of silently dropped before the handler; the invariant (no steer reaches the
agent) is unchanged and asserted directly.
The staleness test regenerates in the Python lane, which has no
node_modules; prettier-dependent output would make the check pass locally
and fail in CI (or the reverse). Single-quoted literals, bare identifier
keys, no trailing commas or whitespace — prettier --check is clean on the
committed file.
apps/shared/src/gateway-events.ts is now a thin layer over
gateway-contract.generated.ts (client-local synthetic events + the
GatewayEvent envelope); gateway-events.json, its two rendezvous tests and
the duplicated BillingBlock / SessionInfo / ProjectInfo hand copies are
gone. Desktop, TUI, web and shared typecheck against the generated
RpcMethods / ServerRequestMap / BackendGatewayEventMap.
What tsc found once the types were honest: three phantom fields the
backend never sent (tool.start.todos, error.reason,
voice.transcript.voice_stopped) - the TUI todo tests were driving the
list through the phantom and are retargeted to tool.complete, where the
wire actually carries it; nullable fields (`None` on the wire) were typed
as plain optionals in eight places and now coerce at the boundary;
SessionResumeResult had a stale generic.
Contract fixes from the consumer pass: TranscriptMessage is the gateway
projection (text/row_id/context/args), not the stored row; SkinPayload
matches HermesSkin (empty-string defaults, never null); SessionLiveInfo
model/tools/skills are required (always emitted); BillingBlock.billing_url
is required-nullable (dataclass asdict).
tui_gateway/AGENTS.md documents the declare -> regenerate -> tsc loop.
Handlers own their documented domain codes (4006 missing session_id, 4015 bad
url, 4009 orphan claim); the contract's job on the way in is the one check no
handler performs — an unknown key (4000 with the key path). Missing/mistyped
fields are re-checked AFTER a successful handler answer under the strict
test policy, so a contract narrower than the wire still fails the suite.
Two models widened from the suite: SeedMessage (clients forward stored rows
verbatim), tool.complete.args (mirrored child rows omit it). Tests that
drove session.activate with prompt params (and vice versa) or stubbed
_live_session_payload with a bare {session_id} now send the real shapes.
215 methods, 13 server→client requests and 67 notifications now have Pydantic
contracts under tui_gateway/contracts/<topic>.py, rendered to
apps/shared/src/gateway-contract.generated.ts (616 types) and
gateway-contract.openrpc.json. tests/contracts/test_generated.py pins both
files to an in-memory regeneration and asserts catalog completeness from the
CODE side (every registered handler / emitted event / sent request has a
contract, nothing orphaned). scripts/ci/classify_changes.py runs the Python
lane when either generated file changes.
Phantom fields the hand-typed TS carried and no emitter ever set:
tool.start.todos, error.reason, voice.transcript.voice_stopped.
The old suites asserted the deleted wire (`*.request` events, `*.respond`
RPCs, `pending_clarify` snapshots, `_pending`/`_answers` teardowns). Each
test keeps its invariant against the new shape: a seeded live request's
`respond` spy receives the answer object, `hasOpenServerRequest` flips, the
`approval.respond` RPC fallback is asserted ONLY for queue entries restored
without a socket, and Bot Mode rooms answer via `request.answer` /
`clarify.lock`. The group-turns test that polled forever for a
`clarify.respond` that no longer exists (20-minute hang) now completes.
The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.
`tui_gateway/server_requests.py` owns the one mechanism:
send() block the agent thread until the response frame
(`srq-<n>` ids; ints belong to the client)
send_async() fire-and-callback variant (bot relay)
cancel*() withdraw with ONE `request.cancel {id, method, reason}`
event (timeout / interrupt / process exit /
answered elsewhere) instead of per-kind *.expire
open_requests() the still-open frames, replayed by session.resume,
session.activate and session.events.since so a
reconnecting client re-renders every kind, not two
clarify.lock stays a real client→server RPC (locks one batch
answer early); locked answers merge into the final
set even when the closing response carries only the
tail the user answered last
A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.
Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.
Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.