Commit Graph

1057 Commits

Author SHA1 Message Date
hermes-seaeye[bot]
c7c2df1a53 fmt(js): npm run fix on merge (#118435)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-21 18:24:39 +00:00
hermes-seaeye[bot]
bc655bfb40 fmt(js): npm run fix on merge (#118250)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-21 14:41:52 +00:00
Siddharth Balyan
afc3b7c6f3 feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation

Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash
merges (e0ef0eb9c3, d105376b21, ee2f5629b8, 1ab32b212b): the branch's history
no longer shared a base with main, so this is the PR's exact delta against
+3762/-1470, identical to the branch tip 05d7f2d3d5.

The fifteen commits it carried, in order:

--- feat(connectors): the wire follows the connector contract (six-state status, required toolkit metadata, connectionId on mint)

Hermes types exactly what the contract page writes: /tmp/magic/CONTRACT-TOOLKIT-METADATA.resolved.md
(carriers: portal PR2 `sid/connection-api` for the enum and account rows, the portal contract branch
on top of it for the list metadata). No optional-for-compat fields, no fallback branch, no seven-state
word left in the tree. A gateway that does not speak this contract fails validation loudly.

Wire (tools/connectors/gateway/wire.py): `ConnectionStatus` is the six values, `pending` covering the
vendor's INITIALIZING and INITIATED; `ConnectorListItem` requires title, description, https iconUrl and
authKind, and carries activeConnectionId when the session binds an account; `ConnectorListResponse`
types a page whole with its total; `ConnectorAccount` and `ConnectorAccountsResponse` type the account
routes; a mint result names its connectionId on `initiated` and never on `failed`; CONNECTION_REQUIRED
carries connectUrl and connectionId together or not at all; `account` exists on the execute call and is
never sent while multi-account is off.

Client: `list_connectors(search, connected)` refuses a search under three characters before any
request and parses every page whole; `list_accounts` and `account_status` read the account routes,
404 `connection_not_found` is None and 429 raises `RateLimited(retry_after)`.

Watcher: the target keeps the account id the mint named and the toolkit's title and icon from one
list read before the card is emitted; `_status_for` is the one seam between the watcher and the
gateway's status source (the list walk today; the per-account route replaces that body only).
`pending` and a missing status move nothing; revoked and inactive are failed.

Desktop: `ConnectorRow` is the strict list item; `ConnectorRowSeed` is what a tool call's args can
say; `connection.request`/`update` targets carry `connection_id`, `title`, `icon_url`; the mark
ladder is glyph, vendor icon as a plain image, favicon, monogram.

Tests run red first: wire rejects the retired words and the missing metadata; client search
minimum, execute omits account, account_status 404/429; managed targets carry the metadata and the
account id, a pending row moves nothing; the logo ladder order; the store passes the fields through.

--- feat(connectors): onboarding connect-first rides the connection operation

The guided first build had its own connector surface: a 2 s connectors.list poll loop, an auto-open
of every minted link, a hidden "[setup] links opened" note that painted as a user bubble, CONNECT
FIRST rows in the transcript, and a "Start with N apps connected" composer pill that sent the go
signal early. Delete all of it. The build session's first manage_connections connect now shows the
same card as any chat, and the settled tool result is the go signal.

Deleted: store/first-build-connectors.ts, assistant-ui/first-build-connectors.tsx (nothing rendered
it any more), lib/first-build-start.ts, onboarding-chat/start.tsx, the first-build session markers in
handoff-receipt.ts, the connectionRows / latestConnectorPart helpers they were the last readers of,
and the five strings only they used.

The runbook (setup-profile.ts::connectFirstRunbook) now describes the operation's real contract: one
connect call with every picked slug, one card with a row per app, blocks until settled, a per-app
result of connected / skipped / not_connected. No wait, no "Start with", no "already active" branch
(the result never carried it), and no offer to re-mint: Try again and Continue live on the card, and a
Continue means the user moved on.

Tests (red first): the runbook names one connect call and its per-app result and never the retired
model actions; the no-account rule survives an empty pick.

--- feat(connectors): the registry is keyed by profile, every watch read is bounded by the deadline, connectionId is optional, connect and execute carry returnTo and op

Registry (P1-14): tools/connectors/live.py keys an operation by (profile home, session key). Two multiplexed
profiles can carry the same timestamp-based session key; the tool thread opens under its turn's profile
override and the RPC side names the session record's profile home, the default profile resolving to the
process home on both sides.

Watch reads (P1-8 residual): list_connectors takes a per-page timeout and the watcher bounds each read by
the operation's remaining deadline (floor one second), so a stalled gateway page cannot hold the operation
past its deadline.

Contract follow-up from the portal handoff (/tmp/magic/HANDOFF-HERMES-PR3-PORTAL-CONTRACT.md): connectionId
on a mint result and on CONNECTION_REQUIRED is plain Optional (a no-auth toolkit answers active with none; a
failed mint answers with neither link nor id); a target without one has nothing to watch. Connect and
execute requests carry returnTo (hermes-desktop | hermes-desktop-dev | portal) and op, so the vendor's done
page can send the browser back to the app; the call sites and the deep-link handler follow in the next
commit.

Tests, red first: two profiles share a session key without seeing each other; each watch read shrinks with
the deadline; the client honours a per-page timeout; the optional id; the two request fields. The fakes'
list_connectors accept the timeout keyword; two wrapper spies that did not were the cause of a 300 s hang
under the per-file harness.

--- feat(connectors): an MCP target runs on the shared connection operation

manage_connections ran MCP targets as a renderer errand: the desktop card ran
the catalog lookup, the install and the OAuth flow, then told the backend what
had happened, and every surface without a card got `unavailable` plus two
terminal commands. That left the outcome in the renderer's word, and left the
model unable to connect an MCP server anywhere but the desktop.

The backend now owns the work, the way it already owns a managed connector.
`prepare` starts the OAuth flow in process and points the browser at the
backend's own callback route; an install that still needs credentials waits
pending and publishes their names as `required_env`; `observe` reads the flow
or the install worker on every tick. The worker itself moved out of
tui_gateway/mcp_oauth_sessions.py into tools/connectors/mcp_oauth.py, and that
module now calls it, so the Capabilities tab keeps its RPC session table.

The card may only say approved, skipped or continue: a claim of any other
state moves nothing. With that, `renderer_flow` and the `unavailable` state
have no producer left and are gone from the contract. Off the desktop there is
no card, so the action runs at once and the result carries the authorization
URL for the user to open.

Try again reaches the same work: `_reissue` dispatches on `target.kind`, so an
MCP row re-runs its own install, enable or OAuth instead of minting a managed
gateway link.

(cherry picked from commit fa0438c4121807604e7c983ba42b901314a1305b)

--- feat(connectors): the MCP card projects the operation instead of running the flow

The MCP setup card used to own the work: it called the catalog, ran
installMcpCatalogEntry, polled the action, drove the OAuth window, and then told
the backend what had happened through connection.respond. That made the renderer
a second authority on target state, so an install that finished after the window
closed, or a card that never mounted, left the operation with a state nobody
could correct. PR3 moves that work to the backend, so the card has one job left:
show the operation and send the user's consent.

McpSetupPending now renders request.targets the way ConnectorOffer renders them:
one row per target, one verb by (action, state), Continue below the shell.
Install and Enable send {status: 'approved'} and nothing else; an authorize row
opens the link the backend minted, with no OAuth RPC of its own; a running
install holds its verb with the busy mark; a failed row sends
connectors.connect {reconnect: true} on the open operation, the same call the
managed card makes. The owner lookup and that call are now shared helpers on
connector-tool.tsx rather than a second copy here.

ConnectionTargetOutcome loses connected / initiated / failed: those were the
renderer reporting state, which it no longer may do. ConnectionActor loses
renderer_flow for the same reason. ConnectionOperationTarget gains required_env,
the credentials an MCP install is still waiting for; the card renders a field per
entry under the pending row and holds Install until every required one has text,
so the values travel with the approval instead of through a separate catalog call.

(cherry picked from commit 2e42ad19f63e41f274a87199d07eb99a7c995cbe)

--- fix(connectors): the toolkit metadata leaves the wire; the vendor logo is derived from the slug

Sid cancelled portal PR3 (the toolkit metadata on the list item) on 2026-09-14, so the fields the contract
commit typed as required come off the wire: no title, description, iconUrl, authKind or activeConnectionId
on the list item, no total on the page, no search or connected query, no title or icon_url on a target
snapshot. The list item is exactly what portal #1220 emits.

The desktop derives the vendor logo from the toolkit slug instead (`connectorIconUrl`,
https://logos.composio.dev/api/{slug}; the gateway slug is the vendor slug, checked for every lead-order
pick), rendered as a plain image because the host sends no CORS header. The title stays
`connectorTitle(slug)`. The mark ladder is unchanged: glyph, derived vendor icon, favicon, monogram.

Everything else in the contract stands: six-state status, optional connectionId on mint results and on
CONNECTION_REQUIRED, returnTo and op, the account routes. The contract commit's message still names the
metadata; this commit is the correction.

--- fix(connectors): an MCP target survives the card's Continue, and the model is not sent to an action that refuses it

The review of the MCP backend found eight defects; each one is a test first.

- The off-desktop note told the model to confirm with action 'status', which refuses MCP
  names outright. It now says the authorization finishes in the background and the tools
  arrive on the next turn.
- The card's existence is a property of the session. A missing callback no longer routes a
  desktop session down the off-desktop path, where the model would be handed a live link.
- A second failure of a backend attempt was published with actor 'user'. ``refresh`` takes
  the actor, so the frame says who produced the text.
- Continue on the RPC thread can settle the operation between any read of a target's state
  and the transition that follows it. A settled operation has a frozen result, so the lost
  move is dropped; an IllegalTransition no longer escapes into the tool result. The same
  rule covers a worker whose outcome arrives late and a Try again that arrives after Continue.
- Try again on an install carried no credentials, so the second install ran with an empty
  env. The approved map is kept on the runner (never on the target: target fields are
  serialised to the model) and reused.
- An approval that does not cover a required credential no longer reaches the worker, where
  ``install_entry``'s prompt would block on stdin forever; the row waits for the card's
  fields instead.
- ``enable`` writes ``mcp_servers`` under the scope and lock the dashboard's toggle route
  uses, so the two read-modify-write paths in one process cannot drop each other's write.
- Several authorize targets start their flows together and share one wait; one wait per
  target kept the card empty for minutes.
- The off-desktop operation is in no session's registry, so it now emits no connection.update.

Try again is also refused for a target the transition table cannot move back to 'initiated'
(an MCP target has no move out of 'expired') and for an operation that settled during the call.

The contract test that froze the Actor enum is replaced by two behaviour tests: no actor but
the backend watcher can connect an MCP target, and the card's word never moves one.

(cherry picked from commit 081c16412b48667758a77a7e7c29eaf2be6a93dd)

--- fix(connectors): the watcher reads one account route per target, and every frame carries its seq

The watch loop walked the whole toolkit list once per tick to learn whether one
target had connected. That read costs a vendor call per page, cannot tell one
account from another, and forced `awaiting_new_attempt`: after a forced
reconnect the list still reported the OLD account `active`, so the row had to be
disbelieved until it read as something else once. The gateway now serves
`GET /v1/connectors/accounts/{connectionId}`, so each pending target reads its
own account: the id the mint named, one read per target per tick at 1 Hz (the
route's bucket is 180/min per principal), each bounded by the operation's
remaining deadline. A 429 parks that one target until its Retry-After passes and
leaves the others reading. A 404 is "not yet" until the deadline. A target the
mint gave no account for has nothing to read, so it is not read. A forced
reconnect watches the new account, which is why `awaiting_new_attempt` and its
two tests are gone.

`mint` and `run_remote` now name where the browser should come back to
(`returnTo`, plus the operation id on a mint), so the vendor's done page can
hand the user back to the desktop app that asked instead of stranding them on a
web page. Only the desktop registers that URL scheme, so no other surface sends
either field.

`connection.update` frames were built by re-reading the operation after the lock
was released, so a second writer could give an older frame a newer state and the
renderer could not tell which frame was last. Every write now advances a
monotonic `seq` and takes its snapshot under the same lock, and the emitter sends
that snapshot; a renderer that keeps the highest seq per operation can drop a
frame that arrives out of order. `connectors.operation.wake` lets the desktop's
deep-link return ask for a read now instead of at the next tick; it only shortens
the wait and trusts nothing else in the link.

(cherry picked from commit f5423cf72302b92b26c043495184ad6426fa2b36)

--- fix(connectors): the desktop card follows the operation, and says so out loud

The MCP card was a second implementation of the connector card with the
review's defects: it painted for any request on the session, kept its
controls after the operation settled, offered an approve verb while the
backend was still minting an authorize link, and re-enabled Install when
the RPC returned rather than when the state frame moved the row. A second
click in that window sent the consent twice.

The renderer now reads the operation's `seq`: an update or a status frame
whose sequence is not greater than the one already applied changes
nothing, and a resume snapshot neither revives a settled card nor puts
back a row a newer frame has moved. Without it the transport's ordering
decided what the user saw.

`hermes://connections/done?op=…` brings the user back from the browser to
the session that opened the operation and wakes its watcher, so the row
moves at once instead of at the watcher's next tick. Only the op id is
used; the link's status moves no row.

Each row's mark and cue are one polite live region, so a row that flips is
heard and not only seen, focus follows the row the backend moved while the
card holds it, and the waiting mark stops spinning under reduced motion.

(cherry picked from commit 7c30d252556d7d496ddb354b50bcecc41bceef27)

--- fix(connectors): the renderer types seq as the wire carries it

The watcher commit made `seq` a required field on every operation frame. The
desktop store still declared its own optional `seq` so it would compile against a
backend that predates the field; that backend no longer exists on this branch, so
the hedge is dead code and the fixtures were short one field. The store now reads
`seq` from the shared types, and the fixtures count the way the backend does.

The repeat-frame test asserted the whole request keeps its reference. With a real
rising `seq` the request must change; the invariant the test guards is that the
target row keeps its identity so open credential inputs do not remount.

--- fix(connectors): a URL that arrives after the shared wait still lands on its row

The shared prepare wait failed every row still pending when it ran out, while
that row's own thread was still waiting on the provider. When the URL arrived a
moment later the thread's move raised inside the daemon thread, the link was
lost, and Try again started a second flow. The wait now bounds only how long
prepare blocks; a row still pending afterwards is left to its own thread, which
is the only writer of that row and ends with the URL or the flow's own failure.

The comment on the approved credentials said every target field is serialised to
the model; it is not (the snapshot names its keys). The reason they live on the
runner is that they are secrets and the runner's life is exactly the operation's.

--- fix(connectors): the review findings the operation must survive before the fold

The MCP prepare threads and the install worker started with an empty context,
so a named-profile turn's home override never reached them: the flow resolved
the process home's `mcp_servers` and stored the token there. Each thread now
runs in a copy of the calling thread's context. `connection.respond` had the
same gap on the RPC thread: an approval ran the enable, which writes
config.yaml, with no profile bound, so the flag landed in the launch home. The
handler binds the session record's profile the way `_connector_rpc` does;
`config_write_scope(None)` keeps that override, so the enable needs no change.

The watcher raised out of the tool when a per-target Skip landed while that
target's read was in flight: the operation stayed open with `live` closed, and
every later answer got 4004. A read for a row that is no longer live is dropped
at debug; only a refusal on a live row is still a contract violation. The same
skip from the card raced a row the backend had just connected and aborted the
rest of the answer; a skip for a resolved row is ignored and every entry, then
the settle check, still runs.

A read was bounded by the whole remaining deadline, so a hung gateway held the
first read for 300 s and Continue could not return the tool; a read now waits
ten seconds at most. A 429 parked only the target that read, but the budget is
the principal's, so the next target's read in the same tick spent it again:
every live target waits out the one Retry-After.

The MCP surface rule read the platform alone, so a desktop call without the
callback (registry dispatch from execute_code) opened an operation nobody
rendered and blocked for the deadline; it now uses the managed rule, surface
and callback. A write after settlement advanced `seq` while emitting the
frozen frame, so the resume snapshot named a seq no frame carried; the counter
stops at the settle frame. A mint that reports `initiated` with no account id
logs that the watcher cannot read the row.

(cherry picked from commit 587b228016bf8ee2b022e7fa58c2d9f6057809e9)

--- fix(connectors): the desktop card holds a verb until the backend answers, and never takes the keyboard from a credential field

The review of the desktop card found five defects; each one is a test first.

- A resume snapshot was refused whenever the cache held a settled operation, whichever
  operation it was, so a session that opened a second operation after settling the first
  never got its card back from a resume. And the refused snapshot handed the caller the
  settled cache as "the request", so the session was flagged as needing input behind a
  summary with no controls. Only the same operation can refuse the snapshot now, and a
  refused one is no pending card.
- The focus handoff picked the row's first button, which after pending -> initiated is the
  disabled working verb; the focus call was a no-op and the keyboard landed on the document
  body. It picks the first control that can take focus, else the row. It also moved focus out
  of a credential field the user was typing in whenever another row moved; it leaves an
  editable alone. The "focus Continue once every row resolved" branch was dead (Continue
  unmounts the moment nothing is unresolved), so it and its ref plumbing are gone.
- The done link navigated to a settled operation's session and rejected when the wake RPC
  did (4004 once the operation left the live registry). A settled request is ignored, and a
  refused wake is nothing: the wake only shortens the wait, the watcher still ticks.
- Install spun forever when the store refused to send the consent (the operation gone or
  settled under the card): `respondToConnectionRequest` resolves false in that case and the
  verb was only released in `catch`.
- After a partial approval (a required credential missing) the backend answers with a
  same-state frame whose detail names what is missing; nothing released the verb because it
  was held until the row's state moved. The row now remembers the seq the click saw and holds
  the verb only until a frame past it arrives, which is the backend's word on the click
  whether or not the row moved.

(cherry picked from commit 4c534d6e2af979778d9a2423cf97d717edeaa1a9)

--- fix(connectors): a skip that loses the race to the watcher is ignored, not raised

The skip guard read the row's state and then moved it; the watcher can connect the
row between the two, and the refused move aborted the rest of the card's answer.
The refusal itself is now the witness: a move refused for a row that is resolved,
or on a settled operation, is the same nothing-to-do as a row resolved earlier.

Two recording fakes in the managed tests kept their lists on the class; they now
start per instance so a lifted fake cannot share reads between tests.

--- fix(connectors): a resume that lands behind a newer live frame still reports the pending card

The refused-snapshot branch answered "no pending card" for both reasons it can
refuse: the operation settled, or a newer frame already moved a row. Only the
first is no card. For the second the live card is still open and blocking the
turn, so the caller must keep the session flagged as waiting on it.

* test(connectors): defer new connection coverage until implementation settles

Remove PR-added test cases and their unused helpers while retaining
existing tests adapted to the changed connection contract. The three
PR-only renderer test files are removed for now.

Focused behavioral coverage will be added as the final implementation
step before verification. Existing main coverage is not being removed
wholesale, and this does not declare the feature merge-ready.

* fix(connectors): commit MCP authorization at initialize, save setup values after success

The OAuth probe treated one exception as one outcome: any failure after the browser
step restored the token snapshot and manager entry, so a server that accepted the
token but failed tools/list discarded a completed consent. Now the probe reports
whether initialize succeeded (details["initialized"], read from the claimed
MCPServerTask). Failure before that point rolls back as before. Failure after it
saves the server config, keeps the tokens, and reports tools unavailable through
flow.discovery_error; the card can retry discovery without repeating consent.

Catalog install wrote the submitted values to .env before install_entry and the
probe ran. The values now live in the secret scope for the duration of the install
(get_env_value reads through get_secret, so install_entry finds them without a
prompt), and .env is written only after the probe returns tools. A failed probe
removes the server block and writes nothing. _probe_tool_names returns None on a
failed probe instead of an empty list, so failure and a valid empty listing are
distinct.

Failure text is redacted before it reaches target detail: every value the card
submitted for that target is replaced by exact match, then the pattern redactor
runs.

required_env now carries the manifest's secret and default flags; Target carries the
manifest's post_install text as instructions. The wire contract gains secret,
default, instructions and discovery_error; generated TS and OpenRPC regenerated.

* fix(connectors): one OAuth callback receiver picker for the connection card

The card's authorize target built its redirect from the dashboard web server and
raised when none was bound in the process, so a standalone hermes --tui session
could never authorize an MCP server. The receiver is now chosen in one place
(choose_callback_receiver): a pinned pre-registered client keeps the SDK's own
listener on the registered port; a client-advertised loopback URI is used as-is and
its callback arrives through the mcp.servers.oauth.callback relay; otherwise the
backend binds a one-shot loopback listener and feeds it into the flow. The dashboard
route stays with the dashboard web page, which cannot bind a port.

tui_gateway/mcp_oauth_sessions.py had a second copy of the loopback listener and a
_worker that referenced _probe_with_rollback, set_hermes_home_override,
reset_hermes_home_override and Path without importing them, so every RPC-started
flow raised NameError. Both are deleted; start_flow spawns run_worker directly and
uses the same receiver picker. Flow registration is shared (register_flow /
finish_flow) so a card-started flow with a client URI is reachable by the relay.

Under an SSH session with no client listener the attempt's detail carries the
existing paste-the-redirect instructions. Nothing on the card path opens a browser.

* feat(desktop): setup-form modal for MCP connection cards

The connection card rendered an MCP server's setup fields inline: every field as a
password input, no default value, no instructions. A URL such as the n8n MCP
server URL was typed blind, and the manifest's setup text never reached the user.

Two components carry the form now. SetupFieldList renders the ordered fields the
backend declares (a plain field as text prefilled with the manifest default, a
secret field masked and empty). SetupFormDialog composes it with a one-line title,
the manifest instructions, an inline error for a failed attempt, and Cancel /
Connect. The row's Install action opens the dialog when the target has fields;
Connect sends {status: approved, env}; Cancel sends {status: skipped}. A failed
attempt keeps the dialog open with the draft intact. The draft lives in the dialog
component only; nothing reaches the store or the resume snapshot.

Once the backend publishes the authorization URL the dialog shows it as text with
an Open in browser button. Nothing opens a browser on a state change: the Try
again path on both cards used to open the re-minted link at once; it now waits for
the row's update frame and the user's click.

The store types gain the wire's secret, default, instructions and discoveryError
fields and normalise them; a connected target with discoveryError renders as
authorized with tools unavailable.

* feat(tui): connection card in the Ink TUI

The Ink TUI had no client for the connection operation: connection.request,
connection.update, connection.respond and the pending_connection resume snapshot
were unhandled, so a manage_connections call in a TUI session could only print a
link through the model.

connectionOperationStore.ts holds the backend snapshot: a request opens only for a
new operation id, an update applies only to the live operation with a higher seq,
a settled update freezes the id so a late request frame cannot reopen the card.
The gateway event handler feeds it; session resume hydrates it from
pending_connection.

connectionSetupOverlay.tsx renders one callout in the prompt zone: a one-line
title, the manifest instructions, every field (plain rows prefilled with the
default, secret rows masked), then a Connect / Cancel selector. Connect sends the
draft through connection.respond; Cancel skips the active target. When the backend
publishes the authorization URL the callout shows it as text under "Press Enter to
open in browser"; Enter is the only thing that opens it. A failed attempt unlocks
the fields with the draft kept. An accepted secret renders as "Set" and is never
echoed. A target authorized without tools shows that state and Continue.

The overlay joins the existing input-owner set in overlayStore so typing and other
prompts are blocked while it is up.

* feat(cli): connection panel in the classic CLI

The classic hermes CLI passed no connection_callback, so a manage_connections
call could only print an authorization link through the model and could not take
a setup value at all.

The CLI now renders the connection operation as a prompt_toolkit panel, on the
same queue mechanism as the clarify panel: the agent thread's callback opens the
panel, blocks until the first decision, then returns so the operation's watch loop
runs; every later state reaches the panel through ConnectionOperation.on_change
(installed only when no gateway hook is set, restored on close).

Panel: one-line title, the manifest instructions, one row per field (plain rows
prefilled with the default, secret rows masked in the input buffer and rendered as
"Set" once accepted), then Connect / Cancel. Connect sends the draft through
apply_answer; Esc skips the active target; Ctrl-C sets the tool-thread interrupt so
the operation settles as interrupted. Once the backend publishes the authorization
URL the panel shows it with the target detail and "Press Enter to open in browser";
Enter is the only thing that opens it. A failed attempt unlocks the fields with the
draft kept; an authorized target without tools offers Retry discovery / Continue.

Up/Down move between rows, Left/Right toggle the action, and the panel joins the
blocking-overlay guards so chat input, history and voice stay out while it is up.
The single-query (headless) mode passes no callback, as it does for clarify.

* feat(connectors): register a connected MCP server and report its tools in the result

After a successful install or authorization the target carried the probe's tool
names and the model was told the tools "become available on your next turn". MCP
tool schemas are deferrable by construction, so nothing about them lives in the
sent tool array; a server registered in the scoped registry is callable through
tool_describe/tool_call in the same turn. The operation now registers the server
(register_mcp_servers under the owner's home scope) once authorization is
committed, records the registered names on the target, and the settled result
carries a tools_listing block in the deferred-catalog format plus a note that the
tools are callable now. A registration failure keeps the target connected with
tools: [] and a sanitized discovery_error. agent.tools and the system prompt are
untouched; the between-turns refresh updates the catalog block as before.

The card gate no longer asks for the desktop platform. Every surface that renders
the card attaches a connection callback (Desktop, the Ink TUI, the classic CLI);
registry dispatch and messaging sessions attach none and keep the link result.

Pre-commit rollback in probe_with_rollback used restore(only_if_absent=True),
which skips the rollback when a token file exists. On a first-ever authorization
the only file is the one this attempt wrote, so a token the resource rejected was
kept. Live E2E (controlled provider answering 403 to the issued token) showed the
row fail and the token survive; the rollback now restores the snapshot outright,
and the same run shows the token file removed.

* fix(connectors): install an OAuth catalog entry through the card's own flow

The card's install ran install_entry and then a plain probe. For an OAuth entry
that probe has no card flow around it: in the desktop backend it failed at once
("non-interactive environment and no cached tokens"), and in the classic CLI it
saw a TTY, opened a browser by itself and drew install_entry's curses tool
checklist over the panel. Since a probe failure is now an error, the failed
install was rolled back with _remove_mcp_server, and authorize refuses a server
that is not configured. A clean home had no path to a connected OAuth entry, which
is 55 of the 65 catalog entries. Found by the live three-surface run.

An install now builds the entry's configuration in memory (card_install_config:
no prompts, no probe, no checklist; a prior tool selection or the manifest's
curated filter applies). An OAuth entry starts the same flow authorize uses with
that configuration and the setup values in the attempt's secret scope, so the row
reaches the URL step, and the configuration and the setup values are saved
together when initialize accepts the token. Every other entry is probed in memory
and saved after the server answers. A failure writes nothing, so a failed
reinstall keeps the previous configuration, and the failed row asks for its fields
again so the card can reopen the form over the draft it kept.

* fix(agent): make a server connected in this turn callable in this turn

tool_describe and tool_call resolve names inside the agent's toolset selection,
which is fixed when the agent is built. A server that manage_connections had just
registered was therefore "not found" for the rest of the turn whose result calls
its tools available, and stayed out of the next turn's catalog too. The live
Desktop run showed it: the result listed mcp__fx_oauth_fields__echo and the
tool_describe that followed answered not_found.

The executor now adds the MCP servers the call connected to the selection. Only
the selection changes; agent.tools does not, so the sent tool schema bytes stay
the same. A selection of None (every toolset) and the no_mcp sentinel are left
alone.

* fix(cli): reopen the connection form when a required field is still empty

When the backend refuses an answer because a required field is missing it keeps
the row pending and names the fields. The panel mapped that frame to its waiting
phase, which draws the detail line and nothing else, so the user saw "waiting for
FX_API_KEY" with no fields and no buttons until the 300 s deadline. The panel now
returns to the form on the first missing field, over the draft it kept.

* test(connectors): the no-card path is the one with no callback attached

The card gate is "a connection callback is attached" since the classic CLI got
its own panel. This test still attached one under platform "cli" and expected no
card, so it opened a real operation, waited out the 300 s deadline and failed.
It now drives the path that has no card: no callback.

* fix(connectors): order a cancel against the commit and stop deleting a working grant

Found by the live runs on Desktop, the Ink TUI and the classic CLI.

A user's skip did not stop the OAuth attempt. With the worker parked in the token
request, the row settled "skipped" and the token was written 33 s later. A skip
now cancels the attempt, and one lock orders that cancel against the commit: the
attempt is either canceled with the earlier tokens restored, or committed and
kept. When the commit won, the skipped row says the authorization was kept. An
interrupted turn cancels its attempts the same way.

Every attempt began by deleting the saved tokens, so retrying discovery for an
authorized server demanded consent again, and a cancel in between left no grant
at all. The card's flow now connects with the saved tokens first, with no browser
step; only when they do not work does it replace them. The RPC session surface
keeps the old behavior because its caller waits for an authorization URL.

A second Connect from the form a failed row reopened was dropped without a frame,
which left the Desktop dialog with Connect loading and Cancel disabled until the
deadline. An approval on a failed or expired row is now a retry with the new
values.

The reported tool names came from the registration call, which returns nothing
for a server the process already holds. A retry after a failed listing therefore
said "no tools" while the server had them. The names are read from the registry,
and a parked server is woken first.

An authorization that commits after its card closed by deadline was never
registered, so the next turn still could not reach it. The runner keeps such
attempts, and the between-turns refresh adopts the ones that were approved.

An error with no message reached the user as a class name ("CancelledError").
The tool description still said a server's tools arrive on the next turn.

* fix(desktop): let an OAuth install open its link, and keep the form's draft

An install of an OAuth catalog entry with no setup fields reached the URL step
and gave the user nothing to click: an initiated install was always drawn as a
disabled spinner, and the only other place the link is shown is the setup dialog,
which opens for entries with fields. That is most of the catalog. An initiated row
that carries a link now offers Open, whatever the action.

The setup dialog reset its draft whenever the field list changed identity. The
backend sends a fresh list with every frame and an empty one while an attempt
runs, so a failed Connect erased what the user had typed. Fields now only fill in
what the draft lacks, and closing the dialog drops the draft.

The settled summary dropped the "tools unavailable" fact, and the live row offered
a Try again for that state which the gateway refuses for a connected target. The
summary keeps the fact and the row offers no dead control.

* fix(tui): keep the typed draft when a Connect fails

The overlay reset its draft whenever required_env changed identity, and every
backend snapshot delivers a freshly parsed array. A failed Connect therefore came
back as a form with the default region and an empty secret. The draft now resets
per target only, and a snapshot's fields fill in what the draft lacks.

* fix(cli): show the URL step and start an install that has no fields

The panel applied the user's answer and then set its phase to "waiting". The
backend applies the answer on the same thread and its change hook had already set
the next phase, so the URL step of an OAuth install and the reopened form for a
missing field were both overwritten, and the card sat on "Waiting…" until the
deadline. The waiting phase is now set before the answer is applied.

A pending install or enable with no fields opened in the waiting phase, which
sends no approval, so the flow never started. It now opens on Connect/Cancel. A
target that already carries its link opens on the URL step.

Connect on a failed row re-ran the attempt without the values now in the draft; it
sends them. The authorized-without-tools phase offered a "Retry discovery" that a
connected target cannot run inside the same operation; it offers Continue.

* fix(connectors): a newer OAuth attempt replaces the older one for the same server

A card that closed by deadline leaves its worker waiting on the browser for up to
300 s, and a tampered callback leaves one waiting too. A retry or a new operation
for the same server then ran beside it: both wrote the same token files, and the
older one's rollback could write over the newer attempt's grant.

The newest attempt per home and server is recorded. Starting one cancels the older
attempt, takes over its pre-attempt snapshot so a later failure still restores the
state from before either, and the older attempt's rollback leaves the files alone.

* fix(desktop): label an install's link control "Open in browser"

The control that hands an install's authorization link to the browser reused the
action's verb, so the row showed "Install" before the click and "Install" again at
the link step, told apart only by the cue. It now reads "Open in browser", the
label the setup dialog already uses for the same act. Authorize keeps its verb.

* docs(mcp): describe the setup card on the desktop, the terminal UI and the CLI

The MCP guide said the chat install exists only in the desktop app and that the
CLI relays commands. All three surfaces now show the same card: fields, Connect or
Cancel, an authorization link the user opens, one save when the server has
accepted the token, and tools the agent can call in the same turn. The tools
reference gains the result fields (tools, tools_listing, discovery_error) and the
no-card behavior of an OAuth install.

* fix(cli): keep the layout hook callable without a connection widget

The CLI panel commit added connection_widget to _build_tui_layout_children as a
required keyword. That method is the documented override point for wrappers, and
five existing tests (extension hooks, prompt stash, subagent dock) call it without
the new argument, so CI failed with a TypeError. The argument is now optional, like
the other widgets added after the hook was published; a missing widget is left out
of the layout.

The settled tool result also carried the target's setup instructions. Those are
the card's text for the user, and a catalog entry's notes can predate this flow
("restart your session so the tools are loaded"), which contradicts a result that
says the tools are callable now. The model-facing result drops them; the card
payloads keep them. This restores test_connector_local_batches.

* fix(connectors): wait for an in-progress registration before reporting a server's tools

Found with the real Vercel MCP server. Saving the configuration wakes the config
watcher, which starts its own connect for the new server. The operation's
registration call then skips the server as "already connecting" and returns at
once, so the card settled "connected" with no tools, and the 214 tools were
registered three seconds later. With no names in the result the model searched,
found the hosted connector of the same vendor and asked the user to connect that
instead.

The registered names are now read once the registration has finished: while
another task is connecting the server the read waits (30 s at most), a parked
server is woken once, and a server that finished registering with no tools is
still a valid empty list.
2026-09-21 10:04:32 +05:30
teknium1
1d30aa5571 feat(tui): Processes block in the live-work dock and /agents overlay
The Ink dock and /agents overlay only merged subagent.* events with
subagent.list. They now poll process.list for the owning session on the
same cadence and paint a Processes block under the agents (⚙ command ·
elapsed · last output; exit verdict for 60 s after exit). The dock budget
is split so neither block hides the other; processes alone still surface
the dock. Glyphs live in lib/processGlyph.ts so panel and overlay match.
2026-09-20 13:55:03 -07:00
teknium1
2c78b9b39e feat(process_registry): stamp exited_at and expose it on process.list
The live-work docks retire a finished background process ~60 s after it
ends; the registry only knew started_at, so the age of an exit was not
observable. _move_to_finished is the single choke point every exit path
(reader loop, reconcile, kill) passes through, so the stamp lives there.
completion_reason rides along so a killed process can read "killed"
instead of "exit -15".
2026-09-20 13:55:03 -07:00
Konstantin Khlopkov
1cf8a9fb41 fix(tui): count streamed frames as heartbeat liveness
The TUI's JsonRpcRequestChannel took the 'response' liveness default, so
only a ping-ack or a response to a pending call refreshed the deadline:
a socket streaming a delta every second through a 134s turn was declared
dead at the 45s deadline and force-closed, splitting sessions that went
on to complete server-side (#115251).

Pass heartbeatLiveness: 'any-inbound' on the TUI channel — the same
contract the desktop/web client already uses — so any inbound frame
counts as life. A silent drop still trips the deadline, so true
dead-transport detection is unchanged; only the false-kill class goes.

The channel's 'response' mode keeps its semantics for callers that
explicitly choose it; the shared wiring test pins the TUI construction
site so the option cannot silently regress.

Fixes #115251
2026-09-20 12:18:44 -07:00
teknium1
d6a69ea5a3 fix(tui): treat every ctrl chord as a binding name, never typed text
Widen #115382 from bare C0 bytes to every ctrl-modified keypress: a kitty
CSI u / xterm modifyOtherKeys Ctrl+L (`ESC [ 108 ; 5 u`) reaches parseKey
with ctrl set and name `l` exactly like the 0x0c redraw byte, so the
composer typed an `l` for it too. `isControlChord` is ctrl && !isPasted;
bracketed pastes stay text. Test trimmed to two invariants (chord is
bindable and refused by the insert gate; typed text and pastes still land).
2026-09-20 12:11:10 -07:00
finn763
addf694901 fix(tui): a ctrl chord's control byte is not typed text (stray l after resume)
The dashboard asks a reused PTY's TUI for a full redraw on every re-attach by
writing the force-redraw byte Ctrl+L (0x0c, `hermes_cli/pty_session.py`
TUI_FORCE_REDRAW) into its stdin (`web_routers/chat_ws.py`:
`session.attach(ws, force_redraw=not _created)`). parse-keypress names a C0
control byte after the letter it encodes (0x0c -> 'l') and `parseKey` hands
that name to `InputEvent.input` so bindings can match ctrl+<letter> on the raw
byte; the composer's insert gate read the name as typed text, so the redraw
byte landed in the input box as a solitary `l` that prefixed the next message
(#115284 — also on tab switch and window restore, both of which re-attach).

Flag the event instead of touching `input` (clearing it would break the
composer's own ctrl+a/e/u/k/w/z/y branches and the global ctrl+c/x/o/t
pass-through): `InputEvent.isControlByteChord` is true only for a single-byte
C0 sequence with ctrl set, so kitty CSI u / xterm modifyOtherKeys chords and
bracketed pastes are unaffected. The composer skips both insert sites for it,
and still flushes a pending key-burst because a chord is not typing.
2026-09-20 12:11:10 -07:00
teknium1
b787fb9128 fix(tui): Ctrl+D exits from an empty composer on macOS too
The exit binding matched isAction(key, ch, "d"), which on macOS means Cmd+D:
Ghostty consumes Cmd+D for split panes and literal Ctrl+D never matched, so
the TUI stayed open. Ctrl+D is the terminal EOF convention, not a Cmd
shortcut, so it now also routes through the existing isMacActionFallback seam
(target union gains "d"), and on every platform it exits only when the
composer holds no text, buffered lines or attachments, matching the classic
CLI. Slim redo of #116454.

Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-20 11:15:12 -07:00
Konstantin Khlopkov
6abbc02228 fix(tui): no gateway respawn after graceful-exit kill (#114987) 2026-09-20 11:04:04 -07:00
teknium1
171a1777b5 fix(desktop,tui): reasoning pill and effort rows say ultra sends max on this route
The Desktop pill, the catalog row meta and the effort radio row rendered a
clamped pick as a plain "Ultra", and the TUI status bar as "ultra" — presenting
a Hermes-internal step as a wire level the route does not have, while the CLI's
`/reasoning` shows "ultra (sends max on this route)" (#115876).

Both surfaces now read `session.info.reasoning_effort_wire`:
- Desktop pill / catalog row meta: compact "Ultra→Max", tooltip and aria-label
  "Effort: Ultra (sends Max on this route)" (new `modelOptions.sendsOnRoute`
  string in every locale, mirroring the CLI wording).
- Effort radio row for the selected clamped level: "Ultra (sends Max on this
  route)".
- TUI status bar: "ultra→max".
- Unknown wire ('' — not stamped yet, or an optimistic pick) or a verbatim
  level makes no claim, so nothing changes on routes that send the level as
  is. `setCurrentReasoningEffort` and the tile optimistic write clear the wire
  so a stale clamp never pairs with a new pick until the gateway re-stamps.

Docs: one sentence in the Desktop guide. Part of #61634.
2026-09-19 11:22:44 -07:00
teknium1
6f6ed01355 fix: WAF 403s stop reading as key rejections; anthropic_messages routes send custom_providers extra_headers
Two gaps for custom providers behind a WAF/CDN:

- `build_anthropic_client` never consulted `custom_providers[].extra_headers`,
  so a relay in `anthropic_messages` mode that rejects the SDK User-Agent kept
  403ing even with `extra_headers: {User-Agent: ...}` configured, while the
  OpenAI-wire clients already applied it. The lookup now lives in
  `_new_sdk_client`, the one constructor every builder path goes through
  (init, /model switch, rebuild, auxiliary), keyed by the caller's raw route
  because entries are keyed by the `/v1` form the normalizer strips.
  Salvaged direction of #46002 (@wait4xx). Fixes #24293, #9721.

- `_status_403` classified every non-billing 403 as `auth`, so a WAF's plain
  "Your request was blocked." or a Cloudflare browser challenge printed "Your
  API key was rejected" and could rotate a healthy credential. A 403 carrying
  established block/challenge markers is now `upstream_blocked`: no rotation,
  no retry, fallback allowed, WAF/User-Agent guidance on every surface (CLI
  loop, chat copy, cli chat error copy, TUI gateway + Ink TUI copy). Generic
  403 and all 401 keep the auth verdict. Salvaged direction of #70567
  (@ooiuuii) and #53114 (@AgenticSpark). Fixes #53099, #70566.
2026-09-19 09:57:21 -07:00
teknium1
5ae7ec8623 fix(setup): a provider configured after boot unblocks the dashboard chat without a restart
The serve process's free-tier boot record (`free_tier_bootstrap._record`) is
built once; when the boot inventory found nothing (mint failed or the tier is
off) it says `provider_configured: false` for the process lifetime and
`setup.status` answers from it, so the Ink chat (dashboard /chat, `hermes
--tui`) parks every new session on "Setup Required" no matter what the user
configures afterwards — the Models page, a picker key save, or `hermes setup`
from a shell all land on disk and change nothing in the running process.

Reconcile the record on read instead of re-minting on write: `reconcile_record`
re-runs the cheap inventory (`_inventory_other_providers`, the resolver ladder
with the free-tier rung hidden) when a `False` record's config files
(`config.yaml` / `.env` / `auth.json`) moved since the inventory that built it,
replaces the record's inventory half and broadcasts `setup.ready`. The mint
verdict is kept as is: only the boot bootstrap and its retries mint.
`wait_for_record` (what `setup.status` reads) reconciles, which also covers
writes this process never saw (`/setup`'s `hermes setup` handoff, `hermes model`
in docker exec, a hand edit); the two in-process write paths from #114708
(`/api/model/set`, `model.save_key`) call it for the immediate broadcast.

Dropped from #114708: `retry_bootstrap_mint(force=True)` on the write paths — a
forced portal mint (network, cooldown bypassed) inside an HTTP write; the record
only needed its inventory refreshed. Not taken from #108773: OR-ing the record
with `_has_any_provider_configured()` — that first-run guard counts host-wide
credentials (other host-wide agent-CLI logins) and answers True on a blank machine (verified
live on this host), which would open the gate with nothing configured.

The 4001 hint for a sessionless `config.set model` named "Settings -> Models",
a Desktop-only surface; the Ink chat has no Settings and the dashboard has a
Models page. One string now names /setup, the Models page and Settings ->
Models. The TUI's "Setup Required" panel no longer advertises `/model` as the
in-place fix (sessionless picks are refused by design, 75b1b43ec1).

Co-authored-by: chelsealong <chelsealong@126.com>
2026-09-18 12:52:36 -07:00
teknium1
9b0ca895b3 fix(tui,desktop): send session_id on commands.catalog, complete.slash and skills.reload
The clients called all three RPCs with `{}`, so the server's session-less
fallback (`_completion_cwd({})`) read the process TERMINAL_CWD, which
`gateway/run.py` rewrites to $HOME at import in the default-config case: the
catalog, popup and reload still missed the session's project skills.

- ui-tui: `useCompletion` adds `session_id` to `complete.slash`;
  `/reload-skills` sends it on `skills.reload` and the follow-up
  `commands.catalog`; the gateway-ready catalog fetch sends it when a session
  already exists (reconnect). Before the first session the gateway binds the
  same workspace it seeds a new session with.
- Desktop: `useSlashCompletions` takes `sessionId`, sends it on both RPCs and
  keys the completion cache per session so two chats in two repos don't
  serve each other's catalog.
2026-09-18 10:22:47 -07:00
teknium1
f7e70e954e test(ui-tui): the no-heartbeat check expects the capability advertisement
Every gateway.ready now puts one client.capabilities frame on the wire, so
"no frames" is no longer the invariant; "no gateway.ping" is.
2026-09-17 09:04:38 -07:00
kshitijk4poor
dc8fe4def8 refactor(shared): a crashed server-request handler reports through an owner hook, not console.error
The deliverRequest catch branch wrote to a module-level console.error sink
(the only direct console.* in apps/shared/src) and fired onUnhandledRequest,
whose contract is "nobody handled it, already answered -32601" — so the TUI
logged a -32603 crash as "unhandled server request".

Add onRequestHandlerError(error, request) to JsonRpcRequestChannelOptions
beside onHeartbeatFailure, call it from the catch after answering -32603,
and drop the console sink. Wire both owners: HermesGateway (desktop, via a
GatewayClientOptions passthrough) logs to console.error like its dial-failure
sink; ui-tui gatewayClient pushes a [protocol] log line. Collapse the two
normalisation arms into the existing `error instanceof Error ? … : new
Error(String(error))` idiom and restore the early `return true` instead of
the handled flag + break — nothing runs after the loop but the -32601
fallthrough.

Test: the crash case now asserts onRequestHandlerError fires once for the
-32603 request and onUnhandledRequest only for the -32601 one.
2026-09-17 21:18:37 +05:30
KoNit-K
9583c8c45a fix(tui): preserve inflight synthetic display metadata 2026-09-16 17:54:17 -07:00
teknium1
04ded3b145 fix: cover the geocoding 'location not found' branch in the weather test
Review noted the error-phase test only exercised the HTTP-503 path; the
branch where geocoding returns `{results: []}` was untested. Extend the
existing error test with a second stub so both failure routes are
asserted without adding a test case.
2026-09-16 16:49:48 -07:00
Vitor Cepeda Lopes
b5821a578d fix(tui): move weather widget to Open-Meteo 2026-09-16 16:49:48 -07:00
teknium1
abdb402701 fix(mcp): carry the lazy status across the TUI wire, tests and docs
Follow-up to the ported status fix:

- `tui_gateway/contracts/tools_mcp_plugins.py::McpRuntimeStatus` is a
  closed wire enum; `mcp.servers.status` would raise `ContractViolation`
  on the new `lazy` value. Declare it and regenerate the TS/OpenRPC
  contract files.
- `ui-tui` session panel: an unknown status fell through to the red
  `failed` branch; render `lazy` with its cached tool count (inline
  branch, no component extraction).
- Two invariant tests, both red on origin/main: the real discovery path
  yields `status: lazy` with the cached tool count and a summary without
  `failed` (eager control stays `configured`, live control stays
  `connected`); a lazy-only run neither warns nor re-arms the startup
  retry, while a configured-only run still does.
- Document the per-server `lazy` key (undocumented until now) in
  `cli-config.yaml.example`, the MCP config reference and the MCP guide.
2026-09-15 19:06:54 -07:00
teknium1
5a4c3b32d0 fix(tui): track the durable session id in ui state instead of on the replaceable info object
Agent-less producers (session.activate of a lazily-resumed session via
_fallback_session_info, session.info events from agent-less cwd switches)
replace state.info with payloads that omit stored_session_id, so the exit
handler's recovery target went stale/null after such a switch. Keep the
durable id in a dedicated ui.storedSid field written at every sid transition
(create/resume/activate), read it from there in the exit handler, and when a
session.info payload omits it, carry the tracked id forward onto info.
2026-09-15 18:53:05 -07:00
teknium1
4ccbc2936d test(tui): events keep flowing and backoff grows across gateway reconnects
Invariant for #111594: after drain() on mount, every later transport
generation still emits gateway.ready live, and the reconnect delay grows
across consecutive failures instead of restarting at the base delay.
2026-09-15 18:53:05 -07:00
gustavosmendes
583dbb534c fix(tui): preserve sessions across gateway reconnects
An attached (dashboard-embedded) Ink TUI whose WebSocket dropped never
recovered even though the backend stayed alive, and a spawned gateway that
crashed resumed the wrong session.

- GatewayClient no longer resets `subscribed` on each transport generation:
  the renderer drain()s once on mount, so every post-reconnect event
  (gateway.ready included) stayed buffered forever.
- clearReconnect() keeps the attempt counter; it is reset on gateway.ready
  (and kill()), so backoff actually grows across failed reconnects.
- useMainApp's exit handler no longer calls start() for an attached socket
  closure — GatewayClient owns that reconnect; it only respawns a still-owned
  child, and plans the resume with the durable stored_session_id (what
  session.resume takes) instead of the process-local runtime sid.
- session.create's stored_session_id is carried into ui state / the active
  session file so recovery and the exit epilogue target the durable id.
- The recovery target is cleared only after resumeById resolves into a live
  sid, so a second disconnect during setup/history loading keeps it.
- Stale-socket identity guards on 'open'/'message'.

Salvaged from #111599 (@gustavosmendes) with trims: kept the client-side
backoff reconnect for spawned children too (the "keeps trying to reconnect
in the background" copy depends on it), kept the RPC-triggered reconnect
during backoff, kept the spawn-mode "reply in progress was lost" wording
(true for a dead child) and added attached-mode copy in userMessages.ts,
dropped the config_warning contract regen and the SessionCreateResponse
re-export (local type gains stored_session_id instead).

Fixes #111594
2026-09-15 18:53:05 -07:00
teknium1
81805a97ef fix(tui): drop a pending key-burst flush when an external value replaces the draft
The composer hands key bursts to the parent on a 16ms timer. When
history navigation, a slash completion or a submit clears/replaces the
draft while such a flush is still armed, the [value] effect reset the
local buffer correctly but left the timer running, so 16ms later the
stale burst was handed to the parent and overwrote the external value
(second hole in #111934).

Disarm the pending flush in the external branch of the [value] effect:
once the parent has replaced the draft, a burst typed against the old
draft can never be the newer value. Second invariant test covers it
alongside the stale own-echo case.
2026-09-15 18:50:50 -07:00
liuhao1024
c9fe50f39d fix(tui): keep stale own-echo flushes from rewinding composer keystrokes
A deferred key-burst flush can still be in flight when the parent's
re-render lands: the echoed value is the one we emitted, older than
vRef because the user typed past it. The [value] effect treated any
non-equal incoming value as an external assignment and rewound local
state — the cursor jumped backward and freshly typed letters were
overwritten (#111934).

Track the last value handed to onChange; an echo matching it stays on
the own-change path, and the pending flush for the newer local value
converges the parent on its next timer.
2026-09-15 18:50:50 -07:00
teknium1
66878996dd fix(ux): plain-language, actionable user-facing messages (desktop-tui)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 03:46:45 -07:00
teknium1
b67441309c fix(contracts): SessionLiveInfo model/tools/skills stay optional — lazy and mirror paths emit session.info without them
The strict suite showed 41 emit sites sending {model} or {} alone; the TUI
banner coerces the missing maps instead of the contract lying about them.
2026-09-14 06:12:19 -07:00
teknium1
f6306d1920 feat(contracts): TypeScript consumes the generated contract; hand-typed wire shapes deleted
apps/shared/src/gateway-events.ts is now a thin layer over
gateway-contract.generated.ts (client-local synthetic events + the
GatewayEvent envelope); gateway-events.json, its two rendezvous tests and
the duplicated BillingBlock / SessionInfo / ProjectInfo hand copies are
gone. Desktop, TUI, web and shared typecheck against the generated
RpcMethods / ServerRequestMap / BackendGatewayEventMap.

What tsc found once the types were honest: three phantom fields the
backend never sent (tool.start.todos, error.reason,
voice.transcript.voice_stopped) - the TUI todo tests were driving the
list through the phantom and are retargeted to tool.complete, where the
wire actually carries it; nullable fields (`None` on the wire) were typed
as plain optionals in eight places and now coerce at the boundary;
SessionResumeResult had a stale generic.

Contract fixes from the consumer pass: TranscriptMessage is the gateway
projection (text/row_id/context/args), not the stored row; SkinPayload
matches HermesSkin (empty-string defaults, never null); SessionLiveInfo
model/tools/skills are required (always emitted); BillingBlock.billing_url
is required-nullable (dataclass asdict).

tui_gateway/AGENTS.md documents the declare -> regenerate -> tsc loop.
2026-09-14 06:12:19 -07:00
teknium1
9f7f2f28c0 feat(gateway): server→client JSON-RPC requests replace the *.request/*.respond event pairs (#110521)
The gateway asked the user questions (approval, clarify, sudo, secret,
vault, MCP setup, the desktop read/act bridges) by emitting a
`<x>.request` EVENT carrying a hand-minted request_id, blocking the
agent thread on a module dict keyed by that id, and exposing a paired
`<x>.respond` METHOD per kind — thirteen pairs, four registries
(`_pending`, `_answers`, `_batch_clarify`, `_EXPIRING_REQUESTS`) and a
per-kind reconnect snapshot (`pending_clarify` / `pending_approval`)
that only two of the thirteen kinds ever got. JSON-RPC already has the
primitive: the server sends a request frame with an id and the client
answers with a response frame bearing the same id.

`tui_gateway/server_requests.py` owns the one mechanism:

  send()          block the agent thread until the response frame
                  (`srq-<n>` ids; ints belong to the client)
  send_async()    fire-and-callback variant (bot relay)
  cancel*()       withdraw with ONE `request.cancel {id, method, reason}`
                  event (timeout / interrupt / process exit /
                  answered elsewhere) instead of per-kind *.expire
  open_requests() the still-open frames, replayed by session.resume,
                  session.activate and session.events.since so a
                  reconnecting client re-renders every kind, not two
  clarify.lock    stays a real client→server RPC (locks one batch
                  answer early); locked answers merge into the final
                  set even when the closing response carries only the
                  tail the user answered last

A client that does not implement a method answers -32601 and the agent
fails fast (the old fixed-timeout "unavailable" probes for tour/preview
still work — a wire error IS an answer). Approval: the queue entry's
settle hook withdraws the request when `/approve` from another surface,
a timeout or an interrupt resolves it first, so no window keeps a dead
card. Compute-host children own their waits; the parent mirrors their
open frames for replay and relays `clarify.lock` + response frames.

Clients: `JsonRpcRequestChannel` gains `onRequest` (unhandled → -32601,
dedup by id) and `JsonRpcGatewayClient` re-delivers `open_requests`
from the replay result. Desktop gets `gateway-event/server-requests.ts`
(one handler per method, replacing the request branches of
`input-requests.ts` / `desktop-bridge.ts`) and a `store/server-requests`
registry so every answer site calls `respondToServerRequest(id, result)`
synchronously; the TUI gets `createServerRequestHandler.ts` +
`serverRequestStore.ts`. `gateway-events.json` now pins both halves
(events + server request methods); the two contract tests check both.

Live (real stdio gateway, real `clarify_callback` on the agent thread):
before, `clarify.request` event + `clarify.respond` RPC, batch final
answers lost ('' returned); after, `{"id":"srq-…","method":"clarify"}`
frame, `session.events.since.open_requests` replays it, response frame
`{"answer":"yes"}` reaches the agent, batch lock + final response
merge to `{"q0":"1","q1":"free text"}`.
2026-09-14 06:02:05 -07:00
teknium1
ebe8cda8ea feat(tui_gateway): real JSON-RPC server→client requests replace the *.request / *.respond notification pair
The backend never sent a JSON-RPC request; when it needed an answer from the
renderer it hand-correlated a `*.request` notification with a later `*.respond`
method through four module-level dicts, a timeout thread and 13 derived
`*.expire` names, plus a separate reconnect snapshot per prompt kind. That is a
second request/response layer built on a protocol that already has one.

`tui_gateway/server_requests.py` sends `{id: "srq-…", method, params}` and
blocks on the response frame with that id (string ids never collide with the
clients' integer ids). One `request.cancel {id, method, reason}` notification
withdraws a request on timeout / interrupt / session close. `open_requests` on
`session.resume` / `session.activate` / `session.events.since` re-delivers
unanswered requests after a reconnect; the shared TypeScript channel does that
itself before the caller sees the result. Batch clarify keeps its per-question
locks as a normal `clarify.lock` RPC (the last lock resolves the request).
Approvals stay queue-backed (`tools.approval` owns the timeout, `/approve all`,
coalescing): the request resolves the queue entry and the entry's own
resolution withdraws the request through `register_gateway_settle`.

Deleted: `_block`, `_respond`, `_pending`, `_answers`,
`_pending_prompt_payloads`, `_batch_clarify`, `_EXPIRING_REQUESTS`, the
`*.respond` methods, every `*.request` / `*.expire` event, `pending_clarify`.
Compute-host (turn isolation) mirrors the child's open request and relays the
response frame / lock to it. Desktop, TUI and shared clients register
`onRequest` handlers where they used to switch on `*.request` events; answers
are response frames over the socket the request arrived on, so #91684's
owner-routing class cannot recur for prompts.
2026-09-14 06:02:05 -07:00
teknium1
2c0bec33f9 feat(model-pickers): reasoning effort selection on every model picker
The Desktop composer got a reasoning-effort pill this morning; every other place a
model is picked still left the effort to a separate command (`/reasoning`) or a
hand edit of config.yaml. `hermes model` had one effort step for Copilot only, and
its auxiliary-model menu had none at all even though every aux block already reads
`auxiliary.<task>.reasoning_effort`.

One request now carries a model pick AND its effort on every surface:

- `hermes_cli/model_switch.py`: the single `/model` parser accepts `--reasoning
  <level>` (validated against `parse_reasoning_effort`; unknown level ->
  `MODEL_SWITCH_ERR_BAD_REASONING`; Unicode-dash normalized like the other flags).
  `ModelSwitchRequest.reasoning_effort` rides with the pick.
- Classic CLI (`cli_model_switch_mixin`, `cli_tui_mixin`): `/model X --reasoning
  high` applies the effort AFTER the agent swap (`switch_model` re-resolves
  `reasoning_config` from config.yaml, so an earlier write is clobbered) with the
  pick's scope (session; config on `--global`; `--once` snapshots and restores it).
  The `/model` picker gains a third stage, "Reasoning effort for <model>", built
  from `VALID_REASONING_EFFORTS` + none + "Keep current effort"; hidden when the
  inventory capability map says the route has no reasoning control.
- TUI gateway (`tui_gateway/model_switch.py`, serves Ink TUI + Desktop):
  `config.set model "X --reasoning high"` applies after the swap; session pin
  (`create_reasoning_override`) by default, `agent.reasoning_effort` on --global,
  one-turn restore carries `reasoning_config`; re-emits `session_info` so the
  status bar shows the new effort.
- Ink TUI `ModelPicker`: step 3/3 (same rows, same capability gate) emitting
  `<model> --provider <slug> --reasoning <level> <scope>`; the new-session draft
  label strips the flag like `--provider`.
- Messaging gateway `/model`: `--reasoning` goes through the existing
  `_apply_reasoning_selection` (the `/reasoning` applier) with the pick's scope.
- `hermes model`: one shared post-pick effort step for the MAIN model (replaces
  the Copilot-only inline prompt; Copilot keeps its per-model level set via
  `github_model_reasoning_efforts`, other routes get the ladder, catalog
  `supports_reasoning=False` skips it) plus a "Reasoning effort for the current
  model..." row. The auxiliary menu's provider->model and custom-endpoint flows end
  with the same step (+ "Provider default"), stored as
  `auxiliary.<task>.reasoning_effort` / `delegation.reasoning_effort`, shown in
  the task list ("openrouter · model · high"), cleared by "Reset all to auto";
  tasks whose block omits the key by design (MoA slots, memory_query_rewrite) skip
  it.

Live (temp HERMES_HOME, stub key, no model call):
- `hermes model` -> aux -> Vision -> OpenRouter -> model: before ends at
  "Vision: openrouter · <m>", no key written; after adds "Select reasoning effort"
  and saves `reasoning_effort: high`.
- `hermes model` -> DeepSeek -> model: before no effort step; after the step
  writes `agent.reasoning_effort: xhigh`.
- tui_gateway stdio: `config.set model "... --reasoning high --session"` before
  errors "Model names cannot contain spaces"; after switches and `config.get
  reasoning` returns high; bad level -> the canonical error text.
- classic CLI `process_command`: before the same spaces error; after "Reasoning
  effort: high" under the switch summary, `--global` writes config.
- `hermes --tui` PTY: /model -> step 1/3 -> 2/3 -> 3/3 -> high; transcript
  "reasoning: high", status bar "fable 5.1 high".
2026-09-13 16:43:50 -07:00
teknium1
f1d5c99fe5 feat: background-process completions paint a compact title, not the raw notification wall
Subagent completions already got this: the model receives the full
`[ASYNC DELEGATION …]` text while the CLI/TUI/Desktop paint a one-line
"Subagent Task Completed: <goal>" event. Background-process completions
(`terminal(background=True, notify=True)`) still echoed the entire
`[IMPORTANT: Background process proc_… completed normally (exit code 0).
Command: … Output: …]` block as if the user had typed it.

Generalise the delegation mechanism: `TimelineNotification` (formerly
`SubagentNotification`) carries `display_kind` + `display_text`;
`ProcessNotificationBatch` renders a `process_complete` one with a
`process_completion_display_text` title ("Background Process Finished:
<cmd>", "Background Process Failed (exit 1): <cmd>", "N Background
Processes Finished"). The TUI gateway stamps the same kind/metadata on
the synthesized turn and emits the title on `status.update`; Ink and
Desktop project `process_complete` rows as timeline events (Desktop keeps
the raw output behind the existing expandable async-result row). Model
content is byte-identical to before.
2026-09-13 14:35:32 -07:00
hermes-seaeye[bot]
61304d5923 fmt(js): npm run fix on merge (#110132)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-13 17:58:46 +00:00
teknium1
988d471479 style(ts): sort imports/exports the way perfectionist wants after the rebase 2026-09-13 06:50:57 -07:00
teknium1
c764d6d354 fix(shared): ensureContrast keeps the desktop's 0.2-step ladder; TUI chain opts into 0.05
The shared ensureContrast shipped the TUI's fine 0.05×20 ladder, which
changed --dt-primary-solid for 7 of 15 desktop presets (nous #3b6acb →
#3f70d8, cyberpunk #00661a → #008021, slate #505457 → #6f7377) while the PR
body said no preset VALUE changed. The ladder is now the desktop's original
algorithm exactly — pole by luminance < 0.5, accumulating 0.2 steps up to
1.0001, re-mixed from the source colour — with `step` as a parameter. The
only pre-refactor TUI caller (ColorChain.ensureContrast) passes 0.05, so
the terminal palette is byte-identical too.

Test: apps/desktop context.test.tsx iterates every builtin preset × mode,
paints it through ThemeProvider and asserts --dt-primary-solid equals the
value a reference copy of the old desktop algorithm computes. Sabotage
(default step 0.05): 11/30 rows fail. Docs: the SDK table now lists
contrastRatio as `number | null` under sRGB measures, not OKLCH.
2026-09-13 06:50:57 -07:00
teknium1
057c2c85fc refactor(slash): delete the dead web slash re-implementation; one slash parser + command.dispatch narrowing in @hermes/shared
web/src/lib/slashExec.ts and web/src/components/SlashPopover.tsx had zero
importers since the React composer was replaced by the PTY-embedded TUI
(f49afd3122) — exactly what web/AGENTS.md forbids, now orphaned. Their
parseSlash still carried the `(.*)` newline bug and lacked the `prefill`
variant. Desktop and the TUI each hand-rolled the same slash split and the
same command.dispatch narrowing; the multi-line fix (#41323, #55510) had to
be applied to each copy separately.

Sites:
  web/src/lib/slashExec.ts::executeSlash/parseSlash/parseCommandDispatch  -> deleted
  web/src/components/SlashPopover.tsx::SlashPopover                        -> deleted
  apps/desktop/src/lib/chat-runtime.ts::parseSlashCommand                  -> apps/shared/src/slash.ts::parseSlashCommand
  apps/desktop/src/lib/chat-runtime.ts::parseCommandDispatch               -> apps/shared/src/slash.ts::parseCommandDispatch
  apps/desktop/src/lib/chat-runtime.ts::SLASH_COMMAND_RE                   -> apps/shared/src/slash.ts::SLASH_COMMAND_RE
  apps/desktop/src/app/types.ts::*CommandDispatchResponse (5 interfaces)   -> apps/shared/src/slash.ts
  ui-tui/src/domain/slash.ts::parseSlashCommand/looksLikeSlashCommand      -> apps/shared/src/slash.ts
  ui-tui/src/lib/rpc.ts::asCommandDispatch                                 -> apps/shared/src/slash.ts::parseCommandDispatch
  ui-tui/src/gatewayTypes.ts::CommandDispatchResponse                      -> apps/shared/src/slash.ts
  9 desktop importers + 3 TUI importers repointed.

Behavior change: desktop `parseSlashCommand` now lower-cases the command
name like the TUI, backend `resolve_command` and `slash.exec` already do
(`/Help` resolved before via the case-insensitive backend; local desktop
action lookups were case-sensitive). TUI's parsed result no longer carries
the redundant `cmd` echo (no consumer read it).

Tests: apps/shared/src/slash.test.ts (parseSlashCommand multi-line /
newline-boundary / degenerate cases; parseCommandDispatch every variant +
malformed rejection). Sabotage: restoring `(.*)` in SLASH_PARTS_RE fails
2 tests; restored -> 7 pass. Desktop chat-runtime.test.ts and TUI
asCommandDispatch.test.ts cases moved here; slashParity.test.ts repointed.
2026-09-13 06:50:57 -07:00
teknium1
35022e02ed refactor(themes): one sRGB color-math module in @hermes/shared; measured readableOn + fine ensureContrast ladder on both surfaces
ui-tui/src/lib/color.ts called itself "the twin of the desktop app's
src/themes/color.ts" and the two had already drifted: the desktop measured
readableOn but used a coarse 0.2x5 ensureContrast ladder and returned 0 for
unparseable luminance; the TUI had the fine 0.05x20 ladder and null-for-garbage
but a luminance>0.5 threshold readableOn. Both now import the primitives from
apps/shared/src/color.ts (`@hermes/shared/color`, also exported from the root
index); each surface keeps only what is specific to it. No palette / preset /
skin VALUE changes anywhere — only math.

Sites (path::symbol → canonical):
  apps/desktop/src/themes/color.ts::hexToRgb           → @hermes/shared/color::parseColor (deleted)
  apps/desktop/src/themes/color.ts::rgbToHex           → @hermes/shared/color::toHex (deleted)
  apps/desktop/src/themes/color.ts::mix                → @hermes/shared/color::mix
  apps/desktop/src/themes/color.ts::relativeLuminance  → @hermes/shared/color::relativeLuminance
  apps/desktop/src/themes/color.ts::contrastRatio      → @hermes/shared/color::contrastRatio
  apps/desktop/src/themes/color.ts::readableOn         → @hermes/shared/color::readableOn (desktop wrapper readableInk pins ['#161616','#ffffff'])
  apps/desktop/src/themes/color.ts::ensureContrast     → @hermes/shared/color::ensureContrast
  ui-tui/src/lib/color.ts::{Rgb,parseColor,toHex,mix,relativeLuminance,contrastRatio,readableOn,ensureContrast,lighten,darken}
                                                       → @hermes/shared/color (same names)
  Stays desktop-only (apps/desktop/src/themes/color.ts): luminance, normalizeHex, readableInk, OKLCH set
    (hexToOklch, oklchToHex, oklchToSrgb255, maxChroma, hueDelta, harmonize, mixOklab, withHue, ensureContrastOklch).
  Stays TUI-only (ui-tui/src/lib/color.ts): liftForContrast, grayOf, desaturate, toHsl, fromHsl, retone,
    boostSaturation, color()/ColorChain.
  Importers repointed (17): apps/desktop/src/{sdk/index.ts, themes/context.tsx, themes/retint.ts,
    themes/retint.test.ts, themes/skin.ts, themes/vscode.ts, themes/vscode.test.ts};
    ui-tui/src/{theme.ts, sdk/index.ts, sdk/apps/weather.tsx, app/createGatewayEventHandler.ts,
    components/agentsPanel.tsx, components/branding.tsx, components/loaders.tsx,
    components/overlayPrimitives.tsx, lib/color.ts, lib/color.test.ts}.
  Wiring: apps/shared/package.json exports './color'; apps/shared/src/index.ts re-exports;
    apps/desktop/tsconfig.json paths + vite.config.ts alias for '@hermes/shared/color'
    (ui-tui resolves the subpath via the workspace package exports, like './billing').

Behavior change (1): relativeLuminance / contrastRatio return null for unparseable
  input on the desktop too (previously 0, which made garbage measure like pure
  black). Desktop SDK export `contrastRatio` therefore widens to `number | null`.
  Only ensureContrastOklch relied on the number: it now treats null as "already
  passing / can't measure" and returns the input unchanged. Every other desktop
  caller passes 6-digit hex.

Behavior change (2): readableOn MEASURES both candidate inks and returns the one
  with the higher contrast ratio (desktop semantics; the threshold version got
  mid-lightness accents wrong: white on #4f9e5e is 3.29:1 vs near-black 5.50:1).
  Signature is readableOn(bg, inks = ['#000000', '#ffffff']); the desktop passes
  its own pair via `readableInk` so desktop output is byte-identical. The TUI
  switches from the luminance>0.5 threshold to measurement: over the 185 distinct
  hexes in ui-tui/src/theme.ts (DARK/LIGHT seeds + built palettes) and
  hermes_cli/skin_engine.py, 56 flip from '#ffffff' to '#000000' — all
  mid-lightness accents (L 0.18–0.49, e.g. #cd7f32, #4caf50, #ef5350, #ffa726,
  #4dabf7) where black measures 4.6–10.8:1 against white's 1.9–4.5:1. Note the
  TUI never called readableOn directly; it only reaches ensureContrast's pole
  choice (below), and ensureContrast is only reachable via the color() chain and
  the theme.ts re-export (no production caller today).

Behavior change (3): ensureContrast steps 0.05 x 20 from the ORIGINAL color toward
  the measured readableOn pole (TUI semantics). The desktop previously stepped
  0.2 x 5 toward a threshold-chosen pole, so desktop-derived accents that needed a
  lift (skin/VS Code imports whose accent fails 4.5:1 on the sidebar, and
  --dt-primary-solid) may now land up to 0.15 closer to their original hue —
  they stop at the first passing rung. Palette VALUES are unchanged; only
  synthesized colors move.

Also: parseColor accepts #rgb shorthand where desktop hexToRgb rejected it —
  strictly more permissive; the only desktop path fed raw user hex is
  normalizeHex, which already expands shorthand itself.

Tests: apps/shared/src/color.test.ts (moved TUI parse/mix/contrast cases +
  two invariants):
  - "readableOn(%s) returns the ink with the higher measured contrast" — computes
    contrastRatio for each candidate in the test and asserts the returned ink is
    the max (a contract, not a hardcoded hex) over #4f9e5e (both ink pairs),
    #cba6f7, #ffffff, #101014.
    Sabotage: reverted readableOn to the luminance threshold → 3 red
    (#4f9e5e x2, #cba6f7); restored → green.
  - "ensureContrast(%s on %s) clears %s" — 5 failing pairs end ≥ min; plus
    "leaves passing and unparseable colors byte-identical".
    Sabotage: truncated the ladder to 3 rungs → 5 red; restored → green.
  ui-tui/src/lib/color.test.ts keeps only the color() chain case.

Validation:
  apps/shared: npx tsc -p . --noEmit (0) && npx vitest run → 3 files, 30 tests passed; npm run lint clean
  apps/desktop: npx tsc -p . --noEmit (0); npx vitest run --project ui → 798/800 files, 7563/7572 tests;
    the 9 failures (src/app/messaging/index.test.tsx x8 12s-timeouts, src/lib/markdown-blocks.test.ts
    property fuzz 36s) are load-induced flakes under the full parallel run: both files pass in
    isolation on this branch (16/16) and on origin/main; neither imports color math. npm run lint 0 errors
  ui-tui: npm run build:ink; npx tsc -p . --noEmit (0) && npx vitest run → 168 files, 1764 tests passed; npm run lint 0 errors
  git diff --check clean; no new gitignored .d.ts.

Handoff: desktop vs web preset palettes diverge for the four shared ids
  (web presets carry a 3-slot palette {background, midground, foreground(alpha 0)}
  + warmGlow, not the desktop's 24-slot set, so only the comparable slots are
  listed; web `foreground` is #ffffff alpha 0 on all four — a glow/overlay
  slot, not text ink). Design call for Teknium; nothing changed here.

    preset     slot        desktop                     web
    cyberpunk  background  #000a00                     #040608
    cyberpunk  accent      #00ff41 (primary/ring/mid)  #9bffcf (midground)
    cyberpunk  foreground  #00ff41                     #ffffff (alpha 0)
    ember      background  #160800                     #1a0a06
    ember      accent      #d97316 (ring/midground)    #ffd8b0 (midground = desktop fg/primary)
    ember      foreground  #ffd8b0                     #ffffff (alpha 0)
    midnight   background  #08081c                     #0a0a1f
    midnight   accent      #8b80e8 (ring/midground)    #d4c8ff (midground)
    midnight   foreground  #ddd6ff                     #ffffff (alpha 0)
    mono       background  #0e0e0e                     #0e0e0e  (match)
    mono       accent      #9a9a9a (ring/midground)    #eaeaea (midground = desktop fg/primary)
    mono       foreground  #eaeaea                     #ffffff (alpha 0)
2026-09-13 06:50:57 -07:00
teknium1
a3d259019b refactor(ts): one stripAnsi in @hermes/shared (TUI's OSC/DCS/partial-CSI coverage); desktop adopts it
Three TS surfaces each carried their own ANSI stripper with different
coverage. The TUI's (OSC, DCS/SOS/PM/APC strings, complete and truncated
CSI, multi-byte non-CSI ESC sequences, stray ESC, C0 controls) is now the
single implementation at apps/shared/src/ansi.ts, exported from the root
index and the new `@hermes/shared/ansi` subpath (ui-tui has no DOM lib, so
it imports the subpath like it does for billing/skin).

Sites (path::symbol → canonical):
  ui-tui/src/lib/text.ts::stripAnsi, sanitizeAnsiForRender, hasAnsi
      → moved to apps/shared/src/ansi.ts (text.ts now imports stripAnsi
        from '@hermes/shared/ansi' for its own trail helpers)
  ui-tui: 13 importers repointed from '../lib/text.js' to
      '@hermes/shared/ansi' (createGatewayEventHandler.ts,
      components/messageLine.tsx, 11 __tests__ files)
  apps/desktop/src/lib/ansi.ts::stripAnsi (2 regexes) → deleted;
      parseAnsi/ansiColorClass/hasAnsiCodes stay (styled-segment parser)
  apps/desktop/src/app/session/hooks/use-prompt-actions/index.ts
      → imports stripAnsi from '@hermes/shared/ansi'
  apps/desktop/src/components/assistant-ui/tool/fallback-model/index.ts
      private SGR-only stripAnsi → deleted; imports the shared one

Tests: the TUI 'ANSI sanitizers' cases move from
ui-tui/src/__tests__/text.test.ts to apps/shared/src/ansi.test.ts, plus
one invariant: an OSC-8 hyperlink + DCS string + SGR + partial CSI tail
strips to exactly the visible text with no ESC/BEL left.

Behavior change: desktop chat system messages (use-prompt-actions) and
inline-diff chrome (stripInlineDiffChrome) now also lose OSC hyperlink
payloads, DCS strings, truncated CSI tails and C0 control bytes that the
weaker regexes let through. TUI behavior is unchanged.
2026-09-13 06:50:57 -07:00
teknium1
172b2a722b refactor(ts): one compactNumber and one reasoning-effort value set in @hermes/shared
Three hand-rolled compact-number formatters and two mirrored copies of the
reasoning-effort value set collapse into apps/shared/src/format.ts and
apps/shared/src/reasoning-effort.ts, exported from the package root and as
the subpaths `@hermes/shared/format` / `@hermes/shared/reasoning-effort`
(the TUI compiles with lib ES2023 and imports subpaths only). Surfaces keep
their own label maps and UI helpers. No re-export shims remain.

Convention for compactNumber (desktop's implementation, moved verbatim):
lowercase 'k', uppercase 'M', promotion-guarded thresholds (>= 999.5 -> k,
>= 999_950 -> M) so rounding can never print "1000k", trailing ".0"
stripped, non-finite / <= 0 -> "0".

Sites (path::symbol -> canonical):

  apps/desktop/src/lib/format.ts::compactNumber          -> apps/shared/src/format.ts::compactNumber (moved; file deleted)
  web/src/lib/format.ts::formatTokenCount                -> deleted
  ui-tui/src/lib/text.ts::fmtK                           -> deleted (text.ts's own callers use compactNumber)
  apps/desktop/src/app/agents/index.tsx                  -> @hermes/shared
  apps/desktop/src/app/chat/sidebar/chrome.tsx           -> @hermes/shared
  apps/desktop/src/app/chat/sidebar/session-row.tsx      -> @hermes/shared
  apps/desktop/src/app/command-center/index.tsx          -> @hermes/shared
  apps/desktop/src/app/shell/context-usage-panel.tsx     -> @hermes/shared
  apps/desktop/src/app/shell/titlebar-controls.tsx       -> @hermes/shared
  apps/desktop/src/app/skills/index.tsx                  -> @hermes/shared
  apps/desktop/src/app/skills/mcp-tab.tsx                -> @hermes/shared
  apps/desktop/src/components/ui/tab-dropdown.tsx        -> @hermes/shared
  apps/desktop/src/lib/statusbar.tsx                     -> @hermes/shared
  apps/desktop/src/sdk/index.ts::compactNumber           -> re-exported from @hermes/shared (plugin SDK surface unchanged)
  apps/desktop/src/plugins/kanban/{board,drawer}.tsx     -> unchanged (import via @hermes/plugin-sdk)
  web/src/components/ModelInfoCard.tsx::formatTokenCount -> @hermes/shared::compactNumber
  web/src/pages/ModelsPage.tsx::formatTokenCount         -> @hermes/shared::compactNumber
  ui-tui/src/components/appChrome.tsx::fmtK              -> @hermes/shared/format::compactNumber
  ui-tui/src/components/thinking.tsx::fmtK               -> @hermes/shared/format::compactNumber
  ui-tui/src/app/slash/commands/session.ts::fmtK         -> @hermes/shared/format::compactNumber
  ui-tui/src/__tests__/text.test.ts::fmtK suite          -> apps/shared/src/format.test.ts (table incl. promotion guard)

  apps/desktop/src/lib/reasoning-effort.ts::REASONING_EFFORTS/REASONING_EFFORT_VALUES/
      DEFAULT_REASONING_EFFORT/ReasoningEffort/isReasoningEffort  -> apps/shared/src/reasoning-effort.ts
      (SHORT_LABELS, reasoningEffortLabel, isThinkingEnabled, resolveReasoningEffort stay local)
  apps/desktop/src/app/settings/constants.ts             -> @hermes/shared
  apps/desktop/src/app/settings/model-settings.tsx       -> @hermes/shared
  apps/desktop/src/app/shell/model-catalog-menu.tsx      -> @hermes/shared (+ local reasoningEffortLabel)
  apps/desktop/src/app/shell/model-edit-submenu.tsx      -> @hermes/shared (+ local UI helpers)
  apps/desktop/src/app/shell/model-menu-panel.tsx        -> @hermes/shared
  apps/desktop/src/lib/model-status-label.ts             -> @hermes/shared (+ local reasoningEffortLabel)
  apps/desktop/src/sdk/index.ts                          -> value set re-exported from @hermes/shared; label helper stays from '@/lib/reasoning-effort'
  apps/desktop/src/lib/reasoning-effort.test.ts          -> value-set + isReasoningEffort cases moved to apps/shared/src/reasoning-effort.test.ts
  web/src/lib/reasoning-effort.ts::EFFORT_OPTIONS        -> labels mapped over shared REASONING_EFFORT_VALUES (same order: none, then 7 levels)
  web/src/lib/reasoning-effort.ts::VALID_EFFORTS         -> Set(REASONING_EFFORT_VALUES); normalizeEffort falls back to DEFAULT_REASONING_EFFORT

Semantics kept: web `none` is selectable; desktop `none` resolves to ''
(thinking off); desktop isReasoningEffort still trims + lowercases.

Behavior change:
  - web: token counts on the Models page and ModelInfoCard now print a
    lowercase 'k' and are promotion-guarded: 128_000 "128K" -> "128k",
    999_999 "1000.0K" -> "1M", 1_500 "1.5K" -> "1.5k". 'M' is unchanged.
  - TUI: fmtK used Intl compact notation; compactNumber differs only in
    suffix case and the guard: 1_000_000 "1m" -> "1M", and billions no
    longer get a 'b' suffix (1_000_000_000 "1b" -> "1000M"). Sub-million
    values are identical ("999", "1k", "1.5k"). Non-positive values now
    print "0" instead of "-1k".
  - desktop: none (its formatter moved verbatim).

Tests: apps/shared/src/format.test.ts::"compactNumber" (table incl.
999_999 -> "1M", 999_949 -> "999.9k"; fails when the promotion guard is
removed) and apps/shared/src/reasoning-effort.test.ts::"reasoning-effort"
(no duplicate values, `none` is the only non-level, default is a member;
fails on a duplicated level or a `none`-accepting isReasoningEffort).
2026-09-13 06:50:57 -07:00
teknium1
a2ae8f229d refactor(ts): one fuzzy + model-search-text helper in @hermes/shared; desktop picker ranks with fuzzyRank
Three byte-identical (modulo prettier and a "keep in sync" header comment)
copies of model-search-text.ts and two of fuzzy.ts collapse into one copy
each under apps/shared/src, exported from the package root and as the
subpaths `@hermes/shared/fuzzy` / `@hermes/shared/model-search-text` (the
TUI compiles with lib ES2023 and imports subpaths, never the DOM-typed
root). The vitest suites move with the code; no re-export shims remain.

Sites (path::symbol -> canonical):

  ui-tui/src/lib/fuzzy.ts::fuzzyScore/fuzzyScoreMulti/fuzzyRank   -> apps/shared/src/fuzzy.ts (moved)
  web/src/lib/fuzzy.ts::fuzzyScore/fuzzyScoreMulti/fuzzyRank      -> deleted
  ui-tui/src/lib/model-search-text.ts::modelSearchText            -> apps/shared/src/model-search-text.ts (moved)
  web/src/lib/model-search-text.ts::modelSearchText               -> deleted
  apps/desktop/src/lib/model-search-text.ts::modelSearchText      -> deleted
  ui-tui/src/lib/fuzzy.test.ts                                    -> apps/shared/src/fuzzy.test.ts (moved)
  ui-tui/src/lib/model-search-text.test.ts                        -> apps/shared/src/model-search-text.test.ts (moved)
  ui-tui/src/components/modelPicker.tsx::fuzzyRank, modelSearchText     -> @hermes/shared/fuzzy, @hermes/shared/model-search-text
  web/src/components/ModelPickerDialog.tsx::fuzzyRank, modelSearchText  -> @hermes/shared
  web/src/lib/model-picker-filter.ts::fuzzyScoreMulti                   -> @hermes/shared
  apps/desktop/src/components/model-picker.tsx::modelSearchText         -> @hermes/shared (+ fuzzyRank, see below)

The header comment now names only the cross-language twin
(hermes_cli/model_search.py) as the thing to keep in sync.

Behavior change (desktop only): the desktop model picker used to filter
model rows with `foldIncludes` substring matching and keep the curated
order; it now ranks them with the same `fuzzyRank(models, query,
modelSearchText)` the web and TUI pickers use. What a user sees
differently while typing a query:

  - subsequence queries match: "g4o" now finds "gpt-4o" (previously only
    a literal substring such as "gpt-4" or "4o" matched);
  - the best match floats to the top instead of rows staying in curated
    order (exact > prefix > word-boundary > contiguous > scattered);
  - a query that matches the provider name/slug still shows that
    provider's full curated list in order, exactly as before;
  - an empty query still shows the curated list verbatim.

The in-row highlight is unchanged (substring emphasis via HighlightMatches),
so a fuzzy-only hit renders without emphasis rather than mis-highlighting.

Tests: apps/desktop/src/components/model-picker.test.tsx::"orders model
rows exactly as the shared fuzzyRank does" asserts the rendered row order
equals the shared fuzzyRank order for the same inputs (fails on both the
old substring filter and a reversed ranking).
2026-09-13 06:50:57 -07:00
teknium1
c1e0fd83f9 fix(shared): GatewayEventMap drops phantom keys and types child_session_id
Re-verified against the tui_gateway emitters:
- SubagentEventPayload.cost_usd / .iteration: not in
  tool_progress.py::_SUBAGENT_FIELDS, never emitted → removed; the TUI's
  turnController no longer copies them (its SubagentProgress keeps the
  fields for spawn-history persistence).
- SubagentEventPayload.child_session_id: emitted (in _SUBAGENT_FIELDS, read
  by agent_callbacks.py::_mirror_subagent_to_child) but untyped → added.
- ToolCompletePayload.error: _on_tool_complete never sets it → removed;
  the TUI's completeTool drops its dead `error` parameter and renders the
  trail line as non-error (which is what it always did on the wire).
- ToolStartPayload.todos: not on the wire either, but the TUI handler and
  its fixtures exercise recordTodos from tool.start; kept with a comment
  saying so rather than churning the handler.
- MessageCompletePayload.failure_reason: prompt_turn.py passes
  result.get("failure_reason") through → `string | null`.
2026-09-13 05:42:31 -07:00
teknium1
6b406f1c89 refactor(ts): ui-tui rides apps/shared's JSON-RPC request channel; one pending map, one heartbeat, typed RPC errors
Two independent JSON-RPC client cores existed for one backend: apps/shared's
JsonRpcGatewayClient (desktop, web) and ui-tui/src/gatewayClient.ts, which
re-implemented request ids, the pending map with timeouts, response->error
mapping, event decoding and the gateway.ping heartbeat (~200 LOC, drifted).

Split the transport-agnostic half out of the shared client into
JsonRpcRequestChannel (apps/shared/src/json-rpc-channel.ts): the owner binds a
JsonRpcTransport { send(text) } per connection generation and feeds inbound
text through handleFrame(). JsonRpcGatewayClient keeps only the WebSocket
lifecycle, seq replay and the typed event hub on top of it; the Ink TUI keeps
only its two transports (spawned child stdio, attached socket) and its
mount-order event buffering, and delegates everything else.

Behavior change:
- TUI RPC errors now carry the JSON-RPC `code` / `data` (JsonRpcGatewayError)
  instead of a bare Error(message); the TUI's timeout text is now the shared
  "request timed out after Ns: <method>" (was "timeout: <method>", matched by
  no caller) and callers may pass a per-call timeout.
- TUI heartbeat liveness counts any inbound frame (shared semantics) rather
  than tracking one in-flight ping id; the interval/deadline are unchanged
  and pings no longer carry the unread `last_activity_ms` param.
- Desktop isMissingRpcMethod reads the -32601 code first and only regexes the
  message for code-less (IPC-flattened) errors, so a tool result that merely
  mentions "unknown method" no longer reads as a capability verdict.
- Shared connect() now settles on a `close` during the handshake (auth-gate
  4401/4403) instead of waiting out the 15s connect timeout, and
  invalidate()/close() drop the socket generation before calling close() so a
  synchronous close event cannot run the closed-path twice.
2026-09-13 05:42:31 -07:00
teknium1
36773e0d78 refactor(ts): one GatewayEventMap in apps/shared typed from tui_gateway emitters; drop never-emitted tool.progress
Three TypeScript clients each declared their own copy of the tui_gateway wire
types and had drifted apart: apps/shared had a partial GatewayEventName union
with a `(string & {})` escape hatch, ui-tui/gatewayTypes.ts a 150-line
discriminated union, and apps/desktop an `RpcEvent<T>` that was field-for-field
the shared GatewayEvent with `type: string`. None matched the emitter:
message.complete lacked warning/status/error/recoverable/error_surface,
tool.start/tool.complete lacked args/result, SessionResumeResponse lacked
session_key/messages_omitted/hydrating/auto_continue/todo_state, three
different ModelOptionProvider shapes disagreed on fields, and all three unions
handled a `tool.progress` event that no Python emitter has ever produced.

Now:

* `apps/shared/src/gateway-events.ts` is the single home: payload interfaces
  typed from the Python emitters (file::symbol cited per interface),
  `BackendGatewayEventMap` (89 backend names) + `ClientLocalGatewayEventMap`
  (5 TUI-synthetic transport events, clearly marked, excluded from the
  contract) merged into `GatewayEventMap`; `GatewayEvent<K>` is discriminated
  on `type` with `seq` typed. RPC shapes shared by 2+ surfaces live beside it
  (ModelOptionProvider = union of every field hermes_cli/inventory.py sets,
  incl. pricing_pending/free_tier_pending; SessionResumeResponse<Info>;
  SessionListItem with resolved_id; Usage).
* `JsonRpcGatewayClient.on<K>` is keyed by event name; the gateway.ready
  heartbeat/replay_epoch and per-frame `seq` reads are typed instead of cast.
* ui-tui and apps/desktop import the shared names; their local duplicates are
  deleted (no re-export shims — importers are repointed; the desktop plugin
  SDK barrel keeps its public `RpcEvent` name as an alias of GatewayEvent).
  web/src repoints ModelOptionProvider/ModelOptionsResponse.
* `tool.progress` handling is removed from the TUI handler/turnController,
  desktop event sets/tools handler, shared union, tests, and two docs
  (`grep '"tool.progress"' tui_gateway/` = 0 hits; the `display.tool_progress`
  config mode is unrelated and untouched).
* `message.complete.warning` (history-commit note from
  prompt_turn.py::_complete_turn_payload) is typed and surfaced on both
  surfaces through their existing notice paths (TUI pushActivity 'warn',
  desktop notify kind 'warning').

Contract: `apps/shared/src/gateway-events.json` is the sorted list of
backend-emitted names. `tests/tui_gateway/test_gateway_event_contract.py`
collects names from the Python emitter side (emit-helper literals, the
`.request → .expire` table, change-watcher table, child delta mirror,
subagent relay, desktop_ui tool emitters, gateway.ready/setup.ready/
browser-controller frames) and asserts emitted == JSON in both directions.
`apps/shared/src/gateway-events.test.ts` asserts BACKEND_EVENT_NAMES (which
the map type is `satisfies`-checked against) == JSON. Sabotage-verified: a
fake JSON name fails both tests; a fake TS name fails tsc + vitest; a fake
Python `_emit("...")` fails pytest.
2026-09-13 05:42:31 -07:00
bixycler
70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
hermes-seaeye[bot]
6f8b8e77dd fmt(js): npm run fix on merge (#107545)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-10 17:41:19 +00:00
Teknium
b51da65258 fix: adapt execute_code cell authority to the widened prompt-callback table
_callback_api() now yields (getter, setter) pairs for every per-thread prompt
(approval, sudo, vault unlock); the kernel cell captured and restored the old
fixed 4-tuple. Iterate the table so a cell carries every callback and a future
addition needs no change here. Test recorder unpacks the new shape.

Also: perfectionist import order in ui-tui interfaces.ts (CI lint).
2026-09-10 10:35:07 -07:00
Teknium
ac33da3c73 fix(tui): hide the composer while the password-manager unlock card is open
Live Ink TUI repro: with the unlock card mounted, keystrokes reached BOTH the
masked prompt and the still-focused composer, so the master password echoed
in clear text in the composer row and was queued as a message. `$isBlocked`
(which unmounts the composer for approval/sudo/secret cards) did not list the
new overlay; the pet's awaiting-input predicate had the same gap. After the
fix the raw PTY stream no longer contains the typed password.

Also tightens the classic-CLI panel copy to fit an 80-column box.
2026-09-10 10:35:07 -07:00
Teknium
92e0de0ac4 feat(vault): Desktop, TUI and CLI surfaces for password-manager unlock
Desktop
- Settings → Credential Vault gains a "Password managers" section: per-manager
  toggle (disabled with a hint when the CLI isn't installed), Locked/Unlocked
  pill, Unlock (masked master-password dialog → vault.unlock) and Lock.
  Items from a manager show a source badge instead of a delete button.
- Mid-turn vault.unlock.request renders a masked card in the chat (same
  contract as the secret/sudo cards: dismiss = keep locked, late answers
  tolerated, blocks the composer, badges background sessions).
- i18n parity en/ar/ja/zh/zh-hant.

Ink TUI (hermes --tui): vault.unlock.request/expire overlay via MaskedPrompt;
Esc keeps the manager locked.

CLI: `hermes vault sources [--enable|--disable NAME]`; `hermes vault list`
shows the source column and names enabled-but-locked managers.

Docs: credential-vault.md covers managers, per-session unlock, and the
headless (cron/webhook/API/-q) no-prompt posture.
2026-09-10 10:35:07 -07:00
Siddharth Balyan
ae43fd6df6 fix(tui): bare URLs render verbatim instead of a derived site label (#106843)
The markdown renderer resolved every link to a label — an authored one when
present, otherwise the fetched HTML <title>, otherwise a slug derived from the
last path segment. For a bare URL that meant the target never reached the
screen: `Connect link: https://connect.example.com/link/lk_...` rendered as
`Connect link: <site name>`, so the connect-link handoff could not be read,
copied or retyped from the TUI.

A bare URL is now its own text. Authored markdown labels still win, because
`Link` emits the OSC 8 hyperlink unconditionally and the renderer records it
per cell, so a label never strands its target. Title resolution no longer
drives any render path.
2026-09-10 02:21:54 +05:30
hermes-seaeye[bot]
d4d4ecfae0 fmt(js): npm run fix on merge (#105739)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-08 10:47:44 +00:00