Commit Graph

799 Commits

Author SHA1 Message Date
teknium1
661d73f06b test+docs: trim doctor exit-status tests to two invariants; document the exit code
Keep the in-process return-code matrix (issues / manual / --fix full and
partial repair) and the real-process exit-status check; drop the exception
passthrough, --ack and --live cases, which pin behaviour this change does not
touch. Document 0/1 exit status under `hermes doctor` in the CLI reference.
2026-09-21 01:57:05 -07:00
teknium1
2520eb7c0d docs(backup): list the browser profile dirs hermes backup excludes 2026-09-21 01:44:42 -07:00
Siddharth Balyan
afc3b7c6f3 feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation

Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash
merges (e0ef0eb9c3, d105376b21, ee2f5629b8, 1ab32b212b): the branch's history
no longer shared a base with main, so this is the PR's exact delta against
+3762/-1470, identical to the branch tip 05d7f2d3d5.

The fifteen commits it carried, in order:

--- feat(connectors): the wire follows the connector contract (six-state status, required toolkit metadata, connectionId on mint)

Hermes types exactly what the contract page writes: /tmp/magic/CONTRACT-TOOLKIT-METADATA.resolved.md
(carriers: portal PR2 `sid/connection-api` for the enum and account rows, the portal contract branch
on top of it for the list metadata). No optional-for-compat fields, no fallback branch, no seven-state
word left in the tree. A gateway that does not speak this contract fails validation loudly.

Wire (tools/connectors/gateway/wire.py): `ConnectionStatus` is the six values, `pending` covering the
vendor's INITIALIZING and INITIATED; `ConnectorListItem` requires title, description, https iconUrl and
authKind, and carries activeConnectionId when the session binds an account; `ConnectorListResponse`
types a page whole with its total; `ConnectorAccount` and `ConnectorAccountsResponse` type the account
routes; a mint result names its connectionId on `initiated` and never on `failed`; CONNECTION_REQUIRED
carries connectUrl and connectionId together or not at all; `account` exists on the execute call and is
never sent while multi-account is off.

Client: `list_connectors(search, connected)` refuses a search under three characters before any
request and parses every page whole; `list_accounts` and `account_status` read the account routes,
404 `connection_not_found` is None and 429 raises `RateLimited(retry_after)`.

Watcher: the target keeps the account id the mint named and the toolkit's title and icon from one
list read before the card is emitted; `_status_for` is the one seam between the watcher and the
gateway's status source (the list walk today; the per-account route replaces that body only).
`pending` and a missing status move nothing; revoked and inactive are failed.

Desktop: `ConnectorRow` is the strict list item; `ConnectorRowSeed` is what a tool call's args can
say; `connection.request`/`update` targets carry `connection_id`, `title`, `icon_url`; the mark
ladder is glyph, vendor icon as a plain image, favicon, monogram.

Tests run red first: wire rejects the retired words and the missing metadata; client search
minimum, execute omits account, account_status 404/429; managed targets carry the metadata and the
account id, a pending row moves nothing; the logo ladder order; the store passes the fields through.

--- feat(connectors): onboarding connect-first rides the connection operation

The guided first build had its own connector surface: a 2 s connectors.list poll loop, an auto-open
of every minted link, a hidden "[setup] links opened" note that painted as a user bubble, CONNECT
FIRST rows in the transcript, and a "Start with N apps connected" composer pill that sent the go
signal early. Delete all of it. The build session's first manage_connections connect now shows the
same card as any chat, and the settled tool result is the go signal.

Deleted: store/first-build-connectors.ts, assistant-ui/first-build-connectors.tsx (nothing rendered
it any more), lib/first-build-start.ts, onboarding-chat/start.tsx, the first-build session markers in
handoff-receipt.ts, the connectionRows / latestConnectorPart helpers they were the last readers of,
and the five strings only they used.

The runbook (setup-profile.ts::connectFirstRunbook) now describes the operation's real contract: one
connect call with every picked slug, one card with a row per app, blocks until settled, a per-app
result of connected / skipped / not_connected. No wait, no "Start with", no "already active" branch
(the result never carried it), and no offer to re-mint: Try again and Continue live on the card, and a
Continue means the user moved on.

Tests (red first): the runbook names one connect call and its per-app result and never the retired
model actions; the no-account rule survives an empty pick.

--- feat(connectors): the registry is keyed by profile, every watch read is bounded by the deadline, connectionId is optional, connect and execute carry returnTo and op

Registry (P1-14): tools/connectors/live.py keys an operation by (profile home, session key). Two multiplexed
profiles can carry the same timestamp-based session key; the tool thread opens under its turn's profile
override and the RPC side names the session record's profile home, the default profile resolving to the
process home on both sides.

Watch reads (P1-8 residual): list_connectors takes a per-page timeout and the watcher bounds each read by
the operation's remaining deadline (floor one second), so a stalled gateway page cannot hold the operation
past its deadline.

Contract follow-up from the portal handoff (/tmp/magic/HANDOFF-HERMES-PR3-PORTAL-CONTRACT.md): connectionId
on a mint result and on CONNECTION_REQUIRED is plain Optional (a no-auth toolkit answers active with none; a
failed mint answers with neither link nor id); a target without one has nothing to watch. Connect and
execute requests carry returnTo (hermes-desktop | hermes-desktop-dev | portal) and op, so the vendor's done
page can send the browser back to the app; the call sites and the deep-link handler follow in the next
commit.

Tests, red first: two profiles share a session key without seeing each other; each watch read shrinks with
the deadline; the client honours a per-page timeout; the optional id; the two request fields. The fakes'
list_connectors accept the timeout keyword; two wrapper spies that did not were the cause of a 300 s hang
under the per-file harness.

--- feat(connectors): an MCP target runs on the shared connection operation

manage_connections ran MCP targets as a renderer errand: the desktop card ran
the catalog lookup, the install and the OAuth flow, then told the backend what
had happened, and every surface without a card got `unavailable` plus two
terminal commands. That left the outcome in the renderer's word, and left the
model unable to connect an MCP server anywhere but the desktop.

The backend now owns the work, the way it already owns a managed connector.
`prepare` starts the OAuth flow in process and points the browser at the
backend's own callback route; an install that still needs credentials waits
pending and publishes their names as `required_env`; `observe` reads the flow
or the install worker on every tick. The worker itself moved out of
tui_gateway/mcp_oauth_sessions.py into tools/connectors/mcp_oauth.py, and that
module now calls it, so the Capabilities tab keeps its RPC session table.

The card may only say approved, skipped or continue: a claim of any other
state moves nothing. With that, `renderer_flow` and the `unavailable` state
have no producer left and are gone from the contract. Off the desktop there is
no card, so the action runs at once and the result carries the authorization
URL for the user to open.

Try again reaches the same work: `_reissue` dispatches on `target.kind`, so an
MCP row re-runs its own install, enable or OAuth instead of minting a managed
gateway link.

(cherry picked from commit fa0438c4121807604e7c983ba42b901314a1305b)

--- feat(connectors): the MCP card projects the operation instead of running the flow

The MCP setup card used to own the work: it called the catalog, ran
installMcpCatalogEntry, polled the action, drove the OAuth window, and then told
the backend what had happened through connection.respond. That made the renderer
a second authority on target state, so an install that finished after the window
closed, or a card that never mounted, left the operation with a state nobody
could correct. PR3 moves that work to the backend, so the card has one job left:
show the operation and send the user's consent.

McpSetupPending now renders request.targets the way ConnectorOffer renders them:
one row per target, one verb by (action, state), Continue below the shell.
Install and Enable send {status: 'approved'} and nothing else; an authorize row
opens the link the backend minted, with no OAuth RPC of its own; a running
install holds its verb with the busy mark; a failed row sends
connectors.connect {reconnect: true} on the open operation, the same call the
managed card makes. The owner lookup and that call are now shared helpers on
connector-tool.tsx rather than a second copy here.

ConnectionTargetOutcome loses connected / initiated / failed: those were the
renderer reporting state, which it no longer may do. ConnectionActor loses
renderer_flow for the same reason. ConnectionOperationTarget gains required_env,
the credentials an MCP install is still waiting for; the card renders a field per
entry under the pending row and holds Install until every required one has text,
so the values travel with the approval instead of through a separate catalog call.

(cherry picked from commit 2e42ad19f63e41f274a87199d07eb99a7c995cbe)

--- fix(connectors): the toolkit metadata leaves the wire; the vendor logo is derived from the slug

Sid cancelled portal PR3 (the toolkit metadata on the list item) on 2026-09-14, so the fields the contract
commit typed as required come off the wire: no title, description, iconUrl, authKind or activeConnectionId
on the list item, no total on the page, no search or connected query, no title or icon_url on a target
snapshot. The list item is exactly what portal #1220 emits.

The desktop derives the vendor logo from the toolkit slug instead (`connectorIconUrl`,
https://logos.composio.dev/api/{slug}; the gateway slug is the vendor slug, checked for every lead-order
pick), rendered as a plain image because the host sends no CORS header. The title stays
`connectorTitle(slug)`. The mark ladder is unchanged: glyph, derived vendor icon, favicon, monogram.

Everything else in the contract stands: six-state status, optional connectionId on mint results and on
CONNECTION_REQUIRED, returnTo and op, the account routes. The contract commit's message still names the
metadata; this commit is the correction.

--- fix(connectors): an MCP target survives the card's Continue, and the model is not sent to an action that refuses it

The review of the MCP backend found eight defects; each one is a test first.

- The off-desktop note told the model to confirm with action 'status', which refuses MCP
  names outright. It now says the authorization finishes in the background and the tools
  arrive on the next turn.
- The card's existence is a property of the session. A missing callback no longer routes a
  desktop session down the off-desktop path, where the model would be handed a live link.
- A second failure of a backend attempt was published with actor 'user'. ``refresh`` takes
  the actor, so the frame says who produced the text.
- Continue on the RPC thread can settle the operation between any read of a target's state
  and the transition that follows it. A settled operation has a frozen result, so the lost
  move is dropped; an IllegalTransition no longer escapes into the tool result. The same
  rule covers a worker whose outcome arrives late and a Try again that arrives after Continue.
- Try again on an install carried no credentials, so the second install ran with an empty
  env. The approved map is kept on the runner (never on the target: target fields are
  serialised to the model) and reused.
- An approval that does not cover a required credential no longer reaches the worker, where
  ``install_entry``'s prompt would block on stdin forever; the row waits for the card's
  fields instead.
- ``enable`` writes ``mcp_servers`` under the scope and lock the dashboard's toggle route
  uses, so the two read-modify-write paths in one process cannot drop each other's write.
- Several authorize targets start their flows together and share one wait; one wait per
  target kept the card empty for minutes.
- The off-desktop operation is in no session's registry, so it now emits no connection.update.

Try again is also refused for a target the transition table cannot move back to 'initiated'
(an MCP target has no move out of 'expired') and for an operation that settled during the call.

The contract test that froze the Actor enum is replaced by two behaviour tests: no actor but
the backend watcher can connect an MCP target, and the card's word never moves one.

(cherry picked from commit 081c16412b48667758a77a7e7c29eaf2be6a93dd)

--- fix(connectors): the watcher reads one account route per target, and every frame carries its seq

The watch loop walked the whole toolkit list once per tick to learn whether one
target had connected. That read costs a vendor call per page, cannot tell one
account from another, and forced `awaiting_new_attempt`: after a forced
reconnect the list still reported the OLD account `active`, so the row had to be
disbelieved until it read as something else once. The gateway now serves
`GET /v1/connectors/accounts/{connectionId}`, so each pending target reads its
own account: the id the mint named, one read per target per tick at 1 Hz (the
route's bucket is 180/min per principal), each bounded by the operation's
remaining deadline. A 429 parks that one target until its Retry-After passes and
leaves the others reading. A 404 is "not yet" until the deadline. A target the
mint gave no account for has nothing to read, so it is not read. A forced
reconnect watches the new account, which is why `awaiting_new_attempt` and its
two tests are gone.

`mint` and `run_remote` now name where the browser should come back to
(`returnTo`, plus the operation id on a mint), so the vendor's done page can
hand the user back to the desktop app that asked instead of stranding them on a
web page. Only the desktop registers that URL scheme, so no other surface sends
either field.

`connection.update` frames were built by re-reading the operation after the lock
was released, so a second writer could give an older frame a newer state and the
renderer could not tell which frame was last. Every write now advances a
monotonic `seq` and takes its snapshot under the same lock, and the emitter sends
that snapshot; a renderer that keeps the highest seq per operation can drop a
frame that arrives out of order. `connectors.operation.wake` lets the desktop's
deep-link return ask for a read now instead of at the next tick; it only shortens
the wait and trusts nothing else in the link.

(cherry picked from commit f5423cf72302b92b26c043495184ad6426fa2b36)

--- fix(connectors): the desktop card follows the operation, and says so out loud

The MCP card was a second implementation of the connector card with the
review's defects: it painted for any request on the session, kept its
controls after the operation settled, offered an approve verb while the
backend was still minting an authorize link, and re-enabled Install when
the RPC returned rather than when the state frame moved the row. A second
click in that window sent the consent twice.

The renderer now reads the operation's `seq`: an update or a status frame
whose sequence is not greater than the one already applied changes
nothing, and a resume snapshot neither revives a settled card nor puts
back a row a newer frame has moved. Without it the transport's ordering
decided what the user saw.

`hermes://connections/done?op=…` brings the user back from the browser to
the session that opened the operation and wakes its watcher, so the row
moves at once instead of at the watcher's next tick. Only the op id is
used; the link's status moves no row.

Each row's mark and cue are one polite live region, so a row that flips is
heard and not only seen, focus follows the row the backend moved while the
card holds it, and the waiting mark stops spinning under reduced motion.

(cherry picked from commit 7c30d252556d7d496ddb354b50bcecc41bceef27)

--- fix(connectors): the renderer types seq as the wire carries it

The watcher commit made `seq` a required field on every operation frame. The
desktop store still declared its own optional `seq` so it would compile against a
backend that predates the field; that backend no longer exists on this branch, so
the hedge is dead code and the fixtures were short one field. The store now reads
`seq` from the shared types, and the fixtures count the way the backend does.

The repeat-frame test asserted the whole request keeps its reference. With a real
rising `seq` the request must change; the invariant the test guards is that the
target row keeps its identity so open credential inputs do not remount.

--- fix(connectors): a URL that arrives after the shared wait still lands on its row

The shared prepare wait failed every row still pending when it ran out, while
that row's own thread was still waiting on the provider. When the URL arrived a
moment later the thread's move raised inside the daemon thread, the link was
lost, and Try again started a second flow. The wait now bounds only how long
prepare blocks; a row still pending afterwards is left to its own thread, which
is the only writer of that row and ends with the URL or the flow's own failure.

The comment on the approved credentials said every target field is serialised to
the model; it is not (the snapshot names its keys). The reason they live on the
runner is that they are secrets and the runner's life is exactly the operation's.

--- fix(connectors): the review findings the operation must survive before the fold

The MCP prepare threads and the install worker started with an empty context,
so a named-profile turn's home override never reached them: the flow resolved
the process home's `mcp_servers` and stored the token there. Each thread now
runs in a copy of the calling thread's context. `connection.respond` had the
same gap on the RPC thread: an approval ran the enable, which writes
config.yaml, with no profile bound, so the flag landed in the launch home. The
handler binds the session record's profile the way `_connector_rpc` does;
`config_write_scope(None)` keeps that override, so the enable needs no change.

The watcher raised out of the tool when a per-target Skip landed while that
target's read was in flight: the operation stayed open with `live` closed, and
every later answer got 4004. A read for a row that is no longer live is dropped
at debug; only a refusal on a live row is still a contract violation. The same
skip from the card raced a row the backend had just connected and aborted the
rest of the answer; a skip for a resolved row is ignored and every entry, then
the settle check, still runs.

A read was bounded by the whole remaining deadline, so a hung gateway held the
first read for 300 s and Continue could not return the tool; a read now waits
ten seconds at most. A 429 parked only the target that read, but the budget is
the principal's, so the next target's read in the same tick spent it again:
every live target waits out the one Retry-After.

The MCP surface rule read the platform alone, so a desktop call without the
callback (registry dispatch from execute_code) opened an operation nobody
rendered and blocked for the deadline; it now uses the managed rule, surface
and callback. A write after settlement advanced `seq` while emitting the
frozen frame, so the resume snapshot named a seq no frame carried; the counter
stops at the settle frame. A mint that reports `initiated` with no account id
logs that the watcher cannot read the row.

(cherry picked from commit 587b228016bf8ee2b022e7fa58c2d9f6057809e9)

--- fix(connectors): the desktop card holds a verb until the backend answers, and never takes the keyboard from a credential field

The review of the desktop card found five defects; each one is a test first.

- A resume snapshot was refused whenever the cache held a settled operation, whichever
  operation it was, so a session that opened a second operation after settling the first
  never got its card back from a resume. And the refused snapshot handed the caller the
  settled cache as "the request", so the session was flagged as needing input behind a
  summary with no controls. Only the same operation can refuse the snapshot now, and a
  refused one is no pending card.
- The focus handoff picked the row's first button, which after pending -> initiated is the
  disabled working verb; the focus call was a no-op and the keyboard landed on the document
  body. It picks the first control that can take focus, else the row. It also moved focus out
  of a credential field the user was typing in whenever another row moved; it leaves an
  editable alone. The "focus Continue once every row resolved" branch was dead (Continue
  unmounts the moment nothing is unresolved), so it and its ref plumbing are gone.
- The done link navigated to a settled operation's session and rejected when the wake RPC
  did (4004 once the operation left the live registry). A settled request is ignored, and a
  refused wake is nothing: the wake only shortens the wait, the watcher still ticks.
- Install spun forever when the store refused to send the consent (the operation gone or
  settled under the card): `respondToConnectionRequest` resolves false in that case and the
  verb was only released in `catch`.
- After a partial approval (a required credential missing) the backend answers with a
  same-state frame whose detail names what is missing; nothing released the verb because it
  was held until the row's state moved. The row now remembers the seq the click saw and holds
  the verb only until a frame past it arrives, which is the backend's word on the click
  whether or not the row moved.

(cherry picked from commit 4c534d6e2af979778d9a2423cf97d717edeaa1a9)

--- fix(connectors): a skip that loses the race to the watcher is ignored, not raised

The skip guard read the row's state and then moved it; the watcher can connect the
row between the two, and the refused move aborted the rest of the card's answer.
The refusal itself is now the witness: a move refused for a row that is resolved,
or on a settled operation, is the same nothing-to-do as a row resolved earlier.

Two recording fakes in the managed tests kept their lists on the class; they now
start per instance so a lifted fake cannot share reads between tests.

--- fix(connectors): a resume that lands behind a newer live frame still reports the pending card

The refused-snapshot branch answered "no pending card" for both reasons it can
refuse: the operation settled, or a newer frame already moved a row. Only the
first is no card. For the second the live card is still open and blocking the
turn, so the caller must keep the session flagged as waiting on it.

* test(connectors): defer new connection coverage until implementation settles

Remove PR-added test cases and their unused helpers while retaining
existing tests adapted to the changed connection contract. The three
PR-only renderer test files are removed for now.

Focused behavioral coverage will be added as the final implementation
step before verification. Existing main coverage is not being removed
wholesale, and this does not declare the feature merge-ready.

* fix(connectors): commit MCP authorization at initialize, save setup values after success

The OAuth probe treated one exception as one outcome: any failure after the browser
step restored the token snapshot and manager entry, so a server that accepted the
token but failed tools/list discarded a completed consent. Now the probe reports
whether initialize succeeded (details["initialized"], read from the claimed
MCPServerTask). Failure before that point rolls back as before. Failure after it
saves the server config, keeps the tokens, and reports tools unavailable through
flow.discovery_error; the card can retry discovery without repeating consent.

Catalog install wrote the submitted values to .env before install_entry and the
probe ran. The values now live in the secret scope for the duration of the install
(get_env_value reads through get_secret, so install_entry finds them without a
prompt), and .env is written only after the probe returns tools. A failed probe
removes the server block and writes nothing. _probe_tool_names returns None on a
failed probe instead of an empty list, so failure and a valid empty listing are
distinct.

Failure text is redacted before it reaches target detail: every value the card
submitted for that target is replaced by exact match, then the pattern redactor
runs.

required_env now carries the manifest's secret and default flags; Target carries the
manifest's post_install text as instructions. The wire contract gains secret,
default, instructions and discovery_error; generated TS and OpenRPC regenerated.

* fix(connectors): one OAuth callback receiver picker for the connection card

The card's authorize target built its redirect from the dashboard web server and
raised when none was bound in the process, so a standalone hermes --tui session
could never authorize an MCP server. The receiver is now chosen in one place
(choose_callback_receiver): a pinned pre-registered client keeps the SDK's own
listener on the registered port; a client-advertised loopback URI is used as-is and
its callback arrives through the mcp.servers.oauth.callback relay; otherwise the
backend binds a one-shot loopback listener and feeds it into the flow. The dashboard
route stays with the dashboard web page, which cannot bind a port.

tui_gateway/mcp_oauth_sessions.py had a second copy of the loopback listener and a
_worker that referenced _probe_with_rollback, set_hermes_home_override,
reset_hermes_home_override and Path without importing them, so every RPC-started
flow raised NameError. Both are deleted; start_flow spawns run_worker directly and
uses the same receiver picker. Flow registration is shared (register_flow /
finish_flow) so a card-started flow with a client URI is reachable by the relay.

Under an SSH session with no client listener the attempt's detail carries the
existing paste-the-redirect instructions. Nothing on the card path opens a browser.

* feat(desktop): setup-form modal for MCP connection cards

The connection card rendered an MCP server's setup fields inline: every field as a
password input, no default value, no instructions. A URL such as the n8n MCP
server URL was typed blind, and the manifest's setup text never reached the user.

Two components carry the form now. SetupFieldList renders the ordered fields the
backend declares (a plain field as text prefilled with the manifest default, a
secret field masked and empty). SetupFormDialog composes it with a one-line title,
the manifest instructions, an inline error for a failed attempt, and Cancel /
Connect. The row's Install action opens the dialog when the target has fields;
Connect sends {status: approved, env}; Cancel sends {status: skipped}. A failed
attempt keeps the dialog open with the draft intact. The draft lives in the dialog
component only; nothing reaches the store or the resume snapshot.

Once the backend publishes the authorization URL the dialog shows it as text with
an Open in browser button. Nothing opens a browser on a state change: the Try
again path on both cards used to open the re-minted link at once; it now waits for
the row's update frame and the user's click.

The store types gain the wire's secret, default, instructions and discoveryError
fields and normalise them; a connected target with discoveryError renders as
authorized with tools unavailable.

* feat(tui): connection card in the Ink TUI

The Ink TUI had no client for the connection operation: connection.request,
connection.update, connection.respond and the pending_connection resume snapshot
were unhandled, so a manage_connections call in a TUI session could only print a
link through the model.

connectionOperationStore.ts holds the backend snapshot: a request opens only for a
new operation id, an update applies only to the live operation with a higher seq,
a settled update freezes the id so a late request frame cannot reopen the card.
The gateway event handler feeds it; session resume hydrates it from
pending_connection.

connectionSetupOverlay.tsx renders one callout in the prompt zone: a one-line
title, the manifest instructions, every field (plain rows prefilled with the
default, secret rows masked), then a Connect / Cancel selector. Connect sends the
draft through connection.respond; Cancel skips the active target. When the backend
publishes the authorization URL the callout shows it as text under "Press Enter to
open in browser"; Enter is the only thing that opens it. A failed attempt unlocks
the fields with the draft kept. An accepted secret renders as "Set" and is never
echoed. A target authorized without tools shows that state and Continue.

The overlay joins the existing input-owner set in overlayStore so typing and other
prompts are blocked while it is up.

* feat(cli): connection panel in the classic CLI

The classic hermes CLI passed no connection_callback, so a manage_connections
call could only print an authorization link through the model and could not take
a setup value at all.

The CLI now renders the connection operation as a prompt_toolkit panel, on the
same queue mechanism as the clarify panel: the agent thread's callback opens the
panel, blocks until the first decision, then returns so the operation's watch loop
runs; every later state reaches the panel through ConnectionOperation.on_change
(installed only when no gateway hook is set, restored on close).

Panel: one-line title, the manifest instructions, one row per field (plain rows
prefilled with the default, secret rows masked in the input buffer and rendered as
"Set" once accepted), then Connect / Cancel. Connect sends the draft through
apply_answer; Esc skips the active target; Ctrl-C sets the tool-thread interrupt so
the operation settles as interrupted. Once the backend publishes the authorization
URL the panel shows it with the target detail and "Press Enter to open in browser";
Enter is the only thing that opens it. A failed attempt unlocks the fields with the
draft kept; an authorized target without tools offers Retry discovery / Continue.

Up/Down move between rows, Left/Right toggle the action, and the panel joins the
blocking-overlay guards so chat input, history and voice stay out while it is up.
The single-query (headless) mode passes no callback, as it does for clarify.

* feat(connectors): register a connected MCP server and report its tools in the result

After a successful install or authorization the target carried the probe's tool
names and the model was told the tools "become available on your next turn". MCP
tool schemas are deferrable by construction, so nothing about them lives in the
sent tool array; a server registered in the scoped registry is callable through
tool_describe/tool_call in the same turn. The operation now registers the server
(register_mcp_servers under the owner's home scope) once authorization is
committed, records the registered names on the target, and the settled result
carries a tools_listing block in the deferred-catalog format plus a note that the
tools are callable now. A registration failure keeps the target connected with
tools: [] and a sanitized discovery_error. agent.tools and the system prompt are
untouched; the between-turns refresh updates the catalog block as before.

The card gate no longer asks for the desktop platform. Every surface that renders
the card attaches a connection callback (Desktop, the Ink TUI, the classic CLI);
registry dispatch and messaging sessions attach none and keep the link result.

Pre-commit rollback in probe_with_rollback used restore(only_if_absent=True),
which skips the rollback when a token file exists. On a first-ever authorization
the only file is the one this attempt wrote, so a token the resource rejected was
kept. Live E2E (controlled provider answering 403 to the issued token) showed the
row fail and the token survive; the rollback now restores the snapshot outright,
and the same run shows the token file removed.

* fix(connectors): install an OAuth catalog entry through the card's own flow

The card's install ran install_entry and then a plain probe. For an OAuth entry
that probe has no card flow around it: in the desktop backend it failed at once
("non-interactive environment and no cached tokens"), and in the classic CLI it
saw a TTY, opened a browser by itself and drew install_entry's curses tool
checklist over the panel. Since a probe failure is now an error, the failed
install was rolled back with _remove_mcp_server, and authorize refuses a server
that is not configured. A clean home had no path to a connected OAuth entry, which
is 55 of the 65 catalog entries. Found by the live three-surface run.

An install now builds the entry's configuration in memory (card_install_config:
no prompts, no probe, no checklist; a prior tool selection or the manifest's
curated filter applies). An OAuth entry starts the same flow authorize uses with
that configuration and the setup values in the attempt's secret scope, so the row
reaches the URL step, and the configuration and the setup values are saved
together when initialize accepts the token. Every other entry is probed in memory
and saved after the server answers. A failure writes nothing, so a failed
reinstall keeps the previous configuration, and the failed row asks for its fields
again so the card can reopen the form over the draft it kept.

* fix(agent): make a server connected in this turn callable in this turn

tool_describe and tool_call resolve names inside the agent's toolset selection,
which is fixed when the agent is built. A server that manage_connections had just
registered was therefore "not found" for the rest of the turn whose result calls
its tools available, and stayed out of the next turn's catalog too. The live
Desktop run showed it: the result listed mcp__fx_oauth_fields__echo and the
tool_describe that followed answered not_found.

The executor now adds the MCP servers the call connected to the selection. Only
the selection changes; agent.tools does not, so the sent tool schema bytes stay
the same. A selection of None (every toolset) and the no_mcp sentinel are left
alone.

* fix(cli): reopen the connection form when a required field is still empty

When the backend refuses an answer because a required field is missing it keeps
the row pending and names the fields. The panel mapped that frame to its waiting
phase, which draws the detail line and nothing else, so the user saw "waiting for
FX_API_KEY" with no fields and no buttons until the 300 s deadline. The panel now
returns to the form on the first missing field, over the draft it kept.

* test(connectors): the no-card path is the one with no callback attached

The card gate is "a connection callback is attached" since the classic CLI got
its own panel. This test still attached one under platform "cli" and expected no
card, so it opened a real operation, waited out the 300 s deadline and failed.
It now drives the path that has no card: no callback.

* fix(connectors): order a cancel against the commit and stop deleting a working grant

Found by the live runs on Desktop, the Ink TUI and the classic CLI.

A user's skip did not stop the OAuth attempt. With the worker parked in the token
request, the row settled "skipped" and the token was written 33 s later. A skip
now cancels the attempt, and one lock orders that cancel against the commit: the
attempt is either canceled with the earlier tokens restored, or committed and
kept. When the commit won, the skipped row says the authorization was kept. An
interrupted turn cancels its attempts the same way.

Every attempt began by deleting the saved tokens, so retrying discovery for an
authorized server demanded consent again, and a cancel in between left no grant
at all. The card's flow now connects with the saved tokens first, with no browser
step; only when they do not work does it replace them. The RPC session surface
keeps the old behavior because its caller waits for an authorization URL.

A second Connect from the form a failed row reopened was dropped without a frame,
which left the Desktop dialog with Connect loading and Cancel disabled until the
deadline. An approval on a failed or expired row is now a retry with the new
values.

The reported tool names came from the registration call, which returns nothing
for a server the process already holds. A retry after a failed listing therefore
said "no tools" while the server had them. The names are read from the registry,
and a parked server is woken first.

An authorization that commits after its card closed by deadline was never
registered, so the next turn still could not reach it. The runner keeps such
attempts, and the between-turns refresh adopts the ones that were approved.

An error with no message reached the user as a class name ("CancelledError").
The tool description still said a server's tools arrive on the next turn.

* fix(desktop): let an OAuth install open its link, and keep the form's draft

An install of an OAuth catalog entry with no setup fields reached the URL step
and gave the user nothing to click: an initiated install was always drawn as a
disabled spinner, and the only other place the link is shown is the setup dialog,
which opens for entries with fields. That is most of the catalog. An initiated row
that carries a link now offers Open, whatever the action.

The setup dialog reset its draft whenever the field list changed identity. The
backend sends a fresh list with every frame and an empty one while an attempt
runs, so a failed Connect erased what the user had typed. Fields now only fill in
what the draft lacks, and closing the dialog drops the draft.

The settled summary dropped the "tools unavailable" fact, and the live row offered
a Try again for that state which the gateway refuses for a connected target. The
summary keeps the fact and the row offers no dead control.

* fix(tui): keep the typed draft when a Connect fails

The overlay reset its draft whenever required_env changed identity, and every
backend snapshot delivers a freshly parsed array. A failed Connect therefore came
back as a form with the default region and an empty secret. The draft now resets
per target only, and a snapshot's fields fill in what the draft lacks.

* fix(cli): show the URL step and start an install that has no fields

The panel applied the user's answer and then set its phase to "waiting". The
backend applies the answer on the same thread and its change hook had already set
the next phase, so the URL step of an OAuth install and the reopened form for a
missing field were both overwritten, and the card sat on "Waiting…" until the
deadline. The waiting phase is now set before the answer is applied.

A pending install or enable with no fields opened in the waiting phase, which
sends no approval, so the flow never started. It now opens on Connect/Cancel. A
target that already carries its link opens on the URL step.

Connect on a failed row re-ran the attempt without the values now in the draft; it
sends them. The authorized-without-tools phase offered a "Retry discovery" that a
connected target cannot run inside the same operation; it offers Continue.

* fix(connectors): a newer OAuth attempt replaces the older one for the same server

A card that closed by deadline leaves its worker waiting on the browser for up to
300 s, and a tampered callback leaves one waiting too. A retry or a new operation
for the same server then ran beside it: both wrote the same token files, and the
older one's rollback could write over the newer attempt's grant.

The newest attempt per home and server is recorded. Starting one cancels the older
attempt, takes over its pre-attempt snapshot so a later failure still restores the
state from before either, and the older attempt's rollback leaves the files alone.

* fix(desktop): label an install's link control "Open in browser"

The control that hands an install's authorization link to the browser reused the
action's verb, so the row showed "Install" before the click and "Install" again at
the link step, told apart only by the cue. It now reads "Open in browser", the
label the setup dialog already uses for the same act. Authorize keeps its verb.

* docs(mcp): describe the setup card on the desktop, the terminal UI and the CLI

The MCP guide said the chat install exists only in the desktop app and that the
CLI relays commands. All three surfaces now show the same card: fields, Connect or
Cancel, an authorization link the user opens, one save when the server has
accepted the token, and tools the agent can call in the same turn. The tools
reference gains the result fields (tools, tools_listing, discovery_error) and the
no-card behavior of an OAuth install.

* fix(cli): keep the layout hook callable without a connection widget

The CLI panel commit added connection_widget to _build_tui_layout_children as a
required keyword. That method is the documented override point for wrappers, and
five existing tests (extension hooks, prompt stash, subagent dock) call it without
the new argument, so CI failed with a TypeError. The argument is now optional, like
the other widgets added after the hook was published; a missing widget is left out
of the layout.

The settled tool result also carried the target's setup instructions. Those are
the card's text for the user, and a catalog entry's notes can predate this flow
("restart your session so the tools are loaded"), which contradicts a result that
says the tools are callable now. The model-facing result drops them; the card
payloads keep them. This restores test_connector_local_batches.

* fix(connectors): wait for an in-progress registration before reporting a server's tools

Found with the real Vercel MCP server. Saving the configuration wakes the config
watcher, which starts its own connect for the new server. The operation's
registration call then skips the server as "already connecting" and returns at
once, so the card settled "connected" with no tools, and the 214 tools were
registered three seconds later. With no names in the result the model searched,
found the hosted connector of the same vendor and asked the user to connect that
instead.

The registered names are now read once the registration has finished: while
another task is connecting the server the read waits (30 s at most), a parked
server is woken once, and a server that finished registering with no tools is
still a valid empty list.
2026-09-21 10:04:32 +05:30
teknium1
86a599cbae fix(gateway): human-delay pacing comes from each profile's human_delay config, not process env
BasePlatformAdapter._get_human_delay read HERMES_HUMAN_DELAY_MODE/_MIN_MS/_MAX_MS from the
process environment at every send, so under multiplexing the launch profile's pacing applied
to every served profile, and the documented `human_delay:` config section (mode/min_ms/max_ms,
already in DEFAULT_CONFIG) was never consulted. The runner now resolves `human_delay` per
profile through the same seam as the busy-text timings (`_human_delay_from_config`,
snapshotted in `_snapshot_profile_busy_modes`, installed by `_wire_adapter_handlers`) and the
adapter only consumes the installed range. Invalid `custom` bounds (non-integer, negative,
inverted) warn naming the key and fall back to the natural range.

Fixes #116895
2026-09-20 17:01:05 -07:00
teknium1
522e121e90 fix(desktop): pin the update-check proxy deps exactly and document the proxy variables
The salvaged commit added https-proxy-agent, proxy-from-env and
@types/proxy-from-env as caret ranges; the repo pins every dependency to an
exact version so `npm ci` resolves the same tree everywhere. Pins are the
versions the lockfile already resolved (7.0.6 / 2.1.0 / 1.0.4), regenerated
with `npm install --package-lock-only`.

Documents that the Desktop update check now follows HTTPS_PROXY / HTTP_PROXY /
NO_PROXY in the environment-variables reference.
2026-09-20 15:08:40 -07:00
teknium1
dbcbd9d9db feat(gateway): /branch opens a sibling thread by default; --here keeps this chat (#66023)
On Discord, Telegram, Slack and Matrix a plain `/branch` used to rebind the
CURRENT chat/thread's session key to the clone, ending the original session
on that surface. The user could not keep the original path live while
exploring an alternate one — the opposite of what a branch is for.

Now the handler opens a sibling thread through the adapter's existing
`create_handoff_thread` BEFORE cloning (a failed create never orphans a
branch row), binds the thread's own session key to the clone with the
thread's routing columns written at create time, and leaves the origin key
untouched. `/branch --here` keeps the legacy in-place switch; platforms
without threads, DMs, unknown Discord parents and adapters that cannot open
a thread fall back to in-place with a one-line note. The CLI strips the
flag through the same parser so `--here` never becomes a session title.

Destination source shapes mirror each adapter's inbound key (Discord keys
threads on their own id; Telegram/Slack/Matrix on the parent chat), the
same rules the CLI->platform handoff uses.

Live repro (real gateway + real Slack adapter against a stand-in Slack
Socket Mode/Web API): base ends the origin session and rebinds its key;
fixed posts the thread seed, replies "this chat stays on it", the origin
thread keeps its session and the follow-up typed in the new thread lands on
the branch (parent_session_id = origin).

Design and first implementation by Angello Picasso (#66014, #66024);
this is a slim port onto the split slash_commands_* layout.

Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
2026-09-20 13:37:10 -07:00
treatux
6ae4cb88d5 fix(docs): reference pages name tools the registry never registered
The shipped tool-surface references still document the pre-consolidation
surface: tools-reference.md lists cronjob/todo/process/project_create/
project_list/project_switch/open_preview/close_preview/read_preview/tour/tip
(6 uncallable, 5 hidden dispatch-only aliases), and toolsets-reference.md
still claims web_search is a member of the browser toolset — membership
decacbac3 deliberately removed (#64503) with a regression test. Both rename
commits (e16ad33a9, 217ab2f8d) left website/ untouched.

Pin the pages to the live registry with a contract test (real
discover_builtin_tools()/resolve_toolset queries over the shipped .md data);
rename the rows to the registered surface (cronjob_manage, todo_list,
process_manage, desktop_project enum, desktop_preview, gui_tour, show_tip)
and add the browser row's actual members (browser_vault_*, browser_exec,
apply_layout) that the docs never mentioned.
2026-09-20 12:56:25 -07:00
treatux
2a1ec15a6f docs(website): regenerate the skill docs from the shipped skill tree
website/scripts/generate-skill-docs.py is the documented source of both catalogs
and of the per-skill pages, but nothing compares its output with what is
committed, so the committed copies drifted:

- optional-skills-catalog.md was missing agent-merge-conflict-arbiter and listed
  pr-lens under blockchain (it lives in software-development);
- skills-catalog.md still listed merge-reconciler, which #98539 moved out of the
  bundled set;
- 196 pages and both catalogs carried Windows path separators — in the Path
  column and inside GitHub blob links, where a backslash is a broken URL.

This commit is the generator's output (re-running it on this branch is a no-op),
plus the orphan page #98539 left behind for merge-reconciler: no skill backs it,
its catalog row is gone, and no page links to it.

tests/skills/test_skill_docs_contract.py is the guard: the next skill that ships
without a catalog row, or a page regenerated on Windows, fails there instead of
on the published page.
2026-09-20 12:54:07 -07:00
chelsealong
e23a033e85 docs(matrix): surface the DM-classification bypass in the operator-facing docs
The prior commit only documented the <=2-member DM auto-classification
in the adapter.py module docstring — source an operator configuring
MATRIX_REQUIRE_MENTION etc. would never read. Issue #114733 explicitly
asks for the callout wherever MATRIX_ALLOWED_ROOMS,
MATRIX_FREE_RESPONSE_ROOMS, MATRIX_REQUIRE_MENTION, and MATRIX_AUTO_THREAD
are documented for operators, i.e. the env-var reference table and the
Matrix user guide. Add the same rule + bypassed vars + escape hatch to
both, matching the precedent set by 08aa67e473 for this kind of env-var
clarification, and extend the doc-content test to cover them.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-20 10:42:25 -07:00
teknium1
d5c2e0fb16 fix(env): a parent-injected dashboard session token survives the .env reload
A launcher that spawns `hermes dashboard` (Desktop shell, link-style
integrations) mints HERMES_DASHBOARD_SESSION_TOKEN into the child's
environment and keeps the same token for its own /api probes. Every
dotenv layer in load_hermes_dotenv() loads with override=True, so a
persisted HERMES_DASHBOARD_SESSION_TOKEN line in ~/.hermes/.env replaced
the injected token: the child authenticated with the persisted value and
the parent got HTTP 401 from its own child.

Treat the token as a spawn credential in the dotenv publisher: a value
that dotenv did not put into os.environ (tracked by _DOTENV_PUBLISHED)
is left alone, while a value an earlier pass published still reloads,
so .env edits and home switches behave as before. Other keys, including
the documented HERMES_DASHBOARD_PUBLIC_URL, keep .env-wins precedence.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-20 10:39:01 -07:00
teknium1
c7e165a068 docs: list HERMES_LANGFUSE_MAX_DEPTH in the environment-variable reference
The knob is user-facing; the README alone is not where operators look.
2026-09-19 23:56:34 -07:00
teknium1
05c4c40219 docs(mcp): oauth.user_agent default, redacted token-error excerpt, login keeps server metadata 2026-09-19 23:46:42 -07:00
teknium1
0818892db3 fix(mcp): accept an origin-issued metadata document for a path-scoped OAuth authorization server
A protected resource may advertise a path-scoped authorization server
(`https://www.strava.com/mcp-issuer`) whose RFC 8414 document, served from
`/.well-known/oauth-authorization-server/mcp-issuer`, declares the origin
(`https://www.strava.com`) as its issuer. The SDK's exact-string check
(`validate_metadata_issuer`, RFC 8414 §3.3) rejected that document with
"Authorization server metadata issuer mismatch" and the connection parked
before registration or login (#116233).

`metadata_issued_by_origin` accepts exactly that shape and nothing else: the
document must have been read from the well-known URL derived from the
advertised identifier (so a redirect target, the root document or an OIDC
fallback never qualify) and its issuer must be the advertised identifier's
origin. Only the origin's operator controls that location, so a party
controlling a path or a sibling host cannot use it to make the client accept
another server's endpoints; whoever could would already control the
exact-match document too. The browser flow applies it in the mixin's request pump
(installs the document, hands the SDK a 204 so its loop stops) without touching
`auth_server_url`, so SEP-2352 credential binding keeps the advertised
identifier while the RFC 9207 `iss` check and refresh-token issuer binding use
the document's issuer. The device flow applies the same rule in its own
discovery and now binds registered credentials the way the SDK's Step 4 does,
so the runtime flow reuses them instead of discarding them on the next 401.

Supersedes #116359: its rule accepted the origin issuer from any discovery URL
(root and OIDC fallbacks, redirect targets) and fabricated a 500 response.

Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
2026-09-19 23:40:35 -07:00
teknium1
30de041b01 docs: explain the three gateway connection-failure replies
The chat reply now differs for an interrupted connection, a refused/unroutable
endpoint and a cause-free SDK connection error; the FAQ names each wording and
what to do about it so a user reading one on Telegram/Discord/Slack knows
whether to restart the model server or just /retry (#116323).
2026-09-19 18:24:11 -07:00
Teknium
b6f8f8eb1f Merge pull request #116345 from NousResearch/boa-w3-small-b
Video generation tools no longer let the agent pick the model; video_gen.model is the only selector (Refs #83080)
2026-09-19 14:32:37 -07:00
Teknium
d180fc4311 Merge pull request #116340 from NousResearch/feat/codex-browser-pkce-login
Codex login gains an opt-in browser PKCE flow on localhost:1455; device code stays default (#95743, salvage #97058)
2026-09-19 14:31:35 -07:00
teknium1
19b29df13b fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.

The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.

Refs #83080
2026-09-19 12:22:44 -07:00
teknium1
72fccf2b20 fix: name host, attempts and request size when connect retries are exhausted
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.

One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.

Part of #97548
2026-09-19 12:21:39 -07:00
teknium1
47ab9adc56 docs: document hermes auth add openai-codex --browser and auth.codex_login_flow
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
2026-09-19 12:11:59 -07:00
teknium1
54c01bc19a chore: merge origin/main (resolve hermes_cli/runtime_provider.py, website/docs/user-guide/features/codex-app-server-runtime.md) 2026-09-19 10:51:04 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
b682a98ab8 test(codex): pin proxy override on rotation and model.base_url; document HERMES_CODEX_BASE_URL
Two invariant tests (red on origin/main): a 401 rotation onto a Codex pool
row keeps the HERMES_CODEX_BASE_URL target, and model.base_url under
model.provider: openai-codex resolves for pool credentials. Adds the
previously undocumented HERMES_CODEX_BASE_URL row to the environment
variables reference so proxy users can find the knob and its reach.
2026-09-19 10:33:08 -07:00
teknium1
85b2a3df6c feat(cli): hermes usage [--json] prints the /usage account limits without a session
Codex 5h/weekly windows (and Anthropic/OpenRouter limits) were only reachable
through the interactive `/usage` slash command, so cron jobs and shell scripts
had no way to read quota state (#33094, #57476). `hermes usage` fetches the
same snapshot through `agent.account_usage.fetch_account_usage` — the credential
resolution a session with no live agent uses — and prints it with the same
renderer; `--json` emits one stable, documented document, exit 1 with a single
stderr line when no credential is configured or the fetch fails.

Slim redo of #81819 (@himanusia): top-level command instead of `hermes auth
usage`, no --all/--account/--reset (the per-entry paths rendered the wrong
account for anthropic and the default path bypassed the runtime resolver).

Co-authored-by: himanusia <himanusia@users.noreply.github.com>
2026-09-19 10:31:14 -07:00
teknium1
03d40f8939 fix(gateway): a --global /model keeps the session override under a channel_overrides model; report a failed stale-override cleanup
Precedence is session /model > channel_overrides > config.yaml, so dropping the session
override after a --global write regressed chats whose channel_overrides names a model: the
confirmation said "switched to gpt-5.5" while the next turn resolved 'channel-model'. The
cleanup now runs only when no channel_overrides entry applies to the source; otherwise the
override stays (and is written through) so the confirmation stays true.

A failed set_model_override(key, None) was logger.debug only while the reply claimed a clean
save and memory had already popped the override — the durable stale copy would shadow
config.yaml on the next restart (the original #100314 symptom). The cleanup failure is now
returned like the config-write error: the in-memory override is kept, the confirmation carries
a warning line instead of "Saved to config.yaml", and a riding --reasoning stays session-scoped.

Tests (tests/gateway/test_model_picker_persist.py, both red on the previous head): the
channel_overrides case drives _handle_model_command with a real JSONL SessionStore and asserts
_resolve_session_agent_runtime(source) yields gpt-5.5; the cleanup-failure case asserts the
warning and the kept in-memory override. Docs note the channel exception and that the CLI/TUI
keep their per-session pin by design (resume restores the model that chat used).
2026-09-19 09:59:32 -07:00
teknium1
1a59519244 fix(gateway): a --global /model pick leaves config.yaml as the only durable model authority
A `/model <m> --global` switch (typed or picker) used to persist twice: the profile
config.yaml AND a per-session `model_override` in the session store. The session copy
has higher precedence on rehydration, so after a later global change (CLI `hermes model`,
another chat's `--global`) and a gateway restart the stale override silently won —
#100314 saw an explicit `gpt-5.6-sol-900k` resume as the base 272K `gpt-5.6-sol`.

Now `_record_model_switch` writes config.yaml FIRST and, on success, drops the redundant
session override from memory and the store. If the config write fails the switch stays a
truthful session override and the confirmation says "config.yaml not updated (...)" plus
the session-only hint instead of claiming "Saved to config.yaml"; a riding `--reasoning`
pin follows the same effective scope. `--session` and `--once` semantics are unchanged.

Tests: two invariants in tests/gateway/test_model_picker_persist.py (typed+picker clear
the durable override and a fresh SessionStore rehydrates nothing; failed config write keeps
the override and an honest reply), red on base. test_model_command_request_overrides now
points get_hermes_home at its own config so the --provider switch resolves session-scoped
as intended instead of the sandbox's fresh-install first-pick rule.

Fixes #100314
Supersedes #99825 (slim redo; the original wrapped the commit boundary through a
sys.modules-swapped mixin).

Co-authored-by: Andrex Ibiza, MBA <andrexibiza@gmail.com>
2026-09-19 09:59:32 -07:00
teknium1
393ffff04f fix(providers): resolve the chatgpt alias in the /model parser and hermes auth login too
The static catalog alias table (models_catalog_static._PROVIDER_ALIASES, consumed by
parse_model_input / normalize_provider in hermes_cli/models.py) had no entry, so
`/model chatgpt:<model>` and `-m chatgpt:<model>` kept the prefix as part of the model
name while `openai-codex:<model>` split correctly. hermes auth login now falls back to
the shared auth alias table instead of its own 6-entry list (custom providers still win).
Pins providers.normalize_provider('chatgpt') as well and documents the aliases in the
--provider row.

Part of #95794
2026-09-19 09:38:01 -07:00
Teknium
99a6ed741a Merge pull request #115757 from NousResearch/fix/boa-codex-app-server-lifecycle-migration-mcp
fix(codex-runtime): same-name MCP servers no longer make ~/.codex/config.toml unloadable; adds `hermes codex-runtime migrate` (#79023, salvage #79182)
2026-09-19 09:35:41 -07:00
kshitijk4poor
1536e1d242 docs(discord): finish the free-response auto-thread surfaces
Adds the key to cli-config.yaml.example and the multi-profile per-key list,
records the env-over-YAML precedence, and drops the duplicated
no_thread_channels clause in the new section.
2026-09-19 21:43:59 +05:30
Xunjin ZHENG
774e070731 feat(discord): add free_response_auto_thread opt-in
Free-response channels skip auto-threading by default so the bot replies
inline (lightweight chat mode). This prevented users who wanted BOTH
mention-free replies AND per-conversation threads from getting either.

Add a new opt-in `discord.free_response_auto_thread` (env:
`DISCORD_FREE_RESPONSE_AUTO_THREAD`, default false) that, when true,
re-enables auto-threading in free-response channels. Voice-linked
channels continue to skip auto-thread regardless, and the flag is
gated behind the global `DISCORD_AUTO_THREAD=true`.

Default behavior is unchanged; all 291 existing discord tests pass.
2026-09-19 21:43:59 +05:30
Siddharth Balyan
8d4abc3ea6 fix(mcp): retire the n8n bridge catalog entry (#116048)
Stop offering the third-party bridge for new catalog installs. Existing
connections keep their saved transport, credentials, and tool selection;
the runtime and configured-server controls do not require a manifest.

Update CLI examples and document that catalog reinstall is unavailable.
Adding n8n's official server remains separate work.
2026-09-19 17:51:30 +05:30
Kevin
8b7caf226f feat(codex): named custom providers work with the codex_app_server runtime
`model.openai_runtime: codex_app_server` only ever admitted `openai` /
`openai-codex`: a named custom provider (`providers.<name>`) resolves to
provider="custom" on the named-custom ladder rung, which never ran the
runtime gate, so `/codex-runtime codex_app_server` silently left the main
turn on Hermes' chat-completions client. And even when routed, thread/start
sent only `cwd`, so codex could not know which of its own providers to use.

- `_maybe_apply_codex_app_server_runtime` takes `requested_provider` and
  admits provider="custom" only when `codex_model_provider_id()` finds a
  configured `providers.<name>` entry (bare `custom`, ollama/vllm aliases
  and unknown names have no stable id -> ineligible, unchanged).
- The named-custom rung applies the same opt-in the pool rung already does
  for openai/openai-codex.
- `CodexAppServerSession(model=, model_provider=)` -> `thread/start.model` /
  `.modelProvider` (fields verified against the codex 0.147 app-server
  schema). `_ensure_codex_session` fills them only for custom agents; codex
  resolves base_url/env_key from its own `[model_providers.<name>]`, so the
  Hermes credential never enters the JSON-RPC payload.
- `tui_gateway/server.py::_make_agent` forwards `requested_provider` so the
  Desktop/TUI agent knows the provider id (it otherwise collapses to
  "custom" and codex would fall back to its default provider).
- Docs: matching `[model_providers.<name>]` + `env_key` requirement and the
  bare-`custom` ineligibility.

Ported and trimmed from #75191 by @cosin2077 (aux-loop `allow_codex_app_server`
plumbing, `cli-config.yaml.example` block and the integration-test suite dropped:
background_review already maps codex_app_server -> codex_responses on main).

Fixes #75186
2026-09-19 01:04:05 -07:00
teknium1
d09183d8d3 feat(cli): add hermes codex-runtime migrate [--dry-run] [--json]
The managed-block marker and the docs have referred to `hermes codex-runtime
migrate` all along, but no such CLI subcommand existed: the only way to run the
~/.codex/config.toml migration outside a chat session was to import the private
hermes_cli.codex_runtime_plugin_migration.migrate (#79023). The new subcommand
group (hermes_cli/subcommands/codex_runtime.py, registered like the other
groups) calls the same migrate() the /codex-runtime slash command uses, on the
selected profile home, with --dry-run (no write) and --json (full report incl.
preserved_user_servers and errors); exit code 1 when the report has errors.

Tests: one invariant for the same-name table (single header, valid TOML, user
command kept, report lists the name) and one for the CLI dry-run/json path.
Docs: conflict policy + command in the codex runtime guide and CLI reference.
2026-09-18 23:38:13 -07:00
teknium1
67757285f6 feat(sessions): hermes sessions repair-profiles settles crossed-profile durable state
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.

`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:

1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
   generations, usage rows, system prompt) to the owning store, parents before
   children so lineage survives, copy-then-delete so a crash leaves a duplicate
   the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
   existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
   -> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
   import re-injects them into routing every boot).

Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).

Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).

Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
2026-09-18 22:36:41 -07:00
teknium1
5e95050608 fix(agent): a server context rejection the transcript cannot explain is no longer "conversation too long"
A single-slot local server (LM Studio, Ollama) returns 500 "Context size has been exceeded."
when ANOTHER request — a background review from an earlier session — holds its context.
The foreground loop classified that as context_overflow, tried to compress a one-sentence
conversation, could not shrink it, and rendered "This conversation has grown too long …
/new … /compress" with compression_exhausted=True (gateway auto-reset, user message dropped
from the transcript).

_recover_context_length now measures first: when the server quoted no count of its own and
the local request estimate (+ output reservation) sits under half the known window, the turn
ends with distinct copy naming the likely cause (another request on the server / smaller
server window), failure_reason=server_error, retryable, no compression_exhausted — so CLI,
TUI/Desktop and the gateway all render a transient failure. Servers that quote their own
measurement ("233153 tokens > 200000 maximum") and requests near the window keep the
compress-and-retry path unchanged. The two buffered "keeping context_length … and
compressing" notices drop the trailing clause so they read true on both paths.

Fixes #114644
2026-09-18 19:18:57 -07:00
kshitijk4poor
1250a3e4eb fix(backup): an incomplete archive never prunes the last complete ones
Gate review: `--keep` pruning ran after the summary regardless of `errors`,
so a timer hitting the same unreadable file every run would exit 1 each time
and still rotate every complete `hermes-backup-*.zip` out after N runs,
leaving only incomplete archives. Prune only after a complete backup; the
test pins a pre-existing good archive surviving an incomplete run with
`--keep 1`. The summary no longer hard-codes the caller's exit code.
2026-09-19 03:43:58 +05:30
leomcamilo
ccb3d968ce fix(backup): exit non-zero when a full backup is incomplete
`hermes backup` recorded per-file failures, printed `Backup incomplete: <path>`
and still returned shell status 0, so a cron job or systemd timer would publish
"successful" archives missing state.db indefinitely.

`run_backup()` now returns whether the archive is complete and `cmd_backup()`
maps False to exit status 1. The zip is kept so the operator can still restore
the rest; hard failures keep their SystemExit(1)/(2). `--quick` is unchanged.

Slim redo of #101096 on current main (the branch predates the run_backup /
_run_backup_locked split and the backup lock); same policy, same exit codes.
Supersedes #68866 (@jbryce) which proposed the policy first.
2026-09-19 03:43:58 +05:30
kshitijk4poor
8925c70a1c docs(whatsapp): group access section says what the gateway admits; env reference rows; trim bridge tests
Groups: policy, group-JID allowlist, and that participants are still authorised by
the gateway sender allowlist or pairing (`open` alone admits nobody without one);
`require_mention` defaults to false; WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
rows in the environment reference. The alt-id node tests collapse to one (the
participantAlt case duplicated the first-contact case; the "still resolves via mapping
files" case only re-asserted matchesAllowedUser).
2026-09-19 03:15:40 +05:30
teknium1
d15208e5f0 docs(website): re-run the link sweep over pages merged since the rebase
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
2026-09-18 14:27:04 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
teknium1
b7b203cda0 fix(skills): built-in name collisions show a note in /skills, /help skills and the palette
A skill whose slug is a core command name or alias (e.g. a skill dir named
`handoff` or `plan`) is deliberately kept out of slash auto-registration —
370ebf2d3 ("guard skill slash commands against core-command and slug
collisions"): the skill map is consulted before built-in handlers in the
gateway dispatch path, so an auto /handoff would shadow the core command.
That guard stays. What users saw until now was only a WARNING in agent.log
repeated every session; the skill sat in `/skills list` as "enabled" with no
hint why `/handoff` ran the built-in instead (#113560).

Now one helper, agent.skill_commands.skill_command_collision_note, is the
single collision predicate: scan_skill_commands() uses it for the skip, and
four surfaces render the note it returns —
  "slash command /<name> unavailable — name taken by built-in; use /skill <name>"

- hermes_cli/skills_hub.py::do_list — the Status cell of `/skills list` /
  `hermes skills list`
- hermes_cli/cli_info_mixin.py::show_help — one dim ⚠ line per colliding
  skill under `/help skills` (also when no skill command is registered)
- tui_gateway/methods_tools.py::_catalog_skills — the commands.catalog RPC
  carries the notice in its existing `warning` field (rendered by the Desktop
  `/commands` output); discovery-failure messages still win over it
- hermes_cli/slash_exec.py::_exec_commands — the messaging-gateway `/commands`
  listing (Telegram/Discord/…) appends one ⚠ line per colliding skill

Fixes #113560
2026-09-18 12:51:22 -07:00
teknium1
250e12e760 fix(config): provider switch via config set drops the previous provider's route (#113719)
`hermes config set model.provider X` re-points the `model:` block at a new
provider but left `model.base_url` / `model.api_mode` from the previous
route in place. The runtime honours a persisted api_mode/base_url for
whatever provider the block names, so X's key was posted to the old
endpoint (e.g. https://chatgpt.com/backend-api/codex + codex_responses)
and every request 401'd with `api_key_not_supported` blaming X.

Reshapes the salvaged clearing from #40869 (which popped base_url on every
provider write) into route-aware syncing, mirroring what a persisted
`/model` switch writes (`model_selection_config_updates`):

- `hermes_cli/route_identity.py::provider_owns_route` decides whose endpoint
  a base_url is: the target's registry/plugin host, a `providers:` /
  `custom_providers:` entry resolving to the target, or bare custom/local
  aliases (configured BY base_url) -> owned; another known provider's host
  or a named entry with a different endpoint -> foreign; unknown host -> None.
- `drop_stale_model_route` pops base_url + api_mode when foreign (api_mode
  alone, with no base_url, is old-route wire state and goes too); keeps an
  owned route with its api_mode; keeps an unknown host.
- `set_config_value` runs it only when model.provider actually changes,
  prints what was cleared and why, or warns that an unrecognised base_url
  still applies (the warn-only shape of #113725).
- Same provider re-set, `model.default`, a target that owns the URL
  (openai-codex + chatgpt.com), a custom entry with that URL, and bare
  `custom` are untouched.

Tests trimmed to two invariants (clear matrix / keep matrix) in
tests/hermes_cli/test_set_config_value.py; docs in cli-commands.md.

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: Tim Richardson <tim@growthpath.com.au>
2026-09-18 12:49:40 -07:00
teknium1
668e505ac7 fix(mcp): OAuth discovery/registration carry a User-Agent; cancelled login frees its callback port
Widen the salvaged fixes to the whole class and add the pieces they missed:

- The default User-Agent for SDK-built OAuth requests moves from the manager's
  bridge into HermesProviderMixin.async_auth_flow, so the legacy build_oauth_auth
  provider gets it too, and the manager's pre-flight metadata discovery (its own
  client.send of a bare Request) stamps it as well. Why: the SDK sends discovery,
  registration and token requests through client.send(), which never merges the
  client's default headers; www.tradingview.com's WAF answers a header-less GET
  with 403 while curl gets 200, so metadata looked unreadable, the SDK guessed
  /register and /authorize on the MCP host, and the login died with
  "Registration failed: 404" (or, with a pre-registered client, "iss mismatch:
  ... != None" because no issuer was ever discovered).

- The callback listener now runs serve_forever() and is shut down before
  server_close(). A thread parked in handle_request()'s select() keeps the
  closed listening socket alive (the kernel holds the file for the duration
  of the poll), so a flow cancelled mid-wait left the port bound and the
  retry on the same pinned/cached port raised "OAuth callback port N is
  already in use" with no external collider.

- When every authorization-server metadata fetch failed, a registration error
  is re-raised leading with those statuses ("Could not read
  authorization-server metadata (403 from ...); dynamic client registration
  then fell back to a guessed endpoint on the MCP host and failed: ...").
  humanize_oauth_registration_error leaves that message alone so the 403 in
  it is not mistaken for a DCR allowlist refusal.

Docs: mcp-config-reference notes the discovery/registration User-Agent and
the new error lead.
2026-09-18 12:47:49 -07:00
teknium1
76c640ff11 fix(web): list, activate and delete legacy custom_providers entries on Custom Endpoints
GET /api/providers/custom-endpoints read only providers:, so a post-migration
custom_providers: list entry (still routed by get_compatible_custom_providers)
had no row and could be deleted nowhere. Build the legacy rows from that same
merged view (source "custom_providers"; entries from providers: carry a
provider_key, legacy ones do not). DELETE removes the matching list entry
when the id is not under providers:; activate promotes the entry to
providers.<key> first, since the main slot names providers by key.

The doctor residue check keeps firing but no longer claims the row is missing;
its rationale, the docs line and the non-list message now talk about the
retired list store ("legacy custom_providers entries are ignored until it is").
2026-09-18 12:44:32 -07:00
teknium1
3e182f46c3 fix(doctor): flag a non-list custom_providers and legacy list entries with no providers: twin; name the edited profile on Custom Endpoints / Local Models
`hermes doctor` (and the startup config-structure warning) now report a
`custom_providers` value that is not a YAML list — naming the key and the
received type — instead of the runtime silently serving "0 endpoints".
Doctor also warns about every legacy `custom_providers` list entry whose
endpoint URL has no `providers:` twin, with the exact move to make: such an
entry is served by the chat picker (dual-read view) but has no row on the
Custom Endpoints settings page, and the one-shot v11→v12 list migration
(config_migrations._migrate_to_12) never re-fires once the version is past 12.
Warn-only on purpose: re-running the migration would mint `<key>-N`
duplicates for entries that DO have a twin.

Desktop: Custom Endpoints and Local Models send unscoped requests and always
edit the app's active profile; they now print the same "Changes on this page
apply to the “X” profile." note the Model page uses (hidden with one profile).

Part of #114471 (items 2, 5, 6).
2026-09-18 12:44:32 -07:00
teknium1
8669e47a60 fix(picker): curated fallback for cold OAuth rows; Z.AI failed-probe negative cache; trim salvage
Salvage follow-up to the previous commit (#114397 by @Finn763):

- Codex/Copilot rows went through cached_provider_model_ids directly, so a
  cold cache on the non-blocking read path rendered an EMPTY Copilot row
  (live repro: copilot:0). Route them through _live_or_curated_ids like
  every other built-in so the curated list fills the first open.
- Drop the catalog_pending row flag, provider_catalogs_refreshing and
  _mark_catalogs_pending: no surface consumes it and it would have needed
  a gateway contract regen. Drop the _spawn_background_warm wrapper: the
  ollama-cloud row's own SWR refresh already warms that cache.
- Z.AI endpoint detection only persists a SUCCESS, so a key that 429s on
  every endpoint re-ran four chat-completion probes on every
  credential-pool load (load_pool("zai") runs several times per picker
  open; the reporter's logs show exactly these repeated POSTs). Memoize
  the failure in-process for 5 minutes. Copilot already has the same
  negative cache for its token exchange.
- Tests trimmed to two invariants (degraded provider cannot stall the
  open + row still renders; explicit refresh still probes) plus one for
  the Z.AI negative cache; a rigid test fake gains **kw for the widened
  cached_provider_model_ids signature.
- Docs: how GUI pickers source per-provider lists and what Refresh does.

Live repro (temp HERMES_HOME, five built-ins pointed at a stalling
/v1/models stand-in, Z.AI key set): refresh=False 50.5s on origin/main ->
3.7s on this head; without Z.AI 43.7s -> 1.3s.
2026-09-18 11:01:52 -07:00
teknium1
09cf1b926f fix(state): reset forks leave the Python lineage walk too; /resume ranks lineages by activity
The SQL chain step (#114287) stopped a `_reset_from` child of a compression-ended parent
from winning tip projection. The Python twin had the same blind spot:
`_is_compression_child_row` / `_compression_lineage_root` treated the reset fork as a
continuation, so `get_compression_lineage(tip)` collapsed to `[tip]` (ancestors lost for
prompt-cache scope and export) and the fork shared the lineage's turn-lease key. Both now
ask `_is_explicit_fork_child_row(include_reset=True)`; `get_compression_lineage`'s own
early return keeps excluding only branch/delegate/tool so a reset child that later
compresses still walks forward to its children.

Gateway bare `/resume` lists with `order_by_last_active=True`: a lineage compressed for
days is projected onto its live tip and belongs where the user last touched it, not at
its root's `started_at` (the reporter's tip, active yesterday, was buried under a
September-12 start). Desktop already requests `order=recent`.

Docs: `/resume` row in slash-commands reference. Tests: one lineage-walk invariant, one
/resume ranking invariant, both red on origin/main.

Part of #114271
2026-09-18 10:39:13 -07:00
teknium1
4590ef8b58 fix(profiles): polled profile lists never walk skill trees; vanished skill dirs no longer abort enumeration
GET /api/profiles, the profiles.list RPC and the /api/profiles/projects/tree
fan-out are polled by the Desktop every few seconds (roster tick, focus,
gateway-open). Each call ran list_profiles() -> _count_skills() ->
Path.rglob("SKILL.md") over EVERY profile once the 30 s TTL expired: ~4 fs
calls per skill, 2.4-4.7 s per walk on 7-84 profile installs, ~half a core
at idle, and on Linux enough to starve the renderer's 60 s API timeout. A
skill dir removed mid-walk raised FileNotFoundError out of rglob and aborted
the whole profile list.

- list_profiles(lazy_skill_count=True): skill_count is the last known value;
  a missing/aged entry schedules ONE background _count_skills per profile per
  60 s recheck window, so the request thread does zero skill-tree I/O and the
  refresh cadence is decoupled from the poll rate. The two polled callers and
  the REST fallback entry use it; the sync default (CLI, detail views) is
  unchanged.
- _walk_skill_count: the repo walker (agent.skill_utils.iter_skill_index_files,
  os.walk with excluded/support dirs pruned) instead of rglob — half the
  fs calls and best-effort on subtrees that vanish mid-walk. profiles.describe
  uses the same walker.
- _profile_targets always uses profiles_to_serve (pure directory read):
  projects/tree and sessions/pull-requests only ever consumed name/path.
- Skill-count TTL 30 s -> 600 s (signature invalidation still catches
  skill add/remove immediately on the next refresh).

Live repro (5 profiles x 200 SKILL.md, temp HERMES_HOME, py3.11):
  before: GET /api/profiles 2175 stat + 2015 scandir per cold call,
          profiles.list 2189 + 2015, projects/tree 2177 + 2015;
          a skill dir removed mid-walk -> FileNotFoundError from list_profiles()
  after:  GET /api/profiles 0 skill-tree stats/scandir on the request thread
          (counts land from the background refresh by the next poll),
          profiles.list 0, projects/tree 0; _count_skills (detail/control)
          still reports 200 with 203 scandir; the vanished-dir case returns
          199 and list_profiles() enumerates all 5 profiles.

Fixes #114041
Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:18:40 -07:00
teknium1
72360ae1d2 fix(doctor): detect a dead IPv6 route and name network.force_ipv4
#114265 secondary finding 1: ``network.force_ipv4`` was undiscoverable (default
off, mentioned only by a rotating tip), so an advertised-but-blackholed IPv6
prefix cost the reporter weeks. ``hermes doctor`` now runs an ``IPv6 route``
probe in the API Connectivity section: one 2 s IPv6 TCP connect to a known
dual-stack host. A timeout is the dead-route signature and is reported as a
warning plus a summary issue naming ``network.force_ipv4: true``; no AAAA /
no IPv6 route at all is healthy (fails fast, no stall) and ``force_ipv4``
already set skips the probe. Two invariant tests over a mocked connect seam.

Docs: doctor reference and the network config section describe the check.
2026-09-18 09:56:28 -07:00
teknium1
fc94ff56e1 fix(cli): expand HERMES_HOME once at process entry; sweep raw readers
Follow-up to the salvaged hermes_constants expansion (#109212): the CLI has
~30 raw `os.environ["HERMES_HOME"]` readers (fast --version path, early
display.interface probe, profile re-home, dotenv loader, subprocess-home
helper) that never go through get_hermes_home(). A literal `~` (fish, or
any quoted value) left them resolving `~/.hermes` against cwd while the
resolver now expands it, so the process would disagree with itself.

Normalize the env var once, at the earliest point of hermes_cli/main.py
(stdlib-only, before the fast paths), and expand it in the one raw
hermes_constants reader (`_profile_home_path`). A relative value that is
not tilde/variable-shaped is deliberately left alone: rejecting it would
break `HERMES_HOME=./tmp-home` in tests and CI for no user-facing gain.

Tests: real-CLI subprocess with HERMES_HOME='~/.x' under a fake HOME asserts
`config path` lands in the fake home and that no literal `~` directory
appears under cwd (the reporter's acceptance criterion, #114353), plus a
unit test for the normalizer. Docs: environment-variables reference.
2026-09-18 09:49:46 -07:00
teknium1
b65ec6a3f0 docs: list the DISCORD_MISSED_MESSAGE_BACKFILL_* env fallbacks
The adapter honours six env fallbacks for discord.missed_message_backfill
(including the new MAX_ATTEMPTS one) but none was in the environment
variables reference; document them next to the other DISCORD_* rows.
2026-09-18 09:44:59 -07:00