Commit Graph

1006 Commits

Author SHA1 Message Date
teknium1
36773e0d78 refactor(ts): one GatewayEventMap in apps/shared typed from tui_gateway emitters; drop never-emitted tool.progress
Three TypeScript clients each declared their own copy of the tui_gateway wire
types and had drifted apart: apps/shared had a partial GatewayEventName union
with a `(string & {})` escape hatch, ui-tui/gatewayTypes.ts a 150-line
discriminated union, and apps/desktop an `RpcEvent<T>` that was field-for-field
the shared GatewayEvent with `type: string`. None matched the emitter:
message.complete lacked warning/status/error/recoverable/error_surface,
tool.start/tool.complete lacked args/result, SessionResumeResponse lacked
session_key/messages_omitted/hydrating/auto_continue/todo_state, three
different ModelOptionProvider shapes disagreed on fields, and all three unions
handled a `tool.progress` event that no Python emitter has ever produced.

Now:

* `apps/shared/src/gateway-events.ts` is the single home: payload interfaces
  typed from the Python emitters (file::symbol cited per interface),
  `BackendGatewayEventMap` (89 backend names) + `ClientLocalGatewayEventMap`
  (5 TUI-synthetic transport events, clearly marked, excluded from the
  contract) merged into `GatewayEventMap`; `GatewayEvent<K>` is discriminated
  on `type` with `seq` typed. RPC shapes shared by 2+ surfaces live beside it
  (ModelOptionProvider = union of every field hermes_cli/inventory.py sets,
  incl. pricing_pending/free_tier_pending; SessionResumeResponse<Info>;
  SessionListItem with resolved_id; Usage).
* `JsonRpcGatewayClient.on<K>` is keyed by event name; the gateway.ready
  heartbeat/replay_epoch and per-frame `seq` reads are typed instead of cast.
* ui-tui and apps/desktop import the shared names; their local duplicates are
  deleted (no re-export shims — importers are repointed; the desktop plugin
  SDK barrel keeps its public `RpcEvent` name as an alias of GatewayEvent).
  web/src repoints ModelOptionProvider/ModelOptionsResponse.
* `tool.progress` handling is removed from the TUI handler/turnController,
  desktop event sets/tools handler, shared union, tests, and two docs
  (`grep '"tool.progress"' tui_gateway/` = 0 hits; the `display.tool_progress`
  config mode is unrelated and untouched).
* `message.complete.warning` (history-commit note from
  prompt_turn.py::_complete_turn_payload) is typed and surfaced on both
  surfaces through their existing notice paths (TUI pushActivity 'warn',
  desktop notify kind 'warning').

Contract: `apps/shared/src/gateway-events.json` is the sorted list of
backend-emitted names. `tests/tui_gateway/test_gateway_event_contract.py`
collects names from the Python emitter side (emit-helper literals, the
`.request → .expire` table, change-watcher table, child delta mirror,
subagent relay, desktop_ui tool emitters, gateway.ready/setup.ready/
browser-controller frames) and asserts emitted == JSON in both directions.
`apps/shared/src/gateway-events.test.ts` asserts BACKEND_EVENT_NAMES (which
the map type is `satisfies`-checked against) == JSON. Sabotage-verified: a
fake JSON name fails both tests; a fake TS name fails tsc + vitest; a fake
Python `_emit("...")` fails pytest.
2026-09-13 05:42:31 -07:00
Teknium
d46f79233e fix: sort imports in markdown-text.tsx per perfectionist rule 2026-09-12 21:17:02 -07:00
Teknium
e1c05ffa32 feat(desktop): persist video playback speed across transcript videos
Port from block/buzz#7336: a playback rate picked in any transcript
video's native controls persists as a device-level preference, so a
viewer who watches at 2x doesn't re-select it for every clip. New
players (and other open windows, via the persistentAtom storage sync)
start at the saved rate; out-of-range or malformed stored values fall
back to 1x, and returning to 1x removes the stored key.

Adapted from Buzz's hand-rolled localStorage module + custom player to
our persistentAtom store and the single <video> render site in
markdown-text.tsx (MediaAttachment), wrapped as TranscriptVideo.
2026-09-12 21:17:02 -07:00
brooklyn!
995e79afe4 feat(desktop): retire tutorials after the first month 2026-09-12 01:35:33 -05:00
hermes-seaeye[bot]
a84a2223f8 fmt(js): npm run fix on merge (#108829)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-12 04:42:23 +00:00
hermes-seaeye[bot]
f7cd8bf084 fmt(js): npm run fix on merge (#108814)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-12 04:36:01 +00:00
Siddharth Balyan
cbe9e5b294 chore(desktop): literal comments across the onboarding flow (#108438)
* chore(desktop): literal comments in the guide script and runbooks

Comment-only change to onboarding-script.ts and setup-profile.ts. The module
headers now state the purpose and the constraints that shaped each file. The
notes beside the runbook strings keep one fact per sentence, or are deleted
when the string beside them says the same thing. The runbook text, the
persona, the option pills and the SOUL text are unchanged.

Both versions transpile to identical output with comments removed.

* chore(desktop): literal comments in the guide chat cards and stores

Comment-only change to the guided chat's cards, directive dispatcher, option
catalog, assembly module and chip. Metaphor and personification are replaced
by the name of the atom, effect or CSS property they stood for. Comments that
restate the code are deleted. Two stale facts are corrected in place: the
mini layout trees point at app/contrib/layout-presets.ts, and the skip button
sets the onboarding phase to skipped rather than done.

One comment line in cards/frame.tsx from bb/connector-ui-e2e-v2 loses a
metaphor and an em dash; its fact is unchanged.

* chore(desktop): literal comments in the handoff and first build

Comment-only change to the handoff wiring, the kickoff, the receipt store,
the first-build check-ins, the handoff tour, the connector rows and the
machine profile store. Every kept comment names the caller, the constraint or
the defect it prevents. The claim that the tour never throws is removed: the
function can reject and its caller does not catch.

Five comment blocks in connector-tool.tsx written on bb/connector-ui-e2e-v2
lose personification, dramatic capitals and em dashes. Every fact in them
stays, and no block moves.

* chore(desktop): literal comments in the intro reveal

Comment-only change to the intro reveal's clock, timeline, cube renderer,
sound, scenes, store and README. Animation comments now name the actual
ramp, easing or offset with its number. Four comments that contradicted the
code are corrected: the first texture slot opens at 3700 ms, the tear settles
from 1 to 0 over 460 ms, the typing weight delays the character it sits on,
and INTRO_EXIT_MS is wall time in index.tsx but score time in the overlay.

* chore(desktop): literal comments in the Electron onboarding windows

Comment-only change to the window growth geometry and the two onboarding
windows. The 768 px floor keeps its one fact: the floor uses Math.ceil where
the deltas round, because rounding 906.24 DIP down leaves the media query
false. The comment that placed the CSS-pixel to DIP conversion at getBounds
now points at growWindowBounds, where it happens.

* chore(gateway): literal docstrings in the onboarding RPCs and the tour tool

Docstring and comment-only change. The module summaries state what each
module does and where authorization comes from, without contrast pairs. The
tool descriptions the model reads are unchanged. Two words in the tour tool's
module docstring lose personification; the rest of that docstring is as it
was.

ast.dump of both versions, with docstrings stripped, is identical for all
three files.

* chore(desktop): literal punctuation in the relaunch and film-end notes

Comment-only change to four lines that bb/connector-ui-e2e-v2 added to the
boot gate, the gate store and the intro gate. Each em dash becomes a colon, a
full stop or a pair of parentheses; one emphasis capital is lowercased. The
facts in the notes are unchanged.
2026-09-12 10:00:08 +05:30
Siddharth Balyan
474a7f8b2a Guided first launch: the first build connects the picked apps before it starts (NS-859) (#108317)
* feat(desktop): the first build connects the picked apps before it starts

When the user picked apps during setup, the build session's runbook now opens with one batched manage_connections connect for every pick, ends the turn, and waits for the app's "links opened" note before calling wait with a 180 s budget. The task starts the moment the wait returns or the user says to start, with whatever connected; pending apps are named once and skipped, not asked about. The no-account rule stays only for builds with no picks and for machine setup, which needs no account. Account data comes from the connected apps and the web only; credentials that happen to be on the machine are off limits to a first task. The deliverable is a page the user can open plus one real reading or action through a connected app.

The picks are gateway slugs from here on, so the runbook names each one with its title and the model never calls status to match them.

* feat(desktop): the onboarding picker shows only the lead-order apps the gateway carries

The picker reads the live catalog through useConnectorCatalog (#108292) and
offers a pick only when the gateway carries it. This commit narrows what it
shows to CONNECTOR_LEAD_ORDER: the everyday apps, in that order, per D89. The
rest of the catalog stays available to the agent; the first-run card does not
list it. Reverting to the full catalog with search is the one filter clause.

* feat(desktop): mark the first build session at handoff

Persist the build stored id with its created receipt so connection automation stays scoped to that session.

* feat(desktop): open first build connection links as one batch

Claim each tool call before opening its links and send one hidden setup note. Share row state by stored session and retain clickable links when the browser bridge is unavailable.

* feat(desktop): reconcile first build rows during the model wait

Poll the owning gateway while the newest wait is pending and preserve settled or interrupted results. Show connection states without the ordinary offer controls.

* feat(desktop): let the first build start with connected apps

Read the shared connection rows in the composer and submit a visible first-person start message. Persist use of the pill so it cannot return for the same session.

* feat(desktop): add connect first setup copy

Use the connector locale keys for catalog checks, sign-in states, and the start action. Other locales inherit these additions from English.

* feat(desktop): the first build waits 120 s, then asks; signed-in tools are fair

The wait budget drops from 180 to 120 seconds. When it runs out with apps still pending, the model stops and asks in one line whether to continue without them or connect again, instead of deciding alone. Tools already signed in on the machine, such as a logged-in gh, are fair to use for the task when it helps, said in one line; the earlier rule against them goes.

* fix(desktop): an untargeted connector status renders as a tool line, not a card

A manage_connections status call with no connectors list describes the whole catalog. The card rendered that as one Connect row per app the gateway knows, 59 of them, right after the user had connected the one they wanted. Such a part now takes the collapsed tool line like any other tool call, and resolves no owner and polls nothing. A status call that names apps keeps its rows.

* feat(desktop): a three-step tour explains the profile switch at handoff

The one-step signpost becomes a three-step tour when the first build's session is on screen: the profile rail, where the task runs on default and the welcome chat lives on the setup profile; the sessions list, which belongs to the selected profile; the rail again, one click from Hermes. The copy is in the locale files. The app runs it, not the guide: the tour bridge answers only the session the user is looking at, and the guide is a background session by then. Its own note tells it not to describe the tour. The tour no longer skips users who declined the look around, since the tour beat itself is mandatory in the next cut.

* fix(desktop): a rejected hidden submit never lands in the composer draft

A submit with displayKind hidden is machine text, a setup note the user never typed. When the gateway rejected one because the model's turn was still running, the composer restored it into the draft like any rejected message, and the user saw "[setup] links opened for gmail, googlecalendar" sitting in their input box. A rejected hidden submit is now dropped.

* fix(desktop): hold the links-opened note until the build session is idle

The card sent the hidden "[setup] links opened" note the moment it had opened the links, while the model's connect turn was still running; the gateway rejected it. In the live run the model had already called wait in that same turn, so the note had nothing left to say. The card now holds the note as pending on the session's state and delivers it only when the session is idle and the connect is still the newest connector part. When a newer part exists the note is dropped. The runbook says the same from the model's side: call wait right after connect; if it bounces as just minted, end the turn and treat the note as the cue to wait again; a note that arrives after a wait needs no reply.

* fix(desktop): the Start pill says one app, not one apps

* fix(desktop): parse connector dictionaries in wait results

Read connected connector slugs from dictionary and string results, excluding entries explicitly marked disconnected. Use the gateway response shape in the wait fixture so completed connections update their rows.

* fix(desktop): end first-build mode when the handshake settles

Persist a completed connection handshake after a settled wait or an accepted start message. Later connector parts use the ordinary card, and completed sessions cannot auto-open links or send a held setup note.

* fix(desktop): route external submits through busy handling

Steer external visible prompts during a running turn and queue them when steering is unavailable or rejected. Preserve hidden submit dispatch and leave the current draft intact so the start pill can hand delivery to the composer.

* fix(desktop): refresh first-build rows before the wait call

Poll connector status while the newest connect result contains initiated links, using the same watcher as waits. Check the active part before each request so replacing the card retires its earlier poll.

* fix(desktop): bound first-build connector polling

Stop polling after 150 seconds, a start decision, a completed handshake or three consecutive request failures. Ignore in-flight answers after those conditions so an expired watcher cannot overwrite the final rows.

* fix(desktop): preserve connected rows on unavailable status

Keep a confirmed connection when the catalog omits or disables its row or reports itself unavailable. Only unresolved rows become unavailable so a transient catalog answer does not reduce the start count.

* fix(desktop): treat an omitted connector action as status

Apply the gateway default when deciding whether a connector part is an untargeted status check. These catalog answers render as tool lines and do not resolve an owner or start a connection flow.

* fix(desktop): keep catalog checks from retiring live connectors

Exclude untargeted status parts when finding the latest connector action, including omitted actions. The card, poll and start pill now remain attached to the last actionable connector part.

* fix(desktop): connector catalog loading recovers, and disabled toolkits stay out of the picker

useConnectorCatalog (#108292) started in `loading` and returned early when
the session ids were missing, so the connectors card could sit on its
skeleton with Continue disabled and no request in flight. It now starts
`unavailable` until both ids exist, probes when they arrive, and gives the
gateway 15 s before it reports unavailable.

orderConnectorPicks drops rows the gateway marks `enabled: false`: a toolkit
the deployment turned off is not something the build chat can connect.

* fix(desktop): clarify first-build connection instructions

Filter connector picks against the offered apps before building the task runbook. Explain when active connections, early start messages, and skipped apps let the task begin so the model follows the card’s handshake.

* fix(desktop): show handoff steps only for visible panes

Require positive bounds and a visible pane before the handoff tour starts. Keep the two profile rail steps when the sessions list stays hidden so a collapsed sidebar does not suppress the tour.

* style(desktop): prettier on the connect-first files

Format the renderer TypeScript files in the branch comparison with Prettier and clear spacing warnings in the edited connector files. Keep the connection behavior changes in their individual commits so each defect remains reviewable.

* test(desktop): hold the connect-first tests back until the onboarding test pass

The onboarding tests return in one pass after the guide script rewrite (NS-853), not piecemeal in each behaviour PR. The eight files added on this branch are removed here and come back then; they are intact at 980172b9dd for that pass. The one existing test file this branch edits, the composer submit test, keeps its changes because the behaviour it covers changed.

* test(desktop): a busy hidden request steers like a visible one

After the stack onto #108292 the composer has one busy rule for every external request, hidden or visible: steer the live turn, queue when the steer is refused. The test that asserted the earlier rule (dispatch a hidden request while busy) now asserts the merged one, and a second case keeps the idle path: a hidden request the gateway rejects is dropped, never restored into the draft.

* fix(desktop): a hidden note mid-turn rides session.steer, never a user turn

A hidden request that landed while the model's turn ran went through the redirect path, which records a real user turn on the gateway and paints one in the transcript. The base branch's "[connectors] The user clicked Connect …" nudge showed up as the user's own bubble in a live run, and our "[setup] links opened" note would have done the same. The gateway already has the right primitive: session.steer injects text into the model's next tool result with no user turn. Both composers now expose it as onSteerHidden, and the composer's busy rule sends hidden requests through it. When the steer is refused the note queues with its hidden kind, the queue panel shows "Setup note" instead of the text, and the drain resubmits it hidden. Visible messages a button sends still redirect the turn as before.
2026-09-12 10:00:07 +05:30
brooklyn!
4f1966edac feat(desktop): connector cards that wait for the sign-in, and a guided first launch that holds together (#108292)
* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift

The label stays in the box, invisible, and the spinner is absolutely
centred over it, so a Connect or Approve button keeps its width while it
works instead of collapsing to a spinner. The approval bar had the same
thrash and moves onto it.

* refactor(desktop): one consent card for connectors and MCP setup

McpSetupTool rendered its own copy of the connector card's markup. It now
renders ConnectorCard for the pending question and ConnectorSummary once
settled, and the card gains what MCP needed: keyboard accelerators, a
source line, a question heading. The card also gets an avatar variant
(40px mark in the left gutter, text and buttons on one column) and a
collapseWhenSettled switch so a connector can stay a full card with a
green Connected pill in the action slot while MCP keeps its one-line
summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and
Spotify; Slack via Tabler because simple-icons dropped the mark.

* feat(desktop): connector card drives the agent through manage_connections wait

The offer used to end in a Continue in chat button, and the agent, seeing
an unconnected status, would improvise around the app. Now the card does
what the TUI does. Clicking Connect opens the browser and sends one hidden
line telling the agent to park in manage_connections action=wait for that
slug and to never call connect again (a second link cancels the one being
signed into). Not now sends its own line. A hidden request that lands
while the turn is busy steers it, or queues if the turn just ended.

Which call owns the live card changes too: consecutive calls naming the
same apps are one exchange (connect, the wait, the status that follows),
and the first of the last exchange is the card, so the agent's wait no
longer demotes the card mid-authorization and mints a fresh one below it.
A targeted ask renders one or two bare cards; only a real catalog gets the
header, search and refresh.

* feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them

The welcome chat knew connectors only as preferences to pick and wire up
later, so asked to connect Gmail it invented a Settings page that does not
exist. Both scripts now carry one rule set: status once, one batched
connect for every app named, the card is the ask so write a line and end
the turn, never route around a declined app with another client or
credential. The build handoff checks real connection status instead of
asserting none are connected, and the first task must be finishable, not
free of, the apps they picked. The connectors card explains what
connecting means and reports the count on its Continue button.

* fix(tools): resolve the Nous identity for share_auth profiles in the connector gate

A profile created with share_auth has no auth.json of its own and signs
in through the root store. Every other credential reader falls back to
the global root; the connector gate read HERMES_HOME/auth.json directly,
saw nothing, and stripped manage_connections from the profile's tool
list, so the welcome chat's agent truthfully reported the tool missing.
The gate now goes through get_provider_auth_state.

* fix(agent): name a provider retry backoff on the live status line

The retry status is buffered and replays only when every retry fails, so
during a 60s backoff after a 5xx the user saw a bare spinner. Right after
a connector sign-in landed this read as the agent going silent. The
backoff now also rewrites the live wait notice, which the desktop already
renders in the thread status row; it is transient and clears on recovery.

* test(desktop): connector rehearsal launcher and flagged connector spec

connector-rehearsal.mjs starts the real desktop and backend under a fresh
HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344,
so the onboarding connector flow can be driven end to end by hand or from
outside. The Playwright spec covers the flagged connector step.

* fix(desktop): send the agent back into wait when the user keeps waiting after a timeout

The card's Keep waiting re-entered the poll but the agent's own wait had
timed out too and nothing told it to go back in, so it would start
talking mid-authorization. keepWaiting now fires onWaiting like connect
does. Tests also pin that an expired or revoked grant asks the gateway
for reconnect, not connect.

* style(desktop): blank lines in connector-flow test per lint

* feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film

The intro is a one-time reveal, so anyone rehearsing the guided chat behind
it sits through it on every fresh HERMES_HOME. The flag rides the existing
launch-flags path (main → preload → renderer) next to guestOnboarding and
only gates isIntroRevealEnabled; the backend never sees it. The rehearsal
launcher sets it.

* fix(desktop): onboarding card Continue stays Done after the transcript rebuilds

The card kept its Done flag in component state. The hidden submit and the
turn-end hydrate both rebuild the message list, so the card remounted with
the flag false and Continue came back live, letting a step be answered
twice. The committed steps now live with the other onboarding answers,
keyed by step, and the first-build chip pick rides the same store.
remember_onboarding projects by key, so the new field never reaches USER.md.

* fix(desktop): no provider picker or free-tier chip over the guided first launch

Two sign-in surfaces leaked into the guide. A credential probe on the
setup profile (a free-tier token mid refresh, a session before its runtime
settled) hit requestDesktopOnboarding and dropped the provider picker over
the chat the user was in; and the statusbar free-tier chip sat there
offering a second sign-in the whole time. Both now yield while the gate
phase is cinematic, guided or handoff. The free tier is the provider for
those phases, and the guide offers sign-in on its own ready screen.

* fix(desktop): onboarding connector picks are real catalog slugs

The picker offered Spotify, GitHub and Stripe, none of which the deployed
connector catalog carries, and spelled Calendar and Drive with hyphens the
gateway does not use. A pick the build chat could not honour ended as
"Spotify isn't in the connector list" after the user had been told to
expect it. The list is now twelve slugs from the live status catalog,
spelled as the gateway spells them; GitHub is out (the terminal has git
and gh), chat channels stay on Messaging. Marks for the new entries; the
Google marks answer both spellings. The build runbook offers the picked
connections in its first turn rather than after the work is underway.

* fix(desktop): the free-tier ready screen never interrupts the guided chat

A readiness round fires when the layout pick assembles the window, and it
raised the free-tier ready screen over the conversation: the user was
dropped into the main app, dismissed it, and came back to a card they had
already answered. The guide is the introduction. The ready screen now
yields while the gate is cinematic, guided or handoff, and the notice is
acked the moment the guided chat takes the screen, not only when the film
does, so a skipped film no longer leaves it pending.

* feat(desktop): tour options that lead to building, and a fork that follows the tour

"Just the basics" and "Show me around" read as a click-through with no
exit; "I'll figure it out" read as declining help. Now Quick tour, Show me
everything, and Skip, let's build something. The script also folds the
fork into the same turn as the tour, so when the user closes the overlay
the next ask is already waiting instead of a transcript that ends on the
tour call.

* feat(desktop): the onboarding connector picker reads the live catalog

A hardcoded list, however carefully copied from today's catalog, is the
next drift. The picker now asks connectors.list through the same
session-owned RPC the connector cards use and offers exactly what the
gateway carries: a curated lead order puts the everyday apps first, chat
channels stay on Messaging, everything else is reachable by search. The
picks are gateway slugs, handed straight to manage_connections. No
catalog (toolset off, gateway unreachable) ends the step honestly with
Skip instead of inventing apps.

* test(desktop): the guided first launch never forces a sign-in

The acceptance criterion the guided onboarding was built to, as a test:
while the gate is cinematic, guided or handoff, the provider picker does
not open and a credential warning is dropped rather than deferred to the
next send. Outside the guide the picker opens as before. Red against the
tree before the guards landed (6 of 9).

* fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape

Closing the app during the guided first launch and reopening it booted the
normal shell around the persisted solo layout: the connecting splash, the
stock composer and model picker, a small window whose sidebars would not
open, while the gate still read guided. The gate now queues a kickoff for
the guided phase too (the kickoff adopts the existing guide chat by title),
takes the solo shape before the gateway opens rather than after, and the
connecting overlay yields to the guide's own opening. A typed reply in the
composer now closes an ask card and the first-build chips the same way a
click does; the layout card's Continue comes back Done.

* style(desktop): one answeredAfter helper for the ask card and first-build chips

* fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows

Between the film and the greeting the full-size shell painted for a beat:
finishIntroReveal showed the main window, then the kickoff shrank it once
the setup profile answered. The listener on the intro's hidden edge now
takes the guide's shape (solo layout + small centred window) synchronously,
so the window is already the guide when it is shown. One takeGuideShape
owns the pair; kickoff and the boot gate call it idempotently.

* style(desktop): the 'nothing connects yet' line reads first on the connectors card
2026-09-12 10:00:06 +05:30
Teknium
0dcadf6f41 revert: remove Collective Wisdom V1 (#94266)
Reverts the in-tree org skill-marketplace: hermes_wisdom package, three
model tools, CLI/gateway/desktop/dashboard/Telegram/Slack surfaces.

Later non-Wisdom work on shared files (guest onboarding i18n, dashboard
startup schema, Slack adapter, tui_gateway) is kept; Wisdom-only call
sites and config were stripped from those files.
2026-09-11 11:54:49 -07:00
emozilla
4e865d42d1 fix(desktop): restore flat glass surfaces without text bleed 2026-09-11 13:41:05 -05:00
joaomarcos
a863bbcc28 fix(desktop): make pool slot timeouts actionable 2026-09-11 06:23:18 -07:00
hermes-seaeye[bot]
efca39279a fmt(js): npm run fix on merge (#108214)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-11 12:53:17 +00:00
Siddharth Balyan
0591da2ba6 Guided first launch: review fixes from #107985 and the free-tier chip badge (NS-848, NS-855) (#108211)
* fix(desktop): centralize guide handoff receipt reads

Resolve the guide receipt key and value together in setup-profile. Use the helper at all four read sites so connection scoping follows one implementation.

* fix(desktop): recover from unreadable handoff receipts

Memoize receipt reads and show Retry only for the error phase. Quarantine corrupt data before retrying, and resolve the guide identity when the failed request did not retain it so a fresh build can start.

Cover preservation of corrupt data and removal from the active receipt key with an invariant test.

* fix(desktop): validate persisted onboarding phases from one list

Derive OnboardingPhase and persisted-value validation from the same phase list so future phases survive relaunch. Verify every persisted phase reloads and an unknown value falls back to idle.

* fix(desktop): share window centering arithmetic

Extract centeredBounds and use it for onboarding boot and window growth. Keep the existing work-area clamps and coordinate rounding unchanged.

* fix(desktop): compute progress steps inline

Remove the ineffective ProgressCard memo because streaming flushes replace the messages array. Keep the same transcript scan and rendered steps.

* fix(desktop): center the free-tier status chip detail

Wrap the model label and sign-in badge in an inline flex span with a shared gap. This centers the badge beside the model text without changing other status-bar details.

* fix(desktop): derive the guide receipt key in one place

The Retry path spelled the key derivation out again because the read helper throws on a corrupt receipt before it can return the key. A separate guideHandoffReceiptKey serves both the reader and the quarantine, so the derivation has one home again.

* fix(desktop): keep the free-tier badge at its intended leading

Badge declares leading-none, but the class merger drops it behind the size variant's font-size class, so the badge inherits a 1.5 leading and renders 16px tall next to an 11px label. That height, not the inline alignment, is what read as a detached badge. Restating leading-none on the chip's badge brings it to 11.6px, inside the label's cap height. The Badge component itself is left alone; every other badge in the app has the same dropped leading and that is a separate decision.
2026-09-11 12:46:48 +00:00
brooklyn!
aa05e5c0f4 fix(desktop): resolve side ownership before workspace registration 2026-09-11 06:12:12 -05:00
brooklyn!
6cc51b407c fix(desktop): keep sidebar tabs and restore controls clear of titlebar chrome 2026-09-11 06:12:12 -05:00
brooklyn!
477194879c fix(desktop): recover minimized sidebars from their existing controls
Co-authored-by: wukangcheng1994 <160389295+wukangcheng1994@users.noreply.github.com>
2026-09-11 06:12:12 -05:00
hermes-seaeye[bot]
8c74118c4a fmt(js): npm run fix on merge (#108127)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-11 10:23:35 +00:00
Siddharth Balyan
b8e8639445 Guided first launch: smaller code and review fixes over PR B1 (NS-848, PR B2) (#107985)
* refactor(desktop): compress intro reveal

Remove the unused inline cinematic fallback and its skip callback plumbing now that the native window owns playback. Keep native timing and exit behavior unchanged, colocate the spinner with text effects, and document the current launch and handoff contract.

Area delta against B1: 47 additions, 53 deletions, net -6 lines across six files. Most fork verdicts were already applied in B1.

* refactor(desktop): compress guided chat surface

Remove random greeting variants and retain one existing opener per locale,
while preserving the banked greeting and machine-name suggestion. Trim
assembly commentary while keeping the reasons for its layout invariants.

Move solo-boot and window-growth IPC into a topical Electron sibling so
onboarding handlers no longer grow main.ts. Preserve sender gating,
reveal ordering and window geometry.

* refactor(desktop): compress onboarding handoff

Split welcome-chat kickoff from durable handoff effects and wire each
hook directly. Remove duplicated option types, a redundant readiness
comparison, nullable receipt-key state and stale prose while preserving
B1 routing and recovery.

* refactor(desktop): compress guided chat back half

Remove the duplicate handoff completion key and its reader/writer helpers.
Use the onboarding phase record for replay guards and settled cards while
keeping accepted receipts as the completion boundary.

Preserve the signpost and plugin plan under ruling 5.

* refactor(desktop): compress stores and transcript integration

Remove unused machine reset and untargeted host composer submission. Trim machine and presence commentary while preserving their live consumers. Wire reasoning through the existing scratchpad surface and memoise progress history without mutating it.

Keep parser and connector rendering under rulings 4 and 5. Area delta: 36 insertions, 101 deletions; net -65 lines.

* fix(desktop): keep skipped onboarding apart from a completed handoff

Make skipGuide() persist skipped and let beginOnboardingHandoff accept
guided or skipped. requestSetupHandoff and HandoffCard derive completion
from the accepted receipt.

The latch merge conflated skipping the guide with starting the first
build. A later handoff therefore claimed "was started" without creating
a session. Preserve skipping as its own terminal phase so a later
handoff can create the build and reach done only after acceptance.

* refactor(desktop): B2 review notes

Correct the layout-growth comment in assembly.ts: growing to preserve
the chat size balloons the window. Restore the WHY clauses in
onboarding-handoff.ts and onboarding-kickoff.ts for pending title metadata
on older backends and the caller's requestGateway reading the create pin.

Move guideSourceConnectionId beside $setupSession in setup-profile.ts
and derive each hook's option types from useSessionActions, so kickoff
no longer imports the heavier handoff leg.

Move $handoffError and retrySetupHandoff beside $setupHandoff in
setup-profile.ts and update the card and handoff hook importers.
This removes the setup-profile/handoff-receipt cycle and leaves receipt
persistence dependent only on storage and its receipt type.

* fix(desktop): derive the first-build receipt key one way

Use guideSourceConnectionId(guide.storedId) for the save and request
receipt keys, matching resume and HandoffCard. Keep guide.connectionId
for RPC routing and preserve the receipt key's string format.

At boot the resume path only knows the guide's stored id. When no owner
hint exists but an active gateway connection does, keying writes by the
resolver's ambient route hides the accepted receipt from resume and
leaves the card on Opening. One derivation lets every path find the same
receipt without changing where the build request is sent.

* feat(desktop): intro type at 150% for legibility

Set the intro window's root font size to 150%. Every measure in the intro
is in rem, so the chat card, its rows, bubbles and gaps scale together.
The brand close uses viewport units; its wordmark and tagline are scaled
by hand to match (6.8 to 10.2 vmin, 1.35 to 2 vmin). The hero card's
width cap rises from 900 to 1350 px so lines keep their length on a
large display; its minimum width is unchanged so the three-column stage
still fits a laptop.

The intro fills a display the user sits back from, and at the app's 16 px
root its text read too small on a large monitor (director ruling).

* fix(desktop): status bar keeps one fill under glass; free-tier chip reads Nous, model, Sign in

Under the glass appearance in sidebar scope, the body paints a hard stop
at the rail's edge (glass mix left, opaque chrome right) and the status
bar was transparent, so the seam ran through the bar and cut whichever
item sat on it: in the 886 px guided window, the free-tier chip. The bar
now belongs to the opaque content column across the full width, the way
Finder's does; window scope has no seam and keeps the transparent bar.

The chip itself read "Nous · free tier · nous/welcome" with the Sign in
badge touching the label. It now reads "Nous", the model id small and
monospace, then a solid Sign in badge set off by a gap; the full
"Nous · model" string moves to the tooltip. "Free tier" is no longer
said in the bar (director ruling).
2026-09-11 15:45:48 +05:30
Siddharth Balyan
0e927c914d Guided first launch behind HERMES_GUEST_ONBOARDING: intro, guided chat, first task in default (NS-848, PR B1) (#107958)
* feat(desktop): port guided onboarding substrate

Add seeded session creation, transcript directives, profile routing, and the shared window and pane primitives needed by the guided flow. Keep later-step mounts deferred and exclude provider selection and retry machinery.

* refactor(desktop): anti-slop cleanup for substrate

Assemble seed parameters in the existing create helper and use the owning transcript attribute type. Read the guaranteed gateway and connection contracts directly to remove runtime type probes and unchecked assertions.

* test(desktop): create-overrides invariants

Verify that reasoning and title overrides do not select a provider or model. Empty overrides and seeds add no parameters.

* feat(desktop): port first-run cinematic window

Play the cinematic behind the guest onboarding launch flag using bundled Collapse and JetBrains Mono. Give the native window its own controller and restore the app on skip, renderer deadman or native watchdog.

Drop the perf scenario because it depends on the removed replay hook. Guided chat kickoff and app-shell gate wiring remain with their later steps.

* refactor(desktop): anti-slop cleanup for cinematic

Preserve audio and canvas behavior through named types and inferred results. Split the viewport node and frame drawing to keep control flow bounded. Cut comments that only repeat the code.

* feat(desktop): add onboarding gate and answers stores

Track cinematic, guided chat, handoff and completion in one phase record. Queue the guide after the intro and share pending kickoff work between callers.

Keep existing saved answers while dropping retired preferences. Leave intro seen-state ownership with the cinematic store.

* feat(desktop): port guided onboarding chat

Add guided setup cards, runbooks, machine context, and onboarding presence. Connect transcript rendering and first-build progress to the desktop behind the onboarding flag. Leave session kickoff and handoff execution for the next step.

* refactor(desktop): anti-slop cleanup for guided chat

Keep directive and layout lookups typed. Remove unsafe test casts and isolate onboarding transcript calculations without changing the flow.

* feat(desktop): connect guided onboarding to durable first-build handoff

Start the guide only after its profile backend confirms bootstrap readiness. Seed or adopt the welcome chat, then transfer the first build to default with a durable receipt and explicit retry.

Wire cinematic completion, screen stand-down, layout growth and progress check-ins. Save agreed preferences before creating the build and release prompt slots after storage refusal.

* refactor(desktop): anti-slop cleanup for onboarding handoff

Reuse the gateway request and error contracts. Isolate guide adoption and snapshot validation while preserving receipt recovery and reasoning overrides.

Validate persisted receipt fields at the JSON boundary without coercion. Keep corrupt identities rejected and retain only the permitted test mocks.

* fix(desktop): guided chat review fixes

Wire the native machine probe so guided setup can suggest a name and offer the right first task. Restore the comments that explain the flow boundaries.

The directive registration uses the launch flag to preserve ordinary chat. Ruling 6 folds active.ts into assembly to keep activity ownership together and removes the second greeting source so the seeded and visible greetings agree.

* fix(desktop): handoff review fixes

Probe the guide backend before switching profiles so a readiness refusal keeps classic onboarding on the current backend.

Restore list-valued personalization coverage and routing rationale. Remove the obsolete setup status fixture.

* chore(desktop): onboarding script cull and rehearsal recipe

Document a temporary-state rehearsal using the existing onboarding flag and optional portal stand-in. Keep the main scripts unchanged and retain window growth for the guided chat.

* fix(connectors): reject incomplete catalog responses

* feat(gateway): scope connector controls to the owning session

* feat(desktop): connect apps through native session-owned controls

* feat(desktop): gate connector cards and enable free-tier access

Use the launch flag before mounting connector controls so classic transcripts add no status requests. Allow existing free-tier identities through the read-only tool gateway gate and test the owning-profile RPC path with A’s launch gate. Keep authorization links out of previews.

* style(desktop): format connector translations

Apply Prettier to the connector copy blocks while preserving upstream translations and free-tier wording.

* refactor(desktop): anti-slop cleanup for connector card

Use the transcript JSON contract and concrete RPC parameters. Preserve malformed-value filtering at one string boundary and make the fixture and row types explicit. Keep connector execution and cancellation behavior unchanged.

* feat(desktop): detect initial language from the OS

Use the native machine locale when no supported language is saved. Preserve explicit choices and leave inferred languages out of config.

* refactor(desktop): anti-slop cleanup for initial locale detection

Keep unvalidated config values at the existing validation boundary. Pass no saved choice after that boundary has ruled it out, preserving locale precedence.

* test(desktop): onboarding port test set

Make native window tests reject duplicate IPC handlers and isolate disabled onboarding. Assert the active gate mock when onboarding re-enables.

Keep the test set limited to behavior carried by the port.

* fix(desktop): recover failed guide kickoff and reveal once

The review found that a failed guide create stranded the solo shell and draft profile, and solo boot faded an already visible window a second time. Restore the prior route and layout, release onboarding through its existing phase record, and surface create failures. Let the film own the reveal while solo boot animates the visible resize.

* fix(desktop): preserve transcript ownership across cards and handoff

The review reproduced answers submitted to the focused chat, repeated questions disabled across sessions, handoff recovery using foreground identity, and mount-dependent progress history. Target each card’s own composer, scope settlement to its message and session, carry the issuing guide through handoff, and derive progress from its transcript with streaming activity. Reuse the existing owner ladder for exact and profile-only routes.

* fix(gateway): preserve connector ownership with profile routing

The review found that shared-primary profile metadata was rejected before connector dispatch, while desktop controls treated a missing registry id as missing ownership. Accept profile only as routing metadata and keep the live transport as authorization. Resolve card ownership through the existing exact/profile ladder, retaining ambient routing only for the single-backend case.

* fix(desktop): resolve plugin roots and gate the Basic layout

The review found that the first plugin build was seeded with a different installation’s fixed path, and the director ruled that flag-off layouts must match main. Resolve the running desktop’s plugin root before seeding a plugin build and register Basic only when onboarding is enabled. Keep the runbook wording and the ordinary four layout presets intact.

* fix(desktop): clear review-fix slop findings

The slop gate flagged an undocumented layout-data assertion and unknown-return types in the new test selectors. Record the layout registry invariant and preserve each selector’s return type. The only remaining production finding is the accepted connector-tools baseline.

* fix(desktop): detect the OS language on a fresh install

The review found that the merged English config default prevented the
desktop from probing the OS language on a fresh install. Add an opt-in
saved-values read so an absent choice remains distinct from saved English.

Preserve default-valued English only for explicit language saves; unrelated
settings saves must not turn a merged default into a language choice.
Older backends ignore the new query options and keep returning merged
English, preserving their existing desktop behavior.

* test(desktop): make the flag-off layout registry test deterministic

The flag-off test awaited the full controller import, pulling in the UI
graph and installing application watchers just to read layout presets.
That import took 9.5 seconds locally and timed out in the director's run.

Move the existing trees and registration into a small layout-presets
module. Production and the synchronous test use the same flag-gated
registration, without starting the controller in the test. Keep the real
registry invariant and dispose the test's contributions after completion.

* fix(desktop): keep the transcript parser and ::ask behind the onboarding flag

Register the guided chat's question card only with onboarding enabled.
Restore main's whole-paragraph parser and contribution rendering when the
flag is off, including its streaming prose behavior. Keep segmentation for
the guided flow until B4 decides the parser's wider use.

Restore main's two parser test files so its existing product and plugin
contracts remain the flag-off check.

* test: drop the onboarding and connector tests pending a later ticket

Apply the director's ruling to remove B1's added test files and restore
main's existing suites. Keep only the gateway route-reader mock contract
that main's profile tests need against the shipped activation behavior;
their cases and assertions stay intact.

The flow's shape is not settled and B3/B4 rewrite it. The connector layer
will also be reworked. The live CDP run is the flow check until a follow-up
ticket brings tests back.

---------

Co-authored-by: brooklyn! <brooklyn.bb.nicholson@gmail.com>
2026-09-11 15:45:43 +05:30
shannonsands
a6ee31f55a feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation

* feat(wisdom): add private contribution loop

* feat(wisdom): add managed consumption workflows

* fix(wisdom): close cross-repository safety gaps

* fix(wisdom): align local package and lifecycle policy

* fix(wisdom): require explicit profile setup

* docs(wisdom): repin reconciled gateway head

* fix(wisdom): fence content downloads and approval receipts

* docs(wisdom): record generation-fenced downloads

* docs(wisdom): record unified delivery PR

* fix(ci): stop passing invalid classifier inputs

* docs(wisdom): remove internal requirements ledger

* feat(wisdom): localize dashboard and desktop copy

* feat(wisdom): complete local contribution and consumption UX

* style(wisdom): satisfy desktop lint

* chore(wisdom): refresh requirements pin

* test(dashboard): allow formatted profile copy

* test(wisdom): stabilize desktop interaction coverage

* fix(wisdom): surface dashboard action failures

* fix(wisdom): add repeatable Portal demo login

* feat(wisdom): add actionable skill notifications

* feat(wisdom): add notification install and update actions

* fix(wisdom): make Telegram skill alerts actionable

* fix(wisdom): always refresh demo Agent login

* feat(wisdom): embed Telegram notification actions

* fix(wisdom): preserve Telegram notifications after actions

* fix(wisdom): keep Telegram notification cards readable

* feat(wisdom): add Telegram candidate approval flow

* feat(wisdom): explain Telegram qualification reasons

* fix(wisdom): reconcile cross-surface candidate actions

* feat(telegram): add Collective Wisdom management command

* chore(wisdom): refresh Gateway contract pin

* chore(wisdom): advance Gateway contract pin

* feat(wisdom): align command UX across clients

* feat(slack): add Collective Wisdom management parity

* feat(wisdom): add security and professionalism reviews

* feat(wisdom): add first-time qualification guidance

* feat(wisdom): simplify qualification sharing choices

* feat(skills): add optional editorial metadata

* feat(wisdom): enrich legacy skill presentation

* fix(wisdom): harden review and update boundaries

* fix(wisdom): emit canonical review timestamps

* fix(wisdom): align with merged gateway and main

* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)

- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
  7-day evidence builder that excludes bundled/hub/managed skills and
  dismissed/handled/recently-suggested content hashes, strict pydantic
  schemas for agent output with repair-or-reject, fixed copy templates
  (Share / Teammate / Published / Update / Mute), idempotent retried
  delivery ledger with stale-action resolution, weekly review job,
  resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.

* wisdom: agent-led renderers and button action dispatcher

- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
  is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
  ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
  packaging flow, Install/Update -> plan command. Never publishes/installs.

* wisdom: CLI verbs, agent_led config default, conversational catalog skill

- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
  verbs, share/install flows and fixed notification templates.

* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons

- gateway housekeeping tick calls maybe_run_weekly_review with a home
  channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
  duration keyboard, send_wisdom_agent_recommendation rich card + fallback.

* fix(wisdom): integrate local mediation and harden model and setup boundaries

* fix(wisdom): honor authoritative recommendation policy and defer on failure

* fix(wisdom): synchronize opaque suppression and recheck delivery preferences

* feat(wisdom): route weekly selection through the session-owned assessment queue

* fix(wisdom): prepare and submit the reviewed generated share package

* feat(wisdom): separate native Share preparation from publication consent

* feat(wisdom): sync native mute choices through a leased preference outbox

* feat(wisdom): bind native mute controls to durable preference choices

* feat(wisdom): add scoped desktop and dashboard notification settings

* fix(wisdom): revalidate feed recommendations before assessment and delivery

* fix(wisdom): persist validated delivery receipts before completing notices

* feat(wisdom): add private notification claim and receipt client

* Persist Wisdom send reservations and recover delivery acknowledgements

* Route legacy Wisdom controls through current native review

* Add typed private Wisdom operation outcome client

* fix(wisdom): make agent-led advice usable in the local demo

* fix(wisdom): keep requested consent outside proactive limits

* fix(wisdom): distinguish unavailable assessments and preserve digest text

* fix(wisdom): assess ongoing usefulness beyond the current task

* fix(wisdom): restore immediate qualification sharing controls

* fix(wisdom): separate qualification review from installation advice

* fix(wisdom): collapse review checklists and simplify sharing copy

* fix(wisdom): show compact sharing progress and publication receipts

* fix(wisdom): require credential prefixes rather than matching skill names

* fix(wisdom): finish package checks before presenting sharing consent

* fix(wisdom): scan local skills before qualification cards

* fix(wisdom): update moderation results on existing sharing cards

* fix(wisdom): keep sharing review accessible from receipt cards

* fix(wisdom): align mediated review cards and collapsible checks

* fix(wisdom): clarify clean security summary wording

* fix(wisdom): normalize consent plans and add explicit recheck

* fix(wisdom): keep install and update receipts concise

* fix(wisdom): collapse assessments and deduplicate operation cards

* fix(wisdom): restore private Portal review from native cards

* fix(wisdom): sync Portal publication to original consent card

* fix(wisdom): show local skill version on sharing cards

* fix(wisdom): skip agent recommendations for self-published versions

* fix(wisdom): simplify candidate notices and local-edit recovery copy

* feat(wisdom): submit locally reviewed packages with one confirmation

* feat(wisdom): expose safe receipt and outcome sync recovery

* wisdom: onboarding notice says detect and share, names the user's own skill

Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark

Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.

* wisdom: one opener, no approval line, ask to share after the skill is shown

Product owner review of the candidate card.

- The Hermes written card now opens with the same sentence as the fixed card
  ("Your organisation has enabled Collective Wisdom, a feature designed to
  automatically detect and share useful skills across all team members.")
  instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
  and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
  It is now the last line, after the skill name, description, why suggested
  and the checks, and reads "Would you like to share it?" (matching the
  agent led template wording).

Tests updated for the new order; proposalNotice removed from all desktop locales.

* wisdom: American spelling, organization

Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.

* wisdom: candidate card copy round 4 (owner review)

Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:

1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
   the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
   card (Telegram rich card and plain fallback, legacy agent-led share
   template).
3. The skill name and description are labelled: "Skill name: <name>" and
   "What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
   inappropriate content found)" with no per-check bullets and no "Pass";
   a failed review reads "Needs a look before sharing at work (possible
   inappropriate content)" and lists only the checks that flagged
   something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
   details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
   days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
   editorial_name, a simple one_line_description and a compelling
   why_coworkers_benefit under 300 characters; "Be concise and
   convincing." becomes "Be concise and compelling: the goal is that the
   user wants to share it."

Tests updated for the new strings; review_text() gains direct coverage.

* wisdom: re-apply owner copy after rebase

- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice

* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors

Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.

* fix(wisdom): reconcile optional SDK tests and frontend lint

* fix(wisdom): default to agent-written notification summaries

* fix(wisdom): restore deferred install review and browse controls

* feat(wisdom): inspect installed setup with exact package provenance

* feat(wisdom): run native-approved installed setup steps with durable evidence

* fix(wisdom): recover interrupted setup with explicit native consent

* feat(wisdom): hand native installs into guided setup review

* fix(wisdom): continue requested setup with fixed notification copy

* fix(wisdom): preserve setup while waiting for a session model

* fix(wisdom): expose canonical setup review controls on desktop

* fix(wisdom): resume setup after recorded automatic updates

* fix(wisdom): make missing setup prerequisites recheckable

* chore(wisdom): align Agent with verified Gateway contract

* fix(wisdom): stop guessing team slugs in portal links

* fix(wisdom): retire pending advice on account sign-out

* fix(wisdom): cancel advice after terminal account revocation

* fix(wisdom): fence feed responses across account sign-out

* fix(wisdom): checkpoint signed-out feed before reactivation

* fix(wisdom): link proactive advice to scoped notification settings

* fix(wisdom): coalesce queued publication recommendations by version

* fix(wisdom): keep package review navigation local and deferable

* fix(wisdom): reflect installed state in discovery controls

* fix(wisdom): show exact checks before command confirmation

* chore(wisdom): pin bounded analytics privacy contract

* chore(wisdom): pin retired legacy notification contract

* feat(wisdom): review publisher usage with exact sharing copy

* fix(wisdom): align discovery and review check summaries

* fix(wisdom): show expired consent before confirmation

* fix(wisdom): require fresh review for legacy install controls

* fix(wisdom): preserve review expiry across check toggles

* fix(wisdom): retain update policy in native install reviews

* fix(wisdom): surface failed native card edits

* fix(wisdom): persist local command approval reviews

* fix(wisdom): use saved approvals for messaging commands

* test(wisdom): provide scan result in setup handoff fixture

* test(wisdom): exercise Telegram approvals with saved review state

* fix(wisdom): retain suppression policy for offline deferral

* fix(wisdom): reconsider candidates after deferred suppression expires

* fix(wisdom): bind review checks and report verified readiness separately

* fix(wisdom): persist accepted publication intent and recover exact outcomes

* fix(sync): pin UTF-8 tree ordering across writers

* chore(wisdom): pin organisation-scoped Gateway authorization

* fix(wisdom): restrict consent delivery to user-facing sessions

* chore(wisdom): refresh reviewed Gateway contract pin

* fix(wisdom): preserve kept tools in Blank Slate exclusions

* test(auth): reset anonymous fixture with a profile-scoped cache

* fix(wisdom): gate local surfaces and work on current profile entitlement

* fix(wisdom): invalidate quiet tool cache on entitlement changes

* test(wisdom): authorize local consent gateway fixtures

* fix(wisdom): keep entitlement decoding free of native crypto imports

* test(wisdom): provide local entitlement to demo CLI subprocess

* ci: leave upstream workflow unchanged in Wisdom PR

* fix(wisdom): ship package and contracts in Nix wheels

---------

Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
2026-09-11 19:04:06 +10:00
hermes-seaeye[bot]
754dd430ce fmt(js): npm run fix on merge (#108028)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-11 07:32:00 +00:00
Siddharth Balyan
4bdd64b334 The free tier is created in one place, at boot, only behind HERMES_GUEST_ONBOARDING=1 (NS-847) (#107697)
* fix(auth): close the free tier's gaps against the gateway's welcome-tier contract

The inference gateway's welcome tier (NousResearch/api DOCS/anon-tier/plan.md) serves an
anonymous account exactly one model on its own host, refuses everything else with a structured
429, cross-refuses a request on the wrong host with a 400 (403 while the tier is dark), and
tells a signed-in account that still asks for `nous/welcome` what to switch to in an
`x-nous-model-switch` header. Four client-side gaps against that contract:

- Auxiliary calls were refused on every session. The auxiliary client asked the welcome host
  for the Portal's recommended compaction/vision model, a guaranteed 429 `model_not_free`
  before each fallback. On the welcome host it now uses `nous/welcome` (its backing model
  covers auxiliary work) and skips Nous for vision, which the welcome model does not take.

- The structured 429 body was never read. The classifier now parses `reason` /
  `retry_after` / `alternates` / `upgrade_url`: `model_not_free` and `feature_not_free` are
  non-retryable gates that fall back; `at_capacity`, `admission_closed` and `rate_limited`
  are rate limits that honour `retry_after` and never rotate the free tier's only credential.
  The wrong-host 400 and the dark-tier 403 are deterministic, so they abort this route and
  fall back instead of retrying or re-exchanging. The terminal paths say what happened and
  name the sign-in (`/login` in a chat, `hermes auth upgrade` in a terminal).

- The `x-nous-model-switch` header was ignored. The chat-completions transport records it
  beside the rate-limit and credits headers; the next call moves the session, and the config
  default when it still names `nous/welcome`, to the backing model the gateway named.

- A guest fell back to the paid host. With `inference_base_url` absent from the exchange or
  outside the host allowlist, routing defaulted to inference-api, where every request is a
  400. A guest now defaults to the welcome literal at the exchange, in the shared store's
  shape, and in effective routing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit fc758aad7efceff6223fc144a9b5c69f13e41bd8)

* feat(auth): the free tier is set up on request; nous.guest_setup decides whether also on first use

A caller that names nous/welcome on a Nous route with no Nous identity in reach — the guided
setup's session (provider=nous, which skips the resolver's nothing-configured rung), the free-tier
picker row, a bare --provider nous pointed at it — is asking for the free tier. The OAuth runtime
rung now sets it up there instead of failing "not logged in", so the guided chat no longer races
the root profile's first-run mint.

nous.guest_setup is the policy seam: "auto" (default) keeps today's first-use setup wherever
nothing else is configured; "on-request" mints only when the free tier is asked for by name
(nous/welcome, /login, hermes auth upgrade, replacing a retired identity). Implicit callers —
the resolver's last rung, the first-run check, free_tier.status, the CLI's background setup, the
connector token path — still adopt what the shared store holds, so every profile follows the one
identity the guided setup created, but never create one on their own.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit ae915ddc65ecdb81b81e29b604671d15cd49233c)
(cherry picked from commit 62ad1ff3ab200ea064975a32c502041b25910165)

* feat(auth): the guided setup provisions the free tier explicitly; nous.guest_setup is auto | explicit

Two questions govern the free tier: may it exist (nous.guest) and who may CREATE the identity
(nous.guest_setup). "auto" (default) keeps today's first-use setup wherever nothing else is
configured. "explicit" means Hermes never creates one on its own: the only creator is the new
provision_free_tier() primitive, exposed as the free_tier.provision RPC, which the guided setup
on Hermes Desktop calls as its first step — on the root gateway, before the setup profile and
before the guided chat exists — so the identity lands in the root store every profile reads
through and is there before any session asks for nous/welcome. That closes the race against the
backend's own setup, and makes "only when the setup-bot flow is used" literally true.

The earlier "on-request" tier is replaced: it minted whenever any caller named nous/welcome
(the hermes model row, --provider nous), which treated a model name as intent and was broader
than the guided setup. Under "explicit" a nous/welcome request with no identity fails "not
logged in" as before the free tier existed, and /login or hermes auth upgrade report nothing to
sign in from. Implicit callers still adopt an identity the shared store holds, and a retired
credential is replaced (a continuation, not a creation).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit c63d2c935c1e59016164fdfb90cf70b4094466a0)

* fix(auth): remove the nous.guest_setup knob; the free tier is created on first use

`nous.guest_setup: auto | explicit` decided who may CREATE the free-tier identity. Under its
default every line it added was inert (`may_mint` always true), nothing in tree set `explicit`,
unknown values read as `auto`, and under `explicit` a CLI-only install could never get an
identity, which contradicts the first-run contract (first command mints, then chats).

The mint race the knob accompanied is already benign: every caller takes the profile lock then
the shared-store lock, and the loser adopts what the winner wrote. What makes the guided setup
win deterministically is `provision_free_tier()` behind the `free_tier.provision` RPC, which
stays. `nous.guest` remains the only free-tier policy.

Removed: `guest_setup_policy()` and its constants, the `explicit=` / `may_mint=` threading through
`ensure_portal_identity` and `_reconcile_and_provision`, the flag at the three replacement call
sites (now no-ops), the config default, the docs section, and the four `guest_setup` test-config
entries. The three policy tests that hold regardless of the knob are kept under
`TestExplicitProvision`; the two that only tested the knob are deleted.

(cherry picked from commit d8a50526d93c374c0067dd935b5a65055e0af261)

* fix(gateway): a server-driven model switch off nous/welcome does not evict the cached agent

When a signed-in account still asks the paid host for `nous/welcome`, the inference gateway
serves the current backing model and names it in `x-nous-model-switch`. `apply_model_switch`
moves the live session to that model and moves `config.yaml`'s default off the alias in the
same step. The messaging gateway's fallback-eviction check compares the agent's model with the
config default and evicts on any mismatch that is not a /model override, so when the config
write did not land (unreadable config, lock) the cached agent was evicted once per turn, and
prompt caching with it.

`apply_model_switch` now stamps the alias it moved the session off on the agent, and
`_is_intentional_model_switch` treats "agent moved off the alias the config still carries" as
deliberate, beside the existing /model override case. The check takes the agent and the config
model instead of a bare model string; its one caller in `_run_agent_evict_on_fallback` passes them.

(cherry picked from commit 696d1ec86b69db28bf002c841e9389b85178a954)

* fix(auth): the free tier outranks implicit host credentials in provider resolution

On a fresh install with a leftover ~/.aws profile, resolve_provider("auto")
reached the Bedrock rung before the free-tier rung, so the first turn ran on
Bedrock and failed 403 while the free tier was still being minted in the
background at agent setup (NS-829). Live on a Mac with ~/.aws present: 28 s,
three retries, no answer; the next process then switched to nous/welcome.

The free-tier rung now sits directly above the Bedrock chain: when nous.guest
is on, an existing free-tier identity answers, else a blocking mint runs, and
only then does the boto chain get a say. Everything above is unchanged and
still wins: CLI creds, config.yaml model.provider, env keys, the OpenRouter
pool, a logged-in active_provider. nous.guest: false skips the rung, and a
failed mint still falls through to Bedrock and the no-provider guidance.

Tests: six precedence cases (identity present, fresh mint, free tier off, env
key still wins, sign-in still wins, failed mint falls through). The opt-out
test now neutralizes the AWS chain like the precedence tests do; on a machine
with ~/.aws it was failing for the same reason as the bug.

Live after the fix, same Mac, AWS credentials visible, isolated shared store:
identity minted 2 s in, turn on model=nous/welcome provider=nous, answer in
11 s.

(cherry picked from commit a04b05260cd334dd7199ad9b6cd5b2538364c75a)

* fix(auth): review follow-ups for the free-tier rung (NS-829)

- tests/agent/test_bedrock_integration.py: the Bedrock auto-detect test switches
  the free tier off; its contract is the boto chain, and the free tier now
  sits above it.
- gateway/run_notifications.py: the free-tier startup line reads auth.json
  before consulting the resolver, so a gateway boot on a machine with AWS
  credentials never mints or refreshes over the network.
- hermes_cli/anon_auth.py: module docstring says where the free tier sits in
  the ladder instead of "the ladder is untouched".
- tests/hermes_cli/test_provider_precedence.py: two invariant tests instead of
  six (parametrized ladder cases; a failed mint that returns None or raises
  falls through to Bedrock).

scripts/run_tests.sh on the five affected files: 147 passed, 0 failed.

(cherry picked from commit 10790d148c60ada11b9ecdde2cd2c836c6a82a11)

* feat(auth): HERMES_GUEST_ONBOARDING=1 is the one launch gate for the free tier; HERMES_FORCE_GUEST is gone

The free tier is pre-GA. Until GA it must not exist for anyone who did not
ask for it: no identity minted, no portal traffic, no free-tier copy on any
surface. One environment variable now decides that, and one function reads it.

`guest_enabled()` returns False unless `HERMES_GUEST_ONBOARDING` is exactly
"1"; only then does `nous.guest` (the user's off switch) get consulted. Every
free-tier site already funnels through `guest_enabled()`, so the gate closes
minting, routing, connector entitlement, status lines and the picker row in
one place. With the variable unset, `resolve_provider("auto")` on a fresh
install raises `no_provider_configured` exactly as upstream does.

`HERMES_FORCE_GUEST` and `force_guest_mode()` are removed. They inverted the
gate (forced the tier ON over `nous.guest: false`), their "new" value re-minted
identities as a side effect of provider resolution, and `_has_any_provider_
configured` read them ahead of every other check, making the CLI a second
reader of a flag that must have exactly one. `_forced_new_done` and the
`force` parameter of `_reconcile_and_provision` go with them.

Supersedes the dev lever introduced in fcf9d11679 (rung 1) and hardened in
b5c162c3ec. Ruling: NS-845 Q1.1 (recorded on NS-847).

Not a user preference: the variable is never written to config.yaml or .env
and never shown in setup. It is deleted at GA together with its comment in
anon_auth.py. This is a deliberate, temporary exception to the "no new
HERMES_* env vars for non-secret config" rule.

Tests: fixtures set the gate instead of deleting the old lever; one new
invariant (`test_launch_gate_off_means_no_free_tier_at_all`) proves that "",
"0", "true" and "new" all leave the tier off with zero portal calls, red on the
previous commit. The `HERMES_FORCE_GUEST=new` re-mint test is deleted with the
feature.

* feat(auth): the free-tier identity is created in one place, at boot; every other site is a read

Before this commit eight sites could create a Nous free-tier identity as a
side effect of something else: resolving a provider, the CLI's first-run
check, the CLI's session setup (in the background beside an own key), a
connector bearer read, the desktop polling `free_tier.status`, the sign-in
precondition, the desktop's `free_tier.provision`, and the dead-credential
re-mint. A poll could mint. Provider resolution could hit the network. Two
of them raced each other on a fresh install.

Now `hermes_cli/free_tier_bootstrap.py::run_bootstrap` is the only creator.
`hermes serve` runs it on a daemon thread from `_lifespan` beside the other
background boots; `cmd_chat` runs it synchronously before the first-run
guard. It inventories credentials first (`resolve_provider("auto",
skip_free_tier=True)`: what would carry inference if the free tier did not
exist), creates the identity only when `guest_enabled()`, resolves inference,
records a `SetupRecord` in process memory and broadcasts ONE `setup.ready`
event. It runs on every boot; only the mint is gated.

`ensure_portal_identity` now requires `explicit=True` and raises otherwise.
Its callers are the bootstrap, the desktop's `free_tier.provision` (the
explicit retry when the boot could not create the identity) and the two
dead-credential replacements (`auth_nous.resolve_nous_runtime_credentials`,
`managed_tool_gateway._replace_dead_guest_token`). The background thread
path and `provision_free_tier` are deleted with their last callers.

Reads that used to mint and now only read: `auth.py::resolve_provider`
rung 7 (an existing identity still outranks the Bedrock chain, NS-829
ordering kept), `main.py::_has_any_provider_configured`,
`cli_agent_setup_mixin._ensure_runtime_credentials`,
`managed_tool_gateway.read_nous_access_token` (no identity -> None),
`anon_sign_in.run_sign_in` (no identity -> Unavailable),
`methods_free_tier` `free_tier.status`.

`setup.status` answers from the record for the launch profile, blocking up
to 8 s while the bootstrap is in flight so a client's first poll lands after
the identity exists rather than racing it; a named profile, or a process
that never ran the bootstrap, keeps today's live probe. The record's fields
ride along additively (`ready`, `free_tier`, `other_providers`,
`inference_provider`).

Identity and inference are decoupled (NS-845 Q1.3): the mint sets
`active_provider="nous"` only when the inventory found nothing else usable
(`_mint_locked(carries_inference=)`); an adopted account always does. A token
refresh no longer re-elects the provider it refreshed
(`_save_provider_state_to_source` writes credentials, not the user's
choice) — that write was how an own-key install ended up on the free tier
after the first connector call.

Supersedes the mint sites in fcf9d11679, a42d0748fc (first-run check),
bbbaa8935a (CLI background setup), 0179efc989 (`free_tier.status` mint),
62ad1ff3ab / c63d2c935c / d8a50526d9 (the `nous.guest_setup` knob and
`provision_free_tier`), and a04b05260c (blocking mint in the resolver).
Ruling: NS-845 Q1.2 + Q1.3, recorded on NS-847.

Tests: `TestBootstrapIsTheOneCreator` (one mint per process; own key keeps
inference; reads never reach the portal; a refused mint is memoised),
`free_tier.status` fails loudly if it ever calls the creator, the resolver
stub fails loudly if resolution ever mints, `setup.status` reads the record,
`skip_free_tier` proves the inventory question. The three sign-in tests for
the deleted pre-mint collapse into one (`no identity -> Unavailable, zero
portal calls`). Live: real `_lifespan` boot with a fake portal, gate on and
off (/tmp/ns847-recon/evidence/e2e-rung5-c2-serve-boot.txt), and the CLI
matrix incl. an own-key cell (e2e-rung5-c2-bootstrap.txt), 20/20.

* fix(credits): the welcome host is free-tier evidence, so a free-tier identity never sees "run /topup"

A free-tier identity carries $0 by design, so the portal seed reports
`paid_access=False` for it. `is_free_tier_model` did not know the welcome
host, read that as a depleted account, and every free-tier turn ended with
the credits-depleted notice telling the user to top up an account they do
not have.

Rule (4) in `is_free_tier_model`: a `base_url` on the Nous welcome host
(`anon_auth.route_is_welcome_host`) is the free tier. The host is the
evidence, not the model name: the paid inference host can serve
`nous/welcome` to a named account and that account's depletion is real, so
`("nous/welcome", <inference host>)` stays False. Local data only, like the
three rules above it.

Restores the two contracts dropped by hermes-magic 674e11d1eaa (the
prototype line ran without unit tests): the welcome host is free without
any pricing evidence; the model name alone is not. The first is red without
this fix.

* fix(copy): free-tier text stops promising a connector transfer and never names the config key

Sign-in copy on every surface said "Sign in to keep your connectors" and
ended with "Your connectors are kept." The transfer registry that would
make that true is empty (NS-821): nothing carries over today. The copy now
says what signing in does give ("unlock more models and tools") and the
completion line names the account, not a transfer. The docs page loses the
"connectors carry over" paragraph for the same reason.

The picker's off-state line exposed `nous.guest: false` and the word
"guest"; user copy names the free tier only (R-USR-1).

The docs page gains the pre-rollout note: until GA nothing on it happens
without `HERMES_GUEST_ONBOARDING=1`. Its "first command mints" and
"replaced on next use" sentences now describe the boot bootstrap.

zh is a strict locale: the `freeTier` block was English placeholder text
copied from `en`; it is now Chinese. `connectorsKept` is renamed
`completedBody` since it no longer talks about connectors.

* feat(desktop): the free-tier launch flag is decided once in Electron and stamped onto every backend spawn

The Python backend reads `HERMES_GUEST_ONBOARDING` and treats exactly "1"
as on. Until now nothing in the desktop set it, so a packaged app could
never turn the free tier on, and a backend spawned by the app could
disagree with the app about whether the tier was live.

`electron/guest-onboarding.ts` owns the decision: `guestOnboardingEnabled`
is true when the launch env has `HERMES_GUEST_ONBOARDING=1` or argv has
`--guest-onboarding` (the packaged-app spelling). It is read ONCE at launch
into a module constant. `desktopBackendSpawnEnv` wraps every backend env
as the outermost call and writes the flag LAST, as "1" or an explicit "0",
so no earlier spread (`process.env`, `backend.env`) can resurrect a stray
value from the parent shell.

Stamped onto all three spawn sites: the primary `serve` spawn, the pooled
per-profile spawn, and the remote SSH `exec env ...` command (which gains
` HERMES_GUEST_ONBOARDING=1` only when on). The embedded terminal PTY and
the backend probes are not backend spawns and do not get it: a
`hermes --tui` typed in the pane must not mint.

The renderer learns the same fact read-only through the existing
`hermes:launch-flags` sync IPC (`guestOnboarding`) and preload
(`window.hermesDesktop.guestOnboardingEnabled`).

Ruling: NS-845 Q1.1 / Q2 (env var is the contract, `--guest-onboarding`
maps to it in main). Two invariant tests on the pure helpers: only "1" or
the argv flag enables; the spawn env carries "1"/"0" as the last word and
preserves every other key.

* feat(desktop): the renderer learns free-tier readiness from one `setup.ready` push, not a 60 s poll

The backend's boot bootstrap now announces `setup.ready` once, after it has
created (or refused) the free-tier identity and resolved the inference
route. The renderer used to discover both by polling `setup.status`,
`setup.runtime_check` and `free_tier.status` every 60 s from
`useStatusSnapshot`; a fresh install's chip, notice strip and onboarding
overlay could sit stale for up to a minute after boot, and three RPCs a
minute per window kept asking a question whose answer changes only at
boundaries the backend already announces.

`handleLifecycleEvent` routes `setup.ready` (active source only, like
`skin.changed`) to `notifySetupReady()`, a one-shot tick atom in
`live-sync.ts` beside the other change ticks. `useStatusSnapshot` listens
to it and runs one readiness round at once (`setup.status` +
`setup.runtime_check` + `free_tier.status`). The readiness legs also run
once on open and on return from another app, as today. The 60 s tick keeps
only `getStatus()`.

`SetupStatusSnapshot` types the record's additive fields (`ready`,
`free_tier`, `other_providers`, `inference_provider`); readiness semantics
are unchanged and still key on `provider_configured` + `runtime_check`.

Ruling: NS-845 Q1.2 (renderer half). Tests: the lifecycle branch fires one
refresh from the active source and none from another; the snapshot hook's
contract is three legs on open, one leg on the tick.

* fix(cli): the banner names the free tier's model instead of "no model configured"

The welcome banner prints before credentials resolve, so on a fresh install
`model` is empty and the banner said, in red, "no model configured — run
/model or hermes setup". Under the free tier that is false: the route is
already known from local state (identity on disk, tier on), and the first
message will run on `nous/welcome`.

`_banner_left_lines` now asks the route the same question when `model` is
empty (`guest_carries_inference()`, a local read) and shows `welcome · Nous
Research`. When nothing resolves the red line stays. Ruling: NS-845 ("the
banner's 'no model configured' line reads the resolved route").

Live: fresh HERMES_HOME + fake portal, gate on -> `welcome · Nous Research`;
gate off -> the red line, zero portal calls.

* fix(aux): vision on the free tier uses nous/welcome too

The text-only modality on the gateway's `nous/welcome` row is DeepSeek V4 Flash's, the
backing model until the repoint; `z-ai/glm-5.3-flash` is natively multimodal and the
repoint declares the welcome row `text+image->text`. Skipping Nous for vision on the
welcome host would have sent every image step past the free tier for no reason, so the
auxiliary client pins the route's one model for every lane. A backing model that takes
no images answers with the upstream's own error, which the ladder handles as it always has.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
(cherry picked from commit 7456e028faba55480db43015dc2c8df3e393a415)

* fix(gateway): hermes gateway run is a boot owner of the free tier too

Rung 5 made every demand-time free-tier site a read: resolve_provider,
the connector token, the /login precondition. That is only correct if
every process that can reach those sites ran the bootstrap first. The
CLI (cmd_chat) and hermes serve (_lifespan) did; the standalone
messaging gateway did not. A fresh HERMES_HOME with the gate on and
`hermes gateway run` reached provider resolution with no identity to
consume, and /login returned Unavailable. Reported by @andrexibiza on
#107697 (P1).

GatewayRunner.start now runs `free_tier_bootstrap.run_bootstrap` on an
executor thread right after startup recovery and BEFORE any adapter
connects, so a fast first DM cannot arrive with nothing to resolve. It
is its own step, not part of the turn-machinery warm-up: the warm-up is
an optimisation with an off switch (HERMES_STARTUP_WARMUP_TIMEOUT<=0);
the bootstrap is correctness and must always run. With the gate unset it
is a local inventory and no network.

Live, real GatewayRunner.start against a fake portal in a fresh home:
  gate on   -> 1 create, identity persisted, resolve_runtime_provider=nous,
               /login precondition sees the identity
  gate off  -> 0 portal calls, no identity, no_provider_configured
Before the fix the gate-on row was identical to the gate-off row.

Test: the bootstrap seam runs before _start_prefilter_platforms and
delegates to the one creator. Red on 5554eb6993 (no seam), green here.

---------

Co-authored-by: Robin Fernandes <robin@soal.org>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-11 03:45:33 +05:30
Siddharth Balyan
3b01b4ce0f feat(desktop): Nous free tier on Hermes Desktop (#105260)
* feat(desktop): free-tier state over RPC, status routes that name it, and a sign-in that keeps connectors

The desktop learns about the Nous free tier by reading local auth state (pull): free_tier.status
answers has_guest / enabled / carries_inference / notice_pending with zero network, and
free_tier.ack_notice persists the one-time notice flag on the identity itself. setup.runtime_check
reports free_tier for the selected route; /api/portal, the Nous card in /api/providers/oauth and
billing.state carry free_tier (billing answers the free tier locally instead of a portal call that
can only fail). The free-tier picker row carries an explicit free_tier_row flag and is never priced or
locked. POST /api/providers/oauth/nous/start over a free-tier identity registers the connector
transfer and returns its code and consent URL; the poller waits for the transfer before the token
grant, persists the account, runs settle_after_upgrade, and the poll response gains reason,
account_email and model.

* feat(desktop): free tier on Hermes Desktop: ready screen, notice strip, status chip, Billing view, one sign-in dialog

The renderer reads the free tier from free_tier.status (pull) into one store; the first-launch
intro is the same state rendered two ways, keyed on the backend's one-time flag: the onboarding
overlay opens on a ready screen when the free tier carries inference, else a one-time strip above
the composer. Settings > Billing gains a free_tier view (notice with one Sign in, Plan / Model /
Connectors summary, plan card, footnote; no payment or usage rows). A status-bar chip names the
tier and model while it carries inference. Every entry point opens one claimed sign-in dialog that
drives the extended oauth/nous route and maps the poll's status and reason to the ruled screens;
Done settles billing, model options, providers and re-homes a session still on nous/welcome. The
picker badge also fires on free_tier_row. Docs: Desktop section in the free-tier guide, AGENTS notes.

* fix(desktop): free_tier.status starts the free tier's background setup when no identity exists

A served backend has no session-setup moment like the CLI's, so beside an explicit provider the free
tier was never set up on the desktop: no connectors, no notice strip. The first status read now
starts the same one-attempt background setup; the call itself never waits.

* fix(desktop): one Sign in on the Billing page; Settings > Providers names the free tier, never Connected

The free-tier plan card is the what-you-get text alone (the notice carries the page's one Sign in).
The Nous provider row reads Nous · free tier with a Free tier tag while the identity is the free
tier, instead of Nous Portal · Connected.

* fix(desktop): Settings > Providers never files the free tier under Connected

* fix(desktop): the intro's shape is keyed on the route, not on the identity

free_tier.status reports available (an identity exists and the tier is on); whether inference
runs on the free tier is setup.runtime_check.free_tier, keyed on the resolved endpoint. The ready
screen shows when that route is the free tier; the composer strip when the user's own provider
carries inference. An own-key install used to get the ready screen.

* docs(desktop): say what the free-tier chip is keyed on

* fix(desktop): the featured Nous row's pitch on the free tier says what signing in adds

* fix(desktop): a cancelled or superseded sign-in attempt can no longer change the identity or hide the intro

Four lifecycle holes from review. The Nous poller checks the session's cancelled flag after the
transfer wait, after the token grant, and once more under the session lock together with the
save, so a sign-in the user abandoned never persists. The renderer's sign-in store carries an
attempt generation that every continuation checks after each await, so a poll from a closed
attempt cannot publish over the one on screen (and its backend session is cancelled). The ready
screen comes down only after the backend recorded the acknowledgement. A composer still mounted
takes over the notice claim when its owner unmounts. One thin test per hole.
2026-09-11 03:45:32 +05:30
Teknium
d9ca9c974d feat(vault): two-factor codes — automatic from a saved authenticator key, otherwise asked for in the user's UI
Follow-up to #106480. Sites that ask for a code after the password stopped
the agent cold: the login classifier excludes one-time-code fields on
purpose (a password must never land in an OTP box) and there was no tool
for the second step, so the only move was to ask in chat.

browser_vault_enter_code
  Fills the one-time code the current page asks for. Two sources, same
  invariant as passwords (the code goes to the page over the supervisor
  socket and never enters model context):
  - a TOTP seed on the login: local vault `otp_secret` (RFC 6238, stdlib,
    verified against the RFC test vectors), 1Password `op item get --otp`,
    Bitwarden `bw get totp`. Nobody is asked.
  - no seed: the surface prompts "Verification code for {site}"; the user
    types what their phone/email/app shows. Enter on empty / Skip declines
    and the tool returns code_declined ("do not ask again this turn").
  no_code_field tells the model the site wants a passkey / hardware key /
  app approval: hand it to the user's device and wait for navigation.
  Per-digit OTP boxes (maxlength=1 pattern) get one digit each in DOM order.

Surfaces
  CLI: sudo-style panel, code shown as typed (not a secret worth masking,
  typos must be visible), Enter submits, ESC/empty skips.
  Desktop: "Verification code for {site}" card via vault.code.request /
  vault.code.respond (gateway), owner-routed like the other vault prompts.
  Settings → Passwords & Logins: optional "Authenticator key" field on the
  add form (base32 or otpauth:// link); items with one show a "2FA auto"
  badge. `hermes vault add` asks for the same optional key.
  browser_vault_fill's result now says what to do next ("if the site asks
  for a verification code, call browser_vault_enter_code with this handle").
  Six locales.

Verified live (real model, local 2FA site that checks the TOTP; CLI PTY):
  A. login saved with authenticator key → signed in through 2FA, zero
     prompts, code/password absent from the transcript
  B. login without key → code panel → user types code → signed in
  C. panel dismissed → agent stops and explains, never asks in chat
Unit: RFC 6238 vectors, seed normalisation, mint-without-asking, per-digit
spread, decline, no-code-field; Desktop card test (owner routing, trim, Skip).
2026-09-10 11:48:01 -07:00
Teknium
77e55b4d1f fix(desktop): show each in-app tip once, never lap the catalog again
The idle tip rotation walked the catalog as a ring: after the last tip it
wrapped to the first, so a user who had already seen every tip kept
getting "Start fresh", "Teach it once", ... again every six hours for as
long as they used the app. Only the X stopped a tip, and letting a bubble
time out (the normal way it leaves) counted for nothing.

The walk is now one lap. nextTip also steps over every tip in the seen
ledger ($tipShownAt, which already recorded every catalog tip that
reached the screen), so a tip shows once however it left, and the
rotation runs dry once every tip has had its moment. Settings > Reset
clears the seen ledger and the cursor as well as the retired set, and its
button counts what a Reset would actually bring back (shown or closed,
counted once). Agent tips carry no catalog id and are untouched.

Live repro (Playwright against the worktree's Vite renderer, all nine
tips seeded as seen, clock fast-forwarded past settle + cooldown):
origin/main re-showed "Start fresh"; fixed renderer shows nothing;
a fresh user (nothing seen) still gets the first tip.
2026-09-10 11:18:16 -07:00
hermes-seaeye[bot]
6f8b8e77dd fmt(js): npm run fix on merge (#107545)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-10 17:41:19 +00:00
Teknium
98d11c95f4 feat(vault): zero-setup UX — save a login on the page that needs it, managers auto-detected, one "Passwords & Logins" surface
Nobody should have to learn `hermes vault add` or find a toggle before "log into GitHub" works.

- browser_vault_save_login: when the agent reaches a sign-in page with no saved login it asks the user
  on THEIR surface (CLI two-step panel on the sudo modal: identifier shown, password masked; Desktop
  card with labelled Email/username + Password fields). The answer goes to the encrypted vault bound to
  the page origin and is filled at once; the model gets back only the handle and identifier. Declining
  returns save_declined; headless sessions get prompt_unavailable. Never a password in chat.
- Vault tools ride with the browser toolset (check_browser_requirements) instead of appearing only once
  the vault has items — an empty vault is exactly when save_login is needed. browser_vault_list hints
  at it when empty.
- 1Password / Bitwarden are login sources as soon as their CLI is installed; `vault.<name>.enabled`
  is opt-OUT only. Settings shows Detected/Locked/Unlocked/Off/Not detected with a switch only for
  installed managers; `hermes vault sources` reports detection, `--disable`/`--enable` flip the opt-out.
- Desktop nav/page renamed "Passwords & Logins"; empty state tells the user they do not need to add
  anything; all five locales updated. Docs rewritten from "how it works" to "say log into X".
- New per-thread SaveLoginPrompt callback (agent/vault_backends/unlock.py) installed beside the unlock
  prompt on every CLI site and the gateway bridge (vault.save_login.request/respond/expire), propagated
  to worker threads via tools.thread_context.

Live: CLI PTY (real model, packaged Chromium, local login server) — panel shown, identifier + masked
password typed, server received the correct password, password absent from terminal transcript and
from every file under HERMES_HOME outside vault/. Native Electron (headless, isolated HOME/HERMES_HOME,
own Vite + CDP port) — card shown, "Save & sign in", server received the password, Settings lists the
saved item, password absent from the rendered UI.
2026-09-10 10:35:07 -07:00
Teknium
5aa9e7f033 fix(vault): independent-review findings — vendor contracts, profile scope, transport, target binding
Bitwarden unlock now uses the CLI's documented non-interactive channel:
`bw unlock --raw --nointeraction --passwordenv VAR`, VAR set on the child
environment only (bw 2026.x rejects a piped password with "Master password
is required"). Verified against the real published binary.

Manager session tokens are keyed by (profile home, backend): a Desktop
gateway hosting several profiles can no longer reuse or lock another
profile's session. Status probes (`vault.sources`, is_unlocked) no longer
refresh the idle TTL; only real manager calls do. Gateway session teardown
locks the profile's managers (a per-session unlock ends with the session).

1Password service-account token comes from the profile-scoped secret store
(get_secret), not ambient os.environ.

`vault.source.set` no longer references a module constant (bind_module
rebinding dropped it → NameError on every Settings toggle).

Fill target binding: inspection stamps each input with a per-inspection
slot attribute; the fill resolves by stamp and requires type=password, then
strips every stamp. A DOM reflow between inspect and fill can no longer
redirect the password into a text field (reproduced in real Chrome before,
0 filled after).

Redaction boundary: no 4-char floor, CR/LF-normalized form registered
(what a text input actually stores), JSON object KEYS scrubbed in both
browser redactors; longest value first. Docs now state the real trust
model: accidental-disclosure protection, not an execution sandbox.

Desktop: the mid-turn card sends the master password through the owning
session's socket (requestForOwnedSession), never the ambient foreground
gateway; `vault.unlock.expire` clears a stale card; Settings keeps the
master password out of react-query mutation variables (ref consumed by the
mutationFn). One renderer invariant test for the routing.
2026-09-10 10:35:07 -07:00
Teknium
92e0de0ac4 feat(vault): Desktop, TUI and CLI surfaces for password-manager unlock
Desktop
- Settings → Credential Vault gains a "Password managers" section: per-manager
  toggle (disabled with a hint when the CLI isn't installed), Locked/Unlocked
  pill, Unlock (masked master-password dialog → vault.unlock) and Lock.
  Items from a manager show a source badge instead of a delete button.
- Mid-turn vault.unlock.request renders a masked card in the chat (same
  contract as the secret/sudo cards: dismiss = keep locked, late answers
  tolerated, blocks the composer, badges background sessions).
- i18n parity en/ar/ja/zh/zh-hant.

Ink TUI (hermes --tui): vault.unlock.request/expire overlay via MaskedPrompt;
Esc keeps the manager locked.

CLI: `hermes vault sources [--enable|--disable NAME]`; `hermes vault list`
shows the source column and names enabled-but-locked managers.

Docs: credential-vault.md covers managers, per-session unlock, and the
headless (cron/webhook/API/-q) no-prompt posture.
2026-09-10 10:35:07 -07:00
686f6c61
f44b0fd342 fix(desktop): keep model picker on the selected provider
Treat the Desktop catalog selection as a (provider, model) pair so a
shared model id on OpenRouter cannot rewrite a custom or first-party
pick after refresh or keyboard highlight.
2026-09-10 03:58:52 -07:00
Teknium
61afcde8f9 refactor(desktop): one Plugins surface — Capabilities → Plugins owns agent + desktop plugins, install, and the catalog
Plugins were split across two pages that each showed half the picture:
Settings → Plugins listed desktop plugins plus "Install from Git" and a
pointer saying agent plugins live elsewhere; Capabilities → Plugins listed
agent plugins plus the catalog picker but knew nothing about desktop
plugins. A user asking "what extends my Hermes and where do I add more?"
had to visit both and still could not see the whole set in one place.

Capabilities → Plugins is now THE plugins page:

- Agent plugins section (scoped to the profile selector) with the
  "Install from Git" button in its header — installs target the scoped
  profile, not whichever one is active.
- Desktop plugins section beneath it (same for every profile), with the
  folder/rescan controls and the "agent half missing here" drift chip,
  whose repair also lands in the SCOPED profile.
- The catalog picker underneath, unchanged.

Settings → Plugins is removed. `/settings?tab=plugins[&plugin=…]` and the
existing `?tab=mcp` redirect share one table (`settings/moved-tabs.ts`) so
old bookmarks and palette links land on the same row on the new page.
Command palette: plugins moved from the Settings group to the Capabilities
group; installed-plugin rows deep-link to `/skills?tab=plugins&plugin=…`.
Dead `settings.plugins.agent.*` and `settings.nav.plugins` i18n keys dropped;
docs and in-code pointers say Capabilities → Plugins.
2026-09-10 02:30:50 -07:00
brooklyn!
e7c819a7e1 fix(desktop): mask text behind sticky user messages 2026-09-10 01:14:30 -05:00
brooklyn!
bbe212de9b feat(desktop): keep panel tabs in the titlebar 2026-09-10 01:14:30 -05:00
brooklyn!
b5c0a7ebe2 test(desktop): cover clarify submit shortcuts and input guards 2026-09-09 21:57:01 -05:00
brooklyn!
ff7a1e8953 fix(desktop): restore Cmd/Ctrl+Enter in clarify forms 2026-09-09 21:57:01 -05:00
Teknium
97ca90f184 fix(desktop): expired OAuth grant shows a one-click 'Sign in again' instead of a retryable Provider error
A rejected OAuth token (HTTP 401 'User not found' from Nous Portal, Codex,
xAI…) reached the desktop error card as 'Provider error' with Retry as the
first action, which just replays the same dead credential.

Backend: nonretryable_client_error_result dropped failure_reason /
failure_retryable, so error_surface classified every non-retryable 4xx as a
retryable provider failure. It now stamps the classifier verdict like the
max-retries path, and auth-layer descriptors carry auth_kind
(oauth|api_key, derived from the provider catalog tab) + provider_label.

Desktop: an auth/oauth surface renders 'Authentication error', explains that
the <provider> sign-in expired/was revoked, and offers 'Sign in to <provider>
again' which launches that provider's existing onboarding OAuth flow scoped
to the failed session's gateway profile. Retry stays as the follow-up click.
Re-login to the provider already in use keeps the current model instead of
swapping in the recommended default.
2026-09-09 16:30:24 -07:00
Teknium
474143da81 fix(desktop): trim the auto Show-earlier gate and prove it end to end
Drop the injectable `topEdgePx` knob (one caller, one constant) so the
threshold has a single home, and fold the two-branch tail of
shouldAutoShowEarlier into one predicate. Comments keep the WHY: wheel is
the only signal at a clamped scrollTop 0, and the gates name the states
where scrollTop is near 0 without meaning "read earlier".

Tests: collapse the 11-case matrix into one invariant test on the pure
gate, and add a rendered-list test that mounts the real Thread against a
transcript heavier than RENDER_BUDGET, escapes the bottom lock, wheels up
at the clamped top and asserts more turn groups mount — while an upward
wheel mid-transcript mounts none. Red on origin/main (28 groups stay 28),
green with the fix.
2026-09-09 12:23:20 -07:00
KoNit-K
3ec8042483 fix(desktop): page earlier transcript on top-edge scroll
Long conversations and branched sessions hide older turns behind the
render budget, but only the Show earlier button loaded them. Auto-page
through the same showEarlier path when the reader is at the viewport top.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-09 12:23:20 -07:00
Teknium
c80003ff57 refactor(desktop): key the live-owner escape off the 4090 reason code, drop the prose sniffer
The gateway already ships the refusal as machine data: prompt.submit returns
JSON-RPC 4090 with error.data.reason = SESSION_NOT_OWNED. Matching the English
sentence in a new lib helper duplicated that contract and would silently miss a
reworded or localized message. submit.ts now classifies the rejection once
(isSessionNotOwnedError beside the sibling isSessionBusyError, reason code
first, prose only for pre-contract backends) and stamps the inline error bubble
with errorSurface {layer: gateway, code: SESSION_NOT_OWNED, retryable: false},
the same descriptor shape the gateway uses for turn failures. The error card
reads surface.code instead of re-sniffing text; retryable=false already hides
Retry through the existing gate.

Dropped from #106248: the stranded-resume overlay button. That overlay fires
when session.resume itself fails through every retry; reading is never fenced
by the lease, so the live-owner refusal never reaches it (the issue's screenshot
is the inline turn-error card). Also drops the resumeStartNewSession locale
keys and the separate detector module + its test; two invariant tests remain
(submit stamps the descriptor; the card offers Start new session / hides Retry).
2026-09-09 10:57:47 -07:00
KoNit-K
6efe3a45c1 fix(desktop): offer Start new session on live-owner exclusivity refusals
Retry alone is a dead end when another surface holds the session lease
(JSON-RPC 4090 / SESSION_NOT_OWNED). Surface an explicit new-session
escape on the turn-error card and the stranded-resume overlay without
weakening the lease guard.
2026-09-09 10:57:47 -07:00
hermes-seaeye[bot]
990473a79c fmt(js): npm run fix on merge (#106237)
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
2026-09-09 03:18:52 +00:00
brooklyn!
5255d7823e test(desktop): cover scrolling message counts and pane isolation 2026-09-08 22:13:30 -05:00
brooklyn!
342f76a7c2 feat(desktop): show messages below the thread viewport 2026-09-08 22:13:30 -05:00
brooklyn!
9e8c1569b6 fix(desktop): keep pets and starmap animated without focus 2026-09-08 21:53:33 -05:00
brooklyn!
3cad0f312b fix(desktop): keep visible renderer animations running on blur 2026-09-08 21:53:33 -05:00
brooklyn!
258b351fc1 fix(desktop): keep background reports behind bounded disclosures 2026-09-08 18:53:48 -05:00
brooklyn!
960dee7fe5 fix(desktop): Hide tabs works on the sessions sidebar
hideOnly chrome pinned the Sessions/Bots strip on at any tab count, so
never was a silent no-op. The panes stay; ⌘⌥T brings the strip back.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-08 15:35:19 -05:00
Teknium
092ed67532 feat(desktop): surface live subagent work above each composer 2026-09-08 03:06:30 -07:00
Teknium
610c869ac6 fix(desktop): task panel follows the todo_list wire name
The core-tool rename shipped `todo_list` on the wire (legacy alias `todo`
kept for old transcripts), but the Desktop renderer still matched the tool
by the literal `todo` in seven places: the live tool.start/tool.complete
mirror into the composer status stack, the todo-stream router, args
carry-over, the transcript hoist, the silent-tool class, the count noun,
and stored-history hydration. Every live task update therefore went into
the transcript as an ordinary tool row while the task panel stayed empty,
and reopening a chat never restored a finished list.

One predicate (`isTodoToolName`) now owns the wire/legacy name pair and
every site reads it.
2026-09-07 11:25:20 -07:00