* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift
The label stays in the box, invisible, and the spinner is absolutely
centred over it, so a Connect or Approve button keeps its width while it
works instead of collapsing to a spinner. The approval bar had the same
thrash and moves onto it.
* refactor(desktop): one consent card for connectors and MCP setup
McpSetupTool rendered its own copy of the connector card's markup. It now
renders ConnectorCard for the pending question and ConnectorSummary once
settled, and the card gains what MCP needed: keyboard accelerators, a
source line, a question heading. The card also gets an avatar variant
(40px mark in the left gutter, text and buttons on one column) and a
collapseWhenSettled switch so a connector can stay a full card with a
green Connected pill in the action slot while MCP keeps its one-line
summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and
Spotify; Slack via Tabler because simple-icons dropped the mark.
* feat(desktop): connector card drives the agent through manage_connections wait
The offer used to end in a Continue in chat button, and the agent, seeing
an unconnected status, would improvise around the app. Now the card does
what the TUI does. Clicking Connect opens the browser and sends one hidden
line telling the agent to park in manage_connections action=wait for that
slug and to never call connect again (a second link cancels the one being
signed into). Not now sends its own line. A hidden request that lands
while the turn is busy steers it, or queues if the turn just ended.
Which call owns the live card changes too: consecutive calls naming the
same apps are one exchange (connect, the wait, the status that follows),
and the first of the last exchange is the card, so the agent's wait no
longer demotes the card mid-authorization and mints a fresh one below it.
A targeted ask renders one or two bare cards; only a real catalog gets the
header, search and refresh.
* feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them
The welcome chat knew connectors only as preferences to pick and wire up
later, so asked to connect Gmail it invented a Settings page that does not
exist. Both scripts now carry one rule set: status once, one batched
connect for every app named, the card is the ask so write a line and end
the turn, never route around a declined app with another client or
credential. The build handoff checks real connection status instead of
asserting none are connected, and the first task must be finishable, not
free of, the apps they picked. The connectors card explains what
connecting means and reports the count on its Continue button.
* fix(tools): resolve the Nous identity for share_auth profiles in the connector gate
A profile created with share_auth has no auth.json of its own and signs
in through the root store. Every other credential reader falls back to
the global root; the connector gate read HERMES_HOME/auth.json directly,
saw nothing, and stripped manage_connections from the profile's tool
list, so the welcome chat's agent truthfully reported the tool missing.
The gate now goes through get_provider_auth_state.
* fix(agent): name a provider retry backoff on the live status line
The retry status is buffered and replays only when every retry fails, so
during a 60s backoff after a 5xx the user saw a bare spinner. Right after
a connector sign-in landed this read as the agent going silent. The
backoff now also rewrites the live wait notice, which the desktop already
renders in the thread status row; it is transient and clears on recovery.
* test(desktop): connector rehearsal launcher and flagged connector spec
connector-rehearsal.mjs starts the real desktop and backend under a fresh
HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344,
so the onboarding connector flow can be driven end to end by hand or from
outside. The Playwright spec covers the flagged connector step.
* fix(desktop): send the agent back into wait when the user keeps waiting after a timeout
The card's Keep waiting re-entered the poll but the agent's own wait had
timed out too and nothing told it to go back in, so it would start
talking mid-authorization. keepWaiting now fires onWaiting like connect
does. Tests also pin that an expired or revoked grant asks the gateway
for reconnect, not connect.
* style(desktop): blank lines in connector-flow test per lint
* feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film
The intro is a one-time reveal, so anyone rehearsing the guided chat behind
it sits through it on every fresh HERMES_HOME. The flag rides the existing
launch-flags path (main → preload → renderer) next to guestOnboarding and
only gates isIntroRevealEnabled; the backend never sees it. The rehearsal
launcher sets it.
* fix(desktop): onboarding card Continue stays Done after the transcript rebuilds
The card kept its Done flag in component state. The hidden submit and the
turn-end hydrate both rebuild the message list, so the card remounted with
the flag false and Continue came back live, letting a step be answered
twice. The committed steps now live with the other onboarding answers,
keyed by step, and the first-build chip pick rides the same store.
remember_onboarding projects by key, so the new field never reaches USER.md.
* fix(desktop): no provider picker or free-tier chip over the guided first launch
Two sign-in surfaces leaked into the guide. A credential probe on the
setup profile (a free-tier token mid refresh, a session before its runtime
settled) hit requestDesktopOnboarding and dropped the provider picker over
the chat the user was in; and the statusbar free-tier chip sat there
offering a second sign-in the whole time. Both now yield while the gate
phase is cinematic, guided or handoff. The free tier is the provider for
those phases, and the guide offers sign-in on its own ready screen.
* fix(desktop): onboarding connector picks are real catalog slugs
The picker offered Spotify, GitHub and Stripe, none of which the deployed
connector catalog carries, and spelled Calendar and Drive with hyphens the
gateway does not use. A pick the build chat could not honour ended as
"Spotify isn't in the connector list" after the user had been told to
expect it. The list is now twelve slugs from the live status catalog,
spelled as the gateway spells them; GitHub is out (the terminal has git
and gh), chat channels stay on Messaging. Marks for the new entries; the
Google marks answer both spellings. The build runbook offers the picked
connections in its first turn rather than after the work is underway.
* fix(desktop): the free-tier ready screen never interrupts the guided chat
A readiness round fires when the layout pick assembles the window, and it
raised the free-tier ready screen over the conversation: the user was
dropped into the main app, dismissed it, and came back to a card they had
already answered. The guide is the introduction. The ready screen now
yields while the gate is cinematic, guided or handoff, and the notice is
acked the moment the guided chat takes the screen, not only when the film
does, so a skipped film no longer leaves it pending.
* feat(desktop): tour options that lead to building, and a fork that follows the tour
"Just the basics" and "Show me around" read as a click-through with no
exit; "I'll figure it out" read as declining help. Now Quick tour, Show me
everything, and Skip, let's build something. The script also folds the
fork into the same turn as the tour, so when the user closes the overlay
the next ask is already waiting instead of a transcript that ends on the
tour call.
* feat(desktop): the onboarding connector picker reads the live catalog
A hardcoded list, however carefully copied from today's catalog, is the
next drift. The picker now asks connectors.list through the same
session-owned RPC the connector cards use and offers exactly what the
gateway carries: a curated lead order puts the everyday apps first, chat
channels stay on Messaging, everything else is reachable by search. The
picks are gateway slugs, handed straight to manage_connections. No
catalog (toolset off, gateway unreachable) ends the step honestly with
Skip instead of inventing apps.
* test(desktop): the guided first launch never forces a sign-in
The acceptance criterion the guided onboarding was built to, as a test:
while the gate is cinematic, guided or handoff, the provider picker does
not open and a credential warning is dropped rather than deferred to the
next send. Outside the guide the picker opens as before. Red against the
tree before the guards landed (6 of 9).
* fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape
Closing the app during the guided first launch and reopening it booted the
normal shell around the persisted solo layout: the connecting splash, the
stock composer and model picker, a small window whose sidebars would not
open, while the gate still read guided. The gate now queues a kickoff for
the guided phase too (the kickoff adopts the existing guide chat by title),
takes the solo shape before the gateway opens rather than after, and the
connecting overlay yields to the guide's own opening. A typed reply in the
composer now closes an ask card and the first-build chips the same way a
click does; the layout card's Continue comes back Done.
* style(desktop): one answeredAfter helper for the ask card and first-build chips
* fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows
Between the film and the greeting the full-size shell painted for a beat:
finishIntroReveal showed the main window, then the kickoff shrank it once
the setup profile answered. The listener on the intro's hidden edge now
takes the guide's shape (solo layout + small centred window) synchronously,
so the window is already the guide when it is shown. One takeGuideShape
owns the pair; kickoff and the boot gate call it idempotently.
* style(desktop): the 'nothing connects yet' line reads first on the connectors card
A delayed browser could miss the 900ms terminal event and spin forever after the updater exited. Retain terminal delivery until the page acknowledges it, bound unavailable-client teardown and failed requests, and preserve a truthful final display.
Fixes#103747. Builds on OutThisLife and Teknium detached handoff work in #83634 and the #75895 quiet-window design. Continues Axl Ibiza Windows update investigation (#60233, #94107, #100763), including source/review contributions carried by merged #93353 and #85170. Existing #102373, #103140, #95719, #97299 and #103632 retain their separate scopes.
One conflict: upstream 6e7c7c7da9 replaced bot-mode-closed-chat-stays-closed.spec.ts with bot-mode-row-click-mirrors-registry.spec.ts while our side had rewired its mock-server import. Kept upstream's replacement and rewired the three new specs importing ./mock-server to the consolidated tests-js copy (symbols verified present).
Widen the salvaged guard from a hand-maintained four-package floor to the
class it stands for: every `dependencies` + `devDependencies` entry in the
desktop workspace manifest. Live probe on this box: a tree holding vite,
katex, electron and electron-builder but missing `@rolldown/plugin-babel`
still passed the floor-only guard, and `vite build` died loading
`vite.config.ts` after `prebuild` had already run. The floor stays as an
unconditional fallback for an unreadable manifest; optionalDependencies
are skipped because npm legitimately omits them (get-windows).
Five new vitest cases (12 total); the two class tests fail when the
manifest union is removed. Refs #86443.
Follow-up to the salvaged #87980: the test kept its own copy of the
build-critical package list (drift hazard) and the module's default
export had no consumer.
Refs #86443
assert-root-install.mjs exists to turn an incomplete root install into one
actionable line instead of a failure deep inside the build. It only ever
checked that vite resolved, so an install covering part of the workspace
graph passed the guard and died later on something else. That is the shape
reported in #86443: the updater's npm install brought in 521 of the 769
packages a full install gives, root node_modules had vite but not katex, and
the build failed on an unresolved katex/dist/katex.min.css with nothing
pointing at the install as the cause. apps/desktop/src/styles.css imports
that stylesheet, so katex is as load-bearing for the renderer bundle as vite
is, and electron / electron-builder are the same for packaging.
Check all four and name every missing one, so a partial install is reported
once and completely rather than one package per build attempt.
Resolution walks node_modules upward the way Node's own lookup does, rather
than going through require.resolve: a package whose exports map does not
expose ./package.json is not resolvable by path even when correctly
installed, and that must not read as missing. It also keeps a dependency
that landed in the app workspace instead of the hoisted root passing.
The guard now runs from prebuild, ahead of npm run clean, so a tree that
cannot build is rejected before the build deletes its own outputs. On this
checkout clean removes build/electron-types and the tsbuildinfo files, not
release/, so this ordering is not by itself what saves a packaged app; it is
the narrow correctness point that a doomed build should not destroy anything
first. build keeps its own call for anyone invoking the build steps directly,
and the check is pure filesystem lookups, so running it twice costs nothing.
The check is extracted as a pure checkRootInstall() returning {ok, error},
matching assert-dist-built.mjs, so it is unit testable without spawning a
process.
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
one-line compat forwarder) and the new implementation already drains
both pipes asynchronously with bounded abandonment, which supersedes
this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
addition inside the consolidated file; rewired the five upstream specs
still importing './mock-server' to the consolidated path (export sets
verified identical) and dropped the superseded apps/desktop/e2e copy.
The packaged app crashed at launch with 'No QueryClient set, use
QueryClientProvider to set one': useQuery in a lazy chunk (session-list-density)
read a second @tanstack/react-query runtime whose QueryClientContext was never
populated by the entry's QueryClientProvider. The source tree was correct — the
duplication happened at build time, because react-query was the one
context-bearing runtime not pinned to a shared vendor chunk, and rolldown's
merge heuristics inline the spare copy into a lazy chunk depending on toolchain
version.
- vite.config.ts: add @tanstack/react-query to the vendor-react
advancedChunks group + dev dedupe list, mirroring the react-router fix.
- assert-dist-built.mjs: fail the build when the 'No QueryClient set'
invariant appears in more than one JS asset (launch-smoke guard).
- assert-dist-built.test.mjs: unit tests for the new invariant check.
- launch-packaged-app.spec.ts: e2e smoke test asserting the packaged app
boots to real UI, not the QueryClient error boundary.
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.
The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.
Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.
The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.
JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.
The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.
The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.
`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.
node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.
The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.
The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.
`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.
The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.
Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
`ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
parallel units and for a plain `npm run check`, in both directions. Against
the 13-leg matrix the count is 13 to 11, and the whole difference is the
three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
Review follow-ups on the compositor spinner and the invalidation scoping.
Spinner CSS:
- Clip each frame to its own box. Braille renders from a system fallback
face (JetBrains Mono has no U+2800 block), whose metrics are not
guaranteed to fit the 1em frame, so neighbouring ink could bleed into
the viewport.
- Name descendants explicitly in the selection guard. The competing
`[data-selectable-text='true'] *` rule has the same (0,1,0)
specificity, so relying on inheritance made the winner depend on
stylesheet order.
- Scope the compositor promotion to spinners that are actually running.
A permanently promoted layer per parked spinner is pure memory at
fan-out breadth, where many sit mounted and paused at once.
- Give every var() the braille default as its fallback, so a missing
custom property degrades to a working spinner rather than an invalid
declaration.
Spinner component: replace the bare `as CSSProperties` cast on the inline
style with an exported GlyphSpinnerVars contract, so a typo in a custom
property name is a compile error rather than a silently dead declaration.
Assistant message:
- Render the inter-agent collapse as a CHILD of the normal body instead
of a competing root. The settled case previously returned a different
element type than the running case, so settling unmounted the whole row
and mounted a fresh one — discarding the DOM the scroll anchor held.
One component, one root, children vary; the truth table is unchanged,
including the collapsed row carrying no tapback listener.
- Collapse AssistantStatusSlot's separate subscriptions into one selector
returning a stable string. The inputs always move together on a status
flip, so reading them separately just multiplied the wake-ups.
- Give StreamingMarker a stable `data-slot` and assert on that rather
than on `span.hidden`.
Repro script: count settled rows by subtracting streaming markers from
message roots instead of `:not(:has(...))`. The selector walked every
row's subtree on each evaluation, inside the very latency window the
probe measures.
Comments: drop the stale translateY(-100%) description, name both pause
triggers, replace hard-coded line-number citations with selector/symbol
ones, note that only the primary window arms the renderer-pause
attribute, and move the forensic trace numbers out of source comments
into the PR.
Delete the three tests that asserted on stylesheet TEXT. AGENTS.md bans
reading source in tests outright, and they demonstrated exactly why: a
var()-fallback edit that changed no rendered pixel broke one of them.
Replacements that exercise the CSS in a real browser follow.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
Removes the two residual invalidators left by the previous commit.
data-streaming: gone from the message root. The flag is not dead -- it is the
settled-row signal for scripts/run-short-session-hang-repro.mjs -- so it moved
to a permanently-mounted, display:none leaf that is a ROOT-LEVEL sibling, and
the repro now matches on the descendant. Placement is load-bearing three ways:
a node inside [data-slot='aui_assistant-message-content'] would steal
:last-child from the stall indicator and change inter-bubble margins mid-stream
(styles.css:1995-2003); keeping it mounted and toggling only the attribute
keeps the per-flip write on a childless node instead of making it a DOM
structure change; display:none costs no layout or paint while querySelectorAll
and :has() still match it.
Renamed to data-message-streaming rather than reusing data-streaming: shiki
puts that exact attribute on deferred code cards, which are descendants of the
message root, so a descendant-matching selector sharing the name would report
any message holding a deferred code card as still streaming.
root isRunning: gone from the standard path. AssistantMessage now dispatches on
interAgentSender, so the collapse gate's live status subscription lives in
InterAgentAssistantMessage and only the rare inter-agent case pays it. The
enter animation captures its enabled flag once off the runtime, non-reactively,
because use-enter-animation.ts parks the value in a ref behind a useCallback([])
identity and consults it only when the callback ref fires at mount -- a live
subscription fed a value the hook already ignores.
Adds inter-agent-collapse.test.tsx: the collapse gate and the marker contract
both had zero coverage, and nothing in the app reads the marker, so a delete
would otherwise look free and silently regress the repro's response gate.
Behavior-identical; invalidation scope only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
require.resolve returns the macOS realpath (/private/var/...) while
os.tmpdir() stays on the /var symlink, so a raw join() deepEqual failed
even though the spawn was correct.
With a GH_TOKEN/GITHUB_TOKEN in the environment, electron-builder auto-selects
the github provider and resolves owner/repo from the repository field, falling
back to reading <projectDir>/.git/config. projectDir is apps/desktop, which has
no .git of its own, and app-builder-lib does not walk up to the workspace root
-- so resolution returned null and threw "Cannot detect repository by
.git/config".
On Linux this fires from onAfterPack for a plain `dir` target: the darwin and
Windows branches return early for non-installer targets, Linux has no such
guard. That is why the same build worked elsewhere.
--publish never keeps `pack` from reaching this at all, but `dist:*` and
test-desktop.mjs still resolve publish config on a machine with a token, so
declare the field too.
Tests call the real app-builder-lib resolver rather than asserting on the text
of package.json, so they track electron-builder's behavior instead of our
formatting.
Co-authored-by: airo7 <airo7@users.noreply.github.com>
Co-authored-by: frankmendes1979 <frankmendes1979@users.noreply.github.com>
The desktop window opened blank white on a fresh install: React threw
"Minified React error #527" before the first paint, from the
`vendor-react-<hash>.js` chunk.
`apps/desktop` pins react and react-dom to the same exact version, but
`vite.config.ts` aliased both to a hardcoded `../../node_modules/<pkg>` —
straight into the monorepo root, where npm is free to hoist a different
react. `@streamdown/math` is a root dependency whose react peer is
`^18.0.0 || ^19.0.0` and which declares no react-dom peer, so npm hoists
the newest react (19.2.8) to the root while react-dom stays at the
version hoisted from the workspaces (19.2.7). react-dom's own peer is
`react: ^19.2.7`, which 19.2.8 satisfies, so the install reports success
and nothing warns. The bundle then shipped react 19.2.8 with react-dom
19.2.7 and React refused to run.
`npm ci` masks this because the lockfile pins the root react to 19.2.7,
which is why CI is green. The recurrence engine is
`_run_npm_install_deterministic()`: when `npm ci` fails it falls back to
`npm install --no-save`, which re-resolves the whole tree and never
records the result — so the split comes back on the next update and
leaves no trace.
Fix the resolution rather than the hoist. The aliases now resolve both
packages from the desktop workspace itself, where npm guarantees the
declared versions are reachable (it nests a copy under the workspace
exactly when the hoisted one differs), so the pair can only ever match.
Pinning react at the npm layer instead was rejected: every manifest-level
pin tried (root dependency, root `overrides`, a scoped override on
`@streamdown/math`) breaks a fresh install with
`ERESOLVE ... peer ink-text-input@"6.0.0" from @hermes/ink@0.0.1`.
Two guards keep it from silently returning:
- `assert-root-install.mjs` (the existing preflight of `build`,
`dev:renderer` and `preview`) now fails the build when the resolved
react and react-dom versions differ, so a split surfaces as an
actionable error instead of a white window.
- A `tests-js` contract test asserts every workspace pins the two to the
same exact version, and that the desktop bundler no longer points at a
hardcoded `node_modules` path.
Verified on a synthetic split tree (root react 19.2.8 / react-dom 19.2.7,
workspace react 19.2.7): the old aliases resolve 19.2.8 + 19.2.7, the new
ones resolve 19.2.7 + 19.2.7, and the preflight exits 1 when the
workspace itself resolves the mismatch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The desktop journal synchronously read, parsed, cloned, and rewrote one
aggregate localStorage value while streamed turns were repainting. Large
tool results and multi-session state could therefore block the renderer and
leave the app unresponsive, while the existing macOS diagnostic path lacked
a real native hide/restore regression check.
Store bounded recovery projections under per-session keys, migrate legacy v1
data once, isolate quota and storage failures, and preserve the newest
recoverable tail without allowing oversized writes to replace valid state.
Add a real Electron/CDP macOS-arm64 A/B harness with native visibility control,
renderer heartbeat and Settings/composer/transcript checks, plus focused
regressions. Keep bulk tool payloads, diagnostics, and the existing recovery
merge behavior out of the hot path.
Fixes#63047
The bench timed two of the six RPCs the fix touches. Extend it to
image.attach, pdf.attach, clipboard.paste and image.detach so every changed
handler carries a number rather than an inference.
Also report which surfaces reach these RPCs at all, since "why was the GUI
special" is the first question the fix invites. CLI attaches inline in its
own turn path with the agent already built, so it cannot reach the stall;
the TUI calls the same RPCs and was equally exposed. The difference was hit
rate, not code path.
gateway_attach_bench.py drives the real dispatcher with a session whose agent
build is still running and times each attach RPC against prompt.submit as the
control — the harness that located the stall and measures it.
image-attach-bench.mjs times the renderer-side transforms (file read, base64,
RPC frame, embedded-image extraction, render-weight walk) across image sizes.
It is what ruled the renderer out: ~26ms total at 3MB.
Review findings: spawnSync('npm', ...) without shell fails on Windows
(npm is npm.cmd; Node >=18.20 throws EINVAL — same handling as
test-desktop.mjs and stage-native-deps.mjs), and a spawn-level failure
exited 1 with no diagnostic. CI is ubuntu-only but the desktop workspace
supports local Windows dev.
Closes the silent-skip hole reviewers flagged: with the index/count
hardcoded in three sibling strings, a copy-paste slip (shard-2of3 running
--shard=1/3) or a partial 3->4 migration would silently skip a third of
the 428-file suite while CI stays green.
scripts/run-ui-shard.mjs parses N/M from npm_lifecycle_event (the script
NAME is the single source of truth), validates the package's shard family
is exactly 1..M for one M, and delegates through 'npm run test:ui' so the
vitest command stays single-sourced.
Mutation-verified: shard-9of3 name -> exit 1 'index out of range';
adding shard-4of4 beside the 3-family -> exit 1 'must form exactly 1..M';
correct invocation runs shard 2/3 (142 files) identically to before.
The dev:mock script duplicated the e2e mock server. The copy had only
the plain chat reply; every scripted path lived only in the e2e version.
A single mock server now lives in tests-js/scripts/mock-server.ts.
The e2e suite imports it as a library. Running the file directly
starts the server, writes a mock config, and launches the desktop app.
The dev:mock script now runs that file.
The e2e tsconfig lists tests-js/scripts in its include, because the
composite project rule requires every imported file to be listed.
* test(desktop): stress long agent sessions in the multitab perf scenario
--tools seeds every transcript with settled tool rounds and drives the live
stream as a working agent turn (tool calls opened and completed between text
chunks), and each tile reveal is timed to next paint (reveal_max_ms) so deep
transcripts report their mount cost.
* perf(desktop): hold the transcript window cut steady while streaming
A fresh weight-walk per store flush slid the cut forward one message at a
time, ~30x/s, and every slide re-indexed the whole windowed transcript —
each row rendered a different message and the runtime repository took its
O(window) rebuild path instead of the one-message update. advanceTranscriptWindow
anchors the cut to a message id and re-cuts once per ~half page of new
content instead of once per flush.
* perf(desktop): stable rows, stepped backfill, pane-shared render budget
Three thread-list fixes for long streaming sessions: memoize the visible-
groups slice and each turn row so a budget-cut advance no longer re-renders
every mounted turn per streamed token; raise the first-paint backfill in
BACKFILL_STEP slices (one bounded commit per frame) instead of a single
20-to-600 transition whose commit landed as a 780ms freeze mid-stream; and
share RENDER_BUDGET across mounted panes so a 4-way grid mounts a quarter
page per pane instead of 4x the fibers.
* fix(desktop): evict settled session states nothing on screen references
Closing a tile never removed its runtime's entry from $sessionStates, so
every tile ever closed parked its full transcript in the map for the life
of the process. Each leftover entry taxes every subsequent stream flush —
the map is spread-copied per delta and the busy/attention/draft projections
walk every entry per publish — so the app got slower the longer it ran,
which users read as "I need to clean my sessions/dbs".
Publish now evicts a settling state when no tile and not the primary view
holds its runtime (transition side effects still fire, so the settle keeps
its unread dot), and closing a tile drops an already-settled state on the
spot. Busy and needs-input states stay: background turns feed the sidebar
dots, and a first publish always lands because a resume can publish a beat
before the surface binds the runtime.
16 tiles streaming in a 2x2 grid with a day's worth of closed-tile residue:
worst-second 34 -> 58 fps, p99 frame 90 -> 28 ms, longtasks 37 -> 0.
* perf(desktop): index lineage aliases per sessions-list reference
lineageAliases scanned the whole recents list per call, and it is called
per cached session state per status projection per message delta — with a
populated sessions DB and a few busy sessions that multiplied out to
millions of row checks a second during streaming. Build the alias index
once per list reference (the list is replaced wholesale, never mutated)
and look aliases up in O(1).
* perf(desktop): journal each in-flight turn under its own storage key
The v1 journal kept every session's tail in one localStorage key, so each
throttled write re-parsed and re-stringified EVERY busy session's snapshot
— a grid of concurrent streams turned that into a whole-store JSON round
trip dozens of times a second, all on the main thread. Per-session keys
make a write O(own tail) no matter how many other sessions are streaming.
A v1 store migrates on first touch; expired/overflow crash residue is
pruned once per renderer.
* perf(desktop): stress the multitab scenario across grid/streaming/DB axes
The one-stack multitab run hid every cost this round of fixes removed: it
drove hook.publish (store only — no journal, no wiring cache), with an
empty recents list and no closed-tile residue. Streaming now routes through
hook.update (the real gateway write path), and the scenario grows axes for
the workloads users actually hit: --zones splits tiles across visible grid
zones, --streaming caps how many sessions are mid-turn (zone leaders
first), --sessions seeds a lived-in recents list, --dead models settled
sessions no surface references. launch.mjs pins HERMES_DESKTOP_CDP_PORT so
a non-default --port survives the app's own dev-CDP flag.
Fixing the allowlist only helps a fresh install. npm will not re-run an
install script for a package already on disk, so every checkout that
installed while get-windows was blocked stays bricked: `hermes update`
pulls the fix, `npm install` skips the script, and the build fails on the
same missing binding.
Run `npm rebuild get-windows` from the staging step when the binding is
absent, and if that still yields nothing, print the two commands that
recover the checkout by hand instead of the previous advice to reinstall
dependencies, which is exactly what the user already tried. Gated to a
win32 host building for win32, since no other host can produce the binding.
Co-authored-by: JoaoMarcos44 <JoaoMarcos44@users.noreply.github.com>
The published tarball ships lib/binding/napi-9-darwin-unknown-arm64 on every
platform, so a real Windows host has both it and the downloaded win32 binding
— the classify-everything gate threw on the darwin dir and killed every
Windows pack. Stage only bindings naming the target platform (classify still
rejects impostors), stop copyGlobByExt from recursing into lib/binding, and
add a version tripwire so a get-windows bump fails the build until the
lib/windows.js rewrite is re-verified.
Also from review: the renderer answers window.read.respond with empty text
when the IPC invoke rejects (older shell / main-side throw) instead of
stalling the tool's 30s timeout; the tool schema discloses that sibling
Hermes windows are skipped; docs gain read_window_below in both references.
get-windows@9.3.0 (MIT, zero runtime deps on macOS/Linux) is external to the
esbuild bundle and staged into dist/node_modules per target platform: the
universal Swift helper on macOS, the prebuilt N-API binding on Windows
(fail-closed magic-byte validation), nothing on Linux (xprop at runtime).
The staged lib/windows.js is rewritten to load the binding directly so
@mapbox/node-pre-gyp's tree stays out of the package.
⌘K is an overlay that is stateful to itself — pressing it owes the user a
frame immediately, whatever else the shell is doing. It was not built that
way.
`CommandPalette` is mounted for the life of the app, and its body ran
unconditionally: a dozen store subscriptions (connection, desktop version,
client + backend update status/apply, keybinds, worktrees, theme, i18n),
three `useQuery`s, and the group builders that assemble a few hundred rows.
`<Portal>` renders nothing while closed, so none of it was ever visible —
but all of it still ran. An in-flight update rewrites `$updateApply` on
every progress line, and each of those rebuilt the entire row set for a
surface nobody could see.
Split the body into `CommandPaletteBody`, mounted only while the palette is
on screen. A closed palette is now one store subscription. The body is keyed
by open count, so per-open state (search, sub-page) resets by remount and
the explicit close-reset effect goes away, and `mounted` lags `open` by the
150ms exit animation so Radix can still play `data-[state=closed]` instead
of the overlay vanishing.
Rows additionally move behind `useDeferredValue` in their own memo
component. Because that component mounts with the portal, the deferred
initial value applies per open: the first commit is the frame + input, and
the several-hundred-row list arrives in an interruptible follow-up render
rather than blocking the frame the keypress asked for. The empty state is
suppressed while rows are still pending so opening doesn't flash "no
results".
The `enabled: open` gates on the three queries are dropped — the component
only exists when open, so they are inherently lazy, and react-query still
serves a reopen from cache while revalidating.
`submit` measures the scroll jump when a turn is appended; nothing measured
the jump when a session is opened, which is the prepend/settle path. Clicks
sidebar rows and tracks how far the bottom turn moves after first paint.
The sidebar stays mounted beneath every overlay/page, and it subscribes to
$sessions + $workingSessionIds — both tick on every streaming token. An
unmemoized SidebarSessionRow re-rendered the whole list (Codicon, labels,
status dots) on each delta, and that churn bled into every overlay opened on
top: Cron, Profiles, Agents, Starmap, Webhooks, Command Center, Settings.
- SidebarSessionRow: memo() with a custom comparator that ignores the pure
id-forwarding callbacks (fresh closures by design) and compares only the
data that changes what the row paints. Rows bail out while siblings stream.
- SessionStatusDot: the 5 $...SessionIds arrays now read via useStoreSelector
returning this session's boolean, so a dot repaints only when ITS OWN
membership flips, not on every array tick.
- Artifacts: stable cellCtx (useMemo) + memoized Primary/Location/Session
cells so a link-title fetch on one row stops re-rendering the whole table.
Measured before/after (2s idle, sessions streaming), sidebar-fed overlays:
Cron 407->~30 wasted, Profiles 732->~80, Agents 132->9, Starmap 188->56,
Webhooks 154->22.
Gating this behind an opt-in was the wrong call. A dev server already
executes arbitrary local JS — vite's module graph, every postinstall in
node_modules — so a loopback debugging port does not meaningfully widen
what a `npm run dev` session can already do, and `perf:serve` has opened
one unconditionally all along.
Requiring the variable also defeated the point: the tooling exists to be
reached for mid-task, and a capability you must remember to enable before
launching is one you don't have when you need it.
So the port opens on 9222 — the same port scripts/eval.mjs and
scripts/perf/lib/cdp.mjs already default to — for any dev-server run.
HERMES_DESKTOP_CDP_PORT stops being an on-switch and becomes an
override: a different port, or `off` to disable.
The hard gate is unchanged and still checked first: a packaged build
never opens the port, and no env value talks it into it. Neither does an
unpackaged `electron .` against dist/, which is how the packaged app
gets smoke tested.
Refusals only log when they contradict something the developer asked for
(a typo'd port, an explicit `off`). Packaged and dist runs are closed by
design and stay quiet.
The renderer is a Chromium page, and apps/desktop already carries a whole
CDP toolkit for it — scripts/eval.mjs, scripts/perf/lib/cdp.mjs with its
shared SELECTORS map, and the diag-*/probe-* family. None of it can
attach to `hgui` or `npm run dev`, because neither passes
--remote-debugging-port. The only launcher that opens one is
`npm run perf:serve`, which is a separate isolated instance rather than
the app you're looking at.
Add HERMES_DESKTOP_CDP_PORT. When set, the shell opens a CDP port on
loopback so that existing tooling can read the live DOM: computed
styles, geometry, which rule actually won.
Three independent gates, all required, resolved by a pure function in
electron/dev-cdp.ts so the policy is testable without an Electron app:
1. not packaged — a shipped build never opens the port, and this is
checked first so no env combination can talk it into doing so;
2. HERMES_DESKTOP_DEV_SERVER present — an unpackaged `electron .`
against dist/ is how the packaged app gets smoke tested, so it
behaves like the packaged app here;
3. the port explicitly requested and a valid integer.
Default `npm run dev` is unchanged and silent: no port, no nag. An
opt-in that gets refused always logs why, so nobody loses an hour
wondering what isn't listening.
The address is pinned to 127.0.0.1 rather than left to Chromium's
default, and is deliberately not configurable — there's no reason to
expose a renderer debugger off-host and offering the knob invites
someone to try.
scripts/eval.mjs hardcoded :9222 and threw a raw ECONNREFUSED stack when
nothing was there. It now honours the same variable and explains itself.
Switching to a STREAMING session took ~1.4s to settle while an idle
session settled in ~50ms. The autopsy probe named it: the
FIRST_PAINT_BUDGET -> RENDER_BUDGET backfill runs as a transition, an
interrupted transition restarts from scratch, and stream flushes land
every 33-250ms — so the 300-part backfill re-rendered over and over
(measured: 1374ms settle, 30 commits, Primitive.div x2237 for one switch).
Gate the backfill on the thread being idle. The user lands on the live
tail immediately either way; older turns backfill the moment the run
ends, and 'Show earlier' remains the manual path meanwhile.
Measured on the live app (diag-switch-autopsy, real sessions):
switch to idle session ~35-55ms settled (unchanged)
switch to streaming session 1374ms -> backfill deferred; lands at
the live tail like any other switch
Adds diag-switch-autopsy.mjs (per-switch settle/commits/top-renders) and
live-drive.mjs (status/fps/drag one-liners against the running app).
Driving HER real instance (real profile, real transcripts, streams live)
via CDP instead of synthetic tiles finally exposed the remaining stall.
The timeline on a real 60-frame sash drag:
style recalc 2736ms | script 1027ms | layout 89ms
top callsite: pin @ fallback.tsx — 927ms
Two pin-to-bottom ResizeObservers (the bounded tool window's and the
reasoning preview's) pinned on EVERY resize delivery. A sash drag changes
every message's WIDTH once per frame, so each frame ran scrollTop write ->
scrollHeight read across every tool group: a forced write-read reflow
cascade that the render counters could never see (zero React involvement).
Both pins are now height-gated off the RO entry (reflow-free): only
content GROWTH pins. Width-only deliveries return immediately.
Measured on the live app, same drag, before -> after:
fps 11.5 -> 59-60
p95 101ms -> 18ms
slow>33 60/60 -> 1/60
Also in this batch (each was verified live before the next was attempted):
- thread/list: split messageSignature into STRUCTURAL (ids/roles — keys
boundaries + row identity) and WEIGHT (part counts — budget only), and
memoize groups + row JSX. A streamed part-append re-rendered every
turn's boundary via its resetKey prop; explain() measured 540-865
wasted Block renders per drag/stream sample, now {}.
- message-render-boundary: document the structural-only resetKey contract.
- tool/fallback: memoize ToolFallback's part object + ToolEntry/ToolTitle/
ToolGlyph (151 renders each, 100% wasted, on real transcripts).
- use-message-stream: ADAPTIVE flush floor — next flush waits 3x the
measured cost of the last one (33ms floor, 250ms cap), so multi-stream
load degrades text update rate instead of input latency.
- tree-split: preview sash drags with inline flex on the two seam
wrappers, committing the store ONCE on release (fixed-zone sides get
flexBasis only, so a hidden sidebar can't leave a phantom gap).
- debug/: perf-live LoAF long-frame attribution, explain() cascade walker
with changed-hook indices, diag-real-loop/key-latency/switch-trace
probes that drive the real app over CDP.
Typing during 2 live streams: keystroke->paint p50 3.3ms, p95 18.4ms,
zero frames over 33ms. Session switch p50 ~35ms settled; the remaining
~1.3s outlier tail is streaming-session switches (React work-loop, not
style/layout) — next target.
Captures medians of 5 runs for multitab and render-churn so tonight's
wins can't silently regress.
idle-cost is deliberately NOT gated. Its render attribution and idle
commit rate are trustworthy and are what the scenario exists for, but the
drag fps it reports (~0.6fps, p95 814ms) contradicts a direct
single-clock probe of the same gesture on the same build (57fps). I ruled
out sash selection, tile setup, render-counter residue, and a 20s soak,
and could not explain the gap — so the metric ships as a report, not a
gate. Gating CI on a number I can't defend would either fire on a phantom
or mask a real stall.
tier: 'report' is outside GATED ('ci','cold'), so the scenario still runs
and prints but neither compares nor writes a baseline.