Throwaway HERMES_HOME, perf JSON/cpuprofile outputs and typing-lag profiles now land in the Hermes scratch dir; the CONTRIBUTING anti-pattern mention keeps its literal with a no-tmp marker.
Eval probes wrote fixtures, receipts and evidence to hard-coded /tmp paths and
two live A/B tasks literally instructed the model to work in /tmp. They now
derive locations from tempfile.gettempdir() / os.tmpdir() (env overrides kept),
usage examples use relative output names, and the desktop e2e screenshot dirs,
perf scripts and the short-session repro fixture stop naming /tmp. Also adds
the explicit encoding= the windows-footgun check wants in the touched files.
An interrupted extract leaves node_modules/get-windows without package.json
and npm never revisits an existing directory, so the tree stayed broken on
every later update until a manual `npm install get-windows`. Delete such a
dir (workspace hoist or app-local copy) in _install_desktop_workspace_deps
before npm runs so the same update re-extracts it; the staging repair hint
now points at `hermes desktop --force-build` instead of the manual install.
Also drops the unused `exists` injection on findHalfInstalledGetWindowsDir
and the stale "exported for tests" note on missingGetWindowsWarning.
The unresolvable-package case already degraded, but a half-extracted install
(#90829) that keeps package.json while losing lib/binding (win32) or the
macOS helper (darwin) still threw from stageGetWindowsInto and before-pack
rethrew, so the Desktop rebuild died on the optional dep anyway. Stage the
fail-soft JS surface without a binding on win32 (the win32-arm64 precedent),
skip staging on darwin without the helper, and treat a failing native
installer the same way; the runtime reports read_window_below as unavailable
in every one of those states. The classify gate still refuses a present
binary compiled for the wrong platform.
Follow-up to the cherry-picked #109251 (which stops a missing optional
get-windows from killing `npm run build` on win32-x64/darwin):
- The reporter's install (#90829) was not "npm skipped an optional dep" but a
Windows in-place update interrupted by a running Desktop/gateway
(TAR_ENTRY_ERROR): node_modules/get-windows exists with its binding but no
package.json, so require.resolve fails while npm never revisits the
directory. Degrading silently would leave read_window_below dead forever.
`findHalfInstalledGetWindowsDir` walks the same node_modules ancestors
require.resolve does; when the dir is there the warning names it and the
repair (`npm install get-windows --save-exact` in apps/desktop, then
`hermes desktop --force-build`).
- Warning text lives in `missingGetWindowsWarning` (pure, exported) and the
locator is injectable, so tests assert behaviour instead of source text.
- Tests: the old per-platform "fails when absent" cases are replaced by one
every-platform degrade invariant and one half-install → repair-hint case
(both red on origin/main). Docs: desktop.md build-troubleshooting note.
Follow-up to the salvaged #106846 commit (@JoaoMarcos44):
- after-extract.mjs: spell out WHY the stamp moved (electron-builder's
beforeCopyExtraFiles rebuilds the PE with resedit for the ELECTRONASAR
resource; rcedit then cannot commit to that exe, deterministically —
#105629), and why disableAsarIntegrity was not taken.
- after-extract.test.mjs: two invariants — the hook wiring (afterExtract set,
afterPack unset, ASAR integrity still on) and the stamp target
(electron.exe on win32, nothing on other platforms). Red on origin/main.
- set-exe-identity.mjs / scripts/install.ps1: comments still named the
afterPack hook / after-pack.mjs.
The builder takes an optional `platform` (default process.platform) so the
argv test can drive the darwin and non-darwin branches explicitly. The test
no longer redefines process.platform via Object.defineProperty.
The salvaged test returned early off macOS, so Linux CI never executed the
assertion. Stub process.platform to 'darwin' (execFileSync is already mocked)
so the `--sdk macosx clang` argv contract is checked wherever vitest runs, and
add the off-macOS no-op as the control case.
A bare xcrun clang inherits the host default SDK, which may be newer than
the active linker (unknown-arch .tbd stubs at link time). Naming --sdk
macosx explicitly forces the driver and linker to agree on one SDK.
Fixes#113708.
(cherry picked from commit d5fcb6c6ab3de96602528dade2d7f75b7e47ddd5)
* refactor(desktop): extract titleBarOverlayOptions into a tested helper
getTitleBarOverlayOptions in main.ts branched inline on mac/windows/wsl.
Move the decision into titlebar-overlay-width.ts so the WSLg → false rule
(renderer paints its own controls there) is unit-tested next to the width
reservation it pairs with. main.ts keeps a thin caller.
Co-authored-by: null-runner <nicholas.mariani@hotmail.it>
* fix(desktop): renderer-drawn window controls on WSLg
Under WSLg the frameless window (titleBarStyle: 'hidden') had no
minimize/maximize/close. getTitleBarOverlayOptions returned false on the
premise that the RDP host paints replacement controls; it does not for a
frameless window. Electron's native overlay is not the fix either: its
cluster's hit-region drifts from the rendered buttons under the RAIL
compositor.
The renderer now paints Windows-style min/max/close (wslg-window-controls)
routed over a hermes:window-control IPC channel (window-controls.ts →
preload → main). getWindowState reports customWindowControls/isMaximized so
the renderer knows when to mount them and which glyph to show. The cluster
is pinned to TITLEBAR_HEIGHT px (the contrib shell zeroes --titlebar-height
for content subtrees) with 46px native-width caption buttons.
Co-authored-by: Austin Pickett <pickett.austin@gmail.com>
* fix(desktop): wire WSLg window-control IPC and snap maximize onto the work area
Register hermes:window-control in main, report windowControlState from
getWindowState, and re-send window state on maximize/unmaximize so the
renderer's maximize/restore glyph tracks the window.
maximizedBoundsCorrection snaps a WSLg-maximized frameless window back onto
the display work area (RAIL can settle it offset — reported on WSLg 1.0.65).
It is a no-op wherever native maximize already fills the work area, so
healthy compositors are never fought and setBounds cannot loop.
Co-authored-by: null-runner <nicholas.mariani@hotmail.it>
* fix(desktop): call WSLg window-control bridge without the click event
contextBridge structured-clones every argument, and a React SyntheticEvent
is not cloneable, so onClick={controls.minimize} threw "An object could
not be cloned" before ipcRenderer.send ran. The buttons rendered but did
nothing. Wrap the handlers so the bridge is called with no arguments, and
pin that in the test.
* fix(desktop): remove recursive WSLg maximize bounds correction
* fix(desktop): select native Wayland before Electron initialization on WSLg
* fix(desktop): preserve ready pipe and debugger across WSLg launch
* fix(desktop): keep WSLg caption controls scoped to every window
* style(desktop): order window chrome event import
---------
Co-authored-by: null-runner <nicholas.mariani@hotmail.it>
apps/desktop/scripts/set-exe-identity.mjs::stampExeIdentity retried every rcedit
rejection with 0.5/1/2 s backoff. A failure to SPAWN rcedit (bin/rcedit-x64.exe
missing or not executable) is permanent, so the only effect was a 3.5 s delay
before after-pack.mjs's warning.
The npm rcedit wrapper (@malept/cross-spawn-promise) reports a spawn failure as
CrossSpawnError with the errno on `originalError.code`; a non-zero rcedit exit
("Unable to commit changes") is an ExitCodeError with a numeric `code`. Skip the
retry loop when `originalError.code ?? code` is ENOENT/EACCES and rethrow as-is.
Probe: stub rcedit rejecting with that shape -> before 4 attempts, sleeps
[500,1000,2000]; after 1 attempt, no sleeps. Transient-lock retry unchanged.
Part of #112544 (optional review atom).
Preview-pane guest pages (Streamlit's traceback "Ask Google" / "Ask …"
buttons, plain `<a target="_blank">` anchors) could not open anything: the
`<webview>` has no `allowpopups`, so Chromium drops the popup before any
handler runs. Opening from `setWindowOpenHandler` is banned
(GHSA-9f4c-93c8-jc8g, window-open-policy.ts), so this adds an explicit
click bridge instead:
- main.ts installs a guest preload through `will-attach-webview`, keyed on
the `persist:hermes-preview` partition only.
- The preload forwards a clicked `_blank` anchor's resolved href to the
host renderer via `ipcRenderer.sendToHost`; it opens nothing itself.
- PreviewPane admits the URL and routes it through the existing audited
`hermes:openExternal` IPC.
Salvaged from PR #112959 (squash of its two commits).
Fixes#112941
The salvaged commits drop the cpSync/rmSync imports, but stageGetWindowsInto
landed on main after the PR was opened and still called both, so the module
threw ReferenceError on every platform (7 vitest failures on the cherry-picked
head). Route its six sites through copyFileSync/removeDirSync so get-windows
staging survives the same non-ASCII profile paths as node-pty.
preserveRollbackBackup had the same rmSync no-op as cleanStaleAppOutDir: a
surviving .bak makes renameSync fail, so the previous working build gets wiped
instead of kept as rollback material. Same existsSync -> removeDirSync fallback.
Earlier fixes for the same crash: #60480 (@liuhao1024), #61832 (@danilofalcao),
#76211, #103458, #109273, #111590.
Co-authored-by: liuhao1024 <liuhao1024@users.noreply.github.com>
Co-authored-by: danilofalcao <danilofalcao@users.noreply.github.com>
Review follow-ups to the previous commit:
- before-pack.mjs: cleanStaleAppOutDir kept the retrying rmSync (its
EBUSY resilience is worth keeping) but now verifies the tree is gone
and falls back to the libuv-backed removeDirSync when the native
rmSync silently no-ops on a non-ASCII path — previously it logged
'removed stale unpacked dir' while the stale tree survived.
- removeDirSync: export it for reuse; handle a plain file or symlink at
the path (rm -rf semantics) instead of throwing ENOTDIR, and tolerate
broken symlinks via lstat.
- Add a source-level guard test: repo CI runs on Linux where the native
implementations happen to work, so the accented staging test cannot
catch a cpSync/rmSync reintroduction there.
- Link the upstream Node issues in the banner (nodejs/node#61878, fixed
in v24.15.0; nodejs/node#56049, fixed in v24.13.1) and stop
attributing the native rewrite to Node 24 — only the observation was
on v24.11.1.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Node 24's native rewrite of fs.cpSync/fs.rmSync mishandles non-ASCII
Windows paths (observed on v24.11.1 with an accented Windows user name,
i.e. the default %LOCALAPPDATA%\hermes install home):
- recursive cpSync fails with EIO "Access is denied" (errno 5) or
hard-crashes the process (0xC0000409),
- cpSync over an existing file fails with a bogus errno-0
'unlink' / "The operation completed successfully" error,
- rmSync({recursive, force}) silently deletes nothing, so the
half-staged dist/node_modules/node-pty tree poisons every retry.
In practice this bricked the desktop build (and thus install/update)
for every Windows user whose profile path contains accented characters:
stage-native-deps.mjs died copying the conpty prebuild, and each rerun
failed earlier on the stale tree the silent rmSync left behind.
Replace all cpSync/rmSync uses with libuv-backed primitives that handle
these paths correctly on the same node.exe: copyFileSync for single
files, a manual mkdirSync+readdirSync+copyFileSync walk for directory
copies, and a manual unlinkSync/rmdirSync walk (verified with existsSync
afterwards, failing loudly instead of silently) for the dest cleanup.
Add a regression test that stages a fake node-pty tree - including a
nested conpty/ payload - into an accented src/dest path twice, so the
restage exercises the delete-then-recopy path.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The first cut derived the package from `binding-${platform}-${arch}` with an
exact-suffix match, which never matches Windows (`-msvc`) or Linux
(`-gnu`/`-musl`) names, so the repair only ever worked on macOS and
install.sh had to gate it there. Rolldown's own loader already resolves
platform, arch and libc and prints the exact `@rolldown/binding-*` it wanted
in its error chain; parse that instead and drop the gate. Also spawn npm
through a shell on Windows (Node refuses to spawn npm.cmd directly) and trim
the tests to the two invariants (no-op when it loads; installs exactly what
the loader asked for, then re-probes).
The field report saw the commit still failing 5 s after a first attempt,
so 100/300 ms could exhaust before the scanner released the exe. Use
500 ms / 1 s / 2 s (four attempts, 3.5 s worst case) and retry every
rcedit failure instead of one error-message shape: a held handle can also
surface as a load/open error, and a permanent failure only costs 3.5 s
before after-pack.mjs swallows it.
Windows desktop rebuilds intermittently hit rcedit's `Fatal error: Unable
to commit changes` while a real-time file scanner holds the freshly packed
Hermes.exe; `stampExeIdentity` called rcedit exactly once, so the stamp
failed on the first lock. Retry the commit with bounded backoff before
surfacing the original error.
Trimmed from #112552: the after-pack.mjs injection seam and its test —
the hook's swallow-and-warn is pre-existing behaviour on main and needs no
change (the reporter's "Could not install the rebuilt desktop app" comes
from hermes_cli/main_desktop.py::_swap_staged_desktop_app's directory
rename, not from the stamp).
* feat(desktop): give Button a loading prop that swaps label for spinner without layout shift
The label stays in the box, invisible, and the spinner is absolutely
centred over it, so a Connect or Approve button keeps its width while it
works instead of collapsing to a spinner. The approval bar had the same
thrash and moves onto it.
* refactor(desktop): one consent card for connectors and MCP setup
McpSetupTool rendered its own copy of the connector card's markup. It now
renders ConnectorCard for the pending question and ConnectorSummary once
settled, and the card gains what MCP needed: keyboard accelerators, a
source line, a question heading. The card also gets an avatar variant
(40px mark in the left gutter, text and buttons on one column) and a
collapseWhenSettled switch so a connector can stay a full card with a
green Connected pill in the action slot while MCP keeps its one-line
summary. Brand marks for Gmail, Calendar, Drive, Discord, Telegram and
Spotify; Slack via Tabler because simple-icons dropped the mark.
* feat(desktop): connector card drives the agent through manage_connections wait
The offer used to end in a Continue in chat button, and the agent, seeing
an unconnected status, would improvise around the app. Now the card does
what the TUI does. Clicking Connect opens the browser and sends one hidden
line telling the agent to park in manage_connections action=wait for that
slug and to never call connect again (a second link cancels the one being
signed into). Not now sends its own line. A hidden request that lands
while the turn is busy steers it, or queues if the turn just ended.
Which call owns the live card changes too: consecutive calls naming the
same apps are one exchange (connect, the wait, the status that follows),
and the first of the last exchange is the card, so the agent's wait no
longer demotes the card mid-authorization and mints a fresh one below it.
A targeted ask renders one or two bare cards; only a real catalog gets the
header, search and refresh.
* feat(desktop): onboarding connects apps in chat and keeps tasks finishable without them
The welcome chat knew connectors only as preferences to pick and wire up
later, so asked to connect Gmail it invented a Settings page that does not
exist. Both scripts now carry one rule set: status once, one batched
connect for every app named, the card is the ask so write a line and end
the turn, never route around a declined app with another client or
credential. The build handoff checks real connection status instead of
asserting none are connected, and the first task must be finishable, not
free of, the apps they picked. The connectors card explains what
connecting means and reports the count on its Continue button.
* fix(tools): resolve the Nous identity for share_auth profiles in the connector gate
A profile created with share_auth has no auth.json of its own and signs
in through the root store. Every other credential reader falls back to
the global root; the connector gate read HERMES_HOME/auth.json directly,
saw nothing, and stripped manage_connections from the profile's tool
list, so the welcome chat's agent truthfully reported the tool missing.
The gate now goes through get_provider_auth_state.
* fix(agent): name a provider retry backoff on the live status line
The retry status is buffered and replays only when every retry fails, so
during a 60s backoff after a 5xx the user saw a bare spinner. Right after
a connector sign-in landed this read as the agent going silent. The
backoff now also rewrites the live wait notice, which the desktop already
renders in the thread status row; it is transient and clears on recovery.
* test(desktop): connector rehearsal launcher and flagged connector spec
connector-rehearsal.mjs starts the real desktop and backend under a fresh
HERMES_HOME with no copied credentials, a fixed Vite port and CDP on 9344,
so the onboarding connector flow can be driven end to end by hand or from
outside. The Playwright spec covers the flagged connector step.
* fix(desktop): send the agent back into wait when the user keeps waiting after a timeout
The card's Keep waiting re-entered the poll but the agent's own wait had
timed out too and nothing told it to go back in, so it would start
talking mid-authorization. keepWaiting now fires onWaiting like connect
does. Tests also pin that an expired or revoked grant asks the gateway
for reconnect, not connect.
* style(desktop): blank lines in connector-flow test per lint
* feat(desktop): HERMES_SKIP_INTRO=1 / --skip-intro skips the first-run film
The intro is a one-time reveal, so anyone rehearsing the guided chat behind
it sits through it on every fresh HERMES_HOME. The flag rides the existing
launch-flags path (main → preload → renderer) next to guestOnboarding and
only gates isIntroRevealEnabled; the backend never sees it. The rehearsal
launcher sets it.
* fix(desktop): onboarding card Continue stays Done after the transcript rebuilds
The card kept its Done flag in component state. The hidden submit and the
turn-end hydrate both rebuild the message list, so the card remounted with
the flag false and Continue came back live, letting a step be answered
twice. The committed steps now live with the other onboarding answers,
keyed by step, and the first-build chip pick rides the same store.
remember_onboarding projects by key, so the new field never reaches USER.md.
* fix(desktop): no provider picker or free-tier chip over the guided first launch
Two sign-in surfaces leaked into the guide. A credential probe on the
setup profile (a free-tier token mid refresh, a session before its runtime
settled) hit requestDesktopOnboarding and dropped the provider picker over
the chat the user was in; and the statusbar free-tier chip sat there
offering a second sign-in the whole time. Both now yield while the gate
phase is cinematic, guided or handoff. The free tier is the provider for
those phases, and the guide offers sign-in on its own ready screen.
* fix(desktop): onboarding connector picks are real catalog slugs
The picker offered Spotify, GitHub and Stripe, none of which the deployed
connector catalog carries, and spelled Calendar and Drive with hyphens the
gateway does not use. A pick the build chat could not honour ended as
"Spotify isn't in the connector list" after the user had been told to
expect it. The list is now twelve slugs from the live status catalog,
spelled as the gateway spells them; GitHub is out (the terminal has git
and gh), chat channels stay on Messaging. Marks for the new entries; the
Google marks answer both spellings. The build runbook offers the picked
connections in its first turn rather than after the work is underway.
* fix(desktop): the free-tier ready screen never interrupts the guided chat
A readiness round fires when the layout pick assembles the window, and it
raised the free-tier ready screen over the conversation: the user was
dropped into the main app, dismissed it, and came back to a card they had
already answered. The guide is the introduction. The ready screen now
yields while the gate is cinematic, guided or handoff, and the notice is
acked the moment the guided chat takes the screen, not only when the film
does, so a skipped film no longer leaves it pending.
* feat(desktop): tour options that lead to building, and a fork that follows the tour
"Just the basics" and "Show me around" read as a click-through with no
exit; "I'll figure it out" read as declining help. Now Quick tour, Show me
everything, and Skip, let's build something. The script also folds the
fork into the same turn as the tour, so when the user closes the overlay
the next ask is already waiting instead of a transcript that ends on the
tour call.
* feat(desktop): the onboarding connector picker reads the live catalog
A hardcoded list, however carefully copied from today's catalog, is the
next drift. The picker now asks connectors.list through the same
session-owned RPC the connector cards use and offers exactly what the
gateway carries: a curated lead order puts the everyday apps first, chat
channels stay on Messaging, everything else is reachable by search. The
picks are gateway slugs, handed straight to manage_connections. No
catalog (toolset off, gateway unreachable) ends the step honestly with
Skip instead of inventing apps.
* test(desktop): the guided first launch never forces a sign-in
The acceptance criterion the guided onboarding was built to, as a test:
while the gate is cinematic, guided or handoff, the provider picker does
not open and a credential warning is dropped rather than deferred to the
next send. Outside the guide the picker opens as before. Red against the
tree before the guards landed (6 of 9).
* fix(desktop): a relaunch mid-guide resumes the guide, in the guide's shape
Closing the app during the guided first launch and reopening it booted the
normal shell around the persisted solo layout: the connecting splash, the
stock composer and model picker, a small window whose sidebars would not
open, while the gate still read guided. The gate now queues a kickoff for
the guided phase too (the kickoff adopts the existing guide chat by title),
takes the solo shape before the gateway opens rather than after, and the
connecting overlay yields to the guide's own opening. A typed reply in the
composer now closes an ask card and the first-build chips the same way a
click does; the layout card's Continue comes back Done.
* style(desktop): one answeredAfter helper for the ask card and first-build chips
* fix(desktop): the guide takes its shape on the tick the film ends, not after the window shows
Between the film and the greeting the full-size shell painted for a beat:
finishIntroReveal showed the main window, then the kickoff shrank it once
the setup profile answered. The listener on the intro's hidden edge now
takes the guide's shape (solo layout + small centred window) synchronously,
so the window is already the guide when it is shown. One takeGuideShape
owns the pair; kickoff and the boot gate call it idempotently.
* style(desktop): the 'nothing connects yet' line reads first on the connectors card
A delayed browser could miss the 900ms terminal event and spin forever after the updater exited. Retain terminal delivery until the page acknowledges it, bound unavailable-client teardown and failed requests, and preserve a truthful final display.
Fixes#103747. Builds on OutThisLife and Teknium detached handoff work in #83634 and the #75895 quiet-window design. Continues Axl Ibiza Windows update investigation (#60233, #94107, #100763), including source/review contributions carried by merged #93353 and #85170. Existing #102373, #103140, #95719, #97299 and #103632 retain their separate scopes.
One conflict: upstream 6e7c7c7da9 replaced bot-mode-closed-chat-stays-closed.spec.ts with bot-mode-row-click-mirrors-registry.spec.ts while our side had rewired its mock-server import. Kept upstream's replacement and rewired the three new specs importing ./mock-server to the consolidated tests-js copy (symbols verified present).
Widen the salvaged guard from a hand-maintained four-package floor to the
class it stands for: every `dependencies` + `devDependencies` entry in the
desktop workspace manifest. Live probe on this box: a tree holding vite,
katex, electron and electron-builder but missing `@rolldown/plugin-babel`
still passed the floor-only guard, and `vite build` died loading
`vite.config.ts` after `prebuild` had already run. The floor stays as an
unconditional fallback for an unreadable manifest; optionalDependencies
are skipped because npm legitimately omits them (get-windows).
Five new vitest cases (12 total); the two class tests fail when the
manifest union is removed. Refs #86443.
Follow-up to the salvaged #87980: the test kept its own copy of the
build-critical package list (drift hazard) and the module's default
export had no consumer.
Refs #86443
assert-root-install.mjs exists to turn an incomplete root install into one
actionable line instead of a failure deep inside the build. It only ever
checked that vite resolved, so an install covering part of the workspace
graph passed the guard and died later on something else. That is the shape
reported in #86443: the updater's npm install brought in 521 of the 769
packages a full install gives, root node_modules had vite but not katex, and
the build failed on an unresolved katex/dist/katex.min.css with nothing
pointing at the install as the cause. apps/desktop/src/styles.css imports
that stylesheet, so katex is as load-bearing for the renderer bundle as vite
is, and electron / electron-builder are the same for packaging.
Check all four and name every missing one, so a partial install is reported
once and completely rather than one package per build attempt.
Resolution walks node_modules upward the way Node's own lookup does, rather
than going through require.resolve: a package whose exports map does not
expose ./package.json is not resolvable by path even when correctly
installed, and that must not read as missing. It also keeps a dependency
that landed in the app workspace instead of the hoisted root passing.
The guard now runs from prebuild, ahead of npm run clean, so a tree that
cannot build is rejected before the build deletes its own outputs. On this
checkout clean removes build/electron-types and the tsbuildinfo files, not
release/, so this ordering is not by itself what saves a packaged app; it is
the narrow correctness point that a doomed build should not destroy anything
first. build keeps its own call for anyone invoking the build steps directly,
and the check is pure filesystem lookups, so running it twice costs nothing.
The check is extracted as a pure checkRootInstall() returning {ok, error},
matching assert-dist-built.mjs, so it is unit testable without spawning a
process.
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
one-line compat forwarder) and the new implementation already drains
both pipes asynchronously with bounded abandonment, which supersedes
this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
addition inside the consolidated file; rewired the five upstream specs
still importing './mock-server' to the consolidated path (export sets
verified identical) and dropped the superseded apps/desktop/e2e copy.
The packaged app crashed at launch with 'No QueryClient set, use
QueryClientProvider to set one': useQuery in a lazy chunk (session-list-density)
read a second @tanstack/react-query runtime whose QueryClientContext was never
populated by the entry's QueryClientProvider. The source tree was correct — the
duplication happened at build time, because react-query was the one
context-bearing runtime not pinned to a shared vendor chunk, and rolldown's
merge heuristics inline the spare copy into a lazy chunk depending on toolchain
version.
- vite.config.ts: add @tanstack/react-query to the vendor-react
advancedChunks group + dev dedupe list, mirroring the react-router fix.
- assert-dist-built.mjs: fail the build when the 'No QueryClient set'
invariant appears in more than one JS asset (launch-smoke guard).
- assert-dist-built.test.mjs: unit tests for the new invariant check.
- launch-packaged-app.spec.ts: e2e smoke test asserting the packaged app
boots to real UI, not the QueryClient error boundary.
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.
The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.
Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.
The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.
JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.
The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.
The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.
`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.
node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.
The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.
The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.
`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.
The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.
Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
`ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
parallel units and for a plain `npm run check`, in both directions. Against
the 13-leg matrix the count is 13 to 11, and the whole difference is the
three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
Review follow-ups on the compositor spinner and the invalidation scoping.
Spinner CSS:
- Clip each frame to its own box. Braille renders from a system fallback
face (JetBrains Mono has no U+2800 block), whose metrics are not
guaranteed to fit the 1em frame, so neighbouring ink could bleed into
the viewport.
- Name descendants explicitly in the selection guard. The competing
`[data-selectable-text='true'] *` rule has the same (0,1,0)
specificity, so relying on inheritance made the winner depend on
stylesheet order.
- Scope the compositor promotion to spinners that are actually running.
A permanently promoted layer per parked spinner is pure memory at
fan-out breadth, where many sit mounted and paused at once.
- Give every var() the braille default as its fallback, so a missing
custom property degrades to a working spinner rather than an invalid
declaration.
Spinner component: replace the bare `as CSSProperties` cast on the inline
style with an exported GlyphSpinnerVars contract, so a typo in a custom
property name is a compile error rather than a silently dead declaration.
Assistant message:
- Render the inter-agent collapse as a CHILD of the normal body instead
of a competing root. The settled case previously returned a different
element type than the running case, so settling unmounted the whole row
and mounted a fresh one — discarding the DOM the scroll anchor held.
One component, one root, children vary; the truth table is unchanged,
including the collapsed row carrying no tapback listener.
- Collapse AssistantStatusSlot's separate subscriptions into one selector
returning a stable string. The inputs always move together on a status
flip, so reading them separately just multiplied the wake-ups.
- Give StreamingMarker a stable `data-slot` and assert on that rather
than on `span.hidden`.
Repro script: count settled rows by subtracting streaming markers from
message roots instead of `:not(:has(...))`. The selector walked every
row's subtree on each evaluation, inside the very latency window the
probe measures.
Comments: drop the stale translateY(-100%) description, name both pause
triggers, replace hard-coded line-number citations with selector/symbol
ones, note that only the primary window arms the renderer-pause
attribute, and move the forensic trace numbers out of source comments
into the PR.
Delete the three tests that asserted on stylesheet TEXT. AGENTS.md bans
reading source in tests outright, and they demonstrated exactly why: a
var()-fallback edit that changed no rendered pixel broke one of them.
Replacements that exercise the CSS in a real browser follow.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
Removes the two residual invalidators left by the previous commit.
data-streaming: gone from the message root. The flag is not dead -- it is the
settled-row signal for scripts/run-short-session-hang-repro.mjs -- so it moved
to a permanently-mounted, display:none leaf that is a ROOT-LEVEL sibling, and
the repro now matches on the descendant. Placement is load-bearing three ways:
a node inside [data-slot='aui_assistant-message-content'] would steal
:last-child from the stall indicator and change inter-bubble margins mid-stream
(styles.css:1995-2003); keeping it mounted and toggling only the attribute
keeps the per-flip write on a childless node instead of making it a DOM
structure change; display:none costs no layout or paint while querySelectorAll
and :has() still match it.
Renamed to data-message-streaming rather than reusing data-streaming: shiki
puts that exact attribute on deferred code cards, which are descendants of the
message root, so a descendant-matching selector sharing the name would report
any message holding a deferred code card as still streaming.
root isRunning: gone from the standard path. AssistantMessage now dispatches on
interAgentSender, so the collapse gate's live status subscription lives in
InterAgentAssistantMessage and only the rare inter-agent case pays it. The
enter animation captures its enabled flag once off the runtime, non-reactively,
because use-enter-animation.ts parks the value in a ref behind a useCallback([])
identity and consults it only when the callback ref fires at mount -- a live
subscription fed a value the hook already ignores.
Adds inter-agent-collapse.test.tsx: the collapse gate and the marker contract
both had zero coverage, and nothing in the app reads the marker, so a delete
would otherwise look free and silently regress the repro's response gate.
Behavior-identical; invalidation scope only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
require.resolve returns the macOS realpath (/private/var/...) while
os.tmpdir() stays on the /var symlink, so a raw join() deepEqual failed
even though the spawn was correct.
With a GH_TOKEN/GITHUB_TOKEN in the environment, electron-builder auto-selects
the github provider and resolves owner/repo from the repository field, falling
back to reading <projectDir>/.git/config. projectDir is apps/desktop, which has
no .git of its own, and app-builder-lib does not walk up to the workspace root
-- so resolution returned null and threw "Cannot detect repository by
.git/config".
On Linux this fires from onAfterPack for a plain `dir` target: the darwin and
Windows branches return early for non-installer targets, Linux has no such
guard. That is why the same build worked elsewhere.
--publish never keeps `pack` from reaching this at all, but `dist:*` and
test-desktop.mjs still resolve publish config on a machine with a token, so
declare the field too.
Tests call the real app-builder-lib resolver rather than asserting on the text
of package.json, so they track electron-builder's behavior instead of our
formatting.
Co-authored-by: airo7 <airo7@users.noreply.github.com>
Co-authored-by: frankmendes1979 <frankmendes1979@users.noreply.github.com>
The desktop window opened blank white on a fresh install: React threw
"Minified React error #527" before the first paint, from the
`vendor-react-<hash>.js` chunk.
`apps/desktop` pins react and react-dom to the same exact version, but
`vite.config.ts` aliased both to a hardcoded `../../node_modules/<pkg>` —
straight into the monorepo root, where npm is free to hoist a different
react. `@streamdown/math` is a root dependency whose react peer is
`^18.0.0 || ^19.0.0` and which declares no react-dom peer, so npm hoists
the newest react (19.2.8) to the root while react-dom stays at the
version hoisted from the workspaces (19.2.7). react-dom's own peer is
`react: ^19.2.7`, which 19.2.8 satisfies, so the install reports success
and nothing warns. The bundle then shipped react 19.2.8 with react-dom
19.2.7 and React refused to run.
`npm ci` masks this because the lockfile pins the root react to 19.2.7,
which is why CI is green. The recurrence engine is
`_run_npm_install_deterministic()`: when `npm ci` fails it falls back to
`npm install --no-save`, which re-resolves the whole tree and never
records the result — so the split comes back on the next update and
leaves no trace.
Fix the resolution rather than the hoist. The aliases now resolve both
packages from the desktop workspace itself, where npm guarantees the
declared versions are reachable (it nests a copy under the workspace
exactly when the hoisted one differs), so the pair can only ever match.
Pinning react at the npm layer instead was rejected: every manifest-level
pin tried (root dependency, root `overrides`, a scoped override on
`@streamdown/math`) breaks a fresh install with
`ERESOLVE ... peer ink-text-input@"6.0.0" from @hermes/ink@0.0.1`.
Two guards keep it from silently returning:
- `assert-root-install.mjs` (the existing preflight of `build`,
`dev:renderer` and `preview`) now fails the build when the resolved
react and react-dom versions differ, so a split surfaces as an
actionable error instead of a white window.
- A `tests-js` contract test asserts every workspace pins the two to the
same exact version, and that the desktop bundler no longer points at a
hardcoded `node_modules` path.
Verified on a synthetic split tree (root react 19.2.8 / react-dom 19.2.7,
workspace react 19.2.7): the old aliases resolve 19.2.8 + 19.2.7, the new
ones resolve 19.2.7 + 19.2.7, and the preflight exits 1 when the
workspace itself resolves the mismatch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>