Route packaged macOS bundles and Light through the updater strategy.
Use electron-updater 6.8.9 and wait for native signature acceptance before
backend teardown. Keep checkout and Store ownership separate.
Share Darwin feed paths between packaging, runtime and publication.
Validate both native feeds, verify streamed artifact hashes, prevent
same-tag artifact replacement, and conditionally update the channel
pointer. Protect live feed references during canary retention.
Use one notarization owner. Require publishing credentials and validate
the stapled app. Keep Windows, Linux and Termux jobs unchanged.
Verified with updater/feed unit and transport tests, release-helper tests,
desktop typechecks, the desktop JS build, and workflow lint. No E2E,
native macOS install, release dispatch or public publication was run.
Microsoft disables ms-appinstaller by default. Download a bounded local descriptor before teardown and open its file association. Read the registered source URI from Windows when no feed override is configured. Remove the invalid default URL and share checker parsing.
Verified real loopback download, size/error boundaries, apply ordering, 35 Electron tests, Electron typecheck, and 8 Python projection tests. Packaged update acceptance remains separate.
Merge upstream 5e645791ac.
Retain the PM feature-flag owner and add upstream connection options.
Use the deny-only window-open policy while trusted external links keep
the existing IPC path. Keep both session-import and external-link copy.
Preserve captured timeout output when adding terminal yield handoff.
Quickstart tests patch the explicit upstream model-assignment owner.
Migrate incoming legacy OS markers to the branch's platforms gate.
Desktop renderer and Electron typechecks passed. Targeted Electron tests
passed (42 tests), Python conflict checks passed (26 tests, 3 skips),
and the plugin-compat import checker passed. CI owns the broad merge gate.
Dependency publication now recovers interrupted config/facts changes before
activation and leases live generations during collection. Receipts retain
update correlation and failed steps across nested command boundaries.
Doctor and desktop surfaces report those failures through shared owners.
Move checkout updates out of the desktop facade. Stage a detached Windows
relaunch waiter before shutdown, with bounded handshake and process-birth
checks. Keep packaged lifecycle tests isolated from the installed app.
Native verification exposed two production races: cron maintenance imported
the interactive CLI and rewrote TERMINAL_CWD, and install-ID reads collided
with first publication. Use the existing owners and locks. Plugin checks
now run at startup and each due-gated housekeeping tick, not after 60 ticks.
Share updater-test mutation boundaries and remove collection-root fixtures.
Separate cold MCP startup from command latency and give the real HTTP drip
test enough time to reach body handling.
Root npm check passed, including packaging. The fixed-tree Windows Python
run reported 44557 passed, one failed, and 1404 skipped, plus one retry-only
HTTP test. Those final failures now pass in a 35-test bounded batch. A real
isolated gateway wrote startup and periodic plugin-check receipts.
Full final-tree CI, bundled Sandbox deployment, and actual App Installer
relaunch remain unverified. docs/pm-audit-status.md records these limits.
A delayed browser could miss the 900ms terminal event and spin forever after the updater exited. Retain terminal delivery until the page acknowledges it, bound unavailable-client teardown and failed requests, and preserve a truthful final display.
Fixes#103747. Builds on OutThisLife and Teknium detached handoff work in #83634 and the #75895 quiet-window design. Continues Axl Ibiza Windows update investigation (#60233, #94107, #100763), including source/review contributions carried by merged #93353 and #85170. Existing #102373, #103140, #95719, #97299 and #103632 retain their separate scopes.
Prepare dependency generations before selecting them. Keep shipped tool
bytes separate from writable additions, and store facts beside their entries.
Validate proposed plugin sets before config publication. Restore the previous
config if the facts write fails.
Consolidate duplicate updater, backup, setup, and voice helpers. Repair
launcher selection, dependency consumers, download ownership, update feeds,
and native Windows process and file handling.
Verification: 206 changed/prior-failing Python files reported 4630 passed,
one failed, and 330 skipped. Fix the remaining Hindsight fixture boundary.
The final targeted rerun reported 234 passed and two skipped. The store
review regression batch reported 83 passed and one skipped. Desktop
TypeScript checks, 56 selected Electron tests, 24 release tests, and the
removed-import/compatibility guards passed.
This is an integration checkpoint, not full audit acceptance. The complete
Python suite has not run on this fixed tree. Crash-atomic plugin publication,
generation cleanup, receipt correlation, and packaged lifecycle acceptance
remain open in docs/pm-audit-status.md.
Task 2 of the gateway-as-MSIX-service plan (settled: user-context,
1903 floor, config-only demand-start):
- before-build.mjs serviceExtensions(): the desktop6:Service fragment
— Executable = the payload launcher hermes.exe (the same distlib
PE serving the AppExecutionAliases; gateway run --service is the
SCM frontend — NO shim binary), Name=HermesGateway,
StartupType=demand (config-only posture: arrives stopped;
`hermes gateway service on` flips to automatic-at-logon),
StartAccount omitted (desktop6 default = installing user's context;
localSystem explicitly rejected). light + store variants render
service-less (store policy is the plan's open risk item). Rides
the existing generated build/msix-extensions.xml via
customExtensionsPath — same mechanism as the aliases + copilot-key
fragments.
- gen-msix-manifest.mjs: minVersion 10.0.17763 → 10.0.18362 (Win10
1903, the desktop6:Service requirement; settled support-matrix
bump — 1809 is a 2018 OS; rides the release notes).
- Verified: the REAL rendered manifest carries the service block
(gen-msix-manifest bundled x64 → desktop6:Service
Name=HermesGateway Executable=...hermes.exe); the staged fragment
regenerates with it (caught + fixed a serviceFragment→
serviceExtensions name bug live via the real beforeBuild import).
A flat-dir makeappx probe fails identically WITH and WITHOUT the
fragment (the probe harness lacks signing identity) — the honest
full validation is the win32 CI lane's real signed pack.
- tests: 4 serviceExtensions contract tests (category/name/exe/
demand-start/one-block/xmlns-root/light-store-empty/name-prop)
in cli-launchers.test.mjs — 15/15 pass; full scripts/ suite: my
files green (1 pre-existing darwin-staging failure is the sibling
agent's uncommitted territory).
The pm bundle deliberately ships uv-cache/ (warm venv rebuilds) and the
arch audit already exempts it — but batchSignAppTree enumerated and
signed every .exe/.dll under it. That wastes hundreds of Azure sign
round-trips on inert sdist/archive artifacts and FAILS when the cache
holds files that were removed between pm bundle staging and the afterPack
walk ("SignTool Error: File not found" — reproduced on a local win32-arm64
build).
Skip agent-payload/uv-cache/ in the batchSignAppTree enumeration (same
exemption class as the audit), pinned by a test asserting a uv-cache exe
is excluded while the rest of the tree is still signed.
Verified: 90 batch-sign tests pass.
A local win32-arm64 build exposed it: audit-bundle-arch fails because
the payload deliberately ships uv-cache/ (pm bundle copies it for warm
rebuilds of the mutable venv) and that cache holds sdists/archives uv
built for ANY arch — x64/ia32 PEs on an arm64 payload. They are inert
cache bytes, never loaded at runtime, the same class as the fetch-*
prune, but the audit treated them as wrong-arch binaries.
Exempt agent-payload/uv-cache/ with a comment tying it to the pm bundle
behavior, and pin it with a test (exempt inside the cache, still audited
outside).
Verified: local arm64 build + audit green (2019 native binaries, all
arm64, 615 exempt stubs); 160 audit tests pass.
Widen the salvaged guard from a hand-maintained four-package floor to the
class it stands for: every `dependencies` + `devDependencies` entry in the
desktop workspace manifest. Live probe on this box: a tree holding vite,
katex, electron and electron-builder but missing `@rolldown/plugin-babel`
still passed the floor-only guard, and `vite build` died loading
`vite.config.ts` after `prebuild` had already run. The floor stays as an
unconditional fallback for an unreadable manifest; optionalDependencies
are skipped because npm legitimately omits them (get-windows).
Five new vitest cases (12 total); the two class tests fail when the
manifest union is removed. Refs #86443.
Follow-up to the salvaged #87980: the test kept its own copy of the
build-critical package list (drift hazard) and the module's default
export had no consumer.
Refs #86443
assert-root-install.mjs exists to turn an incomplete root install into one
actionable line instead of a failure deep inside the build. It only ever
checked that vite resolved, so an install covering part of the workspace
graph passed the guard and died later on something else. That is the shape
reported in #86443: the updater's npm install brought in 521 of the 769
packages a full install gives, root node_modules had vite but not katex, and
the build failed on an unresolved katex/dist/katex.min.css with nothing
pointing at the install as the cause. apps/desktop/src/styles.css imports
that stylesheet, so katex is as load-bearing for the renderer bundle as vite
is, and electron / electron-builder are the same for packaging.
Check all four and name every missing one, so a partial install is reported
once and completely rather than one package per build attempt.
Resolution walks node_modules upward the way Node's own lookup does, rather
than going through require.resolve: a package whose exports map does not
expose ./package.json is not resolvable by path even when correctly
installed, and that must not read as missing. It also keeps a dependency
that landed in the app workspace instead of the hoisted root passing.
The guard now runs from prebuild, ahead of npm run clean, so a tree that
cannot build is rejected before the build deletes its own outputs. On this
checkout clean removes build/electron-types and the tsbuildinfo files, not
release/, so this ordering is not by itself what saves a packaged app; it is
the narrow correctness point that a doomed build should not destroy anything
first. build keeps its own call for anyone invoking the build steps directly,
and the check is pure filesystem lookups, so running it twice costs nothing.
The check is extracted as a pure checkRootInstall() returning {ok, error},
matching assert-dist-built.mjs, so it is unit testable without spawning a
process.
The fast-moving desktop prerelease channel is now "canary" everywhere:
the tag shape (vX.Y.Z-canary.<ts>), the electron-updater/R2 feed dirs
(canary.yml / releases/<os>/canary/), the update-channel consts and CLI
choices, the MSIX build-number derivation, the App Installer channel
paths, and the Windows Store flight var (MS_STORE_CANARY_FLIGHT_ID).
Also renames the scheduled workflow to canary-release.yml and the
release test file to test_release_canary.py, and flips the CLI flags
(--canary / --prune-canaries / prune-canaries subcommand).
Unrelated "nightly" mentions are untouched: Brave's own browser channel
(browser_connect), cron scheduling prose (README, i18n, cron/browser/
kanban docs, zh-Hans), upstream skill docs (comfyui/unsloth/torchtitan),
evals fixtures, Node's node-nightly prereleases, and cron job names in
gateway tests.
Note: MS_STORE_NIGHTLY_FLIGHT_ID was renamed to MS_STORE_CANARY_FLIGHT_ID
in the workflow — the matching repo/org variable on GitHub must be
renamed in repo settings for the Store flight ring to keep working.
process.windowsStore is true for any MSIX package — App Installer
sideloads included — not just Microsoft Store deployments. The updater
mechanism resolver keyed on it, so out-of-store installs resolved to
external and showed the static 'bundled install' line instead of the
app-installer check/apply/relaunch flow.
The store-vs-sideload distinction is a build-time fact. Bake it into
the install stamp (HERMES_DESKTOP_VARIANT=store vs bundled) and have
isWindowsStore() trust the stamp when present, falling back to the
Electron flag only for dev runs / legacy stamps.
- write-build-stamp.mjs: buildStampPayload() emits payload/store/
distribution/updateMechanism/tag for staged desktop builds; dev and
legacy builds keep the old 5-field shape
- install-stamp.ts: InstallStamp gains store?: boolean
- main.ts: isWindowsStore() reads INSTALL_STAMP.store; loadInstallStamp
mirrors the field from disk
- tests: buildStampPayload variant matrix (bundled/store/light/legacy)
stripFetchCache did two things: dropped fetch-<sha> download-cache dirs and
dropped the chromium-* store entries. Both were wrong for the current flow:
- fetch-* is now pruned by pm gc during the bundle (the payload and CI
cache already ship only live store entries), so the function was belt
and suspenders.
- chromium-* was DROPPED while signNestedChromium (the autosign pass
right after it) was trying to SIGN it — the strip ran first, so the
signing found nothing. Chrome's Mach-O stayed unsigned, Apple's notary
rejected the app, and the shipped mac bundle had no browser at all.
Rip stripFetchCache out entirely: the payload keeps chromium, and the
darwin autosign pass now has the chromium trees to sign (--deep over the
.app, per-file over loose Mach-O) so notary accepts them. The Windows
payload's chromium binaries are covered by the batch Authenticode pass.
The payload batch signs ~64k PEs serially: each chunk is a fresh
signtool child that cold-initializes the .NET dlib, auths against Azure,
and does a separate RFC3161 round-trip per file. Both Azure and the
timestamp server are per-file network waits, so the whole tree was one
long serial chain.
Two changes:
- sign chunks concurrently, capped at DEFAULT_CONCURRENCY (4) via a
small worker pool — N children multiply throughput ~Nx.
- split signing from timestamping: the sign pass is Azure-only (no
/tr,/td), then a second pure-RFC3161 timestamp pass (no /dlib,/dmdf)
retries per chunk. A flaky timestamp server can no longer hold signed
files hostage or force a full re-sign; the retry beats a rebuild.
The product exe and .msix/.msixbundle keep their existing per-file
signing paths untouched. afterPack now awaits the batch.
Wire the desktop app onto the pm store for real distribution:
- MSIX bundle: electron-builder config, appx assets, manifest, copilot
key + deep-link routing, App Installer + Windows Store variant
(sign only the msix; inner binaries covered by the package block map)
- Rust CLI shim (apps/desktop/shim) — bundled builds run from the store
python + shim, never the venv; payload symlinks relativized so the
relocatable venv survives relocation
- Cloudflare R2 release pipeline: publish binaries + update feeds,
nightly channels/tags, stamp-first version resolution
- Update system: gate, uninstall steward, boot bootstrap, release
channels, update receipts
- install.ps1 reduced to a 361-line stage-protocol bootstrapper (heavy
deps are pm's job); darwin updater + update-channel mirror ripped
- doctor: main's re-landed TCC anchor kept, termux branches removed
Rebuilt from ethie/pm onto the pm-store stack. 22 hot files hand-merged;
uv.lock + package-lock.json keep main's newer dep tree; test_engines
reads the pm/lock.json pin; lazy_deps.py deleted (all 222 importers
migrated to pm in the foundation commit).
Introduce the pm store: a unified, hash-verified package store that
replaces lazy_deps and the old installer's ad-hoc tool downloads.
Store tools are provisioned on PATH (ffmpeg, node/npm via pinned uv),
with a resumable 8-way downloader, verify() returning failure reasons,
and adopt() made EPERM-safe. chromium ships in the payload for every
target. The 3600-line install.sh is replaced by a staged bootstrapper
(heavy deps are pm's job after this); setup-hermes.sh, Dockerfile and
nix pin tables are rewired onto the store. Old install-script tests,
lazy_deps/managed_uv/build_info, and the ps1/bash installer test
batteries are removed with the machinery they tested.
Rebuilt from ethie/pm onto upstream/main (ac6c8028e0) after the
utf-8-sig sweep. 16 hot files (main also churned them) hand-merged:
platform adapters, main.py, electron/main.ts, tui_gateway/server.py,
cua_backend, installer-tests workflow, install.sh (full rewrite),
setup-hermes.sh, plugins doc.
The packaged app crashed at launch with 'No QueryClient set, use
QueryClientProvider to set one': useQuery in a lazy chunk (session-list-density)
read a second @tanstack/react-query runtime whose QueryClientContext was never
populated by the entry's QueryClientProvider. The source tree was correct — the
duplication happened at build time, because react-query was the one
context-bearing runtime not pinned to a shared vendor chunk, and rolldown's
merge heuristics inline the spare copy into a lazy chunk depending on toolchain
version.
- vite.config.ts: add @tanstack/react-query to the vendor-react
advancedChunks group + dev dedupe list, mirroring the react-router fix.
- assert-dist-built.mjs: fail the build when the 'No QueryClient set'
invariant appears in more than one JS asset (launch-smoke guard).
- assert-dist-built.test.mjs: unit tests for the new invariant check.
- launch-packaged-app.spec.ts: e2e smoke test asserting the packaged app
boots to real UI, not the QueryClient error boundary.
Every Linux lane that does real work ran on a 4-core `ubuntu-latest`. The
Python suite and the JS checks were split into many small jobs to make that
size usable. Each split job repeated the full setup. In most of the JS jobs
the repeated setup cost more than the work.
The work lanes move to larger runners. Then the splits that existed only to
make small runners usable go away.
Python tests: 12 slices become 1 job on a 96-core runner. Slicing cost a
matrix job, a duration cache, a per-slice artifact and a merge job. 96 cores
clear the floor that the slowest single test file sets, which is about 82s. A
second slice divides work that is already at that floor, and adds a second
setup. Duration data from run 32522943054 gives the numbers behind this: 3178
files, 11645s in series.
The worker count is explicit, because `run_tests.sh` defaults to twice the
core count. A later commit sets it from a measurement on this hardware.
JS checks: 14 jobs become 1. The matrix paid about 371s of repeated setup to
spread about 612s of work. One larger runner installs one time. The three UI
shard scripts and `run-ui-shard.mjs` are therefore removed, because the
unsharded `test:ui` covers the same tests.
The unit of parallel work inside that job is a CHECK, and not a workspace.
apps/desktop is most of the payload, and its own `check` is a serial && chain.
A spread across workspaces alone therefore leaves that chain as the long pole.
A package that declares `check:*` sub-scripts gives one unit for each
sub-script. That is the same selection rule the matrix used.
The loop lives in `.github/scripts/run-workspace-checks.mjs`, so the same
sequence runs on a laptop. It runs 11 units together, buffers the output of
each one, and fails at the end with the full list. Children that share one
stdout interleave their lines and make a failure hard to read.
`npm run --ws check` stops at the first workspace that fails.
`check:test:plugins` joins the desktop `check` script. The matrix prefers
`check:*` sub-scripts over the plain `check` script, so `check:test:plugins`
ran only as its own leg. Without this change the merge drops that suite and
the job stays green.
node_modules is cached on the lockfile, and `npm ci` is skipped on an exact
hit. The `cache: npm` option of `setup-node` caches only the ~/.npm tarball
cache, which leaves the extract and the postinstalls to pay again.
The arm64 image build stays on a native arm64 runner. A build of linux/arm64
on an x64 host uses emulation.
The docker test lane caps its workers at the core count. Each of those tests
drives a container, so the docker daemon sets the limit and not the processor.
`.github/actionlint.yaml` declares the runner labels. actionlint knows the
GitHub-hosted labels only, and an undeclared label reads as an error that
hides the real findings.
The `detect` job checks out one file through a sparse checkout, and its
timeout drops to 1 minute. It reads
`scripts/ci/classify_changes.py` and nothing else.
Verification:
- actionlint reports 9 findings across all workflows. An unmodified HEAD with
the same config reports the same 9. This change adds none.
- A wrong label still fails. actionlint reports `ubuntu-latest-32-cor` and
`ubuntu-latest-32-arm-cores`.
- Every changed workflow parses, and `name` parses as a string.
- A replay of the `save-durations` merge step against a three-artifact layout
returns all 3178 entries.
- An expansion of the npm script graph gives the same leaf commands for the
parallel units and for a plain `npm run check`, in both directions. Against
the 13-leg matrix the count is 13 to 11, and the whole difference is the
three UI shards that collapse into one unsharded `check:test:ui`.
- `--list` reports the 11 units, and a full local run completes and reports
the time of each unit.
- The runner labels cannot be verified here. The first real run is the test.
Review follow-ups on the compositor spinner and the invalidation scoping.
Spinner CSS:
- Clip each frame to its own box. Braille renders from a system fallback
face (JetBrains Mono has no U+2800 block), whose metrics are not
guaranteed to fit the 1em frame, so neighbouring ink could bleed into
the viewport.
- Name descendants explicitly in the selection guard. The competing
`[data-selectable-text='true'] *` rule has the same (0,1,0)
specificity, so relying on inheritance made the winner depend on
stylesheet order.
- Scope the compositor promotion to spinners that are actually running.
A permanently promoted layer per parked spinner is pure memory at
fan-out breadth, where many sit mounted and paused at once.
- Give every var() the braille default as its fallback, so a missing
custom property degrades to a working spinner rather than an invalid
declaration.
Spinner component: replace the bare `as CSSProperties` cast on the inline
style with an exported GlyphSpinnerVars contract, so a typo in a custom
property name is a compile error rather than a silently dead declaration.
Assistant message:
- Render the inter-agent collapse as a CHILD of the normal body instead
of a competing root. The settled case previously returned a different
element type than the running case, so settling unmounted the whole row
and mounted a fresh one — discarding the DOM the scroll anchor held.
One component, one root, children vary; the truth table is unchanged,
including the collapsed row carrying no tapback listener.
- Collapse AssistantStatusSlot's separate subscriptions into one selector
returning a stable string. The inputs always move together on a status
flip, so reading them separately just multiplied the wake-ups.
- Give StreamingMarker a stable `data-slot` and assert on that rather
than on `span.hidden`.
Repro script: count settled rows by subtracting streaming markers from
message roots instead of `:not(:has(...))`. The selector walked every
row's subtree on each evaluation, inside the very latency window the
probe measures.
Comments: drop the stale translateY(-100%) description, name both pause
triggers, replace hard-coded line-number citations with selector/symbol
ones, note that only the primary window arms the renderer-pause
attribute, and move the forensic trace numbers out of source comments
into the PR.
Delete the three tests that asserted on stylesheet TEXT. AGENTS.md bans
reading source in tests outright, and they demonstrated exactly why: a
var()-fallback edit that changed no rendered pixel broke one of them.
Replacements that exercise the CSS in a real browser follow.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01QwTc9XqUjhbay446VjugHZ
Removes the two residual invalidators left by the previous commit.
data-streaming: gone from the message root. The flag is not dead -- it is the
settled-row signal for scripts/run-short-session-hang-repro.mjs -- so it moved
to a permanently-mounted, display:none leaf that is a ROOT-LEVEL sibling, and
the repro now matches on the descendant. Placement is load-bearing three ways:
a node inside [data-slot='aui_assistant-message-content'] would steal
:last-child from the stall indicator and change inter-bubble margins mid-stream
(styles.css:1995-2003); keeping it mounted and toggling only the attribute
keeps the per-flip write on a childless node instead of making it a DOM
structure change; display:none costs no layout or paint while querySelectorAll
and :has() still match it.
Renamed to data-message-streaming rather than reusing data-streaming: shiki
puts that exact attribute on deferred code cards, which are descendants of the
message root, so a descendant-matching selector sharing the name would report
any message holding a deferred code card as still streaming.
root isRunning: gone from the standard path. AssistantMessage now dispatches on
interAgentSender, so the collapse gate's live status subscription lives in
InterAgentAssistantMessage and only the rare inter-agent case pays it. The
enter animation captures its enabled flag once off the runtime, non-reactively,
because use-enter-animation.ts parks the value in a ref behind a useCallback([])
identity and consults it only when the callback ref fires at mount -- a live
subscription fed a value the hook already ignores.
Adds inter-agent-collapse.test.tsx: the collapse gate and the marker contract
both had zero coverage, and nothing in the app reads the marker, so a delete
would otherwise look free and silently regress the repro's response gate.
Behavior-identical; invalidation scope only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
require.resolve returns the macOS realpath (/private/var/...) while
os.tmpdir() stays on the /var symlink, so a raw join() deepEqual failed
even though the spawn was correct.
With a GH_TOKEN/GITHUB_TOKEN in the environment, electron-builder auto-selects
the github provider and resolves owner/repo from the repository field, falling
back to reading <projectDir>/.git/config. projectDir is apps/desktop, which has
no .git of its own, and app-builder-lib does not walk up to the workspace root
-- so resolution returned null and threw "Cannot detect repository by
.git/config".
On Linux this fires from onAfterPack for a plain `dir` target: the darwin and
Windows branches return early for non-installer targets, Linux has no such
guard. That is why the same build worked elsewhere.
--publish never keeps `pack` from reaching this at all, but `dist:*` and
test-desktop.mjs still resolve publish config on a machine with a token, so
declare the field too.
Tests call the real app-builder-lib resolver rather than asserting on the text
of package.json, so they track electron-builder's behavior instead of our
formatting.
Co-authored-by: airo7 <airo7@users.noreply.github.com>
Co-authored-by: frankmendes1979 <frankmendes1979@users.noreply.github.com>
The desktop window opened blank white on a fresh install: React threw
"Minified React error #527" before the first paint, from the
`vendor-react-<hash>.js` chunk.
`apps/desktop` pins react and react-dom to the same exact version, but
`vite.config.ts` aliased both to a hardcoded `../../node_modules/<pkg>` —
straight into the monorepo root, where npm is free to hoist a different
react. `@streamdown/math` is a root dependency whose react peer is
`^18.0.0 || ^19.0.0` and which declares no react-dom peer, so npm hoists
the newest react (19.2.8) to the root while react-dom stays at the
version hoisted from the workspaces (19.2.7). react-dom's own peer is
`react: ^19.2.7`, which 19.2.8 satisfies, so the install reports success
and nothing warns. The bundle then shipped react 19.2.8 with react-dom
19.2.7 and React refused to run.
`npm ci` masks this because the lockfile pins the root react to 19.2.7,
which is why CI is green. The recurrence engine is
`_run_npm_install_deterministic()`: when `npm ci` fails it falls back to
`npm install --no-save`, which re-resolves the whole tree and never
records the result — so the split comes back on the next update and
leaves no trace.
Fix the resolution rather than the hoist. The aliases now resolve both
packages from the desktop workspace itself, where npm guarantees the
declared versions are reachable (it nests a copy under the workspace
exactly when the hoisted one differs), so the pair can only ever match.
Pinning react at the npm layer instead was rejected: every manifest-level
pin tried (root dependency, root `overrides`, a scoped override on
`@streamdown/math`) breaks a fresh install with
`ERESOLVE ... peer ink-text-input@"6.0.0" from @hermes/ink@0.0.1`.
Two guards keep it from silently returning:
- `assert-root-install.mjs` (the existing preflight of `build`,
`dev:renderer` and `preview`) now fails the build when the resolved
react and react-dom versions differ, so a split surfaces as an
actionable error instead of a white window.
- A `tests-js` contract test asserts every workspace pins the two to the
same exact version, and that the desktop bundler no longer points at a
hardcoded `node_modules` path.
Verified on a synthetic split tree (root react 19.2.8 / react-dom 19.2.7,
workspace react 19.2.7): the old aliases resolve 19.2.8 + 19.2.7, the new
ones resolve 19.2.7 + 19.2.7, and the preflight exits 1 when the
workspace itself resolves the mismatch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The desktop journal synchronously read, parsed, cloned, and rewrote one
aggregate localStorage value while streamed turns were repainting. Large
tool results and multi-session state could therefore block the renderer and
leave the app unresponsive, while the existing macOS diagnostic path lacked
a real native hide/restore regression check.
Store bounded recovery projections under per-session keys, migrate legacy v1
data once, isolate quota and storage failures, and preserve the newest
recoverable tail without allowing oversized writes to replace valid state.
Add a real Electron/CDP macOS-arm64 A/B harness with native visibility control,
renderer heartbeat and Settings/composer/transcript checks, plus focused
regressions. Keep bulk tool payloads, diagnostics, and the existing recovery
merge behavior out of the hot path.
Fixes#63047
The bench timed two of the six RPCs the fix touches. Extend it to
image.attach, pdf.attach, clipboard.paste and image.detach so every changed
handler carries a number rather than an inference.
Also report which surfaces reach these RPCs at all, since "why was the GUI
special" is the first question the fix invites. CLI attaches inline in its
own turn path with the agent already built, so it cannot reach the stall;
the TUI calls the same RPCs and was equally exposed. The difference was hit
rate, not code path.
gateway_attach_bench.py drives the real dispatcher with a session whose agent
build is still running and times each attach RPC against prompt.submit as the
control — the harness that located the stall and measures it.
image-attach-bench.mjs times the renderer-side transforms (file read, base64,
RPC frame, embedded-image extraction, render-weight walk) across image sizes.
It is what ruled the renderer out: ~26ms total at 3MB.
Review findings: spawnSync('npm', ...) without shell fails on Windows
(npm is npm.cmd; Node >=18.20 throws EINVAL — same handling as
test-desktop.mjs and stage-native-deps.mjs), and a spawn-level failure
exited 1 with no diagnostic. CI is ubuntu-only but the desktop workspace
supports local Windows dev.
Closes the silent-skip hole reviewers flagged: with the index/count
hardcoded in three sibling strings, a copy-paste slip (shard-2of3 running
--shard=1/3) or a partial 3->4 migration would silently skip a third of
the 428-file suite while CI stays green.
scripts/run-ui-shard.mjs parses N/M from npm_lifecycle_event (the script
NAME is the single source of truth), validates the package's shard family
is exactly 1..M for one M, and delegates through 'npm run test:ui' so the
vitest command stays single-sourced.
Mutation-verified: shard-9of3 name -> exit 1 'index out of range';
adding shard-4of4 beside the 3-family -> exit 1 'must form exactly 1..M';
correct invocation runs shard 2/3 (142 files) identically to before.
* test(desktop): stress long agent sessions in the multitab perf scenario
--tools seeds every transcript with settled tool rounds and drives the live
stream as a working agent turn (tool calls opened and completed between text
chunks), and each tile reveal is timed to next paint (reveal_max_ms) so deep
transcripts report their mount cost.
* perf(desktop): hold the transcript window cut steady while streaming
A fresh weight-walk per store flush slid the cut forward one message at a
time, ~30x/s, and every slide re-indexed the whole windowed transcript —
each row rendered a different message and the runtime repository took its
O(window) rebuild path instead of the one-message update. advanceTranscriptWindow
anchors the cut to a message id and re-cuts once per ~half page of new
content instead of once per flush.
* perf(desktop): stable rows, stepped backfill, pane-shared render budget
Three thread-list fixes for long streaming sessions: memoize the visible-
groups slice and each turn row so a budget-cut advance no longer re-renders
every mounted turn per streamed token; raise the first-paint backfill in
BACKFILL_STEP slices (one bounded commit per frame) instead of a single
20-to-600 transition whose commit landed as a 780ms freeze mid-stream; and
share RENDER_BUDGET across mounted panes so a 4-way grid mounts a quarter
page per pane instead of 4x the fibers.
* fix(desktop): evict settled session states nothing on screen references
Closing a tile never removed its runtime's entry from $sessionStates, so
every tile ever closed parked its full transcript in the map for the life
of the process. Each leftover entry taxes every subsequent stream flush —
the map is spread-copied per delta and the busy/attention/draft projections
walk every entry per publish — so the app got slower the longer it ran,
which users read as "I need to clean my sessions/dbs".
Publish now evicts a settling state when no tile and not the primary view
holds its runtime (transition side effects still fire, so the settle keeps
its unread dot), and closing a tile drops an already-settled state on the
spot. Busy and needs-input states stay: background turns feed the sidebar
dots, and a first publish always lands because a resume can publish a beat
before the surface binds the runtime.
16 tiles streaming in a 2x2 grid with a day's worth of closed-tile residue:
worst-second 34 -> 58 fps, p99 frame 90 -> 28 ms, longtasks 37 -> 0.
* perf(desktop): index lineage aliases per sessions-list reference
lineageAliases scanned the whole recents list per call, and it is called
per cached session state per status projection per message delta — with a
populated sessions DB and a few busy sessions that multiplied out to
millions of row checks a second during streaming. Build the alias index
once per list reference (the list is replaced wholesale, never mutated)
and look aliases up in O(1).
* perf(desktop): journal each in-flight turn under its own storage key
The v1 journal kept every session's tail in one localStorage key, so each
throttled write re-parsed and re-stringified EVERY busy session's snapshot
— a grid of concurrent streams turned that into a whole-store JSON round
trip dozens of times a second, all on the main thread. Per-session keys
make a write O(own tail) no matter how many other sessions are streaming.
A v1 store migrates on first touch; expired/overflow crash residue is
pruned once per renderer.
* perf(desktop): stress the multitab scenario across grid/streaming/DB axes
The one-stack multitab run hid every cost this round of fixes removed: it
drove hook.publish (store only — no journal, no wiring cache), with an
empty recents list and no closed-tile residue. Streaming now routes through
hook.update (the real gateway write path), and the scenario grows axes for
the workloads users actually hit: --zones splits tiles across visible grid
zones, --streaming caps how many sessions are mid-turn (zone leaders
first), --sessions seeds a lived-in recents list, --dead models settled
sessions no surface references. launch.mjs pins HERMES_DESKTOP_CDP_PORT so
a non-default --port survives the app's own dev-CDP flag.
Fixing the allowlist only helps a fresh install. npm will not re-run an
install script for a package already on disk, so every checkout that
installed while get-windows was blocked stays bricked: `hermes update`
pulls the fix, `npm install` skips the script, and the build fails on the
same missing binding.
Run `npm rebuild get-windows` from the staging step when the binding is
absent, and if that still yields nothing, print the two commands that
recover the checkout by hand instead of the previous advice to reinstall
dependencies, which is exactly what the user already tried. Gated to a
win32 host building for win32, since no other host can produce the binding.
Co-authored-by: JoaoMarcos44 <JoaoMarcos44@users.noreply.github.com>
The published tarball ships lib/binding/napi-9-darwin-unknown-arm64 on every
platform, so a real Windows host has both it and the downloaded win32 binding
— the classify-everything gate threw on the darwin dir and killed every
Windows pack. Stage only bindings naming the target platform (classify still
rejects impostors), stop copyGlobByExt from recursing into lib/binding, and
add a version tripwire so a get-windows bump fails the build until the
lib/windows.js rewrite is re-verified.
Also from review: the renderer answers window.read.respond with empty text
when the IPC invoke rejects (older shell / main-side throw) instead of
stalling the tool's 30s timeout; the tool schema discloses that sibling
Hermes windows are skipped; docs gain read_window_below in both references.
get-windows@9.3.0 (MIT, zero runtime deps on macOS/Linux) is external to the
esbuild bundle and staged into dist/node_modules per target platform: the
universal Swift helper on macOS, the prebuilt N-API binding on Windows
(fail-closed magic-byte validation), nothing on Linux (xprop at runtime).
The staged lib/windows.js is rewritten to load the binding directly so
@mapbox/node-pre-gyp's tree stays out of the package.
⌘K is an overlay that is stateful to itself — pressing it owes the user a
frame immediately, whatever else the shell is doing. It was not built that
way.
`CommandPalette` is mounted for the life of the app, and its body ran
unconditionally: a dozen store subscriptions (connection, desktop version,
client + backend update status/apply, keybinds, worktrees, theme, i18n),
three `useQuery`s, and the group builders that assemble a few hundred rows.
`<Portal>` renders nothing while closed, so none of it was ever visible —
but all of it still ran. An in-flight update rewrites `$updateApply` on
every progress line, and each of those rebuilt the entire row set for a
surface nobody could see.
Split the body into `CommandPaletteBody`, mounted only while the palette is
on screen. A closed palette is now one store subscription. The body is keyed
by open count, so per-open state (search, sub-page) resets by remount and
the explicit close-reset effect goes away, and `mounted` lags `open` by the
150ms exit animation so Radix can still play `data-[state=closed]` instead
of the overlay vanishing.
Rows additionally move behind `useDeferredValue` in their own memo
component. Because that component mounts with the portal, the deferred
initial value applies per open: the first commit is the frame + input, and
the several-hundred-row list arrives in an interruptible follow-up render
rather than blocking the frame the keypress asked for. The empty state is
suppressed while rows are still pending so opening doesn't flash "no
results".
The `enabled: open` gates on the three queries are dropped — the component
only exists when open, so they are inherently lazy, and react-query still
serves a reopen from cache while revalidating.
`submit` measures the scroll jump when a turn is appended; nothing measured
the jump when a session is opened, which is the prepend/settle path. Clicks
sidebar rows and tracks how far the bottom turn moves after first paint.
The sidebar stays mounted beneath every overlay/page, and it subscribes to
$sessions + $workingSessionIds — both tick on every streaming token. An
unmemoized SidebarSessionRow re-rendered the whole list (Codicon, labels,
status dots) on each delta, and that churn bled into every overlay opened on
top: Cron, Profiles, Agents, Starmap, Webhooks, Command Center, Settings.
- SidebarSessionRow: memo() with a custom comparator that ignores the pure
id-forwarding callbacks (fresh closures by design) and compares only the
data that changes what the row paints. Rows bail out while siblings stream.
- SessionStatusDot: the 5 $...SessionIds arrays now read via useStoreSelector
returning this session's boolean, so a dot repaints only when ITS OWN
membership flips, not on every array tick.
- Artifacts: stable cellCtx (useMemo) + memoized Primary/Location/Session
cells so a link-title fetch on one row stops re-rendering the whole table.
Measured before/after (2s idle, sessions streaming), sidebar-fed overlays:
Cron 407->~30 wasted, Profiles 732->~80, Agents 132->9, Starmap 188->56,
Webhooks 154->22.