Only Log lines were sanitized at the emitter; the manifest-step error string
was still built from the raw child stderr/stdout tail, so a failing
`install.sh --manifest` (print_error writes `${RED}✗${NC}`) put escape bytes
straight into the Setup failure banner. Run stripAnsi / strip_ansi over the
embedded tail on both the Electron and the Tauri bootstrap surfaces (#112675).
Seven near-identical #[test] fns become one table test over the same inputs
(SGR banners, cursor/erase/private modes, OSC with BEL and ST, \r redraws,
sequences cut by the pipe, plain multi-byte text) plus the existing
sanitized_for_ui_only_touches_log_lines. Same coverage, two tests: the
salvage bar is ≤2 invariant tests per fix.
Verified with the function and tests extracted into a standalone
`rustc --test` binary (cargo test on this host fails in libdbus-sys, a
missing system header unrelated to this crate's code).
The macOS setup app drives a shell install script whose output carries
SGR styling, cursor movement, and OSC title commands; the Live output
pane renders log lines as plain text, so those bytes showed up as
mojibake (#112675).
Strip escape sequences at the Rust-to-webview event boundary (bootstrap
and update flows) and collapse carriage-return redraws to the last
visible frame. The on-disk tee keeps the raw bytes for terminal viewing.
Fixes#112675
Trim the Windows-only re-exec matrix from four tests to two: one per
launch builder (std = launcher fast path, tokio = `--update` handoff),
each paired with a different captured stream. stdout and stderr flow
through the same handle-inheritance path (GetStdHandle +
SetHandleInformation on the installer side, Stdio::null() on the child),
so the 2x2 product added runtime without adding a distinct invariant.
Also replace the "(issue TBD)" placeholder on open_macos_app_detached
with the issue number.
`hermes-setup.exe --update` hands off to the rebuilt Desktop right before it
exits. On Windows, CreateProcess duplicates every inheritable handle of the
installer into the Desktop, and the installer's stdout/stderr are inheritable
pipe handles whenever a shell captures its output. The Desktop therefore held
the pipe's write end and the shell kept waiting until the Desktop exited, even
though the update had succeeded (Windows 11, 2026-09-05: still waiting after
18 minutes, while `> file` returned normally). DETACHED_PROCESS detaches the
console only; Rust's std::process passes bInheritHandles=TRUE and has no
handle list (rust-lang/rust#54760), and setting the child's own stdio to null
does not change which stray handles it inherits (measured: same 4 s stall).
Right before the handoff, clear HANDLE_FLAG_INHERIT on the installer's own
std handles (GetStdHandle + SetHandleInformation, best effort, skipping
missing handles of a GUI-launched installer), through one spawn helper used
by both launch sites; the Desktop's stdio is nulled as well so it never uses
the installer's handles at all. The installer's tracing goes to a file, so
nothing is lost. macOS `open` branches get the same nulling for consistency.
Tests (Windows-only; the crate's CI lane runs on Ubuntu): the test binary
re-executes itself as a helper that launches a 4 s sleeper grandchild through
each real builder and the shared spawn helper; the parent asserts the helper's
piped stdout, and separately stderr, reaches EOF in under 2 s and that the
helper's sentinel line arrived. Without the fix all four cases stall ~4 s;
with it, EOF arrives when the helper exits.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The heal decision is extracted into should_heal_self_marker_refusal()
so the contract is testable: heal ONLY on exit 2 + a marker naming this
process. Five tests pin it — self-owned heals, foreign owner (real live
sibling process) never heals, missing/garbage marker never heals,
non-exit-2 never heals, and the full acquire -> refuse -> drop-claim ->
retry-precondition lifecycle with a real UpdateMarkerGuard.
windows-rust-e2e.yml mirrors the wine2e pattern: fires only on
wine2e-rust/** pushes, runs the crate's cargo test --lib on
windows-latest (the shipping platform). The permanent Linux lane stays
authoritative for the unix-gated pipe-drain fixtures.
Main's live_marker_owner has adopted self-owned markers since
160586ff8/dbc2a9c8e (#74761), so the cherry-picked comment's claim that
it 'maps self-ownership to None' is stale. The raw read is still the
right tool — the heal needs the single fact 'does the marker name our
PID' without age/liveness policy folded in.
An updater binary spawns 'hermes update' while holding the update marker
with its own PID. A checkout that predates the HERMES_UPDATE_HANDOFF_PID
env fix (8c76fe19f) and the ancestor-pid fallback runs its pre-pull
update_lock.py, reads that marker as a live foreign update, and exits 2
— and the updater deliberately skips its retry for exit 2, so the
refusal loops forever: the update being refused is the one that ships
the fix, and the failure screen's Retry re-enters the same state.
Detect the case with a raw marker read (live_marker_owner deliberately
maps self-ownership to None, so it cannot answer this), drop our own
claim, and retry the child once with the marker absent. The guard
re-removes on Drop (idempotent) and the desktop is already gone at this
point, so nothing races the brief marker-free window.
Fixes#75788
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Phase 1 leaves on both pipes reaching EOF without an exit status, so the
final wait was the only thing holding the turn -- and it was a bare
child.wait(), which stranded the cancel channel against a child that
closes its pipes and lingers. Poll cancellation there too.
`run_script` and `run_streamed` both left their select loop on stdout EOF
and then ran unbounded post-loop drains, so `child.wait()` sat downstream
of a read that a surviving descendant can hold open forever. Pipe EOF is
not the child's to give: the write end is inherited by every descendant
spawned without its own redirection, and `hermes update` deliberately
runs its build steps with stdout inherited. One resident gateway stranded
the whole update, exit code included.
Both now go through one `pump_child`, which takes the exit status from
waiting on the process and bounds the drain from the moment it exits — a
slow child is not a stuck one, so nothing is metered while it runs. An
abandoned drain says so in the log rather than silently truncating.
Cancelling was not an escape hatch either: `start_kill` reaches the
child, not the grandchild with the handle, so the bounded drain is what
lets a cancel return at all.
Same bound `Invoke-HermesStep` grew in windows.ps1 (#90455), and the same
shape as Go's `exec.Cmd.WaitDelay`.
nanostores 1.4.0-1.4.1 annotate batch() @__NO_SIDE_EFFECTS__. Rollup
(via vite build) honors that and erases a result-unused batch(...) call
as dead code -- callback included. Since d57f94a33/053eb7aab/4e520f085
moved the gateway-switch publication (activate() +
+ ) inside batch(), packaged desktop builds lost the entire
publication: clicking a profile in the rail did nothing at all.
Dev builds and vitest run unminified, so only the packaged app broke.
nanostores 1.4.2 removes the annotation from batch() (it stays on the
creation functions, where it is correct). Bump all three pinned copies
(apps/desktop, apps/bootstrap-installer, ui-tui) and add a regression
test asserting the installed nanostores never re-annotates batch.
* fix(install): time-box the Windows node-deps stage so a stalled npm or Playwright install can't hang setup forever
scripts/install.sh has bounded this same work with run_with_timeout
"$NODE_DEPS_TIMEOUT" (600s default) since #39219, but install.ps1 never got
the guard: Install-NodeDeps ran both `npm install` and `npx playwright
install chromium` unbounded. A stalled registry fetch or a wedged Chromium
archive extraction (#76222, #84614) froze the installer indefinitely -- one
user left it running 12+ hours overnight before asking for help.
Route both invocations through _Invoke-NativeWithTimeout: cmd.exe launches
the native command with its output merged to a log, the parent polls with a
wall-clock deadline and tails new log lines to the console each tick (the
live progress that makes a 3-minute download distinguishable from a hang),
and on timeout taskkill /T /F kills the real process tree and returns 124 --
the same convention as coreutils timeout and bash's run_with_timeout.
Wait-Job was rejected for this: jobs swallow live output and Stop-Job leaves
the npm child running. Windows PowerShell 5.1-safe throughout.
Timeouts surface as a warning with the log path, a note that re-running the
installer resumes (stages are idempotent), and the NODE_DEPS_TIMEOUT env
override for slow links -- mirroring bash.
Fixes#76222.
Closes#84614.
Supersedes #76303.
Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>
* fix(installer): roll stage timers over to hours so an overnight stall doesn't read as "744 hours"
formatElapsed rendered a running stage as m:ss with unbounded minutes: a
node-deps stage left hanging overnight showed "744:38", which the user who
reported the hang understandably read as 744 hours. formatDuration
(completed stages) had the same unbounded-minutes shape.
Move both formatters into src/lib/format.ts (pure, no React) and add the
hour rollover: h:mm:ss live, "Xh Ym" completed. tests-js pins the shapes,
including 744m38s -> 12:24:38.
---------
Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>
Both hand-off sites pre-write the update marker: the in-app Update button
(applyUpdates) and the Windows bootstrap-recovery path
(handOffWindowsBootstrapRecovery). Either one can strand a user on a
pre-#74782 staged installer, and the recovery path is worse — it fires
when the install is already unhealthy, so a refused claim there wedges
the very repair meant to heal it.
Route both through stagedUpdaterSupportsPrewrittenMarker and log the skip
so the reason is visible in desktop.log instead of looking like a missing
write.
Also document on copy_self_to_hermes_home that its --update no-op is what
lets an installer-protocol change strand the entire installed base on a
binary that predates it — the root enabler of this class of bug.
pin brace-expansion to 5.0.8
update concurrently to 10.0.4
update electron-builder to 26.15.3
update eslint to 10.8.0
update eslint-plugin-perfectionist to 5.10.0
update @assistant-ui/react to 0.15.0
update @assistant-ui/react-streamdown to 0.3.8
update radix-ui to 1.6.7
update react-router-dom to react-router@8.3.0 - react-router-dom is no longer a standalone package, it just reexports react-router
remove @radix-ui/react-slot: we import this from `radix-ui`
remove eslint-plugin-react: we imported it, but never actually used it!
Regression for #74761: acquire must succeed when the marker already
names this process (desktop writeUpdateMarker raced ahead), and still
clean up on Drop.
Since #50238 the desktop writes .hermes-update-in-progress with the
spawned updater's PID before UpdateMarkerGuard::acquire runs. Without a
self-PID exclusion, live_marker_owner treated that as a foreign live
owner and every in-app desktop update aborted into a relaunch loop
(#74761). Treat our own PID as adoptable; keep refusing foreign live
updaters.
The cross-process update lock (fe8e4d93d) made the in-progress marker
mutually exclusive across every update entrypoint — but the Tauri
updater holds that marker for its WHOLE run and then spawns
hermes update as a child stage. The child read the marker, found its
own parent's live pid, refused with exit 2, and the GUI mapped that to
"Hermes is still running. Close all Hermes windows and try the update
again." Retry spawns a fresh updater that deadlocks against itself the
same way, so every GUI-driven update dead-ends on the failure screen
with no winnable retry (observed: three consecutive self-refusals in
bootstrap-installer.log within 90 seconds).
Hand the claim off explicitly: update_child_env exports
HERMES_UPDATE_HANDOFF_PID naming the updater's own pid, and
UpdateLock.acquire treats a live holder matching that pid as the lock
we are already running under — run without claiming, and release
leaves the parent's marker untouched. The env var alone grants
nothing: the pid must also be the live marker owner, so a stale or
forged value cannot bypass the lock, and a dashboard-spawned
hermes update (no handoff env) is still refused exactly as before.
The updater keeps a Tauri/Cocoa event loop alive while it relaunches the
desktop, and that loop can outlive app.exit(0). Relying on Drop alone
left a *successful* update looking active -- a live pid holding a fresh
marker -- which blocked desktop startup and, now that the marker is also
the cross-process update lock, every subsequent updater until the age
ceiling expired.
Release explicitly once all install-tree mutations are done, before the
relaunch. complete() is idempotent so Drop still covers the failure and
panic paths. Arms a process-exit fallback so a wedged event loop cannot
leave a finished updater lingering as a live pid.
Co-authored-by: nateEc <nateEc@users.noreply.github.com>
UpdateMarkerGuard::acquire overwrote the in-progress marker
unconditionally, so a Tauri update launched while a dashboard-spawned
"hermes update" was mid-flight simply took the marker and ran a second
updater over the same checkout. That is the race behind the reported
Windows failure: install-mode bootstrap rewound the tree while the
dashboard's updater was still running npm install against it.
acquire now returns Result and refuses when a live foreign owner holds
the marker, and Drop no longer deletes a marker this process does not
own. Liveness matches the Python and Electron readers of the same file:
dead pid or past the shared age ceiling means stale and reclaimable, so
a crashed updater cannot wedge future updates.
Adds a cfg(unix) libc dependency for the signal-0 liveness probe; the
Windows path uses OpenProcess/GetExitCodeProcess.
The macOS launcher fast path gates on hermes_is_installed(), which needs
.hermes-bootstrap-complete next to a built desktop app. Nothing in the Rust
bootstrap pipeline ever wrote that marker -- only install.ps1 did -- so every
reopen of /Applications/Hermes.app re-ran setup instead of launching.
Publish the marker atomically (temp sibling + fsync + rename) because
hermes_is_installed() only checks existence: a torn direct write would arm
the fast path against a half-installed tree. A marker write failure emits
BootstrapEvent::Failed so the installer UI leaves the progress state.
Co-authored-by: giggling-ginger <giggling-ginger@users.noreply.github.com>
Every npm workspace package now defines check:* scripts (check:unit,
check:lint, check:bundle, check:typecheck, etc.) that fan out to
separate matrix runners in CI. The check umbrella script chains all
shards for local dev.
The matrix discovery in the workspaces job queries npm workspaces,
finds check:* scripts (in package.json insertion order), falls back to
check when none exist, and emits an include matrix. No hardcoded
package names — the workflow is fully auto-derived from workspace
metadata.
Previously every package ran a single check script on one worker, and
the fix step (lint:fix + prettier) ran as a separate CI step with
special-cased run_fix gating to avoid running on every shard. Now that
lint is just another check:lint shard, the run_fix field and the fix
step are gone entirely — lint runs in its own runner like everything
else.
Two residual gaps in the #67369 salvage of #67214, found during review:
1. No timeout on the download client. Since #67369, mutable branch pins
hit the network on EVERY run, and the stale-cache fallback only fires
when download() returns Err. A black-holed connection (captive portal,
hung proxy, dropped packets) never errors, so the whole bootstrap hung
at resolve() instead of falling back to the cached script. Verified
live: a request to a non-routable address now errors at the 10s connect
timeout instead of hanging indefinitely.
2. The UTF-8 BOM was only written inside download(). Immutable commit-pin
caches take CachePlan::Reuse and are served untouched forever, and the
stale-fallback path also re-serves the old file - so a BOM-less .ps1
cached by a pre-fix installer kept reproducing the #67193 ANSI-codepage
parse failure on every retry (production builds pin BUILD_PIN_COMMIT,
so this is exactly the retry population). upgrade_cached_script() now
BOM-upgrades legacy .ps1 caches in place (atomic tmp+rename,
best-effort, idempotent) on both reuse paths; .sh untouched.
The salvaged helper cleared its line buffer on entry. Inside run_script's
tokio::select! loop, a stdout line arrival cancels the in-flight stderr
read (and vice versa); read_until had already consumed bytes into the
buffer, and the next call's clear() silently dropped that partial line.
Keep partially-read bytes across cancellation (clear only after a full
line is decoded) and emit an unterminated final line at EOF instead of
swallowing it. Adds a cancellation regression test (fails against the
clear-on-entry version) and an EOF-tail test.
Locks the #67193 invariants: localized PowerShell error bytes survive
decode_console_bytes / read_decoded_line (including CP1252-only 0x91/0x92
punctuation), cached .ps1 files get a single UTF-8 BOM for Windows
PowerShell 5.1 -File, .sh stays BOM-less, and mutable branch caches plan
a refresh with stale-cache fallback.
Windows PowerShell 5.1 emits ParserError text in the console ANSI code page,
but the GUI bootstrap aborted BufReader::lines() on the first non-UTF-8 byte
and Retry kept reusing a poisoned install-main.ps1 for branch pins. Decode
child output with a real Windows-1252 fallback, write a UTF-8 BOM on cached
.ps1 files for -File, and refresh mutable branch/tag caches on each run
(immutable SHAs stay cached).
Fixes#67193
hermes update is a Python CLI writing to a pipe when the Tauri updater or
the desktop's in-app POSIX path spawns it, so CPython block-buffers stdout.
Long quiet steps stream nothing to the progress UI. Worst case is the
pre-update backup (updates.pre_update_backup: true): it can zip multi-GB
archives for minutes while the updater still shows the previous line
('waiting for Hermes to exit...'). Users read that as a hang, cancel a
healthy update, and the orphaned child keeps mutating the install.
Set PYTHONUNBUFFERED=1 in both spawn sites (update_child_env in the Tauri
updater, applyUpdatesPosixInApp in the desktop) so output streams line by
line.
Also make the lock-probe unit test pass on macOS: the packaged payload
lives under Contents/Resources there, and Path::ends_with is
case-sensitive, so the lowercase resources/app.asar assertion only ever
matched the Windows/Linux layouts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Five follow-ups to #57659 from post-merge review:
1. install.ps1: gateway scheduled-task re-enable now runs in a finally
(a thrown Remove-Item/uv venv failure previously stranded the user's
gateway autostart disabled), and tasks that were already disabled
before the install are no longer blindly re-enabled.
2. The venv-python holder guard is no longer bypassed by plain --force
(which the desktop bootstrap passes on every update while its lock
probe only checks hermes.exe/app.asar). New explicit --force-venv is
the escape hatch; --force keeps bypassing only the hermes.exe shim
guard.
3. _detect_venv_python_processes now also catches uv/base-interpreter
trampolines whose exe is outside the venv, via cmdline (venv path or
'-m hermes_cli.main' tied to this install root) and cwd.
4. Missing venv python is now UNHEALTHY on managed installs
(.hermes-bootstrap-complete / .update-incomplete markers) so the
repair lane runs instead of 'Already up to date!'; the repair branch
recreates the venv first when it's gone entirely. Dev checkouts keep
reporting healthy.
5. install.ps1 comment no longer claims a Startup-folder disarm the
code doesn't perform (logon-only, not a mid-install respawner).
Bring Hermes-Setup.exe's UI onto the shared design tokens (self-contained, no
desktop-component coupling) and add two capabilities:
- design: flat stage rows (running step opaque, rest muted), neutral check /
destructive cross, running fourier-flow Loader, hairline --stroke-nous
borders, fill-less log panel; ported BrandMark (nous-girl) + HackeryButton +
Loader standalone; re-synced button variants; de-boxed success/failure.
- theme: follow the OS light/dark via the authoritative Tauri window theme
(theme.ts + onThemeChanged, core🪟allow-theme), with Nous dark seed
colors in styles.css so the --ui-*/--dt-* chain derives correctly.
- updates: split the monolithic "Updating" bar into handoff -> download ->
rebuild (+ install on macOS) stages via a shared update_stages() builder, a
live elapsed timer on the running stage, and a dev-only fake-boot preview
(gated on import.meta.env.DEV, stripped from the shipped bundle).
When a Windows user relaunches Hermes while an in-app update is still
running (the desktop vanished with no progress and looks crashed), the
fresh instance spawns its own dashboard backend. That backend re-locks
the venv shim, the updater's straggler cleanup (force_kill_other_hermes
-> taskkill /F /T /IM hermes.exe) kills it, the launch dies with the 45s
"backend didn't come up" timeout, and the user relaunches into the same
trap -- an infinite respawn/kill loop (#50238).
Root cause: no mutual exclusion between an applying update and a fresh
desktop spawning its own local backend.
Fix: the updater publishes a HERMES_HOME/.hermes-update-in-progress
marker (pid + start time) for the whole run via an RAII drop-guard that
removes it on every exit path (success, early return, panic). A
freshly-launched desktop checks the marker before spawning its local
backend and PARKS until the update finishes -- then brings the backend
up itself (it is the surviving instance; the updater's own relaunch hits
the single-instance lock and quits). A stale marker (dead pid or past a
20-minute ceiling) is pruned so a crashed updater can never strand
future launches. No rogue backend spawns mid-update, so
force_kill_other_hermes has nothing legitimate to kill.
Marker parse/staleness logic is extracted to update-marker.cjs and
unit-tested; the Rust guard has unit tests; the Rust-write <-> JS-read
contract is E2E-verified.
The desktop self-update runs `hermes update` then `hermes desktop
--build-only`, and only relaunches if the rebuild returns 0. The first
`--build-only` can exit nonzero on a still-settling post-update tree or a
network-blocked Electron fetch that the installer's self-heal repaired
mid-run — so both updaters (the Tauri setup binary and the in-app POSIX
path) bailed before the relaunch step. The update landed but the app
never restarted; a manual launch worked because the heal had completed.
Retry `--build-only` once in both paths before failing, mirroring the
retry-once `hermes update` already does (and the CLI `hermes update`'s
own desktop rebuild). A second run builds clean off the healed dist and
is a near-no-op when the first actually succeeded (content-hash stamp).
- update.rs: retry stage 2; add rebuild_needs_retry() + test
- main.cjs: retry via new update-rebuild.cjs helper (behavior-tested)
Upgrade the Vite/esbuild surfaces that kept web, ui-tui, and the bootstrap installer on vulnerable esbuild versions, regenerate the root lockfile, and preserve intentional package+lock dependency edits during update lockfile cleanup.
Make `powershell_under_root` visible under `cfg(test)` so the
%SystemRoot%\System32\WindowsPowerShell\v1.0\powershell.exe layout is
asserted on any host (the rest of the resolution is gated to Windows).
The native Windows installer spawned PowerShell via the bare program name
`powershell.exe`, which trusts PATH to contain
%SystemRoot%\System32\WindowsPowerShell\v1.0. On machines whose PATH was
trimmed or truncated (Windows silently drops entries once the variable
exceeds its length limit), the lookup fails and the spawn dies with
"program not found" before install.ps1 runs at all — the installer then
stalls at "0 of 0 steps".
Resolve PowerShell by absolute path first (%SystemRoot%/%windir%), then
fall back to PATH (powershell 5.1, then pwsh 7), then a bare name as a
last resort. Also include the resolved interpreter in the spawn-failure
context; the old message printed only the script path, which misleadingly
read as if the .ps1 itself was missing.