The suite now runs as one job with high per-file concurrency. Four tests
depend on state that they share with their siblings, or on a timer that
outlives them. That was safe at 8 workers. It is not safe at 96 or more.
Runs 32547184159 and 32551746525 show them.
1. Every pytest subprocess shared one temp root.
pytest puts tmp_path under <temproot>/pytest-of-<user>/. At the end of a
session it walks that directory with cleanup_dead_symlinks(). The walk lists
the directory. Then it asks whether the `pytest-current` symlink resolves.
Then it unlinks the symlink. A second process replaces that symlink between
the question and the unlink. The first process then raises FileNotFoundError
after all of its tests passed. Two files failed this way and passed on retry.
scripts/run_tests_parallel.py now gives each subprocess its own temp root
through PYTEST_DEBUG_TEMPROOT, and deletes it after the attempt. No two
processes share a directory. The race has no shared object to act on.
Proof: a direct driver of _pytest.pathlib.cleanup_dead_symlinks against one
root, with a second thread that replaces the symlink, raises the same
FileNotFoundError on 'pytest-current' as CI. A private root for each
subprocess removes that condition. A separate check confirms that 5
subprocesses receive 5 distinct roots, that tmp_path lands inside the private
root, and that no root survives the attempt.
2. The config read guard walked directories that other tests were writing.
tests/hermes_cli/test_config_read_guard.py scanned the tree with rglob. rglob
descends into every directory and filters after that, so it calls scandir() on
__pycache__ trees that the guard never inspects. Sibling processes create and
delete those entries during the run. A directory that disappears in the middle
of a walk raises FileNotFoundError out of rglob.
The scan now uses os.walk. It prunes excluded directories before it descends,
and it ignores a directory that disappears. __pycache__ joins the excluded
set, because bytecode is not source.
The guard still catches what it exists to catch. With a planted raw
yaml.safe_load of config.yaml in hermes_cli/, the test fails and names the
planted file. With a clean tree it passes.
3. A PTY test waited for a file to exist, and not for its content.
tests/tools/test_process_registry_write_stdin_surrogates.py spawns a child
that runs open(out,'wb').write(sys.stdin.buffer.readline()). open() creates
the file empty. The bytes arrive only after the PTY delivers the line. The
wait stopped at out.exists(), which the empty file already satisfies, so the
read returned b'' when the parent won that gap. This test failed both attempts
in CI, and did not pass on retry.
The test now waits for the expected bytes, with a bounded deadline.
Proof: the old wait loses 6 times in 25 runs on an idle 16-core machine. The
new wait loses 0 times in 25.
4. A dialog close timer outlived the test that started it.
ConfirmDialog holds the "done" beat for 600ms after a successful confirm, then
calls onClose. The timer had no cleanup, so an unmount inside that window left
it armed. It then called onClose on a tree that is gone, which reaches
setState in the parent. vitest can tear the environment down first, and React
then reads `window` during the update:
ReferenceError: window is not defined
at resolveUpdatePriority (react-dom-client.development.js:1308)
at dispatchSetState
at Timeout.t4 [as _onTimeout] session-actions-menu.tsx:574
The frame at session-actions-menu.tsx:574 is the `onClose` prop of
DeleteSessionDialog. The owner of the timer is ConfirmDialog, which now keeps
the handle in a ref and clears it on unmount.
Zoomable had the same fault, with a 1500ms timer that clears a "copied" flag.
copy-button.tsx and tooltip.tsx already clear their timers.
Proof: a new test confirms, unmounts inside the 600ms window, then advances
the clock. Against the old code it fails with "expected onClose to not be
called at all, but actually been called 1 times". Against the new code it
passes.
Verification:
- The affected Python files and the tests of the runner itself pass under
scripts/run_tests.sh.
- The desktop ui suite passes: 566 files, 5382 tests, and no
"window is not defined".
- eslint reports 0 errors on apps/desktop. The 118 warnings are the state
before this change. The two cleanup effects carry an eslint-disable line for
the ref-mirror rule. They write a timer handle, and not a mirror of a
reactive value. The rule permits this, and its own comment names the case.
- The PTY test cannot run on the NixOS development machine. That machine has
no python3 outside the nix store, and the test uses the literal `python3`.
The child exits 127 there. The fix rests on the 25-run measurement above and
on CI.
Follow-ups to the salvaged WSL-bridge gating (#66447):
- wsl-path-bridge.ts: discard wsl.exe stderr so the 'WSL is not installed'
banner can never leak into an attached console on WSL-less machines (#80184).
- scripts/desktop-update/windows.ps1: add an explorer.exe-mediated detached
relaunch rung between the WMI attempt and the tethered Start-Process
fallback. When Win32_Process.Create fails (observed ReturnValue 8), the
Desktop no longer re-attaches to the hand-off console, so its stdout stops
flooding the window and the console can close.
* fix(update): bound the Windows update hand-off's step pipe drain
Invoke-HermesStep collected each step's output with ReadToEndAsync().Result.
That task does not complete when the step exits; it completes when the pipe
reaches EOF. On Windows the write end of a redirected pipe goes to the child as
an inheritable handle, so every descendant spawned without its own redirection
holds a duplicate and EOF waits for the last of them to close it. hermes update
deliberately runs its build steps with stdout inherited, so the tree under a
step is arbitrarily deep and not something this script can enumerate. When one
of those descendants is a resident gateway, the pipe stays open for the life of
the gateway and the hand-off blocks forever.
Everything the hand-off owes the Desktop is downstream of that call:
.hermes-update-result.json is never written, .hermes-update-in-progress is never
cleared, and the Desktop is never relaunched. The app sits on "Updating Hermes"
until the user kills the gateway by hand, and the stale marker then refuses the
next update too.
Read both pipes in chunks into a StringBuilder and bound the drain once the step
process itself has exited. The bound cannot truncate a slow step: the clock only
starts after the process is gone, at which point everything it wrote is already
in the pipe buffer waiting to be read, so the grace only has to cover the final
drain. Chunked reads are what make abandoning safe at all, since .Result cannot
hand back a partial read.
Also switch to the bounded WaitForExit overload. The argument-less one waits on
redirected streams as well, which is the same unbounded wait by another name.
An abandoned drain logs one line to logs/desktop-update-handoff.log naming the
cause, so a truncated step log is never mistaken for a step that printed
nothing.
Measured on Windows 11 / PowerShell 5.1 against a step whose grandchild
inherits its stdout and outlives it by 45s: 47.4s before, 4.3s after, with the
step's exit code and output preserved in both.
Fixes#90455
* test(update): prove the hand-off survives a step that leaks its pipe
Four source-level guards on Invoke-HermesStep, scoped to that function so the
legitimate WaitForExit and .Result uses elsewhere in the script cannot mask a
regression: no ReadToEndAsync, a drain bound keyed on the step having exited,
no argument-less WaitForExit, and a log line when a drain is abandoned. All
four fail against the previous drain. They are source-level for the same
reason the sibling python-handoff guard is: Linux CI cannot execute the
PowerShell hand-off.
Source-level is not enough for a deadlock, though, so the script also grows a
-SelfTestPipeDrain fixture alongside the existing -SelfTestUi one. It needs no
checkout, no install and no update: it starts a step that spawns a grandchild
with UseShellExecute = $false and no redirection, which is exactly the shape
that makes the grandchild inherit the step's stdout and stderr, then exits 7
while the grandchild sleeps on. The fixture asserts the grandchild was still
alive when Invoke-HermesStep returned, so a pass cannot be a timing
coincidence, and that the exit code and the step's output both survived the
abandonment. A windows_only test drives it, so the OS lane runs the real
drain rather than a text match.
Measured on Windows 11 / PowerShell 5.1: 4.3s with the fix, 47.4s (the
grandchild's full lifetime) with the previous drain restored.
The python-handoff guard now reads the script with its -SelfTest* blocks
removed. Those blocks exercise the machinery deliberately and exit before any
marker, venv or desktop work, so the "every step drives python.exe, never the
hermes.exe shim" rule does not apply to them. Scoping the source that way
rather than allow-listing a target keeps that rule absolute for every real
step.
Refs #90455
* fix(update): don't meter the step drain that #90455's bound introduced
Chunked reads make the bounded drain possible, but the loop idled 150ms
after every chunk it consumed, so a step's output moved at one 16 KiB
buffer per tick (~107 KB/s). The pipe then backs up, which is
backpressure on the *running* step rather than a slow read: a chatty
step blocks on write() waiting for the reader.
`hermes update` is exactly that shape -- the Electron/vite build alone
is megabytes -- so the layer that fixed "the hand-off waits forever"
would have shipped "the hand-off is slow" in its place.
Idle only when both pipes came up empty, and idle on the reads
themselves (WaitAny with the same 150ms cap) rather than on the clock:
a freshly issued ReadAsync is rarely complete by the very next pass, so
a bare `if (-not $moved)` still sleeps between chunks. WaitAny expires
on its own, so a silent step keeps the marquee animating and keeps the
abandon deadline advancing.
Measured against the drain as submitted, same harness, one variable:
4 MiB of step stdout 38.99s -> 0.07s
1 MiB stdout + 1 MiB err 18.22s -> 0.27s
leaked grandchild (20s) 3.24s -> 3.20s, exit code + output kept
quiet step, exits at 4s 4.29s -> 4.04s, 29 passes (not spinning)
* test(update): make the pipe-drain fixture cover metering, not just deadlock
The fixture proved the drain returns while a descendant holds the pipe.
It could not have caught the opposite failure -- a drain slow enough to
backpressure the step it is reading -- and that is the regression the
first version of this fix shipped.
Add a flood arm: a step that writes megabytes and holds nothing, with a
wall-clock budget far under what a sleep-per-chunk drain needs. The two
arms bracket the contract from both sides: bounded when a descendant
holds the pipe open, never slower than the step can write.
Few large lines rather than many small ones, deliberately --
Write-HandoffLog is one Add-Content per line and runs inside the
measured window, so line-heavy output would time the logger.
Also drops the four source-grep guards. Reading windows.ps1's text to
assert it contains `$abandonAt` tests the shape of the source, not its
behavior: it passes on a drain that is wired wrong but spelled right,
fails on a correct refactor, and blocks the extraction it should
survive. AGENTS.md bans the pattern outright, and all four pass on the
metered drain. The executable arms cover the same contract and actually
run the code -- the Windows lane is where this is verified either way.
---------
Co-authored-by: Jack Lau <72348727+jackulau@users.noreply.github.com>
Nothing in CI compiled this crate. `.rs` lives under `apps/`, so the
change classifier matched a Rust edit as `frontend` and ran the
TypeScript matrix, which cannot notice a Rust error — the crate's 58 unit
tests had never executed once, and neither would the pipe-drain tests in
the previous commit.
Adds a `rust` lane and a Linux `cargo test --lib` job. Linux on purpose:
the pipe-drain fixtures need a real process tree whose grandchild
inherits the parent's stdout and are `#[cfg(unix)]`, so a Windows runner
would compile them out and report green over zero coverage. The Windows
half of that contract is `-SelfTestPipeDrain` on the existing Windows
lane.
The desktop updater ran `hermes update --yes`, which auto-restored any
uncommitted source-tree edits onto the freshly updated checkout. On dirty
from-source installs this silently carried local modifications across every
update and could break the rebuilt app (field report: Windows update handoff
leaving the app 'crashed').
New `hermes update --keep-stash`: local changes are still autostashed so the
update can proceed, but are never re-applied — they stay parked in git stash
with printed recovery guidance. Both desktop handoff scripts (windows.ps1,
posix.sh) now pass it, probing `update --help` first so older installed
backends without the flag keep working. Failure paths are unchanged (stash
preserved, no restore); updates.non_interactive_local_changes: discard still
wins.
Tests: park/restore/failure-path coverage incl. a sabotage-verified
regression test; docs updated.
The self-test grew a branch that spawns a Python child so pytest could prove
progress advances during one. It doesn't need to: /progress is answered from
its own runspace, so the existing hold already blocks the main thread, and
the spawn only exercised Invoke-HermesStep, which nothing here changes.
The Windows test now asserts the invariant instead of the self-test's stage
string, and the posix half -- previously untested, and the half that broke --
gets real coverage: serve-ui.py's wire shape, and posix.sh driven end to end
with a stub `hermes` that reports which stage was on screen while it ran.
The shim is a shared page, but only windows.ps1 publishes a stage and an
elapsed count. On mac and Linux posix.sh publishes `running` with an empty
message and no clock, so the running branch rendered the h2 back into the
muted line ("Updating Hermes" twice) and started a clock in the browser,
losing "Hermes will open once done." on both platforms.
A clock started in the page measures when the window painted, not how long
the update has been running -- on posix that is the only clock there is, and
it reads zero after the desktop-exit wait has already burned 30s. That is the
hardcoded-milestone problem #75895 removed, in a new costume.
So: elapsed comes from the orchestrator or is not shown. serve-ui.py stamps
it per request from the hand-off start (a value written into the status file
would freeze between publishes, which are minutes apart -- exactly the stall
the line exists to disprove), matching what Windows' in-process listener
already does. posix.sh gets the stages it was missing, at the four gates it
genuinely waits on. Absent a stage the page keeps the settled copy, and an
old orchestrator that sends no clock simply shows no clock.
MOVEFILE_DELAY_UNTIL_REBOOT was the quarantine's last resort, and it is worse
than doing nothing. It writes to HKLM, so a non-elevated update — every
Desktop-driven one, and most terminal ones — gets ERROR_ACCESS_DENIED and
reports nothing. When it does succeed it frees nothing for the install
running right now, and the queued operation outlives that update: at the next
boot it moves aside whatever sits at the shim path, including a shim a later
repair just wrote.
Drops the fallback and sweeps entries older versions queued, matching only
our own <shim> -> <shim>.old.<stamp> pairs so unrelated installers keep
theirs.
Salvaged from #88121 by @fangliquanflq.
llms.txt coverage is asserted in Python, but website/ sat on the Python skip
list, so a PR adding a docs page — or regressing the generator — went green
without ever running the test that checks the page is reachable. That is how
the index drifted to 53% coverage unnoticed.
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.
The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
Appending --daemonized after ORIGINAL_ARGS put it past the `--`
relaunch-args separator on Linux, so it was absorbed into
RELAUNCH_ARGS instead of being parsed as a flag. HANDOFF_DAEMONIZED
never got set, so the one-shot self-detach block re-fired on every
re-exec -- an unbounded self-exec loop (thousands of iterations/sec,
100%+ CPU, argv growing until execve fails with E2BIG) whenever
relaunch args were present, which is the normal invocation shape on
Linux.
Fixes#86957
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Three problems with install.sh --commit:
1. No validation: non-hex or too-short arguments passed through to git,
producing misleading errors.
2. Fetch failure swallowed by || true: abbreviated SHAs are refused by
GitHub's server ("couldn't find remote ref"), but the error was
silently ignored.
3. Checkout failure not checked: git checkout --detach with a missing
object produces a misleading "does not take a path argument" error
and the install continues unpinned, exiting 0.
Fix:
- Validate --commit is a 7-40 hex string up front
- Remove || true from fetch; fail with actionable message directing
users to full 40-char SHAs
- Check git checkout --detach result and fail hard on error
Fixes#87268
Both install-blocking npm install call sites (browser tools and TUI) ran
with --silent and no output capture, so failures printed only a generic
error message with no npm diagnostics.
Apply the same pattern used by the camofox install path: redirect npm
output to a temp file and replay it on failure, so users can see the
actual error (EBADENGINE, ETARGET, network timeout, registry 5xx, etc.).
Fixes#87340
Install-Uv piped the astral installer's entire output to Out-Null, so any
real failure (proxy block, AV quarantine, permissions) surfaced only as the
generic "uv installed but not found" message, and astral.sh was the sole
install source even though corporate proxies commonly block it while the
byte-identical GitHub releases installer downloads fine.
Three-rung ladder, all inside Install-Uv:
1. astral.sh installer with output captured via Tee-Object.
2. GitHub releases installer mirror (same UV_INSTALL_DIR).
3. Salvage an existing uv.exe (Get-Command uv, or the astral default
%USERPROFILE%\.local\bin\uv.exe) by copying it into $HermesHome\bin so
the managed-first invariant holds.
On total failure, print the last 15 lines of captured installer output plus
the existing manual-install pointer.
Reported by @BitBernd; proxy diagnosis by @gakugaku; Out-Null suppression
first identified by @webtecnica in #69366.
Closes#69216
Review finding on PR #83194 (egilewski): Install-Venv committed the venv
transaction as soon as the replacement had a working interpreter, deleting
the parked previous venv. Install-Dependencies is a separate later stage
(a separate process under the stage-per-process bootstrap) and every
dependency tier or the baseline-import gate can still fail after that
point - a failed update could still leave Hermes and the blocker probe
unusable with no rollback source.
Now:
- Install-Venv records the parked backup in venv.pending-backup instead
of deleting it, and excludes it from the venv.stale.* sweep.
- Install-Dependencies wraps the dependency tiers + baseline-import gate
in the transaction: Restore-VenvBackup on failure (parks the failed
replacement as venv.failed.*, renames the previous venv back), and
Complete-VenvTransaction only after the imports prove the replacement
usable.
- Source-contract regression tests for the boundary
(tests/test_install_ps1_venv_transaction_boundary.py).
When Rename-Item on the live venv is denied, do not fall back to an
in-place Remove-Item that can gut site-packages and leave no rollback.
Also mark venv-blocker probe failures with probe_failed so they cannot
be read as a clear scan (#83149).
The relative exclude-newer = "14 days" cutoff bricks installs whenever the
resolver cannot see (or accept) a package's upload date:
- defusedxml / python-olm / unpaddedbase64 (#80387, #79434): ancient frozen
releases (2021-2023) whose upload dates are often absent from mirror
indexes and stale uv HTTP caches. uv then filters them entirely
("there are no versions of defusedxml"), breaking [youtube]/[wecom]/
[matrix] resolution and daily `uv sync --locked` runs.
- setuptools / pillow / mcp (#78227, #75992, #76020): exact-pinned deps.
When the pinned version's upload date is invisible, the resolver filters
the ONLY acceptable candidate — setuptools==83.0.0 in
[build-system].requires meant the project could not even be built from a
git checkout on released v0.20.0. Exempting an exact pin costs nothing:
the version cannot float without a reviewed pin bump.
Changes:
- pyproject.toml: add all six to the existing exclude-newer-package
whitelist, with rationale comments per class.
- uv.lock: regenerated; diff is the whitelist metadata only (verified
zero version drift, still 249 packages).
- tests/test_packaging_metadata.py: new standing guard
test_build_system_requires_exempt_from_exclude_newer — every
[build-system].requires package must be whitelisted while a relative
exclude-newer cutoff is configured. Verified both directions (fails
when setuptools is removed from the whitelist).
- scripts/install.sh: fix the stale tier-name comparison ("all (with
RL/matrix extras)" vs actual "all") that mislabeled every successful
Tier-1 install as a fallback-tier install (#79434 bonus finding).
Verification: uv lock --check green on uv 0.11.19 and 0.12.5;
uv sync --extra all --locked green; uv pip install -e '.[all]' resolves;
whitelist mechanism A/B-proven on a minimal project (unsatisfiable ->
resolves; build-requires variant: uv build fails -> succeeds).
Reported-by: MichaelClawHub (#80387), liujianqiu (#79434), maxonliu (#78227)
The Desktop-spawned hand-off consistently died during Electron's quit
teardown on macOS: the orchestrator process group was terminated right
after `running: hermes update ...`, so no exit code, result file, bundle
swap, or relaunch ever happened, and the loopback shim window surfaced
the death as ERR_CONNECTION_REFUSED or "Aw, Snap!" error code 15
(reproductions in #66753).
- Re-exec the orchestrator through a one-shot setsid child and let the
direct Electron child exit immediately; the real orchestrator is owned
by launchd (PPID 1), outside Electron's teardown, same marker/result
protocol.
- Hold TERM ignored across the `hermes update` invocation and
log-and-ignore the single teardown TERM that can still arrive after
the desktop PID dies (durable SIGNAL breadcrumb for diagnosis).
- Delay start_ui until the desktop PID is gone plus 1s so the shim
server/window are never born inside the teardown window.
- Run both UI processes in their own sessions; keep SIGTERM/SIGHUP
ignored in the shim server and stop it with SIGKILL, so a stray TERM
can no longer leave the progress window on a dead loopback URL while
the update continues.
Verified on a production git install (macOS arm64, Darwin 27.0,
v0.20.1): six consecutive Desktop-triggered/production-shape updates
completed end-to-end including a full desktop rebuild + codesign; the
shim survived a deliberately injected TERM+HUP mid-update and a full
`hermes desktop --force-build --build-only` running alongside it.
Fixes the macOS reproductions in #66753.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adopts two improvements from #81586 (kshitijk4poor), with one correction:
- Test-ManagedNodeInUse now queries Win32_Process (ExecutablePath +
CommandLine substring) instead of Get-Process .Path: a cmd.exe wrapper
running npm.cmd from the tree reports its own exe in System32, so the
tree path shows up only in the command line. Win32_Process.CommandLine
works on Windows PowerShell 5.1 and 7+, unlike the Get-Process
.CommandLine ETS property (7.4+ only); a single CIM query also beats a
per-process property access loop. This closes the gap flagged by the
triage bot on #81500 (Update-ManagedNpm's only in-use protection is
this pre-check).
- Test-Node stages the extracted tree to a sibling node.new-* before the
swap, so the final swap is a same-volume rename (atomic) instead of a
cross-volume Move-Item (copy+delete, non-atomic).
The mtime-touch ordering from #81586 was NOT adopted: touching the backup
only after the swap succeeds leaves it at its old (long-lived-tree) mtime
for the whole swap window, which the age-gated litter sweep (st_mtime <
cutoff) removes — reopening the concurrent-sweep race the touch exists to
close. The touch stays immediately after the backup rename.
Co-authored-by: kshitijk4poor <82637225+kshitijk4poor@users.noreply.github.com>
The Hermes-managed Node tree at %HERMES_HOME%\node is destructively
rewritten while the desktop app's Node processes execute from it:
the Node-26 heal did shutil.rmtree + move, the EBADENGINE repair ran
npm install --global --prefix into the tree, and install.ps1's
Test-Node did Remove-Item + Move-Item. Windows rejects those writes
with PermissionError: [WinError 5] on npm.cmd.
- _heal_managed_node_windows: stage the fully-downloaded tree in a
sibling node.new-* dir, then rename-swap (live tree -> node.old-*,
staged -> node). The live tree is never deleted before its
replacement is ready, so an interrupted heal cannot gut it; a
refused rename is the OS-level in-use signal and defers (returns
None) instead of forcing the write.
- heal_hermes_managed_node: an in-use deferral does not record the
once-per-process attempt, so the heal retries once the tree is free.
- managed_node_tree_in_use: cheap psutil pre-check (Windows only) that
avoids pointless 30-50MB re-downloads in long-lived processes.
- upgrade_managed_npm: defer the in-place npm self-upgrade while the
tree is in use, with a notice.
- install.ps1: Test-ManagedNodeInUse guard around Update-ManagedNpm and
the Test-Node install branch, which now rename-swaps instead of
delete-then-move.
An in-use-but-outdated tree keeps serving the old runnable Node (old
Node beats no Node), and every npm resolution re-evaluates the heal, so
the upgrade applies automatically on the next update with the app
closed.
Everything Browser Use is now managed by Hermes: the canonical binary
is the one install_cli() provisions into HERMES_HOME/bin, and every
resolution and provisioning site prefers it.
- _find_cli(): probe order flipped to managed bin -> PATH ->
user-level tool dir (then uvx across the same order). A user's own
uv tool install can no longer shadow the Hermes-managed copy with a
drifted version; side installs only matter when we have nothing.
- install_cli(): a browser-use on PATH no longer short-circuits the
install — only the managed copy does, so selecting any backend
provisions the copy Hermes controls and updates.
- _ensure_browser_use_cli() (hermes tools): drops its own PATH check
and always delegates to install_cli(), the single owner of the
managed-copy policy.
- install.sh / install.ps1: same short-circuit fix — only
HERMES_HOME/bin/browser-use counts as installed.
Follow-up to #86240 and #86320: with every non-Camofox backend
selection installing the CLI, managed-first closes the remaining
version-drift/shadowing class instead of guarding single sites.
Tests updated to pin the new contract: managed beats PATH and
user-local; PATH install does not satisfy install_cli; helper always
delegates. E2E-verified precedence chain with real files and a real
degraded-PATH install attempt.
`uv pip install -e .` has to replace the console-script shims, so
_quarantine_running_hermes_exe must first rename the running hermes.exe out
of the way. That rename fails whenever any child process spawned from that
hermes.exe is still alive: on Windows a child inherits a handle on the parent
image. It is the inherited handle, not the trampoline, that pins the file --
killing the child makes the identical rename succeed, and the shim flavour
(uv trampoline vs distlib launcher) makes no difference.
The updater spawns such children itself (npx cache warm, memory-provider
refresh -- hindsight-api runs as a daemon with --idle-timeout 300 and outlives
the step that started it), so this presents as a race rather than a hard
failure: the same hand-off succeeds on one run and dies on the next. Step 2's
shim-unlock preflight cannot catch it, because the shim genuinely is unlocked
at that moment; the pinning child appears later, during the update.
When the rename loses that race, _schedule_replace_on_reboot is the last
resort -- and MOVEFILE_DELAY_UNTIL_REBOOT writes to HKLM, so it needs
elevation. A Desktop-driven update is not elevated, so it returns
ERROR_ACCESS_DENIED, `uv pip install -e .` exits 2, and the ZIP fallback
repeats the identical sequence. The desktop build stage is then never reached
while the pre-build clean has already removed apps/desktop/release, leaving an
install whose Start Menu shortcut points at a Hermes.exe that no longer exists.
Running the same code as `python.exe -m hermes_cli.main update` puts the
inherited handles on python.exe, which uv never has to replace.
posix.sh is deliberately untouched: unlinking a running executable is legal
there, so the equivalent call is harmless.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(install): fail when Node dependencies cannot install (#85297)
The POSIX installer converted root and TUI npm failures into warnings, then
printed a dependency-success message and reached the installation-complete
banner with a zero exit status. This left consumers with no usable
node_modules while reporting success.
Treat both required npm installs as fatal: log an error, restore tracked
lockfile churn, return status 1, and propagate the failure from the monolithic
and node-deps stage callers. Successful installs, Termux and missing-Node
skips, missing-manifest skips, and optional Playwright/Browser Use/Computer
Use best-effort behavior remain unchanged. The fix is limited to the POSIX
installer; the PowerShell installer is outside this issue's scope.
Focused and adjacent installer tests passed (32), with bash syntax,
py_compile, and diff checks clean. The broader installer family had 90 passes,
one unrelated pre-existing failure, and two skips; the full suite was
environment-limited by missing dependencies. CodeRabbit, iterative deep
security/compatibility reviews, and final confidence security/compatibility
reviews were clean against the final diff.
Fixes#85297
* fix(install): require npm alongside node in check_node (#77003)
A stray `node` symlink without a sibling `npm` (leftover from a node
version manager) made check_node report "Node.js found"; every later
npm install then failed and the desktop build died with an opaque
"Node.js / npm unavailable". Node now only counts as found when npm
resolves on the same PATH, with an explicit "stray node symlink?" branch
that falls through to the Hermes-managed Node (which bundles npm).
The overlapping success-log honesty half of the original PR is subsumed
by the previous commit, which makes a failed npm install fatal rather
than conditionally-logged; the behavioral tests there cover it, so this
commit keeps only the check_node PATH-gate assertions.
Fixes#77003.
Co-authored-by: criptogus <criptogus@users.noreply.github.com>
---------
Co-authored-by: Eugeniusz Gilewski <egilewski@egilewski.com>
Co-authored-by: CriptoGus <128640021+criptogus@users.noreply.github.com>
Co-authored-by: criptogus <criptogus@users.noreply.github.com>
* fix(install): time-box the Windows node-deps stage so a stalled npm or Playwright install can't hang setup forever
scripts/install.sh has bounded this same work with run_with_timeout
"$NODE_DEPS_TIMEOUT" (600s default) since #39219, but install.ps1 never got
the guard: Install-NodeDeps ran both `npm install` and `npx playwright
install chromium` unbounded. A stalled registry fetch or a wedged Chromium
archive extraction (#76222, #84614) froze the installer indefinitely -- one
user left it running 12+ hours overnight before asking for help.
Route both invocations through _Invoke-NativeWithTimeout: cmd.exe launches
the native command with its output merged to a log, the parent polls with a
wall-clock deadline and tails new log lines to the console each tick (the
live progress that makes a 3-minute download distinguishable from a hang),
and on timeout taskkill /T /F kills the real process tree and returns 124 --
the same convention as coreutils timeout and bash's run_with_timeout.
Wait-Job was rejected for this: jobs swallow live output and Stop-Job leaves
the npm child running. Windows PowerShell 5.1-safe throughout.
Timeouts surface as a warning with the log path, a note that re-running the
installer resumes (stages are idempotent), and the NODE_DEPS_TIMEOUT env
override for slow links -- mirroring bash.
Fixes#76222.
Closes#84614.
Supersedes #76303.
Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>
* fix(installer): roll stage timers over to hours so an overnight stall doesn't read as "744 hours"
formatElapsed rendered a running stage as m:ss with unbounded minutes: a
node-deps stage left hanging overnight showed "744:38", which the user who
reported the hang understandably read as 744 hours. formatDuration
(completed stages) had the same unbounded-minutes shape.
Move both formatters into src/lib/format.ts (pure, no React) and add the
hour rollover: h:mm:ss live, "Xh Ym" completed. tests-js pins the shapes,
including 744m38s -> 12:24:38.
---------
Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>
Choosing Computer Use should be a config flip, not a hunt for
'hermes computer-use install'. Three provisioning rungs:
- install.sh / install.ps1 pre-install cua-driver (best-effort,
non-fatal, time-boxed at 660s above the upstream installer's 600s
lock window; --skip-computer-use / -SkipComputerUse to opt out;
Termux and unwritable-/Applications skipped cleanly)
- PUT /api/tools/toolsets/{name} (dashboard + desktop toggle) spawns
the background 'hermes tools post-setup cua_driver' action when the
toolset is enabled while the binary is missing — previously the
toggle 'saved' but the tool never appeared in the schema because
check_computer_use_requirements() couldn't find the binary
- hermes tools interactive flow already installed via
_toolset_needs_configuration_prompt/_POST_SETUP_INSTALLED (unchanged)
Docs: computer-use.md enabling section rewritten around the new flow;
installation.md documents --skip-computer-use.
ensure_browser() (install.sh) and Install-AgentBrowser (install.ps1)
are reached only via the explicit --ensure browser / -Ensure browser
on-demand mode, itself only triggered by an actual browser-tool call's
lazy-install fallback or `hermes acp --setup-browser`. agent-browser
already resolves via npx in that same fallback before ever reaching
these scripts, so eagerly npm-installing a second, separately
version-pinned copy here was redundant and an extra credential/
supply-chain surface for a path npx already covers. Chromium
acquisition for this on-demand path is now deferred entirely to
_maybe_autoinstall_chromium's existing lazy fallback. camofox's
install and system-browser detection/configuration are unaffected.
install.ps1 also drops the now-dead -SkipChromium switch, confirmed
unused at its one call site.
* fix(windows): SSH ControlMaster gating + stop hijacking the user's python
Two Windows environment-integrity fixes:
1. tools/environments/ssh.py (#73927): Windows OpenSSH has no
Unix-domain-socket ControlMaster support, so unconditionally passing
ControlPath/ControlMaster/ControlPersist failed EVERY tool call on a
Windows-hosted ssh terminal backend with 'getsockname failed: Not a
socket'. Gate the three multiplexing options behind a module-level
_SSH_MULTIPLEX = (os.name != 'nt'); the scp upload path is gated the
same way. On Windows the backend now works without connection pooling
(each command a fresh connection); POSIX behavior is unchanged. The
teardown 'ssh -O exit' is naturally inert because the socket never
exists on Windows.
2. scripts/install.ps1 (#83797): the installer put the whole
venv\Scripts directory on the user PATH, which contains python.exe /
pythonw.exe / pip.exe and so silently hijacked the 'python' command in
every terminal on the machine — unrelated projects started resolving
python to Hermes' runtime interpreter. Now copy only the launchers
(hermes.exe, hermes-acp.exe) into a dedicated $InstallDir\bin and put
THAT on PATH. Existing installs are migrated: the legacy venv\Scripts
entry is stripped from the user PATH on the next install/update. The
new bin dir is under $InstallDir (…\hermes-agent), which the uninstall
PATH sweep already matches via its \hermes-agent marker.
Updated the stale hermes_cli/update_cmd.py docstring that described the
old venv\Scripts-on-PATH layout.
Tests: SSH ControlMaster gating pinned both directions (multiplex on →
flags present; off → absent but BatchMode/StrictHostKeyChecking retained).
install.ps1 parses clean via the PowerShell AST parser.
* docs: update windows-native install docs for the bin\ launcher layout
CI (test_windows_native_docs) pins the docs and installer to the same
PATH layout. The #83797 fix moved the PATH entry from venv\Scripts to a
dedicated $InstallDir\bin holding only the hermes launchers, so update
the Windows-native guide to match: PATH-after-install section, the
install-steps list, the directory-layout table, the Get-Command
verification line, and the 'command not found' pitfall. Test now asserts
the bin\ layout and guards against a regression back to venv\Scripts on
PATH.
* fix: keep install.ps1 pure ASCII (PowerShell 5.1 codepage safety)
The two comments I added in the #83797 PATH-hijack fix used em-dashes,
tripping tests/test_install_ps1_ascii_only.py — Windows PowerShell 5.1
reads a BOM-less .ps1 in the system ANSI codepage (not UTF-8), so a
non-ASCII byte can misdecode into a stray quote and desync the parser
(issues #66994/#67000). Replace the em-dashes with ASCII '--'.
The Browser Use CLI became the default browser backend, but nothing
provisioned it: users without uv/uvx (field report from DongyangHe on
macOS) silently fell back to the built-in browser tools with no notice.
- install_cli() in tools/browser_use_cli.py: uv tool install browser-use
via the managed uv (bootstrapped on demand), linked into
$HERMES_HOME/bin (UV_TOOL_BIN_DIR)
- _find_cli() now also probes $HERMES_HOME/bin for browser-use/uvx —
Hermes' managed uv is not on the user's PATH
- hermes tools post_setup actually installs (Camofox standard) instead
of printing instructions
- install.sh / install.ps1 provision the CLI at install time
(best-effort, non-fatal, honors --skip-browser)
- CLI startup shows a one-line notice (24h rate-limited) when the
default backend downgraded to the built-in tools
Round 4 of helix4u's review — the durable fallback is now real:
- Result protocol gains `manual`: an ok result the user still must act
on (reopen the app, reinstall the GUI package, fix the sandbox helper).
Both orchestrators set it on every DONE_NOTE/downgrade path; the Desktop
consumer surfaces manual results in a real dialog on next boot instead
of a log line — the browserless-Linux disappearance now ends at a
visible dialog, worst case one boot later. Older result files without
the field parse as manual:false (covered).
- notify ladder verifies EXECUTION, not existence: zenity/kdialog must
survive their first second (an instant death means no display and falls
through); the no-surface case is an explicit best-effort contract whose
guaranteed channel is the result dialog.
- mac DONE_NOTE + failed relaunch of the kept/rolled-back bundle is no
longer swallowed (`|| true` dropped): the durable message carries both
facts.
- launch/gate matrices assert `manual` in the result JSON; consumer
round-trip tested in handoff-result.test.ts.