Commit Graph

1887 Commits

Author SHA1 Message Date
Teknium
7a1aafb4e1 fix(update): raise the step-idle watchdog default to 10 minutes
Desktop builds legitimately produce no output for minutes on Windows
(the updater terminal is static until a step completes), so 300s risks
cancelling healthy long steps. 600s keeps the watchdog meaningful for
true stalls while clearing slow builds; HERMES_UPDATE_STEP_IDLE_SECONDS
still overrides. Per Teknium's review.
2026-08-26 18:46:34 -07:00
Teknium
266d8ce2bb fix(update): normalize windows.ps1 to LF and keep fixture here-string braces off column 0
The cherry-pick landed the file with CRLF endings and a fixture
here-string whose col-0 brace prematurely terminated the handoff test's
SelfTest-block strip, tripping the drive-python-not-the-shim guard on
fixture code. LF restored (matching main), child-script loop inlined.
2026-08-26 18:23:33 -07:00
Teknium
50d3c53f6c fix(update): count logs/update.log growth as watchdog progress
The #95625 watchdog cancels a step after StepIdleTimeoutSeconds (300s)
with no stdout/stderr. But a real `hermes update` is stdout-silent for
40+ minutes by design: the Electron/vite build streams to
logs/update.log, not the child's pipes (hermes_cli/update_cmd.py's
update-log tee). An output-only ceiling would therefore kill every
healthy large update at 5 minutes and mark it exit 124.

The drain now fingerprints logs/update.log (size + mtime) and, when the
idle ceiling is otherwise reached, treats growth of that file as
progress: reset the clock instead of terminating the tree. The stat
runs only once the ceiling fires, so the hot drain path never touches
the filesystem. HERMES_UPDATE_STEP_IDLE_SECONDS remains the override;
HERMES_UPDATE_PROGRESS_LOG points the self-test at its own file.

TDD proof: -SelfTestPipeDrain gains a fourth arm, logstall -- a step
that is silent on its pipes but appends to the progress log every
second and must reach its natural exit 3, never 124. Linux CI pins the
same contract at source level (TestIdleWatchdogCountsUpdateLogGrowth);
sabotage-verified: making the log-growth consult inert fails
test_stall_branch_consults_log_growth_before_terminating.
2026-08-26 18:23:33 -07:00
fangliquanflq
91c475bf5f fix(update): assign Windows steps before execution 2026-08-26 18:23:33 -07:00
fangliquanflq
dc91b6b524 fix(update): quiesce stalled Windows updater trees 2026-08-26 18:23:33 -07:00
fangliquanflq
80be4890f9 fix(update): recover stalled Windows desktop handoffs 2026-08-26 18:23:33 -07:00
fred0m
53057f2bc4 fix(update): run config migration on the 'Already up to date' repair path (#91360)
A failed update attempt can pull fresh code onto disk and then die before
the config-migration block (e.g. a PyPI timeout during the dependency
sync). The desktop hand-off retries; the retry takes the commit_count == 0
branch, repairs deps, prints 'Already up to date!' and returns early -
skipping _run_config_check_fresh / migrate_config entirely. The fresh
code (requiring a newer _config_version) then refuses to start against
the old config until 'hermes doctor --fix' is run.

Fix: _maybe_migrate_config_on_current() mirrors the version_bump_only
handling (silent, non-interactive) and is called on both repair-path
completion points before claiming success.

Also: scripts/desktop-update/posix.sh no longer retries when the update
was deliberately SKIPPED (checkout parked on a non-target branch) -, the
retry is deterministic and only wastes time. Uses a dedicated non-
colliding exit code (8) and an honest message instead of 'Update failed'.

New tests: tests/hermes_cli/test_update_config_migration_on_current.py
(5 cases: migrate-when-behind, noop-current, noop-ahead, warning re-
surface, silent check failure).
2026-08-26 17:42:50 -07:00
alt-glitch
a667ab60d9 fix(livetest): render multi-query bridge calls 2026-08-25 16:39:12 -07:00
alt-glitch
e455e4afd0 feat(tool-search): multi-query search, batched describe, Snowball stemming
tool_search now takes queries: string[] (searched independently against
the same catalog, limit applies per query, default 5 / max 25) and
returns the split shape: per-query groups carry tool names only, one
shared tools map holds each matched tool's source, description (400-char
cap) and required parameter names once. When some queries miss, a single
top-level available_sources + hint block replaces the old per-response
fallback.

tool_describe now takes names: string[] and returns a map keyed by name;
unknown names collect in not_found (with the refresh hint) and
non-deferrable names keep their per-name spelling-check error in errors,
so one bad name no longer fails the whole call. Duplicates dedupe
silently.

The shared tokenizer now applies Snowball stemming (english, exact-pinned
snowballstemmer) at both index and query time, closing the measured
plural/singular miss where 'issues' failed to return create_issue. The
inline BM25 is unchanged. Stemmer instances are thread-local (they carry
mutable parse state and bridge dispatch can run on parallel tool-call
threads).

New config knobs under tools.tool_search: max_queries / max_describe_names
(default 10 each, floor 1, no upper clamp) bound the per-call array
inputs; over-cap calls error so the model repairs in one round-trip.

No backward compatibility with the single query/name shapes, by decision.
scripts/analyze_livetest.py renders both shapes since transcripts on disk
may predate this change.
2026-08-25 16:39:12 -07:00
lorzl
9525c0e5b7 fix(desktop-update): use the system default browser for the update shim
The Windows update hand-off shim was hardcoded to Microsoft Edge
(Find-EdgeExe), so machines whose default browser is Chrome still got
an Edge --app progress window, and every run leaked a throwaway
browser profile (hermes-update-ui-<pid>) under %TEMP% that was never
removed.

- Get-DefaultBrowserExe replaces Find-EdgeExe: resolves the OS default
  browser from the UserChoice ProgId (https first, http fallback).
  ChromeHTML -> Chrome, MSEdgeHTM -> Edge; any other ProgId returns
  $null and degrades to the existing WinForms card.
- The dedicated --user-data-dir profile is now removed when the shim
  closes, and stale hermes-update-ui-* leftovers from interrupted
  runs are swept from %TEMP% in the same pass.
- --app + --user-data-dir is Chromium-only, so the whitelist is
  intentionally limited to chrome/msedge; Edge keeps
  --disable-features=msImplicitSignin to suppress the implicit MSA
  sign-in that leaks into shim windows (#88410).
2026-08-24 04:01:49 -07:00
Teknium
da57f49237 fix(desktop-update): shim window only opens in the user's own browser family
A Safari/Firefox/Helium user who merely had Chrome installed watched Chrome open on every desktop update (community report). start_ui now checks the system default browser (LaunchServices https handler on macOS, xdg-settings on Linux) and skips the shim window unless the default is Chromium-family; notify_fallback and the durable result file still carry the outcome. Detection is best-effort: any failure keeps the old behavior.
2026-08-24 03:14:24 -07:00
jeremyrandria-debug
60bb2bb719 fix(update): auto-close desktop-update shim window after error/manual outcomes
On error/manual outcomes stop_ui('leave-window') kept the browser shim
window open indefinitely, so an aborted update left a Chrome window on
screen until the user closed it by hand; repeated update attempts piled
up more windows.

stop_ui now always closes the shim. leave-window paths keep it up for a
short grace period (HERMES_UPDATE_SHIM_GRACE_SECONDS, default 15) so a
watching user can read the message, then close it. The success path is
unchanged. The error/manual outcome is durably written to
.hermes-update-result.json and surfaced in a dialog on the next Desktop
boot, so closing the shim loses no information.
2026-08-24 03:14:24 -07:00
liuhao1024
eb21740b06 Also skip Brave: its P3A bar paints over the update shim window
Brave renders its own P3A privacy-notice bar ("Got it" / "Disable" /
"Learn more") at the top of the throwaway-profile window the posix shim
opens, cramped to unreadability at the shim's small size - the same
window-pollution class as Edge's MSA sync notice, and equally immune to
the throwaway --user-data-dir (#88682). Drop Brave from the candidates
on both platforms; Chrome and Chromium stay.

Covers #88682 on top of #88410
2026-08-24 03:14:24 -07:00
liuhao1024
f329f9e40d fix(desktop-update): never render the posix update shim in Edge
Edge's OS-level Microsoft-account integration signs even a fresh
throwaway profile into the user's MSA and renders its own "syncing
your browsing data" notification — the user's MSA email included —
inside the update window, which is titled "Hermes" (#88410). The
throwaway --user-data-dir start_ui already passes cannot block that
OS-account path, so the only reliable protection is to not pick Edge
at all: drop it from the browser candidates on both macOS and Linux.
The update UI is a best-effort layer — with no other Chromium-family
browser installed, start_ui falls back to its existing
"no renderer; skipping UI" path and the update itself is unaffected.

Fixes #88410
2026-08-24 03:14:24 -07:00
Teknium
2e75862790 fix(install.ps1): record 'skipped-long-path' when ConvertTo-LongPath short-circuits
The ordinary-long-path early return now records why no resolver ran, so
the ResolvedPathReport stays truthful instead of silently inheriting a
stale value from an earlier call in the same session. Diagnostics hunk
taken from PR #93100.

Co-authored-by: aniruddhaadak80 <aniruddhaadak80@users.noreply.github.com>
2026-08-23 18:27:21 -07:00
liuhao1024
203f111c98 fix(install.ps1): initialize LastResolver before the resolved-path report
ConvertTo-LongPath short-circuits for ordinary long paths (no ~\d alias),
so $script:LastResolver is only assigned when a short path actually needs
expansion. The ResolvedPathReport block read it unconditionally, which is
fatal under Set-StrictMode before any install stage runs (#93017: fresh
installs died at line 367 through three different invocation styles).

Initialize it to 'none' — the resolver's own value for "nothing ran" — at
script scope before Set-LongProfileEnvVars can invoke a resolver.
2026-08-23 18:27:21 -07:00
Teknium
a6bec08f35 fix(install): actually invoke check_cxx_compiler in both install stages
Salvage follow-up for #88993: the preflight was defined but never
called from the prerequisites stage or the full-install path, so the
Fedora node-gyp failure (#93063) would still occur. Wire it in after
check_node in both sequences.
2026-08-23 18:25:42 -07:00
osgeek90
60cbc70430 installer: check for a C++ compiler before building native Node modules
npm install inside install_node_deps() builds native addons (e.g. node-pty) via node-gyp, which needs a C/C++ compiler. That was never checked, so a missing g++ only surfaced as a generic "npm install failed or timed out" deep inside npm's own output — and because install_node_deps failing short-circuits the rest of main() via `|| return`, users end up with no `hermes` command and no clue why.
2026-08-23 18:25:42 -07:00
emozilla
fe95ed3930 Merge origin/main: reconcile with PR #92092 (in-checkout launcher restore)
PR #92092 fixed the same vanished-launcher bug by restoring copies into
the legacy in-checkout hermes-agent\bin from the update tail. That
location is what this branch removes: untracked files there are swept
by the update autostash on every cycle (restore/sweep treadmill, plus a
parked stash entry per update under --keep-stash), and unconditional
exe copies break on relocatable venvs ('uv trampoline failed to
canonicalize script path'). This branch's managed-binary-dir layout
supersedes both mechanisms, so the merge resolves to it:

- drop _sync_windows_cli_launchers and its _ensure_acp_launcher call
  (Windows staging/repair lives in ensure_windows_bin_launchers at
  process start and migrate_windows_bin_path in the update tail);
  _ensure_acp_launcher is a Windows no-op again
- keep #92092's genuinely better installer semantics: staging stays in
  a dedicated Install-HermesCommandLaunchers function that throws
  BEFORE any PATH mutation when the required launcher cannot be staged
  and verified -- previously Set-PathVariable could put an empty dir on
  PATH and still print 'hermes command ready'. Reworked for this
  branch's layout: caller passes the destination ($HermesHome\bin),
  launcher form follows the venv (exe copy vs .cmd delegator), and the
  verify step accepts either form
- rework #92092's AST-lifted PowerShell test for the new function
  signature, keeping its fail-before-PATH-mutation assertions and
  adding relocatable-venv form-selection coverage
- drop tests/hermes_cli/test_windows_cli_launcher_repair.py (pinned the
  superseded in-checkout mechanism; equivalent and broader coverage
  lives in tests/hermes_cli/test_ensure_windows_bin_launchers.py)
2026-08-22 22:16:38 -04:00
emozilla
679e9cd294 fix(windows): stage hermes launchers in the managed binary dir, not the git checkout
The installer staged the hermes/hermes-acp launcher copies at
hermes-agent\bin -- inside the git working tree -- and put that dir on
the user PATH (#84452). The update command's pre-pull autostash
(git stash push --include-untracked) swept those untracked, unignored
copies off disk, and once the desktop updater stopped re-applying
stashes (--keep-stash, 5dd221d442) nothing restored them: `hermes`
stopped resolving in every new terminal on every desktop-updated
install.

Move the canonical launcher home to the managed binary dir
(%LOCALAPPDATA%\hermes\bin, next to the managed uv) -- outside the
checkout, where no git operation can ever touch it. The dir is
per-machine and shared by every profile, so all anchoring uses
get_default_hermes_root(), never HERMES_HOME (which points inside
profiles\<name> under `hermes -p`).

The copy design also had a second latent break: managed-uv rebuilds
create relocatable venvs, and a relocatable venv's exe trampoline
resolves relative to its own location -- a copy outside venv\Scripts
dies with 'uv trampoline failed to canonicalize script path'. Launcher
form now depends on the venv (lockstep in install.ps1 and
_install_repair.py): exe copy for normal venvs, a .cmd delegator
invoking the in-venv exe by absolute path for relocatable ones. Either
form counts as present, so pre-rebuild exe copies are left alone.

Delivery to the existing fleet, per cohort:

- already-broken installs cannot run the CLI, so an import-time heal in
  hermes_cli.main (ensure_windows_bin_launchers) re-stages missing
  launchers when the desktop app spawns its backend -- the one channel
  that still reaches them. Gates fail toward inaction: canonical dir
  only for the managed clone, legacy hermes-agent\bin only while the
  user PATH still resolves through it (some pre-managed-uv installs
  have no hermes\bin PATH entry; the legacy re-stage is what fixes
  those). Staging-name + os.replace keeps concurrent process starts
  from tearing a launcher; the helper never raises.
- healthy old-layout installs migrate in the update tail
  (migrate_windows_bin_path): stage canonical launchers, verify them
  BEFORE touching the registry, prepend hermes\bin to the user PATH,
  strip the legacy entries (hermes-agent\bin and venv\Scripts, #83797),
  preserving REG_EXPAND_SZ and raw %VARS%. The legacy dir's files stay
  on purpose -- configs holding absolute launcher paths keep working;
  only the sweepable PATH resolution route goes.
- fresh installs get the new layout from install.ps1 directly.

/bin/ is gitignored so the one update that DELIVERS this fix cannot
sweep pre-migration launchers a final time under the old rules; the
gitignore line, the legacy re-stage branch, and the update-tail call
are transition machinery with a named expiry once the fleet has
migrated.

Also rewrites _ensure_acp_launcher's stale Windows paragraph to match
(raw docstring fixes its invalid \S escape) and updates the Windows
native docs to the new layout, with a docs<->installer parity test.
2026-08-22 13:38:56 -04:00
Gille
a08b909199 fix(windows): preserve launcher layout invariants 2026-08-22 00:05:51 -07:00
Gille
9782275b2a fix(windows): restore dedicated CLI launchers on update 2026-08-22 00:05:51 -07:00
ethernet
969094e4d2 fix(tests): remove four shared-state and lifetime faults at high concurrency
The suite now runs as one job with high per-file concurrency. Four tests
depend on state that they share with their siblings, or on a timer that
outlives them. That was safe at 8 workers. It is not safe at 96 or more.
Runs 32547184159 and 32551746525 show them.

1. Every pytest subprocess shared one temp root.

pytest puts tmp_path under <temproot>/pytest-of-<user>/. At the end of a
session it walks that directory with cleanup_dead_symlinks(). The walk lists
the directory. Then it asks whether the `pytest-current` symlink resolves.
Then it unlinks the symlink. A second process replaces that symlink between
the question and the unlink. The first process then raises FileNotFoundError
after all of its tests passed. Two files failed this way and passed on retry.

scripts/run_tests_parallel.py now gives each subprocess its own temp root
through PYTEST_DEBUG_TEMPROOT, and deletes it after the attempt. No two
processes share a directory. The race has no shared object to act on.

Proof: a direct driver of _pytest.pathlib.cleanup_dead_symlinks against one
root, with a second thread that replaces the symlink, raises the same
FileNotFoundError on 'pytest-current' as CI. A private root for each
subprocess removes that condition. A separate check confirms that 5
subprocesses receive 5 distinct roots, that tmp_path lands inside the private
root, and that no root survives the attempt.

2. The config read guard walked directories that other tests were writing.

tests/hermes_cli/test_config_read_guard.py scanned the tree with rglob. rglob
descends into every directory and filters after that, so it calls scandir() on
__pycache__ trees that the guard never inspects. Sibling processes create and
delete those entries during the run. A directory that disappears in the middle
of a walk raises FileNotFoundError out of rglob.

The scan now uses os.walk. It prunes excluded directories before it descends,
and it ignores a directory that disappears. __pycache__ joins the excluded
set, because bytecode is not source.

The guard still catches what it exists to catch. With a planted raw
yaml.safe_load of config.yaml in hermes_cli/, the test fails and names the
planted file. With a clean tree it passes.

3. A PTY test waited for a file to exist, and not for its content.

tests/tools/test_process_registry_write_stdin_surrogates.py spawns a child
that runs open(out,'wb').write(sys.stdin.buffer.readline()). open() creates
the file empty. The bytes arrive only after the PTY delivers the line. The
wait stopped at out.exists(), which the empty file already satisfies, so the
read returned b'' when the parent won that gap. This test failed both attempts
in CI, and did not pass on retry.

The test now waits for the expected bytes, with a bounded deadline.

Proof: the old wait loses 6 times in 25 runs on an idle 16-core machine. The
new wait loses 0 times in 25.

4. A dialog close timer outlived the test that started it.

ConfirmDialog holds the "done" beat for 600ms after a successful confirm, then
calls onClose. The timer had no cleanup, so an unmount inside that window left
it armed. It then called onClose on a tree that is gone, which reaches
setState in the parent. vitest can tear the environment down first, and React
then reads `window` during the update:

    ReferenceError: window is not defined
     at resolveUpdatePriority (react-dom-client.development.js:1308)
     at dispatchSetState
     at Timeout.t4 [as _onTimeout] session-actions-menu.tsx:574

The frame at session-actions-menu.tsx:574 is the `onClose` prop of
DeleteSessionDialog. The owner of the timer is ConfirmDialog, which now keeps
the handle in a ref and clears it on unmount.

Zoomable had the same fault, with a 1500ms timer that clears a "copied" flag.
copy-button.tsx and tooltip.tsx already clear their timers.

Proof: a new test confirms, unmounts inside the 600ms window, then advances
the clock. Against the old code it fails with "expected onClose to not be
called at all, but actually been called 1 times". Against the new code it
passes.

Verification:
- The affected Python files and the tests of the runner itself pass under
  scripts/run_tests.sh.
- The desktop ui suite passes: 566 files, 5382 tests, and no
  "window is not defined".
- eslint reports 0 errors on apps/desktop. The 118 warnings are the state
  before this change. The two cleanup effects carry an eslint-disable line for
  the ref-mirror rule. They write a timer handle, and not a mirror of a
  reactive value. The rule permits this, and its own comment names the case.
- The PTY test cannot run on the NixOS development machine. That machine has
  no python3 outside the nix store, and the test uses the literal `python3`.
  The child exits 127 there. The fix rests on the 25-run measurement above and
  on CI.
2026-08-22 02:25:12 -04:00
Teknium
0a8cdec697 fix(desktop): silence wsl.exe stderr banner + detached explorer relaunch rung
Follow-ups to the salvaged WSL-bridge gating (#66447):
- wsl-path-bridge.ts: discard wsl.exe stderr so the 'WSL is not installed'
  banner can never leak into an attached console on WSL-less machines (#80184).
- scripts/desktop-update/windows.ps1: add an explorer.exe-mediated detached
  relaunch rung between the WMI attempt and the tethered Start-Process
  fallback. When Win32_Process.Create fails (observed ReturnValue 8), the
  Desktop no longer re-attaches to the hand-off console, so its stdout stops
  flooding the window and the console can close.
2026-08-21 00:49:41 -07:00
fangliquanflq
ee6a9f8326 fix(updater): carry acquisition age through scripts 2026-08-20 12:52:21 -05:00
brooklyn!
5e32e3aecd Merge pull request #90941 from NousResearch/bb/installer-drain-bound
fix(installer): the bootstrap installer stops waiting on pipe EOF
2026-08-20 12:50:43 -05:00
brooklyn!
2b1bff624e fix(update): Windows Desktop updates finish instead of parking on "Updating Hermes" (#90937)
* fix(update): bound the Windows update hand-off's step pipe drain

Invoke-HermesStep collected each step's output with ReadToEndAsync().Result.
That task does not complete when the step exits; it completes when the pipe
reaches EOF. On Windows the write end of a redirected pipe goes to the child as
an inheritable handle, so every descendant spawned without its own redirection
holds a duplicate and EOF waits for the last of them to close it. hermes update
deliberately runs its build steps with stdout inherited, so the tree under a
step is arbitrarily deep and not something this script can enumerate. When one
of those descendants is a resident gateway, the pipe stays open for the life of
the gateway and the hand-off blocks forever.

Everything the hand-off owes the Desktop is downstream of that call:
.hermes-update-result.json is never written, .hermes-update-in-progress is never
cleared, and the Desktop is never relaunched. The app sits on "Updating Hermes"
until the user kills the gateway by hand, and the stale marker then refuses the
next update too.

Read both pipes in chunks into a StringBuilder and bound the drain once the step
process itself has exited. The bound cannot truncate a slow step: the clock only
starts after the process is gone, at which point everything it wrote is already
in the pipe buffer waiting to be read, so the grace only has to cover the final
drain. Chunked reads are what make abandoning safe at all, since .Result cannot
hand back a partial read.

Also switch to the bounded WaitForExit overload. The argument-less one waits on
redirected streams as well, which is the same unbounded wait by another name.

An abandoned drain logs one line to logs/desktop-update-handoff.log naming the
cause, so a truncated step log is never mistaken for a step that printed
nothing.

Measured on Windows 11 / PowerShell 5.1 against a step whose grandchild
inherits its stdout and outlives it by 45s: 47.4s before, 4.3s after, with the
step's exit code and output preserved in both.

Fixes #90455

* test(update): prove the hand-off survives a step that leaks its pipe

Four source-level guards on Invoke-HermesStep, scoped to that function so the
legitimate WaitForExit and .Result uses elsewhere in the script cannot mask a
regression: no ReadToEndAsync, a drain bound keyed on the step having exited,
no argument-less WaitForExit, and a log line when a drain is abandoned. All
four fail against the previous drain. They are source-level for the same
reason the sibling python-handoff guard is: Linux CI cannot execute the
PowerShell hand-off.

Source-level is not enough for a deadlock, though, so the script also grows a
-SelfTestPipeDrain fixture alongside the existing -SelfTestUi one. It needs no
checkout, no install and no update: it starts a step that spawns a grandchild
with UseShellExecute = $false and no redirection, which is exactly the shape
that makes the grandchild inherit the step's stdout and stderr, then exits 7
while the grandchild sleeps on. The fixture asserts the grandchild was still
alive when Invoke-HermesStep returned, so a pass cannot be a timing
coincidence, and that the exit code and the step's output both survived the
abandonment. A windows_only test drives it, so the OS lane runs the real
drain rather than a text match.

Measured on Windows 11 / PowerShell 5.1: 4.3s with the fix, 47.4s (the
grandchild's full lifetime) with the previous drain restored.

The python-handoff guard now reads the script with its -SelfTest* blocks
removed. Those blocks exercise the machinery deliberately and exit before any
marker, venv or desktop work, so the "every step drives python.exe, never the
hermes.exe shim" rule does not apply to them. Scoping the source that way
rather than allow-listing a target keeps that rule absolute for every real
step.

Refs #90455

* fix(update): don't meter the step drain that #90455's bound introduced

Chunked reads make the bounded drain possible, but the loop idled 150ms
after every chunk it consumed, so a step's output moved at one 16 KiB
buffer per tick (~107 KB/s). The pipe then backs up, which is
backpressure on the *running* step rather than a slow read: a chatty
step blocks on write() waiting for the reader.

`hermes update` is exactly that shape -- the Electron/vite build alone
is megabytes -- so the layer that fixed "the hand-off waits forever"
would have shipped "the hand-off is slow" in its place.

Idle only when both pipes came up empty, and idle on the reads
themselves (WaitAny with the same 150ms cap) rather than on the clock:
a freshly issued ReadAsync is rarely complete by the very next pass, so
a bare `if (-not $moved)` still sleeps between chunks. WaitAny expires
on its own, so a silent step keeps the marquee animating and keeps the
abandon deadline advancing.

Measured against the drain as submitted, same harness, one variable:

  4 MiB of step stdout      38.99s -> 0.07s
  1 MiB stdout + 1 MiB err  18.22s -> 0.27s
  leaked grandchild (20s)    3.24s -> 3.20s, exit code + output kept
  quiet step, exits at 4s    4.29s -> 4.04s, 29 passes (not spinning)

* test(update): make the pipe-drain fixture cover metering, not just deadlock

The fixture proved the drain returns while a descendant holds the pipe.
It could not have caught the opposite failure -- a drain slow enough to
backpressure the step it is reading -- and that is the regression the
first version of this fix shipped.

Add a flood arm: a step that writes megabytes and holds nothing, with a
wall-clock budget far under what a sleep-per-chunk drain needs. The two
arms bracket the contract from both sides: bounded when a descendant
holds the pipe open, never slower than the step can write.

Few large lines rather than many small ones, deliberately --
Write-HandoffLog is one Add-Content per line and runs inside the
measured window, so line-heavy output would time the logger.

Also drops the four source-grep guards. Reading windows.ps1's text to
assert it contains `$abandonAt` tests the shape of the source, not its
behavior: it passes on a drain that is wired wrong but spelled right,
fails on a correct refactor, and blocks the extraction it should
survive. AGENTS.md bans the pattern outright, and all four pass on the
metered drain. The executable arms cover the same contract and actually
run the code -- the Windows lane is where this is verified either way.

---------

Co-authored-by: Jack Lau <72348727+jackulau@users.noreply.github.com>
2026-08-20 12:41:17 -05:00
Brooklyn Nicholson
5caea5e501 ci: run cargo test for the bootstrap installer
Nothing in CI compiled this crate. `.rs` lives under `apps/`, so the
change classifier matched a Rust edit as `frontend` and ran the
TypeScript matrix, which cannot notice a Rust error — the crate's 58 unit
tests had never executed once, and neither would the pipe-drain tests in
the previous commit.

Adds a `rust` lane and a Linux `cargo test --lib` job. Linux on purpose:
the pipe-drain fixtures need a real process tree whose grandchild
inherits the parent's stdout and are `#[cfg(unix)]`, so a Windows runner
would compile them out and report green over zero coverage. The Windows
half of that contract is `-SelfTestPipeDrain` on the existing Windows
lane.
2026-08-20 11:40:49 -05:00
Teknium
5dd221d442 feat: desktop updates no longer re-apply local source edits (--keep-stash)
The desktop updater ran `hermes update --yes`, which auto-restored any
uncommitted source-tree edits onto the freshly updated checkout. On dirty
from-source installs this silently carried local modifications across every
update and could break the rebuilt app (field report: Windows update handoff
leaving the app 'crashed').

New `hermes update --keep-stash`: local changes are still autostashed so the
update can proceed, but are never re-applied — they stay parked in git stash
with printed recovery guidance. Both desktop handoff scripts (windows.ps1,
posix.sh) now pass it, probing `update --help` first so older installed
backends without the flag keep working. Failure paths are unchanged (stash
preserved, no restore); updates.non_interactive_local_changes: discard still
wins.

Tests: park/restore/failure-path coverage incl. a sabotage-verified
regression test; docs updated.
2026-08-20 04:54:14 -07:00
Jeffrey Quesnelle
612b3633d2 Merge pull request #77915 from bbednarski9/feat/relay-native-plugin-init
feat(relay)!: initialize static/dynamic plugins via native integration, remove opt-in plugin
2026-08-19 22:11:13 -04:00
brooklyn!
6284afbaab Merge pull request #90194 from NousResearch/review/88767
fix(update): show the current stage and elapsed time while updating
2026-08-19 13:31:12 -05:00
Brooklyn Nicholson
e602225d86 test(update): cover the progress contract without a test hook in the updater
The self-test grew a branch that spawns a Python child so pytest could prove
progress advances during one. It doesn't need to: /progress is answered from
its own runspace, so the existing hold already blocks the main thread, and
the spawn only exercised Invoke-HermesStep, which nothing here changes.

The Windows test now asserts the invariant instead of the self-test's stage
string, and the posix half -- previously untested, and the half that broke --
gets real coverage: serve-ui.py's wire shape, and posix.sh driven end to end
with a stub `hermes` that reports which stage was on screen while it ran.
2026-08-19 13:16:13 -05:00
Brooklyn Nicholson
cf2ab8522a fix(update): one elapsed clock, served to the shim by both orchestrators
The shim is a shared page, but only windows.ps1 publishes a stage and an
elapsed count. On mac and Linux posix.sh publishes `running` with an empty
message and no clock, so the running branch rendered the h2 back into the
muted line ("Updating Hermes" twice) and started a clock in the browser,
losing "Hermes will open once done." on both platforms.

A clock started in the page measures when the window painted, not how long
the update has been running -- on posix that is the only clock there is, and
it reads zero after the desktop-exit wait has already burned 30s. That is the
hardcoded-milestone problem #75895 removed, in a new costume.

So: elapsed comes from the orchestrator or is not shown. serve-ui.py stamps
it per request from the hand-off start (a value written into the status file
would freeze between publishes, which are minutes apart -- exactly the stall
the line exists to disprove), matching what Windows' in-process listener
already does. posix.sh gets the stages it was missing, at the four gates it
genuinely waits on. Absent a stage the page keeps the settled copy, and an
old orchestrator that sends no clock simply shows no clock.
2026-08-19 13:16:13 -05:00
Brooklyn Nicholson
a9eb99d172 fix(update): stop deferring shim renames to next boot
MOVEFILE_DELAY_UNTIL_REBOOT was the quarantine's last resort, and it is worse
than doing nothing. It writes to HKLM, so a non-elevated update — every
Desktop-driven one, and most terminal ones — gets ERROR_ACCESS_DENIED and
reports nothing. When it does succeed it frees nothing for the install
running right now, and the queued operation outlives that update: at the next
boot it moves aside whatever sits at the shim path, including a shim a later
repair just wrote.

Drops the fallback and sweeps entries older versions queued, matching only
our own <shim> -> <shim>.old.<stamp> pairs so unrelated installers keep
theirs.

Salvaged from #88121 by @fangliquanflq.
2026-08-19 13:14:40 -05:00
Gille
107531549a fix(update): show live Desktop update progress 2026-08-19 13:11:27 -05:00
Brooklyn Nicholson
3a034356a2 ci: run the Python lane for docs and website script changes
llms.txt coverage is asserted in Python, but website/ sat on the Python skip
list, so a PR adding a docs page — or regressing the generator — went green
without ever running the test that checks the page is reachable. That is how
the index drifted to 53% coverage unnoticed.
2026-08-19 12:41:45 -05:00
Bryan Bednarski
8afd98ef2a refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski
0a079b946f fix(relay): retain legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:02 -07:00
Bryan Bednarski
e8644e05a3 refactor(relay): remove legacy observability plugin
Signed-off-by: Bryan Bednarski <bbednarski@nvidia.com>
2026-08-19 08:52:01 -07:00
rob-maron
0b588cb3a4 MCP CIMD auth 2026-08-18 20:03:18 -07:00
ethernet
00c3872882 feat(ci): add nix flake check as unrequired job
The workflow owns its triggers and ci.yml does not call it. A
reusable-workflow call holds the caller run in progress for the full
build, and GitHub refuses `gh run rerun` on a run that is still in
progress. A separate run reruns and cancels on its own.

The job restores /nix/store from the GitHub Actions cache and saves from
main only. A cache that a PR writes is visible to that PR alone, so a
save there spends the quota of the repository and helps no later run.
2026-08-18 20:42:06 -04:00
ethernet
1dbe469276 refactor(ci): hoist docker detect-changes into the .py file
The docker.yml gate held its own copy of the build formula, in shell.
classify_changes.py now owns a derived docker lane, and the nix lane in
the next commit derives from the same file. Two formulas in two
languages drift apart, and one Python function with tests does not.
2026-08-18 20:42:06 -04:00
f-trycua
5b010f448f fix(computer-use): verify Windows driver repair 2026-08-16 11:34:40 -07:00
Francesco Bonacci
81af2ef013 fix(computer-use): reconcile existing cua-driver installs 2026-08-16 11:34:40 -07:00
Yasushi Fukutake
b44f956dff fix(desktop-update): put --daemonized ahead of ORIGINAL_ARGS in posix.sh re-exec
Appending --daemonized after ORIGINAL_ARGS put it past the `--`
relaunch-args separator on Linux, so it was absorbed into
RELAUNCH_ARGS instead of being parsed as a flag. HANDOFF_DAEMONIZED
never got set, so the one-shot self-detach block re-fired on every
re-exec -- an unbounded self-exec loop (thousands of iterations/sec,
100%+ CPU, argv growing until execve fails with E2BIG) whenever
relaunch args were present, which is the normal invocation shape on
Linux.

Fixes #86957

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-16 06:25:21 -07:00
RelaxJonh
eac1f65340 fix(install): validate --commit SHA and fail hard on fetch/checkout errors (#87268)
Three problems with install.sh --commit:

1. No validation: non-hex or too-short arguments passed through to git,
   producing misleading errors.

2. Fetch failure swallowed by || true: abbreviated SHAs are refused by
   GitHub's server ("couldn't find remote ref"), but the error was
   silently ignored.

3. Checkout failure not checked: git checkout --detach with a missing
   object produces a misleading "does not take a path argument" error
   and the install continues unpinned, exiting 0.

Fix:
- Validate --commit is a 7-40 hex string up front
- Remove || true from fetch; fail with actionable message directing
  users to full 40-char SHAs
- Check git checkout --detach result and fail hard on error

Fixes #87268
2026-08-16 01:45:10 -07:00
RelaxJonh
f887819421 fix(install): capture npm output on failure for diagnosable errors (#87340)
Both install-blocking npm install call sites (browser tools and TUI) ran
with --silent and no output capture, so failures printed only a generic
error message with no npm diagnostics.

Apply the same pattern used by the camofox install path: redirect npm
output to a temp file and replay it on failure, so users can see the
actual error (EBADENGINE, ETARGET, network timeout, registry 5xx, etc.).

Fixes #87340
2026-08-16 01:44:45 -07:00
Teknium
3af56c2203 fix(install.ps1): surface uv installer errors and add GitHub + existing-uv fallbacks (#69216)
Install-Uv piped the astral installer's entire output to Out-Null, so any
real failure (proxy block, AV quarantine, permissions) surfaced only as the
generic "uv installed but not found" message, and astral.sh was the sole
install source even though corporate proxies commonly block it while the
byte-identical GitHub releases installer downloads fine.

Three-rung ladder, all inside Install-Uv:
1. astral.sh installer with output captured via Tee-Object.
2. GitHub releases installer mirror (same UV_INSTALL_DIR).
3. Salvage an existing uv.exe (Get-Command uv, or the astral default
   %USERPROFILE%\.local\bin\uv.exe) by copying it into $HermesHome\bin so
   the managed-first invariant holds.

On total failure, print the last 15 lines of captured installer output plus
the existing manual-install pointer.

Reported by @BitBernd; proxy diagnosis by @gakugaku; Out-Null suppression
first identified by @webtecnica in #69366.

Closes #69216
2026-08-14 22:33:58 -07:00
konsisumer
4b0c1031db fix(desktop-update): wait for rebuilt executable before relaunch 2026-08-14 22:23:00 -07:00
Teknium
aa5a960675 fix(installer): hold the venv rollback source through dependency validation (#83149)
Review finding on PR #83194 (egilewski): Install-Venv committed the venv
transaction as soon as the replacement had a working interpreter, deleting
the parked previous venv. Install-Dependencies is a separate later stage
(a separate process under the stage-per-process bootstrap) and every
dependency tier or the baseline-import gate can still fail after that
point - a failed update could still leave Hermes and the blocker probe
unusable with no rollback source.

Now:
- Install-Venv records the parked backup in venv.pending-backup instead
  of deleting it, and excludes it from the venv.stale.* sweep.
- Install-Dependencies wraps the dependency tiers + baseline-import gate
  in the transaction: Restore-VenvBackup on failure (parks the failed
  replacement as venv.failed.*, renames the previous venv back), and
  Complete-VenvTransaction only after the imports prove the replacement
  usable.
- Source-contract regression tests for the boundary
  (tests/test_install_ps1_venv_transaction_boundary.py).
2026-08-14 21:58:09 -07:00