A console-less parent (pythonw backend, detached daemons) allocates a visible
console per bare spawn, so Windows startup flashes ~7-8 console windows
(#117781) and the GPU statusbar flashes two more every 5s (#101895). Threads
creationflags=windows_hide_flags() through every hermes_cli spawn site the
desktop reaches: gitlock (tasklist probe + git plumbing), the scheduled-task
PowerShell probe, the updater's _git_run/_no_prompt_git_kwargs/_GIT_TEXT_KW
runners, the safe-directory config probe, and consolidates the statusbar's two
nvidia-smi reads per poll into one cached hidden query. The statusbar also
stops polling while the document is hidden and resumes on visibilitychange
(#120262).
Supersedes #117805, #117833, #102047, #120285 (thanks @JoaoMarcos44, @Finn763,
@monerostar, @szzhoujiarui).
The pre_update_version co_varnames check added in af26acab73 rejected the
current updater's entrypoints, but it also rejected historical updaters
from 2026-07-26..08-16, whose _cmd_update_impl/cmd_update never declared
that local — their in-process dependency installs would have continued on
the old interpreter instead of handing off to the takeover child.
Invert the gate: the CURRENT updater declares the
_hermes_current_updater_frame sentinel local in both entrypoints; no
historical on-disk version can contain it. in_historical_update() now
hands off on entrypoint match unless the sentinel is present, restoring
the handoff for every historical window while still refusing current
frames. Verified live for all four frame classes; mutation check: with
the sentinel check removed, the current-updater refusal test fails.
git 2.53+ `index-pack --promisor` hands every object outside the promisor
packs to pack-objects, which BUG()s (should_include_obj) on the missing
objects those objects lead to. A partial clone with unmarked packs therefore
keeps crashing its fetches.
The retry added for this ran `git -c remote.origin.promisor= fetch`. That
override does nothing: git registers promisor remotes additively, so the
repo's own `true` still wins (GIT_TRACE on 2.50 and 2.55 shows
`index-pack --promisor` and `filter tree:0` under the override). The retry
only passed when the crashed fetch had already stored the pack; a retry that
has to transfer again crashed again.
- On the crash, give every pack without a `.promisor` marker one, then retry
the same fetch once. The clone stays partial, no refetch, no config change.
If git still crashes, the updater points at the docs section.
- The two fetches that turn a full clone partial now mark the clone's old
packs afterwards: the depth-1 `tree:0` unshallow in
`fetch_full_commit_graph`, and `worktree_ops._deepen_shallow_repo`
(`--filter=blob:none`). git writes the partial-clone config before it
fetches, so this also runs when the fetch fails. Without it, a converted
depth-1 install crashed every later fetch on git 2.55 (3 of 3 rounds in a
local repro).
The Desktop-initiator fact called read_live_update(), whose pid liveness probe
(os.kill) and stale-marker unlink are update-lock policy. Running them from a
metrics helper tripped the retired-holder-gate tripwire in
test_update_venv_holder_retirement and could delete a marker as a side effect.
We already hold the lock, so a marker naming another pid is the orchestrator.
The update pipeline already writes one machine-readable receipt per run, so the
run/stage counters are derived from it in the process that finalizes it instead
of instrumenting stages for metrics. The receipt gains minimal stage END marks
(plan, snapshot, apply[mode], deps, build, restart; verify is inferred from the
fleet matrix), an initiator fact (desktop when a live orchestrator claim
names another pid) and the pre-update commit date, which is all the derivation
needs for outcome, failed stage, duration and from-version-age buckets.
A pre-pull interpreter must never import pulled code: when it is the finalizer
(same pid, checkout sha moved) it parks the receipt (stdlib only) and the next
Hermes start records it. Collection off => nothing imported beyond a config read.
The completion-process test stub gains the new record_stage API the real
module now exposes (no assertion changed).
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
A restore that conflicts (or whose untracked baseline is unknown, or whose
untracked files the update replaced) parks the user's patch and the update
still printed 'Update complete!' and exited 0 (#122557). The run's stash now
stays unsettled until it is restored, discarded, or parked because the user
asked (--keep-stash, declined prompt); an unsettled stash becomes the
completion line, which is not a success line, so the update ends partial,
exits 1 and the receipt says partial. Failure verdicts name it too.
Follow-up trim of the #125964 salvage:
- Armer falls back to _current_checkout_sha() (the identity the reader compares the fleet
against) instead of a new receipt helper; at arm time the receipt has no post_update yet.
- Reader: the SHA-less, inventory-less record follows the existing inventory-less rule
(live fleet all current on the checkout) with no extra pid-liveness gate, matching the
invariant the SHA-bearing inventory-less record already has. The gatewayless branch stays
fail-closed for an SHA-less record (checkout_contains("") is not evidence).
- Drop the obligation_fields pid plumbing and update_receipt helper; keep 2 tests:
discharge on live-fleet evidence, and the armer recording the checkout SHA.
A no-op `hermes update` (pre == post SHA) whose head capture failed armed the
host update-restart record with `expected_sha: ""` and no inventory. The reader
bailed out on the empty SHA before the live-fleet reconciliation, so the
"gateways may still be serving pre-update modules" warning could never clear
even when every live gateway was current on the checkout.
Reader (hermes_cli/update_cmd_fleet.py):
- An inventoried obligation without its SHA stays fail-closed.
- An inventory-less, SHA-less record now falls through to the existing
live-fleet check once the process that armed it is gone (pid liveness via
pid_exists_stdlib), so a status call cannot discharge a restart phase that
is still running.
Armer (hermes_cli/update_cmd.py):
- When the completion request carries no SHA, fall back to the receipt's
post_update.sha (then pre_update.sha) instead of arming with "".
obligation_fields() now reports the arming pid alongside expected_sha.
Tests: discharge on all-current fleet, kept when the armer pid is alive, kept
on a stale fleet, kept with an inventory but no SHA; the discharge test fails
on the pre-fix reader with the issue's exact warning.
Review follow-up: compute the parked-branch case once in the rollback,
fold the baseline comments into one, reuse pull() and a new complete()
helper in the tests, and say in the docs that a rollback restores the
commit the checkout ran before the pull on every update path.
Review follow-ups for the branch-switch rollback:
- If the parked branch cannot be checked out again (for example, another
worktree holds it), restore its commit detached. Before this, the install
stayed on the broken update branch, and the recovery hint repeated the
failing command.
- Any `merge-base --is-ancestor` result other than "contained" keeps the
strict baseline. rc 128 used to fall back to the lenient one, which
contradicted its own comment.
- `_pull_updates` returns the baseline it verified, and
`_apply_pulled_update` checks against it. Before, that function
recomputed the baseline with a second rev-parse and merge-base against a
hardcoded `origin/<branch>`, which a fork push in between could move.
- Drop a `commit_count == -1` check that could never be false.
- Update the rollback description in `updating.md`.
- Tests share the fixture `git` helper, a repo-init helper and a `pull()`
wrapper instead of repeating those calls.
Rollback picks its command in one if/elif and names the branch it
restores, and `_pull_updates` passes `rollback_branch` straight through.
`_CheckoutPlan` always has the field, so `_apply_pulled_update` reads it
directly instead of through getattr.
A switch that changes the running code with no new commits to count set
commit_count to -1, which printed "commit count unknown on this shallow
checkout" on non-shallow checkouts. The plan now records that case, and
the update says it switched the running checkout instead.
A regression test pins the stale-local-main case. A merge that returns 0
without moving HEAD must still be refused, which rules out both a
running-code baseline and a "the switch already moved" bypass.
The pre-fetch heal (#123254) and the post-update tag fetch now share
fetch_full_commit_graph, so both follow one clone-mode rule: a partial clone
repeats its own filter, a full clone fetches unfiltered (#122353), and only a
depth-limited clone whose history is really missing unshallows as tree:0. A
full clone grafted by a user's `git fetch --depth 1` already has its history
locally, so it unshallows unfiltered and stays full.
Also: the network-timeout message no longer claims "no response from the
remote" (a large transfer on a slow link hits the same cap), without adding a
classifier rule that would assert the opposite.
A depth-1 installer checkout far behind main could never update: the plain
main fetch pulled nearly the full object graph (merge-heavy history drags in
side-branch ancestry past the shallow boundary), hit the 300s network cap
mid-transfer every time, and the unshallow step that makes later fetches
incremental only ran after a successful update. Unshallow the checkout first
(commit-graph-only fetch with the same 900s budget the post-update tag fetch
uses, pulling the update branch along), non-fatally, and stop reporting a
wall-clock kill as 'no response from the remote'.
Fixes#123254
On a partial clone (clone --filter=tree:0), git 2.53/2.54 crashes every
fetch: index-pack --promisor runs repack_local_links(), which feeds
pack-objects --exclude-promisor-objects-best-effort, and pack-objects
BUG()s (SIGABRT, 'should only be called on existing objects') when the
link traversal hits a legitimately missing promisor object. The crash is
deterministic, so every fetch — and with it the CLI's update/check flows
— stays broken until the user applies the manual workaround themselves.
The recovery is the reporter's verified workaround, wired in: on this
exact crash (all three stderr markers), retry the fetch once with
-c remote.origin.promisor= (per-invocation only; the user's filter
choice stays in their config). Wired into both fetch sites — the update
fetch and the check fetch (which also serves forks via upstream).
Unrelated fetch failures are returned untouched; if the retry still
crashes, the update path prints the manual heal command.
Fixes#124272
SourceTarget carries either a commit or a branch, but types both optional;
the typed tracking_refspec() made ty see the branch arm as str | None.
Assert the constructor invariant where the branch is taken (ty diff vs
main: 9 vs 10 diagnostics, none new).
A checkout with no local copy of the target branch (every tag-pinned narrow
clone, which is detached) reaches the branch with `checkout -B <b> origin/<b>`.
That lands HEAD on the target before `rev-list HEAD..origin/<b>` runs, so the
count was always 0: the update printed "Already up to date!" after moving the
code, and skipped dependency sync, config migration and restarts.
Capture HEAD before that checkout, count from it, and hand it to the pull
phase as pre_sync_sha — the same channel the fork upstream sync already uses —
so the did-HEAD-move guard and syntax rollback compare against the code that
was actually running.
The forced `+refs/heads/<b>:refs/remotes/<r>/<b>` refspec was spelled out at
all three fetch sites (update --check, the apply fetch, the fork upstream
sync). Its leading `+` is load-bearing on depth-1 clones, so one helper,
`update_cmd_check.tracking_refspec`, now owns the string and the reason.
Tag-pinned --single-branch clones configure remote.origin.fetch as a tag-only
refspec, so fetching a branch by name writes FETCH_HEAD and never creates
refs/remotes/origin/<branch>. Every downstream rev-parse of the compare branch
then failed with 'Branch not found' even though the remote plainly has it
(#125112).
Cover the whole class: the --check fetch, the apply-path fetch, and the fork
upstream-sync fetch all pass an explicit +refs/heads/<branch>: mapping.
Verified A/B on a real narrow clone: branch-name fetch leaves origin/main
unresolvable, refspec fetch resolves it.
(cherry picked from commit 6ad2ff3f3f6190ad1fdaed4356d00406019e8dcc)
Passive checks (`hermes --version`, the banner, the Desktop/dashboard
update check) now make one channel-read attempt again. Offline usually
surfaces as DNS EAI_AGAIN or ENETUNREACH, which pm.network.is_transient
classes as transient, so wrapping every read in retry_network added 7 s
of backoff to the synchronous version line, stretched a hung CDN from
30 s to ~127 s, and logged a WARNING to stderr on every retry.
`release_channels.retrying_reads()` opts a block in; update_cmd wraps
the channel resolution of `hermes update` in it. Tests: keep the
Retry-After retry test (now scoped to the update path) and replace the
budget-exhaustion test with one pinning that passive reads and 404s make
exactly one attempt.
A git-less Windows ZIP install has no .git and never runs git, yet
_prepare_git_command called expose_pm_git() before it checked
use_zip_update: offline or with no pin, pm.ensure raised and aborted an
update that never needed git. expose_pm_git now takes the project root
and does nothing unless it is a git checkout (both callers).
When install.ps1 runs as one process, the git it staged into PM's store
is on the inherited PATH, so expose_pm_git returned early and PM's facts
never recorded it; plain hermes processes (plugin git installs, doctor,
version info) had no git until the first `hermes update`. A git found
under PM's store is now treated as that unrecorded copy and ensured, so
the source completion writes the git fact at install time.
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
On a Windows machine with no git, scripts/install.ps1 stages the pinned Git
for Windows into the pm store for its own process only, and pm's facts
never record it. Every later product process that ran a bare `git` failed:
- `hermes update` died before its first step:
"✗ Update failed: [WinError 2] The system cannot find the file specified"
- the source completion that finishes an installer re-run or an update
found no commit, deleted the install stamp, and the next boot wrote an
adoption stamp that names no commit. It also printed
"Could not refresh release history ([WinError 2] ...)".
expose_pm_git(): when Windows resolves no git, ensure PM's git explicitly
(both callers are user-initiated, like ensure_tools_for_sync) and put its
composed PATH (cmd and usr\bin) on the process, so children inherit it.
The update calls it before its first git, and the source completion calls it
before its builds and stamp. A machine with a working git is untouched.
main_desktop already ensures PM's git for its build for the same reason.
Both update shapes move HEAD off a detached commit: the branch update checks
out main, the release update checks out the release commit. A commit made on
the detached HEAD is on no branch, so after the move only the expiring reflog
reached it, and the branch path printed nothing about it (gated on #124643).
One helper, _park_detached_head, now runs at both sites before HEAD moves.
When HEAD is detached at a commit no ref contains (refs/stash excluded: the
autostash is dropped after the update, which is why the branch path parks
before stashing), it writes refs/hermes-update-backups/detached-<branch>-<ts>-<sha>
(the divergence rescue refs' scheme, now pruned on the same terms) and prints
the ref and how to list the commits. If the write fails the update stops
rather than orphan the work. A detached HEAD already on a branch or tag (a
pinned release) gets no ref. It replaces the release path's silent, never
pruned refs/hermes/pre-release/<sha>.
`hermes update` only rebuilt Desktop when apps/desktop/release/ or dist/
existed, so an installed Hermes.app that only the checkout update refreshes
(a bootstrap build) never got newer once release/ was gone or never built.
Count such bundles as a Desktop to keep current. Ownership is read from the
bundle's install-stamp.json: `updateMechanism: self` (or a stamp older than
the field). Self-updating releases and commit builds are never rebuilt or
copied over, and only the checkout an installed app actually runs (the one
under the default Hermes home) claims it. The post-build install now uses
the same set.
_cmd_update_check (CC 30, 224 lines) did debris cleanup, channel
resolution, the scoped upstream/origin fetch, shallow-graft repair and
two verdict printers inline. Those steps move to the topical sibling
hermes_cli/update_cmd_check.py; the facade function (a frozen updater
surface) stays in update_cmd and orchestrates them at CC 10.
Output, exit codes and git argv are unchanged. The sibling reads
_capture_head_sha / _no_prompt_git_kwargs through the facade and
imports gitlock, source_releases, source_check and config per call, so
existing monkeypatch seams keep reaching production.
Print the count of commits not on origin/<branch> next to the rescue ref and
a `git log origin/<branch>..<ref>` command, keep pruning the other ref kind
when one for-each-ref call fails, and drive the regression through the real
_pull_updates path against a temp git repo.
Co-authored-by: Bart Collet <8973350+bartcollet@users.noreply.github.com>
Co-authored-by: nca7777 <213462973+nca7777@users.noreply.github.com>
On the update's target branch a diverged history is resolved with
`reset --hard origin/<branch>`. Divergence there has two causes the checkout
cannot tell apart: an upstream force-push or rebase, where nothing local is
lost, and local commits on that branch, where the reset discards every one of
them. Only the orphan case (no common ancestor) wrote a rescue ref, so the
common case left the reflog as the sole way back: a 90-day expiry the user has
to know to reach for, in a checkout Hermes updates unattended. The custom-branch
path above already treats local commits as worth preserving; this is the same
work on the branch the update targets.
The reset itself is unchanged. `pre_pull_sha` is now anchored for both kinds,
under `refs/hermes-update-backups/diverged-<branch>-…` or `…/orphan-…`, and the
message names the recovery command for the diverged kind.
The #87694 footprint concern does not carry over. There `pre_pull_sha` is an
autostash orphan commit holding a full working-tree snapshot, which is what
could reach multi-GB. Here it is ordinary branch history, whose objects the
reflog already pins for its own expiry window. Both kinds expire under the same
keep-newest and max-age rules, which `_prune_orphan_rescue_refs` now applies per
kind rather than to the orphan prefix alone — otherwise the new refs would
accumulate forever, which is the actual shape of #87694.
Tests: a real two-repo divergence proves the discarded commit stays reachable
through the ref, and the existing mock test for this path now pins the new rule
instead of the old skip.
Signed-off-by: moep90 <volleyballlive@googlemail.com>
(cherry picked from commit 8b9568f9fd2e3811ec7afd8320dc432d135374e2)
pyproject.toml is a static 0.0.0 on source checkouts since install-stamp.json
became the runtime identity, but the "Update complete! (vA → vB)" line still
read it on both sides, so a main → pm-clean update printed "v0.21.4 → v0.0.0"
and every later update would print "(v0.0.0)". Both sides now use the git
identity write_source_stamp publishes moments later, so the line and the stamp
agree; a tagless checkout shows git.<sha> without a bogus "v" prefix.
_read_project_version stays: old updaters import it from the new tree.
Conflict resolutions and semantic fixups:
- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
fence missed, deregister from _live_foreground in a finally) wrapped around
pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
the old early-recovery block stays gone; main's interrupted-pull restore
(auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
and is cleared once git is done. The marker's target is the ref git actually
moves to (a release tag, not always origin/<branch>), since the restore
compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
had dropped and main's auto-merged restore needs (NameError on the first
launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
settings.about block; trim it to `updates` as pm-clean's type and the other
overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
key-leak switch leg needs the SDK, and the api_server two-tenant test needs
aiohttp, both PM runtime extras the test env does not carry.
Review blocker on this PR: the pull marker was removed only on the success
path. When `_reconcile_diverged_checkout` exited (merge conflict on a custom
branch, failed reset) the marker stayed with a dead owner, the updater told
the user to `git stash apply`, and the next `hermes` launch ran
`git restore` over every dirty file upstream also changed, wiping the
re-applied work (and restoring files mid-merge, leaving a dangling
MERGE_HEAD).
- The marker is removed on every exit of the git phase except
KeyboardInterrupt, which reaches git too and can leave a torn tree.
- The marker records the resolved target commit, not `origin/<branch>`,
so a later `git fetch` cannot widen the restore.
- Startup restores only paths whose content is exactly the target blob
(per `git hash-object`, so autocrlf/filters match) or, for deletions,
that are gone; anything else is the user's and stays.
- Startup does nothing while a merge, rebase, cherry-pick or revert is in
progress.
- The marker lives in the real git dir, so linked-worktree installs are
covered too; the bare `except Exception: pass` is narrowed and reports.
- New test pins the main.py wiring: the restore/relaunch runs before any
other hermes_cli module is imported.
Git moves the checkout file by file and HEAD last, so an update killed while
"Pulling updates" (Ctrl-C, closed terminal, power loss, OOM) left HEAD on the
old commit with a prefix of the files already new. That mixed tree failed at
import in every entry point, `hermes update` included
(`cannot import name 'load_yaml_file_readonly' from 'utils'`), and only a
manual `git reset --hard` could bring the install back.
The updater now brackets the fast-forward (and the divergence reconcile) with
`.git/hermes-update-pull` naming its pid, the pre-pull commit and the target.
`hermes_cli.main` calls `_early_recovery.restore_interrupted_pull()` right
after `hermes_bootstrap`, before any other checkout import: when the marker's
owner is gone and HEAD is still the pre-pull commit, every path that differs
between the two commits goes back to HEAD (the tree the venv was built for),
the dead git's index.lock is cleared, and the command re-execs itself (modules
already imported, main.py included, may be the half-written ones), so a rerun
`hermes update` repairs the install and then updates normally. Paths the update does not change keep local edits; the
updater's autostash stays in `git stash list`. Our own pid is never treated as
a live owner (containers hand a retry the killed updater's pid).
Hindsight now ships from its maintainer's repo (vectorize-io/hindsight,
hindsight-integrations/hermes) through plugin-catalog/hindsight.yaml, so the
in-tree copy under plugins/memory/hindsight goes away. Homes that still name
`memory.provider: hindsight` are migrated by hermes_cli/memory_provider_migration.py
(the `hermes update` hook and the first agent start install the catalog plugin);
that path is untouched here.
Core-side special cases that only made sense with the bundled copy go with it:
the `memory.hindsight` LAZY_DEPS feature, `_provider_pip_dependencies`'s
`hindsight-all` expansion in `hermes memory setup` (the plugin's own
`_ensure_local_runtime()` installs the embedded runtime now), the compat-manifest
pointers of the deleted module, and prose that listed hindsight among the in-tree
providers. Generic provider-name lists, the HINDSIGHT_* env hints and the API-key
redaction pattern stay — a catalog-installed hindsight still uses them.
Revert this commit alone to restore the bundled provider.
Every update route now finishes through update_completion._complete_selected,
which restarted the whole fleet unconditionally -- including the "Already up to
date" route that main sent through the pending-restart catch-up. Net effect:
each cron tick and each profile's `hermes update` drained and re-killed the one
multiplexed gateway.
Port the catch-up path's two live guards into the completion tail:
- host_restart_already_completed(checkout sha): a sibling profile attaches to
the restart this host already stamped (#95294); the restart phase now stamps
it via mark_host_restart_completed.
- every planned runtime AND every live fleet row current at the checkout sha
(#117051, d6b0d37ece). Both are required: the live matrix lists gateways
only, so a planned serve still on pre-update code keeps the restart.
The already-current route also arms the host obligation with the checkout sha;
an SHA-less arm replaced the standing record and wiped the restarted proof.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Two distinct defects in the Windows gateway pause/resume path, one
symptom: "hermes update" aborting exit 1 on a checkout that was
already current and needed no code change.
1. Asymmetric error handling between the two finish paths. The pull
path wraps the Windows gateway resume in
_resume_windows_gateways_and_merge_outcome, whose contract is
explicit: "Must never abort the update" -- it catches the failure,
marks the outcome incomplete, warns, and continues (receipt status
"partial", exit 1 only via the existing incomplete-repair gate).
_finish_already_up_to_date ("Already up to date" path) called
_resume_windows_gateways_after_update bare, so the identical
RuntimeError -- e.g. the relaunch-verification race with a parent
Job Object kill, #48820 -- killed the run as an unhandled exception
instead. Fixed by routing through the same merge helper and folding
a failed resume into current_checkout_complete, landing on the
pre-existing "partial" finalize + sys.exit(1) contract the repair
path already uses for other incomplete states.
2. The atexit callback could replay the same failure a second time.
Every foreground call site registers
_resume_windows_gateways_after_update via atexit.register as a
dead-process safety net, but the function only cleared
token["resume_needed"] on its OWN success path -- a raise from
_verify_relaunched_gateways_alive left the registration live, so
interpreter teardown called the function again with the same
still-armed token, repeated the whole relaunch + verify, failed
identically, and printed "Exception ignored in atexit callback".
Fixed by unregistering the function from atexit as soon as
execution reaches it (foreground or the atexit fallback itself) --
ownership transfers to whichever call site got there first, and a
failure below is reported exactly once. atexit.unregister is a
no-op when the function was never registered, so a non-Windows or
no-resume-needed call is unaffected.
Both fixes are surgical: they leave the "a failed relaunch keeps
resume_needed=True so the token is retryable" contract completely
untouched (see the existing
test_resume_windows_gateway_service_failure_stays_retryable /
test_resume_windows_gateway_launcher_refresh_failure_stays_retryable
tests, which this PR does not modify) -- that design is intentional,
not the bug.
Closes#115563.
Testing: Windows-specific behavior verified by code trace + full
resolution-chain tests against the real functions (mocked I/O, real
control flow) on Linux/macOS, per the "Don't fake the host OS" rule
this repo's own AGENTS.md sets -- no sys.platform patching, and the
Windows-only branch (_is_windows() == True) is exercised exactly as
the existing test suite for this module already does (monkeypatching
_is_windows on the shared test fixtures, not faking the platform).
True live-Windows process-topology proof is out of scope for this PR
(the wine2e lane is reserved for process-topology E2E per AGENTS.md);
this fix is control-flow/bookkeeping, not host-specific syscalls.
- 2 new regression tests, both proven red on the pre-fix code via
git stash (both fail with the exact defect signature: the "Already
up to date" path re-raises instead of demoting to partial; the
atexit unregister call never happens) and green on the fix
- scripts/run_tests.sh across the full affected surface (Windows
update/resume/reconciliation + related pause/resume/venv-repair
test files): 228 passed, 0 failed, 4 skipped (pre-existing,
unrelated to this change)
Closes the second #97208 atom (qingchuan-x): after a Windows hand-off child refused the
dependency sync (Desktop backend held the venv), the next `hermes update` found git current,
passed the core-import probe — the old release imports fine — and printed
"✓ Already up to date!" while the installed distribution still read hermes-agent 0.20.6
against a 0.21.2 checkout with lagging pins.
`_venv_dependency_set_stale` asks the venv's own interpreter for the installed hermes-agent
version and compares it with pyproject's; a mismatch on the current-checkout path takes
the same repair route as an unhealthy venv. Unknown states (no venv, not installed as a
distribution, probe failure) are never stale, so dev checkouts are untouched.
Review follow-up on the #101600 fix:
- The wait moves from `_cmd_update_impl` to `cmd_update`, BEFORE `UpdateLock.acquire()`.
The child used to run under the parent's marker (handoff pid / ancestor claim); the
parent's exit — which the child now deliberately waits for — released that marker, so
the whole Windows dependency install and gateway resume ran with no update lock. Waiting
first lets the child claim a marker of its own for the tail.
- `SHIM_PARENT_PID_ENV` names the `hermes.exe` launcher ancestor (the process that holds
the shim image open and exits only after reaping this interpreter) when psutil sees one,
else this pid. `_windows_shim_in_process_chain` is split into `_venv_shim_matcher` +
`_windows_shim_ancestor` so the holder pid shares the matcher; docs sentence adjusted.
- The direct-call wait test is replaced by one driving `cmd_update` in a real child with an
older parent: proves the phase starts after the parent died, the child owns the lock, the
handed-off pause token is adopted (no re-discovery) and the env is consumed. Removing the
wait call, the token adoption, or moving the wait back after the lock each turns it red.
`hermes update` run from `venv\Scripts\hermes.exe` hands the dependency install to a
child under the venv Python because the shim cannot replace itself. Two races were left
in that hand-off (#101600, three independent Windows reproductions):
- The child never waited for the shim-run parent. The shim quarantine is a single
`os.rename` with sub-second retries; when the child reached it while the parent was
still alive the rename failed with PermissionError, `ShimQuarantineError` deferred the
whole install and the receipt was marked failed although the checkout had advanced.
- On the legacy re-exec (`_abort_dependency_sync_if_self_locked` ->
`_reexec_dependency_sync_off_windows_shim`, still live for a current checkout with an
unhealthy venv -- i.e. the retry after any failed update) the PARENT resumed the paused
gateways after spawning the child, holding hermes.exe open for the whole relaunch and
raising `RuntimeError: Windows gateway relaunch after update was not verified alive`;
the child then re-ran pause discovery, found that freshly relaunched gateway before its
pid file existed and force-killed it as "without profile mapping".
Now every child spawned off the shim gets `HERMES_UPDATE_SHIM_PARENT_PID` and, at the top
of `_cmd_update_impl`, waits (bounded, 30 s, `(pid, create_time)` identity via psutil)
for that process to exit before scanning holders, pausing gateways or renaming shims. The
legacy re-exec carries the Windows pause token to the child in
`HERMES_UPDATE_GATEWAY_RESUME` and disarms the parent's copy, so the parent exits at once
and the child resumes exactly the fleet the parent stopped instead of rediscovering it --
the same ownership transfer the post-swap hand-off already does (94ced1a2b2).
The wait is host-independent and proven with real processes on Linux; the shim paths run
on the windows-latest lane through the existing `_is_windows`-patched suite.
Fixes#101600
Refs #98010#103031
Supersedes #101630
Co-authored-by: fangliquanflq <fangliquan@qq.com>
`import pm.ensure` bound the submodule onto the package, shadowing the facade's
`pm.ensure()` function for every later caller in the process (photon's sidecar
start hit `'module' object is not callable`). The module is pm.install now; the
function keeps its name. The facade resolves through `__import__` rather than
`importlib.import_module` so a test that patches import_module globally does not
break attribute access on pm.
tools/lazy_deps.py returns to the 16-line stop_for_relaunch shim the branch wrote
(an origin/main merge had replaced it with main's 775-line implementation); the
project-metadata tests follow. update_cmd re-exports the four old_updater_deps
names the shim tests resolve through hermes_cli.update_cmd.
Both repair paths now run the memory-provider refresh between the tool-dep
restore and the plugin-dep reapply, exactly like _sync_python_dependencies_after_pull
and the ZIP path, so the three routes stay interchangeable.