303 Commits

Author SHA1 Message Date
Brooklyn Nicholson
b9b1328b7d fix(cli,desktop): hide every console-subsystem child spawn on Windows
A console-less parent (pythonw backend, detached daemons) allocates a visible
console per bare spawn, so Windows startup flashes ~7-8 console windows
(#117781) and the GPU statusbar flashes two more every 5s (#101895). Threads
creationflags=windows_hide_flags() through every hermes_cli spawn site the
desktop reaches: gitlock (tasklist probe + git plumbing), the scheduled-task
PowerShell probe, the updater's _git_run/_no_prompt_git_kwargs/_GIT_TEXT_KW
runners, the safe-directory config probe, and consolidates the statusbar's two
nvidia-smi reads per poll into one cached hidden query. The statusbar also
stops polling while the document is hidden and resumes on visibilitychange
(#120262).

Supersedes #117805, #117833, #102047, #120285 (thanks @JoaoMarcos44, @Finn763,
@monerostar, @szzhoujiarui).
2026-09-29 18:49:58 -05:00
kshitijk4poor
db17e07389 fix(update): mark current updater frames with a sentinel, not missing locals
The pre_update_version co_varnames check added in af26acab73 rejected the
current updater's entrypoints, but it also rejected historical updaters
from 2026-07-26..08-16, whose _cmd_update_impl/cmd_update never declared
that local — their in-process dependency installs would have continued on
the old interpreter instead of handing off to the takeover child.

Invert the gate: the CURRENT updater declares the
_hermes_current_updater_frame sentinel local in both entrypoints; no
historical on-disk version can contain it. in_historical_update() now
hands off on entrypoint match unless the sentinel is present, restoring
the handoff for every historical window while still refusing current
frames. Verified live for all four frame classes; mutation check: with
the sentinel check removed, the current-updater refusal test fails.
2026-09-29 16:47:48 +05:30
kshitijk4poor
282ad01cb5 fix(update): mark unmarked packs instead of the inert promisor retry (#124272)
git 2.53+ `index-pack --promisor` hands every object outside the promisor
packs to pack-objects, which BUG()s (should_include_obj) on the missing
objects those objects lead to. A partial clone with unmarked packs therefore
keeps crashing its fetches.

The retry added for this ran `git -c remote.origin.promisor= fetch`. That
override does nothing: git registers promisor remotes additively, so the
repo's own `true` still wins (GIT_TRACE on 2.50 and 2.55 shows
`index-pack --promisor` and `filter tree:0` under the override). The retry
only passed when the crashed fetch had already stored the pack; a retry that
has to transfer again crashed again.

- On the crash, give every pack without a `.promisor` marker one, then retry
  the same fetch once. The clone stays partial, no refetch, no config change.
  If git still crashes, the updater points at the docs section.
- The two fetches that turn a full clone partial now mark the clone's old
  packs afterwards: the depth-1 `tree:0` unshallow in
  `fetch_full_commit_graph`, and `worktree_ops._deepen_shallow_repo`
  (`--filter=blob:none`). git writes the partial-clone config before it
  fetches, so this also runs when the fetch fails. Without it, a converted
  depth-1 install crashed every later fetch on git 2.55 (3 of 3 rounds in a
  local repro).
2026-09-29 15:24:15 +05:30
teknium1
2005a9a817 fix(telemetry): update initiator reads the lock marker raw
The Desktop-initiator fact called read_live_update(), whose pid liveness probe
(os.kill) and stale-marker unlink are update-lock policy. Running them from a
metrics helper tripped the retired-holder-gate tripwire in
test_update_venv_holder_retirement and could delete a marker as a side effect.
We already hold the lock, so a marker naming another pid is the orchestrator.
2026-09-28 12:43:03 -07:00
teknium1
fd14f50b5d feat(metrics): hermes.update.run/stage derived from the final update receipt
The update pipeline already writes one machine-readable receipt per run, so the
run/stage counters are derived from it in the process that finalizes it instead
of instrumenting stages for metrics. The receipt gains minimal stage END marks
(plan, snapshot, apply[mode], deps, build, restart; verify is inferred from the
fleet matrix), an initiator fact (desktop when a live orchestrator claim
names another pid) and the pre-update commit date, which is all the derivation
needs for outcome, failed stage, duration and from-version-age buckets.

A pre-pull interpreter must never import pulled code: when it is the finalizer
(same pid, checkout sha moved) it parks the receipt (stdlib only) and the next
Hermes start records it. Collection off => nothing imported beyond a config read.
The completion-process test stub gains the new record_stage API the real
module now exposes (no assertion changed).
2026-09-28 12:43:03 -07:00
Teknium
a5bd246865 Old pre-decomposition import paths are gone: plugin compat layer removed on schedule (#126164)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule

The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:

- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
  names, restored dead definitions) and the three re-export stub modules
  (gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
  the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
  `plugins.allow_deprecated_imports` escape hatch

An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.

hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).

In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.

* chore: retrigger CI (zero-job startup_failure phantom)

* test: drop resolution allowlist rows for the two deleted which() sites

hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
2026-09-28 10:21:41 -07:00
teknium1
52d814cfd0 fix(update): an autostash the update could not restore ends the update partial, never 'Update complete!'
A restore that conflicts (or whose untracked baseline is unknown, or whose
untracked files the update replaced) parks the user's patch and the update
still printed 'Update complete!' and exited 0 (#122557). The run's stash now
stays unsettled until it is restored, discarded, or parked because the user
asked (--keep-stash, declined prompt); an unsettled stash becomes the
completion line, which is not a success line, so the update ends partial,
exits 1 and the receipt says partial. Failure verdicts name it too.
2026-09-28 05:33:25 -07:00
finn763
f77f1a2a37 fix(update): name this run's unsettled autostash before any update outcome (#122557) 2026-09-28 05:33:25 -07:00
JoaoMarcos44
353ac62b31 fix(update): refuse active Git operations before mutation 2026-09-28 05:33:25 -07:00
JoaoMarcos44
f4b49df25f fix(update): prove divergence before reset fallback 2026-09-28 05:33:25 -07:00
teknium1
f46c7159e4 fix(update): trim SHA-less obligation salvage to the checkout fallback (#125952, salvage #125964)
Follow-up trim of the #125964 salvage:

- Armer falls back to _current_checkout_sha() (the identity the reader compares the fleet
  against) instead of a new receipt helper; at arm time the receipt has no post_update yet.
- Reader: the SHA-less, inventory-less record follows the existing inventory-less rule
  (live fleet all current on the checkout) with no extra pid-liveness gate, matching the
  invariant the SHA-bearing inventory-less record already has. The gatewayless branch stays
  fail-closed for an SHA-less record (checkout_contains("") is not evidence).
- Drop the obligation_fields pid plumbing and update_receipt helper; keep 2 tests:
  discharge on live-fleet evidence, and the armer recording the checkout SHA.
2026-09-28 04:44:29 -07:00
kokhlo
e57f6d0f2d fix(update): discharge an SHA-less host restart obligation on live-fleet evidence
A no-op `hermes update` (pre == post SHA) whose head capture failed armed the
host update-restart record with `expected_sha: ""` and no inventory. The reader
bailed out on the empty SHA before the live-fleet reconciliation, so the
"gateways may still be serving pre-update modules" warning could never clear
even when every live gateway was current on the checkout.

Reader (hermes_cli/update_cmd_fleet.py):
- An inventoried obligation without its SHA stays fail-closed.
- An inventory-less, SHA-less record now falls through to the existing
  live-fleet check once the process that armed it is gone (pid liveness via
  pid_exists_stdlib), so a status call cannot discharge a restart phase that
  is still running.

Armer (hermes_cli/update_cmd.py):
- When the completion request carries no SHA, fall back to the receipt's
  post_update.sha (then pre_update.sha) instead of arming with "".

obligation_fields() now reports the arming pid alongside expected_sha.

Tests: discharge on all-current fleet, kept when the armer pid is alive, kept
on a stale fleet, kept with an inventory but no SHA; the discharge test fails
on the pre-fix reader with the issue's exact warning.
2026-09-28 04:44:29 -07:00
kshitijk4poor
e59df70a12 refactor(update): name the parked rollback once and reuse the test helpers
Review follow-up: compute the parked-branch case once in the rollback,
fold the baseline comments into one, reuse pull() and a new complete()
helper in the tests, and say in the docs that a rollback restores the
commit the checkout ran before the pull on every update path.
2026-09-28 16:00:01 +05:30
kshitijk4poor
9e8202f802 fix(update): never leave a failed branch rollback on broken code
Review follow-ups for the branch-switch rollback:

- If the parked branch cannot be checked out again (for example, another
  worktree holds it), restore its commit detached. Before this, the install
  stayed on the broken update branch, and the recovery hint repeated the
  failing command.
- Any `merge-base --is-ancestor` result other than "contained" keeps the
  strict baseline. rc 128 used to fall back to the lenient one, which
  contradicted its own comment.
- `_pull_updates` returns the baseline it verified, and
  `_apply_pulled_update` checks against it. Before, that function
  recomputed the baseline with a second rev-parse and merge-base against a
  hardcoded `origin/<branch>`, which a fork push in between could move.
- Drop a `commit_count == -1` check that could never be false.
- Update the rollback description in `updating.md`.
- Tests share the fixture `git` helper, a repo-init helper and a `pull()`
  wrapper instead of repeating those calls.
2026-09-28 16:00:01 +05:30
kshitijk4poor
a4ed805290 fix(update): tidy the branch-switch rollback and name a switch-only update
Rollback picks its command in one if/elif and names the branch it
restores, and `_pull_updates` passes `rollback_branch` straight through.
`_CheckoutPlan` always has the field, so `_apply_pulled_update` reads it
directly instead of through getattr.

A switch that changes the running code with no new commits to count set
commit_count to -1, which printed "commit count unknown on this shallow
checkout" on non-shallow checkouts. The plan now records that case, and
the update says it switched the running checkout instead.

A regression test pins the stale-local-main case. A merge that returns 0
without moving HEAD must still be refused, which rules out both a
running-code baseline and a "the switch already moved" bypass.
2026-09-28 16:00:01 +05:30
funky-xamarin
593eaf6c07 fix(update): distinguish early fork sync from pending branch repair 2026-09-28 16:00:01 +05:30
funky-xamarin
a58fc4fa8e fix(update): retain branch repair baseline through fork completion 2026-09-28 16:00:01 +05:30
funky-xamarin
495ab057f1 fix(update): verify stale branch repair against its immediate tip 2026-09-28 16:00:01 +05:30
funky-xamarin
57b3868d52 fix(update): preserve running checkout across branch-switch updates 2026-09-28 16:00:01 +05:30
teknium1
010618c767 fix(update): one commit-graph fetch for the pre-fetch heal and the post-update pass
The pre-fetch heal (#123254) and the post-update tag fetch now share
fetch_full_commit_graph, so both follow one clone-mode rule: a partial clone
repeats its own filter, a full clone fetches unfiltered (#122353), and only a
depth-limited clone whose history is really missing unshallows as tree:0. A
full clone grafted by a user's `git fetch --depth 1` already has its history
locally, so it unshallows unfiltered and stays full.

Also: the network-timeout message no longer claims "no response from the
remote" (a large transfer on a slow link hits the same cap), without adding a
classifier rule that would assert the opposite.
2026-09-28 03:08:56 -07:00
Konstantin Khlopkov
e377309d86 fix(update): heal stale shallow history before the bounded fetch
A depth-1 installer checkout far behind main could never update: the plain
main fetch pulled nearly the full object graph (merge-heavy history drags in
side-branch ancestry past the shallow boundary), hit the 300s network cap
mid-transfer every time, and the unshallow step that makes later fetches
incremental only ran after a successful update. Unshallow the checkout first
(commit-graph-only fetch with the same 900s budget the post-update tag fetch
uses, pulling the update branch along), non-fatally, and stop reporting a
wall-clock kill as 'no response from the remote'.

Fixes #123254
2026-09-28 03:08:56 -07:00
Yuan Li
fc6144e3c1 fix(update): self-heal git pack-objects crash on partial-clone fetches
On a partial clone (clone --filter=tree:0), git 2.53/2.54 crashes every
fetch: index-pack --promisor runs repack_local_links(), which feeds
pack-objects --exclude-promisor-objects-best-effort, and pack-objects
BUG()s (SIGABRT, 'should only be called on existing objects') when the
link traversal hits a legitimately missing promisor object. The crash is
deterministic, so every fetch — and with it the CLI's update/check flows
— stays broken until the user applies the manual workaround themselves.

The recovery is the reporter's verified workaround, wired in: on this
exact crash (all three stderr markers), retry the fetch once with
-c remote.origin.promisor= (per-invocation only; the user's filter
choice stays in their config). Wired into both fetch sites — the update
fetch and the check fetch (which also serves forks via upstream).
Unrelated fetch failures are returned untouched; if the retry still
crashes, the update path prints the manual heal command.

Fixes #124272
2026-09-28 03:08:56 -07:00
kokhlo
c70374b561 fix(update): map SSL_CERT_FILE into GIT_SSL_CAINFO for git children 2026-09-28 02:00:41 -07:00
kshitijk4poor
57fba5f772 fix(update): narrow the source target's branch before building the refspec
SourceTarget carries either a commit or a branch, but types both optional;
the typed tracking_refspec() made ty see the branch arm as str | None.
Assert the constructor invariant where the branch is taken (ty diff vs
main: 9 vs 10 diagnostics, none new).
2026-09-28 01:53:35 +05:30
kshitijk4poor
80d6e395db fix(update): count new commits from where HEAD was when checkout -B moves it
A checkout with no local copy of the target branch (every tag-pinned narrow
clone, which is detached) reaches the branch with `checkout -B <b> origin/<b>`.
That lands HEAD on the target before `rev-list HEAD..origin/<b>` runs, so the
count was always 0: the update printed "Already up to date!" after moving the
code, and skipped dependency sync, config migration and restarts.

Capture HEAD before that checkout, count from it, and hand it to the pull
phase as pre_sync_sha — the same channel the fork upstream sync already uses —
so the did-HEAD-move guard and syntax rollback compare against the code that
was actually running.
2026-09-28 01:53:35 +05:30
kshitijk4poor
0ee82ff4cc refactor(update): build the tracking refspec in one place
The forced `+refs/heads/<b>:refs/remotes/<r>/<b>` refspec was spelled out at
all three fetch sites (update --check, the apply fetch, the fork upstream
sync). Its leading `+` is load-bearing on depth-1 clones, so one helper,
`update_cmd_check.tracking_refspec`, now owns the string and the reason.
2026-09-28 01:53:35 +05:30
Yuan Li
ab6cc2fd84 fix(updater): fetch branches by explicit refspec so narrow clones resolve the compare ref
Tag-pinned --single-branch clones configure remote.origin.fetch as a tag-only
refspec, so fetching a branch by name writes FETCH_HEAD and never creates
refs/remotes/origin/<branch>. Every downstream rev-parse of the compare branch
then failed with 'Branch not found' even though the remote plainly has it
(#125112).

Cover the whole class: the --check fetch, the apply-path fetch, and the fork
upstream-sync fetch all pass an explicit +refs/heads/<branch>: mapping.

Verified A/B on a real narrow clone: branch-name fetch leaves origin/main
unresolvable, refspec fetch resolves it.

(cherry picked from commit 6ad2ff3f3f6190ad1fdaed4356d00406019e8dcc)
2026-09-28 01:53:35 +05:30
teknium1
d1484ab44b fix(update): only an explicit hermes update retries channel reads
Passive checks (`hermes --version`, the banner, the Desktop/dashboard
update check) now make one channel-read attempt again. Offline usually
surfaces as DNS EAI_AGAIN or ENETUNREACH, which pm.network.is_transient
classes as transient, so wrapping every read in retry_network added 7 s
of backoff to the synchronous version line, stretched a hung CDN from
30 s to ~127 s, and logged a WARNING to stderr on every retry.

`release_channels.retrying_reads()` opts a block in; update_cmd wraps
the channel resolution of `hermes update` in it. Tests: keep the
Retry-After retry test (now scoped to the update path) and replace the
budget-exhaustion test with one pinning that passive reads and 404s make
exactly one attempt.
2026-09-27 03:41:42 -07:00
teknium1
2270d3511d fix(update): expose PM's git only for git checkouts, and record the installer's copy
A git-less Windows ZIP install has no .git and never runs git, yet
_prepare_git_command called expose_pm_git() before it checked
use_zip_update: offline or with no pin, pm.ensure raised and aborted an
update that never needed git. expose_pm_git now takes the project root
and does nothing unless it is a git checkout (both callers).

When install.ps1 runs as one process, the git it staged into PM's store
is on the inherited PATH, so expose_pm_git returned early and PM's facts
never recorded it; plain hermes processes (plugin git installs, doctor,
version info) had no git until the first `hermes update`. A git found
under PM's store is now treated as that unrecorded copy and ensured, so
the source completion writes the git fact at install time.

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-27 02:54:03 -07:00
teknium1
305072f4a6 fix(update): give git-less Windows installs PM's git for update and source completion
On a Windows machine with no git, scripts/install.ps1 stages the pinned Git
for Windows into the pm store for its own process only, and pm's facts
never record it. Every later product process that ran a bare `git` failed:

- `hermes update` died before its first step:
  "✗ Update failed: [WinError 2] The system cannot find the file specified"
- the source completion that finishes an installer re-run or an update
  found no commit, deleted the install stamp, and the next boot wrote an
  adoption stamp that names no commit. It also printed
  "Could not refresh release history ([WinError 2] ...)".

expose_pm_git(): when Windows resolves no git, ensure PM's git explicitly
(both callers are user-initiated, like ensure_tools_for_sync) and put its
composed PATH (cmd and usr\bin) on the process, so children inherit it.
The update calls it before its first git, and the source completion calls it
before its builds and stamp. A machine with a working git is untouched.
main_desktop already ensures PM's git for its build for the same reason.
2026-09-27 02:54:03 -07:00
teknium1
ff46826872 fix(update): keep commits made on a detached HEAD behind a named rescue ref
Both update shapes move HEAD off a detached commit: the branch update checks
out main, the release update checks out the release commit. A commit made on
the detached HEAD is on no branch, so after the move only the expiring reflog
reached it, and the branch path printed nothing about it (gated on #124643).

One helper, _park_detached_head, now runs at both sites before HEAD moves.
When HEAD is detached at a commit no ref contains (refs/stash excluded: the
autostash is dropped after the update, which is why the branch path parks
before stashing), it writes refs/hermes-update-backups/detached-<branch>-<ts>-<sha>
(the divergence rescue refs' scheme, now pruned on the same terms) and prints
the ref and how to list the commits. If the write fails the update stops
rather than orphan the work. A detached HEAD already on a branch or tag (a
pinned release) gets no ref. It replaces the release path's silent, never
pruned refs/hermes/pre-release/<sha>.
2026-09-27 01:00:31 -07:00
Brooklyn Nicholson
53a785449f fix(update): rebuild Desktop for an installed app with no checkout build
`hermes update` only rebuilt Desktop when apps/desktop/release/ or dist/
existed, so an installed Hermes.app that only the checkout update refreshes
(a bootstrap build) never got newer once release/ was gone or never built.

Count such bundles as a Desktop to keep current. Ownership is read from the
bundle's install-stamp.json: `updateMechanism: self` (or a stamp older than
the field). Self-updating releases and commit builds are never rebuilt or
copied over, and only the checkout an installed app actually runs (the one
under the default Hermes home) claims it. The post-build install now uses
the same set.
2026-09-26 21:42:48 -05:00
ethernet
a67891b0ee refactor(update): split hermes update --check into update_cmd_check
_cmd_update_check (CC 30, 224 lines) did debris cleanup, channel
resolution, the scoped upstream/origin fetch, shallow-graft repair and
two verdict printers inline. Those steps move to the topical sibling
hermes_cli/update_cmd_check.py; the facade function (a frozen updater
surface) stays in update_cmd and orchestrates them at CC 10.

Output, exit codes and git argv are unchanged. The sibling reads
_capture_head_sha / _no_prompt_git_kwargs through the facade and
imports gitlock, source_releases, source_check and config per call, so
existing monkeypatch seams keep reaching production.
2026-09-24 14:08:22 -04:00
ethernet
cb7b18431a Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/src/app/settings/connections-registry.tsx
#	scripts/install.ps1
#	scripts/install.sh
#	tests/hermes_cli/test_update_autostash.py
2026-09-24 03:48:31 -04:00
brooklyn!
db1ff59937 fix(update): say how many commits the diverged reset drops and where they went
Print the count of commits not on origin/<branch> next to the rescue ref and
a `git log origin/<branch>..<ref>` command, keep pruning the other ref kind
when one for-each-ref call fails, and drive the regression through the real
_pull_updates path against a temp git repo.

Co-authored-by: Bart Collet <8973350+bartcollet@users.noreply.github.com>
Co-authored-by: nca7777 <213462973+nca7777@users.noreply.github.com>
2026-09-24 02:40:53 -05:00
moep90
21cc8bfef3 fix(update): anchor discarded local commits before the diverged reset
On the update's target branch a diverged history is resolved with
`reset --hard origin/<branch>`. Divergence there has two causes the checkout
cannot tell apart: an upstream force-push or rebase, where nothing local is
lost, and local commits on that branch, where the reset discards every one of
them. Only the orphan case (no common ancestor) wrote a rescue ref, so the
common case left the reflog as the sole way back: a 90-day expiry the user has
to know to reach for, in a checkout Hermes updates unattended. The custom-branch
path above already treats local commits as worth preserving; this is the same
work on the branch the update targets.

The reset itself is unchanged. `pre_pull_sha` is now anchored for both kinds,
under `refs/hermes-update-backups/diverged-<branch>-…` or `…/orphan-…`, and the
message names the recovery command for the diverged kind.

The #87694 footprint concern does not carry over. There `pre_pull_sha` is an
autostash orphan commit holding a full working-tree snapshot, which is what
could reach multi-GB. Here it is ordinary branch history, whose objects the
reflog already pins for its own expiry window. Both kinds expire under the same
keep-newest and max-age rules, which `_prune_orphan_rescue_refs` now applies per
kind rather than to the orphan prefix alone — otherwise the new refs would
accumulate forever, which is the actual shape of #87694.

Tests: a real two-repo divergence proves the discarded commit stays reachable
through the ref, and the existing mock test for this path now pins the new rule
instead of the old skip.

Signed-off-by: moep90 <volleyballlive@googlemail.com>
(cherry picked from commit 8b9568f9fd2e3811ec7afd8320dc432d135374e2)
2026-09-24 02:40:53 -05:00
ethernet
4b413254b0 fix(update): completion line reports checkout identity, not pyproject's inert 0.0.0
pyproject.toml is a static 0.0.0 on source checkouts since install-stamp.json
became the runtime identity, but the "Update complete! (vA → vB)" line still
read it on both sides, so a main → pm-clean update printed "v0.21.4 → v0.0.0"
and every later update would print "(v0.0.0)". Both sides now use the git
identity write_source_stamp publishes moments later, so the line and the stamp
agree; a tagless checkout shows git.<sha> without a bogus "v" prefix.
_read_project_version stays: old updaters import it from the new tree.
2026-09-24 00:24:41 -04:00
ethernet
d75d3fe15c Merge origin/main into ethie/pm-clean
Conflict resolutions and semantic fixups:

- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
  fence missed, deregister from _live_foreground in a finally) wrapped around
  pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
  the old early-recovery block stays gone; main's interrupted-pull restore
  (auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
  pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
  and is cleared once git is done. The marker's target is the ref git actually
  moves to (a release tag, not always origin/<branch>), since the restore
  compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
  had dropped and main's auto-merged restore needs (NameError on the first
  launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
  settings.about block; trim it to `updates` as pm-clean's type and the other
  overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
  names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
  key-leak switch leg needs the SDK, and the api_server two-tenant test needs
  aiohttp, both PM runtime extras the test env does not carry.
2026-09-23 21:55:59 -04:00
teknium1
f035bf872e fix(update): interrupted-pull restore never touches the user's own edits
Review blocker on this PR: the pull marker was removed only on the success
path. When `_reconcile_diverged_checkout` exited (merge conflict on a custom
branch, failed reset) the marker stayed with a dead owner, the updater told
the user to `git stash apply`, and the next `hermes` launch ran
`git restore` over every dirty file upstream also changed, wiping the
re-applied work (and restoring files mid-merge, leaving a dangling
MERGE_HEAD).

- The marker is removed on every exit of the git phase except
  KeyboardInterrupt, which reaches git too and can leave a torn tree.
- The marker records the resolved target commit, not `origin/<branch>`,
  so a later `git fetch` cannot widen the restore.
- Startup restores only paths whose content is exactly the target blob
  (per `git hash-object`, so autocrlf/filters match) or, for deletions,
  that are gone; anything else is the user's and stays.
- Startup does nothing while a merge, rebase, cherry-pick or revert is in
  progress.
- The marker lives in the real git dir, so linked-worktree installs are
  covered too; the bare `except Exception: pass` is narrowed and reports.
- New test pins the main.py wiring: the restore/relaunch runs before any
  other hermes_cli module is imported.
2026-09-23 17:42:24 -07:00
teknium1
e30f8c6344 fix(update): a hermes update killed mid-pull no longer bricks every entry point
Git moves the checkout file by file and HEAD last, so an update killed while
"Pulling updates" (Ctrl-C, closed terminal, power loss, OOM) left HEAD on the
old commit with a prefix of the files already new. That mixed tree failed at
import in every entry point, `hermes update` included
(`cannot import name 'load_yaml_file_readonly' from 'utils'`), and only a
manual `git reset --hard` could bring the install back.

The updater now brackets the fast-forward (and the divergence reconcile) with
`.git/hermes-update-pull` naming its pid, the pre-pull commit and the target.
`hermes_cli.main` calls `_early_recovery.restore_interrupted_pull()` right
after `hermes_bootstrap`, before any other checkout import: when the marker's
owner is gone and HEAD is still the pre-pull commit, every path that differs
between the two commits goes back to HEAD (the tree the venv was built for),
the dead git's index.lock is cleared, and the command re-execs itself (modules
already imported, main.py included, may be the half-written ones), so a rerun
`hermes update` repairs the install and then updates normally. Paths the update does not change keep local edits; the
updater's autostash stays in `git stash list`. Our own pid is never treated as
a live owner (containers hand a retry the killed updater's pid).
2026-09-23 17:42:24 -07:00
teknium1
4cbf862abe chore(memory): remove the bundled hindsight provider (moved to the plugin catalog)
Hindsight now ships from its maintainer's repo (vectorize-io/hindsight,
hindsight-integrations/hermes) through plugin-catalog/hindsight.yaml, so the
in-tree copy under plugins/memory/hindsight goes away. Homes that still name
`memory.provider: hindsight` are migrated by hermes_cli/memory_provider_migration.py
(the `hermes update` hook and the first agent start install the catalog plugin);
that path is untouched here.

Core-side special cases that only made sense with the bundled copy go with it:
the `memory.hindsight` LAZY_DEPS feature, `_provider_pip_dependencies`'s
`hindsight-all` expansion in `hermes memory setup` (the plugin's own
`_ensure_local_runtime()` installs the embedded runtime now), the compat-manifest
pointers of the deleted module, and prose that listed hindsight among the in-tree
providers. Generic provider-name lists, the HINDSIGHT_* env hints and the API-key
redaction pattern stay — a catalog-installed hindsight still uses them.

Revert this commit alone to restore the bundled provider.
2026-09-23 01:16:41 -07:00
ethernet
a60bc9cc64 fix(update): completion tail skips the fleet restart it does not owe
Every update route now finishes through update_completion._complete_selected,
which restarted the whole fleet unconditionally -- including the "Already up to
date" route that main sent through the pending-restart catch-up. Net effect:
each cron tick and each profile's `hermes update` drained and re-killed the one
multiplexed gateway.

Port the catch-up path's two live guards into the completion tail:
- host_restart_already_completed(checkout sha): a sibling profile attaches to
  the restart this host already stamped (#95294); the restart phase now stamps
  it via mark_host_restart_completed.
- every planned runtime AND every live fleet row current at the checkout sha
  (#117051, d6b0d37ece). Both are required: the live matrix lists gateways
  only, so a planned serve still on pre-update code keeps the restart.

The already-current route also arms the host obligation with the checkout sha;
an SHA-less arm replaced the standing record and wiped the restarted proof.
2026-09-21 19:03:35 -04:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
Tyler Lyon
18f4b29daf fix(cli): symmetric Windows resume-failure handling + atexit disarm
Two distinct defects in the Windows gateway pause/resume path, one
symptom: "hermes update" aborting exit 1 on a checkout that was
already current and needed no code change.

1. Asymmetric error handling between the two finish paths. The pull
   path wraps the Windows gateway resume in
   _resume_windows_gateways_and_merge_outcome, whose contract is
   explicit: "Must never abort the update" -- it catches the failure,
   marks the outcome incomplete, warns, and continues (receipt status
   "partial", exit 1 only via the existing incomplete-repair gate).
   _finish_already_up_to_date ("Already up to date" path) called
   _resume_windows_gateways_after_update bare, so the identical
   RuntimeError -- e.g. the relaunch-verification race with a parent
   Job Object kill, #48820 -- killed the run as an unhandled exception
   instead. Fixed by routing through the same merge helper and folding
   a failed resume into current_checkout_complete, landing on the
   pre-existing "partial" finalize + sys.exit(1) contract the repair
   path already uses for other incomplete states.

2. The atexit callback could replay the same failure a second time.
   Every foreground call site registers
   _resume_windows_gateways_after_update via atexit.register as a
   dead-process safety net, but the function only cleared
   token["resume_needed"] on its OWN success path -- a raise from
   _verify_relaunched_gateways_alive left the registration live, so
   interpreter teardown called the function again with the same
   still-armed token, repeated the whole relaunch + verify, failed
   identically, and printed "Exception ignored in atexit callback".
   Fixed by unregistering the function from atexit as soon as
   execution reaches it (foreground or the atexit fallback itself) --
   ownership transfers to whichever call site got there first, and a
   failure below is reported exactly once. atexit.unregister is a
   no-op when the function was never registered, so a non-Windows or
   no-resume-needed call is unaffected.

Both fixes are surgical: they leave the "a failed relaunch keeps
resume_needed=True so the token is retryable" contract completely
untouched (see the existing
test_resume_windows_gateway_service_failure_stays_retryable /
test_resume_windows_gateway_launcher_refresh_failure_stays_retryable
tests, which this PR does not modify) -- that design is intentional,
not the bug.

Closes #115563.

Testing: Windows-specific behavior verified by code trace + full
resolution-chain tests against the real functions (mocked I/O, real
control flow) on Linux/macOS, per the "Don't fake the host OS" rule
this repo's own AGENTS.md sets -- no sys.platform patching, and the
Windows-only branch (_is_windows() == True) is exercised exactly as
the existing test suite for this module already does (monkeypatching
_is_windows on the shared test fixtures, not faking the platform).
True live-Windows process-topology proof is out of scope for this PR
(the wine2e lane is reserved for process-topology E2E per AGENTS.md);
this fix is control-flow/bookkeeping, not host-specific syscalls.

- 2 new regression tests, both proven red on the pre-fix code via
  git stash (both fail with the exact defect signature: the "Already
  up to date" path re-raises instead of demoting to partial; the
  atexit unregister call never happens) and green on the fix
- scripts/run_tests.sh across the full affected surface (Windows
  update/resume/reconciliation + related pause/resume/venv-repair
  test files): 228 passed, 0 failed, 4 skipped (pre-existing,
  unrelated to this change)
2026-09-20 10:40:12 -07:00
fangliquan
26e85cfea2 fix(update): pin handoff module before checkout swap 2026-09-20 10:29:35 -07:00
teknium1
44945d224c fix(update): "Already up to date" re-syncs a venv installed from an older release
Closes the second #97208 atom (qingchuan-x): after a Windows hand-off child refused the
dependency sync (Desktop backend held the venv), the next `hermes update` found git current,
passed the core-import probe — the old release imports fine — and printed
"✓ Already up to date!" while the installed distribution still read hermes-agent 0.20.6
against a 0.21.2 checkout with lagging pins.

`_venv_dependency_set_stale` asks the venv's own interpreter for the installed hermes-agent
version and compares it with pyproject's; a mismatch on the current-checkout path takes
the same repair route as an unhealthy venv. Unknown states (no venv, not installed as a
distribution, probe failure) are never stale, so dev checkouts are untouched.
2026-09-19 02:02:14 -07:00
teknium1
6f0dae8b6f fix(update): shim hand-off child outwaits the launcher pid before the update lock and pins the wiring
Review follow-up on the #101600 fix:

- The wait moves from `_cmd_update_impl` to `cmd_update`, BEFORE `UpdateLock.acquire()`.
  The child used to run under the parent's marker (handoff pid / ancestor claim); the
  parent's exit — which the child now deliberately waits for — released that marker, so
  the whole Windows dependency install and gateway resume ran with no update lock. Waiting
  first lets the child claim a marker of its own for the tail.
- `SHIM_PARENT_PID_ENV` names the `hermes.exe` launcher ancestor (the process that holds
  the shim image open and exits only after reaping this interpreter) when psutil sees one,
  else this pid. `_windows_shim_in_process_chain` is split into `_venv_shim_matcher` +
  `_windows_shim_ancestor` so the holder pid shares the matcher; docs sentence adjusted.
- The direct-call wait test is replaced by one driving `cmd_update` in a real child with an
  older parent: proves the phase starts after the parent died, the child owns the lock, the
  handed-off pause token is adopted (no re-discovery) and the env is consumed. Removing the
  wait call, the token adoption, or moving the wait back after the lock each turns it red.
2026-09-19 02:01:36 -07:00
teknium1
168012585e fix(update): Windows shim hand-off child waits for hermes.exe to exit and owns the gateway resume
`hermes update` run from `venv\Scripts\hermes.exe` hands the dependency install to a
child under the venv Python because the shim cannot replace itself. Two races were left
in that hand-off (#101600, three independent Windows reproductions):

- The child never waited for the shim-run parent. The shim quarantine is a single
  `os.rename` with sub-second retries; when the child reached it while the parent was
  still alive the rename failed with PermissionError, `ShimQuarantineError` deferred the
  whole install and the receipt was marked failed although the checkout had advanced.
- On the legacy re-exec (`_abort_dependency_sync_if_self_locked` ->
  `_reexec_dependency_sync_off_windows_shim`, still live for a current checkout with an
  unhealthy venv -- i.e. the retry after any failed update) the PARENT resumed the paused
  gateways after spawning the child, holding hermes.exe open for the whole relaunch and
  raising `RuntimeError: Windows gateway relaunch after update was not verified alive`;
  the child then re-ran pause discovery, found that freshly relaunched gateway before its
  pid file existed and force-killed it as "without profile mapping".

Now every child spawned off the shim gets `HERMES_UPDATE_SHIM_PARENT_PID` and, at the top
of `_cmd_update_impl`, waits (bounded, 30 s, `(pid, create_time)` identity via psutil)
for that process to exit before scanning holders, pausing gateways or renaming shims. The
legacy re-exec carries the Windows pause token to the child in
`HERMES_UPDATE_GATEWAY_RESUME` and disarms the parent's copy, so the parent exits at once
and the child resumes exactly the fleet the parent stopped instead of rediscovering it --
the same ownership transfer the post-swap hand-off already does (94ced1a2b2).

The wait is host-independent and proven with real processes on Linux; the shim paths run
on the windows-latest lane through the existing `_is_windows`-patched suite.

Fixes #101600
Refs #98010 #103031
Supersedes #101630

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-19 02:01:36 -07:00
ethernet
6a5a6a05d2 refactor(pm): rename the pm.ensure submodule to pm.install; lazy_deps back to the shim
`import pm.ensure` bound the submodule onto the package, shadowing the facade's
`pm.ensure()` function for every later caller in the process (photon's sidecar
start hit `'module' object is not callable`). The module is pm.install now; the
function keeps its name. The facade resolves through `__import__` rather than
`importlib.import_module` so a test that patches import_module globally does not
break attribute access on pm.

tools/lazy_deps.py returns to the 16-line stop_for_relaunch shim the branch wrote
(an origin/main merge had replaced it with main's 775-line implementation); the
project-metadata tests follow. update_cmd re-exports the four old_updater_deps
names the shim tests resolve through hermes_cli.update_cmd.
2026-09-18 23:27:05 -04:00
teknium1
6c16233b0a fix(update): refresh memory-provider deps in the same order as the pull path
Both repair paths now run the memory-provider refresh between the tool-dep
restore and the plugin-dep reapply, exactly like _sync_python_dependencies_after_pull
and the ZIP path, so the three routes stay interchangeable.
2026-09-18 19:57:08 -07:00