Commit Graph

226 Commits

Author SHA1 Message Date
Tyler Lyon
18f4b29daf fix(cli): symmetric Windows resume-failure handling + atexit disarm
Two distinct defects in the Windows gateway pause/resume path, one
symptom: "hermes update" aborting exit 1 on a checkout that was
already current and needed no code change.

1. Asymmetric error handling between the two finish paths. The pull
   path wraps the Windows gateway resume in
   _resume_windows_gateways_and_merge_outcome, whose contract is
   explicit: "Must never abort the update" -- it catches the failure,
   marks the outcome incomplete, warns, and continues (receipt status
   "partial", exit 1 only via the existing incomplete-repair gate).
   _finish_already_up_to_date ("Already up to date" path) called
   _resume_windows_gateways_after_update bare, so the identical
   RuntimeError -- e.g. the relaunch-verification race with a parent
   Job Object kill, #48820 -- killed the run as an unhandled exception
   instead. Fixed by routing through the same merge helper and folding
   a failed resume into current_checkout_complete, landing on the
   pre-existing "partial" finalize + sys.exit(1) contract the repair
   path already uses for other incomplete states.

2. The atexit callback could replay the same failure a second time.
   Every foreground call site registers
   _resume_windows_gateways_after_update via atexit.register as a
   dead-process safety net, but the function only cleared
   token["resume_needed"] on its OWN success path -- a raise from
   _verify_relaunched_gateways_alive left the registration live, so
   interpreter teardown called the function again with the same
   still-armed token, repeated the whole relaunch + verify, failed
   identically, and printed "Exception ignored in atexit callback".
   Fixed by unregistering the function from atexit as soon as
   execution reaches it (foreground or the atexit fallback itself) --
   ownership transfers to whichever call site got there first, and a
   failure below is reported exactly once. atexit.unregister is a
   no-op when the function was never registered, so a non-Windows or
   no-resume-needed call is unaffected.

Both fixes are surgical: they leave the "a failed relaunch keeps
resume_needed=True so the token is retryable" contract completely
untouched (see the existing
test_resume_windows_gateway_service_failure_stays_retryable /
test_resume_windows_gateway_launcher_refresh_failure_stays_retryable
tests, which this PR does not modify) -- that design is intentional,
not the bug.

Closes #115563.

Testing: Windows-specific behavior verified by code trace + full
resolution-chain tests against the real functions (mocked I/O, real
control flow) on Linux/macOS, per the "Don't fake the host OS" rule
this repo's own AGENTS.md sets -- no sys.platform patching, and the
Windows-only branch (_is_windows() == True) is exercised exactly as
the existing test suite for this module already does (monkeypatching
_is_windows on the shared test fixtures, not faking the platform).
True live-Windows process-topology proof is out of scope for this PR
(the wine2e lane is reserved for process-topology E2E per AGENTS.md);
this fix is control-flow/bookkeeping, not host-specific syscalls.

- 2 new regression tests, both proven red on the pre-fix code via
  git stash (both fail with the exact defect signature: the "Already
  up to date" path re-raises instead of demoting to partial; the
  atexit unregister call never happens) and green on the fix
- scripts/run_tests.sh across the full affected surface (Windows
  update/resume/reconciliation + related pause/resume/venv-repair
  test files): 228 passed, 0 failed, 4 skipped (pre-existing,
  unrelated to this change)
2026-09-20 10:40:12 -07:00
fangliquan
26e85cfea2 fix(update): pin handoff module before checkout swap 2026-09-20 10:29:35 -07:00
teknium1
44945d224c fix(update): "Already up to date" re-syncs a venv installed from an older release
Closes the second #97208 atom (qingchuan-x): after a Windows hand-off child refused the
dependency sync (Desktop backend held the venv), the next `hermes update` found git current,
passed the core-import probe — the old release imports fine — and printed
"✓ Already up to date!" while the installed distribution still read hermes-agent 0.20.6
against a 0.21.2 checkout with lagging pins.

`_venv_dependency_set_stale` asks the venv's own interpreter for the installed hermes-agent
version and compares it with pyproject's; a mismatch on the current-checkout path takes
the same repair route as an unhealthy venv. Unknown states (no venv, not installed as a
distribution, probe failure) are never stale, so dev checkouts are untouched.
2026-09-19 02:02:14 -07:00
teknium1
6f0dae8b6f fix(update): shim hand-off child outwaits the launcher pid before the update lock and pins the wiring
Review follow-up on the #101600 fix:

- The wait moves from `_cmd_update_impl` to `cmd_update`, BEFORE `UpdateLock.acquire()`.
  The child used to run under the parent's marker (handoff pid / ancestor claim); the
  parent's exit — which the child now deliberately waits for — released that marker, so
  the whole Windows dependency install and gateway resume ran with no update lock. Waiting
  first lets the child claim a marker of its own for the tail.
- `SHIM_PARENT_PID_ENV` names the `hermes.exe` launcher ancestor (the process that holds
  the shim image open and exits only after reaping this interpreter) when psutil sees one,
  else this pid. `_windows_shim_in_process_chain` is split into `_venv_shim_matcher` +
  `_windows_shim_ancestor` so the holder pid shares the matcher; docs sentence adjusted.
- The direct-call wait test is replaced by one driving `cmd_update` in a real child with an
  older parent: proves the phase starts after the parent died, the child owns the lock, the
  handed-off pause token is adopted (no re-discovery) and the env is consumed. Removing the
  wait call, the token adoption, or moving the wait back after the lock each turns it red.
2026-09-19 02:01:36 -07:00
teknium1
168012585e fix(update): Windows shim hand-off child waits for hermes.exe to exit and owns the gateway resume
`hermes update` run from `venv\Scripts\hermes.exe` hands the dependency install to a
child under the venv Python because the shim cannot replace itself. Two races were left
in that hand-off (#101600, three independent Windows reproductions):

- The child never waited for the shim-run parent. The shim quarantine is a single
  `os.rename` with sub-second retries; when the child reached it while the parent was
  still alive the rename failed with PermissionError, `ShimQuarantineError` deferred the
  whole install and the receipt was marked failed although the checkout had advanced.
- On the legacy re-exec (`_abort_dependency_sync_if_self_locked` ->
  `_reexec_dependency_sync_off_windows_shim`, still live for a current checkout with an
  unhealthy venv -- i.e. the retry after any failed update) the PARENT resumed the paused
  gateways after spawning the child, holding hermes.exe open for the whole relaunch and
  raising `RuntimeError: Windows gateway relaunch after update was not verified alive`;
  the child then re-ran pause discovery, found that freshly relaunched gateway before its
  pid file existed and force-killed it as "without profile mapping".

Now every child spawned off the shim gets `HERMES_UPDATE_SHIM_PARENT_PID` and, at the top
of `_cmd_update_impl`, waits (bounded, 30 s, `(pid, create_time)` identity via psutil)
for that process to exit before scanning holders, pausing gateways or renaming shims. The
legacy re-exec carries the Windows pause token to the child in
`HERMES_UPDATE_GATEWAY_RESUME` and disarms the parent's copy, so the parent exits at once
and the child resumes exactly the fleet the parent stopped instead of rediscovering it --
the same ownership transfer the post-swap hand-off already does (94ced1a2b2).

The wait is host-independent and proven with real processes on Linux; the shim paths run
on the windows-latest lane through the existing `_is_windows`-patched suite.

Fixes #101600
Refs #98010 #103031
Supersedes #101630

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-19 02:01:36 -07:00
teknium1
6c16233b0a fix(update): refresh memory-provider deps in the same order as the pull path
Both repair paths now run the memory-provider refresh between the tool-dep
restore and the plugin-dep reapply, exactly like _sync_python_dependencies_after_pull
and the ZIP path, so the three routes stay interchangeable.
2026-09-18 19:57:08 -07:00
Yagna Vudathu
1f6e737ecb fix(cli): heal memory-provider bridge packages on every venv repair path
The venv-repair and runtime-repair paths reinstalled core, lazy and tool
deps but never refreshed memory-provider bridge packages, unlike the
git-pull and ZIP paths. Both repair paths now heal them last.

Fixes #113741.
2026-09-18 19:57:08 -07:00
teknium1
be98c68fa3 fix: call _resolve_pre_update_backup_mode directly in the receipt classifier
The resolver only does getattr() on args and already swallows its own
config-load errors, so the try/except around it could never fire; drop it
(review follow-up, no behaviour change).
2026-09-18 10:15:11 -07:00
ahrazzle
e276fc2157 fix(update): record a disabled pre-update backup as a skip, not a failure
cmd_update classified the pre-update backup with a single predicate
(`pre_update_snapshot_id is not None`) and the detail string
"disabled or failed", which conflates two different outcomes: a deliberate
opt-out (config `updates.pre_update_backup: off`/`false`, or `--no-backup`)
and a backup that was requested but captured nothing. A receipt whose safety
net was switched off therefore read exactly like one whose backup broke.

The outcomes are now distinct: an opt-out is recorded through
update_receipt.record_skip with the reason that caused it (`--no-backup` vs
the config value), and only a requested-but-empty backup is a failed step.
The `skips` channel already carried a `reason` field and was always empty in
practice, so no new receipt surface is introduced.

Concretely: cli-config.yaml.example still ships
`updates.pre_update_backup: false` (#94944), which resolves to mode "off", so
fresh installs run with no pre-update snapshot at all. On such a machine
every update recorded `pre_update_backup: ok=false, "disabled or failed"` and
the receipt gave no way to tell that nothing had failed and nothing had been
attempted -- diagnosing it required reading _resolve_pre_update_backup_mode
by hand.

Tests: tests/hermes_cli/test_update_receipt.py::TestPreUpdateBackupStep
2026-09-18 10:15:11 -07:00
yoniebans
c0aa3ce354 fix(update): keep manual serve restart obligations visible through gateway recovery
Manual serve/dashboard runtimes with no restart mechanism kept re-arming fleet_restart_pending: every CLI start warned about a completed update. This keeps the false alarm out while never hiding a real obligation: reminders are keyed by (pid, create_time), survive receipt rotation and gateway recovery, and clear only when the incarnation is provably gone or explicitly handed off. The pending marker owns its runtime inventory from the moment it is written; settlement derives coverage only from that inventory, so an older receipt can never discharge a newer update's obligation. Markers without an inventory stay pending. Startup keeps the completed-restart evidence path: a receipt whose restart phase finished vouches for the fleet at its post-update SHA.

Reported-by: adamkrawczyk
Diagnosis credit: KoNit-K (#107237)
Mechanism findings: g3org3yo, kokhlo, fmercurio
Reported-by: andrexibiza (P1 generation ownership)
2026-09-18 17:09:28 +02:00
teknium1
96e8a23222 feat(plugins): install declared Python dependencies and re-apply them after hermes update
A directory plugin's pyproject.toml [project].dependencies (or manifest
python_dependencies) are now installed into the Hermes venv on install/enable/
update, resolved under a constraints file built from Hermes' own pinned
dependencies together with every enabled peer's declarations. A candidate that
cannot resolve is refused before its tree is moved into place; nothing is
installed and no other plugin is touched. --no-deps opts a single install out.

hermes update rebuilds the venv from Hermes' lock and strips everything else, so
its git, zip and both repair paths now re-apply the union of every profile's
enabled plugins. When the union no longer resolves, each plugin is resolved on
its own so only culprits are dropped (never an alphabetical neighbour); a
plugin-vs-plugin conflict peels non-memory plugins first because a Hermes that
boots without memory reads as data loss. Dropped plugins are disabled through
the real config writer with a loud message naming the fix.

python_runtime: external lets sidecar-venv plugins (Mnemosyne's shape) opt out
of the union; hermes-agent self-dependencies and direct-URL requirements are
never installed (the latter are surfaced for the user to install by hand).
2026-09-17 00:11:27 -07:00
teknium1
bec8924594 fix(update): hand-off carries sibling snapshots + Windows pause token on both paths; child pinned to the parent's home
Independent review of the post-swap hand-off found state that did not cross or was
handled by the wrong side:

- update_cmd_config._LAST_SIBLING_SNAPSHOTS is rebound by the pre-update backup in the
  parent; the child read an empty dict and the sibling-profile cron/model-settings safety
  nets silently did nothing. It now rides in the payload.
- The ZIP path handed off without the Windows pause token, and the parent's finally/atexit
  resumed the paused gateways while the child was reinstalling the venv. The token now
  crosses on both paths; the parent flips resume_needed off only after a child actually ran.
- The child inherits HERMES_HOME but re-read the sticky active_profile, finishing a
  `-p default` update inside another profile's home. _apply_profile_override keeps the
  parent's resolved home for the post-swap child.
- The child env now also sets HERMES_UPDATE_REEXEC so cmd_update's Windows hard-exit tail
  covers it and the shim re-exec can never fire a second time from inside it; the
  hand-off pid is only claimed when no upstream updater (Tauri/Electron) already named
  one.
- Popen + wait instead of subprocess.run: Ctrl-C no longer kills the child (which owns the
  receipt and the gateway resume) 250 ms after the interrupt.
- A child that cannot start: the parent takes the receipt back, records the failed step,
  writes .update-incomplete and the gateway exit code, exits 1.
- The child consumes the hand-off file and builds git_cmd without re-running the checkout
  preflight (no second fork banner).
2026-09-17 00:02:09 -07:00
teknium1
94ced1a2b2 fix(update): finish hermes update in an interpreter born on the pulled code
`hermes update` started in an interpreter that had imported the PRE-pull tree, then
kept running every post-swap phase (dependency sync, Node/web/Desktop builds,
maintenance, config migration, fleet restart, verification, receipt) in that same
process, lazily importing NEW source into an OLD `sys.modules` graph. Any rename
between the two commits surfaced as an ImportError/AttributeError inside the updater
after the code swap had already succeeded (#87134, #111271, #112465, #112558, #112604).
Each incident added another purge, reload list or per-step isolation, and each moved
the crash to the next module nobody had listed.

The pre-pull process now stops at the swap: it writes the open receipt, the pre-update
fleet plan, the pre-update version/active features and the Windows pause token to a
hand-off file and re-executes `hermes update <same flags> --post-swap <file>` under the
venv interpreter. The child imports exclusively from the pulled tree, resumes the
receipt and owns the rest of the run; the parent relays its exit code. Git and ZIP paths
both hand off. On Windows, when the updater runs from `hermes.exe`, the child is spawned
detached exactly as the shim hand-off already did (the shim cannot be awaited while the
sync must replace it).

With no pulled code ever executing in a pre-pull interpreter, the stale-module layer is
dead and removed: `_purge_stale_hermes_modules`, `_stale_purge_prefixes`,
`_evict_module`, `_STALE_PURGE_*`, `_reload_updated_runtime_modules`,
`_reload_process_scan_modules`, `_reload_config_modules`, `_UPDATE_RUNTIME_RELOAD_MODULES`
and their tests. `_run_config_check_fresh` / `_run_migrate_config_fresh` keep their names
and simply call the config API.

Tests: the hand-off boundary (child argv/env, detached receipt + plan in the payload,
exit-code relay) and the child side (receipt resumed with its history, plan rebuilt,
pre-update snapshots taken from the payload). Mocked updater flows run the tail
in-process through the same payload round-trip (autouse fixture; opt out with
`@pytest.mark.real_post_swap_handoff`).
2026-09-17 00:02:09 -07:00
teknium1
bafbf88032 fix(update): bracket the post-repair lazy restore with the lazy-refresh marker
The healthy-branch restore after a runtime (SQLite) repair called
_refresh_active_lazy_features as a bare statement and dropped its bool, so a
failed restore into the freshly swapped venv still ended in '✓ Already up to
date!' and a rerun could not recover (the next snapshot sees the stripped venv).
Mirror the pull path (update_cmd_deps): write the lazy-refresh-incomplete marker
before the restore, clear it only when the refresh reports success, otherwise
print the same 'recovery incomplete' hint so the next `hermes` run finishes it.

Test: one invariant case (refresh returns False -> marker written, not cleared);
red on the previous head, green here. The marker helpers are stubbed so the test
never writes into the real home.
2026-09-16 17:45:29 -07:00
teknium1
56b9b19a50 refactor(update): fold the post-repair restore into the healthy branch; trim tests
The salvaged fix added a third `elif runtime_repaired` arm that duplicated the
whole `_repair_node_deps_on_current_checkout(...)` call from the healthy `else`.
Guard only the restore step inside that branch instead, so there is one
node-deps call and the healthy path stays byte-identical when no repair ran.
`ensure_uv()` already returns the managed uv (or its bootstrap), so the
`updated_uv or ensured_uv` fallback is unnecessary; use the ensured path like
`_repair_venv_on_current_checkout` does.

Tests: the contributor's test was inserted above the class docstring of
`TestCmdUpdateBranchFallback`, orphaning that docstring. Move it into its own
class with a shared fixture and add the control case (healthy venv, no repair
-> no restore calls). A/B: with origin/main's update_cmd.py the restore test
fails and the control passes; on this head both pass.
2026-09-16 17:45:29 -07:00
KoNit-K
2b76a63578 fix(updater): restore optional deps after sqlite repair 2026-09-16 17:45:29 -07:00
teknium1
659f5ae94f fix(update): keep .venv installs whole through ZIP fallback and venv repair
Follow-up to the cherry-picked #112966 so the uv-default `.venv` layout is
supported end to end, not only at the lookup sites:

- `_ZIP_PRESERVED_TOP_LEVEL` gains `.venv`. The dirty-tree guard runs
  `git status --ignored=matching`, so a gitignored `.venv/` surfaced as
  `!! .venv/` and refused every ZIP fallback on such installs ("the working
  tree has uncommitted changes or untracked files") — the live runtime was
  being treated as user data the overlay would destroy.
- `_repair_venv_on_current_checkout` recreates the venv at the resolved
  directory instead of a literal `venv`, so a broken `.venv` is rebuilt in
  place rather than growing a second environment that `project_venv_dir()`
  then prefers while `bin/hermes.cmd` still launches the old one.
- `_refuse_update_if_venv_foreign_owned` scans the resolved venv (the only
  remaining `PROJECT_ROOT / "venv"` literal on the update path).
- windows.ps1 names the actual shim path in the lock-timeout message.
- Tests: extend the real-git ZIP guard test with the `.venv` case (red
  before this commit); the holder-guard test now uses a kernel-runner child
  whose cmdline lacks `hermes_cli.main`, so only the venv-prefix arm can match
  it (red on origin/main); drop the `process.platform`-override vitest case,
  which exercised the same resolver as the `.venv` case with a different
  directory string.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-16 17:13:56 -07:00
KoNit-K
b6cc751e5f fix(update): support uv default venv in desktop updates 2026-09-16 17:13:56 -07:00
teknium1
2246c245f5 refactor(update): single defer flag for the deferred catch-up; trim tests; document the flag
Salvage of #93649 (@TurgutKural). _apply_pending_fleet_restart_catchup took two
booleans (respect_no_gateway_restart + no_gateway_restart) that were only ever
true together; one `defer` keyword says the same thing. The 13 tests are cut to
the two invariants (pulled path skips restart + verify and keeps the marker;
already-current path defers the catch-up). The user guide gains a section on
running `hermes update` from inside the gateway.
2026-09-15 19:28:39 -07:00
Turgut Kural
7b27ea3639 fix(cli): allow hermes update without gateway restart for cron (rebased on upstream/main)
(cherry picked from commit e70f78e54a96f2e8037f7e385fc31563bbeaf392)
(cherry picked from commit 225f56ab29e91977f1fef2e743c50dd6e2e89b60)
2026-09-15 19:28:39 -07:00
kshitijk4poor
d6ea7aa002 refactor(update): inline the tests exclusion; fix comments the wider purge made stale
`_STALE_PURGE_EXCLUDED_TOP_LEVEL` had one reader and a re-export nothing patched;
the config-check comment claimed root modules stay cached across the purge (no
longer true); the purge test module docstring still described package prefixes.
2026-09-15 21:57:59 +05:30
kshitijk4poor
9395f2f0d4 refactor(update): the purge scan has no fallback; trim to two invariants
The checkout root is where the update just pulled into, so "root unreadable" cannot
happen after a successful pull — drop the OSError fallback and the static five-name
tuple it fell back to (the tuple was the drift that caused the bug). Drop the phantom
`hermes_cli.hermes_logging` protection entry: no such module exists; the real
`hermes_logging` is root-level and now protected by name. Keep two tests: the stale
root `utils` scenario (red on base) and the hermes_logging protection.
2026-09-15 21:57:59 +05:30
Hubert Nimitanakit
29758a4eb2 fix(cli): purge every top-level checkout module, not just five packages
`_STALE_PURGE_PREFIXES` listed five package names, so every top-level module in
the checkout root survived the post-pull purge. `hermes update` then imported
new source against a cached pre-pull `utils`, `hermes_constants`, `plugins` or
`providers`.

Field failure (2026-09-12, macOS): updating 0.20.6 -> 0.21.2 crossed 3145986c20,
which added `base_url_origin` to `utils.py`. The restart phase's up-front
`from hermes_cli.gateway import ...` pulled `agent.auxiliary_client`, whose
`from utils import base_url_origin` hit the cached 0.20.6 `utils`:

    Update incomplete - gateway auto-restart failed: cannot import name
    'base_url_origin' from 'utils' (.../hermes-agent/utils.py)

The purge docstring already claims it evicts EVERY cached Hermes module; the
hardcoded tuple was the same "re-fixed per symptom" shape it replaced. Scan
`PROJECT_ROOT` instead: top-level `.py` files plus directories with an
`__init__.py`. 50 names here, and a newly added module can no longer drift out.

Two exclusions, both deliberate:

- `hermes_logging` joins `_STALE_PURGE_PROTECTED`. Its queue listener, handler
  list and `_logging_initialized` flag are module globals, so a fresh copy
  starts a second QueueListener over the same log files while the first runs.
- `tests` is never purged. pytest resolves fixtures through the identity of its
  already-imported test modules.

Falls back to the old tuple when the root is unreadable.
2026-09-15 21:57:59 +05:30
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
fangliquanflq
4e3165b75a fix(updater): finish Node phase after Windows handoff 2026-09-13 14:44:38 -07:00
bixycler
70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
joaomarcos
2fb87f047f fix(cli): repair shallow boundaries already dropped by stale-graft prune
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.

Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.

This complements PR #108290, which owns the prevention half.

Refs #108286
2026-09-12 15:00:37 +05:30
Teknium
9ce7547faf fix(update): bound network git in hermes update and prune shallow grafts on apply
- _git_run(network=True) now carries a 300s timeout; a dead-stalled fetch
  (HTTP/2 to GitHub on some networks, black-holed proxy) becomes a failed run
  whose stderr names the stall instead of an update pinned on
  'Fetching updates...' forever (#93759, #95777). Local git stays unbounded.
- The apply path prunes stale .git/shallow grafts alongside its existing
  lock/tmp_pack cleanup, so installs that already accumulated grafts from
  past depth-1 checks (#105951: 57 entries) heal on their next update, not
  only on --check.
2026-09-11 02:06:32 -07:00
liuhao1024
5bfa389e8a fix(cli): prune stale shallow grafts left by depth-1 update checks (#105951)
Every 'git fetch --depth 1' in 'hermes update --check' (and the past
banner passive checks, before #107648 moved them to the GitHub API)
appends the fetched tip to .git/shallow as a new graft and git never
removes the previous one, so a long-lived shallow installer checkout
accumulates one graft per check (57 observed). The stale grafts break
merge-base and push 'hermes update' into the orphan-divergence reset
path with a rescue ref on every run.

prune_stale_shallow_grafts() now runs after each successful depth-1
fetch in 'hermes update --check' and clears the grafts already
accumulated by past checks: it keeps only the boundaries still
protecting referenced tips (HEAD, FETCH_HEAD, every ref tip) and
atomically rewrites .git/shallow, restoring the original file if the
trimmed set breaks history walking. The dropped commits are already
unreachable; their objects are left for git gc.

Rebased onto main after #107648: the banner.py hook is dropped (the
passive check no longer git-fetches); the update --check prune and the
cleanup of already-accumulated grafts are kept.

(cherry picked from commit 6174837fc5b9f4cc3d4d46dc1b2d9a2f6b83c120)
2026-09-11 02:06:32 -07:00
Teknium
c0c6b31543 fix(update): stream build progress without concealing silent stalls
Retain partial-line output, UTF-8 decoding, failure output and cancellation cleanup. Based on streaming investigations by Artemonim (#101850) and lEWFkRAD (#104843); gateway tee adapted from fangliquanflq (#97402). Live Linux child/tee probe: withheld or dropped on base, visible in 0.02 seconds after. Campaign-locked tests and native Windows proof are pending.
2026-09-07 05:56:50 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
e311471040 refactor(hermes_cli): update_cmd — _apply_pulled_update/_repair_venv_on_current_checkout phase helpers, _base_git_cmd/_tip_shas/_current_branch_name/_pip_install_prefix/_finalize_receipt, early-return guards, print joins 2026-09-02 23:00:49 -07:00
Teknium
1767ee8119 refactor(hermes_cli): compact docstrings/comments in update_cmd*, update_abort_recovery, uninstall (keep every WHY) 2026-09-02 21:28:27 -07:00
Teknium
a8e8db4fc5 refactor(hermes_cli): AST-neutral string-literal joins and import-blank squeeze in update_cmd*/uninstall 2026-09-02 21:12:29 -07:00
Teknium
882b655f21 refactor(hermes_cli): dedupe update_cmd/update_cmd_deps — _updates_config, _print_update_check_result, unified uv/pip repair+sync branches, _git_run for raw git calls, _probe_failure/_clip/_npm_lock_cache_file 2026-09-02 20:55:00 -07:00
Teknium
adba16445a refactor(hermes_cli): AST-neutral layout compaction of update_cmd*/uninstall (hug brackets, pack args) 2026-09-02 20:27:53 -07:00
Teknium
d66a62469b refactor(update): lift parked-branch guard out of _prepare_checkout_for_update (161 -> 101 LOC) 2026-09-02 17:03:42 -07:00
Teknium
73ca3efea4 refactor(update): lift divergence reconcile and syntax-guard rollback out of _pull_updates (153 -> 47 LOC) 2026-09-02 17:02:23 -07:00
Teknium
2d7fcf636e refactor(update): compact re-export import blocks (AST-identical, -117 lines) 2026-09-02 17:00:07 -07:00
Teknium
60041b787a refactor(update): join short multi-line statements onto one line (AST-identical, -416 lines) 2026-09-02 16:52:38 -07:00
Teknium
1aa9285312 refactor(update): hand-compact comments/docstrings in update_cmd.py, update_cmd_fleet.py, update_cmd_zip.py (AST-identical); rewrite stale module docstring 2026-09-02 16:45:44 -07:00
Teknium
d46eb1f964 refactor(update): collapse 90 try/except-pass and try/except-logger.debug blocks into suppress()/_best_effort() (new leaf update_cmd_common.py) 2026-09-02 16:34:55 -07:00
Teknium
ebf2875049 refactor(update): lift already-up-to-date finish out of _cmd_update_impl (290 -> 257 LOC) 2026-09-02 16:31:17 -07:00
Teknium
bff23bea2c refactor(update): lift option resolution, receipt/plan, git prep, HEAD verification and CalledProcessError handling out of _cmd_update_impl (521 -> 290 LOC) 2026-09-02 16:29:04 -07:00
Teknium
a59680217f refactor(update): hand-compact comments/docstrings in split modules (AST-identical); re-export get_hermes_home/shutil; repoint source-inspection tests 2026-09-02 16:11:02 -07:00
Teknium
4674f8b254 refactor(update): split zip/stash/config/deps/git/maint clusters out of update_cmd.py 2026-09-02 16:00:26 -07:00
Teknium
096826bf7d refactor(update): split gateway fleet restart/verify into update_cmd_fleet.py 2026-09-02 15:55:50 -07:00
Teknium
5206a504ca refactor(update): split Windows gateway lifecycle into update_cmd_windows.py 2026-09-02 15:40:14 -07:00
Teknium
e2d7a6f8b8 refactor(update): drop dead _gateway_restart_recovery_profiles; write sibling-snapshot cache via module attribute 2026-09-02 15:27:33 -07:00