Commit Graph

201 Commits

Author SHA1 Message Date
bixycler
70d0f556d7 fix(branding): use the Caduceus ☤ (U+2624), not the Rod of Asclepius ⚕ (U+2625)
Every inline glyph — CLI banner/status bar/response labels/goodbye, setup
and doctor boxes, gateway update prompts, WhatsApp reply prefix, TUI theme,
locale strings and the docs — used ⚕, the staff of Asclepius (medicine).
Hermes carries the Caduceus ☤. The ASCII-art logo was already correct.

Mechanical swap across 60 files (no logic change); both glyphs are
East-Asian-width Neutral so no layout shifts. Skins that set their own
`response_label` / `goodbye` are unaffected.

Direction from PR #7064 (@bixycler), the earliest of #7064 / #9611 / #15574,
redone against current main.

Fixes #9565
2026-09-12 08:25:54 -07:00
joaomarcos
2fb87f047f fix(cli): repair shallow boundaries already dropped by stale-graft prune
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.

Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.

This complements PR #108290, which owns the prevention half.

Refs #108286
2026-09-12 15:00:37 +05:30
Teknium
9ce7547faf fix(update): bound network git in hermes update and prune shallow grafts on apply
- _git_run(network=True) now carries a 300s timeout; a dead-stalled fetch
  (HTTP/2 to GitHub on some networks, black-holed proxy) becomes a failed run
  whose stderr names the stall instead of an update pinned on
  'Fetching updates...' forever (#93759, #95777). Local git stays unbounded.
- The apply path prunes stale .git/shallow grafts alongside its existing
  lock/tmp_pack cleanup, so installs that already accumulated grafts from
  past depth-1 checks (#105951: 57 entries) heal on their next update, not
  only on --check.
2026-09-11 02:06:32 -07:00
liuhao1024
5bfa389e8a fix(cli): prune stale shallow grafts left by depth-1 update checks (#105951)
Every 'git fetch --depth 1' in 'hermes update --check' (and the past
banner passive checks, before #107648 moved them to the GitHub API)
appends the fetched tip to .git/shallow as a new graft and git never
removes the previous one, so a long-lived shallow installer checkout
accumulates one graft per check (57 observed). The stale grafts break
merge-base and push 'hermes update' into the orphan-divergence reset
path with a rescue ref on every run.

prune_stale_shallow_grafts() now runs after each successful depth-1
fetch in 'hermes update --check' and clears the grafts already
accumulated by past checks: it keeps only the boundaries still
protecting referenced tips (HEAD, FETCH_HEAD, every ref tip) and
atomically rewrites .git/shallow, restoring the original file if the
trimmed set breaks history walking. The dropped commits are already
unreachable; their objects are left for git gc.

Rebased onto main after #107648: the banner.py hook is dropped (the
passive check no longer git-fetches); the update --check prune and the
cleanup of already-accumulated grafts are kept.

(cherry picked from commit 6174837fc5b9f4cc3d4d46dc1b2d9a2f6b83c120)
2026-09-11 02:06:32 -07:00
Teknium
c0c6b31543 fix(update): stream build progress without concealing silent stalls
Retain partial-line output, UTF-8 decoding, failure output and cancellation cleanup. Based on streaming investigations by Artemonim (#101850) and lEWFkRAD (#104843); gateway tee adapted from fangliquanflq (#97402). Live Linux child/tee probe: withheld or dropped on base, visible in 0.02 seconds after. Campaign-locked tests and native Windows proof are pending.
2026-09-07 05:56:50 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
e311471040 refactor(hermes_cli): update_cmd — _apply_pulled_update/_repair_venv_on_current_checkout phase helpers, _base_git_cmd/_tip_shas/_current_branch_name/_pip_install_prefix/_finalize_receipt, early-return guards, print joins 2026-09-02 23:00:49 -07:00
Teknium
1767ee8119 refactor(hermes_cli): compact docstrings/comments in update_cmd*, update_abort_recovery, uninstall (keep every WHY) 2026-09-02 21:28:27 -07:00
Teknium
a8e8db4fc5 refactor(hermes_cli): AST-neutral string-literal joins and import-blank squeeze in update_cmd*/uninstall 2026-09-02 21:12:29 -07:00
Teknium
882b655f21 refactor(hermes_cli): dedupe update_cmd/update_cmd_deps — _updates_config, _print_update_check_result, unified uv/pip repair+sync branches, _git_run for raw git calls, _probe_failure/_clip/_npm_lock_cache_file 2026-09-02 20:55:00 -07:00
Teknium
adba16445a refactor(hermes_cli): AST-neutral layout compaction of update_cmd*/uninstall (hug brackets, pack args) 2026-09-02 20:27:53 -07:00
Teknium
d66a62469b refactor(update): lift parked-branch guard out of _prepare_checkout_for_update (161 -> 101 LOC) 2026-09-02 17:03:42 -07:00
Teknium
73ca3efea4 refactor(update): lift divergence reconcile and syntax-guard rollback out of _pull_updates (153 -> 47 LOC) 2026-09-02 17:02:23 -07:00
Teknium
2d7fcf636e refactor(update): compact re-export import blocks (AST-identical, -117 lines) 2026-09-02 17:00:07 -07:00
Teknium
60041b787a refactor(update): join short multi-line statements onto one line (AST-identical, -416 lines) 2026-09-02 16:52:38 -07:00
Teknium
1aa9285312 refactor(update): hand-compact comments/docstrings in update_cmd.py, update_cmd_fleet.py, update_cmd_zip.py (AST-identical); rewrite stale module docstring 2026-09-02 16:45:44 -07:00
Teknium
d46eb1f964 refactor(update): collapse 90 try/except-pass and try/except-logger.debug blocks into suppress()/_best_effort() (new leaf update_cmd_common.py) 2026-09-02 16:34:55 -07:00
Teknium
ebf2875049 refactor(update): lift already-up-to-date finish out of _cmd_update_impl (290 -> 257 LOC) 2026-09-02 16:31:17 -07:00
Teknium
bff23bea2c refactor(update): lift option resolution, receipt/plan, git prep, HEAD verification and CalledProcessError handling out of _cmd_update_impl (521 -> 290 LOC) 2026-09-02 16:29:04 -07:00
Teknium
a59680217f refactor(update): hand-compact comments/docstrings in split modules (AST-identical); re-export get_hermes_home/shutil; repoint source-inspection tests 2026-09-02 16:11:02 -07:00
Teknium
4674f8b254 refactor(update): split zip/stash/config/deps/git/maint clusters out of update_cmd.py 2026-09-02 16:00:26 -07:00
Teknium
096826bf7d refactor(update): split gateway fleet restart/verify into update_cmd_fleet.py 2026-09-02 15:55:50 -07:00
Teknium
5206a504ca refactor(update): split Windows gateway lifecycle into update_cmd_windows.py 2026-09-02 15:40:14 -07:00
Teknium
e2d7a6f8b8 refactor(update): drop dead _gateway_restart_recovery_profiles; write sibling-snapshot cache via module attribute 2026-09-02 15:27:33 -07:00
Teknium
401efb0b2b refactor(cli): decompose _cmd_update_impl into phase helpers, unify duplicated update helpers, compact comments
hermes_cli/update_cmd.py 11217 -> 10204 LOC; _cmd_update_impl 2797 -> 521 LOC.

Decomposition (behavior-neutral, AST free-name verified) — _cmd_update_impl is
now a thin orchestrator calling, in order:
  _clear_windows_venv_holders_or_exit -> _prepare_checkout_for_update
  (-> _CheckoutPlan) -> _repair_current_checkout | _pull_updates ->
  _sync_python_dependencies_after_pull -> _run_post_update_maintenance ->
  _restart_gateway_fleet_after_update (-> _GatewayRestartOutcome;
  _restart_systemd_gateway_units) -> _resume_windows_gateways_and_merge_outcome
  -> _verify_fleet_after_update. _resolve_manage_cmd hoisted to module level.

Unified helpers: _git_run (49 captured git subprocess.run sites),
_systemctl / _systemctl_reset_and_restart (17 sites + 2 pairs),
_sweep_bytecode_after_update (3x), _print_bundled_skills_sync_report (ZIP+git),
_ensure_venv_pip (ZIP+git), _self_and_non_gateway_ancestor_pids (2x),
_record_update_step (4x), _write_gateway_update_exit_code for the 3 inline
".update_exit_code" writes. Dropped duplicate nested copies of
_wait_for_service_active/_service_restart_sec/_print_items, dead
upstream_exists, a dead if/pass branch and unused imports.

Comments/docstrings hand-compacted (236 blocks, AST-identical with docstrings
normalized); rationale/invariant sentences kept.

Tests: source-inspection guards repointed to the helper that now owns the code
(test_update_self_lock, test_update_fleet_check_fail_closed,
test_update_apply_shallow_count); new regression test drives the real Windows
resume/merge helper.
2026-09-02 13:30:54 -07:00
Teknium
10088c569e fix(update): hermes update no longer hangs on a GitHub username prompt
GitHub answers anonymous fetches with HTTP 401 during outages (and for
renamed/private repos). git then prompts `Username for 'https://github.com':`
on the inherited terminal and `hermes update` sits there — users read it as
Hermes demanding a GitHub login.

Every network git call in the updater (fetch/pull/push, apply + --check +
fork sync) now runs with GIT_TERMINAL_PROMPT=0 / stdin=DEVNULL, so the 401
fails fast into the fetch-failure classifier, which now reports it as a
GitHub-side rejection (likely outage) rather than blaming the user's
credentials. Credential helpers/askpass are left configured so private-fork
origins still authenticate.

Live repro: PTY-attached update --check against a 401 origin hung 15s+ on
the prompt before; exits rc=1 in 0.2s with the diagnosis after.

Same class as #73751 (@Frowtek, pre-main.py decomposition); passive banner
half salvaged from #101421 (@RobbertC5).
2026-09-02 12:34:26 -07:00
Teknium
f9bca5a0d0 fix(update): reconcile serve/dashboard runtimes in their own vocabulary and escalate survivors (#100479)
Widen the two salvaged fixes (#100490, #100493) to the whole class:

- match_runtime_outcomes: serve/dashboard rows never borrow gateway
  bookkeeping at ANY site — not just the bare hermes-gateway unit name
  (#100490) but also relaunched_profiles / externally_supervised_profiles
  and the profile-substring unit match (hermes-gateway-work credited the
  'work' serve). They reconcile against hermes-serve*/hermes-dashboard*
  units (exact names, scope prefix tolerated) or, when the caller passes
  the (pid, create_time) survivor probe result, by incarnation liveness.
- update_cmd success path: the survivor rows from #100493's new call now
  feed the Phase-2 reconciliation, so a surviving unmanaged serve is
  'unaccounted' -> exit 1 + 'partial' receipt, not warn-and-exit-0.
- report_unaccounted_runtimes: a serve/dashboard miss names the serve
  remedy instead of 'hermes gateway restart', which cannot reach it.

Tests: 6 reconciliation cases (sibling sites, unit vocabulary, exact-name
guard, incarnation probe, remedy text) + an end-to-end cmd_update case
asserting warn + unaccounted + exit 1 + receipt runtime_outcomes.
2026-09-02 00:42:53 -07:00
TwotNguyenVN
5733d55f77 fix(update): warn surviving pre-update serve and dashboard runtimes on success (#100479) 2026-09-02 00:42:53 -07:00
kshitijk4poor
d4611ac837 chore: review follow-ups for salvaged #99954
- restore the success-path debug log the old git-pull guard had
- drop the dead 'tag' test-helper param and unused snapshot return
- hoist the repeated get_hermes_home() call
2026-09-02 01:52:11 +05:30
Sahilvishnaliya
d0f0afb009 fix(update): post-update state.db guard covers every profile, not just the root
The #68474 post-update integrity guard verified only the root home's state.db, but the pre-update snapshot already covered every sibling profile (#66140 create_pre_update_snapshots_all_profiles). A profile database corrupted by the update was never detected and never auto-restored - that profile's sessions were silently gone while the update reported success (#97994).

Both guard sites (ZIP path and git-pull path) now route through a shared _verify_and_restore_state_dbs_post_update() that verifies the root DB plus every _sibling_profile_homes() DB, restoring each from its OWN most recent valid snapshot with per-profile operator-visible reporting. Refactors the two near-identical inline guards into one helper - behavior for the root DB is unchanged.

Tests: corrupt-sibling-with-snapshot gets restored while root stays untouched; valid-sibling not touched; corrupt-sibling-without-snapshot reported without raising. Fixes #97994.
2026-09-02 01:52:11 +05:30
kshitijk4poor
6b46725a17 fix(update): surface leftover update autostashes older than 7 days (#63717)
Parked (--keep-stash) and conflict-preserved autostash entries were never
mentioned again after the update run that created them — one persisted 9+
days unnoticed (#63717 problem 6). hermes update now lists
hermes-update-autostash-* entries older than 7 days at the start of the
git update path, with review/restore/drop guidance. Deliberately a warning,
not a GC: a stash entry can be the only copy of uncommitted work, so
nothing is ever dropped automatically.
2026-09-02 01:50:21 +05:30
Teknium
ab9866bc64 fix(gateway): survive Windows Job-Object teardown across gateway restarts (#48820)
Fourth reproduction on #48820: the updater's post-update resume respawned
the gateway through _spawn_gateway_restart_watcher, the process died within
seconds (parent Job Object denying CREATE_BREAKAWAY_FROM_JOB kills the
child on job teardown), and "✓ Restarting Windows gateway profile(s)" was
printed anyway — 12.5h of silent platform downtime, with zero trace because
the watcher respawned with stdout/stderr=DEVNULL.

Three surgical changes:

1. Watcher respawn stdio → logs/gateway-stdio.log (hermes_cli/gateway.py).
   The inlined watcher now routes the respawned gateway's stray
   stdout/stderr to the same sidecar log gateway_windows._spawn_detached
   uses (DEVNULL only as fallback), so a gateway killed moments after
   respawn leaves a trace. Direct implementation of the 4th repro's
   hardening suggestion (1).

2. Watcher respawn stamps _HERMES_GATEWAY_BREAKAWAY=1/0 exactly like the
   canonical _spawn_detached, so the respawned gateway's exit-diag /
   lifecycle records show whether it escaped the parent Job Object — a
   job-teardown kill is no longer indistinguishable from any other silent
   death.

3. Post-update resume verifies liveness before vouching
   (hermes_cli/update_cmd.py). _resume_windows_gateways_after_update now
   runs the same provisional-hit + 2s-confirmation liveness poll every
   other spawn path uses (gateway_windows._wait_for_gateway_ready, widened
   with all_profiles= for the fleet) before printing ✓, writes the #91675
   start attestation for the verified PIDs, and fails the resume with a
   "restart could not be verified" warning + recovery hint when no stable
   gateway appears. Suggestion (2) of the 4th repro; closes the last
   silent-success hole in the family (#84185 fixed the cold-start leg,
   #91675 the direct-start leg; this is the relaunch leg).

Live proof on windows-latest (wine2e lane): real kill-on-close Job Objects
confirm breakaway children survive teardown and non-breakaway children die
(the exact #48820 mechanism); the real watcher respawn cycle leaves the
stdio trace + breakaway stamp; and the resume path refuses to print ✓ for
a dead relaunch.

Fixes the Bug-1 relaunch-trust leg of #48820.
2026-09-01 11:43:36 -07:00
Teknium
e9541d213f fix(gateway): never print ✓ for a Windows gateway that dies after the liveness poll
The 6s post-spawn liveness poll (#86687) returned on the FIRST
process-table hit, so a gateway created and then killed moments later —
e.g. by the parent shell's Job Object teardown when
CREATE_BREAKAWAY_FROM_JOB is denied — still earned a "✓ Gateway started"
line (#91675 hole a). And no poll can ever observe a death that happens
AFTER the CLI process exits, which is exactly when the Job Object
teardown fires.

Two layers:

1. _wait_for_gateway_ready now treats the first hit as provisional: the
   gateway must stay visible through a 2s confirmation window
   (_confirm_gateway_stable) before it is reported ready; a death during
   confirmation resumes polling until the deadline. Failure output is an
   honest ✗ with the Job Object explanation and the schtasks /Run
   recovery command when a Scheduled Task exists.
2. Start attestation (report-async-death): every ✓ persists
   state/gateway.start-attestation.json with the vouched-for PIDs. The
   next `gateway start`/`gateway status` invocation checks it — if the
   attested PIDs are gone with no clean-exit record in the lifecycle
   ledger, the CLI reports (once) that the previous ✓ was false and
   prints the schtasks recovery hint. `gateway stop` and a clean
   lifecycle-ledger exit clear the marker silently.

Also: when _spawn_detached had to retry without
CREATE_BREAKAWAY_FROM_JOB, the ✓ now carries an explicit "could not
break away from this shell's Job Object" warning, and the post-update
cold-start ✓ (update_cmd) writes the same attestation marker.

Sub-symptom (b) of #91675 (post-update cold-start only resumes the
active profile) is handled separately by PR #99685.

Fixes #91675
2026-09-01 10:06:06 -07:00
Teknium
045865377c fix(update): restore user model settings config.yaml rewrites drop during update
Post-update safety net for the config half of #64160: Desktop update/repair
cycles have rewritten user-set model.provider/model.default/model.base_url/
model.api_key and dropped the moa: section entirely — settings the gateway
and unattended cron jobs also consume, so the rewrite silently redirects
paid inference machine-wide.

Mirrors the restore_cron_jobs_if_emptied pattern (#34600): compare the live
config.yaml against the pre-update quick snapshot taken minutes earlier by
the same update run and restore ONLY protected keys the user had set that
were changed or dropped — never the whole file, so version stamps and new
sections the migration legitimately wrote survive. Runs on every update
completion path via _check_and_apply_config_migration, plus the same net for
every sibling profile against its own same-generation snapshot (#66140
pattern). Exception-swallowing: a safety-net failure never breaks an
otherwise-good update.

6 new tests in tests/hermes_cli/test_backup.py cover restore-on-rewrite,
no-op-on-untouched, preservation of legitimate migration writes, no-op when
the user never set the protected keys, unreadable live config, and missing
snapshot id.

Fixes #64160 (config half; the active-profile half is the desktop migration
commits earlier on this branch).
2026-09-01 09:54:22 -07:00
Teknium
ada3c28d45 fix(restore): fail closed on in-process holders before unlinking state.db sidecars
Both destructive restore paths (_safe_restore_db's unlink+move fallback and
update_cmd._restore_state_db_from_snapshot) guarded only against FOREIGN
holders via _foreign_db_holder_pids(), which excludes the calling process by
design. A live tracked connection in the same process (the agent's own
SessionDB during /snapshot restore, a read-pool handle, a second SessionDB
instance) was unprotected: the swap unlinked state.db and its -wal/-shm
under it, leaving the process on deleted-inode fds — the #90837/#90950
split-brain fingerprint, produced first-party. Proven live on main via
/proc/self/fd (state.db-wal (deleted) ghosts after both paths ran under a
tracked connection).

Run the destructive swap inside sqlite_safe_read.offline_file_access(),
which fails CLOSED on any tracked in-process connection and holds the
connection-lifecycle lock across the whole swap so no new connection can
appear mid-replace. Holder-free restores are unchanged (control verified:
restore succeeds, stale sidecars cleared).

Part of the #90837 sidecar-unlink audit (wave 6).
2026-09-01 09:27:47 -07:00
Teknium
9f377264aa refactor(update): consolidate the gateway drain triage into one shared helper
Follow-up on the #100179 deadlock break (cherry-picked from PR #100207 by
@salch-cred): the systemd and bare-process restart paths carried two
duplicated copies of the same three-way decision (ancestor fire-and-forget
#100179 / wedged escalation #81642 / normal graceful drain). Extract it
into _drain_or_signal_gateway_for_update() so both call sites share one
implementation, and add direct unit tests for all three branches.

No behavior change: same prints, same return semantics, same drain budget
handling at both sites.
2026-09-01 08:34:51 -07:00
salch-cred
49a71c9727 fix(update): break the cron-update three-way restart deadlock (#100179)
When hermes-auto-update runs \hermes update\ from cron, the update
process lives INSIDE the gateway's own process tree. Waiting for that
gateway to exit is a circular wait:

  gateway  waits on all in-flight work units (#77184 don't-amputate)
    -> cron agent session waits on the \hermes update\ process to exit
      -> \hermes update\ waits on the gateway to exit  [back to A]

The wedged-loop probe (#81642) cannot break it: the cron session posts
activity every ~180s (process-tool poll return), so it is 'actively
waiting forever' and never marked wedged. The gateway logs
'Restart deferred: waiting on 1 active work unit(s)' every 30s until the
1800s force-drain cap amputates its own updater's session — reported as
a 5+ minute hang with gateway_state.json stuck at draining +
restart_requested (v0.21.0, main @ 530aa7b10f).

Fix (the issue's recommended option 1): at both drain sites in
update_cmd.py — systemd (line ~9862) and the bare-process/launchd path
(line ~10203) — check \_is_pid_ancestor_of_current_process(pid)\ before
drain-waiting. When the target gateway IS an ancestor, use
\_request_gateway_self_restart\ (SIGUSR1, no exit-wait) and return: the
gateway's own restart flow completes normally once this process, and
therefore the cron work unit holding it, exits.

Both helpers already exist in hermes_cli/gateway.py (277-304) and
\_request_gateway_self_restart\ already refuses non-ancestor PIDs, so a
normal out-of-tree \hermes update\ keeps its full drain semantics
(including the #86684 cron floor) untouched.

Tests (tests/hermes_cli/test_update_cron_deadlock_guard.py, 6):
- own PID / parent PID are ancestors; 0 and negative are not
- self-restart refuses a non-ancestor PID [linux]
- ancestor path sends SIGUSR1 and NEVER calls _wait_for_pid_exit
  (the deadlock witness — a wait there is the bug) [linux]
- non-ancestor path still drain-waits with the given budget [linux]
Existing graceful/sigusr1/restart tests pass unchanged (9 passed).

Fixes #100179
2026-09-01 08:34:51 -07:00
Justin Wilson
fb18fedf29 fix(updater): rebuild desktop on Windows hand-off repair path
The HERMES_UPDATE_REEXEC child and the current-checkout Node repair
path printed success without calling _rebuild_desktop_after_update.
A failed rebuild now withholds the success banner the same way the
commits-pulled path does.

Fixes #97343
2026-09-01 08:27:58 -07:00
JoaoMarcos44
86b50fb43a fix(update): back up HEAD to a rescue ref before orphan-history reset
On orphan divergence (no common ancestor with origin/<branch>, #87694),
`hermes update`'s ff-only fallback went straight to `reset --hard`,
silently discarding the entire local commit graph with no recovery path.

Probe `git merge-base HEAD origin/<branch>` before the reset; when no
common ancestor exists, park the pre-pull SHA under
refs/hermes-update-backups/orphan-<branch>-<utc-ts>-<sha12> via a single
`git update-ref`. Ordinary divergence (ancestor exists) is byte-for-byte
unchanged. The update-ref return code is checked so the user is never
told a backup exists when the write failed.

Bounded growth (size-analysis mandate): a rescue ref pins every object
reachable from the parked commit — in the incident shape that includes a
full working-tree snapshot which can be multi-GB. _prune_orphan_rescue_refs
enforces two limits on every orphan incident: keep at most 10 refs
(count cap) and expire any ref older than 30 days (age expiry, parsed
from the ref-name timestamp). The user-facing message states when the
backup expires.

Tests: orphan backup, honest failure messaging, count-cap prune,
age expiry, unparseable-name safety, ordinary-divergence regression
guard, update-ref sabotage (non-fatal), missing pre-pull SHA, reset
failure persistence, real-git merge-base anchor, and a real-git
end-to-end prune test proving pruned refs unpin objects for gc.

Fixes #87694
Salvaged from #87745 with expiry mitigation added.
2026-09-01 07:01:12 -07:00
joaomarcos
27cd0ff4e8 fix(update): keep serve-unit recovery identity scope-qualified
Review on #96235: discovery distinguished `(scope, unit)`, but the skip
payload and the reported outcomes reduced that to the bare service name.
`user/hermes-serve.service` and `system/hermes-serve.service` are two
different processes, so a single unqualified token could suppress recovery
of both: if the user-scope unit was already settled when the restart phase
aborted, the stale system-scope unit was never restarted and nothing
downstream reported it.

Scope now travels with the unit end to end:

- the in-process systemd loop records a scope-qualified twin of
  `restarted_services` (`restarted_scoped_units`) while the bare-name list
  keeps its existing vocabulary for the fleet probe and the receipt;
- the recovery payload carries `{"scope", "unit"}` objects, and the child
  keys discovery, skips, outcomes and accounting by `(scope, base)`;
- `verified` / `failed` — and therefore the receipt and the completion
  predicate — report `user/hermes-serve`, never a bare name;
- an entry with no scope (a payload written by a pre-update interpreter)
  stays unqualified and is read as scope-agnostic, and an unrecognized
  scope drops the skip rather than honouring it: dropping a skip can only
  cost one more restart-and-verify, honouring an unreadable one can leave
  a stale generation running.

Also from review: the survivor probe compared PIDs alone while the plan
discarded the process incarnation, so a new serve that reused the planned
PID read as the pre-update survivor. The inventory now records the ledger's
`create_time` in the serve/dashboard runtime detail and the probe compares
`(pid, create_time)`, still failing closed when either side has none.

Finally, abort recovery moves out of the update monolith into
`hermes_cli/update_abort_recovery.py` (417 lines) with `update_cmd`
re-exporting the names `hermes_cli.main` and the update flow address.
`update_cmd.py` ends up 75 lines smaller than the PR's base commit instead
of 249 lines larger.

Tests: dual-scope same-name regressions in both directions, proof that no
systemctl verb reaches an already-settled scope, per-scope outcomes, the
legacy unqualified shape, the qualified payload shape, scope-qualified
completion accounting, PID-reuse vs. same-incarnation survivors, and the
inventory carrying `create_time`.

Refs #92145

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01YQ9oCBKgAMHSGG8CLEHLMC
2026-09-01 07:00:54 -07:00
joaomarcos
f9fec169b5 fix(update): recover hermes-serve units after an aborted restart phase
The fresh-process recovery boundary added for #92145 only reaches gateway
profiles. `hermes serve` -- the runtime that hosts `tui_gateway.server`,
and the process the original report saw failing every chat turn -- is not a
gateway profile, so no `gateway restart` command can reach it and the
gateway-only `collect_fleet_versions` read-back cannot see it either.

The spawn-ledger collector classifies serve/dashboard runtimes purely by
spawner liveness, and a systemd-launched `hermes serve` sets neither
HERMES_SPAWN nor HERMES_PARENT_PID, so it is recorded as `manual-serve`
and the recovery partition skips it as unrecoverable. The result is an
update that clears its incomplete flag on gateway coverage alone while a
live serve process keeps serving the pre-update module graph.

- restart active `hermes-serve*` systemd units from the fresh child,
  enumerated from systemd rather than from the misclassifying inventory,
  and verify a changed MainPID on an active unit before claiming coverage;
- report any pre-update serve/dashboard process that is still the same
  process, and never kill one -- a manual or Desktop-owned serve has no
  relaunch authority;
- require every runtime family, not just the gateway leg, before a
  fresh-process recovery may clear the incomplete flag;
- persist serve-unit outcomes and surviving runtimes in the update receipt.
2026-09-01 07:00:54 -07:00
Teknium
0f16e413e6 fix(update): run pending fleet-restart catchup before the runtime-verification exit gate
The #95294/#91277 fleet contract requires the pending-restart check to
execute on every already-up-to-date pass; the verification exit(1) now
fires only after the catchup runs, so a vulnerable runtime demotes the
outcome to partial without stranding the fleet on stale code.
2026-08-31 12:58:14 -07:00
Teknium
d6bf89de40 fix(update): treat an unprobeable post-update interpreter as non-blocking
Verification-features-must-not-false-positive-on-rollout rule: only a
POSITIVE vulnerable SQLite probe demotes success to partial. A dev
checkout without a venv (or a failed probe subprocess) keeps the
success banner and the fleet-restart catchup path — the CI-red
test_update_fleet_restart_pending already-up-to-date trio pinned this.
2026-08-31 12:58:14 -07:00
fangliquanflq
9caf31394f fix(update): verify repaired current-checkout runtime 2026-08-31 12:58:14 -07:00
fangliquanflq
4f19808915 fix(update): fail current-checkout runtime verification 2026-08-31 12:58:14 -07:00
fangliquanflq
7e22725bb4 fix(update): verify SQLite runtime remediation 2026-08-31 12:58:14 -07:00
SayHell0W0rld
6cccc2ef7e fix(update): locate PortableGit under the shared root, not profile home
Review feedback on #88136 (monerostar): a profile-scoped `hermes update`
sets HERMES_HOME to <root>/profiles/<name>, but the Hermes-managed
PortableGit tree lives under the SHARED root (<root>/git/...). The locator
checked get_hermes_home() only, so a broken trampoline during a
profile-scoped update was not swapped and fell through to ZIP.

Extract _portable_git_candidates() (shared root first, profile home as
fallback) and add a regression test for the profile layout.
2026-08-31 12:21:46 -07:00
SayHell0W0rld
0bab9ff9e7 fix(update): self-heal broken Git-for-Windows trampoline on Windows
A Git-for-Windows trampoline launcher (bin\git.exe / cmd\git.exe shim,
~46KB) that fails to re-exec the real git-core binary refuses every git
call with a "BUG (fork bomb)" guard instead of running it (#87876).

Detect the trampoline up front via `git --version`, locate a real git
binary (Git for Windows or Hermes-managed PortableGit locations), and
rebuild the git command with it so fetch/pull/checkout keep working with
a real git instead of degrading to the ZIP fallback. When no real binary
can be found, leave the command untouched so the existing fetch-failure
handler still falls back to the ZIP path on Windows (#88046).
2026-08-31 12:21:46 -07:00
andyst-dev
4d3e1e4d13 fix(desktop): prevent venv scan timeout on busy Windows hosts 2026-08-31 12:00:33 -07:00