Commit Graph

30 Commits

Author SHA1 Message Date
Teknium
a5bd246865 Old pre-decomposition import paths are gone: plugin compat layer removed on schedule (#126164)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule

The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:

- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
  names, restored dead definitions) and the three re-export stub modules
  (gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
  the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
  `plugins.allow_deprecated_imports` escape hatch

An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.

hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).

In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.

* chore: retrigger CI (zero-job startup_failure phantom)

* test: drop resolution allowlist rows for the two deleted which() sites

hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
2026-09-28 10:21:41 -07:00
teknium1
010618c767 fix(update): one commit-graph fetch for the pre-fetch heal and the post-update pass
The pre-fetch heal (#123254) and the post-update tag fetch now share
fetch_full_commit_graph, so both follow one clone-mode rule: a partial clone
repeats its own filter, a full clone fetches unfiltered (#122353), and only a
depth-limited clone whose history is really missing unshallows as tree:0. A
full clone grafted by a user's `git fetch --depth 1` already has its history
locally, so it unshallows unfiltered and stays full.

Also: the network-timeout message no longer claims "no response from the
remote" (a large transfer on a slow link hits the same cap), without adding a
classifier rule that would assert the opposite.
2026-09-28 03:08:56 -07:00
Konstantin Khlopkov
e377309d86 fix(update): heal stale shallow history before the bounded fetch
A depth-1 installer checkout far behind main could never update: the plain
main fetch pulled nearly the full object graph (merge-heavy history drags in
side-branch ancestry past the shallow boundary), hit the 300s network cap
mid-transfer every time, and the unshallow step that makes later fetches
incremental only ran after a successful update. Unshallow the checkout first
(commit-graph-only fetch with the same 900s budget the post-update tag fetch
uses, pulling the update branch along), non-fatally, and stop reporting a
wall-clock kill as 'no response from the remote'.

Fixes #123254
2026-09-28 03:08:56 -07:00
liuhao1024
6b51b0fd4e fix(update): expire fetch reflogs that pin stale shallow grafts
prune_stale_shallow_grafts dropped grafts no live ref points at, but the
remote-tracking reflog still naming a dropped fetch tip made the --reflog
fail-safe walk fail (its parent was never fetched), rolling the prune back
on every run — grafts kept accumulating and the next update fell into
orphan divergence (#124645). Expire exactly the fetch reflogs that pin a
dropped graft before writing the trimmed file; user reflogs (HEAD, local
branches) are never touched, so the fail-safe rollback still guards them.

Fixes #124645
2026-09-28 03:08:56 -07:00
Yuan Li
ffa8817345 fix(update): recognize the Windows pack-objects crash stderr in the partial-clone recovery
Git for Windows 2.54 reports the same partial-clone pack-objects BUG()
without the POSIX 'pack-objects died of signal 6' line (#124293 field
evidence): the BUG assertion is followed directly by 'fatal: could not
finish pack-objects to repack local links' and 'fatal: index-pack failed'.
The recognizer demanded the signal wording, so the recovery retry missed
this Windows occurrence entirely.

Require the BUG fingerprint plus EITHER terminator of the same abort —
the POSIX signal line or the Windows repack fatal — and pin both stderr
shapes plus a fingerprint-without-terminator negative in the regression.
2026-09-28 03:08:56 -07:00
Yuan Li
fc6144e3c1 fix(update): self-heal git pack-objects crash on partial-clone fetches
On a partial clone (clone --filter=tree:0), git 2.53/2.54 crashes every
fetch: index-pack --promisor runs repack_local_links(), which feeds
pack-objects --exclude-promisor-objects-best-effort, and pack-objects
BUG()s (SIGABRT, 'should only be called on existing objects') when the
link traversal hits a legitimately missing promisor object. The crash is
deterministic, so every fetch — and with it the CLI's update/check flows
— stays broken until the user applies the manual workaround themselves.

The recovery is the reporter's verified workaround, wired in: on this
exact crash (all three stderr markers), retry the fetch once with
-c remote.origin.promisor= (per-invocation only; the user's filter
choice stays in their config). Wired into both fetch sites — the update
fetch and the check fetch (which also serves forks via upstream).
Unrelated fetch failures are returned untouched; if the retry still
crashes, the update path prints the manual heal command.

Fixes #124272
2026-09-28 03:08:56 -07:00
Yuan Li
5849446b02 fix(update): repeat the clone's own partial-clone filter, not tree:0
Hardcoding tree:0 for partial clones would silently tighten a checkout the
user cloned with --filter=blob:none. Read the existing
remote.origin.partialclonefilter instead and pass it back verbatim; full
clones still fetch unfiltered.
2026-09-28 03:08:56 -07:00
Yuan Li
de98f88909 fix(update): do not re-arm partial-clone config on the post-update tag fetch
git fetch --filter=tree:0 writes remote.origin.promisor=true and
remote.origin.partialclonefilter=tree:0 even when those keys were
deliberately removed, so every hermes update converted a de-partialised
checkout back into a partial clone and re-armed the should_include_obj
fetch failure (#122353). Pass the filter only when the checkout already
is a partial clone; full clones fetch tags unfiltered.
2026-09-28 03:08:56 -07:00
Austin Pickett
b00b4bb7f2 fix(cli): pin utf-8 decoding on all text-mode subprocess readers
On Windows, subprocess text=True without an explicit encoding decodes
child output with the ANSI code page (e.g. 'gbk'); non-ASCII bytes then
raise UnicodeDecodeError inside subprocess._readerthread, killing the
Hermes backend before it becomes ready and surfacing as the desktop boot
timeout.

Sweep every hermes_cli text=True subprocess call to encoding='utf-8',
errors='replace', and add an AST-based regression test that fails when a
future text-mode call omits the encoding.

Fixes #55658
2026-09-25 14:22:12 -04:00
ethernet
97c4fa022c fix(update): refresh release tags before stamping source versions 2026-09-24 23:52:31 -04:00
ethernet
f5caa6c26f fix(update): fetch the commit graph of shallow checkouts before publishing identity
Pre-PM installers cloned --depth 1, so an install updated onto this line
stamped baseVersion "unknown" and printed git.<sha>. The completion tail
(install, update, historical takeover, --finish-update) now turns a
shallow checkout into a treeless partial clone (fetch --unshallow
--filter=tree:0) before the completion line and the stamp read identity.
Commits only: 77M -> 119M in ~5s on the real repo. A failed fetch warns
and the update still completes; the next update retries.
2026-09-24 00:51:07 -04:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
liuhao1024
137cbb7819 fix(checkpoints): sweep tmp_pack debris stranded by timed-out store gcs
A git gc killed by _run_git's timeout strands tmp_pack_* files in the bare
store's objects/pack; gc.auto=0 means git itself never reclaims them (~16 GB
observed on one host). clear_stale_tmp_packs() already sweeps this debris
class for the update checkout/worktree paths — teach it to resolve a bare
repo's objects/pack (no .git/ layer) and call it from prune_checkpoints(),
unconditionally under _store_has_head so the daily auto-prune and
'hermes checkpoints prune' both reap it even when no ref moved (a sweep is a
cheap directory listing, unlike the pack-rewriting gc that stays gated on
refs having moved).
2026-09-20 12:03:06 -07:00
teknium1
d767bfabda fix(update): drop the _IS_WINDOWS test flag; prove the read-only sweep on a real Windows host
The salvaged commit added a module-level `_IS_WINDOWS` flag "so tests can
flip it" and a test that monkeypatched it to True with a faked Path.unlink.
A flag flipped by tests is an OS fake (repo rule: never make the interpreter
believe it is on another OS), so use the inline `os.name == "nt"` check the
rest of this module already uses and replace the test with a
`@pytest.mark.windows_only` invariant: a real aged 0o444 tmp_pack_* under
.git/objects/pack is removed by clear_stale_tmp_packs on native Windows.
The WARNING-level caplog test stays (Linux-runnable, red on base).

Supersedes #116391 (@kokhlo), which carried the same inline-check shape.

Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
2026-09-20 10:02:55 -07:00
liuhao1024
9bd542c4a0 fix(cli): clear the write bit before unlinking read-only pack temps on Windows
git renames its transfer temps (tmp_pack_* and friends) into place
read-only; Windows refuses to unlink a read-only file with EACCES, and
_sweep_stale swallowed that at debug level — so the sweep silently
removed nothing on real debris (#116384: 171 tmp_pack_* files / 8.8 GB
accumulated with zero sweep success lines in the logs).

Clear the write bit before unlinking on Windows and surface every skip
at warning: a cleaner that fails silently is worse than none.
2026-09-20 10:02:55 -07:00
ethernet
b4a294fff9 Merge origin/main; keep PM as plugin dependency owner
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
2026-09-17 13:52:05 -04:00
teknium1
76478fc08d fix: hermes -w repack no longer stampedes a shared clone
Every `hermes -w` launch on a clone past the pack-sprawl threshold started
its own `git repack -a -d` of the whole object store in a daemon thread.
On a multi-agent box that meant dozens of concurrent multi-GB repacks of
the same repo, each too starved (nice 19, under the others) to finish
inside the 1800 s timeout. subprocess.run's timeout killed only `git
repack`, so the `pack-objects` grandchild kept running with ppid 1 for
days; the CLI exiting orphaned it the same way. Observed: 51 pack-objects
processes, load 200 on 20 cores, 143 GB swap, 29 GB of `.tmp-*-pack`
debris, and every `hermes` invocation taking 5-7 s of wall clock for
0.6 s of CPU.

- `_claim_repack_slot`: one repack per clone per 6 h across processes
  (`.git/hermes-repack.lock`, mtime = stamp; stale takeover via `replace`
  so only one of N racers wins).
- `_run_bounded_repack`: own process group + `kill_process_tree` on
  timeout and at exit, so the whole tree dies with the launcher.
- incremental `git repack -d --geometric=2 --write-midx` instead of a
  full `-a` rewrite: consolidates sprawl in seconds instead of rewriting
  10 GB per run.
- `.tmp-<pid>-pack*` (pack-objects debris) joins gitlock's stale tmp-pack
  sweep and is swept before repacking.
2026-09-16 16:43:08 -07:00
ethernet
5623942736 fix(update): tolerate BOM receipts and preserve rollback bytes 2026-09-13 18:21:59 -04:00
kshitijk4poor
eec131b716 refactor(cli): hold shallow.lock across write-verify-restore; share the atomic write
Simplify-pass follow-up on the repair rework: the self-check rev-list walks
and the rollback restore now run inside the same _ShallowLock hold (rev-list
never takes shallow.lock), so lock contention can no longer defeat a failed
self-check's rollback and leave a broken .git/shallow in place. The tmp-write
+ os.replace sequence shared by repair and prune moves into _write_shallow.
Scope note added: fetch-by-SHA install tips (HEAD-reflog-only) are not repair
candidates; corruption of that shape is prevented by the prune's reflog
fail-safe. 68 focused tests green; two-cycle repair->prune E2E re-verified.
2026-09-12 15:00:37 +05:30
kshitijk4poor
966fb375fc fix(cli): make shallow-boundary repair survive the prune and the graph-safety review findings
Rework of the repair pass from #108361 (salvage) addressing the blocking
review findings, verified with real-git probes:

- Sequencing: prune_stale_shallow_grafts' fail-safe now also walks
  rev-list --all --reflog, so a boundary the repair just restored (one a
  reflog-only commit still needs) is never dropped again; previously the
  production repair->prune sequence re-broke the repo on every update run.
- Header-only parent parsing: a "parent <sha>" line inside a commit
  message body is prose; _batch_missing_parents stops at the blank line
  ending the commit header, so healthy history is never truncated.
- Candidates restricted to fetch-recorded tips (refs/remotes/* reflogs),
  not --batch-all-objects: unrelated object loss (a deleted parent of a
  locally-created commit) is no longer re-labelled as shallow history;
  fsck keeps reporting it.
- Concurrent-writer safety: both .git/shallow writers now hold git's own
  shallow.lock, so a depth-1 fetch between read and write fails fast
  instead of being clobbered (or clobbering us).
- Cheap gate: repair runs its subprocess fan-out only when
  rev-list --all --reflog already fails; healthy updates pay one probe.
- --batch-check returncode is now checked; shared helpers
  (_shallow_file_path, _ShallowLock) replace the copy-pasted plumbing;
  test file footguns fixed (encoding=, as_uri()) and the missing
  repair->prune end-to-end regression added, mutation-checked.
2026-09-12 15:00:37 +05:30
joaomarcos
2fb87f047f fix(cli): repair shallow boundaries already dropped by stale-graft prune
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.

Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.

This complements PR #108290, which owns the prevention half.

Refs #108286
2026-09-12 15:00:37 +05:30
liuhao1024
5bfa389e8a fix(cli): prune stale shallow grafts left by depth-1 update checks (#105951)
Every 'git fetch --depth 1' in 'hermes update --check' (and the past
banner passive checks, before #107648 moved them to the GitHub API)
appends the fetched tip to .git/shallow as a new graft and git never
removes the previous one, so a long-lived shallow installer checkout
accumulates one graft per check (57 observed). The stale grafts break
merge-base and push 'hermes update' into the orphan-divergence reset
path with a rescue ref on every run.

prune_stale_shallow_grafts() now runs after each successful depth-1
fetch in 'hermes update --check' and clears the grafts already
accumulated by past checks: it keeps only the boundaries still
protecting referenced tips (HEAD, FETCH_HEAD, every ref tip) and
atomically rewrites .git/shallow, restoring the original file if the
trimmed set breaks history walking. The dropped commits are already
unreachable; their objects are left for git gc.

Rebased onto main after #107648: the banner.py hook is dropped (the
passive check no longer git-fetches); the update --check prune and the
cleanup of already-accumulated grafts are kept.

(cherry picked from commit 6174837fc5b9f4cc3d4d46dc1b2d9a2f6b83c120)
2026-09-11 02:06:32 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
23b9ffc4fa fix(integration): restore subprocess stdin=DEVNULL / utf-8 encoding guards and windows-footgun gates dropped by round-3 compaction
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
2026-09-03 02:46:19 -07:00
Teknium
d4904b8327 refactor(hermes_cli): foreign_sessions _block_text/_message_turn fold, heartbeat/dingtalk/gitlock/install_identity compaction 2026-09-03 00:10:28 -07:00
Teknium
fc51206750 refactor(hermes_cli): second pass on the 15 small modules — compact bodies, drop redundant locals/blanks, tighten module docs 2026-09-02 21:20:42 -07:00
Teknium
416a1c1acd refactor(hermes_cli): simplify 15 small modules — dead wrappers, unified helpers, dict dispatch, compact docs 2026-09-02 20:29:34 -07:00
Teknium
2c7f5d12b5 refactor(hclib): cron/update/install lifecycle — cron status, update_* receipts and recovery, install repair, service manager 2026-09-02 14:45:25 -07:00
Teknium
df3d41ee67 fix(update): sweep aborted-fetch tmp_pack debris before it corrupts the pack directory (#93732)
Every git fetch that dies mid-transfer (timeout, HTTP 429, dropped
line) strands a tmp_pack_* file in .git/objects/pack, and git never
cleans them. The banner's background update check is the main generator
on flaky lines — several aborted fetches a day — and the reporter's
install accumulated hundreds of files / 6.0 GB over 9 days until the
pack directory corrupted outright and every update check hung or
failed permanently.

clear_stale_tmp_packs() in gitlock.py sweeps tmp_pack_/tmp_idx_/
tmp_rev_/tmp_mtimes_ debris with the exact safety contract the lock
sweep already uses: only files past the 10-minute age floor, never
while any git process runs, never raises, real pack-*.pack/.idx files
untouchable by construction (prefix match). Wired into all three
fetch-adjacent sites: _cmd_update_check, the update apply path, and
the banner's passive check (generator = janitor).

Live E2E: 300 aged tmp_pack files (the reported scale-shape) swept
from a real repo; an in-flight fresh tmp and ancient real packs
survived; fsck clean and a real fetch round-trip succeeded after.
2026-08-26 16:45:22 -07:00
RGerrish
7fe3bf042b fix(update): self-heal stale git locks and stop false 'update available' on shallow clones
Two related failure modes after a crashed/interrupted fetch on a shallow
clone (git clone --depth 1 installs):

1. STALE LOCK WEDGES EVERY FETCH. A killed fetch can leave .git/shallow.lock
   behind; every later 'git fetch' then fails with 'Unable to create
   .../shallow.lock: File exists'. 'hermes update --check' reported a hard
   fetch failure, and the passive banner check swallowed the exception and
   compared stale refs. Add hermes_cli.gitlock.clear_stale_git_locks(), a
   guarded sweep (age + git-process check so a live fetch is never yanked)
   wired into the check path, the apply path, and the banner's passive check.

2. SHALLOW TIP-SHA COMPARE FALSE-POSITIVES. On a shallow clone the check
   cannot count commits, so it compares tip SHAs. Local cherry-picks on top
   of the remote tip (e.g. re-applied local patches) make HEAD differ from
   origin/main even though HEAD already contains it — a false 'update
   available' banner. Add hermes_cli.gitlock.is_ancestor_of_head() and use
   'git merge-base --is-ancestor' in the CLI check and banner paths before
   reporting an update. Mirror in the desktop (update-count.ts gains an
   isAncestor input; main.ts probes merge-base --is-ancestor).

Tests: tests/test_gitlock.py (9) covering stale/young/no-lock/no-repo sweeps
and ancestry true/false; update-count.test.ts +3 for the isAncestor path.
2026-08-15 04:33:47 -07:00