Every `hermes -w` launch on a clone past the pack-sprawl threshold started
its own `git repack -a -d` of the whole object store in a daemon thread.
On a multi-agent box that meant dozens of concurrent multi-GB repacks of
the same repo, each too starved (nice 19, under the others) to finish
inside the 1800 s timeout. subprocess.run's timeout killed only `git
repack`, so the `pack-objects` grandchild kept running with ppid 1 for
days; the CLI exiting orphaned it the same way. Observed: 51 pack-objects
processes, load 200 on 20 cores, 143 GB swap, 29 GB of `.tmp-*-pack`
debris, and every `hermes` invocation taking 5-7 s of wall clock for
0.6 s of CPU.
- `_claim_repack_slot`: one repack per clone per 6 h across processes
(`.git/hermes-repack.lock`, mtime = stamp; stale takeover via `replace`
so only one of N racers wins).
- `_run_bounded_repack`: own process group + `kill_process_tree` on
timeout and at exit, so the whole tree dies with the launcher.
- incremental `git repack -d --geometric=2 --write-midx` instead of a
full `-a` rewrite: consolidates sprawl in seconds instead of rewriting
10 GB per run.
- `.tmp-<pid>-pack*` (pack-objects debris) joins gitlock's stale tmp-pack
sweep and is swept before repacking.
Simplify-pass follow-up on the repair rework: the self-check rev-list walks
and the rollback restore now run inside the same _ShallowLock hold (rev-list
never takes shallow.lock), so lock contention can no longer defeat a failed
self-check's rollback and leave a broken .git/shallow in place. The tmp-write
+ os.replace sequence shared by repair and prune moves into _write_shallow.
Scope note added: fetch-by-SHA install tips (HEAD-reflog-only) are not repair
candidates; corruption of that shape is prevented by the prune's reflog
fail-safe. 68 focused tests green; two-cycle repair->prune E2E re-verified.
Rework of the repair pass from #108361 (salvage) addressing the blocking
review findings, verified with real-git probes:
- Sequencing: prune_stale_shallow_grafts' fail-safe now also walks
rev-list --all --reflog, so a boundary the repair just restored (one a
reflog-only commit still needs) is never dropped again; previously the
production repair->prune sequence re-broke the repo on every update run.
- Header-only parent parsing: a "parent <sha>" line inside a commit
message body is prose; _batch_missing_parents stops at the blank line
ending the commit header, so healthy history is never truncated.
- Candidates restricted to fetch-recorded tips (refs/remotes/* reflogs),
not --batch-all-objects: unrelated object loss (a deleted parent of a
locally-created commit) is no longer re-labelled as shallow history;
fsck keeps reporting it.
- Concurrent-writer safety: both .git/shallow writers now hold git's own
shallow.lock, so a depth-1 fetch between read and write fails fast
instead of being clobbered (or clobbering us).
- Cheap gate: repair runs its subprocess fan-out only when
rev-list --all --reflog already fails; healthy updates pay one probe.
- --batch-check returncode is now checked; shared helpers
(_shallow_file_path, _ShallowLock) replace the copy-pasted plumbing;
test file footguns fixed (encoding=, as_uri()) and the missing
repair->prune end-to-end regression added, mutation-checked.
A reflog-only commit can remain present after stale-graft pruning drops the shallow boundary it needs, while its parent was never fetched. That leaves git gc, fsck, and rev-list unable to traverse the repository. Prevention alone is insufficient because a broken gc walk prevents reflogs from expiring.
Repair scans local commit objects without graph traversal, identifies commits with missing parents, and atomically restores their shallow boundaries. It only updates .git/shallow and never expires reflogs, prunes, or deletes objects, so the operation is non-destructive and idempotent.
This complements PR #108290, which owns the prevention half.
Refs #108286
Every 'git fetch --depth 1' in 'hermes update --check' (and the past
banner passive checks, before #107648 moved them to the GitHub API)
appends the fetched tip to .git/shallow as a new graft and git never
removes the previous one, so a long-lived shallow installer checkout
accumulates one graft per check (57 observed). The stale grafts break
merge-base and push 'hermes update' into the orphan-divergence reset
path with a rescue ref on every run.
prune_stale_shallow_grafts() now runs after each successful depth-1
fetch in 'hermes update --check' and clears the grafts already
accumulated by past checks: it keeps only the boundaries still
protecting referenced tips (HEAD, FETCH_HEAD, every ref tip) and
atomically rewrites .git/shallow, restoring the original file if the
trimmed set breaks history walking. The dropped commits are already
unreachable; their objects are left for git gc.
Rebased onto main after #107648: the banner.py hook is dropped (the
passive check no longer git-fetches); the update --check prune and the
cleanup of already-accumulated grafts are kept.
(cherry picked from commit 6174837fc5b9f4cc3d4d46dc1b2d9a2f6b83c120)
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
Every git fetch that dies mid-transfer (timeout, HTTP 429, dropped
line) strands a tmp_pack_* file in .git/objects/pack, and git never
cleans them. The banner's background update check is the main generator
on flaky lines — several aborted fetches a day — and the reporter's
install accumulated hundreds of files / 6.0 GB over 9 days until the
pack directory corrupted outright and every update check hung or
failed permanently.
clear_stale_tmp_packs() in gitlock.py sweeps tmp_pack_/tmp_idx_/
tmp_rev_/tmp_mtimes_ debris with the exact safety contract the lock
sweep already uses: only files past the 10-minute age floor, never
while any git process runs, never raises, real pack-*.pack/.idx files
untouchable by construction (prefix match). Wired into all three
fetch-adjacent sites: _cmd_update_check, the update apply path, and
the banner's passive check (generator = janitor).
Live E2E: 300 aged tmp_pack files (the reported scale-shape) swept
from a real repo; an in-flight fresh tmp and ancient real packs
survived; fsck clean and a real fetch round-trip succeeded after.
Two related failure modes after a crashed/interrupted fetch on a shallow
clone (git clone --depth 1 installs):
1. STALE LOCK WEDGES EVERY FETCH. A killed fetch can leave .git/shallow.lock
behind; every later 'git fetch' then fails with 'Unable to create
.../shallow.lock: File exists'. 'hermes update --check' reported a hard
fetch failure, and the passive banner check swallowed the exception and
compared stale refs. Add hermes_cli.gitlock.clear_stale_git_locks(), a
guarded sweep (age + git-process check so a live fetch is never yanked)
wired into the check path, the apply path, and the banner's passive check.
2. SHALLOW TIP-SHA COMPARE FALSE-POSITIVES. On a shallow clone the check
cannot count commits, so it compares tip SHAs. Local cherry-picks on top
of the remote tip (e.g. re-applied local patches) make HEAD differ from
origin/main even though HEAD already contains it — a false 'update
available' banner. Add hermes_cli.gitlock.is_ancestor_of_head() and use
'git merge-base --is-ancestor' in the CLI check and banner paths before
reporting an update. Mirror in the desktop (update-count.ts gains an
isAncestor input; main.ts probes merge-base --is-ancestor).
Tests: tests/test_gitlock.py (9) covering stale/young/no-lock/no-repo sweeps
and ancestry true/false; update-count.test.ts +3 for the isAncestor path.