`LSPService` kept one client per `(server_id, workspace_root)` for the life
of the process. In a long-running gateway that outlives its coding sessions
the client for a removed worktree stayed registered with its stdio pipes
held open — tsserver heaps of several GiB pointed at trees that no longer
existed (#102345). The idle reaper (d7578018c5) does not cover this: a
client whose root vanished is not idle from the server's point of view.
- `LSPService.release_workspace(path)`: detaches every client whose folders
live under `path` under `_state_lock`, waits for in-flight spawns so a
concurrent `_get_or_spawn` cannot reinsert a client after release, prunes
the delta baselines and broken-set entries beneath the path, and shuts the
clients down on the service loop. Multi-root servers (pyright) only drop
the folder (`LSPClient.remove_workspace_folder`) so siblings keep their
shared process. Idempotent; best-effort; returns the count.
- The reaper sweep now also uses that primitive for clients whose every
workspace folder no longer exists (externally deleted roots).
- `agent.lsp.release_workspace()` reaches every started service without
creating one; `hermes_cli.worktree_ops.release_lsp_clients()` is called
from both worktree-removal paths (`cli._cleanup_worktree`, kanban
`_cleanup_worktree_workspace`) BEFORE `git worktree remove`.
Salvages the direction of #102381 (@Sahilvishnaliya, release_workspace +
cli hook) and the deleted-root eligibility of #95047 (@israellot); both
matched on `key[1]`, which is `""` for multi-root servers and would have
reaped pyright on every sweep — the primitive here matches on
`client.workspace_folders` instead.
Co-authored-by: Sahil Vishnalya <222165401+Sahilvishnaliya@users.noreply.github.com>
Co-authored-by: Israel Lot <840042+israellot@users.noreply.github.com>
Every `hermes -w` launch on a clone past the pack-sprawl threshold started
its own `git repack -a -d` of the whole object store in a daemon thread.
On a multi-agent box that meant dozens of concurrent multi-GB repacks of
the same repo, each too starved (nice 19, under the others) to finish
inside the 1800 s timeout. subprocess.run's timeout killed only `git
repack`, so the `pack-objects` grandchild kept running with ppid 1 for
days; the CLI exiting orphaned it the same way. Observed: 51 pack-objects
processes, load 200 on 20 cores, 143 GB swap, 29 GB of `.tmp-*-pack`
debris, and every `hermes` invocation taking 5-7 s of wall clock for
0.6 s of CPU.
- `_claim_repack_slot`: one repack per clone per 6 h across processes
(`.git/hermes-repack.lock`, mtime = stamp; stale takeover via `replace`
so only one of N racers wins).
- `_run_bounded_repack`: own process group + `kill_process_tree` on
timeout and at exit, so the whole tree dies with the launcher.
- incremental `git repack -d --geometric=2 --write-midx` instead of a
full `-a` rewrite: consolidates sprawl in seconds instead of rewriting
10 GB per run.
- `.tmp-<pid>-pack*` (pack-objects debris) joins gitlock's stale tmp-pack
sweep and is swept before repacking.
`hermes worktree prune` (and the startup/cron pruner) classified every clean tree in a
repository with no remote as "clean and fully merged/pushed" and force-deleted its branch,
even when the branch carried commits that exist nowhere else. `audit_branches` returned []
in the same repos, so a unique local-only branch was invisible to the audit as well.
Root cause: `_worktree_has_unpushed_commits` answered False when `refs/remotes` was empty
("nothing to be unpushed against") and both consumers — `worktree_gc._classify_tree` and
`worktree_ops._classify_prune_candidates` — read False as "safe to reap".
The preceding commit (#111897) flips that branch to True, which is safe but also means a
no-remote repo can never reclaim anything (a tree sitting at trunk, or squash-merged into
it, stays "unpushed" forever because `_worktree_commits_all_merged_upstream` finds no
origin/* base). This commit replaces the unconditional True with a real baseline:
- `_worktree_local_trunk`: `main`/`master`, else the branch checked out in the main
worktree; None when no trunk exists.
- `_worktree_merge_base_ref`: origin/HEAD|origin/main|origin/master, falling back to the
local trunk ONLY when the repo has no remote-tracking refs at all. Single resolver used
by `_worktree_commits_all_merged_upstream` and `worktree_gc.audit_branches`.
- `_worktree_has_unpushed_commits`: with no remote refs, `git log HEAD --not <trunk>`;
no trunk -> True (preserve), matching the docstring's fail-safe promise.
Net effect in a no-remote repo: unique work is kept ("unpushed commits not found
upstream"), trees at/merged into the local trunk still reap, branch audit reports unique
branches as keep and merged ones as delete, and `git branch -D` can only run on a branch
whose every commit is reachable from or patch-equivalent to the trunk.
Tests: the salvaged no-remote keep test now uses a shared `local_repo` fixture; a control
test proves trunk-merged trees still reclaim and the branch audit reports in the same
repo; `test_merged_predicate_fails_safe_without_upstream` now pins both halves of the
contract (trunk resolves -> merged; no trunk at all -> False/preserve).
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).