Twenty-five call sites told users to run
`python -c "from pm import sync_venv; sync_venv(['x'], explicit=True)"`
because `hermes pm install` only took package names. Add `--extra`
(repeatable; syncs the venv with the named extras and nothing else) and
pm.install_hint(extra), the single builder every site now uses, so the
advice stays correct when the command changes.
A cold PM runtime under allow_lazy_installs:false now reports the extra
the caller wanted and the command that provisions both, instead of a
bare "pm-runtime: not installed".
- 15 `MERGE-CHECK:` conflict-resolution comments removed from prod code (two were
TODOs already done: the utf-8-sig sessions.json read lives in session_persistence,
the pm-aware cron script helpers in scheduler_script).
- 49 imports the branch left unused (ruff F401, none present at the merge base,
none inside PLUGIN-COMPAT blocks). update_cmd's frozen-surface re-exports are
trimmed to the names tests/compat/old_updater_surface.json actually lists under
hermes_cli.update_cmd; the rest resolve through hermes_cli.main.__getattr__.
- tools/environments/local_gitbash_probe.py: nothing imported it once _find_bash
delegated to pm.shell().
- Three try/except wrappers around calls that cannot raise (install_truststore,
get_hermes_home, and a duplicated except clause in supermemory).
Merge d86627a7f3 into pm-clean. Keep the shared bounded ripgrep transport and preserve translate_path=False for probe patterns. Targeted canonical search tests: 19 passed, 3 POSIX-only skips on Windows. Full integrated CI has not been run for this merge.
- acquire(): the discard-and-recurse path becomes one more iteration of the existing wait loop;
the two inline 'with lifecycle_lock: _teardown(db)' copies reuse _teardown_generation, and the
type-narrowing asserts go away with the recursion. release() reads generation.path directly.
- Restore the real os.replace inode swaps in test_state_db_file_identity.py and the registry tests:
those files carry no windows_only marker so they never run on Windows, and the monkeypatched
predicate stopped exercising the stat->identity->halt path anywhere.
- Drop the auto-archive change and its 4 tests: on main the sweep gets a bare SessionDB and
db.close() already releases a registry-shared handle, so the described NameError leak only
existed on this branch's earlier head. trace_upload: acquire(None) already defaults.
- Trim the barrier tests to the invariant pair (retired drain must not lift a pending current
teardown; replacement not published before the last close settles) plus the raising-close
settlement; comments say the WHY once.
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.
Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.
Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.
The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.
Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.
Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.
Fixes#102827
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2
Repo scanners (check_subprocess_stdin, check-windows-footguns --all) flagged 21 sites where
the r3 single-line collapses lost stdin=DEVNULL, encoding='utf-8'/errors='replace', the
'# windows-footgun: ok' same-line marker, or the getattr(os, 'geteuid') gate. Each guard is
restored at the call site (real portability/hang fixes, not suppressions).
Introduce the pm store: a unified, hash-verified package store that
replaces lazy_deps and the old installer's ad-hoc tool downloads.
Store tools are provisioned on PATH (ffmpeg, node/npm via pinned uv),
with a resumable 8-way downloader, verify() returning failure reasons,
and adopt() made EPERM-safe. chromium ships in the payload for every
target. The 3600-line install.sh is replaced by a staged bootstrapper
(heavy deps are pm's job after this); setup-hermes.sh, Dockerfile and
nix pin tables are rewired onto the store. Old install-script tests,
lazy_deps/managed_uv/build_info, and the ps1/bash installer test
batteries are removed with the machinery they tested.
Rebuilt from ethie/pm onto upstream/main (ac6c8028e0) after the
utf-8-sig sweep. 16 hot files (main also churned them) hand-merged:
platform adapters, main.py, electron/main.ts, tui_gateway/server.py,
cua_backend, installer-tests workflow, install.sh (full rewrite),
setup-hermes.sh, plugins doc.
SessionDB could leave native SQLite handles open when construction failed
partway through schema/pragma/FTS/repair/lock/interrupt handling. Other
short-lived callers (MCP reads/polling, session search, reactions, trace
upload, insights, shutdown recovery) opened temporary SessionDB handles
without a complete ownership boundary. API-server profile caches and
RetainDB shutdown had similar late-close races. Under sustained load this
exhausted file descriptors (EMFILE).
- Close partially initialized SessionDB connections on every constructor
exception path via a finally block guarded by an initialization-complete
flag.
- Close temporary/cross-profile SessionDB handles in finally blocks across
CLI, MCP, search, trace, reactions, insights, and recovery paths.
- Add API-server per-profile cache ownership and disconnect cleanup.
- Make RetainDB writer-queue shutdown exception-safe: track connections per
thread, close on worker exit, reject new enqueues after shutdown starts,
and sweep any connections left by short-lived threads.
- Add regression coverage for constructor failures, worker-thread readers,
API disconnect failures, shutdown recovery, RetainDB late enqueue, and
foreign-loop async clients.
Salvage notes: the original PR's per-thread WAL-reader ownership changes
were superseded by main's read-connection pool (permits + checkout/return);
its cron timeout-abandon fix is credited separately to #72822's earlier
identical fix.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
AST-driven pass over every subprocess.run/Popen/check_output/check_call/call
with text=True (or universal_newlines=True) and no explicit encoding=:
append encoding='utf-8', errors='replace' at the kwarg site. 136 call
sites across 28 files (cli.py, hermes_cli/main.py, tools_config.py,
environments, computer_use, gateway, scripts, skills helpers, agent/*).
Together with the salvaged #55339/#60741 commits this closes out issue
#53428's bug class; the salvaged #60751 linter rule in
check-windows-footguns.py now enforces it repo-wide (verified: 807 files
scanned, zero findings).
* feat(trace): upload sessions to HF Agent Trace Viewer
Salvage trace upload as a smaller CLI-first feature: deterministic Claude Code JSONL export, fail-closed redaction, lazy Hugging Face dependency, and no gateway slash-command wiring.
* chore(trace): drop external porting references from docstrings
Describe the trace-upload design in Hermes' own terms.
* feat(sessions): fold trace upload into 'sessions export --format trace'
Integrates the HF Agent Trace Viewer exporter (PR #36145) onto the
unified export surface instead of a separate 'hermes trace' subcommand:
- --format trace: Claude Code JSONL to stdout/file, or one
<id>.trace.jsonl per session for filtered bulk export; defaults to
the most recent session when no --session-id/filters given.
- --upload pushes to the user's private HF traces dataset (--public to
opt out of private); reads HF_TOKEN with guided setup when missing.
- traces are secret-redacted by default (force mode); --no-redact opts
out after review; redaction failure blocks export (fail closed).
- hermes_cli/trace.py + subcommands/trace.py removed; agent/trace_upload.py
is the single engine. Docs EN + zh-Hans; 4 new CLI tests.