`snapshot_paths` / `_store_blob` write blobs OUTSIDE `_ledger_lock`, seconds
before the row that references them is appended. Since `_maintain_size` now
runs `gc_blobs()` on any append that trims, a sweep in another process during
that window deleted the in-flight capture: its row then landed with sha256
values `read_blob` cannot resolve and `rollback_entry` failed.
`_gc_blobs_locked` now skips any unreferenced blob file (incl. `.tmp-*`) whose
st_mtime is newer than `_BLOB_GC_GRACE_SECS = 3600`; no new config key.
Docs (curator.md) and the `gc_blobs` docstring say so.
PROOF: probes/s5_final_inflight_blob.py before -> `P1 blob survives P2's
sweep: False ... read_blob: None`; after -> `True ... read_blob: b'P1
in-flight skill file'`. test_concurrent_appends_never_lose_a_middle_row gained
a fresh + 2h-aged unreferenced blob pair: red with the grace set to 0 (fresh
blob deleted), green at 3600 (aged deleted, fresh kept). ruff clean,
check-windows-footguns --all clean, run_tests.sh ledger + AX files 29 passed.
Append-time `_delta` already strips unchanged paths, so on any ledger written
since it landed `compact_ledger` parsed every row and then `_rewrite_ledger`
wrote the identical file (up to 5 MB) on every maintenance sweep. Compare the
compacted payload with the raw bytes and return `(kept, len(raw), len(raw))`
without touching the file when they are equal; return semantics are unchanged.
PROOF: probe (probes/S5_compact_noop.py) with `os.replace` wrapped in a counter —
first sweep over 5 legacy fat rows: (5, 1545, 585), 1 os.replace; second sweep on
the now-compact file: (5, 585, 585), 0 os.replace. Existing ledger tests green
(29 passed in tests/tools/test_skill_ledger.py + test_computer_use_ax_walk_bound.py).
`_maintain_size()` now runs inside every `append_entry`, and `compact_ledger`/
`_trim_oldest` do read-bytes -> `os.replace` over the ledger with no cross-writer
lock, so a concurrent appender's O_APPEND write landing on the replaced inode was
silently lost — and `gc_blobs` then deleted that row's blobs. `_ledger_lock()`
(same `skill_file_lock` idiom as `_skill_mutation_lock`; fcntl/msvcrt, re-entrant,
no-op where neither exists) is held across the append AND the sweep, and inside
`compact_ledger`/`_trim_oldest`/`gc_blobs` for the `curator ledger --compact` path.
PROOF: tests/tools/test_skill_ledger.py::test_concurrent_appends_never_lose_a_middle_row
forces writer B to append between writer A's sweep read and its os.replace; every
row must be present or counted as trimmed. Green 2/2 (3.5 s). With `_ledger_lock`
made `contextlib.nullcontext()`: red 4/4 ("none silently lost").
Hoist the sole remaining `5 * 1024 * 1024` literal to `_DEFAULT_LEDGER_MAX_BYTES`
(fold-3 already collapsed the docstring/except copies into `_skills_cfg`), and
spell `_trim_oldest`'s early return as `if not dropped` — `dropped` is a length
difference and can never be negative. The `if lines else b""` payload guard now
lives in the shared `_rewrite_ledger` and stays: `compact_ledger` can legitimately
reach it with an empty row list.
PROOF: scripts/run_tests.sh tests/tools/test_skill_ledger.py
tests/tools/test_computer_use_ax_walk_bound.py green, no assertion changes;
`_max_ledger_bytes()` returns 5242880 via real import; ruff + footguns clean.
Both gates repeated the try / lazy-import / cfg_get(load_config_readonly(), "skills",
…) / debug-log / default shape. `_skills_cfg(key, default)` owns it; the callers keep
their bool()/int() coercion and defaults. Still a lazy import of the module attribute,
so tests that monkeypatch `hermes_cli.config.load_config_readonly` keep taking effect.
PROOF: tests/tools/test_skill_ledger.py 23 passed unchanged (incl.
test_config_gate_off_no_ledger_writes); mutation — `_skills_cfg` returns the default
without reading config → test_config_gate_off_no_ledger_writes,
test_auto_compact_triggers_at_threshold, test_trim_oldest_when_still_over_cap fail.
`_trim_oldest` and `compact_ledger` each carried the same join → `.<op>.tmp`
→ write_bytes → os.replace tail. Factor it into `_rewrite_ledger(path, lines:
List[bytes], op) -> bytes` at the bytes boundary (compact encodes its str rows
before the call and keeps returning `len(data)`); tmp naming per op is preserved.
PROOF: tests/tools/test_skill_ledger.py 23 passed unchanged; mutation — helper
skips `os.replace` → test_auto_compact_triggers_at_threshold and
test_trim_oldest_when_still_over_cap fail, 21 pass.
`ledger_enabled()` is the first call on every `append_entry`, and it still
paid `load_config()` (a deepcopy of the cached config) after the rest of the
append path had moved to `load_config_readonly()`. It only reads one key via
`cfg_get`, so the read-only load is safe. The one test that patches
`load_config` to turn the ledger off now patches `load_config_readonly` too.
PROOF: PYTHONPATH=<wt> import + `ledger_enabled()` → True;
scripts/run_tests.sh tests/tools/test_skill_ledger.py
tests/tools/test_computer_use_ax_walk_bound.py → 28 passed, 0 failed.
`_trim_oldest` used `str.splitlines()`, which also breaks on U+2028/U+2029/
U+0085. `append_entry` writes rows with `ensure_ascii=False`, so a row whose
evidence contains one of those code points was counted as two lines and
rewritten as two malformed physical lines — contradicting "lines in the
retained tail are never rewritten". Split on b"\n" (dropping the trailing
empty element), re-join with b"\n", and account sizes in bytes.
Sibling readers (`compact_ledger`, `gc_blobs`, `list_entries`) use the same
idiom pre-existing on main and are left for a follow-up.
PROOF: tests/tools/test_skill_ledger.py::test_trim_oldest_when_still_over_cap
gains one assertion (a U+2028 row survives a trim byte-for-byte); it fails
with the old `splitlines()` body ("At index 85 diff: b'\n' != b'\xe2'") and
passes with this change. 23 passed.
_max_ledger_bytes() runs on every ledger append (right after ledger_enabled()
already deep-copied the config) and _computer_use_cfg() now runs twice per
computer-use capture; neither mutates the result, so use
load_config_readonly() as hermes_cli/config.py documents for read-only hot
paths (skips the deepcopy, ~half the cache-hit cost). Tests that patched
load_config now patch load_config_readonly as well (test-only edit).
PROOF: real import under PYTHONSAFEPATH shows load_config() ==
load_config_readonly() (same dict shape; defaults 5242880 / 200 resolve);
tests/tools/test_skill_ledger.py + test_computer_use_ax_walk_bound.py +
test_skill_ledger_delta.py -> 39 passed, 0 failed.
The loop keeps lines[-1] unconditionally on its first iteration and
compact_ledger() (always run first by _maintain_size) strips blank lines,
so 'newest_json not in kept' could never be true; the size=2 slack was
also spurious since the join is exactly sum(len+1). Move _TRIM_LOW_WATER
above its first user.
PROOF: tests/tools/test_skill_ledger.py all green before and after (the
newest-survives assertion in test_trim_oldest_when_still_over_cap still
holds); probes/s5_trim_probe.py output unchanged.
_trim_oldest never parses lines: it drops the OLDEST ones whatever they
are, so 'malformed lines always survive' was false. Reword docstring,
curator.md and the kept test's docstring to the true guarantee: lines in
the retained tail are never rewritten or parsed, so a malformed line there
survives verbatim. No behaviour change.
PROOF: probes/s5_trim_probe.py (cap 4096, malformed first line, 4 fat
appends) -> old malformed survived: False, contradicting the old wording;
the kept test asserts a LAST-line malformed row, which is what the new
wording promises.
At the cap every append ran compact + trim + gc_blobs: trimming to exactly
`skills.ledger_max_bytes` leaves the file over the cap again on the very next
append. Trim to 80% of the cap instead (one rule derived from the cap, no
second config key), and run `gc_blobs()` only when the trim actually dropped
rows — the only path that can orphan a blob. `_trim_oldest` now reports the
dropped-line count and reuses `_read_ledger`, so an unreadable/undecodable
ledger skips the trim the same way it aborts compaction and blob GC.
Malformed lines are still kept verbatim.
The low-water-mark idea credits @fangliquanflq (#118690).
gc_blobs carried two byte-identical `warning("malformed ledger line; blob
GC skipped"); return 0, 0` blocks — one under `except json.JSONDecodeError`,
one under `if not isinstance(row, dict)` (tools/skill_ledger.py:350-357;
simplify reuse #2, efficiency note, re-gate G2 S2). Funnel the decode
failure into the type guard (`row = None`) so the abort exists once.
Behaviour is unchanged for both the [malformed-json] and [non-dict-row]
parametrizations: any line that is not a JSON object still aborts the
sweep with the same warning.
compact_ledger, gc_blobs and list_entries each stamped the same prelude:
read the ledger, catch (OSError, UnicodeError), warn "skill_ledger: ledger
unreadable (%s); <what>", bail (tools/skill_ledger.py:305-310, :341-346,
:412-417 — simplify reuse #1). Fold them into a private sibling
_read_ledger(what, *, quiet_missing=False) -> Optional[bytes] that reads
bytes, validates the UTF-8 decode and warns once unless quiet_missing and
the ledger is merely absent (list_entries: a fresh install has no ledger).
Returns bytes rather than str so compact_ledger keeps reporting the on-disk
size in bytes_before instead of re-encoding. Warning texts ("compaction
skipped", "blob GC skipped", "listing empty") are unchanged, so the
existing invariant tests pass verbatim.
list_entries() (tools/skill_ledger.py:414) caught (OSError, UnicodeError) but
returned [] silently, so a ledger that EXISTS but cannot be decoded made
`hermes curator ledger` print "ledger is empty" and `hermes curator rollback`
report "no ledger entry with id" with no hint that the file is damaged — while
gc_blobs()/compact_ledger() already warn on the same condition (:309, :345).
Now warn for any read failure except FileNotFoundError (a missing ledger is the
normal empty state). `%s`-lazy logger call, same "ledger unreadable" prefix.
Also adds the regression test the UnicodeError widening (2213062284) lacked:
re-gate mutation M8 reverted the except clause to `except OSError:` and the
whole suite stayed green. New tests write b"\xff" to ledger_path() and assert
list_entries() == [], get_entry() is None and "listing empty" in caplog; a
second test pins that a merely missing ledger stays silent.
Red/green: `except OSError:` -> UnicodeDecodeError (1 red); warning dropped ->
caplog assert red; warning made unconditional -> missing-ledger test red.
The gc_blobs() docstring (tools/skill_ledger.py:334) only mentioned malformed lines as a
reason to abort the sweep; this branch added the unreadable/undecodable-ledger abort
without updating it. State both conditions and the invariant (blobs are kept).
list_entries() caught only OSError (tools/skill_ledger.py:409), so a ledger whose bytes
are not valid UTF-8 raised UnicodeDecodeError out of `hermes curator ledger`, get_entry()
and rollback_entry() — the very corrupt state gc_blobs() and compact_ledger() now
tolerate on this branch. Widen to (OSError, UnicodeError) and return [], matching the
existing "unreadable ledger == empty" contract; the read path is read-only, so nothing
is lost by treating the file as having no entries.
Manual probe (HERMES_HOME=<tmp>, ledger bytes b"\xff"):
before: list_entries() -> UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff
after: list_entries() -> []; get_entry('abc') -> None;
rollback_entry('abc') -> (False, "no ledger entry with id 'abc'")
compact_ledger() returns (0, 0, 0) when the ledger cannot be read or decoded
(tools/skill_ledger.py:308-309) but did so silently, so `hermes curator ledger --compact`
reported a no-op compaction as if the ledger were simply empty. gc_blobs() already logs
"ledger unreadable (...); blob GC skipped" for the identical condition; emit the
symmetric warning here so the operator sees the ledger needs attention.
Manual probe (HERMES_HOME=<tmp>, ledger bytes b"\xff"):
WARNING skill_ledger: ledger unreadable ('utf-8' codec can't decode byte 0xff in
position 0: invalid start byte); compaction skipped
compact_ledger() -> (0, 0, 0); ledger bytes still b"\xff".
The regression test that pins this arrives in the following commit.
After json.loads() succeeds, gc_blobs() called row.get(...) unconditionally
(tools/skill_ledger.py:353). A syntactically valid but non-object line such as `[]`
or `"x"` raised AttributeError out of gc_blobs() and up through
`hermes curator ledger --compact`. list_entries() already tolerates such rows with an
isinstance(row, dict) check (~L414); mirror it here and abort the sweep with the same
"malformed ledger line; blob GC skipped" warning — a row we cannot interpret may still
hold blob references, so nothing may be deleted.
Test: `non-dict-row` param on the kept invariant test (ledger + b"[]\n"). RED with the
guard removed (AttributeError: 'list' object has no attribute 'get'), GREEN at head.
gc_blobs() aborts the sweep on a json.JSONDecodeError (tools/skill_ledger.py:351-352)
but did so silently, while the sibling read-failure abort a few lines above logs
"ledger unreadable (...); blob GC skipped". `hermes curator ledger --compact` therefore
printed "0 blobs removed" as if the sweep had run and found nothing. Emit the same
warning family so the operator learns the ledger needs repair before blobs can be GC'd.
Test: `malformed-json` param on the kept invariant test — (0, 0), blobs intact, and
"blob GC skipped" in caplog. RED with the warning removed
(assert 'blob GC skipped' in ''), GREEN at head.
compact_ledger() caught OSError on the read but decoded the bytes outside
the guard, so a ledger with invalid UTF-8 raised UnicodeDecodeError out of
`hermes curator ledger --compact`. Same escape class the gc_blobs() fix
closes; treat it identically (no-op, ledger left untouched).
The new read-failure branch in gc_blobs() returned (0, 0) silently, so
`hermes curator ledger --compact` printed "0 unreferenced blob(s) removed"
as if the sweep had run. Log a warning with the underlying error so an
operator can tell "nothing to reap" from "could not look".
The kept invariant test now also asserts the warning via caplog.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
Second half of #107539 on top of @liuhao1024's snapshot_paths filter: the same
TRANSIENT_DIRS set (venv, node_modules, caches, .git) now gates the whole-tree
snapshot (directory parts only, so a file named `venv` is still skill content),
and rollback carries a live nested venv back the way it already carries `.git`.
Nothing ever pruned the blob store; `hermes curator ledger --compact` now
deletes every blob no ledger entry references (98.9% of 47k blobs on the
reporting install). A malformed ledger line aborts the sweep — an unreadable
entry may still hold references.
Supersedes the curator_backup half of #100545 / #81669.
snapshot_paths() hashed every file under a skill dir with no exclusion
filter, so a stray venv/node_modules/__pycache__/.git under a skill was
copied content-addressed into ~/.hermes/.curator_backups/blobs/ on every
mutation — and nothing ever prunes that store, so the blobs grew without
bound (real install: 47k blobs / 1.3 GB in one day, 98.9% unreferenced).
Filter at capture: skip any file whose parent-chain component matches
_SNAPSHOT_EXCLUDE_DIRS (the same set proposed for the curator tarball in
transient dir is still captured — only files inside those dirs are
dropped.
The unreferenced-blob GC suggested in the issue is intentionally left
out; capture-side filtering stops the growth, and pruning existing
garbage is a separate, riskier change.
Follow-through on the 6-minute startup hang: the curator's storage grew without bound
and every pass paid for it. 1.9 GB skills tree = 1.1 GB snapshots + 653 MB ledger +
131 MB archive + ~50 MB skills.
- Snapshot only before a consolidation pass (the one that rewrites content in place).
The prune-only pass renames directories into .archive/ (its own undo) and every
mutation is ledgered; a whole-tree tarball before a rename was ceremony. Retention
(keep, default 5 -> 2) still runs on every pass, and a snapshot can no longer be its
own prune victim (same-second id reuse after a prune sorted below its sibling).
- Ledger entries drop paths whose hash is identical on both sides. rollback_entry
writes every before-path and removes after-only paths, so unchanged files were dead
weight: a 4,178-file skill cost 1.5 MB per patch. `hermes curator ledger --compact`
rewrites an existing log in place (655 MB -> 4.8 MB on the real one, ids preserved);
pre-rollback safety entries keep their full capture.
- archive_skill flattens under the skill NAME (mlops/training/accelerate is
huggingface-accelerate) and restore_skill falls back to the frontmatter name for
archives older builds flattened by directory name (6 of 52 were unrestorable).
- maybe_run_curator claims the due pass with an O_EXCL lock (stale after 1 h); two CLIs
launched 12 s apart both ran the weekly prune.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
Review finding on #112218 (major): `_skill_lock_path` opened `<skills>/.locks/<name>.lock`
before the name was validated, so `skill_manage(action='create', name='a'*300)` raised
OSError (File name too long) and a NUL name raised ValueError instead of the handler's
JSON error, and every rejected name ('../../etc', '') left a residue lock file.
- tools/skill_manager_tool.py: lock filename is sha256(basename).lock (fixed width, no
filesystem limit reachable; `foo` and `category/foo` still share one lock), the redundant
`_find_skill` rglob is gone, and `skill_manage` runs `_validate_name` on the name
(create) / basename (other actions) before the lock is opened.
- '.locks' joins the skills-dir exclusion sets (EXCLUDED_SKILL_DIRS, ledger
_NON_PACKAGE_TOPS, learning-graph/skill-commands skip parts, curator backup excludes).
- tests: 2 invariants in TestSkillMutationLock (rejected names -> JSON + no .locks residue;
digest-keyed lock shared across name forms), red on the old head.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Consolidation re-homes a skill's references/ / scripts/ out of the tree
before delete/archive, so the ledger captured only what was left
(files: 1 = SKILL.md) and `hermes curator rollback` restored a hollow
skill — the support files were only recoverable by hand out of the
pre-run .curator_backups tar.
The ledger's delete/archive/purge captures now complete themselves from
the newest curator skills.tar.gz: disk hashes win, the backup fills only
missing paths, tar members escaping the package prefix are rejected,
and every fill target stays under skills/ and HERMES_HOME. The same
fill runs at rollback time, so hollow entries recorded before this fix
still restore the complete package.
Wired at the four capture sites (skill_manage delete, archive_skill,
purge, record_mutation) and verified end-to-end: incident shape
(re-home -> delete -> entry has both files -> rollback restores both),
historical hollow entry repair, no-backup degradation, disk-hash
priority, and tar path-traversal rejection.
Read text files with the encoding utf-8-sig so a BOM at the start of a
file does not cause a Unicode decode error (Windows editors add BOMs).
Reconstructed from ethie/pm commits 48a32b135b + 013219e814 onto the
current upstream/main base: only the utf-8 -> utf-8-sig transforms were
carried (370 exact line pairs across 205 files); pm-rename hunks that
rode in the original commit were left to the pm-store commit, and
utf8sig hunks entangled with content changes ride their owning commit.
Rebuilt on ethie/pm-clean off ac6c8028e0 (upstream/main).
Tracker #79686 P3. Every skill mutation — curator, agent, or user — now
appends one entry to the append-only JSONL ledger at
~/.hermes/skills/.curator_ledger.jsonl, with per-file before/after
manifests whose contents are stored content-addressed (sha256-deduped)
under ~/.hermes/.curator_backups/blobs/.
- tools/skill_ledger.py: append/list/get, blob store, actor derivation
(curator|agent|user), single-entry rollback that takes a pre-rollback
safety entry first and FAILS CLOSED when that capture fails (consistent
with the whole-run tarball rollback hardening from #63366). Path
containment check so a hand-edited ledger can't write outside
HERMES_HOME.
- Hooked all three choke points: skill_manage() dispatch (all actors,
delete intent recorded via absorbed_into/archived evidence),
archive_skill()/restore_skill(), and curator auto-transitions (tagged
actor=curator via a ContextVar override).
- Ledger failures never block the mutation — telemetry, not a gate.
Config gate skills.ledger (default true).
- hermes curator ledger [--skill NAME] [--limit N] and
hermes curator rollback <entry-id> (whole-tree snapshot rollback
unchanged).
- Optional TTL purge of skills/.archive/: curator.archive_ttl_days
(default 0 = never) + explicit hermes curator purge, recorded in the
ledger with before-blobs so purges stay recoverable.
- Docs: curator.md sections on the ledger, single-edit rollback, and
archive TTL purge.
Curator invariants unchanged: only created_by:agent skills auto-transition,
never hard-delete autonomously, pinned exempt; foreground user deletes stay
hard-delete (and are now recoverable via the ledger).
Closes#45778, #50875. Tests adapted from #50261 by @yu-xin-c.