The retry passes ``route`` through unchanged, so the caller's ``sort`` (newest/oldest) still
drives ORDER BY; the comment claimed bm25 ranking unconditionally. Say what the code does
rather than force rank order — a user who asked for newest-first should get newest-first
from the relaxed hits too.
Tests: the seven ``_or_relaxed_query`` helper cases become one parametrized test; one DB
recovery test (exact untouched, paraphrase recovered, all partial rows, role_filter honoured)
and one negative (explicit NOT not relaxed, true miss stays empty, CJK route never reaches
the rewrite). Drops the upstream product name from module prose (credit stays in the PR body).
Port from nearai/ironclaw#7553 (Filter::FtsRanked): FTS5's implicit AND
between terms means a paraphrased multi-word query misses a stored
sentence that lacks even one of the words. When the exact-match search
and the substring fallbacks all return zero rows, retry the same
unicode61 FTS index with the terms OR-joined, ranked by bm25 so rows
covering more terms surface first.
Strictly additive: gated on a zero-result miss, so successful searches
keep exact-match semantics and ordering. Queries with explicit OR/NOT,
single-term queries, and CJK-routed queries are left untouched. Quoted
phrases relax as whole units.
Adapted for hermes-agent: implemented inside SessionSearchMixin's
zero-result fallback chain (after the CJK-bigram/trigram substring
retries) rather than as a separate filter variant, reusing the already-
built SQL/params so all source/role/sort filters apply to the retry.
Push the session-start bounds into search_messages so FTS LIMIT
cannot be filled by out-of-range hits. Covers FTS5, CJK, trigram,
LIKE fallback, and the unindexed-gap supplement.
Refs #86021.
gateway/run.py bridges the LAUNCH profile's sessions.cjk_fts / search_slow_ms
into HERMES_CJK_FTS / HERMES_SEARCH_SLOW_MS at import (and re-bridged them per
turn), and hermes_state_fts / hermes_state_search read os.getenv — so a served
secondary always got the default profile's values.
hermes_state_common.routed_sessions_setting() reads the routed profile's
config.yaml under a HERMES_HOME override and the env bridge when unscoped; both
consumers use it. The per-turn re-bridge is skipped inside a secondary's scope
so it can no longer write the default's slots from a routed turn.
optimize_fts() and vacuum() already refuse to run against a quarantined
handle (_db_corrupt / _db_replaced / _db_wal_generation_lost): both would
rewrite index/file pages in place, turning contained, diagnosable
corruption into an amplified one. rebuild_fts() never got the same guard,
despite being the more destructive of the two ("discards and recreates the
index data entirely", per its own docstring, vs. optimize_fts's segment
merge).
It's also independently reachable outside _execute_write's own quarantine
check: gateway/session_transcript.py's _rebuild_fts_once() calls
db.rebuild_fts() directly from the FTS-corruption transcript-retry path,
with no quarantine check of its own (only a WAL split-brain / foreign-holder
check, a different concern). A quarantined handle hitting that retry path
would run a full FTS rebuild — and commit it — on a corrupt, replaced, or
split-WAL-generation file.
Add the same self._raise_if_db_corrupt()/self._raise_if_db_replaced() pair
optimize_fts() already has, at the top of rebuild_fts(), before it enters
the cross-process rebuild admission.
Review feedback on #103657: the sibling clauses in the same statement
declare ESCAPE, and the trash enumeration two blocks up escapes its
underscores via replace. Use the same escaped ESCAPE form so every
underscore in the statement is a literal match instead of a
single-character wildcard. Behavior on the fixed schema is unchanged
(same enumeration split); this is consistency plus defense against
future lookalike table names.
Also reword the exclusion comment to the verified failure mechanism:
fts5 xRename renames the whole shadow family in one step, so sweeping
the cjk vtable aborts the loop on the next shadow entry and drags the
_config table (read by the vtable constructor) into the trash family.
The demote enumeration (name LIKE 'messages_fts_%') sweeps the
messages_fts_cjk vtable and its shadow tables into the fts_v22_trash_*
renames. Renaming the cjk vtable cascades to its shadow tables and breaks
the vtable constructor chain, so 'hermes sessions optimize-storage'
aborts with 'vtable constructor failed: messages_fts_cjk' on every DB
that carries both a legacy inline FTS layout and an established cjk
index (#103647). The cjk family is an independent v23+ index, not part
of the demoted legacy layout: skip it in the enumeration.
Two independent reviews of the seeded-create change found three more
places where the newly durable hidden row, or the new create-time copy,
was not handled by the same rule as the rest of the path:
- Message search (dashboard search and the session_search tool) had no
display_kind filter, so a hidden opening row matched a query the
person never saw. The shared search predicate now skips hidden rows.
- _seed_row left the fresh session row behind when the transcript copy
failed after the row was committed. The first prompt's retry copies
the whole seed, so a kept partial copy would be duplicated. The row
is now deleted when the copy did not complete, the compensation
_persist_branch applies to branch children; the first prompt then
starts clean.
- _live_session_payload (a resume that reuses a live session) reported
message_count as the raw history length while its messages array was
filtered. It now follows _resume_response: the stored size when
messages are omitted, else the wire count.
Tests: the two seeded-create tests now drive the first-submit path
through _persist_session_row_for_submit, the function prompt.submit
calls, and assert search and the reuse-live count; a third test pins
the rollback (no row after a failed copy, one copy after the retry).
A quarantined/replaced/split-generation handle must never run a full-file rewrite or an FTS5
'optimize': both read damaged or foreign pages and commit the result back, turning contained,
diagnosable corruption into an amplified one. Same guard _execute_write applies to every write.
Salvaged from #102092 onto current main: the _try_wal_checkpoint half landed via #106315's
_quarantine_reason(), so only the two rewrite sites remain.
Follow-up to #106317. Typing the steer row (display_kind="steer") for the renderer and the
alternation-repair guard collided with the convention that any display_kind on a user row means
scaffolding: is_user_originated_turn / _is_actionable_user_turn / split_user_originated_turn
returned False for it (tail anchoring, auto-focus, dispatcher views, resume counts) while
_is_real_user_message returned True (anchor restoration) — the two predicate families disagreed
on the same row, and list_recent_user_messages (/undo, /rewind) skipped it in SQL. A steer
carries full user authority; the steer kind is now whitelisted in all four.
Also: the pre-API drain's requeue tail reuses _requeue_pending_steer instead of a copy; the TUI
history projection compares against STEER_DISPLAY_KIND; the steer() docstring describes the row.
Keep projections query-free and timestamp/id neighbor ordering. Bound parameter batches at 500 and avoid scanning complete sessions with LAG/LEAD. Based on the N+1 analysis in #104296; no additional YAML cache or durability/freshness changes.
Co-authored-by: DevvGwardo <25094504+DevvGwardo@users.noreply.github.com>
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
On a fan-out-heavy install state.db reached 3.4 GB; 70% of message bytes
belonged to subagent sessions, and every one of those rows was also
indexed into messages_fts_trigram, whose shadow tables are ~2.6x the
text they cover (1,029 MB trigram vs 350 MB standard FTS on that DB).
session_search already hides source='subagent' sessions, so the
substring/CJK index bought nothing for them.
Extend the v29 cron exclusion: the messages_fts_trigram_src view, the
three sync triggers, and both deferred-backfill INSERT...SELECTs now use
one shared predicate (FTS_TRIGRAM_SESSION_SQL / fts_trigram_session_sql)
that skips sessions with source IN ('cron','subagent') or the
$._delegate_from creation marker (children spawned under a gateway turn
inherit the gateway's source). Compression/branch continuations carry
parent_session_id without the marker and stay indexed. Child rows remain
canonical in `messages` and fully indexed in the standard messages_fts
word index; explicit source_filter=['subagent'] CJK searches route to
LIKE like cron already did.
The v29 migration gate becomes `< 30` and reuses the same view-swap +
admitted rebuild, so existing installs purge historical child postings
once on open. Fresh DB with 2,000 x 2 KB child messages: 22.4 MB ->
12.5 MB (trigram shadow 10.09 MB -> 0.02 MB).
Clear zero-row rebuild markers so empty databases can finish teardown and
stamp the new layout. Extend the migration regression through close/reopen
recovery for both empty and populated databases.
Keep structured tool_calls searchable through the standard FTS index while
removing their repetitive JSON from the trigram projection. Reuse the
existing optimize-storage rebuild path for deployed v1 layouts.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
test_no_locked_readers_gate.py (#97676) parses hermes_state.py's SessionDB
class body with ast and flags any method that holds the writer lock
around a pure-read query — Pattern C, where every concurrent turn's
persistence convoys behind an unrelated read. #97676 converted 39 such
methods and closed detection blind spots for alias/variable-SQL readers.
But SessionDB is declared as
`class SessionDB(SessionSearchMixin, SessionSchemaMixin, SessionPortabilityMixin)`,
and the gate only ever opened hermes_state.py — it never parsed the three
mixin files those base classes are defined in, so a locked reader
declared there was structurally invisible to the scanner regardless of
how good the alias/variable-SQL detection got.
Applying the gate's exact scanning logic to the three mixin files
directly turns up 9 genuine pure-read methods still holding the writer
lock, none in #97676's converted list:
- hermes_state_search.py: _fts_teardown_trash_step, fts_optimize_available,
optimize_fts_storage, list_recent_user_messages
- hermes_state_portability.py: distinct_session_cwds, list_cron_job_runs,
_get_session_rich_rows_batch, list_skill_scaffolded_sessions,
get_first_assistant_text
_get_session_rich_rows_batch is a hot path: it backs list_sessions_rich's
compression-tip resolution and the web server's session-search hydration
across every gateway install — its own docstring already claims "same
read-your-writes guarantee as list_sessions_rich", but list_sessions_rich
was already using _read_ctx() (its guarantee comes from flush_token_counts()
before the read, not from holding the writer lock) while this method's
implementation never caught up to match.
Converted all 9 to `with self._read_ctx() as conn:`, the exact pattern
#97676 used, verified each is a genuine pure read with no hidden writes
by tracing every helper call it makes.
Extended the gate itself (_ALL_STATE_SOURCES) to scan all three mixin
files under their own class names, plus hermes_state.py, so this blind
spot can't silently reopen. Added a regression test
(test_scan_all_state_sources_visits_every_mixin_file) that plants a
synthetic violation in a mixin-shaped file and asserts the scanner still
catches it — a change that reverts the file list back to one file passes
the existing sabotage test but fails this one.
Mutation-verified: with the gate's new scope but the old (unconverted)
mixin sources, test_no_locked_pure_readers fails and names all 9 real
violations with correct file/line. Restored the fix; it passes clean.
Two pieces of PR #100130 (@HexLab98) re-applied on top of the orphaned-flock
break (894fc35337) and fail-closed admission (#100895) that landed since:
* `is_advisory_lock_contention` (hermes_state_common): only EAGAIN /
EWOULDBLOCK / EACCES / EDEADLK mean "another process holds the lock".
ESTALE / ENOTSUP / ENOLCK / EIO from flock or msvcrt.locking are
environment failures that polling cannot fix — `_acquire_db_flock` and
both Windows msvcrt loops (FTS rebuild admission, state.db repair lock)
now defer immediately with the real errno instead of burning the full
120s / holder timeout and then logging a fake "held by another process".
* `retry_deferred_fts_recovery` (hermes_state_schema): a SessionDB whose
open-time `_recover_stale_fts` deferred (foreign holders or busy rebuild
lock) stayed `_fts_stale` — LIKE-only search — until the process
reopened state.db. Short-lived CLIs reopen every run; the gateway opens
once and stays up for days, so the deferral was effectively permanent
(#100108). The retry runs from the EXISTING gateway housekeeping tick
(`_start_gateway_housekeeping`, 60s) against the shared SessionDB
instances via `hermes_state_registry.live_shared_session_dbs()`:
non-blocking admission (`fts_rebuild_admission(timeout_seconds=0)`),
bounded backoff 60s -> 1h, no new thread, still fails closed on live
holders. `fts_rebuild_admission` gains the `timeout_seconds` kwarg.
* WAL-reset warning names `sys.executable` so a "linked SQLite 3.45.1"
line can be matched to the interpreter that actually linked it
(#100108 point 3).
Deliberately NOT carried from #100130: the "leftover lock file = holder"
premise (a 0-byte lock file never blocked flock; the real cause was the
fork-inherited fd, fixed in 894fc35337) and the `_rebuild_fts_once`
one-shot rework.
Co-authored-by: HexLab98 <liruixinch@outlook.com>
Salvaged from PR #90734 by @Kyzcreig onto current main:
- hermes_state_search.py: post-commit FTS incremental merge failures
(including the bare SystemError CPython's sqlite3 layer raises under
cross-thread errmsg scrambling) are contained and logged instead of
escaping and making the caller replay an ambiguous, possibly-durable
write (exactly-once refinement by @yuzilongleif-collab).
- tests/state/test_writer_conn_thread_safety.py: live reader/writer race
hammer + AST sweep freezing the no-unlocked-writer-conn invariant.
On top: the sweep now also flags self._conn PASSED to helpers, which
caught one more live site on main — get_session_delete_targets handed
the shared writer connection to _collect_delegate_child_ids inside a
_read_ctx block, executing on it without self._lock. Routed to the
borrowed read connection.
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
Co-authored-by: yuzilongleif-collab <235949691+yuzilongleif-collab@users.noreply.github.com>
Follow-up to the salvaged #93200 commit. Factors the portable
_cross_process_repair_lock ownership pattern (msvcrt on Windows, flock on
POSIX, bounded 120s wait) into a cycle-safe shared primitive,
fts_rebuild_admission() in hermes_state_common, and routes EVERY full
structural FTS rebuild entry point through it:
- SessionSearchMixin.rebuild_fts() (replaces the POSIX-only, fail-open
30s flock from the original commit)
- _init_schema's trigger-repair rebuilds (_rebuild_fts_indexes /
_rebuild_legacy_fts_indexes) via _run_admitted_startup_rebuild
- _recover_stale_fts()
Fail closed: a caller that cannot acquire the authority DEFERS the rebuild
(FTS detached + durable stale breadcrumb, retried at next startup) instead
of proceeding into the exact concurrent-rebuild interleaving that
structurally corrupted state.db in production. Chunked deferred backfill
(fts_rebuild_step) intentionally stays outside the authority.
Adds spawned-process regression tests (real child process holding the real
lock file): holder blocks contender, deferral fails closed on both the
runtime and schema paths, release/holder-death permits the next owner, and
stale recovery completes after contention clears. Sabotage-verified: 4/6
tests fail with the admission forced open.
When two Hermes processes (e.g. gateway + serve) detect FTS corruption
simultaneously, both run rebuild_fts() on the same database file in
parallel. rebuild_fts() only holds an in-instance threading lock, so the
concurrent rebuilds collide on write and structurally corrupt the
database ('file is not a database' / 'database disk image is malformed').
This happened twice in production (2026-08-15 and 2026-08-23), each time
requiring a full page-level salvage of state.db: sessions b-tree
clobbered, 5507 messages recovered row-by-row.
Fix: acquire an exclusive fcntl.flock on <db_path>.fts_rebuild.lock
before rebuilding, with a bounded 30s wait. The SQLite writer lock
remains the final backstop. POSIX-only; no-op elsewhere.
Every search route in _search_messages_impl (FTS, CJK bigram, trigram,
LIKE fallback, rebuild-gap supplement) selected m.content, then the
result tail popped it unread. On DBs with multi-MB tool rows, each
search read and materialized up to `limit` full rows only to discard
them. Snippets come from snippet()/substr() in SQL and the context
window is re-fetched by id, so no code path ever read the column.
Drop the column from all six SELECT lists. Returned dicts are
unchanged: content was never part of the public result (the pop ran
before return), and tests/test_hermes_state.py already documents that
contract.
(cherry picked from commit d0c3af167e7dd4eb18e1bea29ba107a94911ea24)
SessionDB.close() ran `PRAGMA wal_checkpoint(TRUNCATE)`. Every cron
run_agent opens and closes its own transient SessionDB, so on a busy
fleet this fired a full WAL reset many times an hour, racing the
gateway's long-lived writer on a large WAL database and tearing hot
B-tree pages -- structurally the same corruption this module's own
periodic checkpoint was already switched to PASSIVE to avoid (#45383).
Only close() and two manual-maintenance paths still used TRUNCATE.
Route every checkpoint on the shared state.db through PASSIVE:
- close() (hermes_state.py)
- pre-VACUUM in vacuum() (hermes_state.py)
- post-optimize-storage (hermes_state_search.py)
PASSIVE never resets/truncates the WAL and never takes the exclusive
checkpoint lock, so it cannot lose a transient closer's race with the
live writer. The WAL is instead bounded by `journal_size_limit` and the
writer's natural post-checkpoint reset. TRUNCATE belongs only on a
sole-opener/quiescent connection (e.g. offline maintenance); this change
does not try to detect that -- PASSIVE is the safe default.
Diagnosed as the root cause of three state.db B-tree corruptions in
2026-08: damage localized to the hottest-written pages (gateway_routing
and the sessions indexes), with whole zero-filled pages still live and
off the freelist -- the checkpoint/reset-race signature, not disk or
application SQL.
Tests: tests/test_wal_checkpoint_strategy.py now asserts PASSIVE at
close(), before vacuum(), and after optimize_fts_storage() VACUUM;
tests/test_hermes_state.py asserts close() likewise. Focused run:
226 passed.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Closes the residual the contributor's own triage comment flagged: % was
excluded from the special-char class to protect the CJK LIKE fallback,
but a non-CJK query never reaches that fallback (is_cjk gates it), so
'50%' still hit MATCH raw and silently returned zero results. Strip %
whenever the sanitized query contains no CJK; the CJK path keeps its
pre-existing contract. Regression tests for both directions.
_sanitize_fts5_query's strip step only removed +{}():"^ . Every other
character FTS5's grammar rejects outside a quoted phrase reached MATCH
raw and raised, and — as the step's own comment says about the colon it
was fixed for — the execute site swallows that into zero results. Session
search silently found nothing for ordinary queries:
it's fts5: syntax error near "'"
gateway/run.py fts5: syntax error near "/"
user@host fts5: syntax error near "@"
a,b fts5: syntax error near ","
why? fts5: syntax error near "?"
e=mc2 fts5: syntax error near "="
Complete the class and assemble it with re.escape, because written as a
regex literal the backslash was eaten as an escape and never made it in
(C:\path\file still raised after the first pass).
Measured against a real FTS5 table over 651 realistic queries:
373 unparsable before, 77 after. The remainder is leading/trailing "." and
"-", which #43889 already covers.
% is deliberately left in: the CJK path falls back to a LIKE search that
needs it literal and escapes wildcards itself, so stripping it widened
those queries onto unrelated rows (test_cjk_like_escapes_wildcards).