Commit Graph

34040 Commits

Author SHA1 Message Date
ethernet
86efc1f945 fix(release): allow explicit environment clears in commit bundles
An inherited HERMES_HOME can defeat a test bundle's data-directory suffix.
Older Windows installers also persisted that variable in the user registry.
This gives the app fresh UI state while its backend reads existing sessions.

Add --bundle-unset NAME, encoded as null in the existing bundle environment
object. Apply each clear as an explicit empty value before module startup.
Do not restore an explicitly empty HERMES_HOME from the Windows registry.
Ordinary defaults still preserve runtime overrides.

Verified the release parser, builder handoff, compiled startup ordering,
registry opt-out, and child environment with focused regression tests.
A native Windows probe passed with an inherited home. No MSIX was rebuilt.
2026-09-11 15:59:08 -04:00
ethernet
2b1312d4ca Merge branch 'ethie/pm-binary' into ethie/pm-clean 2026-09-11 15:13:11 -04:00
ethernet
162c5b92d5 feat(desktop): name commit and canary builds in the product display name
Commit builds now show 'Hermes Agent <sha7>' (e.g. Hermes Agent abc1234)
and canary builds 'Hermes Canary' / 'Hermes Light Canary' / 'Hermes Agent
Canary' as the OS-visible product name, so side-by-side installs and
per-commit artifacts are readable at a glance.

Display-only by design: appId, appNamePascal, and msixAppIdWithOrg are
unchanged, so a canary MSIX still updates in place over stable and
userData / single-instance sharing with the stable install is unaffected.
bundle-electron-main.mjs derives the commit from the install stamp
(source='commit-build') so the baked runtime identity matches the
packaging identity.
2026-09-11 14:58:11 -04:00
ethernet
3301c31ff8 fix(desktop): restore backend lifecycle and update build contracts 2026-09-11 14:35:25 -04:00
ethernet
4e45b78311 fix(release): inline Windows reserved-path check for pre-3.13 runners
ntpath.isreserved was added in Python 3.13, but release workflow legs
(commit-builds-summary, builds-table, builds-pending) run bare python3 on
ubuntu-24.04, whose system Python is 3.12. render-builds-table.py crashed
with AttributeError before rendering the expected-binary matrix.

Port the CPython ntpath reserved-name semantics (device stems incl.
superscript COM/LPT forms, trailing dot/space per component) into
_is_windows_reserved() in scripts/releases/r2.py so the release transport
stays self-contained on whatever python3 the runner provides. Verified
byte-parity against real ntpath.isreserved on a 239-case corpus.
2026-09-11 14:30:56 -04:00
ethernet
06ef8ce786 feat(paths): suffix default agent and desktop data directories 2026-09-11 14:25:18 -04:00
ethernet
7d326adf9d feat(release): bake explicit environment defaults into commit bundles 2026-09-11 14:25:12 -04:00
ethernet
fea2858c99 merge: unify shared product builders, caches, and Windows prerequisites
Merge ethie/shared-product-builders with the CI dependency cache and native Windows setup work. Preserve UTF-8 diagnostics in the shared Python environment runner. Pass a persistent cache through isolated native staging and PM-runtime construction. Reuse one Windows prerequisite installer from source setup, native adapters, and CI, preserving Rust homes across HOME isolation.

Verified 85 targeted Python tests (5 host skips), 18 JavaScript tests, workflow validation, and scoped lint/typecheck. On native Windows ARM64, five prerequisite contracts passed and the actual shared provider reused OpenSSL, compiled its header with MSVC, and retained Rust under isolated HOME. Full signed distribution builds and live Actions cache transfer remain CI verification.
2026-09-11 13:45:05 -04:00
ethernet
4cc2b7bab5 fix(ci): remove legacy uv cache compatibility 2026-09-11 13:32:32 -04:00
ethernet
018d2b39d8 fix(pm): prepare platform trust before bootstrap downloads
Standalone Python cannot locate the system CA bundle on this NixOS host.
Use the shell-staged uv to prepare the independent PM runtime before PM
fetches managed Python. Declare and lock truststore in that runtime, then
activate it before CLI and worker imports construct HTTPS clients.

Make setup's awk pin reader follow object nesting rather than indentation.
Use the same reader for tool versions and artifact fields.

Verified cold activation with CA overrides removed, pinned Python and uv
downloads, pm doctor, and a public worker HTTPS install. The targeted suite
passed 56 tests with one Windows-only skip. The TLS regression fails when
truststore is installed but entrypoint activation is removed. Bash syntax
and Ruff checks passed. The full suite was not run.

Two additional setup-toolchain tests fail on unchanged HEAD because their
fixtures reach the real home before home isolation. CI bootstrap unification
is not part of this change.
2026-09-11 13:31:41 -04:00
ethernet
fffccbb2ae fix(setup): prepare native Windows ARM64 build dependencies
Source activation could not build cryptography because the setup shell
could not discover the installed OpenSSL development libraries.

Configure Visual Studio ARM64, Clang, Rust, and static OpenSSL before PM
runs. Reuse installed tools and install missing prerequisites. Select a
classic vcpkg with a ports tree and use an explicit installation root.
Report damaged shared libraries without deleting the shared installation.

Verified native PowerShell activation and deactivation on Promise, then
warm activation with no new dependency generation. Cryptography imported
with static OpenSSL. Real vcpkg checks covered installation into a path
with spaces, warm reuse, manifest mode, and damaged-package rejection.
The canonical Windows runner passed the helper and output-encoding tests.

Fresh Visual Studio and Rust installation were not exercised because
Promise already had those toolchains installed.
2026-09-11 13:29:45 -04:00
ethernet
caf27c01b9 fix(pm): decode captured build output as UTF-8
Windows defaulted captured uv output to CP1252. A UnicodeDecodeError in
the pipe reader hid the OpenSSL build failure and left an empty diagnostic.
Decode uv and npm output as UTF-8, replacing malformed bytes while keeping
the exit status and build error.

Real subprocess tests cover stdout, stderr, legacy locale defaults, and
malformed output. The encoding tests passed on native Windows ARM64 and
Linux through scripts/run_tests.sh.
2026-09-11 13:29:45 -04:00
ethernet
493ae9daa3 fix(ci): reuse uv wheels and save caches after bundle failures 2026-09-11 13:28:15 -04:00
ethernet
0a3a189235 refactor(build): remove superseded launcher templates 2026-09-11 13:17:10 -04:00
ethernet
1bf588234c refactor(build): share product recipes across distributions
Build TUI, web, desktop UI and runnable agent products from explicit
prepared inputs. Keep dependency preparation separate from distribution
packaging, with PM and native builds sharing uv environment construction.

Docker copies compiled frontend products instead of build dependencies.
Nix retains uv2nix environments and consumes shared assembly through store
references. Native desktop and Termux use the same launcher and frontend
contracts. Preserve the independent PM runtime and source imports from
arbitrary working directories.

Keep failed frontend builds from replacing the previous product, reject
source/output overlap, and bound dependency-process output draining.
Include hermes_wisdom in the Nix wheel: real CLI smoke tests exposed its
missing package declaration on the base revision too.

Verified focused Python and JavaScript suites, Docker build/runtime checks,
Nix desktop and CLI/ACP checks, standalone TUI and packaged Electron PTY,
and real full-Chromium interaction. Native signed installers, Android device
installation and the full repository suite remain CI verification.
2026-09-11 13:16:55 -04:00
ethernet
284dbaf537 fix(pm): isolate bootstrap dependencies and unify YAML on ruamel
Activation reaches plugin discovery before the application dependencies
exist. Give PM its own locked Python project and runtime so it can install
or repair the application without importing that dependency tree.

Keep PM outside the application workspace. A shared uv workspace resolves
the application graph and cannot provide this isolation. Route mutations
through an isolated worker and preserve transaction callbacks, cancellation,
custom package registrations, and correlated receipts.

Use the same runtime builder for source installs and packaged payloads.
Keep offline wheelhouse support in that builder. Nix builds the independent
PM lock as a separate derivation. Refuse lazy-disabled bootstrap before
installing tools or dependencies.

Move first-party YAML readers and writers to ruamel. Keep the application
lock's transitive PyYAML requirements for third-party packages.

Verification:
- Focused canonical Python suite: 177 passed, 1 host-gated skip.
- Electron backend probes: 12 passed. Electron typecheck passed.
- Both uv locks, scoped lint, Bash syntax, and whitespace checks passed.
- Cold activation, corrupt-app repair, offline staging, and relocation ran.
- Built and exercised the Nix PM runtime and standalone YAML merge script.

Six broader caller test files retain the same 24 failing test IDs as an
archive of HEAD. The existing real-home guard blocks those tests before
they can exercise the affected paths. No full-suite pass is claimed.
Native Windows signing and full Bionic package execution remain unverified.
2026-09-11 12:23:51 -04:00
ethernet
bfabc23f7f fix(desktop): reconcile runtime wiring after the PM merge
The merge combined old callers with newer lifecycle and update modules.
It also dropped native handlers while keeping their preload methods.
Type declarations alone could not repair those runtime failures.

Restore bounded backend teardown and retain failed-stop ownership.
Use API-only passive checkout checks with a daily disk cache, and pass
manual refresh requests through the updater strategy. Keep the shared
About UI and restore onboarding, feature flags, and notification wiring.

Verified all workspace typechecks, lint on the changed desktop files,
focused UI and Electron tests, and the development bundle. A headless
Electron smoke test exercised the real main process, preload, and native
IPC. A separate test exercised update checks with a linked git worktree,
loopback HTTP, and the disk cache. The full repository suite was not run.
2026-09-11 11:56:35 -04:00
ethernet
8f6d98e4c3 fix activation of devenv, use /usr/bin/env bash everywhere 2026-09-11 11:17:18 -04:00
ethernet
3c2e1bd452 fix(desktop): disable MSIX virtualization 2026-09-11 10:58:02 -04:00
ethernet
b3bfc3afe5 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/electron/backend-connection-state.test.ts
#	apps/desktop/electron/backend-connection-state.ts
#	apps/desktop/electron/backend-exit.test.ts
#	apps/desktop/electron/main.ts
#	apps/desktop/electron/pool-spawn-coordinator.test.ts
#	apps/desktop/electron/pool-stop.ts
#	apps/desktop/electron/preload.ts
#	apps/desktop/src/app/settings/about-settings.tsx
#	apps/desktop/src/app/updates-overlay.tsx
#	apps/desktop/src/global.d.ts
#	apps/desktop/src/store/notifications.ts
#	apps/desktop/src/store/updates.ts
#	gateway/config_loader.py
#	hermes_cli/banner.py
#	plugins/platforms/dingtalk/adapter.py
#	tests/hermes_cli/test_plugins_cmd.py
#	tests/test_live_system_guard.py
#	tui_gateway/server.py
#	website/docs/user-guide/desktop.md
2026-09-11 09:36:04 -04:00
Teknium
a2413495ad fix(state): fts_trigram_session_sql qualifies model_config under the json_valid guard
The _sql_json_extract wrapper removed the `COALESCE(model_config` text the
alias string-replace keyed on, so the deferred-backfill SELECT joined an
unqualified `model_config`. Build the predicate from the alias directly and
derive the unaliased constant from it, so the two can never disagree.
2026-09-11 06:24:54 -07:00
teknium1
46afbfec10 fix(state): route the remaining model_config marker reads through _sql_json_extract
Sibling widening of the #101726 salvage: FTS_TRIGRAM_SESSION_SQL (trigram view/triggers/backfill),
the v16 delegate-tagging data migration and reopen_session's legacy reset-child stamp still called
json_extract() on the raw model_config cell, so one malformed JSON row could still abort FTS
maintenance, a schema migration or /resume of a reset child. Zero raw model_config json_extract
reads remain in hermes_state_*.py.
2026-09-11 06:24:54 -07:00
teknium1
c4a8572d68 test(state): pin corrupt-row robustness — one bad timestamp/oversized IN-list never kills the command
Three invariants over real SQLite fixtures (TEXT and 8.4e252 timestamps written straight into the
REAL columns; 1200 ids under a 999-variable ceiling via setlimit): list/export/insights complete
and name the corrupt session in a WARNING; writers never persist an out-of-window timestamp;
prune/delete_sessions succeed with zero orphaned messages. All three fail on origin/main.
2026-09-11 06:24:54 -07:00
Mi55ed
25670cd9e3 fix(state): batch large session cleanup queries below SQLite's variable limit
prune_sessions(), delete_empty_sessions() and prune_empty_ghost_sessions()
built one IN (?, ..., ?) clause containing every selected session id (the
parent-orphaning UPDATE), so cleaning more than SQLITE_MAX_VARIABLE_NUMBER
sessions failed atomically with "too many SQL variables". The per-row
DELETE loops that followed are folded into the same 900-id batches.

Same single _execute_write() transaction; only the binding is split.

Hand-ported from PR #100658 (targeted the pre-decomposition hermes_state.py
god file; the methods now live in hermes_state_maintenance.py /
hermes_state_sessions.py). The one-pass transcript-directory sweep from
that PR is not ported (out of scope for the variable-limit bug). Authored
by @Mi55ed; ported under --author.
2026-09-11 06:24:54 -07:00
Michael Steuer
acfda37598 fix(sessions): chunk IN-list ids in bulk delete to avoid 'too many SQL variables'
`hermes sessions prune --source cron --older-than 14` on a store with ~60K
cron sessions (~50K matches) died with sqlite3.OperationalError: too many
SQL variables. SessionDB.delete_sessions, _collect_delegate_child_ids and
_delete_delegate_children each bound the full id list into a single
IN (?,?,...). SQLite caps bound parameters at SQLITE_MAX_VARIABLE_NUMBER
(999 on < 3.32, 32766 after), so any bulk delete above that failed outright.

Chunk every IN list (`_id_chunks` / `_SQL_IN_CHUNK` in hermes_state_common,
900 ids; the delegate walk binds each id twice so it chunks at half). Same
transaction, same cascade/orphan contract; only the parameter binding is
split.

Hand-ported from PR #102679 (targeted the pre-decomposition hermes_state.py
god file; the functions now live in hermes_state_sessions.py). Authored by
@mssteuer; ported under --author.
2026-09-11 06:24:54 -07:00
teknium1
496eb13bd7 fix(state): one corrupt timestamp row no longer kills sessions list, export or insights
SQLite dynamic typing lets a TEXT cell ('not-a-timestamp'), inf/nan or a
garbage double (8.4e252 salvaged from a damaged page) sit in a REAL
timestamp column. Every reader called datetime.fromtimestamp()/float
arithmetic on the raw cell, so ONE bad row raised TypeError/OverflowError
out of the row loop and took down the whole `hermes sessions list`/browse
table (#102399), all three exporters — JSONL/MD, QMD, HTML (#102352) —
and `hermes insights` (#99959).

Fix the class with ONE helper, hermes_cli.timefmt.coerce_epoch(): a
stored cell becomes float epoch seconds inside a sane 1970..2103 window
or None after a WARNING that names the session id. Every reader routes
through it — relative_time (list/browse/resume picker), format_epoch
(prune/candidates tables), the three exporters' timestamp formatters,
insights' _get_sessions/_day/period range — so a bad row renders as
'?'/'N/A'/raw text for that one cell and the command completes.

Write side: hermes_state_messages._coerce_timestamp (append_message,
append_messages_batch, import) and the import path's started_at now use
the same window, so a new out-of-range timestamp falls back to now()
instead of being persisted — new bad rows cannot be written by Hermes.

Reported-by: #102399, #102352, #99959 reporters; kokhlo's insights
analysis pointed at every reporting site, not just line 860.
2026-09-11 06:24:54 -07:00
liuhao1024
df1388e77b test(tools): extend the delegation fd-leak injection to the schema replay
The schema initializer now replays SCHEMA_SQL through executescript (the
single-authority path from #94701's follow-up), which bypasses the
execute()-level DDL failure injection — the regression stopped raising.
Bind the same simulated failure onto the executescript path so the
connect-close-on-init-failure contract stays pinned for both replay
mechanisms.
2026-09-11 06:24:54 -07:00
liuhao1024
a6d65cdd09 fix(state): single durable-shape authority for async_delegations
Review follow-up on #94701: the delegation tool's _initialize_schema
still carried its own CREATE TABLE + ALTER column list for
async_delegations, leaving a second durable-shape authority even with
the column declared in SCHEMA_SQL. Its legacy ALTER added
origin_session_id as bare TEXT (nullable, no default); reconciliation
repairs missing column names only, so a database first opened through
the tool kept a non-canonical shape forever (#94691).

Remove the private DDL entirely. The tool's initializer now calls a new
reconcile_state_schema() in hermes_state_schema, which replays the
canonical SCHEMA_SQL (idempotent CREATE IF NOT EXISTS for every table,
canonical indexes included) and reuses SessionDB's declarative
_reconcile_columns for missing-column backfill — one reconciliation
implementation, one authority. Because _parse_schema_columns
reconstructs each column's full constraint expression (type, NOT NULL,
DEFAULT), the tool-first legacy path now adds origin_session_id as
TEXT NOT NULL DEFAULT '' — the canonical shape — and SQLite backfills
existing rows with the '' default.

Opening-order regressions compare FULL PRAGMA table_info metadata
(type, notnull, dflt_value, pk) plus the canonical index set across
fresh SessionDB→tool, legacy→SessionDB, and legacy→tool→SessionDB,
each preserving a pre-existing legacy delegation row.
2026-09-11 06:24:54 -07:00
liuhao1024
0bb27d37cf fix(state): declare origin_session_id in the canonical async_delegations schema
The delegation tool carries its own CREATE TABLE for async_delegations
(tools/async_delegation.py _initialize_schema) plus a lazy ALTER TABLE
ADD COLUMN for the tables it finds already existing. Its column list
had drifted ahead of the canonical SCHEMA_SQL: origin_session_id
(raw api_server session id of the originating request, the wake
self-post target) existed only through the tool's lazy path, so two
databases at the same schema_version had different
async_delegations shapes depending solely on whether the delegation
tool had ever run. Rebuild/replay pipelines that reconstruct state.db
from the canonical schema then hit the column with no version gate to
explain it (#94691).

Declare the column in SCHEMA_SQL with the same TEXT NOT NULL DEFAULT ''
shape the tool uses. Fresh installs now carry it canonically; the
declarative _reconcile_columns backfills it into legacy databases on
the next writable open (same pattern as earlier additive columns); the
tool's lazy ALTER keeps serving pre-reconciliation databases. The two
schema authorities now agree, pinned by a test that runs the tool's
initializer over a canonical database and asserts the shape is
unchanged.

Fixes #94691
2026-09-11 06:24:54 -07:00
Efe Büken
a239c4f811 fix(state): tolerate malformed session marker JSON 2026-09-11 06:24:54 -07:00
Xipong
78f85112d9 perf(state): batch message hydration in export_all 2026-09-11 06:24:54 -07:00
ethernet
e86d31fade fix(pm): refresh artifacts within the same minor version
BtbN publishes new FFmpeg builds without changing the version number.
The shared-minor comparison therefore reported stale artifacts as current.

Compare advertised artifact URLs during resolution and hash changed URLs
when applying the update. Keep dry-run checks metadata-only and preserve
pins for targets without an update source, including Termux.

Verification: 41 focused tests passed. The regression exercises real
archive downloads, installation, retained target pins, and a second
update that performs no writes. Native ARM64 FFmpeg also passed a real
16 kHz audio encode after installation.
2026-09-11 09:24:18 -04:00
teknium1
16e4496d90 fix(agent): corrupt-session recovery guidance names the session's own state.db
The corruption explainer filled `{db_path}` from `_default_db_path()`, the
process default. A Desktop `serve` backend launched on the root home hosts
named-profile sessions whose SessionDB is `profiles/<name>/state.db`, so the
operator was told to inspect/repair a different profile's database. Pass the
agent's own `_session_db.db_path` from the turn finalizer; the process
default remains the fallback for agents without a bound store.

Reported in #105887.
2026-09-11 06:24:11 -07:00
teknium1
df0eed4f6b fix(session_search): a bare session id never reads another profile's state.db
Reading a session by id that missed the caller's store fell through to
_locate_session_db(), which opened every profile's state.db read-only and
returned the first owner's full transcript — no opt-in, no profile named, and
the miss path even fired after an explicit non-matching profile= read. Any
caller holding an id (ids appear in logs and tool output) could read a
foreign profile's conversation. Profiles are isolated islands by design.

A miss now stays a miss, with a hint to name the owning profile
(profile=<name> / @session:<profile>/<id>), which remains the sanctioned,
explicit cross-profile read. The schema eval runner no longer needs to fake
the scan.

Reported by the #106761 filer; reproduced by @kokhlo. Refs #87779.
2026-09-11 06:24:11 -07:00
teknium1
5a720bbb1b test(tui): launch-handle pin test opens a real registry SessionDB
The salvaged #102534 test patched hermes_state.get_shared_session_db, a seam
main dropped (server._get_db now calls hermes_state_registry.acquire), so the
fixture errored at setup and the file reported 0 passed. Assert on the real
handle's db_path and on the foreign home staying untouched instead of on a
fake factory; red with the pin reverted, green with it.
2026-09-11 06:24:11 -07:00
HexLab98
342705135a test(tui): cover launch SessionDB binding under profile override windows
Add a #102526 regression and adjust resume ownership leak filtering now that
the shared launch handle carries an explicit state.db path.
2026-09-11 06:24:11 -07:00
HexLab98
3c3ca061a6 fix(desktop): pin launch SessionDB handle to the launch home
The lazy _get_db() singleton followed get_hermes_home(), so a first touch
inside the multiplex cron ticker's per-profile override window permanently
bound the default backend to another profile's state.db (#102526).
2026-09-11 06:24:11 -07:00
Sora-bluesky
3a350508f7 fix(cli): wait for SessionDB teardown when deleting a profile
close_all_under returned after the last release dropped the generation
and before the physical close finished, so rmtree still saw the open
handle. Wait directory-matching teardown barriers the same way close_all
does.
2026-09-11 06:24:11 -07:00
Sora-bluesky
f1ee8e0f39 fix(cli): release SessionDB handles when deleting a profile
delete_profile already force-closes holographic memory_store.db in this
process, but the shared SessionDB registry kept state.db open. Recreate
then failed with a replaced/locked database. Close every shared handle
under the doomed directory, same contract as MemoryStore.release_all_under.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-11 06:24:11 -07:00
teknium1
d956e05e5e chore: contributor mappings for the statedb-wal salvage (gaoanze888 work email, nikkoxgonzales noreply) 2026-09-11 06:23:23 -07:00
teknium1
e00a21cd39 fix(state): retry a transient SQLITE_IOERR on the pooled read path before surfacing it
Since 0.21.0 reads go through mode=ro pooled connections. A read-only OPEN already
rides out the millisecond WAL transition window (checkpoint / WAL reset / frame flush
by a sibling process; the ro reader cannot rewrite the -shm index) with a bounded retry
(#100436), but a WARM pooled reader hitting the same window while its SELECT executes
propagated `disk I/O error` straight out of get_session(): 37 identical tracebacks on a
multi-process WSL2 ext4-on-vhdx install, each followed by "compression session recovery
failed", with quick_check=ok (#100871). The reporter's A/B shows the operator
workaround (journal_mode=delete) collapses read throughput ~30000x, so the flake has to
be absorbed on the read path.

_read_one/_read_all now replay the idempotent statement within the existing read-only
IOERR budget (3 x 50 ms) on the SAME connection -- close+reopen would cancel this
process's POSIX locks for every sibling connection -- and a persistent IOERR still
propagates. No quarantine: EIO on a read is busy, not broken. Every SELECT in the
SessionDB siblings (63 call sites) reaches the pool through these two helpers, so the
class is covered without a wrapper type.

Same-connection retry per #100882's analysis (@fangliquanflq); #100883
(@Sahilvishnaliya) diagnosed the missing recovery in the 0.21.0 read pool.

Fixes #100871.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: Sahilvishnaliya <222165401+Sahilvishnaliya@users.noreply.github.com>
2026-09-11 06:23:23 -07:00
teknium1
295edf2557 fix(state): gate the lock-free read pool on a confirmed WAL header, not the assumed mode
apply_wal_with_fallback() reports "wal" in two indeterminate cases -- the vulnerable-
SQLite gate (_apply_delete_for_wal_reset_bug) and the non-vulnerable probe-unknown path
(a7f2a593d1) -- meaning "touched nothing, the connection inherits the header's mode".
SessionDB turned that assumption into `_wal_active=True`, which enables the mode=ro read
pool that skips `self._lock`. On a file that is really in rollback-journal mode those
readers race the writer with a 5s busy timeout and no retry: random SQLITE_BUSY read
failures for the instance's lifetime (#86515).

Confirm the header on the freshly opened connection before enabling the pool. When the
probe is still blocked, reads queue on the writer connection under the lock -- slower,
never wrong. Every other apply_wal_with_fallback caller ignores the return value, so the
consumer is the right place to gate; changing the return contract to Optional across
15 call sites (#87044's shape) is not needed.

Live repro: DELETE-mode file, sibling holding BEGIN EXCLUSIVE during open ->
before: _wal_active=True and _checkout_read_conn() hands out a pooled mode=ro conn;
after: _wal_active=False, reads take the locked writer path.

Fixes #86515. Based on the analysis in #87044.
Co-authored-by: QDung210 <dqdung205@gmail.com>
2026-09-11 06:23:23 -07:00
teknium1
9f9ec647f5 fix(gateway): re-check the live store before broadcasting the state.db warning
_send_session_db_warning_notifications() broadcast the error recorded at startup
without asking whether it was still true. A startup `database is locked` routinely
clears while the adapters are still connecting (another profile's open, a `hermes
sessions` one-shot, a slow SMB lock release), so the home channels were told the
store was unavailable when it had already healed — and a warning that is wrong once
is ignored the next time it is right.

The RecoverableHandleCache opener already clears `_session_db_init_error` on
recovery; the broadcast now drives one open attempt through it first and only warns
when the error is still standing. A store that is genuinely still down warns exactly
as before.

Fixes #108031.
2026-09-11 06:23:23 -07:00
teknium1
045704ce00 test(state): port the #107411 dentry tests to the shared WAL-generation harness
main replaced the file-local _make_db/_require_wal/_unlink_sidecars helpers with
tests/hermes_state/_wal_generation_harness (make_db/require_wal/lose_sidecars) after
PR #107411 branched; the cherry-picked tests referenced the old names and failed at
collection with NameError. Fixture-only change; the assertions are unchanged.
2026-09-11 06:23:23 -07:00
nikkoxgonzales
4d9ec9be50 fix(db): survive DELETE fallback failure in WAL setup 2026-09-11 06:23:23 -07:00
gaoanze
031d3760d0 fix(state): distinguish retired WAL recovery guidance
Classify deleted WAL generations separately from main-file replacement and point operators at the captured-generation manifest and mode-aware recovery path.

Co-authored-by: crazyief <8566250+crazyief@users.noreply.github.com>
2026-09-11 06:23:23 -07:00
ca-shrimp
3a3ec6eeaf fix(state): serialize replaced/generation probe with close() in _execute_write
A lock-free _raise_if_db_replaced() probe at the top of the _execute_write
retry loop raced a concurrent close(). close() runs under the same _lock and
ends the WAL generation: it checkpoints, closes the connection (SQLite
unlinks the -wal/-shm sidecars), nulls _conn and clears
_db_sidecar_identity. The probe could observe the mid-teardown state —
sidecars already unlinked while _db_sidecar_identity was not yet cleared —
and misclassify this process's OWN clean close as an externally deleted WAL
generation, raising a sticky DeletedWalGenerationError that permanently
refused every later write on that handle (#105567).

Move the live probe inside the lock, ahead of the close-race reopen
decision, so it only ever observes the stable post-close state (identity
cleared -> the existing adopt/reopen path). The corrupt flag check stays on
the lock-free fast path; external file/generation replacement detection is
unchanged, just serialized with teardown.

Synthetic repro (100 rounds x 40 writes, direct SessionDB handles): before
~9 failing rounds / ~360 DeletedWalGenerationError; after 0 failures,
4000/4000 writes persisted across repeated runs. tests/state (181) plus the
generation/replaced/corrupt guard suites (55) pass.

Fixes #105567
2026-09-11 06:23:23 -07:00
chelsealong
d6b036b726 fix(state): compare fd identity against the watched path, not st_nlink
st_nlink == 0 alone cannot distinguish a genuine orphan from one that
still has a surviving hard link (e.g. a backup) after the watched
sidecar path itself was removed or replaced — that left st_nlink >= 1
on a truly orphaned generation, letting a new opener through while a
live writer still owned the old one. Compare (st_dev, st_ino) between
the fd and the current watched sidecar path instead: only an exact
match means they're the same live file, so any mismatch or unstattable
watched path still fails closed.
2026-09-11 06:23:23 -07:00
chelsealong
84a3c4de74 fix(state): require nlink==0 before treating a /proc fd as an unlinked WAL sidecar
iter_deleted_sqlite_sidecar_holders() and SessionDB._wal_generation_was_lost()
both treated a `` (deleted)`` suffix on a /proc/<pid>/fd/* target as proof that
state.db-wal or state.db-shm was unlinked. On OpenZFS that suffix is not proof:
a live, still-linked file whose dentry was unhashed is reported the same way,
with st_nlink still 1 and the same (dev, ino) as the path. The guard then fires
permanently and the gateway falls back to JSONL forever, because the WAL was
never actually deleted.

Add _fd_is_truly_unlinked(), which confirms via os.stat(fd_path).st_nlink == 0
before a target counts as an orphaned generation. An unstattable descriptor
still counts as deleted, so the guard keeps failing closed. _iter_proc_fd_targets()
and _proc_fd_targets() now also yield the /proc fd path itself so both call
sites (open-path and the sticky write-path probe) can run the check.
2026-09-11 06:23:23 -07:00
Teknium
7dec81568e test(desktop): trim salvaged pool tests to invariants
Drop two PrimaryProfilePin cases that only restate the constructor
defaults and blank-string normalisation, and the wiring-routing test that
froze POOL_LIMITS_SETTINGS_ROUTE to a literal string — a snapshot of the
constant, not a behaviour contract. The two kept pin tests cover the bug
(a live primary keeps answering for its booted profile after the stored
preference moves; teardown releases the pin), and the notifications tests
cover the toast action end-to-end.
2026-09-11 06:23:18 -07:00