Review follow-up for #117505: delivery_ledger, api_server_runs and
async_delegation reconciled a live owner to dead/unknown on the same 1 s
same-host fingerprint drift that cron/executions had just stopped doing.
One comparator (gateway.status.start_time_fingerprints_match) replaces the
private cron tolerance and the three exact-equality sites.
_owner_is_live() required the recomputed start-time fingerprint to be exactly
equal to the one recorded at claim time. On hosts where same-host readings
drift by ~1s (macOS kern.boottime adjustment), every long-running execution
was reconciled to unknown while its owner process was still alive and healthy.
A mismatch is now treated as inconclusive within a small tolerance (both
platform fingerprint scales are x100, so 200 means 2s), and an unreadable
reading is fail-safe, matching the function's existing posture that inability
to prove death must not rewrite durable state.
The bound reads the script timeout, which falls through to load_config() when
HERMES_CRON_SCRIPT_TIMEOUT is unset. Computing it eagerly on every reap put a
config load back on the idle gateway tick (CI: tests/cron/test_idle_tick_config_skip.py).
Resolve it on the first live-owned claimed/running row instead; an idle tick with no
such rows stays config-free.
The carrier used a hardcoded 7200 s wall-clock ceiling. HERMES_CRON_TIMEOUT is
an *inactivity* limit (cron/jobs.py, _run_agent_with_watchdog), so a healthy
agent job in the external-worker topology can legitimately run past any fixed
wall clock; marking its row `unknown` releases the gateway's running guard and
the next tick double-dispatches while the first worker is still running (its
later finish_execution is fenced out). The constant also ignored
HERMES_CRON_TIMEOUT=0 (unlimited) and HERMES_CRON_SCRIPT_TIMEOUT > 7200.
Mirror cron/jobs.py::_oneshot_run_claim_ttl_seconds instead:
stale_after = max(3 × HERMES_CRON_TIMEOUT, script timeout, 7200), reusing the
existing accessors (_cron_inactivity_seconds, _get_script_timeout). When the
inactivity timeout is 0/unlimited or not a finite positive number there is no
bound to derive, so live owners are skipped entirely (fail closed — today's
behaviour). Drop the `except (ValueError, TypeError): pass` swallow (rows are
always aware ISO from hermes_time.now), fix the stray whitespace in the
SELECT, record a distinct `error` reason for the wedged-owner case, and keep
the comment honest: the deadlocked worker process is not terminated, and
in-process rows (process_id == _PROCESS_ID) remain out of scope.
A claimed/running execution whose owner process is alive but permanently
deadlocked (e.g. futex_wait after a route/proxy flip) passes _owner_is_live()
and is never reclaimed — the execution row stays 'running' forever and the
job rejects every subsequent fire with 'Job is already running'.
recover_interrupted_executions() now also checks the wall clock: a
claimed/running row older than STALE_RUNNING_CLAIM_TIMEOUT_SECONDS (7200 s,
covering the 3600 s script timeout + 600 s inactivity watchdog + margin)
is marked unknown even though the owner PID is alive.
Tests: stale claim recovered, recent claim NOT recovered.
The suppression feature needs one cron cell: a failure notice whose target hides warning
notifications is recorded as `suppressed` (bot_chat_pending record, deliveries queue row,
execution delivery_outcome) instead of being sent. That is kept.
Everything else the PR added to cron/ is an independent ledger rework and is removed here:
execution delivery manifests + filesystem manifest journal, `schema_meta` fence and the
"legacy intent adoption" layer, incident occurrence generations + trigger, jobs projection
CAS (`bind_delivery_execution`/`update_delivery_projection`), the per-tick projection
reconciler, and the `_deliver_targets`/`_settle_manifest` split. origin/main has none of
the state that layer migrates from (zero occurrences of delivery_manifest / manifest_journal
/ schema_meta); the "pre-flag old writer" the r5-r9 tests simulate is a vendored snapshot of
this PR's own earlier revision (tests/cron/_r5_prior_executions.py). Those fixes may have
merit on their own and should land as separate PRs with a main-reproducing test each.
Also restores main's `_classify_delivery_outcome` precedence (`failed` before `queued`)
and drops the SHA-pinned `git show <PR commit>` test, which would go red the moment a
rebase-merge rewrote that commit.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
cron/ledger.py (e24c8499) existed so a long-running scheduler that lazily imports
notepad/incidents AFTER `hermes update` never needs new names from a module it already has
cached. The dedup deleted it and imported open_db/transaction from hermes_cli.sqlite_util at
module level; a pre-upgrade daemon has the OLD sqlite_util cached (executions imported
add_column_if_missing from it), so the first job tick after an upgrade would ImportError in
scheduler_prompt._build_job_prompt until restart.
- cron/{notepad,incidents,executions,delivery_queue}: import open_db/transaction/
add_column_if_missing and cron.jobs._ensure_cron_dir inside _connect/_transaction/
_initialize_schema. This also stops the 3.8k-line cron.jobs being pulled eagerly by
importing a store (it was lazy in cron/ledger.open_ledger).
- gateway/hosted_rooms_common, hosted_room_policy_checkpoint: same treatment; the gateway
imports hosted_rooms lazily from request handlers, so it has the same skew exposure.
- tests/cron/test_upgrade_module_skew.py: simulate the real skew (delete open_db/transaction
from the cached sqlite_util, then import each store). The previous repoint deleted names
from cron.executions, which notepad/incidents do not import from, so it passed regardless.
Sabotage: a module-level `from hermes_cli.sqlite_util import open_db` in notepad fails it
with "cannot import name 'open_db'".
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.
hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.
Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
`ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
hermes_state_messages._scrub_surrogates (0 callers) deleted.
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.
Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
Capture the exact UTC scheduled instant before either due scanning or the
external fire claim advances jobs.json. Bind it to the durable attempt
before worker handoff; manual and unclassified direct attempts stay null.
Consult any retained completed matching row, independently of stale stamps,
claim-time windows, and newer failed attempts. Preserve unknown and legacy
attempt eligibility rather than guessing that a side effect completed.
Real isolated restart probes reproduce duplicate script writes on base and
suppress them on the fix for builtin tick and provider fire. Distinct and
manual occurrences still execute. The campaign-serialized cron suite is
queued; this progressive commit preserves the verified integration step.
Credit holny's issue #104790 and guard proposal #104323; exact identity
replaces the approximation rather than importing its legacy heuristic.
Co-authored-by: holny <holny@foxmail.com>
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
Follow-up to the salvaged restart-safe worker (#101877):
- delivery_queue: a row still `pending` at the worker's wait timeout was
marked `failed` and never drained, so any gateway outage longer than the
300s budget (e.g. a restart that runs `hermes update`) silently lost the
delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
every transaction (each `get_status` poll paid for it; terminalizing
paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
(bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
`hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
process with a 1s timeout instead — the worker commits its terminal row
before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
`mark_execution_running -> None`, which now means "ownership lost, return
before run_job" — the test passed without ever reaching the path it
names. Mocking `{}` restores it (mutation-checked).
- scheduler_provider.py: _profile_entry/_profile_cron_scope replace four hand-rolled
home-override+store blocks in the multiplex ticker; comments compacted to the WHY.
- incidents/executions/notepad/monitor/suggestions/blueprint_catalog/__init__: docstrings
and comments compacted; no code change (AST-identical).
- claim_job_for_fire returns the atomically claimed snapshot with a unique
fire owner; heartbeat_fire_claim renews the lease; mark_job_run fences
terminal writes by expected_fire_owner so a stale worker cannot record
over a replacement claim.
- run_one_job heartbeats the fire claim and forwards a combined cancel
event (ownership loss OR external cancel) into run_job; the agent path
is interrupted cooperatively and script-based jobs (no_agent + pre-run
scripts) are hard-stopped with a process-tree kill (POSIX killpg
SIGTERM then SIGKILL for surviving group members; Windows
taskkill /T /F), with a bounded pipe drain so a SIGTERM-ignoring
descendant cannot wedge the worker on communicate() EOF.
- Shutdown interruption is scoped to the exact execution token instead of
the bare job ID, so a replacement run of the same job never consumes a
stale interrupted flag.
- fire_claim_fence serializes save/deliver side effects per profile+job
with a cross-process flock; remove_job prunes the fence-lock entry.
- Preserves upstream BaseException terminal recording (#73973),
completed one-shot retention (#80624), blocked_config preflight
(T1-26), and the advance_next_runs batch on top of current main.
Use database.journal_mode as the sole non-secret operator setting, preserve the vulnerable-SQLite safety gate and existing WAL databases, validate explicit DELETE results, document the active config path, and cover real SQLite openers with behavioral tests.
Add HERMES_JOURNAL_MODE env / database.journal_mode config for
virtiofs/NFS/SMB where WAL is not crash-safe. Route 5 bypass openers
through apply_wal_with_fallback so a single setting covers every .db
(#68545).
On vulnerable SQLite (e.g. 3.50.4), do not enable WAL for fresh/non-WAL
shared databases — prefer DELETE instead. Leave existing on-disk WAL
alone (no live downgrade under concurrent gateway/cron openers). Surface
Python/SQLite version details as a doctor warning (#69784).