12 Commits

Author SHA1 Message Date
Christopher
eb2781dd2f fix(gateway): drain a restart-safe worker's cron delivery when it is queued, not on the next tick
A restart-safe cron worker runs outside the gateway, queues its final send in
cron/deliveries.db and polls that row once a second. The gateway drained the
queue only as a housekeeping chore at the 60 s tick, so a reminder whose turn
ended in 8 s reached the user a minute later (the issue's 60.86 s silence).

The housekeeping thread now watches the queue file (and its WAL) between
ticks and drains as soon as a stamp moved; the tick's drain stays as the
fallback. The watch never resets itself after draining: the drain's own status
write costs one empty pass, a worker that enqueued during the drain still
wakes it, and an empty pass writes nothing, so the stamp settles.

Fixes #117307
2026-09-20 15:30:45 -07:00
kshitijk4poor
55b254a2c7 refactor(cron): keep the suppressed disposition, drop the delivery-manifest ledger rework
The suppression feature needs one cron cell: a failure notice whose target hides warning
notifications is recorded as `suppressed` (bot_chat_pending record, deliveries queue row,
execution delivery_outcome) instead of being sent. That is kept.

Everything else the PR added to cron/ is an independent ledger rework and is removed here:
execution delivery manifests + filesystem manifest journal, `schema_meta` fence and the
"legacy intent adoption" layer, incident occurrence generations + trigger, jobs projection
CAS (`bind_delivery_execution`/`update_delivery_projection`), the per-tick projection
reconciler, and the `_deliver_targets`/`_settle_manifest` split. origin/main has none of
the state that layer migrates from (zero occurrences of delivery_manifest / manifest_journal
/ schema_meta); the "pre-flag old writer" the r5-r9 tests simulate is a vendored snapshot of
this PR's own earlier revision (tests/cron/_r5_prior_executions.py). Those fixes may have
merit on their own and should land as separate PRs with a main-reproducing test each.

Also restores main's `_classify_delivery_outcome` precedence (`failed` before `queued`)
and drops the SHA-pinned `git show <PR commit>` test, which would go red the moment a
rebase-merge rewrote that commit.
2026-09-18 01:43:35 +05:30
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
teknium1
579dbe0c71 fix(cron): sqlite_util imports are late so a running scheduler survives an on-disk upgrade
cron/ledger.py (e24c8499) existed so a long-running scheduler that lazily imports
notepad/incidents AFTER `hermes update` never needs new names from a module it already has
cached. The dedup deleted it and imported open_db/transaction from hermes_cli.sqlite_util at
module level; a pre-upgrade daemon has the OLD sqlite_util cached (executions imported
add_column_if_missing from it), so the first job tick after an upgrade would ImportError in
scheduler_prompt._build_job_prompt until restart.

- cron/{notepad,incidents,executions,delivery_queue}: import open_db/transaction/
  add_column_if_missing and cron.jobs._ensure_cron_dir inside _connect/_transaction/
  _initialize_schema. This also stops the 3.8k-line cron.jobs being pulled eagerly by
  importing a store (it was lazy in cron/ledger.open_ledger).
- gateway/hosted_rooms_common, hosted_room_policy_checkpoint: same treatment; the gateway
  imports hosted_rooms lazily from request handlers, so it has the same skew exposure.
- tests/cron/test_upgrade_module_skew.py: simulate the real skew (delete open_db/transaction
  from the cached sqlite_util, then import each store). The previous repoint deleted names
  from cron.executions, which notepad/incidents do not import from, so it passed regardless.
  Sabotage: a module-level `from hermes_cli.sqlite_util import open_db` in notepad fails it
  with "cannot import name 'open_db'".
2026-09-13 05:08:29 -07:00
teknium1
576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
liuzikaii
5b2c417db8 fix(cron): serialize delivery deduplication with terminal retention 2026-09-07 05:58:24 -07:00
Teknium
53db597201 simplify(compat): hermes_state — drop 81 re-exports + 3 registry aliases + 3 shims, repoint 45 callers + 60 test files
hermes_state.py: delete every '# noqa: F401 (re-exported...)' import block (hermes_state_common/errors/guard/
readpool/sessions/fts/dbfile/wal/repair/registry + agent.context_compressor _DB_PERSISTED_MARKER_KEY); keep
only the names hermes_state.py itself uses, without noqa.
hermes_state_registry.py: drop get_shared_session_db/release_shared_session_db/close_shared_session_dbs
aliases; every caller (gateway/, tools/, tui_gateway/, cron/, mcp_serve, run_agent, tests) now imports
acquire/release/close_all/release_or_close from hermes_state_registry.
hermes_state_titles.py: drop set_auto_title_if_empty shim (title_generator keeps its getattr fallback).
Re-remove shim-only names restored by 34abf954bd: latest_user_message_row_id (tests call
latest_message_row_id(key, role='user'); role-targeting assertions kept) and get_session_activity (tests
build the snapshot via agent.session_activity.build_activity_snapshot over db.get_session(sid)).
hermes_state_wal._log_once resolves its dedupe sets as module globals instead of via hermes_state;
hermes_state_repair helpers call module globals directly (tests patch hermes_state_repair.<name>).
Frozen updater surface untouched (update_cmd_maint imports only SessionDB from hermes_state).
2026-09-03 13:46:50 -07:00
kshitijk4poor
c3e9b28a42 fix(cron): key worker deliveries by the job's own attempt; deferred send is not a failure
Review fold-in on the salvage of #101877:

- `_deliver_result` routed to the durable queue whenever the worker's
  `_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was
  delivering. A worker whose script dispatches another job in-process
  (`hermes cron run <other>`) inherits that env and would have queued the
  nested job's message under the outer execution id — `INSERT OR IGNORE`
  then drops it silently. Match the marker against the delivering job's own
  `execution_id`, as `run_one_job` already does. Regression test added
  (mutation-checked: fails with the guard removed).
- A `pending` row left queued at the worker's wait timeout was still
  reported as a delivery error, so `mark_job_run` recorded
  `last_status=delivery_failed` for a message the next gateway's drain goes
  on to send, and nothing ever corrects the job record. Log and return
  success instead; the deliveries row is the authority for the send.
- Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead
  of a second hardcoded terminal set.
2026-09-03 12:40:26 +05:30
kshitijk4poor
b440a492b3 fix(cron): keep unsent worker deliveries queued, cheapen handoff polling
Follow-up to the salvaged restart-safe worker (#101877):

- delivery_queue: a row still `pending` at the worker's wait timeout was
  marked `failed` and never drained, so any gateway outage longer than the
  300s budget (e.g. a restart that runs `hermes update`) silently lost the
  delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
  queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
  every transaction (each `get_status` poll paid for it; terminalizing
  paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
  (bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
  `hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
  owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
  ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
  process with a 1s timeout instead — the worker commits its terminal row
  before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
  created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
  runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
  `mark_execution_running -> None`, which now means "ownership lost, return
  before run_job" — the test passed without ever reaching the path it
  names. Mocking `{}` restores it (mutation-checked).
2026-09-03 12:40:26 +05:30
Brooklyn Nicholson
83efdf5e5e [verified] fix(cron): close restart handoff races 2026-09-03 12:40:26 +05:30
Brooklyn Nicholson
29e5172487 [verified] fix(cron): harden gateway restart handoff 2026-09-03 12:40:26 +05:30
Brooklyn Nicholson
3373e97693 [verified] fix(cron): preserve active runs across gateway restart 2026-09-03 12:40:26 +05:30