Every defaults-free config reader (gateway runtime, TUI gateway, cron
scheduler + job snapshot, `hermes send` env bridge, doctor memory section,
hermes_cli/main early parse, hermes_time, hermes_logging, the gateway
fallback-chain refresh) re-implemented "read config.yaml + managed overlay +
${VAR} expansion" by hand, in three different orders, and none of them
replayed the model-key canonicalization or the last-known-good recovery that
load_config() gained. An admin-pinned `${VAR}` expanded on one surface and was
bridged literally on another; `model: {name: x}` resolved to an empty model
everywhere except the gateway.
hermes_cli/config_effective.py::load_user_config_effective is the one
primitive: user file → ${VAR} → managed overlay → _normalize_root_model_keys,
no DEFAULT_CONFIG merge, sharing read_raw_config's parse cache and serving the
last good parse (in-process, then backups/config/*.good.*) on torn YAML;
`fail_closed=True` raises for the one caller that keeps its own last-good
state (the fallback-chain refresh). gateway/run.py::_load_gateway_runtime_config
is deleted — it was _load_gateway_config plus expansion, and _load_gateway_config
now expands.
Behavior change: _load_bridge_config, send_cmd._load_hermes_env and
doctor_state._doctor_memory_config expanded BEFORE the overlay; they now match
load_config (managed `${VAR}` expands against the process env only). All nine
sites gain model-key canonicalization and last-good recovery.
send_cmd._load_hermes_env now routes its .env read through
env_loader._load_dotenv_with_fallback so the credential sanitizer runs.
cron/ledger.py (e24c8499) existed so a long-running scheduler that lazily imports
notepad/incidents AFTER `hermes update` never needs new names from a module it already has
cached. The dedup deleted it and imported open_db/transaction from hermes_cli.sqlite_util at
module level; a pre-upgrade daemon has the OLD sqlite_util cached (executions imported
add_column_if_missing from it), so the first job tick after an upgrade would ImportError in
scheduler_prompt._build_job_prompt until restart.
- cron/{notepad,incidents,executions,delivery_queue}: import open_db/transaction/
add_column_if_missing and cron.jobs._ensure_cron_dir inside _connect/_transaction/
_initialize_schema. This also stops the 3.8k-line cron.jobs being pulled eagerly by
importing a store (it was lazy in cron/ledger.open_ledger).
- gateway/hosted_rooms_common, hosted_room_policy_checkpoint: same treatment; the gateway
imports hosted_rooms lazily from request handlers, so it has the same skew exposure.
- tests/cron/test_upgrade_module_skew.py: simulate the real skew (delete open_db/transaction
from the cached sqlite_util, then import each store). The previous repoint deleted names
from cron.executions, which notepad/incidents do not import from, so it passed regardless.
Sabotage: a module-level `from hermes_cli.sqlite_util import open_db` in notepad fails it
with "cannot import name 'open_db'".
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.
hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.
Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
`ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
hermes_state_messages._scrub_surrogates (0 callers) deleted.
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.
Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.
CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.
Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).
Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).
Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
Inspired by Claude Cowork (desktop changelog v1.46388.1, 2026-09-04), which
added "automatic re-runs (after 5, 15, and 30 minutes) for a scheduled task
that could not reach the model at all, for example right after the computer
wakes behind a VPN."
A recurring cron job whose run fails with a transient network/DNS error
before ANY model call previously sat out a full period (a daily job fired
into a reconnecting VPN silently skipped a day). Now the scheduler pulls
next_run_at earlier along a bounded 5/15/30-minute ladder, suppresses the
interim failure notice while a re-run is pending, and resets the ladder on
any run that reaches the model.
Deliberately narrower than a generic retry (cf. PR #16512): zero API calls +
transient classification means nothing executed and nothing was spent, so a
re-run cannot duplicate side effects. One-shots are excluded (at-most-times
dispatch accounting, #38758); retries never fire past the schedule's own
next occurrence; `cron.retry_unreachable: false` disables.
- cron/unreachable_retry.py: ladder, classification, plan/clear/will_retry
- cron/scheduler.py: flag unreachable failures in run_job; suppress interim
notice; thread model_unreachable through the fenced bookkeeping write
- cron/jobs.py: mark_job_run schedules/clears the ladder under the jobs lock
- docs: website/docs/user-guide/features/cron.md
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.
Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
Follow-up to the cherry-picked #102431 fix, addressing the review findings:
- The two real-helper scheduler tests ran the Linux-only helper unmarked
and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
scripts/check-windows-footguns.py --all (lint lane red). The remedy text
is now built once in the helper and passed into the warning, so the
"scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
once (helper level); default config still Popens externally and
`require_restart_safe_scope: true` raises (scheduler level, real helper).
Dropped the stubbed duplicate, the standalone config-raise test and the
in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
with the same `except Exception` guard as the sibling
`failure_nudge_threshold` read, so a config error no longer escapes
the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
`cron.require_restart_safe_scope` documented in the cron user guide.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).
Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.
The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.
Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.
Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
Replace the POSIX-only jobs-flock contention test (skipped off-POSIX,
~120 LOC of monkeypatched flock plumbing) with a single invariant test
that fails on pre-fix code in <1s: hold the per-job fire fence from a
worker thread, assert the heartbeat still returns True on the calling
thread, and that a takeover is still detected (False). The docstring on
heartbeat_fire_claim now records WHY it is not under the fence, so the
next refactor does not put it back.
Co-authored-by: Oliver Heckmann <46627487+oheckmann74@users.noreply.github.com>
Co-authored-by: salch-cred <141555468+salch-cred@users.noreply.github.com>
heartbeat_fire_claim only CAS-refreshes claim.at via _with_job; wrapping
_under_fire_fence across save_jobs let a blocked .jobs.lock pin the fence
and cause mark_job_run to fail closed on completed jobs.
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.
The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
(new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
(`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
updated, MCP discovery + log routing run for it. Other profiles' adapters are never
touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
that did not pick the profile up (older build / signal failed).
Every start (and every `ensure_hermes_home` from a sibling CLI invocation)
chmod'd the data directory to 0700, wiping group/other bits and the ACL
mask on a bind mount shared with other containers (hermes-webui, Nix
desktop + dashboard). _secure_file already skipped containers for this
reason; _secure_dir did not.
In a container the directory mode is now left to the operator unless
HERMES_HOME_MODE is set explicitly, which is still applied. cron/jobs.py
had its own 0700/0600 copies that bypassed the managed/container rules;
they now delegate to the shared helpers so cron/output stops re-locking
the mount as well.
Fixes#10757
Same bug class as the wake-gate call fixed in the picked commit: when a caller
invokes _build_job_prompt without a cached prerun_script, the sibling ran the
job's script inline via _run_job_script(script_path) and dropped the job's
configured workdir, so the script ran from the scripts-dir parent. Route it
through _resolve_job_workdir like the no_agent and wake-gate sites already do.
Sweep of every _run_job_script / _run_job_script_with_claim_heartbeat caller:
_run_no_agent_job and cron/monitor.py already pass workdir; the wake-gate site
is fixed by the salvaged commit; this was the last one.
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch
advance is provisional, restore once, never twice, grace, opt-out, paused never
catches up, same on standalone and multiplexed); the user guide describes the
catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md
lists it as a hardening invariant.
tick() advances a recurring job's next_run_at BEFORE dispatch so a crash
mid-run cannot re-fire it on every restart (at-most-once). That leaves a
window — advance persisted, fire claim not yet taken — in which the process
dying (interpreter already finalizing, executor refusing new futures, SIGKILL,
the Desktop idle-exit from #107485) loses the occurrence silently: the
restarted scan sees only the future next_run_at, writes no execution row and
no log line, and a daily job skips a day.
Contract (documented in cron-internals.md): every recurring occurrence is
accounted for — it runs once, or its skip is logged with a reason.
- The due scan stamps `pending_slot = {scheduled_at, at, by}` in the same save
that records `last_dispatch`; claim_job_for_fire and mark_job_run clear it,
and an explicit schedule / next_run_at / enabled / state rewrite (edit,
pause, resume, run-now) drops it.
- A later scan that finds the stamp with a provably gone owner (this process
and the job not in its running set, or another owner past the fire-claim
lease / dead pid) restores scheduled_at as next_run_at ONCE and logs a
WARNING (cron/occurrences.py::unclaimed_pending_slot). The restored instant
then meets the ordinary policy — completed_occurrence() blocks a second fire
of a slot that already ran, the grace window classifies it late, past-grace
collapses the backlog into one run, cron.catch_up_missed: false skips it
with a logged reason. Never a replay of N slots.
Same store fields on both topologies: a standalone `hermes -p X gateway run`
and a profile served by the default multiplexer evaluate the identical record.
Under gateway.multiplex_profiles a secondary profile X is ticked, dispatched
and notified from the default profile's process, where os.environ holds the
DEFAULT profile's .env and X's values live only in the per-turn secret scope /
HERMES_HOME override. Every remaining read that skipped that scope made X
behave differently from `hermes -p X gateway run`:
- cron: HERMES_CRON_TIMEOUT, HERMES_MODEL (job/preflight fallback),
HERMES_CRON_MAX_PARALLEL, inflight allowance, prefill file and the script
timeout were bare os.getenv → the default profile's values; a job without a
model silently ran on the default's HERMES_MODEL instead of refusing.
cron/env_settings.py::cron_env_setting reads the scope (fire) or the ticked
home's .env (tick thread), plain environ when multiplexing is off.
- child env: the restart-safe cron worker, the Bot Chat delivery child and the
kanban worker inherited the launch profile's non-credential .env settings
and bridged TERMINAL_* policy (TERMINAL_ENV=docker, default's image,
HERMES_MODEL) — X's worker ran in the default's docker image on the
default's model. tools/environments/local.py::strip_launch_profile_env drops
them when the child targets another served profile.
- kanban: the worker --toolsets pin was silently dropped for every served
assignee (toolset probes call get_secret without a scope → swallowed
UnscopedSecretError); notifier pings, artifact uploads and the wake text ran
under the default's media policy / display language (only wake() was scoped).
- /loop: _post_turn_loop_completion hopped to the executor without contextvars,
writing the completed tick into the DEFAULT profile's state.db and leaving
X's row awaiting_response forever; the --until judge ran with the default's
aux credentials.
- background processes: a secondary's processes.json (scope-relative since
adf23550f5) was never read at startup; its processes were not re-adopted
and notify_on_complete notices were lost. Startup recovers every served
home under its scope; recovery adopts each session once.
- completion delivery: background_process_notifications was evaluated once per
drain for the ambient profile (default's mode for everyone; X's `off`
dropped a sibling's `all` event), recovered watchers used the default's
mode, HERMES_BACKGROUND_NOTIFICATIONS was read raw from environ;
_deliver_platform_notice used the default's GatewayConfig so a secondary's
notice_delivery: private went public.
Not changed: gateway/run.py and tools/async_delegation.py (PR #106742
rewrites both). Known residue left for the env-bridge lane:
HERMES_SESSION_STALL_TIMEOUT is bridged once from the launch config.
scheduler bug, not a parity gap; unchanged here.
SharedRouteAdapters.get called ProfileRoute.matches without guild_id, so the
documented Discord route shape (guild_id + chat_id) never authorized a cron
target and the satellite fell to standalone delivery ("DISCORD_BOT_TOKEN is
not set" every fire). A cron target has no inbound guild anchor; the route's
own guild_id is passed so its target-exact discriminators decide.
_resolve_target_transport then vetoed the authorized shared transport on the
SATELLITE's platforms.<p>.enabled (absent block or enabled: false), although
that block describes a connector the satellite never runs. The shared hit now
builds the transport directly (keeping the satellite's non-credential platform
settings), and a live native adapter with no config block is no longer read as
"disabled" (#89302) — same normalization the relay path already had.
Fixes#89302
Co-authored-by: web3blind <264741654+web3blind@users.noreply.github.com>
Symptom (#102041): under a multiplex gateway the default profile's vault/1Password/
Bitwarden/plugin-sourced credentials vanished for the rest of the process after the first
cron fire or the post-discovery plugin refresh; with the key already in the process env
(systemd EnvironmentFile=) the scope was empty from boot. Every get_secret() read then
failed closed ("No usable credentials", every Telegram sender rejected).
Why: _apply_external_secret_sources marked the home applied after any real fetch, but only
snapshotted names in report.provenance — the NEWLY applied ones. On a re-apply the previous
apply's own write-back makes every key `skipped_existing`, so the snapshot latched to {} and
_hydrate_profile_secret_sources returned that empty snapshot forever. Separately,
reset_secret_source_cache() was process-wide, so one profile's cron re-pull dropped every
sibling's hydrated snapshot (1aa62ceb45 isolated the routed reload but not the reset).
Change:
- env_loader: snapshot every name a source supplied (provenance + skipped_existing) from
the home's effective environment, so a shadowed re-apply keeps the values it had.
- reset_secret_source_cache(hermes_home=None): optional per-home reset; global clear kept
for tests/config edits.
- cron per-fire re-pull and plugins._refresh_secret_sources_after_discovery reset + reload
only the home they resolve to.
- Two invariant tests (red on base) in tests/test_env_loader_secret_sources.py; docs note in
secret-source-plugin.md.
Reported-by: luochen1990
Addresses #102041
- `_target_matches_origin` docstring still claimed fan-out targets are
"deliberately NOT mirrored" and mirroring is origin-only; that has been
false since origin_fallback landed and is doubly so with the home lane.
Point to `_target_mirror_eligible` for the policy instead.
- cron.md DM-only bullet said "origin DM session"; the config comment in
the same change already says "target DM". Align.
- Remove two `_expand_routing_tokens` unit assertions bolted into the
eligibility test; the `" ALL , slack "` parametrize covers expansion
end-to-end through `_resolve_delivery_targets`.
A managed cron with `deliver: "slack"` (no captured origin) delivers to the
Slack home channel, but the brief was never mirrored into that session and
the in_channel seed never fired: bare-platform targets resolved with no
`_resolved_from` provenance, so `_target_mirror_eligible` treated them like
`all` broadcast expansions.
- `_resolve_delivery_targets` derives `from_broadcast` from the raw token
and passes it to `_resolve_single_delivery_target`; a user-written bare
platform token is tagged `_resolved_from: "home"`, `all` expansions stay
untagged (fan-out is never continuable).
- `_target_mirror_eligible`: `home` eligible under the same flags as
`origin_fallback` (per-job attach_to_session wins, else cron.mirror_delivery).
- `_MIRROR_PROVENANCE_RANK`: `home` == `origin_fallback` so token order
through dedup cannot strip eligibility.
- Docs, tool schema text, cron/gateway AGENTS.md updated; existing
exact-dict pins carry the new key.
Squashed from PR #101819 (commits 5fe612c9, 65e36926, 10e0af1d, f2cc4639):
they predate the cron/scheduler.py -> scheduler_delivery.py split and do
not cherry-pick individually onto current main.
claim_job_for_fire refused any claim younger than FIRE_CLAIM_TTL_SECONDS (300s)
regardless of whether its owner was still alive, so a `hermes cron run` killed
mid-flight (timeout, Ctrl-C, OOM) blocked the next manual run with "already being
fired" for up to five minutes. The executions table already reaps dead owners on
sight; the job-record fire_claim now does the same via gateway.status._pid_exists
when the claim's `by` names a pid on this host. Foreign hosts, explicit
HERMES_MACHINE_ID values, and probe failures keep the TTL (fail safe).
A global model/provider change must never stop a cron job. The #44585 guard
raised [drift_skip] for every unpinned job whose provider_snapshot /
model_snapshot no longer matched the live global default, so one `hermes model`
switch silently killed whole fleets (reported by fastfinge, nitinthewiz,
Dr-ilies; 13 of 60 jobs on the project lead's box after
claude-fable-5 -> claude-fable-5.1).
The snapshot is now the job's effective pin: _load_cron_job_config prefers
job['model_snapshot'] over the global default and _resolve_job_runtime passes
job['provider_snapshot'] as `requested` when neither a per-job pin nor a
cron.model / cron.model_provider fleet default covers the axis. One INFO line
per differing axis tells the operator what the job is running on and how to
move it. Jobs without a snapshot (legacy records) still follow the global
default; the existing fallback chain still handles a snapshot provider that
fails to resolve.
Both goals of #44585 hold: no silent inherit of a paid default (the job runs on
what it was created under) and no outage. Owner decision (Teknium): "main agent
model changing should not stop crons from executing, ever".
Removed as unreachable: _check_model_drift, DRIFT_SKIP markers, the
drift_alerted alert-once bit (mark_drift_alerted + the _record_run_outcome pop),
the drift special-cases in _compose_run_delivery and
_summarize_cron_failure_for_delivery, cron_model_drift_guard_enabled and the
cron.model_drift_guard config key (v42 migration drops it from existing
configs). The PLUGIN-COMPAT clear_drift_alerted block is untouched (scheduled
revert).
The `hermes config set model.default` notice and the Desktop model-change toast
are reworded from "will fail closed / will be skipped" to "keep running on the
model they were created under"; the impact payload drops guard_enabled (all six
desktop locales updated).
Second entry of the same bug class: POST /api/cron/jobs/{id}/trigger →
_fire_cron_job_for_profile → CronScheduler.fire_due → claim_fire built its claim
without `manual`, so an off-tick run from the web UI stamped the future slot exactly
like the tools path #105704 fixes. fire_due/claim_fire gain `manual` (forwarded only
when set, mirroring `force`, so third-party providers keep working) and the dashboard
trigger passes it when the provider's signature accepts it. Webhook and misfire
catch-up fires run the slot that is due and keep the stamp.
Also drops the base-green tick-stamp test (the same contract is pinned by
tests/cron/test_scheduled_occurrence.py) and documents `manual` vs `force`.
claim_job_for_fire() derives the occurrence identity from next_run_at
before the same function advances it. On a scheduler tick next_run_at is
the occurrence being run, which is correct; on an off-tick manual run it
is the NEXT occurrence, so the execution is stamped with the identity of
a slot that has not happened yet. _job_is_due() then finds a completed
execution carrying that identity and skips the real slot, returning
before the last_dispatch write — no error, no log line, no dispatch
record.
The manual flag already guards this and both _job_is_due() and
claim_job_for_fire() honour it; the agent-facing run-now path never
declared itself. Add a keyword-only manual= parameter and pass it from
_claim_for_manual_run(). Deliberately not force=True: force also calls
_activate_job_record(), which would resume a paused or disabled job, and
the run-now tool depends on continuing to refuse those.
The local flag is renamed to manual_fire so the new parameter is not
shadowed inside the apply closure, which would raise UnboundLocalError.
Three existing tests in tests/tools/ pinned the old call signature via
assert_called_once_with; they now pin manual=True, so dropping the flag
again fails loudly rather than silently reintroducing the skip.
Restores the intent stated in #104790 — the column records the scheduled
instant an execution was claimed for, and an off-tick manual run was
claimed for none.
Fixes#105690
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A system-level gateway unit has no ordering against user@<uid>.service and
linger may be enabled after boot, so the bus can appear after the one-shot
adoption in run_gateway() ran. Derive XDG_RUNTIME_DIR/DBUS_SESSION_BUS_ADDRESS
fresh for the availability probe and every scoped spawn (cron worker, Kanban
worker, PTY/pipe terminal spawns, scope cleanup) so the 60s failure TTL can
actually recover. Refs #104893.
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.
Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
Persist paused state, timestamp, reason and no first trigger in the original
locked creation write. Forward the same boolean contract across CLI, tool,
gateway API and dashboard API, validating at the store boundary. Preserve
explicit operator force-run behavior and normal enabled creation.
The live CLI probe also caught the command shim dropping failure return codes;
forward them so invalid creation reports exit 1 rather than success.
Credit earlier atomic-creation work in #78935 and #94952 and the focused
implementation in #104578. The broader manifest staging layer is not imported.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: Chloé DuPont <321112755+misschloedupont@users.noreply.github.com>
Persist a redacted chained traceback in the private run output and expose
redacted last_error in tool and slash listings, including historical errors.
Keep the run_job concise error return unchanged for delivery classification.
Slim redo of liuhao1024's earliest #104545; adds forced redaction and keeps
formatting in a topical sibling. Local SDK/socket A/B verifies diagnosis
visibility plus healthy-script, clearing, and private-file controls.
Canonical tests queued under the campaign lock at commit time.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Slim redo of #104546 and #104551: scan newest-first, match suppression only before payload separators, and keep error context. Covers wake gates and empty outputs without reading every historical file twice.
Co-authored-by: PRATHAMESH75 <prathamesh290504@gmail.com>
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Carry the existing write fence across Hermes-owned spawn boundaries without
dropping board routing or changing credential policy. Grant dispatcher and
managed tool runtimes explicit task scope; align CLI task mutations with tools.
Verify real shell/CLI descendants, dispatcher startup, and supervised stdio
transport against isolated SQLite boards. This is cooperative runtime scoping,
not OS confinement.
Refs #103974, #104058, #104904
Canonicalize the accepted raw next_run_at value before fast-forward rather
than its timezone-interpreted datetime. Legacy naive values remain runnable
but cannot establish an exact UTC identity. Preserve repaired aware slots.
Extend the existing identity invariant with the naive-slot control, and use
real ledger creation in the provider ordering test instead of an invented
execution ID that cannot pass the owner-fenced occurrence setter.
Live validation: actual builtin script run was RED (invented UTC identity)
and is now GREEN (NULL identity). Repeated builtin/provider/worker rollback
A/B remains 2 writes on base versus 1 on head, with distinct/manual controls.
Canonical cron regression rerun is queued under the campaign lock.
Capture the exact UTC scheduled instant before either due scanning or the
external fire claim advances jobs.json. Bind it to the durable attempt
before worker handoff; manual and unclassified direct attempts stay null.
Consult any retained completed matching row, independently of stale stamps,
claim-time windows, and newer failed attempts. Preserve unknown and legacy
attempt eligibility rather than guessing that a side effect completed.
Real isolated restart probes reproduce duplicate script writes on base and
suppress them on the fix for builtin tick and provider fire. Distinct and
manual occurrences still execute. The campaign-serialized cron suite is
queued; this progressive commit preserves the verified integration step.
Credit holny's issue #104790 and guard proposal #104323; exact identity
replaces the approximation rather than importing its legacy heuristic.
Co-authored-by: holny <holny@foxmail.com>
The TTL was three copies of the literal 300 (claim_job_for_fire default,
rearm_oneshot, and the new stale-error guard). Hoisted so the three lanes
cannot drift; incident narrative in the guard comment cut to the WHY.
Multi-process schedulers sharing one jobs store (gateway + Desktop serve tabs)
re-armed a job every tick while a long run in another process was still
heartbeating its fire_claim, producing claim-fight churn and killing the live
run (brain, 2026-09-02). Treat a fresh fire_claim as 'running elsewhere'.
The deferral helpers from the previous commit were appended to the cron/scheduler.py facade and
carried ~70 lines of fallbacks (getattr/callable checks, result() waits, a teardown_registered
flag) for hypothetical Future doubles; the only producer is _cron_pool.submit, always a real
concurrent.futures.Future whose done()/add_done_callback() cannot raise. Also drops the
_finalize_cron_session_db passthrough. Behaviour unchanged; the test now asserts the
finalize+teardown contract on the real Future instead of a patched wrapper.
The #102827 corruption is pure zero holes -- frames lost across a WAL
generation. SessionDB.close() produces exactly that when it runs against a
file another live handle is still writing: PRAGMA wal_checkpoint(PASSIVE),
then the connection close that lets SQLite unlink -wal/-shm. The dangerous
event is a physical close overlapping any other live physical lifetime for
the same path, so both sides of it are closed here.
Late write vs. close: a cron watchdog timeout only stops waiting, and
ThreadPoolExecutor.shutdown(wait=False) cannot interrupt a worker already
inside run_conversation. The agent and its registry reference are now held
until that worker's Future completes, so its last frames land before any
checkpoint.
Close vs. open: the per-path barrier now COUNTS admitted teardowns. A path
can own several closes at once -- the current generation's final release and
a retired generation's drain are admitted independently under the registry
lock, and the per-path mutex only serializes teardowns that already entered
it. With one bare event per path, a releasing thread descheduled between
generation removal and the mutex let the next teardown to settle remove and
signal the shared event: close_all() returned over a pending close and
acquire() published a replacement writer on top of a handle still inside
checkpoint/unlink. _TeardownBarrier tracks event + pending count,
_admit_teardown_locked registers each close in the same lock section that
removes the generation, and only the last settled teardown lifts the
barrier. Physical I/O stays outside the registry lock and unrelated paths
still progress independently.
The auto-archive sweep called release_or_close in its finally while the
import was local to a different function, so every eligible sweep raised
NameError, the outer except Exception swallowed it at debug level, and the
borrowed registry reference was never returned -- a holder leak that pins a
retired generation open. The helper is now bound in the calling scope.
Remaining in-process writable SessionDB() call sites (trace upload, the
API-server profile cache, the web-server writable paths, startup schema
reconcile) go through the canonical registry acquire/release_or_close, and
gateway maintenance borrows pinned handles instead of iterating an unpinned
snapshot.
Regressions: overlapping final releases of the current and retired
generations in both orderings with the first paused before the lifecycle
mutex, teardown-error settlement, an unrelated-path control, and refcount
assertions for the auto-archive sweep on success, on failure, across
repeated sweeps and with auto-archive disabled.
Fixes#102827
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BzxCWw6SuHXhXMdkiEwMa2