Commit Graph

714 Commits

Author SHA1 Message Date
kshitijk4poor
840c00c124 refactor(cron): reuse _install_fire_secret_scope for the external-worker env build
_launch_external_cron_worker hand-rolled the same hydrate -> set_secret_scope ->
set_multiplex_context(routed) sequence, and the same context-then-scope reset, that
_install_fire_secret_scope/_reset_fire_secret_scope already encode for the in-process
fire. Two copies of the ordering is how the two paths drift apart; the helper is the
one place that owns it. The only observable difference (re-setting an already-active
multiplex context for the span) is a no-op the token reset undoes.

Rename _scope_token in _run_one_job_body to _fire_scope_tokens: it has held the helper's
(scope, context) tuple since the routed-fire change, not a single token.
2026-09-15 11:03:39 +05:30
kshitijk4poor
0fd59b447a fix(cron): scope the handoff multiplex context to the worker env build; gate the no_agent overlay
_launch_external_cron_worker wrapped fifty lines of dispatch, payload
write and scope hydration in the routed-fire multiplex context, though
only build_subprocess_env / strip_launch_profile_env read it. Compute
`multiplex_active` once (process flag OR routed fire), serialize it into
the payload, and set the context for exactly the env build inside the
existing secret-scope try/finally. Drop restore_managed_env on this path:
the worker re-runs load_hermes_dotenv -> _apply_managed_env at import
and strip_launch_profile_env already leaves managed keys in place.

_run_job_script now overlays the installed scope and re-applies managed
keys only under multiplex. In a single-profile process the scope is
os.environ, so the overlay could only re-sanitize values the child
already inherits; the "no-op outside multiplex" comment is now literal.
2026-09-15 11:03:39 +05:30
kshitijk4poor
f367ebeb5d refactor(cron): strip external-source residue in strip_launch_profile_env itself
_run_job_script popped the launch profile's secret-source names in an
inline loop right above strip_launch_profile_env, so only the no_agent
child got that protection; the four other callers of the same helper
(the external cron worker, scheduler_delivery, kanban dispatch and the
byterover plugin) still handed a served profile the launch vault or
1Password names. Fold the names into the helper's residue set, which is
already gated on multiplex and on the target not being the launch
profile, and keeps administrator-managed keys.
2026-09-15 11:03:39 +05:30
kshitijk4poor
fb4eaa59ee refactor(cron): inline the multiplex-context span into _launch_external_cron_worker
The `_inner` wrapper existed only to set and reset the multiplex contextvar
around the whole launch. Every consumer of that context (the payload flag,
`hydrate_profile_secret_sources`, `strip_launch_profile_env`, the scrubbed
`build_subprocess_env`) runs before the worker is spawned, so the span now
covers exactly the handoff-environment build and ends before `Popen` and the
acknowledgement wait. One function, one try/finally, same behaviour.
2026-09-15 11:03:39 +05:30
John Paul Soliva
62b4488cb5 fix(cron): routed fires are multiplexed at the worker handoff; managed keys keep policy precedence
Review findings on f5f88d5058. Three are defects the previous round introduced.

Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.

Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.

Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.

Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.

Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.

(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
2026-09-15 11:03:39 +05:30
John Paul Soliva
3dedff6a12 fix(cron): only strip launch external-source names when multiplexing
The source-name strip added alongside the external-source fix ran
unconditionally. Outside multiplexing there is no other profile to leak from --
os.environ IS this profile's environment -- so popping those names relied on the
routed scope overlay putting each one back, which in turn relies on the
per-home snapshot recorded at boot. Correct today, but it made a single-profile
child's credentials depend on bookkeeping that has nothing to do with isolation.

Guard it the way strip_launch_profile_env guards itself: no multiplexing, no
strip. A single-profile no_agent child now keeps a byte-identical env even if a
source's snapshot were ever missing. Pinned by a regression that runs a real
child with a source-owned name in os.environ and no multiplex context; making
the strip unconditional fails it.

(cherry picked from commit afa429b30a1c9c9b4c011f7d2426a9f099911596)
2026-09-15 11:03:39 +05:30
John Paul Soliva
0943e77136 fix(cron): strip launch external-source names too, and make the scope refresh replace
Two credential-isolation gaps found in review of the previous head.

1. strip_launch_profile_env() only knows dotenv- and terminal-config-owned names,
but external secret sources (vault, 1Password, ...) also write their names into
the shared os.environ and are tracked in secret_source_names(). A name the LAUNCH
profile's source supplied therefore still reached a routed no_agent child. Drop
every non-global source-owned name from the base; the routed scope overlay that
follows puts back exactly the ones that profile's OWN sources supply, since
build_profile_secret_scope folds get_secret_source_values(home) in.

2. refresh_installed_secret_scope() merged the rebuild with dict.update(), so a
name a source had stopped supplying -- rotated, revoked, source removed -- kept
its old value for the rest of the fire. Replace the mapping contents instead: the
rebuild is the profile's current truth.

Regressions: a routed child sees <unset> for a launch-source name while its own
source value comes through, and a refresh whose rebuild omits a name drops it.
Both fail if the corresponding change is reverted.

(cherry picked from commit ecd51517c4828a75acb5f458ed99aea0bc3e5e9f)
2026-09-15 11:03:39 +05:30
John Paul Soliva
023e4f997f fix(cron): drop the launch profile's dotenv residue before overlaying a routed no_agent scope
The routed no_agent child env started from all of os.environ and only overwrote the
names present in the installed scope. A name defined only by the LAUNCH profile's .env
and absent from the routed profile therefore reached the routed child with the launch
value instead of unset -- the secret scrub only knows classified names, so a custom or
unclassified secret crossed the profile boundary (review finding on the first head).

strip_launch_profile_env (main, 284d220ba4) is the primitive the external-worker path
already uses for exactly this: it drops the launch profile's dotenv-owned keys and the
bridged TERMINAL_* settings, and is a no-op outside multiplex or when the target IS the
launch profile. Apply it to the base BEFORE the scope overlay (so a shared name keeps its
routed value) and BEFORE the sanitizer (so routed values still pass the scrub and
passthrough rules). Pinned by a child-process negative control: the launch-only name
arrives <unset>, the shared name arrives routed, and the parent process is unchanged.

(cherry picked from commit 69349b527b5144c1f1540a4c10035d5ec798c1db)
2026-09-15 11:03:39 +05:30
John Paul Soliva
dbede34f6e fix(cron): a routed profile's cron fire in the desktop backend runs under multiplex semantics
The desktop backend ticks EVERY local profile's cron store from one process — its own docstring
says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global
multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every
isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared
`os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban
subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary
profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed
there after the tick, and a scope miss read the launch profile's tokens (#107692).

Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py)
is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not
the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the
override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the
profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that
span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are
therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs
before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam,
left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation
applies inside the routed fire with no per-site patching; the launch profile's own fires and the
backend's turns keep single-profile semantics; marker and override both reach the pool worker via
`copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through
`is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970).

Two consequences of suppressing the write are handled rather than left as regressions:
- a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`;
  the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub /
  passthrough rules apply to those values and the parent process is never mutated;
- plugin secret sources are discovered on the fire's first agent build, after the scope froze,
  and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds
  the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern
  `_publish_env_value` already uses for `.env` writes under multiplex).
And the profile's external secret sources are hydrated before the scope is frozen, the order
gateway/run.py and the external cron worker already use.

Tests pin each direction: the marker without the semantics before the scope, the semantics on and
off exactly with it, the marker reaching a copy_context worker; the process's own profile staying
single-profile; the restart-safe handoff's child env building without raising under a routed tick
with a passthrough key registered; a real child process receiving the routed values while
`os.environ` keeps the launch value; a source registered after the freeze reaching the fire
through the real PluginManager refresh. Reverting any one direction fails a distinct test.

(cherry picked from commit 2f87677425d2cca19286ac83bc45cab23e546669)
2026-09-15 11:03:39 +05:30
teknium1
8f7853188f fix(cron): keep deferred delivery exceptions from aborting ticks
Catch unexpected delivery exceptions after claim, retain diagnostics and continue
sibling admissions without authorizing replay. Preserve indefinite retention.

Reproduced PermissionError at target traversal after discovery. Native Electron
controlled-fault A/B confirms the healthy sibling settles and renders once.
2026-09-14 17:29:32 -07:00
teknium1
002ee41cfc fix(cron): keep unowned Bot Chat delivery on its resolved home
Extend deferred dispatch's destination pin to ordinary CLI fallback, so
custom-root and active-profile changes cannot redirect a checked target.
Refuse a missing destination before launch and name the target on failure.
Replace the old env-clearing expectation with two behavioral invariants
and retain the native Electron custom-root reproduction.

Adapted from the root-boundary fix and diagnosis in #104066.
Related #104055, #104066.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-14 17:29:32 -07:00
teknium1
3b0fe0cc2b fix(cron): keep deferred Bot Chat delivery bound to admission
Carry the original destination home and delivery ID into deferred drain and
its child, rather than re-resolving a mutable profile/root. Missing destinations
fail closed; supported-owner handoffs remain transferred, not ambiguous failures.
Capture the producer root before the background thread starts, and retain/log
malformed JSON without stopping healthy admissions or the whole cron tick.

Two invariants reproduced failures on the published head. Real Electron root
change and malformed-record cases are red before and green after; nested DM
control remains passing. No automatic retry of claimed or uncertain turns.
2026-09-14 17:29:32 -07:00
teknium1
6af84ead60 test(cron): retain native Bot Chat delivery reproduction 2026-09-14 17:29:32 -07:00
teknium1
5d8390d1a4 fix(cron): retain Bot Chat output while a CLI owner is open
Keep never-started output behind unsupported owners and drain in admission
order after release. Persist claims before execution and never replay uncertain
started turns. Existing supported-owner receipts keep their authority.

Credits 686f6c61's residual queue proposal in #100319. This is a scoped
implementation, not general retry of failed CLI subprocesses.

Native Electron before/after: CLI-owned target previously returned
SESSION_NOT_OWNED and remained empty after release/tick; now its queued
output and reply appear once in the target Bot Chat. Nested quiet CLI
message_agent delivery to a named Desktop owner also passes on base.
2026-09-14 17:29:32 -07:00
teknium1
45b42202fa fix(cron): a raising profile gate ticks nothing; housekeeping respawns a dead ticker (#111010)
Follow-up on the salvaged #111034:

- `_start_multiplex` published the enumerated home list before the gate had
  filtered it, so a raising `profile_gate` (the Desktop stand-down probe from
  #100489) kept the thread alive but ticked every profile UNGATED — racing the
  gateway that owns them for the same cron store. The list is now assigned only
  after gating; a gate failure yields zero ticks for that cycle.
- `cron/scheduler_thread.py::SupervisedTickerThread` wraps the gateway ticker
  thread; `_start_gateway_housekeeping` gets a per-tick "Cron ticker supervisor"
  chore that respawns a ticker that ended without a stop request and logs the
  outage at ERROR. Every guard inside `start()` keeps the loop alive, but nothing
  outside it could notice a thread that had already ended.
- Tests trimmed to the invariants, proven red on origin/main: a REAL corrupt
  `executions.db` (no patched recover) no longer kills the ticker; a raising gate
  keeps the thread alive with zero ticks; housekeeping restarts a dead ticker and
  leaves a stopped one alone.

Root cause: the unguarded pre-loop `recover_interrupted_executions()` +
`record_ticker_heartbeat()` were added by d9dd05b69d (#61791, "truthful
execution ledger", 2026-07-09). The reporter's build (e440bf35) also carried
#107485's `completed_occurrence()` in the due scan, which opens the same ledger
on every tick — the first traceback in their errors.log is that in-loop hit
(caught); the restart then hit the SAME corrupt ledger from the pre-loop
recovery scan, which nothing caught: thread dead, no heartbeat, no error marker.
2026-09-14 16:14:33 -07:00
Konstantin Khlopkov
8fce4baf8c fix(cron): keep the ticker thread alive through startup recovery and marker writes (#111010)
The gateway runs InProcessCronScheduler on an unsupervised daemon thread, so any exception escaping outside the guarded tick body ends cron silently while the gateway keeps serving. Guard the three remaining escape windows: startup recovery + initial heartbeat in start(), per-cycle profile enumeration/gating in _start_multiplex, and every status-marker store write (heartbeat/error/clear) in both loops. A failing store now degrades to a logged warning and a failed cycle instead of thread death.
2026-09-14 16:14:33 -07:00
Jony
45aebc11b0 fix(cron): preserve cleanup profile scope 2026-09-14 16:13:51 -07:00
teknium1
498abb677e fix(cron): a successful run resolves the job's open incidents; a repeat re-opens them
The incident ledger only ever grew. A one-off failure (a drift skip after a
global model bump, a provider outage) stayed `detected`/`alerted` forever
after the job recovered, so `hermes cron incidents` listed 32 "open"
incidents on an install where all 32 jobs had since run OK, and the list
stopped saying anything about current health.

A successful run now marks that job's `detected`/`alerted` incidents
`resolved` (new state). `resolved` is distinct from the operator's `closed`
ack on purpose: `upsert_incident` re-opens a resolved incident as
`detected` when the same error signature recurs, so the operator is alerted
again for a job that broke a second time, while `closed` keeps the
signature silent as before. Wired from `_compose_run_delivery` next to the
failure-side upsert; best-effort, store errors never affect delivery.

CLI: `--state resolved` filter and a green `resolved` colour; `closed` is
now dim. Docs updated in the same change.
2026-09-14 09:19:32 -07:00
Teknium
0877decd1b Port from paradigmxyz/centaur#1479: cron runs no longer re-schedule themselves from recurring prompt language
A scheduled job whose prompt carries its own cadence phrasing ('Each
Monday, review...') can convince the agent to create ANOTHER cron job
at execution time instead of just doing the work — each run spawning a
sibling job. Hermes policy-denies the cronjob toolset in cron context
by default, but cron.allow_agent_scheduling: true re-enables it and
opens exactly this loop.

_build_job_prompt now extends the always-injected cron hint with a
RECURSION clause: this is a run of an existing job; never create or
update a cron job from schedule language in the task prompt; treat
cadence phrasing as context for this run.

Adapted from paradigmxyz/centaur#1479 (same failure mode in their
scheduled-task workflow runner).
2026-09-13 20:52:58 -07:00
Teknium
f72e79a111 fix(fallback): named custom providers keep their configured identity after automatic fallback (#98739)
resolve_runtime_provider returns the bare billing class 'custom' for every
named providers:/custom_providers: entry; the configured id only survives in
requested_provider. All three fallback resolvers (gateway, TUI/desktop, cron)
persisted runtime['provider'] as the agent identity, so an automatic fallback
labeled the session 'custom' in the UI and billing rows, while a manual
/model switch to the same provider showed the configured name.

New shared helper hermes_cli.fallback_config.effective_runtime_provider()
upgrades the bare class back to the entry's configured identity (ad-hoc
provider: custom entries stay unchanged), applied at all three sites —
same class as the delegate_tool fix.
2026-09-13 20:51:06 -07:00
teknium1
95c7e9a0cb fix(cron): return a strictly-later next run when the base sits in the DST fall-back hour
Port from QwenLM/qwen-code#11723: attaching the configured zone to croniter's
naive wall-clock result resolves the repeated autumn hour to its earlier
occurrence (fold=0), so a base inside the second occurrence received a
next_run_at up to an hour in the past — the fire path would treat it as due
immediately and loop. Try both folds of each candidate and return the earliest
instant strictly after the base; wall-clock jobs still fire exactly once on
the repeated hour (Vixie cron semantics).
2026-09-13 16:46:45 -07:00
Jaimin
5a05332d72 fix(cron): preserve elapsed durations across DST changes 2026-09-13 16:46:45 -07:00
JF Lemieux
d6d29b0010 fix(cron): anchor croniter to the configured IANA timezone
croniter 6.x ignores tzinfo on its start time and instead uses the start's
UTC *offset* as its working offset. compute_next_run passed a tz-aware
last_run_at straight into croniter(expr, base_time), which produced two bugs:

- a last_run_at stored in UTC (+00:00) shifted the next fire to the cron
  hour in UTC rather than local time (09:00 UTC = 05:00 America/Toronto)
- on DST transition days the wall-clock hour drifted one hour off (08:00 on
  spring-forward, 10:00 on fall-back)

Render the base as the configured zone's naive wall clock for croniter,
then re-attach the zone to the result, so the wall-clock hour stays correct
every calendar day including DST boundaries. Fall back to the base's own
zone only when no timezone is configured.

Adds regression tests covering a UTC-stored last_run_at, spring-forward,
fall-back, and a full-year walk across both DST transitions.
2026-09-13 16:46:45 -07:00
teknium1
24c136e8af fix(cron): block a job whose requested MCP server resolves to zero tools
Under a multiplexer MCP tools are registered per profile overlay while the
server toolset alias is process-global, so a cron job naming a server in
enabled_toolsets that is connected only for another profile validated as a
known toolset, resolved to zero tools, and ran tool-less with quiet_mode
hiding the only diagnostic; the run was booked success (#109050).

After cron MCP discovery, an explicitly requested enabled MCP server that
resolves empty in this profile scope now takes the existing blocked_config
path (incident, alert-once, visible last_status). The implicit merge of
all enabled servers is not judged; only servers the job asked for.
2026-09-13 14:43:19 -07:00
teknium1
157aa716b7 fix(cron): external worker ack deadline equals the handoff adoption grace
The dispatch path abandoned a handoff after a fixed 5s while the dead-owner
recovery ledger already tolerates HANDOFF_ADOPTION_GRACE_SECONDS (30s) for the
same pending handoff. Field measurements (issue #109243: p90 claimed->started
10.7s, cold worker starts 9-12s dominated by imports plus secret hydration)
put the cold mode squarely inside the old deadline, so healthy handoffs were
logged as ownership-uncertain and never had their worker pid recorded. Use the
one constant the ledger already defines instead of a second, separately tuned
number, so the two guards around one event cannot disagree again.

Reshapes the salvaged test from #109252 to the current
restart_safe_gateway_child_argv signature and asserts the acknowledged-path
side effect (worker pid recorded) rather than clock progress.

Fixes #109243
2026-09-13 14:37:49 -07:00
KoNit-K
a7810a7a6e fix(cron): allow cold external worker startup 2026-09-13 14:37:49 -07:00
teknium1
b91088d768 refactor(config): one effective-user-config loader replaces 9 hand-rolled raw→overlay→expand pipelines
Every defaults-free config reader (gateway runtime, TUI gateway, cron
scheduler + job snapshot, `hermes send` env bridge, doctor memory section,
hermes_cli/main early parse, hermes_time, hermes_logging, the gateway
fallback-chain refresh) re-implemented "read config.yaml + managed overlay +
${VAR} expansion" by hand, in three different orders, and none of them
replayed the model-key canonicalization or the last-known-good recovery that
load_config() gained. An admin-pinned `${VAR}` expanded on one surface and was
bridged literally on another; `model: {name: x}` resolved to an empty model
everywhere except the gateway.

hermes_cli/config_effective.py::load_user_config_effective is the one
primitive: user file → ${VAR} → managed overlay → _normalize_root_model_keys,
no DEFAULT_CONFIG merge, sharing read_raw_config's parse cache and serving the
last good parse (in-process, then backups/config/*.good.*) on torn YAML;
`fail_closed=True` raises for the one caller that keeps its own last-good
state (the fallback-chain refresh). gateway/run.py::_load_gateway_runtime_config
is deleted — it was _load_gateway_config plus expansion, and _load_gateway_config
now expands.

Behavior change: _load_bridge_config, send_cmd._load_hermes_env and
doctor_state._doctor_memory_config expanded BEFORE the overlay; they now match
load_config (managed `${VAR}` expands against the process env only). All nine
sites gain model-key canonicalization and last-good recovery.
send_cmd._load_hermes_env now routes its .env read through
env_loader._load_dotenv_with_fallback so the credential sanitizer runs.
2026-09-13 05:09:06 -07:00
teknium1
579dbe0c71 fix(cron): sqlite_util imports are late so a running scheduler survives an on-disk upgrade
cron/ledger.py (e24c8499) existed so a long-running scheduler that lazily imports
notepad/incidents AFTER `hermes update` never needs new names from a module it already has
cached. The dedup deleted it and imported open_db/transaction from hermes_cli.sqlite_util at
module level; a pre-upgrade daemon has the OLD sqlite_util cached (executions imported
add_column_if_missing from it), so the first job tick after an upgrade would ImportError in
scheduler_prompt._build_job_prompt until restart.

- cron/{notepad,incidents,executions,delivery_queue}: import open_db/transaction/
  add_column_if_missing and cron.jobs._ensure_cron_dir inside _connect/_transaction/
  _initialize_schema. This also stops the 3.8k-line cron.jobs being pulled eagerly by
  importing a store (it was lazy in cron/ledger.open_ledger).
- gateway/hosted_rooms_common, hosted_room_policy_checkpoint: same treatment; the gateway
  imports hosted_rooms lazily from request handlers, so it has the same skew exposure.
- tests/cron/test_upgrade_module_skew.py: simulate the real skew (delete open_db/transaction
  from the cached sqlite_util, then import each store). The previous repoint deleted names
  from cron.executions, which notepad/incidents do not import from, so it passed regardless.
  Sabotage: a module-level `from hermes_cli.sqlite_util import open_db` in notepad fails it
  with "cannot import name 'open_db'".
2026-09-13 05:08:29 -07:00
teknium1
576accd92b refactor(sqlite): one open_db/transaction layer for every small store; plugin DBs use the WAL fallback
Twelve modules each carried their own sqlite3.connect + PRAGMA + `with conn:`
stack. The #69567 fd-leak fix (a `with conn:` commits but never closes, so each
call leaked a connection and its WAL/SHM fds until GC) was pasted as code plus
docstring into six of them and hosted_room_policy_checkpoint never received
it; plugins/plugin_storage.plugin_db was the only production caller issuing a
raw `PRAGMA journal_mode=WAL`, bypassing the network-FS fallback, the
WAL-reset-bug gate and the never-live-downgrade invariant that
hermes_state_wal.apply_wal_with_fallback carries.

hermes_cli/sqlite_util.py (already home to add_column_if_missing/write_txn,
imported by cron, gateway and hermes_cli alike) gains `open_db(path, *,
db_label, busy_timeout_ms, wal, foreign_keys, synchronous_full, row_factory,
check_same_thread, wal_lock_retries, initialize)` and `transaction(conn,
immediate=)`; cron/ledger.py is deleted and hosted_rooms_common's
open_sqlite/connect/transaction become 1-3 line forwarders. Migrated:
agent/verification_evidence, cron/{executions,incidents,notepad,
delivery_queue}, gateway/{delivery_ledger,hosted_room_policy_checkpoint,
hosted_rooms_common (-> hosted_rooms, hosted_room_driver)}, hermes_cli/
projects_db, tools/async_delegation, plugins/plugin_storage.

Behavior changes (each module keeps its effective PRAGMA set otherwise):
- hosted_room_policy_checkpoint: connection now closed after every use and
  on init failure (was leaked per call), busy_timeout PRAGMA set explicitly.
- projects_db: gains busy_timeout=5000 (was the sqlite3 default 5 s connect
  timeout with no PRAGMA); explicit and observable.
- delivery_ledger / async_delegation: busy_timeout PRAGMA now mirrors the
  10 s connect timeout they already had.
- plugin_storage.plugin_db: WAL through apply_wal_with_fallback (DELETE on
  network filesystems / WAL-reset-vulnerable builds instead of raw WAL);
  busy_timeout=5000.
- cron/incidents._redact_error: redact_sensitive_text(force=True) — the
  error text is persisted to disk.
- delivery_ledger's private duplicate-column guard and the unguarded
  `ALTER TABLE ADD COLUMN` sites (shared_metrics, api_server_run_idempotency,
  holographic store, kanban model_override) go through add_column_if_missing.
- hermes_state.py::_scrub_surrogates: dead byte-copy of
  hermes_state_messages._scrub_surrogates (0 callers) deleted.
2026-09-13 05:08:29 -07:00
teknium1
6a312fba54 feat(update): name the work a draining gateway is waiting on
`hermes update` printed "draining (up to 1875s)..." and then nothing for up
to 30 minutes while the gateway's in-band restart waited on in-flight work
(agent.restart_after_turn_timeout). Neither the updater nor the gateway log
said WHAT was being waited on, so a single long cron job read as a hung
update.

Gateway side: GatewayShutdownMixin._describe_active_work() enumerates each
unit the restart wait holds for — chat turns (session key, model, current
tool, elapsed), cron jobs (job id, elapsed, and the restart-safe external
worker pid when the run was handed off; cron/scheduler now records that pid
next to the running id), api/deferred runs by count. It is written to
gateway_state.json as `active_work` while the state is `draining` (cleared
otherwise) and appended to the 30s "Restart deferred" log line.

CLI side: hermes_cli/update_cmd_drain_report.py reads `active_work` and
prints a progress block every 30s during the SIGUSR1 exit wait — the
holder(s), their pids, elapsed time, seconds left before the forced
restart, and the config knob that caps the wait. Wired into the systemd,
launchd and manual gateway restart paths of `hermes update` and into
`hermes gateway restart`; `hermes gateway status` lists the same units
while draining. A pre-fix gateway (no `active_work` field) gets an explicit
"gateway did not report" line rather than silence.

Live A/B (real gateway, 90s no-agent cron job in flight, SIGUSR1 from the
caller): base = 79s of silence, no `active_work` in the state file; head =
the job named with pid/elapsed/remaining every interval, log line carries
the same detail.
2026-09-13 05:08:20 -07:00
teknium1
9b6dcad91d fix(utils): writers that published through mkstemp on main keep NEW files at 0600
0dfb4234 made every mode-less atomic write follow the process umask for NEW
targets, restoring what open("w")-based writers did. Ten of the folded sites
were not open("w") writers: they created the file through mkstemp and never
chmod'd, so on main a fresh file was 0600 regardless of umask (bot mailboxes,
relay inbox, turn markers, sessions.json, cron jobs/output, banner snapshot,
plugin toolset cache, presets, shell hooks, install id). CI caught the loosening
in tests/tools/test_bot_live_owner_delivery.py (st_mode 0o077 bits set).

Pass mode=0o600 explicitly at those ten sites; the umask default stays for the
sites that were open("w") on main. Invariant test exercises two real writers.
2026-09-13 05:07:11 -07:00
teknium1
3ef8b384a9 refactor(persistence): 24 hand-rolled atomic JSON/text writers go through utils.atomic_json_write / atomic_write_text
Each copy re-implemented temp+replace by hand and lacked one or more of
fsync, symlink preservation, atomic_replace's Windows-contention retry and
EXDEV/bind-mount fallback, mode preservation, or interrupt-safe temp
cleanup. Three (gateway/session_persistence, cron/suggestions,
agent/shell_hooks) were verbatim inlines of utils._atomic_write; two
modules defined their own directory-fsync helper, now utils.fsync_directory.
plugins/google_meet/_jsonfile.write_json_atomic is deleted (callers use the
canonical helper directly).

Behavior change: every one of these writers now fsyncs the payload, keeps a
pre-existing target's mode, cleans its temp file on BaseException, and
survives Windows AV/indexer contention and cross-device renames the way
config writes already did. cron/suggestions.json is 0600 from creation
(previously chmod'ed after the replace). Skipped on purpose: cron/jobs.py
two-phase staging, gateway/status._write_json_excl (create-only lock),
kanban_transfer staging (not atomic writers); tools/skill_usage.
_write_suppressed_names lives inside a PLUGIN-COMPAT block.
2026-09-13 05:07:11 -07:00
Teknium
ec58e08a35 feat(cron): automatic bounded re-runs when a fire never reached the model
Inspired by Claude Cowork (desktop changelog v1.46388.1, 2026-09-04), which
added "automatic re-runs (after 5, 15, and 30 minutes) for a scheduled task
that could not reach the model at all, for example right after the computer
wakes behind a VPN."

A recurring cron job whose run fails with a transient network/DNS error
before ANY model call previously sat out a full period (a daily job fired
into a reconnecting VPN silently skipped a day). Now the scheduler pulls
next_run_at earlier along a bounded 5/15/30-minute ladder, suppresses the
interim failure notice while a re-run is pending, and resets the ladder on
any run that reaches the model.

Deliberately narrower than a generic retry (cf. PR #16512): zero API calls +
transient classification means nothing executed and nothing was spent, so a
re-run cannot duplicate side effects. One-shots are excluded (at-most-times
dispatch accounting, #38758); retries never fire past the schedule's own
next occurrence; `cron.retry_unreachable: false` disables.

- cron/unreachable_retry.py: ladder, classification, plan/clear/will_retry
- cron/scheduler.py: flag unreachable failures in run_job; suppress interim
  notice; thread model_unreachable through the fenced bookkeeping write
- cron/jobs.py: mark_job_run schedules/clears the ladder under the jobs lock
- docs: website/docs/user-guide/features/cron.md
2026-09-12 22:13:33 -07:00
Søren L. Hansen
0037a4b17a feat(cron): add resnap action to adopt the current global inference default
Unpinned cron jobs snapshot the global provider/model at creation and fail
closed when the global default drifts (#44585). Pinning was the only way
forward, but it makes a job stop tracking the global default forever.

Add resnap: refresh an unpinned job's provider/model snapshot to the CURRENT
global resolution without pinning it, so it adopts the user's deliberately
changed default while keeping tracking future changes. Single job via
cronjob(action='resnap', job_id=...) or hermes cron resnap <id>; bulk via
cronjob(action='resnap', all=true) or hermes cron resnap --all. Refuses to
guess scope when neither is given. The drift-guard alert now points at both
options (pin vs resnap). No inference call is made — it recomputes the
snapshot string from config.
2026-09-12 20:57:21 -07:00
teknium1
819988acb7 docs(kanban): per-platform opt-in is the documented path (hermes tools enable kanban --platform X) 2026-09-12 12:32:55 -07:00
kshitijk4poor
f003e449be refactor(cron): trim the scope-degrade dispatch to its invariants
Follow-up to the cherry-picked #102431 fix, addressing the review findings:

- The two real-helper scheduler tests ran the Linux-only helper unmarked
  and failed on macOS/Windows; the surviving one is now `linux_only`.
- `_warn_scope_degraded_once` used a bare `os.getuid()` that tripped
  scripts/check-windows-footguns.py --all (lint lane red). The remedy text
  is now built once in the helper and passed into the warning, so the
  "scope binary vanished" case no longer warns about a missing D-Bus.
- Tests trimmed to the invariant bar: degraded != in_process and warns
  once (helper level); default config still Popens externally and
  `require_restart_safe_scope: true` raises (scheduler level, real helper).
  Dropped the stubbed duplicate, the standalone config-raise test and the
  in_process half already covered by the existing passthrough test.
- `GatewayChildDispatch.reason` had no reader outside a test; removed.
- Both degrade branches share one local `_degrade(detail)`.
- The per-fire config read uses `load_config_readonly()` (no deepcopy)
  with the same `except Exception` guard as the sibling
  `failure_nudge_threshold` read, so a config error no longer escapes
  the launcher.
- Kanban's no-run-id guard fails closed for any non-`in_process` mode
  instead of matching one enum value.
- Rationale restated in six places collapsed to the helper docstring;
  `cron.require_restart_safe_scope` documented in the cron user guide.
2026-09-12 23:17:12 +05:30
Paul Robertson
560b6d2e81 fix(cron): degrade gracefully when systemd user scopes are unavailable
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).

Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.

The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.

Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.

Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
2026-09-12 23:17:12 +05:30
kshitijk4poor
9a60a7f316 test(cron): one fail-fast guard for the heartbeat vs its own run's fence
Replace the POSIX-only jobs-flock contention test (skipped off-POSIX,
~120 LOC of monkeypatched flock plumbing) with a single invariant test
that fails on pre-fix code in <1s: hold the per-job fire fence from a
worker thread, assert the heartbeat still returns True on the calling
thread, and that a takeover is still detected (False). The docstring on
heartbeat_fire_claim now records WHY it is not under the fence, so the
next refactor does not put it back.

Co-authored-by: Oliver Heckmann <46627487+oheckmann74@users.noreply.github.com>
Co-authored-by: salch-cred <141555468+salch-cred@users.noreply.github.com>
2026-09-12 22:57:43 +05:30
KoNit-K
87b013b21e fix(cron): do not hold fire fence during heartbeat save_jobs
heartbeat_fire_claim only CAS-refreshes claim.at via _with_job; wrapping
_under_fire_fence across save_jobs let a blocked .jobs.lock pin the fence
and cause mark_job_run to fail closed on completed jobs.
2026-09-12 22:57:43 +05:30
teknium1
d1dbb0ac9e feat(gateway): multiplexer hot-serves profiles created while it runs, unroutes deleted ones
A `gateway.multiplex_profiles` gateway enumerated `profiles/` once at boot, so a profile
created afterwards (CLI, dashboard, Desktop, TUI) was never served until `hermes gateway
restart`; Desktop and the dashboard gave no reminder, so a new profile's bot simply never
connected.

The served set is now reconciled at runtime (`gateway/run_profile_reconcile.py`):
- `hermes_cli/profiles.py` create/delete ping the multiplexer over its control socket
  (new `rescan-profiles` verb); a supervised watcher rescans every 30s as the safety net.
- A new profile gets its adapters under its own runtime scope from its config/.env
  (`_start_one_profile_adapters`, same duplicate-credential guard as boot, now seeded
  with the LIVE secondaries' claims), `served_profiles` in gateway_state.json is
  updated, MCP discovery + log routing run for it. Other profiles' adapters are never
  touched.
- A served profile whose config.yaml/.env changed is re-scanned so a token added after
  create builds the adapter; already-live/queued platforms are skipped (no second poller).
- A deleted profile (tombstone) has its reconnects cancelled, adapters torn down,
  pairing/busy bookkeeping and cached agents dropped, and this process's SQLite /
  memory-store handles released so the deleter's rmtree succeeds.
- The in-process cron ticker takes a live enumerator so new profiles' jobs fire.
- PUT /api/messaging/platforms/<id>?profile=X returns `hot_served` when a live
  multiplexer rebuilt X's adapters; Desktop/dashboard skip the restart banner then.
- `hermes profile create` confirms hot-serve; the restart reminder stays for a gateway
  that did not pick the profile up (older build / signal failed).
2026-09-12 08:49:16 -07:00
teknium1
b78abb4710 fix(config): stop forcing 0700 on HERMES_HOME inside containers
Every start (and every `ensure_hermes_home` from a sibling CLI invocation)
chmod'd the data directory to 0700, wiping group/other bits and the ACL
mask on a bind mount shared with other containers (hermes-webui, Nix
desktop + dashboard). _secure_file already skipped containers for this
reason; _secure_dir did not.

In a container the directory mode is now left to the operator unless
HERMES_HOME_MODE is set explicitly, which is still applied. cron/jobs.py
had its own 0700/0600 copies that bypassed the managed/container rules;
they now delegate to the shared helpers so cron/output stops re-locking
the mount as well.

Fixes #10757
2026-09-12 08:43:43 -07:00
teknium1
66884c4397 fix(cron): honour job workdir on the inline _build_job_prompt script path too
Same bug class as the wake-gate call fixed in the picked commit: when a caller
invokes _build_job_prompt without a cached prerun_script, the sibling ran the
job's script inline via _run_job_script(script_path) and dropped the job's
configured workdir, so the script ran from the scripts-dir parent. Route it
through _resolve_job_workdir like the no_agent and wake-gate sites already do.

Sweep of every _run_job_script / _run_job_script_with_claim_heartbeat caller:
_run_no_agent_job and cron/monitor.py already pass workdir; the wake-gate site
is fixed by the salvaged commit; this was the last one.
2026-09-12 08:21:08 -07:00
Yong Chang Yi
62b0235d45 fix(cron): pass workdir to agent pre-run scripts 2026-09-12 08:21:08 -07:00
Teknium
6cd4fbd640 docs(cron): state the missed-occurrence contract for restart gaps
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch
advance is provisional, restore once, never twice, grace, opt-out, paused never
catches up, same on standalone and multiplexed); the user guide describes the
catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md
lists it as a hardening invariant.
2026-09-12 01:51:59 -07:00
Teknium
eaa5ac94ea fix(cron): fire one catch-up for a slot missed during a restart gap (#107485)
tick() advances a recurring job's next_run_at BEFORE dispatch so a crash
mid-run cannot re-fire it on every restart (at-most-once). That leaves a
window — advance persisted, fire claim not yet taken — in which the process
dying (interpreter already finalizing, executor refusing new futures, SIGKILL,
the Desktop idle-exit from #107485) loses the occurrence silently: the
restarted scan sees only the future next_run_at, writes no execution row and
no log line, and a daily job skips a day.

Contract (documented in cron-internals.md): every recurring occurrence is
accounted for — it runs once, or its skip is logged with a reason.

- The due scan stamps `pending_slot = {scheduled_at, at, by}` in the same save
  that records `last_dispatch`; claim_job_for_fire and mark_job_run clear it,
  and an explicit schedule / next_run_at / enabled / state rewrite (edit,
  pause, resume, run-now) drops it.
- A later scan that finds the stamp with a provably gone owner (this process
  and the job not in its running set, or another owner past the fire-claim
  lease / dead pid) restores scheduled_at as next_run_at ONCE and logs a
  WARNING (cron/occurrences.py::unclaimed_pending_slot). The restored instant
  then meets the ordinary policy — completed_occurrence() blocks a second fire
  of a slot that already ran, the grace window classifies it late, past-grace
  collapses the backlog into one run, cron.catch_up_missed: false skips it
  with a logged reason. Never a replay of N slots.

Same store fields on both topologies: a standalone `hermes -p X gateway run`
and a profile served by the default multiplexer evaluate the identical record.
2026-09-12 01:51:59 -07:00
Teknium
2bc471cb35 fix(cron): skip catch-up only with a future anchor 2026-09-12 01:51:59 -07:00
Teknium
df51797e2e feat(cron): let planned downtime skip missed recurring runs 2026-09-12 01:51:59 -07:00
Teknium
284d220ba4 fix(multiplex): cron, kanban, /loop and completion paths for a served profile match its standalone gateway
Under gateway.multiplex_profiles a secondary profile X is ticked, dispatched
and notified from the default profile's process, where os.environ holds the
DEFAULT profile's .env and X's values live only in the per-turn secret scope /
HERMES_HOME override. Every remaining read that skipped that scope made X
behave differently from `hermes -p X gateway run`:

- cron: HERMES_CRON_TIMEOUT, HERMES_MODEL (job/preflight fallback),
  HERMES_CRON_MAX_PARALLEL, inflight allowance, prefill file and the script
  timeout were bare os.getenv → the default profile's values; a job without a
  model silently ran on the default's HERMES_MODEL instead of refusing.
  cron/env_settings.py::cron_env_setting reads the scope (fire) or the ticked
  home's .env (tick thread), plain environ when multiplexing is off.
- child env: the restart-safe cron worker, the Bot Chat delivery child and the
  kanban worker inherited the launch profile's non-credential .env settings
  and bridged TERMINAL_* policy (TERMINAL_ENV=docker, default's image,
  HERMES_MODEL) — X's worker ran in the default's docker image on the
  default's model. tools/environments/local.py::strip_launch_profile_env drops
  them when the child targets another served profile.
- kanban: the worker --toolsets pin was silently dropped for every served
  assignee (toolset probes call get_secret without a scope → swallowed
  UnscopedSecretError); notifier pings, artifact uploads and the wake text ran
  under the default's media policy / display language (only wake() was scoped).
- /loop: _post_turn_loop_completion hopped to the executor without contextvars,
  writing the completed tick into the DEFAULT profile's state.db and leaving
  X's row awaiting_response forever; the --until judge ran with the default's
  aux credentials.
- background processes: a secondary's processes.json (scope-relative since
  adf23550f5) was never read at startup; its processes were not re-adopted
  and notify_on_complete notices were lost. Startup recovers every served
  home under its scope; recovery adopts each session once.
- completion delivery: background_process_notifications was evaluated once per
  drain for the ambient profile (default's mode for everyone; X's `off`
  dropped a sibling's `all` event), recovered watchers used the default's
  mode, HERMES_BACKGROUND_NOTIFICATIONS was read raw from environ;
  _deliver_platform_notice used the default's GatewayConfig so a secondary's
  notice_delivery: private went public.

Not changed: gateway/run.py and tools/async_delegation.py (PR #106742
rewrites both). Known residue left for the env-bridge lane:
HERMES_SESSION_STALL_TIMEOUT is bridged once from the launch config.
scheduler bug, not a parity gap; unchanged here.
2026-09-11 19:58:07 -07:00
Teknium
444fa8166a fix(cron): routed-profile cron delivers through the shared bot for guild-scoped routes and profiles without a platforms block
SharedRouteAdapters.get called ProfileRoute.matches without guild_id, so the
documented Discord route shape (guild_id + chat_id) never authorized a cron
target and the satellite fell to standalone delivery ("DISCORD_BOT_TOKEN is
not set" every fire). A cron target has no inbound guild anchor; the route's
own guild_id is passed so its target-exact discriminators decide.

_resolve_target_transport then vetoed the authorized shared transport on the
SATELLITE's platforms.<p>.enabled (absent block or enabled: false), although
that block describes a connector the satellite never runs. The shared hit now
builds the transport directly (keeping the satellite's non-credential platform
settings), and a live native adapter with no config block is no longer read as
"disabled" (#89302) — same normalization the relay path already had.

Fixes #89302
Co-authored-by: web3blind <264741654+web3blind@users.noreply.github.com>
2026-09-11 15:28:37 -07:00
fangliquanflq
c17629a0a2 fix(cron): scope restart-safe worker environment
(cherry picked from commit e57f719f942736104ee7ff79999f41b8f2a23d63)
2026-09-11 15:28:37 -07:00