Review finding (minor): after the module-first reorder in cron/scheduler_delivery.py the delivery.shutil.which -> /bin/hermes monkeypatches were unreachable; both tests already accept the module argv.
The fixture adaptation for the `python -m hermes_cli.main` launcher located the
CLI argv via `argv.index("-p")`, but profile homes (`profiles/<name>`) never get
a `-p` flag appended, so the `beta` parametrization raised ValueError inside the
fake subprocess.run and the delivery reported failure. Strip the launcher prefix
by shape instead (3 tokens for `python -m hermes_cli.main`, 1 for a binary).
The cron scheduler runs inside the long-lived gateway process and spawned
`hermes ... chat` for Bot Chat delivery through `shutil.which("hermes")`
first, falling back to `sys.executable -m hermes_cli.main` only when PATH
had no `hermes`. That is the same resolution order gateway.run.
_resolve_hermes_bin just flipped for /update and /restart (#111569): the
running interpreter's module argv is exactly this install, PATH is not.
Delivery now resolves the running install first and uses PATH only as the
fallback. Tests that asserted the PATH argv shape or armed on the `which`
seam are moved to the module-argv shape / the `find_spec` seam.
_start_desktop_cron_ticker now hands the built-in scheduler a callable that
re-enumerates profiles every tick (so a deleted profile stops being ticked
without an app restart). The sibling scope test in tests/cron still compared
the kwarg to a list snapshot and went red in CI; it now asserts the callable
resolves to the same homes, which is the contract the scheduler consumes.
Follow-up to the salvaged assertion: the elapsed-grace contract is what
the implementation promises, but the cancellation must still come from
the renewal-error path (>= 1 failed renewal observed), or a stub that
never gets called would pass. Docstring names why a minimum attempt
count was the flake (#111471).
_resolve_cron_enabled_toolsets returned None when _get_platform_tools
raised, and AIAgent reads None as "load every toolset": a malformed
platform_toolsets block (or a stale-module import error after an update)
turned the operator's cron restriction into the full default set, with
only a log warning. Unattended jobs process untrusted text, so that is a
privilege widening, not a safety net (#111380).
The resolver now raises a RuntimeError naming the cause; run_job's
existing failure path records it on the job (last_error, failure streak,
incident) and the agent is never constructed. Per-job enabled_toolsets
(unknown names included) and the MCP merge path never touch the
platform resolver and are unchanged; the disabled-toolset resolver has no
fail-open branch.
Live: platform_toolsets: oops -> before: run ok, enabled_toolsets=None,
98 tool names selected; after: run fails "Cron toolset resolution
failed, so this run was refused rather than given every tool", agent
never constructed. normal / unknown-per-job / mcp-merge shapes: identical
before and after.
Co-authored-by: Austin Bell <10687162+robertaustinbell@users.noreply.github.com>
The completed-occurrence dedup gate (due scan in _evaluate_due_job and the
fire claim in claim_job_for_fire) consumes the due slot and advances
next_run_at without a run and without a ledger row, so before this change
a skip left zero trace: no log line, no execution, last_status untouched
(#111414 reported exactly that silhouette). The gate now logs a WARNING
naming the job, the skipped instant and the completed execution that
already covers it, from the one seam both gates share.
The reporter's root mechanism — an off-tick run stamping the NEXT
occurrence's identity onto its row, which the dedup later honoured — is
already closed on main (ac10770894, a73b750391, cd685a22e6; 82ae78cc11
ignores such backdated rows), so a stale-stamped row fires normally; only
a genuinely completed slot is skipped, and now says so.
Co-authored-by: holny <holny@foxmail.com>
Follow-up to the salvaged #111528 hunk: _prepare_job_prompt passed the monitor
block through a conditional kwargs dict although _build_job_prompt already
treats a falsy runtime_data_prompt as absent. Drop the duplicate unit-level
positive test — the run_job-level test in test_monitor_kind.py covers the
same invariant through the real path — and keep the strict-user-prompt
control.
One TTL for both probe verdicts (the success-TTL constant collapses into
`_SYSTEMD_SCOPE_PROBE_TTL_SECONDS`), and the ack-wait loop consults
`scoped_spawn_lost_user_bus()` when a scoped dispatch exits before the
worker acknowledged: with `/run/user/<uid>/bus` gone the job error names
the missing bus and the enable-linger remedy instead of the wrapper's bare
`exit 1`, and the cached True flips so the next fire degrades to a direct
external subprocess rather than consuming another occurrence on a dead
wrapper. Contributor test trimmed to the revalidation invariant.
The future-instant guard in claim_job_for_fire dropped the occurrence
identity for ANY claim ahead of the stored next_run_at. A hosted/webhook
fire for the armed slot that arrives a few seconds early (the fire
scheduler's clock runs ahead of ours) was therefore treated as an
off-tick run: it ran occurrence-free, mark_job_run recomputed the same
cron slot from a now still before it, and the tick/misfire backstop then
ran the slot a second time.
Only claims at least FIRE_CLAIM_SKEW_SECONDS (60 s) ahead of the slot are
now classified off-tick, so dashboard/manual far-future fires stay
occurrence-free while a skewed early fire keeps the slot identity.
completed_occurrence honours the same window so the early run's
completion row (finished just before the slot) still proves the slot
done instead of being discarded as poison.
Review finding: claim_job_for_fire future-instant guard had no skew tolerance; an early hosted fire for the armed slot ran twice.
tests/cron/test_manual_fire_occurrence.py (from #106970) overlapped the two smaller
invariants salvaged from #110419: an unclassified off-tick claim stays occurrence-free
(test_claim_job_for_fire.py) and a completion recorded before its occurrence is not
proof (test_scheduled_occurrence.py). On-time / late binding is already pinned by
test_scheduled_occurrence.py::test_ledger_migration_and_completion_identity.
Direction (ii) of the t_a662ca7a fix menu, ported from the live install
(ffffab1bd6 on the deployed checkout) and layered on top of the already-
landed direction (i) (ac10770894, manual= flag): claim_job_for_fire() now
also declines to bind an occurrence whose instant is still in the future.
A scheduled tick only fires when now >= next_run_at (_evaluate_due_job
returns False while the stored occurrence is in the future), so a claim
arriving BEFORE the stored next occurrence cannot be the tick that owns
it: it is a manual / dashboard / webhook fire (including the
claim_ttl_seconds reclaim after an expired lease, which arrives with no
manual flag) and stays occurrence-free - ledger row scheduled_instant=NULL,
next_run_at untouched. On-time and late (catch-up) ticks keep the exact
at-most-once occurrence identity.
Live case 2026-09-08: job d88fde170fb4 [bot:carl] Daily Plan Evening
(0 19 * * *), manual run at 19:53 stamped scheduled_instant=
2026-09-10T02:00:00+00:00 (= next day 19:00 PDT) status='completed';
completed_occurrence() then refused every later manual fire ("Job is
already being fired by the scheduler; not run again.") and the next
scheduled tick dedupe-skipped its real delivery. Pre-fix poison rows in
existing ledgers are not repaired by this commit; repair remains
audit-preserving UPDATE-only via the retired reference script.
Regression test (tests/cron/test_manual_fire_occurrence.py): two
invariant tests proven red on base b88e6776f4 (mid-cycle manual fire and
webhook claim_fire bind the future instant) plus three green-on-base
guard nets (on-time tick binding, late catch-up binding, forced-fire
occurrence-free). Red on base: 2 failed / 3 passed; green post-fix:
5 passed. tests/cron suite: 1232 passed, 1 skipped, 1 pre-existing
environmental failure in
test_run_job_cron_execute_code_deny_does_not_pollute_later_gateway_execute_code
that fails identically on base.
Kanban: t_d36f3b54 / t_2a050d21 (defect t_a662ca7a)
Cron scripts resolve only against the job's own profile scripts/ dir (by design,
profiles never share files), so a job copied between profiles fails with a path
that looks plausible and no hint about why. The runtime error now names the
profile folder and the two fixes (copy the script, or `hermes cron edit`).
Closes#94821. Builds on #105775 (creation-time existence check, credited).
_validate_cron_script_path only checked that a no_agent/monitor script path
was relative and contained within scripts_dir, never that the file actually
existed — so a misplaced script registered fine and only failed at every
fire with a generic "Script not found" from scheduler_script.py, dropping
once-only jobs from jobs.json on failure.
The error messages also hardcoded "~/.hermes/scripts/" even though
resolution goes through get_hermes_home(), which is per-profile — actively
misleading users running non-default profiles into placing scripts in the
wrong directory.
Now the validator resolves the file and returns a "Script file not found"
error naming the actual resolved scripts_dir, surfacing the mistake at
creation time instead of at every scheduled fire.
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
Masking inert heredoc bodies before the referenced-script walk (previous
commit) also hid every path the body names, so a Python body that does
os.system('/x/restart.sh'), a bare '/x/restart.sh' line, or a nested cat
heredoc naming the script all went undetected — main caught them because
the walk read the script and found the lifecycle command inside.
The interpreter does execute those paths at runtime, so the walk must still
read them. Only the fail-closed verdicts (cloud placeholder, oversized or
binary file) stay restricted to the masked view: a >1 MiB data file merely
mentioned inside an inert body is not a script, which is the false positive
the previous commit fixed and its tests keep pinning.
Review finding: masking the heredoc body from the walk widened fail-open —
scripts executed by path from within the body were no longer scanned.
Keep the inert-body path (false) and the unquoted-body path (still true);
the `sh -c`-in-`cat` heredoc and plain script-reference cases are already
pinned by the existing lifecycle-guard suites. Comment shortened to the WHY.
The direct lifecycle scan masks provably-inert heredoc bodies
(strip_inert_heredoc_bodies), but the referenced-script and -c payload
walks ran on the unmasked command. A path inside such a body is never
shell-executed, so walking it was a pure false positive: a >1 MiB path
mentioned in a quoted heredoc body (e.g. python3 - <<'PY') failed closed
and hard-blocked an innocent command.
The walks now use the same masked view as the direct scan. Masking only
fires under the stripper's conservative contract (quoted delimiters,
exact terminator, single simple command, allowlisted consumer, no command
substitution); unquoted/shell-consumed/ambiguous bodies stay visible and
fail-closed.
Fixes#110422.
authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction
remove_job() deletes <cron>/output/<job_id>/ together with the record, but the
finishing run then called save_job_output(), re-creating the directory and
writing the final run into it. Every self-removing job leaked an orphan
directory the store no longer knew about, and the docs claim that "only the
job record is gone afterwards" was false. Skip the save when
self_removal_delivery_allowed() is true (the same check that already excuses
the missing record on the delivery and mark paths); delivery composes without
an output_file, as it already does for non-file paths.
The BaseException handler in _run_one_job_body still called mark_job_run on
the missing record after a self-removal crash. Guard it with the same check so
the crash path matches the completion path instead of probing a deleted record.
Docs: state that the record and its output directory are both gone.
Review finding: self-removed run re-creates the rmtree'd output dir (orphan leak); crash path marks a missing record.
Follow-up to the salvaged #111044 commits:
- self_removal_delivery_allowed() now also requires that no record currently
holds the job id. The marker alone said "this run removed its record"; it did
not say the id is still empty. A replacement record (another owner reclaiming
the id) must be treated as a stolen claim, not a self-removal.
- Drop the allow_self_removed kwarg on fire_claim_fence: the fence already has
the job_id and the ContextVar marker, so it can decide on its own; the caller
no longer threads a flag it computed from the same predicate.
- _FireOwnership.lost(): keep the explicit lost event and the no-owner short
circuit ahead of the self-removal check so an interrupted run is still
reported as lost even after it removed its record.
- _finish_completed_run: skip mark_job_run entirely for a self-removed job
(nothing to mark) instead of calling it and then excusing the False.
- Tests trimmed to two invariants, both A/B'd against origin/main: the
self-removing run delivers after a post-removal heartbeat tick (RED on main),
and a self-removal followed by a replacement record is still discarded
(GREEN on main, guards the new predicate).
- Docs: user-guide cron.md notes that a job may remove itself and still report.
A run that deletes its own job after the first heartbeat interval was still marked stale because the fire-claim loop treated a missing record as lost ownership before the self-removal marker could win.
Co-authored-by: Cursor <cursoragent@cursor.com>
Trim the salvaged suite from four tests to the two invariants that were red on
main: (1) with the sentinel engaged POST /api/cron/fire answers 503 +
Retry-After 60 and never calls claim_fire, and the same job is admitted (202,
claimed, fired) once the sentinel is removed; (2) fire_overdue_jobs dispatches
nothing and leaves next_run_at untouched while engaged, and the first sweep
after resume catches the job up through claim_fire. The webhook test lives
beside the other cron-fire webhook tests (test_cron_fire_webhook.py) and uses
their real spy provider instead of a MagicMock resolver; the "verifier crashes
-> 401" case was already covered there.
Docs: cron.md gains a "Pausing everything: hermes pause" section stating that
all three automated doors honour pause, that in-flight runs are never killed,
and that manual runs are an operator override; the CLI reference table lists
hermes pause / hermes resume.
Two inline test-hardening suggestions + two follow-up coverage cases
(Xipong called them non-blocking but they're cheap and prove the
boundaries).
- tests/gateway/test_api_server_jobs.py::TestCronFireEstop::
- test_fire_webhook_returns_503_when_estop_engaged: also patch
`cron.scheduler_provider.resolve_cron_scheduler` and assert that
`claim_fire`, `fire_claimed`, and `fire_due` are never reached.
A 503 by itself does not prove admission never happened.
- test_fire_webhook_401_when_verifier_crashes (new): crashing
verifier → 401, ESTOP is never consulted. Proves auth runs before
the ESTOP check so the sentinel state cannot leak to unauth callers.
- tests/cron/test_misfire_catchup.py::TestFireOverdueJobs::
- test_estop_engaged_skips_backstop: spy on `provider.claim_fire`
and `cron.scheduler_provider.threading.Thread`. Deterministic
proof that no claim was attempted and no worker thread was
constructed (the prior `wait_fired(timeout=0.5)` was timing-based).
- test_estop_release_restores_backstop: spy on `provider.claim_fire`
on the recovery path with an explicit `assert_called_once` so a
regression that fires without claiming still fails the test.
All 58 tests across the affected files pass.
`agent/estop.py:1-9` documents that while the sentinel exists "the cron
scheduler ... skips work." The built-in ticker honors this
(`cron/scheduler.py:3749-3754`), but the managed-cron paths do not:
- Door 2 (NAS fire webhook `_handle_cron_fire`,
`gateway/platforms/api_server.py:3480-3557`): no ESTOP check between
the JWT/drain guards and `provider.claim_fire`. Added a
`check_paused("cron-webhook")` guard inside the reservation block,
returning 503 + Retry-After so the NAS retries after `hermes resume`
rather than silently dropping the run.
- Door 3 (misfire backstop `fire_overdue_jobs`,
`cron/scheduler_provider.py:253-344`): no ESTOP check at the top of
the function. Added an early-return `check_paused("cron-misfire")`
guard. Self-healing — the next sweep after `hermes resume` catches
everything up via the existing claim_fire path.
Both guards use `suppress(ImportError)` matching the ticker idiom so a
broken estop module fails open rather than killing cron. Distinct
component names keep the existing log-once mechanism independent per
surface.
Manual runs (`hermes cron run`, dashboard Trigger) deliberately
unchanged: operator override is arguably a feature, and PR #105144
already rewrites that path.
Tests (3 new, 49 pre-existing in affected files all pass):
- tests/cron/test_misfire_catchup.py::test_estop_engaged_skips_backstop
- tests/cron/test_misfire_catchup.py::test_estop_release_restores_backstop
- tests/gateway/test_api_server_jobs.py::test_fire_webhook_returns_503_when_estop_engaged
Answers the P1 review on #111187: build_profile_secret_scope() held only
<profile>/.env plus that profile's external-source snapshot, never the
administrator-managed .env. The launch process applies that file LAST with
override (_apply_managed_env), so a managed key beats the user's own value in
os.environ. Under multiplex semantics get_secret() stops falling back to
os.environ on a scope miss, so inside a routed cron fire (and equally inside a
real multiplex gateway turn, which builds its scope through the same function
via gateway/run.py::_load_profile_secret_scope) a managed-only credential
resolved as absent and a managed-vs-user collision resolved to the USER value:
reversed precedence.
Fix at the source: build_profile_secret_scope() overlays load_managed_env()
last, after the profile .env and external sources, skipping process-global
names exactly as it does for the other two layers. Every multiplex-authoritative
scope (gateway turn, routed desktop fire, external worker env build) is built
here, so managed authority is composed once instead of restored per consumer.
No generic ambient-env fallback is reintroduced: only the managed file's own
keys enter the scope, and only with the managed file's values.
Regression (parametrized, two invariants): inside a routed fire a managed-only
key resolves through get_secret(); a managed-vs-user collision yields the
managed value.
The earlier trim dropped the only test of the `routed_profile_fire() and
not is_multiplex_active()` branch in the external-worker handoff: the
process flag is OFF (a desktop tick is not a multiplexer) yet a fire routed
to a sibling profile must serialize `multiplex_active=True` and hand the
worker an env without the launch profile's residue, and the context must
not outlive the handoff span. Without it that branch could regress to the
pre-fix behaviour unnoticed.
The PR shipped nineteen tests across six files, most of them variations of
one boundary. Keep the three that pin distinct behaviour:
- a routed desktop-ticker fire runs under multiplex semantics for exactly
its scope: a scope miss returns None instead of the launch credential and
the parent os.environ is byte-identical afterwards;
- a routed no_agent child never sees a launch-only name, whether the launch
.env defined it or a launch external source supplied it (applied or lost
to a pre-existing process value), while its own values come through;
- administrator-managed keys keep policy precedence over the routed
profile's own value.
Everything else was either a positive control of the same seam, a
set-membership check on a module-level constant, or a re-statement through
a different entry point.
Review findings on f5f88d5058. Three are defects the previous round introduced.
Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.
Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.
Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.
Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.
Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.
(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
Review findings on d8c467f223, each reproduced through its production path.
Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.
Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.
Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).
Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.
(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
The source-name strip added alongside the external-source fix ran
unconditionally. Outside multiplexing there is no other profile to leak from --
os.environ IS this profile's environment -- so popping those names relied on the
routed scope overlay putting each one back, which in turn relies on the
per-home snapshot recorded at boot. Correct today, but it made a single-profile
child's credentials depend on bookkeeping that has nothing to do with isolation.
Guard it the way strip_launch_profile_env guards itself: no multiplexing, no
strip. A single-profile no_agent child now keeps a byte-identical env even if a
source's snapshot were ever missing. Pinned by a regression that runs a real
child with a source-owned name in os.environ and no multiplex context; making
the strip unconditional fails it.
(cherry picked from commit afa429b30a1c9c9b4c011f7d2426a9f099911596)
Two credential-isolation gaps found in review of the previous head.
1. strip_launch_profile_env() only knows dotenv- and terminal-config-owned names,
but external secret sources (vault, 1Password, ...) also write their names into
the shared os.environ and are tracked in secret_source_names(). A name the LAUNCH
profile's source supplied therefore still reached a routed no_agent child. Drop
every non-global source-owned name from the base; the routed scope overlay that
follows puts back exactly the ones that profile's OWN sources supply, since
build_profile_secret_scope folds get_secret_source_values(home) in.
2. refresh_installed_secret_scope() merged the rebuild with dict.update(), so a
name a source had stopped supplying -- rotated, revoked, source removed -- kept
its old value for the rest of the fire. Replace the mapping contents instead: the
rebuild is the profile's current truth.
Regressions: a routed child sees <unset> for a launch-source name while its own
source value comes through, and a refresh whose rebuild omits a name drops it.
Both fail if the corresponding change is reverted.
(cherry picked from commit ecd51517c4828a75acb5f458ed99aea0bc3e5e9f)
The routed no_agent child env started from all of os.environ and only overwrote the
names present in the installed scope. A name defined only by the LAUNCH profile's .env
and absent from the routed profile therefore reached the routed child with the launch
value instead of unset -- the secret scrub only knows classified names, so a custom or
unclassified secret crossed the profile boundary (review finding on the first head).
strip_launch_profile_env (main, 284d220ba4) is the primitive the external-worker path
already uses for exactly this: it drops the launch profile's dotenv-owned keys and the
bridged TERMINAL_* settings, and is a no-op outside multiplex or when the target IS the
launch profile. Apply it to the base BEFORE the scope overlay (so a shared name keeps its
routed value) and BEFORE the sanitizer (so routed values still pass the scrub and
passthrough rules). Pinned by a child-process negative control: the launch-only name
arrives <unset>, the shared name arrives routed, and the parent process is unchanged.
(cherry picked from commit 69349b527b5144c1f1540a4c10035d5ec798c1db)
The desktop backend ticks EVERY local profile's cron store from one process — its own docstring
says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global
multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every
isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared
`os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban
subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary
profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed
there after the tick, and a scope miss read the launch profile's tokens (#107692).
Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py)
is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not
the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the
override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the
profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that
span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are
therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs
before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam,
left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation
applies inside the routed fire with no per-site patching; the launch profile's own fires and the
backend's turns keep single-profile semantics; marker and override both reach the pool worker via
`copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through
`is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970).
Two consequences of suppressing the write are handled rather than left as regressions:
- a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`;
the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub /
passthrough rules apply to those values and the parent process is never mutated;
- plugin secret sources are discovered on the fire's first agent build, after the scope froze,
and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds
the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern
`_publish_env_value` already uses for `.env` writes under multiplex).
And the profile's external secret sources are hydrated before the scope is frozen, the order
gateway/run.py and the external cron worker already use.
Tests pin each direction: the marker without the semantics before the scope, the semantics on and
off exactly with it, the marker reaching a copy_context worker; the process's own profile staying
single-profile; the restart-safe handoff's child env building without raising under a routed tick
with a passthrough key registered; a real child process receiving the routed values while
`os.environ` keeps the launch value; a source registered after the freeze reaching the fire
through the real PluginManager refresh. Reverting any one direction fails a distinct test.
(cherry picked from commit 2f87677425d2cca19286ac83bc45cab23e546669)
Catch unexpected delivery exceptions after claim, retain diagnostics and continue
sibling admissions without authorizing replay. Preserve indefinite retention.
Reproduced PermissionError at target traversal after discovery. Native Electron
controlled-fault A/B confirms the healthy sibling settles and renders once.
Extend deferred dispatch's destination pin to ordinary CLI fallback, so
custom-root and active-profile changes cannot redirect a checked target.
Refuse a missing destination before launch and name the target on failure.
Replace the old env-clearing expectation with two behavioral invariants
and retain the native Electron custom-root reproduction.
Adapted from the root-boundary fix and diagnosis in #104066.
Related #104055, #104066.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Carry the original destination home and delivery ID into deferred drain and
its child, rather than re-resolving a mutable profile/root. Missing destinations
fail closed; supported-owner handoffs remain transferred, not ambiguous failures.
Capture the producer root before the background thread starts, and retain/log
malformed JSON without stopping healthy admissions or the whole cron tick.
Two invariants reproduced failures on the published head. Real Electron root
change and malformed-record cases are red before and green after; nested DM
control remains passing. No automatic retry of claimed or uncertain turns.
Keep never-started output behind unsupported owners and drain in admission
order after release. Persist claims before execution and never replay uncertain
started turns. Existing supported-owner receipts keep their authority.
Credits 686f6c61's residual queue proposal in #100319. This is a scoped
implementation, not general retry of failed CLI subprocesses.
Native Electron before/after: CLI-owned target previously returned
SESSION_NOT_OWNED and remained empty after release/tick; now its queued
output and reply appear once in the target Bot Chat. Nested quiet CLI
message_agent delivery to a named Desktop owner also passes on base.
Follow-up on the salvaged #111034:
- `_start_multiplex` published the enumerated home list before the gate had
filtered it, so a raising `profile_gate` (the Desktop stand-down probe from
#100489) kept the thread alive but ticked every profile UNGATED — racing the
gateway that owns them for the same cron store. The list is now assigned only
after gating; a gate failure yields zero ticks for that cycle.
- `cron/scheduler_thread.py::SupervisedTickerThread` wraps the gateway ticker
thread; `_start_gateway_housekeeping` gets a per-tick "Cron ticker supervisor"
chore that respawns a ticker that ended without a stop request and logs the
outage at ERROR. Every guard inside `start()` keeps the loop alive, but nothing
outside it could notice a thread that had already ended.
- Tests trimmed to the invariants, proven red on origin/main: a REAL corrupt
`executions.db` (no patched recover) no longer kills the ticker; a raising gate
keeps the thread alive with zero ticks; housekeeping restarts a dead ticker and
leaves a stopped one alone.
Root cause: the unguarded pre-loop `recover_interrupted_executions()` +
`record_ticker_heartbeat()` were added by d9dd05b69d (#61791, "truthful
execution ledger", 2026-07-09). The reporter's build (e440bf35) also carried
#107485's `completed_occurrence()` in the due scan, which opens the same ledger
on every tick — the first traceback in their errors.log is that in-loop hit
(caught); the restart then hit the SAME corrupt ledger from the pre-loop
recovery scan, which nothing caught: thread dead, no heartbeat, no error marker.
The gateway runs InProcessCronScheduler on an unsupervised daemon thread, so any exception escaping outside the guarded tick body ends cron silently while the gateway keeps serving. Guard the three remaining escape windows: startup recovery + initial heartbeat in start(), per-cycle profile enumeration/gating in _start_multiplex, and every status-marker store write (heartbeat/error/clear) in both loops. A failing store now degrades to a logged warning and a failed cycle instead of thread death.
The incident ledger only ever grew. A one-off failure (a drift skip after a
global model bump, a provider outage) stayed `detected`/`alerted` forever
after the job recovered, so `hermes cron incidents` listed 32 "open"
incidents on an install where all 32 jobs had since run OK, and the list
stopped saying anything about current health.
A successful run now marks that job's `detected`/`alerted` incidents
`resolved` (new state). `resolved` is distinct from the operator's `closed`
ack on purpose: `upsert_incident` re-opens a resolved incident as
`detected` when the same error signature recurs, so the operator is alerted
again for a job that broke a second time, while `closed` keeps the
signature silent as before. Wired from `_compose_run_delivery` next to the
failure-side upsert; best-effort, store errors never affect delivery.
CLI: `--state resolved` filter and a green `resolved` colour; `closed` is
now dim. Docs updated in the same change.
Rebased onto main (~9700 commits): the branch predates the god-file
decomposition and the tests/state -> tests/hermes_state move.
Fixes for seams that moved:
- cron memory contract patches cron.scheduler_delivery._resolve_origin and
hermes_state_registry.acquire (where run_job now reads them).
- state.db conformance imports _live_writer_holds_db from
hermes_state_repair and drives the public create_quick_snapshot.
- update receipt: serve runtimes are reconciled in their own unit
vocabulary since #100479, so the "full accounting" row names the serve
unit instead of borrowing a gateway relaunch.
Deleted (change-detectors / source-greps / duplicates / dead code):
- source-text scan of hermes_state*.py for "maintenance-shaped" defs and
the symbol REGISTRY it fed (renames are not regressions).
- per-op exact-outcome table under a live writer -> one invariant: refuse
with a lock error or report zero work, DB stays intact.
- copy_db_and_verify pins: the symbol is a revert-scheduled plugin-compat
pointer, not production code (check_compat_pointers.py).
- cron ON-direction tests already pinned by tests/cron/test_scheduler.py,
plus a tautology that never called production code.
- exact warning-wording asserts in the env deprecation truth table.
- same-file duplicates: fixed-point (implied by idempotence), tripwire
round-trips, absent-registry, sequential "race" re-enactment.
Two weeks of closed issues/merged PRs show the same areas regenerating:
each salvage pinned its instance while the class invariant had no test.
These suites pin the invariants themselves:
- tests/conformance/test_profile_write_tripwire.py — no writes to the
default profile tree while a profile is active (#88532#92662#89190#89625#92156); reusable tripwire fixture, 4 surfaces
- tests/hermes_cli/test_env_deprecation_truthtable.py — 18-row truth
table for the Deprecated-.env warning (#88829#89016#89389#90299)
- tests/cron/test_cron_memory_contract.py — cron<->memory contract that
flipped twice in Aug (#91269 -> #91384 -> #91447)
- tests/agent/test_injected_param_strip_retry_registry.py — every
strippable injected param x real 400 shapes must strip-and-retry;
unknown params must still fail (#90257#89897#91164#89503)
- tests/agent/test_transcript_decoration_idempotence.py — f(f(x))==f(x)
law + 4-breakpoint budget for apply_anthropic_cache_control (#90971)
- tests/state/test_state_db_maintenance_conformance.py — registry-
enumerated maintenance ops refuse/degrade under a live writer; copies
of corrupt DBs are refused or flagged (#91839#90806#90613#88235)
- tests/tools/test_bot_mode_canonical_chat_resolution.py — canonical
Bot Chat resolution is idempotent, never mints, unique per profile,
race-safe (#92040#90705#92692#90005#90732, PR #92129)
- tests/hermes_cli/test_update_receipt_truthfulness.py — receipts:
crash never claims success; success requires full fleet accounting;
refusal != failure (#91283#91439#92902#92780)
117 tests, all sabotage-verified (each suite proven to FAIL when its
bug class is reintroduced).
A scheduled job whose prompt carries its own cadence phrasing ('Each
Monday, review...') can convince the agent to create ANOTHER cron job
at execution time instead of just doing the work — each run spawning a
sibling job. Hermes policy-denies the cronjob toolset in cron context
by default, but cron.allow_agent_scheduling: true re-enables it and
opens exactly this loop.
_build_job_prompt now extends the always-injected cron hint with a
RECURSION clause: this is a run of an existing job; never create or
update a cron job from schedule language in the task prompt; treat
cadence phrasing as context for this run.
Adapted from paradigmxyz/centaur#1479 (same failure mode in their
scheduled-task workflow runner).
Port from QwenLM/qwen-code#11723: attaching the configured zone to croniter's
naive wall-clock result resolves the repeated autumn hour to its earlier
occurrence (fold=0), so a base inside the second occurrence received a
next_run_at up to an hour in the past — the fire path would treat it as due
immediately and loop. Try both folds of each candidate and return the earliest
instant strictly after the base; wall-clock jobs still fire exactly once on
the repeated hour (Vixie cron semantics).