Commit Graph

552 Commits

Author SHA1 Message Date
teknium1
173ccfe2dd fix: drop dead shutil.which patches in bot-chat delivery tests
Review finding (minor): after the module-first reorder in cron/scheduler_delivery.py the delivery.shutil.which -> /bin/hermes monkeypatches were unreachable; both tests already accept the module argv.
2026-09-15 19:03:08 -07:00
teknium1
b0bde32959 test(cron): bot-chat CLI-home test strips the launcher prefix without requiring -p
The fixture adaptation for the `python -m hermes_cli.main` launcher located the
CLI argv via `argv.index("-p")`, but profile homes (`profiles/<name>`) never get
a `-p` flag appended, so the `beta` parametrization raised ValueError inside the
fake subprocess.run and the delivery reported failure. Strip the launcher prefix
by shape instead (3 tokens for `python -m hermes_cli.main`, 1 for a binary).
2026-09-15 19:03:08 -07:00
teknium1
f336048b08 fix(cron): bot-chat delivery launches the running install, not whatever hermes PATH names
The cron scheduler runs inside the long-lived gateway process and spawned
`hermes ... chat` for Bot Chat delivery through `shutil.which("hermes")`
first, falling back to `sys.executable -m hermes_cli.main` only when PATH
had no `hermes`. That is the same resolution order gateway.run.
_resolve_hermes_bin just flipped for /update and /restart (#111569): the
running interpreter's module argv is exactly this install, PATH is not.

Delivery now resolves the running install first and uses PATH only as the
fallback. Tests that asserted the PATH argv shape or armed on the `which`
seam are moved to the module-argv shape / the `find_spec` seam.
2026-09-15 19:03:08 -07:00
teknium1
aa03d38612 test(cron): desktop ticker scope test reads the live profile enumerator
_start_desktop_cron_ticker now hands the built-in scheduler a callable that
re-enumerates profiles every tick (so a deleted profile stops being ticked
without an app restart). The sibling scope test in tests/cron still compared
the kwarg to a list snapshot and went red in CI; it now asserts the callable
resolves to the same homes, which is the contract the scheduler consumes.
2026-09-15 18:54:28 -07:00
teknium1
e33693c932 test(cron): the heartbeat grace test also pins that a failed renewal preceded cancellation
Follow-up to the salvaged assertion: the elapsed-grace contract is what
the implementation promises, but the cancellation must still come from
the renewal-error path (>= 1 failed renewal observed), or a stub that
never gets called would pass. Docstring names why a minimum attempt
count was the flake (#111471).
2026-09-15 18:32:25 -07:00
KoNit-K
ea92bbc5c2 test(cron): stabilize heartbeat grace assertion 2026-09-15 18:32:25 -07:00
teknium1
9013fcdc87 fix(cron): an unreadable cron toolset restriction fails the run instead of granting every tool
_resolve_cron_enabled_toolsets returned None when _get_platform_tools
raised, and AIAgent reads None as "load every toolset": a malformed
platform_toolsets block (or a stale-module import error after an update)
turned the operator's cron restriction into the full default set, with
only a log warning. Unattended jobs process untrusted text, so that is a
privilege widening, not a safety net (#111380).

The resolver now raises a RuntimeError naming the cause; run_job's
existing failure path records it on the job (last_error, failure streak,
incident) and the agent is never constructed. Per-job enabled_toolsets
(unknown names included) and the MCP merge path never touch the
platform resolver and are unchanged; the disabled-toolset resolver has no
fail-open branch.

Live: platform_toolsets: oops -> before: run ok, enabled_toolsets=None,
98 tool names selected; after: run fails "Cron toolset resolution
failed, so this run was refused rather than given every tool", agent
never constructed. normal / unknown-per-job / mcp-merge shapes: identical
before and after.

Co-authored-by: Austin Bell <10687162+robertaustinbell@users.noreply.github.com>
2026-09-15 18:31:55 -07:00
teknium1
f7d2bb95e1 fix(cron): a due slot skipped as already completed is logged, not silent
The completed-occurrence dedup gate (due scan in _evaluate_due_job and the
fire claim in claim_job_for_fire) consumes the due slot and advances
next_run_at without a run and without a ledger row, so before this change
a skip left zero trace: no log line, no execution, last_status untouched
(#111414 reported exactly that silhouette). The gate now logs a WARNING
naming the job, the skipped instant and the completed execution that
already covers it, from the one seam both gates share.

The reporter's root mechanism — an off-tick run stamping the NEXT
occurrence's identity onto its row, which the dedup later honoured — is
already closed on main (ac10770894, a73b750391, cd685a22e6; 82ae78cc11
ignores such backdated rows), so a stale-stamped row fires normally; only
a genuinely completed slot is skipped, and now says so.

Co-authored-by: holny <holny@foxmail.com>
2026-09-15 18:31:26 -07:00
teknium1
719a420ec6 refactor(cron): pass monitor context as a plain kwarg; trim salvage tests to two
Follow-up to the salvaged #111528 hunk: _prepare_job_prompt passed the monitor
block through a conditional kwargs dict although _build_job_prompt already
treats a falsy runtime_data_prompt as absent. Drop the duplicate unit-level
positive test — the run_job-level test in test_monitor_kind.py covers the
same invariant through the real path — and keep the strict-user-prompt
control.
2026-09-15 18:30:59 -07:00
KoNit-K
e7cb478db6 fix(cron): sanitize monitor runtime prompt data 2026-09-15 18:30:59 -07:00
teknium1
a313a211d7 fix(cron): a scoped worker whose user bus vanished names the cause and re-probes
One TTL for both probe verdicts (the success-TTL constant collapses into
`_SYSTEMD_SCOPE_PROBE_TTL_SECONDS`), and the ack-wait loop consults
`scoped_spawn_lost_user_bus()` when a scoped dispatch exits before the
worker acknowledged: with `/run/user/<uid>/bus` gone the job error names
the missing bus and the enable-linger remedy instead of the wrapper's bare
`exit 1`, and the cached True flips so the next fire degrades to a direct
external subprocess rather than consuming another occurrence on a dead
wrapper. Contributor test trimmed to the revalidation invariant.
2026-09-15 06:28:07 -07:00
teknium1
fcbdcdb428 fix(cron): let a skewed early fire own its armed slot
The future-instant guard in claim_job_for_fire dropped the occurrence
identity for ANY claim ahead of the stored next_run_at. A hosted/webhook
fire for the armed slot that arrives a few seconds early (the fire
scheduler's clock runs ahead of ours) was therefore treated as an
off-tick run: it ran occurrence-free, mark_job_run recomputed the same
cron slot from a now still before it, and the tick/misfire backstop then
ran the slot a second time.

Only claims at least FIRE_CLAIM_SKEW_SECONDS (60 s) ahead of the slot are
now classified off-tick, so dashboard/manual far-future fires stay
occurrence-free while a skewed early fire keeps the slot identity.
completed_occurrence honours the same window so the early run's
completion row (finished just before the slot) still proves the slot
done instead of being discarded as poison.

Review finding: claim_job_for_fire future-instant guard had no skew tolerance; an early hosted fire for the armed slot ran twice.
2026-09-15 06:07:32 -07:00
teknium1
7cecc9eb3d test(cron): keep two invariant tests for the future-instant guard
tests/cron/test_manual_fire_occurrence.py (from #106970) overlapped the two smaller
invariants salvaged from #110419: an unclassified off-tick claim stays occurrence-free
(test_claim_job_for_fire.py) and a completion recorded before its occurrence is not
proof (test_scheduled_occurrence.py). On-time / late binding is already pinned by
test_scheduled_occurrence.py::test_ledger_migration_and_completion_identity.
2026-09-15 06:07:32 -07:00
fangliquan
ced1bf29a8 test(cron): cover future occurrence poisoning 2026-09-15 06:07:32 -07:00
Chris Herrera
cd685a22e6 fix(cron): manual fires before the next occurrence must not consume it
Direction (ii) of the t_a662ca7a fix menu, ported from the live install
(ffffab1bd6 on the deployed checkout) and layered on top of the already-
landed direction (i) (ac10770894, manual= flag): claim_job_for_fire() now
also declines to bind an occurrence whose instant is still in the future.

A scheduled tick only fires when now >= next_run_at (_evaluate_due_job
returns False while the stored occurrence is in the future), so a claim
arriving BEFORE the stored next occurrence cannot be the tick that owns
it: it is a manual / dashboard / webhook fire (including the
claim_ttl_seconds reclaim after an expired lease, which arrives with no
manual flag) and stays occurrence-free - ledger row scheduled_instant=NULL,
next_run_at untouched. On-time and late (catch-up) ticks keep the exact
at-most-once occurrence identity.

Live case 2026-09-08: job d88fde170fb4 [bot:carl] Daily Plan Evening
(0 19 * * *), manual run at 19:53 stamped scheduled_instant=
2026-09-10T02:00:00+00:00 (= next day 19:00 PDT) status='completed';
completed_occurrence() then refused every later manual fire ("Job is
already being fired by the scheduler; not run again.") and the next
scheduled tick dedupe-skipped its real delivery. Pre-fix poison rows in
existing ledgers are not repaired by this commit; repair remains
audit-preserving UPDATE-only via the retired reference script.

Regression test (tests/cron/test_manual_fire_occurrence.py): two
invariant tests proven red on base b88e6776f4 (mid-cycle manual fire and
webhook claim_fire bind the future instant) plus three green-on-base
guard nets (on-time tick binding, late catch-up binding, forced-fire
occurrence-free). Red on base: 2 failed / 3 passed; green post-fix:
5 passed. tests/cron suite: 1232 passed, 1 skipped, 1 pre-existing
environmental failure in
test_run_job_cron_execute_code_deny_does_not_pollute_later_gateway_execute_code
that fails identically on base.

Kanban: t_d36f3b54 / t_2a050d21 (defect t_a662ca7a)
2026-09-15 06:07:32 -07:00
teknium1
248938b4e6 fix(cron): "Script not found" says scripts are per-profile and how to fix it
Cron scripts resolve only against the job's own profile scripts/ dir (by design,
profiles never share files), so a job copied between profiles fails with a path
that looks plausible and no hint about why. The runtime error now names the
profile folder and the two fixes (copy the script, or `hermes cron edit`).

Closes #94821. Builds on #105775 (creation-time existence check, credited).
2026-09-15 04:15:16 -07:00
chelsealong
7ecbe50d98 fix(cron): validate script existence and fix profile-aware path in error messages
_validate_cron_script_path only checked that a no_agent/monitor script path
was relative and contained within scripts_dir, never that the file actually
existed — so a misplaced script registered fine and only failed at every
fire with a generic "Script not found" from scheduler_script.py, dropping
once-only jobs from jobs.json on failure.

The error messages also hardcoded "~/.hermes/scripts/" even though
resolution goes through get_hermes_home(), which is per-profile — actively
misleading users running non-default profiles into placing scripts in the
wrong directory.

Now the validator resolves the file and returns a "Script file not found"
error naming the actual resolved scripts_dir, surfacing the mistake at
creation time instead of at every scheduled fire.
2026-09-15 04:13:47 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1
56e563d3a4 fix: keep reading scripts named inside masked heredoc bodies
Masking inert heredoc bodies before the referenced-script walk (previous
commit) also hid every path the body names, so a Python body that does
os.system('/x/restart.sh'), a bare '/x/restart.sh' line, or a nested cat
heredoc naming the script all went undetected — main caught them because
the walk read the script and found the lifecycle command inside.

The interpreter does execute those paths at runtime, so the walk must still
read them. Only the fail-closed verdicts (cloud placeholder, oversized or
binary file) stay restricted to the masked view: a >1 MiB data file merely
mentioned inside an inert body is not a script, which is the false positive
the previous commit fixed and its tests keep pinning.

Review finding: masking the heredoc body from the walk widened fail-open —
scripts executed by path from within the body were no longer scanned.
2026-09-15 04:00:34 -07:00
teknium1
1fd0ef7520 test(cron): trim the heredoc-walk regression to its two invariants
Keep the inert-body path (false) and the unquoted-body path (still true);
the `sh -c`-in-`cat` heredoc and plain script-reference cases are already
pinned by the existing lifecycle-guard suites. Comment shortened to the WHY.
2026-09-15 04:00:34 -07:00
Kevin Rajan
59a1403fa2 fix(cron): mask inert heredoc bodies before the referenced-script walk
The direct lifecycle scan masks provably-inert heredoc bodies
(strip_inert_heredoc_bodies), but the referenced-script and -c payload
walks ran on the unmasked command. A path inside such a body is never
shell-executed, so walking it was a pure false positive: a >1 MiB path
mentioned in a quoted heredoc body (e.g. python3 - <<'PY') failed closed
and hard-blocked an innocent command.

The walks now use the same masked view as the direct scan. Masking only
fires under the stripper's conservative contract (quoted delimiters,
exact terminator, single simple command, allowlisted consumer, no command
substitution); unquoted/shell-consumed/ambiguous bodies stay visible and
fail-closed.

Fixes #110422.

authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction
2026-09-15 04:00:34 -07:00
Hukla
43891e9f8e fix(cron): inject skill config into scheduled runs 2026-09-15 03:58:27 -07:00
teknium1
134ef6454d fix(cron): self-removed runs leave no output directory and skip mark_job_run on crash
remove_job() deletes <cron>/output/<job_id>/ together with the record, but the
finishing run then called save_job_output(), re-creating the directory and
writing the final run into it. Every self-removing job leaked an orphan
directory the store no longer knew about, and the docs claim that "only the
job record is gone afterwards" was false. Skip the save when
self_removal_delivery_allowed() is true (the same check that already excuses
the missing record on the delivery and mark paths); delivery composes without
an output_file, as it already does for non-file paths.

The BaseException handler in _run_one_job_body still called mark_job_run on
the missing record after a self-removal crash. Guard it with the same check so
the crash path matches the completion path instead of probing a deleted record.

Docs: state that the record and its output directory are both gone.

Review finding: self-removed run re-creates the rmtree'd output dir (orphan leak); crash path marks a missing record.
2026-09-15 03:42:40 -07:00
teknium1
acb2c45e35 fix(cron): self-removal excuses only a missing record; replacement records stay fail-closed
Follow-up to the salvaged #111044 commits:

- self_removal_delivery_allowed() now also requires that no record currently
  holds the job id. The marker alone said "this run removed its record"; it did
  not say the id is still empty. A replacement record (another owner reclaiming
  the id) must be treated as a stolen claim, not a self-removal.
- Drop the allow_self_removed kwarg on fire_claim_fence: the fence already has
  the job_id and the ContextVar marker, so it can decide on its own; the caller
  no longer threads a flag it computed from the same predicate.
- _FireOwnership.lost(): keep the explicit lost event and the no-owner short
  circuit ahead of the self-removal check so an interrupted run is still
  reported as lost even after it removed its record.
- _finish_completed_run: skip mark_job_run entirely for a self-removed job
  (nothing to mark) instead of calling it and then excusing the False.
- Tests trimmed to two invariants, both A/B'd against origin/main: the
  self-removing run delivers after a post-removal heartbeat tick (RED on main),
  and a self-removal followed by a replacement record is still discarded
  (GREEN on main, guards the new predicate).
- Docs: user-guide cron.md notes that a job may remove itself and still report.
2026-09-15 03:42:40 -07:00
KoNit-K
8b3059fc6e fix(cron): keep self-removed runs alive across post-removal heartbeats
A run that deletes its own job after the first heartbeat interval was still marked stale because the fire-claim loop treated a missing record as lost ownership before the self-removal marker could win.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-15 03:42:40 -07:00
KoNit-K
31b4001167 fix(cron): deliver completed self-removing runs 2026-09-15 03:42:40 -07:00
teknium1
931387dd42 test(cron): two invariant ESTOP tests for the fire webhook and misfire backstop; docs
Trim the salvaged suite from four tests to the two invariants that were red on
main: (1) with the sentinel engaged POST /api/cron/fire answers 503 +
Retry-After 60 and never calls claim_fire, and the same job is admitted (202,
claimed, fired) once the sentinel is removed; (2) fire_overdue_jobs dispatches
nothing and leaves next_run_at untouched while engaged, and the first sweep
after resume catches the job up through claim_fire. The webhook test lives
beside the other cron-fire webhook tests (test_cron_fire_webhook.py) and uses
their real spy provider instead of a MagicMock resolver; the "verifier crashes
-> 401" case was already covered there.

Docs: cron.md gains a "Pausing everything: hermes pause" section stating that
all three automated doors honour pause, that in-flight runs are never killed,
and that manual runs are an operator override; the CLI reference table lists
hermes pause / hermes resume.
2026-09-15 03:33:13 -07:00
JulianCruzet
05e13a1fcc fix(cron): address Xipong's review nits on #110563
Two inline test-hardening suggestions + two follow-up coverage cases
(Xipong called them non-blocking but they're cheap and prove the
boundaries).

- tests/gateway/test_api_server_jobs.py::TestCronFireEstop::
  - test_fire_webhook_returns_503_when_estop_engaged: also patch
    `cron.scheduler_provider.resolve_cron_scheduler` and assert that
    `claim_fire`, `fire_claimed`, and `fire_due` are never reached.
    A 503 by itself does not prove admission never happened.
  - test_fire_webhook_401_when_verifier_crashes (new): crashing
    verifier → 401, ESTOP is never consulted. Proves auth runs before
    the ESTOP check so the sentinel state cannot leak to unauth callers.

- tests/cron/test_misfire_catchup.py::TestFireOverdueJobs::
  - test_estop_engaged_skips_backstop: spy on `provider.claim_fire`
    and `cron.scheduler_provider.threading.Thread`. Deterministic
    proof that no claim was attempted and no worker thread was
    constructed (the prior `wait_fired(timeout=0.5)` was timing-based).
  - test_estop_release_restores_backstop: spy on `provider.claim_fire`
    on the recovery path with an explicit `assert_called_once` so a
    regression that fires without claiming still fails the test.

All 58 tests across the affected files pass.
2026-09-15 03:33:13 -07:00
JulianCruzet
95fd5d6816 fix(cron): honor ESTOP in NAS fire webhook and misfire backstop
`agent/estop.py:1-9` documents that while the sentinel exists "the cron
scheduler ... skips work." The built-in ticker honors this
(`cron/scheduler.py:3749-3754`), but the managed-cron paths do not:

- Door 2 (NAS fire webhook `_handle_cron_fire`,
  `gateway/platforms/api_server.py:3480-3557`): no ESTOP check between
  the JWT/drain guards and `provider.claim_fire`. Added a
  `check_paused("cron-webhook")` guard inside the reservation block,
  returning 503 + Retry-After so the NAS retries after `hermes resume`
  rather than silently dropping the run.

- Door 3 (misfire backstop `fire_overdue_jobs`,
  `cron/scheduler_provider.py:253-344`): no ESTOP check at the top of
  the function. Added an early-return `check_paused("cron-misfire")`
  guard. Self-healing — the next sweep after `hermes resume` catches
  everything up via the existing claim_fire path.

Both guards use `suppress(ImportError)` matching the ticker idiom so a
broken estop module fails open rather than killing cron. Distinct
component names keep the existing log-once mechanism independent per
surface.

Manual runs (`hermes cron run`, dashboard Trigger) deliberately
unchanged: operator override is arguably a feature, and PR #105144
already rewrites that path.

Tests (3 new, 49 pre-existing in affected files all pass):
- tests/cron/test_misfire_catchup.py::test_estop_engaged_skips_backstop
- tests/cron/test_misfire_catchup.py::test_estop_release_restores_backstop
- tests/gateway/test_api_server_jobs.py::test_fire_webhook_returns_503_when_estop_engaged
2026-09-15 03:33:13 -07:00
kshitijk4poor
cedf4a3d78 fix(secret-scope): compose the managed .env into every profile secret scope
Answers the P1 review on #111187: build_profile_secret_scope() held only
<profile>/.env plus that profile's external-source snapshot, never the
administrator-managed .env. The launch process applies that file LAST with
override (_apply_managed_env), so a managed key beats the user's own value in
os.environ. Under multiplex semantics get_secret() stops falling back to
os.environ on a scope miss, so inside a routed cron fire (and equally inside a
real multiplex gateway turn, which builds its scope through the same function
via gateway/run.py::_load_profile_secret_scope) a managed-only credential
resolved as absent and a managed-vs-user collision resolved to the USER value:
reversed precedence.

Fix at the source: build_profile_secret_scope() overlays load_managed_env()
last, after the profile .env and external sources, skipping process-global
names exactly as it does for the other two layers. Every multiplex-authoritative
scope (gateway turn, routed desktop fire, external worker env build) is built
here, so managed authority is composed once instead of restored per consumer.
No generic ambient-env fallback is reintroduced: only the managed file's own
keys enter the scope, and only with the managed file's values.

Regression (parametrized, two invariants): inside a routed fire a managed-only
key resolves through get_secret(); a managed-vs-user collision yields the
managed value.
2026-09-15 11:03:39 +05:30
kshitijk4poor
5850a50a81 test(cron): restore the routed-fire handoff test for _launch_external_cron_worker
The earlier trim dropped the only test of the `routed_profile_fire() and
not is_multiplex_active()` branch in the external-worker handoff: the
process flag is OFF (a desktop tick is not a multiplexer) yet a fire routed
to a sibling profile must serialize `multiplex_active=True` and hand the
worker an env without the launch profile's residue, and the context must
not outlive the handoff span. Without it that branch could regress to the
pre-fix behaviour unnoticed.
2026-09-15 11:03:39 +05:30
kshitijk4poor
042b9830b4 test(cron): keep three invariants for the routed-fire isolation
The PR shipped nineteen tests across six files, most of them variations of
one boundary. Keep the three that pin distinct behaviour:

- a routed desktop-ticker fire runs under multiplex semantics for exactly
  its scope: a scope miss returns None instead of the launch credential and
  the parent os.environ is byte-identical afterwards;
- a routed no_agent child never sees a launch-only name, whether the launch
  .env defined it or a launch external source supplied it (applied or lost
  to a pre-existing process value), while its own values come through;
- administrator-managed keys keep policy precedence over the routed
  profile's own value.

Everything else was either a positive control of the same seam, a
set-membership check on a module-level constant, or a re-statement through
a different entry point.
2026-09-15 11:03:39 +05:30
John Paul Soliva
62b4488cb5 fix(cron): routed fires are multiplexed at the worker handoff; managed keys keep policy precedence
Review findings on f5f88d5058. Three are defects the previous round introduced.

Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.

Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.

Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.

Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.

Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.

(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
2026-09-15 11:03:39 +05:30
John Paul Soliva
9d7de6c140 fix(cron): close three launch-residue leaks into a routed no_agent child
Review findings on d8c467f223, each reproduced through its production path.

Stale launch key. `strip_launch_profile_env` built its residue set from a
re-parse of the launch `.env`. A key removed or renamed in that file after
boot is still in `os.environ` with the old value (dotenv never unsets), and
the current file no longer names it, so it survived into the routed child.
`_load_dotenv_with_fallback` — the one chokepoint every dotenv load goes
through — now records the KEY names it put into the process env, additive for
the process lifetime (`launch_dotenv_keys()`), and the strip unions that record
with the current file.

Source name that lost to the process env. `_apply_external_secret_sources`
snapshots every name a source SUPPLIED (`provenance` + `skipped_existing`),
but `secret_source_names()` only exposed `_SECRET_SOURCES`, which is
provenance metadata and names applied values alone. A launch-profile source
that supplied `CUSTOM_VAULT_SECRET` while the process already had it was
therefore invisible to the scrub, and a routed child with an empty scope got
the launch value. Supplied names are tracked separately
(`_SOURCE_SUPPLIED_NAMES`) so the provenance labels stay honest, and
`secret_source_names()` returns the union.

Last plugin source removed. `_refresh_secret_sources_after_discovery`
returned before the cache reset and the installed-scope refresh whenever no
plugin source was enabled — and `discover_and_load(force=True)` unloads the
old registration first, so removing the final plugin source hit exactly that
return with the removed plugin's names still in the per-home snapshot and the
current scope. The manager now remembers that a discovery re-applied plugin
sources and, on the next discovery that finds none, reconciles once. A home
that never had a plugin source is still a no-op (pinned by the existing tests).

Regressions: the stale-key lifecycle and the skipped-existing case through
`_run_job_script` against a real routed child, and the removal case through
the manager. Each checked by reverting its fix and confirming the test fails.

(cherry picked from commit d464f5f6126a394cfb47937f683d3a5e2f141840)
2026-09-15 11:03:39 +05:30
John Paul Soliva
3dedff6a12 fix(cron): only strip launch external-source names when multiplexing
The source-name strip added alongside the external-source fix ran
unconditionally. Outside multiplexing there is no other profile to leak from --
os.environ IS this profile's environment -- so popping those names relied on the
routed scope overlay putting each one back, which in turn relies on the
per-home snapshot recorded at boot. Correct today, but it made a single-profile
child's credentials depend on bookkeeping that has nothing to do with isolation.

Guard it the way strip_launch_profile_env guards itself: no multiplexing, no
strip. A single-profile no_agent child now keeps a byte-identical env even if a
source's snapshot were ever missing. Pinned by a regression that runs a real
child with a source-owned name in os.environ and no multiplex context; making
the strip unconditional fails it.

(cherry picked from commit afa429b30a1c9c9b4c011f7d2426a9f099911596)
2026-09-15 11:03:39 +05:30
John Paul Soliva
0943e77136 fix(cron): strip launch external-source names too, and make the scope refresh replace
Two credential-isolation gaps found in review of the previous head.

1. strip_launch_profile_env() only knows dotenv- and terminal-config-owned names,
but external secret sources (vault, 1Password, ...) also write their names into
the shared os.environ and are tracked in secret_source_names(). A name the LAUNCH
profile's source supplied therefore still reached a routed no_agent child. Drop
every non-global source-owned name from the base; the routed scope overlay that
follows puts back exactly the ones that profile's OWN sources supply, since
build_profile_secret_scope folds get_secret_source_values(home) in.

2. refresh_installed_secret_scope() merged the rebuild with dict.update(), so a
name a source had stopped supplying -- rotated, revoked, source removed -- kept
its old value for the rest of the fire. Replace the mapping contents instead: the
rebuild is the profile's current truth.

Regressions: a routed child sees <unset> for a launch-source name while its own
source value comes through, and a refresh whose rebuild omits a name drops it.
Both fail if the corresponding change is reverted.

(cherry picked from commit ecd51517c4828a75acb5f458ed99aea0bc3e5e9f)
2026-09-15 11:03:39 +05:30
John Paul Soliva
023e4f997f fix(cron): drop the launch profile's dotenv residue before overlaying a routed no_agent scope
The routed no_agent child env started from all of os.environ and only overwrote the
names present in the installed scope. A name defined only by the LAUNCH profile's .env
and absent from the routed profile therefore reached the routed child with the launch
value instead of unset -- the secret scrub only knows classified names, so a custom or
unclassified secret crossed the profile boundary (review finding on the first head).

strip_launch_profile_env (main, 284d220ba4) is the primitive the external-worker path
already uses for exactly this: it drops the launch profile's dotenv-owned keys and the
bridged TERMINAL_* settings, and is a no-op outside multiplex or when the target IS the
launch profile. Apply it to the base BEFORE the scope overlay (so a shared name keeps its
routed value) and BEFORE the sanitizer (so routed values still pass the scrub and
passthrough rules). Pinned by a child-process negative control: the launch-only name
arrives <unset>, the shared name arrives routed, and the parent process is unchanged.

(cherry picked from commit 69349b527b5144c1f1540a4c10035d5ec798c1db)
2026-09-15 11:03:39 +05:30
John Paul Soliva
dbede34f6e fix(cron): a routed profile's cron fire in the desktop backend runs under multiplex semantics
The desktop backend ticks EVERY local profile's cron store from one process — its own docstring
says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global
multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every
isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared
`os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban
subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary
profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed
there after the tick, and a scope miss read the launch profile's tokens (#107692).

Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py)
is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not
the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the
override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the
profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that
span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are
therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs
before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam,
left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation
applies inside the routed fire with no per-site patching; the launch profile's own fires and the
backend's turns keep single-profile semantics; marker and override both reach the pool worker via
`copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through
`is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970).

Two consequences of suppressing the write are handled rather than left as regressions:
- a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`;
  the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub /
  passthrough rules apply to those values and the parent process is never mutated;
- plugin secret sources are discovered on the fire's first agent build, after the scope froze,
  and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds
  the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern
  `_publish_env_value` already uses for `.env` writes under multiplex).
And the profile's external secret sources are hydrated before the scope is frozen, the order
gateway/run.py and the external cron worker already use.

Tests pin each direction: the marker without the semantics before the scope, the semantics on and
off exactly with it, the marker reaching a copy_context worker; the process's own profile staying
single-profile; the restart-safe handoff's child env building without raising under a routed tick
with a passthrough key registered; a real child process receiving the routed values while
`os.environ` keeps the launch value; a source registered after the freeze reaching the fire
through the real PluginManager refresh. Reverting any one direction fails a distinct test.

(cherry picked from commit 2f87677425d2cca19286ac83bc45cab23e546669)
2026-09-15 11:03:39 +05:30
teknium1
8f7853188f fix(cron): keep deferred delivery exceptions from aborting ticks
Catch unexpected delivery exceptions after claim, retain diagnostics and continue
sibling admissions without authorizing replay. Preserve indefinite retention.

Reproduced PermissionError at target traversal after discovery. Native Electron
controlled-fault A/B confirms the healthy sibling settles and renders once.
2026-09-14 17:29:32 -07:00
teknium1
002ee41cfc fix(cron): keep unowned Bot Chat delivery on its resolved home
Extend deferred dispatch's destination pin to ordinary CLI fallback, so
custom-root and active-profile changes cannot redirect a checked target.
Refuse a missing destination before launch and name the target on failure.
Replace the old env-clearing expectation with two behavioral invariants
and retain the native Electron custom-root reproduction.

Adapted from the root-boundary fix and diagnosis in #104066.
Related #104055, #104066.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-14 17:29:32 -07:00
teknium1
3b0fe0cc2b fix(cron): keep deferred Bot Chat delivery bound to admission
Carry the original destination home and delivery ID into deferred drain and
its child, rather than re-resolving a mutable profile/root. Missing destinations
fail closed; supported-owner handoffs remain transferred, not ambiguous failures.
Capture the producer root before the background thread starts, and retain/log
malformed JSON without stopping healthy admissions or the whole cron tick.

Two invariants reproduced failures on the published head. Real Electron root
change and malformed-record cases are red before and green after; nested DM
control remains passing. No automatic retry of claimed or uncertain turns.
2026-09-14 17:29:32 -07:00
teknium1
5d8390d1a4 fix(cron): retain Bot Chat output while a CLI owner is open
Keep never-started output behind unsupported owners and drain in admission
order after release. Persist claims before execution and never replay uncertain
started turns. Existing supported-owner receipts keep their authority.

Credits 686f6c61's residual queue proposal in #100319. This is a scoped
implementation, not general retry of failed CLI subprocesses.

Native Electron before/after: CLI-owned target previously returned
SESSION_NOT_OWNED and remained empty after release/tick; now its queued
output and reply appear once in the target Bot Chat. Nested quiet CLI
message_agent delivery to a named Desktop owner also passes on base.
2026-09-14 17:29:32 -07:00
teknium1
45b42202fa fix(cron): a raising profile gate ticks nothing; housekeeping respawns a dead ticker (#111010)
Follow-up on the salvaged #111034:

- `_start_multiplex` published the enumerated home list before the gate had
  filtered it, so a raising `profile_gate` (the Desktop stand-down probe from
  #100489) kept the thread alive but ticked every profile UNGATED — racing the
  gateway that owns them for the same cron store. The list is now assigned only
  after gating; a gate failure yields zero ticks for that cycle.
- `cron/scheduler_thread.py::SupervisedTickerThread` wraps the gateway ticker
  thread; `_start_gateway_housekeeping` gets a per-tick "Cron ticker supervisor"
  chore that respawns a ticker that ended without a stop request and logs the
  outage at ERROR. Every guard inside `start()` keeps the loop alive, but nothing
  outside it could notice a thread that had already ended.
- Tests trimmed to the invariants, proven red on origin/main: a REAL corrupt
  `executions.db` (no patched recover) no longer kills the ticker; a raising gate
  keeps the thread alive with zero ticks; housekeeping restarts a dead ticker and
  leaves a stopped one alone.

Root cause: the unguarded pre-loop `recover_interrupted_executions()` +
`record_ticker_heartbeat()` were added by d9dd05b69d (#61791, "truthful
execution ledger", 2026-07-09). The reporter's build (e440bf35) also carried
#107485's `completed_occurrence()` in the due scan, which opens the same ledger
on every tick — the first traceback in their errors.log is that in-loop hit
(caught); the restart then hit the SAME corrupt ledger from the pre-loop
recovery scan, which nothing caught: thread dead, no heartbeat, no error marker.
2026-09-14 16:14:33 -07:00
Konstantin Khlopkov
8fce4baf8c fix(cron): keep the ticker thread alive through startup recovery and marker writes (#111010)
The gateway runs InProcessCronScheduler on an unsupervised daemon thread, so any exception escaping outside the guarded tick body ends cron silently while the gateway keeps serving. Guard the three remaining escape windows: startup recovery + initial heartbeat in start(), per-cycle profile enumeration/gating in _start_multiplex, and every status-marker store write (heartbeat/error/clear) in both loops. A failing store now degrades to a logged warning and a failed cycle instead of thread death.
2026-09-14 16:14:33 -07:00
Jony
45aebc11b0 fix(cron): preserve cleanup profile scope 2026-09-14 16:13:51 -07:00
teknium1
498abb677e fix(cron): a successful run resolves the job's open incidents; a repeat re-opens them
The incident ledger only ever grew. A one-off failure (a drift skip after a
global model bump, a provider outage) stayed `detected`/`alerted` forever
after the job recovered, so `hermes cron incidents` listed 32 "open"
incidents on an install where all 32 jobs had since run OK, and the list
stopped saying anything about current health.

A successful run now marks that job's `detected`/`alerted` incidents
`resolved` (new state). `resolved` is distinct from the operator's `closed`
ack on purpose: `upsert_incident` re-opens a resolved incident as
`detected` when the same error signature recurs, so the operator is alerted
again for a job that broke a second time, while `closed` keeps the
signature silent as before. Wired from `_compose_run_delivery` next to the
failure-side upsert; best-effort, store errors never affect delivery.

CLI: `--state resolved` filter and a green `resolved` colour; `closed` is
now dim. Docs updated in the same change.
2026-09-14 09:19:32 -07:00
teknium1
8a5a66d0d8 test: trim the Aug-2026 conformance suites to behaviour invariants on main
Rebased onto main (~9700 commits): the branch predates the god-file
decomposition and the tests/state -> tests/hermes_state move.

Fixes for seams that moved:
- cron memory contract patches cron.scheduler_delivery._resolve_origin and
  hermes_state_registry.acquire (where run_job now reads them).
- state.db conformance imports _live_writer_holds_db from
  hermes_state_repair and drives the public create_quick_snapshot.
- update receipt: serve runtimes are reconciled in their own unit
  vocabulary since #100479, so the "full accounting" row names the serve
  unit instead of borrowing a gateway relaunch.

Deleted (change-detectors / source-greps / duplicates / dead code):
- source-text scan of hermes_state*.py for "maintenance-shaped" defs and
  the symbol REGISTRY it fed (renames are not regressions).
- per-op exact-outcome table under a live writer -> one invariant: refuse
  with a lock error or report zero work, DB stays intact.
- copy_db_and_verify pins: the symbol is a revert-scheduled plugin-compat
  pointer, not production code (check_compat_pointers.py).
- cron ON-direction tests already pinned by tests/cron/test_scheduler.py,
  plus a tautology that never called production code.
- exact warning-wording asserts in the env deprecation truth table.
- same-file duplicates: fixed-point (implied by idempotence), tripwire
  round-trips, absent-registry, sequential "race" re-enactment.
2026-09-13 21:13:23 -07:00
Teknium
d15f4a7826 test: regression conformance suites for the 7 recurring Aug-2026 bug classes
Two weeks of closed issues/merged PRs show the same areas regenerating:
each salvage pinned its instance while the class invariant had no test.
These suites pin the invariants themselves:

- tests/conformance/test_profile_write_tripwire.py — no writes to the
  default profile tree while a profile is active (#88532 #92662 #89190
  #89625 #92156); reusable tripwire fixture, 4 surfaces
- tests/hermes_cli/test_env_deprecation_truthtable.py — 18-row truth
  table for the Deprecated-.env warning (#88829 #89016 #89389 #90299)
- tests/cron/test_cron_memory_contract.py — cron<->memory contract that
  flipped twice in Aug (#91269 -> #91384 -> #91447)
- tests/agent/test_injected_param_strip_retry_registry.py — every
  strippable injected param x real 400 shapes must strip-and-retry;
  unknown params must still fail (#90257 #89897 #91164 #89503)
- tests/agent/test_transcript_decoration_idempotence.py — f(f(x))==f(x)
  law + 4-breakpoint budget for apply_anthropic_cache_control (#90971)
- tests/state/test_state_db_maintenance_conformance.py — registry-
  enumerated maintenance ops refuse/degrade under a live writer; copies
  of corrupt DBs are refused or flagged (#91839 #90806 #90613 #88235)
- tests/tools/test_bot_mode_canonical_chat_resolution.py — canonical
  Bot Chat resolution is idempotent, never mints, unique per profile,
  race-safe (#92040 #90705 #92692 #90005 #90732, PR #92129)
- tests/hermes_cli/test_update_receipt_truthfulness.py — receipts:
  crash never claims success; success requires full fleet accounting;
  refusal != failure (#91283 #91439 #92902 #92780)

117 tests, all sabotage-verified (each suite proven to FAIL when its
bug class is reintroduced).
2026-09-13 21:13:23 -07:00
Teknium
0877decd1b Port from paradigmxyz/centaur#1479: cron runs no longer re-schedule themselves from recurring prompt language
A scheduled job whose prompt carries its own cadence phrasing ('Each
Monday, review...') can convince the agent to create ANOTHER cron job
at execution time instead of just doing the work — each run spawning a
sibling job. Hermes policy-denies the cronjob toolset in cron context
by default, but cron.allow_agent_scheduling: true re-enables it and
opens exactly this loop.

_build_job_prompt now extends the always-injected cron hint with a
RECURSION clause: this is a run of an existing job; never create or
update a cron job from schedule language in the task prompt; treat
cadence phrasing as context for this run.

Adapted from paradigmxyz/centaur#1479 (same failure mode in their
scheduled-task workflow runner).
2026-09-13 20:52:58 -07:00
teknium1
95c7e9a0cb fix(cron): return a strictly-later next run when the base sits in the DST fall-back hour
Port from QwenLM/qwen-code#11723: attaching the configured zone to croniter's
naive wall-clock result resolves the repeated autumn hour to its earlier
occurrence (fold=0), so a base inside the second occurrence received a
next_run_at up to an hour in the past — the fire path would treat it as due
immediately and loop. Try both folds of each candidate and return the earliest
instant strictly after the base; wall-clock jobs still fire exactly once on
the repeated hour (Vixie cron semantics).
2026-09-13 16:46:45 -07:00