Commit Graph

743 Commits

Author SHA1 Message Date
teknium1
f336048b08 fix(cron): bot-chat delivery launches the running install, not whatever hermes PATH names
The cron scheduler runs inside the long-lived gateway process and spawned
`hermes ... chat` for Bot Chat delivery through `shutil.which("hermes")`
first, falling back to `sys.executable -m hermes_cli.main` only when PATH
had no `hermes`. That is the same resolution order gateway.run.
_resolve_hermes_bin just flipped for /update and /restart (#111569): the
running interpreter's module argv is exactly this install, PATH is not.

Delivery now resolves the running install first and uses PATH only as the
fallback. Tests that asserted the PATH argv shape or armed on the `which`
seam are moved to the module-argv shape / the `find_spec` seam.
2026-09-15 19:03:08 -07:00
teknium1
9c0633b631 docs(cron): name idle-vs-wall-clock and the separate script timeout
The corrected invariant still read as if "inactive loops" were the thing
being bounded. Say what the watchdog measures (idle time, not wall-clock),
what that means for operators sizing work (an active job is never cut off),
and that attached scripts have their own bound (_DEFAULT_SCRIPT_TIMEOUT),
which is the second misreading #111633 calls out.
2026-09-15 18:46:54 -07:00
KoNit-K
6c9a665de2 docs(cron): clarify inactivity watchdog 2026-09-15 18:46:54 -07:00
teknium1
9013fcdc87 fix(cron): an unreadable cron toolset restriction fails the run instead of granting every tool
_resolve_cron_enabled_toolsets returned None when _get_platform_tools
raised, and AIAgent reads None as "load every toolset": a malformed
platform_toolsets block (or a stale-module import error after an update)
turned the operator's cron restriction into the full default set, with
only a log warning. Unattended jobs process untrusted text, so that is a
privilege widening, not a safety net (#111380).

The resolver now raises a RuntimeError naming the cause; run_job's
existing failure path records it on the job (last_error, failure streak,
incident) and the agent is never constructed. Per-job enabled_toolsets
(unknown names included) and the MCP merge path never touch the
platform resolver and are unchanged; the disabled-toolset resolver has no
fail-open branch.

Live: platform_toolsets: oops -> before: run ok, enabled_toolsets=None,
98 tool names selected; after: run fails "Cron toolset resolution
failed, so this run was refused rather than given every tool", agent
never constructed. normal / unknown-per-job / mcp-merge shapes: identical
before and after.

Co-authored-by: Austin Bell <10687162+robertaustinbell@users.noreply.github.com>
2026-09-15 18:31:55 -07:00
teknium1
f7d2bb95e1 fix(cron): a due slot skipped as already completed is logged, not silent
The completed-occurrence dedup gate (due scan in _evaluate_due_job and the
fire claim in claim_job_for_fire) consumes the due slot and advances
next_run_at without a run and without a ledger row, so before this change
a skip left zero trace: no log line, no execution, last_status untouched
(#111414 reported exactly that silhouette). The gate now logs a WARNING
naming the job, the skipped instant and the completed execution that
already covers it, from the one seam both gates share.

The reporter's root mechanism — an off-tick run stamping the NEXT
occurrence's identity onto its row, which the dedup later honoured — is
already closed on main (ac10770894, a73b750391, cd685a22e6; 82ae78cc11
ignores such backdated rows), so a stale-stamped row fires normally; only
a genuinely completed slot is skipped, and now says so.

Co-authored-by: holny <holny@foxmail.com>
2026-09-15 18:31:26 -07:00
teknium1
719a420ec6 refactor(cron): pass monitor context as a plain kwarg; trim salvage tests to two
Follow-up to the salvaged #111528 hunk: _prepare_job_prompt passed the monitor
block through a conditional kwargs dict although _build_job_prompt already
treats a falsy runtime_data_prompt as absent. Drop the duplicate unit-level
positive test — the run_job-level test in test_monitor_kind.py covers the
same invariant through the real path — and keep the strict-user-prompt
control.
2026-09-15 18:30:59 -07:00
KoNit-K
e7cb478db6 fix(cron): sanitize monitor runtime prompt data 2026-09-15 18:30:59 -07:00
teknium1
3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
teknium1
14ebd48c64 fix: derive the lost-bus check from the env the scoped worker was spawned with
scoped_spawn_lost_user_bus() re-derived the user bus from an empty base env,
so it only ever looked under /run/user/<uid>. The worker itself is launched
with systemd_user_bus_env(worker_env), which honours a configured
XDG_RUNTIME_DIR. On a host whose bus lives outside the default runtime dir,
any unrelated systemd-run exit was therefore misread as "bus gone": the job
error named a missing bus that was still there, the cached scope verdict
flipped to False, and the next 60s of cron fires dispatched without cgroup
isolation.

The check now takes the spawn env, drops the bus address the spawn already
carried, re-derives from that, and decides on the DBUS_SESSION_BUS_ADDRESS
key rather than on dict truthiness (a non-empty env with no bus was never
"bus present").

Review finding: scoped_spawn_lost_user_bus used systemd_user_bus_env({}) instead of the worker's spawn env, so a configured XDG_RUNTIME_DIR made unrelated wrapper exits flip the scope cache to unscoped dispatch.
2026-09-15 06:28:07 -07:00
teknium1
a313a211d7 fix(cron): a scoped worker whose user bus vanished names the cause and re-probes
One TTL for both probe verdicts (the success-TTL constant collapses into
`_SYSTEMD_SCOPE_PROBE_TTL_SECONDS`), and the ack-wait loop consults
`scoped_spawn_lost_user_bus()` when a scoped dispatch exits before the
worker acknowledged: with `/run/user/<uid>/bus` gone the job error names
the missing bus and the enable-linger remedy instead of the wrapper's bare
`exit 1`, and the cached True flips so the next fire degrades to a direct
external subprocess rather than consuming another occurrence on a dead
wrapper. Contributor test trimmed to the revalidation invariant.
2026-09-15 06:28:07 -07:00
teknium1
fcbdcdb428 fix(cron): let a skewed early fire own its armed slot
The future-instant guard in claim_job_for_fire dropped the occurrence
identity for ANY claim ahead of the stored next_run_at. A hosted/webhook
fire for the armed slot that arrives a few seconds early (the fire
scheduler's clock runs ahead of ours) was therefore treated as an
off-tick run: it ran occurrence-free, mark_job_run recomputed the same
cron slot from a now still before it, and the tick/misfire backstop then
ran the slot a second time.

Only claims at least FIRE_CLAIM_SKEW_SECONDS (60 s) ahead of the slot are
now classified off-tick, so dashboard/manual far-future fires stay
occurrence-free while a skewed early fire keeps the slot identity.
completed_occurrence honours the same window so the early run's
completion row (finished just before the slot) still proves the slot
done instead of being discarded as poison.

Review finding: claim_job_for_fire future-instant guard had no skew tolerance; an early hosted fire for the armed slot ran twice.
2026-09-15 06:07:32 -07:00
fangliquan
82ae78cc11 fix(cron): reject impossible occurrence completions 2026-09-15 06:07:32 -07:00
Chris Herrera
cd685a22e6 fix(cron): manual fires before the next occurrence must not consume it
Direction (ii) of the t_a662ca7a fix menu, ported from the live install
(ffffab1bd6 on the deployed checkout) and layered on top of the already-
landed direction (i) (ac10770894, manual= flag): claim_job_for_fire() now
also declines to bind an occurrence whose instant is still in the future.

A scheduled tick only fires when now >= next_run_at (_evaluate_due_job
returns False while the stored occurrence is in the future), so a claim
arriving BEFORE the stored next occurrence cannot be the tick that owns
it: it is a manual / dashboard / webhook fire (including the
claim_ttl_seconds reclaim after an expired lease, which arrives with no
manual flag) and stays occurrence-free - ledger row scheduled_instant=NULL,
next_run_at untouched. On-time and late (catch-up) ticks keep the exact
at-most-once occurrence identity.

Live case 2026-09-08: job d88fde170fb4 [bot:carl] Daily Plan Evening
(0 19 * * *), manual run at 19:53 stamped scheduled_instant=
2026-09-10T02:00:00+00:00 (= next day 19:00 PDT) status='completed';
completed_occurrence() then refused every later manual fire ("Job is
already being fired by the scheduler; not run again.") and the next
scheduled tick dedupe-skipped its real delivery. Pre-fix poison rows in
existing ledgers are not repaired by this commit; repair remains
audit-preserving UPDATE-only via the retired reference script.

Regression test (tests/cron/test_manual_fire_occurrence.py): two
invariant tests proven red on base b88e6776f4 (mid-cycle manual fire and
webhook claim_fire bind the future instant) plus three green-on-base
guard nets (on-time tick binding, late catch-up binding, forced-fire
occurrence-free). Red on base: 2 failed / 3 passed; green post-fix:
5 passed. tests/cron suite: 1232 passed, 1 skipped, 1 pre-existing
environmental failure in
test_run_job_cron_execute_code_deny_does_not_pollute_later_gateway_execute_code
that fails identically on base.

Kanban: t_d36f3b54 / t_2a050d21 (defect t_a662ca7a)
2026-09-15 06:07:32 -07:00
teknium1
71d229f5cd fix(cron): the silence instruction names [SILENT] as an untranslatable ASCII token; docs + trim tests
Follow-up to the salvaged #110940 commit:

- `cron/scheduler_prompt.py::_CRON_HINT` tells the model the sentinel is a literal
  ASCII control token that must never be translated or rephrased — the prompt-side
  half of the fix, so a lane answering in any other language is steered to the
  canonical token instead of relying on the filter knowing that language.
- Docs: the supported-token list ('Intentional Silence Tokens') gains the zh-Hans
  forms in the English page and the zh-Hans mirror gets the section it lacked.
- Tests trimmed to two invariants (translated forms match in every shape the English
  ones do; prose that mentions the word is still delivered), proven red on origin/main.
2026-09-15 04:30:31 -07:00
teknium1
248938b4e6 fix(cron): "Script not found" says scripts are per-profile and how to fix it
Cron scripts resolve only against the job's own profile scripts/ dir (by design,
profiles never share files), so a job copied between profiles fails with a path
that looks plausible and no hint about why. The runtime error now names the
profile folder and the two fixes (copy the script, or `hermes cron edit`).

Closes #94821. Builds on #105775 (creation-time existence check, credited).
2026-09-15 04:15:16 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1
56e563d3a4 fix: keep reading scripts named inside masked heredoc bodies
Masking inert heredoc bodies before the referenced-script walk (previous
commit) also hid every path the body names, so a Python body that does
os.system('/x/restart.sh'), a bare '/x/restart.sh' line, or a nested cat
heredoc naming the script all went undetected — main caught them because
the walk read the script and found the lifecycle command inside.

The interpreter does execute those paths at runtime, so the walk must still
read them. Only the fail-closed verdicts (cloud placeholder, oversized or
binary file) stay restricted to the masked view: a >1 MiB data file merely
mentioned inside an inert body is not a script, which is the false positive
the previous commit fixed and its tests keep pinning.

Review finding: masking the heredoc body from the walk widened fail-open —
scripts executed by path from within the body were no longer scanned.
2026-09-15 04:00:34 -07:00
teknium1
1fd0ef7520 test(cron): trim the heredoc-walk regression to its two invariants
Keep the inert-body path (false) and the unquoted-body path (still true);
the `sh -c`-in-`cat` heredoc and plain script-reference cases are already
pinned by the existing lifecycle-guard suites. Comment shortened to the WHY.
2026-09-15 04:00:34 -07:00
Kevin Rajan
59a1403fa2 fix(cron): mask inert heredoc bodies before the referenced-script walk
The direct lifecycle scan masks provably-inert heredoc bodies
(strip_inert_heredoc_bodies), but the referenced-script and -c payload
walks ran on the unmasked command. A path inside such a body is never
shell-executed, so walking it was a pure false positive: a >1 MiB path
mentioned in a quoted heredoc body (e.g. python3 - <<'PY') failed closed
and hard-blocked an innocent command.

The walks now use the same masked view as the direct scan. Masking only
fires under the stripper's conservative contract (quoted delimiters,
exact terminator, single simple command, allowlisted consumer, no command
substitution); unquoted/shell-consumed/ambiguous bodies stay visible and
fail-closed.

Fixes #110422.

authored with AI assistance (Muse, Meta's Muse Spark) under the contributor's direction
2026-09-15 04:00:34 -07:00
Hukla
43891e9f8e fix(cron): inject skill config into scheduled runs 2026-09-15 03:58:27 -07:00
teknium1
b6b7802447 fix(approval): every unattended context clears leaked presence vars
Widen the cron-only clearing to `_unattended_contexts()`: a webhook /
api_server session running inside a gateway inherits HERMES_EXEC_ASK=1
exactly like an external cron worker does, and `_presence()` returning
is_ask=True sent it to the gateway-decision branch with no notifier — a
pending card nobody can answer — instead of `approvals.unattended_mode`.
Same class as #110932, one predicate.

Test trimmed to two invariants (cron / webhook leak → cleared; interactive
keeps presence); the launch-path comment in cron/scheduler.py names the
env-fallback consumers instead of an internal incident log.
2026-09-15 03:56:17 -07:00
fabiantax
2db47cc9fb fix(cron): strip interactive presence vars from external worker env
Gateway sets HERMES_EXEC_ASK=1 (interactive launches set HERMES_INTERACTIVE / HERMES_GATEWAY_SESSION) at runtime; systemd-run cron workers inherited them, _is_interactive_cli() then bypassed approvals.cron_mode for terminal and every run hung on a pending card nobody could answer (fab-swarm #105: ms197 lane left 6 claims stranded, 10-30s hangs). Local mitigation; upstream report to follow.
2026-09-15 03:56:17 -07:00
teknium1
134ef6454d fix(cron): self-removed runs leave no output directory and skip mark_job_run on crash
remove_job() deletes <cron>/output/<job_id>/ together with the record, but the
finishing run then called save_job_output(), re-creating the directory and
writing the final run into it. Every self-removing job leaked an orphan
directory the store no longer knew about, and the docs claim that "only the
job record is gone afterwards" was false. Skip the save when
self_removal_delivery_allowed() is true (the same check that already excuses
the missing record on the delivery and mark paths); delivery composes without
an output_file, as it already does for non-file paths.

The BaseException handler in _run_one_job_body still called mark_job_run on
the missing record after a self-removal crash. Guard it with the same check so
the crash path matches the completion path instead of probing a deleted record.

Docs: state that the record and its output directory are both gone.

Review finding: self-removed run re-creates the rmtree'd output dir (orphan leak); crash path marks a missing record.
2026-09-15 03:42:40 -07:00
teknium1
acb2c45e35 fix(cron): self-removal excuses only a missing record; replacement records stay fail-closed
Follow-up to the salvaged #111044 commits:

- self_removal_delivery_allowed() now also requires that no record currently
  holds the job id. The marker alone said "this run removed its record"; it did
  not say the id is still empty. A replacement record (another owner reclaiming
  the id) must be treated as a stolen claim, not a self-removal.
- Drop the allow_self_removed kwarg on fire_claim_fence: the fence already has
  the job_id and the ContextVar marker, so it can decide on its own; the caller
  no longer threads a flag it computed from the same predicate.
- _FireOwnership.lost(): keep the explicit lost event and the no-owner short
  circuit ahead of the self-removal check so an interrupted run is still
  reported as lost even after it removed its record.
- _finish_completed_run: skip mark_job_run entirely for a self-removed job
  (nothing to mark) instead of calling it and then excusing the False.
- Tests trimmed to two invariants, both A/B'd against origin/main: the
  self-removing run delivers after a post-removal heartbeat tick (RED on main),
  and a self-removal followed by a replacement record is still discarded
  (GREEN on main, guards the new predicate).
- Docs: user-guide cron.md notes that a job may remove itself and still report.
2026-09-15 03:42:40 -07:00
KoNit-K
8b3059fc6e fix(cron): keep self-removed runs alive across post-removal heartbeats
A run that deletes its own job after the first heartbeat interval was still marked stale because the fire-claim loop treated a missing record as lost ownership before the self-removal marker could win.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-15 03:42:40 -07:00
KoNit-K
31b4001167 fix(cron): deliver completed self-removing runs 2026-09-15 03:42:40 -07:00
teknium1
788601358d refactor(skills): live-dashboard becomes an optional skill reconfigured for the Desktop app
Why: the web dashboard is being deprecated in favour of the Electron Desktop
app, and a skill that ships a bespoke cron blueprint should not be bundled by
default. Reconfigure instead of just rebasing:

- Move skills/productivity/live-dashboard -> optional-skills/productivity/
  live-dashboard (install with `hermes skills install
  official/productivity/live-dashboard`); register it like every other
  optional skill: per-skill docs page under user-guide/skills/optional/,
  optional-skills-catalog row, sidebars entry.
- Desktop reality: add a "Show the dashboard" step — when `desktop_preview`
  is in the toolset (Desktop/GUI sessions) render index.html in the in-app
  preview pane after every build/tick and on request; otherwise report the
  absolute file path. Prerequisites section added (Enough1122 review).
- Drop the hard-wired cron/blueprint_catalog.py entry: the curated catalog
  is for bundled skills and would preload a skill that may not be
  installed. Use the skills-pipeline blueprint instead —
  `metadata.hermes.blueprint` on the SKILL.md registers a daily
  all-dashboards sweep as a /suggestions entry at install time (opt-in,
  never auto-scheduled), which is exactly the mechanism main provides for
  optional skills.
- Never hardcode ~/.hermes in prose the agent executes: refer to the Hermes
  home directory's dashboards/<slug>/ and write absolute paths into cron
  prompts (also answers the review's "tick prompt must name the state-file
  path" point).
- Tests follow the skill to optional-skills/, the catalog-blueprint tests
  are replaced by one parse_blueprint/blueprint_to_job_spec invariant and
  one desktop_preview-with-path-fallback invariant.
2026-09-15 03:35:33 -07:00
Teknium
2154458685 Inspired by Energy: one-sentence live dashboards — bundled skill + automation blueprint
Energy (getenergy.com) ships natural-language persistent dashboards:
describe what you want to see in one sentence and the agent builds a
self-updating status page fed by email threads, signed-in websites, and
files. This ports the concept onto Hermes's existing cron + connector
architecture:

- skills/productivity/live-dashboard: setup/tick split skill — pin the
  dashboard contract, verify one live read per source before scheduling,
  keep dashboard.json as source of truth with a self-contained HTML
  projection, stale-read discipline, deliver only on material change.
- cron/blueprint_catalog.py: live-dashboard automation blueprint
  (purpose/sources/time/recurrence/deliver slots) rendering to the
  dashboard form, /blueprint command, and hermes:// deep-link.
- tests/skills/test_live_dashboard_skill.py: skill standards + blueprint
  registration + real fill_blueprint E2E.
- docs: per-skill page, skills catalog row, sidebar entry.
2026-09-15 03:35:33 -07:00
JulianCruzet
95fd5d6816 fix(cron): honor ESTOP in NAS fire webhook and misfire backstop
`agent/estop.py:1-9` documents that while the sentinel exists "the cron
scheduler ... skips work." The built-in ticker honors this
(`cron/scheduler.py:3749-3754`), but the managed-cron paths do not:

- Door 2 (NAS fire webhook `_handle_cron_fire`,
  `gateway/platforms/api_server.py:3480-3557`): no ESTOP check between
  the JWT/drain guards and `provider.claim_fire`. Added a
  `check_paused("cron-webhook")` guard inside the reservation block,
  returning 503 + Retry-After so the NAS retries after `hermes resume`
  rather than silently dropping the run.

- Door 3 (misfire backstop `fire_overdue_jobs`,
  `cron/scheduler_provider.py:253-344`): no ESTOP check at the top of
  the function. Added an early-return `check_paused("cron-misfire")`
  guard. Self-healing — the next sweep after `hermes resume` catches
  everything up via the existing claim_fire path.

Both guards use `suppress(ImportError)` matching the ticker idiom so a
broken estop module fails open rather than killing cron. Distinct
component names keep the existing log-once mechanism independent per
surface.

Manual runs (`hermes cron run`, dashboard Trigger) deliberately
unchanged: operator override is arguably a feature, and PR #105144
already rewrites that path.

Tests (3 new, 49 pre-existing in affected files all pass):
- tests/cron/test_misfire_catchup.py::test_estop_engaged_skips_backstop
- tests/cron/test_misfire_catchup.py::test_estop_release_restores_backstop
- tests/gateway/test_api_server_jobs.py::test_fire_webhook_returns_503_when_estop_engaged
2026-09-15 03:33:13 -07:00
kshitijk4poor
840c00c124 refactor(cron): reuse _install_fire_secret_scope for the external-worker env build
_launch_external_cron_worker hand-rolled the same hydrate -> set_secret_scope ->
set_multiplex_context(routed) sequence, and the same context-then-scope reset, that
_install_fire_secret_scope/_reset_fire_secret_scope already encode for the in-process
fire. Two copies of the ordering is how the two paths drift apart; the helper is the
one place that owns it. The only observable difference (re-setting an already-active
multiplex context for the span) is a no-op the token reset undoes.

Rename _scope_token in _run_one_job_body to _fire_scope_tokens: it has held the helper's
(scope, context) tuple since the routed-fire change, not a single token.
2026-09-15 11:03:39 +05:30
kshitijk4poor
0fd59b447a fix(cron): scope the handoff multiplex context to the worker env build; gate the no_agent overlay
_launch_external_cron_worker wrapped fifty lines of dispatch, payload
write and scope hydration in the routed-fire multiplex context, though
only build_subprocess_env / strip_launch_profile_env read it. Compute
`multiplex_active` once (process flag OR routed fire), serialize it into
the payload, and set the context for exactly the env build inside the
existing secret-scope try/finally. Drop restore_managed_env on this path:
the worker re-runs load_hermes_dotenv -> _apply_managed_env at import
and strip_launch_profile_env already leaves managed keys in place.

_run_job_script now overlays the installed scope and re-applies managed
keys only under multiplex. In a single-profile process the scope is
os.environ, so the overlay could only re-sanitize values the child
already inherits; the "no-op outside multiplex" comment is now literal.
2026-09-15 11:03:39 +05:30
kshitijk4poor
f367ebeb5d refactor(cron): strip external-source residue in strip_launch_profile_env itself
_run_job_script popped the launch profile's secret-source names in an
inline loop right above strip_launch_profile_env, so only the no_agent
child got that protection; the four other callers of the same helper
(the external cron worker, scheduler_delivery, kanban dispatch and the
byterover plugin) still handed a served profile the launch vault or
1Password names. Fold the names into the helper's residue set, which is
already gated on multiplex and on the target not being the launch
profile, and keeps administrator-managed keys.
2026-09-15 11:03:39 +05:30
kshitijk4poor
fb4eaa59ee refactor(cron): inline the multiplex-context span into _launch_external_cron_worker
The `_inner` wrapper existed only to set and reset the multiplex contextvar
around the whole launch. Every consumer of that context (the payload flag,
`hydrate_profile_secret_sources`, `strip_launch_profile_env`, the scrubbed
`build_subprocess_env`) runs before the worker is spawned, so the span now
covers exactly the handoff-environment build and ends before `Popen` and the
acknowledgement wait. One function, one try/finally, same behaviour.
2026-09-15 11:03:39 +05:30
John Paul Soliva
62b4488cb5 fix(cron): routed fires are multiplexed at the worker handoff; managed keys keep policy precedence
Review findings on f5f88d5058. Three are defects the previous round introduced.

Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.

Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.

Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.

Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.

Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.

(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
2026-09-15 11:03:39 +05:30
John Paul Soliva
3dedff6a12 fix(cron): only strip launch external-source names when multiplexing
The source-name strip added alongside the external-source fix ran
unconditionally. Outside multiplexing there is no other profile to leak from --
os.environ IS this profile's environment -- so popping those names relied on the
routed scope overlay putting each one back, which in turn relies on the
per-home snapshot recorded at boot. Correct today, but it made a single-profile
child's credentials depend on bookkeeping that has nothing to do with isolation.

Guard it the way strip_launch_profile_env guards itself: no multiplexing, no
strip. A single-profile no_agent child now keeps a byte-identical env even if a
source's snapshot were ever missing. Pinned by a regression that runs a real
child with a source-owned name in os.environ and no multiplex context; making
the strip unconditional fails it.

(cherry picked from commit afa429b30a1c9c9b4c011f7d2426a9f099911596)
2026-09-15 11:03:39 +05:30
John Paul Soliva
0943e77136 fix(cron): strip launch external-source names too, and make the scope refresh replace
Two credential-isolation gaps found in review of the previous head.

1. strip_launch_profile_env() only knows dotenv- and terminal-config-owned names,
but external secret sources (vault, 1Password, ...) also write their names into
the shared os.environ and are tracked in secret_source_names(). A name the LAUNCH
profile's source supplied therefore still reached a routed no_agent child. Drop
every non-global source-owned name from the base; the routed scope overlay that
follows puts back exactly the ones that profile's OWN sources supply, since
build_profile_secret_scope folds get_secret_source_values(home) in.

2. refresh_installed_secret_scope() merged the rebuild with dict.update(), so a
name a source had stopped supplying -- rotated, revoked, source removed -- kept
its old value for the rest of the fire. Replace the mapping contents instead: the
rebuild is the profile's current truth.

Regressions: a routed child sees <unset> for a launch-source name while its own
source value comes through, and a refresh whose rebuild omits a name drops it.
Both fail if the corresponding change is reverted.

(cherry picked from commit ecd51517c4828a75acb5f458ed99aea0bc3e5e9f)
2026-09-15 11:03:39 +05:30
John Paul Soliva
023e4f997f fix(cron): drop the launch profile's dotenv residue before overlaying a routed no_agent scope
The routed no_agent child env started from all of os.environ and only overwrote the
names present in the installed scope. A name defined only by the LAUNCH profile's .env
and absent from the routed profile therefore reached the routed child with the launch
value instead of unset -- the secret scrub only knows classified names, so a custom or
unclassified secret crossed the profile boundary (review finding on the first head).

strip_launch_profile_env (main, 284d220ba4) is the primitive the external-worker path
already uses for exactly this: it drops the launch profile's dotenv-owned keys and the
bridged TERMINAL_* settings, and is a no-op outside multiplex or when the target IS the
launch profile. Apply it to the base BEFORE the scope overlay (so a shared name keeps its
routed value) and BEFORE the sanitizer (so routed values still pass the scrub and
passthrough rules). Pinned by a child-process negative control: the launch-only name
arrives <unset>, the shared name arrives routed, and the parent process is unchanged.

(cherry picked from commit 69349b527b5144c1f1540a4c10035d5ec798c1db)
2026-09-15 11:03:39 +05:30
John Paul Soliva
dbede34f6e fix(cron): a routed profile's cron fire in the desktop backend runs under multiplex semantics
The desktop backend ticks EVERY local profile's cron store from one process — its own docstring
says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global
multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every
isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared
`os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban
subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary
profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed
there after the tick, and a scope miss read the launch profile's tokens (#107692).

Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py)
is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not
the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the
override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the
profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that
span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are
therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs
before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam,
left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation
applies inside the routed fire with no per-site patching; the launch profile's own fires and the
backend's turns keep single-profile semantics; marker and override both reach the pool worker via
`copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through
`is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970).

Two consequences of suppressing the write are handled rather than left as regressions:
- a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`;
  the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub /
  passthrough rules apply to those values and the parent process is never mutated;
- plugin secret sources are discovered on the fire's first agent build, after the scope froze,
  and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds
  the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern
  `_publish_env_value` already uses for `.env` writes under multiplex).
And the profile's external secret sources are hydrated before the scope is frozen, the order
gateway/run.py and the external cron worker already use.

Tests pin each direction: the marker without the semantics before the scope, the semantics on and
off exactly with it, the marker reaching a copy_context worker; the process's own profile staying
single-profile; the restart-safe handoff's child env building without raising under a routed tick
with a passthrough key registered; a real child process receiving the routed values while
`os.environ` keeps the launch value; a source registered after the freeze reaching the fire
through the real PluginManager refresh. Reverting any one direction fails a distinct test.

(cherry picked from commit 2f87677425d2cca19286ac83bc45cab23e546669)
2026-09-15 11:03:39 +05:30
teknium1
8f7853188f fix(cron): keep deferred delivery exceptions from aborting ticks
Catch unexpected delivery exceptions after claim, retain diagnostics and continue
sibling admissions without authorizing replay. Preserve indefinite retention.

Reproduced PermissionError at target traversal after discovery. Native Electron
controlled-fault A/B confirms the healthy sibling settles and renders once.
2026-09-14 17:29:32 -07:00
teknium1
002ee41cfc fix(cron): keep unowned Bot Chat delivery on its resolved home
Extend deferred dispatch's destination pin to ordinary CLI fallback, so
custom-root and active-profile changes cannot redirect a checked target.
Refuse a missing destination before launch and name the target on failure.
Replace the old env-clearing expectation with two behavioral invariants
and retain the native Electron custom-root reproduction.

Adapted from the root-boundary fix and diagnosis in #104066.
Related #104055, #104066.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
2026-09-14 17:29:32 -07:00
teknium1
3b0fe0cc2b fix(cron): keep deferred Bot Chat delivery bound to admission
Carry the original destination home and delivery ID into deferred drain and
its child, rather than re-resolving a mutable profile/root. Missing destinations
fail closed; supported-owner handoffs remain transferred, not ambiguous failures.
Capture the producer root before the background thread starts, and retain/log
malformed JSON without stopping healthy admissions or the whole cron tick.

Two invariants reproduced failures on the published head. Real Electron root
change and malformed-record cases are red before and green after; nested DM
control remains passing. No automatic retry of claimed or uncertain turns.
2026-09-14 17:29:32 -07:00
teknium1
6af84ead60 test(cron): retain native Bot Chat delivery reproduction 2026-09-14 17:29:32 -07:00
teknium1
5d8390d1a4 fix(cron): retain Bot Chat output while a CLI owner is open
Keep never-started output behind unsupported owners and drain in admission
order after release. Persist claims before execution and never replay uncertain
started turns. Existing supported-owner receipts keep their authority.

Credits 686f6c61's residual queue proposal in #100319. This is a scoped
implementation, not general retry of failed CLI subprocesses.

Native Electron before/after: CLI-owned target previously returned
SESSION_NOT_OWNED and remained empty after release/tick; now its queued
output and reply appear once in the target Bot Chat. Nested quiet CLI
message_agent delivery to a named Desktop owner also passes on base.
2026-09-14 17:29:32 -07:00
teknium1
45b42202fa fix(cron): a raising profile gate ticks nothing; housekeeping respawns a dead ticker (#111010)
Follow-up on the salvaged #111034:

- `_start_multiplex` published the enumerated home list before the gate had
  filtered it, so a raising `profile_gate` (the Desktop stand-down probe from
  #100489) kept the thread alive but ticked every profile UNGATED — racing the
  gateway that owns them for the same cron store. The list is now assigned only
  after gating; a gate failure yields zero ticks for that cycle.
- `cron/scheduler_thread.py::SupervisedTickerThread` wraps the gateway ticker
  thread; `_start_gateway_housekeeping` gets a per-tick "Cron ticker supervisor"
  chore that respawns a ticker that ended without a stop request and logs the
  outage at ERROR. Every guard inside `start()` keeps the loop alive, but nothing
  outside it could notice a thread that had already ended.
- Tests trimmed to the invariants, proven red on origin/main: a REAL corrupt
  `executions.db` (no patched recover) no longer kills the ticker; a raising gate
  keeps the thread alive with zero ticks; housekeeping restarts a dead ticker and
  leaves a stopped one alone.

Root cause: the unguarded pre-loop `recover_interrupted_executions()` +
`record_ticker_heartbeat()` were added by d9dd05b69d (#61791, "truthful
execution ledger", 2026-07-09). The reporter's build (e440bf35) also carried
#107485's `completed_occurrence()` in the due scan, which opens the same ledger
on every tick — the first traceback in their errors.log is that in-loop hit
(caught); the restart then hit the SAME corrupt ledger from the pre-loop
recovery scan, which nothing caught: thread dead, no heartbeat, no error marker.
2026-09-14 16:14:33 -07:00
Konstantin Khlopkov
8fce4baf8c fix(cron): keep the ticker thread alive through startup recovery and marker writes (#111010)
The gateway runs InProcessCronScheduler on an unsupervised daemon thread, so any exception escaping outside the guarded tick body ends cron silently while the gateway keeps serving. Guard the three remaining escape windows: startup recovery + initial heartbeat in start(), per-cycle profile enumeration/gating in _start_multiplex, and every status-marker store write (heartbeat/error/clear) in both loops. A failing store now degrades to a logged warning and a failed cycle instead of thread death.
2026-09-14 16:14:33 -07:00
Jony
45aebc11b0 fix(cron): preserve cleanup profile scope 2026-09-14 16:13:51 -07:00
teknium1
498abb677e fix(cron): a successful run resolves the job's open incidents; a repeat re-opens them
The incident ledger only ever grew. A one-off failure (a drift skip after a
global model bump, a provider outage) stayed `detected`/`alerted` forever
after the job recovered, so `hermes cron incidents` listed 32 "open"
incidents on an install where all 32 jobs had since run OK, and the list
stopped saying anything about current health.

A successful run now marks that job's `detected`/`alerted` incidents
`resolved` (new state). `resolved` is distinct from the operator's `closed`
ack on purpose: `upsert_incident` re-opens a resolved incident as
`detected` when the same error signature recurs, so the operator is alerted
again for a job that broke a second time, while `closed` keeps the
signature silent as before. Wired from `_compose_run_delivery` next to the
failure-side upsert; best-effort, store errors never affect delivery.

CLI: `--state resolved` filter and a green `resolved` colour; `closed` is
now dim. Docs updated in the same change.
2026-09-14 09:19:32 -07:00
Teknium
0877decd1b Port from paradigmxyz/centaur#1479: cron runs no longer re-schedule themselves from recurring prompt language
A scheduled job whose prompt carries its own cadence phrasing ('Each
Monday, review...') can convince the agent to create ANOTHER cron job
at execution time instead of just doing the work — each run spawning a
sibling job. Hermes policy-denies the cronjob toolset in cron context
by default, but cron.allow_agent_scheduling: true re-enables it and
opens exactly this loop.

_build_job_prompt now extends the always-injected cron hint with a
RECURSION clause: this is a run of an existing job; never create or
update a cron job from schedule language in the task prompt; treat
cadence phrasing as context for this run.

Adapted from paradigmxyz/centaur#1479 (same failure mode in their
scheduled-task workflow runner).
2026-09-13 20:52:58 -07:00
Teknium
f72e79a111 fix(fallback): named custom providers keep their configured identity after automatic fallback (#98739)
resolve_runtime_provider returns the bare billing class 'custom' for every
named providers:/custom_providers: entry; the configured id only survives in
requested_provider. All three fallback resolvers (gateway, TUI/desktop, cron)
persisted runtime['provider'] as the agent identity, so an automatic fallback
labeled the session 'custom' in the UI and billing rows, while a manual
/model switch to the same provider showed the configured name.

New shared helper hermes_cli.fallback_config.effective_runtime_provider()
upgrades the bare class back to the entry's configured identity (ad-hoc
provider: custom entries stay unchanged), applied at all three sites —
same class as the delegate_tool fix.
2026-09-13 20:51:06 -07:00
teknium1
95c7e9a0cb fix(cron): return a strictly-later next run when the base sits in the DST fall-back hour
Port from QwenLM/qwen-code#11723: attaching the configured zone to croniter's
naive wall-clock result resolves the repeated autumn hour to its earlier
occurrence (fold=0), so a base inside the second occurrence received a
next_run_at up to an hour in the past — the fire path would treat it as due
immediately and loop. Try both folds of each candidate and return the earliest
instant strictly after the base; wall-clock jobs still fire exactly once on
the repeated hour (Vixie cron semantics).
2026-09-13 16:46:45 -07:00