Commit Graph

16 Commits

Author SHA1 Message Date
teknium1
248938b4e6 fix(cron): "Script not found" says scripts are per-profile and how to fix it
Cron scripts resolve only against the job's own profile scripts/ dir (by design,
profiles never share files), so a job copied between profiles fails with a path
that looks plausible and no hint about why. The runtime error now names the
profile folder and the two fixes (copy the script, or `hermes cron edit`).

Closes #94821. Builds on #105775 (creation-time existence check, credited).
2026-09-15 04:15:16 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
kshitijk4poor
0fd59b447a fix(cron): scope the handoff multiplex context to the worker env build; gate the no_agent overlay
_launch_external_cron_worker wrapped fifty lines of dispatch, payload
write and scope hydration in the routed-fire multiplex context, though
only build_subprocess_env / strip_launch_profile_env read it. Compute
`multiplex_active` once (process flag OR routed fire), serialize it into
the payload, and set the context for exactly the env build inside the
existing secret-scope try/finally. Drop restore_managed_env on this path:
the worker re-runs load_hermes_dotenv -> _apply_managed_env at import
and strip_launch_profile_env already leaves managed keys in place.

_run_job_script now overlays the installed scope and re-applies managed
keys only under multiplex. In a single-profile process the scope is
os.environ, so the overlay could only re-sanitize values the child
already inherits; the "no-op outside multiplex" comment is now literal.
2026-09-15 11:03:39 +05:30
kshitijk4poor
f367ebeb5d refactor(cron): strip external-source residue in strip_launch_profile_env itself
_run_job_script popped the launch profile's secret-source names in an
inline loop right above strip_launch_profile_env, so only the no_agent
child got that protection; the four other callers of the same helper
(the external cron worker, scheduler_delivery, kanban dispatch and the
byterover plugin) still handed a served profile the launch vault or
1Password names. Fold the names into the helper's residue set, which is
already gated on multiplex and on the target not being the launch
profile, and keeps administrator-managed keys.
2026-09-15 11:03:39 +05:30
John Paul Soliva
62b4488cb5 fix(cron): routed fires are multiplexed at the worker handoff; managed keys keep policy precedence
Review findings on f5f88d5058. Three are defects the previous round introduced.

Managed keys were stripped as launch residue. Recording every dotenv load as
residue swept in the administrator-managed `.env`, which `_apply_managed_env`
applies LAST with override precisely so it beats the user's own `.env`. A
routed child then lost `ORG_POLICY_FLAG=managed-value` to the routed user's
`user-value`. Managed keys are now recorded separately, never enter the
residue set, and are re-applied over the routed scope in both child builders
(`scheduler_script`, the restart-safe handoff) so the child sees the same
precedence the launch process does. `kanban_db_dispatch` and
`scheduler_delivery` strip without any overlay, so for them the exclusion
alone is the guarantee; the test pins the case that exercises it — the same
key defined in both the user and the managed file.

Private hydration did not record supplied names. `_hydrate_profile_secret_sources`
now feeds `provenance` plus `skipped_existing` into the same ownership set the
process-global path uses; the provenance label map stays applied-only.

Removal cleanup cleared its marker before the fallible work. A raising
reload left the removed plugin's credential active with no retry, because the
next no-source discovery saw the flag already false. The marker is cleared
only after reset, reload and installed-scope refresh succeed.

Routed fire not multiplexed at the handoff. `run_one_job` enables the
context in `_install_fire_secret_scope`, which runs AFTER
`_launch_external_cron_worker`, so a routed desktop fire on the managed path
serialized `multiplex_active=False` and built the worker env with launch
residue and no scrub. The handoff now treats `routed_profile_fire()` as
multiplexed for exactly its own span; the worker re-establishes the state from
the payload as before.

Each fix was checked by reverting it and confirming its regression fails,
including the overlay half and the exclusion half of the managed fix
separately.

(cherry picked from commit 329cbd8963d68c45b425e95a5b11ade59f513960)
2026-09-15 11:03:39 +05:30
John Paul Soliva
3dedff6a12 fix(cron): only strip launch external-source names when multiplexing
The source-name strip added alongside the external-source fix ran
unconditionally. Outside multiplexing there is no other profile to leak from --
os.environ IS this profile's environment -- so popping those names relied on the
routed scope overlay putting each one back, which in turn relies on the
per-home snapshot recorded at boot. Correct today, but it made a single-profile
child's credentials depend on bookkeeping that has nothing to do with isolation.

Guard it the way strip_launch_profile_env guards itself: no multiplexing, no
strip. A single-profile no_agent child now keeps a byte-identical env even if a
source's snapshot were ever missing. Pinned by a regression that runs a real
child with a source-owned name in os.environ and no multiplex context; making
the strip unconditional fails it.

(cherry picked from commit afa429b30a1c9c9b4c011f7d2426a9f099911596)
2026-09-15 11:03:39 +05:30
John Paul Soliva
0943e77136 fix(cron): strip launch external-source names too, and make the scope refresh replace
Two credential-isolation gaps found in review of the previous head.

1. strip_launch_profile_env() only knows dotenv- and terminal-config-owned names,
but external secret sources (vault, 1Password, ...) also write their names into
the shared os.environ and are tracked in secret_source_names(). A name the LAUNCH
profile's source supplied therefore still reached a routed no_agent child. Drop
every non-global source-owned name from the base; the routed scope overlay that
follows puts back exactly the ones that profile's OWN sources supply, since
build_profile_secret_scope folds get_secret_source_values(home) in.

2. refresh_installed_secret_scope() merged the rebuild with dict.update(), so a
name a source had stopped supplying -- rotated, revoked, source removed -- kept
its old value for the rest of the fire. Replace the mapping contents instead: the
rebuild is the profile's current truth.

Regressions: a routed child sees <unset> for a launch-source name while its own
source value comes through, and a refresh whose rebuild omits a name drops it.
Both fail if the corresponding change is reverted.

(cherry picked from commit ecd51517c4828a75acb5f458ed99aea0bc3e5e9f)
2026-09-15 11:03:39 +05:30
John Paul Soliva
023e4f997f fix(cron): drop the launch profile's dotenv residue before overlaying a routed no_agent scope
The routed no_agent child env started from all of os.environ and only overwrote the
names present in the installed scope. A name defined only by the LAUNCH profile's .env
and absent from the routed profile therefore reached the routed child with the launch
value instead of unset -- the secret scrub only knows classified names, so a custom or
unclassified secret crossed the profile boundary (review finding on the first head).

strip_launch_profile_env (main, 284d220ba4) is the primitive the external-worker path
already uses for exactly this: it drops the launch profile's dotenv-owned keys and the
bridged TERMINAL_* settings, and is a no-op outside multiplex or when the target IS the
launch profile. Apply it to the base BEFORE the scope overlay (so a shared name keeps its
routed value) and BEFORE the sanitizer (so routed values still pass the scrub and
passthrough rules). Pinned by a child-process negative control: the launch-only name
arrives <unset>, the shared name arrives routed, and the parent process is unchanged.

(cherry picked from commit 69349b527b5144c1f1540a4c10035d5ec798c1db)
2026-09-15 11:03:39 +05:30
John Paul Soliva
dbede34f6e fix(cron): a routed profile's cron fire in the desktop backend runs under multiplex semantics
The desktop backend ticks EVERY local profile's cron store from one process — its own docstring
says "like a multiplex gateway" (hermes_cli/web_server.py) — but never sets the process-global
multiplex flag, and cannot: its own chat turns are unscoped and would fail closed. Every
isolation in the tree keys on that flag — the guard that keeps a routed `.env` out of the shared
`os.environ`, `get_secret`'s fail-closed miss, passthrough resolution, the MCP and kanban
subprocess scrubs — so all of it was inert for a sibling profile's fire. Verified: a secondary
profile's API keys replaced the launch profile's in `os.environ` with `override=True` and stayed
there after the tick, and a scope miss read the launch profile's tokens (#107692).

Give multiplex mode a context-local counterpart. `set_multiplex_context` (agent/secret_scope.py)
is OR'd into `is_multiplex_active()`. `_profile_cron_scope` only MARKS a fire whose home is not
the process's own (`routed_profile_fire`, decided against `get_process_hermes_home()`, the
override-immune resolver); `_install_fire_secret_scope` in cron/scheduler.py installs the
profile's hydrated secret scope and, for a marked fire, the multiplex context — for exactly that
span, dropped again before the scope by `_reset_fire_secret_scope`. Multiplex semantics are
therefore never active in cron without a scope to read: `run_one_job`'s restart-safe handoff runs
before the body's scope and keeps today's semantics (its own scope is #107413 / #106050's seam,
left untouched so this composes with whichever lands). Every existing multiplex-keyed isolation
applies inside the routed fire with no per-site patching; the launch profile's own fires and the
backend's turns keep single-profile semantics; marker and override both reach the pool worker via
`copy_context()`. `get_secret` read the raw global in its miss branch; it now goes through
`is_multiplex_active()`. The dotenv guard keeps its pinned flag-only form (#77970).

Two consequences of suppressing the write are handled rather than left as regressions:
- a `no_agent` script's env is `os.environ.copy()`, which no longer carries the routed `.env`;
  the runner overlays the installed scope onto the base BEFORE sanitizing, so the same scrub /
  passthrough rules apply to those values and the parent process is never mutated;
- plugin secret sources are discovered on the fire's first agent build, after the scope froze,
  and the post-discovery reload is hydrate-only under multiplex semantics; the refresh now folds
  the values into the installed scope in place (`refresh_installed_secret_scope`, the pattern
  `_publish_env_value` already uses for `.env` writes under multiplex).
And the profile's external secret sources are hydrated before the scope is frozen, the order
gateway/run.py and the external cron worker already use.

Tests pin each direction: the marker without the semantics before the scope, the semantics on and
off exactly with it, the marker reaching a copy_context worker; the process's own profile staying
single-profile; the restart-safe handoff's child env building without raising under a routed tick
with a passthrough key registered; a real child process receiving the routed values while
`os.environ` keeps the launch value; a source registered after the freeze reaching the fire
through the real PluginManager refresh. Reverting any one direction fails a distinct test.

(cherry picked from commit 2f87677425d2cca19286ac83bc45cab23e546669)
2026-09-15 11:03:39 +05:30
Teknium
284d220ba4 fix(multiplex): cron, kanban, /loop and completion paths for a served profile match its standalone gateway
Under gateway.multiplex_profiles a secondary profile X is ticked, dispatched
and notified from the default profile's process, where os.environ holds the
DEFAULT profile's .env and X's values live only in the per-turn secret scope /
HERMES_HOME override. Every remaining read that skipped that scope made X
behave differently from `hermes -p X gateway run`:

- cron: HERMES_CRON_TIMEOUT, HERMES_MODEL (job/preflight fallback),
  HERMES_CRON_MAX_PARALLEL, inflight allowance, prefill file and the script
  timeout were bare os.getenv → the default profile's values; a job without a
  model silently ran on the default's HERMES_MODEL instead of refusing.
  cron/env_settings.py::cron_env_setting reads the scope (fire) or the ticked
  home's .env (tick thread), plain environ when multiplexing is off.
- child env: the restart-safe cron worker, the Bot Chat delivery child and the
  kanban worker inherited the launch profile's non-credential .env settings
  and bridged TERMINAL_* policy (TERMINAL_ENV=docker, default's image,
  HERMES_MODEL) — X's worker ran in the default's docker image on the
  default's model. tools/environments/local.py::strip_launch_profile_env drops
  them when the child targets another served profile.
- kanban: the worker --toolsets pin was silently dropped for every served
  assignee (toolset probes call get_secret without a scope → swallowed
  UnscopedSecretError); notifier pings, artifact uploads and the wake text ran
  under the default's media policy / display language (only wake() was scoped).
- /loop: _post_turn_loop_completion hopped to the executor without contextvars,
  writing the completed tick into the DEFAULT profile's state.db and leaving
  X's row awaiting_response forever; the --until judge ran with the default's
  aux credentials.
- background processes: a secondary's processes.json (scope-relative since
  adf23550f5) was never read at startup; its processes were not re-adopted
  and notify_on_complete notices were lost. Startup recovers every served
  home under its scope; recovery adopts each session once.
- completion delivery: background_process_notifications was evaluated once per
  drain for the ambient profile (default's mode for everyone; X's `off`
  dropped a sibling's `all` event), recovered watchers used the default's
  mode, HERMES_BACKGROUND_NOTIFICATIONS was read raw from environ;
  _deliver_platform_notice used the default's GatewayConfig so a secondary's
  notice_delivery: private went public.

Not changed: gateway/run.py and tools/async_delegation.py (PR #106742
rewrites both). Known residue left for the env-bridge lane:
HERMES_SESSION_STALL_TIMEOUT is bridged once from the launch config.
scheduler bug, not a parity gap; unchanged here.
2026-09-11 19:58:07 -07:00
Teknium
a9b0dd6742 simplify(compat): cron scheduler — drop 82 re-exports, repoint 8 callers + 33 test files
cron/scheduler.py no longer re-exports the split modules (scheduler_delivery /
_script / _prompt / _preflight); it imports only the 19 names it calls itself
(bottom-of-file, E402 kept for the import cycle). Dropped the shim-only
`import shutil` and the F401 note on windows_hide_flags (still used by
scheduler.py). Split modules now call same-module helpers directly, reach
sibling split modules via late-bound module refs (_delivery/_script/_preflight)
next to _sched, and import windows_hide_flags themselves; origin-resident names
(load_config, Path, _SCRIPT_TIMEOUT, heartbeat_run_claim, ...) still go through
_sched. Callers/tests import + patch the defining module.
2026-09-03 13:38:36 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
2c0961b334 refactor(cron): table-dispatch media send, unify tree-kill fallback + home env lookup, squeeze local-import blanks 2026-09-02 20:12:25 -07:00
Teknium
829d7331c7 refactor(cron): pack brackets, fold closers, compact docstrings in scheduler_* modules (AST-neutral) 2026-09-02 19:51:39 -07:00
Teknium
76a1ebdb4c refactor(cron): extract per-target delivery prologue, unify timeout/env/skill-name helpers in scheduler_* modules 2026-09-02 19:31:14 -07:00
Teknium
5a8fbbf16c refactor(cron): split scheduler.py into delivery/script/prompt/preflight modules
cron/scheduler.py 6386 -> 3423. Four sibling modules hold the free-function clusters,
re-exported from cron.scheduler so every existing import/monkeypatch target keeps resolving;
origin-resident helpers are reached late-bound via `_sched` (same-object visibility, no
import cycle). Every moved node is AST-identical to the base. Source-inspection test for the
bash resolver repointed to cron/scheduler_script.py.
2026-09-02 18:34:26 -07:00