Commit Graph

952 Commits

Author SHA1 Message Date
teknium1
f4075a20d6 feat(telemetry): gateway platform health, delivery, first-reply latency and cron run metrics
The gateway and cron ticker were blind spots in shared metrics: we could not
tell which messaging platforms fail to connect or drop, how often replies fail
to reach users, how long users wait for a first reply, or whether scheduled
jobs run, fail, get skipped or get missed while Hermes was down.

New opt-in counters (all through the existing enabled() / record_process_mark
gate, recorded fire-and-forget on one background worker that runs in a copy of
the caller's context so the owning profile gets the row):

- hermes.platform.health: platform, event (connect_ok/connect_failed/
  reconnect/disconnect), error_class. One seam in the runner
  (_connect_adapter_with_timeout, used by cold start, multiplex secondaries and
  the reconnect watcher) plus the fatal-error handler; classified from
  exception types, HTTP statuses and Hermes's own fatal codes, never text.
- hermes.platform.delivery: platform, outcome, failure_class. One row per
  logical BasePlatformAdapter._send_with_retry call (retries and plain-text
  fallback included), from SendResult's typed fields or the exception type.
- hermes.gateway.reply_latency: platform, first_response_bucket. Clock starts
  when _handle_message accepts a non-internal turn; stops on the stream
  consumer's first delivered text or send_final_ledgered. Busy acks and
  command replies never stop it.
- hermes.cron.run: outcome (success/failed/missed/skipped), delivery_kind,
  duration_bucket. Recorded at the write-once terminal execution row
  (finish_execution) and where the due scan drops an occurrence (catch-up
  disabled, expired one-shot). Job names, prompts and targets never leave.

Platform names: platforms Hermes ships under plugins/platforms/ now report by
name, and a plugin-catalog platform reports its catalog entry name only when
the installer-owned .install-metadata.json record proves a catalog install of
the plugin dir that defines the registered adapter factory (a URL install
cannot claim a catalog name via its tree or manifest). Everything else stays
"plugin". GATEWAY_PLATFORMS becomes catalog-backed (_CatalogValues), and the
v3 schema's platform fields accept published catalog identifiers.
2026-09-28 12:43:03 -07:00
Teknium
a5bd246865 Old pre-decomposition import paths are gone: plugin compat layer removed on schedule (#126164)
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule

The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:

- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
  names, restored dead definitions) and the three re-export stub modules
  (gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
  the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
  `plugins.allow_deprecated_imports` escape hatch

An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.

hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).

In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.

* chore: retrigger CI (zero-job startup_failure phantom)

* test: drop resolution allowlist rows for the two deleted which() sites

hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
2026-09-28 10:21:41 -07:00
alt-glitch
318f6ac63f fix(cron): order execution, delivery and incident history by instant, not by text
The cron SQLite stores write hermes_time.now().isoformat(), an ISO string
that carries the local UTC offset. That offset changes at a DST transition
and after a timezone config change, so the text order of these columns is
not their time order. Across a fall-back, 01:10-05:00 is 20 minutes after
01:50-04:00 but sorts before it. latest_execution(), latest_executions()
(which fills job["latest_execution"] for `hermes cron list` and the
dashboard), list_executions() and its before_claimed_at cursor, the
delivery queue's pending-claim order, delivery and execution pruning, and
list_incidents() all returned the older row first.

Every ORDER BY and the cursor comparison on these columns now use
julianday(col), which SQLite resolves to the same instant for any offset.
The raw text stays as the next sort key, so rows within the same
millisecond (julianday's resolution) keep their microsecond order, and the
cursor compares the same (instant, text) key on both sides. Stored values
and the schema are unchanged, so existing rows with mixed offsets order
correctly without a migration.

latest_executions() moves from a correlated per-row subquery to one
ROW_NUMBER() window. With julianday() in the ORDER BY the index on
claimed_at no longer serves the per-row LIMIT 1, and the correlated form
took about 90 ms at 1000 rows; the window form takes about 1 ms.

Refs #86520
Thanks to #86889 (zhao0112) for the executions-ledger diagnosis.
2026-09-28 19:07:43 +05:30
teknium1
17957f5014 fix(review): ticker heartbeat proves liveness only while its writer is alive; routed hand-off never trusts an unknown probe; satellite manual runs get the ticker's routed grant
MAJOR — `_builtin_gateway_liveness` trusted a bare-epoch ticker heartbeat for
~200 s after its writer died, for every profile, and the multiplexer pid gate
had been dropped: a killed serve/Desktop ticker read as alive and
`hermes -p X cron run` queued a primary-routed run "for the gateway's next
tick" with no ticker present.

- `cron/jobs.py::record_ticker_heartbeat` stamps `<epoch> <pid>`;
  `get_ticker_heartbeat_age` parses the first field (legacy bare stamps still
  yield an age); new `ticker_heartbeat_writer_alive` requires the stamped pid
  to be alive and treats a bare stamp as NOT proof by itself.
- The heartbeat-only rung is now `fresh AND (served-by-multiplexer OR writer
  alive)`: the multiplexer record proves the host process for a served named
  profile (and covers a stale-code multiplexer still writing bare stamps), so
  only the in-process serve/Desktop ticker relies on the heartbeat, and then
  only with a live writer. `cron status`'s in-process-ticker rung applies the
  same rule.

MINOR (a) — `_hand_off_primary_routed_run` gated on `is True`; an unknown
(None) probe returns an error naming the uncertainty instead of "queued …
runs and delivers it".

MINOR (b) — `_run_claimed_job` wraps a shared-bot satellite's resolved map in
`SharedRouteAdapters(primary, _primary_profile_routes_for_current_home())`,
the same grant the ticker's `tick_adapters_for` makes, instead of handing it
the full primary adapter map.

Tests (each red on the previous head a18977090be):
- test_cron_satellite_diagnostics.py::test_in_process_ticker_heartbeat_counts_only_while_its_writer_lives
- test_cronjob_run_primary_routed.py::test_routed_run_without_a_serving_gateway_fails_before_the_turn[None]
- test_cronjob_run_immediate.py::test_execute_job_now_grants_a_shared_bot_satellite_only_its_routed_targets
2026-09-28 05:28:05 -07:00
teknium1
50b4c07c98 fix(cron): trim external-worker plugin discovery salvage to the bar (#121929, salvage #121950)
Keeps #121950's one-line fix — `discover_plugins()` under the payload's home override before
`hydrate_profile_secret_sources()` in `_run_external_worker_payload` (L3 ticker -> L5 child
process; T2/T4) — and trims around it: the comment keeps only the WHY, and the two salvaged
tests collapse into one invariant that proves the owning-profile scoping the issue asks for:
the worker's launch HERMES_HOME has no plugins, the payload's profile home ships a
`register_secret_source` plugin, and the plugin's value reaches the secret scope seen by
`run_one_job`. The bare-profile control case was already covered by
`tests/cron/test_restart_safe_worker.py::test_external_worker_adopts_execution_and_runs_payload_once`.
2026-09-28 05:28:05 -07:00
teknium1
b9372c409f fix(cron): satellite route lookup reads the primary's layered config (#121212, salvage #121595)
Trim of #121595: instead of consulting the managed file first and the raw user file second
through two new helpers, `_primary_profile_routes_for_current_home` now reads the primary home
through `gateway.config_loader.read_yaml_layers` — the exact user-file + managed-overlay layering
`load_gateway_config()` uses, so the satellite side cannot drift from the gateway's own routing
view (L8: behavioural read through the surface's loader, never `read_user_config_raw`). The
absent-user-file early return goes too: a fleet host with only `/etc/hermes/config.yaml` must
still see its routes, exactly as the gateway loader does.

Both consumers ride the one helper — preflight rescue (`_delivery_platform_routed_from_primary_gateway`)
and the delivery-time `SharedRouteAdapters` fallback in `scheduler_provider.tick_adapters_for` and
the `gateway/run.py` delivery-queue drain — so managed-scope routes stop false-blocking and stop
failing delivery closed on every satellite (T2/T4).

Tests trimmed to two invariants on real files: managed-only routes reach both preflight and the
delivery adapter map; managed routes replace user-file routes (list precedence).
2026-09-28 05:28:05 -07:00
D Yiapanis
411c0405b6 fix(cron): discover plugins in external workers so plugin secret sources hydrate
External cron workers start with the builtin secret-source registry alone —
plugin-registered sources (the documented path for third-party vaults,
developer-guide/secret-source-plugin) only exist after plugin discovery runs.
Since the restart-safe topology hands every scheduled agent-mode job to a
worker process, jobs on secondary profiles under a multiplexed gateway died
at runtime resolution with 'No usable credentials found for provider ...'
while identical direct runs in the gateway succeeded.

Run the idempotent discover_plugins() under the profile home override before
hydrating: discovery registers the profile's plugin sources and re-pulls
enabled ones (#64177 bootstrap parity), the same lazy-load pattern the
delivery path already uses for plugin platforms.

Fixes #121929

(cherry picked from commit 248d441559478978cbdae1cc2178856997bef5f3)
2026-09-28 05:28:05 -07:00
ahisblessed
40715d95a0 fix(cron): read managed-scope gateway.profile_routes in preflight (#121212)
_primary_profile_routes_for_current_home() read only the primary home's
raw config.yaml, so routes pinned in the managed scope (e.g.
gateway.profile_routes in /etc/hermes/config.yaml) never reached it:
preflight false-blocked satellite agent jobs and routed delivery failed
closed. The managed overlay is now consulted first (it wins over the user
file, as in the merged config), with the raw user read as fallback.

Fixes #121212

(cherry picked from commit 7b05243220e98a4c8e22587d51ca786492cb148b)
2026-09-28 05:28:05 -07:00
Jay Zhou
49bc08d272 fix(cron): record the live rejection when the router raises it
DeliveryRouter._deliver_to_platform raises a failed SendResult's error,
so send_path_degraded reached _live_send_text's except arm and never set
live_error. The reconnect queue then saw None and dropped the payload.

(cherry picked from commit 4c40b7ea687386668d34a28a96672bbe3460ebc2)
2026-09-28 05:28:05 -07:00
Jay Zhou
63d8a0d5b2 fix(cron): keep a reconnect-only rejection for the adapter when standalone fails
A cron target whose live-lane send came back `send_path_degraded` (the
adapter will deliver after it reconnects) fell through to the standalone
sender. On a satellite profile the cron worker has no platform token, so
standalone failed deterministically and the payload was lost, though
the adapter holding the token was one reconnect away.

When standalone fails after a reconnect-only live rejection, the payload
is now recorded in the delivery ledger as a failed reconnect-only row
owned by the adapter's profile, so the existing post-reconnect sweep
redelivers it. Standalone still runs first: a profile that can send
delivers immediately, attachments included, and nothing is queued. The
ledger carries text only; dropped attachments are reported. Failure
reports now include the target's thread id.

Fixes #125363

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit b47ecf14646189882ce1689bf88f13511b69a7ce)
2026-09-28 05:28:05 -07:00
teknium1
5392a83ccd fix(kanban): module-form worker carries the install root on PYTHONPATH (#122299, #122487, #122500, salvage #122420)
_resolve_hermes_argv proves hermes_cli importable in the GATEWAY, where the
store-python shim / launcher bootstrap put the repo root on sys.path in-process.
The worker env scrub (build_subprocess_env / served_profile_child_env) strips
Hermes-owned PYTHONPATH entries, so the bare `sys.executable -m hermes_cli.main`
child could not import the package the parent just proved importable and every
worker died at birth (ModuleNotFoundError: hermes_cli) -> crashed x2 -> gave_up.

Reuse cron's pin_hermes_tree_on_pythonpath (#112729) on the module argv only:
a resolved shim path owns its imports. hermes_cli.main's bootstrap then
activates the committed PM dependency generation as usual, so a store Python
with an empty site-packages still boots. Same class in the Bot Chat delivery
runner (cron/scheduler_delivery.py), which builds the module argv after
served_profile_child_env.
2026-09-28 03:37:09 -07:00
kshitijk4poor
6e69a8933a refactor(cron): simplify the interpreter salvage after review
- Refuse pythonw: it discards captured output, and the user-interpreter
  path skips the Windows helper that used to swap it for python.exe, so
  an agent job would go silently quiet.
- Drop the one-off key pop on clear: like workdir/monitor_script, a
  cleared interpreter is stored as null and _script_argv already treats
  it as unset.
- Drop an unused import in the no-agent interpreter test.
- Docstring says what the check is (name-based), parametrize ids.
2026-09-28 02:32:05 +05:30
kshitijk4poor
4450cc9d90 fix(cron): only accept a Python executable as a job interpreter
The lifecycle guard classifies a `.py` script as Python and skips its
shell reference walk, so `interpreter=/bin/bash` on a `.py` whose body is
`bash restart.sh` was created and then executed as shell. Refuse any
interpreter whose name (or symlink target's name) is not a Python image,
reusing the guard's own `_INTERPRETER_IMAGE_RE`.

One parametrized test covers every refusal (bare name, missing,
directory, non-executable, /bin/bash, python -> /bin/bash symlink); the
two run-path tests now use a `python3`-named wrapper.
2026-09-28 02:32:05 +05:30
M1racleShih
6f7cc7e74c feat(cron): allow Python scripts to use an external interpreter
Add an optional per-job `interpreter` field so a cron Python `script` /
`monitor_script` can run under a user-managed venv instead of Hermes' own
Python, letting scripts import packages the Hermes runtime does not carry
(#8714). Nothing is installed, frozen, or restored automatically.

- cron/jobs.py: persist + normalize the field (absent => record unchanged;
  empty string clears it on update).
- cron/scheduler_script.py: _resolve_cron_interpreter() validates the path
  at run time (absolute/~ required, regular file, executable on POSIX);
  _script_argv runs [interpreter, script] and skips the managed-store
  bootstrap/PYTHONPATH overlays, which exist for Hermes' own venv.
  Threaded through _run_job_script, the claim-heartbeat wrapper, the
  pre-run prompt path and monitor scripts.
- hermes_cli: --interpreter on `cron create` / `cron edit`; shown in
  details and `cron list`.
- tools/cronjob_tools.py: programmatic/CLI lane only, like model and
  reasoning_effort — absent from the model-facing schema.

Shell scripts (.sh/.bash) still always run under bash. Revives #8741.

Ported onto current main from #70500 (the scheduler moved to
cron/scheduler_script.py and the CLI/tool became table-driven since the
PR's base).

Co-authored-by: MestreY0d4-Uninter <241404605+MestreY0d4-Uninter@users.noreply.github.com>
2026-09-28 02:32:05 +05:30
kshitijk4poor
758ad514eb fix(cron): re-check disk before an unmergeable repair replaces the store
A degraded-lock sibling may rewrite jobs.json between load_jobs' read of an
id-keyed or invalid-jobs store and its repair save; only force the replace
while disk is still a shape the merge cannot read.

Co-authored-by: Ayushman Padhi <208280836+ayushmanpadhi@users.noreply.github.com>
2026-09-27 18:30:44 +05:30
kshitijk4poor
c642f42bc3 fix(cron): fail closed inside the merge peek; keep merge on mergeable repairs
Review fold for the corrupt-store refusal:

- Move the refusal into _unmerged_disk_jobs (after the #80703 stat-stamp
  fast path) instead of a separate pre-check in _save_jobs_unlocked. The
  pre-check parsed jobs.json on every save, doubling the parse and
  defeating the stamp fast path on the scheduler's per-fire saves. A stamp
  match already proves disk is the file load_jobs parsed cleanly, and the
  in-merge raise also covers the verify-after-stage re-peek (TOCTOU the
  pre-check left open). replace=True never reaches it.
- load_jobs' auto-repair uses replace=True only for the two shapes the
  peek cannot read (id-keyed map, non-list "jobs" field). Every other
  repair keeps the shrink-merge, so a sibling's create that lands during
  a repair under the degraded flock-timeout lock is preserved again
  (#80624), as it was before the refusal.
- fsync the directory atomic_replace actually renamed into (it resolves a
  symlinked jobs.json), matching utils._atomic_write; refresh the stale
  comments/docstrings and pin the refusal message in the test.

Co-authored-by: Ayushman Padhi <208280836+ayushmanpadhi@users.noreply.github.com>
2026-09-27 18:30:44 +05:30
kshitijk4poor
553388b320 fix(cron): refuse to overwrite a corrupt jobs store on save
A merging save_jobs() over an unreadable jobs.json treated the store as
empty (the non-repairing peek returns None and the shrink-merge skips it),
so any save that landed after corruption — e.g. behind a degraded-lock
sibling's non-atomic copy fallback — silently replaced every job on disk.
Fail closed instead: raise and leave the bytes untouched. replace=True
remains the explicit disaster-recovery rewrite, and load_jobs' own repair
(a locked full-store read, so its repaired list is authoritative, incl.
id-keyed maps the peek deliberately refuses to flatten) now uses it.

Also fsync the parent directory after the atomic rename so the published
store survives power loss, via the existing utils.fsync_directory.

Lock-timeout and EXDEV fallback policy are unchanged.

Co-authored-by: Ayushman Padhi <208280836+ayushmanpadhi@users.noreply.github.com>
2026-09-27 18:30:44 +05:30
kshitijk4poor
009699b8a0 fix(cron): a missing dependency interpreter fails the script run
_posix_cron_script_argv fell back to sys.executable when the selected venv's
interpreter was gone. On a managed-store install that is the bare store Python,
so the run logged a warning and then died with the same ModuleNotFoundError as
#123044. A half-deleted venv reaches this branch because _recorded_venv only
checks pyvenv.cfg. Raise instead; _run_job_script's try reports it as a failed
run naming the missing interpreter, matching how a broken PM record is handled.
The existing broken-selection test is parametrized over both causes.
2026-09-27 15:04:50 +05:30
kshitijk4poor
4d4aed0f98 fix(cron): POSIX script bootstrap keeps plain-run __main__ semantics
Review follow-up on the live-checkout bootstrap.

- runpy.run_path runs the script under a temporary __main__ that is
  swapped back once the body returns, so atexit handlers or non-daemon
  threads that pickle script-defined classes failed ("not found as
  __main__.Foo") where a plain `python script.py` works. The bootstrap
  now installs a real __main__ module (SourceFileLoader, __cached__ as
  in a plain run) and execs the compiled script in it. Tracebacks keep
  one bootstrap frame instead of runpy's three.
- Under -P / PYTHONSAFEPATH there is no cwd entry to replace, so the
  repo is prepended instead of overwriting a stdlib path.
- Interpreter via pm.environments.project_python (the existing
  venv_python(selected_venv()) helper).
- A missing venv interpreter still falls back to the caller's Python
  but now logs a warning naming it, like the Windows bootstrap fallback.
- Docstrings: the lazy-install gate's real reason (a script importing
  hermes_bootstrap could complete a source update and execv itself onto
  the bare store Python via sys.orig_argv); _script_argv no longer
  claims "else sys.executable".
2026-09-27 15:04:50 +05:30
kshitijk4poor
3cdb0c80a6 fix(cron): POSIX cron scripts import the live checkout; broken PM selection fails the run
Follow-up to the venv-interpreter change above.

- The selected venv resolves Hermes from its generation's workspace
  snapshot, which only a dependency change (uv.lock / extras / Python /
  plugins, pm/packages.py::expected_stamp) rebuilds. After a code-only
  update, scripts imported older Hermes code than the gateway runs. A
  `python -c` bootstrap now puts the live checkout right after the
  script's directory, in-process, so nothing is inherited by the
  script's children (no PYTHONPATH, #123440).
- POSIX dispatch moves into _script_argv (platform branch), so
  _windows_cron_python_invocation is Windows-only again and the
  PYTHONPATH-keyed bootstrap gate goes back to its original form.
- The interpreter path comes from pm.environments.venv_python.
- _script_argv now runs inside _run_job_script's try. PM record reads
  (store manifest, facts, selection) can raise ValueError/KeyError as
  well as RuntimeError; before, those escaped the runner and stranded
  the execution row, and on POSIX the local fallback handed the script
  to the bare store interpreter (the #123044 symptom). A broken
  selection is now a failed run with the PM error, per
  selected_venv's contract. This also covers the same pre-existing
  gap on the Windows committed_venv path.
- Tests: two invariant tests replace the three change-detector tests
  (live checkout beats a snapshot on sys.path, venv site-packages
  resolve, script-dir sys.path[0], no PYTHONPATH; broken selection
  fails the run instead of escaping).
2026-09-27 15:04:50 +05:30
liuhao1024
03156afc7b fix(cron): degrade to the caller's interpreter when selected_venv raises
_posix_cron_python_invocation called selected_venv(repo) unguarded; the
four RuntimeError cases it is documented to raise would escape through
_script_argv (which runs before _run_job_script's try), crash the tick,
and leave the execution row in running forever — the exact half-migrated
install the test docstring already claimed was supported. Mirror the
Windows bootstrap's degrade-don't-crash contract: warn and run on the
caller's interpreter. Adds the raising-selection case to the POSIX
invocation test (red on the previous head), a sealed-payload caveat on
the e2e test's hand-written .pth, and a pointer from the Windows
invocation docstring to its POSIX counterpart.

(cherry picked from commit 8d5b3b5e9727d612d91b2a7485e0f4bb245358f3)
2026-09-27 15:04:50 +05:30
liuhao1024
cb11c994f7 fix(cron): run POSIX .py job scripts on the selected venv interpreter
On managed-store installs the gateway's store Python carries the repo and
the managed site-packages only in-process; cron .py scripts spawned with
sys.executable re-resolve imports from scratch and die with
ModuleNotFoundError (#123044). Run them on the selected dependency venv's
interpreter instead — its pyvenv.cfg site-packages (editable installs
included) need no PYTHONPATH overlay, so children the script spawns never
inherit the store's paths and import CPython-3.14 extension modules on a
foreign interpreter (#123440). Lazy installs are disabled for script
children so hermes_bootstrap imports off the store-record venv cannot
republish launchers. The .pth bootstrap routing in _script_argv is
tightened to PYTHONPATH-carrying overlays, which the POSIX path no
longer produces.

(cherry picked from commit 2d6976c3d6d8b877ccf73672fda2e963025ec30b)
2026-09-27 15:04:50 +05:30
kshitijk4poor
7ca5cca50a fix(cron): an Infinity or negative repeat.completed no longer breaks load_jobs
06a495cc5b normalized non-int counters but caught only TypeError/ValueError.
json.loads turns a hand-edited Infinity / -Infinity / 1e999 into float inf,
int(inf) raises OverflowError, and that escaped load_jobs, so every job
(list, tick, mark_job_run, hermes cron list) failed, not just the bad one.
Catch OverflowError (-> 0), and also clamp a negative int count, which
granted extra runs. The docs sentence now says non-negative.
2026-09-26 23:00:54 +05:30
kshitijk4poor
06a495cc5b fix(cron): normalize any non-int repeat.completed, not only null
A hand-edited "completed": "2" still crashed every recorded run ("2" + 1),
and 1.0 was stored as 2.0 ("2.0/3"). load_jobs now coerces any non-int
counter to a non-negative int (0 when unparseable). Document the load-time
repair next to the direct-edit tip.

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-26 22:04:17 +05:30
kshitijk4poor
d2b0b7ea82 fix(cron): normalize a null repeat.completed once in load_jobs
A hand-edited "completed": null was patched with `.get("completed") or 0`
at five arithmetic readers. The display readers were still missed: the
cronjob tool listed a never-run one-shot as "1/1", and `hermes cron`
printed "None/3". Any new reader would bring the bug back.

Every reader gets its jobs through load_jobs: mark_job_run, update_job,
claim_dispatch, the due scan, merge_job_definition, list_jobs/get_job for
cronjob_job_args._repeat_display and hermes_cli/cron._job_rows. So reset a
null to 0 once in the load_jobs repair pass, which also persists the fix,
and change the per-site `or 0` patches back to their base form.

Co-authored-by: mochamgx <1114149@qq.com>
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 22:04:17 +05:30
kshitijk4poor
a7a2684677 fix(cron): log each jobs.json repair warning once, from the locked pass only
Since load_jobs re-runs itself under the job-store lock whenever it finds
a repair, every detection-time warning ran twice: once on the unlocked
pass and again on the locked re-read. The id-keyed "Skipping N non-dict
entries" warning had logged once on base, so this was a new regression.
The scalar 'jobs' warning doubled too. Only the list-junk warning had a
lock-depth guard.

Collect the repair details as notes and emit them next to the existing
"Auto-repaired jobs.json" line, which only the locked pass reaches. With
that, the per-warning depth guard is no longer needed. Also build the
junk list once instead of scanning jobs three times; the
isinstance(jobs, list) check was always true at that point.

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-26 22:04:17 +05:30
kshitijk4poor
ca7eded74c fix(cron): treat null repeat.completed as 0 in update_job and job definitions
#122114 guarded mark_job_run/claim paths against a hand-edited
"completed": null, but update_job and the job_definition merge still used
.get("completed", 0), which returns None when the key is present, so the
null was carried forward into the store. Use `or 0` there too, and pin the
behaviour with one test covering mark_job_run and update_job.

Fixes #123281

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Co-authored-by: mochamgx <1114149@qq.com>
2026-09-26 22:04:17 +05:30
mochamgx
b653fd724a fix(cron): guard None repeat.completed when recording a job run
`repeat.get("completed", 0)` only falls back to 0 when the KEY is
absent. When the key exists with value None (a job created but never
successfully recorded a run), it returns None, and the following
`completed += 1` raises:

    unsupported operand type(s) for +=: 'NoneType' and 'int'

Effect: the run is never recorded (last_run_at/last_status stay null)
and every fire logs a scheduling error. The job itself executes fine.

Same latent hazard at the `completed < times` comparison, which a None
value would also fail on.

Fix all three sibling sites with `repeat.get("completed") or 0`.

(cherry picked from commit 0c5024de901900a747849ca44750327b56477f3b)
2026-09-26 22:04:17 +05:30
kshitijk4poor
2c54299eda fix(cron): log the jobs.json junk-skip warning once per repair
An unlocked load_jobs that finds junk entries hands off to a locked re-read,
which runs the same analysis again, so every repair logged the "Skipping N
non-object entries" warning twice. Log it only on the locked pass (the one
that actually saves the repair).
2026-09-26 22:04:17 +05:30
John Paul Soliva
a51d7f7bea fix(cron): run the load_jobs repair save under the job-store lock
An unlocked reader (list_jobs) that found a repairable store saved its parse-time snapshot. The shrink-merge restores only missing ids, so a locked writer's update to a job already in that snapshot (enabled, next_run_at, run claims) was reverted. Outside the lock, the repair now re-reads and saves under _jobs_lock(); this covers every repair kind, including the pre-existing bare-list, id-map and control-character repairs.

(cherry picked from commit 0f90a10212dbc360007a59bb09acc1e0ce676f61)
2026-09-26 22:04:17 +05:30
JoaoMarcos44
883f5422a1 fix(cron): repair a scalar 'jobs' field instead of returning it from load_jobs
{"jobs": null} (or a string, number or bool) escaped load_jobs unchanged, so every reader crashed the same way the non-object list entries did. Replace it with an empty list and persist the repair. Ported from #123405.

(cherry picked from commit 1f380ec62c2ce65fd1d31656850b611234e3385e)
2026-09-26 22:04:17 +05:30
John Paul Soliva
e66e351cf5 fix(cron): persist the junk-entry repair when no valid job remains; log types only
An all-invalid list filtered to [] skipped save_jobs (the 'if jobs and repair' guard), so every tick repeated the warning. Save whenever a repair ran (the shrink-merge still keeps concurrent valid jobs), and log the dropped entries' types instead of their raw values.

(cherry picked from commit 3b5b1aa61332cf79554fc2a24ae925908c37fe5d)
2026-09-26 22:04:17 +05:30
John Paul Soliva
695766e3f9 fix(cron): skip non-object entries in jobs.json instead of halting every tick
A null, string or number in the canonical {"jobs": [...]} list reached
every reader: the due scan raised AttributeError on each tick, so no job
fired, and list/resolve crashed too. Drop the junk with a warning and
self-heal the file, as the id-keyed map flatten already does.

(cherry picked from commit 2e2438d160967392184b2193cfc2e2ca124c2b7a)
2026-09-26 22:04:17 +05:30
kshitijk4poor
e9a01046eb fix(cron): label post-handoff external-worker failures distinctly
run_one_job's dispatch-failure handler already routes an
_ExternalWorkerPostHandoffError through bookkeeping only (no incident, no
ping), but the recorded error still read "Restart-safe cron worker dispatch
failed", which is misleading in last_error / executions.db for a worker that
was spawned and may have run. Compute post_handoff first and label it
"Restart-safe cron worker failed after handoff: ..."; the pre-handoff prefix
is unchanged.
2026-09-26 21:13:12 +05:30
kshitijk4poor
8c60444222 fix(cron): only alert on external-worker failures before the handoff
_launch_external_cron_worker returns through _wait_for_external_cron_worker
once the worker is spawned/acknowledged, and that waiter can raise (e.g. the
adopted worker died and recovery could not terminalize the row). Such an
error landed in the same run_one_job except branch as a genuine dispatch
failure, so a job that actually ran got a false "Restart-safe cron worker
dispatch failed" incident and failure ping on top of whatever the worker
itself delivered.

Wrap waiter failures in a dedicated _ExternalWorkerPostHandoffError and skip
_deliver_crash_failure for it, keeping the base bookkeeping-only behaviour
(mark_job_run + finish_execution, the latter a no-op on a worker-owned row).
Pre-handoff dispatch failures still open an incident and notify.
2026-09-26 21:13:12 +05:30
kshitijk4poor
fb32e96d07 fix(cron): keep finish_execution guaranteed when dispatch-failure delivery raises
The dispatch-failure branch called _deliver_crash_failure before the
try/finally that closes the execution row. If delivery ever raised,
finish_execution was skipped and the row stayed claimed until dead-owner
recovery. Run delivery inside the try so the pre-existing finally still
terminalizes the attempt; delivery_error/outcome default to None when it
raises.
2026-09-26 21:13:12 +05:30
kshitijk4poor
66bca15070 fix(cron): call _deliver_crash_failure unwrapped on dispatch failure
_deliver_crash_failure is already best-effort internally (incident upsert,
_deliver_result and _mark_incident_alerted each swallow their own errors),
and the sibling crash path in run_one_job calls it without a wrapper. The
extra try/except was a defensive layer that could only mask a genuine bug
in the notice helper; drop it so both call sites match. The existing
try/finally still guarantees finish_execution runs after mark_job_run.

webtecnica submitted an equivalent independent fix in #123433.

Fixes #123401

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-26 21:13:12 +05:30
liuhao1024
ecc5c263ec fix(cron): surface external-worker dispatch failures through the incident path
A failed restart-safe handoff in run_one_job() recorded the failure on the
job and in the executions ledger, then returned before any incident or
delivery path ran: no cron_incidents row, no failure-lane notice. Route the
dispatch-failure branch through _deliver_crash_failure() so it opens the
same job+signature incident and delivers the same failure notice as any
other job failure, with the existing alerted-cooldown withholding repeats.
A notice-path exception no longer loses the bookkeeping: mark_job_run and
finish_execution still run with a "failed" delivery outcome (#123401).

(cherry picked from commit e5b5969df1b7ca212c5a6e27d30f4778fb835c52)
2026-09-26 21:13:12 +05:30
kshitijk4poor
fecaf3afae refactor(cron): name the retry ladder applicability gate once
plan_retry's exhausted-ladder warning re-listed _ladder_instant's gates
(recurring, not paused, retry enabled) by hand, so the two could drift and
the warning fire for a job the ladder never applied to. Both now call
_ladder_applies(job). _ladder_instant checks the attempt count first, so the
exhausted path loads config exactly once (in the warning gate).

Behaviour unchanged: old-vs-new equivalence over 1,728 job states
(will_retry, plan_retry result, mutated job, log levels) shows 0 diffs.
2026-09-26 20:52:42 +05:30
kshitijk4poor
9de3cb2a0e refactor(cron): share one ladder decision between plan_retry and will_retry
WHY: will_retry hand-copied plan_retry's decision (recurring/paused, enabled,
ladder exhausted, next-rung instant, yield to the natural run) - the exact
drift that let fast jobs hold failure notices forever. Both now call a pure
_ladder_instant(job, natural_next, now); will_retry keeps only its own checks
(final finite repeat, uncomputable natural run) and reads the clock once.
The final-repeat check is spelled `times is not None and times > 0` as in
_advance_after_run. Behaviour is unchanged: an old-vs-new differential over
1728 job states matches on will_retry, plan_retry result, job state and logs.

Co-authored-by: Yuan Li <dskwelmcy@163.com>
2026-09-26 20:52:42 +05:30
kshitijk4poor
5693204f10 fix(cron): will_retry answers False on a finite repeat's final run
WHY: on the last run of a finite repeat, _advance_after_run completes the
job and mark_job_run skips plan_retry (is_terminal_job), but will_retry still
predicted a re-run, so that final failure notice was held forever. Mirror
the terminal guard. Folds in the edge reported by #109991 (Liuzikaii).
2026-09-26 20:52:42 +05:30
kshitijk4poor
27760e6715 refactor(cron): trim will_retry docstring and rename its yield test
WHY: the salvaged docstring narrated the incident and referenced another PR;
replace it with the invariant the predictor must hold. The test only covers
the yield branch, so drop "terminal_paths" from its name.
2026-09-26 20:52:42 +05:30
Yuan Li
a31e3a6b94 fix(cron): mirror plan_retry's yield branch in will_retry so fast jobs stop holding failure notices
will_retry gated notice suppression on recurring/paused/attempt/config only.
plan_retry has a yield branch: when the schedule's own next occurrence is at or
before the pending ladder rung it schedules nothing and clears state without
consuming an attempt. For a job on a cadence at or under a rung (<=5m, and the
15m/30m rungs for faster cadences) every failure hit that branch, the attempt
counter never advanced, and will_retry kept answering True — so during a
sustained outage every failure notice was held forever. The documented escape
('once the ladder is exhausted, the next failure alerts normally') was
unreachable: the ladder could never exhaust.

will_retry now recomputes the natural next occurrence exactly as
_advance_after_run will and answers True only when the rung precedes it — i.e.
exactly when plan_retry will actually park a re-run. Notices now go out on the
first failure for cadences the ladder cannot help, and slow jobs keep their
silent bounded re-runs.

(cherry picked from commit 9d3d6006204269103188a5ee366d8eb7c9f482a2)
2026-09-26 20:52:42 +05:30
kshitijk4poor
63cdaebf53 docs(cron): cut the committed-generation comment to the rule (#122183)
The 19-line block narrated review history and still carried the base claim
"rather than interpreting relocated pyvenv.cfg", which is now wrong: with no
generation committed the code falls through to the handed venv's pyvenv.cfg.
Four lines state the rule. Comment-only.
2026-09-26 19:57:55 +05:30
Halldrix
74c5832fe7 fix(cron): fall through to the handed venv when no generation is committed (#122183)
The no-generation arm returned the bare store Python with a repo-only
overlay, which is worse than the main behavior it replaced: a store
Python with nothing committed is refused by _require_own_dependencies
with "hermes pm repair", and a cron script has no hermes_bootstrap to
refuse cleanly, so it runs and dies on its first third-party import. It
also stripped the overlay's site-packages entry, so
_windows_cron_bootstrap_argv warned on every spawn.

Fall through to the uv overlay instead: the handed venv keeps its own
base interpreter and its own packages, which is what worked on main.
The committed-generation arm is unchanged and still returns the store
Python with that generation's site-packages.

Also corrects the comment: site_packages is a pure path join on Windows
(pm/environments.py:250-255) and never raises there; the raise is the
pyvenv.cfg probe inside _recorded_venv:192, already covered here, and
dependency_site() is outside this try. (#123668 review)

(cherry picked from commit b3760eff7a1ea8207b0f90c4c9cbfdbc6ab753a2)
2026-09-26 19:57:55 +05:30
Halldrix
afc450d178 fix(cron): read the committed generation for Windows cron scripts (#122183)
_windows_cron_python_invocation selected the dependency tree with
selected_venv, which falls back to base_venv and answers the leftover
pre-PM <root>/venv when no generation is recorded. That tree belongs to
whichever interpreter created it, so on a PM-managed install the managed
store Python 3.14 got a cp311 site-packages on PYTHONPATH and every cron
script died with "No module named 'pydantic_core._pydantic_core'" -- the
cron sibling of the gateway crash in #122183, which janviernine flagged
there and left out of scope.

committed_venv never answers with the in-tree venv. With nothing committed
this install provisioned no tree, so overlay none and let the child keep
the interpreter it was handed rather than borrowing a foreign ABI.

Every other selected_venv caller (tools/environments/local_pythonpath,
hermes_cli/_early_recovery, pm/extras, pm/environments_adopt) already gates
on runtime_facts_path().is_file(); this was the only unguarded one.

(cherry picked from commit dd2b7181ab0aec671ff1f8d0946e0d963cb50f52)
2026-09-26 19:57:55 +05:30
ethernet
43c105d5ba Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 07:39:33 -04:00
kshitijk4poor
348938f84e docs(cron): list the retry state's expr fingerprint 2026-09-24 16:59:23 +05:30
kshitijk4poor
3a5bf8c03d fix(cron): fingerprint the unreachable-model retry's schedule expression
A direct jobs.json expression edit while a retry was parked still fired once
at the old ladder instant; record the expression the retry was planned under,
as the quota-hold recovery fire does. Also drop an unreachable None guard in
_recovery_worthwhile (its only call is behind 'not blocked').
2026-09-24 16:59:23 +05:30
kshitijk4poor
3e3cfec49d fix(cron): fire a cron job's unreachable-model retry instead of re-anchoring it
unreachable_retry.plan_retry parks a cron job at now + ladder delay, an
instant that is off the expression's lattice. The stale-cron guard on the
due scan classified it as a direct schedule edit (STALE_CRON_EXPR_EDIT) and
re-anchored it to the natural occurrence without firing, so the 5/15/30-minute
ladder never ran for cron jobs — the same failure mode the quota-hold recovery
fire had to exempt itself from. Record the ladder instant in the retry state
and let _reanchor_stale_cron accept either planner's parked instant as an
authorized one-shot off-lattice fire; interval jobs and legacy state without
the instant keep the previous behaviour.
2026-09-24 16:59:23 +05:30