739 Commits

Author SHA1 Message Date
kshitijk4poor
6e69a8933a refactor(cron): simplify the interpreter salvage after review
- Refuse pythonw: it discards captured output, and the user-interpreter
  path skips the Windows helper that used to swap it for python.exe, so
  an agent job would go silently quiet.
- Drop the one-off key pop on clear: like workdir/monitor_script, a
  cleared interpreter is stored as null and _script_argv already treats
  it as unset.
- Drop an unused import in the no-agent interpreter test.
- Docstring says what the check is (name-based), parametrize ids.
2026-09-28 02:32:05 +05:30
kshitijk4poor
4450cc9d90 fix(cron): only accept a Python executable as a job interpreter
The lifecycle guard classifies a `.py` script as Python and skips its
shell reference walk, so `interpreter=/bin/bash` on a `.py` whose body is
`bash restart.sh` was created and then executed as shell. Refuse any
interpreter whose name (or symlink target's name) is not a Python image,
reusing the guard's own `_INTERPRETER_IMAGE_RE`.

One parametrized test covers every refusal (bare name, missing,
directory, non-executable, /bin/bash, python -> /bin/bash symlink); the
two run-path tests now use a `python3`-named wrapper.
2026-09-28 02:32:05 +05:30
M1racleShih
6f7cc7e74c feat(cron): allow Python scripts to use an external interpreter
Add an optional per-job `interpreter` field so a cron Python `script` /
`monitor_script` can run under a user-managed venv instead of Hermes' own
Python, letting scripts import packages the Hermes runtime does not carry
(#8714). Nothing is installed, frozen, or restored automatically.

- cron/jobs.py: persist + normalize the field (absent => record unchanged;
  empty string clears it on update).
- cron/scheduler_script.py: _resolve_cron_interpreter() validates the path
  at run time (absolute/~ required, regular file, executable on POSIX);
  _script_argv runs [interpreter, script] and skips the managed-store
  bootstrap/PYTHONPATH overlays, which exist for Hermes' own venv.
  Threaded through _run_job_script, the claim-heartbeat wrapper, the
  pre-run prompt path and monitor scripts.
- hermes_cli: --interpreter on `cron create` / `cron edit`; shown in
  details and `cron list`.
- tools/cronjob_tools.py: programmatic/CLI lane only, like model and
  reasoning_effort — absent from the model-facing schema.

Shell scripts (.sh/.bash) still always run under bash. Revives #8741.

Ported onto current main from #70500 (the scheduler moved to
cron/scheduler_script.py and the CLI/tool became table-driven since the
PR's base).

Co-authored-by: MestreY0d4-Uninter <241404605+MestreY0d4-Uninter@users.noreply.github.com>
2026-09-28 02:32:05 +05:30
kshitijk4poor
c642f42bc3 fix(cron): fail closed inside the merge peek; keep merge on mergeable repairs
Review fold for the corrupt-store refusal:

- Move the refusal into _unmerged_disk_jobs (after the #80703 stat-stamp
  fast path) instead of a separate pre-check in _save_jobs_unlocked. The
  pre-check parsed jobs.json on every save, doubling the parse and
  defeating the stamp fast path on the scheduler's per-fire saves. A stamp
  match already proves disk is the file load_jobs parsed cleanly, and the
  in-merge raise also covers the verify-after-stage re-peek (TOCTOU the
  pre-check left open). replace=True never reaches it.
- load_jobs' auto-repair uses replace=True only for the two shapes the
  peek cannot read (id-keyed map, non-list "jobs" field). Every other
  repair keeps the shrink-merge, so a sibling's create that lands during
  a repair under the degraded flock-timeout lock is preserved again
  (#80624), as it was before the refusal.
- fsync the directory atomic_replace actually renamed into (it resolves a
  symlinked jobs.json), matching utils._atomic_write; refresh the stale
  comments/docstrings and pin the refusal message in the test.

Co-authored-by: Ayushman Padhi <208280836+ayushmanpadhi@users.noreply.github.com>
2026-09-27 18:30:44 +05:30
kshitijk4poor
553388b320 fix(cron): refuse to overwrite a corrupt jobs store on save
A merging save_jobs() over an unreadable jobs.json treated the store as
empty (the non-repairing peek returns None and the shrink-merge skips it),
so any save that landed after corruption — e.g. behind a degraded-lock
sibling's non-atomic copy fallback — silently replaced every job on disk.
Fail closed instead: raise and leave the bytes untouched. replace=True
remains the explicit disaster-recovery rewrite, and load_jobs' own repair
(a locked full-store read, so its repaired list is authoritative, incl.
id-keyed maps the peek deliberately refuses to flatten) now uses it.

Also fsync the parent directory after the atomic rename so the published
store survives power loss, via the existing utils.fsync_directory.

Lock-timeout and EXDEV fallback policy are unchanged.

Co-authored-by: Ayushman Padhi <208280836+ayushmanpadhi@users.noreply.github.com>
2026-09-27 18:30:44 +05:30
kshitijk4poor
009699b8a0 fix(cron): a missing dependency interpreter fails the script run
_posix_cron_script_argv fell back to sys.executable when the selected venv's
interpreter was gone. On a managed-store install that is the bare store Python,
so the run logged a warning and then died with the same ModuleNotFoundError as
#123044. A half-deleted venv reaches this branch because _recorded_venv only
checks pyvenv.cfg. Raise instead; _run_job_script's try reports it as a failed
run naming the missing interpreter, matching how a broken PM record is handled.
The existing broken-selection test is parametrized over both causes.
2026-09-27 15:04:50 +05:30
kshitijk4poor
4d4aed0f98 fix(cron): POSIX script bootstrap keeps plain-run __main__ semantics
Review follow-up on the live-checkout bootstrap.

- runpy.run_path runs the script under a temporary __main__ that is
  swapped back once the body returns, so atexit handlers or non-daemon
  threads that pickle script-defined classes failed ("not found as
  __main__.Foo") where a plain `python script.py` works. The bootstrap
  now installs a real __main__ module (SourceFileLoader, __cached__ as
  in a plain run) and execs the compiled script in it. Tracebacks keep
  one bootstrap frame instead of runpy's three.
- Under -P / PYTHONSAFEPATH there is no cwd entry to replace, so the
  repo is prepended instead of overwriting a stdlib path.
- Interpreter via pm.environments.project_python (the existing
  venv_python(selected_venv()) helper).
- A missing venv interpreter still falls back to the caller's Python
  but now logs a warning naming it, like the Windows bootstrap fallback.
- Docstrings: the lazy-install gate's real reason (a script importing
  hermes_bootstrap could complete a source update and execv itself onto
  the bare store Python via sys.orig_argv); _script_argv no longer
  claims "else sys.executable".
2026-09-27 15:04:50 +05:30
kshitijk4poor
f4057533aa test(cron): keep script-run tests off the host's PM store
POSIX cron scripts now consult the managed store. Without isolation, a
test run from a PM-managed checkout (the standard ~/.hermes/hermes-agent)
reads the host's store and install records, trips the home I/O guard,
and would run scripts on the host's dependency venv. Point
HERMES_RUNTIME_DIR at an empty per-test dir for tests/cron.
2026-09-27 15:04:50 +05:30
kshitijk4poor
3cdb0c80a6 fix(cron): POSIX cron scripts import the live checkout; broken PM selection fails the run
Follow-up to the venv-interpreter change above.

- The selected venv resolves Hermes from its generation's workspace
  snapshot, which only a dependency change (uv.lock / extras / Python /
  plugins, pm/packages.py::expected_stamp) rebuilds. After a code-only
  update, scripts imported older Hermes code than the gateway runs. A
  `python -c` bootstrap now puts the live checkout right after the
  script's directory, in-process, so nothing is inherited by the
  script's children (no PYTHONPATH, #123440).
- POSIX dispatch moves into _script_argv (platform branch), so
  _windows_cron_python_invocation is Windows-only again and the
  PYTHONPATH-keyed bootstrap gate goes back to its original form.
- The interpreter path comes from pm.environments.venv_python.
- _script_argv now runs inside _run_job_script's try. PM record reads
  (store manifest, facts, selection) can raise ValueError/KeyError as
  well as RuntimeError; before, those escaped the runner and stranded
  the execution row, and on POSIX the local fallback handed the script
  to the bare store interpreter (the #123044 symptom). A broken
  selection is now a failed run with the PM error, per
  selected_venv's contract. This also covers the same pre-existing
  gap on the Windows committed_venv path.
- Tests: two invariant tests replace the three change-detector tests
  (live checkout beats a snapshot on sys.path, venv site-packages
  resolve, script-dir sys.path[0], no PYTHONPATH; broken selection
  fails the run instead of escaping).
2026-09-27 15:04:50 +05:30
liuhao1024
03156afc7b fix(cron): degrade to the caller's interpreter when selected_venv raises
_posix_cron_python_invocation called selected_venv(repo) unguarded; the
four RuntimeError cases it is documented to raise would escape through
_script_argv (which runs before _run_job_script's try), crash the tick,
and leave the execution row in running forever — the exact half-migrated
install the test docstring already claimed was supported. Mirror the
Windows bootstrap's degrade-don't-crash contract: warn and run on the
caller's interpreter. Adds the raising-selection case to the POSIX
invocation test (red on the previous head), a sealed-payload caveat on
the e2e test's hand-written .pth, and a pointer from the Windows
invocation docstring to its POSIX counterpart.

(cherry picked from commit 8d5b3b5e9727d612d91b2a7485e0f4bb245358f3)
2026-09-27 15:04:50 +05:30
liuhao1024
cb11c994f7 fix(cron): run POSIX .py job scripts on the selected venv interpreter
On managed-store installs the gateway's store Python carries the repo and
the managed site-packages only in-process; cron .py scripts spawned with
sys.executable re-resolve imports from scratch and die with
ModuleNotFoundError (#123044). Run them on the selected dependency venv's
interpreter instead — its pyvenv.cfg site-packages (editable installs
included) need no PYTHONPATH overlay, so children the script spawns never
inherit the store's paths and import CPython-3.14 extension modules on a
foreign interpreter (#123440). Lazy installs are disabled for script
children so hermes_bootstrap imports off the store-record venv cannot
republish launchers. The .pth bootstrap routing in _script_argv is
tightened to PYTHONPATH-carrying overlays, which the POSIX path no
longer produces.

(cherry picked from commit 2d6976c3d6d8b877ccf73672fda2e963025ec30b)
2026-09-27 15:04:50 +05:30
kshitijk4poor
7ca5cca50a fix(cron): an Infinity or negative repeat.completed no longer breaks load_jobs
06a495cc5b normalized non-int counters but caught only TypeError/ValueError.
json.loads turns a hand-edited Infinity / -Infinity / 1e999 into float inf,
int(inf) raises OverflowError, and that escaped load_jobs, so every job
(list, tick, mark_job_run, hermes cron list) failed, not just the bad one.
Catch OverflowError (-> 0), and also clamp a negative int count, which
granted extra runs. The docs sentence now says non-negative.
2026-09-26 23:00:54 +05:30
kshitijk4poor
06a495cc5b fix(cron): normalize any non-int repeat.completed, not only null
A hand-edited "completed": "2" still crashed every recorded run ("2" + 1),
and 1.0 was stored as 2.0 ("2.0/3"). load_jobs now coerces any non-int
counter to a non-negative int (0 when unparseable). Document the load-time
repair next to the direct-edit tip.

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-26 22:04:17 +05:30
kshitijk4poor
02ddb846a1 test(cron): pin that each jobs.json repair warning is logged exactly once
42d76f5611 moved the repair-detail warnings to the locked pass so an
unlocked load_jobs (which re-runs itself under the lock) no longer logs
them twice, but nothing pinned it: reverting that hunk left every test
green. The folded junk-entries test already builds the list-junk and
scalar 'jobs' shapes through an unlocked load_jobs, so count the records
there: one detail warning and one "Auto-repaired" line per shape. Reverting
42d76f5611's cron/jobs.py hunk (scalar warning doubles) or re-emitting the
notes before the locked re-dispatch (list-junk warning doubles) now fails.
The raw-value check moves into the loop so it covers the list case after
caplog.clear().
2026-09-26 22:04:17 +05:30
kshitijk4poor
dbee807a07 test(cron): fold the unlocked-reader race case into the jobs.json junk test
The stack added three test functions and the budget is two. The race test
and the junk-entry test use the same fixture: create a job, then append
junk to jobs.json. Run the race on the junk test's first unlocked
list_jobs repair instead. It races a name update, which does not change
due-ness, and asserts the update survives on disk. The due-scan and
scalar-shape assertions are unchanged. It still fails if the locked
re-dispatch is removed from load_jobs.
2026-09-26 22:04:17 +05:30
kshitijk4poor
ca7eded74c fix(cron): treat null repeat.completed as 0 in update_job and job definitions
#122114 guarded mark_job_run/claim paths against a hand-edited
"completed": null, but update_job and the job_definition merge still used
.get("completed", 0), which returns None when the key is present, so the
null was carried forward into the store. Use `or 0` there too, and pin the
behaviour with one test covering mark_job_run and update_job.

Fixes #123281

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
Co-authored-by: mochamgx <1114149@qq.com>
2026-09-26 22:04:17 +05:30
kshitijk4poor
71e764b4c4 test(cron): fold jobs.json junk-shape cases into one invariant test
The salvaged PR added four tests (one parametrized) for the load_jobs
boundary. Keep the stack to its invariant-test budget: the all-junk list and
scalar 'jobs' field cases now run inside the junk-entry test, alongside the
separate concurrent-update race test.
2026-09-26 22:04:17 +05:30
John Paul Soliva
a51d7f7bea fix(cron): run the load_jobs repair save under the job-store lock
An unlocked reader (list_jobs) that found a repairable store saved its parse-time snapshot. The shrink-merge restores only missing ids, so a locked writer's update to a job already in that snapshot (enabled, next_run_at, run claims) was reverted. Outside the lock, the repair now re-reads and saves under _jobs_lock(); this covers every repair kind, including the pre-existing bare-list, id-map and control-character repairs.

(cherry picked from commit 0f90a10212dbc360007a59bb09acc1e0ce676f61)
2026-09-26 22:04:17 +05:30
JoaoMarcos44
883f5422a1 fix(cron): repair a scalar 'jobs' field instead of returning it from load_jobs
{"jobs": null} (or a string, number or bool) escaped load_jobs unchanged, so every reader crashed the same way the non-object list entries did. Replace it with an empty list and persist the repair. Ported from #123405.

(cherry picked from commit 1f380ec62c2ce65fd1d31656850b611234e3385e)
2026-09-26 22:04:17 +05:30
John Paul Soliva
e66e351cf5 fix(cron): persist the junk-entry repair when no valid job remains; log types only
An all-invalid list filtered to [] skipped save_jobs (the 'if jobs and repair' guard), so every tick repeated the warning. Save whenever a repair ran (the shrink-merge still keeps concurrent valid jobs), and log the dropped entries' types instead of their raw values.

(cherry picked from commit 3b5b1aa61332cf79554fc2a24ae925908c37fe5d)
2026-09-26 22:04:17 +05:30
John Paul Soliva
695766e3f9 fix(cron): skip non-object entries in jobs.json instead of halting every tick
A null, string or number in the canonical {"jobs": [...]} list reached
every reader: the due scan raised AttributeError on each tick, so no job
fired, and list/resolve crashed too. Drop the junk with a warning and
self-heal the file, as the id-keyed map flatten already does.

(cherry picked from commit 2e2438d160967392184b2193cfc2e2ca124c2b7a)
2026-09-26 22:04:17 +05:30
kshitijk4poor
ffe5cf049d test(cron): pin post-handoff waiter failure as bookkeeping-only
The fold that stops post-handoff waiter failures from raising a false
"dispatch failed" incident/ping had no regression coverage. Drive
run_one_job through a real _wait_for_external_cron_worker whose body raises
and assert: no cron incident, no delivery, one failed mark_job_run carrying
the post-handoff label, and the execution row terminalized as failed.
Red with the post-handoff routing reverted (an incident is opened).
2026-09-26 21:13:12 +05:30
liuhao1024
ecc5c263ec fix(cron): surface external-worker dispatch failures through the incident path
A failed restart-safe handoff in run_one_job() recorded the failure on the
job and in the executions ledger, then returned before any incident or
delivery path ran: no cron_incidents row, no failure-lane notice. Route the
dispatch-failure branch through _deliver_crash_failure() so it opens the
same job+signature incident and delivers the same failure notice as any
other job failure, with the existing alerted-cooldown withholding repeats.
A notice-path exception no longer loses the bookkeeping: mark_job_run and
finish_execution still run with a "failed" delivery outcome (#123401).

(cherry picked from commit e5b5969df1b7ca212c5a6e27d30f4778fb835c52)
2026-09-26 21:13:12 +05:30
kshitijk4poor
cb399b56d9 test(cron): say which failure's notice the yield case releases
The comment claimed the notice goes out immediately. In production
will_retry runs before mark_job_run, so the first failure (5m rung beats
the 10m run) is held; only the attempt-1 failure asserted here yields.
2026-09-26 20:52:42 +05:30
kshitijk4poor
a7131a73d9 test(cron): trim the will_retry yield test docstring
WHY: the ten-line docstring restated the commit narrative. Keep the
invariant and note why calling will_retry after mark_job_run is valid: the
predictor reads only persisted job state.
2026-09-26 20:52:42 +05:30
kshitijk4poor
5693204f10 fix(cron): will_retry answers False on a finite repeat's final run
WHY: on the last run of a finite repeat, _advance_after_run completes the
job and mark_job_run skips plan_retry (is_terminal_job), but will_retry still
predicted a re-run, so that final failure notice was held forever. Mirror
the terminal guard. Folds in the edge reported by #109991 (Liuzikaii).
2026-09-26 20:52:42 +05:30
kshitijk4poor
27760e6715 refactor(cron): trim will_retry docstring and rename its yield test
WHY: the salvaged docstring narrated the incident and referenced another PR;
replace it with the invariant the predictor must hold. The test only covers
the yield branch, so drop "terminal_paths" from its name.
2026-09-26 20:52:42 +05:30
Yuan Li
a31e3a6b94 fix(cron): mirror plan_retry's yield branch in will_retry so fast jobs stop holding failure notices
will_retry gated notice suppression on recurring/paused/attempt/config only.
plan_retry has a yield branch: when the schedule's own next occurrence is at or
before the pending ladder rung it schedules nothing and clears state without
consuming an attempt. For a job on a cadence at or under a rung (<=5m, and the
15m/30m rungs for faster cadences) every failure hit that branch, the attempt
counter never advanced, and will_retry kept answering True — so during a
sustained outage every failure notice was held forever. The documented escape
('once the ladder is exhausted, the next failure alerts normally') was
unreachable: the ladder could never exhaust.

will_retry now recomputes the natural next occurrence exactly as
_advance_after_run will and answers True only when the rung precedes it — i.e.
exactly when plan_retry will actually park a re-run. Notices now go out on the
first failure for cadences the ladder cannot help, and slow jobs keep their
silent bounded re-runs.

(cherry picked from commit 9d3d6006204269103188a5ee366d8eb7c9f482a2)
2026-09-26 20:52:42 +05:30
kshitijk4poor
86f449b172 test(cron,gateway): drop dead Windows venv setup; say why selected_venv is patched (#122183)
Neither kept test reaches a pyvenv.cfg read: the committed-generation path
returns before the cron fall-through, and the gateway overlay never parses
it. The `version=` arg, base/home dirs and the legacy cfg content were setup
nothing consumed. The cron test's selected_venv patch looks inert on head but
is what makes it red on base, so it now says so. Both tests still pass on head
and fail with base prod files (fake-win32 harness).
2026-09-26 19:57:55 +05:30
kshitijk4poor
dd39f83b07 test(cron): keep only the committed-generation invariant (#122183)
The salvage drops the try/except wrapper around committed_venv (a corrupt
facts.json raises as it did on main), so the "unreadable record falls
through" arm no longer describes the code. Keep the one invariant that
pins the bug: a committed generation, never the stale in-tree venv, is
what the cron child gets. Stack budget is two invariant tests.

Co-authored-by: Halldrix <halldrix@users.noreply.github.com>
2026-09-26 19:57:55 +05:30
Halldrix
6d44d9bd2f test(cron): two invariants plus the Windows lane marker (#122183)
platforms("windows") instead of a bare skipif: list_os_marked_tests.py
gates the Windows lane on that marker, so the skipif version ran on no
CI lane at all -- the antipattern root AGENTS.md calls out.

Four tests collapse to the two invariants that actually matter: the
committed generation wins, and no-generation/unreadable-record both
fall through to the handed venv. The legacy non-managed test is gone --
it passes on main by design, and the fall-through test's positive
control now covers that path inside the managed arm.

(cherry picked from commit 30c649f16bce08eeaabd5a407ae0429cf422556a)
2026-09-26 19:57:55 +05:30
kshitijk4poor
82ca84c52d fix(agent): keep openai-codex stale floor/hard cap for inline cron Codex (#69734)
Cron Codex now runs inline via direct_api_call, whose stale budget skipped the
openai-codex large-context floor (600/900/1200s) and HERMES_CODEX_HARD_TIMEOUT_SECONDS
cap applied on the worker path, so healthy >10k-token cron turns were killed at 90s.
Extract the floor+cap into _bound_openai_codex_stale_timeout, used by both paths.
Correct the should_use_direct_api_call docstring (Codex streaming goes through
_stream_codex_passthrough -> _interruptible_api_call, not _StreamingCall) and document
that the worker-only TTFB/idle watchdogs don't run inline. Replace the non-guarding
inline Codex watchdog test with one asserting the >=600s budget.
2026-09-25 21:28:32 +05:30
Ali Ahmed
0213b8b77b fix(agent): keep Codex cron calls inline
Cron Codex Responses calls still went through the spawned interrupt worker
(the #62151 nested-pool deadlock path). Route them through direct_api_call;
the Codex dispatch builds its client via make_client, so the inline stale
watchdog aborts a silent Codex stream even though the worker-only TTFB/idle
watchdogs are bypassed. Delegated children stay chat_completions-only.

Ported onto main's _InlineRequest refactor from PR #70087 (cherry picked
from e0028a5d73). Fixes #69734.
2026-09-25 21:28:32 +05:30
ethernet
50cb6807e4 test(cron): publish the descendant pid atomically
The parent polls for the file to exist, and write_text creates it before
writing, so it could read '' (int() ValueError on the detached-timeout row).
2026-09-24 12:47:41 -04:00
ethernet
43c105d5ba Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 07:39:33 -04:00
kshitijk4poor
3a5bf8c03d fix(cron): fingerprint the unreachable-model retry's schedule expression
A direct jobs.json expression edit while a retry was parked still fired once
at the old ladder instant; record the expression the retry was planned under,
as the quota-hold recovery fire does. Also drop an unreachable None guard in
_recovery_worthwhile (its only call is behind 'not blocked').
2026-09-24 16:59:23 +05:30
kshitijk4poor
3e3cfec49d fix(cron): fire a cron job's unreachable-model retry instead of re-anchoring it
unreachable_retry.plan_retry parks a cron job at now + ladder delay, an
instant that is off the expression's lattice. The stale-cron guard on the
due scan classified it as a direct schedule edit (STALE_CRON_EXPR_EDIT) and
re-anchored it to the natural occurrence without firing, so the 5/15/30-minute
ladder never ran for cron jobs — the same failure mode the quota-hold recovery
fire had to exempt itself from. Record the ladder instant in the retry state
and let _reanchor_stale_cron accept either planner's parked instant as an
authorized one-shot off-lattice fire; interval jobs and legacy state without
the instant keep the previous behaviour.
2026-09-24 16:59:23 +05:30
kshitijk4poor
da46661e4a refactor(cron): simplify quota-hold recovery plumbing
_recovery_worthwhile subtracted same-tz aware datetimes directly, which is
wall-clock arithmetic and off by an hour when the hold-to-occurrence span
crosses a DST transition (weekly jobs make that reachable); use the module's
canonical _elapsed_seconds. Derive the schedule from the job and gate the
"scheduled fire" boolean at the call site under one name
(recover_consumed_fire) from scheduler -> mark_job_run -> plan_hold instead
of three. Collapse plan_hold's interval/cron arms, which duplicated the
not-blocked early return and the boundary park. Persist the expr fingerprint
only for the off-lattice recovery fire: the coalesce branch parks on the
lattice (STALE_CRON_MATCH, also under the croniter-missing fallback), so
is_recovery_fire is never consulted there. Drop a comment restating the
docstring.
2026-09-24 16:59:23 +05:30
kshitijk4poor
ce1ad9addf test(cron): replace hold-policy change detector with dense/repeat invariant
test_hold_policy_coalesces_dense_cron_without_accelerating_sparse_interval
passed on main before the PR (it only re-asserted existing behaviour), so
it detected nothing. Replace it with the two invariants the bounded
recovery adds: a dense hourly schedule keeps its natural :00 instead of an
off-lattice near-duplicate, and a job already carrying quota_hold_until
that fails again is not re-parked.

Co-authored-by: Marko Niskala <manis@aivelho.local>
2026-09-24 16:59:23 +05:30
Marko Niskala
ac10d34f61 fix(cron): recover sparse fires after quota cooldown
(cherry picked from commit 2ffdca938fb73da66d949a0685fa1f827b9088b7)
2026-09-24 16:59:23 +05:30
ethernet
97de4b4ef8 Merge origin/main into ethie/pm-clean
Conflict resolutions and semantic fixups:

- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
  double-quoted scalar is never folded after an escaped backslash. pm-clean builds
  every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
  (ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
  GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
  read-only probe (source_git_env) now refuses promisor lazy fetches, and the
  partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
  Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
  PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
  --include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
  restore joins the darwin branch of the existing after-pack.mjs, and its test
  loads the hook from electron-builder.config.cjs and imports PlatformPackager
  from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
  win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
  pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
2026-09-23 19:51:34 -04:00
brooklyn!
58694737e2 fix(time): repair surrogate-bearing locale zone names at every strftime site
On Windows (fr-FR, es-AR, de-DE reports) the zone name arrives in the ANSI
code page but is decoded under a UTF-8 LC_CTYPE (UTF-8 mode, or Piper/espeak
flipping the process locale mid-run) with surrogateescape. datetime.strftime
splices tzname() in as UTF-8, so "%Z" raised UnicodeEncodeError while the
system prompt was being built and every new/compressed conversation died.

Rework safe_strftime into a small repair: output is untouched for valid text
(the system prompt stays byte-identical), surrogateescape'd bytes decode back
through the ANSI code page ("heure d'été"), anything else degrades to U+FFFD,
and a raising "%Z" is rendered from the repaired tzname(). hermes_time no
longer imports agent.* at module load.

Route the remaining locale-name sites through it: cron quota-hold notice,
auxiliary cooldown notice, cron session titles, session_search dates,
insights, learning graph and billing renew dates. Tests use real datetimes
with a surrogate zone name instead of stubbed strftime.

Fixes #102910

Co-authored-by: Aniruddha Adak <127435065+aniruddhaadak80@users.noreply.github.com>
2026-09-23 18:29:35 -05:00
fangliquan
85cb540e6d fix(time): tolerate Windows locale strftime errors 2026-09-23 18:29:35 -05:00
ethernet
16652eea18 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	gateway/config.py
#	gateway/config_loader.py
#	gateway/readiness.py
#	hermes_cli/managed_scope.py
#	hermes_cli/plugin_python_deps.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/update_cmd_maint.py
#	plugin-catalog/hindsight.yaml
#	plugins/plugin_loader.py
#	providers/__init__.py
#	scripts/run_tests.sh
#	tests/gateway/test_control_socket_windows_live.py
#	tests/gateway/test_gateway_streaming_nested_config.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_plan_reconciliation_windows_live.py
#	tests/hermes_cli/test_update_apply_shallow_count.py
#	tests/hermes_cli/test_update_concurrent_quarantine.py
#	tests/hermes_cli/test_update_shim_self_lock.py
#	tests/hermes_cli/test_verify_console_scripts.py
#	tests/tools/test_lazy_deps.py
#	tests/tui_gateway/test_subprocess_encoding.py
#	tools/lazy_deps.py
2026-09-23 15:26:34 -04:00
teknium1
3cf26c82b4 test(cron): keep two DST fall-back invariants (due gate, real elapsed timers)
Trim the salvaged fold regressions to one due-gate case and one elapsed-time case covering
claim age and the unreachable re-run delay; both are red on origin/main. The broader
virtual-clock soak lands with the delivery E2E suite.
2026-09-23 10:37:15 -07:00
Teo | Nexcore
e3618cdec2 fix(cron): compare due instants across DST folds
Use UTC-normalized comparisons and elapsed durations across due scans, claim ages, recovery, and the external misfire backstop. Add deterministic fall-back regressions for both folds and ordinary intervals.

Co-authored-by: Jaimin <95100522+Jaiminp007@users.noreply.github.com>
2026-09-23 10:37:15 -07:00
teknium1
3f2c86d627 test: restore cron and session-store guards dropped by #120071
- test_cron_bot_chat_delivery::test_deliver_failure_banner_only_stdout_names_exit_code_not_banner:
  a banner-only failed delivery records the exit code, not the resume banner (#104056).
- test_cron_script::TestScriptTimeoutTreeKill (3): pid 0 never reaches
  kill_process_tree (os.kill(0) would kill the scheduler's own group); a
  tree-kill that signals nothing still falls back to group termination; an
  already-exited script is never signalled (pid reuse).
- test_oneshot_guard_warning::test_guard_warns_on_rearmed_consumed_record:
  removal of a re-armed consumed one-shot logs a WARNING naming `cron resume` (#93524).
- test_hermes_state_conn_lock_audit: repo-wide AST lint, every self._conn call
  outside construction holds self._lock (SIGSEGV race with close(), #99349).
- test_pattern_b_scaling::test_statement_count_does_not_scale_with_sessions:
  list_sessions_rich SQL count stays within the N+1 budget (#95380); repaired to
  trace pooled read connections and fail if the trace sees nothing.
2026-09-23 10:34:54 -07:00
teknium1
6dc6fb593a test(cron): own-profile bot-chat turn from a multiplexed tick runs in the ticking home
Regression guard for #119858. The fix itself is already on main: 3b0fe0cc2b pinned
the child HERMES_HOME to the override-aware source home and 786c0e3f9d moved the
spawn onto served_profile_child_env(target_home=home, inherit_credentials=True).
No existing test asserted the OWN-profile (no -p) leg A->B->A under multiplex; these
two do (red at 786c0e3f9dc~1, green on main).
2026-09-23 08:09:47 -07:00
ethernet
3176602021 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-23 10:48:07 -04:00
Austin Pickett
e7bff4b6d8 fix(cron): a pinned job never falls back to the global fallback chain (#120312)
* refactor(fallback): share the pinned-owner chain rule

delegate_task's _resolve_child_fallback_chain decides which fallback chain
a child may walk: a pinned child never borrows the parent chain, an explicit
[] disables fallback, a declared list is the child's own. Cron needs the
same rule for pinned jobs (#100437), so the body moves to
hermes_cli.fallback_config.scoped_fallback_chain and the delegation helper
becomes a thin caller. Behaviour is unchanged; the delegation matrix test
still pins every cell.

* fix(cron): a pinned job never falls back to the global chain

A job with its own provider, model or base_url is an explicit operator pin
(since 0469740ab3 unpinned jobs store none of these). It still walked the
global fallback_providers chain in two places, so a pinned job could run
on a different provider and model than the one chosen:

- _resolve_job_runtime walked the chain on an AuthError or transient
  network failure while resolving the pinned primary;
- _resolve_cron_agent_setup handed the global chain to every cron agent as
  fallback_model, so the conversation loop's provider ladder could swap a
  pinned job mid-run.

Both now read _job_fallback_chain(job, cfg), which returns no chain for a
pinned job through the same scoped_fallback_chain rule delegate_task uses
for pinned children. The pre-dispatch key check reads it too: the global
chain used to skip that check for every job, so a pinned job with a
missing key now blocks before the agent is built instead of failing in the
resolver. The transient-failure notice for a pinned job says it does not
fall back and names --unpin, instead of "No backup provider succeeded".

Unpinned jobs (including legacy *_snapshot records) and same-provider
credential-pool rotation are unchanged. The two scheduler tests that
asserted atomic provider+model fallback swaps used pinned jobs; they now
use unpinned jobs and keep the same assertions.

No per-job fallback_providers list: jobs have no generic override field
(create_job/update_job, the cronjob tool schema and the CLI enumerate each
field), so an opt-in chain would be a new surface on all of them. The
escape hatch is to leave the job unpinned and pick its model with
cron.model / cron.model_provider.

Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>

* docs(cron): pinned jobs do not use fallback_providers

cron.md "Provider recovery" and the pre-dispatch key check, the cron rows
and section in fallback-providers.md, and the developer notes in
cron-internals.md / provider-runtime.md said every cron job inherits the
global chain. State the new rule, the compatibility note for users who
relied on a pinned job landing on the chain, and the unpinned + cron.model
alternative.

---------

Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
2026-09-23 10:42:18 -04:00