Commit Graph

841 Commits

Author SHA1 Message Date
teknium1
95b1f4c855 fix(cron): scope the yield predicate's claims to what the record proves
Review follow-ups on the stale-code tick yield gate:

- `_gate_mocks` patches `_current_gateway_code_sha` with `raising=False` so the
  base tree (where the symbol does not exist yet) fails only on the two new
  invariants instead of erroring the whole module out with AttributeError.
- Note at the pid comparison that `get_running_pid(pid_path=None)` can fall back
  to `get_runtime_status_running_pid()`, which derives the pid from the same
  record — the equality is tautological in that branch and the real proof there
  is `runtime_status_is_stale` + `code_sha`.
- Docstring says what the record actually is: `gateway_state.json` is per-HOME
  and last-writer-wins, the pid equality binds it to the lock holder, and the
  `--replace` takeover window fails open by design.
2026-09-20 20:13:19 -07:00
fangliquanflq
153b5f0145 fix(cron): compare full gateway code identities 2026-09-20 20:13:19 -07:00
fangliquanflq
bc04c35f6a fix(cron): verify fresh gateway before yielding ticks 2026-09-20 20:13:19 -07:00
teknium1
36ede56a66 fix(gateway): share the drift-tolerant start-time comparator across every owner-liveness check
Review follow-up for #117505: delivery_ledger, api_server_runs and
async_delegation reconciled a live owner to dead/unknown on the same 1 s
same-host fingerprint drift that cron/executions had just stopped doing.
One comparator (gateway.status.start_time_fingerprints_match) replaces the
private cron tolerance and the three exact-equality sites.
2026-09-20 16:33:19 -07:00
liuhao1024
52abb64c8f fix(cron): treat a drifted start-time fingerprint as a live owner in recovery
_owner_is_live() required the recomputed start-time fingerprint to be exactly
equal to the one recorded at claim time. On hosts where same-host readings
drift by ~1s (macOS kern.boottime adjustment), every long-running execution
was reconciled to unknown while its owner process was still alive and healthy.
A mismatch is now treated as inconclusive within a small tolerance (both
platform fingerprint scales are x100, so 200 means 2s), and an unreadable
reading is fail-safe, matching the function's existing posture that inability
to prove death must not rewrite durable state.
2026-09-20 16:33:19 -07:00
John Paul Soliva
786c0e3f9d fix(cron): a bot-chat delivery builds its child env for the DESTINATION profile, not the sender
`deliver: bot-chat:<other profile>` spawned the destination's agent turn with the sending
gateway's whole environment: its `.env` settings, bridged `TERMINAL_*` policy, platform
authorization gates and provider credentials. The lane called
`strip_launch_profile_env(env)` with no target, so the strip resolved against the ambient home
override — which is the SENDER's home, never the destination's. On an ordinary root-profile
gateway (`hermes gateway run`, no `-p`) `_is_routed_home` is then false and the strip is a
complete no-op, including the #113270 gate strip that lives after its early return.

This is the only cron child built for a profile other than the one whose tick spawned it; the
worker lane (`scheduler.py`) targets its own home, so its no-target call is correct. Build this
one through `served_profile_child_env(target_home=home, inherit_credentials=True)` — the helper
`kanban_db_dispatch` and `web_server_gateway` already use for cross-profile spawns: it strips the
launch residue against the real target, scrubs credentials the launch process was given by
systemd/Compose/the shell (which no name-based strip can see), points TMPDIR at the destination's
scratch, and overlays the destination's own secrets, as a standalone `hermes -p <profile>` has.

A failure to build that environment (an unreadable target home under per-user 0700, a broken
secret source) is reported as a refusal string like every other failure in this lane rather than
raised: `_deliver_result`'s fan-out does not catch, unlike the deferred drain.

Regressions drive the real `_deliver_to_bot_chat`; removing the fix fails the two leak witnesses
(`HERMES_MODEL leaked from the launch profile`, and the firing profile's gate reaching another
profile's turn under an active override) and leaves the four guard tests green.

Fixes #117220

(cherry picked from commit 6cc81d7ddad8c9793f21fdda7c2e260c3ee1ca44)
2026-09-20 15:33:28 -07:00
teknium1
6bbfd63ae4 fix(cron): drop the is_cron_silence_response alias; trim tests to three invariants
The sibling module already reaches the scheduler via the late-bound _sched
reference, so the public alias was a shim (shape gate: no re-export for
internal moves). Tests kept: answer survives a huge prompt, silent run falls
through to an older archive, headingless (script-mode) archive injects whole.
2026-09-20 15:31:21 -07:00
liuhao1024
ce4d0ee92d refactor(cron): expose is_cron_silence_response for the prompt module
scheduler_prompt reached into the private scheduler helper; a rename would
have broken it silently. Add a public module-level alias and call it
late-bound instead (review feedback).
2026-09-20 15:31:21 -07:00
liuhao1024
a11b77fe39 fix(cron): split archives on the last heading; match lane silence forms
Follow-up to the review on #117302:

- _archive_answer: rpartition — the assembled prompt half can itself carry
  a literal "## Response" heading (a skill documenting its response format,
  an injected previous answer quoting it), so the writer's boundary is the
  LAST occurrence; an early split re-injects the prompt noise this
  extraction exists to drop.
- silence check: reuse _sched._is_cron_silence_response instead of the
  literal [SILENT] comparison, so the bracketless SILENT / NO_REPLY forms
  the delivery lane itself suppresses (#46917, #51438) also fall through
  to an older archive.
- _SELF_CONTEXT_INTRO: "most recent non-silent output" — a silent latest
  run now deliberately resolves to an older archive.
2026-09-20 15:31:21 -07:00
liuhao1024
387e8e1e16 fix(cron): inject the previous run's answer, not its prompt
context_from head-truncation amputated the ## Response section of long
archives (skill-bearing prompts routinely exceed the 8000-char budget),
so self-continuity silently became a no-op: the job was handed its own
old prompt plus a truncation marker. Reduce each agent-run archive to
its answer (text after ## Response), fall through to an older archive
on [SILENT]/blank responses, inject nothing when no archive has a
usable answer, and clip oversized answers head+tail with an explicit
chars-omitted marker since conclusions sit at the end. Script-mode
archives without the heading keep injecting whole.

Fixes #117290
2026-09-20 15:31:21 -07:00
Christopher
eb2781dd2f fix(gateway): drain a restart-safe worker's cron delivery when it is queued, not on the next tick
A restart-safe cron worker runs outside the gateway, queues its final send in
cron/deliveries.db and polls that row once a second. The gateway drained the
queue only as a housekeeping chore at the 60 s tick, so a reminder whose turn
ended in 8 s reached the user a minute later (the issue's 60.86 s silence).

The housekeeping thread now watches the queue file (and its WAL) between
ticks and drains as soon as a stamp moved; the tick's drain stays as the
fallback. The watch never resets itself after draining: the drain's own status
write costs one empty pass, a worker that enqueued during the drain still
wakes it, and an empty pass writes nothing, so the stamp settles.

Fixes #117307
2026-09-20 15:30:45 -07:00
teknium1
27144dd59a fix: the wedged-cron drain check parses jobs.json once per run (review follow-up)
get_wedged_job_ids called cron.jobs.get_job (a full jobs.json load) per in-flight
job, and the restart drain invokes it on the event loop every 0.1 s for the whole
wait: 3 running jobs x 600 s = ~18,000 parses. Resolve each run's allowance once
(cached per job_id while it is registered, dropped on release) from a single
load_jobs() call.
2026-09-20 12:24:10 -07:00
teknium1
69266ecb42 fix(gateway): the restart drain skips a cron run wedged past its in-flight allowance
`_wedged_agent_count` only ever looked at chat agents, so a cron run that would
never finish (a no-agent job whose delivery hung on a dead transport) was
structurally un-skippable: `hermes update` sat in "draining" for the full
`agent.restart_after_turn_timeout` printing "0 wedged and excluded" while the
script had finished 8 seconds in.

Cron has no per-turn activity clock, but the scheduler already defines when an
in-flight claim can no longer be making progress: `sweep_stale_inflight`'s
`max(2 * interval, cron.inflight_max_minutes)` allowance. It cannot release a
claim whose worker thread is still alive, so expose that judgement as
`cron.scheduler.get_wedged_job_ids()` and let the drain count those runs as
wedged (restart is their remedy), the way it already treats idle chat turns.
`_describe_active_work` marks the cron unit `wedged` so the status line names it.

Fixes #115469 (Defect B; Defect A is the bounded standalone send this branch
stacks on).
2026-09-20 12:24:10 -07:00
John Paul Soliva
76a486a1eb fix(bot-relay): a relayed turn is booked from its turn report at the cap, not killed with its handoff
The relay's delivery child shared one 600s deadline between the target's turn
and the one-shot exit linger, whose own budget is the same 600s — so a turn that
answered in seconds and then handed off to a teammate was killed mid-linger,
reported to the sender as delivery_timeout (auto-retried: the turn ran twice),
its reply lost, and the handoff delivery the linger protected destroyed.

The -Q child's turn report (#113608) now carries the answer the run will print,
rewritten when a follow-up turn displaces it, and the poll loop that books a
child from that report moves next to the contract as
quiet_single_query.run_reported_turn. The cron lane keeps its policy (book after
a 2s exit grace); the relay waits for exit under the cap as before — a teammate's
reply during the linger may still become the printed answer — and at the cap
books a reported child from its latest report and leaves it to finish. Only a
turn that never ends is a timeout. The report is 0600 from creation now that it
carries the answer.

Fixes #114980
2026-09-20 11:55:07 -07:00
686f6c61
c89be4c52d fix(cron): spawn bot-chat delivery from a live cwd
The cron bot-chat delivery child inherited the scheduler's cwd; when that
directory had been removed (a kanban worker whose scratch workspace was reaped
by completion cleanup) the child died in `hermes_cli/_startup_fast.py::
ensure_project_root_on_path` — a relative `sys.path` entry goes through
`os.getcwd()` inside `realpath`, which raises FileNotFoundError — before it
could parse argv, and the finished job was booked `delivery_failed`.

- `_run_bot_chat_turn` pins the child's `cwd` to the target home the lane has
  already verified exists (`env["HERMES_HOME"]`).
- `ensure_project_root_on_path` resolves entries through a `realpath` that
  tolerates a gone cwd, so any `hermes` invocation from a dead directory still
  starts (the second half of #102941).

Salvaged from #102967 (@686f6c61), resolved onto the report-driven Popen lane
(#113608); tests drive the CLI entry point and the delivery spawn seam from a
deleted cwd.
2026-09-20 11:30:36 -07:00
teknium1
0469740ab3 feat(cron): jobs follow the main agent model at fire time; pinned locks it on request
An unpinned cron job used to snapshot the global provider/model at creation and treat that
snapshot as its effective pin (#44585), so `hermes model` / `/model` never moved the fleet and
`hermes cron resnap` existed to catch jobs up. New ruling: jobs run on whatever the main agent
model is when they fire. Resolution is per-job pin > cron.model / cron.model_provider (the cron
fleet default) > model.default.

`pinned` replaces the implicit snapshot with an explicit lock: create/update with pinned=true
writes the CURRENT main provider+model onto the job as an ordinary per-job pin; pinned=false
releases both. The cronjob tool exposes it (schema: only when the user asks; it can only lock
the main model, never point spend at a different one) and reports `pinned` per job; the CLI
gets `--pin` / `--unpin`. Legacy records that still carry *_snapshot keys follow the main model.

Removed with the snapshot: `hermes cron resnap`, the tool's resnap action + `all` param, the
"N unpinned jobs keep running on ..." notice in `hermes model` / `hermes config set` / the
dashboard model assignment, and the Desktop cron-model-impact card (setMainModelAssignment
keeps the expensive-model confirm flow in store/model-assignment.ts).

Live A/B (real store + run_job against a temp HERMES_HOME): main-model X -> Y, unpinned job
fires on X before, Y after; pinned job stays on X; unpin -> Y; legacy snapshot record -> Y.
2026-09-20 09:20:51 -07:00
kshitijk4poor
2d0867695d refactor(cron): one claim-TTL headroom constant shared by the job store and the ledger 2026-09-20 17:59:57 +05:30
kshitijk4poor
2eb40f71d8 fix(cron): derive the live-owner stale bound only when a live-owned row is seen
The bound reads the script timeout, which falls through to load_config() when
HERMES_CRON_SCRIPT_TIMEOUT is unset. Computing it eagerly on every reap put a
config load back on the idle gateway tick (CI: tests/cron/test_idle_tick_config_skip.py).
Resolve it on the first live-owned claimed/running row instead; an idle tick with no
such rows stays config-free.
2026-09-20 17:59:57 +05:30
kshitijk4poor
678c79eaa0 refactor(cron): reap throttle keyed by hermes_home_key 2026-09-20 17:59:57 +05:30
kshitijk4poor
77592f4ec2 refactor(cron): claim age from the aware timestamp only 2026-09-20 17:59:57 +05:30
kshitijk4poor
f4b4032211 docs(cron): reap docstrings describe the live-owner stale bound 2026-09-20 17:59:57 +05:30
kshitijk4poor
b552dc9a8c fix(cron): derive the live-owner stale-claim bound from the configured timeouts
The carrier used a hardcoded 7200 s wall-clock ceiling. HERMES_CRON_TIMEOUT is
an *inactivity* limit (cron/jobs.py, _run_agent_with_watchdog), so a healthy
agent job in the external-worker topology can legitimately run past any fixed
wall clock; marking its row `unknown` releases the gateway's running guard and
the next tick double-dispatches while the first worker is still running (its
later finish_execution is fenced out). The constant also ignored
HERMES_CRON_TIMEOUT=0 (unlimited) and HERMES_CRON_SCRIPT_TIMEOUT > 7200.

Mirror cron/jobs.py::_oneshot_run_claim_ttl_seconds instead:
stale_after = max(3 × HERMES_CRON_TIMEOUT, script timeout, 7200), reusing the
existing accessors (_cron_inactivity_seconds, _get_script_timeout). When the
inactivity timeout is 0/unlimited or not a finite positive number there is no
bound to derive, so live owners are skipped entirely (fail closed — today's
behaviour). Drop the `except (ValueError, TypeError): pass` swallow (rows are
always aware ISO from hermes_time.now), fix the stray whitespace in the
SELECT, record a distinct `error` reason for the wedged-owner case, and keep
the comment honest: the deadlocked worker process is not terminated, and
in-process rows (process_id == _PROCESS_ID) remain out of scope.
2026-09-20 17:59:57 +05:30
Yagna Vudathu
df753d4fc3 fix(cron): key the dead-owner reap throttle by profile home
`_last_dead_owner_reap_at` was a process-global float, so under
`gateway.multiplex_profiles` the first profile ticked in a cycle consumed
the 300 s reap window for every other profile: a stale claim in profile B
was not noticed until profile A's window elapsed. Key the throttle by the
resolved profile home so each profile gets its own dead-owner scan.

Carried from #115696 (whyyagswhy, first submitter for #115692); only the
throttle hunk and `test_reap_throttle_is_profile_scoped` are taken — the
warn-only stuck-claim helpers are superseded by the live-owner stale-claim
guard in this stack.
2026-09-20 17:59:57 +05:30
holny
086da1e072 fix(cron): wall-clock stale-claim guard for live-but-deadlocked workers (#115692)
A claimed/running execution whose owner process is alive but permanently
deadlocked (e.g. futex_wait after a route/proxy flip) passes _owner_is_live()
and is never reclaimed — the execution row stays 'running' forever and the
job rejects every subsequent fire with 'Job is already running'.

recover_interrupted_executions() now also checks the wall clock: a
claimed/running row older than STALE_RUNNING_CLAIM_TIMEOUT_SECONDS (7200 s,
covering the 3600 s script timeout + 600 s inactivity watchdog + margin)
is marked unknown even though the owner PID is alive.

Tests: stale claim recovered, recent claim NOT recovered.
2026-09-20 17:59:57 +05:30
kshitijk4poor
715f21ffd9 refactor(cron): degraded marker is a plain body wrapped by the standard cronjob header 2026-09-20 17:55:54 +05:30
kshitijk4poor
15a3691404 fix(cron): degraded marker uses the redacted job name 2026-09-20 17:55:54 +05:30
kshitijk4poor
4d65eca72b refactor(cron): read the bot-chat delivery timeout once per delivery 2026-09-20 17:55:54 +05:30
kshitijk4poor
fba9b930f7 fix(cron): drained degraded marker is posted without the generic cronjob header 2026-09-20 17:55:54 +05:30
kshitijk4poor
619370c0bd fix(cron): degraded-delivery marker is recognised by its record, not its text
The recursion guard for the timeout marker matched 'DELIVERY DEGRADED' in the
payload, so a job whose own output contained that phrase would silently lose its
notice. The deferred record now carries an explicit degraded flag and the guard
reads it. The marker's id is a fresh sha256 digest of '<key>:degraded' instead of
'<key>-degraded': deferred ids double as live-owner delivery ids, which
tools.bot_live_delivery._delivery_id requires to be 32-64 hex characters.

test_bot_chat_pending documents the deliberate extra turn: a timeout drains one
short marker turn, whose own timeout queues nothing more.
2026-09-20 17:55:54 +05:30
kshitijk4poor
bdd89b1ec8 refactor(cron): drop the fire-fence ERROR dedupe from cron/jobs.py
The 10-minute per-key complaint cache (_fire_fence_complaints /
_fence_timeout_log) only rate-limited the log line. A contender that fails
closed on the fire fence every tick is a real contention class (cross-process
`hermes cron run`, a second gateway, or the forced-release sweep evicting the
running-set entry mid-delivery); demoting repeats to DEBUG hides it for ten
minutes instead of fixing it, and the helper was a module-level appendage on
the jobs facade. cron/jobs.py is back to main; the test that poked the private
cache goes with it.
2026-09-20 17:55:54 +05:30
Emir Saffar
fa939860ca fix(cron): never lose a bot-chat alert to a delivery timeout
A bot-chat delivery child that times out is killed mid-turn. The killed
child cannot complete, but its turn may have persisted the query into the
session before the kill — so replaying the full payload risks a duplicate,
while staying silent loses the alert entirely (the message was only saved
to job output; the Bot Chat consumer never saw it).

- Timeout on the CLI fallback lane queues a short degraded-delivery marker
  through the existing deferred lane: it references `hermes cron runs`
  and carries a 280-char excerpt instead of the full payload, so a turn
  that persisted the query before the kill cannot duplicate content.
  Recursion-guarded: a marker that itself times out is never re-marked.
- `_run_bot_chat_turn`: a turn report landing in the kill window books
  the delivery instead of raising TimeoutExpired (delivered turns were
  misreported as timeouts).
- `jobs.py`: fire-fence timeout logs ONE ERROR per episode (600s
  interval; repeats log DEBUG, a successful acquire clears the episode).
  One delivery turn holding the fence for minutes previously produced an
  8-ERROR storm plus restore warnings from every failed tick.
2026-09-20 17:55:54 +05:30
liuhao1024
00ecaf7721 fix(cron): decode the Windows bot-chat lane as the UTF-8 the child writes
The bot-chat delivery child is guaranteed UTF-8 on Windows (hermes_cli
reconfigures its streams via hermes_bootstrap even under
PYTHONIOENCODING=cp1252), but the gateway parent is not started in UTF-8
mode — its env overlay sets only PYTHONIOENCODING, which rebinds the
parent's own streams but does not change the locale codec text=True picks.
The delivery pipes were therefore decoded with the ANSI code page (cp1252):
an accented reply came back mojibake'd, and bytes undefined in cp1252
killed the reader thread, losing the captured reply and the failure
diagnostic while the delivery still booked as delivered (#115894).

Pin encoding='utf-8' on the win32 branch only — on POSIX the child keeps
the locale codec, so the locale default stays correct there (#66566) —
mirroring the win32 branch of the no_agent script lane (#45099).
2026-09-20 00:06:46 -07:00
liuhao1024
8be901ac0f fix(cron): decode cron child output lossily (script + bot-chat delivery)
A stray non-UTF-8 byte in a cron child's stdout/stderr (e.g. a grandchild
sharing the pipe interleaving a partial multi-byte write) raised
UnicodeDecodeError in the decode step and failed a run that in fact
completed: the no_agent script lane died inside communicate() with the
output discarded, and the bot-chat delivery lane lost both the captured
reply and the failure diagnostic when the drain thread died (#105582).

Decode lossily on the POSIX side of both lanes: keep the platform-default
(locale) encoding — gating encoding= to win32 was deliberate (#66566) —
but relax errors= to 'replace', matching the lossy UTF-8 decode the win32
branch of the script lane already applies (#45099).
2026-09-20 00:06:46 -07:00
teknium1
729b1c3d56 fix(cron): bound the standalone send inside the coroutine so the thread fallback keeps it
Wrapping the coroutine in `asyncio.wait_for` outside `asyncio.run` left the
running-loop fallback (which closes the unstarted coroutine and retries in a
thread) with a never-awaited wait_for wrapper and an unbounded inner send.
Awaiting `wait_for` inside `_send` bounds every runner of the coroutine.

The invariant tests now drive the production entry point `_deliver_result`
with no live adapters (the standalone lane) instead of `_standalone_send`
directly, and the user-visible knob `cron.standalone_send_timeout_seconds`
is documented (#115469).
2026-09-19 23:50:50 -07:00
liuhao1024
c53eaf0d82 fix(cron): bound the standalone send lane with a wall-clock timeout
The standalone fallback ran _send_to_platform under a bare asyncio.run,
while the gateway-loop dispatch inside it awaits its future with a
deliberate no-timeout shield. The shield's comment assumes an outer
_run_async bound that never exists on this lane, so a mid-reconnect
transport pinned the run (and the restart drain behind it) indefinitely
(#115469). Bound the wait with asyncio.wait_for — default 60s, matching
the live lane's future.result(timeout=60); configurable via
cron.standalone_send_timeout_seconds.
2026-09-19 23:50:50 -07:00
keenvc
24cfa762ca fix(cron): publish worker ack file via atomic rename
_run_external_worker_payload wrote the acknowledgement JSON directly
into ack_path after creating it with O_CREAT|O_EXCL. That made an
empty 0-byte file visible to the scheduler's exists()-then-read
polling loop (every 50ms) before the JSON body was flushed, so the
poller would occasionally read an empty file and crash with
json.decoder.JSONDecodeError: Expecting value: line 1 column 1 (char 0)
(observed 3x on 2026-09-09 in production logs).

Fix: write the ack payload to a per-pid temp file in the same
directory, then os.replace() it into place. Readers now only ever
see ack_path fully absent or fully written.
2026-09-19 23:44:07 -07:00
whyyagswhy
ffae0825ef fix(cron): live-lane text send uses the already-authorized transport (#115656)
_live_send_text rebuilt DeliveryRouter over the plain target_adapters dict
and re-resolved, discarding the SharedRouteAdapters-authorized transport from
_resolve_target_transport. Under satellite config (own platforms.<p> block
disabled) the second resolution yields None and delivery falls through to the
credentialless standalone lane and fails.

DeliveryRouter._deliver_to_platform takes an optional transport; the live
lane passes t.transport. All router pipeline behavior (oversize cap, relay
home stamping, thread routing) is unchanged; other callers still resolve.

(cherry picked from commit 9a2a67647cd823fe9120891d863a6fc6981fec71)
2026-09-19 22:34:21 -07:00
teknium1
fc49f7619d fix(cron): a missing-credential preflight verdict names the profile and HERMES_HOME it read
The blocked_config reason for a missing provider credential now carries
"[profile '<name>', HERMES_HOME <path>]" — the home the scheduler actually
read auth.json/.env from — under the ticker's profile scope, so a
multiplexed satellite profile reports its own home, not the gateway's
launch home.

Why: #116213 reports an openai-codex cron job blocked with "No Codex
credentials stored" while an interactive session under "the same"
HERMES_HOME resolves the credential. A 5-shape x 5-scope live matrix
(singleton, expired+refreshable, pool-only, ~/.codex only, none; root,
named profile, root-only auth, multiplex default/named) on origin/main
and on the reporter's build 345cd2b0 shows interactive and cron
preflight agree in every cell — both call the same
resolve_runtime_provider ladder and read the same store. The remaining
explanation is a scheduler process reading a different home than the
shell (Docker HOME vs HERMES_HOME, a service unit without the shell's
env, a satellite profile), which the bare verdict could not reveal.
Naming the store the verdict judged makes that mismatch visible in the
one alert the user receives.

Part of #116213
2026-09-19 12:14:32 -07:00
teknium1
a8cc840372 chore: merge origin/main (resolve gateway/run.py) 2026-09-19 10:48:53 -07:00
teknium1
24f57317da fix(cron): drop a stale quota hold on schedule edit; scope the docs to the usage-probe case
update_job recomputes next_run_at from the new schedule but left quota_hold_until behind, so
the marker kept shielding a record that was no longer parked where it said. Clear it with the
schedule edit; the next fire re-parks (with a fresh notice) if the window is still closed.

cron.md claimed every 429 with a retry-after hint holds the job; only the provider-resolve
usage probe (a rate-limited AuthError) is classified. Say so.
2026-09-19 10:12:47 -07:00
teknium1
b680dcd5d4 fix(cron): hold fires through a provider's closed usage window instead of re-firing every tick
A quota-exhausted subscription provider answers every request with a 429 and an explicit
`retry after <N>s` (Codex: ~33 h). The scheduler walked the fallback chain, failed, and
re-fired the job on its normal cadence into the same closed window — one guaranteed failure
and one delivered alert per tick until the window reopened (#89376: ~460 failed runs and
alerts across four profiles in one exhaustion).

Why: the retry-after was available at the failure site (a rate-limited AuthError in the
RuntimeError's cause chain) but nothing consumed it, and the stale-error re-arm (#62002)
would have pulled any manually parked next_run_at back after one cadence anyway.

What: cron/quota_hold.py mirrors cron/unreachable_retry.py in the opposite direction —
run_job flags the failure with the provider's remaining seconds (read only from a
rate-limited AuthError, never from arbitrary text), the one failure alert says the job is
held, and mark_job_run parks next_run_at at the first scheduled occurrence after the window
while stamping quota_hold_until; _job_is_stale_error_recurring treats an active hold as
deliberate. Any run that reaches the model clears it. Preflight no longer mislabels a
rate-limited AuthError as "provider credential missing" (blocked_config) so the hold applies
with or without a fallback chain.

Persisted marker shape (quota_hold_until) and the hold direction follow PR #89395.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Fixes #89376
2026-09-19 10:12:47 -07:00
teknium1
cb062f7bd3 fix: WAF-block follow-ups — honest hint, cron action, anthropic debug dump (#24293)
- turn_recovery: the upstream_blocked hint no longer advertises model.default_headers,
  which _apply_user_default_headers skips on anthropic_messages/bedrock_converse.
- cron: an upstream_blocked run says to set a User-Agent via extra_headers or pin
  another provider, instead of 'run it again' — a retry never heals a firewall block.
- agent_runtime_helpers: in anthropic_messages the request dump read agent.client
  (None) and printed 'Authorization: Bearer None' with a /chat/completions URL; it now
  reads the live _anthropic_client key and uses /messages.
2026-09-19 09:57:21 -07:00
teknium1
d5d653c2f0 fix(cron): deliver=origin from an api_server session reaches the home channel instead of failing silently
A cron job created from an api_server turn captured origin.platform="api_server".
That adapter's send() is a stub (request/response only, supports_async_delivery=False),
so every fire ran fine (last_status=ok) and then recorded
"API server uses HTTP request/response, not send()" — nothing reached the creator and
nothing warned them at creation time (#69304).

Producer: _origin_from_env() stamps no origin when the session cannot receive async
delivery, so the job takes the existing no-origin path — home-channel fallback at fire
time and the creation-time "local-only" notice (its wording now names stateless HTTP API
sessions alongside CLI/TUI). Fire time: _resolve_origin() treats an already-stamped
api_server origin as missing so existing jobs.json entries fall back too.

Co-authored-by: Chen Jin <243284244@qq.com>
2026-09-19 09:54:20 -07:00
teknium1
adf33e3110 fix: surface the pre-agent provider fallback on the TUI/Desktop gateway and cron too
The same credential-resolution fallback the messaging gateway now announces also runs in
tui_gateway/server.py (_resolve_runtime_with_fallback, used_fallback=True only logged 'Primary
auth failed, falling back to ...') and in cron/scheduler.py (_resolve_job_runtime, 'fallback
resolved to ...'), with no user-visible notice on either surface.

- hermes_cli/fallback_config.pre_agent_fallback_notice: shared text builder (gateway/run.py now
  calls it) so the three pre-agent paths cannot drift in wording.
- TUI/Desktop: _resolve_agent_model_runtime tags the fallback runtime; _make_agent pops the tag
  and sets agent._pending_fallback_notice, emitted once via status_callback on the first reply.
- cron: _resolve_job_runtime tags the runtime; the setup pops it and run_job prepends the notice
  to a deliverable final_response ([SILENT] and the [CRON_FAILURE] first-line marker keep their
  contracts) — a cron agent has no status rail, so the report is the only place it can surface.

Tests drive the production entry points (_make_agent, run_job) with the resolver and AIAgent
stubbed; each goes red with its attach/prepend removed.
2026-09-19 01:27:29 -07:00
teknium1
2b94b0d40f fix(kanban): dispatcher blocks a card on the first terminal provider error
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).

Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.

`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.

No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.

Part of #114587

Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
2026-09-18 20:34:16 -07:00
teknium1
c6e3a77577 fix(cron): the ticker supervisor respawns only a ticker that crashed, not one that returned
SupervisedTickerThread (#111010) treated any ended thread as dead. An external provider's
start() (Chronos) arms remote one-shots and returns by design, so every housekeeping tick
logged "Cron ticker thread died without a stop request; restarting" at ERROR and re-ran
start() - a fresh recover_interrupted + NAS list/arm reconcile once a minute on every hosted
instance. Track whether the target escaped with an exception and respawn only then; the
built-in ticker returns normally only on stop_event.
2026-09-18 20:07:21 -07:00
teknium1
5d027bbd06 fix(cron): -f patterns must match the gateway cmdline, not any "hermes" substring; ledger names agent-issued kills
Review follow-up on #113667:

- _pattern_reaches_host_interpreter: when the `-f` head token is not an
  interpreter image, require the pattern to plausibly reach the gateway
  cmdline (hermes_cli / hermes+gateway tokens, mirroring Branch D). On
  the previous head `pkill -f 'hermes-polis/run.sh'` and
  `pkill -f my_hermes_bot.py` were hard-blocked although both are allowed
  on main and cannot match the gateway; both spellings are now negative
  controls in test_kill_forms_that_do_not_reach_the_gateway.
- lifecycle_ledger: the unclean-exit warning enumerated only OS causes
  (SIGKILL / OOM / VM death); an agent- or descendant-issued kill of the
  host interpreter leaves identical evidence and is now named in the
  cause family. One invariant test.
- test file: encoding="utf-8" on the bare write_text calls the footgun
  scanner flags.
2026-09-18 12:49:13 -07:00
teknium1
726bccd28d fix(cron): lifecycle guard blocks kills aimed at the gateway's own interpreter image
An agent-issued `taskkill /F /IM python.exe` (or `pkill -9 python3`, `killall python`,
`Stop-Process -Name python`, `taskkill /FI "IMAGENAME eq python.exe"`, `pgrep python | xargs
kill`) from inside the supervised gateway killed the gateway: every branch of
_GATEWAY_LIFECYCLE_PATTERN was anchored on a hermes/gateway token, and the supervised gateway is
literally a `python` process. Branch E is token-aware (not a line regex) so option values are
never read as targets, `-f` cmdline patterns are judged as patterns (`pkill -f 'python
my_script.py'` passes, `pkill -f 'python -m hermes_cli.main'` does not), and other image names
(`taskkill /F /IM agent-browser.exe`) stay killable. Numeric-PID kills stay out of scope: the
explicit PID / `proc_*` id is the ownership-scoped route the terminal rejection now names.

The guard never ran on the Windows Scheduled-Task topology either: the launcher exports only the
generalized HERMES_SUPERVISED_CHILD marker, which gateway/restart.py never read.
is_supervised_gateway_launch() reads it and gates the self-kill guards;
is_gateway_supervisor_process() deliberately keeps ignoring it because it also selects the
exit-75 restart route, which the task has no restart policy to honour (#113670).

Supersedes the narrow `/IM python.exe` regex from #113671 (kept for authorship); the Windows
spellings from #94379 (`hermes.exe gateway restart`, `taskkill`/`Stop-Process` on hermes-gateway
tokens) ride along.

Fixes #113667
2026-09-18 12:49:13 -07:00
hotragn
d966b34cc3 fix(security): gateway lifecycle guard recognises Windows command spellings
The guard that stops a supervised gateway from restarting, stopping or
uninstalling itself only knew POSIX spellings of those commands, so the
Windows spellings of the very same operations walked straight past it.

Branch A anchored on the bare CLI name, but on Windows the CLI is spelled
with an executable suffix, and npm-style shims install `hermes.cmd` and
`hermes.ps1` beside the `hermes.exe`. Every one of these was allowed
while the unsuffixed form was blocked:

    hermes.exe gateway restart      -> ALLOWED (now blocked)
    hermes.cmd gateway stop         -> ALLOWED (now blocked)
    hermes.bat gateway restart      -> ALLOWED (now blocked)
    hermes.ps1 gateway uninstall    -> ALLOWED (now blocked)
    C:\tools\hermes.exe gateway stop-> ALLOWED (now blocked)

Branch D had the same gap for process termination: `\bp?kill\b` cannot
reach inside `taskkill` (there is no word boundary between the two `k`s),
so `taskkill /F /IM hermes-gateway.exe` and `Stop-Process -Name
hermes-gateway` were allowed while `pkill -f hermes gateway` was blocked.

This is reachable. The guard is gated on `_is_supervised_gateway_process()`
— PID-file ownership — not on platform, and its callers (terminal_tool,
code_execution_tool, approval, cron) all run on Windows.

Service-control spellings (`net stop`, `sc stop`, `Stop-Service`) are
deliberately left out: they presuppose a service install this guard has
no evidence of, and guessing at one risks blocking unrelated services.
`C:/tools/hermes.exe` (forward slashes) also stays allowed — the `/` in
the Branch A lookbehind is the deliberate #77173 path false-positive
fix, and narrowing it is a separate decision.

19 new cases pin the caught spellings, the quoted/wrapped forms that
reach the same place through the tokenizing rescan, and that the suffix
does not widen the match — `hermes.exe gateway start` stays benign,
`my-hermes.exe` is still a different binary. 13 of the 19 fail without
this change.

(cherry picked from commit e81c42e4482cb72a93f6468d7571f98aa5fcb848)
2026-09-18 12:49:13 -07:00
KoNit-K
339ee5b9dc fix(cron): block taskkill gateway interpreter
(cherry picked from commit 72f033c5823af00176c3b730b8e8502ad9b2b772)
2026-09-18 12:49:13 -07:00