Commit Graph

797 Commits

Author SHA1 Message Date
teknium1
2b94b0d40f fix(kanban): dispatcher blocks a card on the first terminal provider error
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).

Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.

`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.

No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.

Part of #114587

Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
2026-09-18 20:34:16 -07:00
teknium1
c6e3a77577 fix(cron): the ticker supervisor respawns only a ticker that crashed, not one that returned
SupervisedTickerThread (#111010) treated any ended thread as dead. An external provider's
start() (Chronos) arms remote one-shots and returns by design, so every housekeeping tick
logged "Cron ticker thread died without a stop request; restarting" at ERROR and re-ran
start() - a fresh recover_interrupted + NAS list/arm reconcile once a minute on every hosted
instance. Track whether the target escaped with an exception and respawn only then; the
built-in ticker returns normally only on stop_event.
2026-09-18 20:07:21 -07:00
teknium1
5d027bbd06 fix(cron): -f patterns must match the gateway cmdline, not any "hermes" substring; ledger names agent-issued kills
Review follow-up on #113667:

- _pattern_reaches_host_interpreter: when the `-f` head token is not an
  interpreter image, require the pattern to plausibly reach the gateway
  cmdline (hermes_cli / hermes+gateway tokens, mirroring Branch D). On
  the previous head `pkill -f 'hermes-polis/run.sh'` and
  `pkill -f my_hermes_bot.py` were hard-blocked although both are allowed
  on main and cannot match the gateway; both spellings are now negative
  controls in test_kill_forms_that_do_not_reach_the_gateway.
- lifecycle_ledger: the unclean-exit warning enumerated only OS causes
  (SIGKILL / OOM / VM death); an agent- or descendant-issued kill of the
  host interpreter leaves identical evidence and is now named in the
  cause family. One invariant test.
- test file: encoding="utf-8" on the bare write_text calls the footgun
  scanner flags.
2026-09-18 12:49:13 -07:00
teknium1
726bccd28d fix(cron): lifecycle guard blocks kills aimed at the gateway's own interpreter image
An agent-issued `taskkill /F /IM python.exe` (or `pkill -9 python3`, `killall python`,
`Stop-Process -Name python`, `taskkill /FI "IMAGENAME eq python.exe"`, `pgrep python | xargs
kill`) from inside the supervised gateway killed the gateway: every branch of
_GATEWAY_LIFECYCLE_PATTERN was anchored on a hermes/gateway token, and the supervised gateway is
literally a `python` process. Branch E is token-aware (not a line regex) so option values are
never read as targets, `-f` cmdline patterns are judged as patterns (`pkill -f 'python
my_script.py'` passes, `pkill -f 'python -m hermes_cli.main'` does not), and other image names
(`taskkill /F /IM agent-browser.exe`) stay killable. Numeric-PID kills stay out of scope: the
explicit PID / `proc_*` id is the ownership-scoped route the terminal rejection now names.

The guard never ran on the Windows Scheduled-Task topology either: the launcher exports only the
generalized HERMES_SUPERVISED_CHILD marker, which gateway/restart.py never read.
is_supervised_gateway_launch() reads it and gates the self-kill guards;
is_gateway_supervisor_process() deliberately keeps ignoring it because it also selects the
exit-75 restart route, which the task has no restart policy to honour (#113670).

Supersedes the narrow `/IM python.exe` regex from #113671 (kept for authorship); the Windows
spellings from #94379 (`hermes.exe gateway restart`, `taskkill`/`Stop-Process` on hermes-gateway
tokens) ride along.

Fixes #113667
2026-09-18 12:49:13 -07:00
hotragn
d966b34cc3 fix(security): gateway lifecycle guard recognises Windows command spellings
The guard that stops a supervised gateway from restarting, stopping or
uninstalling itself only knew POSIX spellings of those commands, so the
Windows spellings of the very same operations walked straight past it.

Branch A anchored on the bare CLI name, but on Windows the CLI is spelled
with an executable suffix, and npm-style shims install `hermes.cmd` and
`hermes.ps1` beside the `hermes.exe`. Every one of these was allowed
while the unsuffixed form was blocked:

    hermes.exe gateway restart      -> ALLOWED (now blocked)
    hermes.cmd gateway stop         -> ALLOWED (now blocked)
    hermes.bat gateway restart      -> ALLOWED (now blocked)
    hermes.ps1 gateway uninstall    -> ALLOWED (now blocked)
    C:\tools\hermes.exe gateway stop-> ALLOWED (now blocked)

Branch D had the same gap for process termination: `\bp?kill\b` cannot
reach inside `taskkill` (there is no word boundary between the two `k`s),
so `taskkill /F /IM hermes-gateway.exe` and `Stop-Process -Name
hermes-gateway` were allowed while `pkill -f hermes gateway` was blocked.

This is reachable. The guard is gated on `_is_supervised_gateway_process()`
— PID-file ownership — not on platform, and its callers (terminal_tool,
code_execution_tool, approval, cron) all run on Windows.

Service-control spellings (`net stop`, `sc stop`, `Stop-Service`) are
deliberately left out: they presuppose a service install this guard has
no evidence of, and guessing at one risks blocking unrelated services.
`C:/tools/hermes.exe` (forward slashes) also stays allowed — the `/` in
the Branch A lookbehind is the deliberate #77173 path false-positive
fix, and narrowing it is a separate decision.

19 new cases pin the caught spellings, the quoted/wrapped forms that
reach the same place through the tokenizing rescan, and that the suffix
does not widen the match — `hermes.exe gateway start` stays benign,
`my-hermes.exe` is still a different binary. 13 of the 19 fail without
this change.

(cherry picked from commit e81c42e4482cb72a93f6468d7571f98aa5fcb848)
2026-09-18 12:49:13 -07:00
KoNit-K
339ee5b9dc fix(cron): block taskkill gateway interpreter
(cherry picked from commit 72f033c5823af00176c3b730b8e8502ad9b2b772)
2026-09-18 12:49:13 -07:00
teknium1
4ce832eeb9 fix(cron): bot-chat delivery cap bounds the bot's turn, not its exit linger
A cron job delivering to Bot Chat booked "timed out after 600s" for turns
that finished in seconds: cron/scheduler_delivery.py::_deliver_to_bot_chat
waited for the `hermes chat -Q` child to EXIT, but when the bot's turn had
messaged a teammate (message_agent -> notify_on_complete runner) the child
then runs the one-shot exit linger, bounded by
terminal.oneshot_completion_wait_seconds (default 600) — the same default as
cron.bot_chat_delivery_timeout_seconds. The cap counts from the claim, the
linger only starts when the turn ends, so the cap expired first on every
such delivery, booked a completed turn as a timeout, held the job's fire
fence for the full cap, and killed the child mid-linger — tearing down the
reply the linger exists to protect (#90879).

Ordering the two bounds cannot fix it (the turn has no duration bound; the
linger has its own contract), so the cap stops competing with the linger:

- hermes_cli/quiet_single_query.py: the -Q child accepts a per-process
  report path (HERMES_QUIET_TURN_REPORT_FILE), popped before the turn like
  HERMES_TURN_AUTHOR so nothing the turn spawns inherits it; cli.py writes
  {pid, exit_code, error} there the moment the turn ends, BEFORE the linger.
- cron: _run_bot_chat_turn polls the child and that report under the cap.
  Report present -> the delivery is booked from it (real exit code and
  stream tails when the child exits within a short grace) and the still-
  lingering child is left running, drained and reaped by a daemon thread.
  No report by the cap -> the turn never ended: killed and booked as a
  timeout, exactly as before. The linger itself is untouched.

Live repro (real _deliver_to_bot_chat, real `hermes chat -Q` against a
loopback provider whose turn spawns `sleep 90` with notify_on_complete,
cap 45s): base books the timeout at 45.7s for a turn that ended at +5s and
kills the child; fixed head books success at 11.2s, the child lingers, the
teammate follow-up turn runs at +97s and the child exits on its own.

Supersedes #113649's cap = delivery + linger (the thread shows a headroom
only moves the race and lengthens the fence hold); analysis credit to the
reporter and the thread's independent verification.

Fixes #113608

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:29:13 -07:00
teknium1
e1c0896518 fix(cron): no_agent script env comes from the factory's own snapshot, not a raw copy at the spawn site
tests/agent/test_subprocess_env_guard.py flagged cron/scheduler_script.py:362 as a new raw
os.environ.copy() spawn-env site. build_subprocess_env gains strip_launch_profile=True so the
launch profile's .env residue is still dropped from the base before the secret scrub (same
order and semantics as before: strip first, then scope-overlay the owning profile's declared
names), and the spawn site no longer snapshots the environ itself.
2026-09-18 10:04:57 -07:00
teknium1
802a9975d2 fix(cron): no_agent scripts get the owning profile's declared secret, never the launch profile's
A no_agent cron script owned by a profile served by a multi-profile gateway (or the
Desktop/dashboard backend) could not receive its own credential: the served profile's
.env never enters the process env (load_hermes_dotenv skips the process-global load
for a routed home), and the child-env sanitizer only resolved a terminal.env_passthrough
name through the profile secret scope when the launch environment already carried that
name. So the documented declaration mechanism (SECURITY.md 2.3) could not supply a
value that exists only in the served profile's scope, while the launch profile's own
.env credentials rode into the served profile's script child unstripped.

- tools/env_passthrough.scoped_passthrough_additions: declared names the bound scope
  holds but the env being filtered lacks; reads the scope alone (no os.environ, no
  other profile), empty without a scope so single-profile spawns are byte-identical.
- _scrubbed_env (terminal, cron scripts, bg processes, search workers) and
  execute_code's _scrub_child_env overlay those names after filtering.
- cron _run_job_script and terminal _make_run_env strip the launch profile's .env
  residue from the base (strip_launch_profile_env, a no-op for the launch profile's
  own jobs) before the filter, so the scope overlay lands after the strip.

Why not the 1Password-only shape of #114218: the gap is the declaration mechanism, not
one vendor's token, and an unconditional token pop broke the single-profile .env flow.
Docs: cron no_agent credential section, secrets child-process section, security
passthrough note.

Fixes #114209
Supersedes #114218
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-18 10:04:57 -07:00
teknium1
d6a58adbc5 fix(gateway,cron): 401 sign-in hints name the provider and the failing profile
The gateway exception reply (_STATUS_HINTS[401]) and the cron auth failure
notice still said a literal `hermes auth add <provider>` with no profile
selector — the placeholder this PR removed from the chat copy. Both now
format relogin_command_hint(provider): the exact OAuth command for a known
OAuth slug (turn agent's provider / the job's pinned provider), `hermes auth
add <slug>` for an API-key slug, and a profile-pinned placeholder when the
slug is unknown at the call site.

Part of #114012
2026-09-18 10:00:29 -07:00
teknium1
eafed27cf0 fix(cron): a completed run keeps its result when a fire-claim heartbeat sample misses
Long-running cron jobs (>60 s, i.e. past one heartbeat interval) intermittently
ended with last_status=error / "Interrupted by shutdown before terminal
completion." while their output was complete. The fire-claim heartbeat thread
took ONE sample that read the claim as not ours, latched lost_ownership, and
_FireOwnership.lost() then trusted that latch without asking the store again —
so a run whose claim still validated (the very condition under which
_record_fire_ownership_lost writes that message) was recorded as interrupted,
its output never saved. The reporter's own error string proves the miss was
transient: it is only ever written when the owner re-validates True.

Why this shape:
- _FireOwnership.lost(): an explicit transport cancel stays terminal; a latched
  heartbeat miss is re-checked against the store, and a claim that still
  validates keeps the run's real outcome. The owner-fenced mark_job_run and
  fire_claim_fence remain the authority, so a genuinely re-owned claim still
  yields (no delivery, no terminal write over the new owner, ledger discard).
  An unreachable store after a latched miss stays fail-closed.
- _heartbeat_loop: a miss is re-sampled once after
  _FIRE_CLAIM_MISS_CONFIRM_SECONDS before it latches. Latching cancels the live
  agent run and ends the lease refresh, so a single sample must not do that; a
  genuinely re-owned claim misses twice and latches ~1 s later.

The issue's other hypothesis (worker process identity under a systemd-run
--scope worker) is falsified by the code: the owner token is stored in
fire_claim.by at claim time and compared to the stored value, never recomputed
from _machine_id(), and the heartbeat thread runs in a copied context so it
reads the same store.

Slimmer redo of #113364 (same direction: revalidate before trusting the latch)
without its production-dead isinstance(_CombinedCancelEvent) branch.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:57:02 -07:00
teknium1
454f9c1421 fix(cron): remind once per cooldown after the alert-once gate; migrate alerted_at
Builds on the previous commit (alerted treated like closed): a permanently
silent `alerted` incident would hide a job that stays broken for days, so the
gate now withholds only inside `cron.failure_repeat_alert_hours` (default 6,
0 = re-alert on every failing run) and lets exactly one reminder through, which
re-stamps the window.

- cron/incidents.py: `alerted_at` column (added in place to existing ledgers via
  add_column_if_missing); set_incident_state(..., "alerted") stamps it every time
  so a reminder restarts the cooldown; resolved->detected re-open clears it so
  the same error after a green run alerts immediately.
- cron/scheduler.py: `_repeat_alert_withheld` reads the stamp; a missing or
  unparseable stamp (pre-migration row) delivers rather than swallowing the
  alert; `closed` still wins; the unreadable-ledger fail-open is unchanged. The
  crash path (`_deliver_crash_failure`) shares the gate through
  `_upsert_incident_for_failure`.
- hermes_cli/config_defaults.py + website/docs cron page: document the key.
- tests: fold the contributor's two tests and the old "unacked failures keep
  alerting per run" change-detector into two invariants (unit gate incl. 0 /
  legacy row / closed; end-to-end alert once -> reminder once -> recovery re-arms).

Fixes #113665
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 09:40:00 -07:00
Yagna Vudathu
7aa466665b fix(cron): alert once per failure incident instead of every run
A failing job re-pinged the operator on every run: _upsert_incident_for_failure
only suppressed acked (closed) signatures, while the post-delivery alerted mark
had no reader. Treat alerted like closed in the upsert gate so the same
signature goes silent after the first ping. Recovery still re-arms via resolved
-> detected, and a changed error still mints a new incident.

Fixes #113665.
2026-09-18 09:40:00 -07:00
teknium1
5ebbcbc8dc fix(cron): an oversized or live-SQLite cron script raises the named refusal, not the lifecycle verdict
`_read_script_for_scanning` returned a lifecycle-shaped sentinel ("hermes gateway
restart") when the job's OWN script was oversized, a device, or a SQLite database
open in this process, so `check_gateway_lifecycle` raised the misattributed
"cron job contains a gateway lifecycle command" error. It now returns
`(text, refusal)` via `_unreadable_reason` and the caller raises the existing
named "could not scan ... <reason>" refusal, matching the referenced-script walk
(#113944 atom 4).

Also adds encoding="utf-8" to the PR's own test file where the Windows footgun
scanner flagged bare write_text()/open() calls.
2026-09-18 09:35:16 -07:00
teknium1
d46abd445f fix(cron): a data file only mentioned in an inert heredoc never fails the lifecycle guard closed
The referenced-script walk reads paths named inside a masked (provably inert)
interpreter body so `os.system('/x/restart.sh')` is still caught, but it
turned every "could not scan" outcome for those mentions into a block: a
<1 MiB JSON with one >64 KiB line exhausted the text budget at depth 1, a
markdown table of paths pulled 64 remote-read misses and exhausted that
budget, and a live SQLite database (`state.db` inside a running gateway)
failed closed via LiveConnectionError with nothing logged. 22 of 36
in-gateway terminal blocks in one deployment were this class (#113944).

Thread `executed` through `_contains_unsafe_gateway_action`: content reached
only through a mention is still scanned for a literal lifecycle command, but
budget/depth/size/device/live-DB/cloud there is "nothing to scan", never a
verdict. Executed candidates still fail closed exactly as before.

When the walk does fail closed for a non-lifecycle reason, record it on the
budget (`scan_gateway_lifecycle` returns `(unsafe, refusal)`), log a WARNING
naming the path and reason, and tell the model that reason in both the
terminal guard and the cron `check_gateway_lifecycle` path instead of the
generic "cannot restart, stop, or uninstall the gateway" text that made it
reword and retry the same command.
2026-09-18 09:35:16 -07:00
lEWFkRAD
cb8d652b49 fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
2026-09-18 09:32:34 -07:00
teknium1
0db097dc22 fix(cron): scrub the bot-chat deferred record too; trim redaction tests to two invariants; docs
Follow-up to the salvaged #73026 (@ryangu00):

- `_deliver_to_bot_chat` now rebinds `content` to the redacted copy before anything reads it, so
  the durable deferred record (`cron/bot_chat_delivery.defer`) and its later replay carry the
  scrubbed payload as well as the live-owner and CLI lanes. Previously only the assembled
  `message` was scrubbed while the raw output was persisted for replay. Redaction is idempotent,
  so the replay's second pass is a no-op.
- The seven contributor tests collapse into two parametrized invariants: every outward lane
  (platform send / session mirror incl. job-name splice / bot-chat turn) masks the secret while
  the payload and framing survive; and egress redaction is forced (independent of
  `security.redact_secrets`) and fails closed (a raising redactor replaces the payload).
- Cron docs: note what is redacted on delivery, that it ignores `security.redact_secrets`, and
  what is deliberately not stripped (credential-named URL params, shapeless user secrets).

The failure-notice lane #55271 targeted (`_summarize_cron_failure_for_delivery`) exits through
the same `_deliver_result` chokepoint, so it is covered without a second redaction site.

Co-authored-by: necoweb3 <sswdarius@gmail.com>
2026-09-18 09:30:54 -07:00
Ryan Gu
c0362da9a6 fix(cron): redact secrets from delivery content before sending
Shell-job stdout/stderr is redacted where it is captured, but an LLM cron job's
response text reached delivery unscanned — so a job that surfaced a credential
(echoed a failing curl with an API key, summarised a config file) sent it
verbatim to the chat.

_redact_cron_payload() is the single fail-closed policy: force=True because
this is a safety boundary, not logging — the security.redact_secrets
preference governs the user's own logs and must not be able to turn scrubbing
off on the way out to a chat (same reasoning as tools/delegation_live_log.py).
A raising redactor replaces the payload rather than letting it through.

Every outward lane now routes through it:

- delivery content, at the single chokepoint right after media extraction, so
  the live-adapter and standalone send lanes consume the same redacted string;
- the session-mirror payload, which is derived from the raw content and so did
  not inherit the chokepoint — and a transcript outlives the message;
- the job NAME, which the mirror sinks and the continuable thread title splice
  around the redacted body; the name is user-controlled config, so a name
  embedding a credential re-leaked it right next to the scrubbed text;
- the bot-chat lane, which upstream added after this PR was opened: it hands
  the raw output to another profile's canonical Bot Chat as a real inbound
  turn, i.e. into a second transcript.

Rebased onto current main after cron/scheduler.py was split — the delivery
paths now live in cron/scheduler_delivery.py, so the patch moved with them.

URL-credential redaction (redact_url_credentials=True) is deliberately left
off: in this tree it guards context-persistence boundaries, not chat egress,
and cron legitimately delivers magic-link / pre-signed URLs to the user's own
chat.
2026-09-18 09:30:54 -07:00
teknium1
e16fc212ec fix(cron,tools): exact-id receipt reads fail closed on a non-dict JSON payload
Follow-up to the previous commit: rather than an isinstance check in every
caller (defer, deliver_to_live_owner, complete_delivery, the scheduler's
receipt comparison, bot_mode_dm's admit/wait), the two exact-id readers —
cron/bot_chat_delivery.py::read_pending and tools/bot_live_delivery.py::_read
— now raise ValueError for a payload that parses but is not a JSON object,
the same way they already propagate a PermissionError: a possibly-live
receipt is never overwritten, and the bulk-scan wrapper _scan_read turns the
same ValueError into its existing warn-once-and-skip path, so one malformed
ticket no longer wedges _next_sequence or claim_pending_delivery (#109820's
rule for the live-delivery dir).

Co-authored-by: beardthelion <beardthelion@users.noreply.github.com>
2026-09-18 09:19:04 -07:00
beardthelion
8235afad5d fix(cron): a parseable non-dict Bot Chat receipt no longer wedges the deferred drain
A receipt file that parses as JSON but is not an object (`42`, `"oops"`,
`[1]` — corruption, a truncated write, a foreign writer) slipped past the
"bad JSON" guard in `cron/bot_chat_delivery.py::_records` and raised
TypeError in the `sorted(..., key=item[1]["sequence"])` of `_drain` and in
`defer`'s sequence allocation, before a single healthy sibling was
delivered — on every scheduler tick, until someone deleted the file.

`_records` now rejects a non-dict payload on the same warn-once-and-preserve
path as unparseable JSON (the file's own rule: one bad file must not wedge
the dir), and the `_drain` re-read under the lock skips a record that turned
non-dict between the scan and the claim.

Ported from PR #114241 (hunks for cron/bot_chat_delivery.py::_records and
::_drain; the exact-id read hunks are superseded by the follow-up commit
that rejects non-dict payloads once at the readers).
2026-09-18 09:19:04 -07:00
Mohamad Kanso
2968b363bf fix(cron): sleep to deadline in ticker loops to eliminate clock drift (#114467)
Co-authored-by: yahuo <yahuoking@gmail.com>
2026-09-18 09:17:29 -07:00
teknium1
d6d3ee8584 fix(cron): terminal-worker reaper lives in scheduler_detached_worker, not the scheduler facade
cron/scheduler.py is already past the facade line-count gate; the reap helper is a detached-worker teardown concern with an existing topical home. The waiter late-imports it so behaviour and the sync run_one_job contract are unchanged.
2026-09-18 09:16:57 -07:00
liuhao1024
21d7283f3d fix(cron): reap the external worker when the waiter returns on terminal ledger
The waiter returns as soon as the execution ledger turns terminal, but
the worker process may still be in final teardown at that point. The
gateway stays the worker's parent, so nobody calling wait() afterwards
pins a zombie (STAT=Z) under it until the gateway is restarted. The
terminal early-return now hands the process to a short-lived daemon
thread whose single job is the final wait().
2026-09-18 09:16:57 -07:00
kshitijk4poor
c62bd9f207 refactor(cron): probe name covers both bounds; delattr fails loudly; one-line facade import
Review follow-ups on the salvage stack: the probe now names the fire-claim pair it exercises,
strips the constants with delattr so a cron.jobs that stops exporting them breaks the test's
premise loudly instead of silently, and the three-name facade import in unclaimed_pending_slot
no longer wraps now that the TTL left it.
2026-09-18 12:54:24 +05:30
Konstantin Khlopkov
e4e63e325a fix(cron): move FIRE_CLAIM_TTL_SECONDS to the same leaf as the skew bound
FIRE_CLAIM_SKEW_SECONDS now lives in cron/constants.py so a sibling loaded from a newer
on-disk tree never resolves it through the cron.jobs object a long-lived process cached at
boot. FIRE_CLAIM_TTL_SECONDS is the other half of the fire-claim pair and
cron/occurrences.py::unclaimed_pending_slot reached it the same way, through the facade.
Move it too, so both bounds have one home and the next sibling cannot re-open the trap by
importing the TTL from cron.jobs.

jobs.py keeps importing both from the leaf because it uses them (claim TTL default, live-claim
checks, skewed early-fire ownership); that import is a consumer, not a re-export shim.

Salvaged from #114692 on top of #114675; only the TTL delta and the leaf docstring are kept.
2026-09-18 12:54:24 +05:30
KoNit-K
f967d9b8f1 fix(cron): isolate fire claim skew constant 2026-09-18 12:54:24 +05:30
kshitijk4poor
3da502b67a refactor(cron): name the Bot Chat policy platform; drop a no-op pop
Bot Chat is the TUI/Desktop transcript, so its warning policy is `display.platforms.tui`;
the bare `"tui"` literal at two sites read as a typo beside `BOT_CHAT_PLATFORM = "bot-chat"`.

`_deliver_to_bot_chat` popped `_notification_all_targets_suppressed` on entry, but both
callers already guarantee the key is absent (`_deliver_result` pops it before and after the
call; `_drain`'s record JSON was snapshotted before the flag is ever set).
2026-09-18 01:43:35 +05:30
kshitijk4poor
5b6876ccc8 refactor(cron,gateway): count suppressed targets locally; import the wake predicate once
`_notification_suppressed_targets` lived on the job dict but had no reader outside
`_deliver_result`, and `_deliver_to_bot_chat` snapshots `dict(job)` into the durable
deferred record mid-loop, so a suppressed native target earlier in the loop leaked the list
into bot_chat_pending. A local counter has the same meaning and no durable footprint.

`diagnostic_wake_muted` is imported at module level in base.py instead of inside the
per-message handler and the turn-error path (no cycle: the resolver imports only
gateway.display_config).
2026-09-18 01:43:35 +05:30
kshitijk4poor
55b254a2c7 refactor(cron): keep the suppressed disposition, drop the delivery-manifest ledger rework
The suppression feature needs one cron cell: a failure notice whose target hides warning
notifications is recorded as `suppressed` (bot_chat_pending record, deliveries queue row,
execution delivery_outcome) instead of being sent. That is kept.

Everything else the PR added to cron/ is an independent ledger rework and is removed here:
execution delivery manifests + filesystem manifest journal, `schema_meta` fence and the
"legacy intent adoption" layer, incident occurrence generations + trigger, jobs projection
CAS (`bind_delivery_execution`/`update_delivery_projection`), the per-tick projection
reconciler, and the `_deliver_targets`/`_settle_manifest` split. origin/main has none of
the state that layer migrates from (zero occurrences of delivery_manifest / manifest_journal
/ schema_meta); the "pre-flag old writer" the r5-r9 tests simulate is a vendored snapshot of
this PR's own earlier revision (tests/cron/_r5_prior_executions.py). Those fixes may have
merit on their own and should land as separate PRs with a main-reproducing test each.

Also restores main's `_classify_delivery_outcome` precedence (`failed` before `queued`)
and drops the SHA-pinned `git show <PR commit>` test, which would go red the moment a
rebase-merge rewrote that commit.
2026-09-18 01:43:35 +05:30
Victor Kyriazakos
cd3de040ab feat(notifications): opt-in suppression of user-channel warning notifications
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.

Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
2026-09-18 01:43:35 +05:30
kshitijk4poor
d150fc2024 refactor(cron): one fail-closed tail; the transport-cancel log no longer claims a successful delivery
The exemption now also covers a delivered failure notice, so "during a successful delivery" was stale, and the elif/else bodies differed only by that log line.
2026-09-17 23:55:37 +05:30
kshitijk4poor
4c0ad43dfe fix(cron): a delivered failure notice keeps its real error when the claim sample misses after delivery
The post-delivery exemption required d.success, so a run whose FAILURE notice was delivered (delivery_attempted, no delivery_error) followed by a sampled claim miss still took _record_fire_ownership_lost and overwrote the real error with the interrupted-by-shutdown text. Widen the fall-through to 'delivery completed and not transport-cancelled' regardless of success: _finish_completed_run's owner-fenced mark_job_run (expected_fire_owner set on both the ok and failure paths) remains the authority, so a genuine loss still records nothing.
2026-09-17 23:55:37 +05:30
kshitijk4poor
8a6daa7702 refactor(cron): _FireOwnership owns the combined cancel event; two handles, not three
_run_one_job_body took a pre-built _CombinedCancelEvent plus the two raw events it was built from; run_one_job is the sole caller, so the combined view is derivable. _FireOwnership now takes claim_lost + transport_cancel, builds cancel_event (what run_job receives) itself, and exposes transport_cancelled(). The sampled-miss latch stays on the raw claim_lost event. The empty-final_response soft-failure normalisation moves above the ownership-lost check so the exemption reads off d.success alone. Behaviour identical; the two tests that inspected the old kwarg follow the rename.
2026-09-17 23:55:37 +05:30
kshitijk4poor
a0844a931f refactor(cron): trim the fire-claim exemption comments to the WHY 2026-09-17 23:55:37 +05:30
finn763
132a181d80 fix(cron): a delivered run keeps its ok status when the fire claim sample misses after delivery (#105861)
The post-delivery fire-claim check is a single sample of jobs.json. When it
missed on a run whose notice had already reached the platform,
_record_fire_ownership_lost() overwrote the delivered success with
'Interrupted by shutdown before terminal completion.' -> last_status: error,
and the operator's health watchdog then alerted on every tick for a job that
was working (#105861, 3 confirmed incidents).

Suggested fix #1: when the run's delivery actually completed (delivery
attempted, no delivery error, non-empty terminal response), log the ownership
loss as a warning and fall through to _finish_completed_run() -- whose
owner-fenced mark_job_run() is the authoritative claim check, recording ok
while the claim is still held and nothing at all when it really is gone.

A claim lost *during* a side-effect fence (delivery did not complete), a failed
delivery, a run that never delivered, and a claim lost before delivery all keep
the ownership-loss error path unchanged.

Review follow-up: fire_claim_lost ORs the sampled claim with the transport-level
cancel_event, so the flag alone can't say which one fired. _run_one_job_body now
also receives transport_cancel and the post-delivery exemption is scoped to the
sampled path only -- an explicit transport cancel (dashboard drain / shutdown)
stays fail-closed and still records the interrupted run. Regression added:
test_transport_cancel_during_delivery_stays_fail_closed sets the external event
during delivery and asserts last_status=error /
'Interrupted by shutdown before terminal completion.'

Review follow-up 2: the transport side of that flag can't be read off the event
either. _FireOwnership.lost() latched a sampled loss by calling set() on
fire_claim_lost, which with a transport cancel in play is
_CombinedCancelEvent(lost_ownership, cancel_event) -- and _CombinedCancelEvent.set()
sets every source it ORs. A pure sampled miss therefore made the transport event
read as cancelled, the exemption above was unreachable, and the drain event the
*caller* owns was mutated by this run's bookkeeping. _FireOwnership now also
receives the sampled source itself (sampled_claim_lost, threaded from
run_one_job's lost_ownership) and latches that one event: the combined wrapper
keeps its OR semantics and is_set() still reports the loss, so the in-flight
agent / pre-run script is interrupted exactly as before -- only the source
identity is preserved. Regression added:
test_sampled_miss_leaves_an_idle_transport_event_alone (real claimed job,
successful delivery, miss_at_sample=2, initially unset threading.Event) asserts
last_status=ok, last_error=None and the transport event untouched; it fails on
the previous head with the reported symptom.
2026-09-17 23:55:37 +05:30
teknium1
d4dfbba485 fix(cron): pin the worker's PYTHONPATH from the sanitized env in a sibling helper; skip under a wheel install
The pin landed inline in the cron/scheduler.py facade and rebuilt PYTHONPATH from
raw os.environ. That resurrected the Hermes-owned entries build_subprocess_env
had just stripped (runtime site-packages, launcher spellings of the repo root)
instead of extending what the sanitizer kept. It also ran unconditionally: under
a wheel / pipx / uv-tool install `Path(__file__).parent.parent` IS purelib, so the
pin hoisted site-packages above the stdlib on the worker's sys.path instead of
being a no-op.

cron/scheduler_worker_env.py::pin_hermes_tree_on_pythonpath prepends repo_root
to worker_env's own PYTHONPATH and returns the env untouched when repo_root is
sysconfig purelib (cron/ is already importable there). Test: A/B with the
previous inline pin -> the raw-environ-only entry leaked into the spawn env.

Part of #112729: this hardens the PYTHONSAFEPATH / cwd-not-checkout case; the
reporter's failure did not reproduce on main from cwd=checkout, so the real
worker stderr tail (now captured by e094e25b26 / f857a4ed99) is still needed.
2026-09-17 08:54:39 -07:00
teknium1
9269b19e4f fix(cron): restart-safe worker imports the gateway's own checkout instead of relying on cwd
The external cron worker is spawned as `sys.executable -m cron.scheduler`. Its
entry module is `cron.scheduler`, not `hermes_cli.main`, so it never runs the
bootstrap that puts the checkout on the gateway's sys.path; it only imported
`cron` at all through the implicit `-m` cwd entry (cwd was already the repo
root). That implicit path breaks on real hosts: a venv whose editable install
maps a moved or deleted checkout (the finder in this repo's own venv points at a
worktree that no longer exists), or a host that sets PYTHONSAFEPATH so `-m`
ignores cwd. The worker then dies with "No module named 'cron'" before its
ownership ack and every fire records
"cron external worker exited before ownership acknowledgement (exit 1)" (#112729).

The shared subprocess sanitizer strips Hermes-owned PYTHONPATH entries because
user children must not see our tree; this child IS Hermes, so after the env is
built the worker gets an explicit PYTHONPATH: the checkout `cron/scheduler.py`
lives in first, then whatever PYTHONPATH the gateway itself was started with.
The cwd stays the same checkout. Kanban workers spawn `-m hermes_cli.main` and
already get the bootstrap, so this is the only affected entry point.

Live probe (worktree python from cwd=/tmp with PYTHONSAFEPATH=1, real
_launch_external_cron_worker + real Popen): before, the worker stderr tail is
"Error while finding module specification for 'cron.scheduler'
(ModuleNotFoundError: No module named 'cron')"; after, the worker imports
cron.scheduler, loads the payload and reaches the durable-ownership check.

The two review minors on the first pass (worker unlinks its own stderr capture;
comments no longer claim DEVNULL) already landed in f857a4ed99 and are covered
by that PR's tests.
2026-09-17 08:54:39 -07:00
teknium1
66a91b7ce8 fix(cron): a server parked on a permanent error blocks the job again; warn once per outage
The reconnecting exemption keyed only on `_ever_connected`, so a server that
connected once and then parked on a PERMANENT error (revoked credentials,
endpoint gone: `_park_and_rearm('from parked state (permanent error)')`) looked
identical to a router-reboot park. Its self-probe fails the same way every
interval, so the job ran tool-less on every tick forever with only a gateway
WARNING and never received the one-shot blocked_config alert it used to get.

Record the revival reason `_park` was given on MCPServerTask (`_park_reason`,
cleared when a session proves healthy) and have `mcp_server_reconnecting`
return False for a permanent-error park, restoring the block and its alert.
Transient parks (network blip, rapid-drop budget exhausted) still return True.

Also dedupe the exemption's WARNING per job+server for the length of one
outage, like the one-shot blocked_config alert, instead of logging every tick;
the entry drops once the server resolves tools again so the next outage warns.
2026-09-17 08:51:30 -07:00
teknium1
6924cc66fb fix(cron): run a job whose enabled_toolsets MCP server is only reconnecting, instead of blocking it
A cron job naming an MCP server in its enabled_toolsets was hard-blocked
(blocked_config, one-shot alert, no LLM call) whenever that server resolved
to zero tools — including the minute a router reboot or DNS blip left an
otherwise healthy server parked and self-probing. The preflight could not
tell "the network blinked" from "wrong name / other profile's server", and
the job lost a whole tick to a transient outage the MCP layer was already
recovering from on its own (#112871).

The MCP run task keeps that distinction: a server that connected once in
this process (_ever_connected) and is currently sessionless with a live
task is degraded/parked and will revive. tools.mcp_tool_discovery grows a
read-only mcp_server_reconnecting(name) predicate on that state (resolved
through the profile's own or adopted connection key, never connecting),
and _empty_requested_mcp_toolsets skips such servers with a WARNING that
names them — the job runs with whatever tools did resolve, as the reporter
asked. A server that never connected for this profile keeps the block,
which is the misconfiguration case #109050 built the check for. The
blocked reason now says "never connected for this profile" so the
remaining block reads as what it is.

Docs: the pre-dispatch validation section lists the MCP check and the
reconnecting exemption.
2026-09-17 08:51:30 -07:00
brooklyn!
98745df3b9 fix(serve): hold work admission through cooperative retirement
Fence accepted RPCs, side agents, continuations, delegation and cron before issuing process-bound retirement permits.

Co-authored-by: chelsealong <chelsealong@126.com>
Co-authored-by: Mark Sheppard <mark@orchardstreetpress.com>
2026-09-17 00:34:50 -05:00
teknium1
7ad51527d4 fix(cron): blocked-config notice no longer conditions the retry on the user fixing something
The MCP zero-tools reason now says the block "clears by itself" when the server
comes back, which contradicted the fixed suffix "will try again at the next
scheduled time once this is fixed". Drop "once this is fixed": the retry happens
at the next scheduled time regardless of who or what resolves the block.
2026-09-16 17:53:04 -07:00
teknium1
fad867648d fix(cron): say the MCP zero-tools preflight block clears itself when the server reconnects
The block from `_empty_requested_mcp_toolsets` is re-evaluated on every dispatch and the
alert fires once, but its reason read as terminal ("Fix the server or remove it from the
job's toolsets"). After a one-minute DNS outage parked two remote MCP servers, two
operators independently misdiagnosed the block as a config error they had introduced,
while the job resumed on its own four minutes later (#112871).

Spell out both readings in the reason text: temporarily unreachable → clears by itself on
the next dispatch; wrong name / other profile → fix or remove. The blocked run document
already said this; the WARNING line and the alert — what the operator actually reads — did
not.

Part of #112871 (the fail-closed decision for a parked-but-recoverable server is separate).
2026-09-16 17:53:04 -07:00
teknium1
f857a4ed99 fix(cron): acknowledged worker removes its own stderr capture; comments match captured stderr
The `<execution_id>.stderr` capture was unlinked only by the launching gateway
(pre-ack exit branch and the `handoff_files` finally). In the restart-safe
topology the worker deliberately outlives a gateway restart, so a gateway that
died mid-run left one orphan file per surviving run under
cron/external-workers/. After the ack the gateway never reads the capture
(it only serves the pre-ack death report), so the worker now unlinks the
sibling derived from its ack path in the post-run finally; the pre-ack paths
are unchanged so a worker dying before its ack still gets its stderr reported.

Two comments still described stderr as DEVNULL (the systemd-run failure branch
and the `__main__` setup_logging motivation); reword them to the captured-stderr
behaviour while keeping the WHY.
2026-09-16 17:52:40 -07:00
teknium1
e094e25b26 fix(cron): a restart-safe worker that dies before its ack reports its own stderr
`_launch_external_cron_worker` spawned the worker with `stderr=DEVNULL`, so a worker
that crashed before publishing its acknowledgement — an import error, a module missing
from the interpreter it was handed, a bad payload — surfaced only as
"exited before ownership acknowledgement (exit 1)". #112729 shows the cost: the
reporter, a claimant and a reviewer spent three rounds guessing at the mechanism
(sys.executable symlinks — refuted) because the worker's actual traceback was thrown
away. The worker's own log handler is installed only after `cron.scheduler` imports,
so an import-time failure never reaches agent.log either.

Capture stderr into `<home>/cron/external-workers/<execution_id>.stderr` (0o600), and
on a pre-ack exit append a redacted tail of it to the dispatch error (which lands in
`last_error`, the executions ledger and the gateway log). The file rides in
`handoff_files` so every exit path removes it; the tail helper lives in
`scheduler_diagnostics` next to the run-error redaction.

Part of #112729
2026-09-16 17:52:40 -07:00
teknium1
f9745c1e3b docs(kanban): worker guidance is scoped to the owning worker
User docs say who receives the worker protocol (only the dispatcher-spawned
worker that owns the task) and the kanban AGENTS section names the helper
the guidance and stop-nudge injection sites use, so the next prompt-side
HERMES_KANBAN_TASK reader does not grow its own predicate. Other readers on
main pair the env read with is_dispatcher_owned_worker_context() and are
described as such rather than rerouted.
2026-09-16 17:48:48 -07:00
teknium1
6a99766424 fix(agent): system watchdog interrupts are attributed to their issuer, not the user
Every hard interrupt reached `begin_iteration` through the same `_interrupt_requested`
flag, so the turn loop booked all of them as `interrupted_by_user` — a cron run killed
by the scheduler's inactivity watchdog, a turn aborted by the liveness watchdog, a lost
session turn lease and a gateway inactivity timeout all read as a human pressing stop,
and the investigation went to the wrong subsystem (#112647).

`interrupt()` already records a trusted category per interrupt (`_tool_interrupt_reason`,
fed by the `tool_reason` every producer can pass through `request_hard_interrupt`). The
exit reason is now derived from it: the three categories `interrupt()` itself mints for
human stops keep `interrupted_by_user` / `interrupted_during_api_call`; any other
category names its producer — `interrupted_by_system(cron_inactivity_watchdog)`,
`interrupted_during_api_call(turn_liveness_watchdog)`. The system producers that passed
no `tool_reason` (cron inactivity watchdog, turn liveness watchdog, lease loss, gateway
inactivity timeout) now name themselves, so the model-visible tool-cancellation text
says the same thing. `_publish_interrupt_state` logs ONE line naming the source so the
turn record and the log agree.

No change to WHEN anything interrupts. `interrupted_during_api_call` moves to the
prefix-matched explanation table so the parameterised form keeps its user copy.

Slim redo of #112652 by @KoNit-K, which added a parallel `issuer` attribute and
keyword; this reuses the existing `tool_reason` plumbing instead.

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
2026-09-16 17:48:17 -07:00
teknium1
3c9431a876 fix(cron): seed Telegram forum-topic cron threads on the group slot topic replies key
_THREAD_REPLY_CHAT_TYPE covered Slack and Matrix but left Telegram on the default "thread"
slot, while the Telegram adapter's _build_message_event types every supergroup message
"group" (forum topics included), so a forum-supergroup continuable cron seed landed in
agent:main:telegram🧵<chat>:<topic> and the user's topic reply resolved to
agent:main:telegram:group:<chat>:<topic> — the same seed/reply split the Matrix fix closes,
on the untouched sibling platform (and the slot the /handoff path already binds).

Add telegram to the table; retarget the two tests that pinned the wrong "thread" contract
for a Telegram seed (the backward-compat default now uses Discord, the one platform whose
in-thread replies key "thread"); add one seed-key == reply-key invariant test for the
Telegram forum topic. Mirror the three one-line Matrix doc edits into the zh-Hans pages.
2026-09-16 17:18:44 -07:00
teknium1
f41ad3bc0a fix(gateway,cron): key Matrix handoff/cron thread seeds the way Matrix replies are keyed
A Matrix `attach_to_session` cron delivery and a CLI→Matrix `/handoff` now land in their own
thread AND a human reply in that thread continues the seeded session (#112918).

The salvaged `MatrixAdapter.create_handoff_thread` (ea-s21, #108367) gives both seeders a
thread root. The second half of the bug is a key-shape mismatch: the handoff watcher and
`_seed_cron_thread_session` seeded `matrix🧵<room>:<root>` while the adapter keys every
in-thread reply on the ROOM's type (`matrix:group|dm:<room>:<root>`), so seed and reply never
met. Following the direction landed for Slack in 1f3f45e87b (#111896), the seeders now mirror
the adapter instead of the adapter moving onto a `thread` slot: rekeying inbound Matrix threads
would orphan every existing Matrix thread session and drop `is_group` in authz for in-thread
messages (`_GROUP_CHAT_TYPES` has no `thread`).

- gateway/run_startup.py: Matrix destinations key `dm`/`group` from the adapter's
  `get_chat_info` (`_handoff_home_is_dm`).
- cron/scheduler_delivery.py: `_THREAD_REPLY_CHAT_TYPE` — Slack and Matrix non-DM thread seeds
  use `group`. Slack channel cron threads had the same mismatch (adapter `build_source` keys
  `group`; the seed said `thread`); the two cron tests that pinned `thread` for Slack channels
  asserted the wrong shape and now assert the adapter's.
- Dropped from #108367: the `chat_type="thread"` inbound rekey (see above) and its tests;
  contributor tests trimmed to two invariants.
- Docs: Matrix listed among thread-capable handoff/cron platforms.

Live probe (in-process, real MatrixAdapter + real handoff destination + real cron seeder):
before  create_handoff_thread -> None; handoff key matrix🧵… ≠ inbound matrix:group:…;
        cron seed matrix🧵… ≠ inbound
after   create_handoff_thread -> '$seed'; handoff == inbound == cron seed (room and DM cases)

Co-authored-by: ea-s21 <190767603+ea-s21@users.noreply.github.com>
2026-09-16 17:18:44 -07:00
teknium1
9f48f4ed04 fix(cron): deliver [CRON_FAILURE] evidence verbatim, not through provider heuristics
The agent-declared failure evidence was routed through
_summarize_cron_failure_for_delivery, whose substring heuristics re-diagnose
natural-language text: a child that "timed out after 30 minutes waiting on the
database" was delivered as "the AI model service did not respond in time … add
one with `hermes fallback add`", and a child that saw a vendor 401 became "the
AI model service rejected the sign-in. Sign in again with /login". The operator
got a wrong diagnosis and a wrong remediation; the real evidence only survived
in the saved run output.

The marker is now carried as a flag (_RunDelivery.agent_declared ->
_compose_run_delivery(agent_declared=...)) and composed like blocked_config
already is for the same reason: `⚠️ Cron '<name>' failed: <evidence>` plus the
run-log / re-run hint. Incident ledger, acked-suppression and the streak nudge
are unchanged.

Review follow-up on #113155.
2026-09-16 17:11:51 -07:00
KoNit-K
f26e436d35 fix(cron): record agent-declared cron failures 2026-09-16 17:11:51 -07:00