33 Commits

Author SHA1 Message Date
Austin Pickett
e7bff4b6d8 fix(cron): a pinned job never falls back to the global fallback chain (#120312)
* refactor(fallback): share the pinned-owner chain rule

delegate_task's _resolve_child_fallback_chain decides which fallback chain
a child may walk: a pinned child never borrows the parent chain, an explicit
[] disables fallback, a declared list is the child's own. Cron needs the
same rule for pinned jobs (#100437), so the body moves to
hermes_cli.fallback_config.scoped_fallback_chain and the delegation helper
becomes a thin caller. Behaviour is unchanged; the delegation matrix test
still pins every cell.

* fix(cron): a pinned job never falls back to the global chain

A job with its own provider, model or base_url is an explicit operator pin
(since 0469740ab3 unpinned jobs store none of these). It still walked the
global fallback_providers chain in two places, so a pinned job could run
on a different provider and model than the one chosen:

- _resolve_job_runtime walked the chain on an AuthError or transient
  network failure while resolving the pinned primary;
- _resolve_cron_agent_setup handed the global chain to every cron agent as
  fallback_model, so the conversation loop's provider ladder could swap a
  pinned job mid-run.

Both now read _job_fallback_chain(job, cfg), which returns no chain for a
pinned job through the same scoped_fallback_chain rule delegate_task uses
for pinned children. The pre-dispatch key check reads it too: the global
chain used to skip that check for every job, so a pinned job with a
missing key now blocks before the agent is built instead of failing in the
resolver. The transient-failure notice for a pinned job says it does not
fall back and names --unpin, instead of "No backup provider succeeded".

Unpinned jobs (including legacy *_snapshot records) and same-provider
credential-pool rotation are unchanged. The two scheduler tests that
asserted atomic provider+model fallback swaps used pinned jobs; they now
use unpinned jobs and keep the same assertions.

No per-job fallback_providers list: jobs have no generic override field
(create_job/update_job, the cronjob tool schema and the CLI enumerate each
field), so an opt-in chain would be a new surface on all of them. The
escape hatch is to leave the job unpinned and pick its model with
cron.model / cron.model_provider.

Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>

* docs(cron): pinned jobs do not use fallback_providers

cron.md "Provider recovery" and the pre-dispatch key check, the cron rows
and section in fallback-providers.md, and the developer notes in
cron-internals.md / provider-runtime.md said every cron job inherits the
global chain. State the new rule, the compatibility note for users who
relied on a pinned job landing on the chain, and the unpinned + cron.model
alternative.

---------

Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
2026-09-23 10:42:18 -04:00
teknium1
9ecd22e5da fix(cron): release a profile's in-flight claim under the key it registered with
Under one multiplexing ticker, every non-launch profile's in-flight claim
leaked on every run. The cron home scope is a ContextVar: the claim is taken on
the ticker thread inside `_profile_cron_scope`, but the pool worker's `finally`
sits outside `ctx.run`, where the worker thread resolves the LAUNCH home — so
the discard missed the real key. The job then skipped a fire window until the
force-release backstop swept it, and the shutdown drain plus
`hermes.cron.jobs.running` saw phantom work.

`_submit_with_guard` now captures the registering home and passes it to
`release_running_job(job_id, home=...)` on every release path.

Also in this pass:

- Stand down for a profile that runs its OWN gateway (`_cron_profile_gate`,
  the gate the serve/Desktop ticker already passes). On a host pinned to
  per-profile gateways both processes raced that profile's tick lock, and when
  the launch gateway won, delivery went through SharedRouteAdapters/fail-closed
  instead of the profile's live adapters. The gate compares the liveness PID
  against `os.getpid()`: this process holds the launch `gateway.pid` and
  publishes every served profile in `served_profiles`, so a bare liveness answer
  would have stood cron down host-wide.
- Job-liveness consumers ask `is_job_running(job_id, home=...)` instead of the
  host-wide bare-id union, which let profile A's running `daily-brief` report
  profile B's idle one as running and keep B's stale one-shot alive.
- `register_ticked_homes` reaps the parallel pools of homes that leave the
  ticked set; pools lived until `atexit`, so every home ever ticked kept a
  ThreadPoolExecutor and its worker threads.
- `mark_running_jobs_interrupted` reads the real home Path from
  `_inflight_home_path` instead of rebuilding it from the normcased key half.
- The ownership test imports `cron.scheduler_ownership` per test, so the file
  fails behaviourally on base instead of as a collection error.
2026-09-21 02:52:15 -07:00
teknium1
cb647a018f fix(cron): one host ticker owns every profile's cron, per profile
The cron ticker multiplexes N profiles from one process while its ownership
predicates and its in-flight bookkeeping still assumed one profile per process.

- `_should_yield_tick_to_fresh_gateway` asked a process-global boolean
  (`owns_gateway_runtime_lock`) and a launch-home lock probe, so one answer
  covered every profile ticked. It now asks `scheduler_ownership`:
  `owns_cron_tick_for(home)` (this process is the host gateway AND ticks that
  home) and `live_gateway_ticking(home)` (another live host gateway whose
  published served set covers that home).
- In-flight state (`_running_job_ids`, `_running_since`, `_running_futures`,
  `_running_allowance_s`, `_running_worker_pids`, `_running_fire_owners`,
  `_restart_safe_waiter_job_ids`, `_interrupted_job_ids`) is keyed by
  `_inflight_key(job_id)` = `(home key, job id)`; two profiles carrying a
  `daily-brief` no longer read as one job. The public accessors still report
  the host-wide union of bare job ids for the shutdown drain.
- The parallel worker pool is keyed by home: `cron.max_parallel_jobs` is a
  per-profile key, and the single global pool was sized by whichever profile
  ticked first and torn down by the next one.
- `gateway/run.py` no longer gates the cron tick set on
  `gateway.multiplex_profiles`: that flag gates adapters, and with it off every
  non-launch profile's jobs sat in a store no ticker visited.
2026-09-21 02:52:15 -07:00
teknium1
8863b36fd6 fix(cron): stale-code yield reads as an outage in cron status; hermes update restarts proven-stale gateways
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:

* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
  loop) and the success marker; with no success marker on disk it printed
  "✓ Gateway is running — cron jobs will fire automatically", and with a stale one
  it pointed at "Check the gateway log" instead of naming the cause. The persisted
  `CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
  and reported as "Gateway is running STALE code — fires NOTHING", with both
  revisions and the restart command. A fresh heartbeat with a recorded error and no
  success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
  left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
  to the existing drain-first `request_restart` path (SIGUSR1) via
  `hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
  code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
  contract the restart phase already uses for unmapped manual gateways. The drain
  budget computation is shared (`_gateway_drain_budget`).

The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).

Fixes #117275
2026-09-21 01:57:31 -07:00
teknium1
bdde0b0c28 docs(cron): describe when a stale-code ticker yields and when it keeps dispatching
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped
code_sha is the on-disk revision; a lock held by an equally stale process never
counts. Placed under "Gateway Integration" so it does not collide with the
"Stale-code yield" section #117501 adds under "Locking".
2026-09-20 20:13:19 -07:00
teknium1
41b6ba09d9 docs: name the registered tool in prose that still says cronjob/todo/process
The same rename that left the two reference pages stale (cronjob->cronjob_manage,
todo->todo_list, process->process_manage) left the guides and developer docs
referring to the old tool names; a reader following them gets "no such tool".
Toolset names (`cronjob`, `todo`) are unchanged and left alone.
2026-09-20 12:56:25 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
teknium1
eafed27cf0 fix(cron): a completed run keeps its result when a fire-claim heartbeat sample misses
Long-running cron jobs (>60 s, i.e. past one heartbeat interval) intermittently
ended with last_status=error / "Interrupted by shutdown before terminal
completion." while their output was complete. The fire-claim heartbeat thread
took ONE sample that read the claim as not ours, latched lost_ownership, and
_FireOwnership.lost() then trusted that latch without asking the store again —
so a run whose claim still validated (the very condition under which
_record_fire_ownership_lost writes that message) was recorded as interrupted,
its output never saved. The reporter's own error string proves the miss was
transient: it is only ever written when the owner re-validates True.

Why this shape:
- _FireOwnership.lost(): an explicit transport cancel stays terminal; a latched
  heartbeat miss is re-checked against the store, and a claim that still
  validates keeps the run's real outcome. The owner-fenced mark_job_run and
  fire_claim_fence remain the authority, so a genuinely re-owned claim still
  yields (no delivery, no terminal write over the new owner, ledger discard).
  An unreachable store after a latched miss stays fail-closed.
- _heartbeat_loop: a miss is re-sampled once after
  _FIRE_CLAIM_MISS_CONFIRM_SECONDS before it latches. Latching cancels the live
  agent run and ends the lease refresh, so a single sample must not do that; a
  genuinely re-owned claim misses twice and latches ~1 s later.

The issue's other hypothesis (worker process identity under a systemd-run
--scope worker) is falsified by the code: the owner token is stored in
fire_claim.by at claim time and compared to the stored value, never recomputed
from _machine_id(), and the heartbeat thread runs in a copied context so it
reads the same store.

Slimmer redo of #113364 (same direction: revalidate before trusting the latch)
without its production-dead isinstance(_CombinedCancelEvent) branch.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:57:02 -07:00
lEWFkRAD
cb8d652b49 fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
2026-09-18 09:32:34 -07:00
teknium1
e89605b4e6 docs(cron): record the due-only occurrence identity in the missed-occurrence contract
Also map the two salvaged contributor emails (kleros109, fangliquanflq).
2026-09-15 06:07:32 -07:00
teknium1
0b40f5a790 docs: fold the root docs/ tree into the Docusaurus site and delete it
docs/ was not the documentation site; it was a grab bag of long-form
design notes, wire contracts and observability guides that landed with
feature PRs because their authors needed somewhere to put them. Root
AGENTS.md already says long-form dev docs live in
website/docs/developer-guide/; this moves the 14 living documents there
(or to the matching user-guide section) so they are published, searchable
and linked from the sidebar instead of being found by grep only.

Developer guide: micro-compaction, gateway-session-lifecycle (was
session-lifecycle), state-db-recovery, multiplexing-gateway,
chronos-managed-cron-contract, relay-connector-contract, observer-hooks
(was observability/README), gateway-monitoring (observability/monitoring),
relay-shared-metrics, middleware, streaming-tts, billing-lifecycle.
User guide: egress/network-isolation (was security/network-egress-
isolation), features/kanban-multi-gateway (was kanban/multi-gateway).

Each page got title/description frontmatter and a sidebar entry; repo-
relative links became site links or GitHub blob URLs; two MDX brace
hazards escaped. Every in-tree pointer (module docstrings, config
comments, the relay conformance test's Path, the monitoring-doc test,
gateway-internals, cron-internals, kanban docs, .dockerignore, AGENTS.md)
now names the new location. `docusaurus build` passes with no unresolved
links on the moved pages.
2026-09-13 06:06:46 -07:00
Teknium
6cd4fbd640 docs(cron): state the missed-occurrence contract for restart gaps
cron-internals.md gets a 'Missed-occurrence contract' section (pre-dispatch
advance is provisional, restore once, never twice, grace, opt-out, paused never
catches up, same on standalone and multiplexed); the user guide describes the
catch-up-once behaviour above the cron.catch_up_missed opt-out; cron/AGENTS.md
lists it as a hardening invariant.
2026-09-12 01:51:59 -07:00
Teknium
5280fe9987 fix: cron and local DMs reach an open Desktop Bot Chat
Route local producers to durable owner ingress before attempting the unowned
CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling;
never fall back after ambiguous admission. Report cron admission as queued,
not completed or failed, in job status, the execution ledger and CLI/tool UX.

Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for
both idle and busy owners. Fixed owner consumes idle cron, busy cron, local
DM and mounted-chat cron exactly once, keeps its lease, yields to queued
human input, and preserves the prior model-request prefix and tool schema.
Inference alone used a deterministic loopback wire stub; no paid model call.
2026-09-07 16:48:29 -07:00
Teknium
bbbcde8173 fix(process): stop late forks escaping deadline tree cleanup 2026-09-07 08:19:32 -07:00
Teknium
df4b3733ba fix(cron): every last_status consumer renders delivery_failed explicitly (dashboard badge, Desktop inspector, /cron list, docs)
Audit of every last_status reader outside the scheduler (rg last_status across
web/, apps/desktop/, hermes_cli/, tui_gateway/, tools/, scripts/, website/):

- web dashboard CronPage: last_status was never rendered at all — a
  delivery_failed job showed a green 'scheduled' badge and only a small red
  'delivery: ...' line. New pure cronLastResult() helper maps the closed
  literal set to tones (ok=success, delivery_failed/blocked_config=warning,
  error/unknown=destructive) and the card now shows an amber
  'delivery_failed' badge (title = last_delivery_error).
- Desktop hermes-bots routine inspector: 'Last result' printed the raw
  literal; routineLastResult() spells out each one ('Ran, but delivery
  failed', 'Blocked by configuration (not run)', ...), unknown passes through.
- /cron list (cli_commands_mixin): 'Last run: <ts> (delivery_failed)' now
  appends the delivery reason, since last_error is None for those runs.
- hermes cron list/doctor and the cronjob tool already handled the literal
  on this branch; no consumer compared == 'ok' for success apart from the
  cronjob manual-run path, which the branch already fixed.
- developer-guide/cron-internals.md: table of last_status literals + which
  detail field carries the reason.

Live repro (real 'hermes dashboard' on a temp HERMES_HOME with a
delivery_failed job, CronPage rendered against the live /api/cron/jobs):
before — badges [scheduled, default, telegram:123]; after — badges
[scheduled, delivery_failed (warning tone, title 'telegram: 502 Bad
Gateway'), default, telegram:123].
2026-09-02 00:52:58 -07:00
Teknium
a2da0ab797 feat(cron): bot-chat delivery target — cron output lands in a bot's canonical Bot Chat and the bot responds
deliver='bot-chat[:<profile>]' is a machine-local pseudo-platform: the
scheduler delivers job output as a real inbound turn in the target
profile's canonical Bot Chat via the chat CLI lane (--in ~ -c "Bot Chat"
--create-if-missing -Q --query-file), the same lane Bot Mode
agent-to-agent messages use. The bot reads the output, acts on it, and
responds in its chat — instead of the output only landing in Run history.

- cron/scheduler.py: token parsing, target resolution (own profile /
  named local profile / unknown -> skipped with warning), subprocess
  delivery lane with cron.bot_chat_delivery_timeout_seconds (default
  600s), preflight exemption, and bot-chat entries in
  cron_delivery_targets() for UI pickers. Excluded from 'all' by design.
- tools/cronjob_tools.py: create/update-time validation — named profiles
  must exist on this machine (fail at create, not at 3am); deliver schema
  documents the new token.
- tui_gateway/methods_tools.py: cron.manage add forwards deliver.
- hermes_cli/profiles.py: list_profile_names() cheap name-only scan.
- hermes-bots plugin: Create Cronjob dialog gains a 'Send results to'
  picker (Run history only / <bot>'s chat); bot-chat jobs send the BARE
  token on the profile-scoped create so Desktop-side aliases can never
  name a profile the backend doesn't have.
- Docs: user cron guide, automate-with-cron, cron-internals.

Machine-local by construction: names resolve only against the executing
machine's ~/.hermes/profiles/, so overlapping profile names across
multiple connected gateways are unambiguous.
2026-08-21 12:48:53 -07:00
Teknium
ef04d846e9 feat(cron): cron agents now run with memory enabled like every other agent
Cron jobs were constructed with skip_memory=True and a hard 'memory'
toolset denial, so MEMORY.md/USER.md never loaded and the memory tool was
stripped even from per-job enabled_toolsets. That was inconsistent with
kanban/delegate/gateway agents (which all get memory) and forced users
into hacky bypasses.

- cron/scheduler.py: skip_memory=False on the cron AIAgent; drop 'memory'
  from _resolve_cron_disabled_toolsets; remove _strip_cron_memory_toolset
  and its call sites
- agent/agent_init.py: update stale comment referencing the cron denylist
- tests: flip pinning tests to the new contract (memory enabled, per-job
  memory toolset kept, user-level denylist still wins)
- docs: cron-internals + automate-with-cron no longer claim cron has no
  persistent memory
2026-08-21 03:46:37 -07:00
Teknium
43874d1a96 docs: accuracy sweep + coverage for 2 months of shipped features
Accuracy pass (all 373 pages audited against code, 13 parallel audits):
- configuration.md: 12 stale defaults/keys (file-sync rewrite, clarify
  timeout, streaming knobs, iteration budget, TTS/STT enums)
- reference/: commands/env-vars/toolsets/tools synced with
  COMMAND_REGISTRY, argparse tree, OPTIONAL_ENV_VARS, TOOLSETS
  (28 env vars added, 3 phantom removed, mcp__ naming, webhook
  platform restricted toolset)
- features/, messaging/, developer-guide/, guides/: ~60 factual fixes
  (web_extract truncation, dashboard auth fail-closed, delegation
  blocked tools, adapter signatures, session schema v23, phantom
  Matrix env vars, hermes setup tts, auth spotify, webhook --skills)
- zh-Hans: explicit heading IDs fix 2 broken WSL2 anchors

New coverage for features shipped in the last 2 months (verified
against code before writing):
- compression.in_place, verify-on-stop (+v31/v32 migration reality),
  ${env:VAR} SecretRef, display.timestamp_format, session:compress
  hook + thread_id/chat_type fields
- /journey learning timeline, per-channel model/system-prompt
  overrides, /sessions search, clarify multi-select, -z --usage-file,
  uninstall --dry-run, config get/unset
- MCP elicitation, extra_headers, discover_models, api-server run cap,
  Bedrock cachePoint, Discord reasoning_style, Google Chat clarify
  cards, vibe reactions, resume cwd restore, Yuanbao forwarded
  messages, api_content sidecar, roaming pet, tool_progress log mode,
  WhatsApp polls/locations, kanban per-task model + lifecycle hooks
2026-07-29 08:48:05 -07:00
Teknium
643b0dc678 fix(cron): raise default pre-run script timeout from 120s to 1h (#55489)
Cron pre-run scripts were capped at 120s by default, which surprised
users running long data-collection scripts on crons (the whole point of
crons being to offload long work). Raise _DEFAULT_SCRIPT_TIMEOUT to 3600s
(1 hour).

This bounds the script only — skill/agent jobs already run on a separate
inactivity budget (HERMES_CRON_TIMEOUT, default 600s idle, 0=unlimited),
not a wall-clock cap. Scripts dispatch to a persistent thread pool and do
not hold the tick lock, so a long script doesn't starve other due jobs.

Docs clarified to make the script-vs-agent timeout distinction explicit.

env/config overrides (HERMES_CRON_SCRIPT_TIMEOUT,
cron.script_timeout_seconds) unchanged and still take precedence.
2026-06-30 01:00:39 -07:00
Ben Barclay
e1f4098b9f docs(cron): document explicit per-channel delivery targets for all platforms (#54630)
The cron delivery table only showed Discord/Telegram with explicit
target syntax and described Slack and every other platform as
home-channel-only. In fact the generic platform:<target> routing in
_resolve_single_delivery_target resolves explicit targets for every
platform: Slack (#channel / channel ID / channel:thread_ts), Matrix
(room/user IDs), Feishu (chat:thread), WhatsApp (JID / E.164), Signal
(group / E.164), SMS, Email, and Weixin all have dedicated explicit-
target branches in _parse_target_ref; the remaining platforms accept a
generic platform:<chat_id> passthrough.

Update the Delivery Model table (en + zh-Hans) to show the real
per-platform syntax, document #channel name resolution via the channel
directory, and note the Slack thread_ts nuance. Docs-only.
2026-06-29 15:23:16 +10:00
Ben
b75757d4aa feat(cron): wire on_jobs_changed, cron.chronos config, docs + agent↔NAS contract
Phase 4F (F.1 + F.2 + F.3, agent side). F.4 is the operator-run live smoke
(needs a NAS deployment); recorded in the PR, not code.

F.1 — on_jobs_changed wiring:
- cron/scheduler.py: _notify_provider_jobs_changed() — resolve the active
  provider, call on_jobs_changed(), swallow errors. Lives in scheduler.py (not
  jobs.py) so the store stays free of provider imports (no import cycle).
- Wired at the consumer surfaces AFTER a successful mutation: the cronjob model
  tool (tools/cronjob_tools.py, create/update/remove/pause/resume) — which the
  `hermes cron` CLI also routes through — and the REST handlers
  (gateway/platforms/api_server.py, same five). Built-in's no-op default = zero
  behavior change on the default path. Sleeping-agent direct jobs.json writes
  (no tool/CLI/REST) are covered by reconcile-on-wake in start().

F.2 — config: cron.chronos.{portal_url,callback_url,expected_audience,
nas_jwks_url}. All non-secret; the agent holds no scheduler creds and the
outbound provision call reuses the existing Nous token (no token key). Additive
deep-merge key, no version literal.

F.3 — docs:
- docs/chronos-managed-cron-contract.md: authoritative agent↔NAS wire contract
  (the three agent-cron endpoints + inbound /api/cron/fire + the 3-hop trust
  model + at-most-once/re-arm semantics). This is what the NAS-side agent builds
  against.
- cron-internals.md: "Managed cron (Chronos) for scale-to-zero" section.
- cli-commands.md: cron.provider accepts chronos + the cron.chronos.* keys.
- User docs name no scheduler vendor (QStash is a NAS-internal detail).

INVARIANT re-verified: zero qstash/upstash hits across plugins/cron, gateway,
hermes_cli, tools, website/docs (the one remaining repo hit is an unrelated
Context7 MCP comment in tools/mcp_tool.py).

Tests: test_jobs_changed_notify (5) — notify calls provider hook, swallows
errors, built-in harmless, tool create/remove notify. Full cron + chronos +
webhook + config + api_server_jobs suites green (504 in the cron+chronos+webhook
run).
2026-06-18 15:11:32 +10:00
Ben
bfb6e0bb33 docs(cron): document CronScheduler provider + cron.provider key
Phase 3.5. cron-internals.md gateway-integration section now describes the
pluggable trigger (resolve_cron_scheduler, built-in default, plugins/cron
discovery, the never-without-a-trigger fallback, and the trigger-vs-execution
split). cli-commands.md notes cron.provider near the hermes cron entry.
2026-06-18 14:18:31 +10:00
Teknium
1d5deac346 fix(website): cross-locale doc links + drop empty ko locale (#31895)
The locale switcher appeared broken because hardcoded markdown links
(`](/docs/X)`) got double-prefixed by Docusaurus to `/docs/<locale>/docs/X`
(404) in non-English locales, and the MDX hero `<a href>` on the index
page escaped locale routing entirely.

Changes:
- Rewrite 922 `](/docs/X)` -> `](/X)` across 166 docs files (strip trailing
  .md too). Docusaurus prepends locale + baseUrl itself.
- docs/index.md -> index.mdx; hero "Get Started" anchor -> Docusaurus
  <Link> so it stays inside the active locale.
- Drop `ko` locale entirely from docusaurus.config.ts + delete i18n/ko/
  (4 stale auto-translated kanban pages, <2% coverage, misleading).

Verified `npm run build` succeeds for both en and zh-Hans; `build/zh-Hans/
index.html` has no /docs/zh-Hans/docs/... double-prefixed paths.

PR2 will translate the 335 English docs into i18n/zh-Hans/.
2026-05-24 23:16:20 -07:00
Teknium
289cc47631 docs: resync reference, user-guide, developer-guide, and messaging pages against code (#17738)
Broad drift audit against origin/main (b52b63396).

Reference pages (most user-visible drift):
- slash-commands: add /busy, /curator, /footer, /indicator, /redraw, /steer
  that were missing; drop non-existent /terminal-setup; fix /q footnote
  (resolves to /queue, not /quit); extend CLI-only list with all 24
  CLI-only commands in the registry
- cli-commands: add dedicated sections for hermes curator / fallback /
  hooks (new subcommands not previously documented); remove stale
  hermes honcho standalone section (the plugin registers dynamically
  via hermes memory); list curator/fallback/hooks in top-level table;
  fix completion to include fish
- toolsets-reference: document the real 52-toolset count; split browser
  vs browser-cdp; add discord / discord_admin / spotify / yuanbao;
  correct hermes-cli tool count from 36 to 38; fix misleading claim
  that hermes-homeassistant adds tools (it's identical to hermes-cli)
- tools-reference: bump tool count 55 -> 68; add 7 Spotify, 5 Yuanbao,
  2 Discord toolsets; move browser_cdp/browser_dialog to their own
  browser-cdp toolset section
- environment-variables: add 40+ user-facing HERMES_* vars that were
  undocumented (--yolo, --accept-hooks, --ignore-*, inference model
  override, agent/stream/checkpoint timeouts, OAuth trace, per-platform
  batch tuning for Telegram/Discord/Matrix/Feishu/WeCom, cron knobs,
  gateway restart/connect timeouts); dedupe the Cron Scheduler section;
  replace stale QQ_SANDBOX with QQ_PORTAL_HOST

User-guide (top level):
- cli.md: compression preserves last 20 turns, not 4 (protect_last_n: 20)
- configuration.md: display.platforms is the canonical per-platform
  override key; tool_progress_overrides is deprecated and auto-migrated
- profiles.md: model.default is the config key, not model.model
- sessions.md: CLI/TUI session IDs use 6-char hex, gateway uses 8
- checkpoints-and-rollback.md: destructive-command list now matches
  _DESTRUCTIVE_PATTERNS (adds rmdir, cp, install, dd)
- docker.md: the container runs as non-root hermes (UID 10000) via
  gosu; fix install command (uv pip); add missing --insecure on the
  dashboard compose example (required for non-loopback bind)
- security.md: systemctl danger pattern also matches 'restart'
- index.md: built-in tool count 47 -> 68
- integrations/index.md: 6 STT providers, 8 memory providers
- integrations/providers.md: drop fictional dashscope/qwen aliases

Features:
- overview.md: 9 image models (not 8), 9 TTS providers (not 5),
  8 memory providers (Supermemory was missing)
- tool-gateway.md: 9 image models
- tools.md: extend common-toolsets list with search / messaging /
  spotify / discord / debugging / safe
- fallback-providers.md: add 6 real providers from PROVIDER_REGISTRY
  (lmstudio, kimi-coding-cn, stepfun, alibaba-coding-plan,
  tencent-tokenhub, azure-foundry)
- plugins.md: Available Hooks table now includes on_session_finalize,
  on_session_reset, subagent_stop
- built-in-plugins.md: add the 7 bundled plugins the page didn't
  mention (spotify, google_meet, three image_gen providers, two
  dashboard examples)
- web-dashboard.md: add --insecure and --tui flags
- cron.md: hermes cron create takes positional schedule/prompt, not
  flags

Messaging:
- telegram.md: TELEGRAM_WEBHOOK_SECRET is now REQUIRED when
  TELEGRAM_WEBHOOK_URL is set (gateway refuses to start without it
  per GHSA-3vpc-7q5r-276h). Biggest user-visible drift in the batch.
- discord.md: HERMES_DISCORD_TEXT_BATCH_SPLIT_DELAY_SECONDS default
  is 2.0, not 0.1
- dingtalk.md: document DINGTALK_REQUIRE_MENTION /
  FREE_RESPONSE_CHATS / MENTION_PATTERNS / HOME_CHANNEL /
  ALLOW_ALL_USERS that the adapter supports
- bluebubbles.md: drop fictional BLUEBUBBLES_SEND_READ_RECEIPTS env
  var; the setting lives in platforms.bluebubbles.extra only
- qqbot.md: drop dead QQ_SANDBOX; add real QQ_PORTAL_HOST and
  QQ_GROUP_ALLOWED_USERS
- wecom-callback.md: replace 'hermes gateway start' (service-only)
  with 'hermes gateway' for first-time setup

Developer-guide:
- architecture.md: refresh tool/toolset counts (61/52), terminal
  backend count (7), line counts for run_agent.py (~13.7k), cli.py
  (~11.5k), main.py (~10.4k), setup.py (~3.5k), gateway/run.py
  (~12.2k), mcp_tool.py (~3.1k); add yuanbao adapter, bump platform
  adapter count 18 -> 20
- agent-loop.md: run_agent.py line count 10.7k -> 13.7k
- tools-runtime.md: add vercel_sandbox backend
- adding-tools.md: remove stale 'Discovery import added to
  model_tools.py' checklist item (registry auto-discovery)
- adding-platform-adapters.md: mark send_typing / get_chat_info as
  concrete base methods; only connect/disconnect/send are abstract
- acp-internals.md: ACP sessions now persist to SessionDB
  (~/.hermes/state.db); acp.run_agent call uses
  use_unstable_protocol=True
- cron-internals.md: gateway runs scheduler in a dedicated background
  thread via _start_cron_ticker, not on a maintenance cycle; locking
  is cross-process via fcntl.flock (Unix) / msvcrt.locking (Windows)
- gateway-internals.md: gateway/run.py ~12k lines
- provider-runtime.md: cron DOES support fallback (run_job reads
  fallback_providers from config)
- session-storage.md: SCHEMA_VERSION = 11 (not 9); add migrations
  10 and 11 (trigram FTS, inline-mode FTS5 re-index); add
  api_call_count column to Sessions DDL; document messages_fts_trigram
  and state_meta in the architecture tree
- context-compression-and-caching.md: remove the obsolete 'context
  pressure warnings' section (warnings were removed for causing
  models to give up early)
- context-engine-plugin.md: compress() signature now includes
  focus_topic param
- extending-the-cli.md: _build_tui_layout_children signature now
  includes model_picker_widget; add to default layout

Also fixed three pre-existing broken links/anchors the build warned
about (docker.md -> api-server.md, yuanbao.md -> cron-jobs.md and
tips#background-tasks, nix-setup.md -> #container-aware-cli).

Regenerated per-skill pages via website/scripts/generate-skill-docs.py
so catalog tables and sidebar are consistent with current SKILL.md
frontmatter.

docusaurus build: clean, no broken links or anchors.
2026-04-29 20:55:59 -07:00
Teknium
1acf81fdf5 docs: add QQBot to all 14 docs pages (full platform parity)
- sidebars.ts: sidebar navigation entry
- webhooks.md: deliver field routing table
- configuration.md: platform keys list
- sessions.md: platform identifiers table
- features/cron.md: delivery target table
- developer-guide/architecture.md: adapter listing
- developer-guide/cron-internals.md: delivery target table
- developer-guide/gateway-internals.md: file tree listing
- guides/cron-troubleshooting.md: supported platforms list
- integrations/index.md: platform links list
- reference/toolsets-reference.md: toolset table

(qqbot.md, environment-variables.md, and messaging/index.md were
already included in the contributor's original PR)
2026-04-14 00:11:49 -07:00
Teknium
ba50fa3035 docs: fix 30+ inaccuracies across documentation (#9023)
Cross-referenced all docs pages against the actual codebase and fixed:

Reference docs (cli-commands.md, slash-commands.md, profile-commands.md):
- Fix: hermes web -> hermes dashboard (correct subparser name)
- Fix: Wrong provider list (removed deepseek, ai-gateway, opencode-zen,
  opencode-go, alibaba; added gemini)
- Fix: Missing tts in hermes setup section choices
- Add: Missing --image flag for hermes chat
- Add: Missing --component flag for hermes logs
- Add: Missing CLI commands: debug, backup, import
- Fix: /status incorrectly marked as messaging-only (available everywhere)
- Fix: /statusbar moved from Session to Configuration category
- Add: Missing slash commands: /fast, /snapshot, /image, /debug
- Add: Missing /restart from messaging commands table
- Fix: /compress description to match COMMAND_REGISTRY
- Add: --no-alias flag to profile create docs

Configuration docs (configuration.md, environment-variables.md):
- Fix: Vision timeout default 30s -> 120s
- Fix: TTS providers missing minimax and mistral
- Fix: STT providers missing mistral
- Fix: TTS openai base_url shown with wrong default
- Fix: Compression config showing stale summary_model/provider/base_url
  keys (migrated out in config v17) -> target_ratio/protect_last_n

Getting-started docs:
- Fix: Redundant faster-whisper install (already in voice extra)
- Fix: Messaging extra description missing Slack

Developer guide:
- Fix: architecture.md tool count 48 -> 47, toolset count 40 -> 19
- Fix: run_agent.py line count 9,200 -> 10,700
- Fix: cli.py line count 8,500 -> 10,000
- Fix: main.py line count 5,500 -> 6,000
- Fix: gateway/run.py line count 7,500 -> 9,000
- Fix: Browser tools count 11 -> 10
- Fix: Platform adapter count 15 -> 18 (add wecom_callback, api_server)
- Fix: agent-loop.md wrong budget sharing (not shared, independent)
- Fix: agent-loop.md non-existent _get_budget_warning() reference
- Fix: context-compression-and-caching.md non-existent function name
- Fix: toolsets-reference.md safe toolset includes mixture_of_agents (it doesn't)
- Fix: toolsets-reference.md hermes-cli tool count 38 -> 36

Guides:
- Fix: automate-with-cron.md claims daily at 9am is valid (it's not)
- Fix: delegation-patterns.md Max 3 presented as hard cap (configurable)
- Fix: sessions.md group thread key format (shared by default, not per-user)
- Fix: cron-internals.md job ID format and JSON structure
2026-04-13 10:53:10 -07:00
Teknium
7cec784b64 fix: complete Weixin platform parity audit — 16 missing integration points
Systematic audit found Weixin missing from:

Code:
- gateway/run.py: early WEIXIN_ALLOW_ALL_USERS env check
- gateway/platforms/webhook.py: cross-platform delivery routing
- hermes_cli/dump.py: platform detection for config export
- hermes_cli/setup.py: hermes setup wizard platform list + _setup_weixin
- hermes_cli/skills_config.py: platform labels for skills config UI

Docs (11 pages):
- developer-guide/architecture.md: platform adapter listing
- developer-guide/cron-internals.md: delivery target table
- developer-guide/gateway-internals.md: file tree
- guides/cron-troubleshooting.md: supported platforms list
- integrations/index.md: platform links
- reference/toolsets-reference.md: toolset table
- user-guide/configuration.md: platform keys for tool_progress
- user-guide/features/cron.md: delivery target table
- user-guide/messaging/index.md: intro text, feature table,
  mermaid diagram, toolset table, setup links
- user-guide/messaging/webhooks.md: deliver field + routing table
- user-guide/sessions.md: platform identifiers table
2026-04-10 05:54:37 -07:00
Teknium
95ee453bc0 docs: add cron script timeout and provider recovery documentation
- Add HERMES_CRON_TIMEOUT and HERMES_CRON_SCRIPT_TIMEOUT to env vars reference
- Add script timeout and provider recovery sections to cron features page
- Add timeout resolution chain and credential pool details to cron internals
2026-04-10 02:57:57 -07:00
Teknium
7120d6cdd6 fix(bluebubbles): add missing integration points and documentation (#6460)
- hermes_cli/skills_config.py: add platform label for per-platform skill config
- gateway/session.py: add to PII-safe platforms (no mention system)
- website/docs/user-guide/messaging/bluebubbles.md: full setup guide
- website/sidebars.ts: sidebar navigation entry
- 10 docs pages: add BlueBubbles to all platform enumerations
  (env vars, toolsets, cron delivery, gateway internals, etc.)
2026-04-09 00:19:05 -07:00
Teknium
c58e16757a docs: fix 40+ discrepancies between documentation and codebase (#5818)
Comprehensive audit of all ~100 doc pages against the actual code, fixing:

Reference docs:
- HERMES_API_TIMEOUT default 900 -> 1800 (env-vars)
- TERMINAL_DOCKER_IMAGE default python:3.11 -> nikolaik/python-nodejs (env-vars)
- compression.summary_model default shown as gemini -> actually empty string (env-vars)
- Add missing GOOGLE_API_KEY, GEMINI_API_KEY, GEMINI_BASE_URL env vars (env-vars)
- Add missing /branch (/fork) slash command (slash-commands)
- Fix hermes-cli tool count 39 -> 38 (toolsets-reference)
- Fix hermes-api-server drop list to include text_to_speech (toolsets-reference)
- Fix total tool count 47 -> 48, standalone 14 -> 15 (tools-reference)

User guide:
- web_extract.timeout default 30 -> 360 (configuration)
- Remove display.theme_mode (not implemented in code) (configuration)
- Remove display.background_process_notifications (not in defaults) (configuration)
- Browser inactivity timeout 300/5min -> 120/2min (browser)
- Screenshot path browser_screenshots -> cache/screenshots (browser)
- batch_runner default model claude-sonnet-4-20250514 -> claude-sonnet-4.6
- Add minimax to TTS provider list (voice-mode)
- Remove credential_pool_strategies from auth.json example (credential-pools)
- Fix Slack token path platforms/slack/ -> root ~/.hermes/ (slack)
- Fix Matrix store path for new installs (matrix)
- Fix WhatsApp session path for new installs (whatsapp)
- Fix HomeAssistant config from gateway.json to config.yaml (homeassistant)
- Fix WeCom gateway start command (wecom)

Developer guide:
- Fix tool/toolset counts in architecture overview
- Update line counts: main.py ~5500, setup.py ~3100, run.py ~7500, mcp_tool ~2200
- Replace nonexistent agent/memory_store.py with memory_manager.py + memory_provider.py
- Update _discover_tools() list: remove honcho_tools, add skill_manager_tool
- Add session_search and delegate_task to intercepted tools list (agent-loop)
- Fix budget warning: two-tier system (70% caution, 90% warning) (agent-loop)
- Fix gateway auth order (per-platform first, global last) (gateway-internals)
- Fix email_adapter.py -> email.py, add webhook.py + api_server.py (gateway-internals)
- Add 7 missing providers to provider-runtime list

Other:
- Add Docker --cap-add entries to security doc
- Fix Python version 3.10+ -> 3.11+ (contributing)
- Fix AGENTS.md discovery claim (not hierarchical walk) (tips)
- Fix cron 'add' -> canonical 'create' (cron-internals)
- Add pre_api_request/post_api_request hooks to plugin guide
- Add Google/Gemini provider to providers page
- Clarify OPENAI_BASE_URL deprecation (providers)
2026-04-07 10:17:44 -07:00
Teknium
43d468cea8 docs: comprehensive documentation audit — fix stale info, expand thin pages, add depth (#5393)
Major changes across 20 documentation pages:

Staleness fixes:
- Fix FAQ: wrong import path (hermes.agent → run_agent)
- Fix FAQ: stale Gemini 2.0 model → Gemini 3 Flash
- Fix integrations/index: missing MiniMax TTS provider
- Fix integrations/index: web_crawl is not a registered tool
- Fix sessions: add all 19 session sources (was only 5)
- Fix cron: add all 18 delivery targets (was only telegram/discord)
- Fix webhooks: add all delivery targets
- Fix overview: add missing MCP, memory providers, credential pools
- Fix all line-number references → use function name searches instead
- Update file size estimates (run_agent ~9200, gateway ~7200, cli ~8500)

Expanded thin pages (< 150 lines → substantial depth):
- honcho.md: 43 → 108 lines — added feature comparison, tools, config, CLI
- overview.md: 49 → 55 lines — added MCP, memory providers, credential pools
- toolsets-reference.md: 57 → 175 lines — added explanations, config examples,
  custom toolsets, wildcards, platform differences table
- optional-skills-catalog.md: 74 → 153 lines — added 25+ missing skills across
  communication, devops, mlops (18!), productivity, research categories
- integrations/index.md: 82 → 115 lines — added messaging, HA, plugins sections
- cron-internals.md: 90 → 195 lines — added job JSON example, lifecycle states,
  tick cycle, delivery targets, script-backed jobs, CLI interface
- gateway-internals.md: 111 → 250 lines — added architecture diagram, message
  flow, two-level guard, platform adapters, token locks, process management
- agent-loop.md: 112 → 235 lines — added entry points, API mode resolution,
  turn lifecycle detail, message alternation rules, tool execution flow,
  callback table, budget tracking, compression details
- architecture.md: 152 → 295 lines — added system overview diagram, data flow
  diagrams, design principles table, dependency chain

Other depth additions:
- context-references.md: added platform availability, compression interaction,
  common patterns sections
- slash-commands.md: added quick commands config example, alias resolution
- image-generation.md: added platform delivery table
- tools-reference.md: added tool counts, MCP tools note
- index.md: updated platform count (5 → 14+), tool count (40+ → 47)
2026-04-05 19:45:50 -07:00
teknium1
c3ea620796 feat: add multi-skill cron editing and docs 2026-03-14 19:18:10 -07:00
teknium1
d87a1615ce docs: add ACP and internal systems implementation guides
- add ACP user and developer docs covering setup, lifecycle, callbacks,
  permissions, tool rendering, and runtime behavior
- add developer guides for agent loop, provider runtime resolution,
  prompt assembly, context caching/compression, gateway internals,
  session storage, tools runtime, trajectories, and cron internals
- refresh architecture, quickstart, installation, CLI reference, and
  environments docs to link the new implementation pages and ACP support
2026-03-14 00:29:48 -07:00