Commit Graph

1194 Commits

Author SHA1 Message Date
teknium1
004c8a51f0 feat(api-server): reasoning on non-streaming replies; echoed reasoning input items are ignored
Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:

- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
  `choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
  completed `reasoning` output item ahead of the message (and ahead of that step's
  `function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
  from the assistant messages the agent already persisted (`build_assistant_message`
  stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
  in non-stream mode the agent fires the callback per provider delta AND once more with
  the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
  `conversation_history`. Responses SDK clients replay a prior response's `output`
  list as the next `input`; before, the item became an empty `user` message (in
  `input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
  (`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
  `output_item.done` item and the non-streaming shape.

Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
2026-09-19 00:46:31 -07:00
teknium1
b957782917 feat(api-server): stream model reasoning on /v1/chat/completions and /v1/responses
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).

- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
  `_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
  keeping them distinct from answer text. The lossy 500-char `reasoning.available`
  progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
  (the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
  (`output_item.added`, `reasoning_summary_part.added`,
  `reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
  `output_item.done`), closed before the next message/function_call item opens
  and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.

Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
2026-09-18 23:58:24 -07:00
teknium1
40778c74e2 fix: MoA aggregator merges adjacent same-role messages only for destinations that reject them
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.

Reactive, destination-scoped recovery instead:

- error_classifier: new `FailoverReason.role_alternation` for the vendor
  alternation wordings (checked before the request-validation table since
  the body also carries `invalid_request_error`); same abort+fallback hints
  as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
  loop's `_merge_user_content`), `destination_key` (base_url|provider,
  model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
  adjacent user turns merged, remember the destination on the facade for
  the session so later iterations pre-merge, never touch destinations that
  accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.

Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.

Fixes #112358
2026-09-18 20:56:35 -07:00
teknium1
b71359066c fix: /stop halts background subagents and returns their partial results as interrupted completions
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:

- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
  already did this in `_handle_reset_command`, the second request is
  idempotent) and the idle `_handle_stop_command` tail, which replied
  "No active task to stop." while a background child was running; it now
  stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
  spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.

The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.

An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.

Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.

Part of #114456
2026-09-18 20:34:42 -07:00
teknium1
2b94b0d40f fix(kanban): dispatcher blocks a card on the first terminal provider error
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).

Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.

`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.

No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.

Part of #114587

Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
2026-09-18 20:34:16 -07:00
teknium1
6babdc96b8 fix(image-gen): every token-billed image backend records session usage (chat path, Image API, OpenAI gpt-image)
Widen the salvaged Image API hunk (#114340) to the whole class:

- plugins/image_gen/_common.py: record_token_usage() — one helper feeding the
  aux accounting chokepoint (agent.aux_accounting.record_aux_usage) with task
  "image_generation", the billing provider and the priced model id. Dict and
  SDK usage objects alike; a body without tokens is a no-op, as is a call
  outside a turn.
- plugins/image_gen/openrouter: the /chat/completions path (the DEFAULT model
  chain — openai/gpt-5.4-image-2, google/gemini-3-pro-image — is chat-only and
  token-billed) now records too; the Image API path uses the shared helper
  (the contributor's local _record_image_api_usage is folded into it) and both
  pass base_url so pricing resolves the route. Task renamed image_gen ->
  image_generation to match the other aux task names.
- plugins/image_gen/openai: gpt-image bills per text/image token; record the
  Images API usage block under the API model (gpt-image-2), not the Hermes
  quality-tier label.
- FAL, xAI, Krea, DeepInfra, Meta, openai-codex return no token usage and stay
  unrecorded (nothing to bill per token).
- tests: the contributor's two tests folded into one invariant parametrized
  over chat / Image API / no-usage control; one OpenAI invariant.
- docs: image-generation.md "How It Works Internally" gains the accounting step.

Fixes #114324
2026-09-18 19:56:51 -07:00
teknium1
a6ad3cc3ff fix(review): derive the default input budget from the fork's resolved context window
The salvaged commit (#114651) re-probed the endpoint through get_model_context_length
at spawn time. The review fork is a full AIAgent whose context_compressor.context_length
is already resolved by the same chain (config override, catalog, persisted provider limit,
endpoint probe), so read that instead: no second lookup, and the budget tracks exactly
the window the fork's own compaction uses. Tests trimmed to two invariants on the new seam.

Docs: `auxiliary.background_review.max_input_tokens` and its derived default in
website/docs/user-guide/features/memory.md, including the note that a top-level
`background_review:` block is not read (#114645 secondary).

Co-authored-by: Artie Fisher <artiefisher123@gmail.com>
2026-09-18 19:18:57 -07:00
teknium1
420feb6307 docs: session-bound Desktop/TUI settings writes and dashboard launch-profile fallback are per profile 2026-09-18 15:12:27 -07:00
teknium1
d15208e5f0 docs(website): re-run the link sweep over pages merged since the rebase
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
2026-09-18 14:27:04 -07:00
teknium1
2fbcd8b0ea docs(website): link pages by relative Markdown path so they open on GitHub (#114428)
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.

Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).

Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
2026-09-18 14:27:04 -07:00
teknium1
f0c26f0554 fix(mcp): name the HTTP status, URL and body behind "Server returned an error response"
mcp >= 2.0's Streamable HTTP client folds any non-2xx whose body it cannot parse
as a JSON-RPC error into the opaque `-32603 Server returned an error response`.
Hermes printed that text verbatim in the SSE-fallback warning and in the
"both transports failed" ConnectionError, so users saw no status, no URL and
none of the server's own words (e.g. `400 {"code":-32020,"message":"Unsupported
MCP-Protocol-Version"}`) and had to reach for curl to learn what the server
actually said (#114350, #113359).

- `_make_http_rejection_recorder`: response hook on the owned SDK-httpx client
  that remembers the last 4xx/5xx (status, method, URL, head of the body; SSE
  bodies are never read). Sibling of the redirect-header stripper hook.
- `_describe_http_failure`: appends that detail only when the root cause is
  the SDK's opaque -32603 text, so a real JSON-RPC error or an httpx status
  error is never duplicated.
- `_run_http`: the fallback warning and both raise paths (both-transports
  ConnectionError; the no-fallback re-raise after a proven session, strict
  redirect headers or a non-rejection) carry the detail. Debug-logs the
  endpoint each connect attempt uses.
- Docs: troubleshooting entry for reading the new message.

Verified live against a real Streamable HTTP server (`hermes mcp test`, temp
HERMES_HOME): base prints `Streamable HTTP: Server returned an error response`;
fixed head prints `... (HTTP 400 from POST http://127.0.0.1:PORT/mcp:
{"jsonrpc":"2.0","id":null,"error":{"code":-32020,...}})`. Control: a 400
served as application/json already surfaces the JSON-RPC message and gets no
appendix; servers negotiating `initialize` down to 2025-06-18 (fixture and a
real FastMCP on mcp 1.12.4) connect and list tools on base and fix alike — the
pinned mcp 2.0.0 stamps the negotiated version on every post-handshake request
(wire-recorded), so the sticky seed in `_run_http` is not the cause of the
reported 400 and stays as designed (#14816).

Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
2026-09-18 12:48:42 -07:00
teknium1
e21a9e502d fix(kanban): a breaker trip is not promoted back to ready in the same tick
``_account_crashes`` trips the protocol-violation budget with
``force_trip`` after three consecutive clean exits, but the task's
``consecutive_failures`` is still 1 — so ``recompute_ready``, which
re-derives the threshold from its caller's ``failure_limit`` (default 2),
promoted the just-blocked card straight back to ``ready`` in the same
tick: ``protocol_violation -> gave_up -> promoted -> claimed`` forever
whenever ``failure_limit`` exceeds the violation count (the systemic
same-error trip at limit 1 had the same hole).

``_has_sticky_block`` now also honours the breaker's own verdict: the
newest ``gave_up`` since the last ``unblocked`` holds the card when it
recorded a violation-streak trip or ``failures >= effective_limit``.
``hermes kanban unblock`` remains the release path and still grants a
fresh budget. Tests cover the fresh-process classification (rc=0 and
rc=75), the held trip + operator unblock, and the trailer emission gate.
2026-09-18 12:48:16 -07:00
teknium1
edd4ec3124 fix(kanban): book a dead worker the same way whichever process notices it
The dead-worker sweep learned a worker's exit status only from
``_recent_worker_exits``, which ``reap_worker_zombies`` fills via
``os.waitpid`` — so only the process that spawned the worker ever knows
how it exited. With ``kanban.dispatch_in_gateway: false`` every
``hermes kanban dispatch`` tick is a fresh process, the registry is
empty, and a worker that exited rc=0 without a terminal board call was
booked as a bare ``crashed`` / ``pid N not alive``: no
``protocol_violation`` marker, no corrective error text for the retry
worker, no violation streak — and a rc=75 quota wall was counted as a
failure instead of a neutral ``rate_limited`` requeue.

A Kanban worker (``HERMES_KANBAN_TASK`` set) now writes
``[kanban-worker-exit] rc=<code>`` as the last line of its own log on
every one-shot exit path; when the registry has no entry for a dead PID
the sweep reads that trailer and books the exit through the same
code -> kind mapping. A worker killed before its exit epilogue leaves no
trailer and stays a plain crash. ``_worker_final_output`` strips the
trailer so it never leaks into the board diagnostic.

Direction (a durable, process-independent witness in the worker log)
from PR #113638; its predicate keyed on the ``Resume this session with:``
summary, which the CLI prints before ``sys.exit`` for rc 0, 1, 75 and 130
alike and so would have booked quota walls and failed turns as protocol
violations.

Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
2026-09-18 12:48:16 -07:00
teknium1
5179b5177d fix(dashboard): say when the scheduler last ticked on the Cron page
#114309 expects the dashboard, like `hermes cron status`, to say "scheduler last
ticked X hours ago" when a stopped ticker leaves next runs stranded in the past.
The Cron page only had the per-row `Overdue since` label and no way to date the
outage: GET /api/cron/jobs carried nothing about the ticker.

Every job the dashboard cron endpoints return now carries
`scheduler_heartbeat_age_s` — its own profile's ticker heartbeat age read inside
the same store scope as the job list (None when it cannot be dated) — and the Cron
page renders "Scheduler last ticked 7h ago — jobs that came due since then have not
fired" above the list when the oldest heartbeat among jobs expected to fire is
past the CLI's STALE_AFTER threshold (three missed 60s iterations plus slack).
Paused, disabled and completed jobs never raise the banner, matching the overdue
label's rule. The field is additive, so the response shape and every existing
consumer are unchanged.

Tests: one router test lists two profiles with one stale heartbeat file and checks
each job reports its own age (None for the profile without a heartbeat); one vitest
pins the stale-age selection and the relative label.
2026-09-18 12:46:54 -07:00
teknium1
00b7997f47 fix(desktop): label an overdue next run on the Bot Mode routine card
The hermes-bots plugin's Routines card (and its inspector's `Next run` row)
still read `Next: 7 hr. ago` for a `next_run_at` parked past the scheduler
grace — the one Desktop surface the overdue labelling left out (#114309).
The card now switches to `t.cron.overdueSince` and the inspector row to
`Overdue since`, using the same `nextRunOverdueMs` decision as the core cron
panel and sidebar, so paused/completed/disabled jobs keep the plain label.

`nextRunOverdueMs` is exported through `@hermes/plugin-sdk` (the plugin only
imports from the SDK) and its parameter is widened to the optional-field shape
the plugin's `RoutineJob` carries; no plugin i18n key is needed because the
card already renders from core `t.cron`, which every locale resolves.

Test: one card test renders an overdue vs an upcoming job and checks both the
card label and the inspector rows; the fixture's fixed 2026-08-23 next_run_at
is made clock-relative so it never ages into "overdue" on its own.
2026-09-18 12:46:54 -07:00
teknium1
eddf7283e4 fix(cron): label overdue next runs in cron list, dashboard and Desktop; one parser for the instant
`hermes cron list`, the web dashboard Cron page and the Desktop cron panel/sidebar all
rendered a `next_run_at` parked hours in the past as an ordinary upcoming "Next run" —
the only user-visible trace of a scheduler that stopped ticking (#114309). Every
surface now labels a slot past `cron doctor`'s 15-minute grace as overdue (CLI
`Overdue:` row with the lateness, web `Overdue since`, Desktop `Overdue since` label
on the detail panel and sidebar meta) while paused/disabled/completed jobs keep the
plain label because they are not expected to fire.

The CLI's overdue check now parses through `cron.jobs._parse_aware` and `hermes_time.now`
so status/list/doctor and the ticker agree on the instant (mixed offsets, DST folds,
legacy naive stamps read as system-local like the scheduler does), and both the
ordering and the subtraction normalise to UTC: Python compares same-tzinfo datetimes by
wall clock, which is wrong across a DST fold. Status still orders the soonest run by
instant (#113874) and prints the stored stamp.

Tests: the salvaged status tests are trimmed to two invariants (overdue + stale
heartbeat is loud on status AND list; within-grace stays plain on both), the DST-fold
ordering tests freeze the CLI clock as well so their 2026-11 fixtures never start
reading as overdue, and one vitest each pins the web and Desktop helpers.

Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-18 12:46:54 -07:00
teknium1
63fb1a7d36 fix(kanban): warn once when an inherited delegation fence rejects the worker's auto-heartbeat
A worker process that carries HERMES_DELEGATED_CHILD_CONTEXT next to its
HERMES_KANBAN_TASK is fenced by kanban_path_is_fenced's env-marker branch: both
auto-heartbeat writes raised PermissionError at DEBUG only, so the board showed a
worker that never beats while the process was alive (the salvaged commit already
made the return value honest, but _touch_activity ignores it). Log the refusal once
per process at WARNING with the cause, and skip the bridge for an in-process
delegate child BEFORE stamping the rate-limit window so a chatty child cannot starve
the worker's own heartbeat.

No self-write grant is added: the marker + TASK combination means "descendant, not
worker" by design (b578261584, 8a8c3634e8), and the only in-tree grant path,
kanban_db_dispatch._default_spawn, already pops the marker. The real-spawn grant
test now proves that from a dispatcher that itself carries the marker: the worker's
auto-heartbeat lands and its handoff completes.

Docs: worker guide notes the automatic claim extension and what the refusal warning
means.
2026-09-18 12:46:28 -07:00
teknium1
63f09c2d2c fix(cli): /cron run reports a refused run instead of "Triggered … next scheduler tick"
`_cron_job_action` printed `(^_^)b Triggered job … It will run on the next
scheduler tick.` unconditionally for `run`, so a run-now the tool refused
(`execution_skipped`: paused job, claim held by another run) was
indistinguishable from an accepted one. Print the refusal, and otherwise
reuse `hermes_cli.cron._run_outcome` — the same verdict line `hermes cron run`
already prints (ran now / background / skipped).

Tests: one invariant per atom — the real one-shot CLI shape
(`_SESSION_ASYNC_DELIVERY` scoped False) reaps a dead-owner 'running' ledger
row before the sync fallback; `/cron run` on a refused claim prints the reason
and never "Triggered". Both red on origin/main. Replaces the #113938 test hunk
(it relocated the prior test's assertion). Docs: execution-history section
notes the pre-manual-run reap.

Part of #113923.
2026-09-18 11:07:56 -07:00
teknium1
e052500722 docs(desktop): the Kanban board switcher sits in the page header, not the window title bar
/kanban is a workspace page route: the titlebar slot is hidden there and the
switcher is projected into the Kanban page's header row (WORKSPACE_PAGE_HEADER_AREA
after #114960). Point readers, the i18n comment and the test comment at that row.
2026-09-18 10:53:00 -07:00
teknium1
d235ae0991 fix(desktop): Kanban board switcher reads as a control — icon, tooltip, i18n
Follow-up to the salvaged #114647 hunk (visible "Board" label +
`Board: <name>` accessible name):

- replace the native `title` with the app's `Tip` ("Switch board"), the
  same tooltip every other dropdown trigger uses, so hover names the
  ACTION instead of repeating the board name
- lead the trigger with the plugin's `project` icon so the title-bar
  text reads as the Kanban board control, not a static window title
- add `switchBoard` to the kanban plugin bundle (en/ja/zh/zh-hant)
- test registers the real locale bundle and asserts the rendered label,
  accessible name and tooltip (the raw-key fallback matched
  `/board:/i` by coincidence)
- docs: describe the Desktop switcher and its local selection

Fixes #114642
2026-09-18 10:53:00 -07:00
teknium1
92bb5b92b8 fix: revert a quota-benched credential through the pool, not a private probe; cover /model and chained rotations
Reshape of the salvaged fix from #114513 (@whyyagswhy):

- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
  benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
  refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
  or round-robin order. The contributor's version called the private
  ``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
  and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
  429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
  the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
  the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
  by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
  hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
  ``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
  fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
  and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.
2026-09-18 10:35:37 -07:00
teknium1
9cc7886afe fix(docs): kanban attachment removal keeps a blob other rows still reference
The File attachments section still promised that removing an attachment always deletes the on-disk file. With reference-counted blobs the file is only unlinked once no other attachment row points at it, so the user guide now states that condition and names the operator CLI path (hermes kanban attach-rm ATTACHMENT_ID).
2026-09-18 10:34:31 -07:00
teknium1
4ce832eeb9 fix(cron): bot-chat delivery cap bounds the bot's turn, not its exit linger
A cron job delivering to Bot Chat booked "timed out after 600s" for turns
that finished in seconds: cron/scheduler_delivery.py::_deliver_to_bot_chat
waited for the `hermes chat -Q` child to EXIT, but when the bot's turn had
messaged a teammate (message_agent -> notify_on_complete runner) the child
then runs the one-shot exit linger, bounded by
terminal.oneshot_completion_wait_seconds (default 600) — the same default as
cron.bot_chat_delivery_timeout_seconds. The cap counts from the claim, the
linger only starts when the turn ends, so the cap expired first on every
such delivery, booked a completed turn as a timeout, held the job's fire
fence for the full cap, and killed the child mid-linger — tearing down the
reply the linger exists to protect (#90879).

Ordering the two bounds cannot fix it (the turn has no duration bound; the
linger has its own contract), so the cap stops competing with the linger:

- hermes_cli/quiet_single_query.py: the -Q child accepts a per-process
  report path (HERMES_QUIET_TURN_REPORT_FILE), popped before the turn like
  HERMES_TURN_AUTHOR so nothing the turn spawns inherits it; cli.py writes
  {pid, exit_code, error} there the moment the turn ends, BEFORE the linger.
- cron: _run_bot_chat_turn polls the child and that report under the cap.
  Report present -> the delivery is booked from it (real exit code and
  stream tails when the child exits within a short grace) and the still-
  lingering child is left running, drained and reaped by a daemon thread.
  No report by the cap -> the turn never ended: killed and booked as a
  timeout, exactly as before. The linger itself is untouched.

Live repro (real _deliver_to_bot_chat, real `hermes chat -Q` against a
loopback provider whose turn spawns `sleep 90` with notify_on_complete,
cap 45s): base books the timeout at 45.7s for a turn that ended at +5s and
kills the child; fixed head books success at 11.2s, the child lingers, the
teammate follow-up turn runs at +97s and the child exits on its own.

Supersedes #113649's cap = delivery + linger (the thread shows a headroom
only moves the race and lengthens the fence hold); analysis credit to the
reporter and the thread's independent verification.

Fixes #113608

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 10:29:13 -07:00
teknium1
d98b1480b9 fix(cli): a benched or signed-out credential prints its reason instead of the first-run provider wizard
`_runtime_credentials_ready()` collapsed "nothing is configured" and "the
configured credential is unusable right now" into one False, so a profile
whose only Nous credential was benched by a failed refresh (or quarantined
DEAD) was told "No inference provider is configured yet" and offered the
provider picker — re-running setup on an OAuth provider with single-use
refresh tokens can rotate the grant away from the session that was working
(#113720).

The readiness probe now returns the resolver's exception too
(`_probe_runtime_credentials`). At CLI startup only the resolver's
`no_provider_configured` verdict reaches the wizard; any other failure is
explained by `_explain_unusable_credentials`: `format_auth_error(exc)`
plus, from the provider's pool, "cooling down ... re-enters rotation in
about Nm" or "sign-in was lost (<reason>); run `hermes auth add <provider>`".
An actually-empty profile still gets the wizard.

Fixes #113720
Supersedes #113732 (@whyyagswhy): its `_inventory_other_providers()` gate
returns False for a Nous-configured profile, so the reported Nous case would
still have reached the wizard, and it printed no reason.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
2026-09-18 10:24:15 -07:00
teknium1
ab59c81f32 docs(skills): project skills follow the TUI/Desktop session workspace
User-visible note for #114359: a placeholder terminal.cwd no longer hides a trusted repo's skills in the TUI/Desktop.
2026-09-18 10:22:47 -07:00
teknium1
eb28bc1bf7 fix(kanban): host spawn refusals stop charging cards; oneshot-unit dispatch scope-wraps workers
Two defects at the same spawn boundary (hermes_cli/kanban_db_dispatch.py::
_restart_safe_worker_argv -> tools/process_registry.py::restart_safe_gateway_child_argv).

#114720 — a managed-gateway dispatcher whose user D-Bus session had gone away
raised the restart-safe-scope RuntimeError at spawn; the ordinary spawn
except-clause fed it to _record_task_failure, which advanced
consecutive_failures and at failure_limit parked the card as a bare `blocked`
(block_kind NULL) with a `gave_up` run. Nothing about the card had run.

- restart_safe_gateway_child_argv raises RestartSafeScopeUnavailable
  (RuntimeError subclass) so the dispatcher can tell a host refusal from a
  card failure.
- _record_task_failure(infrastructure=True): run + event carry
  `infrastructure: true`, consecutive_failures is left alone, the breaker
  never trips; the card stays `ready` with the linger remedy in
  last_failure_error and a warning in the dispatcher log.
- check_respawn_guard returns `infrastructure_cooldown` while the latest run
  is such a refusal inside the rate-limit cooldown window, so a dead bus is
  retried spaced instead of every tick.
- require_restart_safe_scope=True keeps its hard-fail semantics under the
  managed gateway.

#113612 — `hermes kanban dispatch` from an operator's Type=oneshot systemd unit
spawned unmanaged workers: the gate `_is_supervised_gateway_process()` is
False for a CLI dispatcher, so the helper returned the bare argv and the
workers died at the unit's cgroup teardown with empty logs, while the
`running` rows waited for reclaim.

- New outlives_parent=True (kanban only; cron blocks its caller and keeps the
  old gate): under any systemd unit (INVOCATION_ID) a fire-and-forget worker
  is scope-wrapped when the user bus is reachable. Without a bus it degrades
  with a loud once-per-process warning naming the consequence and both
  remedies (enable-linger, KillMode=process) rather than refusing — the
  unit's KillMode/lifetime is unknowable and a long-lived sequencer without
  linger must keep working.

Docs: kanban.md "Workers and systemd cgroups" + guard reasons; cron.md note.

Based on analysis from PR #113624.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-18 10:21:42 -07:00
teknium1
c2bf40d1ce fix(kanban): docs say a bound Project keeps the directory when the field is cleared; UI test pins only the unbind seam
With a bound project the settings dialog omits a blank directory from the PATCH and rename_board mirrors the project's primary folder, so clearing the field cannot fall back to scratch until the project is unbound; the Settings bullet now says so. The bundle test drops five wording-level substring asserts that pinned the payload shape already covered behaviourally by test_kanban_board_project_api.py, keeping only the switcher's unbind seam.
2026-09-18 10:19:54 -07:00
teknium1
0985aeb732 fix(kanban): board switcher shows the bound project with an unbind action
The create/settings dialogs (salvaged from #114664) let a user set or clear
a board's project_id, but the binding was invisible outside Settings.
GET /boards already annotates every board with project_id + project_name,
so the switcher now renders a "Project: <name>" badge whose × sends
PATCH {project_id: ""} — the same clear the settings dialog uses — without
touching default_workdir.

Settings also stops sending default_workdir: "" next to a chosen project
when the directory field is blank: the explicit "" suppressed the server's
project → default_workdir mirror, so binding via Settings left the board
without the workspace default the create dialog would have seeded.

Trim the salvaged test to the invariants (selector wiring, payload shapes,
badge unbind) and document the control on the Kanban docs page.
2026-09-18 10:19:54 -07:00
teknium1
89c6a8a145 fix(auth): Nous "not logged in" refresh failures leave rotation instead of a silent hour-long bench
A Nous pool entry whose refresh reached the resolver's state-shape raises
("Hermes is not logged into Nous Portal.", "No access token found ...",
"No refresh token is available ...") carried relogin_required=True but no
OAuth code, so `OAuthProviderFlow.is_terminal_refresh_error` classified it
transient: `_recover_failed_refresh` benched the only credential for an
hour with null last_error_* fields and nothing at WARNING (#113718).

Give those raises codes (`nous_auth_missing`, `nous_auth_missing_access_token`,
`nous_auth_missing_refresh_token`) and list them in the Nous flow's
terminal_refresh_codes, mirroring the Codex/xAI `*_auth_missing_refresh_token`
convention. The pool then quarantines, marks the row DEAD with the reason,
and logs the existing WARNING naming `hermes auth add nous`; every consumer
of the classifier (pool, singleton `_refresh_nous_or_quarantine`) agrees.
A failure that says nothing about the login (network error, non-JSON 5xx
body) still benches as before.

Fixes #113718
Salvages #113726 (@whyyagswhy) — tests kept, consumer-side hunk replaced
by the classifier fix.
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
2026-09-18 10:17:59 -07:00
teknium1
fe6b5cfd15 docs: explicit auxiliary providers walk their own fallback_chain on auth errors 2026-09-18 10:13:32 -07:00
teknium1
964fbaae2c fix(auth): status snapshot peeks the credential pool instead of leasing it
`get_codex_auth_status` / `get_xai_oauth_auth_status` back every credential-gated
listing (`/model` picker rows, `hermes doctor`, dashboard auth cards). The shared
`_pool_first_oauth_status` read the pool with `select()`, which is a runtime lease:
it refreshes an expiring single-use token and, when that speculative POST fails
transiently (503, endpoint unreachable), benches the entry with a persisted
exhaustion cooldown. The picker then rendered the provider as an unconfigured
skeleton ("needs setup" / "0 models") while the runtime resolver kept serving
the same credential. On a round-robin pool the read also rotated and persisted
the priority order.

Read the pool with `peek()`: no refresh, no accounting, no rotation, no persist.
Refreshing stays with the runtime resolver reached through the `resolve`
fallback, whose failures persist nothing (same design as the Nous status).

Slimmer redo of #114379 by @Finn763 (observe flag threaded through four pool
methods plus non-refreshing resolvers); fixtures adapted from that PR.

Live: pool-only openai-codex entry, expired access token, stubbed token endpoint
returning 503 — before: 1 refresh POST, entry persisted `exhausted`,
`has_available()` False, picker row gone; after: 0 POSTs, entry untouched,
`has_available()` True, picker row present with the catalog. Runtime `select()`
still refreshes (control).

Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
2026-09-18 10:12:25 -07:00
teknium1
802a9975d2 fix(cron): no_agent scripts get the owning profile's declared secret, never the launch profile's
A no_agent cron script owned by a profile served by a multi-profile gateway (or the
Desktop/dashboard backend) could not receive its own credential: the served profile's
.env never enters the process env (load_hermes_dotenv skips the process-global load
for a routed home), and the child-env sanitizer only resolved a terminal.env_passthrough
name through the profile secret scope when the launch environment already carried that
name. So the documented declaration mechanism (SECURITY.md 2.3) could not supply a
value that exists only in the served profile's scope, while the launch profile's own
.env credentials rode into the served profile's script child unstripped.

- tools/env_passthrough.scoped_passthrough_additions: declared names the bound scope
  holds but the env being filtered lacks; reads the scope alone (no os.environ, no
  other profile), empty without a scope so single-profile spawns are byte-identical.
- _scrubbed_env (terminal, cron scripts, bg processes, search workers) and
  execute_code's _scrub_child_env overlay those names after filtering.
- cron _run_job_script and terminal _make_run_env strip the launch profile's .env
  residue from the base (strip_launch_profile_env, a no-op for the launch profile's
  own jobs) before the filter, so the scope overlay lands after the strip.

Why not the 1Password-only shape of #114218: the gap is the declaration mechanism, not
one vendor's token, and an unconditional token pop broke the single-profile .env flow.
Docs: cron no_agent credential section, secrets child-process section, security
passthrough note.

Fixes #114209
Supersedes #114218
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-18 10:04:57 -07:00
teknium1
021a4bfa9b fix(acp): pin the session cwd while registering IDE-provided MCP servers
The default stdio child cwd now reads agent.runtime_cwd.resolve_context_cwd()
at spawn time (previous commit), but ACP new_session/load_session register the
client's MCP servers outside the per-turn cwd pin, so a hosted ACP session with
logical cwd /workspace/a — the reporter's exact scenario — still spawned its
stdio servers in the hermes-acp process directory. Run register_mcp_servers in
a copied context with set_session_cwd(state.cwd); run_coroutine_threadsafe
carries that context onto the MCP loop task and the long-lived server task
copies it, so reconnects respawn in the same directory.

Also: trim the salvaged tests to the two invariants (session pin becomes the
default; explicit config cwd wins) — the TERMINAL_CWD fallback and the
missing-directory→None cases are runtime_cwd's own contract, pinned in
tests/agent/test_runtime_cwd.py, and the native default (no anchor → None) is
already pinned by test_start_preserves_native_default_cwd. Document the `cwd`
key and its default in the MCP docs.

Shared-process caveat recorded in the PR body: MCP connections are per
process/registry scope, not per chat session, so the anchor is read once at
connect time (profile-level terminal.cwd for gateway/cron; the owning session
for ACP-provided servers).
2026-09-18 10:03:40 -07:00
teknium1
9b55e7e6ac fix(plugins): attach stored git credentials only after an anonymous refusal, on every network verb
Widen the anonymous-first decision from #114545 to the whole class. The
clone-only wrapper (_clone_with_auth_fallback) folds into one runner,
git_credentials.run_git_with_credential_fallback, used by every site that
attached a stored credential unconditionally:

  - plugins_cmd._run_plugin_git -> clone, the pinned --ref fetch in
    _checkout_exact_revision, and the pull in _git_pull_plugin_dir
    (`hermes plugins update`), which the thread on #114526 showed still
    failing with "could not read Username ... terminal prompts disabled"
    after the clone succeeded anonymously
  - mcp_catalog._do_git_install (catalog MCP git installs)
  - profile_distribution._git_clone (profile distributions from a git URL)

Why the runner and not per-caller guards: the decision belongs to the
credential layer once; a rejected credential sent pre-emptively is what
turns a working anonymous clone into a 401 that git can only answer with
the prompt noninteractive_git_env disables. The credential-required
classifier also matches git's "returned error: 401" spelling (the picked
version needed a trailing space that git never prints) and bytes output.

Side effect: `gh auth token` is now shelled out only after a refusal, so
public installs no longer pay that subprocess at all.

Tests: the sibling verbs (fetch, pull) x (public, refused) and the negative
that a non-credential failure never resolves a credential; the SSH no-op
case from the pick is already asserted by the private-remote test. Docs:
website/docs/user-guide/features/plugins.md.

Fixes #114526
Co-authored-by: holny <holny@users.noreply.github.com>
2026-09-18 10:01:01 -07:00
teknium1
3f3fe7d763 docs: describe stop_details in the Anthropic refusal message and log line 2026-09-18 09:52:36 -07:00
teknium1
6545fb8246 fix(kanban): show the resolved dispatch_profiles allowlist in hermes kanban diagnostics
Closes the third ask of #113620: an operator on a shared board can see what
this home believes it may claim. Text output gains a leading
`kanban.dispatch_profiles: <any | names | none (fail-closed: …)>` line;
--json now returns {"dispatch_profiles": ..., "tasks": [...]} with the
per-task rows unchanged. Resolution reuses _dispatch_profile_allowlist so the
CLI and the dispatcher can never disagree.
2026-09-18 09:50:19 -07:00
teknium1
c432cd2723 fix(kanban): warn when the dispatch allowlist is unreadable; document the fail-closed spellings
Follow-up to the salvaged #113623 hunk: a config read that raises now logs a
WARNING naming the error before claiming nothing, so a corrupt config on a
shared board is visible instead of silently changing this home's claim
scope. Docs: the example no longer shows `dispatch_profiles: null` as the
"any profile" default (null is now fail-closed like `[]`); omit the key for
"any". Trimmed the contributor's positive-control test (the two invariant
tests cover the fixed class; the listed-profile control is the live probe).

Fixes #113620
Salvages #113623
2026-09-18 09:50:19 -07:00
liuhao1024
1ea94b3745 fix(kanban): decomposed cards fall back to the root task's assignee, never the dispatcher's own profile
The decomposer resolves both `kanban.default_assignee` (unrouted children)
and `kanban.orchestrator_profile` (the root after fan-out) as
"explicit config, else the active profile". The active profile is whatever
HERMES_HOME hosts the dispatcher — in the report an incognito `private`
profile with no credentials — so every unrouted child AND the root card
itself were silently re-owned by a profile that can never do the work.

The resolver now tries the root task's own assignee before the active
profile: explicit config → root assignee (if it names an existing profile)
→ active profile. Slim redo of #114303's decompose hunk at the resolver so
it covers the orchestrator fallback too; tests and docs from #114303.

Fixes #114294
Salvages #114303

Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
2026-09-18 09:48:41 -07:00
Phoebie Builder
5579fd5cdf fix(kanban): wake a served profile's api_server session in-process under its own scope
A Kanban notify+wake subscription whose destination is a raw session id on the
shared api_server (platform=api_server, chat_id=<session id>, notifier_profile=
served secondary) was never delivered under gateway.multiplex_profiles: no
profile_routes entry can anchor a session id, so _adapter_for_subscription()
returned None and _claim_for_sub skipped the row before claiming (cursor frozen,
no delivery attempt). A platform-wide api_server route is not the fix — it
matches every api_server destination, so the default profile's own api_server
subscriptions would fail closed instead.

The wake leg had the same shape: _self_post_chat_completion() always POSTs to
the unprefixed /v1/chat/completions with the PRIMARY adapter's key, so a served
profile's wake turn would resume the session in the default profile's store.
/p/<profile>/v1/... requires that profile's own API_SERVER_KEY, which a
route-only profile legitimately does not have.

Authorize the shared adapter only when the subscription's session id is a row in
the served profile's own state.db stamped for that profile (NULL legacy stamp =
the store's own profile, the same rule the dashboard session routes apply) and
run the wake in-process under that profile's runtime scope through a new
APIServerAdapter.run_internal_session_turn(). Unknown session, foreign stamp,
unserved/deleted profile and unreadable store all fail closed; the concurrent-run
cap defers the wake (cursor rewind, retry) instead of bypassing it.
2026-09-18 09:46:35 -07:00
teknium1
b0cd35e259 fix(mcp): dashboard/Desktop Authorize honours a pre-registered client's pinned loopback redirect
A no-DCR entry (client_id + redirect_port, e.g. the Asana manifest) has
http://localhost:<port>/callback registered with the vendor, which matches
redirect URLs exactly; the dashboard flow overrode it with its own callback
URL, so the in-app Authorize button could never complete for such entries.
The pinned loopback listener now wins (over the dashboard URL and any cached
registration URI); the dashboard flow only publishes the authorization URL,
and the stdin paste reader stays off under a dashboard flow. Docs/post_install
say the approving browser must run on the Hermes machine.
2026-09-18 09:45:32 -07:00
teknium1
c553df915c fix(mcp): Asana catalog installs a working V2 pre-registered OAuth client
The V1 beta server https://mcp.asana.com/sse is retired and Asana's V2 server
(https://mcp.asana.com/v2/mcp, Streamable HTTP) has no Dynamic Client
Registration: every client must be an Asana "MCP app" the user registers in
the developer console. The salvaged manifest fixed the URL and documented the
manual `oauth:` block; this commit makes the catalog install itself produce
that block so no hand-edit of config.yaml is needed.

- hermes_cli/mcp_catalog.py: manifests may pin a closed `auth.oauth` mapping
  (client_id, client_secret, redirect_host, redirect_port, scope) that
  `_build_server_config` copies to `mcp_servers.<name>.oauth`; every `${VAR}`
  it references must be declared in `auth.env`, mirroring the api_key header
  contract, so a placeholder can never reach the token endpoint as a literal.
  `install_entry` now prompts `auth.env` for OAuth entries too (dashboard and
  Desktop already render `required_env` regardless of auth type).
- optional-mcps/asana/manifest.yaml: declare ASANA_CLIENT_ID/SECRET, pin the
  client + `http://localhost:27890/callback` (Asana matches the redirect URL
  exactly), and rewrite post_install around the MCP-app registration steps.
- tests: shipped-catalog invariant (no manifest installs the retired `/sse`
  URL; the Asana client credentials are declared `${VAR}` refs; callback
  pinned) + install-path invariant for the new `auth.oauth` block, including
  the undeclared-reference rejection.
- docs: catalog section on entries that need a user-owned OAuth app.

Live: `hermes mcp install asana` + `hermes mcp login asana` under a temp
HERMES_HOME. Before: config `url: …/sse`, no oauth block, authorize URL on the
V1 server (mcp.asana.com/authorize) with a DCR client and 127.0.0.1 redirect.
After: config `url: …/v2/mcp` + oauth `${ASANA_CLIENT_ID}` refs, .env holds
the values, authorize URL on app.asana.com/-/oauth_authorize with the
configured client_id, redirect_uri=http://localhost:27890/callback and
resource=https://mcp.asana.com/v2/mcp — the flow Asana's guide documents.
2026-09-18 09:45:32 -07:00
teknium1
a89ef71b10 refactor(plugins): kill-list annotation takes the pre-resolved list directly; trim to two invariants
Follow-up to the salvaged #113687 / #113682 commits:

- `removed_annotation(name, dir_path, removed_entries=None)` resolves the kill list once itself
  when no list is passed and matches against it; the `removed_annotation_batch()` wrapper and its
  name-keyed dict are gone (a hub row keyed by `name` could shadow a same-named nested plugin).
  Both surfaces (`plugins list`, `_merged_plugins_hub`) call `resolved_removed_entries()` once
  and pass the list per row.
- The negative cache is one module float deadline instead of a URL-keyed dict plus two helpers;
  the duplicated stale-cache read is one `_stale_live_cache()`.
- Tests trimmed to two invariants: failure memory honours the TTL and still serves a stale copy;
  hub rebuild + `plugins list` each cost one network attempt and still annotate an in-tree removal.
  The #113682 route test's fake gains the new third argument.
- Docs: the live-refresh section says a failed fetch is remembered for a minute.
2026-09-18 09:43:19 -07:00
teknium1
454f9c1421 fix(cron): remind once per cooldown after the alert-once gate; migrate alerted_at
Builds on the previous commit (alerted treated like closed): a permanently
silent `alerted` incident would hide a job that stays broken for days, so the
gate now withholds only inside `cron.failure_repeat_alert_hours` (default 6,
0 = re-alert on every failing run) and lets exactly one reminder through, which
re-stamps the window.

- cron/incidents.py: `alerted_at` column (added in place to existing ledgers via
  add_column_if_missing); set_incident_state(..., "alerted") stamps it every time
  so a reminder restarts the cooldown; resolved->detected re-open clears it so
  the same error after a green run alerts immediately.
- cron/scheduler.py: `_repeat_alert_withheld` reads the stamp; a missing or
  unparseable stamp (pre-migration row) delivers rather than swallowing the
  alert; `closed` still wins; the unreadable-ledger fail-open is unchanged. The
  crash path (`_deliver_crash_failure`) shares the gate through
  `_upsert_incident_for_failure`.
- hermes_cli/config_defaults.py + website/docs cron page: document the key.
- tests: fold the contributor's two tests and the old "unacked failures keep
  alerting per run" change-detector into two invariants (unit gate incl. 0 /
  legacy row / closed; end-to-end alert once -> reminder once -> recovery re-arms).

Fixes #113665
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 09:40:00 -07:00
teknium1
8e1b4216b5 fix(cli): startup auth fallback re-resolves reasoning effort through the CLI chokepoint
_resolve_fallback_runtime swapped requested_provider/model after the primary's
AuthError but left reasoning_config at the launch model's effort, so the first
request to an always-thinking fallback model (zai/glm-5.3-flash with a
reasoning_overrides entry) went out with the primary's `medium` and 400'd.
Kanban workers hit it on every spawn since each is a fresh `hermes chat`.

Route the swap through `_resolve_cli_reasoning`, the CLI-level chokepoint that
/model, /new and --resume already use (77f5de23dd), instead of an inline
try/except-pass copy: it reads the same in-memory CLI_CONFIG startup resolved
the primary against, so nothing can fail and the swap is never at risk. Keep
the `_explicit_reasoning_config` guard from #113497: an explicit
`--reasoning` is the user's intent for this run and outranks the fallback
model's config (kanban passes it per task the same way).

Test: one real-HermesCLI invariant, parametrized over the two cases
(fallback override applies; explicit flag survives), replacing the
SimpleNamespace stub test. Docs: fallback-providers.md.

Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
2026-09-18 09:35:49 -07:00
lEWFkRAD
cb8d652b49 fix(cron): resume keeps a recurring slot that elapsed while paused due (#113603)
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).

resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.

Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
2026-09-18 09:32:34 -07:00
teknium1
093559f5a9 docs(kanban): goal judge gate fails open on transport failure; headless affinity key 2026-09-18 09:32:02 -07:00
teknium1
0db097dc22 fix(cron): scrub the bot-chat deferred record too; trim redaction tests to two invariants; docs
Follow-up to the salvaged #73026 (@ryangu00):

- `_deliver_to_bot_chat` now rebinds `content` to the redacted copy before anything reads it, so
  the durable deferred record (`cron/bot_chat_delivery.defer`) and its later replay carry the
  scrubbed payload as well as the live-owner and CLI lanes. Previously only the assembled
  `message` was scrubbed while the raw output was persisted for replay. Redaction is idempotent,
  so the replay's second pass is a no-op.
- The seven contributor tests collapse into two parametrized invariants: every outward lane
  (platform send / session mirror incl. job-name splice / bot-chat turn) masks the secret while
  the payload and framing survive; and egress redaction is forced (independent of
  `security.redact_secrets`) and fails closed (a raising redactor replaces the payload).
- Cron docs: note what is redacted on delivery, that it ignores `security.redact_secrets`, and
  what is deliberately not stripped (credential-named URL params, shapeless user secrets).

The failure-notice lane #55271 targeted (`_summarize_cron_failure_for_delivery`) exits through
the same `_deliver_result` chokepoint, so it is covered without a second redaction site.

Co-authored-by: necoweb3 <sswdarius@gmail.com>
2026-09-18 09:30:54 -07:00
teknium1
06350c79d5 fix(kanban): dashboard done/review refusals name the open parents
complete_task and request_review return bare False when the dependency
gate refuses, and _patch_status only enriched the 409 for status=ready,
so PATCH status=done/review on a gated card answered "not valid from
current state" and the bulk entry said "transition refused". Consult
unsatisfied_parents on a refused done/review and name the parents in
the 409 detail and the bulk entry error, matching the CLI/tool wording.
2026-09-18 09:29:44 -07:00
teknium1
0cdc325b46 fix(kanban): dependency-gate refusals name their cause on every surface
Salvage follow-up on the cherry-picked #113398 (@fangliquanflq) and #113452
(@JoaoMarcos44) commits, trimmed to the salvage bar and widened to the class.

- link: drop the HERMES_KANBAN_BOARD binding on the run-id ownership proof
  (tools._worker_run_id / kanban._worker_run_id_for). complete/block prove
  ownership with the same task-scoped run id and never board-bind it; the
  guarded case is a forged cross-board task-id collision, not a live path.
  Dropped the two cross-board tests and the duplicate db owner test.
- complete: move the open-parents enumeration into kanban_db.unsatisfied_parents
  next to _parents_satisfied so the tool, the CLI (`hermes kanban complete`
  printed "unknown id or terminal state" for the same refusal), kanban_show
  (`unsatisfied_parents`) and diagnostics share one query.
- show: new `running_with_open_parents` diagnostic (hermes kanban show /
  diagnostics / dashboard) for a running card whose parent is not done —
  the "surface unresolved parents on a running card" ask of #113374.
- tests trimmed to invariants; docs for the refusal text, the show field and
  the diagnostic.

Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-18 09:29:44 -07:00
fangliquanflq
117b77965d fix(kanban): preserve owned dependency handoffs 2026-09-18 09:29:44 -07:00