Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:
- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
`choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
completed `reasoning` output item ahead of the message (and ahead of that step's
`function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
from the assistant messages the agent already persisted (`build_assistant_message`
stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
in non-stream mode the agent fires the callback per provider delta AND once more with
the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
`conversation_history`. Responses SDK clients replay a prior response's `output`
list as the next `input`; before, the item became an empty `user` message (in
`input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
(`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
`output_item.done` item and the non-streaming shape.
Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).
- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
`_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
keeping them distinct from answer text. The lossy 500-char `reasoning.available`
progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
(the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
(`output_item.added`, `reasoning_summary_part.added`,
`reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
`output_item.done`), closed before the next message/function_call item opens
and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.
Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.
Reactive, destination-scoped recovery instead:
- error_classifier: new `FailoverReason.role_alternation` for the vendor
alternation wordings (checked before the request-validation table since
the body also carries `invalid_request_error`); same abort+fallback hints
as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
loop's `_merge_user_content`), `destination_key` (base_url|provider,
model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
adjacent user turns merged, remember the destination on the facade for
the session so later iterations pre-merge, never touch destinations that
accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.
Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.
Fixes#112358
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:
- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
already did this in `_handle_reset_command`, the second request is
idempotent) and the idle `_handle_stop_command` tail, which replied
"No active task to stop." while a background child was running; it now
stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.
The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.
An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.
Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.
Part of #114456
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).
Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.
`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.
No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.
Part of #114587
Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
Widen the salvaged Image API hunk (#114340) to the whole class:
- plugins/image_gen/_common.py: record_token_usage() — one helper feeding the
aux accounting chokepoint (agent.aux_accounting.record_aux_usage) with task
"image_generation", the billing provider and the priced model id. Dict and
SDK usage objects alike; a body without tokens is a no-op, as is a call
outside a turn.
- plugins/image_gen/openrouter: the /chat/completions path (the DEFAULT model
chain — openai/gpt-5.4-image-2, google/gemini-3-pro-image — is chat-only and
token-billed) now records too; the Image API path uses the shared helper
(the contributor's local _record_image_api_usage is folded into it) and both
pass base_url so pricing resolves the route. Task renamed image_gen ->
image_generation to match the other aux task names.
- plugins/image_gen/openai: gpt-image bills per text/image token; record the
Images API usage block under the API model (gpt-image-2), not the Hermes
quality-tier label.
- FAL, xAI, Krea, DeepInfra, Meta, openai-codex return no token usage and stay
unrecorded (nothing to bill per token).
- tests: the contributor's two tests folded into one invariant parametrized
over chat / Image API / no-usage control; one OpenAI invariant.
- docs: image-generation.md "How It Works Internally" gains the accounting step.
Fixes#114324
The salvaged commit (#114651) re-probed the endpoint through get_model_context_length
at spawn time. The review fork is a full AIAgent whose context_compressor.context_length
is already resolved by the same chain (config override, catalog, persisted provider limit,
endpoint probe), so read that instead: no second lookup, and the budget tracks exactly
the window the fork's own compaction uses. Tests trimmed to two invariants on the new seam.
Docs: `auxiliary.background_review.max_input_tokens` and its derived default in
website/docs/user-guide/features/memory.md, including the note that a top-level
`background_review:` block is not read (#114645 secondary).
Co-authored-by: Artie Fisher <artiefisher123@gmail.com>
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
mcp >= 2.0's Streamable HTTP client folds any non-2xx whose body it cannot parse
as a JSON-RPC error into the opaque `-32603 Server returned an error response`.
Hermes printed that text verbatim in the SSE-fallback warning and in the
"both transports failed" ConnectionError, so users saw no status, no URL and
none of the server's own words (e.g. `400 {"code":-32020,"message":"Unsupported
MCP-Protocol-Version"}`) and had to reach for curl to learn what the server
actually said (#114350, #113359).
- `_make_http_rejection_recorder`: response hook on the owned SDK-httpx client
that remembers the last 4xx/5xx (status, method, URL, head of the body; SSE
bodies are never read). Sibling of the redirect-header stripper hook.
- `_describe_http_failure`: appends that detail only when the root cause is
the SDK's opaque -32603 text, so a real JSON-RPC error or an httpx status
error is never duplicated.
- `_run_http`: the fallback warning and both raise paths (both-transports
ConnectionError; the no-fallback re-raise after a proven session, strict
redirect headers or a non-rejection) carry the detail. Debug-logs the
endpoint each connect attempt uses.
- Docs: troubleshooting entry for reading the new message.
Verified live against a real Streamable HTTP server (`hermes mcp test`, temp
HERMES_HOME): base prints `Streamable HTTP: Server returned an error response`;
fixed head prints `... (HTTP 400 from POST http://127.0.0.1:PORT/mcp:
{"jsonrpc":"2.0","id":null,"error":{"code":-32020,...}})`. Control: a 400
served as application/json already surfaces the JSON-RPC message and gets no
appendix; servers negotiating `initialize` down to 2025-06-18 (fixture and a
real FastMCP on mcp 1.12.4) connect and list tools on base and fix alike — the
pinned mcp 2.0.0 stamps the negotiated version on every post-handshake request
(wire-recorded), so the sticky seed in `_run_http` is not the cause of the
reported 400 and stays as designed (#14816).
Co-authored-by: JoaoMarcos44 <joaomarcosdias444@gmail.com>
``_account_crashes`` trips the protocol-violation budget with
``force_trip`` after three consecutive clean exits, but the task's
``consecutive_failures`` is still 1 — so ``recompute_ready``, which
re-derives the threshold from its caller's ``failure_limit`` (default 2),
promoted the just-blocked card straight back to ``ready`` in the same
tick: ``protocol_violation -> gave_up -> promoted -> claimed`` forever
whenever ``failure_limit`` exceeds the violation count (the systemic
same-error trip at limit 1 had the same hole).
``_has_sticky_block`` now also honours the breaker's own verdict: the
newest ``gave_up`` since the last ``unblocked`` holds the card when it
recorded a violation-streak trip or ``failures >= effective_limit``.
``hermes kanban unblock`` remains the release path and still grants a
fresh budget. Tests cover the fresh-process classification (rc=0 and
rc=75), the held trip + operator unblock, and the trailer emission gate.
The dead-worker sweep learned a worker's exit status only from
``_recent_worker_exits``, which ``reap_worker_zombies`` fills via
``os.waitpid`` — so only the process that spawned the worker ever knows
how it exited. With ``kanban.dispatch_in_gateway: false`` every
``hermes kanban dispatch`` tick is a fresh process, the registry is
empty, and a worker that exited rc=0 without a terminal board call was
booked as a bare ``crashed`` / ``pid N not alive``: no
``protocol_violation`` marker, no corrective error text for the retry
worker, no violation streak — and a rc=75 quota wall was counted as a
failure instead of a neutral ``rate_limited`` requeue.
A Kanban worker (``HERMES_KANBAN_TASK`` set) now writes
``[kanban-worker-exit] rc=<code>`` as the last line of its own log on
every one-shot exit path; when the registry has no entry for a dead PID
the sweep reads that trailer and books the exit through the same
code -> kind mapping. A worker killed before its exit epilogue leaves no
trailer and stays a plain crash. ``_worker_final_output`` strips the
trailer so it never leaks into the board diagnostic.
Direction (a durable, process-independent witness in the worker log)
from PR #113638; its predicate keyed on the ``Resume this session with:``
summary, which the CLI prints before ``sys.exit`` for rc 0, 1, 75 and 130
alike and so would have booked quota walls and failed turns as protocol
violations.
Co-authored-by: kokhlo <konstantin.khlopkov93@gmail.com>
#114309 expects the dashboard, like `hermes cron status`, to say "scheduler last
ticked X hours ago" when a stopped ticker leaves next runs stranded in the past.
The Cron page only had the per-row `Overdue since` label and no way to date the
outage: GET /api/cron/jobs carried nothing about the ticker.
Every job the dashboard cron endpoints return now carries
`scheduler_heartbeat_age_s` — its own profile's ticker heartbeat age read inside
the same store scope as the job list (None when it cannot be dated) — and the Cron
page renders "Scheduler last ticked 7h ago — jobs that came due since then have not
fired" above the list when the oldest heartbeat among jobs expected to fire is
past the CLI's STALE_AFTER threshold (three missed 60s iterations plus slack).
Paused, disabled and completed jobs never raise the banner, matching the overdue
label's rule. The field is additive, so the response shape and every existing
consumer are unchanged.
Tests: one router test lists two profiles with one stale heartbeat file and checks
each job reports its own age (None for the profile without a heartbeat); one vitest
pins the stale-age selection and the relative label.
The hermes-bots plugin's Routines card (and its inspector's `Next run` row)
still read `Next: 7 hr. ago` for a `next_run_at` parked past the scheduler
grace — the one Desktop surface the overdue labelling left out (#114309).
The card now switches to `t.cron.overdueSince` and the inspector row to
`Overdue since`, using the same `nextRunOverdueMs` decision as the core cron
panel and sidebar, so paused/completed/disabled jobs keep the plain label.
`nextRunOverdueMs` is exported through `@hermes/plugin-sdk` (the plugin only
imports from the SDK) and its parameter is widened to the optional-field shape
the plugin's `RoutineJob` carries; no plugin i18n key is needed because the
card already renders from core `t.cron`, which every locale resolves.
Test: one card test renders an overdue vs an upcoming job and checks both the
card label and the inspector rows; the fixture's fixed 2026-08-23 next_run_at
is made clock-relative so it never ages into "overdue" on its own.
`hermes cron list`, the web dashboard Cron page and the Desktop cron panel/sidebar all
rendered a `next_run_at` parked hours in the past as an ordinary upcoming "Next run" —
the only user-visible trace of a scheduler that stopped ticking (#114309). Every
surface now labels a slot past `cron doctor`'s 15-minute grace as overdue (CLI
`Overdue:` row with the lateness, web `Overdue since`, Desktop `Overdue since` label
on the detail panel and sidebar meta) while paused/disabled/completed jobs keep the
plain label because they are not expected to fire.
The CLI's overdue check now parses through `cron.jobs._parse_aware` and `hermes_time.now`
so status/list/doctor and the ticker agree on the instant (mixed offsets, DST folds,
legacy naive stamps read as system-local like the scheduler does), and both the
ordering and the subtraction normalise to UTC: Python compares same-tzinfo datetimes by
wall clock, which is wrong across a DST fold. Status still orders the soonest run by
instant (#113874) and prints the stored stamp.
Tests: the salvaged status tests are trimmed to two invariants (overdue + stale
heartbeat is loud on status AND list; within-grace stays plain on both), the DST-fold
ordering tests freeze the CLI clock as well so their 2026-11 fixtures never start
reading as overdue, and one vitest each pins the web and Desktop helpers.
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
A worker process that carries HERMES_DELEGATED_CHILD_CONTEXT next to its
HERMES_KANBAN_TASK is fenced by kanban_path_is_fenced's env-marker branch: both
auto-heartbeat writes raised PermissionError at DEBUG only, so the board showed a
worker that never beats while the process was alive (the salvaged commit already
made the return value honest, but _touch_activity ignores it). Log the refusal once
per process at WARNING with the cause, and skip the bridge for an in-process
delegate child BEFORE stamping the rate-limit window so a chatty child cannot starve
the worker's own heartbeat.
No self-write grant is added: the marker + TASK combination means "descendant, not
worker" by design (b578261584, 8a8c3634e8), and the only in-tree grant path,
kanban_db_dispatch._default_spawn, already pops the marker. The real-spawn grant
test now proves that from a dispatcher that itself carries the marker: the worker's
auto-heartbeat lands and its handoff completes.
Docs: worker guide notes the automatic claim extension and what the refusal warning
means.
`_cron_job_action` printed `(^_^)b Triggered job … It will run on the next
scheduler tick.` unconditionally for `run`, so a run-now the tool refused
(`execution_skipped`: paused job, claim held by another run) was
indistinguishable from an accepted one. Print the refusal, and otherwise
reuse `hermes_cli.cron._run_outcome` — the same verdict line `hermes cron run`
already prints (ran now / background / skipped).
Tests: one invariant per atom — the real one-shot CLI shape
(`_SESSION_ASYNC_DELIVERY` scoped False) reaps a dead-owner 'running' ledger
row before the sync fallback; `/cron run` on a refused claim prints the reason
and never "Triggered". Both red on origin/main. Replaces the #113938 test hunk
(it relocated the prior test's assertion). Docs: execution-history section
notes the pre-manual-run reap.
Part of #113923.
/kanban is a workspace page route: the titlebar slot is hidden there and the
switcher is projected into the Kanban page's header row (WORKSPACE_PAGE_HEADER_AREA
after #114960). Point readers, the i18n comment and the test comment at that row.
Follow-up to the salvaged #114647 hunk (visible "Board" label +
`Board: <name>` accessible name):
- replace the native `title` with the app's `Tip` ("Switch board"), the
same tooltip every other dropdown trigger uses, so hover names the
ACTION instead of repeating the board name
- lead the trigger with the plugin's `project` icon so the title-bar
text reads as the Kanban board control, not a static window title
- add `switchBoard` to the kanban plugin bundle (en/ja/zh/zh-hant)
- test registers the real locale bundle and asserts the rendered label,
accessible name and tooltip (the raw-key fallback matched
`/board:/i` by coincidence)
- docs: describe the Desktop switcher and its local selection
Fixes#114642
Reshape of the salvaged fix from #114513 (@whyyagswhy):
- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
or round-robin order. The contributor's version called the private
``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.
The File attachments section still promised that removing an attachment always deletes the on-disk file. With reference-counted blobs the file is only unlinked once no other attachment row points at it, so the user guide now states that condition and names the operator CLI path (hermes kanban attach-rm ATTACHMENT_ID).
A cron job delivering to Bot Chat booked "timed out after 600s" for turns
that finished in seconds: cron/scheduler_delivery.py::_deliver_to_bot_chat
waited for the `hermes chat -Q` child to EXIT, but when the bot's turn had
messaged a teammate (message_agent -> notify_on_complete runner) the child
then runs the one-shot exit linger, bounded by
terminal.oneshot_completion_wait_seconds (default 600) — the same default as
cron.bot_chat_delivery_timeout_seconds. The cap counts from the claim, the
linger only starts when the turn ends, so the cap expired first on every
such delivery, booked a completed turn as a timeout, held the job's fire
fence for the full cap, and killed the child mid-linger — tearing down the
reply the linger exists to protect (#90879).
Ordering the two bounds cannot fix it (the turn has no duration bound; the
linger has its own contract), so the cap stops competing with the linger:
- hermes_cli/quiet_single_query.py: the -Q child accepts a per-process
report path (HERMES_QUIET_TURN_REPORT_FILE), popped before the turn like
HERMES_TURN_AUTHOR so nothing the turn spawns inherits it; cli.py writes
{pid, exit_code, error} there the moment the turn ends, BEFORE the linger.
- cron: _run_bot_chat_turn polls the child and that report under the cap.
Report present -> the delivery is booked from it (real exit code and
stream tails when the child exits within a short grace) and the still-
lingering child is left running, drained and reaped by a daemon thread.
No report by the cap -> the turn never ended: killed and booked as a
timeout, exactly as before. The linger itself is untouched.
Live repro (real _deliver_to_bot_chat, real `hermes chat -Q` against a
loopback provider whose turn spawns `sleep 90` with notify_on_complete,
cap 45s): base books the timeout at 45.7s for a turn that ended at +5s and
kills the child; fixed head books success at 11.2s, the child lingers, the
teammate follow-up turn runs at +97s and the child exits on its own.
Supersedes #113649's cap = delivery + linger (the thread shows a headroom
only moves the race and lengthens the fence hold); analysis credit to the
reporter and the thread's independent verification.
Fixes#113608
Co-authored-by: KoNit-K <konit.block@protonmail.com>
`_runtime_credentials_ready()` collapsed "nothing is configured" and "the
configured credential is unusable right now" into one False, so a profile
whose only Nous credential was benched by a failed refresh (or quarantined
DEAD) was told "No inference provider is configured yet" and offered the
provider picker — re-running setup on an OAuth provider with single-use
refresh tokens can rotate the grant away from the session that was working
(#113720).
The readiness probe now returns the resolver's exception too
(`_probe_runtime_credentials`). At CLI startup only the resolver's
`no_provider_configured` verdict reaches the wizard; any other failure is
explained by `_explain_unusable_credentials`: `format_auth_error(exc)`
plus, from the provider's pool, "cooling down ... re-enters rotation in
about Nm" or "sign-in was lost (<reason>); run `hermes auth add <provider>`".
An actually-empty profile still gets the wizard.
Fixes#113720
Supersedes #113732 (@whyyagswhy): its `_inventory_other_providers()` gate
returns False for a Nous-configured profile, so the reported Nous case would
still have reached the wizard, and it printed no reason.
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
Two defects at the same spawn boundary (hermes_cli/kanban_db_dispatch.py::
_restart_safe_worker_argv -> tools/process_registry.py::restart_safe_gateway_child_argv).
#114720 — a managed-gateway dispatcher whose user D-Bus session had gone away
raised the restart-safe-scope RuntimeError at spawn; the ordinary spawn
except-clause fed it to _record_task_failure, which advanced
consecutive_failures and at failure_limit parked the card as a bare `blocked`
(block_kind NULL) with a `gave_up` run. Nothing about the card had run.
- restart_safe_gateway_child_argv raises RestartSafeScopeUnavailable
(RuntimeError subclass) so the dispatcher can tell a host refusal from a
card failure.
- _record_task_failure(infrastructure=True): run + event carry
`infrastructure: true`, consecutive_failures is left alone, the breaker
never trips; the card stays `ready` with the linger remedy in
last_failure_error and a warning in the dispatcher log.
- check_respawn_guard returns `infrastructure_cooldown` while the latest run
is such a refusal inside the rate-limit cooldown window, so a dead bus is
retried spaced instead of every tick.
- require_restart_safe_scope=True keeps its hard-fail semantics under the
managed gateway.
#113612 — `hermes kanban dispatch` from an operator's Type=oneshot systemd unit
spawned unmanaged workers: the gate `_is_supervised_gateway_process()` is
False for a CLI dispatcher, so the helper returned the bare argv and the
workers died at the unit's cgroup teardown with empty logs, while the
`running` rows waited for reclaim.
- New outlives_parent=True (kanban only; cron blocks its caller and keeps the
old gate): under any systemd unit (INVOCATION_ID) a fire-and-forget worker
is scope-wrapped when the user bus is reachable. Without a bus it degrades
with a loud once-per-process warning naming the consequence and both
remedies (enable-linger, KillMode=process) rather than refusing — the
unit's KillMode/lifetime is unknowable and a long-lived sequencer without
linger must keep working.
Docs: kanban.md "Workers and systemd cgroups" + guard reasons; cron.md note.
Based on analysis from PR #113624.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
With a bound project the settings dialog omits a blank directory from the PATCH and rename_board mirrors the project's primary folder, so clearing the field cannot fall back to scratch until the project is unbound; the Settings bullet now says so. The bundle test drops five wording-level substring asserts that pinned the payload shape already covered behaviourally by test_kanban_board_project_api.py, keeping only the switcher's unbind seam.
The create/settings dialogs (salvaged from #114664) let a user set or clear
a board's project_id, but the binding was invisible outside Settings.
GET /boards already annotates every board with project_id + project_name,
so the switcher now renders a "Project: <name>" badge whose × sends
PATCH {project_id: ""} — the same clear the settings dialog uses — without
touching default_workdir.
Settings also stops sending default_workdir: "" next to a chosen project
when the directory field is blank: the explicit "" suppressed the server's
project → default_workdir mirror, so binding via Settings left the board
without the workspace default the create dialog would have seeded.
Trim the salvaged test to the invariants (selector wiring, payload shapes,
badge unbind) and document the control on the Kanban docs page.
A Nous pool entry whose refresh reached the resolver's state-shape raises
("Hermes is not logged into Nous Portal.", "No access token found ...",
"No refresh token is available ...") carried relogin_required=True but no
OAuth code, so `OAuthProviderFlow.is_terminal_refresh_error` classified it
transient: `_recover_failed_refresh` benched the only credential for an
hour with null last_error_* fields and nothing at WARNING (#113718).
Give those raises codes (`nous_auth_missing`, `nous_auth_missing_access_token`,
`nous_auth_missing_refresh_token`) and list them in the Nous flow's
terminal_refresh_codes, mirroring the Codex/xAI `*_auth_missing_refresh_token`
convention. The pool then quarantines, marks the row DEAD with the reason,
and logs the existing WARNING naming `hermes auth add nous`; every consumer
of the classifier (pool, singleton `_refresh_nous_or_quarantine`) agrees.
A failure that says nothing about the login (network error, non-JSON 5xx
body) still benches as before.
Fixes#113718
Salvages #113726 (@whyyagswhy) — tests kept, consumer-side hunk replaced
by the classifier fix.
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
`get_codex_auth_status` / `get_xai_oauth_auth_status` back every credential-gated
listing (`/model` picker rows, `hermes doctor`, dashboard auth cards). The shared
`_pool_first_oauth_status` read the pool with `select()`, which is a runtime lease:
it refreshes an expiring single-use token and, when that speculative POST fails
transiently (503, endpoint unreachable), benches the entry with a persisted
exhaustion cooldown. The picker then rendered the provider as an unconfigured
skeleton ("needs setup" / "0 models") while the runtime resolver kept serving
the same credential. On a round-robin pool the read also rotated and persisted
the priority order.
Read the pool with `peek()`: no refresh, no accounting, no rotation, no persist.
Refreshing stays with the runtime resolver reached through the `resolve`
fallback, whose failures persist nothing (same design as the Nous status).
Slimmer redo of #114379 by @Finn763 (observe flag threaded through four pool
methods plus non-refreshing resolvers); fixtures adapted from that PR.
Live: pool-only openai-codex entry, expired access token, stubbed token endpoint
returning 503 — before: 1 refresh POST, entry persisted `exhausted`,
`has_available()` False, picker row gone; after: 0 POSTs, entry untouched,
`has_available()` True, picker row present with the catalog. Runtime `select()`
still refreshes (control).
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
A no_agent cron script owned by a profile served by a multi-profile gateway (or the
Desktop/dashboard backend) could not receive its own credential: the served profile's
.env never enters the process env (load_hermes_dotenv skips the process-global load
for a routed home), and the child-env sanitizer only resolved a terminal.env_passthrough
name through the profile secret scope when the launch environment already carried that
name. So the documented declaration mechanism (SECURITY.md 2.3) could not supply a
value that exists only in the served profile's scope, while the launch profile's own
.env credentials rode into the served profile's script child unstripped.
- tools/env_passthrough.scoped_passthrough_additions: declared names the bound scope
holds but the env being filtered lacks; reads the scope alone (no os.environ, no
other profile), empty without a scope so single-profile spawns are byte-identical.
- _scrubbed_env (terminal, cron scripts, bg processes, search workers) and
execute_code's _scrub_child_env overlay those names after filtering.
- cron _run_job_script and terminal _make_run_env strip the launch profile's .env
residue from the base (strip_launch_profile_env, a no-op for the launch profile's
own jobs) before the filter, so the scope overlay lands after the strip.
Why not the 1Password-only shape of #114218: the gap is the declaration mechanism, not
one vendor's token, and an unconditional token pop broke the single-profile .env flow.
Docs: cron no_agent credential section, secrets child-process section, security
passthrough note.
Fixes#114209
Supersedes #114218
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
The default stdio child cwd now reads agent.runtime_cwd.resolve_context_cwd()
at spawn time (previous commit), but ACP new_session/load_session register the
client's MCP servers outside the per-turn cwd pin, so a hosted ACP session with
logical cwd /workspace/a — the reporter's exact scenario — still spawned its
stdio servers in the hermes-acp process directory. Run register_mcp_servers in
a copied context with set_session_cwd(state.cwd); run_coroutine_threadsafe
carries that context onto the MCP loop task and the long-lived server task
copies it, so reconnects respawn in the same directory.
Also: trim the salvaged tests to the two invariants (session pin becomes the
default; explicit config cwd wins) — the TERMINAL_CWD fallback and the
missing-directory→None cases are runtime_cwd's own contract, pinned in
tests/agent/test_runtime_cwd.py, and the native default (no anchor → None) is
already pinned by test_start_preserves_native_default_cwd. Document the `cwd`
key and its default in the MCP docs.
Shared-process caveat recorded in the PR body: MCP connections are per
process/registry scope, not per chat session, so the anchor is read once at
connect time (profile-level terminal.cwd for gateway/cron; the owning session
for ACP-provided servers).
Widen the anonymous-first decision from #114545 to the whole class. The
clone-only wrapper (_clone_with_auth_fallback) folds into one runner,
git_credentials.run_git_with_credential_fallback, used by every site that
attached a stored credential unconditionally:
- plugins_cmd._run_plugin_git -> clone, the pinned --ref fetch in
_checkout_exact_revision, and the pull in _git_pull_plugin_dir
(`hermes plugins update`), which the thread on #114526 showed still
failing with "could not read Username ... terminal prompts disabled"
after the clone succeeded anonymously
- mcp_catalog._do_git_install (catalog MCP git installs)
- profile_distribution._git_clone (profile distributions from a git URL)
Why the runner and not per-caller guards: the decision belongs to the
credential layer once; a rejected credential sent pre-emptively is what
turns a working anonymous clone into a 401 that git can only answer with
the prompt noninteractive_git_env disables. The credential-required
classifier also matches git's "returned error: 401" spelling (the picked
version needed a trailing space that git never prints) and bytes output.
Side effect: `gh auth token` is now shelled out only after a refusal, so
public installs no longer pay that subprocess at all.
Tests: the sibling verbs (fetch, pull) x (public, refused) and the negative
that a non-credential failure never resolves a credential; the SSH no-op
case from the pick is already asserted by the private-remote test. Docs:
website/docs/user-guide/features/plugins.md.
Fixes#114526
Co-authored-by: holny <holny@users.noreply.github.com>
Closes the third ask of #113620: an operator on a shared board can see what
this home believes it may claim. Text output gains a leading
`kanban.dispatch_profiles: <any | names | none (fail-closed: …)>` line;
--json now returns {"dispatch_profiles": ..., "tasks": [...]} with the
per-task rows unchanged. Resolution reuses _dispatch_profile_allowlist so the
CLI and the dispatcher can never disagree.
Follow-up to the salvaged #113623 hunk: a config read that raises now logs a
WARNING naming the error before claiming nothing, so a corrupt config on a
shared board is visible instead of silently changing this home's claim
scope. Docs: the example no longer shows `dispatch_profiles: null` as the
"any profile" default (null is now fail-closed like `[]`); omit the key for
"any". Trimmed the contributor's positive-control test (the two invariant
tests cover the fixed class; the listed-profile control is the live probe).
Fixes#113620
Salvages #113623
The decomposer resolves both `kanban.default_assignee` (unrouted children)
and `kanban.orchestrator_profile` (the root after fan-out) as
"explicit config, else the active profile". The active profile is whatever
HERMES_HOME hosts the dispatcher — in the report an incognito `private`
profile with no credentials — so every unrouted child AND the root card
itself were silently re-owned by a profile that can never do the work.
The resolver now tries the root task's own assignee before the active
profile: explicit config → root assignee (if it names an existing profile)
→ active profile. Slim redo of #114303's decompose hunk at the resolver so
it covers the orchestrator fallback too; tests and docs from #114303.
Fixes#114294
Salvages #114303
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
A Kanban notify+wake subscription whose destination is a raw session id on the
shared api_server (platform=api_server, chat_id=<session id>, notifier_profile=
served secondary) was never delivered under gateway.multiplex_profiles: no
profile_routes entry can anchor a session id, so _adapter_for_subscription()
returned None and _claim_for_sub skipped the row before claiming (cursor frozen,
no delivery attempt). A platform-wide api_server route is not the fix — it
matches every api_server destination, so the default profile's own api_server
subscriptions would fail closed instead.
The wake leg had the same shape: _self_post_chat_completion() always POSTs to
the unprefixed /v1/chat/completions with the PRIMARY adapter's key, so a served
profile's wake turn would resume the session in the default profile's store.
/p/<profile>/v1/... requires that profile's own API_SERVER_KEY, which a
route-only profile legitimately does not have.
Authorize the shared adapter only when the subscription's session id is a row in
the served profile's own state.db stamped for that profile (NULL legacy stamp =
the store's own profile, the same rule the dashboard session routes apply) and
run the wake in-process under that profile's runtime scope through a new
APIServerAdapter.run_internal_session_turn(). Unknown session, foreign stamp,
unserved/deleted profile and unreadable store all fail closed; the concurrent-run
cap defers the wake (cursor rewind, retry) instead of bypassing it.
A no-DCR entry (client_id + redirect_port, e.g. the Asana manifest) has
http://localhost:<port>/callback registered with the vendor, which matches
redirect URLs exactly; the dashboard flow overrode it with its own callback
URL, so the in-app Authorize button could never complete for such entries.
The pinned loopback listener now wins (over the dashboard URL and any cached
registration URI); the dashboard flow only publishes the authorization URL,
and the stdin paste reader stays off under a dashboard flow. Docs/post_install
say the approving browser must run on the Hermes machine.
The V1 beta server https://mcp.asana.com/sse is retired and Asana's V2 server
(https://mcp.asana.com/v2/mcp, Streamable HTTP) has no Dynamic Client
Registration: every client must be an Asana "MCP app" the user registers in
the developer console. The salvaged manifest fixed the URL and documented the
manual `oauth:` block; this commit makes the catalog install itself produce
that block so no hand-edit of config.yaml is needed.
- hermes_cli/mcp_catalog.py: manifests may pin a closed `auth.oauth` mapping
(client_id, client_secret, redirect_host, redirect_port, scope) that
`_build_server_config` copies to `mcp_servers.<name>.oauth`; every `${VAR}`
it references must be declared in `auth.env`, mirroring the api_key header
contract, so a placeholder can never reach the token endpoint as a literal.
`install_entry` now prompts `auth.env` for OAuth entries too (dashboard and
Desktop already render `required_env` regardless of auth type).
- optional-mcps/asana/manifest.yaml: declare ASANA_CLIENT_ID/SECRET, pin the
client + `http://localhost:27890/callback` (Asana matches the redirect URL
exactly), and rewrite post_install around the MCP-app registration steps.
- tests: shipped-catalog invariant (no manifest installs the retired `/sse`
URL; the Asana client credentials are declared `${VAR}` refs; callback
pinned) + install-path invariant for the new `auth.oauth` block, including
the undeclared-reference rejection.
- docs: catalog section on entries that need a user-owned OAuth app.
Live: `hermes mcp install asana` + `hermes mcp login asana` under a temp
HERMES_HOME. Before: config `url: …/sse`, no oauth block, authorize URL on the
V1 server (mcp.asana.com/authorize) with a DCR client and 127.0.0.1 redirect.
After: config `url: …/v2/mcp` + oauth `${ASANA_CLIENT_ID}` refs, .env holds
the values, authorize URL on app.asana.com/-/oauth_authorize with the
configured client_id, redirect_uri=http://localhost:27890/callback and
resource=https://mcp.asana.com/v2/mcp — the flow Asana's guide documents.
Follow-up to the salvaged #113687 / #113682 commits:
- `removed_annotation(name, dir_path, removed_entries=None)` resolves the kill list once itself
when no list is passed and matches against it; the `removed_annotation_batch()` wrapper and its
name-keyed dict are gone (a hub row keyed by `name` could shadow a same-named nested plugin).
Both surfaces (`plugins list`, `_merged_plugins_hub`) call `resolved_removed_entries()` once
and pass the list per row.
- The negative cache is one module float deadline instead of a URL-keyed dict plus two helpers;
the duplicated stale-cache read is one `_stale_live_cache()`.
- Tests trimmed to two invariants: failure memory honours the TTL and still serves a stale copy;
hub rebuild + `plugins list` each cost one network attempt and still annotate an in-tree removal.
The #113682 route test's fake gains the new third argument.
- Docs: the live-refresh section says a failed fetch is remembered for a minute.
Builds on the previous commit (alerted treated like closed): a permanently
silent `alerted` incident would hide a job that stays broken for days, so the
gate now withholds only inside `cron.failure_repeat_alert_hours` (default 6,
0 = re-alert on every failing run) and lets exactly one reminder through, which
re-stamps the window.
- cron/incidents.py: `alerted_at` column (added in place to existing ledgers via
add_column_if_missing); set_incident_state(..., "alerted") stamps it every time
so a reminder restarts the cooldown; resolved->detected re-open clears it so
the same error after a green run alerts immediately.
- cron/scheduler.py: `_repeat_alert_withheld` reads the stamp; a missing or
unparseable stamp (pre-migration row) delivers rather than swallowing the
alert; `closed` still wins; the unreadable-ledger fail-open is unchanged. The
crash path (`_deliver_crash_failure`) shares the gate through
`_upsert_incident_for_failure`.
- hermes_cli/config_defaults.py + website/docs cron page: document the key.
- tests: fold the contributor's two tests and the old "unacked failures keep
alerting per run" change-detector into two invariants (unit gate incl. 0 /
legacy row / closed; end-to-end alert once -> reminder once -> recovery re-arms).
Fixes#113665
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
_resolve_fallback_runtime swapped requested_provider/model after the primary's
AuthError but left reasoning_config at the launch model's effort, so the first
request to an always-thinking fallback model (zai/glm-5.3-flash with a
reasoning_overrides entry) went out with the primary's `medium` and 400'd.
Kanban workers hit it on every spawn since each is a fresh `hermes chat`.
Route the swap through `_resolve_cli_reasoning`, the CLI-level chokepoint that
/model, /new and --resume already use (77f5de23dd), instead of an inline
try/except-pass copy: it reads the same in-memory CLI_CONFIG startup resolved
the primary against, so nothing can fail and the swap is never at risk. Keep
the `_explicit_reasoning_config` guard from #113497: an explicit
`--reasoning` is the user's intent for this run and outranks the fallback
model's config (kanban passes it per task the same way).
Test: one real-HermesCLI invariant, parametrized over the two cases
(fallback override applies; explicit flag survives), replacing the
SimpleNamespace stub test. Docs: fallback-providers.md.
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
A recurring job paused before one of its slots and resumed after it lost
that occurrence silently: resume_job recomputed next_run_at from now, so the
elapsed slot was neither fired nor recorded — no execution row, no incident,
no log line, and last_dispatch stayed on the previous run (the reporter's
daily job showed next_run_at jumping two cadences with nothing in between).
resume_job now leaves a past stored next_run_at in place for cron/interval
jobs and logs that it did. The first tick after resume then applies the
existing occurrence policy to that instant — late fire within grace, one
collapsed catch-up run past grace, or the loud "missed its scheduled time"
skip when cron.catch_up_missed is false — so the slot is accounted for the
same way a restart-gap slot is (#107485 contract: every recurring occurrence
runs once or its skip is logged). One-shots, future instants and jobs
created --paused (next_run_at null) still recompute from now.
Salvaged from PR #114296 (resume_job hunk only; its ride-along copies of
main's self-removal/fire-claim-skew code and issue-numbered test were
dropped).
Follow-up to the salvaged #73026 (@ryangu00):
- `_deliver_to_bot_chat` now rebinds `content` to the redacted copy before anything reads it, so
the durable deferred record (`cron/bot_chat_delivery.defer`) and its later replay carry the
scrubbed payload as well as the live-owner and CLI lanes. Previously only the assembled
`message` was scrubbed while the raw output was persisted for replay. Redaction is idempotent,
so the replay's second pass is a no-op.
- The seven contributor tests collapse into two parametrized invariants: every outward lane
(platform send / session mirror incl. job-name splice / bot-chat turn) masks the secret while
the payload and framing survive; and egress redaction is forced (independent of
`security.redact_secrets`) and fails closed (a raising redactor replaces the payload).
- Cron docs: note what is redacted on delivery, that it ignores `security.redact_secrets`, and
what is deliberately not stripped (credential-named URL params, shapeless user secrets).
The failure-notice lane #55271 targeted (`_summarize_cron_failure_for_delivery`) exits through
the same `_deliver_result` chokepoint, so it is covered without a second redaction site.
Co-authored-by: necoweb3 <sswdarius@gmail.com>
complete_task and request_review return bare False when the dependency
gate refuses, and _patch_status only enriched the 409 for status=ready,
so PATCH status=done/review on a gated card answered "not valid from
current state" and the bulk entry said "transition refused". Consult
unsatisfied_parents on a refused done/review and name the parents in
the 409 detail and the bulk entry error, matching the CLI/tool wording.
Salvage follow-up on the cherry-picked #113398 (@fangliquanflq) and #113452
(@JoaoMarcos44) commits, trimmed to the salvage bar and widened to the class.
- link: drop the HERMES_KANBAN_BOARD binding on the run-id ownership proof
(tools._worker_run_id / kanban._worker_run_id_for). complete/block prove
ownership with the same task-scoped run id and never board-bind it; the
guarded case is a forged cross-board task-id collision, not a live path.
Dropped the two cross-board tests and the duplicate db owner test.
- complete: move the open-parents enumeration into kanban_db.unsatisfied_parents
next to _parents_satisfied so the tool, the CLI (`hermes kanban complete`
printed "unknown id or terminal state" for the same refusal), kanban_show
(`unsatisfied_parents`) and diagnostics share one query.
- show: new `running_with_open_parents` diagnostic (hermes kanban show /
diagnostics / dashboard) for a running card whose parent is not done —
the "surface unresolved parents on a running card" ask of #113374.
- tests trimmed to invariants; docs for the refusal text, the show field and
the diagnostic.
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>