Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:
- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
`choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
completed `reasoning` output item ahead of the message (and ahead of that step's
`function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
from the assistant messages the agent already persisted (`build_assistant_message`
stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
in non-stream mode the agent fires the callback per provider delta AND once more with
the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
`conversation_history`. Responses SDK clients replay a prior response's `output`
list as the next `input`; before, the item became an empty `user` message (in
`input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
(`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
`output_item.done` item and the non-streaming shape.
Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).
- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
`_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
keeping them distinct from answer text. The lossy 500-char `reasoning.available`
progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
(the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
(`output_item.added`, `reasoning_summary_part.added`,
`reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
`output_item.done`), closed before the next message/function_call item opens
and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.
Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
Review findings on the first cut:
* A hosted provider with a stale model.ollama_num_ctx passed the floor for
a 40K model although nothing ever raises a hosted window. The served
window now counts only when the endpoint is local (is_local_endpoint),
the same gate the num_ctx probe uses.
* Moving the whole num_ctx phase ahead of the floor also moved tek's
compressor clamp ahead of it, which flipped two observables: a
model.context_length above a sub-64K num_ctx was rejected instead of
constructed-and-clamped, and the floor's message reported the clamped
value with advice (set model.context_length) that could not help. The
phase is split: resolution runs before the floor, the clamp
(_clamp_compressor_to_ollama_num_ctx) stays at its original position,
so every case main constructed still constructs with identical
compressor numbers and the floor's message is unchanged.
Tests: the harness is a module-level helper so the new class no longer
re-collects the parent's tests (13 -> 11 collected); the negative pins
the local-endpoint gate (red when the gate is dropped) instead of a case
main already rejected.
An Ollama server serves num_ctx. A Modelfile or model.ollama_num_ctx at
65536 is a usable window even when the model metadata advertises 40960,
yet the floor ran before num_ctx was resolved and read only the probed
window, so the agent (a cron job reaching a local fallback in the report)
refused to construct with "context window of 40,960 tokens".
num_ctx resolution now runs before the floor and the floor takes
max(probed, served). The compressor keeps tek's one-directional clamp
(cb71d5f1b1): it still targets the smaller probed window, so nothing about
compaction thresholds changes; a served window below 64K is still rejected.
Rebuilt from PR 100475 by fangliquan (the agent_init it targeted was
decomposed since); the unrelated cron pin contract test there is not taken.
Co-authored-by: fangliquan <fangliquan@qq.com>
Review findings on the mirror, all verified live against throwaway items:
* Identity gate. The service name is fixed but the file path honours
CLAUDE_CONFIG_DIR, and Claude Code rotates on its own schedule, so the
item under the service can hold a different login's pair or a newer
rotation. Overwriting it would be the bug in the other direction. The
refresh token that was just POSTed is threaded through
`_write_claude_code_credentials(spent_refresh_token=...)` from both
callers (the singleton refresher and the pool commit) and the mirror
only updates an item whose `refreshToken` equals it.
* Attribute parsing. `security` prints an attribute as `"text"` when it
is plain printable ASCII — UNescaped, an embedded `"` appears raw — and
as `0x<HEX> "<echo>"` otherwise. The old regex modelled `\"` escaping
that never happens: a `"` in the account truncated it and the mirror
created a second item; any non-ASCII byte made it return "" and the
mirror silently no-oped. One `find-generic-password -g` call now yields
account and payload together, parsed line-anchored in both encodings.
* Fail-soft is now total (`except Exception`): the file commit already
succeeded when the mirror runs, and a raise here made the refresher
mark a landed rotation as consumed-uncommitted.
* `quoted` lambda -> nested def; `ensure_ascii=True` made explicit since
`-w` returns non-ASCII payloads as hex.
Two pre-existing test doubles for the writer accept the new keyword.
Two defects in the cherry-picked mechanism, both found live on macOS:
* `add-generic-password -w` with no value prompts on /dev/tty when a
terminal exists, so the CLI refresh path hung for the 10 s timeout and
wrote nothing; with no terminal it read one line from stdin, hit EOF on
the confirmation read and stored an EMPTY password — bricking the very
item this exists to keep fresh. The command line now goes to
`security -i` on stdin with the payload hex-encoded (`-X`): no argv, no
tty prompt, no quoting of the JSON.
* `-a getpass.getuser()` assumed the item's account is the login user;
`-U` matches on account + service, so a mismatch would have created a
second item Claude Code never reads. The account is read from the
existing item's `acct` attribute and the mirror is skipped without one.
The command builder is a pure function so the no-argv / hex / account
contract is tested on every lane; the Darwin-gated no-entry no-op test is
`macos_only` instead of faking `platform.system`. Tests trimmed to the
invariant bar; the merge test pins that `mcpOAuth` siblings survive.
Live E2E on a throwaway Keychain item seeded with 24 mcpOAuth entries:
triple rotated, scopes/subscriptionType and all 24 siblings preserved,
one item, updated under the seeded account.
On macOS the Keychain is Claude Code's authoritative credential store, but
Hermes only ever wrote ~/.claude/.credentials.json. Since the refresh token is
single-use and rotating, every Hermes-initiated refresh left the Keychain
holding an already-invalidated token, which Claude Code then spent into
invalid_grant and discarded ("Login: Expired").
_write_claude_code_credentials now mirrors the committed refresh into the
existing "Claude Code-credentials" entry via security add-generic-password,
merging the rotated token triple over the existing payload so
subscriptionType / rateLimitTier / scopes survive. The payload is fed on stdin
(bare -w), never argv. Fail-soft: a mirror failure is logged, never raised —
the file commit already succeeded and the resolver resolves from it. No-op off
Darwin and when no entry exists (never create one the user has not).
Add a raw payload reader (metadata preserved), a pure merge helper, and the
mirror; extend the conftest keychain guard to neutralize the new writer in any
test that hasn't opted in.
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.
`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:
1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
generations, usage rows, system prompt) to the owning store, parents before
children so lineage survives, copy-then-delete so a crash leaves a duplicate
the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
-> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
import re-injects them into routing every boot).
Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).
Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).
Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
`_profile_name_for_source` already accepts `adapter_profile=None`, so the
branch in `_stamp_routed_profile` was two spellings of the same call.
Collapse it and give the two test doubles that replace the method the
same keyword so they keep matching the production signature.
Salvage follow-up to #106019.
A matcher exception fell through to the default profile, so a transient
failure while resolving a sender route silently served the message from the
default profile's runtime. Raise the existing rejection instead, which the
ingress gate already drops on.
Only the failure path changes: a source that simply matches no route keeps
falling through to the default/active profile as before.
Voice input reused the bound text channel's cached source and replaced only
`user_id`, so a second speaker inherited the profile resolved for the first.
The ingress gate does not re-resolve a source that already carries a profile.
Re-resolve at the voice call site instead of clearing the profile: clearing
would drop the receiving bot and could re-home voice arriving on a secondary
profile's own bot. `_voice_input_source` reattaches the transport provenance
`from_dict` discards, and `_stamp_routed_profile` takes the receiving bot's
profile as the fallback when no route matches.
Kanban re-subscription could not repair a row created before sender capture:
`user_id` was only written by the INSERT, so a legacy row stayed senderless and
the notifier's conservative fallback left it undeliverable for good. It now
self-heals like `user_id_alt`.
An explicit `user_id: null` or empty string is rejected instead of widening the
route to every sender, and `to_dict` omits the field when unset so a round-trip
cannot reintroduce it. Numeric `0` from an adapter normalizes to "0" rather than
being dropped, without changing the shared coercion used by the other fields.
Closes#33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.
Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.
Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.
The switched-provider comment implied the guard catches the resolver raise;
it is suppress() leaving st.* at the session values that does. The
stay-on-provider comment now states the one semantic widening of moving to
the shared helper: an empty resolver result keeps the session endpoint on
any host, not only off openrouter.ai.
Gate finding: resolve_runtime_provider(requested="custom") never returns an
empty base_url — with no model.base_url but an OpenRouter key present it
lands on OpenRouter's default host, so the "keep the current endpoint"
fallback was dead and an Anthropic session adopting `provider: custom`
hopped to openrouter.ai. The guard _creds_for_current_provider already
carries for this (#74143) is now a shared helper used by both branches;
the branch assigns once instead of assign-then-undo; the import sits with
the other from-imports; the new test's write_text carries an encoding.
Second invariant test pins the no-endpoint case.
The bare-custom credential branch kept the session's current base_url and
key whenever the target was `custom`, regardless of where the session was.
That is right for a bare-custom session picking another model on its own
endpoint, but when the TUI's per-turn config sync adopts `provider: custom`
from an OpenRouter (or any other) session, the new model was paired with
OpenRouter's URL and key and every request 400/404'd (Bug 2 in issue 73680).
Arriving from another provider now resolves the configured custom endpoint
through resolve_runtime_provider; with none configured the current endpoint
is kept, so direct aliases still supply their own URL afterwards and the
bare-custom same-endpoint case is unchanged.
Gate fold: the optional-prefix regex was spelled three times in the agent
stripper; it is one constant now (_NS_PREFIX). The cli.py docstring pointed
at run_agent._strip_think_blocks, which moved to
agent.agent_runtime_helpers.strip_think_blocks; stray blank lines dropped
from the test file.
muse-spark (opencode-go, Responses wire) occasionally serializes its next native
tool call onto the text channel as <atem:function_calls>…</atem:function_calls>
instead of a function_call item. With real tool work earlier in the turn and
finish_reason=stop, the literal XML is then delivered as the turn's final answer
(observed on a production cron worker: the turn ended on the full XML block; the
task itself had executed).
Teach the tool-call XML patterns an optional namespace prefix (?:[\w.-]+:)? on
the open/close tags and the unterminated-opener tail, in every copy of the
stripper: agent/agent_runtime_helpers.py (_TOOL_CALL_BLOCK_PATTERNS,
_STRAY_TOOL_CALL_CLOSER_PATTERN, _UNTERMINATED_TOOL_CALL_PATTERN) and the
display-side cli.py::_strip_reasoning_tags, which the test suite requires to stay
in sync. Covered positions: closed block, stray closer, unterminated opener
(stream cut mid-serialization).
With the block stripped to nothing, the existing empty-response recovery path
(agent/turn_final_response.py) re-prompts the model instead of shipping XML as
the answer.
Refs #103483 (the native-call XML leaking as text is listed there as not covered yet).
Tests: tests/agent/test_strip_reasoning_tags_cli.py — full block and
cut-mid-serialization tail, each asserted against both strippers. Proven red on
base: 2 failed / 4 passed without the pattern change; 6 passed with it.
A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.
Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.
Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.
Phase 2 of #88715.
A multiplexed gateway answered "which bot received this / who may admit it /
where does it run" in three places (`_transport_owner`,
`_authorization_home_for_source`, `_resolve_profile_home_for_source` +
`_session_key_profile`) that agreed only because they read the same fallback
chain. `gateway/session_identity.py` answers them once: `resolve_identity()`
folds `_admit_primary_source` + `_stamp_routed_profile` + the transport-owner
lookup and pins a frozen `RoutingIdentity` (transport_profile, runtime_profile,
authorization_home, runtime_home, weak transport ref) on the source as a
wire-invisible attribute, like `_transport_adapter_ref`. Under multiplexing a
route to an unserved profile raises `IdentityUnresolved` instead of a
`None`-means-default return; `"default"` is spelled out inside the object.
Additive: the existing helpers become thin readers of the identity when it is
present and keep their fallback chain when it is not, `source.profile` stays
the serialized runtime profile (None on the wire ⇔ default) and every
historical `agent:main` key is byte-identical. `replace_source()` copies a
source without losing its provenance (run_topics used to hand-copy the
transport ref).
Phase 1 of #88715; the gateway rows of #90142 / #93943.
The independent-account test drives the real load_pool() -> _refresh_entry() path
and asserts the second account POSTs its own refresh token and keeps its principal
(red on base: the row silently became account A). The alias test pins both halves
of the same-account rule: never fall back onto an older singleton, always follow a
newer re-auth.
A manual:device_code Codex pool entry never writes its rotation back to the
auth.json singleton (independent-credential contract, #39236). The singleton
sync adopted differing singleton tokens with no staleness proof, so after a
pool-side rotation the stale singleton was re-adopted over the pool's fresh
chain and the already-consumed refresh token was POSTed again
(refresh_token_reused).
Gate adoption on the singleton's last_refresh not predating the entry's own
rotation; missing stamps on either side keep the historical
adopt-on-difference behavior (#70111).
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.
Reactive, destination-scoped recovery instead:
- error_classifier: new `FailoverReason.role_alternation` for the vendor
alternation wordings (checked before the request-validation table since
the body also carries `invalid_request_error`); same abort+fallback hints
as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
loop's `_merge_user_content`), `destination_key` (base_url|provider,
model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
adjacent user turns merged, remember the destination on the facade for
the session so later iterations pre-merge, never touch destinations that
accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.
Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.
Fixes#112358
When a Desktop tab references a session id that exists in no profile DB
(deleted elsewhere, or a stale id after a profile rename / wiped backend),
the resume path already drops the window to a fresh draft without toasting
or hot-looping (62af32efe7, bounded by goneSessionVerdict). But the text
the user had typed into that session stayed stashed under the dead key in
`hermes:composer-drafts:v3` — not lost, yet invisible, because nothing ever
opens that key again. From the user's side the message just vanished.
The gone verdict now announces the dead key (`announceGoneSessionDraft`),
and the composer's swap onto the fresh-draft scope consumes it once
(`adoptGoneSessionDraft`): the stash moves into the new-chat bucket, the
composer is seeded with it, and an inline "Restored your unsent message"
strip with Undo appears above the input. Undo returns the text to the dead
key while it is unchanged (still recoverable by the same path); once edited
it only dismisses. Keyed on the one-shot announcement rather than on the
composer observing an id→fresh transition, because the composer remounts
across the drop (a loading route mounts no composer). A fresh draft that
already has text is never clobbered; an empty stash shows nothing.
Offer, don't hijack (apps/desktop/AGENTS.md): no navigation beyond the drop
that already happened, no focus move, no toast.
Live (headless Electron over CDP, worktree backend): type into a live
session → delete its row + restart the backend → reload at the dead id.
Before: route drops to #/, composer empty, text only in localStorage under
the dead key. After: same drop, composer holds the text, strip + Undo shown;
Undo empties the composer and puts the text back under the key; a second
open of the dead id shows no strip; a dead id with no stash shows nothing.
Fixes#111868
`display.show_reasoning` is a client-side display decision since 40f2368875,
and the Desktop transcript honors it through `$showReasoning` (#115335). The
composer path did not: `/reasoning` was marked `desktop="advanced"` in the
Python slash registry, so the Desktop refused it as "not available", while a
gateway-routed `/reasoning hide` would only have written config.yaml and left
the atom waiting for the next config refresh.
Route `/reasoning` to the gateway's `config.set key=reasoning` — the Ink TUI's
path (ui-tui/src/app/slash/commands/session.ts) — and mirror the answer:
`hide`/`show` flip `$showReasoning` on the spot, effort levels leave the gate
alone and get session scope (a throwaway slash-worker CLI could never pin the
live session's effort). No arg reports the current effort and display, like
Ink. Drops the registry's `advanced` marker and regenerates the desktop dump.
Live (real `hermes serve`, scratch HERMES_HOME, real WebSocket): config.set
key=reasoning value=hide -> {value: hide}, config.yaml show_reasoning=False;
show -> True; high -> effort only, display untouched.
Fixes#111761
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.
- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
current users). When false:
- `read_claude_code_credentials()` — the only reader of the borrowed Claude
Code login — returns None, so the resolver fallback, the expired-token
refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
file; the pool prunes a `claude_code` row an earlier adopting process
persisted.
- `_recover_codex_tokens_from_cli` returns None for both automatic recovery
paths (rejected refresh, half-empty singleton); the real AuthError is
surfaced instead. The interactive import offer in `hermes auth add
openai-codex` still asks first and is unaffected.
- One INFO line per process the first time adoption would have happened;
`hermes auth list` / `hermes auth status anthropic|openai-codex` print the
same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.
Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
The ui_meta mirror of a room is head-trimmed to 16 messages and then to
the 48 KB envelope, silently. It is the only on-disk copy of the room, so
an agent that consults it to check what the user said reads a bounded
window with no sign that it is bounded (#114341, secondary). Stamp
`omitted: N` on the projected room whenever entries fall off the head,
recomputed after every byte-budget shift so the envelope bound stays
honest, and carry the larger writer's count through the publish merge as
a lower bound. Local room state never receives the field: the merge-back
into $groupChats builds from explicit fields.
The ui_meta projection sliced messages at 1200 chars with no
signal that the body continued. Stamp truncated and keep the
marker inside the existing per-message budget.
The salvaged tests exercised the formatGroupDeltaLines helper in
isolation, so a call-site regression (a bare slice reappearing in
prepareGroupRoundMember) would have stayed green. Drive the room through
runGroupChatRounds instead and assert on the prompt the gateway actually
received: over-long delta -> marker naming the dropped count and the
dropped head absent; delta that fits -> no marker. The helper-level pair
is dropped so the file keeps two invariant tests for this fix.
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:
- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
already did this in `_handle_reset_command`, the second request is
idempotent) and the idle `_handle_stop_command` tail, which replied
"No active task to stop." while a background child was running; it now
stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.
The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.
An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.
Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.
Part of #114456
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).
Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.
`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.
No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.
Part of #114587
Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
A dispatcher-spawned worker whose turn failed on a provider error that no
retry can heal (credential rejected: auth / auth_permanent, model_not_found,
ssl_cert_verification) now exits KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78,
BSD EX_CONFIG) instead of a plain 1, so the dispatcher can tell "the
provider will reject every further spawn" from "this attempt crashed".
The classification is the agent's own FailoverReason verdict already on
the turn result (failure_reason) — no error-text regex. billing stays in
the transient set: credit comes back, a revoked key does not.
Shared by both one-shot paths (-q and -Q) and gated on HERMES_KANBAN_TASK,
so a person's `hermes chat -q` keeps exiting 1 on the same error.
Part of #114587
SupervisedTickerThread (#111010) treated any ended thread as dead. An external provider's
start() (Chronos) arms remote one-shots and returns by design, so every housekeeping tick
logged "Cron ticker thread died without a stop request; restarting" at ERROR and re-ran
start() - a fresh recover_interrupted + NAS list/arm reconcile once a minute on every hosted
instance. Track whether the target escaped with an exception and respawn only then; the
built-in ticker returns normally only on stop_event.
NAS maps the agent-cron bearer to a provisioned instance via an `agent:*` client or the
`hermes-cli-vps` bootstrap session (hermes-portal `server/agent-cron/instance-auth.ts`). A
container whose auth.json holds a plain `hermes-cli` user login is refused with 403
invalid_client on every arm, re-arm and list, for the life of that credential - and the
re-login users try first (`hermes auth logout nous` + device code) replaces the bootstrap
session, making it permanent. Chronos previously logged one bare warning per job and left
the jobs with no trigger at all: they only ran through the misfire sweep, minutes late.
NasCronClientError now carries the HTTP status and the OAuth `error` code. On an identity
rejection the provider logs ONE warning that names the real remedy (restore the hosted
credential from the Nous Portal; re-login cannot fix it), stops calling NAS, and starts the
built-in ticker with the gateway's own adapters/loop so scheduled jobs keep firing on time.
Transient 5xx/transport failures keep retrying on the next reconcile.
Supersedes #97566 (@wesleysimplicio), whose runtime-credential swap resolves to the same
bearer on main (`agent_key` is the access token) and so could not change the 403.
Both repair paths now run the memory-provider refresh between the tool-dep
restore and the plugin-dep reapply, exactly like _sync_python_dependencies_after_pull
and the ZIP path, so the three routes stay interchangeable.
The venv-repair and runtime-repair paths reinstalled core, lazy and tool
deps but never refreshed memory-provider bridge packages, unlike the
git-pull and ZIP paths. Both repair paths now heal them last.
Fixes#113741.
OpenRouter (_generate_via_chat, _generate_via_image_api) and OpenAI gpt-image
called record_token_usage only after image extraction and save succeeded. A
token-billed 200 that returned text but no image (the "returned no image"
fallback case), an empty `data` list, or a failed save therefore consumed
billed tokens that never reached session_model_usage; on the fallback chain
only the fallback model's tokens were recorded.
Move the call to immediately after a successful post_json / SDK call, before
extraction/save, in all three paths. One call per HTTP success, so no double
count; failure responses (non-2xx, timeouts) still record nothing because
there is no body to read usage from.
Tests: the parametrized invariant in test_openrouter_compat_provider gains the
no-image-with-usage case for both surfaces (result is `empty_response`, one row
still recorded); the OpenAI test is parametrized on has_image the same way.
Widen the salvaged Image API hunk (#114340) to the whole class:
- plugins/image_gen/_common.py: record_token_usage() — one helper feeding the
aux accounting chokepoint (agent.aux_accounting.record_aux_usage) with task
"image_generation", the billing provider and the priced model id. Dict and
SDK usage objects alike; a body without tokens is a no-op, as is a call
outside a turn.
- plugins/image_gen/openrouter: the /chat/completions path (the DEFAULT model
chain — openai/gpt-5.4-image-2, google/gemini-3-pro-image — is chat-only and
token-billed) now records too; the Image API path uses the shared helper
(the contributor's local _record_image_api_usage is folded into it) and both
pass base_url so pricing resolves the route. Task renamed image_gen ->
image_generation to match the other aux task names.
- plugins/image_gen/openai: gpt-image bills per text/image token; record the
Images API usage block under the API model (gpt-image-2), not the Hermes
quality-tier label.
- FAL, xAI, Krea, DeepInfra, Meta, openai-codex return no token usage and stay
unrecorded (nothing to bill per token).
- tests: the contributor's two tests folded into one invariant parametrized
over chat / Image API / no-usage control; one OpenAI invariant.
- docs: image-generation.md "How It Works Internally" gains the accounting step.
Fixes#114324
OpenRouter token-billed image calls parsed usage into extra only, so
session_model_usage never saw them. Record prompt/completion tokens via the
ambient aux accounting on success; flat-fee responses without token usage
write nothing.
Fixes#114324.
When a per-file pytest subprocess dies by signal (the sqlite cross-thread
close in #113186 was a SIGSEGV after every test had passed), faulthandler
prints "Fatal Python error: Segmentation fault" and no summary line, so
every count parses to 0. The runner filed that under "1 file where no
tests ran (collection/import error, ...)" beneath a summary that read
"0 failed" and exited 1 — two wrong diagnoses for one real bug, and it
was misread as a runner problem twice on main.
The runner now detects a signal death or a "Fatal Python error:" banner,
prefixes the captured output with the diagnosis (same convention as the
timeout path), marks the progress line CRASHED, counts "N files CRASHED"
on the summary line, lists the file in its own failure bucket, and no
longer trips the "NO TESTS RAN" guard for a crash that ran tests. The
flake retry already covers crashes (any non-zero rc), so nothing changes
there.
Moves sha and docs_url to a4843d82c0b2c8443830b79f5e0a19b8a78a8303 (v0.19.2, peeled
commit). Capabilities block unchanged and re-verified at that commit.
Signed-off-by: SmokeDev <degensmoke@gmail.com>
Moves the pin from 8aeac6a6 to 4764f687, the peeled commit of tag v0.19.1.
The only source change on the auth surface is #146 (read-only Codex lane,
authored by @teknium1), merged as-is. plugin.yaml changed only its version
field; hook registrations, tools, middleware and required env are unchanged
against the declared capabilities block.
Signed-off-by: SmokeDev <degensmoke@gmail.com>