Commit Graph

37633 Commits

Author SHA1 Message Date
teknium1
004c8a51f0 feat(api-server): reasoning on non-streaming replies; echoed reasoning input items are ignored
Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:

- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
  `choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
  completed `reasoning` output item ahead of the message (and ahead of that step's
  `function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
  from the assistant messages the agent already persisted (`build_assistant_message`
  stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
  in non-stream mode the agent fires the callback per provider delta AND once more with
  the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
  `conversation_history`. Responses SDK clients replay a prior response's `output`
  list as the next `input`; before, the item became an empty `user` message (in
  `input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
  (`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
  `output_item.done` item and the non-streaming shape.

Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
2026-09-19 00:46:31 -07:00
teknium1
b957782917 feat(api-server): stream model reasoning on /v1/chat/completions and /v1/responses
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).

- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
  `_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
  keeping them distinct from answer text. The lossy 500-char `reasoning.available`
  progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
  (the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
  (`output_item.added`, `reasoning_summary_part.added`,
  `reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
  `output_item.done`), closed before the next message/function_call item opens
  and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.

Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
2026-09-18 23:58:24 -07:00
kshitijk4poor
afaa53e5fe fix(agent_init): judge the served window only for local endpoints; clamp after the floor as before
Review findings on the first cut:

* A hosted provider with a stale model.ollama_num_ctx passed the floor for
  a 40K model although nothing ever raises a hosted window. The served
  window now counts only when the endpoint is local (is_local_endpoint),
  the same gate the num_ctx probe uses.
* Moving the whole num_ctx phase ahead of the floor also moved tek's
  compressor clamp ahead of it, which flipped two observables: a
  model.context_length above a sub-64K num_ctx was rejected instead of
  constructed-and-clamped, and the floor's message reported the clamped
  value with advice (set model.context_length) that could not help. The
  phase is split: resolution runs before the floor, the clamp
  (_clamp_compressor_to_ollama_num_ctx) stays at its original position,
  so every case main constructed still constructs with identical
  compressor numbers and the floor's message is unchanged.

Tests: the harness is a module-level helper so the new class no longer
re-collects the parent's tests (13 -> 11 collected); the negative pins
the local-endpoint gate (red when the gate is dropped) instead of a case
main already rejected.
2026-09-19 11:44:15 +05:30
kshitijk4poor
3de7140bcd fix(agent_init): the 64K context floor judges the window Ollama serves, not the GGUF metadata
An Ollama server serves num_ctx. A Modelfile or model.ollama_num_ctx at
65536 is a usable window even when the model metadata advertises 40960,
yet the floor ran before num_ctx was resolved and read only the probed
window, so the agent (a cron job reaching a local fallback in the report)
refused to construct with "context window of 40,960 tokens".

num_ctx resolution now runs before the floor and the floor takes
max(probed, served). The compressor keeps tek's one-directional clamp
(cb71d5f1b1): it still targets the smaller probed window, so nothing about
compaction thresholds changes; a served window below 64K is still rejected.

Rebuilt from PR 100475 by fangliquan (the agent_init it targeted was
decomposed since); the unrelated cron pin contract test there is not taken.

Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-19 11:44:15 +05:30
kshitijk4poor
14a3463454 fix(anthropic): mirror only the Keychain item that held the spent pair; parse both security attribute encodings
Review findings on the mirror, all verified live against throwaway items:

* Identity gate. The service name is fixed but the file path honours
  CLAUDE_CONFIG_DIR, and Claude Code rotates on its own schedule, so the
  item under the service can hold a different login's pair or a newer
  rotation. Overwriting it would be the bug in the other direction. The
  refresh token that was just POSTed is threaded through
  `_write_claude_code_credentials(spent_refresh_token=...)` from both
  callers (the singleton refresher and the pool commit) and the mirror
  only updates an item whose `refreshToken` equals it.
* Attribute parsing. `security` prints an attribute as `"text"` when it
  is plain printable ASCII — UNescaped, an embedded `"` appears raw — and
  as `0x<HEX>  "<echo>"` otherwise. The old regex modelled `\"` escaping
  that never happens: a `"` in the account truncated it and the mirror
  created a second item; any non-ASCII byte made it return "" and the
  mirror silently no-oped. One `find-generic-password -g` call now yields
  account and payload together, parsed line-anchored in both encodings.
* Fail-soft is now total (`except Exception`): the file commit already
  succeeded when the mirror runs, and a raise here made the refresher
  mark a landed rotation as consumed-uncommitted.
* `quoted` lambda -> nested def; `ensure_ascii=True` made explicit since
  `-w` returns non-ASCII payloads as hex.

Two pre-existing test doubles for the writer accept the new keyword.
2026-09-19 11:25:23 +05:30
kshitijk4poor
bf5a6f6ae6 fix(anthropic): mirror the Keychain item through security -i with a hex payload, under its own account
Two defects in the cherry-picked mechanism, both found live on macOS:

* `add-generic-password -w` with no value prompts on /dev/tty when a
  terminal exists, so the CLI refresh path hung for the 10 s timeout and
  wrote nothing; with no terminal it read one line from stdin, hit EOF on
  the confirmation read and stored an EMPTY password — bricking the very
  item this exists to keep fresh. The command line now goes to
  `security -i` on stdin with the payload hex-encoded (`-X`): no argv, no
  tty prompt, no quoting of the JSON.
* `-a getpass.getuser()` assumed the item's account is the login user;
  `-U` matches on account + service, so a mismatch would have created a
  second item Claude Code never reads. The account is read from the
  existing item's `acct` attribute and the mirror is skipped without one.

The command builder is a pure function so the no-argv / hex / account
contract is tested on every lane; the Darwin-gated no-entry no-op test is
`macos_only` instead of faking `platform.system`. Tests trimmed to the
invariant bar; the merge test pins that `mcpOAuth` siblings survive.

Live E2E on a throwaway Keychain item seeded with 24 mcpOAuth entries:
triple rotated, scopes/subscriptionType and all 24 siblings preserved,
one item, updated under the seeded account.
2026-09-19 11:25:23 +05:30
salch-cred
6d4ca15d38 fix(agent): mirror Claude Code OAuth refresh into the macOS Keychain (#98334)
On macOS the Keychain is Claude Code's authoritative credential store, but
Hermes only ever wrote ~/.claude/.credentials.json. Since the refresh token is
single-use and rotating, every Hermes-initiated refresh left the Keychain
holding an already-invalidated token, which Claude Code then spent into
invalid_grant and discarded ("Login: Expired").

_write_claude_code_credentials now mirrors the committed refresh into the
existing "Claude Code-credentials" entry via security add-generic-password,
merging the rotated token triple over the existing payload so
subscriptionType / rateLimitTier / scopes survive. The payload is fed on stdin
(bare -w), never argv. Fail-soft: a mirror failure is logged, never raised —
the file commit already succeeded and the resolver resolves from it. No-op off
Darwin and when no entry exists (never create one the user has not).

Add a raw payload reader (metadata preserved), a pure merge helper, and the
mirror; extend the conftest keychain guard to neutralize the new writer in any
test that hasn't opted in.
2026-09-19 11:25:23 +05:30
teknium1
67757285f6 feat(sessions): hermes sessions repair-profiles settles crossed-profile durable state
The per-profile store model (#88734), the parent-inheritance fence (#88381),
profile-stamped topic rows (#76423) and profile-prefixed voice keys (#75198)
are all forward-only: they put NEW state under the right profile and refuse to
widen existing damage, but nothing walks the stores and settles what earlier
releases left crossed. #113884 found 246 sessions stranded that way and could
only warn.

`hermes sessions repair-profiles` scans every profile's state.db plus the
gateway's voice-mode and sessions.json files and names six kinds of crossing:

1. `profile_name` disagreeing with the row's own session key -> relabel;
2. rows physically in another profile's store -> move (all message
   generations, usage rows, system prompt) to the owning store, parents before
   children so lineage survives, copy-then-delete so a crash leaves a duplicate
   the next run settles;
3. `parent_session_id` crossing namespaces -> sever (own identity kept);
4. routing rows outside the default store under multiplexing -> move (an
   existing row wins); routing rows for a profile that no longer exists -> drop;
5. Telegram topic bindings and voice-mode entries missing their bot's profile
   -> relabel from the sessions that hold the chat (ambiguous chats reported);
6. sessions.json mirror entries for an unclaimed namespace -> drop (the legacy
   import re-injects them into routing every boot).

Report-only by default. `--apply` refuses while a gateway owns any store, takes
a quick snapshot of every store first, and is idempotent. Two cases are
reported but never guessed: rows keyed to a profile that does not exist, and
`agent:main` rows inside a named profile's store (`--legacy-main rekey|move`
says which of the two histories they are).

Storage side lives in `hermes_state_profile_repair.py` (SessionDB mixin);
orchestration across stores in `hermes_cli/sessions_repair_profiles.py`; the
CLI face in `hermes_cli/sessions_cmd_repair_profiles.py` (pre-DB handler: it
opens every store itself).

Part of #88715 (PR-6). Closes the remediation gap #113884 only warns about.
2026-09-18 22:36:41 -07:00
teknium1
29c98eecc6 refactor(gateway): one call shape for the re-stamp seam
`_profile_name_for_source` already accepts `adapter_profile=None`, so the
branch in `_stamp_routed_profile` was two spellings of the same call.
Collapse it and give the two test doubles that replace the method the
same keyword so they keep matching the production signature.

Salvage follow-up to #106019.
2026-09-18 22:36:24 -07:00
Italo Fernandes
032b91d9b9 fix(gateway): reject inbound events when route matching fails
A matcher exception fell through to the default profile, so a transient
failure while resolving a sender route silently served the message from the
default profile's runtime. Raise the existing rejection instead, which the
ingress gate already drops on.

Only the failure path changes: a source that simply matches no route keeps
falling through to the default/active profile as before.
2026-09-18 22:36:24 -07:00
Italo Fernandes
e864bedf95 fix(gateway): close sender-routing gaps found in review
Voice input reused the bound text channel's cached source and replaced only
`user_id`, so a second speaker inherited the profile resolved for the first.
The ingress gate does not re-resolve a source that already carries a profile.
Re-resolve at the voice call site instead of clearing the profile: clearing
would drop the receiving bot and could re-home voice arriving on a secondary
profile's own bot. `_voice_input_source` reattaches the transport provenance
`from_dict` discards, and `_stamp_routed_profile` takes the receiving bot's
profile as the fallback when no route matches.

Kanban re-subscription could not repair a row created before sender capture:
`user_id` was only written by the INSERT, so a legacy row stayed senderless and
the notifier's conservative fallback left it undeliverable for good. It now
self-heals like `user_id_alt`.

An explicit `user_id: null` or empty string is rejected instead of widening the
route to every sender, and `to_dict` omits the field when unset so a round-trip
cannot reintroduce it. Numeric `0` from an adapter normalizes to "0" rather than
being dropped, without changing the shared coercion used by the other fields.
2026-09-18 22:36:24 -07:00
Italo Fernandes
d6f3285b81 feat(gateway): route inbound messages to profiles by sender user_id
Closes #33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.

Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.

Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.
2026-09-18 22:36:24 -07:00
kshitijk4poor
9dee8863e1 docs(model_switch): say what the fallback actually relies on
The switched-provider comment implied the guard catches the resolver raise;
it is suppress() leaving st.* at the session values that does. The
stay-on-provider comment now states the one semantic widening of moving to
the shared helper: an empty resolver result keeps the session endpoint on
any host, not only off openrouter.ai.
2026-09-19 10:43:50 +05:30
kshitijk4poor
0cb3760383 fix(model_switch): treat a fall-through to OpenRouter's default as "no custom endpoint configured"
Gate finding: resolve_runtime_provider(requested="custom") never returns an
empty base_url — with no model.base_url but an OpenRouter key present it
lands on OpenRouter's default host, so the "keep the current endpoint"
fallback was dead and an Anthropic session adopting `provider: custom`
hopped to openrouter.ai. The guard _creds_for_current_provider already
carries for this (#74143) is now a shared helper used by both branches;
the branch assigns once instead of assign-then-undo; the import sits with
the other from-imports; the new test's write_text carries an encoding.
Second invariant test pins the no-endpoint case.
2026-09-19 10:43:50 +05:30
kshitijk4poor
72749b2482 test: read config with encoding=utf-8 in the custom-provider switch tests
Pre-existing windows-footgun hits in the file this PR touches; the lint
lane wants the whole file clean.
2026-09-19 10:43:50 +05:30
kshitijk4poor
daf882e3c2 fix(model_switch): switching to a bare custom endpoint from another provider resolves the configured one
The bare-custom credential branch kept the session's current base_url and
key whenever the target was `custom`, regardless of where the session was.
That is right for a bare-custom session picking another model on its own
endpoint, but when the TUI's per-turn config sync adopts `provider: custom`
from an OpenRouter (or any other) session, the new model was paired with
OpenRouter's URL and key and every request 400/404'd (Bug 2 in issue 73680).

Arriving from another provider now resolves the configured custom endpoint
through resolve_runtime_provider; with none configured the current endpoint
is kept, so direct aliases still supply their own URL afterwards and the
bare-custom same-endpoint case is unchanged.
2026-09-19 10:43:50 +05:30
kshitijk4poor
cf13448580 refactor(agent): name the tool-call XML namespace prefix once
Gate fold: the optional-prefix regex was spelled three times in the agent
stripper; it is one constant now (_NS_PREFIX). The cli.py docstring pointed
at run_agent._strip_think_blocks, which moved to
agent.agent_runtime_helpers.strip_think_blocks; stray blank lines dropped
from the test file.
2026-09-19 10:41:08 +05:30
DoGMaTiiC
60735cbf81 fix(agent): strip namespace-prefixed text-channel tool-call XML from visible text
muse-spark (opencode-go, Responses wire) occasionally serializes its next native
tool call onto the text channel as <atem:function_calls>…</atem:function_calls>
instead of a function_call item. With real tool work earlier in the turn and
finish_reason=stop, the literal XML is then delivered as the turn's final answer
(observed on a production cron worker: the turn ended on the full XML block; the
task itself had executed).

Teach the tool-call XML patterns an optional namespace prefix (?:[\w.-]+:)? on
the open/close tags and the unterminated-opener tail, in every copy of the
stripper: agent/agent_runtime_helpers.py (_TOOL_CALL_BLOCK_PATTERNS,
_STRAY_TOOL_CALL_CLOSER_PATTERN, _UNTERMINATED_TOOL_CALL_PATTERN) and the
display-side cli.py::_strip_reasoning_tags, which the test suite requires to stay
in sync. Covered positions: closed block, stray closer, unterminated opener
(stream cut mid-serialization).

With the block stripped to nothing, the existing empty-response recovery path
(agent/turn_final_response.py) re-prompts the model instead of shipping XML as
the answer.

Refs #103483 (the native-call XML leaking as text is listed there as not covered yet).

Tests: tests/agent/test_strip_reasoning_tags_cli.py — full block and
cut-mid-serialization tail, each asserted against both strippers. Proven red on
base: 2 failed / 4 passed without the pattern change; 6 passed with it.
2026-09-19 10:41:08 +05:30
kshitijk4poor
956dadf0ad chore: map DoGMaTiiC's contributor email for the #114136 salvage 2026-09-19 10:41:08 +05:30
teknium1
c07708671d fix(gateway): every adapter session key goes through one seam (+ lint)
A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.

Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.

Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.

Phase 2 of #88715.
2026-09-18 22:04:43 -07:00
teknium1
ad651b8250 feat(gateway): RoutingIdentity — one frozen identity per inbound event
A multiplexed gateway answered "which bot received this / who may admit it /
where does it run" in three places (`_transport_owner`,
`_authorization_home_for_source`, `_resolve_profile_home_for_source` +
`_session_key_profile`) that agreed only because they read the same fallback
chain. `gateway/session_identity.py` answers them once: `resolve_identity()`
folds `_admit_primary_source` + `_stamp_routed_profile` + the transport-owner
lookup and pins a frozen `RoutingIdentity` (transport_profile, runtime_profile,
authorization_home, runtime_home, weak transport ref) on the source as a
wire-invisible attribute, like `_transport_adapter_ref`. Under multiplexing a
route to an unserved profile raises `IdentityUnresolved` instead of a
`None`-means-default return; `"default"` is spelled out inside the object.

Additive: the existing helpers become thin readers of the identity when it is
present and keep their fallback chain when it is not, `source.profile` stays
the serialized runtime profile (None on the wire ⇔ default) and every
historical `agent:main` key is byte-identical. `replace_source()` copies a
source without losing its provenance (run_topics used to hand-copy the
transport ref).

Phase 1 of #88715; the gateway rows of #90142 / #93943.
2026-09-18 22:04:26 -07:00
teknium1
03fee43ca3 chore: map contributor email for @Caelier (#100423 salvage) 2026-09-18 20:57:07 -07:00
teknium1
c1ad9daa8f test: two invariant tests for Codex singleton adoption, replacing the salvaged suites
The independent-account test drives the real load_pool() -> _refresh_entry() path
and asserts the second account POSTs its own refresh token and keeps its principal
(red on base: the row silently became account A). The alias test pins both halves
of the same-account rule: never fall back onto an older singleton, always follow a
newer re-auth.
2026-09-18 20:57:07 -07:00
Tranquil-Flow
20512a47a6 fix(agent): refuse stale singleton adoption over rotated manual:device_code entries (#106705)
A manual:device_code Codex pool entry never writes its rotation back to the
auth.json singleton (independent-credential contract, #39236). The singleton
sync adopted differing singleton tokens with no staleness proof, so after a
pool-side rotation the stale singleton was re-adopted over the pool's fresh
chain and the already-consumed refresh token was POSTed again
(refresh_token_reused).

Gate adoption on the singleton's last_refresh not predating the entry's own
rotation; missing stamps on either side keep the historical
adopt-on-difference behavior (#70111).
2026-09-18 20:57:07 -07:00
Marc Caelier
46503f1672 fix(auth): isolate Codex singleton sync by principal 2026-09-18 20:57:07 -07:00
teknium1
76724f3dc2 fix(agent): terminal copy for the new role_alternation failure reason
Reclassifying alternation 400s away from format_error left non-MoA callers
on the generic default text.
2026-09-18 20:56:35 -07:00
teknium1
40778c74e2 fix: MoA aggregator merges adjacent same-role messages only for destinations that reject them
The aggregator request deliberately ends `user(task), user(guidance)` on
iteration 1 of every turn (#113175) so the whole prefix stays byte-stable
for the provider prompt cache. Strict-alternation chat templates
(llama.cpp / vLLM Jinja templates, Mistral, some OpenRouter routes) 400 on
that adjacency ("Conversation roles must alternate ..."), and the turn
failed. Merging proactively for everyone was declined because it brings
back the byte divergence #113175/#113784 removed and only moves the 400.

Reactive, destination-scoped recovery instead:

- error_classifier: new `FailoverReason.role_alternation` for the vendor
  alternation wordings (checked before the request-validation table since
  the body also carries `invalid_request_error`); same abort+fallback hints
  as format_error so non-MoA consumers behave exactly as before.
- moa_alternation (new sibling): `merge_same_role_messages` (reuses the
  loop's `_merge_user_content`), `destination_key` (base_url|provider,
  model), `is_role_alternation_rejection`.
- moa_loop._call_prepared_aggregator: on that 400, retry ONCE with the
  adjacent user turns merged, remember the destination on the facade for
  the session so later iterations pre-merge, never touch destinations that
  accepted the split shape. The trace records the messages actually sent.
- docs: caching section explains the reactive merge.

Live loopback (real call_llm -> SDK -> HTTP stub that 400s on same-role
adjacency): before, iteration 1 fails with BadRequestError after 1 request;
after, 2 requests (split -> 400 -> merged -> 200), next turn pre-merged in
1 request; the accepting-stub control sends byte-identical requests.

Fixes #112358
2026-09-18 20:56:35 -07:00
teknium1
acccd54c67 fix(desktop): restore the unsent composer draft when a session is gone
When a Desktop tab references a session id that exists in no profile DB
(deleted elsewhere, or a stale id after a profile rename / wiped backend),
the resume path already drops the window to a fresh draft without toasting
or hot-looping (62af32efe7, bounded by goneSessionVerdict). But the text
the user had typed into that session stayed stashed under the dead key in
`hermes:composer-drafts:v3` — not lost, yet invisible, because nothing ever
opens that key again. From the user's side the message just vanished.

The gone verdict now announces the dead key (`announceGoneSessionDraft`),
and the composer's swap onto the fresh-draft scope consumes it once
(`adoptGoneSessionDraft`): the stash moves into the new-chat bucket, the
composer is seeded with it, and an inline "Restored your unsent message"
strip with Undo appears above the input. Undo returns the text to the dead
key while it is unchanged (still recoverable by the same path); once edited
it only dismisses. Keyed on the one-shot announcement rather than on the
composer observing an id→fresh transition, because the composer remounts
across the drop (a loading route mounts no composer). A fresh draft that
already has text is never clobbered; an empty stash shows nothing.

Offer, don't hijack (apps/desktop/AGENTS.md): no navigation beyond the drop
that already happened, no focus move, no toast.

Live (headless Electron over CDP, worktree backend): type into a live
session → delete its row + restart the backend → reload at the dead id.
Before: route drops to #/, composer empty, text only in localStorage under
the dead key. After: same drop, composer holds the text, strip + Undo shown;
Undo empties the composer and puts the text back under the key; a second
open of the dead id shows no strip; a dead id with no stash shows nothing.

Fixes #111868
2026-09-18 20:56:12 -07:00
teknium1
0c3ebefb5e fix(desktop): /reasoning hide|show gates Thinking blocks immediately
`display.show_reasoning` is a client-side display decision since 40f2368875,
and the Desktop transcript honors it through `$showReasoning` (#115335). The
composer path did not: `/reasoning` was marked `desktop="advanced"` in the
Python slash registry, so the Desktop refused it as "not available", while a
gateway-routed `/reasoning hide` would only have written config.yaml and left
the atom waiting for the next config refresh.

Route `/reasoning` to the gateway's `config.set key=reasoning` — the Ink TUI's
path (ui-tui/src/app/slash/commands/session.ts) — and mirror the answer:
`hide`/`show` flip `$showReasoning` on the spot, effort levels leave the gate
alone and get session scope (a throwaway slash-worker CLI could never pin the
live session's effort). No arg reports the current effort and display, like
Ink. Drops the registry's `advanced` marker and regenerates the desktop dump.

Live (real `hermes serve`, scratch HERMES_HOME, real WebSocket): config.set
key=reasoning value=hide -> {value: hide}, config.yaml show_reasoning=False;
show -> True; high -> effort only, display untouched.

Fixes #111761
2026-09-18 20:55:49 -07:00
teknium1
0cd13aaf74 docs: relative link to the borrowed-logins section (route-style links fail the docs check) 2026-09-18 20:55:24 -07:00
teknium1
6c7f693473 feat(auth): opt out of borrowing Codex CLI / Claude Code logins (auth.adopt_external_logins)
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.

- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
  current users). When false:
  - `read_claude_code_credentials()` — the only reader of the borrowed Claude
    Code login — returns None, so the resolver fallback, the expired-token
    refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
    file; the pool prunes a `claude_code` row an earlier adopting process
    persisted.
  - `_recover_codex_tokens_from_cli` returns None for both automatic recovery
    paths (rejected refresh, half-empty singleton); the real AuthError is
    surfaced instead. The interactive import offer in `hermes auth add
    openai-codex` still asks first and is unaffected.
  - One INFO line per process the first time adoption would have happened;
    `hermes auth list` / `hermes auth status anthropic|openai-codex` print the
    same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.

Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
2026-09-18 20:55:24 -07:00
teknium1
d99226015d fix(desktop): count the head entries the group-chat mirror drops
The ui_meta mirror of a room is head-trimmed to 16 messages and then to
the 48 KB envelope, silently. It is the only on-disk copy of the room, so
an agent that consults it to check what the user said reads a bounded
window with no sign that it is bounded (#114341, secondary). Stamp
`omitted: N` on the projected room whenever entries fall off the head,
recomputed after every byte-budget shift so the envelope bound stays
honest, and carry the larger writer's count through the publish merge as
a lower bound. Local room state never receives the field: the merge-back
into $groupChats builds from explicit fields.
2026-09-18 20:37:28 -07:00
686f6c61
d32dacf855 fix(desktop): mark truncated group-chat sync payloads
The ui_meta projection sliced messages at 1200 chars with no
signal that the body continued. Stamp truncated and keep the
marker inside the existing per-message budget.
2026-09-18 20:37:28 -07:00
teknium1
0a3fce8846 test(desktop): pin the omitted-head marker on the real group turn path
The salvaged tests exercised the formatGroupDeltaLines helper in
isolation, so a call-site regression (a bare slice reappearing in
prepareGroupRoundMember) would have stayed green. Drive the room through
runGroupChatRounds instead and assert on the prompt the gateway actually
received: over-long delta -> marker naming the dropped count and the
dropped head absent; delta that fits -> no marker. The helper-level pair
is dropped so the file keeps two invariant tests for this fix.
2026-09-18 20:37:28 -07:00
liuhao1024
c91d955ff7 fix(desktop): mark the truncated head of a bot group-chat turn delta 2026-09-18 20:37:28 -07:00
teknium1
b71359066c fix: /stop halts background subagents and returns their partial results as interrupted completions
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:

- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
  already did this in `_handle_reset_command`, the second request is
  idempotent) and the idle `_handle_stop_command` tail, which replied
  "No active task to stop." while a background child was running; it now
  stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
  spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.

The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.

An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.

Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.

Part of #114456
2026-09-18 20:34:42 -07:00
teknium1
2b94b0d40f fix(kanban): dispatcher blocks a card on the first terminal provider error
Before, a worker killed by a revoked credential or a missing model was
booked as an ordinary crash and re-spawned into the identical failure
until kanban.failure_limit / max_retries was spent — burning worker slots
and the retry budget on something a retry cannot fix (#114587).

Now KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78) is its own exit kind,
`terminal_provider`: `_classify_dead_worker_exit` books the run crashed
with the provider's words appended, and `_account_crashes` force-trips
the breaker on that first death (sticky, so recompute_ready does not
resume it before the operator fixes the provider). Same booking in the
implementation and the review lane — the review worker dies through the
same sweep. Transient failures (429 / 5xx / timeout) keep the existing
rate_limited requeue and the consecutive_failures budget.

`hermes kanban show` / the dashboard diagnostics now fire for that trip
below the repeated-failure threshold ("Provider rejected this profile's
credential or model — blocked after one attempt") with the fix path.

No new columns, no separate review-lane counter, no new config: option B
of #114587. The terminal-vs-transient split was proposed in #114589 by
@TFOjojo; its regex-on-error-text classifier is replaced by the worker's
own FailoverReason verdict.

Part of #114587

Co-authored-by: TFOjojo <279183633+TFOjojo@users.noreply.github.com>
2026-09-18 20:34:16 -07:00
teknium1
c20432fffa fix(kanban): worker exits EX_CONFIG on a terminal provider error
A dispatcher-spawned worker whose turn failed on a provider error that no
retry can heal (credential rejected: auth / auth_permanent, model_not_found,
ssl_cert_verification) now exits KANBAN_TERMINAL_PROVIDER_EXIT_CODE (78,
BSD EX_CONFIG) instead of a plain 1, so the dispatcher can tell "the
provider will reject every further spawn" from "this attempt crashed".
The classification is the agent's own FailoverReason verdict already on
the turn result (failure_reason) — no error-text regex. billing stays in
the transient set: credit comes back, a revoked key does not.

Shared by both one-shot paths (-q and -Q) and gated on HERMES_KANBAN_TASK,
so a person's `hermes chat -q` keeps exiting 1 on the same error.

Part of #114587
2026-09-18 20:34:16 -07:00
teknium1
c6e3a77577 fix(cron): the ticker supervisor respawns only a ticker that crashed, not one that returned
SupervisedTickerThread (#111010) treated any ended thread as dead. An external provider's
start() (Chronos) arms remote one-shots and returns by design, so every housekeeping tick
logged "Cron ticker thread died without a stop request; restarting" at ERROR and re-ran
start() - a fresh recover_interrupted + NAS list/arm reconcile once a minute on every hosted
instance. Track whether the target escaped with an exception and respawn only then; the
built-in ticker returns normally only on stop_event.
2026-09-18 20:07:21 -07:00
teknium1
e56ec8c9f7 fix(chronos): a 403 invalid_client from NAS hands cron fires to the built-in ticker (#97494)
NAS maps the agent-cron bearer to a provisioned instance via an `agent:*` client or the
`hermes-cli-vps` bootstrap session (hermes-portal `server/agent-cron/instance-auth.ts`). A
container whose auth.json holds a plain `hermes-cli` user login is refused with 403
invalid_client on every arm, re-arm and list, for the life of that credential - and the
re-login users try first (`hermes auth logout nous` + device code) replaces the bootstrap
session, making it permanent. Chronos previously logged one bare warning per job and left
the jobs with no trigger at all: they only ran through the misfire sweep, minutes late.

NasCronClientError now carries the HTTP status and the OAuth `error` code. On an identity
rejection the provider logs ONE warning that names the real remedy (restore the hosted
credential from the Nous Portal; re-login cannot fix it), stops calling NAS, and starts the
built-in ticker with the gateway's own adapters/loop so scheduled jobs keep firing on time.
Transient 5xx/transport failures keep retrying on the next reconcile.

Supersedes #97566 (@wesleysimplicio), whose runtime-credential swap resolves to the same
bearer on main (`agent_key` is the access token) and so could not change the 403.
2026-09-18 20:07:21 -07:00
teknium1
6c16233b0a fix(update): refresh memory-provider deps in the same order as the pull path
Both repair paths now run the memory-provider refresh between the tool-dep
restore and the plugin-dep reapply, exactly like _sync_python_dependencies_after_pull
and the ZIP path, so the three routes stay interchangeable.
2026-09-18 19:57:08 -07:00
Yagna Vudathu
1f6e737ecb fix(cli): heal memory-provider bridge packages on every venv repair path
The venv-repair and runtime-repair paths reinstalled core, lazy and tool
deps but never refreshed memory-provider bridge packages, unlike the
git-pull and ZIP paths. Both repair paths now heal them last.

Fixes #113741.
2026-09-18 19:57:08 -07:00
teknium1
255b4fd9da fix(image-gen): record token usage as soon as the billed HTTP 200 lands
OpenRouter (_generate_via_chat, _generate_via_image_api) and OpenAI gpt-image
called record_token_usage only after image extraction and save succeeded. A
token-billed 200 that returned text but no image (the "returned no image"
fallback case), an empty `data` list, or a failed save therefore consumed
billed tokens that never reached session_model_usage; on the fallback chain
only the fallback model's tokens were recorded.

Move the call to immediately after a successful post_json / SDK call, before
extraction/save, in all three paths. One call per HTTP success, so no double
count; failure responses (non-2xx, timeouts) still record nothing because
there is no body to read usage from.

Tests: the parametrized invariant in test_openrouter_compat_provider gains the
no-image-with-usage case for both surfaces (result is `empty_response`, one row
still recorded); the OpenAI test is parametrized on has_image the same way.
2026-09-18 19:56:51 -07:00
teknium1
6babdc96b8 fix(image-gen): every token-billed image backend records session usage (chat path, Image API, OpenAI gpt-image)
Widen the salvaged Image API hunk (#114340) to the whole class:

- plugins/image_gen/_common.py: record_token_usage() — one helper feeding the
  aux accounting chokepoint (agent.aux_accounting.record_aux_usage) with task
  "image_generation", the billing provider and the priced model id. Dict and
  SDK usage objects alike; a body without tokens is a no-op, as is a call
  outside a turn.
- plugins/image_gen/openrouter: the /chat/completions path (the DEFAULT model
  chain — openai/gpt-5.4-image-2, google/gemini-3-pro-image — is chat-only and
  token-billed) now records too; the Image API path uses the shared helper
  (the contributor's local _record_image_api_usage is folded into it) and both
  pass base_url so pricing resolves the route. Task renamed image_gen ->
  image_generation to match the other aux task names.
- plugins/image_gen/openai: gpt-image bills per text/image token; record the
  Images API usage block under the API model (gpt-image-2), not the Hermes
  quality-tier label.
- FAL, xAI, Krea, DeepInfra, Meta, openai-codex return no token usage and stay
  unrecorded (nothing to bill per token).
- tests: the contributor's two tests folded into one invariant parametrized
  over chat / Image API / no-usage control; one OpenAI invariant.
- docs: image-generation.md "How It Works Internally" gains the accounting step.

Fixes #114324
2026-09-18 19:56:51 -07:00
Yagna Vudathu
98ad0d8148 fix(image-gen): record token-billed OpenRouter calls to session usage
OpenRouter token-billed image calls parsed usage into extra only, so
session_model_usage never saw them. Record prompt/completion tokens via the
ambient aux accounting on success; flat-fee responses without token usage
write nothing.

Fixes #114324.
2026-09-18 19:56:51 -07:00
teknium1
0ddba07ad7 fix(ci): report an interpreter crash as CRASHED, not "no tests ran"
When a per-file pytest subprocess dies by signal (the sqlite cross-thread
close in #113186 was a SIGSEGV after every test had passed), faulthandler
prints "Fatal Python error: Segmentation fault" and no summary line, so
every count parses to 0. The runner filed that under "1 file where no
tests ran (collection/import error, ...)" beneath a summary that read
"0 failed" and exited 1 — two wrong diagnoses for one real bug, and it
was misread as a runner problem twice on main.

The runner now detects a signal death or a "Fatal Python error:" banner,
prefixes the captured output with the diagnosis (same convention as the
timeout path), marks the progress line CRASHED, counts "N files CRASHED"
on the summary line, lists the file in its own failure bucket, and no
longer trips the "NO TESTS RAN" guard for a crash that ran tests. The
flake retry already covers crashes (any non-zero rc), so nothing changes
there.
2026-09-18 19:41:51 -07:00
teknium
1af1db281e plugin-catalog: hermes-talk disclosure line for the Codex realtime lane + contributor map 2026-09-18 19:38:09 -07:00
TheSmokeDev
a3baf5135d plugin-catalog: bump hermes-talk pin to v0.21.0, add voice category 2026-09-18 19:38:09 -07:00
SmokeDev
e8e712e7dc feat(plugins): pin hermes-talk catalog entry to v0.19.2
Moves sha and docs_url to a4843d82c0b2c8443830b79f5e0a19b8a78a8303 (v0.19.2, peeled
commit). Capabilities block unchanged and re-verified at that commit.

Signed-off-by: SmokeDev <degensmoke@gmail.com>
2026-09-18 19:38:09 -07:00
SmokeDev
dbb9dcc200 catalog: pin hermes-talk to v0.19.1 (read-only Codex lane)
Moves the pin from 8aeac6a6 to 4764f687, the peeled commit of tag v0.19.1.
The only source change on the auth surface is #146 (read-only Codex lane,
authored by @teknium1), merged as-is. plugin.yaml changed only its version
field; hook registrations, tools, middleware and required env are unchanged
against the declared capabilities block.

Signed-off-by: SmokeDev <degensmoke@gmail.com>
2026-09-18 19:38:09 -07:00