Commit Graph

5548 Commits

Author SHA1 Message Date
teknium1
2632229bcf fix(auth): named profiles read the root auth.json again (revert #111724)
Reverts 93889b770d ("named profiles no longer inherit the root
profile's auth.json"). After the Desktop update every bot profile that had
relied on the root OpenAI Codex login failed with "No Codex credentials
stored. Run `hermes -p <bot> auth add openai-codex --type oauth`", and users
had to re-run the device-code flow once per bot (5-6 times in the field
report). Sharing one grant across profiles is the intended design: OAuth
refresh tokens are single-use, so ONE grant lives at the root, profiles
resolve it read-only, and a refresh under a profile writes the rotated chain
back to root (Codex / xAI write-through, borrowed-row pool bookkeeping,
forked-grant heal) — never a per-profile copy.

Restored: `_global_auth_file_path` / `_load_global_auth_store` fallback in
`_load_provider_state*` / `read_credential_pool` / `_provider_state_transaction`,
Codex + xAI root write-through, `credential_pool` borrowed-root persistence,
`heal_forked_single_use_oauth_grants`, `share_auth` on profile creation
(Desktop create dialog checkbox), and the docs. `profile_credential_audit.py`
(the `hermes update` "profiles without a provider" notice) is removed with it.

Kept from after #111724: `_save_codex_tokens(set_active=...)` for image gen,
the plugin-auth `status` dispatch and the external-login notice in
`hermes auth list`, and the registry-derived env-var hint in agent_init.
2026-09-21 09:55:40 -07:00
kshitijk4poor
fe4294802e fix: apply the per-model image strip to the max-iterations summary request
agent/chat_completion_helpers.py::_iteration_summary_api_messages hand-builds
api_messages from canonical history (shallow row copies) and calls
agent._build_api_kwargs directly, bypassing turn_api_request.build_api_request
— the sibling send path where strip_images_for_rejecting_model runs. On main
this went unnoticed because the image-rejection recovery stripped history
itself, so the summary request was image-free by accident. This stack keeps
history intact, so a model recorded in agent._image_rejecting_models would
receive images in the summary request → 4xx → 'max_iterations_no_summary'.

Call strip_images_for_rejecting_model right after the vision eviction, mirroring
build_api_request. Hazard-checked: the strip rebinds the row's content to a new
list (never mutates the nested list shared with history), so the shallow copies
keep history untouched — the new test asserts both the text-only output and the
unchanged history.
2026-09-21 21:43:51 +05:30
kshitijk4poor
6540224f69 docs: fix three statements the per-model image strip left stale
The image-rejection recovery no longer switches the session to
text-only or strips history; it records the (provider, model) and
build_api_request strips images from that model's requests only. Three
places still described the old behaviour:

- recover_before_classification docstring and the adjacent comment said
  "switch session to text-only" / "mark session vision-unsupported".
- TestStripImagesDropsStaleApiContent's rationale claimed the strip runs
  on persistent history and that leaving api_content would replay
  rejected images every turn because the recovery gates on
  _image_rejecting_models — false on both counts: current callers pass
  per-call clones. Reworded as the generic helper contract (a rewritten
  persisted row must drop its sidecar) with the no-op note.
- _strip_images_from_messages docstring, same fix.

Comments and docstrings only; no code change.
2026-09-21 21:43:51 +05:30
kshitijk4poor
565b2ce206 refactor: dedupe the corrupt-image recovery and split the phrase lists
Two branches in turn_recovery.py carried the same isinstance +
_strip_images_from_messages guard and the same "Provider rejected a
corrupted image" notice: the new pre-classification corrupt branch and
the FailoverReason.image_corrupt branch. Both now call one module
helper, _strip_request_images_and_retry(agent, api_messages) -> bool,
so the strip-this-attempt-only policy lives in one place.

_IMAGE_REJECTION_PHRASES was rebound as unsupported + corrupt, which
made _looks_like_image_content_rejection silently cover corrupt payloads
and forced the recovery to test the corrupt list a second time to undo
that. The name stays (external references) but is now the
unsupported-only tuple; _IMAGE_CORRUPT_PHRASES is disjoint, and the
recovery asks the two questions explicitly:
`_corrupt or (model not yet rejected and unsupported)`.

Tests: the phrase-isolation matcher asks the same disjoint question the
recovery does; the image_corrupt source-contract check looks for the
helper call instead of the inlined strip.
2026-09-21 21:43:51 +05:30
kshitijk4poor
b5ec7edbc4 refactor: drop the redundant strip from the image-rejection recovery
recover_before_classification's capability branch stripped api_messages
itself right after recording the (provider, model) in
_image_rejecting_models. That strip is dead work: the verdict is
`continue`, which re-enters build_api_request with the SAME api_messages
object, and strip_images_for_rejecting_model runs there — before
_build_api_kwargs on every attempt — and strips them because the key is
now in the set. Keeping both meant two places had to agree on the send
policy. The corrupt-image branch keeps its local strip on purpose: it
deliberately does not record the model, so the send path would not.

test_canonical_history_keeps_its_images now asserts the text-only wire
copy through strip_images_for_rejecting_model (the real send path) and
checks build_api_request still calls it ahead of _build_api_kwargs;
commenting that call out turns the test red.
2026-09-21 21:43:51 +05:30
kshitijk4poor
7ef78c1278 refactor: read agent._image_rejecting_models directly
init_agent seeds `_image_rejecting_models = set()` at the same site as
`_force_ascii_payload`, which the adjacent sanitize_outbound_kwargs already
reads as a plain attribute. The getattr/isinstance guards and the lazy
re-create in recover_before_classification implied the attribute could be
missing or mistyped on a real agent; it cannot, and defensive fallbacks for
impossible states hide wiring bugs instead of surfacing them. Every test
fixture that reaches these paths seeds the attribute already.
2026-09-21 21:43:51 +05:30
kshitijk4poor
26f9cb03c4 refactor: reuse _provider_model_key for the image-rejection model key
message_sanitization.image_model_key duplicated vision_message_prep's
_provider_model_key — the (provider, model) key the same mixin already uses
for its per-model vision bookkeeping. Two keying functions for the same
concept drift: one normalised the provider (.strip().lower()), the other did
not, so a provider spelled "OpenAI" in one place and "openai" in another
would have been tracked as two models. Key `_image_rejecting_models` on the
existing helper and delete the duplicate (and its __all__ entry).

No import cycle: vision_message_prep imports only lazy_forward,
tool_dispatch_helpers and utils, none of which import message_sanitization
or turn_recovery.
2026-09-21 21:43:51 +05:30
kshitijk4poor
b5dacdb125 fix: a corrupt-image rejection must not blind the model for the session
recover_before_classification matches _IMAGE_REJECTION_PHRASES, which
mixes two kinds of body: capability rejections ("does not support
images", "only text content type is supported", ...) and bad-payload
rejections. After the per-model tracking landed, BOTH added the
(provider, model) to agent._image_rejecting_models, so one truncated
screenshot rejected with "failed to decode image" made every later
request to that model text-only for the rest of the session, even
though the model can see fine.

Split the bad-payload phrases into _IMAGE_CORRUPT_PHRASES:
  - "image data you provided does not represent a valid image"
    (ChatGPT-account Codex backend)
  - "failed to decode image" (Kimi / Moonshot and other
    OpenAI-compatible providers)
_IMAGE_REJECTION_PHRASES stays the union so the turn still recovers on
them. For a corrupt match the recovery now strips the current attempt
only and retries iff something was stripped (mirroring the existing
image_corrupt branch), and leaves the model unmarked; only a capability
rejection records the model.

One test: a "failed to decode image" body strips the wire copy, keeps
history, and leaves _image_rejecting_models empty so the next request
to the same model carries its images.
2026-09-21 21:43:51 +05:30
kshitijk4poor
a955e90422 refactor: hoist strip_images_for_rejecting_model import to module level
turn_api_request already imports from agent.message_sanitization at the
top of the module, so the function-local import added by the salvage is
an inconsistency, not a cycle guard. Hoist it onto the existing import
line; `python -c 'import agent.turn_api_request'` confirms no cycle.
2026-09-21 21:43:51 +05:30
kshitijk4poor
28a75285c4 refactor: drop the dead _vision_supported flag
Image rejections are now tracked per (provider, model) in
agent._image_rejecting_models, and recover_before_classification gates
on that set. Nothing reads agent._vision_supported any more, so the
write in turn_recovery and the per-turn reset entry in turn_context are
dead state. Remove both so the next reader doesn't assume a turn-global
vision gate still exists; update the two tests that asserted or seeded
the attribute.
2026-09-21 21:43:51 +05:30
rodricksz4h5
7299015092 fix(agent): track image rejections per model across a fallback chain
The first head stored a single rejecting (provider, model) and kept the
turn-global `_vision_supported` as the recovery guard. In a fallback
chain that fails: model A rejects images and retries text-only, a later
error activates model B, the restart rebuilds api_messages from history
so B receives the images, and when B rejects them too the branch is
skipped because `_vision_supported` is already False — the request
falls through to generic error handling. Recording B also overwrote A,
so A was no longer treated as text-only on later turns.

`_image_rejecting_models` is now a set of every rejecting model, and it
is also the guard: each model's first rejection runs the recovery and a
repeat rejection from the same model still falls through, so the retry
cannot loop. image_model_key() names the key in one place.

Adds a test for the two-model sequence (fails on the previous head) and
one pinning that a repeat rejection from the same model does not retry.

Thanks to @ehz0ah for the review.

(cherry picked from commit 225fd76ccf90cd3ec509a9327f725ac02d996f24)
2026-09-21 21:43:51 +05:30
rodricksz4h5
1ef306f68d fix(agent): an image rejection strips the request, never the session history
When a provider 4xx's on image content, recover_before_classification
ran _strip_images_from_messages on the canonical `messages` list and
reset _db_flush_scan_prefix. Since #117569 that function pops
_db_persisted on every rewritten dict, so the next flush rewrote those
rows: every image in the session — and every image-only message, which
the stripper deletes outright — was removed from state.db for good.

The rejection describes what the CURRENT model accepts, not what the
conversation holds. An automatic fallback to a text-only provider, or a
single /model switch, was enough to erase images the user had sent to a
vision model, and switching back found them gone. It is the same failure
as the ASCII strip in #117802, on the image path; neither open fix for
that issue touches this branch.

Keep the repair on the send path, where the per-call copy already lives
(_clone_message_for_send exists so send-path rewrites never reach the
persisted transcript, #80498):

- record the rejecting (provider, model) on the agent, a session-scoped
  flag initialised beside _force_ascii_payload;
- strip the in-flight api_messages copy for the immediate retry;
- build_api_request calls strip_images_for_rejecting_model() on each
  attempt's api_messages BEFORE provider conversion. The stripper knows
  Hermes's own part types; a converted payload would slip past it
  (Bedrock Converse image blocks carry no `type`). Keyed on the model,
  so one that accepts images gets them again.

_strip_images_from_messages itself is unchanged, so its role-alternation
and sidecar guarantees still hold on the wire copy. The notice no longer
claims "text-only mode for this session" (_vision_supported resets every
turn) or that images were stripped from history.

(cherry picked from commit cbdb184c27f915ab138b2087f878aed7fcc7c76b)
2026-09-21 21:43:51 +05:30
kshitijk4poor
68f7b90eeb fix: return a raced-successful compression from _await_worker_within_budget
The `future.done()` guard added for dead workers (#117261 / #63892)
unconditionally logged `future.exception()` and returned `(False, None)`.
If the worker finished SUCCESSFULLY in the window between
`result(timeout=)` expiring and the `done()` check, that discarded a
completed compression and sent the caller down the stall/fallback path —
`future.cancel()` becomes a no-op on a settled future and the fallback
route costs a second LLM call. Siblings `_await_in_flight_commit` and
tool_executor's `_poll_sequential_future` already re-read the result on
a settled future; do the same here: `exception() is None` → return
`(True, result)`, otherwise take the stall path as before.

The compression-seam test grows a settled-successful case through the
same helper (a Future subclass whose first timed `result()` still raises
TimeoutError, modelling the race); it fails on the previous guard and
passes with this one.
2026-09-21 21:34:44 +05:30
kshitijk4poor
a8a81f4cb9 refactor: log the dead compression worker's exception once; trim guard comments
_await_worker_within_budget now emits one INFO line with future.exception()
when it takes the stall path for a worker that already died, so the log shows
WHY the fallback chain was entered instead of silently returning (False, None)
— previously the only trace was the absence of "still streaming" lines.

The three 5-7-line comment blocks added by the #117261 pick restated the same
alias fact each time; each is now 2 lines citing #63892 and the 3.11 alias once.
No behaviour change beyond the log line.
2026-09-21 21:34:44 +05:30
kshitijk4poor
7740a4ac20 fix: report a worker that died with TimeoutError as exited in _join_cancelled_worker
_join_cancelled_worker returned False from its `except
concurrent.futures.TimeoutError:` arm. On 3.11+ that class IS the builtin
TimeoutError, so the arm also fires when the cancelled worker itself DIED
raising a timeout-class error (the aux client raises bare TimeoutError on a
stalled summary stream). The caller, _release_cancelled_worker, treats False
as "still running": it logs 'did not exit within grace' and skips
fence.allow_cancelled_lock_release(), so the session compression lease of a
provably-dead worker was retained/orphaned.

Return future.done() instead: a settled future never becomes unsettled, so a
done future means the thread exited and the lease can be released. A live
worker that merely outlasted the grace still yields False.

Same guard family as the sibling loops fixed in the preceding pick (#117261).
2026-09-21 21:34:44 +05:30
Trevor Nash-Keller
8305113328 fix(compression): don't mistake a worker's TimeoutError for a poll timeout
On Python 3.11+ concurrent.futures.TimeoutError IS the builtin TimeoutError
(asyncio.TimeoutError and socket.timeout alias it too). Poll loops shaped like

    try:
        return future.result(timeout=slice)
    except concurrent.futures.TimeoutError:
        ...keep waiting...

therefore cannot distinguish "the wait slice expired" (worker alive) from "the
worker raised TimeoutError" (worker dead). auxiliary_client raises a bare
TimeoutError when a summary stream stalls, so this is reachable in production.

When it happened the host re-waited on an already-settled future. result() then
returned instantly every iteration, spinning at ~2k iterations/sec and logging
"Context compression still streaming" about a dead worker, until the entire idle
budget elapsed. One session burned 535s and wrote ~90k duplicate log lines
(15.6MiB) before failing with context_compression_timeout, and every later turn
re-entered the same path.

Guard each loop with future.done(): a settled future never becomes unsettled.
- _await_worker_within_budget: take the stall path at once, so the configured
  fallback chain is actually reached instead of after a 120s false stall.
- _await_in_flight_commit: re-raise the worker's exception. This loop had no
  ceiling, so a dead worker spun forever.
- tool_executor._poll_sequential_future: same, and with deadline=None it also
  span indefinitely.

Non-timeout worker exceptions still propagate unchanged, and a live worker still
polls exactly as before.

Regression tests pin all three. Against unpatched code the two wait tests fail
and the commit-wait test hangs until the 300s harness SIGKILL, reproducing the
infinite loop directly.

(cherry picked from commit 274457304ce0393407574fe4e43c4a450f20bac3)
2026-09-21 21:34:44 +05:30
kshitijk4poor
c27219adc8 refactor: detach aliased tools with _clone_message_for_send, not deepcopy
sanitize_outbound_kwargs detached the agent.tools alias with copy.deepcopy
before the ASCII strip. The repo's send-path detach idiom is the structural
clone _clone_message_for_send (dicts/lists recursively, immutable leaves
shared), which is what every other outbound copy uses and is cheaper on
JSON-shaped, acyclic payloads. It is sufficient here because
_sanitize_structure only rebinds str leaves inside dict/list containers, so
the clone fully isolates the canonical tool schemas. Imported lazily inside
the function (conversation_loop imports this module) exactly as
turn_finalizer does. Drops the now-unused `import copy`.
2026-09-21 21:33:44 +05:30
kshitijk4poor
a0dd43cc8c refactor: return a single bool from _repair_transport_credentials
Both callers immediately reduced the (headers_sanitized, credential_sanitized)
tuple with `or`; nobody distinguished the two. Returning one
`transport_repaired` bool removes the unpacking at both call sites and the
redundant flag bookkeeping inside the helper without changing behaviour.
2026-09-21 21:33:44 +05:30
kshitijk4poor
307428377b refactor: drop dead api_kwargs strip from the ASCII-codec recovery branch
The ASCII branch of _recover_unicode_encode_error popped/stripped/restored
`tools` on the failed attempt's api_kwargs and tracked `_tools_sanitized`,
but that dict is discarded: build_api_request rebuilds api_kwargs from
agent.tools on every retry iteration and sanitize_outbound_kwargs already
strips the whole payload once `_force_ascii_payload` is set. Only
api_messages (reused across retries) and the local active_system_prompt
still need stripping here.

Also removes the `_force_ascii_payload = False` assignment in the UTF-8
branch (the flag is initialised False in agent_init and only ever set True
in this function), the now-unused `_sanitize_tools_non_ascii` import and
the dead `import copy`. The test drops its assertion on the recovery
dict's tools identity; the chokepoint assertion (agent.tools byte-stable
after sanitize_outbound_kwargs) is kept.
2026-09-21 21:33:44 +05:30
kshitijk4poor
d47af1421b refactor: share the header/api-key repair between both ASCII-codec branches
_recover_unicode_encode_error carried two copies of the same block that
strips non-ASCII from _client_kwargs["default_headers"] and the API key
(agent.api_key, _client_kwargs["api_key"], client.api_key). The UTF-8
runtime branch was a silent copy; the ASCII runtime branch also printed
the "bad copy-paste?" hint. Two copies drift, and a user whose key was
repaired under a UTF-8 runtime got no hint about why auth might now fail.

Extract one module-level _repair_transport_credentials(agent) returning
(headers_sanitized, credential_sanitized) and call it from both branches.
The hint is emitted whenever the key changed, in both runtimes; the retry
decisions in each branch are unchanged. Net -7 LOC.
2026-09-21 21:33:44 +05:30
kshitijk4poor
336d9662cc fix: detach aliased agent.tools at the outbound sanitization chokepoint
api_kwargs["tools"] is rebuilt from agent.tools on every attempt
(_build_api_kwargs_for_mode: tools_for_api = agent.tools; transports set
api_kwargs["tools"] = tools without copying), so the list usually aliases
the canonical tool schemas. The ASCII retry path then runs
sanitize_outbound_kwargs with _force_ascii_payload set, and its in-place
strip rewrote agent.tools for the rest of the session.

Move the guard to the chokepoint: when the flag is set and tools IS
agent.tools, deepcopy before stripping. The deepcopy the contributor pick
added inside _recover_unicode_encode_error only protected the failed
request's kwargs, which are discarded before the retry; drop it and have
recovery skip an aliased tools list entirely (request-local lists are
still stripped for the diagnostic message).

The kept test now exercises the chokepoint directly: with the flag set on
kwargs whose tools aliases agent.tools, agent.tools must be byte-stable
afterwards. Verified red against the previous chokepoint.
2026-09-21 21:33:44 +05:30
kshitijk4poor
05287a5369 fix: surface ascii-codec error when UTF-8 runtime repairs nothing
Under a UTF-8 runtime, _recover_unicode_encode_error only repairs the
api_key and _client_kwargs['default_headers']. When neither changed, the
retry is byte-identical to the failed request and cannot succeed; it just
consumed both _unicode_sanitization_passes before the error finally
surfaced. Return False in that case so recover_before_classification
falls through to the normal classification/error path immediately.

The existing api_key + default_headers sanitization is unchanged; the
"retrying unchanged request content" message is removed because that
branch no longer retries. The PR's UTF-8 invariant test keeps asserting
the request copy is not stripped, and now asserts the False return.
2026-09-21 21:33:44 +05:30
joaomarcos
17dfeb1c55 fix: preserve conversation history during unicode recovery
(cherry picked from commit d2cec3b33dc728d702b037b0b9c678672322ce3c)
2026-09-21 21:33:44 +05:30
teknium1
af66d5db15 fix(agent): remote backend probe no longer puts user, $HOME and cwd into the system prompt
The non-local terminal-backend probe (`_BACKEND_PROBE_CMD` /
`_format_backend_probe`) ran `whoami`, `$HOME` and `pwd` inside the sandbox
and rendered `User:` / `Home:` / `Working directory:` lines into the system
prompt on every turn. Nothing downstream consumes those values — the only
reader is the model, which can `whoami && pwd` when a task needs them — so
they were user-identifying metadata sent to the provider for no behavioural
gain. The probe now asks for and renders only `OS: <uname -s> <uname -r>`,
and the prompt block tells the model how to fetch the rest on demand.

Fixes #117262
2026-09-21 01:04:37 -07:00
teknium1
b7803a1763 fix(compression): cap the protected tail at 20% of the context window
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole
rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the
"protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every
compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B
timed out before compaction ever changed anything, and protect_last_n read as an uncompressed
tail rather than a minimum.

TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk /
pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool
groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of
32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%).

Probe (12 tool-heavy turns, 49 rows, 12.8K tokens):
  ctx    8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows;  window [4,10) -> [4,42)
  ctx   16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34)
  ctx   32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22)
  ctx  128K: identical before/after
2026-09-21 01:04:16 -07:00
Siddharth Balyan
afc3b7c6f3 feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation

Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash
merges (e0ef0eb9c3, d105376b21, ee2f5629b8, 1ab32b212b): the branch's history
no longer shared a base with main, so this is the PR's exact delta against
+3762/-1470, identical to the branch tip 05d7f2d3d5.

The fifteen commits it carried, in order:

--- feat(connectors): the wire follows the connector contract (six-state status, required toolkit metadata, connectionId on mint)

Hermes types exactly what the contract page writes: /tmp/magic/CONTRACT-TOOLKIT-METADATA.resolved.md
(carriers: portal PR2 `sid/connection-api` for the enum and account rows, the portal contract branch
on top of it for the list metadata). No optional-for-compat fields, no fallback branch, no seven-state
word left in the tree. A gateway that does not speak this contract fails validation loudly.

Wire (tools/connectors/gateway/wire.py): `ConnectionStatus` is the six values, `pending` covering the
vendor's INITIALIZING and INITIATED; `ConnectorListItem` requires title, description, https iconUrl and
authKind, and carries activeConnectionId when the session binds an account; `ConnectorListResponse`
types a page whole with its total; `ConnectorAccount` and `ConnectorAccountsResponse` type the account
routes; a mint result names its connectionId on `initiated` and never on `failed`; CONNECTION_REQUIRED
carries connectUrl and connectionId together or not at all; `account` exists on the execute call and is
never sent while multi-account is off.

Client: `list_connectors(search, connected)` refuses a search under three characters before any
request and parses every page whole; `list_accounts` and `account_status` read the account routes,
404 `connection_not_found` is None and 429 raises `RateLimited(retry_after)`.

Watcher: the target keeps the account id the mint named and the toolkit's title and icon from one
list read before the card is emitted; `_status_for` is the one seam between the watcher and the
gateway's status source (the list walk today; the per-account route replaces that body only).
`pending` and a missing status move nothing; revoked and inactive are failed.

Desktop: `ConnectorRow` is the strict list item; `ConnectorRowSeed` is what a tool call's args can
say; `connection.request`/`update` targets carry `connection_id`, `title`, `icon_url`; the mark
ladder is glyph, vendor icon as a plain image, favicon, monogram.

Tests run red first: wire rejects the retired words and the missing metadata; client search
minimum, execute omits account, account_status 404/429; managed targets carry the metadata and the
account id, a pending row moves nothing; the logo ladder order; the store passes the fields through.

--- feat(connectors): onboarding connect-first rides the connection operation

The guided first build had its own connector surface: a 2 s connectors.list poll loop, an auto-open
of every minted link, a hidden "[setup] links opened" note that painted as a user bubble, CONNECT
FIRST rows in the transcript, and a "Start with N apps connected" composer pill that sent the go
signal early. Delete all of it. The build session's first manage_connections connect now shows the
same card as any chat, and the settled tool result is the go signal.

Deleted: store/first-build-connectors.ts, assistant-ui/first-build-connectors.tsx (nothing rendered
it any more), lib/first-build-start.ts, onboarding-chat/start.tsx, the first-build session markers in
handoff-receipt.ts, the connectionRows / latestConnectorPart helpers they were the last readers of,
and the five strings only they used.

The runbook (setup-profile.ts::connectFirstRunbook) now describes the operation's real contract: one
connect call with every picked slug, one card with a row per app, blocks until settled, a per-app
result of connected / skipped / not_connected. No wait, no "Start with", no "already active" branch
(the result never carried it), and no offer to re-mint: Try again and Continue live on the card, and a
Continue means the user moved on.

Tests (red first): the runbook names one connect call and its per-app result and never the retired
model actions; the no-account rule survives an empty pick.

--- feat(connectors): the registry is keyed by profile, every watch read is bounded by the deadline, connectionId is optional, connect and execute carry returnTo and op

Registry (P1-14): tools/connectors/live.py keys an operation by (profile home, session key). Two multiplexed
profiles can carry the same timestamp-based session key; the tool thread opens under its turn's profile
override and the RPC side names the session record's profile home, the default profile resolving to the
process home on both sides.

Watch reads (P1-8 residual): list_connectors takes a per-page timeout and the watcher bounds each read by
the operation's remaining deadline (floor one second), so a stalled gateway page cannot hold the operation
past its deadline.

Contract follow-up from the portal handoff (/tmp/magic/HANDOFF-HERMES-PR3-PORTAL-CONTRACT.md): connectionId
on a mint result and on CONNECTION_REQUIRED is plain Optional (a no-auth toolkit answers active with none; a
failed mint answers with neither link nor id); a target without one has nothing to watch. Connect and
execute requests carry returnTo (hermes-desktop | hermes-desktop-dev | portal) and op, so the vendor's done
page can send the browser back to the app; the call sites and the deep-link handler follow in the next
commit.

Tests, red first: two profiles share a session key without seeing each other; each watch read shrinks with
the deadline; the client honours a per-page timeout; the optional id; the two request fields. The fakes'
list_connectors accept the timeout keyword; two wrapper spies that did not were the cause of a 300 s hang
under the per-file harness.

--- feat(connectors): an MCP target runs on the shared connection operation

manage_connections ran MCP targets as a renderer errand: the desktop card ran
the catalog lookup, the install and the OAuth flow, then told the backend what
had happened, and every surface without a card got `unavailable` plus two
terminal commands. That left the outcome in the renderer's word, and left the
model unable to connect an MCP server anywhere but the desktop.

The backend now owns the work, the way it already owns a managed connector.
`prepare` starts the OAuth flow in process and points the browser at the
backend's own callback route; an install that still needs credentials waits
pending and publishes their names as `required_env`; `observe` reads the flow
or the install worker on every tick. The worker itself moved out of
tui_gateway/mcp_oauth_sessions.py into tools/connectors/mcp_oauth.py, and that
module now calls it, so the Capabilities tab keeps its RPC session table.

The card may only say approved, skipped or continue: a claim of any other
state moves nothing. With that, `renderer_flow` and the `unavailable` state
have no producer left and are gone from the contract. Off the desktop there is
no card, so the action runs at once and the result carries the authorization
URL for the user to open.

Try again reaches the same work: `_reissue` dispatches on `target.kind`, so an
MCP row re-runs its own install, enable or OAuth instead of minting a managed
gateway link.

(cherry picked from commit fa0438c4121807604e7c983ba42b901314a1305b)

--- feat(connectors): the MCP card projects the operation instead of running the flow

The MCP setup card used to own the work: it called the catalog, ran
installMcpCatalogEntry, polled the action, drove the OAuth window, and then told
the backend what had happened through connection.respond. That made the renderer
a second authority on target state, so an install that finished after the window
closed, or a card that never mounted, left the operation with a state nobody
could correct. PR3 moves that work to the backend, so the card has one job left:
show the operation and send the user's consent.

McpSetupPending now renders request.targets the way ConnectorOffer renders them:
one row per target, one verb by (action, state), Continue below the shell.
Install and Enable send {status: 'approved'} and nothing else; an authorize row
opens the link the backend minted, with no OAuth RPC of its own; a running
install holds its verb with the busy mark; a failed row sends
connectors.connect {reconnect: true} on the open operation, the same call the
managed card makes. The owner lookup and that call are now shared helpers on
connector-tool.tsx rather than a second copy here.

ConnectionTargetOutcome loses connected / initiated / failed: those were the
renderer reporting state, which it no longer may do. ConnectionActor loses
renderer_flow for the same reason. ConnectionOperationTarget gains required_env,
the credentials an MCP install is still waiting for; the card renders a field per
entry under the pending row and holds Install until every required one has text,
so the values travel with the approval instead of through a separate catalog call.

(cherry picked from commit 2e42ad19f63e41f274a87199d07eb99a7c995cbe)

--- fix(connectors): the toolkit metadata leaves the wire; the vendor logo is derived from the slug

Sid cancelled portal PR3 (the toolkit metadata on the list item) on 2026-09-14, so the fields the contract
commit typed as required come off the wire: no title, description, iconUrl, authKind or activeConnectionId
on the list item, no total on the page, no search or connected query, no title or icon_url on a target
snapshot. The list item is exactly what portal #1220 emits.

The desktop derives the vendor logo from the toolkit slug instead (`connectorIconUrl`,
https://logos.composio.dev/api/{slug}; the gateway slug is the vendor slug, checked for every lead-order
pick), rendered as a plain image because the host sends no CORS header. The title stays
`connectorTitle(slug)`. The mark ladder is unchanged: glyph, derived vendor icon, favicon, monogram.

Everything else in the contract stands: six-state status, optional connectionId on mint results and on
CONNECTION_REQUIRED, returnTo and op, the account routes. The contract commit's message still names the
metadata; this commit is the correction.

--- fix(connectors): an MCP target survives the card's Continue, and the model is not sent to an action that refuses it

The review of the MCP backend found eight defects; each one is a test first.

- The off-desktop note told the model to confirm with action 'status', which refuses MCP
  names outright. It now says the authorization finishes in the background and the tools
  arrive on the next turn.
- The card's existence is a property of the session. A missing callback no longer routes a
  desktop session down the off-desktop path, where the model would be handed a live link.
- A second failure of a backend attempt was published with actor 'user'. ``refresh`` takes
  the actor, so the frame says who produced the text.
- Continue on the RPC thread can settle the operation between any read of a target's state
  and the transition that follows it. A settled operation has a frozen result, so the lost
  move is dropped; an IllegalTransition no longer escapes into the tool result. The same
  rule covers a worker whose outcome arrives late and a Try again that arrives after Continue.
- Try again on an install carried no credentials, so the second install ran with an empty
  env. The approved map is kept on the runner (never on the target: target fields are
  serialised to the model) and reused.
- An approval that does not cover a required credential no longer reaches the worker, where
  ``install_entry``'s prompt would block on stdin forever; the row waits for the card's
  fields instead.
- ``enable`` writes ``mcp_servers`` under the scope and lock the dashboard's toggle route
  uses, so the two read-modify-write paths in one process cannot drop each other's write.
- Several authorize targets start their flows together and share one wait; one wait per
  target kept the card empty for minutes.
- The off-desktop operation is in no session's registry, so it now emits no connection.update.

Try again is also refused for a target the transition table cannot move back to 'initiated'
(an MCP target has no move out of 'expired') and for an operation that settled during the call.

The contract test that froze the Actor enum is replaced by two behaviour tests: no actor but
the backend watcher can connect an MCP target, and the card's word never moves one.

(cherry picked from commit 081c16412b48667758a77a7e7c29eaf2be6a93dd)

--- fix(connectors): the watcher reads one account route per target, and every frame carries its seq

The watch loop walked the whole toolkit list once per tick to learn whether one
target had connected. That read costs a vendor call per page, cannot tell one
account from another, and forced `awaiting_new_attempt`: after a forced
reconnect the list still reported the OLD account `active`, so the row had to be
disbelieved until it read as something else once. The gateway now serves
`GET /v1/connectors/accounts/{connectionId}`, so each pending target reads its
own account: the id the mint named, one read per target per tick at 1 Hz (the
route's bucket is 180/min per principal), each bounded by the operation's
remaining deadline. A 429 parks that one target until its Retry-After passes and
leaves the others reading. A 404 is "not yet" until the deadline. A target the
mint gave no account for has nothing to read, so it is not read. A forced
reconnect watches the new account, which is why `awaiting_new_attempt` and its
two tests are gone.

`mint` and `run_remote` now name where the browser should come back to
(`returnTo`, plus the operation id on a mint), so the vendor's done page can
hand the user back to the desktop app that asked instead of stranding them on a
web page. Only the desktop registers that URL scheme, so no other surface sends
either field.

`connection.update` frames were built by re-reading the operation after the lock
was released, so a second writer could give an older frame a newer state and the
renderer could not tell which frame was last. Every write now advances a
monotonic `seq` and takes its snapshot under the same lock, and the emitter sends
that snapshot; a renderer that keeps the highest seq per operation can drop a
frame that arrives out of order. `connectors.operation.wake` lets the desktop's
deep-link return ask for a read now instead of at the next tick; it only shortens
the wait and trusts nothing else in the link.

(cherry picked from commit f5423cf72302b92b26c043495184ad6426fa2b36)

--- fix(connectors): the desktop card follows the operation, and says so out loud

The MCP card was a second implementation of the connector card with the
review's defects: it painted for any request on the session, kept its
controls after the operation settled, offered an approve verb while the
backend was still minting an authorize link, and re-enabled Install when
the RPC returned rather than when the state frame moved the row. A second
click in that window sent the consent twice.

The renderer now reads the operation's `seq`: an update or a status frame
whose sequence is not greater than the one already applied changes
nothing, and a resume snapshot neither revives a settled card nor puts
back a row a newer frame has moved. Without it the transport's ordering
decided what the user saw.

`hermes://connections/done?op=…` brings the user back from the browser to
the session that opened the operation and wakes its watcher, so the row
moves at once instead of at the watcher's next tick. Only the op id is
used; the link's status moves no row.

Each row's mark and cue are one polite live region, so a row that flips is
heard and not only seen, focus follows the row the backend moved while the
card holds it, and the waiting mark stops spinning under reduced motion.

(cherry picked from commit 7c30d252556d7d496ddb354b50bcecc41bceef27)

--- fix(connectors): the renderer types seq as the wire carries it

The watcher commit made `seq` a required field on every operation frame. The
desktop store still declared its own optional `seq` so it would compile against a
backend that predates the field; that backend no longer exists on this branch, so
the hedge is dead code and the fixtures were short one field. The store now reads
`seq` from the shared types, and the fixtures count the way the backend does.

The repeat-frame test asserted the whole request keeps its reference. With a real
rising `seq` the request must change; the invariant the test guards is that the
target row keeps its identity so open credential inputs do not remount.

--- fix(connectors): a URL that arrives after the shared wait still lands on its row

The shared prepare wait failed every row still pending when it ran out, while
that row's own thread was still waiting on the provider. When the URL arrived a
moment later the thread's move raised inside the daemon thread, the link was
lost, and Try again started a second flow. The wait now bounds only how long
prepare blocks; a row still pending afterwards is left to its own thread, which
is the only writer of that row and ends with the URL or the flow's own failure.

The comment on the approved credentials said every target field is serialised to
the model; it is not (the snapshot names its keys). The reason they live on the
runner is that they are secrets and the runner's life is exactly the operation's.

--- fix(connectors): the review findings the operation must survive before the fold

The MCP prepare threads and the install worker started with an empty context,
so a named-profile turn's home override never reached them: the flow resolved
the process home's `mcp_servers` and stored the token there. Each thread now
runs in a copy of the calling thread's context. `connection.respond` had the
same gap on the RPC thread: an approval ran the enable, which writes
config.yaml, with no profile bound, so the flag landed in the launch home. The
handler binds the session record's profile the way `_connector_rpc` does;
`config_write_scope(None)` keeps that override, so the enable needs no change.

The watcher raised out of the tool when a per-target Skip landed while that
target's read was in flight: the operation stayed open with `live` closed, and
every later answer got 4004. A read for a row that is no longer live is dropped
at debug; only a refusal on a live row is still a contract violation. The same
skip from the card raced a row the backend had just connected and aborted the
rest of the answer; a skip for a resolved row is ignored and every entry, then
the settle check, still runs.

A read was bounded by the whole remaining deadline, so a hung gateway held the
first read for 300 s and Continue could not return the tool; a read now waits
ten seconds at most. A 429 parked only the target that read, but the budget is
the principal's, so the next target's read in the same tick spent it again:
every live target waits out the one Retry-After.

The MCP surface rule read the platform alone, so a desktop call without the
callback (registry dispatch from execute_code) opened an operation nobody
rendered and blocked for the deadline; it now uses the managed rule, surface
and callback. A write after settlement advanced `seq` while emitting the
frozen frame, so the resume snapshot named a seq no frame carried; the counter
stops at the settle frame. A mint that reports `initiated` with no account id
logs that the watcher cannot read the row.

(cherry picked from commit 587b228016bf8ee2b022e7fa58c2d9f6057809e9)

--- fix(connectors): the desktop card holds a verb until the backend answers, and never takes the keyboard from a credential field

The review of the desktop card found five defects; each one is a test first.

- A resume snapshot was refused whenever the cache held a settled operation, whichever
  operation it was, so a session that opened a second operation after settling the first
  never got its card back from a resume. And the refused snapshot handed the caller the
  settled cache as "the request", so the session was flagged as needing input behind a
  summary with no controls. Only the same operation can refuse the snapshot now, and a
  refused one is no pending card.
- The focus handoff picked the row's first button, which after pending -> initiated is the
  disabled working verb; the focus call was a no-op and the keyboard landed on the document
  body. It picks the first control that can take focus, else the row. It also moved focus out
  of a credential field the user was typing in whenever another row moved; it leaves an
  editable alone. The "focus Continue once every row resolved" branch was dead (Continue
  unmounts the moment nothing is unresolved), so it and its ref plumbing are gone.
- The done link navigated to a settled operation's session and rejected when the wake RPC
  did (4004 once the operation left the live registry). A settled request is ignored, and a
  refused wake is nothing: the wake only shortens the wait, the watcher still ticks.
- Install spun forever when the store refused to send the consent (the operation gone or
  settled under the card): `respondToConnectionRequest` resolves false in that case and the
  verb was only released in `catch`.
- After a partial approval (a required credential missing) the backend answers with a
  same-state frame whose detail names what is missing; nothing released the verb because it
  was held until the row's state moved. The row now remembers the seq the click saw and holds
  the verb only until a frame past it arrives, which is the backend's word on the click
  whether or not the row moved.

(cherry picked from commit 4c534d6e2af979778d9a2423cf97d717edeaa1a9)

--- fix(connectors): a skip that loses the race to the watcher is ignored, not raised

The skip guard read the row's state and then moved it; the watcher can connect the
row between the two, and the refused move aborted the rest of the card's answer.
The refusal itself is now the witness: a move refused for a row that is resolved,
or on a settled operation, is the same nothing-to-do as a row resolved earlier.

Two recording fakes in the managed tests kept their lists on the class; they now
start per instance so a lifted fake cannot share reads between tests.

--- fix(connectors): a resume that lands behind a newer live frame still reports the pending card

The refused-snapshot branch answered "no pending card" for both reasons it can
refuse: the operation settled, or a newer frame already moved a row. Only the
first is no card. For the second the live card is still open and blocking the
turn, so the caller must keep the session flagged as waiting on it.

* test(connectors): defer new connection coverage until implementation settles

Remove PR-added test cases and their unused helpers while retaining
existing tests adapted to the changed connection contract. The three
PR-only renderer test files are removed for now.

Focused behavioral coverage will be added as the final implementation
step before verification. Existing main coverage is not being removed
wholesale, and this does not declare the feature merge-ready.

* fix(connectors): commit MCP authorization at initialize, save setup values after success

The OAuth probe treated one exception as one outcome: any failure after the browser
step restored the token snapshot and manager entry, so a server that accepted the
token but failed tools/list discarded a completed consent. Now the probe reports
whether initialize succeeded (details["initialized"], read from the claimed
MCPServerTask). Failure before that point rolls back as before. Failure after it
saves the server config, keeps the tokens, and reports tools unavailable through
flow.discovery_error; the card can retry discovery without repeating consent.

Catalog install wrote the submitted values to .env before install_entry and the
probe ran. The values now live in the secret scope for the duration of the install
(get_env_value reads through get_secret, so install_entry finds them without a
prompt), and .env is written only after the probe returns tools. A failed probe
removes the server block and writes nothing. _probe_tool_names returns None on a
failed probe instead of an empty list, so failure and a valid empty listing are
distinct.

Failure text is redacted before it reaches target detail: every value the card
submitted for that target is replaced by exact match, then the pattern redactor
runs.

required_env now carries the manifest's secret and default flags; Target carries the
manifest's post_install text as instructions. The wire contract gains secret,
default, instructions and discovery_error; generated TS and OpenRPC regenerated.

* fix(connectors): one OAuth callback receiver picker for the connection card

The card's authorize target built its redirect from the dashboard web server and
raised when none was bound in the process, so a standalone hermes --tui session
could never authorize an MCP server. The receiver is now chosen in one place
(choose_callback_receiver): a pinned pre-registered client keeps the SDK's own
listener on the registered port; a client-advertised loopback URI is used as-is and
its callback arrives through the mcp.servers.oauth.callback relay; otherwise the
backend binds a one-shot loopback listener and feeds it into the flow. The dashboard
route stays with the dashboard web page, which cannot bind a port.

tui_gateway/mcp_oauth_sessions.py had a second copy of the loopback listener and a
_worker that referenced _probe_with_rollback, set_hermes_home_override,
reset_hermes_home_override and Path without importing them, so every RPC-started
flow raised NameError. Both are deleted; start_flow spawns run_worker directly and
uses the same receiver picker. Flow registration is shared (register_flow /
finish_flow) so a card-started flow with a client URI is reachable by the relay.

Under an SSH session with no client listener the attempt's detail carries the
existing paste-the-redirect instructions. Nothing on the card path opens a browser.

* feat(desktop): setup-form modal for MCP connection cards

The connection card rendered an MCP server's setup fields inline: every field as a
password input, no default value, no instructions. A URL such as the n8n MCP
server URL was typed blind, and the manifest's setup text never reached the user.

Two components carry the form now. SetupFieldList renders the ordered fields the
backend declares (a plain field as text prefilled with the manifest default, a
secret field masked and empty). SetupFormDialog composes it with a one-line title,
the manifest instructions, an inline error for a failed attempt, and Cancel /
Connect. The row's Install action opens the dialog when the target has fields;
Connect sends {status: approved, env}; Cancel sends {status: skipped}. A failed
attempt keeps the dialog open with the draft intact. The draft lives in the dialog
component only; nothing reaches the store or the resume snapshot.

Once the backend publishes the authorization URL the dialog shows it as text with
an Open in browser button. Nothing opens a browser on a state change: the Try
again path on both cards used to open the re-minted link at once; it now waits for
the row's update frame and the user's click.

The store types gain the wire's secret, default, instructions and discoveryError
fields and normalise them; a connected target with discoveryError renders as
authorized with tools unavailable.

* feat(tui): connection card in the Ink TUI

The Ink TUI had no client for the connection operation: connection.request,
connection.update, connection.respond and the pending_connection resume snapshot
were unhandled, so a manage_connections call in a TUI session could only print a
link through the model.

connectionOperationStore.ts holds the backend snapshot: a request opens only for a
new operation id, an update applies only to the live operation with a higher seq,
a settled update freezes the id so a late request frame cannot reopen the card.
The gateway event handler feeds it; session resume hydrates it from
pending_connection.

connectionSetupOverlay.tsx renders one callout in the prompt zone: a one-line
title, the manifest instructions, every field (plain rows prefilled with the
default, secret rows masked), then a Connect / Cancel selector. Connect sends the
draft through connection.respond; Cancel skips the active target. When the backend
publishes the authorization URL the callout shows it as text under "Press Enter to
open in browser"; Enter is the only thing that opens it. A failed attempt unlocks
the fields with the draft kept. An accepted secret renders as "Set" and is never
echoed. A target authorized without tools shows that state and Continue.

The overlay joins the existing input-owner set in overlayStore so typing and other
prompts are blocked while it is up.

* feat(cli): connection panel in the classic CLI

The classic hermes CLI passed no connection_callback, so a manage_connections
call could only print an authorization link through the model and could not take
a setup value at all.

The CLI now renders the connection operation as a prompt_toolkit panel, on the
same queue mechanism as the clarify panel: the agent thread's callback opens the
panel, blocks until the first decision, then returns so the operation's watch loop
runs; every later state reaches the panel through ConnectionOperation.on_change
(installed only when no gateway hook is set, restored on close).

Panel: one-line title, the manifest instructions, one row per field (plain rows
prefilled with the default, secret rows masked in the input buffer and rendered as
"Set" once accepted), then Connect / Cancel. Connect sends the draft through
apply_answer; Esc skips the active target; Ctrl-C sets the tool-thread interrupt so
the operation settles as interrupted. Once the backend publishes the authorization
URL the panel shows it with the target detail and "Press Enter to open in browser";
Enter is the only thing that opens it. A failed attempt unlocks the fields with the
draft kept; an authorized target without tools offers Retry discovery / Continue.

Up/Down move between rows, Left/Right toggle the action, and the panel joins the
blocking-overlay guards so chat input, history and voice stay out while it is up.
The single-query (headless) mode passes no callback, as it does for clarify.

* feat(connectors): register a connected MCP server and report its tools in the result

After a successful install or authorization the target carried the probe's tool
names and the model was told the tools "become available on your next turn". MCP
tool schemas are deferrable by construction, so nothing about them lives in the
sent tool array; a server registered in the scoped registry is callable through
tool_describe/tool_call in the same turn. The operation now registers the server
(register_mcp_servers under the owner's home scope) once authorization is
committed, records the registered names on the target, and the settled result
carries a tools_listing block in the deferred-catalog format plus a note that the
tools are callable now. A registration failure keeps the target connected with
tools: [] and a sanitized discovery_error. agent.tools and the system prompt are
untouched; the between-turns refresh updates the catalog block as before.

The card gate no longer asks for the desktop platform. Every surface that renders
the card attaches a connection callback (Desktop, the Ink TUI, the classic CLI);
registry dispatch and messaging sessions attach none and keep the link result.

Pre-commit rollback in probe_with_rollback used restore(only_if_absent=True),
which skips the rollback when a token file exists. On a first-ever authorization
the only file is the one this attempt wrote, so a token the resource rejected was
kept. Live E2E (controlled provider answering 403 to the issued token) showed the
row fail and the token survive; the rollback now restores the snapshot outright,
and the same run shows the token file removed.

* fix(connectors): install an OAuth catalog entry through the card's own flow

The card's install ran install_entry and then a plain probe. For an OAuth entry
that probe has no card flow around it: in the desktop backend it failed at once
("non-interactive environment and no cached tokens"), and in the classic CLI it
saw a TTY, opened a browser by itself and drew install_entry's curses tool
checklist over the panel. Since a probe failure is now an error, the failed
install was rolled back with _remove_mcp_server, and authorize refuses a server
that is not configured. A clean home had no path to a connected OAuth entry, which
is 55 of the 65 catalog entries. Found by the live three-surface run.

An install now builds the entry's configuration in memory (card_install_config:
no prompts, no probe, no checklist; a prior tool selection or the manifest's
curated filter applies). An OAuth entry starts the same flow authorize uses with
that configuration and the setup values in the attempt's secret scope, so the row
reaches the URL step, and the configuration and the setup values are saved
together when initialize accepts the token. Every other entry is probed in memory
and saved after the server answers. A failure writes nothing, so a failed
reinstall keeps the previous configuration, and the failed row asks for its fields
again so the card can reopen the form over the draft it kept.

* fix(agent): make a server connected in this turn callable in this turn

tool_describe and tool_call resolve names inside the agent's toolset selection,
which is fixed when the agent is built. A server that manage_connections had just
registered was therefore "not found" for the rest of the turn whose result calls
its tools available, and stayed out of the next turn's catalog too. The live
Desktop run showed it: the result listed mcp__fx_oauth_fields__echo and the
tool_describe that followed answered not_found.

The executor now adds the MCP servers the call connected to the selection. Only
the selection changes; agent.tools does not, so the sent tool schema bytes stay
the same. A selection of None (every toolset) and the no_mcp sentinel are left
alone.

* fix(cli): reopen the connection form when a required field is still empty

When the backend refuses an answer because a required field is missing it keeps
the row pending and names the fields. The panel mapped that frame to its waiting
phase, which draws the detail line and nothing else, so the user saw "waiting for
FX_API_KEY" with no fields and no buttons until the 300 s deadline. The panel now
returns to the form on the first missing field, over the draft it kept.

* test(connectors): the no-card path is the one with no callback attached

The card gate is "a connection callback is attached" since the classic CLI got
its own panel. This test still attached one under platform "cli" and expected no
card, so it opened a real operation, waited out the 300 s deadline and failed.
It now drives the path that has no card: no callback.

* fix(connectors): order a cancel against the commit and stop deleting a working grant

Found by the live runs on Desktop, the Ink TUI and the classic CLI.

A user's skip did not stop the OAuth attempt. With the worker parked in the token
request, the row settled "skipped" and the token was written 33 s later. A skip
now cancels the attempt, and one lock orders that cancel against the commit: the
attempt is either canceled with the earlier tokens restored, or committed and
kept. When the commit won, the skipped row says the authorization was kept. An
interrupted turn cancels its attempts the same way.

Every attempt began by deleting the saved tokens, so retrying discovery for an
authorized server demanded consent again, and a cancel in between left no grant
at all. The card's flow now connects with the saved tokens first, with no browser
step; only when they do not work does it replace them. The RPC session surface
keeps the old behavior because its caller waits for an authorization URL.

A second Connect from the form a failed row reopened was dropped without a frame,
which left the Desktop dialog with Connect loading and Cancel disabled until the
deadline. An approval on a failed or expired row is now a retry with the new
values.

The reported tool names came from the registration call, which returns nothing
for a server the process already holds. A retry after a failed listing therefore
said "no tools" while the server had them. The names are read from the registry,
and a parked server is woken first.

An authorization that commits after its card closed by deadline was never
registered, so the next turn still could not reach it. The runner keeps such
attempts, and the between-turns refresh adopts the ones that were approved.

An error with no message reached the user as a class name ("CancelledError").
The tool description still said a server's tools arrive on the next turn.

* fix(desktop): let an OAuth install open its link, and keep the form's draft

An install of an OAuth catalog entry with no setup fields reached the URL step
and gave the user nothing to click: an initiated install was always drawn as a
disabled spinner, and the only other place the link is shown is the setup dialog,
which opens for entries with fields. That is most of the catalog. An initiated row
that carries a link now offers Open, whatever the action.

The setup dialog reset its draft whenever the field list changed identity. The
backend sends a fresh list with every frame and an empty one while an attempt
runs, so a failed Connect erased what the user had typed. Fields now only fill in
what the draft lacks, and closing the dialog drops the draft.

The settled summary dropped the "tools unavailable" fact, and the live row offered
a Try again for that state which the gateway refuses for a connected target. The
summary keeps the fact and the row offers no dead control.

* fix(tui): keep the typed draft when a Connect fails

The overlay reset its draft whenever required_env changed identity, and every
backend snapshot delivers a freshly parsed array. A failed Connect therefore came
back as a form with the default region and an empty secret. The draft now resets
per target only, and a snapshot's fields fill in what the draft lacks.

* fix(cli): show the URL step and start an install that has no fields

The panel applied the user's answer and then set its phase to "waiting". The
backend applies the answer on the same thread and its change hook had already set
the next phase, so the URL step of an OAuth install and the reopened form for a
missing field were both overwritten, and the card sat on "Waiting…" until the
deadline. The waiting phase is now set before the answer is applied.

A pending install or enable with no fields opened in the waiting phase, which
sends no approval, so the flow never started. It now opens on Connect/Cancel. A
target that already carries its link opens on the URL step.

Connect on a failed row re-ran the attempt without the values now in the draft; it
sends them. The authorized-without-tools phase offered a "Retry discovery" that a
connected target cannot run inside the same operation; it offers Continue.

* fix(connectors): a newer OAuth attempt replaces the older one for the same server

A card that closed by deadline leaves its worker waiting on the browser for up to
300 s, and a tampered callback leaves one waiting too. A retry or a new operation
for the same server then ran beside it: both wrote the same token files, and the
older one's rollback could write over the newer attempt's grant.

The newest attempt per home and server is recorded. Starting one cancels the older
attempt, takes over its pre-attempt snapshot so a later failure still restores the
state from before either, and the older attempt's rollback leaves the files alone.

* fix(desktop): label an install's link control "Open in browser"

The control that hands an install's authorization link to the browser reused the
action's verb, so the row showed "Install" before the click and "Install" again at
the link step, told apart only by the cue. It now reads "Open in browser", the
label the setup dialog already uses for the same act. Authorize keeps its verb.

* docs(mcp): describe the setup card on the desktop, the terminal UI and the CLI

The MCP guide said the chat install exists only in the desktop app and that the
CLI relays commands. All three surfaces now show the same card: fields, Connect or
Cancel, an authorization link the user opens, one save when the server has
accepted the token, and tools the agent can call in the same turn. The tools
reference gains the result fields (tools, tools_listing, discovery_error) and the
no-card behavior of an OAuth install.

* fix(cli): keep the layout hook callable without a connection widget

The CLI panel commit added connection_widget to _build_tui_layout_children as a
required keyword. That method is the documented override point for wrappers, and
five existing tests (extension hooks, prompt stash, subagent dock) call it without
the new argument, so CI failed with a TypeError. The argument is now optional, like
the other widgets added after the hook was published; a missing widget is left out
of the layout.

The settled tool result also carried the target's setup instructions. Those are
the card's text for the user, and a catalog entry's notes can predate this flow
("restart your session so the tools are loaded"), which contradicts a result that
says the tools are callable now. The model-facing result drops them; the card
payloads keep them. This restores test_connector_local_batches.

* fix(connectors): wait for an in-progress registration before reporting a server's tools

Found with the real Vercel MCP server. Saving the configuration wakes the config
watcher, which starts its own connect for the new server. The operation's
registration call then skips the server as "already connecting" and returns at
once, so the card settled "connected" with no tools, and the 214 tools were
registered three seconds later. With no names in the result the model searched,
found the hosted connector of the same vendor and asked the user to connect that
instead.

The registered names are now read once the registration has finished: while
another task is connecting the server the read waits (30 s at most), a parked
server is woken once, and a server that finished registering with no tools is
still a valid empty list.
2026-09-21 10:04:32 +05:30
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
60dd9778c6 fix(desktop): large text pastes attach from HERMES_HOME/composer-pastes when the chat cwd is elsewhere
Desktop persisted a large paste under Electron's userData dir and attached
it as `@file:<abs path>`; the tui_gateway prompt path expands that reference
with `allowed_root=cwd`, so `_resolve_path` refused it with "path is outside
the allowed workspace" whenever the chat's cwd was not an ancestor of the
paste dir (always on Linux/macOS, and on Windows for any project cwd).

Electron now writes pastes under `<HERMES_HOME>/composer-pastes`, and
`agent/context_references.py::_resolve_path` admits exactly that anchored
directory (active profile home and global root, via file_safety._hermes_dirs)
as the one root besides `allowed_root`. A sibling directory that merely
contains the substring stays refused; the credential deny-list still runs
on the admitted path.

Supersedes the substring whitelist proposed in #117150.

Fixes #117149
2026-09-20 19:45:30 -07:00
teknium1
c545272568 fix(lsp): one stalled request no longer silences a workspace for good — retry window, cold-root warm-up budget, per-root exclusion
A language server that missed its budget once marked its (server, root) pair broken for the
process lifetime, the same 5 s steady-state budget was applied to a cold server that also had
to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched
the server off for every workspace.  Three new keys under the existing `lsp` block, all defaulting
to today's behaviour:

- lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per
  pair; an expired pair gets one more try, and the INFO skip line names the retry time.
- lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client
  waits up to this budget (outer join budget follows); warm requests keep wait_timeout.
- lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path
  also covers everything beneath it); a matching root never spawns, logged once at INFO.  A
  non-list value fails closed — WARNING naming the expected shape, every root skipped — because
  silently excluding nothing would re-pay the stall the key was meant to avoid.

Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo).
2026-09-20 18:54:24 -07:00
Chukuwebuka-2003
7d3c0b2f94 fix(compression): an auto-resolved summary model that fails falls back to the main model and is named in the warning
`provider: auto` resolves a compression summary model per call WITHOUT setting
`summary_model`, so the main-model retry gate saw "no separate model" and
re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty
content) on every attempt, and the user-visible aux-failure warning had no
model to name. Record the model the aux lane actually resolved
(`_last_aux_resolved_model`), use it in the retry gate, and pass it into
`_fallback_to_main_for_compression` so the warning names it.

Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor /
_last_aux_resolved_model hunks, the attempt-state field, and its test; the
over-window wait cap, preflight fail-closed and Desktop renderer hunks are
carried by #117084 / #117140.

Part of #116472 (request 4).
2026-09-20 18:53:10 -07:00
fangliquan
bb9058d7e8 fix(compression): preserve steer display identity
(cherry picked from commit a8570878bc848a42cc8029fc45219a57e76be4c4)
2026-09-20 18:24:07 -07:00
teknium1
f568b860d7 fix(agent): pin the reset_at reach through try_activate_fallback and add opt-in fallback.min_switch_reset_seconds
- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback
  (production entry) on a real AIAgent with a one-entry chain, so dropping the
  reset_at forwarding goes red (3 failures before, 8 green after).
- #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the
  rate-limited primary's declared reset is sooner than N seconds,
  try_activate_fallback returns False and no cooldown is armed; docs row added.
2026-09-20 17:00:43 -07:00
fangliquanflq
125bdff601 fix(agent): honor provider reset for fallback cooldown 2026-09-20 17:00:43 -07:00
finn763
36b6efc297 fix(context): re-derive model.context_length on model/provider change
model.context_length is the user's profile-wide ceiling. It was read from
config.yaml in exactly one place — agent construction — and cached twice:
agent._config_context_length (switch/fallback resolution plus every display and
/usage surface) and context_compressor._config_context_length (the compressor's
own re-resolution).

Every live path that re-resolved a runtime then touched only one copy, or cleared
it without re-reading the config:

- switch_model nulled agent._config_context_length and re-derived the intent from
  custom_providers metadata alone, so a ceiling that only exists as
  model.context_length was dropped for the rest of the process;
- the Desktop/TUI compression hot-reload updated the compressor's copy only, so an
  open session showed a pinned ceiling while compressing against provider
  metadata / the 256K fallback.

Both now route through one pair of helpers in agent/agent_init.py:
set_config_context_length (one place that knows where the pin is cached) and
config_context_length_for_runtime (re-read from live config, scoped exactly like
construction, so an unrelated route still never inherits the pin).

(cherry picked from commit 986ff16dadb9966f7328e55f295af5cfa1eb5c88)
2026-09-20 16:31:10 -07:00
chelsealong
de622b291d fix(agent): detect Thai plan tails in promoted-reasoning stall guard
promoted_reasoning_announces_action()'s tail detector only matched
English (plus CJK punctuation boundaries), so a reasoning-only clean
stop ending on a Thai first-person plan (e.g. "จะให้ผม...") was not
recognized as a stall and got delivered to the user as the final
answer instead of nudging continuation.

Add Thai first-person future-action triggers, extend the boundary
class with em/en dash (a common Thai clause separator), and accept
multi-dot ellipsis tails.

(cherry picked from commit 87061e46c423859cf738d4541df6594233f40e7e)
2026-09-20 16:20:52 -07:00
teknium1
53815e24dc fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.

The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.

Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
  before  req_reasoning: {}
  after   req_reasoning: {'reasoning_effort': 'medium'}
  agent.reasoning_effort: low  ->  {'reasoning_effort': 'low'}  (unchanged)
2026-09-20 16:03:38 -07:00
Moep90
3a37a24efd fix(moa): mark reused advisor guidance as predating the tool results
With `fanout: user_turn` (and off-cadence `every_n` iterations) the advisors run
once per user turn and their guidance is replayed verbatim into every later
iteration of that turn. The block reads as fresh instruction, so an advisor that
proposes a tool call keeps proposing it after the acting model already ran it and
has the result in the transcript.

Observed with the clarify tool: the card was answered, the next iteration replayed
the same guidance, the model issued a second identical card, the user dismissed it,
and the turn then held two contradicting results for one question (an answer and an
empty skip). The following aggregation degenerated into a repetition loop until it
hit the output cap.

- `_STALE_GUIDANCE_NOTE` is appended when cached guidance is reused on an iteration
  that already has tool activity since the last real user turn. The advice text is
  handed over unchanged; only the framing says it may be out of date.
- The advisor system prompt now rules out emitting a tool call or a JSON tool-call
  object. Advisors hold no tools, and a tool-call object in advisory text is what
  the aggregator replays.

No change to fanout cadence, caching or accounting. The cadence test that pinned
byte-identical reuse now asserts the advice text is reused and carries the marker.

Signed-off-by: Moep90 <3042152+Moep90@users.noreply.github.com>
(cherry picked from commit 343787f287ad9915345090fba351df5ffa758c9e)
2026-09-20 15:51:19 -07:00
teknium1
0765099ff4 fix(compression): fence the durable cooldown rollback per compressor; one stale-attempt helper
The SQLite rollback no longer runs under the process-wide claim lock: a per-compressor serial lock (taken by _claim_compressor_attempt too) serializes it against claims on that compressor only. The seven pasted working-attempt checks call _raise_if_stale_attempt/_caller_attempt_is_current. Drops the unused _run_as_attempt test helper.
2026-09-20 15:50:49 -07:00
beardthelion
0e33dc9ebc fix(compression): stop detached stale attempts writing shared compressor state
The stall-fallback detaches a timed-out primary worker and reuses the
same ContextCompressor, but the existing attempt-generation guards only
covered the unwind-time snapshot restore. Every other summary-state
write stayed reachable by the still-running primary after the fallback
took over: a late successful summary published _previous_summary and
cleared the fallback's cooldown, a late failure armed a shared failure
cooldown and stamped error state, the cancel rollback and the abort
rollback reverted _previous_summary to the primary's snapshot, and the
durable cooldown rollback row could be overwritten mid-restore.

Compressor code could not fix this with the shared attributes alone:
those cells only name the current owner, never the calling attempt.
The calling attempt's generation now rides a ContextVar bound inside
_run_summary_dispatch, which every attempt's compress_fn passes through
in its own thread, so each attempt reads its own generation. Gates on
the working-attempt marker (not the entry claim, so lock sit-outs do
not suppress the owner) now cover the cancel rollback, late-success
writes, _on_summary_failure, the abort rollback, the deterministic
pin, compress() entry, and a Phase-3 choke point. The durable cooldown
rollback moved inside the claim lock so the DB row and the in-memory
restore are atomic against _claim_compressor_attempt.

Regression tests drive the real interleavings deterministically,
including two threaded end-to-end arms through _run_summary_dispatch
and a real ContextCompressor.

(cherry picked from commit 902bfc229e140becfb36679dc33bad550c2ae1e8)
2026-09-20 15:50:49 -07:00
teknium1
85564321be fix(windows): route every bare-bash spawn through _find_bash and surface silent interpreter failures
CreateProcess resolves a bare "bash" to System32 WSL launcher before PATH,
and shutil.which("bash") inherits PATH order (#115124), so node bootstrap,
the TUI node probe and webhook filter scripts ran the wrong interpreter on
Windows. All four sites now use tools.environments.local._find_bash (Git Bash
first, probed). rc!=0 with no output at all is now a WARNING in webhook
filters and an explicit [inline-shell exit N with no output] marker in skills.

Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
2026-09-20 15:50:26 -07:00
funky-xamarin
1a24b851e5 fix(skills): resolve native Git Bash for Windows inline shell 2026-09-20 15:50:26 -07:00
beardthelion
275002d7a4 fix(context_references): @folder: listing works outside cwd under a widened allowed_root
@folder: targets resolve against allowed_root, which callers may widen
beyond cwd, but _build_folder_listing and _iter_visible_entries assumed
the resolved folder was under cwd: path.relative_to(cwd) raised
ValueError and the blanket except in _expand_reference surfaced it as a
confusing "not in the subpath of" warning instead of a listing.

The listing header now renders cwd-relative when possible, then
allowed_root-relative, else the absolute path. rg --files gets the
absolute folder path (rg echoes the arg as the output prefix, so the
lines parse correctly anywhere), the parent-dir walk drops only the
cwd stop-condition so in-cwd output is unchanged, and entry indentation
is computed relative to the target folder rather than cwd. The os.walk
fallback never assumed cwd.

Regression tests cover the widened-root target through the real
preprocess_context_references entry on both the rg and rg-blocked
paths, plus @file: parity and in-cwd display controls.
2026-09-20 15:49:22 -07:00
teknium1
30d14a3f73 fix: name the CommandCode upstream-outage pattern
The cherry-pick conflict dropped the contributor comment hunk; keep the WHY next to the entry.
2026-09-20 15:28:21 -07:00
fangliquan
1f3c882d81 fix(agent): narrow upstream outage matching 2026-09-20 15:28:21 -07:00
beardthelion
09b72bc6d2 fix(compression): track commit fences as a registration stack
_compress_context published the active commit fence with a
save/restore cell: registration order was serialized by the fence
lock, but completion order is not. When attempt B registered over A
and A finished first, A's finally popped the slot, deleting B's live
fence mid-attempt (hard_interrupt lost the handle serializing cancel
admission against B's begin_commit). B's finally then republished A's
dead fence, which lingered until the next compression. The same
clobber existed in _publish_new_fence, which overwrote the slot
unconditionally when minting the stall-fallback retry fence.

Replace the cell with a stack of per-attempt registrations. The
finally removes only its own registration and republishes the newest
live entry (or clears the slot), so a dead fence can never be
restored over a live newer attempt. The stall-fallback retry swaps
its fence inside the owning registration and publishes only while
that attempt still holds the top registration. Registration moved
inside the try so an early exception cannot strand an entry.

(cherry picked from commit 574e9945cf186071c3da23c4bb517c0cbdaf73e6)
2026-09-20 15:27:11 -07:00
beardthelion
a48b4c7d25 fix(agent): pop _db_persisted on in-place mutations of stamped live dicts
The _db_persisted marker asserts that a message dict's persisted row is
durable as written; any in-place mutation must pop it or the flush scan
identity-skips the dict and state.db keeps the stale row forever. Six
mutation sites violated the contract:

- micro_compaction._merge_adjacent_user_turns rewrote content on a
  carried-forward dict after superseding a stale micro marker. On the
  archive-failure path nothing re-stamps, so the merged text never
  reached state.db.
- repair_message_sequence passes mutated stamped survivors in place:
  _merge_assistant_into (tool_calls union, content join,
  reasoning_content carry), _prune_unanswered_tool_calls (tool_calls
  rewrite), _merge_consecutive_users (content join).
- sanitize_tool_call_arguments rewrote corrupted/blank
  function.arguments and prepended the corruption marker onto stamped
  resumed rows, leaving the corrupt bytes durable and self-perpetuating
  across resumes.
- _sanitize_messages (surrogate and non-ASCII recovery) and
  _strip_images_from_messages (image-rejection recovery) rewrote live
  dicts on the recovery path.

Each site now pops the marker when a persisted field actually changes,
and the agent-aware callers (repair_message_sequence_with_cursor,
turn_iteration_prep, turn_recovery) invalidate the bounded flush-scan
prefix so repaired rows are rewritten on the next flush.
2026-09-20 15:26:34 -07:00
teknium1
13fe9c7171 feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:

- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
  `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
  declaring profile; standard records still replay on OpenRouter-style routes, strict routes
  drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
  status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
  and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
  one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
  without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".

The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.

Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
efc947d72a fix: send the title model call after the turn on a shared custom endpoint
On a `custom` main route (llama.cpp, Ollama, vLLM, ...) whose
auxiliary.title_generation is not pinned elsewhere, the turn prologue fired the
`response_format: json_schema` title request on a daemon thread at the same
instant as the turn's own streaming request, against the same self-hosted
server. A single-slot server can decode the title grammar/completion into the
main reply: the user then receives `{"title": ...}` as the assistant turn, the
main loop persists it as a genuine assistant row, replays it, and the model
adopts the format (#117296). No Hermes writer routes the aux response into the
transcript; the leaked JSON is the main completion itself.

`maybe_auto_title` now returns the upgrade thread and leaves it UNSTARTED when
`title_upgrade_must_wait_for_turn(main_runtime)`; the prologue parks it on
`agent._deferred_title_upgrade` and `finalize_turn` starts it once the model
has answered. Hosted providers keep the turn-start timing. Usage accounting
(`task='title_generation'`) and `sessions.title` are unchanged.
2026-09-20 14:09:57 -07:00
teknium1
3e579ee7af fix: capped @-reference child output always reports a nonzero returncode
_run_quiet drained each pipe up to _MAX_QUIET_OUTPUT_BYTES and relied on
proc.kill() to make the returncode nonzero. A child that flushed past the
cap and exited 0 before the drain thread crossed it (a fast writer on a
loaded runner) was not killable, so the result came back returncode=0 with
truncated stdout: the caller's fallback path keys on the returncode and
treated the truncation as success. This is also why
test_run_quiet_caps_child_output failed on main's CI (`assert 0 != 0`)
while passing locally.

The drain now records that the cap was crossed and the result is forced to
returncode 137 (128 + SIGKILL) when the child exited 0, so the contract in
the docstring holds regardless of scheduling.
2026-09-20 13:57:43 -07:00
teknium1
81faca2f1c fix: failed initialize keeps the LSP error type instead of raising TypeError (review follow-up)
The failure-details rewrap re-instantiated the caught exception with a single
string; LSPRequestError takes (code, message, data), so a JSON-RPC error to
`initialize` surfaced to log_spawn_failed as a TypeError with none of the exit
status / stderr details. Attach the details to the original exception in place
and re-raise. Adds an `init_error` mock-server script and one invariant test
(red on the previous head).
2026-09-20 13:54:58 -07:00