Manager tokens: lock generation fence (a Lock acknowledged while `bw unlock`
/ `op signin` is still running discards the late token); tokens record the
unlocking gateway session and are released when THAT session ends, not when
any sibling session in the profile is torn down.
1Password: OP_CONNECT_HOST/TOKEN come from the profile's scoped secret store
like the service token (Connect outranks a service token inside op), never
from the launch environment.
Vault RPCs bind params.profile (home + secret scope) so a shared remote
backend serving several profiles locks/lists/unlocks the requested one;
unknown profile → RPC error, not a crash.
Fill target: inspection stamps are `<nonce>:<index>`; a fill resolves only
its own inspection's stamps, so an interleaved second inspection can no
longer redirect A's password into a newly mounted field (real Chrome: 0
filled, both fields empty).
Desktop Settings: every RPC goes through the owner profile's socket
(requestGatewayForProfile), query keys carry (connection, profile), an owner
change closes dialogs and wipes drafts (a master password typed for A is
never submitted to B; a late list from A never paints under B), and vault.add
secrets travel in a ref consumed by the mutationFn instead of mutation
variables. Three owner-routing invariant tests on the real component.
Docs/PR body: session-scoped release, lock-race semantics, bw --passwordenv.
Bitwarden unlock now uses the CLI's documented non-interactive channel:
`bw unlock --raw --nointeraction --passwordenv VAR`, VAR set on the child
environment only (bw 2026.x rejects a piped password with "Master password
is required"). Verified against the real published binary.
Manager session tokens are keyed by (profile home, backend): a Desktop
gateway hosting several profiles can no longer reuse or lock another
profile's session. Status probes (`vault.sources`, is_unlocked) no longer
refresh the idle TTL; only real manager calls do. Gateway session teardown
locks the profile's managers (a per-session unlock ends with the session).
1Password service-account token comes from the profile-scoped secret store
(get_secret), not ambient os.environ.
`vault.source.set` no longer references a module constant (bind_module
rebinding dropped it → NameError on every Settings toggle).
Fill target binding: inspection stamps each input with a per-inspection
slot attribute; the fill resolves by stamp and requires type=password, then
strips every stamp. A DOM reflow between inspect and fill can no longer
redirect the password into a text field (reproduced in real Chrome before,
0 filled after).
Redaction boundary: no 4-char floor, CR/LF-normalized form registered
(what a text input actually stores), JSON object KEYS scrubbed in both
browser redactors; longest value first. Docs now state the real trust
model: accidental-disclosure protection, not an execution sandbox.
Desktop: the mid-turn card sends the master password through the owning
session's socket (requestForOwnedSession), never the ambient foreground
gateway; `vault.unlock.expire` clears a stale card; Settings keeps the
master password out of react-query mutation variables (ref consumed by the
mutationFn). One renderer invariant test for the routing.
Live Desktop repro: the fixture model called browser_vault_unlock and got
unlock_unavailable although the renderer was interactive. tool_executor runs
handlers on a propagated worker thread; thread_context only copied the
approval and sudo thread-local callbacks, so the unlock prompt registered by
_wire_callbacks was invisible there and can_prompt_here() said nobody could
answer. The callback table in thread_context now lists every per-thread
prompt (approval, sudo, vault unlock) so a new one cannot silently drop off
worker threads again. After the fix the same turn shows the masked card and
completes.
vault.sources / `hermes vault sources` report a manager as installed when
its configured binary_path exists, not only when it is on PATH.
The browser vault now draws from three login sources behind one handle
shape: the local encrypted vault (vault_…), 1Password Login items (op:…)
and Bitwarden Password Manager logins (bw:…). browser_vault_list aggregates
metadata across them; browser_vault_fill routes by prefix and resolves the
password at fill time only, through the manager CLI.
External managers are locked until the user unlocks them for the current
session. The new browser_vault_unlock tool (and the fill path, implicitly)
asks the surface to show a masked master-password prompt — CLI panel
(reuses the sudo panel state), TUI/Desktop via a vault.unlock.request
blocking card. The password goes to `op signin --raw` / `bw unlock --raw`
on stdin, never argv or env; only the session token is kept, in memory,
with a 30-minute idle TTL, cleared on session close or `vault.lock`.
Headless contexts (cron, webhook, api_server, -q) can never prompt: the
manager is reported as locked with unlock=unavailable_in_this_session and
fill refuses — the same posture approvals take where nobody can answer.
Config: vault.onepassword / vault.bitwarden {enabled, binary_path, …};
a 1Password service-account token skips the prompt for headless use.
RPC: vault.sources, vault.source.set, vault.unlock, vault.lock for Settings.
Tests (2, real subprocess against a fake bw; each proven red by sabotage):
headless never prompts or spawns; unlock feeds stdin only, token never
enters os.environ, fill routes by prefix and the password only reaches the
fill script.
Consolidated re-apply of #96988 onto current main. Ported from
Merit-Systems/OpenInstinct (MIT) opaque-handle autofill design: the model
sees vault handles + login metadata, the password is resolved and filled
server-side over the supervised CDP socket, and filled values are scrubbed
from every browser tool result by an unconditional redaction registry.
Rebase adaptations to the Sep-2026 facade/sibling layout:
- toolsets: one _HERMES_CORE_TOOLS entry (the browser toolset derives from it)
- hermes_cli/main.py: vault parser registered via the subcommand owner table
- file_safety: vault/ joins the _READ_DENIED_DIRS credential-dir table
- redact: registry scrub runs before the redact_secrets early-return
- browser_vault_tool: _run_browser_command now lives in browser_tool_session
The facade forwarded turn_author= to the loop's public entry point, which did not
declare it: every real AIAgent.run_conversation() turn raised TypeError (16 CI
failures across provider, sidecar, cron and finite-chat suites). The PR's tests
only exercised build_turn_context directly, so the missing hop was invisible.
Adds one facade-through-loop test that goes red when the kwarg is dropped.
Keep one or two behaviour tests per seam (author reset on a cached agent, forged
_turn_author refused, guard trips and cools, single charge on the busy path) and
drop the parser/setting enumerations. a2a_key goes with them: nothing in this PR
reads it; the honcho follow-up that does can bring it back with its consumer.
message_agent sent a bare bot:<profile> id through hermes peer dm, so a remote coder and the recipient's own coder shared one author id. The peer branch now sends bot:<hostname>/<profile>, with the hostname cleaned like any author field and slashes dropped. The direct local path keeps the bare id.
A relayed DM stamped bot:<profile> on the recipient turn, so an ops profile on another machine and the local ops profile shared one author id. The Desktop now forwards from_connection with each bot_relay.deliver, and delivery_turn_author builds bot:<connection>/<profile> for it while the Desktop's own gateway ("local") keeps the bare id. An api author object accepts an optional origin string that yields the same shape.
MemoryManager.on_turn_start forwarded the author kwargs to every provider and _each_provider swallowed the TypeError, so a provider with the two-positional on_turn_start(n, text) stopped running. The kwargs are now filtered against the provider's signature the way _provider_sync_accepts filters sync_turn.
The Telegram adapter asks the authorization check before dispatch, the ingress gate asks it
again, and the busy path asks a third time. Each call counted one loop-guard event, so a
Telegram bot tripped the budget after a third of the configured messages. The verdict now only
refuses a chat that is cooling down. The ingress gate counts an admitted bot message once.
`parse_turn_author` treats only booleans, integers and the strings true/1/yes as a bot flag,
and returns None for an author with neither id nor name. Names keep format characters and
non-breaking spaces so emoji sequences survive. The quiet one-shot pops HERMES_TURN_AUTHOR
before the turn so tool subprocesses do not inherit it. `max_events` must be a whole positive
number. Issue numbers move out of code comments.
on_turn_start already received the author trio. sync_turn did not, so a
provider that wanted to write the turn under its author had to stash state
between the two hooks. sync_turn now takes turn_author as a keyword-only
argument, and MemoryManager sends it only to providers whose signature
accepts it, so existing providers keep working unchanged.
build_turn_context resets the author on the agent at the start of every turn
so a cached gateway agent never carries a bot author into the next human
turn. agent/turn_author.py holds the parsing and the HERMES_TURN_AUTHOR
carrier.
MemoryProvider.identity_signature() is a new optional hook: the identity
values a provider writes under, declared by the provider itself, for the
gateway's agent cache to key on.
`on_turn_start` documents a per-turn kwargs channel — "kwargs may include:
remaining_tokens, model, platform, tool_count" — and `MemoryManager`
forwards whatever it receives. Its only caller passed nothing, so a memory
provider had no way to learn who wrote the turn it was being told about.
Providers that key durable state on identity resolve one identity when the
session is created. A shared session does not work that way: threads are
shared by default (`thread_sessions_per_user` is False), so alice, bob, and
another agent all write turns into a session whose peer is whoever spoke
first. The gateway's answer today is the `[name]` prefix it prepends to the
message text, which the model reads and a provider cannot.
`turn_author` now travels from the gateway through `run_conversation` into
`build_turn_context`, which forwards `author_id`, `author_name`, and
`author_is_bot` to every provider. It stops there — the trio never reaches
the model, and providers that ignore the kwargs are unaffected.
The bot flag is sent on every transport, not only shared sessions: a
provider deciding whether a turn may write to durable memory needs it in a
DM too.
`SessionSource.is_bot` is only as good as its producers. `build_source`
defaults it to False and 3 of 32 adapter call sites pass it, so most
platforms still report every author as human. Populating the rest is
follow-up work; nothing here depends on the flag being right yet.
Add deepseek/deepseek-v4.1-flash to OPENROUTER_MODELS (Nous list derives from it),
regenerate the docs manifest, and give the slug its own 1M context entry and 600s
reasoning-stale floor — the longest-key-first scan otherwise lands the new slug on
the 128K `deepseek` catch-all and no floor. Live probed on both routes: echoed
model matches, usage.cost billed.
Three defects found in review of the first head (@ehz0ah):
* The mid-request hop read the PERSISTED provider from disk. A live
`/model xai-oauth` session over `model.provider: auto` therefore still
reached the discovery chain and billed Nous. _try_payment_fallback now
takes the route's main_runtime snapshot; disk is the fallback only when
no session runtime exists.
* After a configured fallback was quarantined mid-request (401, refresh
failed), the second pass went straight to the discovery chain, which
the new gate refuses — so later CONFIGURED entries never ran and the
original error was re-raised. The second pass now re-walks the task
chain and main chain (the quarantined entry is unhealthy and skipped)
before discovery.
* current_provider_owns_vendor dropped ids detect_vendor could not
classify, so Bedrock's 15-id catalog (14 unclassified `us.anthropic…`)
looked exclusively DeepSeek and `/model deepseek-v4-pro` stuck on
Bedrock. An unclassified id now counts as evidence of a multi-vendor
catalog: ownership requires every id to classify to the one vendor.
With a main provider selected, an unusable main route (expired xAI/Codex
OAuth token, 401/402/429 mid-session) fell through the built-in discovery
chain (OpenRouter -> Nous -> custom -> api-key) and quietly ran every
compression, title and memory-flush call on whichever OTHER account was
still logged in. Reported as "using Grok on my Premium+ sub, my Nous
Portal balance kept draining" — the chat visibly stayed on Grok while the
side tasks were billed elsewhere, and re-logging into X did not help
because the aux side never consulted the selected provider.
The discovery chain is now reserved for installs with no selected main
provider (`model.provider: auto` / unset). Otherwise the ladder is
main -> auxiliary.<task>.fallback_chain -> fallback_providers -> refuse
with a warning naming the dead provider and the fix. Both entry points
gate on the same predicate: the resolve-time route and the mid-request
payment/auth hop (_try_payment_fallback).
Existing chain tests that asserted the hop now pin `provider=auto`, the
one case where discovery is still the contract.
The publish-side taxonomy (_end_stamp_class, _compression_parent_obstacle,
compression_parent_deliberately_ended) and the agent-guard delegation produced
exactly main's verdict -- every non-automatic stamp still fails closed -- so
they only changed an error string. Inline the "explicit close with no
continuation" test into reopen_if_explicitly_closed(), restore main's publish
branch and agent guard untouched, and keep two tests: the field shape rotates
after the host clears the stamp; boundary/compression/automatic stamps and a
session already claimed for teardown are never cleared.
Consolidated from PR #106543 (5 commits, final tree d2c4d908) by @Totoro-qaq.
publish_compression_child() fails closed on any non-automatic end stamp and
end_session() is first-stamp-wins, so a stale tui_close on a session the TUI
still routes turned every rotation into "compute the summary, then discard it"
(#106459). The host that still routes the session clears the stale explicit
close via SessionDB.reopen_if_explicitly_closed() before the turn starts;
publication never heals explicit closes. Review probes by @ehz0ah.
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.
Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.
Builds on YipTszkwan's #107126 (earliest fix in the cluster).
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:
* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
omitted `extra_body.thinking`. The server then defaults to thinking-on, so
the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
V-series regex), so the id a user picked never reached the wire and the
config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
instead of the real 1M window, capping the model at an eighth of its
context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
detector at its 180s default instead of the 600s reasoning-model floor.
Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".
Adds the id to all four gates plus regression coverage for each site.
Every GLM slug that missed a DEFAULT_CONTEXT_LENGTHS key fell to the "glm" 202,752 catch-all, so
compression fired at ~20% of the real window and the aux feasibility check auto-lowered the session
threshold to that number (the 202,752 in the Coatue compression report).
- Catalog: GLM entries from Nous + OpenRouter /v1/models (2026-09-09): 5.3 / 5.3-flash 1,310,720
(:batch/:US 1,048,576), 5.2 1,048,576, 5 / 5.1 / 4.7 / 4.6 204,800, *-turbo / 4.7-flash 202,752.
- _catalog_key_matches: version separators normalised on both sides so relay slugs like z-ai-glm-5-3
hit glm-5.3 (#97398; approach from PR #97412 by @shellybotmoyer, re-implemented on current main).
- get_custom_provider_context_length: entry-level custom_providers[].context_length backs every model the
entry serves when no per-model override exists, so /model switch stops dropping it (#98387; from
PR #98396 by @liuhao1024, re-implemented on the config_providers sibling).
- check_compression_model_feasibility: when aux compression is the main model on the main route, reuse the
main model's resolved window instead of re-resolving without its pin (#89500 mechanism 2, #45519).
Closes#97398, #98387, #89500, #87825, #97595, #97820 (suffix strip already on main; catalog values now match).
A rejected OAuth token (HTTP 401 'User not found' from Nous Portal, Codex,
xAI…) reached the desktop error card as 'Provider error' with Retry as the
first action, which just replays the same dead credential.
Backend: nonretryable_client_error_result dropped failure_reason /
failure_retryable, so error_surface classified every non-retryable 4xx as a
retryable provider failure. It now stamps the classifier verdict like the
max-retries path, and auth-layer descriptors carry auth_kind
(oauth|api_key, derived from the provider catalog tab) + provider_label.
Desktop: an auth/oauth surface renders 'Authentication error', explains that
the <provider> sign-in expired/was revoked, and offers 'Sign in to <provider>
again' which launches that provider's existing onboarding OAuth flow scoped
to the failed session's gateway profile. Retry stays as the follow-up click.
Re-login to the provider already in use keeps the current model instead of
swapping in the recommended default.
Follow-up to #106866 from the post-merge audits (@ehz0ah, @victor-kyriazakos):
- _recoverable_pool_provider: the "endpoint mismatch, not a dead key" shield now (a) compares the full
origin via base_url_origin() (scheme+host+port — a different port or an HTTPS->HTTP downgrade is a
different trust boundary) and (b) applies only when the rejected client carried the SESSION's key, so
an independently owned auxiliary pool keeps rotating at its own configured origin.
- _resolve_named_custom_branch: explicit/per-task base_url and api_key compose OVER the providers.<name>
entry field-by-field instead of the entry silently replacing them (URL-only, key-only, both, neither).
- _acreate_with_progress: plain-create fallback only for a rejected stream NEGOTIATION (mirrors the sync
wrapper); a mid-stream failure propagates to the classified recovery ladder instead of re-sending the
whole prompt non-streaming.
Tests (all red on main): foreign origin x {host, port, scheme} shields the session key; same origin and
independent aux pool still rotate; named-provider override matrix through resolve_provider_client on a
real temp HERMES_HOME; async content-then-error is not retried non-streaming.
`auxiliary.<task>.provider: openai` was rewritten to custom + api.openai.com/v1 unconditionally, and the
api-key discovery rung used the registry default endpoint, so proxy/gateway users (OPENAI_BASE_URL or
providers.openai) had compression hop to the public endpoint with a proxy-issued key -> 401 -> the
key/pool was quarantined and the session sat over the compression threshold.
- _expand_direct_api_alias: a providers.openai entry keeps its name (named-custom branch applies its
base_url/key); otherwise OPENAI_BASE_URL wins over the public default.
- _resolve_api_key_provider: when the bound session runtime is that provider, use its endpoint + key.
- _recoverable_pool_provider: a rejection at a host other than the session's configured endpoint for
the same provider is an endpoint mismatch, not a dead key -> no rotation / unhealthy mark on the pool.
Complete the seam default from the salvaged commit: _relay_async_completion needs an async twin of
_create_with_progress, and the async primary attempt previously streamed only for stream-only
providers, so a hooked async compression call ticked the watchdog zero times.
Adds one invariant test: with a hook installed, BOTH relay defaults stream and tick per chunk;
without a hook they are byte-identical plain creates. Red on origin/main.
* feat: add session-scoped connector access for onboarding
* fix(connectors): availability is the config flag AND the portal entitlement — no free-tier leg
The port carried a third availability leg from hermes-magic: a stored guest
(free-tier) identity short-circuits the managed-tool entitlement check. That
leg reads hermes_cli.anon_auth, which does not exist on hermes-agent main, so
connectors_available() raised ImportError inside its fail-closed try and the
whole connector surface was silently dark on a plain upstream checkout.
On this tree availability is the two-leg AND the design started with:
tools.connectors.enabled AND managed_nous_tools_enabled(). The free-tier leg
is a hermes-magic concern and belongs in hermes-magic's own delta over this
branch, next to the identity it depends on. Its integration test goes with it.
* docs(tool-search): connectors section — remote tools through the bridge
The squashed port carried the code but not the user-facing docs. Restores the
Connectors section of the Tool Search page and the connector-gateway host /
CONNECTOR_GATEWAY_URL override on the Tool Gateway page, updated for the
manage_connections tool and the pure-connector batch rule.
* fix(tool-search): connector tools rank with local tools in one pass instead of taking leftover slots
dispatch_tool_search ran BM25 over the local catalog, filled `limit` slots,
then appended connector hits only into slots left empty. On a 300-tool
catalog no slot was ever empty, so with Gmail and Google Calendar connected
"send gmail email" returned five betterstack tools and zero connector tools.
The gateway's hits for a query now become catalog entries (connector name,
slug words, description as the search text) and join the local catalog for
that query's BM25 pass. One ranking, one rarest-token admission rule for both
sources, `limit` as the total per query. The merge loop and the separate
record builder for connector hits are gone; `_shared_tool_record` serves both
sources.
The gateway search timeout rises from 8 s to 30 s. One request with six
use_cases measured 7 s, so 8 s sat on the edge and cut real answers off; the
failure path is unchanged (local-only results, no error to the model).
Live, 311 local tools + gateway, before -> after:
"send gmail email": 5 betterstack tools -> gmail SEND_EMAIL, CREATE_EMAIL_DRAFT
"read google calendar events": 5 betterstack tools -> googlecalendar EVENTS_LIST_ALL_CALENDARS
"linear create issue", "betterstack incident": unchanged
Benchmark (25 labelled queries): connector recall 0.09 -> 0.82, precision@5
0.18 -> 0.59, false positives on absent intents 17 -> 2.
* refactor(tool-search): connector leg into tools/connector_search.py
tools/tool_search.py is a facade. The connector leg (gateway hits as catalog
entries for tool_search, remote schemas for tool_describe, the
connections_in_scope gate) was appended to it by the port. It now lives in
its own sibling, tools/connector_search.py, and the facade imports the three
entry points: connections_in_scope, connector_entries_by_group,
remote_schemas_for.
No behaviour change. The tool_describe remote block became
remote_schemas_for(names, current_tool_defs, connector_describe) with the
same inputs, the same silent-degradation contract and the same injection
seam the tests already use.
* fix(tool-search): at most 7 queries per call, the gateway's search limit
One tool_search call sends all its queries to the connector gateway as one
search request. The gateway answers 7 use_cases per request and returns
HTTP 502 for 8 or more (measured 2026-09-09, re-measured with one-word
use_cases: it is a count limit, not a size limit). With the client cap at
10, a model sending 8 to 10 queries lost every connector hit for that call
and saw local-only results with no error.
The shared constant splits: _MAX_QUERIES_PER_CALL = 7 for search,
_MAX_DESCRIBE_NAMES_PER_CALL = 10 for describe, which has no remote count
limit. Eight or more queries now get the existing "too many queries" retry
hint before any request is made. No chunking: one call, one request.
* fix(tool-search): the model is told that connectors__ names are manage_connections accounts
tool_search results carry names like connectors__gmail__CREATE_EMAIL_DRAFT and
manage_connections is the tool that checks and connects those accounts, but
nothing told the model the two are the same thing. A model that hit
CONNECTION_REQUIRED had to infer the fix on its own.
The tool_search description gains one sentence making the link, added at
assembly only when manage_connections is in the session's tools. Signed out
or with connectors off the tool is absent and the description is unchanged,
so it never names a tool the model cannot call. This follows the existing
rule for cross-tool references (tools/AGENTS.md): they are added dynamically
from the session's actual tool set, never hardcoded in a schema.
Tool defs are fixed for the life of a conversation, so the description is
byte-stable per conversation; this is a one-time prefix change.
Live, real get_tool_definitions() against a signed-in home: sentence present.
Same home with auth.json removed: manage_connections absent, sentence absent.
* fix(connectors): /stop halts a connector batch before the next remote call
dispatch_connector_batch runs every remote entry of a tool_call batch in
sequence. The executor only checks the interrupt flag between tools, and
the whole batch is one tool to it, so a /stop landing during entry 1 of
20 still sent the other 19 to the gateway.
The loop now reads tools.interrupt.is_interrupted before each dispatch.
Once set, it stops calling handle_function_call and fills every unstarted
slot with the loop's existing error-slot shape, code INTERRUPTED and the
message "Stopped by the user before this call was made.", so the result
envelope stays valid and the counts stay honest. Entries already
dispatched keep their real results.
Test: three connector calls where the fake client sets the interrupt on
the first execute. The client sees exactly one call and slots 2 and 3
carry INTERRUPTED. Red on the base branch, green with the fix.
* test(connections): schema assertions become dispatch contracts
test_schema_documents_wait_and_its_timeout froze description fragments
("REQUIRED", "can NOT disconnect", "Nous Portal"). A wording edit fails
it while a real regression (a disconnect that reaches the gateway) does
not. That is a snapshot of prose, not a behaviour contract.
Delete it. The requirement that wait needs connectors is already covered
by test_wait_requires_connectors. The user-only disconnect boundary is
now asserted as behaviour: action disconnect with a connector returns an
error and the fake client records no call. That replaces the earlier
de-authenticate test, which only checked that the word "dashboard"
appeared in the error text.
Test count in the file goes from 26 to 25.
* docs(tool-search): connector batches are one gateway request per entry
The user guide said a connector batch travels as one gateway request. It
does not: model_tools_connectors.dispatch_connector_batch re-enters core
dispatch per entry, and each entry becomes its own execute request in
bridge._run_remote (plus at most one literal-slug retry when the gateway
reports TOOL_NOT_FOUND under the conventional slug). The docstrings in
tools/tool_gateway/bridge.py and tools/tool_gateway/__init__.py still
described the abandoned V1 plan and claimed nothing outside the package
imports it.
Rewrite those sentences to match the code: one request per entry, in
input order, dispatched from model_tools_connectors.py, with the per-entry
approval and interrupt behaviour that motivated the split. The guide also
still showed the single-call shape tool_call(name, arguments); both
places now show the `calls: [{name, arguments}]` array the schema
advertises and note that a single local call is an array of one.
Docs only, no test.
* fix(tools): the between-turns refresh never rewrites the bridge tools
The per-turn MCP refresh folds a fresh tool snapshot into the live array
with preserve_prefix: order and membership stay, but a name present in both
takes the fresh schema. That is right for ordinary tools, whose schema is a
constant. tool_search is the one tool whose description is derived from the
session: the deferred-tool count, the embedded listing, and, on this branch,
whether manage_connections was present. A late MCP server or one failed
portal lookup (manage_connections' check_fn fails closed) changed those bytes
on the next turn, and every byte after tool_search in the cached prefix was
re-prefilled. The array also contradicted itself in that case: the flapping
manage_connections was carried forward while the description lost its hint.
The bridge entries now keep the bytes they were built with for the life of
the conversation. Nothing is lost: tool_search reads the live catalog at
dispatch, so tools that arrived late are still found; connector availability
is checked at dispatch too. The compaction-boundary rebuild (content_aware,
the one sanctioned cache break) still refreshes the description.
Consequence: connector exposure in the prompt is decided once, at agent
build, by whether the user was signed in then. That is the intended
contract.
* refactor(tool-search): normalize_tool_call_entries lives with the other argument validation
The port appended the tool_call argument parser to the tool_search facade.
The family already has tools/tool_search_validation.py for exactly this
work (schema validation of deferred call arguments), so the parser moves
there and the facade imports it. No behaviour change; the one test that
imported it now imports from the defining module.
* refactor(connectors): delete the unused batch dispatcher; _run_remote becomes run_remote
bridge.dispatch_calls and its helpers (_dispatch_calls_inner, _run_pre_dispatch,
_run_local, _error_slot, _maybe_parse_json) and the LocalDispatch / PreDispatch
seams had no production caller. Connector dispatch runs through
model_tools_connectors: dispatch_connector_batch re-enters handle_function_call
once per entry, so scope, hook, approval and middleware policy fire against each
composed name inside core dispatch, and dispatch_connector_call hands the single
planned entry to the bridge's transport function. Only tests called the batch
dispatcher, and they exercised policy seams that production never wires.
The transport function is the module's real entry point, so it drops the
underscore: _run_remote becomes run_remote, body unchanged. The module
docstring now describes the two legs that exist (availability with D32 silent
degradation, and run_remote) instead of the injected seams. Imports that only
the deleted code used are gone; merge.py is untouched because every export
still has a caller.
Tests that drove dispatch_calls are deleted where they covered the removed
seams (pre_dispatch blocks and rewrites, local_dispatch classification, mixed
batches). The literal-slug fallback, the per-entry transport failure, and the
hook rewrite reaching the gateway request body are re-targeted at
handle_function_call('tool_call', ...) with the fake client swapped in at
bridge._default_client_factory, the same seam test_connector_dispatch_policy
uses. Each re-targeted test fails when the retry is disabled in run_remote.
* fix(connectors): search keeps the twin a colliding name reaches, and says so
format_connector_name strips the toolkit prefix, so GMAIL_FETCH_PROFILE and a
literal FETCH_PROFILE on gmail both compose to connectors__gmail__FETCH_PROFILE.
describe and execute decode that name to the prefixed slug first, so the
literal twin is unreachable under it. If a vendor ever shipped both, search
could describe the literal under a name that runs the prefixed tool.
Search is the one place that sees both twins in one response. It now keeps
the twin the name reaches and drops the other with a WARNING that names both
slugs, whichever the gateway listed first. Short names stay; no marker, no
per-process map, no change to describe or execute. No such pair exists in the
live catalog today; the guard turns a silent alias into a logged one.
_is_azure_foundry_responses matches the azure-foundry provider or the
project-scoped services.ai.azure.com host. #101243 proposes narrowing it
to that host alone. The multi-item rejection this branch handles happens
on resource-level *.openai.azure.com hosts too, so the pruning gate now
uses _is_azure_responses: the provider, or either Azure host.
Refs #105369
The first request of every turn after the first replays the encrypted
reasoning item of every earlier response. Azure Foundry accepts one such
item and rejects two or more with HTTP 400 "Conflicting authenticated
continuation identities". In-turn requests never hit this because
_is_post_tool_replay already suppresses replay on Azure when the request
ends on a tool result.
On Azure Foundry, build_kwargs now keeps codex_reasoning_items only on the
newest assistant row that has any. include still requests
reasoning.encrypted_content, compaction checkpoints stay on every row, and
the canonical messages are not mutated. Other Responses endpoints are
unchanged.
Rebuilding a failed session from state.db through the transport and
sending it live: unpatched, 5 reasoning items, 400; patched, 1 item, 200.
Fixes#105369
Azure Foundry (gpt-6-astra) answers a request that replays encrypted
reasoning from more than one prior response with HTTP 400 "Conflicting
authenticated continuation identities", code invalid_value. The classifier
mapped that to a non-retryable client error, so the turn aborted and every
later turn in the session failed the same way.
Map the message to invalid_encrypted_content. The existing recovery in
turn_recovery then strips the replay state and retries once.
Refs #105369
The classifier fix routes Codex patch-budget 400s into the shrink
recovery, but _image_error_max_dimension still returned None for the
Codex wording, so the recovery fell back to the 8000 px default cap
and skipped images between ~5542 and 8000 px that already exceed the
30000-tile budget — burning the single shrink retry without shrinking
anything. Parse the reported patch limit and convert it to a per-side
pixel cap of isqrt(limit)*32 (5536 px for 30000 patches), which keeps
a square image under the budget.
OpenAI Codex Responses rejects an image whose tile-patch budget exceeds
its 30000-patch ceiling with wording ("requires N patches after
processing, exceeding the limit") that contains none of the existing
image-size vocabulary, so the 400 fell through to format_error
(non-retryable). The reactive image-shrink recovery in turn_recovery was
therefore bypassed and the session kept failing — or failover re-sent
the identical oversized image to another model.
Route the patch-budget wording to image_too_large (retryable) so the
existing shrink pass re-encodes the image under the ceiling and retries.
Fixes#106337
finish_reason='length' has two causes: the answer was long (max_tokens reached),
or the prompt itself left no room to generate. _continue_text treated both the
same: append the fragment + a continuation nudge and retry, up to 4 times. In the
second case every retry sends a strictly longer prompt, so each attempt is worse
(Ollama n_ctx=32768: 32,638 -> 32,685 -> 32,732 prompt tokens, all truncated), the
user is told "model hit max output tokens", and max_tokens is not the lever.
The response's usage already carries prompt_tokens and the compressor already
resolves the model's context window; compare them once per truncation. Under
_MIN_CONTINUATION_HEADROOM (512) free tokens the turn ends on the first
truncation, keeps the partial text, names the context window as the cause and
points at /compress or a larger window. Unknown usage or window keeps today's
behaviour. max_tokens semantics untouched.
Co-authored-by: gaoanze888 <214786078+gaoanze888@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The two facades post-#102117 (chat_completion_helpers, agent_runtime_helpers) must not grow
behaviour; the marker/predicate belong with the sibling that already owns primary-cooldown
state shared by the fallback walk and restore_primary_runtime. Also drops the two
try/except-return-False wrappers around pool.entries() and normalize_model_for_provider
(neither raises for a live pool / the never-raising normalizer).
Tests trimmed to the two invariants (#106475): the marker fires only on the single-credential
Codex entitlement 400 and the walk skips the rejected slug in either form; restore is gated on
a rejected primary and still restores an unrelated one.
A Codex ChatGPT-account 400 ('The X model is not supported when using Codex
with a ChatGPT account.') names the model, so with a single credential the
slug is dead for that account. The fallback walk still re-selected it and
restore_primary_runtime switched back to the primary at the start of every
turn, announcing an unverified 'Primary model restored' — the two warnings
alternated forever with zero delivered answers (#106475).
Record the rejected (provider, model) pair on the non-retryable client-error
path (only when no multi-credential pool exists — rotation covers that case,
#71970), skip rejected entries during the fallback walk, and gate
restore_primary_runtime on the primary's slug so the session fails closed
with the terminal entitlement error instead of oscillating. Fixes#106475.
(cherry picked from commit 471435b10288f15387b2549a8b02551d82c0f670)
Hermes Console ran each command on a ThreadPoolExecutor and cancelled only
the asyncio waiter. A command that forks an AIAgent (`curator run
--consolidate`) kept its worker thread and its in-flight provider request
alive after the prompt said "cancelled" — a llama.cpp generation kept
decoding for 30+ minutes and held the inference slot (#106179).
Root cause: the host owning the thread never knew about the agent created
deep inside the synchronous command, so it could not call the existing
cooperative `interrupt()` path that closes the request sockets.
Fix: `agent/interrupt_scope.py` gives the host an `InterruptScope`; the
console binds it around the worker (ContextVar), and every
`AIAgent.run_conversation()` registers itself with the bound scope for the
turn. On cancel, timeout and disconnect the console calls `scope.cancel()`
(hard-interrupts every registered agent; an agent registering after the
cancel is interrupted on entry so the cancel cannot lose the race with a
turn that has not started) and awaits the worker Future with a 10s bound
before reporting cancelled/timeout. Queued-but-unstarted work is dropped via
`Future.cancel()` alone.
Live repro (fake OpenAI-compatible provider blocking like llama.cpp, real
/api/console, real curator dispatch, real AIAgent + direct request path):
origin/main at the "cancelled" frame -> request_exited=false,
worker_exited=false; with this change -> both true, provider saw the peer
close, prompt reported cancelled 0.27s after the frame.
Salvage of #106197 by @kyssta-exe (executor-future handle, cancel/timeout/
disconnect propagation) and #106320 by @Xixiartemis (deterministic
lifecycle regression: terminal(cancelled) => no owned request or worker
remains live; interrupt-on-late-registration). Both rebuilt slimmer: #106197 keyed its fallback on Future.cancel()
returning False, but the handle it held was run_in_executor's asyncio
wrapper, whose cancel() returns True while the thread keeps running, so its
thread-name abort registry was never consulted; #106320's command-scoped
ownership model is folded into one small module hooked at the turn facade
instead of a per-caller `bind_agent`.
Co-authored-by: kyssta-exe <218078013+kyssta-exe@users.noreply.github.com>
Co-authored-by: Xixiartemis <182932319+Xixiartemis@users.noreply.github.com>
Novita returns HTTP 429 with message 'server overload, please try again
later' and error type 'server_overload' when its server is genuinely busy.
Neither phrase was in _OVERLOADED_PATTERNS, so the 429 fell through to
_V_RATE_LIMIT and set should_fallback=True + should_rotate_credential=True —
rotating the credential / falling back early instead of retrying the same
key. Add 'server overload' and 'server_overload' to the overload tuple so
this reaches the existing _V_OVERLOADED verdict (retryable, no rotation).
Closes#106205
An aux task pinned to a private endpoint via `base_url` (a home Ollama
box at `byron.local`, a LAN IP, localhost) is the intended per-task
endpoint feature and can never bill a provider. The Settings → Model
banner still counted it as "still run on openai" forever and offered
"Reset all to main", which would wipe the working local setup; the
post-switch `stale_aux` report had the same blind spot; and the aux row
never showed the `base_url` the backend already sends, so the pin was
indistinguishable from a paid-provider pin.
- `GET /api/model/auxiliary` now stamps each task with `local_endpoint`,
the verdict of the one canonical classifier
(`agent/model_metadata.py::is_local_endpoint`) — no TS mirror of the
private-range rules, so frontend and runtime cannot drift.
- Desktop: the persistent banner filter is the pure
`staleAuxAssignments()` and skips `local_endpoint` pins; the pinned row
appends ` · <base_url>` when one is set.
- `_stale_aux_pins` (post-switch report) skips local pins the same way.
- `is_local_endpoint`: `*.local` (RFC 6762 mDNS) now counts as local, and
IPv6 literals no longer ride the "no dots ⇒ unqualified host" rule, so
a global-scope address (`2607:f8b0::1`) is not local while `::1`,
ULA and link-local still are via the `ipaddress` scope checks.
Slim redo of #106236 (@webtecnica) and #106234 (@huklaa), which fixed the
same symptom with a client-side classifier copy; the bug class, row
display and mDNS/IPv6 classifier corrections are theirs.
Refs #106228
Co-authored-by: Hukla <129692708+huklaa@users.noreply.github.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
When the redirect cap trips, the correction that cancelled the final attempt is
still sitting in _pending_redirect; finalize_turn's clear_interrupt() would drop
it silently. Drain it into the steer slot so it rides result["pending_steer"] and
becomes the next user turn on every surface that already honours that key.
Both new exit reasons get a turn-completion explanation so the user sees why the
turn stopped instead of an empty reply.
The redirect and rebuilt-for-fallback restart paths in apply_retry_restarts
refund the iteration budget and re-issue the iteration with no per-turn
bound. A redirect/interrupt that keeps re-arming the flag refunds forever,
so the turn loop never exits and the durable session turn lease is held
indefinitely (concurrent processes block up to LEASE_WAIT_SECONDS).
Add a per-turn restart_count accumulator (threaded through _run_phase like
the other loop locals) and break out once it exceeds max_retries, matching
the bound the compression path already has.
_sum_clarify only extracted the top-level ``user_response`` key, so batch clarify
results (questions=[...] -> responses[].user_response) fell through to the generic
placeholder and the summarizer never saw the user's answer/permission decision.
Closes#106077.
The overflow-terminal path ends the turn without reaching finalize_turn, so
a transcript that overflowed right after a tool batch ended on a raw tool
result; strict providers reject the next user turn (tool -> user). Close it
with the same final text, mirroring the truncated-tool-call terminal above.
Also: classify once before either log so an overflow no longer emits a
"so the loop can continue" WARNING followed by the contradicting "NOT
seeding" one; reset the stale-streak breaker once for both branches; drop
the "compression could not recover it" wording (this path never reached
compression); trim the test file to the three tests that bind behaviour
(stream -> terminal stub; 413 stays non-terminal; the terminal ends the
turn, closes the tool tail, carries compression_exhausted).
Address review P1s (andrexibiza) on #106266:
1. The overflow-terminal exit in recover_from_truncation now forwards the
#98722 typed compression_exhausted bit (partial_result/end_turn gained the
flag) so the gateway resets/moves future input to a clean session instead
of leaving the bloated durable session authoritative for the next turn.
2. _overflow_terminal is scoped to FailoverReason.context_overflow ONLY.
payload_too_large (413) has its own byte-scored recovery owner
(turn_overflow._recover_payload_too_large, #88960/#47339) that must not be
bypassed; a post-delta 413 keeps its normal continuation stub. Regression
covers both lanes.
Tests: unit asserts result compression_exhausted=True on the marker; a
real streamed partial hitting a 413 payload-too-large error keeps content and
is not terminal. 50 streaming/continuation/gateway regressions pass.
When a stream delivered text and then died on a context-overflow /
payload-too-large error, the partial content (often tens of KB) was seeded as
a length-continuation stub, growing the transcript monotonically. In a
session whose transcript cannot be compressed back under budget
(protect_last_n covers everything -> no_progress, or the summary would
itself be larger -> would_grow), every later request is larger than the one
that just failed — an unrecoverable loop where the user sees a 30+ minute
fake hang and the only remedy is killing the session (#106260).
classify_api_error already labels these errors context_overflow /
payload_too_large (should_compress=True). _partial_stream_stub now returns
an EMPTY stub marked _overflow_terminal for that class instead of seeding
the recovered text, and recover_from_truncation treats the marker as
terminal: the turn ends via the recovery contract with a clear message
(start /new) and the transcript is not polluted with the partial.
Normal partials (network stall, output-cap truncation, tool-call drops) are
unchanged — only the overflow error class changes behavior.
Tests: stub marker + empty content; a real streamed partial hitting a
'maximum context length' error returns the terminal stub; recover_from_
truncation ends the turn (no fragment/nudge appended) on the marker while a
normal stub still runs the continuation path. 67 streaming/continuation
regressions pass.
_maybe_inject_run_budget_wrapup() appends its wrap-up notice to the newest
role:"tool" message in place, with no _DB_PERSISTED_MARKER check. Its sibling,
_maybe_inject_iteration_budget_warning(), got exactly this guard added in the
same recent saga (turn_iteration_prep.py), with the comment "an older turn may
already be cached."
The reachability is structural, not an edge case: _maybe_inject_run_budget_wrapup
is only ever called from prepare_iteration(), at the START of the next iteration
-- strictly after tool_executor.py's _flush_session_db_after_tool_progress has
already flushed and marked the previous iteration's tool row persisted. So every
successful injection was mutating an already-persisted row: the wire request for
that turn carried the notice, but the durable transcript never did, diverging
replay from the live bytes and invalidating the provider's prompt-cache prefix
from that row onward.
Fix:
- Add the same _DB_PERSISTED_MARKER guard to _maybe_inject_run_budget_wrapup,
scoped to the specific tool row the reversed scan lands on (not just
messages[-1], since this function -- unlike its sibling -- scans backward for
the newest tool row rather than only checking the tail).
- Wire _maybe_inject_run_budget_wrapup into _flush_session_db_after_tool_progress
(pre-flush), mirroring exactly how _maybe_inject_iteration_budget_warning is
wired in both places. Without this, the guard alone would make the notice stop
firing in the common case, since prepare_iteration's call site almost always
hits an already-persisted row -- the pre-flush call site is what actually lets
it land in durable bytes.
Verified empirically: read the real call graph (tool_executor.py's three
_flush_session_db_after_tool_progress call sites cover every tool-completion
path) to confirm the guard's premise, then added an end-to-end test using a real
AIAgent + SessionDB that flushes and checks the persisted row for the notice
text. Mutation-verified: reverting the two production files drops exactly the 2
new/updated assertions (28 pass, 2 fail); reapplying restores green (30 passed).
Also ran the sibling iteration-budget-warning and /steer suites (71 passed) to
check for interaction regressions -- none.
Drop the bare "tool.content" pattern (any 400 mentioning tool.content in a
non-list context would be sent through the image-strip path) and the profile
flag snapshot test; the behaviour tests (classifier verdict + proactive
downgrade) already pin the contract.
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`