The non-streaming 5xx unmask probe runs on the stream worker thread while
StreamingWaitMonitor keeps polling last_chunk_time, which nothing refreshes
during the probe. A probe longer than the stale timeout fired
_kill_stale_stream every window: a false "Reconnecting..." status and a
_bump_stale_streak strike toward HERMES_STREAM_STALE_GIVEUP. Suspend the
stale check (timeout = inf) for the probe's duration; the probe keeps its
own non-streaming watchdog.
The probe also went through _adopt_final_response, which latches
_disable_streaming and logs "switching to non-streaming for this session",
then undid the latch in a finally. Split the pure replay into
_replay_final_response and call that instead.
Also: status via error_classifier._extract_status_code, monotonic probe
window seeded in agent_init, docstring says "per 60s window".
Issue #119533 is the misleading "No LLM provider configured" when the
primary's credential pool is burned. The refusal WARNING only fired when
fallback entries existed and never said why the primary failed, so the
reporter's trigger with no ladder configured still left no trace.
resolve_provider_client returns a bare (None, None), but the pool state
is a cheap local read: ask load_pool(primary) whether it has credentials
yet none available, include "credential pool exhausted" in the WARNING,
and emit it even with no fallback entries. A pool-less, ladder-less
primary (genuine first-run) stays silent to avoid duplicating the raise.
The init-time fallback loop re-derived str(_fb.get("provider")) four
times per refused entry even though _fallback_entries guarantees the
key; bind it once per iteration and reuse it for logs, the refusal
list and the moa check.
A bare exception (e.g. KeyError()) stringifies to "", so the refusal
WARNING rendered "kimi ()" with no reason at all; fall back to the
exception type name.
The explicit-provider branch raises missing_provider_credentials_message,
not the generic 'No LLM provider configured' error, so the WARNING that
summarises refused fallback entries must not assert that verdict. Pin it
with a test on the explicit-provider path.
Refs #119533
An organization with no fast allocation for a model gets a 429 whose
anthropic-fast-*-tokens-limit header is 0, with no retry-after. Hermes
treated it as a rate limit: backoff, credential rotation that benched a
key that works at standard speed, then provider fallback, so /fast
failed every turn.
The pre-classification recovery now stops sending `speed` to that model
for the rest of the session and retries once. A 429 with a real fast
limit still takes the retry-after path.
The api_server platform builds a fresh AIAgent per request (per-request
callbacks, model route, ephemeral prompt), so the memory provider was
re-initialised on every request. External providers deliver recall as the
PREVIOUS turn's background prefetch held on the provider instance, so a
continued session (X-Hermes-Session-Id, previous_response_id, declared
session key) never received automatic recall, and for hindsight
local_embedded each init also restarted the embedded daemon, killing the
retain still in flight. Pre-existing: the same probe fails on main before
the hindsight catalog migration (526d135a96, bundled provider).
ApiServerMemorySessions parks the session's initialised MemoryManager
between requests (exclusive check-out/check-in, keyed by profile home +
session id, LRU/idle eviction under the owning profile's scope) and
AIAgent(memory_manager=...) adopts it instead of loading and initialising
the provider again. /v1/chat/completions, /v1/responses, session chat and
/v1/runs all go through the same two seams (_create_agent, turn finally).
`_cap_binds` was the last inline copy of the cap-clamped-to-window predicate that #119406
centralised; the banner still prints the raw configured cap. Plugin engines without the
helper report "not binding", as before when no cap was set. The threshold_tokens comment
and the new test module stop narrating the incident; the parametrize takes the whole agent
config so the `{}` failure path is spelled out rather than derived from a ternary.
`_parse_compression_config` documents "Defaults here MUST match DEFAULT_CONFIG" because the
config-load failure path hands it `{}`. Every key had an inline fallback except the cap
added in #115986, so a broken config.yaml kept `threshold=0.50` but silently dropped the
256K cap — the pre-#115986 500K trigger on 1M-window models. An explicit null still
means ratio-only.
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours
The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.
- `tools/connectors/portal/`: a client for the portal's tool-list route and a
JSON cache under the Hermes home, one file per portal origin and connector.
An entry is fresh for 24 hours. After that the read revalidates with the
stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
`tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
live in `tui_gateway/methods_connectors_account.py`.
The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.
* feat(connectors): catalog, accounts and member tool rules by RPC
The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.
- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
first. The body is a union on `mode`, so a reader can name who turned a
tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
connector, or one connector on or off), with the revision the user saw. A
stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
page can show one card per app.
* feat(connectors): connect an app without a chat session
Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.
- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
`connectors.operation.wake` and `connection.respond` take `owner`, a union
on `type`: `session` (today's behaviour and authorization) or `account`
(routed by `profile`, authorized by the live transport like `mcp.*`).
`session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
thread, under the profile's scope, so the watcher reads the account and
settles the operation. A second connect for an app that is already
connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
address, so its updates go out on the session-less broadcast path.
* feat(mcp-catalog): eighteen more bundled entries name their hosted connector
A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.
* refactor(connectors): the account handlers share one gate, one params model and one write table
The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.
The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.
The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.
An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.
Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.
* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"
The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.
Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.
* fix(connectors): a connect from the page returns to the app after sign-in
The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.
Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.
* test(connectors): defer the new connector RPC coverage
The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.
Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.
Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.
* fix(cli): the connection panel hands the tool thread back at once
The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.
For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.
The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.
Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.
* fix(connectors): "run it again" lives in the library, so the classic CLI can use it
Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.
`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.
Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.
* feat(connectors): the account list and disconnect go through the portal
`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.
The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.
`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.
* fix(connectors): the account RPCs answer what the portal really sends
Checked against the portal source and against the staging and production
services.
- Errors are read from the upstream error code, not the HTTP status. A rule
write answered 409 for a stale revision and for a user with no organisation;
both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
portal's own result for this user, with its stamp and without provider or
subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.
Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.
* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link
Found by two adversarial reviews of the RPC layer and its types.
- `connectors.connect` from a chat session with no open operation is refused
(`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
registry with no card: it made a link nobody watched, returned a reply
without the required `settled` field, and named an operation that was never
registered. There is one way into an operation: the agent's call, or the
account owner's `connectors.connect`. "Run it again" inside an open
operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
profiles with the same key no longer cross-deliver a sign-in link. The event
payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
`connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
The last one is display data: the gateway enforces the rules, the backend
only passes the list on. The phantom `name` and `description` are gone, and
the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
produces it. The contract generator now fails when a contract enum and its
domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
`optional-mcps/` are untouched by this PR again.
anti-slop: no net-new findings (15 touched files).
* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect
The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.
- The agent turn now declares how a link can reach the user
(`tools/connectors/turn.py`): CARD when the agent was built with a
connection callback, SIDE for a subagent or a background turn, LINK for a
headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
tool batch in the agent loop and read by the connector dispatch path, which
never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
tool on any path that derives the tool list, and a connector call on an
unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
does, so no `connection.update` is emitted for an operation no client asked
for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
`respondToConnectionRequest` share one guard and send nothing for a settled
or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
repeated it to the user.
Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.
* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients
A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.
- `tools/tool_labels.py` is the one place that turns a bridged call into a
label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
→ "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
keeps its own emoji, verb and primary-argument preview. A batch gets exactly
one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
failure text on the row of the call that failed. With friendly labels off
it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
carry a typed `labels` field. It does not depend on the classic CLI's
display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
labels, one row per call. The labels reach the row under a key no tool
argument can use. The connect card it drew under a failed tool result is
gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
`manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
"Reading tool details · N tools".
Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.
* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly
- A failed hosted search or describe used to return nothing, by design, so the
model saw only local tools and told the user that a connected app was
missing. The local results are unchanged; when the hosted leg failed, the
`tool_search` and `tool_describe` results carry
`connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
and one hint line. A rejected token is `sign_in_expired`; an entitlement
refusal or a shut gate adds nothing. `tool_describe` no longer lists those
names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
is a hosted connector account; `mcp: true` only when the user asks for an MCP
server, a local server or an install, or when the name exists only in the
catalog; connect and reconnect are hosted verbs, install, enable and
authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
gateway does not know the connector (confirmed on that failure path) and the
name is a catalog entry does the target fail with "X is a local MCP server.
Call manage_connections with action install ...". It is a per-target
outcome: other targets of the same call keep their links and their card. A
vendor failure on a name both sides know stays an ordinary failed row. The
MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
USER asks for that app again; the description and the settled-result notes
say so. A builder saw the model refuse a direct user request.
Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.
* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled
Reproduced on the real Ink TUI with the rig, then fixed:
- The keyboard was dead during the sign-in wait: the card kept a `submitting`
flag that the normal OAuth path never cleared, and Esc went through the same
guard. The in-flight state now belongs to the answered row and clears when
that row moves, when any later frame of the operation arrives, or after
five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
(the input handler had no branch for this overlay); Shift+arrows scroll the
transcript and the card ignores them; arrow keys no longer move the text
cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
operation stayed in the store, and a resume dropped the pending card. The
flag survives idle, a resume shows the pending card again, a session switch
clears it.
- States with no branch: `not_connected` and a row with no link fell into the
credential form; `expired` vanished with no note. The title and the row text
now name the action (connect, reconnect, install, enable, authorize); a
failed or expired row with no fields offers Try again / Skip; a failed row
WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
per app states the outcome. A settled or dismissed operation id is
remembered, so no replay or resume can reopen its card. Esc in the last
"Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
the card in one sentence.
Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.
* chore(connectors): remove the comments and docstrings this branch added
Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.
Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.
* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures
Found by the end-to-end runs on the pushed head.
- Desktop: after a window reload, Continue on the restored card sent nothing.
The answer looked up the backend that holds the session with the runtime
session id, the lookup wants the stored id, and a failed lookup returned
silently. When the lookup gives no owner the answer now goes out on the
window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
a refused scope, a non-member and a missing organisation alike: the handler
runs with the gateway's globals and did not import the reason enum, so its
own error mapping raised. `connectors.accounts.remove` caught auth failures
in its generic branch. `org_required` was mapped on `policy.set` only. All
six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
`ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
A general-plugin context engine is one shared instance; agent init copied it per agent
with copy.deepcopy() only. Engines that hold a SQLite connection or lock (hermes-lcm)
already expose clone_for_agent() for exactly this, but it was never called, so every
init logged "could not be safely copied … falling back to built-in compressor" and the
engine was unusable through the plugin system.
ContextEngine grows clone_for_agent() (default: deepcopy, the previous behaviour) and
_select_context_engine calls it; the failure message now names the hook to override.
Docs: context-engine-plugin.md documents the per-agent clone contract.
Test change (existing on main): tests/agent/test_context_engine.py::
test_agent_init_source_deepcopies_singleton_not_aliases was a source-reading pin on the
literal `copy.deepcopy(_candidate)` line, which this fix intentionally replaces. It is
superseded by tests/agent/test_plugin_context_engine_clone.py, which drives the real
_select_context_engine seam and asserts the invariant it guarded (child update_model()
never mutates the shared singleton) plus the new clone_for_agent() path.
Fixes#99640
credit: @stephenschoettler #62374
credit: @686f6c61 #99677
`memory.provider: builtin` (also `default`, `built-in`, `none`) selects the built-in
store, yet doctor short-circuited only on an empty name and printed
"⚠ builtin plugin not found run: hermes memory setup", and the provider migration
(`hermes update`, agent init) printed "⚠ Memory provider 'builtin' is configured but
not installed and not in the plugin catalog". update_cmd_deps and web_server_memory
each carried their own (differing) inline sentinel set.
One predicate, `agent.memory_provider.is_core_memory_provider`, now owns the sentinel
set next to the MemoryProvider ABC; doctor_state, memory_provider_migration,
update_cmd_deps, web_server_memory and agent_init's provider activation all use it.
Fixes#75647Fixes#115113
credit: @Christopher-Schulze #75679
credit: @Luna161 #83317
credit: @kokhlo #115117
credit: @KnowBotDev #115302
The 1h Anthropic cache tier writes at 2x base (5m: 1.25x) and only pays off when
turns are more than five minutes apart. That is exactly the shape of an interactive
session a person parks and resumes, and exactly not the shape of a subagent, cron
run, one-shot or webhook that calls every few seconds and is gone. A single global
`cache_ttl` cannot be right for both, so operators leave it on 5m and pay a full
context re-write every time they come back to a CLI session after a coffee.
Measured on one install (2 days of per-call API logs, Claude via the Nous route):
63% of interactive cache-write tokens were cold re-writes after a 5-60 minute idle
gap; 1h would cut interactive write cost ~42% while costing ~49% more on subagents
and ~23% more on cron. `auto` resolves once per session from the session source
(`_session_source_for_agent`): 1h for cli/tui/desktop/messaging platforms, 5m for
subagent, cron, oneshot, webhook, kanban, api. Auxiliary/stub calls keep 5m; the
delegate_tool child clamp (#104168) still applies. Default stays "5m".
Reverts 93889b770d ("named profiles no longer inherit the root
profile's auth.json"). After the Desktop update every bot profile that had
relied on the root OpenAI Codex login failed with "No Codex credentials
stored. Run `hermes -p <bot> auth add openai-codex --type oauth`", and users
had to re-run the device-code flow once per bot (5-6 times in the field
report). Sharing one grant across profiles is the intended design: OAuth
refresh tokens are single-use, so ONE grant lives at the root, profiles
resolve it read-only, and a refresh under a profile writes the rotated chain
back to root (Codex / xAI write-through, borrowed-row pool bookkeeping,
forked-grant heal) — never a per-profile copy.
Restored: `_global_auth_file_path` / `_load_global_auth_store` fallback in
`_load_provider_state*` / `read_credential_pool` / `_provider_state_transaction`,
Codex + xAI root write-through, `credential_pool` borrowed-root persistence,
`heal_forked_single_use_oauth_grants`, `share_auth` on profile creation
(Desktop create dialog checkbox), and the docs. `profile_credential_audit.py`
(the `hermes update` "profiles without a provider" notice) is removed with it.
Kept from after #111724: `_save_codex_tokens(set_active=...)` for image gen,
the plugin-auth `status` dispatch and the external-login notice in
`hermes auth list`, and the registry-derived env-var hint in agent_init.
The first head stored a single rejecting (provider, model) and kept the
turn-global `_vision_supported` as the recovery guard. In a fallback
chain that fails: model A rejects images and retries text-only, a later
error activates model B, the restart rebuilds api_messages from history
so B receives the images, and when B rejects them too the branch is
skipped because `_vision_supported` is already False — the request
falls through to generic error handling. Recording B also overwrote A,
so A was no longer treated as text-only on later turns.
`_image_rejecting_models` is now a set of every rejecting model, and it
is also the guard: each model's first rejection runs the recovery and a
repeat rejection from the same model still falls through, so the retry
cannot loop. image_model_key() names the key in one place.
Adds a test for the two-model sequence (fails on the previous head) and
one pinning that a repeat rejection from the same model does not retry.
Thanks to @ehz0ah for the review.
(cherry picked from commit 225fd76ccf90cd3ec509a9327f725ac02d996f24)
When a provider 4xx's on image content, recover_before_classification
ran _strip_images_from_messages on the canonical `messages` list and
reset _db_flush_scan_prefix. Since #117569 that function pops
_db_persisted on every rewritten dict, so the next flush rewrote those
rows: every image in the session — and every image-only message, which
the stripper deletes outright — was removed from state.db for good.
The rejection describes what the CURRENT model accepts, not what the
conversation holds. An automatic fallback to a text-only provider, or a
single /model switch, was enough to erase images the user had sent to a
vision model, and switching back found them gone. It is the same failure
as the ASCII strip in #117802, on the image path; neither open fix for
that issue touches this branch.
Keep the repair on the send path, where the per-call copy already lives
(_clone_message_for_send exists so send-path rewrites never reach the
persisted transcript, #80498):
- record the rejecting (provider, model) on the agent, a session-scoped
flag initialised beside _force_ascii_payload;
- strip the in-flight api_messages copy for the immediate retry;
- build_api_request calls strip_images_for_rejecting_model() on each
attempt's api_messages BEFORE provider conversion. The stripper knows
Hermes's own part types; a converted payload would slip past it
(Bedrock Converse image blocks carry no `type`). Keyed on the model,
so one that accepts images gets them again.
_strip_images_from_messages itself is unchanged, so its role-alternation
and sidecar guarantees still hold on the wire copy. The notice no longer
claims "text-only mode for this session" (_vision_supported resets every
turn) or that images were stripped from history.
(cherry picked from commit cbdb184c27f915ab138b2087f878aed7fcc7c76b)
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
model.context_length is the user's profile-wide ceiling. It was read from
config.yaml in exactly one place — agent construction — and cached twice:
agent._config_context_length (switch/fallback resolution plus every display and
/usage surface) and context_compressor._config_context_length (the compressor's
own re-resolution).
Every live path that re-resolved a runtime then touched only one copy, or cleared
it without re-reading the config:
- switch_model nulled agent._config_context_length and re-derived the intent from
custom_providers metadata alone, so a ceiling that only exists as
model.context_length was dropped for the rest of the process;
- the Desktop/TUI compression hot-reload updated the compressor's copy only, so an
open session showed a pinned ceiling while compressing against provider
metadata / the 256K fallback.
Both now route through one pair of helpers in agent/agent_init.py:
set_config_context_length (one place that knows where the pin is cached) and
config_context_length_for_runtime (re-read from live config, scoped exactly like
construction, so an unrelated route still never inherits the pin).
(cherry picked from commit 986ff16dadb9966f7328e55f295af5cfa1eb5c88)
A named custom provider at api.anthropic.com (api_mode anthropic_messages) whose token comes from a
key_cmd callable lost the Claude Code OAuth identity: the aux custom routes hard-coded
is_oauth=False, agent_init/agent_runtime_helpers/client_lifecycle gated OAuth on provider=="anthropic"
and isinstance(key, str), and the callable-token client builder never added the OAuth betas or the
claude-code user-agent. Anthropic answers such a bare Bearer with 429 rate_limit_error "Error"
(#114967). One resolver, anthropic_credentials.anthropic_route_is_oauth(base_url, credential,
provider=), decides at every site: the route qualifies for the anthropic provider (unchanged) or an
exact api.anthropic.com host, the credential is a string or a callable materialized once
(CommandTokenSource caches), third-party hosts never qualify. model_metadata
_query_anthropic_context_length skips a callable credential instead of crashing agent init
(AttributeError on the same key_cmd route on current main).
Slimmer redo of #115007 by @liuhao1024 (same direction; one shared helper instead of per-site copies).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Trim the salvaged fix to the shape main wants:
* Every gate asks ``agent.transports.registered_api_modes()`` directly. The three helper
spellings (``_has_registered_transport``, ``_registry_knows``, ``is_registered_api_mode``)
and the ``sys.modules`` peek are gone: no transport module imports ``providers`` or
``hermes_cli`` at module level, so a plain import cannot re-enter provider discovery.
* ``hermes_cli/auth.py`` late-registration pass dropped — main already re-syncs plugin
profiles into ``PROVIDER_REGISTRY`` on every registry miss
(``auth_plugin_providers.registry_lookup`` / ``sync_plugin_provider_registry``, #102123);
the probe shows a profile registered after the import-time mirror resolves and reaches
the wire on base.
* ``ProviderTransport.normalize_stream_delta`` and the streaming-assembler hook dropped —
legacy ``delta.function_call`` translation is a separate concern from api_mode
propagation and has no in-tree consumer.
* Tests: 15 gate-by-gate unit tests replaced by two invariants that install a REAL plugin
under a temp HERMES_HOME and walk profile → determine_api_mode → resolve_runtime_provider
→ agent ladder → delegation resolver (positive: red on origin/main; negative: an
unregistered mode still degrades to chat_completions).
A provider plugin ships a transport via `register_transport(api_mode, cls)` and
declares that same string as its profile's `api_mode`. Transports are selected by
that string, but every gate that validates an api_mode compared it against a
closed literal, so a plugin's mode was rejected at each one and rewritten to
`chat_completions`. The plugin's transport was then never selected: no error, no
tool call, the turn silently degraded to prose. `register_transport` was a public
seam with no way through.
Accept a mode when the transport registry knows it, via a new
`agent.transports.registered_api_modes()`, at each gate:
* `agent_init._resolve_api_mode` - the agent's mode ladder;
* `runtime_provider._parse_api_mode` - the config gate;
* `delegate_tool_config` - the delegation resolver;
* `providers.get_provider` - the reverse `TRANSPORT_TO_API_MODE` lookup recorded
an unknown mode as `openai_chat`, which made `determine_api_mode` report
`chat_completions` for a provider that has a dialect transport. This one is the
most deceptive: the other gates already pass, and the transport still is not used.
* `providers.determine_api_mode` - the same table lookup at the other end.
The registry read is deliberately lazy (`sys.modules.get("agent.transports")`,
never an import): this code is reached from `determine_api_mode`, which provider
discovery itself calls while the registry is being populated, and importing the
transport package there re-enters discovery.
Also add `ProviderTransport.normalize_stream_delta()`, the response-side twin of
`convert_messages()`: a provider that streams a tool call on the legacy OpenAI
`delta.function_call` pair instead of indexed `delta.tool_calls` had nowhere to
translate it, and the streaming assembler dropped the call. The default returns
the delta unchanged, so existing transports are untouched; the assembler now asks
the transport instead of hardcoding one provider's shape.
Finally, make plugin-provider registration repeatable. `hermes_cli.auth` registered
plugin profiles once, at import, from a list `hermes_cli.config` had already
partially discovered while importing itself. A profile that sorts LAST in discovery
was absent from that snapshot and never reached `PROVIDER_REGISTRY`, so every
consumer reported it unauthenticated and it silently vanished from the model
picker while working fine from the CLI. `ensure_plugin_providers_registered()` is
now called from `get_auth_status()` and `resolve_provider()`, so a late profile is
picked up instead of staying invisible.
Unregistered modes are still rejected everywhere, and the in-tree literal sets are
unchanged - they are simply no longer the only way in.
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
Gate Lows: an in-flight exception pins the executor frames through its traceback, so the
finally-placed trim could not release the result on that path and would burn the cooldown;
the flag now survives to the next completed batch. The default lives beside _executing_tools.
The Responses auto-upgrade guard keyed only on the ``acp://`` / ``acp+tcp://`` scheme after
the copilot-acp slug check was dropped (#116408). ``COPILOT_ACP_BASE_URL=https://...`` flows
through ``_external_process_spec`` verbatim, so a gpt-5 model on copilot-acp flipped api_mode
to codex_responses against an ACP client. Key the guard on the profile's external_process
auth_type as well, reusing ``runtime_provider_backends._is_external_process_provider`` (CLI
registry first, then the profile registry) at both the routing and launch-kwargs sites.
Picker admission (#116552) listed out-of-tree external-process and OAuth
plugin providers, but selection and status still dispatched through
provider-name tables:
- `hermes model`: no `_PROVIDER_MODEL_FLOWS` entry and no generic flow, so
picking an admitted plugin row was a silent no-op. One generic flow in
model_setup_flows.py, credential step keyed by the profile's auth_type
(external_process -> launch check; oauth_* -> live pool row, else the
`hermes auth add <name>` hint), catalog via merge_profile_catalog; main.py
falls back to it for any registered profile missing from the table.
- `_STATUS_BY_AUTH_TYPE` had no builder for oauth_device_code/oauth_external,
so `get_auth_status`/`list_available_providers().authenticated` stayed False
with a live pool entry. `get_plugin_oauth_auth_status` (auth_plugin_providers
sibling) reads the pool; gated on PLUGIN_MIRRORED_PROVIDERS so bundled OAuth
providers keep their bespoke status bytes.
- `_external_process_auth_evidence` computed evidence for copilot-acp only, so
`inventory._external_process_signed_in` hid every other ACP row from the
Desktop explicit_only picker. Generic evidence = the binary resolves; the
bundled CLI keeps its token-store chain.
- agent_init Responses-upgrade guard dropped the vendor literal redundant
with the acp:// scheme check.
- `fetch_account_usage` bounds the plugin hook with a shared 10 s deadline
(previously only the CLI wrapped the call; gateway/TUI awaited unbounded),
contextvars-propagated so scoped secrets resolve; overrun -> None.
Part of #116408
(cherry picked from commit 536a4e7fcf2d1ac2b8a70bd62a69707d58ea5d1e)
_explicit_client_kwargs hardcoded copilot-acp for command/args launch
kwargs; out-of-tree external_process plugin providers got no launch path
and failed at client construction. Key on the provider profile's
auth_type instead, so every ACP/subprocess provider launches the same
way (same approach as #111194, folded in with credit).
Also harden _profile_live_catalog: a signature-strict external_process
profile (fetch_models requiring keyword-only api_key/base_url) now falls
back to credential kwargs on TypeError instead of crashing discovery.
Tests proven red on base for both behaviors.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
When api_max_retries and the fallback chain are both exhausted on a transient
outage (5xx, overloaded/529, connect/read timeout) and no answer text has been
delivered yet, the turn used to end with "API failed after N retries" even
though the provider would be back a minute later, leaving the user to notice
and re-send. settle_unrecovered_error now hands that case to
agent/turn_recovery_autorecover.py: up to agent.auto_recovery_cycles (default 5)
wait-and-retry cycles on a jittered 15/30/60/60/60 s schedule, a provider
Retry-After winning up to 120 s, each cycle announced on the status rail AND
the live wait line ("Provider temporarily unavailable — retrying automatically
in Ns (cycle k/5); press Esc to stop", with a per-surface stop hint) plus a
log line for cron. The interruptible wait is the existing one, so Esc/stop
cancels cleanly and a steering correction still rebuilds the turn.
Fallback stays first: the ladder engages only when _try_activate_fallback has
nothing left. Overload-class errors ride this schedule instead of growing a
separate overload backoff path (#107307). Non-retryable classes never enter
because they exit through the client-error branch above. TurnRetryState
carries the cycle counter; jittered_backoff supplies the schedule — no second
retry framework.
Credit: @MilevskyYakov's #85441 established the shape (reuse
try_recover_primary_transport / jittered_backoff / TurnRetryState, interrupt
mid-wait, never replay delivered text); this lands it at the exhaustion seam
main has today with a bounded default.
Conflicts resolved toward the branch: PM owns dependency preparation, the
Windows shim re-exec/hand-off path stays retired (main's shim-parent wait,
gateway-resume env token and update_cmd_deps tests dropped), docs describe
the PM update flow. The docker workflow parks install-stamp.json around the
toolchain step instead of deleting it so tests/docker can compare provenance.
model.context_length short-circuits get_model_context_length ahead of every
provider source, but nothing told the user the number they saw was their own
pin rather than provider metadata (#66168). Keep the pin semantics (custom
endpoints depend on it) and make it visible instead:
- agent/context_pin.py: is_context_pinned / context_pin_suffix for renderers,
and warn_once_on_pin_disagreement, called from _resolve_context_length at
agent init. The advertised value comes from LOCAL sources only (endpoint
cache, models.dev disk cache, hardcoded catalog) so the check never adds a
network probe or latency to startup; one warning per (model, pin) per process.
- "(pinned)" label on every CLI surface that renders the window: welcome
banner, /model switch summary, /usage current-context line, wide status bar.
- No new config key or env var; the numeric value used is unchanged.
Fixes#66168
When an OpenAI-compatible local server (llama.cpp, vLLM, ...) serves a
window below MINIMUM_CONTEXT_LENGTH, the startup refusal, the CLI banner
warning and the num_ctx log line were worded for Ollama (/api/show,
model.context_length advice that cannot raise a served window). Name the
served window and the server-agnostic remedies: start the server with a
>=64K context or set model.ollama_num_ctx (honoured on every local
endpoint) to the window it really serves. Hosted routes keep the
model.context_length advice. The floor itself is unchanged.
Part of #87075
A finite `hermes chat -q` / `--oneshot` run has no later session in its HERMES_HOME
to learn for, yet it ran the full interactive self-improvement loop. Measured over
21 one-shot benchmark trajectories: 7 skills created and a bundled one patched
mid-task, 37 of ~215 tool calls on skill_view/skill_manage, skill text = 34% of all
tool-result bytes fed back into context, plus reviewer subagents spawned on the
agent's own diff (one task: 5 delegations, 62 subagent API calls, each re-paying a
cold system prompt).
Keyed on the existing HERMES_SINGLE_QUERY_SESSION marker (approval gate, delegation
dispatcher), so interactive and gateway sessions are byte-identical:
* agent/oneshot_footprint.py (new sibling): skill_manage is pruned from the tool
set; the ## Skills block keeps the index + skill_view but drops the record/patch/
offer-to-save coaching and the "load process skills for work you already know"
push (SKILLS_GUIDANCE follows because it is gated on skill_manage).
* delegation.oneshot_max_children (default 2, 0 = unlimited): total children a
one-shot run may spawn; past it delegate_task returns a tool error telling the
model to finish inline.
* requesting-code-review skill: reviewer/fixer subagents (Steps 5 and 7) are
interactive-only; one-shot applies the checklist inline.
Live: `hermes chat -q "list tools starting with skill_"` on the portal — base
"skill_manage, skill_view, skills_list", fix "skill_view, skills_list"; a real
AIAgent under a temp HERMES_HOME shows the interactive prompt unchanged (6643 chars
both) and the one-shot prompt without skill_manage / offer-to-save.
`preview_threshold_tokens` restated the resolve -> floor -> compute -> cap chain that `update_model`
runs; two copies of the trigger math drift the next time a step is added — the exact bug class
#83450 fixes (the guard quoting a number the compressor will not install). `_derive_trigger` is the
single pure derivation; the auxiliary-summariser ceiling stays in `_apply_threshold_tokens_cap`
because it is per-runtime, not per-model.
The startup banner names the cap only when it set the trigger; on windows where the ratio already
sits below it "(capped at 256,000)" was noise. Comments no longer repeat the default literal.
New config key `agent.text_verbosity` ("" | low | medium | high, default ""
= not sent). When set, the Responses-family transport emits the top-level
`text: {"verbosity": ...}` field so GPT-5-family models can be asked for
terser or fuller final answers independently of reasoning effort (#20203).
Why this shape: the value is parsed once in agent_init._apply_agent_section
(unknown values warn and are ignored, so unset/"" can never flip the
provider default), passed to ResponsesApiTransport.build_kwargs like the
other per-request params, and never reaches chat_completions / Anthropic;
xAI's /responses is skipped the same way service_tier is. An explicit
request_overrides["text"] (e.g. structured-output format) still wins because
overrides merge after it. The Codex preflight whitelist entry that lets the
field through landed in the previous commit.
DEFAULT_CONFIG entry + docs row under Reasoning Effort.
Salvaged from PR #20258 (config key, docs and adapter direction); the
separate agent/text_verbosity.py module, constructor kwarg and the
cli/gateway/cron/tui_gateway plumbing were dropped in favour of the existing
agent-section config seam.
Review findings on the first cut:
* A hosted provider with a stale model.ollama_num_ctx passed the floor for
a 40K model although nothing ever raises a hosted window. The served
window now counts only when the endpoint is local (is_local_endpoint),
the same gate the num_ctx probe uses.
* Moving the whole num_ctx phase ahead of the floor also moved tek's
compressor clamp ahead of it, which flipped two observables: a
model.context_length above a sub-64K num_ctx was rejected instead of
constructed-and-clamped, and the floor's message reported the clamped
value with advice (set model.context_length) that could not help. The
phase is split: resolution runs before the floor, the clamp
(_clamp_compressor_to_ollama_num_ctx) stays at its original position,
so every case main constructed still constructs with identical
compressor numbers and the floor's message is unchanged.
Tests: the harness is a module-level helper so the new class no longer
re-collects the parent's tests (13 -> 11 collected); the negative pins
the local-endpoint gate (red when the gate is dropped) instead of a case
main already rejected.
An Ollama server serves num_ctx. A Modelfile or model.ollama_num_ctx at
65536 is a usable window even when the model metadata advertises 40960,
yet the floor ran before num_ctx was resolved and read only the probed
window, so the agent (a cron job reaching a local fallback in the report)
refused to construct with "context window of 40,960 tokens".
num_ctx resolution now runs before the floor and the floor takes
max(probed, served). The compressor keeps tek's one-directional clamp
(cb71d5f1b1): it still targets the smaller probed window, so nothing about
compaction thresholds changes; a served window below 64K is still rejected.
Rebuilt from PR 100475 by fangliquan (the agent_init it targeted was
decomposed since); the unrelated cron pin contract test there is not taken.
Co-authored-by: fangliquan <fangliquan@qq.com>