A `$HERMES_HOME/plugins/model-providers/<name>/` plugin re-registering a
bundled provider (stepfun with a regional base_url, gmi at a staging host)
wins in `providers._REGISTRY` — register_provider() is last-writer-wins and
the plugin guide promises exactly this — but the runtime reads its endpoint
from `hermes_cli.auth.PROVIDER_REGISTRY`, whose mirror loop skipped every
name already present, so inference kept going to the built-in URL (#48450).
The mirror now applies one explicit precedence rule: when a row core wrote
(built-in or plugin-mirrored) belongs to a name whose profile
currently registered came from a USER plugin,
the row's profile-derived fields are rewritten in place (inference_base_url;
api_key_env_vars / base_url_env_var on api-key rows when the profile declares
env_vars). `providers` records the discovery source per registration
(`provider_source()`), because a bundled profile must never rewrite a
built-in row: several bundled profiles omit the row's `*_BASE_URL` env var
and one differs in auth_type, so an unconditional "profile wins" would have
changed built-in behaviour. With no user plugin PROVIDER_REGISTRY is
byte-identical before/after (78 rows probed). copilot/kimi/zai keep their
bespoke resolution via the existing skip set.
Co-authored-by: xiaoxinova <xiaoxinova@users.noreply.github.com>
"database is locked (another Hermes process held the state.db write lock for
over 60s)" identified the victim only. The open-descriptor scan cannot single
out the writer because every Hermes process (gateway, CLI sessions, worktree
agents, cron) has the DB open, so an operator hit repeatedly by
session_persistence_failed:locked had nothing to act on.
SQLite's unix VFS takes fcntl byte-range locks whose offset encodes the lock
kind (state.db-shm byte 120 = WAL write, 121 = checkpoint; the pending-byte page
on state.db = PENDING/RESERVED), and the kernel exports them with the owning pid
in /proc/locks. hermes_state_lockowners reads that table at the moment the
patience budget runs out and logs one WARNING per write-class holder with
describe_holder_pid()'s argv summary, for both the transcript write path and
open+init lock patience. The holder stays out of the exception text on purpose:
classify_persistence_error() buckets by phrase and a holder argv such as a
worktree named fix-corrupt-db would flip the bucket.
Docs: the Write Contention section still described attempt-counted retries
(_WRITE_MAX_RETRIES = 15); updated to the time budgets in force and the new log line.
turn_api_error enters the fallback walk only for ``retryable=False`` verdicts outside the
retryable-client reasons; the built-in terminal verdicts pin retryable=False while the
rate-limit family stays retryable and cascades after backoff. A ``classify_api_error`` hook
returning ``{"reason": "billing", "should_fallback": True}`` therefore retried the dead
route instead of cascading (the #116408 test passed only because its hook also set
retryable=False). Default retryable to False for such verdicts, leaving rate-limit reasons on
the built-in retry-then-fallback shape; ``RETRYABLE_CLIENT_REASONS`` moves next to the
verdict table so both modules read one set. Contract documented in the plugin guide.
A kind: model-provider plugin is loaded by providers/ discovery and never enters the
PluginManager hook lifecycle, so transform_api_error_classification was unreachable for it
without shipping a second plugin component. The profile now carries an optional
classify_api_error(error, *, status_code, error_code, message, body, model) callable,
consulted as a classifier stage right after the generic plugin hooks and only for the
provider that produced the error. None or an unknown reason leaves the built-in verdict;
built-in providers are untouched (no name table, no lifecycle change).
Also: a plugin refresh_credential returning None/empty was treated as a successful refresh
(row marked ok, stale bearer replayed up to the refresh cap). It now benches the row like a
failed refresh POST, so the loop rotates or falls to the generic sign-in copy.
Part of #116408
(cherry picked from commit b8129fd6fd6a0cf4eeee6d95a5d5d823668a306d)
Minimal OAuthPKCEConfig example plus what Hermes owns and the enforced
security boundary, so plugin authors do not hand-roll a browser flow.
(cherry picked from commit fc4e73d5d2067102ab08823aafbb4f31b3cba785)
Picker admission (#116552) listed out-of-tree external-process and OAuth
plugin providers, but selection and status still dispatched through
provider-name tables:
- `hermes model`: no `_PROVIDER_MODEL_FLOWS` entry and no generic flow, so
picking an admitted plugin row was a silent no-op. One generic flow in
model_setup_flows.py, credential step keyed by the profile's auth_type
(external_process -> launch check; oauth_* -> live pool row, else the
`hermes auth add <name>` hint), catalog via merge_profile_catalog; main.py
falls back to it for any registered profile missing from the table.
- `_STATUS_BY_AUTH_TYPE` had no builder for oauth_device_code/oauth_external,
so `get_auth_status`/`list_available_providers().authenticated` stayed False
with a live pool entry. `get_plugin_oauth_auth_status` (auth_plugin_providers
sibling) reads the pool; gated on PLUGIN_MIRRORED_PROVIDERS so bundled OAuth
providers keep their bespoke status bytes.
- `_external_process_auth_evidence` computed evidence for copilot-acp only, so
`inventory._external_process_signed_in` hid every other ACP row from the
Desktop explicit_only picker. Generic evidence = the binary resolves; the
bundled CLI keeps its token-store chain.
- agent_init Responses-upgrade guard dropped the vendor literal redundant
with the acp:// scheme check.
- `fetch_account_usage` bounds the plugin hook with a shared 10 s deadline
(previously only the CLI wrapped the call; gateway/TUI awaited unbounded),
contextvars-propagated so scoped secrets resolve; overrun -> None.
Part of #116408
(cherry picked from commit 536a4e7fcf2d1ac2b8a70bd62a69707d58ea5d1e)
Model provider plugins can now declare per-model metadata (capability
booleans, context_window, model_family) through the existing
ProviderProfile registry using the canonical model_overrides schema.
Declarations patch catalog metadata without erasing unknown fields; on a
catalog miss, undeclared capability booleans stay unknown rather than
false.
Precedence (tested): explicit user model_overrides > plugin declarations
> catalog/builtin > _default fill-gap. lookup_models_dev_context shares
the same seam so auto-detected context reflects declarations, including
the dashboard /api/model/info response.
Provider aliases resolve declarations; user override lookup keeps its
existing provider-key rules. Metadata lookup may trigger the existing
lazy provider discovery (imports plugin code); the registry is
process-global and discovered once per process.
Progress on #102115.
(cherry picked from commit a2bf725597efddf5238a8d0447705cf31ac71cc8)
ProviderProfile.supports_vision is documented as "the API accepts image content
inside tool-result messages" -- a provider-wide wire capability, not a per-model
user-image verdict. Treating it as the latter flipped every model on the bundled
`router` relay (and models.dev-unknown meta-ai/xiaomi models) from text to native
image parts. Per-model vision now comes from ProviderProfile.model_capabilities
(#116570) through the existing models.dev probe.
The setup-catalog test mirrored the plugin through auth._register_plugin_provider,
which #116553 renames; build the ProviderConfig from public types instead so the
test passes on both trees.
`decide_image_input_mode` consulted config overrides, the managed runtime,
models.dev and Ollama, but never the registered profile's `supports_vision`
— the field the tool-result media path already trusts. A plugin that
declared vision was therefore native for tool results and text-only for
user-attached images. The declaration is now the last probe in
`_VISION_PROBES`: only an explicit True is a verdict, and per-model catalog
entries still win.
Part of #116408
The CANONICAL_PROVIDERS auto-extend skipped plugin profiles by auth_type
("Non-api-key flows need bespoke picker UX"). Every in-tree non-api-key
profile already owns a hand-written row, so the set never excluded an
in-tree provider - it only hid out-of-tree external-process and OAuth
plugins from /model, hermes model and list_available_providers().
Admission is now by slug (dedupe against the built-in rows); visibility
stays gated downstream by credentials (binary resolves / auth.json /
pool entry), so an admitted row reads authenticated=False until sign-in.
- _profile_live_catalog: external_process profiles use fetch_models(),
then fallback_models; every other non-api-key profile returns its
fallback_models instead of None (in-tree ones declare none, so the
built-in rows are byte-identical).
- _credential_fingerprint: external-process command/argv env overrides
key the catalog cache, as an API key does for HTTP providers.
- Drop the "[slug] as the model" shim from the canonical picker lap;
the profile's fallback_models is the honest catalog.
- Tests trimmed to two invariants proven red on base; docs describe how
picker rows and the authenticated flag are derived.
Co-authored-by: ericmaddox <eric.maddox@outlook.com>
Co-authored-by: Denis Araujo <d.araujo@tjrs.jus.br>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: tobenwarrior <lucasyee999@gmail.com>
First in-tree consumer of ProviderProfile.fetch_account_usage: the OpenCode Go plan's rolling /
weekly / monthly windows (GET /zen/go/v1/usage) render in every /usage surface without adding the
provider name to the core _USAGE_FETCHERS table. Ported from #113418, which implemented the same
fetch as a core table entry; the literal endpoint (not the runtime base_url, which loses /v1 in
anthropic_messages mode) and the window mapping are theirs.
Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
Independent review of the plugin refresh branch (#116553) found four gaps
between what model-provider-plugin.md promises for `refresh_credential`
and what `_refresh_entry_impl` did:
1. `replace(entry, **hook_result)` raised TypeError on any non-field key
(`expires_in`, `token_type`, `scope` — the natural token-endpoint shape),
the except benched the row EXHAUSTED and the pair the server had already
rotated was dropped: for single-use refresh tokens that is a lost login.
Field keys now go through `replace()`, everything else merges into
`entry.extra` (mirrors `from_dict`); `None` = no rotation, mark ok.
2. Plugin providers skipped the locked single-use path, so a gateway and a
CLI could both POST the same refresh token (`refresh_token_reused`).
Providers with a hook now take the `_auth_store_lock` path: re-read the
pool store, adopt a peer's usable rotation and skip the hook, else call
it and write through. Eligibility derives from `plugin_refresh_hook()`,
not from extending the built-in name tuple.
3. A raising hook re-benched EXHAUSTED every cooldown forever at DEBUG.
`AuthError(relogin_required=True)` (or a grant-dead OAuth code) is now
terminal: the row goes DEAD with a WARNING naming `hermes auth add`.
Any other exception stays a transient bench (negative test kept).
4. `from_dict`'s extra sweep round-tripped a stray row-level `provider` key
back onto the row on `to_dict()`; it is bookkeeping, not metadata.
New logic lives in `agent/credential_pool_plugin.py` — credential_pool.py
is at the size cap; the facade only dispatches.
Part of #116408
An out-of-tree ProviderProfile with auth_type oauth_device_code/oauth_external loaded and inferred
but was invisible to hermes auth: _register_plugin_provider skipped every auth_type except
external_process and api_key, so resolve_provider() said "Unknown provider", `hermes auth add`
had nothing to dispatch to, the pool could not refresh its rows (REFRESHABLE_OAUTH_PROVIDERS is a
name set) and PooledCredential.from_dict dropped every extra key outside _EXTRA_KEYS on reload.
- hermes_cli/auth_plugin_providers.py (new sibling; auth.py is at the size cap): the registry
mirror now registers every profile under the auth_type it declares, re-syncs after discovery /
on a miss (salvaged from #101768), and owns the seam lookups: auth_handler dispatch, the
fail-loud error for a non-api-key profile that ships no handler, and refresh eligibility
derived from the profile's refresh_credential hook (never a name set).
- providers/base.py: ProviderProfile.auth_handler(action, args) (salvaged from #111610, sync only)
and refresh_credential(entry) -> rotated fields. A separate hook rather than
auth_handler("refresh", ...) because the pool holds a credential row, not an argparse namespace,
and needs tokens back rather than a bool.
- hermes_cli/auth_commands.py: add/status/logout/refresh (incl. interactive add) consult the
plugin handler before the built-in path; `auth refresh` admits plugin rows via the predicate.
- agent/credential_pool.py: _refresh_entry_impl calls the profile hook; from_dict keeps every
non-field key in extra so plugin metadata survives load -> save -> load (to_dict already wrote
it all; sanitize_borrowed_credential_payload semantics unchanged).
- hermes_cli/auth.py: config import moved below PROVIDER_REGISTRY (salvaged from #94231) so a
plugin imported during discovery never sees a partial auth module.
Built-in providers are untouched: only custom/openrouter were unregistered profiles before and
both stay in the skip list; anthropic/nous/openai-codex auth add/status/refresh output is
byte-identical in the before/after probe.
Part of #116408. Salvages #111610 (@Finn763), #101768 (@zihaofeng2001, absorbing #106361 by
@Finn763) and #94231 (@Kyzcreig).
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: zihaofeng2001 <zihaofeng2001@gmail.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
A `kind: model-provider` plugin can register a ProviderProfile and a custom
client but could not contribute an interactive login flow: its module is
imported by provider discovery, and the generic command-plugin loader skips
model-provider manifests on purpose, so `register(ctx)` is not a supported way
to add commands. The core login implementations are also provider-name keyed, so
a standards-based device-code provider had to ship a second standalone command
plugin just for login/status/logout.
Add an optional ProviderProfile.auth_handler(action, args) seam and consult it
first from the existing `hermes auth` actions (add/status/logout/refresh) —
before the known-provider gate and before the credential pool:
- truthy return = the provider owned the action; falsy = per-action fallback.
- sync or async handlers (a returned awaitable is awaited).
- a handler exception becomes a readable SystemExit naming provider + action.
- no handler at all = byte-for-byte unchanged built-in behavior, including the
existing "Unknown provider" exit.
- a registered profile that does not handle the requested action now says so
instead of reporting its provider as unknown.
No new top-level command, no manifest surface; credentials stay provider-owned.
Tests: tests/hermes_cli/test_provider_auth_seam.py drives the real `hermes auth`
parser with a fixture model-provider plugin — dispatch + argument pass-through
for all four actions, per-action fallback, async handler, handler failure,
registry lookup failure, duplicate registration (last-writer-wins), and an
untouched built-in provider. Docs gain the new hook under the model-provider
plugin guide.
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.
The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.
Refs #83080
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).
Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).
Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.
Co-authored-by: KoNit-K <konit.block@protonmail.com>
The compression config table and the yaml example described
`codex_responses_compact_threshold` as "the server compaction trigger"
without saying it is read only while `codex_responses_native: true`
(agent/native_compaction.py::native_compaction_context_management returns
early otherwise). Users set it expecting local compaction to move and
filed #101867. Name the gate in the table row, the yaml comment and the
native-compaction prose, and point at `threshold` / `threshold_tokens`
as the local trigger.
Fixes#101867
With a multi-entry openai-codex pool, a 401 `token_expired` caused by a stale
replayed `encrypted_content` blob went pool-first: each healthy entry was
force-refreshed (single-use refresh token) or benched STATUS_EXHAUSTED before
the strip in `_recover_format_errors` finally ran, so pooled users lost every
account for the bench window over a session-state problem. The single-credential
path likewise burned a forced OAuth refresh on a bearer that was fine.
`recover_after_classification` now takes the Codex stale-reasoning strip first
(after the Nous welcome-tier repair) when the 401 carries `token_expired` and
the transcript still holds `codex_reasoning_items`; the same one-shot latch
bounds it, so a real expiry pays one extra round-trip and then takes the pool /
refresh path exactly as before. The strip body moved into
`_recover_stale_codex_reasoning`, shared with the 400 `invalid_encrypted_content`
branch. Docs: one line in the OpenAI Codex path section describing the
self-heal.
Tests: pooled control with a real two-entry CredentialPool (red on the previous
ordering: pool rotated, entry benched; green now: strip runs, nobody benched);
the single-credential test now asserts no refresh is burned before the strip.
Replaces the direct transport.build_kwargs test with one driving prepare_chat_messages
(groq client strips, OpenRouter client keeps), so dropping the base_url kwarg goes red.
Three passages assumed an uncapped 1M trigger (grok 375K example, legacy tail size, the 850K -> 512K
feasibility example) and the delegation doc said children compact at the ratio only.
Composing PR-4 (intake/delivery split) with PR-5 (persisted transport_profile): the delivery
fallback to the runtime profile's unique adapter is only for sources with NO identity. A pinned
identity names the receiving bot; if that bot has no adapter it is offline and the lane fails
closed, never answering from the runtime profile's bot. Also: tests and docs reference the split
helper, not the removed _adapter_for_source.
The connector stamps `profile` on inbound and passthrough_forward frames but the
gateway never sent it back, so the connector had nothing to stamp on the NEXT
interaction of a routed chat. `_capture_scope` now remembers the routed profile
per chat, `_with_scope` echoes it as `metadata.profile` on chat-addressed frames,
and `send_follow_up` derives it from the `agent:<profile>:` key namespace. A
single-profile gateway emits no key — frames stay byte-identical. Contract §4
documents the round-trip.
Five topology rows: per-credential, shared credential → satellite, shared bot → profile with its own bot (reply via the receiving bot), secondary bot → default, and restored/synthetic (intake fails closed; delivery via the unique owner). Part of #88715 (phase 4).
Configured fallback_providers now also engage on a Codex reasoning-only stall
(incomplete_response) with one bounded grace call; the developer guide listed only the
429/5xx/401/403 triggers.
Store the catalog's max_context_window in a parallel per-token dict populated
by the same fetch instead of widening the fetch's return tuple and versioning
the cache key ("v2:"), so the one external caller and every existing cache
reset keep working. The cap moves into _apply_verified_bump as
min(900K, live max): the bump still fires only for an opted-in -900k alias
whose advertised window is exactly 272K, 900K stays the offline/absent
fallback and a catalog max above 900K does not raise it.
Trim the salvaged tests to two invariants (parametrized alias cap incl. the
absent and above-cap controls; base slug never inherits the max) and document
the live-catalog cap next to the -900k opt-in.
Two red-on-base invariants with the real GatewayRunner resolvers and real
BasePlatformAdapter ingress against a temp home with three served profiles:
per-credential collision (two bots, same chat.id == user.id: distinct lanes,
/stop on A cannot touch B's run, clarify reply on A resolves A's pending) and
shared-credential routed chat (runtime satellite, transport the receiving bot,
busy path keyed on the satellite lane; an unserved route dropped at fresh,
batched, busy and control ingress alike).
SentenceChunker hard-coded min_len=20 and all three construction sites
(gateway StreamingTTSConsumer, CLI/TUI stream_tts_to_speaker, dashboard
/api/audio/speak-stream) called SentenceChunker() with no arguments, so a
short CJK opener ("记得,叫团团。", 7 chars) was always buffered behind the
second sentence, delaying the first audible audio by roughly one LLM
sentence in voice setups. Add tts.streaming.min_len (default 20, unchanged
behaviour) to DEFAULT_CONFIG and a SentenceChunker.from_config(tts_config)
constructor that every construction site now uses, so the knob applies
identically on every speech surface. Invalid values keep the default; 0
floors to 1 so chunking cannot be disabled by accident.
Slim redo of PR #96933 by @liuhao1024 (added a second
first_sentence_min_chars key and per-emission state), credited as author.
Fixes#96927
Closes#33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.
Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.
Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.