Commit Graph

439 Commits

Author SHA1 Message Date
teknium1
e089be80a1 docs(desktop-sdk): note DecodeText loop is opt-in
DecodeText is a published plugin-SDK export; flipping its `loop` default
to false silently changes third-party plugin visuals, so the SDK doc says
so where the component is listed.
2026-09-20 20:14:20 -07:00
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
bdde0b0c28 docs(cron): describe when a stale-code ticker yields and when it keeps dispatching
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped
code_sha is the on-disk revision; a lock held by an equally stale process never
counts. Placed under "Gateway Integration" so it does not collide with the
"Stale-code yield" section #117501 adds under "Locking".
2026-09-20 20:13:19 -07:00
teknium1
13fe9c7171 feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:

- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
  `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
  declaring profile; standard records still replay on OpenRouter-style routes, strict routes
  drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
  status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
  and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
  one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
  without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".

The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.

Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
41b6ba09d9 docs: name the registered tool in prose that still says cronjob/todo/process
The same rename that left the two reference pages stale (cronjob->cronjob_manage,
todo->todo_list, process->process_manage) left the guides and developer docs
referring to the old tool names; a reader following them gets "no such tool".
Toolset names (`cronjob`, `todo`) are unchanged and left alone.
2026-09-20 12:56:25 -07:00
teknium1
297ebd506e test: trim the ambient-chain refusal tests to two two-home invariants
The 19 cherry-picked tests covered each refusal branch separately. One
A -> B -> A test per adapter over two real homes now proves the whole
contract at the production entry (build_credential / _cached_client): the
launch profile keeps its own credential, the cred-less served profile is
refused before the SDK chain (or boto3) is touched, the launch profile is
unaffected afterwards, and the standalone run keeps today's ambient chain.
Docs: the Azure guide and the multiplexing design page name the refusal.

Superseded #116370 (@JoaoMarcos44) proposed the same mechanism.
2026-09-20 12:09:47 -07:00
teknium1
948753d0f8 docs(plugins): enable prompts for the override grant only when the manifest declares capabilities 2026-09-20 11:52:45 -07:00
teknium1
967a7c94f7 docs(plugins): a provider plugin can register its own api_mode transport 2026-09-20 10:11:40 -07:00
teknium1
e5e7fbcd27 fix(auth): a user provider plugin's endpoint overrides the bundled row too
A `$HERMES_HOME/plugins/model-providers/<name>/` plugin re-registering a
bundled provider (stepfun with a regional base_url, gmi at a staging host)
wins in `providers._REGISTRY` — register_provider() is last-writer-wins and
the plugin guide promises exactly this — but the runtime reads its endpoint
from `hermes_cli.auth.PROVIDER_REGISTRY`, whose mirror loop skipped every
name already present, so inference kept going to the built-in URL (#48450).

The mirror now applies one explicit precedence rule: when a row core wrote
(built-in or plugin-mirrored) belongs to a name whose profile
currently registered came from a USER plugin,
the row's profile-derived fields are rewritten in place (inference_base_url;
api_key_env_vars / base_url_env_var on api-key rows when the profile declares
env_vars). `providers` records the discovery source per registration
(`provider_source()`), because a bundled profile must never rewrite a
built-in row: several bundled profiles omit the row's `*_BASE_URL` env var
and one differs in auth_type, so an unconditional "profile wins" would have
changed built-in behaviour. With no user plugin PROVIDER_REGISTRY is
byte-identical before/after (78 rows probed). copilot/kimi/zai keep their
bespoke resolution via the existing skip set.

Co-authored-by: xiaoxinova <xiaoxinova@users.noreply.github.com>
2026-09-20 10:06:29 -07:00
teknium1
9f0cd9d773 docs: plugin profiles and their aliases resolve in /model --provider and the model picker 2026-09-20 10:01:58 -07:00
teknium1
2182f51d7c fix: name the process holding the state.db write lock when a writer times out
"database is locked (another Hermes process held the state.db write lock for
over 60s)" identified the victim only. The open-descriptor scan cannot single
out the writer because every Hermes process (gateway, CLI sessions, worktree
agents, cron) has the DB open, so an operator hit repeatedly by
session_persistence_failed:locked had nothing to act on.

SQLite's unix VFS takes fcntl byte-range locks whose offset encodes the lock
kind (state.db-shm byte 120 = WAL write, 121 = checkpoint; the pending-byte page
on state.db = PENDING/RESERVED), and the kernel exports them with the owning pid
in /proc/locks. hermes_state_lockowners reads that table at the moment the
patience budget runs out and logs one WARNING per write-class holder with
describe_holder_pid()'s argv summary, for both the transcript write path and
open+init lock patience. The holder stays out of the exception text on purpose:
classify_persistence_error() buckets by phrase and a holder argv such as a
worktree named fix-corrupt-db would flip the bucket.

Docs: the Write Contention section still described attempt-counted retries
(_WRITE_MAX_RETRIES = 15); updated to the time budgets in force and the new log line.
2026-09-20 10:01:42 -07:00
kshitijk4poor
838214c8f7 fix(error-classifier): a profile hook asking for fallback on a terminal reason is non-retryable
turn_api_error enters the fallback walk only for ``retryable=False`` verdicts outside the
retryable-client reasons; the built-in terminal verdicts pin retryable=False while the
rate-limit family stays retryable and cascades after backoff. A ``classify_api_error`` hook
returning ``{"reason": "billing", "should_fallback": True}`` therefore retried the dead
route instead of cascading (the #116408 test passed only because its hook also set
retryable=False). Default retryable to False for such verdicts, leaving rate-limit reasons on
the built-in retry-then-fallback shape; ``RETRYABLE_CLIENT_REASONS`` moves next to the
verdict table so both modules read one set. Contract documented in the plugin guide.
2026-09-20 12:17:52 +05:30
teknium1
bce20d0b1f feat: model-provider plugins classify their own API errors via ProviderProfile.classify_api_error
A kind: model-provider plugin is loaded by providers/ discovery and never enters the
PluginManager hook lifecycle, so transform_api_error_classification was unreachable for it
without shipping a second plugin component. The profile now carries an optional
classify_api_error(error, *, status_code, error_code, message, body, model) callable,
consulted as a classifier stage right after the generic plugin hooks and only for the
provider that produced the error. None or an unknown reason leaves the built-in verdict;
built-in providers are untouched (no name table, no lifecycle change).

Also: a plugin refresh_credential returning None/empty was treated as a successful refresh
(row marked ok, stale bearer replayed up to the refresh cap). It now benches the row like a
failed refresh POST, so the loop rotates or falls to the generic sign-in copy.

Part of #116408

(cherry picked from commit b8129fd6fd6a0cf4eeee6d95a5d5d823668a306d)
2026-09-19 21:24:05 -07:00
teknium1
3e67877e1b docs: declarative OAuth (PKCE) subsection in the model-provider plugin guide
Minimal OAuthPKCEConfig example plus what Hermes owns and the enforced
security boundary, so plugin authors do not hand-roll a browser flow.

(cherry picked from commit fc4e73d5d2067102ab08823aafbb4f31b3cba785)
2026-09-19 21:15:16 -07:00
teknium1
f7240ca980 fix: admitted plugin providers are selectable and status-accurate on every surface
Picker admission (#116552) listed out-of-tree external-process and OAuth
plugin providers, but selection and status still dispatched through
provider-name tables:

- `hermes model`: no `_PROVIDER_MODEL_FLOWS` entry and no generic flow, so
  picking an admitted plugin row was a silent no-op. One generic flow in
  model_setup_flows.py, credential step keyed by the profile's auth_type
  (external_process -> launch check; oauth_* -> live pool row, else the
  `hermes auth add <name>` hint), catalog via merge_profile_catalog; main.py
  falls back to it for any registered profile missing from the table.
- `_STATUS_BY_AUTH_TYPE` had no builder for oauth_device_code/oauth_external,
  so `get_auth_status`/`list_available_providers().authenticated` stayed False
  with a live pool entry. `get_plugin_oauth_auth_status` (auth_plugin_providers
  sibling) reads the pool; gated on PLUGIN_MIRRORED_PROVIDERS so bundled OAuth
  providers keep their bespoke status bytes.
- `_external_process_auth_evidence` computed evidence for copilot-acp only, so
  `inventory._external_process_signed_in` hid every other ACP row from the
  Desktop explicit_only picker. Generic evidence = the binary resolves; the
  bundled CLI keeps its token-store chain.
- agent_init Responses-upgrade guard dropped the vendor literal redundant
  with the acp:// scheme check.
- `fetch_account_usage` bounds the plugin hook with a shared 10 s deadline
  (previously only the CLI wrapped the call; gateway/TUI awaited unbounded),
  contextvars-propagated so scoped secrets resolve; overrun -> None.

Part of #116408

(cherry picked from commit 536a4e7fcf2d1ac2b8a70bd62a69707d58ea5d1e)
2026-09-19 21:11:30 -07:00
teknium1
8b42b6e020 docs: restore the model_capabilities field row dropped in the rebase 2026-09-19 21:08:09 -07:00
Ayush Nangia
ef0bac6cdd feat(providers): plugin-declared per-model capability metadata
Model provider plugins can now declare per-model metadata (capability
booleans, context_window, model_family) through the existing
ProviderProfile registry using the canonical model_overrides schema.
Declarations patch catalog metadata without erasing unknown fields; on a
catalog miss, undeclared capability booleans stay unknown rather than
false.

Precedence (tested): explicit user model_overrides > plugin declarations
> catalog/builtin > _default fill-gap. lookup_models_dev_context shares
the same seam so auto-detected context reflects declarations, including
the dashboard /api/model/info response.

Provider aliases resolve declarations; user override lookup keeps its
existing provider-key rules. Metadata lookup may trigger the existing
lazy provider discovery (imports plugin code); the registry is
process-global and discovered once per process.

Progress on #102115.

(cherry picked from commit a2bf725597efddf5238a8d0447705cf31ac71cc8)
2026-09-19 21:08:09 -07:00
teknium1
3a045623bc docs: resolve field-table rebase conflict (auth hooks rows + catalog/vision rows) 2026-09-19 20:56:31 -07:00
teknium1
08feaa98b5 fix(image-routing): drop the provider-wide supports_vision probe; mirror registry via public types in setup test
ProviderProfile.supports_vision is documented as "the API accepts image content
inside tool-result messages" -- a provider-wide wire capability, not a per-model
user-image verdict. Treating it as the latter flipped every model on the bundled
`router` relay (and models.dev-unknown meta-ai/xiaomi models) from text to native
image parts. Per-model vision now comes from ProviderProfile.model_capabilities
(#116570) through the existing models.dev probe.

The setup-catalog test mirrored the plugin through auth._register_plugin_provider,
which #116553 renames; build the ProviderConfig from public types instead so the
test passes on both trees.
2026-09-19 20:56:31 -07:00
teknium1
9de45d9595 fix(vision): image routing honours ProviderProfile.supports_vision
`decide_image_input_mode` consulted config overrides, the managed runtime,
models.dev and Ollama, but never the registered profile's `supports_vision`
— the field the tool-result media path already trusts. A plugin that
declared vision was therefore native for tool results and text-only for
user-attached images. The declaration is now the last probe in
`_VISION_PROBES`: only an explicit True is a verdict, and per-model catalog
entries still win.

Part of #116408
2026-09-19 20:56:31 -07:00
teknium1
6f12165a94 fix(models): admit every plugin provider to the picker by slug; catalogs fall back to the profile
The CANONICAL_PROVIDERS auto-extend skipped plugin profiles by auth_type
("Non-api-key flows need bespoke picker UX"). Every in-tree non-api-key
profile already owns a hand-written row, so the set never excluded an
in-tree provider - it only hid out-of-tree external-process and OAuth
plugins from /model, hermes model and list_available_providers().
Admission is now by slug (dedupe against the built-in rows); visibility
stays gated downstream by credentials (binary resolves / auth.json /
pool entry), so an admitted row reads authenticated=False until sign-in.

- _profile_live_catalog: external_process profiles use fetch_models(),
  then fallback_models; every other non-api-key profile returns its
  fallback_models instead of None (in-tree ones declare none, so the
  built-in rows are byte-identical).
- _credential_fingerprint: external-process command/argv env overrides
  key the catalog cache, as an API key does for HTTP providers.
- Drop the "[slug] as the model" shim from the canonical picker lap;
  the profile's fallback_models is the honest catalog.
- Tests trimmed to two invariants proven red on base; docs describe how
  picker rows and the authenticated flag are derived.

Co-authored-by: ericmaddox <eric.maddox@outlook.com>
Co-authored-by: Denis Araujo <d.araujo@tjrs.jus.br>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: tobenwarrior <lucasyee999@gmail.com>
2026-09-19 20:54:06 -07:00
teknium1
627a12a90e feat(opencode-go): show Go plan windows in /usage through the profile hook
First in-tree consumer of ProviderProfile.fetch_account_usage: the OpenCode Go plan's rolling /
weekly / monthly windows (GET /zen/go/v1/usage) render in every /usage surface without adding the
provider name to the core _USAGE_FETCHERS table. Ported from #113418, which implemented the same
fetch as a core table entry; the literal endpoint (not the runtime base_url, which loses /v1 in
anthropic_messages mode) and the window mapping are theirs.

Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
2026-09-19 20:52:57 -07:00
KoNit-K
400b6c41cd feat(providers): expose plugin account usage hook 2026-09-19 20:52:57 -07:00
teknium1
6621e8aa98 docs: state the two registry-mirror exclusions, the interactive-add namespace shape, and the aux 401 boundary 2026-09-19 20:45:06 -07:00
teknium1
ec1238fa65 fix(credential-pool): plugin refresh keeps rotated tokens, runs locked, and quarantines dead grants
Independent review of the plugin refresh branch (#116553) found four gaps
between what model-provider-plugin.md promises for `refresh_credential`
and what `_refresh_entry_impl` did:

1. `replace(entry, **hook_result)` raised TypeError on any non-field key
   (`expires_in`, `token_type`, `scope` — the natural token-endpoint shape),
   the except benched the row EXHAUSTED and the pair the server had already
   rotated was dropped: for single-use refresh tokens that is a lost login.
   Field keys now go through `replace()`, everything else merges into
   `entry.extra` (mirrors `from_dict`); `None` = no rotation, mark ok.
2. Plugin providers skipped the locked single-use path, so a gateway and a
   CLI could both POST the same refresh token (`refresh_token_reused`).
   Providers with a hook now take the `_auth_store_lock` path: re-read the
   pool store, adopt a peer's usable rotation and skip the hook, else call
   it and write through. Eligibility derives from `plugin_refresh_hook()`,
   not from extending the built-in name tuple.
3. A raising hook re-benched EXHAUSTED every cooldown forever at DEBUG.
   `AuthError(relogin_required=True)` (or a grant-dead OAuth code) is now
   terminal: the row goes DEAD with a WARNING naming `hermes auth add`.
   Any other exception stays a transient bench (negative test kept).
4. `from_dict`'s extra sweep round-tripped a stray row-level `provider` key
   back onto the row on `to_dict()`; it is bookkeeping, not metadata.

New logic lives in `agent/credential_pool_plugin.py` — credential_pool.py
is at the size cap; the facade only dispatches.

Part of #116408
2026-09-19 20:45:06 -07:00
teknium1
8e5a9b16a6 docs+chore: escape bare angle-bracket placeholders in MDX table; map zihaofeng2001 email (#101768 salvage) 2026-09-19 20:45:06 -07:00
teknium1
dce233e809 feat(auth): OAuth-shaped provider plugins register, log in and refresh through their profile
An out-of-tree ProviderProfile with auth_type oauth_device_code/oauth_external loaded and inferred
but was invisible to hermes auth: _register_plugin_provider skipped every auth_type except
external_process and api_key, so resolve_provider() said "Unknown provider", `hermes auth add`
had nothing to dispatch to, the pool could not refresh its rows (REFRESHABLE_OAUTH_PROVIDERS is a
name set) and PooledCredential.from_dict dropped every extra key outside _EXTRA_KEYS on reload.

- hermes_cli/auth_plugin_providers.py (new sibling; auth.py is at the size cap): the registry
  mirror now registers every profile under the auth_type it declares, re-syncs after discovery /
  on a miss (salvaged from #101768), and owns the seam lookups: auth_handler dispatch, the
  fail-loud error for a non-api-key profile that ships no handler, and refresh eligibility
  derived from the profile's refresh_credential hook (never a name set).
- providers/base.py: ProviderProfile.auth_handler(action, args) (salvaged from #111610, sync only)
  and refresh_credential(entry) -> rotated fields. A separate hook rather than
  auth_handler("refresh", ...) because the pool holds a credential row, not an argparse namespace,
  and needs tokens back rather than a bool.
- hermes_cli/auth_commands.py: add/status/logout/refresh (incl. interactive add) consult the
  plugin handler before the built-in path; `auth refresh` admits plugin rows via the predicate.
- agent/credential_pool.py: _refresh_entry_impl calls the profile hook; from_dict keeps every
  non-field key in extra so plugin metadata survives load -> save -> load (to_dict already wrote
  it all; sanitize_borrowed_credential_payload semantics unchanged).
- hermes_cli/auth.py: config import moved below PROVIDER_REGISTRY (salvaged from #94231) so a
  plugin imported during discovery never sees a partial auth module.

Built-in providers are untouched: only custom/openrouter were unregistered profiles before and
both stay in the skip list; anthropic/nous/openai-codex auth add/status/refresh output is
byte-identical in the before/after probe.

Part of #116408. Salvages #111610 (@Finn763), #101768 (@zihaofeng2001, absorbing #106361 by
@Finn763) and #94231 (@Kyzcreig).

Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: zihaofeng2001 <zihaofeng2001@gmail.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
2026-09-19 20:45:06 -07:00
finn763
749443e292 feat(auth): let model-provider plugins own hermes auth add/status/logout/refresh
A `kind: model-provider` plugin can register a ProviderProfile and a custom
client but could not contribute an interactive login flow: its module is
imported by provider discovery, and the generic command-plugin loader skips
model-provider manifests on purpose, so `register(ctx)` is not a supported way
to add commands. The core login implementations are also provider-name keyed, so
a standards-based device-code provider had to ship a second standalone command
plugin just for login/status/logout.

Add an optional ProviderProfile.auth_handler(action, args) seam and consult it
first from the existing `hermes auth` actions (add/status/logout/refresh) —
before the known-provider gate and before the credential pool:

- truthy return = the provider owned the action; falsy = per-action fallback.
- sync or async handlers (a returned awaitable is awaited).
- a handler exception becomes a readable SystemExit naming provider + action.
- no handler at all = byte-for-byte unchanged built-in behavior, including the
  existing "Unknown provider" exit.
- a registered profile that does not handle the requested action now says so
  instead of reporting its provider as unknown.

No new top-level command, no manifest surface; credentials stay provider-owned.

Tests: tests/hermes_cli/test_provider_auth_seam.py drives the real `hermes auth`
parser with a fixture model-provider plugin — dispatch + argument pass-through
for all four actions, per-action fallback, async handler, handler failure,
registry lookup failure, duplicate registration (last-writer-wins), and an
untouched built-in provider. Docs gain the new hook under the model-provider
plugin guide.
2026-09-19 20:45:06 -07:00
Teknium
b6f8f8eb1f Merge pull request #116345 from NousResearch/boa-w3-small-b
Video generation tools no longer let the agent pick the model; video_gen.model is the only selector (Refs #83080)
2026-09-19 14:32:37 -07:00
teknium1
19b29df13b fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.

The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.

Refs #83080
2026-09-19 12:22:44 -07:00
teknium1
03cfc36f51 docs: describe the session.create model×provider refusal
Hosts speaking the gateway protocol need to know the new -32602 shape
(error.data model/provider/suggestions) and which pairs stay permissive.
2026-09-19 12:12:05 -07:00
Teknium
4fda3bbdca Merge pull request #115846 from NousResearch/fix/boa-tts-stt-voice-tts-config-minlen
feat(tts): streaming TTS speaks a short first sentence sooner via tts.streaming.min_len on every surface (#96927, salvage #96933)
2026-09-19 11:33:15 -07:00
Teknium
0d2cb2cacc Merge pull request #115890 from NousResearch/fix/boa-res-R6-streaming-compaction-astra-oauth-gate
fix(compression): gpt-6-astra on Codex OAuth gets native server-side compaction when opted in (#103720, salvage #103718)
2026-09-19 11:28:10 -07:00
teknium1
a169438178 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	hermes_cli/config_defaults.py
2026-09-19 11:22:01 -07:00
Teknium
7feb1af028 Merge pull request #115872 from NousResearch/fix/boa-res-R4-routing-catalog-azure-codex-900k-cap
fix(codex): opted-in -900k aliases cap at the live catalog max_context_window (#105443, salvage #105445)
2026-09-19 11:12:31 -07:00
teknium1
ec2c4eeb80 chore: merge origin/main (resolve apps/desktop/src/lib/voice-client-direct.ts, hermes_cli/config_defaults.py, tools/voice_client_config.py, website/docs/user-guide/features/tts.md) 2026-09-19 10:54:38 -07:00
teknium1
e85cb94da1 chore: merge origin/main (resolve agent/error_classifier.py,tests/agent/test_error_classifier.py) 2026-09-19 10:51:40 -07:00
teknium1
8e3a833365 chore: merge origin/main (resolve website/docs/developer-guide/context-compression-and-caching.md) 2026-09-19 10:47:48 -07:00
teknium1
a516db44e7 Merge remote-tracking branch 'origin/main' into HEAD 2026-09-19 10:46:59 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
4c4b0b748e docs: describe the summary-provider-overload abort in the compression failure ladder 2026-09-19 10:35:44 -07:00
teknium1
801e3fa6dd fix: truncated compaction summaries back off 60s/300s/900s across turns
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).

Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-19 10:30:09 -07:00
teknium1
26675fbcb7 docs: state that codex_responses_compact_threshold applies only under codex_responses_native
The compression config table and the yaml example described
`codex_responses_compact_threshold` as "the server compaction trigger"
without saying it is read only while `codex_responses_native: true`
(agent/native_compaction.py::native_compaction_context_management returns
early otherwise). Users set it expecting local compaction to move and
filed #101867. Name the gate in the table row, the yaml comment and the
native-compaction prose, and point at `threshold` / `threshold_tokens`
as the local trigger.

Fixes #101867
2026-09-19 10:27:28 -07:00
teknium1
fbbc7dfe5d fix: strip stale Codex reasoning on 401 token_expired before the credential pool
With a multi-entry openai-codex pool, a 401 `token_expired` caused by a stale
replayed `encrypted_content` blob went pool-first: each healthy entry was
force-refreshed (single-use refresh token) or benched STATUS_EXHAUSTED before
the strip in `_recover_format_errors` finally ran, so pooled users lost every
account for the bench window over a session-state problem. The single-credential
path likewise burned a forced OAuth refresh on a bearer that was fine.

`recover_after_classification` now takes the Codex stale-reasoning strip first
(after the Nous welcome-tier repair) when the 401 carries `token_expired` and
the transcript still holds `codex_reasoning_items`; the same one-shot latch
bounds it, so a real expiry pays one extra round-trip and then takes the pool /
refresh path exactly as before. The strip body moved into
`_recover_stale_codex_reasoning`, shared with the 400 `invalid_encrypted_content`
branch. Docs: one line in the OpenAI Codex path section describing the
self-heal.

Tests: pooled control with a real two-entry CredentialPool (red on the previous
ordering: pool rotated, entry benched; green now: strip runs, nobody benched);
the single-credential test now asserts no refresh is burned before the strip.
2026-09-19 09:32:00 -07:00
teknium1
0c24dd56cb fix: pin the auxiliary wire's route-scoped reasoning_details strip; document OpenRouter/Nous-only replay
Replaces the direct transport.build_kwargs test with one driving prepare_chat_messages
(groq client strips, OpenRouter client keeps), so dropping the base_url kwarg goes red.
2026-09-19 09:28:45 -07:00
kshitijk4poor
7c6f21a5e1 docs(compression): examples and the delegation trigger note follow the default threshold_tokens cap
Three passages assumed an uncapped 1M trigger (grok 375K example, legacy tail size, the 850K -> 512K
feasibility example) and the delegation doc said children compact at the ratio only.
2026-09-19 16:28:57 +05:30
kshitijk4poor
cfbc6e51f4 docs(compression): developer guide lists the threshold_tokens cap next to threshold 2026-09-19 16:28:57 +05:30
teknium1
d899d890f8 fix(gateway): a restored lane whose receiving bot is offline delivers nowhere
Composing PR-4 (intake/delivery split) with PR-5 (persisted transport_profile): the delivery
fallback to the runtime profile's unique adapter is only for sources with NO identity. A pinned
identity names the receiving bot; if that bot has no adapter it is offline and the lane fails
closed, never answering from the runtime profile's bot. Also: tests and docs reference the split
helper, not the removed _adapter_for_source.
2026-09-19 02:28:50 -07:00
teknium1
04a151c3e0 docs(gateway): identity survives restore, relay, callbacks and thread hops 2026-09-19 02:28:50 -07:00
teknium1
e09cd0d2bc fix(relay): echo the routed profile on every outbound frame and follow_up
The connector stamps `profile` on inbound and passthrough_forward frames but the
gateway never sent it back, so the connector had nothing to stamp on the NEXT
interaction of a routed chat. `_capture_scope` now remembers the routed profile
per chat, `_with_scope` echoes it as `metadata.profile` on chat-addressed frames,
and `send_follow_up` derives it from the `agent:<profile>:` key namespace. A
single-profile gateway emits no key — frames stay byte-identical. Contract §4
documents the round-trip.
2026-09-19 02:28:50 -07:00