Commit Graph

432 Commits

Author SHA1 Message Date
teknium1
967a7c94f7 docs(plugins): a provider plugin can register its own api_mode transport 2026-09-20 10:11:40 -07:00
teknium1
e5e7fbcd27 fix(auth): a user provider plugin's endpoint overrides the bundled row too
A `$HERMES_HOME/plugins/model-providers/<name>/` plugin re-registering a
bundled provider (stepfun with a regional base_url, gmi at a staging host)
wins in `providers._REGISTRY` — register_provider() is last-writer-wins and
the plugin guide promises exactly this — but the runtime reads its endpoint
from `hermes_cli.auth.PROVIDER_REGISTRY`, whose mirror loop skipped every
name already present, so inference kept going to the built-in URL (#48450).

The mirror now applies one explicit precedence rule: when a row core wrote
(built-in or plugin-mirrored) belongs to a name whose profile
currently registered came from a USER plugin,
the row's profile-derived fields are rewritten in place (inference_base_url;
api_key_env_vars / base_url_env_var on api-key rows when the profile declares
env_vars). `providers` records the discovery source per registration
(`provider_source()`), because a bundled profile must never rewrite a
built-in row: several bundled profiles omit the row's `*_BASE_URL` env var
and one differs in auth_type, so an unconditional "profile wins" would have
changed built-in behaviour. With no user plugin PROVIDER_REGISTRY is
byte-identical before/after (78 rows probed). copilot/kimi/zai keep their
bespoke resolution via the existing skip set.

Co-authored-by: xiaoxinova <xiaoxinova@users.noreply.github.com>
2026-09-20 10:06:29 -07:00
teknium1
9f0cd9d773 docs: plugin profiles and their aliases resolve in /model --provider and the model picker 2026-09-20 10:01:58 -07:00
teknium1
2182f51d7c fix: name the process holding the state.db write lock when a writer times out
"database is locked (another Hermes process held the state.db write lock for
over 60s)" identified the victim only. The open-descriptor scan cannot single
out the writer because every Hermes process (gateway, CLI sessions, worktree
agents, cron) has the DB open, so an operator hit repeatedly by
session_persistence_failed:locked had nothing to act on.

SQLite's unix VFS takes fcntl byte-range locks whose offset encodes the lock
kind (state.db-shm byte 120 = WAL write, 121 = checkpoint; the pending-byte page
on state.db = PENDING/RESERVED), and the kernel exports them with the owning pid
in /proc/locks. hermes_state_lockowners reads that table at the moment the
patience budget runs out and logs one WARNING per write-class holder with
describe_holder_pid()'s argv summary, for both the transcript write path and
open+init lock patience. The holder stays out of the exception text on purpose:
classify_persistence_error() buckets by phrase and a holder argv such as a
worktree named fix-corrupt-db would flip the bucket.

Docs: the Write Contention section still described attempt-counted retries
(_WRITE_MAX_RETRIES = 15); updated to the time budgets in force and the new log line.
2026-09-20 10:01:42 -07:00
kshitijk4poor
838214c8f7 fix(error-classifier): a profile hook asking for fallback on a terminal reason is non-retryable
turn_api_error enters the fallback walk only for ``retryable=False`` verdicts outside the
retryable-client reasons; the built-in terminal verdicts pin retryable=False while the
rate-limit family stays retryable and cascades after backoff. A ``classify_api_error`` hook
returning ``{"reason": "billing", "should_fallback": True}`` therefore retried the dead
route instead of cascading (the #116408 test passed only because its hook also set
retryable=False). Default retryable to False for such verdicts, leaving rate-limit reasons on
the built-in retry-then-fallback shape; ``RETRYABLE_CLIENT_REASONS`` moves next to the
verdict table so both modules read one set. Contract documented in the plugin guide.
2026-09-20 12:17:52 +05:30
teknium1
bce20d0b1f feat: model-provider plugins classify their own API errors via ProviderProfile.classify_api_error
A kind: model-provider plugin is loaded by providers/ discovery and never enters the
PluginManager hook lifecycle, so transform_api_error_classification was unreachable for it
without shipping a second plugin component. The profile now carries an optional
classify_api_error(error, *, status_code, error_code, message, body, model) callable,
consulted as a classifier stage right after the generic plugin hooks and only for the
provider that produced the error. None or an unknown reason leaves the built-in verdict;
built-in providers are untouched (no name table, no lifecycle change).

Also: a plugin refresh_credential returning None/empty was treated as a successful refresh
(row marked ok, stale bearer replayed up to the refresh cap). It now benches the row like a
failed refresh POST, so the loop rotates or falls to the generic sign-in copy.

Part of #116408

(cherry picked from commit b8129fd6fd6a0cf4eeee6d95a5d5d823668a306d)
2026-09-19 21:24:05 -07:00
teknium1
3e67877e1b docs: declarative OAuth (PKCE) subsection in the model-provider plugin guide
Minimal OAuthPKCEConfig example plus what Hermes owns and the enforced
security boundary, so plugin authors do not hand-roll a browser flow.

(cherry picked from commit fc4e73d5d2067102ab08823aafbb4f31b3cba785)
2026-09-19 21:15:16 -07:00
teknium1
f7240ca980 fix: admitted plugin providers are selectable and status-accurate on every surface
Picker admission (#116552) listed out-of-tree external-process and OAuth
plugin providers, but selection and status still dispatched through
provider-name tables:

- `hermes model`: no `_PROVIDER_MODEL_FLOWS` entry and no generic flow, so
  picking an admitted plugin row was a silent no-op. One generic flow in
  model_setup_flows.py, credential step keyed by the profile's auth_type
  (external_process -> launch check; oauth_* -> live pool row, else the
  `hermes auth add <name>` hint), catalog via merge_profile_catalog; main.py
  falls back to it for any registered profile missing from the table.
- `_STATUS_BY_AUTH_TYPE` had no builder for oauth_device_code/oauth_external,
  so `get_auth_status`/`list_available_providers().authenticated` stayed False
  with a live pool entry. `get_plugin_oauth_auth_status` (auth_plugin_providers
  sibling) reads the pool; gated on PLUGIN_MIRRORED_PROVIDERS so bundled OAuth
  providers keep their bespoke status bytes.
- `_external_process_auth_evidence` computed evidence for copilot-acp only, so
  `inventory._external_process_signed_in` hid every other ACP row from the
  Desktop explicit_only picker. Generic evidence = the binary resolves; the
  bundled CLI keeps its token-store chain.
- agent_init Responses-upgrade guard dropped the vendor literal redundant
  with the acp:// scheme check.
- `fetch_account_usage` bounds the plugin hook with a shared 10 s deadline
  (previously only the CLI wrapped the call; gateway/TUI awaited unbounded),
  contextvars-propagated so scoped secrets resolve; overrun -> None.

Part of #116408

(cherry picked from commit 536a4e7fcf2d1ac2b8a70bd62a69707d58ea5d1e)
2026-09-19 21:11:30 -07:00
teknium1
8b42b6e020 docs: restore the model_capabilities field row dropped in the rebase 2026-09-19 21:08:09 -07:00
Ayush Nangia
ef0bac6cdd feat(providers): plugin-declared per-model capability metadata
Model provider plugins can now declare per-model metadata (capability
booleans, context_window, model_family) through the existing
ProviderProfile registry using the canonical model_overrides schema.
Declarations patch catalog metadata without erasing unknown fields; on a
catalog miss, undeclared capability booleans stay unknown rather than
false.

Precedence (tested): explicit user model_overrides > plugin declarations
> catalog/builtin > _default fill-gap. lookup_models_dev_context shares
the same seam so auto-detected context reflects declarations, including
the dashboard /api/model/info response.

Provider aliases resolve declarations; user override lookup keeps its
existing provider-key rules. Metadata lookup may trigger the existing
lazy provider discovery (imports plugin code); the registry is
process-global and discovered once per process.

Progress on #102115.

(cherry picked from commit a2bf725597efddf5238a8d0447705cf31ac71cc8)
2026-09-19 21:08:09 -07:00
teknium1
3a045623bc docs: resolve field-table rebase conflict (auth hooks rows + catalog/vision rows) 2026-09-19 20:56:31 -07:00
teknium1
08feaa98b5 fix(image-routing): drop the provider-wide supports_vision probe; mirror registry via public types in setup test
ProviderProfile.supports_vision is documented as "the API accepts image content
inside tool-result messages" -- a provider-wide wire capability, not a per-model
user-image verdict. Treating it as the latter flipped every model on the bundled
`router` relay (and models.dev-unknown meta-ai/xiaomi models) from text to native
image parts. Per-model vision now comes from ProviderProfile.model_capabilities
(#116570) through the existing models.dev probe.

The setup-catalog test mirrored the plugin through auth._register_plugin_provider,
which #116553 renames; build the ProviderConfig from public types instead so the
test passes on both trees.
2026-09-19 20:56:31 -07:00
teknium1
9de45d9595 fix(vision): image routing honours ProviderProfile.supports_vision
`decide_image_input_mode` consulted config overrides, the managed runtime,
models.dev and Ollama, but never the registered profile's `supports_vision`
— the field the tool-result media path already trusts. A plugin that
declared vision was therefore native for tool results and text-only for
user-attached images. The declaration is now the last probe in
`_VISION_PROBES`: only an explicit True is a verdict, and per-model catalog
entries still win.

Part of #116408
2026-09-19 20:56:31 -07:00
teknium1
6f12165a94 fix(models): admit every plugin provider to the picker by slug; catalogs fall back to the profile
The CANONICAL_PROVIDERS auto-extend skipped plugin profiles by auth_type
("Non-api-key flows need bespoke picker UX"). Every in-tree non-api-key
profile already owns a hand-written row, so the set never excluded an
in-tree provider - it only hid out-of-tree external-process and OAuth
plugins from /model, hermes model and list_available_providers().
Admission is now by slug (dedupe against the built-in rows); visibility
stays gated downstream by credentials (binary resolves / auth.json /
pool entry), so an admitted row reads authenticated=False until sign-in.

- _profile_live_catalog: external_process profiles use fetch_models(),
  then fallback_models; every other non-api-key profile returns its
  fallback_models instead of None (in-tree ones declare none, so the
  built-in rows are byte-identical).
- _credential_fingerprint: external-process command/argv env overrides
  key the catalog cache, as an API key does for HTTP providers.
- Drop the "[slug] as the model" shim from the canonical picker lap;
  the profile's fallback_models is the honest catalog.
- Tests trimmed to two invariants proven red on base; docs describe how
  picker rows and the authenticated flag are derived.

Co-authored-by: ericmaddox <eric.maddox@outlook.com>
Co-authored-by: Denis Araujo <d.araujo@tjrs.jus.br>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: tobenwarrior <lucasyee999@gmail.com>
2026-09-19 20:54:06 -07:00
teknium1
627a12a90e feat(opencode-go): show Go plan windows in /usage through the profile hook
First in-tree consumer of ProviderProfile.fetch_account_usage: the OpenCode Go plan's rolling /
weekly / monthly windows (GET /zen/go/v1/usage) render in every /usage surface without adding the
provider name to the core _USAGE_FETCHERS table. Ported from #113418, which implemented the same
fetch as a core table entry; the literal endpoint (not the runtime base_url, which loses /v1 in
anthropic_messages mode) and the window mapping are theirs.

Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
2026-09-19 20:52:57 -07:00
KoNit-K
400b6c41cd feat(providers): expose plugin account usage hook 2026-09-19 20:52:57 -07:00
teknium1
6621e8aa98 docs: state the two registry-mirror exclusions, the interactive-add namespace shape, and the aux 401 boundary 2026-09-19 20:45:06 -07:00
teknium1
ec1238fa65 fix(credential-pool): plugin refresh keeps rotated tokens, runs locked, and quarantines dead grants
Independent review of the plugin refresh branch (#116553) found four gaps
between what model-provider-plugin.md promises for `refresh_credential`
and what `_refresh_entry_impl` did:

1. `replace(entry, **hook_result)` raised TypeError on any non-field key
   (`expires_in`, `token_type`, `scope` — the natural token-endpoint shape),
   the except benched the row EXHAUSTED and the pair the server had already
   rotated was dropped: for single-use refresh tokens that is a lost login.
   Field keys now go through `replace()`, everything else merges into
   `entry.extra` (mirrors `from_dict`); `None` = no rotation, mark ok.
2. Plugin providers skipped the locked single-use path, so a gateway and a
   CLI could both POST the same refresh token (`refresh_token_reused`).
   Providers with a hook now take the `_auth_store_lock` path: re-read the
   pool store, adopt a peer's usable rotation and skip the hook, else call
   it and write through. Eligibility derives from `plugin_refresh_hook()`,
   not from extending the built-in name tuple.
3. A raising hook re-benched EXHAUSTED every cooldown forever at DEBUG.
   `AuthError(relogin_required=True)` (or a grant-dead OAuth code) is now
   terminal: the row goes DEAD with a WARNING naming `hermes auth add`.
   Any other exception stays a transient bench (negative test kept).
4. `from_dict`'s extra sweep round-tripped a stray row-level `provider` key
   back onto the row on `to_dict()`; it is bookkeeping, not metadata.

New logic lives in `agent/credential_pool_plugin.py` — credential_pool.py
is at the size cap; the facade only dispatches.

Part of #116408
2026-09-19 20:45:06 -07:00
teknium1
8e5a9b16a6 docs+chore: escape bare angle-bracket placeholders in MDX table; map zihaofeng2001 email (#101768 salvage) 2026-09-19 20:45:06 -07:00
teknium1
dce233e809 feat(auth): OAuth-shaped provider plugins register, log in and refresh through their profile
An out-of-tree ProviderProfile with auth_type oauth_device_code/oauth_external loaded and inferred
but was invisible to hermes auth: _register_plugin_provider skipped every auth_type except
external_process and api_key, so resolve_provider() said "Unknown provider", `hermes auth add`
had nothing to dispatch to, the pool could not refresh its rows (REFRESHABLE_OAUTH_PROVIDERS is a
name set) and PooledCredential.from_dict dropped every extra key outside _EXTRA_KEYS on reload.

- hermes_cli/auth_plugin_providers.py (new sibling; auth.py is at the size cap): the registry
  mirror now registers every profile under the auth_type it declares, re-syncs after discovery /
  on a miss (salvaged from #101768), and owns the seam lookups: auth_handler dispatch, the
  fail-loud error for a non-api-key profile that ships no handler, and refresh eligibility
  derived from the profile's refresh_credential hook (never a name set).
- providers/base.py: ProviderProfile.auth_handler(action, args) (salvaged from #111610, sync only)
  and refresh_credential(entry) -> rotated fields. A separate hook rather than
  auth_handler("refresh", ...) because the pool holds a credential row, not an argparse namespace,
  and needs tokens back rather than a bool.
- hermes_cli/auth_commands.py: add/status/logout/refresh (incl. interactive add) consult the
  plugin handler before the built-in path; `auth refresh` admits plugin rows via the predicate.
- agent/credential_pool.py: _refresh_entry_impl calls the profile hook; from_dict keeps every
  non-field key in extra so plugin metadata survives load -> save -> load (to_dict already wrote
  it all; sanitize_borrowed_credential_payload semantics unchanged).
- hermes_cli/auth.py: config import moved below PROVIDER_REGISTRY (salvaged from #94231) so a
  plugin imported during discovery never sees a partial auth module.

Built-in providers are untouched: only custom/openrouter were unregistered profiles before and
both stay in the skip list; anthropic/nous/openai-codex auth add/status/refresh output is
byte-identical in the before/after probe.

Part of #116408. Salvages #111610 (@Finn763), #101768 (@zihaofeng2001, absorbing #106361 by
@Finn763) and #94231 (@Kyzcreig).

Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: zihaofeng2001 <zihaofeng2001@gmail.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
2026-09-19 20:45:06 -07:00
finn763
749443e292 feat(auth): let model-provider plugins own hermes auth add/status/logout/refresh
A `kind: model-provider` plugin can register a ProviderProfile and a custom
client but could not contribute an interactive login flow: its module is
imported by provider discovery, and the generic command-plugin loader skips
model-provider manifests on purpose, so `register(ctx)` is not a supported way
to add commands. The core login implementations are also provider-name keyed, so
a standards-based device-code provider had to ship a second standalone command
plugin just for login/status/logout.

Add an optional ProviderProfile.auth_handler(action, args) seam and consult it
first from the existing `hermes auth` actions (add/status/logout/refresh) —
before the known-provider gate and before the credential pool:

- truthy return = the provider owned the action; falsy = per-action fallback.
- sync or async handlers (a returned awaitable is awaited).
- a handler exception becomes a readable SystemExit naming provider + action.
- no handler at all = byte-for-byte unchanged built-in behavior, including the
  existing "Unknown provider" exit.
- a registered profile that does not handle the requested action now says so
  instead of reporting its provider as unknown.

No new top-level command, no manifest surface; credentials stay provider-owned.

Tests: tests/hermes_cli/test_provider_auth_seam.py drives the real `hermes auth`
parser with a fixture model-provider plugin — dispatch + argument pass-through
for all four actions, per-action fallback, async handler, handler failure,
registry lookup failure, duplicate registration (last-writer-wins), and an
untouched built-in provider. Docs gain the new hook under the model-provider
plugin guide.
2026-09-19 20:45:06 -07:00
Teknium
b6f8f8eb1f Merge pull request #116345 from NousResearch/boa-w3-small-b
Video generation tools no longer let the agent pick the model; video_gen.model is the only selector (Refs #83080)
2026-09-19 14:32:37 -07:00
teknium1
19b29df13b fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.

The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.

Refs #83080
2026-09-19 12:22:44 -07:00
teknium1
03cfc36f51 docs: describe the session.create model×provider refusal
Hosts speaking the gateway protocol need to know the new -32602 shape
(error.data model/provider/suggestions) and which pairs stay permissive.
2026-09-19 12:12:05 -07:00
Teknium
4fda3bbdca Merge pull request #115846 from NousResearch/fix/boa-tts-stt-voice-tts-config-minlen
feat(tts): streaming TTS speaks a short first sentence sooner via tts.streaming.min_len on every surface (#96927, salvage #96933)
2026-09-19 11:33:15 -07:00
Teknium
0d2cb2cacc Merge pull request #115890 from NousResearch/fix/boa-res-R6-streaming-compaction-astra-oauth-gate
fix(compression): gpt-6-astra on Codex OAuth gets native server-side compaction when opted in (#103720, salvage #103718)
2026-09-19 11:28:10 -07:00
teknium1
a169438178 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	hermes_cli/config_defaults.py
2026-09-19 11:22:01 -07:00
Teknium
7feb1af028 Merge pull request #115872 from NousResearch/fix/boa-res-R4-routing-catalog-azure-codex-900k-cap
fix(codex): opted-in -900k aliases cap at the live catalog max_context_window (#105443, salvage #105445)
2026-09-19 11:12:31 -07:00
teknium1
ec2c4eeb80 chore: merge origin/main (resolve apps/desktop/src/lib/voice-client-direct.ts, hermes_cli/config_defaults.py, tools/voice_client_config.py, website/docs/user-guide/features/tts.md) 2026-09-19 10:54:38 -07:00
teknium1
e85cb94da1 chore: merge origin/main (resolve agent/error_classifier.py,tests/agent/test_error_classifier.py) 2026-09-19 10:51:40 -07:00
teknium1
8e3a833365 chore: merge origin/main (resolve website/docs/developer-guide/context-compression-and-caching.md) 2026-09-19 10:47:48 -07:00
teknium1
a516db44e7 Merge remote-tracking branch 'origin/main' into HEAD 2026-09-19 10:46:59 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
4c4b0b748e docs: describe the summary-provider-overload abort in the compression failure ladder 2026-09-19 10:35:44 -07:00
teknium1
801e3fa6dd fix: truncated compaction summaries back off 60s/300s/900s across turns
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).

Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-19 10:30:09 -07:00
teknium1
26675fbcb7 docs: state that codex_responses_compact_threshold applies only under codex_responses_native
The compression config table and the yaml example described
`codex_responses_compact_threshold` as "the server compaction trigger"
without saying it is read only while `codex_responses_native: true`
(agent/native_compaction.py::native_compaction_context_management returns
early otherwise). Users set it expecting local compaction to move and
filed #101867. Name the gate in the table row, the yaml comment and the
native-compaction prose, and point at `threshold` / `threshold_tokens`
as the local trigger.

Fixes #101867
2026-09-19 10:27:28 -07:00
teknium1
fbbc7dfe5d fix: strip stale Codex reasoning on 401 token_expired before the credential pool
With a multi-entry openai-codex pool, a 401 `token_expired` caused by a stale
replayed `encrypted_content` blob went pool-first: each healthy entry was
force-refreshed (single-use refresh token) or benched STATUS_EXHAUSTED before
the strip in `_recover_format_errors` finally ran, so pooled users lost every
account for the bench window over a session-state problem. The single-credential
path likewise burned a forced OAuth refresh on a bearer that was fine.

`recover_after_classification` now takes the Codex stale-reasoning strip first
(after the Nous welcome-tier repair) when the 401 carries `token_expired` and
the transcript still holds `codex_reasoning_items`; the same one-shot latch
bounds it, so a real expiry pays one extra round-trip and then takes the pool /
refresh path exactly as before. The strip body moved into
`_recover_stale_codex_reasoning`, shared with the 400 `invalid_encrypted_content`
branch. Docs: one line in the OpenAI Codex path section describing the
self-heal.

Tests: pooled control with a real two-entry CredentialPool (red on the previous
ordering: pool rotated, entry benched; green now: strip runs, nobody benched);
the single-credential test now asserts no refresh is burned before the strip.
2026-09-19 09:32:00 -07:00
teknium1
0c24dd56cb fix: pin the auxiliary wire's route-scoped reasoning_details strip; document OpenRouter/Nous-only replay
Replaces the direct transport.build_kwargs test with one driving prepare_chat_messages
(groq client strips, OpenRouter client keeps), so dropping the base_url kwarg goes red.
2026-09-19 09:28:45 -07:00
kshitijk4poor
7c6f21a5e1 docs(compression): examples and the delegation trigger note follow the default threshold_tokens cap
Three passages assumed an uncapped 1M trigger (grok 375K example, legacy tail size, the 850K -> 512K
feasibility example) and the delegation doc said children compact at the ratio only.
2026-09-19 16:28:57 +05:30
kshitijk4poor
cfbc6e51f4 docs(compression): developer guide lists the threshold_tokens cap next to threshold 2026-09-19 16:28:57 +05:30
teknium1
d899d890f8 fix(gateway): a restored lane whose receiving bot is offline delivers nowhere
Composing PR-4 (intake/delivery split) with PR-5 (persisted transport_profile): the delivery
fallback to the runtime profile's unique adapter is only for sources with NO identity. A pinned
identity names the receiving bot; if that bot has no adapter it is offline and the lane fails
closed, never answering from the runtime profile's bot. Also: tests and docs reference the split
helper, not the removed _adapter_for_source.
2026-09-19 02:28:50 -07:00
teknium1
04a151c3e0 docs(gateway): identity survives restore, relay, callbacks and thread hops 2026-09-19 02:28:50 -07:00
teknium1
e09cd0d2bc fix(relay): echo the routed profile on every outbound frame and follow_up
The connector stamps `profile` on inbound and passthrough_forward frames but the
gateway never sent it back, so the connector had nothing to stamp on the NEXT
interaction of a routed chat. `_capture_scope` now remembers the routed profile
per chat, `_with_scope` echoes it as `metadata.profile` on chat-addressed frames,
and `send_follow_up` derives it from the `agent:<profile>:` key namespace. A
single-profile gateway emits no key — frames stay byte-identical. Contract §4
documents the round-trip.
2026-09-19 02:28:50 -07:00
teknium1
19ca2e505a docs(gateway): document the intake vs delivery transport matrix
Five topology rows: per-credential, shared credential → satellite, shared bot → profile with its own bot (reply via the receiving bot), secondary bot → default, and restored/synthetic (intake fails closed; delivery via the unique owner). Part of #88715 (phase 4).
2026-09-19 01:28:23 -07:00
teknium1
2c4d1f79e2 fix: document the reasoning-only stall failover under Fallback Model
Configured fallback_providers now also engage on a Codex reasoning-only stall
(incomplete_response) with one bounded grace call; the developer guide listed only the
429/5xx/401/403 triggers.
2026-09-19 01:15:11 -07:00
GodsBoy
33b2461be4 fix(compression): enable Astra native compaction on official Codex OAuth 2026-09-19 00:53:26 -07:00
teknium1
a585ad83b2 fix: keep the Codex catalog fetch shape while capping -900k at max_context_window
Store the catalog's max_context_window in a parallel per-token dict populated
by the same fetch instead of widening the fetch's return tuple and versioning
the cache key ("v2:"), so the one external caller and every existing cache
reset keep working. The cap moves into _apply_verified_bump as
min(900K, live max): the bump still fires only for an opted-in -900k alias
whose advertised window is exactly 272K, 900K stays the offline/absent
fallback and a catalog max above 900K does not raise it.

Trim the salvaged tests to two invariants (parametrized alias cap incl. the
absent and above-cap controls; base slug never inherits the max) and document
the live-catalog cap next to the -900k opt-in.
2026-09-19 00:40:24 -07:00
teknium1
b408544510 test(gateway): identity-canonical ingress invariants (#88715 rows A/B) + docs
Two red-on-base invariants with the real GatewayRunner resolvers and real
BasePlatformAdapter ingress against a temp home with three served profiles:
per-credential collision (two bots, same chat.id == user.id: distinct lanes,
/stop on A cannot touch B's run, clarify reply on A resolves A's pending) and
shared-credential routed chat (runtime satellite, transport the receiving bot,
busy path keyed on the satellite lane; an unserved route dropped at fresh,
batched, busy and control ingress alike).
2026-09-19 00:30:29 -07:00
liuhao1024
9823ce4ff1 feat: make the streaming TTS first-sentence threshold configurable (tts.streaming.min_len)
SentenceChunker hard-coded min_len=20 and all three construction sites
(gateway StreamingTTSConsumer, CLI/TUI stream_tts_to_speaker, dashboard
/api/audio/speak-stream) called SentenceChunker() with no arguments, so a
short CJK opener ("记得,叫团团。", 7 chars) was always buffered behind the
second sentence, delaying the first audible audio by roughly one LLM
sentence in voice setups. Add tts.streaming.min_len (default 20, unchanged
behaviour) to DEFAULT_CONFIG and a SentenceChunker.from_config(tts_config)
constructor that every construction site now uses, so the knob applies
identically on every speech surface. Invalid values keep the default; 0
floors to 1 so chunking cannot be disabled by accident.

Slim redo of PR #96933 by @liuhao1024 (added a second
first_sentence_min_chars key and per-emission state), credited as author.

Fixes #96927
2026-09-19 00:16:47 -07:00
Italo Fernandes
d6f3285b81 feat(gateway): route inbound messages to profiles by sender user_id
Closes #33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.

Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.

Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.
2026-09-18 22:36:24 -07:00