Commit Graph

5494 Commits

Author SHA1 Message Date
teknium1
1e9d402e07 fix: LSP tree-kill runs off the event loop (review follow-up)
kill_process_tree is synchronous (taskkill /T /F with a 15s timeout on
Windows), so the hard-kill in _cleanup_process blocked the loop for the
duration. Run it via asyncio.to_thread. The graceful path is unchanged:
shutdown() already waits SHUTDOWN_GRACE on proc.wait() after `exit`, so a
well-behaved server exits 0 before cleanup ever reaches the kill (probe:
returncode 0, no kill_process_tree call).
2026-09-20 12:54:18 -07:00
fangliquan
b773bcf271 fix(lsp): hard-kill failed server trees before reaping 2026-09-20 12:54:18 -07:00
fangliquan
bf52a9519a fix(lsp): reap servers cancelled during startup 2026-09-20 12:54:18 -07:00
teknium1
d03d6c2b39 fix(compression): an over-window session that cannot shrink ends the turn with /new guidance and waits one idle budget, not the ceiling
A session far above the model window (~356k tokens on a 131k window in
#116472) re-ran context compression on every turn: a preflight pass that
reclaimed nothing still let the request go to the provider (400 -> overflow
handler -> another pass), and a summary stream that kept emitting tokens
while never committing held the pre-commit wait to the full 600s ceiling.
On the Desktop that blocked the gateway event loop for 10-20 minutes per
turn and the renderer was eventually killed.

- agent/turn_context.py::_fail_closed_on_insufficient_progress: when a
  preflight pass makes no (or sub-5%) progress and the request provably
  exceeds the model window, raise PreflightCompressionTimedOut with
  "start a new session (/new)" guidance so no provider call is sent. An
  unknown window or a fitting request keeps the send-as-is behaviour; a
  pass that no-op'd on a transient guard (summary-failure cooldown) keeps
  its typed cooldown result. Called from both insufficient-progress
  branches of turn_context_compaction._run_preflight_passes.
- agent/conversation_compression.py::run_compress_context_with_progress_timeout:
  an over-window request's pre-commit wait is bounded by one inactivity
  budget (compression.context_timeout_seconds) instead of
  context_total_ceiling_seconds; the existing first-stall deterministic
  fallback then carries the compaction. Config-derived, no new knob.

Slim slice of #116592's Python half.

Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
2026-09-20 12:52:07 -07:00
teknium1
2b3bab0a5c fix(anthropic): key_cmd Claude Code OAuth identity survives on custom api.anthropic.com routes (main + aux)
A named custom provider at api.anthropic.com (api_mode anthropic_messages) whose token comes from a
key_cmd callable lost the Claude Code OAuth identity: the aux custom routes hard-coded
is_oauth=False, agent_init/agent_runtime_helpers/client_lifecycle gated OAuth on provider=="anthropic"
and isinstance(key, str), and the callable-token client builder never added the OAuth betas or the
claude-code user-agent. Anthropic answers such a bare Bearer with 429 rate_limit_error "Error"
(#114967). One resolver, anthropic_credentials.anthropic_route_is_oauth(base_url, credential,
provider=), decides at every site: the route qualifies for the anthropic provider (unchanged) or an
exact api.anthropic.com host, the credential is a string or a callable materialized once
(CommandTokenSource caches), third-party hosts never qualify. model_metadata
_query_anthropic_context_length skips a callable credential instead of crashing agent init
(AttributeError on the same key_cmd route on current main).

Slimmer redo of #115007 by @liuhao1024 (same direction; one shared helper instead of per-site copies).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 12:49:40 -07:00
teknium1
8d153b26aa feat: add z-ai/glm-5.3-flashx to the OpenRouter and Nous Portal catalogs
Both routes serve the slug (tools supported, 1,048,576 context, $0.37/$1.25 per M).
The Nous list is derived from OPENROUTER_MODELS so one tuple edit covers both; the docs
manifest is regenerated in the same commit.

DEFAULT_CONTEXT_LENGTHS gets its own key: substring matching would otherwise land the
slug on the glm-5.3-flash entry (1,310,720) and overstate the window by 25%.
2026-09-20 12:16:06 -07:00
beardthelion
6087b40932 fix(agent): refuse ambient credential chains for multiplex profiles
Under gateway multiplexing, a served profile with no credential of its own
fell through to the SDKs' ambient default chains, which read the process
environment — the launch profile's Azure service principal, az CLI caches,
host managed identity, or AWS keys — and the minted identity rode the served
profile's configured base_url. vertex_adapter already refuses the same
pattern for google.auth.default().

azure_identity_adapter._scoped_credential now raises under multiplexing
when the profile scope has no complete AZURE_* set; AZURE_CLIENT_ID alone
routes to ManagedIdentityCredential as the explicit user-assigned-MI
opt-in. _probe_token propagates the caller's contextvars into its daemon
thread so doctor/probe checks run scoped (a bare Thread ran unscoped and
masked the refusal). describe_active_credential reads all credential
predicates through the scope.

bedrock_adapter.scoped_aws_session_kwargs now requires a complete
credential under multiplexing (key pair, or AWS_PROFILE as the shared-
config opt-in) instead of returning {} and letting boto3.Session() resolve
the ambient chain; the guard runs before the boto3 import at both call
sites. resolve_bedrock_bearer_token reads AWS_BEARER_TOKEN_BEDROCK through
the profile scope under a HERMES_HOME override.
2026-09-20 12:09:47 -07:00
teknium1
6d1341a13b fix: ranged @file read bails at the char budget mid-line (review follow-up)
_next_line collected every readline(line_cap) piece until the newline, so a
one-line giant (minified JSON) was still fully materialized before the
total_chars > char_budget gate. Pass the remaining budget into the helper and
stop collecting as soon as it is exceeded; the caller returns the oversized
block and never needs the rest of the line.
2026-09-20 11:54:43 -07:00
beardthelion
0accda1f76 fix(agent): bound file/process reads in @ context reference expansion
@file:, @folder:, @diff, @staged, and @git: expansion materialized
unbounded amounts of data before any size gate ran:

- _is_binary_file called read_bytes() and sliced [:4096], loading the
  whole file to sniff it.
- _expand_path_reference read_text()ed the whole file before the
  max_inline_tokens check, so a refused file was fully loaded and
  token-scanned anyway; ranged refs read the whole file to serve a
  slice.
- _file_metadata read_text()ed every folder-listing entry (up to 200)
  just to count lines.
- _run_quiet buffered the child's entire stdout, so git diff / rg
  --files on a large tree loaded an unbounded stream.
- gather expanded every parsed reference in a message with no cap.

Remote-triggerable through the gateway: any inbound message can carry
@-references, so a single message could force GB-scale transient
allocations and full-file scans before the token gate refused them.

Bounds applied:

- Binary sniff reads a 4KB prefix via open().
- Whole-file refs stat() first; st_size > max_inline_tokens *
  CHARS_PER_TOKEN is certainly oversized (the estimator is >= bytes/4
  for every encoding mix) and refuses without reading.
- Ranged refs stream the requested window with readline() pieces
  capped at the char budget; lines outside the window are skipped
  without materializing, so single-line giants (minified JSON,
  one-line logs) cannot expand a window read into a full read.
- Folder metadata streams the line count in 1MiB chunks and reports
  byte size past 4MiB.
- _run_quiet drains pipes on threads up to a 4MiB ceiling and kills
  the child on overflow; the nonzero returncode routes callers to
  their existing fallback paths.
- At most 16 references expand per message; the rest get a warning.

Measured: a refused 21MB file cost 42.4MB peak traced allocation
before and ~3KB after; an 85MB mixed hostile message costs ~6KB.
2026-09-20 11:54:43 -07:00
Aaron08140
829c91aa00 fix(agent): run bare script-path hooks on Windows through their interpreter
A hook declared as `command: "~/.hermes/agent-hooks/x.sh"` — the shape every
example in website/docs/user-guide/features/hooks.md uses — cannot start on
Windows. _spawn() shlex.splits the command and Popen()s it with shell=False, so
the kernel reads the shebang on POSIX but CreateProcess on Windows receives a
text file and answers WinError 193. The hook then reports no returncode, which
a fail_closed gate treats as a failure and every other consumer silently skips.

Route a first argument that is an existing file with a mapped suffix through its
interpreter, reusing tools.environments.local._find_bash() so the resolution
keeps the ordering that avoids WSL's bash.exe (#115124) and surfaces Git-for-
Windows' own guidance when it is absent. POSIX argv is untouched. Suffixes we
cannot resolve an interpreter for still fail, but the diagnostic now names the
remedy instead of the OS's localized complaint.

Repairs five of this file's tests that have been red on native Windows
(TestCallbackSubprocess x4, hooks TestHooksTest::test_fires_real_subprocess_and_parses_block);
the two that stay red are drive-letter/`~` tokenization, which #68508 owns.

Verified on Windows 11 26200 / cp936 with real subprocesses: bare .sh, .bash and
.py hooks execute and carry their exit code; a missing path still reads
"command not found"; `git --version` is unaffected.
2026-09-20 11:27:39 -07:00
beardthelion
5aca5406f0 fix(agent): reject subdirectory hint files that resolve outside the tree
The subdirectory-hint loader resolves each visited directory but never the
hint file itself, so a checked-in sub/AGENTS.md symlinked to an out-of-tree
file (~/.aws/credentials, another agent's config) was followed and injected
into the tool result. Resolve each candidate and require its target to stay
inside the working dir and pass the canonical read deny-list - the same
policy @-references already apply - and read the resolved inode so a link
swapped between check and read still lands on the vetted file. In-tree
symlinks keep working.
2026-09-20 10:51:03 -07:00
liuhao1024
c19246557a fix(gemini): stop rerouting AQ. keys off the default Studio surface
Google now issues AQ. keys for both Google AI Studio and Vertex AI
express mode, so the AQ. prefix no longer identifies the key family
(#115306): auto-routing every AQ. key to aiplatform.googleapis.com 403s
the whole AI Studio fleet (6293fca019 / df53cae72d).

- normalize_gemini_base_url no longer rewrites by key shape; an express
  key reaches aiplatform only through an explicitly configured base,
  which is still completed to the publishers/google form (#114335 path)
- gemini_http_error appends two-way 403 PERMISSION_DENIED guidance: an
  AQ. key rejected on the Studio host learns about the express base_url,
  a key rejected on an explicit aiplatform base learns about the default
- doctor's explicitly configured aiplatform base now also gets the
  publishers completion; the OAuth Vertex .../endpoints/openapi base
  stays untouched

Fixes #115306
2026-09-20 10:50:27 -07:00
liuzikaii
d23d6e8218 fix(lsp): use UTF-16 units for document replacement ranges 2026-09-20 10:24:17 -07:00
teknium1
5195c13873 fix: verify-on-stop recognises python.exe / py launcher interpreters (review follow-up)
_is_interpreter_token matched Path(token).name against the bare-name regex,
so a Windows venv path `...\Scripts\python.exe` (and the `py` launcher)
recorded no ad-hoc evidence and the nudge loop the PR closes stayed open on
that platform. Strip a case-insensitive .exe/.bat/.cmd suffix, take the
basename across backslashes, and accept `py`. Parametrized invariant test,
red before.
2026-09-20 10:23:04 -07:00
Uttkarsh Tiwari
35accdbc30 fix(agent): record ad-hoc verify evidence run through a versioned or absolute interpreter
The ad-hoc matcher accepted an interpreter only when the command word was
literally in `_INTERPRETERS`, so `python3.12`, `/usr/bin/python3.12` and
`/usr/bin/env python3` recorded no evidence for the temp `hermes-verify-*`
script the stop-gate nudge asks the agent to run. The workspace stayed
unverified and the nudge re-fired on every stop, so verify-on-stop could not
be satisfied by following its own instructions.

Recognise the interpreter by name (versioned/absolute paths included) and
treat a leading `env` as transparent. Commands that merely name the script
(`rm`, `chmod`, `cat`) stay non-evidence — the same nudge tells the agent to
clean the script up, and a cleanup command must not clear the gate.
2026-09-20 10:23:04 -07:00
beardthelion
04dc1907be harden vault error handling for malformed state
A vault file that decrypts but holds malformed, non-dict, or non-UTF8
JSON escaped the VaultError contract: _read_all raised raw
JSONDecodeError where callers only catch VaultError, so
hermes vault add/list/rm produced tracebacks. Wrap the decode and
parse in VaultError, and catch VaultError in vault_command so every
subcommand prints the clean error line.

vault.source.set and vault sources --enable/--disable also crashed on
a non-dict vault section (vault: true, vault: {bitwarden: true}) via
an unguarded setdefault chain. Coerce through _ensure_dict, the same
shape guard _voice_cfg_dict documents for voice.*.

Fixes #115867
2026-09-20 10:21:52 -07:00
fangliquan
9f7df273a7 fix(auxiliary): prioritize explicit reasoning config 2026-09-20 10:17:39 -07:00
fangliquan
0a407e1651 fix(auxiliary): consume promoted reasoning config 2026-09-20 10:17:39 -07:00
fangliquan
c86de9c44c fix(auxiliary): preserve raw reasoning body shapes 2026-09-20 10:17:39 -07:00
fangliquan
ab3448e075 fix(auxiliary): honor provider reasoning disable controls 2026-09-20 10:17:39 -07:00
finn763
effcf3af06 fix(gateway): a peer DM retries a transient turn failure once, like the other two lanes
`hermes peer dm` posts to POST /api/sessions/{id}/chat, the third Bot-DM
transport. The local (`tools.bot_mode_dm`) and relayed
(`tui_gateway.methods_bot_relay`) lanes both re-run a transiently failed turn
once under the shared policy (`tools.bot_failure_reasons.retry_action`) and
resume the row the failed attempt left as the transcript's unanswered tail;
this lane ran the turn once and handed the provider's 429 paragraph to the
sender as the reply (#115325).

The policy is asked about a result dict now, not two streams: `result_retry_action`
joins `error` + `failure_reason` (the turn loop's own typed verdict) so one
classifier serves every lane, and the server-error rule accepts the providers'
`server_error` / `overloaded_error` spellings — the codes the in-process lanes
key on instead of a status number.

The resume half is the CLI lane's rule extracted to `agent.session_persistence.
adopt_unanswered_turn`, which `quiet_single_query` (env-gated dispatcher re-run)
and the API lane (in-process re-run, on the agent it just built) now share.

The regression drives the real route and the real `_run_agent` over a real
store: a 429 re-runs the same DM once with the persisted row adopted as this
turn's user message (so no second copy), a 401 still reports one attempt.

(cherry picked from commit 8fe6d46ada5b8064bc7132ee956fda998244c10a)
2026-09-20 10:16:51 -07:00
teknium1
0f3d32ec58 refactor(transports): one registry read per api_mode gate; drop the stream-delta hook
Trim the salvaged fix to the shape main wants:

* Every gate asks ``agent.transports.registered_api_modes()`` directly. The three helper
  spellings (``_has_registered_transport``, ``_registry_knows``, ``is_registered_api_mode``)
  and the ``sys.modules`` peek are gone: no transport module imports ``providers`` or
  ``hermes_cli`` at module level, so a plain import cannot re-enter provider discovery.
* ``hermes_cli/auth.py`` late-registration pass dropped — main already re-syncs plugin
  profiles into ``PROVIDER_REGISTRY`` on every registry miss
  (``auth_plugin_providers.registry_lookup`` / ``sync_plugin_provider_registry``, #102123);
  the probe shows a profile registered after the import-time mirror resolves and reaches
  the wire on base.
* ``ProviderTransport.normalize_stream_delta`` and the streaming-assembler hook dropped —
  legacy ``delta.function_call`` translation is a separate concern from api_mode
  propagation and has no in-tree consumer.
* Tests: 15 gate-by-gate unit tests replaced by two invariants that install a REAL plugin
  under a temp HERMES_HOME and walk profile → determine_api_mode → resolve_runtime_provider
  → agent ladder → delegation resolver (positive: red on origin/main; negative: an
  unregistered mode still degrades to chat_completions).
2026-09-20 10:11:40 -07:00
valerdoskin
ef8cdfc389 fix(transports): accept a provider plugin's own api_mode
A provider plugin ships a transport via `register_transport(api_mode, cls)` and
declares that same string as its profile's `api_mode`. Transports are selected by
that string, but every gate that validates an api_mode compared it against a
closed literal, so a plugin's mode was rejected at each one and rewritten to
`chat_completions`. The plugin's transport was then never selected: no error, no
tool call, the turn silently degraded to prose. `register_transport` was a public
seam with no way through.

Accept a mode when the transport registry knows it, via a new
`agent.transports.registered_api_modes()`, at each gate:

* `agent_init._resolve_api_mode` - the agent's mode ladder;
* `runtime_provider._parse_api_mode` - the config gate;
* `delegate_tool_config` - the delegation resolver;
* `providers.get_provider` - the reverse `TRANSPORT_TO_API_MODE` lookup recorded
  an unknown mode as `openai_chat`, which made `determine_api_mode` report
  `chat_completions` for a provider that has a dialect transport. This one is the
  most deceptive: the other gates already pass, and the transport still is not used.
* `providers.determine_api_mode` - the same table lookup at the other end.

The registry read is deliberately lazy (`sys.modules.get("agent.transports")`,
never an import): this code is reached from `determine_api_mode`, which provider
discovery itself calls while the registry is being populated, and importing the
transport package there re-enters discovery.

Also add `ProviderTransport.normalize_stream_delta()`, the response-side twin of
`convert_messages()`: a provider that streams a tool call on the legacy OpenAI
`delta.function_call` pair instead of indexed `delta.tool_calls` had nowhere to
translate it, and the streaming assembler dropped the call. The default returns
the delta unchanged, so existing transports are untouched; the assembler now asks
the transport instead of hardcoding one provider's shape.

Finally, make plugin-provider registration repeatable. `hermes_cli.auth` registered
plugin profiles once, at import, from a list `hermes_cli.config` had already
partially discovered while importing itself. A profile that sorts LAST in discovery
was absent from that snapshot and never reached `PROVIDER_REGISTRY`, so every
consumer reported it unauthenticated and it silently vanished from the model
picker while working fine from the CLI. `ensure_plugin_providers_registered()` is
now called from `get_auth_status()` and `resolve_provider()`, so a late profile is
picked up instead of staying invisible.

Unregistered modes are still rejected everywhere, and the in-tree literal sets are
unchanged - they are simply no longer the only way in.
2026-09-20 10:11:40 -07:00
kshitijk4poor
e74a84f0b6 refactor(agent): declare the trim flag in _CONTROL_STATE; trim only on a completed batch
Gate Lows: an in-flight exception pins the executor frames through its traceback, so the
finally-placed trim could not release the result on that path and would burn the cooldown;
the flag now survives to the next completed batch. The default lives beside _executing_tools.
2026-09-20 18:09:19 +05:30
kshitijk4poor
05e09aee68 fix(agent): trim after the tool batch unwinds, not inside the commit that still holds the result
Gate review (three lenses converged): inside _commit_tool_result the >=1 MB string is still
referenced by the publish frames (the returned tuple, managed.result / batch.results, the
tool.completed callback), so gc.collect + malloc_trim could not release it and merely spent
the 60 s cooldown. The commit now only sets agent._trim_after_tool_batch; the finally of
AIAgent._execute_tool_calls consumes the flag once every executor frame is gone, coalescing
N large results in one batch into one trim. The import is lazy like every sibling call site
(keeps ctypes out of the agent import chain).
2026-09-20 18:09:19 +05:30
kshitijk4poor
f32f651788 perf(tool_executor): trim memory after publishing a >=1 MB tool result
Post-compression already calls trim_memory (#77356); a huge tool result (raw
stdout, file dumps) is the other allocation a turn drops and was published
with no collection. _commit_tool_result is the one point both the sequential
and concurrent publish paths go through, so the trim lives there, after the
spill + session flush, measured on the string already in hand (multimodal
dicts are never re-serialised). trim_memory's own cooldown/kill-switch apply.

Salvages the intent of #80974 without its bare gc.collect(), re-serialisation
and 186 LOC. Closes #70684 (tool-result half).

Co-authored-by: Christopher-Schulze <210261288+Christopher-Schulze@users.noreply.github.com>
2026-09-20 18:09:19 +05:30
kshitijk4poor
a4ec7d63c7 refactor(anthropic): which() already guarantees an executable file 2026-09-20 17:58:45 +05:30
kshitijk4poor
e58d7df339 fix(anthropic): probe every PATH hit before the install prefixes
The nested loop probed every prefix candidate for 'claude' (e.g. a stale
~/.local/bin/claude) before which('claude-code'), so a current PATH
'claude-code' lost to a stale non-PATH 'claude' — the exact stale-identity
failure the prefix list exists to fix. Two passes: all PATH hits, then the
prefixes; os.path.join and module-level imports.
2026-09-20 17:58:45 +05:30
r00tedbrain-backup
0eb04ef51f fix(anthropic): detect Claude Code version outside PATH on GUI launches
Claude Code version detection resolved the CLI by bare name, which only
searches PATH. GUI-launched processes on macOS (the Electron desktop app,
LaunchAgents) inherit the bare /usr/bin:/bin:/usr/sbin:/sbin, which carries
none of the CLI's install prefixes. Detection found nothing there even with a
current CLI installed and fell through to the stale fallback constant, which
Anthropic rejects:

    HTTP 400: Claude Code 2.1.74 does not support this model;
    version 2.1.251 or newer is required.

Same machine, same CLI, same account: fine from a terminal, rejected from the
desktop app. Verified by running the unmodified function under each PATH —
GUI PATH returned 2.1.74, terminal PATH returned the installed 2.1.276.

Probe the well-known install prefixes in addition to PATH. Additive only: a
PATH hit still wins and is tried first, and on Windows no prefix resolves to a
file so detection falls back to the PATH lookup exactly as before.
2026-09-20 17:58:45 +05:30
kshitijk4poor
a4f34b90c6 fix(bedrock): log the swallowed boto3 lazy-install failure instead of passing silently 2026-09-20 16:41:38 +05:30
kshitijk4poor
ee17fff193 refactor(agent): one bypassing create() in _relay_sync_stream; bypass hoisted above the dispatch tick; dead-path hunk dropped
Gate review: the bypass line had landed between the #114938 comment and the branch it explains;
_acreate_with_stream has no production caller.
2026-09-20 16:33:18 +05:30
kshitijk4poor
8bca1f7167 fix(agent): relay-managed aux streams run the SDK bypass inside the provider callback
Gate review: _relay_sync_stream applied the bypass before handing kwargs to
relay_llm.stream_current, so Relay's tracing saw `messages: []` for every managed streaming
aux call and an intercept that rewrote messages was silently discarded (the SDK merges
extra_body after the transform). The bypass now runs in the provider callback like every
other site; _create_with_progress_once / _acreate_with_progress apply it once at entry
instead of at each of three create() calls. New test: Relay sees the full conversation while
the SDK still gets only the placeholder (red on the previous head).
2026-09-20 16:33:18 +05:30
kshitijk4poor
d66bed6a8d perf(summary): bypass chat SDK request transform on the iteration-limit summary call
The iteration-limit summary rebuilds the full main-loop kwargs via _build_api_kwargs
and calls chat.completions.create itself, so it walked the whole conversation through
the SDK transform (#106776). Route it through agent.sdk_transform_bypass like the
main loop; pin the aux relay path and the summary path with placeholder-only tests.
2026-09-20 16:33:18 +05:30
ericmaddox
fc67cbaab3 perf(agent): bypass chat SDK request transform in auxiliary and summary calls
OpenAI SDK's maybe_transform walks messages and tools trees against type unions with the GIL held, causing 30ms+ CPU stalls on multi-MB conversations before hitting the network. By moving wire-format bulk fields into extra_body (and setting messages=[] at top level), requests produce byte-identical wire JSON bodies while bypassing client-side SDK transform overhead. Applied to auxiliary client create helpers and iteration summary attempts.

Closes #106776
2026-09-20 16:33:18 +05:30
kshitijk4poor
b729bc43bc refactor(config_providers): keep main's self-resolve, read the config through the read-only loader
Gate review: the salvaged config_providers.py hunk was behaviour-neutral churn (the helper
already self-resolved a None route list on main) that also dropped the raw-config fallback.
Restore main's shape; since step 0c now runs for every route with a base_url, resolve via
load_config_readonly() so the per-call deepcopy of load_config() is not paid on every
context-length lookup. Comment trimmed to the WHY; the invariant test moves next to its
siblings in test_custom_provider_context_length.py.
2026-09-20 15:17:49 +05:30
Stephen Cuppett
f4ffd48274 fix(model_metadata): honor per-model context_length when the caller doesn't pass custom_providers
get_custom_provider_context_length only found the user's
custom_providers[].models.<id>.context_length override when the caller
had already loaded the route list and passed it in. Several real call
paths never do — auxiliary fallback screening (auxiliary_client),
vision auto-detect, the CLI/TUI @-context-reference estimators, and
gateway /status all pass custom_providers=None — and step 0c of
get_model_context_length was additionally gated on that list being
truthy. Those paths skipped the override and fell to the 256K/272K
probe-down defaults, while the startup path honored the same setting,
so the same config 'took' in some places and not in others.

Fix: get_custom_provider_context_length self-resolves the route from
config when neither custom_providers nor config was supplied (same
pattern as _resolve_moa_context_length; load_config() is cached on the
config file signature), and the step-0c gate now checks only
(base_url AND model).

Regression test asserts the no-args call shape returns the override
(proven red on base).
2026-09-20 15:17:49 +05:30
kshitijk4poor
59f9ff8dbc refactor(credits): gate the re-warm on the pricing cache's own failure window
The started_at throttle kept a second clock for a deadline the cache
already holds: a failed fetch caches {} with _pricing_cache_retry_after,
and the two clocks drift by the fetch duration. pricing_fetch_suppressed
reads that state directly, so the tracker drops the ad-hoc Thread
attribute and the private-constant import; the guard test now fails the
way the real fetch does (cached {}) instead of a bare no-op.
2026-09-20 14:02:17 +05:30
kshitijk4poor
6cfbc5f891 refactor(credits): throttle the re-warm, scope the seed thread, one subscription predicate
- rewarm_pricing_before_depleted_notice: a failed fetch caches {} for
  _FAILED_CATALOG_TTL_SECONDS and the peek reads that as cold, so every
  header in that window spawned a thread that read the auth store and hit
  the cached {}. Remember when the last warm started and decide inline
  until the window passes. Drop the dead try/except around the pure peek.
- _bg_seed now runs under spawn_context_thread: the warm it gained reads
  the profile's auth store, so the thread must carry the profile scope.
- _rerun_notice_policy replaces the idiom copied at three sites.
- _is_subscription_billed: the free-tier default filtered on any truthy
  billing_mode while _is_model_free keyed on == 'subscription'.
- The no-respawn guard test counts warm calls instead of enumerating
  finished threads (which always read 0).
2026-09-20 14:02:17 +05:30
kshitijk4poor
21cf53b888 fix(credits): re-warm the Nous catalog when a header finds it cold after the TTL
The session-start warm is one-shot, but the Nous catalog expires after
_NOUS_CATALOG_TTL_SECONDS (300s) and peek_cached_pricing deliberately skips an
expired entry. Every inference header after that re-ran the notice policy
against a cold peek, so a subscription-billed model on a depleted account
drew the depleted banner again five minutes into the session — and kept it.

When the account is depleted, the model is not locally known to be free and
the peek is cold, _emit_credits_notices now starts a background warm (via
spawn_context_thread) and leaves the decision to the warm's own re-run instead
of flashing a banner the warm catalog would suppress. Fail-open: a failed warm
leaves the peek cold and that re-run shows the banner; an in-flight warm (the
re-run itself included) never spawns another.

Tests: TTL-expired header stays silent (red without the mixin call); cold
catalog + failed warm still warns and does not re-spawn.
2026-09-20 14:02:17 +05:30
Robin Fernandes
a728c55455 fix(credits): warm the pricing catalog before the session-open notice policy
The depleted-banner exemption for a subscription-billed model reads the
catalog through `is_free_tier_model`, which only PEEKS the pricing cache
and never fetches. Nothing warms that cache during chat startup, so the
cold-start seed evaluated against an empty catalog: a session opening on
a subscription-billed row showed "Credit access paused · run /topup"
even though the row spends no credits. Observed on `openai/gpt-5.6-luna`.

Warm the Nous catalog from the seed's background thread, which is already
off the critical path, before either policy branch runs. When a live
inference header beat the seed, keep its state but re-run the policy —
that evaluation saw the cold catalog and may be showing a banner the warm
one suppresses.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-20 14:02:17 +05:30
teknium1
9573f44ca5 fix(billing): the /billing portal fetch releases the caller at its wall-clock bound too
The same `with ThreadPoolExecutor` join that #115982 reports for
`account_usage._fetch_portal_account` lived in
`billing_usage.fetch_nous_account` (the `/billing` CLI + TUI surface).
Delegate to the single bounded implementation instead of carrying a
second copy, and cover both entry points with one parametrized test
(contributor tests trimmed to two invariants).

Co-authored-by: fangliquan <fangliquan@qq.com>
2026-09-20 00:35:04 -07:00
liuhao1024
7bc7937c08 fix(agent): release the portal account fetch caller at the wall-clock bound 2026-09-20 00:35:04 -07:00
teknium1
29826ca5b6 fix(bedrock): capture reasoning signatures, parse sync reasoningText, keep untouched turns on sealed-blob resend
Follow-up to the cherry-picked #115875 salvage (#115865):

- `_ResponseParts.absorb_reasoning` reads the sync response's nested
  `reasoningText: {text, signature}` (it only understood the flat stream-delta
  shape, so non-streaming Converse thinking was silently dropped) and records
  the `signature` on the ordered block; the stream gate no longer discards
  signature-only deltas.
- `_replay_ordered_blocks` replays that signature inside `reasoningText`, so
  models that sign their thinking accept the replayed turn.
- `strip_redacted_reasoning` left untouched turns with `content: None` (the
  cleaned-content sentinel leaked into the rebuilt message), which the resend
  then failed on with ParamValidationError; untouched turns are now replayed
  verbatim.
- Tests trimmed to two invariants driven through `call_converse` (botocore's
  real Converse input validation as the stand-in client) plus one nested
  sync-shape check and the re-raise control.
2026-09-20 00:22:10 -07:00
Yagna Vudathu
8437310aed fix(bedrock): replay thinking as reasoningText; resend once without redacted blocks on encrypted-content rejection (#115865) 2026-09-20 00:22:10 -07:00
Diamond Hands Dig
c2560f4dcc fix(compression): dispatch alone must not reset the summary idle fence (#114938)
Problem
-------
`_create_with_progress_once()` (and its async twin `_acreate_with_progress`)
ticked the aux forward-progress hook unconditionally at dispatch time, before
the provider had produced any payload. Compression wires that hook to
`CompressionCommitFence.touch_progress`, so a dispatch that dies before any
output — an auth refresh, a retry, a fallback — reset the 120s summary
inactivity timer. A zero-output attempt could therefore run until the 600s
total ceiling instead of idling out, and the fence's `progress_observed`
flipped True on dispatch alone.

Reproduction
------------
Install `aux_progress_hook(fence.touch_progress)`, use a fake client whose
`create()` immediately raises HTTP 401, then call `_create_with_progress_once()`.
On current main `fence.progress_observed` becomes True from dispatch alone.

Fix
---
Keep `_notify_aux_dispatch()` for dispatch telemetry, but drop the two
unconditional `_notify_aux_progress()` calls (sync + async). Progress now
ticks only for substantive stream payloads (`_ChatStreamAccumulator.feed`) or
a completed usable response (`_notify_aux_provider_response` on the plain-call
and shim paths) — the same contract the adapter event hooks already enforce.

Why not an alternative
----------------------
Keeping the dispatch tick and instead making the fence ignore it would leave
every other progress-hook consumer (gateway session hygiene) with the same
false-liveness signal; the hook contract is "forward progress", and dispatch
is not progress. The change is two deletions; the streamed path already ticks
per substantive chunk, so no liveness is lost for genuinely streaming calls.

Validation
----------
- New regression tests (sync + async): a 401 dispatch leaves
  `fence.progress_observed == False` while dispatch telemetry still fires;
  both fail on the pre-fix code and pass after.
- Updated the two tests that pinned the old dispatch tick
  (`test_aux_progress_streaming.py`) and the relay-seam comment.
- `pytest tests/agent/test_aux_progress_streaming.py
  tests/agent/test_aux_relay_progress_seam.py
  tests/agent/test_auxiliary_explicit_cancellation.py
  tests/agent/test_aux_affordable_402_retry.py` → 82 passed.
- Compression suite (review/progress/stall/worker-isolation/attempt-lifecycle
  + timeout-floor + progress-timeout) → 75 passed; the one intermittent
  failure in `test_second_consecutive_stall_commits_the_deterministic_fallback_summary`
  reproduces identically on pristine main (pre-existing flake, see #115204).
- `ruff check` clean on all three changed files.
2026-09-20 00:15:32 -07:00
satoshi
2f69687889 fix(agent): _should_stream keys ACP on the provider profile, not one vendor's name
An out-of-tree external-process provider whose base_url marker is not ``acp://`` still tried
to stream a completion object that is not iterable; the check named copilot-acp alone.
Gate on the profile's external_process auth_type as well.

Salvaged from #107754 (the agent_init half already landed via #116552 / #116958).
2026-09-20 12:36:34 +05:30
teknium1
80154cf3cf fix(streaming): explicit stale_timeout_seconds wins over the context-size tier and reasoning floor
The reasoning-floor gating (#99707) stopped the floor from overriding an
explicit providers.<id>.stale_timeout_seconds on the non-stream path, but the
streaming resolvers still fed the explicit value through _cloud_stale_timeout,
where the context-size tier (240s/300s) and the reasoning floor are max()ed on
top. A 60s explicit deadline on an 75k-token request became 240s, so the
operator could never SHORTEN patience for a hung stream (#115024).

_cloud_stale_timeout_for returns an explicit provider/model value as-is and
scales+floors only the 180s default; both stream resolvers (_StreamingCall.
_resolve_stale_timeout and the Bedrock _derive_stream_stale_timeout) use it.
2026-09-19 23:49:03 -07:00
kshitijk4poor
144890829b fix(usage): plugin usage hook runs under run_bounded_sync; base no-op spawns no thread
``_call_plugin_usage_hook`` re-implemented ``agent.deadline.run_bounded_sync`` with a bare
thread whose exceptions never reached ``fetch_account_usage``'s fail-open ``except`` —
``threading.excepthook`` printed the traceback to stderr on every ``/usage``. Reuse the
shared helper (exceptions re-raise in the caller, timeout logged) and return None without
a thread when the profile inherits ``ProviderProfile.fetch_account_usage``.
2026-09-20 12:17:52 +05:30
kshitijk4poor
838214c8f7 fix(error-classifier): a profile hook asking for fallback on a terminal reason is non-retryable
turn_api_error enters the fallback walk only for ``retryable=False`` verdicts outside the
retryable-client reasons; the built-in terminal verdicts pin retryable=False while the
rate-limit family stays retryable and cascades after backoff. A ``classify_api_error`` hook
returning ``{"reason": "billing", "should_fallback": True}`` therefore retried the dead
route instead of cascading (the #116408 test passed only because its hook also set
retryable=False). Default retryable to False for such verdicts, leaving rate-limit reasons on
the built-in retry-then-fallback shape; ``RETRYABLE_CLIENT_REASONS`` moves next to the
verdict table so both modules read one set. Contract documented in the plugin guide.
2026-09-20 12:17:52 +05:30
kshitijk4poor
925a5992d6 fix(agent): ACP providers never auto-upgrade to Responses, whatever their base_url marker
The Responses auto-upgrade guard keyed only on the ``acp://`` / ``acp+tcp://`` scheme after
the copilot-acp slug check was dropped (#116408). ``COPILOT_ACP_BASE_URL=https://...`` flows
through ``_external_process_spec`` verbatim, so a gpt-5 model on copilot-acp flipped api_mode
to codex_responses against an ACP client. Key the guard on the profile's external_process
auth_type as well, reusing ``runtime_provider_backends._is_external_process_provider`` (CLI
registry first, then the profile registry) at both the routing and launch-kwargs sites.
2026-09-20 12:17:52 +05:30