Commit Graph

5518 Commits

Author SHA1 Message Date
fangliquan
bb9058d7e8 fix(compression): preserve steer display identity
(cherry picked from commit a8570878bc848a42cc8029fc45219a57e76be4c4)
2026-09-20 18:24:07 -07:00
teknium1
f568b860d7 fix(agent): pin the reset_at reach through try_activate_fallback and add opt-in fallback.min_switch_reset_seconds
- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback
  (production entry) on a real AIAgent with a one-entry chain, so dropping the
  reset_at forwarding goes red (3 failures before, 8 green after).
- #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the
  rate-limited primary's declared reset is sooner than N seconds,
  try_activate_fallback returns False and no cooldown is armed; docs row added.
2026-09-20 17:00:43 -07:00
fangliquanflq
125bdff601 fix(agent): honor provider reset for fallback cooldown 2026-09-20 17:00:43 -07:00
finn763
36b6efc297 fix(context): re-derive model.context_length on model/provider change
model.context_length is the user's profile-wide ceiling. It was read from
config.yaml in exactly one place — agent construction — and cached twice:
agent._config_context_length (switch/fallback resolution plus every display and
/usage surface) and context_compressor._config_context_length (the compressor's
own re-resolution).

Every live path that re-resolved a runtime then touched only one copy, or cleared
it without re-reading the config:

- switch_model nulled agent._config_context_length and re-derived the intent from
  custom_providers metadata alone, so a ceiling that only exists as
  model.context_length was dropped for the rest of the process;
- the Desktop/TUI compression hot-reload updated the compressor's copy only, so an
  open session showed a pinned ceiling while compressing against provider
  metadata / the 256K fallback.

Both now route through one pair of helpers in agent/agent_init.py:
set_config_context_length (one place that knows where the pin is cached) and
config_context_length_for_runtime (re-read from live config, scoped exactly like
construction, so an unrelated route still never inherits the pin).

(cherry picked from commit 986ff16dadb9966f7328e55f295af5cfa1eb5c88)
2026-09-20 16:31:10 -07:00
chelsealong
de622b291d fix(agent): detect Thai plan tails in promoted-reasoning stall guard
promoted_reasoning_announces_action()'s tail detector only matched
English (plus CJK punctuation boundaries), so a reasoning-only clean
stop ending on a Thai first-person plan (e.g. "จะให้ผม...") was not
recognized as a stall and got delivered to the user as the final
answer instead of nudging continuation.

Add Thai first-person future-action triggers, extend the boundary
class with em/en dash (a common Thai clause separator), and accept
multi-dot ellipsis tails.

(cherry picked from commit 87061e46c423859cf738d4541df6594233f40e7e)
2026-09-20 16:20:52 -07:00
teknium1
53815e24dc fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.

The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.

Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
  before  req_reasoning: {}
  after   req_reasoning: {'reasoning_effort': 'medium'}
  agent.reasoning_effort: low  ->  {'reasoning_effort': 'low'}  (unchanged)
2026-09-20 16:03:38 -07:00
Moep90
3a37a24efd fix(moa): mark reused advisor guidance as predating the tool results
With `fanout: user_turn` (and off-cadence `every_n` iterations) the advisors run
once per user turn and their guidance is replayed verbatim into every later
iteration of that turn. The block reads as fresh instruction, so an advisor that
proposes a tool call keeps proposing it after the acting model already ran it and
has the result in the transcript.

Observed with the clarify tool: the card was answered, the next iteration replayed
the same guidance, the model issued a second identical card, the user dismissed it,
and the turn then held two contradicting results for one question (an answer and an
empty skip). The following aggregation degenerated into a repetition loop until it
hit the output cap.

- `_STALE_GUIDANCE_NOTE` is appended when cached guidance is reused on an iteration
  that already has tool activity since the last real user turn. The advice text is
  handed over unchanged; only the framing says it may be out of date.
- The advisor system prompt now rules out emitting a tool call or a JSON tool-call
  object. Advisors hold no tools, and a tool-call object in advisory text is what
  the aggregator replays.

No change to fanout cadence, caching or accounting. The cadence test that pinned
byte-identical reuse now asserts the advice text is reused and carries the marker.

Signed-off-by: Moep90 <3042152+Moep90@users.noreply.github.com>
(cherry picked from commit 343787f287ad9915345090fba351df5ffa758c9e)
2026-09-20 15:51:19 -07:00
teknium1
0765099ff4 fix(compression): fence the durable cooldown rollback per compressor; one stale-attempt helper
The SQLite rollback no longer runs under the process-wide claim lock: a per-compressor serial lock (taken by _claim_compressor_attempt too) serializes it against claims on that compressor only. The seven pasted working-attempt checks call _raise_if_stale_attempt/_caller_attempt_is_current. Drops the unused _run_as_attempt test helper.
2026-09-20 15:50:49 -07:00
beardthelion
0e33dc9ebc fix(compression): stop detached stale attempts writing shared compressor state
The stall-fallback detaches a timed-out primary worker and reuses the
same ContextCompressor, but the existing attempt-generation guards only
covered the unwind-time snapshot restore. Every other summary-state
write stayed reachable by the still-running primary after the fallback
took over: a late successful summary published _previous_summary and
cleared the fallback's cooldown, a late failure armed a shared failure
cooldown and stamped error state, the cancel rollback and the abort
rollback reverted _previous_summary to the primary's snapshot, and the
durable cooldown rollback row could be overwritten mid-restore.

Compressor code could not fix this with the shared attributes alone:
those cells only name the current owner, never the calling attempt.
The calling attempt's generation now rides a ContextVar bound inside
_run_summary_dispatch, which every attempt's compress_fn passes through
in its own thread, so each attempt reads its own generation. Gates on
the working-attempt marker (not the entry claim, so lock sit-outs do
not suppress the owner) now cover the cancel rollback, late-success
writes, _on_summary_failure, the abort rollback, the deterministic
pin, compress() entry, and a Phase-3 choke point. The durable cooldown
rollback moved inside the claim lock so the DB row and the in-memory
restore are atomic against _claim_compressor_attempt.

Regression tests drive the real interleavings deterministically,
including two threaded end-to-end arms through _run_summary_dispatch
and a real ContextCompressor.

(cherry picked from commit 902bfc229e140becfb36679dc33bad550c2ae1e8)
2026-09-20 15:50:49 -07:00
teknium1
85564321be fix(windows): route every bare-bash spawn through _find_bash and surface silent interpreter failures
CreateProcess resolves a bare "bash" to System32 WSL launcher before PATH,
and shutil.which("bash") inherits PATH order (#115124), so node bootstrap,
the TUI node probe and webhook filter scripts ran the wrong interpreter on
Windows. All four sites now use tools.environments.local._find_bash (Git Bash
first, probed). rc!=0 with no output at all is now a WARNING in webhook
filters and an explicit [inline-shell exit N with no output] marker in skills.

Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
2026-09-20 15:50:26 -07:00
funky-xamarin
1a24b851e5 fix(skills): resolve native Git Bash for Windows inline shell 2026-09-20 15:50:26 -07:00
beardthelion
275002d7a4 fix(context_references): @folder: listing works outside cwd under a widened allowed_root
@folder: targets resolve against allowed_root, which callers may widen
beyond cwd, but _build_folder_listing and _iter_visible_entries assumed
the resolved folder was under cwd: path.relative_to(cwd) raised
ValueError and the blanket except in _expand_reference surfaced it as a
confusing "not in the subpath of" warning instead of a listing.

The listing header now renders cwd-relative when possible, then
allowed_root-relative, else the absolute path. rg --files gets the
absolute folder path (rg echoes the arg as the output prefix, so the
lines parse correctly anywhere), the parent-dir walk drops only the
cwd stop-condition so in-cwd output is unchanged, and entry indentation
is computed relative to the target folder rather than cwd. The os.walk
fallback never assumed cwd.

Regression tests cover the widened-root target through the real
preprocess_context_references entry on both the rg and rg-blocked
paths, plus @file: parity and in-cwd display controls.
2026-09-20 15:49:22 -07:00
teknium1
30d14a3f73 fix: name the CommandCode upstream-outage pattern
The cherry-pick conflict dropped the contributor comment hunk; keep the WHY next to the entry.
2026-09-20 15:28:21 -07:00
fangliquan
1f3c882d81 fix(agent): narrow upstream outage matching 2026-09-20 15:28:21 -07:00
beardthelion
09b72bc6d2 fix(compression): track commit fences as a registration stack
_compress_context published the active commit fence with a
save/restore cell: registration order was serialized by the fence
lock, but completion order is not. When attempt B registered over A
and A finished first, A's finally popped the slot, deleting B's live
fence mid-attempt (hard_interrupt lost the handle serializing cancel
admission against B's begin_commit). B's finally then republished A's
dead fence, which lingered until the next compression. The same
clobber existed in _publish_new_fence, which overwrote the slot
unconditionally when minting the stall-fallback retry fence.

Replace the cell with a stack of per-attempt registrations. The
finally removes only its own registration and republishes the newest
live entry (or clears the slot), so a dead fence can never be
restored over a live newer attempt. The stall-fallback retry swaps
its fence inside the owning registration and publishes only while
that attempt still holds the top registration. Registration moved
inside the try so an early exception cannot strand an entry.

(cherry picked from commit 574e9945cf186071c3da23c4bb517c0cbdaf73e6)
2026-09-20 15:27:11 -07:00
beardthelion
a48b4c7d25 fix(agent): pop _db_persisted on in-place mutations of stamped live dicts
The _db_persisted marker asserts that a message dict's persisted row is
durable as written; any in-place mutation must pop it or the flush scan
identity-skips the dict and state.db keeps the stale row forever. Six
mutation sites violated the contract:

- micro_compaction._merge_adjacent_user_turns rewrote content on a
  carried-forward dict after superseding a stale micro marker. On the
  archive-failure path nothing re-stamps, so the merged text never
  reached state.db.
- repair_message_sequence passes mutated stamped survivors in place:
  _merge_assistant_into (tool_calls union, content join,
  reasoning_content carry), _prune_unanswered_tool_calls (tool_calls
  rewrite), _merge_consecutive_users (content join).
- sanitize_tool_call_arguments rewrote corrupted/blank
  function.arguments and prepended the corruption marker onto stamped
  resumed rows, leaving the corrupt bytes durable and self-perpetuating
  across resumes.
- _sanitize_messages (surrogate and non-ASCII recovery) and
  _strip_images_from_messages (image-rejection recovery) rewrote live
  dicts on the recovery path.

Each site now pops the marker when a persisted field actually changes,
and the agent-aware callers (repair_message_sequence_with_cursor,
turn_iteration_prep, turn_recovery) invalidate the bounded flush-scan
prefix so repaired rows are rewritten on the next flush.
2026-09-20 15:26:34 -07:00
teknium1
13fe9c7171 feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:

- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
  `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
  declaring profile; standard records still replay on OpenRouter-style routes, strict routes
  drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
  status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
  and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
  one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
  without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".

The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.

Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
efc947d72a fix: send the title model call after the turn on a shared custom endpoint
On a `custom` main route (llama.cpp, Ollama, vLLM, ...) whose
auxiliary.title_generation is not pinned elsewhere, the turn prologue fired the
`response_format: json_schema` title request on a daemon thread at the same
instant as the turn's own streaming request, against the same self-hosted
server. A single-slot server can decode the title grammar/completion into the
main reply: the user then receives `{"title": ...}` as the assistant turn, the
main loop persists it as a genuine assistant row, replays it, and the model
adopts the format (#117296). No Hermes writer routes the aux response into the
transcript; the leaked JSON is the main completion itself.

`maybe_auto_title` now returns the upgrade thread and leaves it UNSTARTED when
`title_upgrade_must_wait_for_turn(main_runtime)`; the prologue parks it on
`agent._deferred_title_upgrade` and `finalize_turn` starts it once the model
has answered. Hosted providers keep the turn-start timing. Usage accounting
(`task='title_generation'`) and `sessions.title` are unchanged.
2026-09-20 14:09:57 -07:00
teknium1
3e579ee7af fix: capped @-reference child output always reports a nonzero returncode
_run_quiet drained each pipe up to _MAX_QUIET_OUTPUT_BYTES and relied on
proc.kill() to make the returncode nonzero. A child that flushed past the
cap and exited 0 before the drain thread crossed it (a fast writer on a
loaded runner) was not killable, so the result came back returncode=0 with
truncated stdout: the caller's fallback path keys on the returncode and
treated the truncation as success. This is also why
test_run_quiet_caps_child_output failed on main's CI (`assert 0 != 0`)
while passing locally.

The drain now records that the cap was crossed and the result is forced to
returncode 137 (128 + SIGKILL) when the child exited 0, so the contract in
the docstring holds regardless of scheduling.
2026-09-20 13:57:43 -07:00
teknium1
81faca2f1c fix: failed initialize keeps the LSP error type instead of raising TypeError (review follow-up)
The failure-details rewrap re-instantiated the caught exception with a single
string; LSPRequestError takes (code, message, data), so a JSON-RPC error to
`initialize` surfaced to log_spawn_failed as a TypeError with none of the exit
status / stderr details. Attach the details to the original exception in place
and re-raise. Adds an `init_error` mock-server script and one invariant test
(red on the previous head).
2026-09-20 13:54:58 -07:00
teknium1
d15b17adf5 fix(lsp): log at INFO when a request is skipped because its root is marked broken
After one timeout `_mark_broken_for_file` disables the whole (server, root)
pair for the life of the process, and every later request returned [] with no
log line at all — at default levels a skipped file and a clean file
(`log_clean` is DEBUG) looked identical, so a workspace that silently lost
TypeScript feedback was indistinguishable from a healthy one (#116446, ask 3).

`LSPService.enabled_for` — the gate every production path (snapshot,
get_diagnostics_sync, file_operations_lint) runs through — now emits
`eventlog.log_skipped_broken`: INFO once per (server, root), DEBUG on repeats,
same dedup bucket pattern as the other announce-once events.
2026-09-20 13:54:58 -07:00
Konstantin Khlopkov
b4f0553f8e fix(lsp): report server exit status and stderr tail on a failed initialize (#116446) 2026-09-20 13:54:58 -07:00
teknium1
d03b5f3770 fix(auth): skip the Copilot token exchange while copilot is only an ambient gh-CLI credential
load_pool("copilot") exchanged the gh CLI token on every load — on installs where copilot is
never selected (main provider deepseek, aux slots auto) that is a network round-trip plus the
"degraded to RAW token" warning on every pool load, ~2 per turn, 1000+ times in five days for
the reporter (#114740). The exchange result only matters once a model is routed to copilot, so
the seeder now keeps the raw token and skips the exchange (and the warning) until
is_provider_explicitly_configured("copilot") is true; the load that follows the user selecting
copilot re-seeds and exchanges as before. auxiliary.<task>.provider now counts as selecting a
provider in _config_selects_provider, like a MoA slot, so an aux slot pinned to copilot still
gets the exchanged token. Contributor tests trimmed to two invariants.
2026-09-20 13:22:30 -07:00
Mohamad Kanso
bd2b8124f4 fix(auth): prevent repeated copilot raw token exchange warnings (#114740) 2026-09-20 13:22:30 -07:00
teknium1
1e9d402e07 fix: LSP tree-kill runs off the event loop (review follow-up)
kill_process_tree is synchronous (taskkill /T /F with a 15s timeout on
Windows), so the hard-kill in _cleanup_process blocked the loop for the
duration. Run it via asyncio.to_thread. The graceful path is unchanged:
shutdown() already waits SHUTDOWN_GRACE on proc.wait() after `exit`, so a
well-behaved server exits 0 before cleanup ever reaches the kill (probe:
returncode 0, no kill_process_tree call).
2026-09-20 12:54:18 -07:00
fangliquan
b773bcf271 fix(lsp): hard-kill failed server trees before reaping 2026-09-20 12:54:18 -07:00
fangliquan
bf52a9519a fix(lsp): reap servers cancelled during startup 2026-09-20 12:54:18 -07:00
teknium1
d03d6c2b39 fix(compression): an over-window session that cannot shrink ends the turn with /new guidance and waits one idle budget, not the ceiling
A session far above the model window (~356k tokens on a 131k window in
#116472) re-ran context compression on every turn: a preflight pass that
reclaimed nothing still let the request go to the provider (400 -> overflow
handler -> another pass), and a summary stream that kept emitting tokens
while never committing held the pre-commit wait to the full 600s ceiling.
On the Desktop that blocked the gateway event loop for 10-20 minutes per
turn and the renderer was eventually killed.

- agent/turn_context.py::_fail_closed_on_insufficient_progress: when a
  preflight pass makes no (or sub-5%) progress and the request provably
  exceeds the model window, raise PreflightCompressionTimedOut with
  "start a new session (/new)" guidance so no provider call is sent. An
  unknown window or a fitting request keeps the send-as-is behaviour; a
  pass that no-op'd on a transient guard (summary-failure cooldown) keeps
  its typed cooldown result. Called from both insufficient-progress
  branches of turn_context_compaction._run_preflight_passes.
- agent/conversation_compression.py::run_compress_context_with_progress_timeout:
  an over-window request's pre-commit wait is bounded by one inactivity
  budget (compression.context_timeout_seconds) instead of
  context_total_ceiling_seconds; the existing first-stall deterministic
  fallback then carries the compaction. Config-derived, no new knob.

Slim slice of #116592's Python half.

Co-authored-by: Chukuwebuka-2003 <ebulamicheal@gmail.com>
2026-09-20 12:52:07 -07:00
teknium1
2b3bab0a5c fix(anthropic): key_cmd Claude Code OAuth identity survives on custom api.anthropic.com routes (main + aux)
A named custom provider at api.anthropic.com (api_mode anthropic_messages) whose token comes from a
key_cmd callable lost the Claude Code OAuth identity: the aux custom routes hard-coded
is_oauth=False, agent_init/agent_runtime_helpers/client_lifecycle gated OAuth on provider=="anthropic"
and isinstance(key, str), and the callable-token client builder never added the OAuth betas or the
claude-code user-agent. Anthropic answers such a bare Bearer with 429 rate_limit_error "Error"
(#114967). One resolver, anthropic_credentials.anthropic_route_is_oauth(base_url, credential,
provider=), decides at every site: the route qualifies for the anthropic provider (unchanged) or an
exact api.anthropic.com host, the credential is a string or a callable materialized once
(CommandTokenSource caches), third-party hosts never qualify. model_metadata
_query_anthropic_context_length skips a callable credential instead of crashing agent init
(AttributeError on the same key_cmd route on current main).

Slimmer redo of #115007 by @liuhao1024 (same direction; one shared helper instead of per-site copies).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 12:49:40 -07:00
teknium1
8d153b26aa feat: add z-ai/glm-5.3-flashx to the OpenRouter and Nous Portal catalogs
Both routes serve the slug (tools supported, 1,048,576 context, $0.37/$1.25 per M).
The Nous list is derived from OPENROUTER_MODELS so one tuple edit covers both; the docs
manifest is regenerated in the same commit.

DEFAULT_CONTEXT_LENGTHS gets its own key: substring matching would otherwise land the
slug on the glm-5.3-flash entry (1,310,720) and overstate the window by 25%.
2026-09-20 12:16:06 -07:00
beardthelion
6087b40932 fix(agent): refuse ambient credential chains for multiplex profiles
Under gateway multiplexing, a served profile with no credential of its own
fell through to the SDKs' ambient default chains, which read the process
environment — the launch profile's Azure service principal, az CLI caches,
host managed identity, or AWS keys — and the minted identity rode the served
profile's configured base_url. vertex_adapter already refuses the same
pattern for google.auth.default().

azure_identity_adapter._scoped_credential now raises under multiplexing
when the profile scope has no complete AZURE_* set; AZURE_CLIENT_ID alone
routes to ManagedIdentityCredential as the explicit user-assigned-MI
opt-in. _probe_token propagates the caller's contextvars into its daemon
thread so doctor/probe checks run scoped (a bare Thread ran unscoped and
masked the refusal). describe_active_credential reads all credential
predicates through the scope.

bedrock_adapter.scoped_aws_session_kwargs now requires a complete
credential under multiplexing (key pair, or AWS_PROFILE as the shared-
config opt-in) instead of returning {} and letting boto3.Session() resolve
the ambient chain; the guard runs before the boto3 import at both call
sites. resolve_bedrock_bearer_token reads AWS_BEARER_TOKEN_BEDROCK through
the profile scope under a HERMES_HOME override.
2026-09-20 12:09:47 -07:00
teknium1
6d1341a13b fix: ranged @file read bails at the char budget mid-line (review follow-up)
_next_line collected every readline(line_cap) piece until the newline, so a
one-line giant (minified JSON) was still fully materialized before the
total_chars > char_budget gate. Pass the remaining budget into the helper and
stop collecting as soon as it is exceeded; the caller returns the oversized
block and never needs the rest of the line.
2026-09-20 11:54:43 -07:00
beardthelion
0accda1f76 fix(agent): bound file/process reads in @ context reference expansion
@file:, @folder:, @diff, @staged, and @git: expansion materialized
unbounded amounts of data before any size gate ran:

- _is_binary_file called read_bytes() and sliced [:4096], loading the
  whole file to sniff it.
- _expand_path_reference read_text()ed the whole file before the
  max_inline_tokens check, so a refused file was fully loaded and
  token-scanned anyway; ranged refs read the whole file to serve a
  slice.
- _file_metadata read_text()ed every folder-listing entry (up to 200)
  just to count lines.
- _run_quiet buffered the child's entire stdout, so git diff / rg
  --files on a large tree loaded an unbounded stream.
- gather expanded every parsed reference in a message with no cap.

Remote-triggerable through the gateway: any inbound message can carry
@-references, so a single message could force GB-scale transient
allocations and full-file scans before the token gate refused them.

Bounds applied:

- Binary sniff reads a 4KB prefix via open().
- Whole-file refs stat() first; st_size > max_inline_tokens *
  CHARS_PER_TOKEN is certainly oversized (the estimator is >= bytes/4
  for every encoding mix) and refuses without reading.
- Ranged refs stream the requested window with readline() pieces
  capped at the char budget; lines outside the window are skipped
  without materializing, so single-line giants (minified JSON,
  one-line logs) cannot expand a window read into a full read.
- Folder metadata streams the line count in 1MiB chunks and reports
  byte size past 4MiB.
- _run_quiet drains pipes on threads up to a 4MiB ceiling and kills
  the child on overflow; the nonzero returncode routes callers to
  their existing fallback paths.
- At most 16 references expand per message; the rest get a warning.

Measured: a refused 21MB file cost 42.4MB peak traced allocation
before and ~3KB after; an 85MB mixed hostile message costs ~6KB.
2026-09-20 11:54:43 -07:00
Aaron08140
829c91aa00 fix(agent): run bare script-path hooks on Windows through their interpreter
A hook declared as `command: "~/.hermes/agent-hooks/x.sh"` — the shape every
example in website/docs/user-guide/features/hooks.md uses — cannot start on
Windows. _spawn() shlex.splits the command and Popen()s it with shell=False, so
the kernel reads the shebang on POSIX but CreateProcess on Windows receives a
text file and answers WinError 193. The hook then reports no returncode, which
a fail_closed gate treats as a failure and every other consumer silently skips.

Route a first argument that is an existing file with a mapped suffix through its
interpreter, reusing tools.environments.local._find_bash() so the resolution
keeps the ordering that avoids WSL's bash.exe (#115124) and surfaces Git-for-
Windows' own guidance when it is absent. POSIX argv is untouched. Suffixes we
cannot resolve an interpreter for still fail, but the diagnostic now names the
remedy instead of the OS's localized complaint.

Repairs five of this file's tests that have been red on native Windows
(TestCallbackSubprocess x4, hooks TestHooksTest::test_fires_real_subprocess_and_parses_block);
the two that stay red are drive-letter/`~` tokenization, which #68508 owns.

Verified on Windows 11 26200 / cp936 with real subprocesses: bare .sh, .bash and
.py hooks execute and carry their exit code; a missing path still reads
"command not found"; `git --version` is unaffected.
2026-09-20 11:27:39 -07:00
beardthelion
5aca5406f0 fix(agent): reject subdirectory hint files that resolve outside the tree
The subdirectory-hint loader resolves each visited directory but never the
hint file itself, so a checked-in sub/AGENTS.md symlinked to an out-of-tree
file (~/.aws/credentials, another agent's config) was followed and injected
into the tool result. Resolve each candidate and require its target to stay
inside the working dir and pass the canonical read deny-list - the same
policy @-references already apply - and read the resolved inode so a link
swapped between check and read still lands on the vetted file. In-tree
symlinks keep working.
2026-09-20 10:51:03 -07:00
liuhao1024
c19246557a fix(gemini): stop rerouting AQ. keys off the default Studio surface
Google now issues AQ. keys for both Google AI Studio and Vertex AI
express mode, so the AQ. prefix no longer identifies the key family
(#115306): auto-routing every AQ. key to aiplatform.googleapis.com 403s
the whole AI Studio fleet (6293fca019 / df53cae72d).

- normalize_gemini_base_url no longer rewrites by key shape; an express
  key reaches aiplatform only through an explicitly configured base,
  which is still completed to the publishers/google form (#114335 path)
- gemini_http_error appends two-way 403 PERMISSION_DENIED guidance: an
  AQ. key rejected on the Studio host learns about the express base_url,
  a key rejected on an explicit aiplatform base learns about the default
- doctor's explicitly configured aiplatform base now also gets the
  publishers completion; the OAuth Vertex .../endpoints/openapi base
  stays untouched

Fixes #115306
2026-09-20 10:50:27 -07:00
liuzikaii
d23d6e8218 fix(lsp): use UTF-16 units for document replacement ranges 2026-09-20 10:24:17 -07:00
teknium1
5195c13873 fix: verify-on-stop recognises python.exe / py launcher interpreters (review follow-up)
_is_interpreter_token matched Path(token).name against the bare-name regex,
so a Windows venv path `...\Scripts\python.exe` (and the `py` launcher)
recorded no ad-hoc evidence and the nudge loop the PR closes stayed open on
that platform. Strip a case-insensitive .exe/.bat/.cmd suffix, take the
basename across backslashes, and accept `py`. Parametrized invariant test,
red before.
2026-09-20 10:23:04 -07:00
Uttkarsh Tiwari
35accdbc30 fix(agent): record ad-hoc verify evidence run through a versioned or absolute interpreter
The ad-hoc matcher accepted an interpreter only when the command word was
literally in `_INTERPRETERS`, so `python3.12`, `/usr/bin/python3.12` and
`/usr/bin/env python3` recorded no evidence for the temp `hermes-verify-*`
script the stop-gate nudge asks the agent to run. The workspace stayed
unverified and the nudge re-fired on every stop, so verify-on-stop could not
be satisfied by following its own instructions.

Recognise the interpreter by name (versioned/absolute paths included) and
treat a leading `env` as transparent. Commands that merely name the script
(`rm`, `chmod`, `cat`) stay non-evidence — the same nudge tells the agent to
clean the script up, and a cleanup command must not clear the gate.
2026-09-20 10:23:04 -07:00
beardthelion
04dc1907be harden vault error handling for malformed state
A vault file that decrypts but holds malformed, non-dict, or non-UTF8
JSON escaped the VaultError contract: _read_all raised raw
JSONDecodeError where callers only catch VaultError, so
hermes vault add/list/rm produced tracebacks. Wrap the decode and
parse in VaultError, and catch VaultError in vault_command so every
subcommand prints the clean error line.

vault.source.set and vault sources --enable/--disable also crashed on
a non-dict vault section (vault: true, vault: {bitwarden: true}) via
an unguarded setdefault chain. Coerce through _ensure_dict, the same
shape guard _voice_cfg_dict documents for voice.*.

Fixes #115867
2026-09-20 10:21:52 -07:00
fangliquan
9f7df273a7 fix(auxiliary): prioritize explicit reasoning config 2026-09-20 10:17:39 -07:00
fangliquan
0a407e1651 fix(auxiliary): consume promoted reasoning config 2026-09-20 10:17:39 -07:00
fangliquan
c86de9c44c fix(auxiliary): preserve raw reasoning body shapes 2026-09-20 10:17:39 -07:00
fangliquan
ab3448e075 fix(auxiliary): honor provider reasoning disable controls 2026-09-20 10:17:39 -07:00
finn763
effcf3af06 fix(gateway): a peer DM retries a transient turn failure once, like the other two lanes
`hermes peer dm` posts to POST /api/sessions/{id}/chat, the third Bot-DM
transport. The local (`tools.bot_mode_dm`) and relayed
(`tui_gateway.methods_bot_relay`) lanes both re-run a transiently failed turn
once under the shared policy (`tools.bot_failure_reasons.retry_action`) and
resume the row the failed attempt left as the transcript's unanswered tail;
this lane ran the turn once and handed the provider's 429 paragraph to the
sender as the reply (#115325).

The policy is asked about a result dict now, not two streams: `result_retry_action`
joins `error` + `failure_reason` (the turn loop's own typed verdict) so one
classifier serves every lane, and the server-error rule accepts the providers'
`server_error` / `overloaded_error` spellings — the codes the in-process lanes
key on instead of a status number.

The resume half is the CLI lane's rule extracted to `agent.session_persistence.
adopt_unanswered_turn`, which `quiet_single_query` (env-gated dispatcher re-run)
and the API lane (in-process re-run, on the agent it just built) now share.

The regression drives the real route and the real `_run_agent` over a real
store: a 429 re-runs the same DM once with the persisted row adopted as this
turn's user message (so no second copy), a 401 still reports one attempt.

(cherry picked from commit 8fe6d46ada5b8064bc7132ee956fda998244c10a)
2026-09-20 10:16:51 -07:00
teknium1
0f3d32ec58 refactor(transports): one registry read per api_mode gate; drop the stream-delta hook
Trim the salvaged fix to the shape main wants:

* Every gate asks ``agent.transports.registered_api_modes()`` directly. The three helper
  spellings (``_has_registered_transport``, ``_registry_knows``, ``is_registered_api_mode``)
  and the ``sys.modules`` peek are gone: no transport module imports ``providers`` or
  ``hermes_cli`` at module level, so a plain import cannot re-enter provider discovery.
* ``hermes_cli/auth.py`` late-registration pass dropped — main already re-syncs plugin
  profiles into ``PROVIDER_REGISTRY`` on every registry miss
  (``auth_plugin_providers.registry_lookup`` / ``sync_plugin_provider_registry``, #102123);
  the probe shows a profile registered after the import-time mirror resolves and reaches
  the wire on base.
* ``ProviderTransport.normalize_stream_delta`` and the streaming-assembler hook dropped —
  legacy ``delta.function_call`` translation is a separate concern from api_mode
  propagation and has no in-tree consumer.
* Tests: 15 gate-by-gate unit tests replaced by two invariants that install a REAL plugin
  under a temp HERMES_HOME and walk profile → determine_api_mode → resolve_runtime_provider
  → agent ladder → delegation resolver (positive: red on origin/main; negative: an
  unregistered mode still degrades to chat_completions).
2026-09-20 10:11:40 -07:00
valerdoskin
ef8cdfc389 fix(transports): accept a provider plugin's own api_mode
A provider plugin ships a transport via `register_transport(api_mode, cls)` and
declares that same string as its profile's `api_mode`. Transports are selected by
that string, but every gate that validates an api_mode compared it against a
closed literal, so a plugin's mode was rejected at each one and rewritten to
`chat_completions`. The plugin's transport was then never selected: no error, no
tool call, the turn silently degraded to prose. `register_transport` was a public
seam with no way through.

Accept a mode when the transport registry knows it, via a new
`agent.transports.registered_api_modes()`, at each gate:

* `agent_init._resolve_api_mode` - the agent's mode ladder;
* `runtime_provider._parse_api_mode` - the config gate;
* `delegate_tool_config` - the delegation resolver;
* `providers.get_provider` - the reverse `TRANSPORT_TO_API_MODE` lookup recorded
  an unknown mode as `openai_chat`, which made `determine_api_mode` report
  `chat_completions` for a provider that has a dialect transport. This one is the
  most deceptive: the other gates already pass, and the transport still is not used.
* `providers.determine_api_mode` - the same table lookup at the other end.

The registry read is deliberately lazy (`sys.modules.get("agent.transports")`,
never an import): this code is reached from `determine_api_mode`, which provider
discovery itself calls while the registry is being populated, and importing the
transport package there re-enters discovery.

Also add `ProviderTransport.normalize_stream_delta()`, the response-side twin of
`convert_messages()`: a provider that streams a tool call on the legacy OpenAI
`delta.function_call` pair instead of indexed `delta.tool_calls` had nowhere to
translate it, and the streaming assembler dropped the call. The default returns
the delta unchanged, so existing transports are untouched; the assembler now asks
the transport instead of hardcoding one provider's shape.

Finally, make plugin-provider registration repeatable. `hermes_cli.auth` registered
plugin profiles once, at import, from a list `hermes_cli.config` had already
partially discovered while importing itself. A profile that sorts LAST in discovery
was absent from that snapshot and never reached `PROVIDER_REGISTRY`, so every
consumer reported it unauthenticated and it silently vanished from the model
picker while working fine from the CLI. `ensure_plugin_providers_registered()` is
now called from `get_auth_status()` and `resolve_provider()`, so a late profile is
picked up instead of staying invisible.

Unregistered modes are still rejected everywhere, and the in-tree literal sets are
unchanged - they are simply no longer the only way in.
2026-09-20 10:11:40 -07:00
kshitijk4poor
e74a84f0b6 refactor(agent): declare the trim flag in _CONTROL_STATE; trim only on a completed batch
Gate Lows: an in-flight exception pins the executor frames through its traceback, so the
finally-placed trim could not release the result on that path and would burn the cooldown;
the flag now survives to the next completed batch. The default lives beside _executing_tools.
2026-09-20 18:09:19 +05:30
kshitijk4poor
05e09aee68 fix(agent): trim after the tool batch unwinds, not inside the commit that still holds the result
Gate review (three lenses converged): inside _commit_tool_result the >=1 MB string is still
referenced by the publish frames (the returned tuple, managed.result / batch.results, the
tool.completed callback), so gc.collect + malloc_trim could not release it and merely spent
the 60 s cooldown. The commit now only sets agent._trim_after_tool_batch; the finally of
AIAgent._execute_tool_calls consumes the flag once every executor frame is gone, coalescing
N large results in one batch into one trim. The import is lazy like every sibling call site
(keeps ctypes out of the agent import chain).
2026-09-20 18:09:19 +05:30
kshitijk4poor
f32f651788 perf(tool_executor): trim memory after publishing a >=1 MB tool result
Post-compression already calls trim_memory (#77356); a huge tool result (raw
stdout, file dumps) is the other allocation a turn drops and was published
with no collection. _commit_tool_result is the one point both the sequential
and concurrent publish paths go through, so the trim lives there, after the
spill + session flush, measured on the string already in hand (multimodal
dicts are never re-serialised). trim_memory's own cooldown/kill-switch apply.

Salvages the intent of #80974 without its bare gc.collect(), re-serialisation
and 186 LOC. Closes #70684 (tool-result half).

Co-authored-by: Christopher-Schulze <210261288+Christopher-Schulze@users.noreply.github.com>
2026-09-20 18:09:19 +05:30