v2026.9.21
5548 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2632229bcf |
fix(auth): named profiles read the root auth.json again (revert #111724)
Reverts
|
||
|
|
fe4294802e |
fix: apply the per-model image strip to the max-iterations summary request
agent/chat_completion_helpers.py::_iteration_summary_api_messages hand-builds api_messages from canonical history (shallow row copies) and calls agent._build_api_kwargs directly, bypassing turn_api_request.build_api_request — the sibling send path where strip_images_for_rejecting_model runs. On main this went unnoticed because the image-rejection recovery stripped history itself, so the summary request was image-free by accident. This stack keeps history intact, so a model recorded in agent._image_rejecting_models would receive images in the summary request → 4xx → 'max_iterations_no_summary'. Call strip_images_for_rejecting_model right after the vision eviction, mirroring build_api_request. Hazard-checked: the strip rebinds the row's content to a new list (never mutates the nested list shared with history), so the shallow copies keep history untouched — the new test asserts both the text-only output and the unchanged history. |
||
|
|
6540224f69 |
docs: fix three statements the per-model image strip left stale
The image-rejection recovery no longer switches the session to text-only or strips history; it records the (provider, model) and build_api_request strips images from that model's requests only. Three places still described the old behaviour: - recover_before_classification docstring and the adjacent comment said "switch session to text-only" / "mark session vision-unsupported". - TestStripImagesDropsStaleApiContent's rationale claimed the strip runs on persistent history and that leaving api_content would replay rejected images every turn because the recovery gates on _image_rejecting_models — false on both counts: current callers pass per-call clones. Reworded as the generic helper contract (a rewritten persisted row must drop its sidecar) with the no-op note. - _strip_images_from_messages docstring, same fix. Comments and docstrings only; no code change. |
||
|
|
565b2ce206 |
refactor: dedupe the corrupt-image recovery and split the phrase lists
Two branches in turn_recovery.py carried the same isinstance + _strip_images_from_messages guard and the same "Provider rejected a corrupted image" notice: the new pre-classification corrupt branch and the FailoverReason.image_corrupt branch. Both now call one module helper, _strip_request_images_and_retry(agent, api_messages) -> bool, so the strip-this-attempt-only policy lives in one place. _IMAGE_REJECTION_PHRASES was rebound as unsupported + corrupt, which made _looks_like_image_content_rejection silently cover corrupt payloads and forced the recovery to test the corrupt list a second time to undo that. The name stays (external references) but is now the unsupported-only tuple; _IMAGE_CORRUPT_PHRASES is disjoint, and the recovery asks the two questions explicitly: `_corrupt or (model not yet rejected and unsupported)`. Tests: the phrase-isolation matcher asks the same disjoint question the recovery does; the image_corrupt source-contract check looks for the helper call instead of the inlined strip. |
||
|
|
b5ec7edbc4 |
refactor: drop the redundant strip from the image-rejection recovery
recover_before_classification's capability branch stripped api_messages itself right after recording the (provider, model) in _image_rejecting_models. That strip is dead work: the verdict is `continue`, which re-enters build_api_request with the SAME api_messages object, and strip_images_for_rejecting_model runs there — before _build_api_kwargs on every attempt — and strips them because the key is now in the set. Keeping both meant two places had to agree on the send policy. The corrupt-image branch keeps its local strip on purpose: it deliberately does not record the model, so the send path would not. test_canonical_history_keeps_its_images now asserts the text-only wire copy through strip_images_for_rejecting_model (the real send path) and checks build_api_request still calls it ahead of _build_api_kwargs; commenting that call out turns the test red. |
||
|
|
7ef78c1278 |
refactor: read agent._image_rejecting_models directly
init_agent seeds `_image_rejecting_models = set()` at the same site as `_force_ascii_payload`, which the adjacent sanitize_outbound_kwargs already reads as a plain attribute. The getattr/isinstance guards and the lazy re-create in recover_before_classification implied the attribute could be missing or mistyped on a real agent; it cannot, and defensive fallbacks for impossible states hide wiring bugs instead of surfacing them. Every test fixture that reaches these paths seeds the attribute already. |
||
|
|
26f9cb03c4 |
refactor: reuse _provider_model_key for the image-rejection model key
message_sanitization.image_model_key duplicated vision_message_prep's _provider_model_key — the (provider, model) key the same mixin already uses for its per-model vision bookkeeping. Two keying functions for the same concept drift: one normalised the provider (.strip().lower()), the other did not, so a provider spelled "OpenAI" in one place and "openai" in another would have been tracked as two models. Key `_image_rejecting_models` on the existing helper and delete the duplicate (and its __all__ entry). No import cycle: vision_message_prep imports only lazy_forward, tool_dispatch_helpers and utils, none of which import message_sanitization or turn_recovery. |
||
|
|
b5dacdb125 |
fix: a corrupt-image rejection must not blind the model for the session
recover_before_classification matches _IMAGE_REJECTION_PHRASES, which
mixes two kinds of body: capability rejections ("does not support
images", "only text content type is supported", ...) and bad-payload
rejections. After the per-model tracking landed, BOTH added the
(provider, model) to agent._image_rejecting_models, so one truncated
screenshot rejected with "failed to decode image" made every later
request to that model text-only for the rest of the session, even
though the model can see fine.
Split the bad-payload phrases into _IMAGE_CORRUPT_PHRASES:
- "image data you provided does not represent a valid image"
(ChatGPT-account Codex backend)
- "failed to decode image" (Kimi / Moonshot and other
OpenAI-compatible providers)
_IMAGE_REJECTION_PHRASES stays the union so the turn still recovers on
them. For a corrupt match the recovery now strips the current attempt
only and retries iff something was stripped (mirroring the existing
image_corrupt branch), and leaves the model unmarked; only a capability
rejection records the model.
One test: a "failed to decode image" body strips the wire copy, keeps
history, and leaves _image_rejecting_models empty so the next request
to the same model carries its images.
|
||
|
|
a955e90422 |
refactor: hoist strip_images_for_rejecting_model import to module level
turn_api_request already imports from agent.message_sanitization at the top of the module, so the function-local import added by the salvage is an inconsistency, not a cycle guard. Hoist it onto the existing import line; `python -c 'import agent.turn_api_request'` confirms no cycle. |
||
|
|
28a75285c4 |
refactor: drop the dead _vision_supported flag
Image rejections are now tracked per (provider, model) in agent._image_rejecting_models, and recover_before_classification gates on that set. Nothing reads agent._vision_supported any more, so the write in turn_recovery and the per-turn reset entry in turn_context are dead state. Remove both so the next reader doesn't assume a turn-global vision gate still exists; update the two tests that asserted or seeded the attribute. |
||
|
|
7299015092 |
fix(agent): track image rejections per model across a fallback chain
The first head stored a single rejecting (provider, model) and kept the turn-global `_vision_supported` as the recovery guard. In a fallback chain that fails: model A rejects images and retries text-only, a later error activates model B, the restart rebuilds api_messages from history so B receives the images, and when B rejects them too the branch is skipped because `_vision_supported` is already False — the request falls through to generic error handling. Recording B also overwrote A, so A was no longer treated as text-only on later turns. `_image_rejecting_models` is now a set of every rejecting model, and it is also the guard: each model's first rejection runs the recovery and a repeat rejection from the same model still falls through, so the retry cannot loop. image_model_key() names the key in one place. Adds a test for the two-model sequence (fails on the previous head) and one pinning that a repeat rejection from the same model does not retry. Thanks to @ehz0ah for the review. (cherry picked from commit 225fd76ccf90cd3ec509a9327f725ac02d996f24) |
||
|
|
1ef306f68d |
fix(agent): an image rejection strips the request, never the session history
When a provider 4xx's on image content, recover_before_classification ran _strip_images_from_messages on the canonical `messages` list and reset _db_flush_scan_prefix. Since #117569 that function pops _db_persisted on every rewritten dict, so the next flush rewrote those rows: every image in the session — and every image-only message, which the stripper deletes outright — was removed from state.db for good. The rejection describes what the CURRENT model accepts, not what the conversation holds. An automatic fallback to a text-only provider, or a single /model switch, was enough to erase images the user had sent to a vision model, and switching back found them gone. It is the same failure as the ASCII strip in #117802, on the image path; neither open fix for that issue touches this branch. Keep the repair on the send path, where the per-call copy already lives (_clone_message_for_send exists so send-path rewrites never reach the persisted transcript, #80498): - record the rejecting (provider, model) on the agent, a session-scoped flag initialised beside _force_ascii_payload; - strip the in-flight api_messages copy for the immediate retry; - build_api_request calls strip_images_for_rejecting_model() on each attempt's api_messages BEFORE provider conversion. The stripper knows Hermes's own part types; a converted payload would slip past it (Bedrock Converse image blocks carry no `type`). Keyed on the model, so one that accepts images gets them again. _strip_images_from_messages itself is unchanged, so its role-alternation and sidecar guarantees still hold on the wire copy. The notice no longer claims "text-only mode for this session" (_vision_supported resets every turn) or that images were stripped from history. (cherry picked from commit cbdb184c27f915ab138b2087f878aed7fcc7c76b) |
||
|
|
68f7b90eeb |
fix: return a raced-successful compression from _await_worker_within_budget
The `future.done()` guard added for dead workers (#117261 / #63892) unconditionally logged `future.exception()` and returned `(False, None)`. If the worker finished SUCCESSFULLY in the window between `result(timeout=)` expiring and the `done()` check, that discarded a completed compression and sent the caller down the stall/fallback path — `future.cancel()` becomes a no-op on a settled future and the fallback route costs a second LLM call. Siblings `_await_in_flight_commit` and tool_executor's `_poll_sequential_future` already re-read the result on a settled future; do the same here: `exception() is None` → return `(True, result)`, otherwise take the stall path as before. The compression-seam test grows a settled-successful case through the same helper (a Future subclass whose first timed `result()` still raises TimeoutError, modelling the race); it fails on the previous guard and passes with this one. |
||
|
|
a8a81f4cb9 |
refactor: log the dead compression worker's exception once; trim guard comments
_await_worker_within_budget now emits one INFO line with future.exception() when it takes the stall path for a worker that already died, so the log shows WHY the fallback chain was entered instead of silently returning (False, None) — previously the only trace was the absence of "still streaming" lines. The three 5-7-line comment blocks added by the #117261 pick restated the same alias fact each time; each is now 2 lines citing #63892 and the 3.11 alias once. No behaviour change beyond the log line. |
||
|
|
7740a4ac20 |
fix: report a worker that died with TimeoutError as exited in _join_cancelled_worker
_join_cancelled_worker returned False from its `except concurrent.futures.TimeoutError:` arm. On 3.11+ that class IS the builtin TimeoutError, so the arm also fires when the cancelled worker itself DIED raising a timeout-class error (the aux client raises bare TimeoutError on a stalled summary stream). The caller, _release_cancelled_worker, treats False as "still running": it logs 'did not exit within grace' and skips fence.allow_cancelled_lock_release(), so the session compression lease of a provably-dead worker was retained/orphaned. Return future.done() instead: a settled future never becomes unsettled, so a done future means the thread exited and the lease can be released. A live worker that merely outlasted the grace still yields False. Same guard family as the sibling loops fixed in the preceding pick (#117261). |
||
|
|
8305113328 |
fix(compression): don't mistake a worker's TimeoutError for a poll timeout
On Python 3.11+ concurrent.futures.TimeoutError IS the builtin TimeoutError
(asyncio.TimeoutError and socket.timeout alias it too). Poll loops shaped like
try:
return future.result(timeout=slice)
except concurrent.futures.TimeoutError:
...keep waiting...
therefore cannot distinguish "the wait slice expired" (worker alive) from "the
worker raised TimeoutError" (worker dead). auxiliary_client raises a bare
TimeoutError when a summary stream stalls, so this is reachable in production.
When it happened the host re-waited on an already-settled future. result() then
returned instantly every iteration, spinning at ~2k iterations/sec and logging
"Context compression still streaming" about a dead worker, until the entire idle
budget elapsed. One session burned 535s and wrote ~90k duplicate log lines
(15.6MiB) before failing with context_compression_timeout, and every later turn
re-entered the same path.
Guard each loop with future.done(): a settled future never becomes unsettled.
- _await_worker_within_budget: take the stall path at once, so the configured
fallback chain is actually reached instead of after a 120s false stall.
- _await_in_flight_commit: re-raise the worker's exception. This loop had no
ceiling, so a dead worker spun forever.
- tool_executor._poll_sequential_future: same, and with deadline=None it also
span indefinitely.
Non-timeout worker exceptions still propagate unchanged, and a live worker still
polls exactly as before.
Regression tests pin all three. Against unpatched code the two wait tests fail
and the commit-wait test hangs until the 300s harness SIGKILL, reproducing the
infinite loop directly.
(cherry picked from commit 274457304ce0393407574fe4e43c4a450f20bac3)
|
||
|
|
c27219adc8 |
refactor: detach aliased tools with _clone_message_for_send, not deepcopy
sanitize_outbound_kwargs detached the agent.tools alias with copy.deepcopy before the ASCII strip. The repo's send-path detach idiom is the structural clone _clone_message_for_send (dicts/lists recursively, immutable leaves shared), which is what every other outbound copy uses and is cheaper on JSON-shaped, acyclic payloads. It is sufficient here because _sanitize_structure only rebinds str leaves inside dict/list containers, so the clone fully isolates the canonical tool schemas. Imported lazily inside the function (conversation_loop imports this module) exactly as turn_finalizer does. Drops the now-unused `import copy`. |
||
|
|
a0dd43cc8c |
refactor: return a single bool from _repair_transport_credentials
Both callers immediately reduced the (headers_sanitized, credential_sanitized) tuple with `or`; nobody distinguished the two. Returning one `transport_repaired` bool removes the unpacking at both call sites and the redundant flag bookkeeping inside the helper without changing behaviour. |
||
|
|
307428377b |
refactor: drop dead api_kwargs strip from the ASCII-codec recovery branch
The ASCII branch of _recover_unicode_encode_error popped/stripped/restored `tools` on the failed attempt's api_kwargs and tracked `_tools_sanitized`, but that dict is discarded: build_api_request rebuilds api_kwargs from agent.tools on every retry iteration and sanitize_outbound_kwargs already strips the whole payload once `_force_ascii_payload` is set. Only api_messages (reused across retries) and the local active_system_prompt still need stripping here. Also removes the `_force_ascii_payload = False` assignment in the UTF-8 branch (the flag is initialised False in agent_init and only ever set True in this function), the now-unused `_sanitize_tools_non_ascii` import and the dead `import copy`. The test drops its assertion on the recovery dict's tools identity; the chokepoint assertion (agent.tools byte-stable after sanitize_outbound_kwargs) is kept. |
||
|
|
d47af1421b |
refactor: share the header/api-key repair between both ASCII-codec branches
_recover_unicode_encode_error carried two copies of the same block that strips non-ASCII from _client_kwargs["default_headers"] and the API key (agent.api_key, _client_kwargs["api_key"], client.api_key). The UTF-8 runtime branch was a silent copy; the ASCII runtime branch also printed the "bad copy-paste?" hint. Two copies drift, and a user whose key was repaired under a UTF-8 runtime got no hint about why auth might now fail. Extract one module-level _repair_transport_credentials(agent) returning (headers_sanitized, credential_sanitized) and call it from both branches. The hint is emitted whenever the key changed, in both runtimes; the retry decisions in each branch are unchanged. Net -7 LOC. |
||
|
|
336d9662cc |
fix: detach aliased agent.tools at the outbound sanitization chokepoint
api_kwargs["tools"] is rebuilt from agent.tools on every attempt (_build_api_kwargs_for_mode: tools_for_api = agent.tools; transports set api_kwargs["tools"] = tools without copying), so the list usually aliases the canonical tool schemas. The ASCII retry path then runs sanitize_outbound_kwargs with _force_ascii_payload set, and its in-place strip rewrote agent.tools for the rest of the session. Move the guard to the chokepoint: when the flag is set and tools IS agent.tools, deepcopy before stripping. The deepcopy the contributor pick added inside _recover_unicode_encode_error only protected the failed request's kwargs, which are discarded before the retry; drop it and have recovery skip an aliased tools list entirely (request-local lists are still stripped for the diagnostic message). The kept test now exercises the chokepoint directly: with the flag set on kwargs whose tools aliases agent.tools, agent.tools must be byte-stable afterwards. Verified red against the previous chokepoint. |
||
|
|
05287a5369 |
fix: surface ascii-codec error when UTF-8 runtime repairs nothing
Under a UTF-8 runtime, _recover_unicode_encode_error only repairs the api_key and _client_kwargs['default_headers']. When neither changed, the retry is byte-identical to the failed request and cannot succeed; it just consumed both _unicode_sanitization_passes before the error finally surfaced. Return False in that case so recover_before_classification falls through to the normal classification/error path immediately. The existing api_key + default_headers sanitization is unchanged; the "retrying unchanged request content" message is removed because that branch no longer retries. The PR's UTF-8 invariant test keeps asserting the request copy is not stripped, and now asserts the False return. |
||
|
|
17dfeb1c55 |
fix: preserve conversation history during unicode recovery
(cherry picked from commit d2cec3b33dc728d702b037b0b9c678672322ce3c) |
||
|
|
af66d5db15 |
fix(agent): remote backend probe no longer puts user, $HOME and cwd into the system prompt
The non-local terminal-backend probe (`_BACKEND_PROBE_CMD` / `_format_backend_probe`) ran `whoami`, `$HOME` and `pwd` inside the sandbox and rendered `User:` / `Home:` / `Working directory:` lines into the system prompt on every turn. Nothing downstream consumes those values — the only reader is the model, which can `whoami && pwd` when a task needs them — so they were user-identifying metadata sent to the provider for no behavioural gain. The probe now asks for and renders only `OS: <uname -s> <uname -r>`, and the prompt block tells the model how to fetch the rest on demand. Fixes #117262 |
||
|
|
b7803a1763 |
fix(compression): cap the protected tail at 20% of the context window
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the "protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B timed out before compaction ever changed anything, and protect_last_n read as an uncompressed tail rather than a minimum. TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk / pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of 32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%). Probe (12 tool-heavy turns, 49 rows, 12.8K tokens): ctx 8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows; window [4,10) -> [4,42) ctx 16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34) ctx 32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22) ctx 128K: identical before/after |
||
|
|
afc3b7c6f3 |
feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash merges ( |
||
|
|
6ba45b0e06 |
fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed holder scan doctor and repair use before rewriting the store. While a gateway, Desktop, dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as `PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning, `--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and every agent answered every turn with the retired-WAL refusal until all writers were stopped by hand (#110054, maintainer follow-up 09-20). The DeletedWalGenerationError text is now two layers: a first sentence for the person reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete files while they run, docs link), then the operator detail. The classifier fingerprint "deleted state.db-wal or state.db-shm" is unchanged. The cause table (`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway home-channel notice) and the chat explainer carry the same first steps; the gateway notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause, which for a held retired generation is the second-writer trap. New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the guard text, the developer state-db-recovery page and the sessions guide): the three steps, the do-nots, why maintenance refuses, and what the files beside state.db are (retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups, snapshots). |
||
|
|
60dd9778c6 |
fix(desktop): large text pastes attach from HERMES_HOME/composer-pastes when the chat cwd is elsewhere
Desktop persisted a large paste under Electron's userData dir and attached it as `@file:<abs path>`; the tui_gateway prompt path expands that reference with `allowed_root=cwd`, so `_resolve_path` refused it with "path is outside the allowed workspace" whenever the chat's cwd was not an ancestor of the paste dir (always on Linux/macOS, and on Windows for any project cwd). Electron now writes pastes under `<HERMES_HOME>/composer-pastes`, and `agent/context_references.py::_resolve_path` admits exactly that anchored directory (active profile home and global root, via file_safety._hermes_dirs) as the one root besides `allowed_root`. A sibling directory that merely contains the substring stays refused; the credential deny-list still runs on the admitted path. Supersedes the substring whitelist proposed in #117150. Fixes #117149 |
||
|
|
c545272568 |
fix(lsp): one stalled request no longer silences a workspace for good — retry window, cold-root warm-up budget, per-root exclusion
A language server that missed its budget once marked its (server, root) pair broken for the process lifetime, the same 5 s steady-state budget was applied to a cold server that also had to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched the server off for every workspace. Three new keys under the existing `lsp` block, all defaulting to today's behaviour: - lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per pair; an expired pair gets one more try, and the INFO skip line names the retry time. - lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client waits up to this budget (outer join budget follows); warm requests keep wait_timeout. - lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path also covers everything beneath it); a matching root never spawns, logged once at INFO. A non-list value fails closed — WARNING naming the expected shape, every root skipped — because silently excluding nothing would re-pay the stall the key was meant to avoid. Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo). |
||
|
|
7d3c0b2f94 |
fix(compression): an auto-resolved summary model that fails falls back to the main model and is named in the warning
`provider: auto` resolves a compression summary model per call WITHOUT setting `summary_model`, so the main-model retry gate saw "no separate model" and re-hit the same bad route (e.g. a proxy channel answering HTTP 200 with empty content) on every attempt, and the user-visible aux-failure warning had no model to name. Record the model the aux lane actually resolved (`_last_aux_resolved_model`), use it in the retry gate, and pass it into `_fallback_to_main_for_compression` so the warning names it. Cherry-picked from #116592 (9bb16bd6aee6): only the context_compressor / _last_aux_resolved_model hunks, the attempt-state field, and its test; the over-window wait cap, preflight fail-closed and Desktop renderer hunks are carried by #117084 / #117140. Part of #116472 (request 4). |
||
|
|
bb9058d7e8 |
fix(compression): preserve steer display identity
(cherry picked from commit a8570878bc848a42cc8029fc45219a57e76be4c4) |
||
|
|
f568b860d7 |
fix(agent): pin the reset_at reach through try_activate_fallback and add opt-in fallback.min_switch_reset_seconds
- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback (production entry) on a real AIAgent with a one-entry chain, so dropping the reset_at forwarding goes red (3 failures before, 8 green after). - #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the rate-limited primary's declared reset is sooner than N seconds, try_activate_fallback returns False and no cooldown is armed; docs row added. |
||
|
|
125bdff601 | fix(agent): honor provider reset for fallback cooldown | ||
|
|
36b6efc297 |
fix(context): re-derive model.context_length on model/provider change
model.context_length is the user's profile-wide ceiling. It was read from config.yaml in exactly one place — agent construction — and cached twice: agent._config_context_length (switch/fallback resolution plus every display and /usage surface) and context_compressor._config_context_length (the compressor's own re-resolution). Every live path that re-resolved a runtime then touched only one copy, or cleared it without re-reading the config: - switch_model nulled agent._config_context_length and re-derived the intent from custom_providers metadata alone, so a ceiling that only exists as model.context_length was dropped for the rest of the process; - the Desktop/TUI compression hot-reload updated the compressor's copy only, so an open session showed a pinned ceiling while compressing against provider metadata / the 256K fallback. Both now route through one pair of helpers in agent/agent_init.py: set_config_context_length (one place that knows where the pin is cached) and config_context_length_for_runtime (re-read from live config, scoped exactly like construction, so an unrelated route still never inherits the pin). (cherry picked from commit 986ff16dadb9966f7328e55f295af5cfa1eb5c88) |
||
|
|
de622b291d |
fix(agent): detect Thai plan tails in promoted-reasoning stall guard
promoted_reasoning_announces_action()'s tail detector only matched English (plus CJK punctuation boundaries), so a reasoning-only clean stop ending on a Thai first-person plan (e.g. "จะให้ผม...") was not recognized as a stall and got delivered to the user as the final answer instead of nudging continuation. Add Thai first-person future-action triggers, extend the boundary class with em/en dash (a common Thai clause separator), and accept multi-dot ellipsis tails. (cherry picked from commit 87061e46c423859cf738d4541df6594233f40e7e) |
||
|
|
53815e24dc |
fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.
The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.
Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
before req_reasoning: {}
after req_reasoning: {'reasoning_effort': 'medium'}
agent.reasoning_effort: low -> {'reasoning_effort': 'low'} (unchanged)
|
||
|
|
3a37a24efd |
fix(moa): mark reused advisor guidance as predating the tool results
With `fanout: user_turn` (and off-cadence `every_n` iterations) the advisors run once per user turn and their guidance is replayed verbatim into every later iteration of that turn. The block reads as fresh instruction, so an advisor that proposes a tool call keeps proposing it after the acting model already ran it and has the result in the transcript. Observed with the clarify tool: the card was answered, the next iteration replayed the same guidance, the model issued a second identical card, the user dismissed it, and the turn then held two contradicting results for one question (an answer and an empty skip). The following aggregation degenerated into a repetition loop until it hit the output cap. - `_STALE_GUIDANCE_NOTE` is appended when cached guidance is reused on an iteration that already has tool activity since the last real user turn. The advice text is handed over unchanged; only the framing says it may be out of date. - The advisor system prompt now rules out emitting a tool call or a JSON tool-call object. Advisors hold no tools, and a tool-call object in advisory text is what the aggregator replays. No change to fanout cadence, caching or accounting. The cadence test that pinned byte-identical reuse now asserts the advice text is reused and carries the marker. Signed-off-by: Moep90 <3042152+Moep90@users.noreply.github.com> (cherry picked from commit 343787f287ad9915345090fba351df5ffa758c9e) |
||
|
|
0765099ff4 |
fix(compression): fence the durable cooldown rollback per compressor; one stale-attempt helper
The SQLite rollback no longer runs under the process-wide claim lock: a per-compressor serial lock (taken by _claim_compressor_attempt too) serializes it against claims on that compressor only. The seven pasted working-attempt checks call _raise_if_stale_attempt/_caller_attempt_is_current. Drops the unused _run_as_attempt test helper. |
||
|
|
0e33dc9ebc |
fix(compression): stop detached stale attempts writing shared compressor state
The stall-fallback detaches a timed-out primary worker and reuses the same ContextCompressor, but the existing attempt-generation guards only covered the unwind-time snapshot restore. Every other summary-state write stayed reachable by the still-running primary after the fallback took over: a late successful summary published _previous_summary and cleared the fallback's cooldown, a late failure armed a shared failure cooldown and stamped error state, the cancel rollback and the abort rollback reverted _previous_summary to the primary's snapshot, and the durable cooldown rollback row could be overwritten mid-restore. Compressor code could not fix this with the shared attributes alone: those cells only name the current owner, never the calling attempt. The calling attempt's generation now rides a ContextVar bound inside _run_summary_dispatch, which every attempt's compress_fn passes through in its own thread, so each attempt reads its own generation. Gates on the working-attempt marker (not the entry claim, so lock sit-outs do not suppress the owner) now cover the cancel rollback, late-success writes, _on_summary_failure, the abort rollback, the deterministic pin, compress() entry, and a Phase-3 choke point. The durable cooldown rollback moved inside the claim lock so the DB row and the in-memory restore are atomic against _claim_compressor_attempt. Regression tests drive the real interleavings deterministically, including two threaded end-to-end arms through _run_summary_dispatch and a real ContextCompressor. (cherry picked from commit 902bfc229e140becfb36679dc33bad550c2ae1e8) |
||
|
|
85564321be |
fix(windows): route every bare-bash spawn through _find_bash and surface silent interpreter failures
CreateProcess resolves a bare "bash" to System32 WSL launcher before PATH,
and shutil.which("bash") inherits PATH order (#115124), so node bootstrap,
the TUI node probe and webhook filter scripts ran the wrong interpreter on
Windows. All four sites now use tools.environments.local._find_bash (Git Bash
first, probed). rc!=0 with no output at all is now a WARNING in webhook
filters and an explicit [inline-shell exit N with no output] marker in skills.
Co-authored-by: funky-xamarin <30426178+Wenfengcheng@users.noreply.github.com>
|
||
|
|
1a24b851e5 | fix(skills): resolve native Git Bash for Windows inline shell | ||
|
|
275002d7a4 |
fix(context_references): @folder: listing works outside cwd under a widened allowed_root
@folder: targets resolve against allowed_root, which callers may widen beyond cwd, but _build_folder_listing and _iter_visible_entries assumed the resolved folder was under cwd: path.relative_to(cwd) raised ValueError and the blanket except in _expand_reference surfaced it as a confusing "not in the subpath of" warning instead of a listing. The listing header now renders cwd-relative when possible, then allowed_root-relative, else the absolute path. rg --files gets the absolute folder path (rg echoes the arg as the output prefix, so the lines parse correctly anywhere), the parent-dir walk drops only the cwd stop-condition so in-cwd output is unchanged, and entry indentation is computed relative to the target folder rather than cwd. The os.walk fallback never assumed cwd. Regression tests cover the widened-root target through the real preprocess_context_references entry on both the rg and rg-blocked paths, plus @file: parity and in-cwd display controls. |
||
|
|
30d14a3f73 |
fix: name the CommandCode upstream-outage pattern
The cherry-pick conflict dropped the contributor comment hunk; keep the WHY next to the entry. |
||
|
|
1f3c882d81 | fix(agent): narrow upstream outage matching | ||
|
|
09b72bc6d2 |
fix(compression): track commit fences as a registration stack
_compress_context published the active commit fence with a save/restore cell: registration order was serialized by the fence lock, but completion order is not. When attempt B registered over A and A finished first, A's finally popped the slot, deleting B's live fence mid-attempt (hard_interrupt lost the handle serializing cancel admission against B's begin_commit). B's finally then republished A's dead fence, which lingered until the next compression. The same clobber existed in _publish_new_fence, which overwrote the slot unconditionally when minting the stall-fallback retry fence. Replace the cell with a stack of per-attempt registrations. The finally removes only its own registration and republishes the newest live entry (or clears the slot), so a dead fence can never be restored over a live newer attempt. The stall-fallback retry swaps its fence inside the owning registration and publishes only while that attempt still holds the top registration. Registration moved inside the try so an early exception cannot strand an entry. (cherry picked from commit 574e9945cf186071c3da23c4bb517c0cbdaf73e6) |
||
|
|
a48b4c7d25 |
fix(agent): pop _db_persisted on in-place mutations of stamped live dicts
The _db_persisted marker asserts that a message dict's persisted row is durable as written; any in-place mutation must pop it or the flush scan identity-skips the dict and state.db keeps the stale row forever. Six mutation sites violated the contract: - micro_compaction._merge_adjacent_user_turns rewrote content on a carried-forward dict after superseding a stale micro marker. On the archive-failure path nothing re-stamps, so the merged text never reached state.db. - repair_message_sequence passes mutated stamped survivors in place: _merge_assistant_into (tool_calls union, content join, reasoning_content carry), _prune_unanswered_tool_calls (tool_calls rewrite), _merge_consecutive_users (content join). - sanitize_tool_call_arguments rewrote corrupted/blank function.arguments and prepended the corruption marker onto stamped resumed rows, leaving the corrupt bytes durable and self-perpetuating across resumes. - _sanitize_messages (surrogate and non-ASCII recovery) and _strip_images_from_messages (image-rejection recovery) rewrote live dicts on the recovery path. Each site now pops the marker when a persisted field actually changes, and the agent-aware callers (repair_message_sequence_with_cursor, turn_iteration_prep, turn_recovery) invalidate the bounded flush-scan prefix so repaired rows are rewritten on the next flush. |
||
|
|
13fe9c7171 |
feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a standalone `kind: model-provider` plugin instead of a bundled one: - ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`, `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op). - Chat Completions transport: provider-native `reasoning_details` carriers follow only their declaring profile; standard records still replay on OpenRouter-style routes, strict routes drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim. - `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes, and never writes config when the executable is missing. - `/model` and the pickers: process providers list their live catalog merged with the pinned one, declared aliases/ids resolve inside the provider, and validation accepts a listed id without probing `process://`. - Delegation keeps the selected external-process provider and protocol for the child. - Model metadata / usage pricing consult the profile's bound and cost hooks first. - Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5". The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the PR were dropped in favour of main's `_model_flow_plugin_provider`. Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com> |
||
|
|
efc947d72a |
fix: send the title model call after the turn on a shared custom endpoint
On a `custom` main route (llama.cpp, Ollama, vLLM, ...) whose
auxiliary.title_generation is not pinned elsewhere, the turn prologue fired the
`response_format: json_schema` title request on a daemon thread at the same
instant as the turn's own streaming request, against the same self-hosted
server. A single-slot server can decode the title grammar/completion into the
main reply: the user then receives `{"title": ...}` as the assistant turn, the
main loop persists it as a genuine assistant row, replays it, and the model
adopts the format (#117296). No Hermes writer routes the aux response into the
transcript; the leaked JSON is the main completion itself.
`maybe_auto_title` now returns the upgrade thread and leaves it UNSTARTED when
`title_upgrade_must_wait_for_turn(main_runtime)`; the prologue parks it on
`agent._deferred_title_upgrade` and `finalize_turn` starts it once the model
has answered. Hosted providers keep the turn-start timing. Usage accounting
(`task='title_generation'`) and `sessions.title` are unchanged.
|
||
|
|
3e579ee7af |
fix: capped @-reference child output always reports a nonzero returncode
_run_quiet drained each pipe up to _MAX_QUIET_OUTPUT_BYTES and relied on proc.kill() to make the returncode nonzero. A child that flushed past the cap and exited 0 before the drain thread crossed it (a fast writer on a loaded runner) was not killable, so the result came back returncode=0 with truncated stdout: the caller's fallback path keys on the returncode and treated the truncation as success. This is also why test_run_quiet_caps_child_output failed on main's CI (`assert 0 != 0`) while passing locally. The drain now records that the cap was crossed and the result is forced to returncode 137 (128 + SIGKILL) when the child exited 0, so the contract in the docstring holds regardless of scheduling. |
||
|
|
81faca2f1c |
fix: failed initialize keeps the LSP error type instead of raising TypeError (review follow-up)
The failure-details rewrap re-instantiated the caught exception with a single string; LSPRequestError takes (code, message, data), so a JSON-RPC error to `initialize` surfaced to log_spawn_failed as a TypeError with none of the exit status / stderr details. Attach the details to the original exception in place and re-raise. Adds an `init_error` mock-server script and one invariant test (red on the previous head). |