Commit Graph

5972 Commits

Author SHA1 Message Date
kshitijk4poor
11c50f05d0 refactor(compression): Astra native-compaction gate reuses is_astra_model
The -900k alias fix hand-rolled a second copy of the Astra slug set and its
vendor-prefix normalization. agent/reasoning_effort.py::is_astra_model is the
documented single home for that set (picker, effort vocabulary and request
sanitizer already key off it), so the gate now calls it and a future Astra
alias stays a one-line edit. The gpt-5.6 marker check is back to main's exact
form.

Tests move into the existing parametrized Astra gate table, which checks both
the capability resolver and the per-request gate: -900k on official Codex OAuth
is eligible; -900k through a relay or on provider openai is not. Docs and the
config example no longer say "exact gpt-6-astra".
2026-09-26 19:27:41 +05:30
qqx
60b99e1932 fix(compression): enable native compaction for Astra 900K alias
(cherry picked from commit 8ce2851f0495db8afb44118928f4c3695bb12e06)
2026-09-26 19:27:41 +05:30
ryrenz
6d42b876f2 fix(bedrock): add Claude Opus 5 to the 1M context table
DEFAULT_CONTEXT_LENGTHS declares claude-opus-5 as 1M, but the Bedrock
static table never got the entry, so the offline path resolved the 128K
default and the agent compressed context ~8x early. Add the key and pin
the table pairing with a test so the next 1M generation cannot drift.

Co-authored-by: JiaDe-Wu <JiaDe-Wu@users.noreply.github.com>
(cherry picked from commit c1be16ecbdd267b14124dd35c0b92ccfcd201818)
2026-09-26 07:14:26 +05:30
teknium1
9863980c61 fix: a stale stream attempt advances the cross-turn breaker once, not per timer re-fire
The streaming stale monitor re-fires every stale window while its worker
has not started/dispatched the attempt yet or is still unwinding a kill,
and each re-fire bumped agent._consecutive_stale_streams. One provider
attempt could therefore count twice (or more), so a turn with four hung
attempts reached HERMES_STREAM_STALE_GIVEUP=5 and the breaker refused the
NEXT turn although the provider only ever saw four stale requests. The
non-streaming and inline watchdogs already count once per call.

Count each started stream attempt once (attempt 0, i.e. nothing started,
never counts), and let the interrupted-wait bump honour the same ledger.

This is the root cause of the intermittent
test_agent_turn_liveness[provider_hang] red on CI (fault calls 4, five
"Stream stale for 3s" kills, probe requests 0, breaker text "5
consecutive stale attempts"): on a busy runner the first attempt's
worker takes >3 s to reach the wire, the monitor kills it before
dispatch (tcp_force_closed=0) and again 3 s after dispatch.

Repro: a 3.6 s sleep before the worker opens its first stream
reproduces the CI signature 4/4 on origin/main and 0/6 with the fix;
under taskset CPU starvation the unmodified module goes red 2/3 on base.
2026-09-26 06:26:48 +05:30
kshitijk4poor
1f041552d7 fix(stream): tidy clean-EOF diagnostics and truncation copy
- _mark_finish_seen helper on the local _diag; also set on the Anthropic
  path when message_delta carries stop_reason
- emit_stream_drop status: 'attempt N/M dropped, reconnecting'
- clean-EOF log wording: server or proxy closed the stream cleanly
- truncated_unreported copy drops the raw finish_reason placeholder
- turn_tool_validation compares against FINISH_REASON_LENGTH
2026-09-26 06:10:53 +05:30
kshitijk4poor
a142d963e0 fix(agent): clean-EOF tool-call retry exhaustion no longer blames the network (#102766)
When the 4 truncated-tool-call retries are exhausted on a clean-EOF stub
(no transport error, no finish_reason), use a dedicated stream_closed_tool_call
copy and failure_reason=truncated instead of stream_dropped_tool_call
('check your network') stamped as timeout. Also compute the truncation
banner in a plain if/elif chain instead of a 4-arm conditional expression.
2026-09-26 06:10:53 +05:30
rainbowgits
d43d97d1c3 fix(agent): stop reporting transport/router truncation as an output-length limit (#91717)
Hand-applied from PR #91738 (1da9001929): the truncation detector moved from
conversation_loop.py to the _truncated branch of agent/turn_tool_validation.py.
Only claim the output cap when finish_reason == 'length'; otherwise use new
hedged site copy 'truncated_unreported' (agent/turn_failure_copy.py).
2026-09-26 06:10:53 +05:30
codexbt
7c10f13f5c fix(stream): report 1-indexed drop attempt in _retry_after_drop (#90215)
MessageStream._retry_after_drop passed attempt + 2 to _emit_stream_drop, so
the first drop was announced as attempt 2/N. Hand-applied from PR #90254
onto main's MessageStream split (original hunks targeted the pre-split
interruptible_streaming_api_call).
2026-09-26 06:10:53 +05:30
chelsealong
a2a19bcc77 fix(stream): distinguish clean-EOF no-finish_reason from a transport drop
Response truncated - stream ended before completion" and the log line
"...treating as a mid-stream drop" fired identically for two distinct
failure classes: a genuine transport exception mid-stream, and the
provider closing the connection cleanly (no exception) without ever
sending a terminal finish_reason chunk. The second case is not a
network drop, but the shared wording sent users chasing a network
problem that never existed (#102766).

_build_partial_stream_stub() now tags its stub _clean_eof=True (every
caller is a clean-EOF site: the streaming loop ended without an
exception). The conversation loop uses that tag to pick distinct
wording; the exception-driven stub is unchanged. The two clean-EOF
log lines in chat_completion_helpers.py no longer say "drop".

stream_diag's per-attempt dict also gains finish_reason_seen, set the
moment a terminal chunk is observed, so log_stream_retry's retry/
exhaustion WARNING can say whether the dying attempt had already seen
a finish_reason before the exception hit.

(cherry picked from commit 2c1855ed6dd0037eebe0b1068628999832617109)
2026-09-26 06:10:53 +05:30
Cursor Agent
6814eb4c08 fix(desktop): stage first-contact profile-build onboarding in tui_gateway
Desktop chat goes through tui_gateway, which never applied the gateway's
first-message profile-build sidecar note. Add shared first_contact_turn_note
helper and wire it into the prompt turn so fresh installs get the same
opt-in profile setup offer as messaging surfaces.

Adapted to current main (salvage of open PR #82765, fixes #82750):
_run_prompt_submit now lives in tui_gateway/prompt_turn.py; the note is
staged on agent._gateway_turn_context_notes (the existing sidecar channel,
consumed by agent.turn_context on the user message) so the system prompt
stays byte-stable; the prior-session probe rides _session_db(session).
The low-value onboarding test classes purged from main were not re-added —
only new coverage for first_contact_turn_note and the staging path.

Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
(cherry picked from commit b843900ca818912f0480f6926a059d292b63ffe6)
2026-09-25 19:10:04 -05:00
Hermes Agent
541e4dc0ba fix(agent): protect the runtime's own interpreter from agent deletes
A session asked to clean up older Pythons removed the uv-managed base
interpreter its own venv depended on; the next boot died with 'uv
trampoline failed to spawn Python child process' and no agent tool could
repair it, because the agent itself no longer started (#58748). Prior
uninstall detection (85ce25687e) only flagged package-manager commands.

Add agent/runtime_self_protection.py and wire it into both layers:

- The approval floor (_floor_block) now blocks shell commands that
  delete the running interpreter, its own venv, the pyvenv.cfg base, or
  the uv-managed install directory — rm/rmdir/rd/del/erase/Remove-Item
  with any flags, find <root> -delete, and uv python uninstall of the
  running version (including --all). The floor runs before yolo /
  approvals.mode=off / cron approve mode, so no session setting can
  bypass it.
- The file-safety write classifier denies write/patch/move/delete to the
  same paths, so the file tools cannot overwrite the interpreter either.

Only the runtime the process itself boots from is protected; every other
venv and interpreter on the machine stays manageable.

Fixes #58748
2026-09-25 16:24:21 -05:00
Hermes Agent
ec5c9c738a fix(processes): persist_on_release keeps background jobs alive across lifecycle kill sweeps (#41225)
Background processes spawned with terminal(background=true) are killed from
three agent-lifecycle sweeps: agent release()'s kill_all, a gateway turn
timeout's kill_started_since, and agent close's owned-process loop. Jobs the
user explicitly wants to outlive the session (overnight batches, watchful
daemons) had no way to opt out.

Add terminal(background=true, persist_on_release=true):
- ProcessSession.persist_on_release, stamped by spawn_local/spawn_via_env,
  carried in crash-recovery checkpoints and exposed via list_sessions()
- kill_all skips persisted sessions only for lifecycle sources
  (_LIFECYCLE_KILL_SOURCES: kill_all, gateway_turn_timeout, agent_close);
  explicit operator stops (process_manage kill, /stop slash + RPC mirror,
  CLI /stop) now pass distinct sources so they still reach persisted jobs
- the agent_close owned-process loop in _close_task_resources skips
  persisted sessions the same way
- gateway shutdown keeps killing persisted jobs (source=gateway_shutdown):
  the host is going away and survivors would become PPID=1 orphans

Co-authored-by: salvaged from #109846 (persist_on_release plumbing) and
extended to the turn-timeout and agent_close paths.
2026-09-25 13:49:31 -05:00
Austin Pickett
9df628615c Merge pull request #121614 from NousResearch/austin/fix/121347-base-url-resolution
fix(providers): honor a configured base_url instead of the vendor host across runtime, fallback, listing and validation
2026-09-25 14:15:01 -04:00
Hermes Agent
a3a85a3143 fix(agent): close non-interrupted tool-tail turns with a visible response
A turn that falls out of the loop after a tool result with no follow-up
assistant text left the durable transcript ending at a raw tool row and
returned a silent result: Desktop/TUI showed a ready composer (or kept
spinning) with no final message. The pending_tool_result explainer copy
already existed but nothing minted the exit reason.

finalize_turn now detects the non-interrupted tool tail, fails the turn
with turn_exit_reason=pending_tool_result, and synthesizes the visible
assistant close before persistence so the durable tail is alternation-
safe. Stream-recovered turns (#95514) and interrupted tails keep their
existing paths.

Fixes #55316
Fixes #54756

Co-authored-by: blakehermes9 <blakehermes9@users.noreply.github.com>
2026-09-25 12:27:24 -05:00
Austin Pickett
4661892752 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution
# Conflicts:
#	hermes_cli/models.py
2026-09-25 13:20:26 -04:00
brooklyn!
296ec08dcf fix(models): keep generation models out of chat
Chat catalogs and the session switch treated image and video generation
models as chat. Exclude them by the capability type and name shape the
catalog already publishes, reject selecting one as the session model, and
do not restore a primary already known to be non-chat. Desktop shows the
fallback switch in the transcript.
2026-09-25 12:07:32 -05:00
Austin Pickett
6926997f05 fix(providers): make a configured relay terminal for catalog egress; refuse a foreign Anthropic endpoint
Addresses two P1 review findings on #121614.

1. `provider_model_ids`: a failed or empty relay probe fell through to the
   canonical per-provider fetcher, sending the provider credential to exactly
   the vendor host the user routed away from — recreating #121387 on the
   failure path. A configured `model.base_url` relay is now TERMINAL for live
   catalog egress and degrades to the local curated list instead. The curated
   tail is extracted as `_static_catalog` and shared by both paths.

   Fetchers that already resolve `model.base_url` themselves and degrade
   locally (`_anthropic_catalog`, `_custom_catalog`, `_openai_catalog`, the
   simple api-key fetchers) are excluded from interception via
   `_RELAY_AWARE_CATALOG_FETCHERS` — they already satisfy the invariant and
   produce a better-merged catalog.

2. `_try_anthropic`: an `explicit_base_url` that failed
   `_is_anthropic_compatible_host` was silently dropped, leaving `base_url` at
   the ambient/canonical host and continuing with the explicit credential —
   a silent retarget of authority, not the refusal the PR body claimed. It now
   returns unavailable before client construction.

Regressions: an egress-sentinel test pins that no vendor fetcher, profile
catalog or models.dev merge is reached after a relay 404/hang; the Anthropic
test now asserts no client and zero SDK builder calls.
2026-09-25 12:56:32 -04:00
Austin Pickett
0cc91bb396 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution 2026-09-25 12:50:14 -04:00
kshitijk4poor
82ca84c52d fix(agent): keep openai-codex stale floor/hard cap for inline cron Codex (#69734)
Cron Codex now runs inline via direct_api_call, whose stale budget skipped the
openai-codex large-context floor (600/900/1200s) and HERMES_CODEX_HARD_TIMEOUT_SECONDS
cap applied on the worker path, so healthy >10k-token cron turns were killed at 90s.
Extract the floor+cap into _bound_openai_codex_stale_timeout, used by both paths.
Correct the should_use_direct_api_call docstring (Codex streaming goes through
_stream_codex_passthrough -> _interruptible_api_call, not _StreamingCall) and document
that the worker-only TTFB/idle watchdogs don't run inline. Replace the non-guarding
inline Codex watchdog test with one asserting the >=600s budget.
2026-09-25 21:28:32 +05:30
Ali Ahmed
0213b8b77b fix(agent): keep Codex cron calls inline
Cron Codex Responses calls still went through the spawned interrupt worker
(the #62151 nested-pool deadlock path). Route them through direct_api_call;
the Codex dispatch builds its client via make_client, so the inline stale
watchdog aborts a silent Codex stream even though the worker-only TTFB/idle
watchdogs are bypassed. Delegated children stay chat_completions-only.

Ported onto main's _InlineRequest refactor from PR #70087 (cherry picked
from e0028a5d73). Fixes #69734.
2026-09-25 21:28:32 +05:30
kshitijk4poor
65a3f97622 refactor(agent): share one run-budget stale-timeout cap across stream and non-stream (#97968) 2026-09-25 21:28:32 +05:30
liuhao1024
40d6529dc8 fix(agent): add MiniMax M2.x to reasoning stale-timeout floor
MiniMax M2.x reasoning models emit reasoning_content blocks before
their first content token (#17924). During extended thinking phases,
they routinely exceed the default 180s chat-model stale-stream
timeout, causing the stale-stream detector to kill the connection
mid-think.

This adds minimax-m2 to _REASONING_STALE_TIMEOUT_FLOORS with a
300s floor — generous enough to cover the documented 240s stall
(test_streaming.py:1270-1278) with margin.

Fixes #62353

(cherry picked from commit d66b2a75dee78f583edb75b699ba055ffcb93013)
2026-09-25 21:28:32 +05:30
1052326311
0421036206 fix(agent): cap implicit cloud stream patience to run budget
(cherry picked from commit f9f3c502b535ea536d2a54207cfba1d4fdf2d146)
2026-09-25 21:28:32 +05:30
kshitijk4poor
56490ca109 refactor(aux): drop the now-uncalled _read_codex_access_token and retarget its test seams
The fold routed every aux Codex read through _resolve_codex_credential_and_base, leaving
_read_codex_access_token with no production callers; three test patches on it had gone inert
(including the 'should use pool token' guard). Point them at the live seams instead.
2026-09-25 21:27:06 +05:30
kshitijk4poor
5e2d55ca9c fix(codex): quota probe and /usage pool paths use the pool route base
Pool rows keep the canonical chatgpt.com URL, so the quota-restored probe
(auth_codex + CredentialPool) and the /usage tier-3 and forced-refresh
paths paired a gateway key with chatgpt.com/backend-api/wham/usage. Route
them through _codex_pool_route_base_url, the chat route's rule
(HERMES_CODEX_BASE_URL > model.base_url > row URL).

Refs #121486
2026-09-25 21:27:06 +05:30
kshitijk4poor
e326520d50 refactor(codex): one credential/route authority for aux + image paths
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
  fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
  directly when the pool yields no token (no second uncached pool load,
  no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
  is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
  helper's error fallback reads the profile-scoped override, not the raw
  process env.
- model setup flow: the confirm guards get the resolved Codex base, not
  the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
2026-09-25 21:27:06 +05:30
kshitijk4poor
e87f673faa fix(codex): send catalog/image credentials only to their own route
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.

- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
  credential actually routes to (runtime_provider._pool_entry_mode_and_url:
  env > model.base_url while the row is canonical > row URL) instead of the
  ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
  resolved with the token from hermes_cli/models.py, the CLI default-model
  swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
  one pool selection; the image plugin, _build_codex_client and the raw
  Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
  chatgpt.com; a gateway key may probe its own gateway's /models.

Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.

Addresses @andrexibiza's review on #121508.
2026-09-25 21:27:06 +05:30
liuhao1024
602aa9a55b fix(agent): keep custom-Codex-base credentials off chatgpt.com
Behind a custom Codex base URL (HERMES_CODEX_BASE_URL / model.base_url
gateway) three paths still hit the hard-coded chatgpt.com host with the
gateway's credential-pool key (#121486):

- the OAuth context-length probe (agent/model_metadata.py) and the
  /model picker's live discovery (hermes_cli/codex_models.py) both GET
  https://chatgpt.com/backend-api/codex/models with
  Authorization: Bearer <gateway key> whenever model.context_length is
  not pinned — the key is sent to a service it does not belong to and
  cannot answer for;
- the openai-codex image_gen plugin posts to the same hard-coded base.

Fix, mirroring the quota probe's existing gate in auth_codex:

- both catalog sites now decline to probe non-JWT credentials (real
  Codex access tokens are JWTs; a gateway key is not one) and fall
  back to the static table / offline sources — same outcome as the
  doomed request today, minus the credential leak;
- a JWT reached through a custom base now probes that base's own
  /models instead of chatgpt.com (catalog URLs are built from the
  resolved base; the per-token cache key includes the base);
- the image plugin resolves its base from HERMES_CODEX_BASE_URL the
  same way the text client does.

Fast-mode host gating in the /fast picker is intentionally left
untouched: lifting it needs an explicit opt-in design decision, not a
bug fix.

(cherry picked from commit 5d76ec525674d7b103ab53955ba5279605457ca1)
[salvage: plugins/image_gen/openai-codex/__init__.py hunk dropped in favour of #121497 (first submitter, profile-scoped override + base-aware Cloudflare headers)]
2026-09-25 21:27:06 +05:30
Austin Pickett
5af33f8375 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution 2026-09-25 09:09:44 -04:00
kshitijk4poor
e62a47ab68 fix(curator): floor interval_hours at 1 and warn once per bad value
interval_hours <= 0 made should_run_now() true on every idle tick,
re-running the review pass each time. Route it through the same
floor-with-default helper as the day counts (renamed _bounded_count),
and log the fallback warning once per (key, value) since the dashboard
status endpoint polls these getters.
2026-09-25 14:08:03 +05:30
AhmetArif0
342c29250d fix(curator): bound automatic transition days like curator prune
get_stale_after_days()/get_archive_after_days() accepted any int from
curator.stale_after_days/archive_after_days. archive_after_days: 0 sets
archive_cutoff to now, so apply_automatic_transitions() (runs unconfirmed
on an idle tick, curator on by default) archives every skill with any past
activity on the next pass; a negative value builds a future cutoff. The
manual path already refuses the same value (_cmd_prune: "--days must be
>= 1"), and "0 disables" is the repo convention elsewhere.

Fix: a value < 1 falls back to the default with one warning naming the
key, the same bound _cmd_prune enforces. Same class of fix as b01b1c8b
(bound kanban gc retention so -N/0 cannot mass-delete).

(cherry picked from commit 6685565a8c5147c8d7c23602f371f60571c44074)
2026-09-25 14:08:03 +05:30
Hermes Agent
658e6c885d revert(agent): drop the no-op side-channel length-stop change
The reasoning half only touched a docstring, and its tests pinned main's
existing continuation (reasoning-off retry, then the 'No visible answer'
ceiling). Nothing changed, so the PR stays on the Windows close/stop fix.
2026-09-25 01:08:23 -05:00
brooklyn!
2c42fcb507 fix(agent): continue a length stop that only has side-channel reasoning
That stop is unfinished thought, not a finished thinking-budget abort.
Continuation still owns it, and the side channel is not the answer.
2026-09-25 01:08:23 -05:00
brooklyn!
8b17394cc6 fix: abort reasoning-field length stops and surface close taskkill failures
finish_reason=length with empty visible content and a non-empty reasoning
or reasoning_content field uses the existing thinking-budget abort. No
model id is consulted. Empty content with no side channel still continues.

Desktop close/stop no longer discards Windows taskkill failures. After the
same tree-kill, owned PIDs are inventoried and only unheld gateway locks
are cleared.
2026-09-25 01:08:23 -05:00
ethernet
39c0a39100 fix(providers): report why a lazy SDK install did not land
_get_anthropic_sdk() swallowed every ensure_import("anthropic") failure and
_require_sdk() then told the user to "Install it with: hermes pm install
--extra anthropic". PM often HAS installed it: sync_venv succeeds into a new
dependency environment that only activates at process boot, and
ensure_import raises "installed; restart Hermes". A lazy-install guard
("this process is not running from the install's dependency environment")
was flattened the same way. Users were told to install something that was
installed, or given a command that doesn't address the actual refusal.

Keep the import as the decider, but remember the InstallError and put its
text in the ImportError. bedrock_adapter._require_boto3 had the identical
shape; azure_identity_adapter already propagates str(exc) and is the model.
2026-09-24 23:37:27 -04:00
ethernet
3c18ea542d Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 15:32:42 -04:00
brooklyn!
f213f755ee fix(ci): join continuation fragments as pairs
The retry path stored fragment text as strings, and the joiner unpacks
(text, stub) pairs. Also sort the composer-clamp props so desktop lint
stops failing every merge.
2026-09-24 15:31:44 -04:00
ethernet
05a78671a7 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 13:13:35 -04:00
kshitijk4poor
c43eb76d86 refactor(streaming): share interim visible/streamed normalization between prefix and exact verdicts
Both interim predicates now compare one (visible, streamed) pair from
_interim_visible_and_streamed so their normalization cannot drift; the two
#88954 codex tests collapse into one parametrized test.
2026-09-24 22:41:10 +05:30
kshitijk4poor
1f61f7087b docs(agent): pin prefix semantics at the previewed-mark call site (#88954)
Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com>
2026-09-24 22:41:10 +05:30
liuhao1024
167f0ef5fa fix(streaming): stop losing the tail of commentary truncated at a tool-call boundary
With streaming enabled on Telegram, the streamed commentary is routinely
cut mid-text when the tool_calls finish arrives; the partial stream is a
prefix of the full commentary. The interim-message path used the
prefix-based _interim_content_was_streamed verdict, so the gateway's
_interim_assistant_cb called on_segment_break() — finalizing the
truncated bubble — and the tail ("...pick it u" vs "...pick it up") was
permanently lost (#88954).

That prefix semantics stays correct for the conversation-loop
"previewed" marks (the streamed prefix IS on the user's screen there,
and the contract test pins it per the #65919 review). The gateway
decision needs the stricter test: add _interim_content_fully_streamed
(exact normalized equality) and use it for the interim-message verdict.
Only an exact match may skip the full-text resend; a partial prefix
falls through to on_commentary() and re-delivers the complete text —
a benign duplicate, never lost text.

(cherry picked from commit 0b06a6660a1a4d47b4974f21ae42d7aeb1cbea15)
2026-09-24 22:41:10 +05:30
kshitijk4poor
2ab2270b32 fix(agent): hide MiniMax-M3 Chinese reasoning tags (#43827)
Add 思考/反思/推理/推敲 to THINK_TAG_NAMES so the streaming scrubber, CLI and
gateway stream filters and the final-response stripper all hide them, and
derive the auxiliary-client reasoning strip from the same list instead of a
hard-coded copy. Bare bracketless markers (unverified) are not covered.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-24 22:41:10 +05:30
ethernet
0f65d698d0 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 13:10:14 -04:00
kshitijk4poor
7fb8736a1e refactor(codex): single source for app-server delta methods and delta text
The event bridge listed the text-delta methods twice (display handlers and
the #118410 liveness set) and duplicated _fire_delta's text extraction.
Derive both from _CODEX_TEXT_DELTA_METHODS, share _delta_text(), and hoist
the progress item-type set to a module constant (no per-event frozenset).
2026-09-24 22:38:35 +05:30
kshitijk4poor
6c94f52ef1 fix(agent): keep local Responses endpoints out of hosted codex watchdog clamps
#86444 widened the large-context stale floor, the 1500s hard ceiling and the
TTFB scale-up/cap gate from OpenAI-Codex to every codex_responses route. That
made the #92302 local-endpoint TTFB branch unreachable (local TTFB fell back to
120s instead of agent.local_stream_stale_timeout) and silently clamped a local
server's configured stale timeout to 1500s. Evaluate is_local_endpoint once and
exclude local endpoints from the hosted clamps; xAI keeps the #86444 behaviour.

Note: xAI large requests now get the raised stale floor but keep first-event
idle semantics (progress gating stays OpenAI-Codex only).
2026-09-24 22:38:35 +05:30
Kevin Yin
fe14267769 fix(cron): bound Codex max-iteration summaries
(cherry picked from commit f8d6ae81fd73ae56d95e84672a7cc9d0365d5c3c)
2026-09-24 22:38:35 +05:30
Christopher
606eb2382c fix(agent): apply large-context watchdog accommodations to all codex_responses transports
(cherry picked from commit 51a60661a4c777e20df7edfab5af9f3c9956aac5)
2026-09-24 22:38:35 +05:30
zxd
94727893b2 fix(codex): refresh turn liveness from app-server progress
(cherry picked from commit 677fde44763a0db4cba457de831bfd049f6c2b17)
2026-09-24 22:38:35 +05:30
kshitijk4poor
572446308e fix(agent): feed inline <think> text to the live reasoning pane (#89647)
The desktop/TUI reasoning pane is driven by reasoning.delta via
reasoning_callback; the scrubber-side collector only filled the final
reasoning_content, which extract_reasoning already recovers from the raw
content, so the pane stayed dead.

- Drop the scrubber reasoning collector (_reasoning_parts, reasoning(),
  clear_reasoning(), \x00 sentinel, _THINK_TAG_RE) and its per-request
  reset hook.
- StreamingThinkScrubber.feed() exposes the text it stripped from inside
  think blocks as last_hidden; _fire_stream_delta forwards it through
  _fire_reasoning_delta(inline=True) while no native reasoning delta has
  arrived for this model response (reset per request) — no double
  reasoning. CLI gating is unchanged: its reasoning_callback is None
  unless show_reasoning/verbose.
- _finish_chat_stream fills reasoning_content from the raw content via
  the existing extract_reasoning when no reasoning delta arrived.
- Replace the collector tests with two guards (live forwarding + native
  suppression; _finish_chat_stream fallback), both red on origin/main.

Co-authored-by: SayHell0W0rld <852938468@qq.com>
2026-09-24 22:37:10 +05:30
kshitijk4poor
363a998f7d fix(agent): scope recovered inline reasoning per model response
Rejoin streamed pieces of one <think> block verbatim (the per-delta newline
join split sentences on every token) and clear collected reasoning in the
per-request _reset_stream_delivery_tracking so a tool-loop's earlier API call
does not bleed into the next call's reasoning_content (the scrubber is only
reset per turn). Follow-up to #90417 (#89647).
2026-09-24 22:37:10 +05:30