The -900k alias fix hand-rolled a second copy of the Astra slug set and its
vendor-prefix normalization. agent/reasoning_effort.py::is_astra_model is the
documented single home for that set (picker, effort vocabulary and request
sanitizer already key off it), so the gate now calls it and a future Astra
alias stays a one-line edit. The gpt-5.6 marker check is back to main's exact
form.
Tests move into the existing parametrized Astra gate table, which checks both
the capability resolver and the per-request gate: -900k on official Codex OAuth
is eligible; -900k through a relay or on provider openai is not. Docs and the
config example no longer say "exact gpt-6-astra".
DEFAULT_CONTEXT_LENGTHS declares claude-opus-5 as 1M, but the Bedrock
static table never got the entry, so the offline path resolved the 128K
default and the agent compressed context ~8x early. Add the key and pin
the table pairing with a test so the next 1M generation cannot drift.
Co-authored-by: JiaDe-Wu <JiaDe-Wu@users.noreply.github.com>
(cherry picked from commit c1be16ecbdd267b14124dd35c0b92ccfcd201818)
The streaming stale monitor re-fires every stale window while its worker
has not started/dispatched the attempt yet or is still unwinding a kill,
and each re-fire bumped agent._consecutive_stale_streams. One provider
attempt could therefore count twice (or more), so a turn with four hung
attempts reached HERMES_STREAM_STALE_GIVEUP=5 and the breaker refused the
NEXT turn although the provider only ever saw four stale requests. The
non-streaming and inline watchdogs already count once per call.
Count each started stream attempt once (attempt 0, i.e. nothing started,
never counts), and let the interrupted-wait bump honour the same ledger.
This is the root cause of the intermittent
test_agent_turn_liveness[provider_hang] red on CI (fault calls 4, five
"Stream stale for 3s" kills, probe requests 0, breaker text "5
consecutive stale attempts"): on a busy runner the first attempt's
worker takes >3 s to reach the wire, the monitor kills it before
dispatch (tcp_force_closed=0) and again 3 s after dispatch.
Repro: a 3.6 s sleep before the worker opens its first stream
reproduces the CI signature 4/4 on origin/main and 0/6 with the fix;
under taskset CPU starvation the unmodified module goes red 2/3 on base.
- _mark_finish_seen helper on the local _diag; also set on the Anthropic
path when message_delta carries stop_reason
- emit_stream_drop status: 'attempt N/M dropped, reconnecting'
- clean-EOF log wording: server or proxy closed the stream cleanly
- truncated_unreported copy drops the raw finish_reason placeholder
- turn_tool_validation compares against FINISH_REASON_LENGTH
When the 4 truncated-tool-call retries are exhausted on a clean-EOF stub
(no transport error, no finish_reason), use a dedicated stream_closed_tool_call
copy and failure_reason=truncated instead of stream_dropped_tool_call
('check your network') stamped as timeout. Also compute the truncation
banner in a plain if/elif chain instead of a 4-arm conditional expression.
Hand-applied from PR #91738 (1da9001929): the truncation detector moved from
conversation_loop.py to the _truncated branch of agent/turn_tool_validation.py.
Only claim the output cap when finish_reason == 'length'; otherwise use new
hedged site copy 'truncated_unreported' (agent/turn_failure_copy.py).
MessageStream._retry_after_drop passed attempt + 2 to _emit_stream_drop, so
the first drop was announced as attempt 2/N. Hand-applied from PR #90254
onto main's MessageStream split (original hunks targeted the pre-split
interruptible_streaming_api_call).
Response truncated - stream ended before completion" and the log line
"...treating as a mid-stream drop" fired identically for two distinct
failure classes: a genuine transport exception mid-stream, and the
provider closing the connection cleanly (no exception) without ever
sending a terminal finish_reason chunk. The second case is not a
network drop, but the shared wording sent users chasing a network
problem that never existed (#102766).
_build_partial_stream_stub() now tags its stub _clean_eof=True (every
caller is a clean-EOF site: the streaming loop ended without an
exception). The conversation loop uses that tag to pick distinct
wording; the exception-driven stub is unchanged. The two clean-EOF
log lines in chat_completion_helpers.py no longer say "drop".
stream_diag's per-attempt dict also gains finish_reason_seen, set the
moment a terminal chunk is observed, so log_stream_retry's retry/
exhaustion WARNING can say whether the dying attempt had already seen
a finish_reason before the exception hit.
(cherry picked from commit 2c1855ed6dd0037eebe0b1068628999832617109)
Desktop chat goes through tui_gateway, which never applied the gateway's
first-message profile-build sidecar note. Add shared first_contact_turn_note
helper and wire it into the prompt turn so fresh installs get the same
opt-in profile setup offer as messaging surfaces.
Adapted to current main (salvage of open PR #82765, fixes#82750):
_run_prompt_submit now lives in tui_gateway/prompt_turn.py; the note is
staged on agent._gateway_turn_context_notes (the existing sidecar channel,
consumed by agent.turn_context on the user message) so the system prompt
stays byte-stable; the prior-session probe rides _session_db(session).
The low-value onboarding test classes purged from main were not re-added —
only new coverage for first_contact_turn_note and the staging path.
Co-authored-by: Noa <rainbowgore@users.noreply.github.com>
(cherry picked from commit b843900ca818912f0480f6926a059d292b63ffe6)
A session asked to clean up older Pythons removed the uv-managed base
interpreter its own venv depended on; the next boot died with 'uv
trampoline failed to spawn Python child process' and no agent tool could
repair it, because the agent itself no longer started (#58748). Prior
uninstall detection (85ce25687e) only flagged package-manager commands.
Add agent/runtime_self_protection.py and wire it into both layers:
- The approval floor (_floor_block) now blocks shell commands that
delete the running interpreter, its own venv, the pyvenv.cfg base, or
the uv-managed install directory — rm/rmdir/rd/del/erase/Remove-Item
with any flags, find <root> -delete, and uv python uninstall of the
running version (including --all). The floor runs before yolo /
approvals.mode=off / cron approve mode, so no session setting can
bypass it.
- The file-safety write classifier denies write/patch/move/delete to the
same paths, so the file tools cannot overwrite the interpreter either.
Only the runtime the process itself boots from is protected; every other
venv and interpreter on the machine stays manageable.
Fixes#58748
Background processes spawned with terminal(background=true) are killed from
three agent-lifecycle sweeps: agent release()'s kill_all, a gateway turn
timeout's kill_started_since, and agent close's owned-process loop. Jobs the
user explicitly wants to outlive the session (overnight batches, watchful
daemons) had no way to opt out.
Add terminal(background=true, persist_on_release=true):
- ProcessSession.persist_on_release, stamped by spawn_local/spawn_via_env,
carried in crash-recovery checkpoints and exposed via list_sessions()
- kill_all skips persisted sessions only for lifecycle sources
(_LIFECYCLE_KILL_SOURCES: kill_all, gateway_turn_timeout, agent_close);
explicit operator stops (process_manage kill, /stop slash + RPC mirror,
CLI /stop) now pass distinct sources so they still reach persisted jobs
- the agent_close owned-process loop in _close_task_resources skips
persisted sessions the same way
- gateway shutdown keeps killing persisted jobs (source=gateway_shutdown):
the host is going away and survivors would become PPID=1 orphans
Co-authored-by: salvaged from #109846 (persist_on_release plumbing) and
extended to the turn-timeout and agent_close paths.
A turn that falls out of the loop after a tool result with no follow-up
assistant text left the durable transcript ending at a raw tool row and
returned a silent result: Desktop/TUI showed a ready composer (or kept
spinning) with no final message. The pending_tool_result explainer copy
already existed but nothing minted the exit reason.
finalize_turn now detects the non-interrupted tool tail, fails the turn
with turn_exit_reason=pending_tool_result, and synthesizes the visible
assistant close before persistence so the durable tail is alternation-
safe. Stream-recovered turns (#95514) and interrupted tails keep their
existing paths.
Fixes#55316Fixes#54756
Co-authored-by: blakehermes9 <blakehermes9@users.noreply.github.com>
Chat catalogs and the session switch treated image and video generation
models as chat. Exclude them by the capability type and name shape the
catalog already publishes, reject selecting one as the session model, and
do not restore a primary already known to be non-chat. Desktop shows the
fallback switch in the transcript.
Addresses two P1 review findings on #121614.
1. `provider_model_ids`: a failed or empty relay probe fell through to the
canonical per-provider fetcher, sending the provider credential to exactly
the vendor host the user routed away from — recreating #121387 on the
failure path. A configured `model.base_url` relay is now TERMINAL for live
catalog egress and degrades to the local curated list instead. The curated
tail is extracted as `_static_catalog` and shared by both paths.
Fetchers that already resolve `model.base_url` themselves and degrade
locally (`_anthropic_catalog`, `_custom_catalog`, `_openai_catalog`, the
simple api-key fetchers) are excluded from interception via
`_RELAY_AWARE_CATALOG_FETCHERS` — they already satisfy the invariant and
produce a better-merged catalog.
2. `_try_anthropic`: an `explicit_base_url` that failed
`_is_anthropic_compatible_host` was silently dropped, leaving `base_url` at
the ambient/canonical host and continuing with the explicit credential —
a silent retarget of authority, not the refusal the PR body claimed. It now
returns unavailable before client construction.
Regressions: an egress-sentinel test pins that no vendor fetcher, profile
catalog or models.dev merge is reached after a relay 404/hang; the Anthropic
test now asserts no client and zero SDK builder calls.
Cron Codex now runs inline via direct_api_call, whose stale budget skipped the
openai-codex large-context floor (600/900/1200s) and HERMES_CODEX_HARD_TIMEOUT_SECONDS
cap applied on the worker path, so healthy >10k-token cron turns were killed at 90s.
Extract the floor+cap into _bound_openai_codex_stale_timeout, used by both paths.
Correct the should_use_direct_api_call docstring (Codex streaming goes through
_stream_codex_passthrough -> _interruptible_api_call, not _StreamingCall) and document
that the worker-only TTFB/idle watchdogs don't run inline. Replace the non-guarding
inline Codex watchdog test with one asserting the >=600s budget.
Cron Codex Responses calls still went through the spawned interrupt worker
(the #62151 nested-pool deadlock path). Route them through direct_api_call;
the Codex dispatch builds its client via make_client, so the inline stale
watchdog aborts a silent Codex stream even though the worker-only TTFB/idle
watchdogs are bypassed. Delegated children stay chat_completions-only.
Ported onto main's _InlineRequest refactor from PR #70087 (cherry picked
from e0028a5d73). Fixes#69734.
MiniMax M2.x reasoning models emit reasoning_content blocks before
their first content token (#17924). During extended thinking phases,
they routinely exceed the default 180s chat-model stale-stream
timeout, causing the stale-stream detector to kill the connection
mid-think.
This adds minimax-m2 to _REASONING_STALE_TIMEOUT_FLOORS with a
300s floor — generous enough to cover the documented 240s stall
(test_streaming.py:1270-1278) with margin.
Fixes#62353
(cherry picked from commit d66b2a75dee78f583edb75b699ba055ffcb93013)
The fold routed every aux Codex read through _resolve_codex_credential_and_base, leaving
_read_codex_access_token with no production callers; three test patches on it had gone inert
(including the 'should use pool token' guard). Point them at the live seams instead.
Pool rows keep the canonical chatgpt.com URL, so the quota-restored probe
(auth_codex + CredentialPool) and the /usage tier-3 and forced-refresh
paths paired a gateway key with chatgpt.com/backend-api/wham/usage. Route
them through _codex_pool_route_base_url, the chat route's rule
(HERMES_CODEX_BASE_URL > model.base_url > row URL).
Refs #121486
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
directly when the pool yields no token (no second uncached pool load,
no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
helper's error fallback reads the profile-scoped override, not the raw
process env.
- model setup flow: the confirm guards get the resolved Codex base, not
the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
Behind a custom Codex base URL (HERMES_CODEX_BASE_URL / model.base_url
gateway) three paths still hit the hard-coded chatgpt.com host with the
gateway's credential-pool key (#121486):
- the OAuth context-length probe (agent/model_metadata.py) and the
/model picker's live discovery (hermes_cli/codex_models.py) both GET
https://chatgpt.com/backend-api/codex/models with
Authorization: Bearer <gateway key> whenever model.context_length is
not pinned — the key is sent to a service it does not belong to and
cannot answer for;
- the openai-codex image_gen plugin posts to the same hard-coded base.
Fix, mirroring the quota probe's existing gate in auth_codex:
- both catalog sites now decline to probe non-JWT credentials (real
Codex access tokens are JWTs; a gateway key is not one) and fall
back to the static table / offline sources — same outcome as the
doomed request today, minus the credential leak;
- a JWT reached through a custom base now probes that base's own
/models instead of chatgpt.com (catalog URLs are built from the
resolved base; the per-token cache key includes the base);
- the image plugin resolves its base from HERMES_CODEX_BASE_URL the
same way the text client does.
Fast-mode host gating in the /fast picker is intentionally left
untouched: lifting it needs an explicit opt-in design decision, not a
bug fix.
(cherry picked from commit 5d76ec525674d7b103ab53955ba5279605457ca1)
[salvage: plugins/image_gen/openai-codex/__init__.py hunk dropped in favour of #121497 (first submitter, profile-scoped override + base-aware Cloudflare headers)]
interval_hours <= 0 made should_run_now() true on every idle tick,
re-running the review pass each time. Route it through the same
floor-with-default helper as the day counts (renamed _bounded_count),
and log the fallback warning once per (key, value) since the dashboard
status endpoint polls these getters.
get_stale_after_days()/get_archive_after_days() accepted any int from
curator.stale_after_days/archive_after_days. archive_after_days: 0 sets
archive_cutoff to now, so apply_automatic_transitions() (runs unconfirmed
on an idle tick, curator on by default) archives every skill with any past
activity on the next pass; a negative value builds a future cutoff. The
manual path already refuses the same value (_cmd_prune: "--days must be
>= 1"), and "0 disables" is the repo convention elsewhere.
Fix: a value < 1 falls back to the default with one warning naming the
key, the same bound _cmd_prune enforces. Same class of fix as b01b1c8b
(bound kanban gc retention so -N/0 cannot mass-delete).
(cherry picked from commit 6685565a8c5147c8d7c23602f371f60571c44074)
The reasoning half only touched a docstring, and its tests pinned main's
existing continuation (reasoning-off retry, then the 'No visible answer'
ceiling). Nothing changed, so the PR stays on the Windows close/stop fix.
finish_reason=length with empty visible content and a non-empty reasoning
or reasoning_content field uses the existing thinking-budget abort. No
model id is consulted. Empty content with no side channel still continues.
Desktop close/stop no longer discards Windows taskkill failures. After the
same tree-kill, owned PIDs are inventoried and only unheld gateway locks
are cleared.
_get_anthropic_sdk() swallowed every ensure_import("anthropic") failure and
_require_sdk() then told the user to "Install it with: hermes pm install
--extra anthropic". PM often HAS installed it: sync_venv succeeds into a new
dependency environment that only activates at process boot, and
ensure_import raises "installed; restart Hermes". A lazy-install guard
("this process is not running from the install's dependency environment")
was flattened the same way. Users were told to install something that was
installed, or given a command that doesn't address the actual refusal.
Keep the import as the decider, but remember the InstallError and put its
text in the ImportError. bedrock_adapter._require_boto3 had the identical
shape; azure_identity_adapter already propagates str(exc) and is the model.
The retry path stored fragment text as strings, and the joiner unpacks
(text, stub) pairs. Also sort the composer-clamp props so desktop lint
stops failing every merge.
Both interim predicates now compare one (visible, streamed) pair from
_interim_visible_and_streamed so their normalization cannot drift; the two
#88954 codex tests collapse into one parametrized test.
With streaming enabled on Telegram, the streamed commentary is routinely
cut mid-text when the tool_calls finish arrives; the partial stream is a
prefix of the full commentary. The interim-message path used the
prefix-based _interim_content_was_streamed verdict, so the gateway's
_interim_assistant_cb called on_segment_break() — finalizing the
truncated bubble — and the tail ("...pick it u" vs "...pick it up") was
permanently lost (#88954).
That prefix semantics stays correct for the conversation-loop
"previewed" marks (the streamed prefix IS on the user's screen there,
and the contract test pins it per the #65919 review). The gateway
decision needs the stricter test: add _interim_content_fully_streamed
(exact normalized equality) and use it for the interim-message verdict.
Only an exact match may skip the full-text resend; a partial prefix
falls through to on_commentary() and re-delivers the complete text —
a benign duplicate, never lost text.
(cherry picked from commit 0b06a6660a1a4d47b4974f21ae42d7aeb1cbea15)
Add 思考/反思/推理/推敲 to THINK_TAG_NAMES so the streaming scrubber, CLI and
gateway stream filters and the final-response stripper all hide them, and
derive the auxiliary-client reasoning strip from the same list instead of a
hard-coded copy. Bare bracketless markers (unverified) are not covered.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The event bridge listed the text-delta methods twice (display handlers and
the #118410 liveness set) and duplicated _fire_delta's text extraction.
Derive both from _CODEX_TEXT_DELTA_METHODS, share _delta_text(), and hoist
the progress item-type set to a module constant (no per-event frozenset).
#86444 widened the large-context stale floor, the 1500s hard ceiling and the
TTFB scale-up/cap gate from OpenAI-Codex to every codex_responses route. That
made the #92302 local-endpoint TTFB branch unreachable (local TTFB fell back to
120s instead of agent.local_stream_stale_timeout) and silently clamped a local
server's configured stale timeout to 1500s. Evaluate is_local_endpoint once and
exclude local endpoints from the hosted clamps; xAI keeps the #86444 behaviour.
Note: xAI large requests now get the raised stale floor but keep first-event
idle semantics (progress gating stays OpenAI-Codex only).
The desktop/TUI reasoning pane is driven by reasoning.delta via
reasoning_callback; the scrubber-side collector only filled the final
reasoning_content, which extract_reasoning already recovers from the raw
content, so the pane stayed dead.
- Drop the scrubber reasoning collector (_reasoning_parts, reasoning(),
clear_reasoning(), \x00 sentinel, _THINK_TAG_RE) and its per-request
reset hook.
- StreamingThinkScrubber.feed() exposes the text it stripped from inside
think blocks as last_hidden; _fire_stream_delta forwards it through
_fire_reasoning_delta(inline=True) while no native reasoning delta has
arrived for this model response (reset per request) — no double
reasoning. CLI gating is unchanged: its reasoning_callback is None
unless show_reasoning/verbose.
- _finish_chat_stream fills reasoning_content from the raw content via
the existing extract_reasoning when no reasoning delta arrived.
- Replace the collector tests with two guards (live forwarding + native
suppression; _finish_chat_stream fallback), both red on origin/main.
Co-authored-by: SayHell0W0rld <852938468@qq.com>
Rejoin streamed pieces of one <think> block verbatim (the per-delta newline
join split sentences on every token) and clear collected reasoning in the
per-request _reset_stream_delivery_tracking so a tool-loop's earlier API call
does not bleed into the next call's reasoning_content (the scrubber is only
reset per turn). Follow-up to #90417 (#89647).