Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
Behind a custom Codex base URL (HERMES_CODEX_BASE_URL / model.base_url
gateway) three paths still hit the hard-coded chatgpt.com host with the
gateway's credential-pool key (#121486):
- the OAuth context-length probe (agent/model_metadata.py) and the
/model picker's live discovery (hermes_cli/codex_models.py) both GET
https://chatgpt.com/backend-api/codex/models with
Authorization: Bearer <gateway key> whenever model.context_length is
not pinned — the key is sent to a service it does not belong to and
cannot answer for;
- the openai-codex image_gen plugin posts to the same hard-coded base.
Fix, mirroring the quota probe's existing gate in auth_codex:
- both catalog sites now decline to probe non-JWT credentials (real
Codex access tokens are JWTs; a gateway key is not one) and fall
back to the static table / offline sources — same outcome as the
doomed request today, minus the credential leak;
- a JWT reached through a custom base now probes that base's own
/models instead of chatgpt.com (catalog URLs are built from the
resolved base; the per-token cache key includes the base);
- the image plugin resolves its base from HERMES_CODEX_BASE_URL the
same way the text client does.
Fast-mode host gating in the /fast picker is intentionally left
untouched: lifting it needs an explicit opt-in design decision, not a
bug fix.
(cherry picked from commit 5d76ec525674d7b103ab53955ba5279605457ca1)
[salvage: plugins/image_gen/openai-codex/__init__.py hunk dropped in favour of #121497 (first submitter, profile-scoped override + base-aware Cloudflare headers)]
interval_hours <= 0 made should_run_now() true on every idle tick,
re-running the review pass each time. Route it through the same
floor-with-default helper as the day counts (renamed _bounded_count),
and log the fallback warning once per (key, value) since the dashboard
status endpoint polls these getters.
get_stale_after_days()/get_archive_after_days() accepted any int from
curator.stale_after_days/archive_after_days. archive_after_days: 0 sets
archive_cutoff to now, so apply_automatic_transitions() (runs unconfirmed
on an idle tick, curator on by default) archives every skill with any past
activity on the next pass; a negative value builds a future cutoff. The
manual path already refuses the same value (_cmd_prune: "--days must be
>= 1"), and "0 disables" is the repo convention elsewhere.
Fix: a value < 1 falls back to the default with one warning naming the
key, the same bound _cmd_prune enforces. Same class of fix as b01b1c8b
(bound kanban gc retention so -N/0 cannot mass-delete).
(cherry picked from commit 6685565a8c5147c8d7c23602f371f60571c44074)
The reasoning half only touched a docstring, and its tests pinned main's
existing continuation (reasoning-off retry, then the 'No visible answer'
ceiling). Nothing changed, so the PR stays on the Windows close/stop fix.
finish_reason=length with empty visible content and a non-empty reasoning
or reasoning_content field uses the existing thinking-budget abort. No
model id is consulted. Empty content with no side channel still continues.
Desktop close/stop no longer discards Windows taskkill failures. After the
same tree-kill, owned PIDs are inventoried and only unheld gateway locks
are cleared.
_get_anthropic_sdk() swallowed every ensure_import("anthropic") failure and
_require_sdk() then told the user to "Install it with: hermes pm install
--extra anthropic". PM often HAS installed it: sync_venv succeeds into a new
dependency environment that only activates at process boot, and
ensure_import raises "installed; restart Hermes". A lazy-install guard
("this process is not running from the install's dependency environment")
was flattened the same way. Users were told to install something that was
installed, or given a command that doesn't address the actual refusal.
Keep the import as the decider, but remember the InstallError and put its
text in the ImportError. bedrock_adapter._require_boto3 had the identical
shape; azure_identity_adapter already propagates str(exc) and is the model.
The retry path stored fragment text as strings, and the joiner unpacks
(text, stub) pairs. Also sort the composer-clamp props so desktop lint
stops failing every merge.
Both interim predicates now compare one (visible, streamed) pair from
_interim_visible_and_streamed so their normalization cannot drift; the two
#88954 codex tests collapse into one parametrized test.
With streaming enabled on Telegram, the streamed commentary is routinely
cut mid-text when the tool_calls finish arrives; the partial stream is a
prefix of the full commentary. The interim-message path used the
prefix-based _interim_content_was_streamed verdict, so the gateway's
_interim_assistant_cb called on_segment_break() — finalizing the
truncated bubble — and the tail ("...pick it u" vs "...pick it up") was
permanently lost (#88954).
That prefix semantics stays correct for the conversation-loop
"previewed" marks (the streamed prefix IS on the user's screen there,
and the contract test pins it per the #65919 review). The gateway
decision needs the stricter test: add _interim_content_fully_streamed
(exact normalized equality) and use it for the interim-message verdict.
Only an exact match may skip the full-text resend; a partial prefix
falls through to on_commentary() and re-delivers the complete text —
a benign duplicate, never lost text.
(cherry picked from commit 0b06a6660a1a4d47b4974f21ae42d7aeb1cbea15)
Add 思考/反思/推理/推敲 to THINK_TAG_NAMES so the streaming scrubber, CLI and
gateway stream filters and the final-response stripper all hide them, and
derive the auxiliary-client reasoning strip from the same list instead of a
hard-coded copy. Bare bracketless markers (unverified) are not covered.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
The event bridge listed the text-delta methods twice (display handlers and
the #118410 liveness set) and duplicated _fire_delta's text extraction.
Derive both from _CODEX_TEXT_DELTA_METHODS, share _delta_text(), and hoist
the progress item-type set to a module constant (no per-event frozenset).
#86444 widened the large-context stale floor, the 1500s hard ceiling and the
TTFB scale-up/cap gate from OpenAI-Codex to every codex_responses route. That
made the #92302 local-endpoint TTFB branch unreachable (local TTFB fell back to
120s instead of agent.local_stream_stale_timeout) and silently clamped a local
server's configured stale timeout to 1500s. Evaluate is_local_endpoint once and
exclude local endpoints from the hosted clamps; xAI keeps the #86444 behaviour.
Note: xAI large requests now get the raised stale floor but keep first-event
idle semantics (progress gating stays OpenAI-Codex only).
The desktop/TUI reasoning pane is driven by reasoning.delta via
reasoning_callback; the scrubber-side collector only filled the final
reasoning_content, which extract_reasoning already recovers from the raw
content, so the pane stayed dead.
- Drop the scrubber reasoning collector (_reasoning_parts, reasoning(),
clear_reasoning(), \x00 sentinel, _THINK_TAG_RE) and its per-request
reset hook.
- StreamingThinkScrubber.feed() exposes the text it stripped from inside
think blocks as last_hidden; _fire_stream_delta forwards it through
_fire_reasoning_delta(inline=True) while no native reasoning delta has
arrived for this model response (reset per request) — no double
reasoning. CLI gating is unchanged: its reasoning_callback is None
unless show_reasoning/verbose.
- _finish_chat_stream fills reasoning_content from the raw content via
the existing extract_reasoning when no reasoning delta arrived.
- Replace the collector tests with two guards (live forwarding + native
suppression; _finish_chat_stream fallback), both red on origin/main.
Co-authored-by: SayHell0W0rld <852938468@qq.com>
Rejoin streamed pieces of one <think> block verbatim (the per-delta newline
join split sentences on every token) and clear collected reasoning in the
per-request _reset_stream_delivery_tracking so a tool-loop's earlier API call
does not bleed into the next call's reasoning_content (the scrubber is only
reset per turn). Follow-up to #90417 (#89647).
Providers that inline reasoning (MiniMax-M3 streams <think>…</think> in
content) never send a reasoning delta, so the desktop reasoning pane stays
dead even though the reasoning happened (#89647).
StreamingThinkScrubber (#17924) already strips inline blocks from streamed
content but discarded the text. It now collects the stripped text
(reasoning() accessor, tag markup removed), and the stream finisher
populates reasoning_content from the scrubber when the provider returned
no reasoning delta.
Hand-grafted from 759b65a8a4 (PR #90417) onto main's refactored scrubber and
_StreamingCall._finish_chat_stream; credential_pool.py commit dropped (out of scope).
openai/_streaming.py raises APIError(body=data["error"]), so exc.body is the
inner error object or a bare string, never {"error": ...}; the predicate was
False for every real stream error. Accept the inner object (any of
type/code/message), a non-empty string, or the wrapper. Share one
_exc_http_status helper with _is_transient_transport_error.
Tests now build the error through openai.OpenAI + httpx.MockTransport (HTTP 200
SSE error event) and exercise the parameter-strip rung's fall-through.
A relay wrapping an error as {"code": 200, "success": false} must not route to
_by_status as a 200. Also reuse _error_obj for the error.* candidate walk and
point at the string-parsing text-SSE sibling.
Drop the bespoke _streaming_render_error stage and its hand-built verdict: the
status-less message path already has a head table and _V_FORMAT_ERROR is the
canonical format_error/no-retry/fallback verdict. _by_status returns None when
there is no status, so the stage position bought nothing.
OpenAI-compatible relays can commit SSE with HTTP 200 and then emit an
OpenAI-style error event; the SDK raises a status-less APIError with the
event only on exc.body. Classify that shape as a route failure in the
recovery ladder's provider-fallback rung (_FALLBACK_REASONS, reason
"structured provider error") and let it fall through the parameter-strip
rung. Grafted from #101540 onto main's recovery-ladder layer.
LM Studio / llama.cpp raise a bare status-less APIError when chat-template
Jinja rendering fails mid-stream; classify as format_error with fallback
instead of burning same-provider retries. Stage fires only when status_code
is None. Hand-applied from #64834 (supersedes #62673) onto the staged
classifier pipeline.
Fixes#62662
An aggregator can deliver an upstream failure as {"error": {"code": 403}}
inside an HTTP-200 SSE stream; the SDK then raises a status-less APIError
whose body carries the status. _extract_status_code only read
exc.status_code/.status, so the ban climbed the transient retry ladder
(api_max_retries resends, then 'temporarily unavailable, wait a minute')
which can never succeed for e.g. a banned account.
Fall back to a body-carried numeric code (100-599, int only — symbolic
string codes stay with _code_from_payload), mirroring the text-SSE path's
_status_code_from_payload. A 403 now classifies as auth: exactly one
request, fallback-eligible, no wait-and-retry advice.
Fixes#121270
(cherry picked from commit 8a09fc53614ea3b954faebb4bfcccc176116338c)
Ports #119027::finalize_continuation_partial: when a stream drop entered the
continuation path and the next request exhausted retries before a token, the
_length_continuation_fragment/_nudge rows were persisted as-is, so resume
replayed a dangling synthetic user nudge. Collapse them into one assistant row
before persistence and feed the collapsed text to the #119081 partial-retention
path (the fragment rows are gone by then).
Co-authored-by: fangliquan <fangliquan@qq.com>
Use is_runaway_repetition on completed replies so asked-for repetition with
distinct lines is delivered; stamp the truncated/retryable failure verdict,
log as diagnostic, and derive both repetition copies from one builder.
The periodic-run detector only matched exact repeats, so counter/noise loops
(#86581 shape) passed. OR it with main's window-count scan, and memoise
rejected runs so anchors inside them skip multiple-period re-expansion.
_maybe_disable_streaming now reuses anthropic_adapter._is_stream_unavailable_error
(its stream-unsupported + Bedrock IAM checks were a verbatim duplicate), so a custom
anthropic_messages provider raising 'Unexpected event order' retries without
streaming instead of repeating the same broken stream. Bedrock keeps
turn_recovery's sticky Converse switch.
Also drive the #60683 usage:null test through _call_anthropic so dropping the
normalize_stream_usage wiring in _open_anthropic_stream turns it red.
create_anthropic_message now runs normalize_stream_usage on the entered
MessageStream, so MiniMax usage:null no longer crashes the SDK after the
full generation and re-bills the request via messages.create(). Drops the
AttributeError/'nonetype'+'usage' substring classifier, which also turned
unrelated AttributeErrors into silent create() retries.
MiniMax-style Anthropic-compatible endpoints send usage: null on message_start
and message_delta; the SDK's accumulate_event() then raises AttributeError on
usage.output_tokens mid-iteration in _call_anthropic. Patch the raw SSE events
before the SDK accumulates them.
Co-authored-by: hejuntt1014 <252818347@qq.com>
Hand-applied from 5ab566fea1 (#45919): _build_partial_stream_stub now
takes api_mode and returns a Messages-shaped stub in anthropic_messages
mode, including the overflow_terminal and content-filter stub sites.
(cherry picked from commit 5ab566fea1, grafted onto the split helper)
When Anthropic-compatible providers (e.g. MiniMax) return usage: null on
message_start, the SDK crashes in get_final_message() with AttributeError
on output_tokens. Extend the existing stream-unavailable classifier and
apply the same create() fallback in the main streaming path.
Fixes#60683
Co-authored-by: Cursor <cursoragent@cursor.com>
(cherry picked from commit d7194c282a5359dc69436c84615997377d7d21c0)
A gap below the newest held id, and turns appended above an unpersisted current
turn, were summarized away without being read. Name those held ids and clone
the rest.
turn_context_compaction imports neither conversation_loop nor
turn_empty_response, so the function-local import is not guarding a
cycle. The helper's docstring now also covers the provider-switch
fallback hop that reached the provider.
The 2^n output-budget ladder was open-coded in three places. Only the
tool-call retry got the #72770 fix; the length-continuation retry
(ceiling max(32768, cap)) and the Codex-incomplete retry (ceiling
max(32768, base)) still re-sent the same budget once the cap reached
32768.
boosted_output_cap() now serves all three sites: it ladders from the
budget actually in use, its ceiling is max(32768, 2×cap), and it is
clamped to the model's known output limit (_get_anthropic_max_output on
anthropic_messages). When the request already sits at that limit it is
not doubled, so the retry no longer earns a provider 400 (#79715).
The length-continuation fix mirrors #72802.
Co-authored-by: webtecnica <contato@webtecnica.com.br>
Fallback activation after empty-response retries re-entered the outer loop
without refunding the provisional call, so a provider hop consumed an
iteration. Refund via _refund_api_call and carry the count back through
EmptyResponseVerdict -> FinalResponseVerdict -> s.api_call_count.
Salvaged from #88867.
Co-authored-by: RelaxJonh <92573950+RelaxJonh@users.noreply.github.com>
The dropped-tools partial-stream continuation now also states the cut was a
transport interruption and tools remain available. The pre-rewording
network-stub text stays in the compressor's synthetic-turn set so
crash-persisted nudges from older sessions are not mistaken for user turns.
Re-port of 6f01f7476b (PR #72273) onto the _LENGTH_CONTINUATION_NETWORK_STUB
constant that _get_continuation_prompt now returns: state the cut was a
transport interruption, that tools remain available, and drop 'Finish the
answer directly' which read as a text-only instruction.
Re-port of f3c6307d22 (PR #72802) formula hunk into agent/turn_truncation.py:
ladder from the requested cap when max_tokens is unset, and cap at 2x that cap
so a request already at >=32768 is not re-sent unchanged. Bundled
skills/email/himalaya files from the PR are dropped.