The chat reply now differs for an interrupted connection, a refused/unroutable
endpoint and a cause-free SDK connection error; the FAQ names each wording and
what to do about it so a user reading one on Telegram/Discord/Slack knows
whether to restart the model server or just /retry (#116323).
Transient provider outages no longer end the turn: bounded auto-recovery ladder after retries and fallback, visible on CLI/TUI/gateway/API/cron (#85426, #107307)
Codex app-server thread survives an API-server restart: thread id persisted per session and thread/resume'd, fail-closed to a fresh thread (#100531, salvage #103352)
Chained /v1/responses turns no longer replay earlier tool calls or double the stored history after history repair/compaction (#89891, rebuild of #70695)
When api_max_retries and the fallback chain are both exhausted on a transient
outage (5xx, overloaded/529, connect/read timeout) and no answer text has been
delivered yet, the turn used to end with "API failed after N retries" even
though the provider would be back a minute later, leaving the user to notice
and re-send. settle_unrecovered_error now hands that case to
agent/turn_recovery_autorecover.py: up to agent.auto_recovery_cycles (default 5)
wait-and-retry cycles on a jittered 15/30/60/60/60 s schedule, a provider
Retry-After winning up to 120 s, each cycle announced on the status rail AND
the live wait line ("Provider temporarily unavailable — retrying automatically
in Ns (cycle k/5); press Esc to stop", with a per-surface stop hint) plus a
log line for cron. The interruptible wait is the existing one, so Esc/stop
cancels cleanly and a steering correction still rebuilds the turn.
Fallback stays first: the ladder engages only when _try_activate_fallback has
nothing left. Overload-class errors ride this schedule instead of growing a
separate overload backoff path (#107307). Non-retryable classes never enter
because they exit through the client-error branch above. TurnRetryState
carries the cycle counter; jittered_backoff supplies the schedule — no second
retry framework.
Credit: @MilevskyYakov's #85441 established the shape (reuse
try_recover_primary_transport / jittered_backoff / TurnRetryState, interrupt
mid-wait, never replay delivered text); this lands it at the exhaustion seam
main has today with a bounded default.
hermes doctor validated model.provider but never the auxiliary blocks, so a
task whose provider could not be resolved (and therefore silently ran on the
main model) passed as healthy. The config check now runs each routed block
through resolve_runtime_provider — the entry point the tasks themselves use —
and turns a resolver error into a finding with the resolver's reason; the
green line shows the resolved provider@host so a block that fell to the public
default endpoint is visible too. Docs: the openai direct-API alias, its
endpoint precedence, and the new warning/doctor behaviour.
After a codex app-server turn's projected rows are durable in the session DB,
store the codex thread id as ``codex_thread_id`` in the session row's
model_config (atomic merge via patch_session_model_config; never for a retired
thread). The FIRST CodexAppServerSession an AIAgent builds for that session
passes the stored id as resume_thread_id, so a rebuilt agent — the next
/api/sessions/{id}/chat request, or the first turn after the API server or
gateway restarts — resumes the model-side thread before turn/start instead of
starting an empty one while Hermes' own transcript continues.
Fail closed when codex cannot hand the thread back (rollout gone, CODEX_HOME
changed, previous app-server killed mid-write): drop the stored id, start a
fresh thread on the same client, and say so once —
"Codex thread could not be resumed; starting a new one." — through
_emit_diagnostic_status, the lifecycle status rail every surface renders (CLI
vprint, TUI/Desktop and gateway status_callback). No other lifecycle change:
a retired or prompt-recreated session in the same process keeps today's
fresh-thread behaviour and overwrites the binding once its turn commits.
Why: CodexAppServerSession kept the thread id in memory only, so every
API-server restart (and every per-request agent) silently reset the model's
memory of the conversation (#100531). Supersedes the persistence half of
#100528 (_persist_projected_messages now reports durability) and the
refuse-and-raise policy of #103352 with the maintainer-approved bounded slice.
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.
The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.
Refs #83080
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.
One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.
Part of #97548
A Codex OAuth model (gpt-6-astra, gpt-5.6-sol/terra/luna, gpt-5.5, ...) served
through a proxy — a custom provider with api_mode: codex_responses at
http://127.0.0.1:8317/v1, or the openai-codex provider behind
HERMES_CODEX_BASE_URL / model.base_url — resolved 1,050,000 from
DEFAULT_CONTEXT_LENGTHS instead of the 272,000 Codex enforces. The compressor
then set its trigger at 525,000 tokens, ~2x past the Codex limit and its 272K
billing tier, and a day of Astra sessions burned two Pro weekly windows.
Root cause: step 2 of get_model_context_length treated any non-known hostname
as a generic custom endpoint and never looked at the route's transport, so the
Codex table two steps away was skipped. The decision key is now the route:
_is_codex_route() is true for provider openai-codex on any URL and for a custom
entry whose canonical api_mode is codex_responses. A custom Codex route answers
from the Codex OAuth table first (with the opted-in -900k bump; unknown slugs
still take the endpoint probes), the native provider skips the custom-endpoint
step and keeps its live catalog probe in step 5, and both bypass the persistent
cache so a stale 1.05M entry cannot outlive the fix. Explicit
providers.<name>.models[].context_length / context_length / model.context_length
overrides still win (step 0 runs first); chat_completions proxies are unchanged.
Reported with the exact resolver path by @0xble (#116199); route-keyed lookup
per @JoaoMarcos44 (#116262).
Fixes#116191
The blocked_config reason for a missing provider credential now carries
"[profile '<name>', HERMES_HOME <path>]" — the home the scheduler actually
read auth.json/.env from — under the ticker's profile scope, so a
multiplexed satellite profile reports its own home, not the gateway's
launch home.
Why: #116213 reports an openai-codex cron job blocked with "No Codex
credentials stored" while an interactive session under "the same"
HERMES_HOME resolves the credential. A 5-shape x 5-scope live matrix
(singleton, expired+refreshable, pool-only, ~/.codex only, none; root,
named profile, root-only auth, multiplex default/named) on origin/main
and on the reporter's build 345cd2b0 shows interactive and cron
preflight agree in every cell — both call the same
resolve_runtime_provider ladder and read the same store. The remaining
explanation is a scheduler process reading a different home than the
shell (Docker HOME vs HERMES_HOME, a service unit without the shell's
env, a satellite profile), which the bare verdict could not reveal.
Naming the store the verdict judged makes that mismatch visible in the
one alert the user receives.
Part of #116213
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
`_response_messages_turn_start_index` located the current turn by matching
`result["messages"]` against `conversation_history + [user]` row by row. The
loop repairs host-fed history before its first call (consecutive assistant or
user rows merge, orphan tool results drop) and compaction rewrites it, so the
transcript legitimately stops sharing a prefix with the client history; the
match then returned 0 and the WHOLE transcript was treated as the current
turn: earlier turns' function_call / function_call_output items were replayed
as this turn's `output` / `run.completed` turn_messages, and the stored
`previous_response_id` chain grew by another copy of the history every turn.
Anchor on the loop's canonical current-turn user index instead
(`agent.turn_context.reanchor_current_turn_user_idx`: last user row that
says this turn's text, else the last user-originated row), and keep the
semantic prefix match only for suffix-only results without a user row
(mocked/legacy paths). The boundary logic moves to the topical sibling
`api_server_turn_boundary.py`; the routes mixin delegates.
The dedupe test that asserted "divergent transcript => append history +
transcript" encoded the duplication itself; it now asserts the fallback for
the case it actually exists for (a suffix-only result).
Fixes#89891
Co-authored-by: heyf <tonyheyifan@gmail.com>
The Desktop pill, the catalog row meta and the effort radio row rendered a
clamped pick as a plain "Ultra", and the TUI status bar as "ultra" — presenting
a Hermes-internal step as a wire level the route does not have, while the CLI's
`/reasoning` shows "ultra (sends max on this route)" (#115876).
Both surfaces now read `session.info.reasoning_effort_wire`:
- Desktop pill / catalog row meta: compact "Ultra→Max", tooltip and aria-label
"Effort: Ultra (sends Max on this route)" (new `modelOptions.sendsOnRoute`
string in every locale, mirroring the CLI wording).
- Effort radio row for the selected clamped level: "Ultra (sends Max on this
route)".
- TUI status bar: "ultra→max".
- Unknown wire ('' — not stamped yet, or an optimistic pick) or a verbatim
level makes no claim, so nothing changes on routes that send the level as
is. `setCurrentReasoningEffort` and the tile optimistic write clear the wire
so a stale clamp never pairs with a new pick until the gateway re-stamps.
Docs: one sentence in the Desktop guide. Part of #61634.
The 429 error card already tells the user when the limit lifts (#115924:
`Limit resets at HH:mm (in 1h 05m)`), but they still had to be at the
keyboard at that moment to press Retry (#98852). The card now also offers
ONE `Retry when the limit resets (HH:mm)` button, driven by the same
`resets_at` the backend stamps on the failed turn.
Clicking it arms a client-side timer for the reset moment and swaps the
button for a live countdown (`Retrying at HH:mm — in 12m 03s`) plus a
Cancel control. When it fires it calls `aui.message().reload()` — the exact
call behind the Retry button — exactly once. If that retry 429s again the
card (and the button) simply reappear; nothing repeats unattended.
Deliberately minimal (maintainer option B): no `retry in 1h/3h/6h` menu,
no custom time picker, no persisted scheduler, no backend auto-resume. The
timer lives in the mounted card only: cancel, unmount, session switch, a
manual Retry or any new message (thread running / message no longer the
tail) all clear it, and closing the app fires nothing. The button is shown
only while the reset is ahead and inside `setTimeout`'s 2^31-1 ms ceiling
(a larger delay would fire immediately in browsers).
- apps/desktop/.../assistant-message.tsx::ScheduledRetryAction
- lib/error-surface.ts: formatResetClock / scheduledRetryDelayMs /
formatCountdown (formatLimitReset now reuses the clock helper)
- i18n keys errorRetryAtReset / errorRetryScheduled /
errorRetryScheduledCancel in types + en/zh/zh-hant/ja/ar
- docs: website/docs/user-guide/desktop.md error-card section
- tests: two invariant vitest cases with fake timers (fires once at
resets_at and not before; Cancel/unmount never fire; past reset hides
the button) — red on the base component, green here.