Commit Graph

3332 Commits

Author SHA1 Message Date
teknium1
30de041b01 docs: explain the three gateway connection-failure replies
The chat reply now differs for an interrupted connection, a refused/unroutable
endpoint and a cause-free SDK connection error; the FAQ names each wording and
what to do about it so a user reading one on Telegram/Discord/Slack knows
whether to restart the model server or just /retry (#116323).
2026-09-19 18:24:11 -07:00
Teknium
8a92051f20 Merge pull request #116328 from NousResearch/fix/boa-w3-new-reports-cron-codex
fix(cron): missing-credential preflight verdict names the profile and HERMES_HOME it read (#116213)
2026-09-19 14:33:58 -07:00
Teknium
059b3d416f Merge pull request #116353 from NousResearch/boa-w3-recovery
Transient provider outages no longer end the turn: bounded auto-recovery ladder after retries and fallback, visible on CLI/TUI/gateway/API/cron (#85426, #107307)
2026-09-19 14:33:20 -07:00
Teknium
8e0828a39c Merge pull request #116346 from NousResearch/fix/boa-w3-new-reports-aux-openai
auxiliary provider "openai" routes identically on every aux path; dead review routes are reported (#116055, salvage #116083)
2026-09-19 14:32:59 -07:00
Teknium
b6f8f8eb1f Merge pull request #116345 from NousResearch/boa-w3-small-b
Video generation tools no longer let the agent pick the model; video_gen.model is the only selector (Refs #83080)
2026-09-19 14:32:37 -07:00
Teknium
abc21bb427 Merge pull request #116344 from NousResearch/boa-w3-session-guard
session.create refuses a model its provider cannot serve instead of a dead first turn (#96817, salvage #96845)
2026-09-19 14:32:15 -07:00
Teknium
b5fcf635dc Merge pull request #116343 from NousResearch/boa-w3-codex-thread
Codex app-server thread survives an API-server restart: thread id persisted per session and thread/resume'd, fail-closed to a fresh thread (#100531, salvage #103352)
2026-09-19 14:31:55 -07:00
Teknium
d180fc4311 Merge pull request #116340 from NousResearch/feat/codex-browser-pkce-login
Codex login gains an opt-in browser PKCE flow on localhost:1455; device code stays default (#95743, salvage #97058)
2026-09-19 14:31:35 -07:00
Teknium
5322b4e637 Merge pull request #116335 from NousResearch/boa-w3-small
Exhausted connect retries now say which host, how many attempts and how large the request was (CLI, TUI/Desktop, gateway)
2026-09-19 14:31:14 -07:00
Teknium
1372de375a Merge pull request #116329 from NousResearch/fix/boa-w3-new-reports-codex-proxy-ctx
fix(context): proxied Codex routes (custom codex_responses provider, HERMES_CODEX_BASE_URL) resolve the 272K Codex window, not the 1.05M direct-API catalog (#116191, credit #116199 #116262)
2026-09-19 14:30:54 -07:00
Teknium
60e166cbf5 Merge pull request #116326 from NousResearch/fix/responses-turn-boundary-anchor
Chained /v1/responses turns no longer replay earlier tool calls or double the stored history after history repair/compaction (#89891, rebuild of #70695)
2026-09-19 14:30:05 -07:00
Teknium
a80ec24fb0 Merge pull request #116291 from NousResearch/fix/boa-partial-ultra-pill
fix(desktop,tui): reasoning pill says ultra sends max on this route instead of a distinct Ultra level (#61634)
2026-09-19 14:29:44 -07:00
Teknium
f22b7cedf9 Merge pull request #116288 from NousResearch/feat/boa-partial-429-retry-at-reset
feat(desktop): 429 usage-limit card can schedule one retry for when the limit resets (#98852)
2026-09-19 14:29:23 -07:00
teknium1
0752127c5e feat(agent): bounded auto-recovery ladder after retries and fallback are spent (#85426, #107307)
When api_max_retries and the fallback chain are both exhausted on a transient
outage (5xx, overloaded/529, connect/read timeout) and no answer text has been
delivered yet, the turn used to end with "API failed after N retries" even
though the provider would be back a minute later, leaving the user to notice
and re-send. settle_unrecovered_error now hands that case to
agent/turn_recovery_autorecover.py: up to agent.auto_recovery_cycles (default 5)
wait-and-retry cycles on a jittered 15/30/60/60/60 s schedule, a provider
Retry-After winning up to 120 s, each cycle announced on the status rail AND
the live wait line ("Provider temporarily unavailable — retrying automatically
in Ns (cycle k/5); press Esc to stop", with a per-surface stop hint) plus a
log line for cron. The interruptible wait is the existing one, so Esc/stop
cancels cleanly and a steering correction still rebuilds the turn.

Fallback stays first: the ladder engages only when _try_activate_fallback has
nothing left. Overload-class errors ride this schedule instead of growing a
separate overload backoff path (#107307). Non-retryable classes never enter
because they exit through the client-error branch above. TurnRetryState
carries the cycle counter; jittered_backoff supplies the schedule — no second
retry framework.

Credit: @MilevskyYakov's #85441 established the shape (reuse
try_recover_primary_transport / jittered_backoff / TurnRetryState, interrupt
mid-wait, never replay delivered text); this lands it at the exhaustion seam
main has today with a bounded default.
2026-09-19 12:38:55 -07:00
teknium1
15fec561db feat(doctor): resolve every routed auxiliary.<task> block and report the dead ones
hermes doctor validated model.provider but never the auxiliary blocks, so a
task whose provider could not be resolved (and therefore silently ran on the
main model) passed as healthy. The config check now runs each routed block
through resolve_runtime_provider — the entry point the tasks themselves use —
and turns a resolver error into a finding with the resolver's reason; the
green line shows the resolved provider@host so a block that fell to the public
default endpoint is visible too. Docs: the openai direct-API alias, its
endpoint precedence, and the new warning/doctor behaviour.
2026-09-19 12:30:32 -07:00
teknium1
e77e24a6a6 fix: persist the codex thread id per session and thread/resume it across an API-server restart
After a codex app-server turn's projected rows are durable in the session DB,
store the codex thread id as ``codex_thread_id`` in the session row's
model_config (atomic merge via patch_session_model_config; never for a retired
thread). The FIRST CodexAppServerSession an AIAgent builds for that session
passes the stored id as resume_thread_id, so a rebuilt agent — the next
/api/sessions/{id}/chat request, or the first turn after the API server or
gateway restarts — resumes the model-side thread before turn/start instead of
starting an empty one while Hermes' own transcript continues.

Fail closed when codex cannot hand the thread back (rollout gone, CODEX_HOME
changed, previous app-server killed mid-write): drop the stored id, start a
fresh thread on the same client, and say so once —
"Codex thread could not be resumed; starting a new one." — through
_emit_diagnostic_status, the lifecycle status rail every surface renders (CLI
vprint, TUI/Desktop and gateway status_callback). No other lifecycle change:
a retired or prompt-recreated session in the same process keeps today's
fresh-thread behaviour and overwrites the binding once its turn commits.

Why: CodexAppServerSession kept the thread id in memory only, so every
API-server restart (and every per-request agent) silently reset the model's
memory of the conversation (#100531). Supersedes the persistence half of
#100528 (_persist_projected_messages now reports durability) and the
refuse-and-raise policy of #103352 with the maintainer-approved bounded slice.
2026-09-19 12:27:01 -07:00
teknium1
19b29df13b fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.

The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.

Refs #83080
2026-09-19 12:22:44 -07:00
teknium1
72fccf2b20 fix: name host, attempts and request size when connect retries are exhausted
When every pre-stream connect attempt to an endpoint fails, the user only saw
"Connection error." repeated per outer retry; the host, the attempt count and
the serialized request size lived in agent.log alone. #97548's reporter had an
~829 KB Codex Responses request fail twice before the stream opened while short
chats went through, which is the request-size-limit signature, and nothing on
screen said so.

One buffered diagnostic line now flows through the existing retry-status path
(flushed on terminal failure, dropped on recovery) on both the Codex Responses
runtime and the Chat Completions stream worker: "Could not open a stream to
<host> after N attempts (request X KB); ...". The host comes from the failed
request's URL (the endpoint actually contacted, proxies included) and the size
from the buffered httpx request body. Re-entering the stream call from the
outer retry/fallback loop does not add another copy.

Part of #97548
2026-09-19 12:21:39 -07:00
teknium1
21ef1b97f9 fix(context): proxied Codex routes resolve the Codex OAuth window, not the direct-API catalog
A Codex OAuth model (gpt-6-astra, gpt-5.6-sol/terra/luna, gpt-5.5, ...) served
through a proxy — a custom provider with api_mode: codex_responses at
http://127.0.0.1:8317/v1, or the openai-codex provider behind
HERMES_CODEX_BASE_URL / model.base_url — resolved 1,050,000 from
DEFAULT_CONTEXT_LENGTHS instead of the 272,000 Codex enforces. The compressor
then set its trigger at 525,000 tokens, ~2x past the Codex limit and its 272K
billing tier, and a day of Astra sessions burned two Pro weekly windows.

Root cause: step 2 of get_model_context_length treated any non-known hostname
as a generic custom endpoint and never looked at the route's transport, so the
Codex table two steps away was skipped. The decision key is now the route:
_is_codex_route() is true for provider openai-codex on any URL and for a custom
entry whose canonical api_mode is codex_responses. A custom Codex route answers
from the Codex OAuth table first (with the opted-in -900k bump; unknown slugs
still take the endpoint probes), the native provider skips the custom-endpoint
step and keeps its live catalog probe in step 5, and both bypass the persistent
cache so a stale 1.05M entry cannot outlive the fix. Explicit
providers.<name>.models[].context_length / context_length / model.context_length
overrides still win (step 0 runs first); chat_completions proxies are unchanged.

Reported with the exact resolver path by @0xble (#116199); route-keyed lookup
per @JoaoMarcos44 (#116262).

Fixes #116191
2026-09-19 12:17:05 -07:00
teknium1
fc49f7619d fix(cron): a missing-credential preflight verdict names the profile and HERMES_HOME it read
The blocked_config reason for a missing provider credential now carries
"[profile '<name>', HERMES_HOME <path>]" — the home the scheduler actually
read auth.json/.env from — under the ticker's profile scope, so a
multiplexed satellite profile reports its own home, not the gateway's
launch home.

Why: #116213 reports an openai-codex cron job blocked with "No Codex
credentials stored" while an interactive session under "the same"
HERMES_HOME resolves the credential. A 5-shape x 5-scope live matrix
(singleton, expired+refreshable, pool-only, ~/.codex only, none; root,
named profile, root-only auth, multiplex default/named) on origin/main
and on the reporter's build 345cd2b0 shows interactive and cron
preflight agree in every cell — both call the same
resolve_runtime_provider ladder and read the same store. The remaining
explanation is a scheduler process reading a different home than the
shell (Docker HOME vs HERMES_HOME, a service unit without the shell's
env, a satellite profile), which the bare verdict could not reveal.
Naming the store the verdict judged makes that mismatch visible in the
one alert the user receives.

Part of #116213
2026-09-19 12:14:32 -07:00
teknium1
03cfc36f51 docs: describe the session.create model×provider refusal
Hosts speaking the gateway protocol need to know the new -32602 shape
(error.data model/provider/suggestions) and which pairs stay permissive.
2026-09-19 12:12:05 -07:00
teknium1
47ab9adc56 docs: document hermes auth add openai-codex --browser and auth.codex_login_flow
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
2026-09-19 12:11:59 -07:00
teknium1
d19963782b fix(api-server): anchor the Responses current turn on this turn's user row, not a history prefix match
`_response_messages_turn_start_index` located the current turn by matching
`result["messages"]` against `conversation_history + [user]` row by row. The
loop repairs host-fed history before its first call (consecutive assistant or
user rows merge, orphan tool results drop) and compaction rewrites it, so the
transcript legitimately stops sharing a prefix with the client history; the
match then returned 0 and the WHOLE transcript was treated as the current
turn: earlier turns' function_call / function_call_output items were replayed
as this turn's `output` / `run.completed` turn_messages, and the stored
`previous_response_id` chain grew by another copy of the history every turn.

Anchor on the loop's canonical current-turn user index instead
(`agent.turn_context.reanchor_current_turn_user_idx`: last user row that
says this turn's text, else the last user-originated row), and keep the
semantic prefix match only for suffix-only results without a user row
(mocked/legacy paths). The boundary logic moves to the topical sibling
`api_server_turn_boundary.py`; the routes mixin delegates.

The dedupe test that asserted "divergent transcript => append history +
transcript" encoded the duplication itself; it now asserts the fallback for
the case it actually exists for (a suffix-only result).

Fixes #89891
Co-authored-by: heyf <tonyheyifan@gmail.com>
2026-09-19 12:11:23 -07:00
justcarlosm
a4a0447b04 feat(telegram): drop_pending_on_cold_boot knob to preserve offline queue 2026-09-20 00:19:28 +05:30
Teknium
f0d8efe4e7 Merge pull request #115905 from NousResearch/fix/boa-res-R2-app-server-customprov
feat(codex): named custom providers work with the codex_app_server runtime (#75186, salvage #75191)
2026-09-19 11:35:36 -07:00
Teknium
4fda3bbdca Merge pull request #115846 from NousResearch/fix/boa-tts-stt-voice-tts-config-minlen
feat(tts): streaming TTS speaks a short first sentence sooner via tts.streaming.min_len on every surface (#96927, salvage #96933)
2026-09-19 11:33:15 -07:00
Teknium
e9f9a8d40f Merge pull request #115901 from NousResearch/fix/boa-res-R1-codex-auth-cli-adopt-guard
fix(auth): Codex CLI recovery no longer replaces a Hermes credential from another workspace (#73667, salvage #73677)
2026-09-19 11:28:31 -07:00
Teknium
0d2cb2cacc Merge pull request #115890 from NousResearch/fix/boa-res-R6-streaming-compaction-astra-oauth-gate
fix(compression): gpt-6-astra on Codex OAuth gets native server-side compaction when opted in (#103720, salvage #103718)
2026-09-19 11:28:10 -07:00
teknium1
049a62ab3d Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	website/docs/user-guide/features/codex-app-server-runtime.md
2026-09-19 11:23:47 -07:00
teknium1
171a1777b5 fix(desktop,tui): reasoning pill and effort rows say ultra sends max on this route
The Desktop pill, the catalog row meta and the effort radio row rendered a
clamped pick as a plain "Ultra", and the TUI status bar as "ultra" — presenting
a Hermes-internal step as a wire level the route does not have, while the CLI's
`/reasoning` shows "ultra (sends max on this route)" (#115876).

Both surfaces now read `session.info.reasoning_effort_wire`:
- Desktop pill / catalog row meta: compact "Ultra→Max", tooltip and aria-label
  "Effort: Ultra (sends Max on this route)" (new `modelOptions.sendsOnRoute`
  string in every locale, mirroring the CLI wording).
- Effort radio row for the selected clamped level: "Ultra (sends Max on this
  route)".
- TUI status bar: "ultra→max".
- Unknown wire ('' — not stamped yet, or an optimistic pick) or a verbatim
  level makes no claim, so nothing changes on routes that send the level as
  is. `setCurrentReasoningEffort` and the tile optimistic write clear the wire
  so a stale clamp never pairs with a new pick until the gateway re-stamps.

Docs: one sentence in the Desktop guide. Part of #61634.
2026-09-19 11:22:44 -07:00
teknium1
a169438178 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	hermes_cli/config_defaults.py
2026-09-19 11:22:01 -07:00
teknium1
6e1f5c9713 feat(desktop): 429 usage-limit card can schedule one retry for when the limit resets
The 429 error card already tells the user when the limit lifts (#115924:
`Limit resets at HH:mm (in 1h 05m)`), but they still had to be at the
keyboard at that moment to press Retry (#98852). The card now also offers
ONE `Retry when the limit resets (HH:mm)` button, driven by the same
`resets_at` the backend stamps on the failed turn.

Clicking it arms a client-side timer for the reset moment and swaps the
button for a live countdown (`Retrying at HH:mm — in 12m 03s`) plus a
Cancel control. When it fires it calls `aui.message().reload()` — the exact
call behind the Retry button — exactly once. If that retry 429s again the
card (and the button) simply reappear; nothing repeats unattended.

Deliberately minimal (maintainer option B): no `retry in 1h/3h/6h` menu,
no custom time picker, no persisted scheduler, no backend auto-resume. The
timer lives in the mounted card only: cancel, unmount, session switch, a
manual Retry or any new message (thread running / message no longer the
tail) all clear it, and closing the app fires nothing. The button is shown
only while the reset is ahead and inside `setTimeout`'s 2^31-1 ms ceiling
(a larger delay would fire immediately in browsers).

- apps/desktop/.../assistant-message.tsx::ScheduledRetryAction
- lib/error-surface.ts: formatResetClock / scheduledRetryDelayMs /
  formatCountdown (formatLimitReset now reuses the clock helper)
- i18n keys errorRetryAtReset / errorRetryScheduled /
  errorRetryScheduledCancel in types + en/zh/zh-hant/ja/ar
- docs: website/docs/user-guide/desktop.md error-card section
- tests: two invariant vitest cases with fake timers (fires once at
  resets_at and not before; Cancel/unmount never fire; past reset hides
  the button) — red on the base component, green here.
2026-09-19 11:20:10 -07:00
Teknium
ea94d88e25 Merge pull request #115851 from NousResearch/fix/boa-desktop-openai-custom-api-mode
fix(desktop): Custom Endpoints pick an API mode and keep /v1/models alias metadata (#93622, salvage #69824)
2026-09-19 11:16:54 -07:00
Teknium
271cf9e2d5 Merge pull request #115938 from NousResearch/fix/boa-res-R8-openai-native-search
feat(web): openai-native backend lets Codex Responses turns use the server-side web_search built-in (#19320, salvage #107377)
2026-09-19 11:15:31 -07:00
Teknium
8ea37a35d9 Merge pull request #115924 from NousResearch/fix/boa-res-R8-reset429-card
feat(desktop): usage-limit 429 error card shows when the limit resets (#98852)
2026-09-19 11:15:09 -07:00
Teknium
d6d9fac694 Merge pull request #115904 from NousResearch/fix/boa-res-R3-responses-reasoning-session-header
feat(providers): custom providers can carry the Hermes session id as an opt-in header (#86241, salvage #104566, supersedes #114496)
2026-09-19 11:14:31 -07:00
Teknium
3df8ad0d7f Merge pull request #115903 from NousResearch/fix/boa-res-R2-app-server-interim
fix(api-server): codex commentary reaches streaming clients on session SSE, /v1/runs and /v1/responses (#67580, salvage #67593 #67613)
2026-09-19 11:14:10 -07:00
Teknium
7612dc7ccb Merge pull request #115886 from NousResearch/fix/boa-res-R5-custom-endpoints-errors-local-ttfb
fix(streaming): local Responses endpoints get the local first-token grace instead of the 120 s hosted cutoff (#92302, salvage #92395)
2026-09-19 11:13:12 -07:00
Teknium
1285de7bbc Merge pull request #115880 from NousResearch/fix/boa-res-R4-routing-catalog-azure-text-verbosity
feat(openai): agent.text_verbosity controls Responses answer length (#20203, salvage #20258)
2026-09-19 11:12:52 -07:00
Teknium
7feb1af028 Merge pull request #115872 from NousResearch/fix/boa-res-R4-routing-catalog-azure-codex-900k-cap
fix(codex): opted-in -900k aliases cap at the live catalog max_context_window (#105443, salvage #105445)
2026-09-19 11:12:31 -07:00
Teknium
c99d016388 Merge pull request #115857 from NousResearch/fix/boa-response-store-api-server
fix(api-server): chained /v1/responses turns store each message once instead of doubling history every turn (#95137, #101644, salvage #85678)
2026-09-19 11:11:48 -07:00
Teknium
d6ff3bcfee Merge pull request #115840 from NousResearch/fix/boa-tts-stt-voice-sample-rate
fix(tts): OpenAI-compatible streaming TTS plays at the endpoint's reported sample rate (#76466, salvage #76501)
2026-09-19 11:10:21 -07:00
Teknium
61db846f19 Merge pull request #115831 from NousResearch/fix/boa-desktop-openai-codex-fallback
fix(desktop): open chats switch to a fallback provider added after they were opened (#95066, salvage #95139)
2026-09-19 11:09:59 -07:00
Teknium
699d46baea Merge pull request #115827 from NousResearch/fix/boa-vision-image-openai-opencodevision
fix(providers): opencode-go vision models attach images natively and resumed sessions stay on the Go endpoint (#96066, salvage #96116)
2026-09-19 11:09:37 -07:00
Teknium
656eb28309 Merge pull request #115822 from NousResearch/fix/boa-streaming-stalls-timeouts-reasoning-streak
fix(agent): Codex reasoning-only stalls switch to the fallback provider instead of ending incomplete (#67321, salvage #67336)
2026-09-19 11:09:15 -07:00
Teknium
709e7adcc4 Merge pull request #115791 from NousResearch/fix/boa-custom-provider-key-resolution
fix(providers): bare custom provider sends the key named by model.key_env; unset key_env is warned about (#67453, salvage #67554)
2026-09-19 11:08:09 -07:00
Teknium
438d2f061a Merge pull request #115763 from NousResearch/fix/boa-codex-oauth-refresh-401-soft-failure
fix(codex): usage-limit soft failures rotate the credential pool before provider fallback (#24159, salvage #24173)
2026-09-19 11:07:49 -07:00
teknium1
7f9e1453a3 chore: merge origin/main (resolve agent/opencode_affinity.py, tests/hermes_cli/test_web_server_idle_proof.py) 2026-09-19 10:56:19 -07:00
teknium1
7c2b81d320 chore: merge origin/main (resolve apps/desktop/src/components/assistant-ui/thread/assistant-message.tsx, website/docs/user-guide/desktop.md) 2026-09-19 10:54:43 -07:00
teknium1
4ff5d91ac7 chore: merge origin/main (resolve apps/desktop/src/app/settings/custom-endpoints-settings.tsx, apps/desktop/src/types/hermes.ts, hermes_cli/web_routers/config_env.py, tests/hermes_cli/test_web_server.py) 2026-09-19 10:54:40 -07:00