Commit Graph

38165 Commits

Author SHA1 Message Date
teknium1
626bc82f13 docs(skills): bundled and optional skills stop pointing the model at /tmp
Every SKILL.md, reference and helper script that told the agent to write
scratch files, clones, worktrees, logs or screenshots under /tmp now uses the
Hermes scratch dir (~/.hermes/cache/scratch, or $TMPDIR / the
${TMPDIR:-${HERMES_HOME:-$HOME/.hermes}/cache/scratch} fallback in scripts):
/tmp is RAM-backed tmpfs on most distros and fills under agent load, and it
does not exist on Termux or native Windows. Python snippets that take the
path expand ~ explicitly. Placeholder-only examples (MCP filesystem root,
media path) use /path/to/... instead. The four remaining literals name /tmp
as the anti-pattern or mirror upstream PyTorch docs and carry a no-tmp marker.
2026-09-19 10:44:26 -07:00
teknium1
0b909a595e docs(prompts): model-facing tool descriptions stop suggesting /tmp
The kanban artifact examples, MEDIA: example, pdftoppm OCR hint, the
background-poller tee hint and the Windows path note all showed the model a
literal /tmp path, which it copied into its own commands on every platform.
They now point at $TMPDIR (Hermes sets it to its scratch dir) or
~/.hermes/cache/scratch; the two Windows-note mentions that describe why /tmp
is the wrong choice keep the literal with a no-tmp marker.
2026-09-19 10:44:26 -07:00
teknium1
471ef5f4c3 fix(desktop): direct dictation requests time out per stt.openai.timeout instead of hanging
The client-direct STT path in the Desktop (voice-client-direct.ts) issued a
bare fetch to the provider with no AbortSignal, so a slow or wedged
transcription endpoint left the microphone stuck on "transcribing" forever;
the gateway's own transcription client already has a deadline.

- tools/voice_client_config.py::_resolve_stt_client_config adds `timeout_s`
  (from `stt.openai.timeout`, default 60; groq/deepinfra riders inherit) to
  the direct STT config the gateway hands the Desktop.
- voice-client-direct.ts routes all three direct STT fetches
  (openai-multipart, xai-stt, elevenlabs-stt) through sttFetch, which aborts
  at that deadline and surfaces "Transcription timed out after Ns".
- Docs: desktop.md dictation paragraph notes the shared budget.

Part of #112939
2026-09-19 10:40:04 -07:00
teknium1
0a1bc07f15 docs(honcho): say that initOnSessionStart blocks tools-mode startup until Honcho answers
What: the `initOnSessionStart` setup-schema description and the Honcho docs page
now state that in `recallMode: tools` the eager init runs synchronously during
agent construction, that it should stay false for Desktop or a local Honcho that
may be down, and that `timeout` in honcho.json caps each SDK call (both keys
added to the Full Config Reference table).

Why: #51492 reports Desktop `session.resume` / `prompt.submit` timeouts when a
local Honcho is down with `initOnSessionStart: true`. The synchronous eager path
is intentional design (#51562 was closed for that reason:
"ready-before-first-tool-call semantics"), so the fix for users is knowing the
trade-off and the two knobs that bound it.
2026-09-19 10:39:32 -07:00
teknium1
005c746d2c docs(config): document the gateway/proxy path to a fast tier via custom-provider extra_body
What: a "Fast tiers behind a gateway or proxy" subsection under Fast Mode in the
configuration guide plus a commented example in cli-config.yaml.example showing
`providers.<name>.extra_body: {service_tier: priority}`.

Why: agent.service_tier / `/fast` deliberately reach only first-party billing
endpoints (hermes_cli/models.py::_fast_mode_route_supported, c7e2e0b779), so
gateway and proxy users asked how to get a priority tier (#78097). The
per-provider extra_body is merged into every chat-completions request for that
endpoint (agent/agent_init.py::_merge_custom_provider_extra_body →
agent/fast_mode.py::effective_request_overrides), which already gives them a
supported path; it was just undocumented next to Fast Mode.
2026-09-19 10:39:01 -07:00
teknium1
bafb778edc feat(gateway): opt-in served_model footer field shows the model that really answered
What: a new opt-in gateway runtime-footer field `served_model` rendered as
`alias → served`. It is populated from the `x-litellm-model-id` response header
(fallback `x-litellm-model-api-base`) that routing proxies send on every chat
completion, captured through an httpx response hook installed on the agent's
OpenAI client (agent/served_model.py, wired in create_openai_client), and from
Hermes' own provider fallback (primary runtime model → active model) when no
header is present. The turn result carries `requested_model` / `served_model`;
gateway/run_turn.py passes them to the footer. Off unless listed in
`display.runtime_footer.fields`; the default field set renders exactly as before.

Why: behind a routing proxy (or during a silent Hermes fallback) every reply
shows the configured alias, so operators cannot see which deployment actually
served a request (#54864). The SDK's parsed objects drop response headers, so
the capture has to sit on the transport.
2026-09-19 10:38:30 -07:00
teknium1
f30fe581a8 fix(desktop): Model settings label the main-model context window and expose the compression model timeout
The Model page field "Context Window" read like a MoA/auxiliary setting,
so a reporter changed it expecting the main model to keep 1,048,576
tokens of context and could not tell which model it governed. Reword the
label and description so it is unambiguously the MAIN chat model's
context-window override (tokens; 0 = detected value; does not affect
auxiliary/MoA models) in en and the ja/ru/zh/zh-hant/ar copies.

Expose auxiliary.compression.timeout (default 120 s, see
hermes_cli/config_defaults.py::_aux) in the Memory & Context section next
to the other compression fields so the /compress timeout can be raised
from the Desktop instead of only via config.yaml. The backend schema
already flattens the nested key (web_server_config.py::
_build_schema_from_config), and field-copy lookup handles 3-segment keys
via defineFieldCopy/fieldCopyForSchemaKey, so only SECTIONS, FIELD_LABELS,
FIELD_DESCRIPTIONS and the locale copies change.

Invariant vitest in helpers.test.ts (red on base): the Memory & Context
section carries auxiliary.compression.timeout with label+description, and
the model_context_length label names the main model.

Part of #69912
2026-09-19 10:37:22 -07:00
teknium1
74c2439af2 fix(desktop): error card 'Switch provider' opens the live session model menu
The provider error card's "Switch provider" button navigated to
Settings → Models, which only changes the default provider/model for NEW
sessions — the failed chat kept its broken provider, so the user had to
find the composer pill on their own to actually recover the turn.

The button now calls requestModelMenuToggle(), the same bus request the
`composer.modelPicker` hotkey uses: it opens the composer pill's live
model menu (pane under the pointer, else the active composer), whose
picks go through model.switch on THIS session. When no chat surface is
on screen (requestModelMenuToggle returns false) it falls back to the
Settings → Models deep link as before. No new RPC.

Also refreshes the ErrorRecoveryPlan.switchProvider doc comment and the
Desktop user-guide bullet describing the button.

Tests: two invariant vitests on the error card (menu opened, no
navigation / menu unavailable → Settings deep link); both fail on base
where the click always navigates.

Part of #95066
2026-09-19 10:36:50 -07:00
teknium1
b056e1f36e fix(codex): send image attachments natively in app-server turn/start (#51053)
The codex_app_server runtime flattened every rich user turn into one text
item and replaced image parts with a literal "[image attached]" marker, so
screenshots and pasted images never reached the model. The app-server
`turn/start` protocol (schema v2/UserInput) accepts image inputs natively:
{type: text}, {type: image, url} for data:/http URLs and
{type: localImage, path} for local files.

_build_turn_input now maps Hermes content parts onto that list (text parts
stay text; image_url/{url} data or http refs become `image`, bare paths
become `localImage`) and run_turn sends the whole list. submitted_user_text
keeps the text portion only, which is what the wire echoes back, so the
echo-ownership dedup in codex_runtime is unchanged. Plain-string input and
the image-only default prompt behave as before.

Docs: trade-off table row for image attachments under the app-server runtime.
2026-09-19 10:36:18 -07:00
teknium1
4c4b0b748e docs: describe the summary-provider-overload abort in the compression failure ladder 2026-09-19 10:35:44 -07:00
fangliquanflq
af93d57ca7 fix(agent): preserve context when summary provider overloads 2026-09-19 10:35:44 -07:00
teknium1
b682a98ab8 test(codex): pin proxy override on rotation and model.base_url; document HERMES_CODEX_BASE_URL
Two invariant tests (red on origin/main): a 401 rotation onto a Codex pool
row keeps the HERMES_CODEX_BASE_URL target, and model.base_url under
model.provider: openai-codex resolves for pool credentials. Adds the
previously undocumented HERMES_CODEX_BASE_URL row to the environment
variables reference so proxy users can find the knob and its reach.
2026-09-19 10:33:08 -07:00
rjshrjndrn
7034628c92 fix(codex): proxy override survives credential rotation and model.base_url is honoured
HERMES_CODEX_BASE_URL was applied at pool resolution and on the auxiliary
clients, but two paths still sent the openai-codex provider back to the
default ChatGPT backend:

- credential rotation: client_lifecycle._swap_credential adopts
  PooledCredential.runtime_base_url, which for openai-codex was the pool
  row's stored canonical URL, so the first 401/429 rotation silently left
  the proxy. The override now lives in runtime_base_url, the one place every
  reader of a Codex pool row (resolution and rotation) goes through.
- model.base_url: the openai-codex branch of _pool_entry_mode_and_url
  returned before the generic model.base_url block. It now honours
  model.base_url under model.provider: openai-codex when the pool row still
  carries the canonical URL (env override keeps precedence).

Slim port of #40924 onto the current layout (the original patched
run_agent._swap_credential, the pool seeder and a new auth.py helper; the
seeder half landed in b62bb2a3d5, the helper is replaced by the
profile-scoped get_secret_str read the landed fixes already use).

Fixes #40913
2026-09-19 10:33:08 -07:00
teknium1
85b2a3df6c feat(cli): hermes usage [--json] prints the /usage account limits without a session
Codex 5h/weekly windows (and Anthropic/OpenRouter limits) were only reachable
through the interactive `/usage` slash command, so cron jobs and shell scripts
had no way to read quota state (#33094, #57476). `hermes usage` fetches the
same snapshot through `agent.account_usage.fetch_account_usage` — the credential
resolution a session with no live agent uses — and prints it with the same
renderer; `--json` emits one stable, documented document, exit 1 with a single
stderr line when no credential is configured or the fetch fails.

Slim redo of #81819 (@himanusia): top-level command instead of `hermes auth
usage`, no --all/--account/--reset (the per-entry paths rendered the wrong
account for anthropic and the default path bypassed the runtime resolver).

Co-authored-by: himanusia <himanusia@users.noreply.github.com>
2026-09-19 10:31:14 -07:00
teknium1
801e3fa6dd fix: truncated compaction summaries back off 60s/300s/900s across turns
A compression summary that ends in finish_reason=length is rejected (the
transcript is preserved) but was re-armed on a flat 30 s cooldown. Because the
compression attempt budget is per turn, every async delegation-completion
turn that arrived after the 30 s lapsed refilled the budget and re-issued the
same deterministic, capped summary request (#69637, reporter follow-up on
afc3d9d3: four identical truncations, one per turn).

Truncations now walk the existing _TIMEOUT_COOLDOWN_LADDER (60 -> 300 -> 900 s)
on their own _consecutive_truncation_failures counter, reset by a healthy
summary and carried across the compression-attempt ownership boundary like
the timeout streak. The counter is deliberately separate from
_consecutive_timeout_failures: that streak also arms the deterministic stall
fallback (_prior_timeout_failures), which a truncation must not trigger.
JSON-decode, closed-stream and empty-content failures keep the 30 s rung.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-19 10:30:09 -07:00
teknium1
65f6e31ffe fix: warn when a new Codex login is the same OpenAI account as a pooled credential
`hermes auth add openai-codex` now prints, after the "Added" line, which existing
openai-codex credential the fresh login duplicates when both tokens carry the
same principal (chatgpt_account_id + sub), and what to do about it.

Why: two logins of one OpenAI account share a single token family upstream, so
the provider revokes the older grant and the second pool entry adds no quota
while silently killing the first (the live datapoint in #47096). Only distinct
accounts rotate independently; the pool cannot keep both alive, so the best it
can do is tell the user at the moment they can still act on it. The identity
helper (_codex_principal_identity) already exists on main; the add path simply
did not consult it. The warning never blocks the add.

Part of #47096
2026-09-19 10:29:37 -07:00
teknium1
6992af1907 docs: explain manual demotion, dead Codex pools and launchd credential reload
Three gaps operators hit with pooled openai-codex logins on always-on gateways:

- `hermes auth priority` existed in the command table but nothing said when to use it.
  Document pre-emptive demotion of a healthy credential (issue #71095): a demoted entry
  stays OK, so the Codex quota-reset probe and `hermes auth reset` never touch it.
- Automatic adoption of ~/.codex/auth.json only repairs an existing Hermes login
  (hermes_cli/auth_codex.py::_recover_codex_tokens_from_cli). A pool whose Codex entries
  are all dead needs `hermes auth add openai-codex` in that home; say so where the dead
  state is documented (issue #63413).
- On the macOS LaunchAgent page, state what `launchctl kickstart -k` does to in-flight
  work (SIGTERM drain per agent.restart_drain_timeout / cron_drain_timeout, tool
  subprocesses killed, launchd relaunch) and that `hermes gateway restart` is the
  drain-aware path; agents are threads, not child processes, so no worker can outlive
  the gateway with stale credentials (issue #38261).
2026-09-19 10:29:06 -07:00
teknium1
8e180c69f6 feat: label a model.context_length pin and warn once when it disagrees with the provider
model.context_length short-circuits get_model_context_length ahead of every
provider source, but nothing told the user the number they saw was their own
pin rather than provider metadata (#66168). Keep the pin semantics (custom
endpoints depend on it) and make it visible instead:

- agent/context_pin.py: is_context_pinned / context_pin_suffix for renderers,
  and warn_once_on_pin_disagreement, called from _resolve_context_length at
  agent init. The advertised value comes from LOCAL sources only (endpoint
  cache, models.dev disk cache, hardcoded catalog) so the check never adds a
  network probe or latency to startup; one warning per (model, pin) per process.
- "(pinned)" label on every CLI surface that renders the window: welcome
  banner, /model switch summary, /usage current-context line, wide status bar.
- No new config key or env var; the numeric value used is unchanged.

Fixes #66168
2026-09-19 10:28:34 -07:00
teknium1
26675fbcb7 docs: state that codex_responses_compact_threshold applies only under codex_responses_native
The compression config table and the yaml example described
`codex_responses_compact_threshold` as "the server compaction trigger"
without saying it is read only while `codex_responses_native: true`
(agent/native_compaction.py::native_compaction_context_management returns
early otherwise). Users set it expecting local compaction to move and
filed #101867. Name the gate in the table row, the yaml comment and the
native-compaction prose, and point at `threshold` / `threshold_tokens`
as the local trigger.

Fixes #101867
2026-09-19 10:27:28 -07:00
teknium1
41ead01d6d fix: custom-provider key_env in the CLI picker/catalog helpers reads through the profile secret scope
hermes_cli/models_local.py::_api_key_from_provider_config and
hermes_cli/model_setup_flows_custom.py::_model_flow_named_custom still read
the variable named by key_env with raw os.getenv / os.environ.get. Under the
multiplexed gateway one process serves many profiles from one os.environ, so
the picker's endpoint probe and the native Ollama catalog headers could carry
a stale or another profile's value, while a key present only in the profile's
.env was not found (#67935). Every other credential read on the custom
provider path (runtime_provider_custom, runtime_provider_backends) already
goes through agent.secret_scope.get_secret_str, which falls back to
os.environ for single-profile processes, so the CLI behaviour is unchanged.
2026-09-19 10:26:08 -07:00
teknium1
1710f8413d fix: exhausted OAuth credential pool reports the cooldown and reset time, not "no credentials"
resolve_provider_client() returns None both when no credential exists and
when every pool entry is benched by a 429 / quota cooldown, so the two raise
sites (main-agent init and the auxiliary ladder) could only say "no
credentials were found. Run `hermes auth add …`" — wrong on both counts for
a valid OAuth grant that is merely rate-limited (#56810).

missing_provider_credentials_message() now first reads the persisted pool
state (read-only, no seeding) through pool_cooldown_message(): when every
live entry has a future cooldown, the message names how many credentials
are cooling down and the earliest reset time, and points at waiting, adding
a credential, or `hermes model`. Empty pools and elapsed cooldowns keep the
missing-credential text.
2026-09-19 10:25:30 -07:00
teknium1
b0952eac0c fix(agent): retry a truncated tool call with a boosted budget on the Responses wire
A ``status=incomplete`` (max_output_tokens) Responses reply whose function_call item was
cut mid-arguments used to end the turn on the first API call: ``_codex_finish_reason``
mapped it to ``incomplete``, the normalizer produced ``tool_calls``, and the tool-argument
validator refused the half-written JSON as "reply was cut off". Chat modes retry that same
shape up to 4x with a 2x/4x/8x/16x ``max_tokens`` boost; ``codex_responses`` was outside
``_CONTINUABLE_MODES`` so it never did.

``_derive_finish_reason`` now routes an incomplete Responses reply that carries a tool call
to ``finish_reason == "length"``, and ``codex_responses`` joins ``_CONTINUABLE_MODES`` so
``recover_from_truncation`` re-issues the same call with the boosted output cap (reasoning
kept, no interim row, no nudge). Text-only incompletes still return ``incomplete`` and stay
on the Codex continuation path, so the length text branch never double-continues them.

Fixes #91770
Co-authored-by: StanleyStetson <24758295+StanleyStetson@users.noreply.github.com>
2026-09-19 10:24:57 -07:00
teknium1
52705fba6c fix(delegate): child inherits the parent's live endpoint and key together
Without a delegation override the child took base_url from the parent's
live client (_client_kwargs / client.base_url) but api_key from the
surface attribute parent_agent.api_key. After a fallback or any runtime
swap the surface key can lag the live client, so the child was built with
a (base_url, key) pair that was never valid on the parent and died on an
instant, non-retryable 401 (#90009, reporter's log shape).

_inherit_parent_base_url becomes _inherit_parent_endpoint and returns the
(base_url, api_key) pair from one source: the live client kwargs when the
parent has an OpenAI-wire client, the surface attributes otherwise (native
Anthropic/Bedrock runtimes keep client=None). The override branches keep
the salvaged all-or-nothing rule from #111823.

Live probe (real AIAgent, real try_activate_fallback, two fake servers):
before, a parent on fallback B with a lagging surface key sent
`Bearer FAKE-KEY-PRIMARY-A` to server B from the child; after, the child
sends B's key to B. Natural fallback (no lag) is unchanged.

Tests trimmed to three invariants (override never borrows the parent
endpoint, provider without base_url is refused, no-override inherits
live endpoint+key together); all red on origin/main.

Fixes #90009
Salvages #111823
2026-09-19 10:24:21 -07:00
Luc Van
aa5b452bfc fix(delegation): assemble subagent credential bundle atomically
_resolve_child_runtime resolves provider and base_url with independent
per-field fallbacks, so a delegation.provider override whose runtime
comes back without a base_url silently inherits the PARENT's endpoint.
The child then sends e.g. copilot credentials to the parent's Codex URL:
every request 404s, and the fallback chain can't rescue it because its
dedup matches provider+model and skips the entry as a self-loop.

Build the bundle all-or-nothing: with a provider override, base_url
comes from the override only; otherwise everything is inherited from
the parent as before. _runtime_provider_credentials now refuses a
provider that resolves without a base_url, unless it is an ACP
transport (addressed by command) or a native-SDK provider
(bedrock/vertex/google), which legitimately have none.
2026-09-19 10:24:21 -07:00
teknium1
05c95d216e fix: read a top-level detail error body so a descriptive 400 is not a bare 400
`agent/error_classifier.py::_body_message_candidates` never yielded the
FastAPI-style top-level `detail` key (string, or nested `{"message": ...}`),
so the Codex gateway's `{"detail": "The '<model>' model is not supported when
using Codex with a ChatGPT account."}` 400 read as a *bare* 400. On a large
session `_classify_400`'s generic-400 heuristic then classified it as
context_overflow: the loop burned compression attempts ("Context length
exceeded (109,962 tokens). Cannot compress further.") and, because
`is_client_error` excludes overflow, the entitlement marker from #106549 never
ran. Reading `detail` makes it a descriptive rejection (format_error: abort +
fall back, no compression) and surfaces the provider's text as the message
instead of `Error code: 400 - {...}`. The pydantic list shape of `detail` is
still handled by `_oversized_message_content_rejection` and is not yielded.

Slim re-port of #100783's detail-body half onto the rule-table classifier;
the session/weekly usage-limit half is a separate class and was dropped.

Refs #81558
Refs #106475
Co-authored-by: Oleg Nagornyy <nagornyy.o@gmail.com>
2026-09-19 10:23:44 -07:00
teknium1
3934fe551c fix(codex): sweep the retired gpt-5.3-codex slug off every Codex-OAuth surface
What: the contributor's pick removes the slug from DEFAULT_CODEX_MODELS and the
forward-compat templates. This follow-up finishes the class on the sibling
surfaces that still taught users the retired slug:

- hermes_cli/cli_model_switch_mixin.py: `/model` -> openai-codex replaced an
  untouched default with the literal "gpt-5.3-codex" when live discovery failed;
  it now uses DEFAULT_CODEX_MODELS[0] so the fallback can never drift from the
  curated list again.
- run_agent.py: the Codex silent-hang hint no longer recommends the retired slug.
- website/docs (+ zh-Hans): the Codex OAuth vision/fallback examples used the
  retired slug and claimed a default that no longer exists (openai-codex has no
  implicit auxiliary model); examples now use gpt-5.4 and say to set `model`.
- tests: one invariant test pins the slug out of the curated list and every
  template tuple (#52492); hint/watchdog tests assert the retired slug is absent;
  catalog fixtures switched to live slugs.

Why: the ChatGPT Codex backend returns HTTP 400 "not supported when using Codex
with a ChatGPT account" for gpt-5.3-codex (#52492; second field report incl. the
official CLI on the #81558 thread, 2026-08-10). Same shape as e8955f222c, which
dropped gpt-5.2-codex / gpt-5.1-codex-max / gpt-5.1-codex-mini. Live discovery
still surfaces the slug if the backend re-enables it.
2026-09-19 10:22:33 -07:00
liuhao1024
82796e06d7 fix(codex): remove dead gpt-5.3-codex from curated fallback list
The chatgpt.com Codex backend now returns HTTP 400 for gpt-5.3-codex
on ChatGPT Pro accounts, matching the same pattern that previously
killed gpt-5.2-codex, gpt-5.1-codex-max, and gpt-5.1-codex-mini.

- Remove gpt-5.3-codex from DEFAULT_CODEX_MODELS
- Update _FORWARD_COMPAT_TEMPLATE_MODELS to re-anchor gpt-5.3-codex-spark
  on gpt-5.4/gpt-5.5 templates instead of the dead gpt-5.3-codex
- Update forward-compat test to use gpt-5.5 as trigger model

Fixes #52492
2026-09-19 10:22:33 -07:00
teknium1
ff4399a0d7 fix: route a slug shared by several catalogs to the provider the user can use
`detect_provider_for_model` took the FIRST static-catalog hit as the only
guess. `gpt-5.6-luna` (and the rest of the gpt-5.6 family) is listed by both
`openai-api` and `openai-codex`, so a user with a Codex OAuth grant and no
OPENAI_API_KEY was routed to a keyless openai-api on a fresh (`auto`)
session, or — after the credential gate — left on the current provider with
the request silently ignored, while the grant they hold was never
considered.

`_static_catalog_matches` now yields every catalog that lists the slug in
ladder order; `detect_provider_for_model` keeps its existing credential gate
and takes the first sibling the user actually has credentials for. A fresh
session with no usable provider anywhere still fails loudly on the first
guess, and a user holding both keys keeps today's openai-api routing.

Fixes #102775
Co-authored-by: Konstantin Khlopkov <konstantin.khlopkov93@gmail.com>
2026-09-19 10:22:01 -07:00
teknium1
55db0a9847 fix: local sub-64K context refusal stops assuming Ollama
When an OpenAI-compatible local server (llama.cpp, vLLM, ...) serves a
window below MINIMUM_CONTEXT_LENGTH, the startup refusal, the CLI banner
warning and the num_ctx log line were worded for Ollama (/api/show,
model.context_length advice that cannot raise a served window). Name the
served window and the server-agnostic remedies: start the server with a
>=64K context or set model.ollama_num_ctx (honoured on every local
endpoint) to the window it really serves. Hosted routes keep the
model.context_length advice. The floor itself is unchanged.

Part of #87075
2026-09-19 10:21:28 -07:00
teknium1
39726f8bb7 fix: effort pickers and /reasoning status say what ultra really sends
`ultra` is a Hermes-internal ladder step (agent/reasoning_effort.py):
no provider wire accepts it and every clamp maps it to the route's
strongest level (max on GPT-5.6 Codex / OpenAI-compatible routes, xhigh
on older Codex models), yet the CLI /reasoning status and the gateway
inline picker presented it as a distinct level. Add one helper,
effort_display_label(), that renders "<level> (sends <clamped> on this
route)" for any clamped level and call it from both surfaces; document
the clamp on the reasoning-effort page. The wire is untouched.

Part of #61634
2026-09-19 10:20:55 -07:00
teknium1
60e698fd2a test: explicit-key and env-key rungs agree on the OpenCode route per model
Invariant for #100854: resolving opencode-go with an explicit api key yields the same
(api_mode, base_url) as the env-key rung for both an Anthropic-routed model
(qwen3.8-flash) and a chat_completions model (glm-5.3-flash). Red on base for the
qwen case (chat_completions + /zen/go/v1 vs anthropic_messages + /zen/go).
2026-09-19 10:20:22 -07:00
fangliquanflq
07f2041e67 fix: explicit --api-key keeps OpenCode's per-model api_mode and base URL
The explicit-credential rung (hermes_cli/runtime_provider.py::_explicit_api_key_provider)
derived api_mode from the persisted config (opencode_by_model=False) and never ran
_finalize_base_url, so `hermes chat --provider opencode-go --api-key … -m qwen3.8-flash`
resolved to chat_completions + https://opencode.ai/zen/go/v1 and 404'd on the relay's SPA
page, while the same model through the env/config-key rung resolved to anthropic_messages +
https://opencode.ai/zen/go. OpenCode Zen/Go serve both wire formats behind one key, so the
route must follow the model on every credential rung, exactly as the env-key rung already
does (opencode_by_model=True + _finalize_base_url).

Ported from PR #100873 onto the post-refactor rung layout (the original hunk targeted
_resolve_explicit_runtime before 3ffd44acd3 split it into per-rung helpers).

Refs #100854
2026-09-19 10:20:22 -07:00
berg
c2b590841c fix(doctor): accept vendor/model slugs for openai-api on a custom endpoint
`hermes doctor` warned that `nvidia/z-ai/glm-5.2` "uses a vendor/model slug
but provider is 'openai-api'" even though model.base_url pointed at a local
OpenAI-compatible router, where the router owns the model namespace and the
vendor-prefixed id is the correct one. The warning now applies only when
openai-api targets api.openai.com (or has no base_url); the rest of the
vendor-slug policy is unchanged.

Ported from #73810 (@bergusdz) onto the doctor_config.py split; the
contributor's two parametrized cases are kept as the invariant test.

Part of #69912
Salvages #73810
2026-09-19 10:18:03 -07:00
teknium1
1f4fbd5145 fix(codex): alias the tool_search bridge on OpenAI Responses so Codex requests are not rejected
OpenAI Responses (api.openai.com and the ChatGPT Codex backend) now reserves the
``tool_search`` namespace for its native Tool Search. Progressive tool
disclosure advertises a client function literally named ``tool_search``, so
every request failed at validation with HTTP 400 "Function
'tool_search.tool_search' not allowed in reserved namespace 'tool_search'"
before any model output.

The xAI fix (#95003) already renames the bridge to ``hermes_tool_search`` on
the wire and maps it back in normalize_response; apply the same request-local
alias when the endpoint is the Codex backend or an OpenAI host. Other Responses
proxies keep the bare name.
2026-09-19 10:17:32 -07:00
teknium1
3dbaffd814 fix(tools): coerce a missing required to [] and pin the MCP normalizer entry (#56123)
Review follow-up. Object nodes that never had a `required` key sanitized to exactly
{type: object, properties: {}} (built-in read_window_below, read_terminal, ...), the shape
strict OpenAI-compatible backends reject as `null is not of type "array"`. Both sanitizer
sites (tools/schema_sanitizer.py::_sanitize_node and the top-level/default shapes, and the MCP
mirror tools/mcp_tool_schema.py::_repair_object_shape) now always emit a `required` list.

The MCP test now drives `_normalize_mcp_input_schema` (the production entry) instead of the
private repair helper, so removing the repair call goes red; the docstring describes the real
base behaviour (the key was dropped only when every entry was invalid).
2026-09-19 10:16:59 -07:00
teknium1
bb7d010cba chore: map jexbow's commit email to their GitHub login
Attribution gate requires every non-noreply author email on the branch to
have a contributors/emails mapping.
2026-09-19 10:16:59 -07:00
liangliang luo
4e1edd9220 fix: correct indentation of repaired['required'] assignment
The assignment was over-indented by one level (20 spaces instead of 16), causing it to be nested inside the if block rather than at the function body level.

Reported by @kvnloo in PR review.
2026-09-19 10:16:59 -07:00
liangliang luo
84861b187b test: add regression test for required array filtering in _repair_object_shape
Covers three scenarios:
- required entries not in properties are filtered out
- all valid entries: required unchanged
- all invalid entries: required becomes empty list

Requested by @kvnloo in PR review.
2026-09-19 10:16:59 -07:00
liangliang luo
9f2a1ad8f3 fix(mcp_tool_schema): preserve empty required arrays in _repair_object_shape
The _repair_object_shape() helper in tools/mcp_tool_schema.py mirrors the
schema_sanitizer._sanitize_node() logic that strips empty 'required' arrays.
When required is [], valid computes to [], the else branch fires, and
repaired.pop('required', None) drops the key entirely. DeepSeek-V4-Pro/Flash
then treats the missing key as null and rejects with 400.

Fix: preserve required: [] instead of popping — same approach as #20151
for the sanitizer site.

Refs #56123. Complements #20151. Supersedes #96151.
2026-09-19 10:16:59 -07:00
teknium1
1d6b56c6e0 fix(tests): drive the lookaround strip-and-retry through recover_after_classification
The lookaround regression test only exercised classify_api_error, so the
production branch it exists to protect (turn_recovery._recover_format_errors
-> strip_pattern_and_format(agent.tools) for llama_cpp_grammar_pattern) could
be disabled without any test going red.

Replace it with a test that runs the real entry point,
recover_after_classification, against a fake agent whose tool schema carries
a lookaround ``pattern``: the pattern must be stripped from agent.tools and
the result must request an immediate retry. Still asserts the classifier
mapping and the generic-schema-400 negative control. A/B: sabotaging the
grammar branch (``if False:``) turns this test red; restoring it goes green.
2026-09-19 10:16:27 -07:00
liuhao1024
3306e105b3 fix(agent): recover from strict OpenAI-compatible "regex lookaround is not supported" 400s
Strict OpenAI-compatible schema validators reject tool JSON Schemas whose
``pattern`` contains a lookahead/lookbehind with
"Invalid JSON schema: regex lookaround is not supported" (HTTP 400,
code invalid_json_schema). The classifier only knew the llama.cpp grammar
sentences, so the turn failed as a non-retryable client error even though
the existing strip-pattern/format retry path already fixes the request.

Match the lookaround sentence in the same grammar guard so it routes through
FailoverReason.llama_cpp_grammar_pattern and retries once with the sanitized
schemas. Ported from #42635 onto the current grammar_hit shape.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-19 10:16:27 -07:00
teknium1
6b5a8fad38 fix(api_server): emit the run record's runtime in the canonical _sanitize_runtime_metadata shape
/v1/runs published a thinner {provider, model} twin under the same wire key that
/v1/chat/completions, /v1/responses and the session chat stream fill via
_sanitize_runtime_metadata (route_source, requested, cleaned ids). Route the served pair through
the same classmethod so one concept has one schema; requested comes from the run's
requested_model/requested_provider overrides, route_source from the model_routes/raw/global
vocabulary the other endpoints use. Doc example updated.
2026-09-19 10:14:38 -07:00
Tranquil-Flow
1b02df86e3 fix(api_server): GET /v1/runs/{id} reports the served fallback runtime and cache-read tokens
GET /v1/runs/{run_id} and the run.completed SSE event only echoed the
requested model and three token counters. After a fallback_providers
switch the run record still named the requested model and usage had no
cache figures, so a supervisor polling /v1/runs for cost attribution
booked the whole run to the wrong provider at the wrong price, with
cache reads counted as full-price input.

agent.provider / agent.model still hold the fallback pair when
run_conversation() returns (the primary is restored only at the start of
the next turn), and agent.session_cache_read_tokens has the cache reads.
Surface them on the completed run status and the terminal event:

- usage gains cache_read_tokens / cache_write_tokens (_USAGE_FIELDS)
- _run_agent_sync returns the served {provider, model} as a third value
  and _finish stamps it as `runtime` on both the pollable status and the
  run.completed event (the persisted idempotent record inherits it).

Ported onto the refactored _run_agent_sync / _finish shape from #102161;
the served pair is read from the agent (the reporter's diff in #102101)
rather than the turn record, so a run whose result dict lacks the keys is
still attributed. Tests trimmed to two invariants.

Fixes #102101
2026-09-19 10:14:38 -07:00
teknium1
62277ebe29 fix(agent): copy-on-write for SDK-object tool calls in _repair_invalid_tool_call_names
An SDK-object tool call with an invalid function.name was coerced by
mutating fn.name in place, so the per-call copy edited the stored object.
Replace it with a dict copy (id/type/function) like the dict path, so the
persisted history object is never touched.
2026-09-19 10:13:20 -07:00
teknium1
7154ac036c fix(agent): coerce invalid stored tool-call names on every outbound request
A tool call whose function.name violates the provider pattern
^[A-Za-z0-9_-]{1,64}$ — OpenAI's synthetic `multi_tool_use.parallel`, or a
whole shell command a weak fallback model put into `name` (372 chars) — is
persisted once and then 400s every later request on a strict endpoint, so
the session silently pins itself to the lenient fallback model (#51944).

`agent/message_sanitization.py::coerce_tool_name` is now the single owner of
the coercion (valid → identity, invalid runs → `_`, cut at 64, empty →
fallback); the Codex Responses adapter uses it in place of its private copy,
and the pre-call sanitizer's nameless-call repair becomes
`_repair_invalid_tool_call_names`, so both outbound builders that already
call `sanitize_api_messages` — the main loop (turn_request_assembly) and the
iteration-limit summary (chat_completion_helpers) — send valid names. Dict
tool calls are rewritten copy-on-write, so the summary path's shallow message
copy never edits persisted history; tool results follow through
`_realign_tool_result_names`. Deterministic, so identical stored bytes always
render identical wire bytes (prompt-cache prefix stays stable).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-19 10:13:20 -07:00
teknium1
24f57317da fix(cron): drop a stale quota hold on schedule edit; scope the docs to the usage-probe case
update_job recomputes next_run_at from the new schedule but left quota_hold_until behind, so
the marker kept shielding a record that was no longer parked where it said. Clear it with the
schedule edit; the next fire re-parks (with a fresh notice) if the window is still closed.

cron.md claimed every 429 with a retry-after hint holds the job; only the provider-resolve
usage probe (a rate-limited AuthError) is classified. Say so.
2026-09-19 10:12:47 -07:00
teknium1
1d140a01ea fix(cron): pin the quota-hold scheduler wiring through a real run_one_job tick
The quota-hold tests only called mark_job_run(quota_hold_seconds=...) directly, so reverting
cron/scheduler.py and cron/scheduler_preflight.py to base left them green. Replace the direct
mark_job_run test with one that drives run_one_job (preflight on) with resolve_runtime_provider
raising the Codex quota AuthError: preflight must not report it as a missing credential, the one
delivered alert must carry the hold notice, and the job must be parked with quota_hold_until.
Reverting either scheduler file now goes red.
2026-09-19 10:12:47 -07:00
teknium1
b680dcd5d4 fix(cron): hold fires through a provider's closed usage window instead of re-firing every tick
A quota-exhausted subscription provider answers every request with a 429 and an explicit
`retry after <N>s` (Codex: ~33 h). The scheduler walked the fallback chain, failed, and
re-fired the job on its normal cadence into the same closed window — one guaranteed failure
and one delivered alert per tick until the window reopened (#89376: ~460 failed runs and
alerts across four profiles in one exhaustion).

Why: the retry-after was available at the failure site (a rate-limited AuthError in the
RuntimeError's cause chain) but nothing consumed it, and the stale-error re-arm (#62002)
would have pulled any manually parked next_run_at back after one cadence anyway.

What: cron/quota_hold.py mirrors cron/unreachable_retry.py in the opposite direction —
run_job flags the failure with the provider's remaining seconds (read only from a
rate-limited AuthError, never from arbitrary text), the one failure alert says the job is
held, and mark_job_run parks next_run_at at the first scheduled occurrence after the window
while stamping quota_hold_until; _job_is_stale_error_recurring treats an active hold as
deliberate. Any run that reaches the model clears it. Preflight no longer mislabels a
rate-limited AuthError as "provider credential missing" (blocked_config) so the hold applies
with or without a fallback chain.

Persisted marker shape (quota_hold_until) and the hold direction follow PR #89395.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Fixes #89376
2026-09-19 10:12:47 -07:00
teknium1
0830b7f741 fix(desktop): cap client-direct STT uploads at 60s
The three direct STT fetch() calls in voice-client-direct.ts carried no
AbortSignal, so a hung provider stalled dictation forever and the gateway
side stt timeout never applied to this surface. Each upload now carries
AbortSignal.timeout(60_000), mirroring the shared 60s STT default; the
configured stt timeout is not part of the voice client config today, so
the constant stands in for it (noted as follow-up in the PR body). One
existing vitest now asserts the signal is present (red without it).
2026-09-19 10:12:15 -07:00
teknium1
aa50383030 fix(qqbot): test the STT timeout at the HTTP call-site; share the 60s default and number parsing
The previous test only checked _resolve_stt_config, so restoring the old
hardcoded timeout=30.0 in _call_stt stayed green. The replacement drives
_call_stt with a fake HTTP client and asserts the timeout kwarg actually
posted (default 60.0 and a configured "95"); proven red with timeout=30.0.

_resolve_stt_config now reuses tools.transcription_common._config_number
(cast=float) instead of an inline try/float/except, and the 60s default
lives once as transcription_common.DEFAULT_STT_TIMEOUT, used by both the
OpenAI-SDK STT path and the QQ adapter. The import is lazy inside the
resolver (0.02s, two light modules, no gateway import cycle).
2026-09-19 10:12:15 -07:00