Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.
The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
Follow-up to the #103857 salvage. The Copilot pattern fallback now uses
is_astra_model (the shared exact set) instead of a gpt-6-astra prefix, so
speed-tier or unknown suffixes such as gpt-6-astra-pro stay off the Astra
ladder, matching every other Astra gate. Also drops the stale
'"max" is gpt-5.6-only' comment in the auxiliary Responses builder.
Addresses two P1 review findings on #121614.
1. `provider_model_ids`: a failed or empty relay probe fell through to the
canonical per-provider fetcher, sending the provider credential to exactly
the vendor host the user routed away from — recreating #121387 on the
failure path. A configured `model.base_url` relay is now TERMINAL for live
catalog egress and degrades to the local curated list instead. The curated
tail is extracted as `_static_catalog` and shared by both paths.
Fetchers that already resolve `model.base_url` themselves and degrade
locally (`_anthropic_catalog`, `_custom_catalog`, `_openai_catalog`, the
simple api-key fetchers) are excluded from interception via
`_RELAY_AWARE_CATALOG_FETCHERS` — they already satisfy the invariant and
produce a better-merged catalog.
2. `_try_anthropic`: an `explicit_base_url` that failed
`_is_anthropic_compatible_host` was silently dropped, leaving `base_url` at
the ambient/canonical host and continuing with the explicit credential —
a silent retarget of authority, not the refusal the PR body claimed. It now
returns unavailable before client construction.
Regressions: an egress-sentinel test pins that no vendor fetcher, profile
catalog or models.dev merge is reached after a relay 404/hang; the Anthropic
test now asserts no client and zero SDK builder calls.
The fold routed every aux Codex read through _resolve_codex_credential_and_base, leaving
_read_codex_access_token with no production callers; three test patches on it had gone inert
(including the 'should use pool token' guard). Point them at the live seams instead.
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
directly when the pool yields no token (no second uncached pool load,
no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
helper's error fallback reads the profile-scoped override, not the raw
process env.
- model setup flow: the confirm guards get the resolved Codex base, not
the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
Add 思考/反思/推理/推敲 to THINK_TAG_NAMES so the streaming scrubber, CLI and
gateway stream filters and the final-response stripper all hide them, and
derive the auxiliary-client reasoning strip from the same list instead of a
hard-coded copy. Bare bracketless markers (unverified) are not covered.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
openai/_streaming.py raises APIError(body=data["error"]), so exc.body is the
inner error object or a bare string, never {"error": ...}; the predicate was
False for every real stream error. Accept the inner object (any of
type/code/message), a non-empty string, or the wrapper. Share one
_exc_http_status helper with _is_transient_transport_error.
Tests now build the error through openai.OpenAI + httpx.MockTransport (HTTP 200
SSE error event) and exercise the parameter-strip rung's fall-through.
OpenAI-compatible relays can commit SSE with HTTP 200 and then emit an
OpenAI-style error event; the SDK raises a status-less APIError with the
event only on exc.body. Classify that shape as a route failure in the
recovery ladder's provider-fallback rung (_FALLBACK_REASONS, reason
"structured provider error") and let it fall through the parameter-strip
rung. Grafted from #101540 onto main's recovery-ladder layer.
try_activate_fallback already forwards the entry's base_url as
explicit_base_url, but the two providers with their own resolver branches
dropped it and built the client on the vendor's canonical host, sending the
fallback turn and the API key somewhere the user never configured.
Thread explicit_base_url into _try_openrouter and _try_anthropic and forward
it from both branches. Anthropic applies the same _is_anthropic_compatible_host
guard as the primary path: a foreign host would 401 every call.
Fixes#121359
Gemini rejects an unsupported schema by naming its own generationConfig
keys ("Unknown name response_json_schema", "Invalid value at
generation_config.response_schema...", "response mime type ... is
unsupported"), never OpenAI's response_format, so such 400s bypassed the
one-retry-without-format rung and surfaced as hard failures now that the
native adapter actually sends the schema. Teach _is_structured_output_rejection
the three Gemini phrasings; the status gate (400/422 only) is unchanged.
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
Eight secret readers wrapped the scoped get_secret() call in a broad
except Exception / contextlib.suppress and fell through to os.environ.
Under multiplex that env holds the default profile's value, so a bound
scope whose resolution fails silently borrowed another profile's
credential: the pairing allowlist reader could then persist the foreign
list into the served profile's .env, and the proxy key, tool gateway
token, OpenRouter and aux provider keys, ElevenLabs key, and Slack token
probe had the same shape.
Keep the deliberate UnscopedSecretError -> os.environ fallback (the
unscoped default-profile path legitimately reads its own env) and let
every other scoped-read failure propagate; the two availability probes
fail closed instead. config._scoped_environ_get now propagates as its
docstring already claimed.
#119410 registered gpt-6-terra alongside sol and luna from the request text, and the Codex
forward-compat synthesis put it (and -900k) in the live /model picker on every surface.
Nothing serves it: not the Codex account catalog (astra, sol, luna), not OpenRouter (same
three, plus -pro), and OpenAI's model page 404s. Forward-compat is for published tiers the
account catalog has not listed yet, not for guessed names. Removed from the Codex fallback
list and template chain, the context/900k tables, the effort ladder prefixes, the aux-client
family list, and the Nous/OpenRouter static catalogs; catalog JSON regenerated. A contract
test pins the synthesized GPT-6 set to the published tiers.
OpenAI shipped gpt-6-sol / gpt-6-terra / gpt-6-luna as the successors of the
gpt-5.6 tier line (Sol and Luna live on OpenRouter + the Nous Portal today).
The curated aggregator catalogs (OPENROUTER_MODELS and the derived nous list,
plus the published website model-catalog.json) now carry the gpt-6 tiers and
their -pro variants instead of the 5.6 ones; the openai-api curated fallback
lists them ahead of 5.6.
Codex OAuth support mirrors the 5.6 + Astra contract for every gpt-6 tier:
curated fallback + forward-compat synthesis (from the 5.6 twin or 5.5),
272K advertised fallback, the opt-in -900k picker variants with the
live-verified 900K bump (still capped by the catalog's max_context_window),
dated-snapshot eligibility, wire-suffix stripping, the compaction auto-raise
on the base slug, and the gpt-5.6 effort ladder (max allowed, minimal
rejected). Pricing rows for gpt-6-sol / gpt-6-luna come from OpenAI's model
pages (272K whole-request tier like Astra); Terra has no published page yet
so it deliberately has none.
/model gpt keeps resolving to the flagship: "astra" joins the rank-0 suffix
set so gpt-6-astra sorts above gpt-6-sol.
test_chat_sdk_transform_bypass stubs _relay_auxiliary_metadata with an empty
metadata dict; read aux_task/api_mode with .get so the seam keeps working.
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.
- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
post_api_request payload shape plus `aux_task`, `api_request_id`
(`aux-...`, shared by every attempt of one logical call), `retry_count`,
`streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
turn is in flight; fail-open (a raising/hung subscriber is logged and
the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
attempt shares (_relay_sync_completion / _relay_async_completion /
_relay_sync_stream) run under the hook pair — retries and fallbacks
included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
aux_task and no api_request events; raising subscriber never breaks
the call). First is red on origin/main.
Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.
Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
A named custom provider at api.anthropic.com (api_mode anthropic_messages) whose token comes from a
key_cmd callable lost the Claude Code OAuth identity: the aux custom routes hard-coded
is_oauth=False, agent_init/agent_runtime_helpers/client_lifecycle gated OAuth on provider=="anthropic"
and isinstance(key, str), and the callable-token client builder never added the OAuth betas or the
claude-code user-agent. Anthropic answers such a bare Bearer with 429 rate_limit_error "Error"
(#114967). One resolver, anthropic_credentials.anthropic_route_is_oauth(base_url, credential,
provider=), decides at every site: the route qualifies for the anthropic provider (unchanged) or an
exact api.anthropic.com host, the credential is a string or a callable materialized once
(CommandTokenSource caches), third-party hosts never qualify. model_metadata
_query_anthropic_context_length skips a callable credential instead of crashing agent init
(AttributeError on the same key_cmd route on current main).
Slimmer redo of #115007 by @liuhao1024 (same direction; one shared helper instead of per-site copies).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
Gate review: _relay_sync_stream applied the bypass before handing kwargs to
relay_llm.stream_current, so Relay's tracing saw `messages: []` for every managed streaming
aux call and an intercept that rewrote messages was silently discarded (the SDK merges
extra_body after the transform). The bypass now runs in the provider callback like every
other site; _create_with_progress_once / _acreate_with_progress apply it once at entry
instead of at each of three create() calls. New test: Relay sees the full conversation while
the SDK still gets only the placeholder (red on the previous head).
OpenAI SDK's maybe_transform walks messages and tools trees against type unions with the GIL held, causing 30ms+ CPU stalls on multi-MB conversations before hitting the network. By moving wire-format bulk fields into extra_body (and setting messages=[] at top level), requests produce byte-identical wire JSON bodies while bypassing client-side SDK transform overhead. Applied to auxiliary client create helpers and iteration summary attempts.
Closes#106776
Problem
-------
`_create_with_progress_once()` (and its async twin `_acreate_with_progress`)
ticked the aux forward-progress hook unconditionally at dispatch time, before
the provider had produced any payload. Compression wires that hook to
`CompressionCommitFence.touch_progress`, so a dispatch that dies before any
output — an auth refresh, a retry, a fallback — reset the 120s summary
inactivity timer. A zero-output attempt could therefore run until the 600s
total ceiling instead of idling out, and the fence's `progress_observed`
flipped True on dispatch alone.
Reproduction
------------
Install `aux_progress_hook(fence.touch_progress)`, use a fake client whose
`create()` immediately raises HTTP 401, then call `_create_with_progress_once()`.
On current main `fence.progress_observed` becomes True from dispatch alone.
Fix
---
Keep `_notify_aux_dispatch()` for dispatch telemetry, but drop the two
unconditional `_notify_aux_progress()` calls (sync + async). Progress now
ticks only for substantive stream payloads (`_ChatStreamAccumulator.feed`) or
a completed usable response (`_notify_aux_provider_response` on the plain-call
and shim paths) — the same contract the adapter event hooks already enforce.
Why not an alternative
----------------------
Keeping the dispatch tick and instead making the fence ignore it would leave
every other progress-hook consumer (gateway session hygiene) with the same
false-liveness signal; the hook contract is "forward progress", and dispatch
is not progress. The change is two deletions; the streamed path already ticks
per substantive chunk, so no liveness is lost for genuinely streaming calls.
Validation
----------
- New regression tests (sync + async): a 401 dispatch leaves
`fence.progress_observed == False` while dispatch telemetry still fires;
both fail on the pre-fix code and pass after.
- Updated the two tests that pinned the old dispatch tick
(`test_aux_progress_streaming.py`) and the relay-seam comment.
- `pytest tests/agent/test_aux_progress_streaming.py
tests/agent/test_aux_relay_progress_seam.py
tests/agent/test_auxiliary_explicit_cancellation.py
tests/agent/test_aux_affordable_402_retry.py` → 82 passed.
- Compression suite (review/progress/stall/worker-isolation/attempt-lifecycle
+ timeout-floor + progress-timeout) → 75 passed; the one intermittent
failure in `test_second_consecutive_stall_commits_the_deterministic_fallback_summary`
reproduces identically on pristine main (pre-existing flake, see #115204).
- `ruff check` clean on all three changed files.
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
auxiliary.<task>.provider: openai was expanded to custom + the user's OpenAI
endpoint only by agent/auxiliary_client.py (compression, vision, title
generation). hermes_cli/runtime_provider.py::resolve_runtime_provider — the
path background_review, curator, MoA slots and delegation use — had no such
expansion, so "openai" hit auth.resolve_provider's registry lookup and raised
"Unknown provider 'openai'". The alias table now lives once in
runtime_provider_custom.py (the direct-alias/custom sibling) and both paths
call it; resolve_runtime_provider applies it before the ladder.
_host_gated_env_key_candidates also pairs OPENAI_API_KEY with a base_url that
is exactly OPENAI_BASE_URL: the alias lands on that proxy when no block
base_url is set, and the key was issued for it — the host gate otherwise sent
the "no-key-required" placeholder there while the aux-client path used the key.
Slim redo of #116083 (same direction: shared alias, applied in the runtime
resolver) without the extra key gate and effective_provider threading.
Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
Some OpenAI-compatible routers answer an upstream connect timeout with a 200
ChatCompletion whose only assistant text is "Connect timeout, please try again
later." and zero completion tokens. The cherry-picked predicate only guarded
ChatCompletionsTransport.validate_response(); the streaming assembler emitted
the text to live callbacks before that check, the iteration-limit summary
normalized the response directly, and the auxiliary _validate_llm_response()
only checked the message shape — so the router failure surfaced as the answer
(#68396, maintainer keep_open review on #68433).
Share the predicate (is_router_timeout_shim) and apply it where each path reads
the response:
- streaming: hold shim-prefix text in the existing pending-text seam used for
echoed SSE, and release it at assembly only when the finished response is not
a shim; the assembled object then fails validate_response and the loop retries
- iteration summary: a shim reads as an empty summary, taking the retry slot
- auxiliary: a shim raises like a malformed response so the fallback chain moves on
Live probe (fake OpenAI-compatible server, first reply shim, real AIAgent, temp
HERMES_HOME): before, non-stream / stream / summary all returned the shim text
(stream also emitted it live); after, all three retry and return the real answer;
control (no shim) still makes one call.
Follow-up to the salvaged #108119 hunk that adds
auxiliary.<task>.no_progress_timeout: a value that is not a positive number
used to fall back to the 60s default silently, which is exactly the
"my 600s request aborted after 60s and nothing told me why" confusion the
key exists to remove. The resolver now logs a warning naming the task and
the rejected value before applying the default.
Documents the key on the configuration page (independent of
auxiliary.<task>.timeout; keepalive frames do not re-arm; capped at the
request timeout; host deadline/cancel still win; per-task scope) and adds a
real-config invariant test that drives call_llm through the genuine
CodexAuxiliaryClient path on a temp HERMES_HOME.
Fixes#108104