694 Commits

Author SHA1 Message Date
kshitijk4poor
8c9f2162c7 refactor(compressor): key the Codex stall on the stream guard's shared marker
Co-authored-by: happy5318 <5318happy@users.noreply.github.com>
(cherry picked from commit 83c0a8bb9c7d0e611bafe436d2e4da77144c26ec)
2026-09-27 18:19:07 +05:30
kshitijk4poor
d24aadfdd1 refactor(copilot): one GitHub effort clamp for the profile and main agent
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.

The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
2026-09-26 22:21:47 +05:30
kshitijk4poor
e991bfe24a fix(copilot): offline Astra effort fallback matches the exact Astra slugs
Follow-up to the #103857 salvage. The Copilot pattern fallback now uses
is_astra_model (the shared exact set) instead of a gpt-6-astra prefix, so
speed-tier or unknown suffixes such as gpt-6-astra-pro stay off the Astra
ladder, matching every other Astra gate. Also drops the stale
'"max" is gpt-5.6-only' comment in the auxiliary Responses builder.
2026-09-26 22:21:47 +05:30
Austin Pickett
6926997f05 fix(providers): make a configured relay terminal for catalog egress; refuse a foreign Anthropic endpoint
Addresses two P1 review findings on #121614.

1. `provider_model_ids`: a failed or empty relay probe fell through to the
   canonical per-provider fetcher, sending the provider credential to exactly
   the vendor host the user routed away from — recreating #121387 on the
   failure path. A configured `model.base_url` relay is now TERMINAL for live
   catalog egress and degrades to the local curated list instead. The curated
   tail is extracted as `_static_catalog` and shared by both paths.

   Fetchers that already resolve `model.base_url` themselves and degrade
   locally (`_anthropic_catalog`, `_custom_catalog`, `_openai_catalog`, the
   simple api-key fetchers) are excluded from interception via
   `_RELAY_AWARE_CATALOG_FETCHERS` — they already satisfy the invariant and
   produce a better-merged catalog.

2. `_try_anthropic`: an `explicit_base_url` that failed
   `_is_anthropic_compatible_host` was silently dropped, leaving `base_url` at
   the ambient/canonical host and continuing with the explicit credential —
   a silent retarget of authority, not the refusal the PR body claimed. It now
   returns unavailable before client construction.

Regressions: an egress-sentinel test pins that no vendor fetcher, profile
catalog or models.dev merge is reached after a relay 404/hang; the Anthropic
test now asserts no client and zero SDK builder calls.
2026-09-25 12:56:32 -04:00
Austin Pickett
0cc91bb396 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution 2026-09-25 12:50:14 -04:00
kshitijk4poor
56490ca109 refactor(aux): drop the now-uncalled _read_codex_access_token and retarget its test seams
The fold routed every aux Codex read through _resolve_codex_credential_and_base, leaving
_read_codex_access_token with no production callers; three test patches on it had gone inert
(including the 'should use pool token' guard). Point them at the live seams instead.
2026-09-25 21:27:06 +05:30
kshitijk4poor
e326520d50 refactor(codex): one credential/route authority for aux + image paths
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
  fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
  directly when the pool yields no token (no second uncached pool load,
  no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
  is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
  helper's error fallback reads the profile-scoped override, not the raw
  process env.
- model setup flow: the confirm guards get the resolved Codex base, not
  the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
2026-09-25 21:27:06 +05:30
kshitijk4poor
e87f673faa fix(codex): send catalog/image credentials only to their own route
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.

- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
  credential actually routes to (runtime_provider._pool_entry_mode_and_url:
  env > model.base_url while the row is canonical > row URL) instead of the
  ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
  resolved with the token from hermes_cli/models.py, the CLI default-model
  swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
  one pool selection; the image plugin, _build_codex_client and the raw
  Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
  chatgpt.com; a gateway key may probe its own gateway's /models.

Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.

Addresses @andrexibiza's review on #121508.
2026-09-25 21:27:06 +05:30
Austin Pickett
5af33f8375 Merge remote-tracking branch 'origin/main' into austin/fix/121347-base-url-resolution 2026-09-25 09:09:44 -04:00
ethernet
05a78671a7 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 13:13:35 -04:00
kshitijk4poor
2ab2270b32 fix(agent): hide MiniMax-M3 Chinese reasoning tags (#43827)
Add 思考/反思/推理/推敲 to THINK_TAG_NAMES so the streaming scrubber, CLI and
gateway stream filters and the final-response stripper all hide them, and
derive the auxiliary-client reasoning strip from the same list instead of a
hard-coded copy. Bare bracketless markers (unverified) are not covered.

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-24 22:41:10 +05:30
ethernet
0f65d698d0 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 13:10:14 -04:00
kshitijk4poor
12a9c4c694 fix(aux): detect the SDK's real in-stream error body shape (#101538)
openai/_streaming.py raises APIError(body=data["error"]), so exc.body is the
inner error object or a bare string, never {"error": ...}; the predicate was
False for every real stream error. Accept the inner object (any of
type/code/message), a non-empty string, or the wrapper. Share one
_exc_http_status helper with _is_transient_transport_error.

Tests now build the error through openai.OpenAI + httpx.MockTransport (HTTP 200
SSE error event) and exercise the parameter-strip rung's fall-through.
2026-09-24 22:36:08 +05:30
TheBlueHouse75
1ba8ebe50e fix(aux): fall back on status-less structured provider errors (#101538)
OpenAI-compatible relays can commit SSE with HTTP 200 and then emit an
OpenAI-style error event; the SDK raises a status-less APIError with the
event only on exc.body. Classify that shape as a route failure in the
recovery ladder's provider-fallback rung (_FALLBACK_REASONS, reason
"structured provider error") and let it fall through the parameter-strip
rung. Grafted from #101540 onto main's recovery-ladder layer.
2026-09-24 22:36:08 +05:30
Austin Pickett
3112afeaff fix(agent): honor a fallback entry's base_url for anthropic and openrouter
try_activate_fallback already forwards the entry's base_url as
explicit_base_url, but the two providers with their own resolver branches
dropped it and built the client on the vendor's canonical host, sending the
fallback turn and the API key somewhere the user never configured.

Thread explicit_base_url into _try_openrouter and _try_anthropic and forward
it from both branches. Anthropic applies the same _is_anthropic_compatible_host
guard as the primary path: a foreign host would 401 every call.

Fixes #121359
2026-09-24 09:47:00 -04:00
ethernet
66d54ebf51 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.github/workflows/docker.yml
#	Dockerfile
#	apps/desktop/src/app/settings/about-settings.tsx
#	docker/stage2-hook.sh
2026-09-24 02:23:44 -04:00
teknium1
236aab3e8f fix(auxiliary): recognise Gemini's structured-output 400 wording
Gemini rejects an unsupported schema by naming its own generationConfig
keys ("Unknown name response_json_schema", "Invalid value at
generation_config.response_schema...", "response mime type ... is
unsupported"), never OpenAI's response_format, so such 400s bypassed the
one-retry-without-format rung and surfaced as hard failures now that the
native adapter actually sends the schema. Teach _is_structured_output_rejection
the three Gemini phrasings; the status gate (400/422 only) is unchanged.
2026-09-23 22:42:14 -07:00
ethernet
c13ea774e6 refactor: make install-stamp.json the single runtime version identity
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.

Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).

Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
2026-09-23 11:41:01 -04:00
ethernet
f4a38b1ee8 Merge origin/main into ethie/pm-clean 2026-09-23 10:13:26 -04:00
beardthelion
0cbe552888 fix(scope): failed profile-scoped secret reads must not borrow os.environ
Eight secret readers wrapped the scoped get_secret() call in a broad
except Exception / contextlib.suppress and fell through to os.environ.
Under multiplex that env holds the default profile's value, so a bound
scope whose resolution fails silently borrowed another profile's
credential: the pairing allowlist reader could then persist the foreign
list into the served profile's .env, and the proxy key, tool gateway
token, OpenRouter and aux provider keys, ElevenLabs key, and Slack token
probe had the same shape.

Keep the deliberate UnscopedSecretError -> os.environ fallback (the
unscoped default-profile path legitimately reads its own env) and let
every other scoped-read failure propagate; the two availability probes
fail closed instead. config._scoped_environ_get now propagates as its
docstring already claimed.
2026-09-23 06:46:02 -07:00
ethernet
339229490c Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.github/workflows/tests-os.yml
#	.github/workflows/tests.yml
#	Dockerfile
#	hermes_cli/memory_setup.py
#	hermes_cli/update_cmd.py
#	hermes_cli/web_server_memory.py
#	plugins/memory/hindsight/README.md
#	plugins/memory/hindsight/__init__.py
#	plugins/memory/hindsight/embedded.py
#	plugins/memory/hindsight/plugin.yaml
#	plugins/memory/hindsight/settings.py
#	plugins/memory/hindsight/setup.py
#	tests/plugins/memory/test_bom_tolerant_config_reads.py
#	tests/plugins/memory/test_hindsight_env_perms.py
#	tests/plugins/memory/test_hindsight_provider.py
#	tests/tools/test_lazy_deps.py
#	tools/lazy_deps.py
#	uv.lock
#	website/docs/user-guide/docker.md
#	website/docs/user-guide/features/memory-providers.md
2026-09-23 05:45:02 -04:00
teknium1
38c289c014 fix(models): drop gpt-6-terra, a tier OpenAI never published
#119410 registered gpt-6-terra alongside sol and luna from the request text, and the Codex
forward-compat synthesis put it (and -900k) in the live /model picker on every surface.
Nothing serves it: not the Codex account catalog (astra, sol, luna), not OpenRouter (same
three, plus -pro), and OpenAI's model page 404s. Forward-compat is for published tiers the
account catalog has not listed yet, not for guessed names. Removed from the Codex fallback
list and template chain, the context/900k tables, the effort ladder prefixes, the aux-client
family list, and the Nous/OpenRouter static catalogs; catalog JSON regenerated. A contract
test pins the synthesized GPT-6 set to the published tiers.
2026-09-23 00:44:33 -07:00
ethernet
4b803147d3 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/src/app/contrib/onboarding-kickoff.ts
#	hermes_cli/profiles.py
2026-09-22 17:16:06 -04:00
teknium1
79ec1f2a34 feat(models): GPT-6 Sol/Terra/Luna replace the 5.6 tiers in Nous/OpenRouter catalogs, with Codex -900k variants
OpenAI shipped gpt-6-sol / gpt-6-terra / gpt-6-luna as the successors of the
gpt-5.6 tier line (Sol and Luna live on OpenRouter + the Nous Portal today).
The curated aggregator catalogs (OPENROUTER_MODELS and the derived nous list,
plus the published website model-catalog.json) now carry the gpt-6 tiers and
their -pro variants instead of the 5.6 ones; the openai-api curated fallback
lists them ahead of 5.6.

Codex OAuth support mirrors the 5.6 + Astra contract for every gpt-6 tier:
curated fallback + forward-compat synthesis (from the 5.6 twin or 5.5),
272K advertised fallback, the opt-in -900k picker variants with the
live-verified 900K bump (still capped by the catalog's max_context_window),
dated-snapshot eligibility, wire-suffix stripping, the compaction auto-raise
on the base slug, and the gpt-5.6 effort ladder (max allowed, minimal
rejected). Pricing rows for gpt-6-sol / gpt-6-luna come from OpenAI's model
pages (272K whole-request tier like Astra); Terra has no published page yet
so it deliberately has none.

/model gpt keeps resolving to the flagship: "astra" joins the rank-0 suffix
set so gpt-6-astra sorts above gpt-6-sol.
2026-09-22 11:50:39 -07:00
ethernet
c13287c915 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/electron/main.ts
#	hermes_cli/backup.py
#	hermes_cli/config.py
#	hermes_cli/plugin_catalog.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/plugins_cmd_catalog.py
#	hermes_cli/plugins_discovery.py
#	hermes_cli/profiles.py
#	hermes_cli/update_cmd_deps.py
#	pyproject.toml
#	tests/gateway/test_dm_topics.py
#	tests/hermes_cli/test_config.py
#	tests/hermes_cli/test_plugins_cmd.py
#	tests/hermes_cli/test_update_autostash.py
#	tests/tools/test_lazy_deps.py
#	tools/lazy_deps.py
#	tools/skill_ledger.py
#	utils.py
#	website/docs/user-guide/security.md
2026-09-22 05:16:50 -04:00
teknium1
966d091d6a fix(aux-hooks): tolerate test seams that stub relay metadata without task keys
test_chat_sdk_transform_bypass stubs _relay_auxiliary_metadata with an empty
metadata dict; read aux_task/api_mode with .get so the seam keeps working.
2026-09-22 01:19:12 -07:00
teknium1
0e5809566f feat(plugins): fire pre/post_auxiliary_call events on every auxiliary LLM call (#79733)
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.

- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
  post_api_request payload shape plus `aux_task`, `api_request_id`
  (`aux-...`, shared by every attempt of one logical call), `retry_count`,
  `streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
  turn is in flight; fail-open (a raising/hung subscriber is logged and
  the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
  attempt shares (_relay_sync_completion / _relay_async_completion /
  _relay_sync_stream) run under the hook pair — retries and fallbacks
  included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
  sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
  tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
  aux_task and no api_request events; raising subscriber never breaks
  the call). First is red on origin/main.

Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.

Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-22 01:19:12 -07:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
Mohamad Kanso
bd2b8124f4 fix(auth): prevent repeated copilot raw token exchange warnings (#114740) 2026-09-20 13:22:30 -07:00
teknium1
2b3bab0a5c fix(anthropic): key_cmd Claude Code OAuth identity survives on custom api.anthropic.com routes (main + aux)
A named custom provider at api.anthropic.com (api_mode anthropic_messages) whose token comes from a
key_cmd callable lost the Claude Code OAuth identity: the aux custom routes hard-coded
is_oauth=False, agent_init/agent_runtime_helpers/client_lifecycle gated OAuth on provider=="anthropic"
and isinstance(key, str), and the callable-token client builder never added the OAuth betas or the
claude-code user-agent. Anthropic answers such a bare Bearer with 429 rate_limit_error "Error"
(#114967). One resolver, anthropic_credentials.anthropic_route_is_oauth(base_url, credential,
provider=), decides at every site: the route qualifies for the anthropic provider (unchanged) or an
exact api.anthropic.com host, the credential is a string or a callable materialized once
(CommandTokenSource caches), third-party hosts never qualify. model_metadata
_query_anthropic_context_length skips a callable credential instead of crashing agent init
(AttributeError on the same key_cmd route on current main).

Slimmer redo of #115007 by @liuhao1024 (same direction; one shared helper instead of per-site copies).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 12:49:40 -07:00
fangliquan
9f7df273a7 fix(auxiliary): prioritize explicit reasoning config 2026-09-20 10:17:39 -07:00
fangliquan
0a407e1651 fix(auxiliary): consume promoted reasoning config 2026-09-20 10:17:39 -07:00
fangliquan
c86de9c44c fix(auxiliary): preserve raw reasoning body shapes 2026-09-20 10:17:39 -07:00
fangliquan
ab3448e075 fix(auxiliary): honor provider reasoning disable controls 2026-09-20 10:17:39 -07:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00
kshitijk4poor
ee17fff193 refactor(agent): one bypassing create() in _relay_sync_stream; bypass hoisted above the dispatch tick; dead-path hunk dropped
Gate review: the bypass line had landed between the #114938 comment and the branch it explains;
_acreate_with_stream has no production caller.
2026-09-20 16:33:18 +05:30
kshitijk4poor
8bca1f7167 fix(agent): relay-managed aux streams run the SDK bypass inside the provider callback
Gate review: _relay_sync_stream applied the bypass before handing kwargs to
relay_llm.stream_current, so Relay's tracing saw `messages: []` for every managed streaming
aux call and an intercept that rewrote messages was silently discarded (the SDK merges
extra_body after the transform). The bypass now runs in the provider callback like every
other site; _create_with_progress_once / _acreate_with_progress apply it once at entry
instead of at each of three create() calls. New test: Relay sees the full conversation while
the SDK still gets only the placeholder (red on the previous head).
2026-09-20 16:33:18 +05:30
ericmaddox
fc67cbaab3 perf(agent): bypass chat SDK request transform in auxiliary and summary calls
OpenAI SDK's maybe_transform walks messages and tools trees against type unions with the GIL held, causing 30ms+ CPU stalls on multi-MB conversations before hitting the network. By moving wire-format bulk fields into extra_body (and setting messages=[] at top level), requests produce byte-identical wire JSON bodies while bypassing client-side SDK transform overhead. Applied to auxiliary client create helpers and iteration summary attempts.

Closes #106776
2026-09-20 16:33:18 +05:30
Diamond Hands Dig
c2560f4dcc fix(compression): dispatch alone must not reset the summary idle fence (#114938)
Problem
-------
`_create_with_progress_once()` (and its async twin `_acreate_with_progress`)
ticked the aux forward-progress hook unconditionally at dispatch time, before
the provider had produced any payload. Compression wires that hook to
`CompressionCommitFence.touch_progress`, so a dispatch that dies before any
output — an auth refresh, a retry, a fallback — reset the 120s summary
inactivity timer. A zero-output attempt could therefore run until the 600s
total ceiling instead of idling out, and the fence's `progress_observed`
flipped True on dispatch alone.

Reproduction
------------
Install `aux_progress_hook(fence.touch_progress)`, use a fake client whose
`create()` immediately raises HTTP 401, then call `_create_with_progress_once()`.
On current main `fence.progress_observed` becomes True from dispatch alone.

Fix
---
Keep `_notify_aux_dispatch()` for dispatch telemetry, but drop the two
unconditional `_notify_aux_progress()` calls (sync + async). Progress now
ticks only for substantive stream payloads (`_ChatStreamAccumulator.feed`) or
a completed usable response (`_notify_aux_provider_response` on the plain-call
and shim paths) — the same contract the adapter event hooks already enforce.

Why not an alternative
----------------------
Keeping the dispatch tick and instead making the fence ignore it would leave
every other progress-hook consumer (gateway session hygiene) with the same
false-liveness signal; the hook contract is "forward progress", and dispatch
is not progress. The change is two deletions; the streamed path already ticks
per substantive chunk, so no liveness is lost for genuinely streaming calls.

Validation
----------
- New regression tests (sync + async): a 401 dispatch leaves
  `fence.progress_observed == False` while dispatch telemetry still fires;
  both fail on the pre-fix code and pass after.
- Updated the two tests that pinned the old dispatch tick
  (`test_aux_progress_streaming.py`) and the relay-seam comment.
- `pytest tests/agent/test_aux_progress_streaming.py
  tests/agent/test_aux_relay_progress_seam.py
  tests/agent/test_auxiliary_explicit_cancellation.py
  tests/agent/test_aux_affordable_402_retry.py` → 82 passed.
- Compression suite (review/progress/stall/worker-isolation/attempt-lifecycle
  + timeout-floor + progress-timeout) → 75 passed; the one intermittent
  failure in `test_second_consecutive_stall_commits_the_deterministic_fallback_summary`
  reproduces identically on pristine main (pre-existing flake, see #115204).
- `ruff check` clean on all three changed files.
2026-09-20 00:15:32 -07:00
fangliquan
8658d15957 fix(provider): honor explicit local provider names 2026-09-19 23:43:32 -07:00
fangliquan
204ebafb3d fix(provider): resolve named fallback before local alias 2026-09-19 23:43:32 -07:00
Konstantin Khlopkov
dacfd0efaa fix(auxiliary): resolve provider aliases from the authoritative auth table 2026-09-19 23:42:57 -07:00
ethernet
ce3d042f0c Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-19 23:08:38 -04:00
ethernet
e1576d06a6 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
2026-09-19 22:57:07 -04:00
teknium1
956a8c843d fix(aux): provider "openai" resolves the same on the runtime and aux-client paths
auxiliary.<task>.provider: openai was expanded to custom + the user's OpenAI
endpoint only by agent/auxiliary_client.py (compression, vision, title
generation). hermes_cli/runtime_provider.py::resolve_runtime_provider — the
path background_review, curator, MoA slots and delegation use — had no such
expansion, so "openai" hit auth.resolve_provider's registry lookup and raised
"Unknown provider 'openai'". The alias table now lives once in
runtime_provider_custom.py (the direct-alias/custom sibling) and both paths
call it; resolve_runtime_provider applies it before the ladder.

_host_gated_env_key_candidates also pairs OPENAI_API_KEY with a base_url that
is exactly OPENAI_BASE_URL: the alias lands on that proxy when no block
base_url is set, and the key was issued for it — the host gate otherwise sent
the "no-key-required" placeholder there while the aux-client path used the key.

Slim redo of #116083 (same direction: shared alias, applied in the runtime
resolver) without the extra key gate and effective_provider threading.

Co-authored-by: Mohamad Kanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-19 12:30:31 -07:00
Teknium
d6d9fac694 Merge pull request #115904 from NousResearch/fix/boa-res-R3-responses-reasoning-session-header
feat(providers): custom providers can carry the Hermes session id as an opt-in header (#86241, salvage #104566, supersedes #114496)
2026-09-19 11:14:31 -07:00
teknium1
7f9e1453a3 chore: merge origin/main (resolve agent/opencode_affinity.py, tests/hermes_cli/test_web_server_idle_proof.py) 2026-09-19 10:56:19 -07:00
teknium1
ffba66970a chore: merge origin/main (resolve agent/auxiliary_client.py) 2026-09-19 10:51:46 -07:00
teknium1
22753744fb fix: reject the HTTP-200 router timeout shim in every OpenAI-compatible consumer
Some OpenAI-compatible routers answer an upstream connect timeout with a 200
ChatCompletion whose only assistant text is "Connect timeout, please try again
later." and zero completion tokens. The cherry-picked predicate only guarded
ChatCompletionsTransport.validate_response(); the streaming assembler emitted
the text to live callbacks before that check, the iteration-limit summary
normalized the response directly, and the auxiliary _validate_llm_response()
only checked the message shape — so the router failure surfaced as the answer
(#68396, maintainer keep_open review on #68433).

Share the predicate (is_router_timeout_shim) and apply it where each path reads
the response:
- streaming: hold shim-prefix text in the existing pending-text seam used for
  echoed SSE, and release it at assembly only when the finished response is not
  a shim; the assembled object then fails validate_response and the loop retries
- iteration summary: a shim reads as an empty summary, taking the retry slot
- auxiliary: a shim raises like a malformed response so the fallback chain moves on

Live probe (fake OpenAI-compatible server, first reply shim, real AIAgent, temp
HERMES_HOME): before, non-stream / stream / summary all returned the shim text
(stream also emitted it live); after, all three retry and return the real answer;
control (no shim) still makes one call.
2026-09-19 10:03:43 -07:00
teknium1
c13d16eb35 feat(compression): warn on an invalid no_progress_timeout and document the key
Follow-up to the salvaged #108119 hunk that adds
auxiliary.<task>.no_progress_timeout: a value that is not a positive number
used to fall back to the 60s default silently, which is exactly the
"my 600s request aborted after 60s and nothing told me why" confusion the
key exists to remove. The resolver now logs a warning naming the task and
the rejected value before applying the default.

Documents the key on the configuration page (independent of
auxiliary.<task>.timeout; keepalive frames do not re-arm; capped at the
request timeout; host deadline/cancel still win; per-task scope) and adds a
real-config invariant test that drives call_llm through the genuine
CodexAuxiliaryClient path on a temp HERMES_HOME.

Fixes #108104
2026-09-19 09:51:45 -07:00