Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>
`provider_routing.models.<model-id>` now takes the same only/ignore/order/sort/
require_parameters/data_collection keys and overlays the flat provider_routing
values whenever the agent is on that model. Resolution lives in the one
chokepoint every request path already uses (_provider_preferences_for_agent),
so CLI, gateway, TUI/Desktop, cron, /model switches, fallback activation and
delegated children on another model all honour it with no per-surface plumbing.
Matching is spelling-tolerant, sharing _canonical_model_variants with
agent.reasoning_overrides.
The OpenRouter profile's speed-tier pin no longer overwrites an explicit user
`only` on the BASE gpt-6-astra slug: the pin exists to keep default routing off
flex/fast, and a user pin is the stronger intent (only: [openai] stays [openai]
instead of becoming [openai, azure, azure/us]). Tier slugs (-fast/-flex) keep
owning `only`.
Live A/B (config only: {gpt-6-astra: [openai], claude-fable-5.1: [anthropic]}):
main sent {"sort":"price"} for fable and OpenRouter served it from Azure; with
this change it sends {"only":["anthropic"],"sort":"price"} and Anthropic serves it.
Schema proposed in #24495 (samplesabotage) and #100711 (Artemonim); this is a
slim chokepoint implementation of that design.
Co-authored-by: samplesabotage <samplesabotage@users.noreply.github.com>
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:
- search: POST https://api.perplexity.ai/search (documented Search API),
search_context_size=low so `snippet` stays description-sized;
results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
route behind `pplx content snippets` (the CLI's `content fetch` is
deprecated upstream). web_extract has no query, so the URLs' path words
serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
backend set + credential ladder + availability probe (web_tools),
registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
dump key lists, nous_subscription direct-credential detection, setup
summary, test conftests, docs.
Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.
OpenCode pins requests sharing an x-opencode-session value to one upstream
backend, which is what keeps its prompt cache warm across a conversation.
Hermes never sent it, so cache ratios on OpenCode traffic were poor.
- agent/opencode_affinity.py: single owner of the header — target detection
(built-in zen/go/free, custom opencode-* providers, any opencode.ai URL)
and the key (affinity scope → conversation root → session id, cron
timestamp stripped), same resolution as OpenRouter/xAI affinity hints.
- build_api_kwargs: merged once after the per-mode builder, so
chat_completions, codex_responses and anthropic_messages all carry it.
- auxiliary _build_call_kwargs: same key from the runtime-main session so
compression/title/vision calls stay on the conversation's backend; the
aux Codex and Anthropic adapters now forward extra_headers.
Closes#81584, #81832 (deepseek-v4-flash 400 without the header).
- alibaba-coding-plan-cn / alibaba-token-plan-cn keep the shared intl key vars
as ordered fallbacks after their dedicated *_CN_API_KEY, so users who set
ALIBABA_CODING_PLAN_API_KEY / ALIBABA_TOKEN_PLAN_API_KEY for the CN endpoint
keep working (the PR as filed dropped them).
- list_authenticated_providers hides a '-cn' row whose only lit key vars are
ones it shares with its non-CN sibling, unless that CN provider is the
configured model.provider. With only the shared key: one row, not two;
DASHSCOPE_API_KEY alone: 3 alibaba rows, not 4.
- Docs: environment-variables.md, providers.md.
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
Sweeper review, all three points:
- Placement: no new plugins/model-providers/ directory. The Token Plan
profiles register from the existing alibaba plugin module — one module
per vendor, matching how the kimi module carries both of its endpoint
variants. Token Plan is the same vendor/service (Model Studio), same
OpenAI-compatible protocol, its own key + endpoints; splitting to a
standalone repo remains a 5-minute change if maintainers prefer.
- Runtime coverage: TestRuntimeAlibabaRegionalAndTokenPlan exercises
resolve_runtime_provider() for all four variants — provider, api_key,
api_mode, base_url — alongside the existing zai/minimax/kilocode
runtime regressions.
- Docs: providers.md, environment-variables.md, cli-commands.md updated
with the bundled variants and their env keys/base-url overrides.
- Widen opencode_zen_free_runtime healing to the union of the static floor,
the in-process live memo, and the SWR disk cache — a newly-live free model
now heals opencode-go/zen selections without a release (sibling site the
original PR missed).
- Memoize _fetch_opencode_free_models() in-process (5 min, negative caching
included) so direct provider_model_ids() validation callers don't each
block on a network round-trip or timeout.
- Drop delisted x-preview-f-free from the offline floor and setup.py sample
list (offline fallback must not offer a model that 401s); add the newly
live deepseek-v4-flash-free / mimo-v2.5-free to setup.py.
- Update stale test fixtures to a live exemplar; add regression tests for
memoization, negative caching, and union healing; docs note in providers.md.
GSC page-level data shows configuration (171K impr, pos 5.3), quickstart
(135K, 5.3), providers (123K, 5.3), web-dashboard (68K, 6.5), docker
(37K, 6.3) and the desktop app page (98K, 3.9) all losing ranking
headroom because their title/H1 are bare nouns instead of the terms
people search.
- Frontmatter title + H1 now carry the query terms on all six pages
- Desktop docs page links back to the new marketing /desktop product
page, joining the two official properties Google sees for the query
Done by Hermes Agent (deepseek-v4-pro via nous), Nous Research.
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
- Updated the Tavily API key description to clarify that it is optional and keyless access is supported.
- Modified the Tavily plugin and provider to handle requests with or without an API key, using Bearer authentication when the key is provided.
- Enhanced documentation to reflect the new keyless functionality and updated environment variable descriptions.
- Added tests to ensure correct behavior for both keyed and keyless requests.
- plugins/model-providers/meta-ai/__init__.py: drop out-of-tree install
instructions from the module docstring (now bundled)
- tests/providers/test_meta_ai_profile.py: port the plugin's test suite
into the repo (registry discovery instead of file-location import)
- website/docs/integrations/providers.md: meta-ai in the first-class
API-key provider list, META_BASE_URL override, contributor-tier
data-training note
Follow-ups on top of the salvaged CommandCode provider plugin (PR #32909):
- hermes_cli/config_defaults.py: COMMANDCODE_API_KEY setup-wizard entry
- hermes_cli/doctor.py: add key to the doctor env-var scan list
(health check comes free via the pluggable-profile loop)
- hermes_cli/dump.py: include commandcode in debug-dump api_keys
- docs: provider table row, fallback-provider table + supported lists
- tests: doctor dedicated-skip test now uses exact-name checks so
Bearer-authed Anthropic-COMPATIBLE gateways (CommandCode (Anthropic))
are allowed in the generic loop while native anthropic stays skipped
E2E verified with real imports: profile registration, aliases,
PROVIDER_REGISTRY auto-extension, bearer-auth host match
(positive + negative), live /models fetch (55 models).
Follow-ups on the #85006 salvage:
- A key_cmd token with no advertised expiry was cached for the life of the
process. The "refresh on 401" contract it relied on has no implementation
(SDK retries cover 429/5xx only), so an expired no-TTL token would 401
every request until restart. Cache on a bounded 15-minute window instead;
helpers that want a longer cache can advertise their real expiry.
- Test for the no-TTL path updated to pin the bounded-window contract;
the remint test's $RANDOM (bash-only, empty under dash) replaced with
date +%s%N so it exercises remint under any /bin/sh.
- website/docs/integrations/providers.md: document key_cmd in the named
custom providers section (contract, precedence, secrets.command contrast).
- optional-skills/devops/actual-setup: field-tested setup skill contributed
by shl0ms, updated for the first-class 'actual' provider (the original
targeted a custom-provider config that now collides with the built-in name)
- docs: providers.md section + tables, environment-variables.md, quickstart.md
- tests/skills: frontmatter + first-class-provider conformance checks
Users with Claude Pro/Max, ChatGPT/Codex, SuperGrok or Gemini plans
cannot find what their plan pays for in Hermes in one place (e.g.
#15291, #27228). One comparison table + per-provider notes; cells the
docs do not yet specify are marked 'not currently documented' rather
than guessed.
Fixes#67278. Since config v12 (see #8776), custom_providers: list is
legacy — the migration converts it to the providers: dict and the
resolver reads the dict first (runtime_provider.py). Docs still taught
the legacy list as primary.
- integrations/providers.md: all 8 YAML examples converted to
providers: dict (field mapping code-verified: api, default_model,
transport), one consolidated legacy-format note
- configuring-models.md, configuration.md, credential-pools.md,
migrate-from-openclaw.md, provider-runtime.md, faq.md: examples and
prose flipped to dict-first; legacy list noted as still-read
- adding-providers.md untouched (sole mention is a literal test
filename)
One page consolidating all three Hermes×Buzz integration paths —
Desktop managed runtime, buzz-acp relay bridge, and the native gateway
platform — with a comparison table, per-path pointers into the detailed
docs, identity guidance, and contributor credits. Registered in
sidebars.ts under Integrations; Buzz added to the messaging platform
list and a new Collaboration Workspaces section on the integrations
index. en + zh-Hans.
The Nous Portal docs claimed routing 'happens through OpenRouter under
the hood' with OpenRouter-equivalent failover, and that the catalog
'mirrors OpenRouter's model list'. That is not the Portal's contract:
some models route through OpenRouter, others through proprietary or
secondary providers, and per-model routing can change over time.
The stale wording licensed users to expect OpenRouter-proprietary
request extensions (top-level cache_control, session_id sticky
routing, provider preferences) to work through the Portal, producing
misfiled bug reports like #71576. Reworded both pages (en + zh-Hans)
and added an explicit note that OpenRouter-specific extensions are not
part of the Portal API contract.
Removes Homebrew and PyPI wheel/sdist as Hermes distribution paths while
preserving the supported source, Docker, and Nix workflows.
Changes:
- Removes the Homebrew formula, PyPI publish workflow, sdist manifest
(MANIFEST.in), and wheel/sdist release-attachment logic from scripts/release.py.
- Keeps setuptools metadata and entry points required by editable installs
and Docker/Nix builds, but adds a setup.py guard that rejects wheel/sdist
builds outside a sealed Nix derivation (HERMES_NIX_BUILD=1).
- Removes pip/Homebrew install detection, PyPI update checks, the pip
self-update path, the deprecation-banner state, the postinstall subcommand,
wheel data-directory fallbacks in agent/i18n.py and hermes_constants.py,
and the ACP Registry manifest/version-lockstep release logic.
- Adds /nix/store/ path detection so `nix run` / `nix profile install`
installs (which don't set HERMES_MANAGED) are correctly identified as
"nix" rather than falling through to "git"/"unknown".
- Retired install-method values ("pip", "homebrew") in existing
.install_method stamps (both code-scoped and home-scoped) are ignored by
the allowlist reader and fall through to "unknown" instead of resurrecting
a retired enum value.
- Updates Nix packaging to ship bare runtime data (locales, optional-mcps)
through store symlinks and wrapper env vars instead of wheel data-files.
- Removes the ACP Registry manifest/icon and their version-lockstep tests.
- Deletes or rewrites packaging, pip-update, Homebrew, and ACP Registry
tests; adds parametrized coverage for the packaging build guard covering
BOTH sdist and wheel paths (the guards live in separate cmdclass entries
— a passing sdist test proves nothing about the wheel path).
- Updates installation/platform documentation and related user-facing copy.
- Adjusts the supply-chain scan so deleted install-hook files do not trigger
a finding, while additions or modifications still require the existing
ci-reviewed label gate.
Supported installation paths (unchanged):
- git installer (install.sh)
- Docker
- Nix/NixOS
- editable development installs (uv sync, uv pip install -e ., pip install -e .)
Cross-checked website/docs against the source at main HEAD and corrected
documented commands, env vars, config keys, headers, and default values
that don't match the code. Docs-only; no behavioral changes.
Refs #36048
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* docs(secrets): secret-source plugin developer guide + sidebar registration for 1Password page
- New developer-guide/secret-source-plugin.md: SecretSource contract
(never raises/prompts, fetch-only, timeout budget), framework-vs-plugin
ownership table, mapped-vs-bulk shape guidance, run_secret_cli()
subprocess-safety, registration + timing note, conformance kit usage,
ErrorKind reference.
- Register user-guide/secrets/onepassword in the sidebar (page shipped
in #59498 but was not listed, so it was unreachable from nav).
- Cross-link the user-guide plugin section to the new dev guide.
* docs: group all plugin guides under a Plugins subcategory in Extending
- Move guides/build-a-hermes-plugin.md -> developer-guide/plugins/index.md
(both locales) and make it the category landing page (slug pinned to
/developer-guide/plugins).
- New sidebar subcategory Developer Guide > Extending > Plugins holding
the general guide + all 8 provider-plugin docs (llm-access, memory,
context-engine, secret-source, model, image-gen, video-gen, web-search);
provider-doc URLs unchanged.
- Client redirect /guides/build-a-hermes-plugin -> /developer-guide/plugins.
- Update 30 cross-links across both locales.
Replace the loopback/PKCE-callback server and manual-paste fallback with
the RFC 8628 device-code flow as the only xAI Grok OAuth login path. The
flow works in headless/SSH/container sessions with no 127.0.0.1 listener,
shrinking the local attack surface.
- Poll the token endpoint with server-provided interval, honoring
slow_down and expires_in; store tokens with auth_mode
oauth_device_code.
- Adaptive proactive refresh skew for short-lived device-code JWTs;
rotated tokens sync back to auth.json, the global root store, and the
credential pool (no refresh-token replay).
- Clear source suppression on successful re-login (CLI + dashboard) and
drop the duplicate dashboard pool entry so exactly one seeded
device_code entry exists.
- Use the shared device_code source name for consistency with the
nous/codex device-code providers.
- Desktop: remove the loopback OAuth flow states and dead type variants;
pkce providers' sign-in URL selection is unchanged.
- Docs (EN + zh-Hans) rewritten for device-code login; drop the deleted
--manual-paste flag from documented commands.
Adds Vertex AI as a first-class provider for Gemini models via Vertex's
OpenAI-compatible endpoint. Vertex authenticates with short-lived OAuth2
access tokens (service-account JSON or ADC), not a static API key — the
missing piece behind the recurring requests (#13484, #12639, #56259).
- agent/vertex_adapter.py: OAuth2 token minting + refresh-on-expiry
(5-min margin), ADC->service-account fallback, global vs regional
endpoint URLs. Config precedence: env var > config.yaml > default.
- plugins/model-providers/vertex/: provider profile (auth_type=vertex),
reuses Gemini's extra_body.google.thinking_config translation.
- runtime_provider: vertex short-circuit BEFORE the credential pool so a
credentials-file path is never mistaken for a static API key; mints a
fresh token + computes base_url per resolve.
- run_agent + conversation_loop: _try_refresh_vertex_client_credentials()
re-mints the token and rebuilds the client on a mid-session 401, so a
long-lived gateway agent survives token expiry (~1h).
- auxiliary_client: vertex auth_type branch for side-LLM tasks.
- config.yaml: vertex.project_id / vertex.region (non-secret, bridged to
env); credential path stays in .env (VERTEX_CREDENTIALS_PATH).
- setup wizard + model picker: dedicated _model_flow_vertex; curated
google/gemini-* model list; --provider choices.
- pricing/metadata: Vertex prices off the gemini docs snapshot; endpoint
host auto-maps to the vertex provider (no probe spam).
- lazy_deps + pyproject [vertex] extra: google-auth, opt-in only.
- docs: guides/google-vertex.md + providers page; tests for adapter +
runtime resolution.
Salvages and modernizes #8427 by @slawt onto current main: rewired from
the legacy PROVIDER_REGISTRY path to the provider-profile architecture,
moved non-secret config out of .env into config.yaml, and added the
per-turn 401 token-refresh the original lacked.
* feat(providers): remove google-gemini-cli + google-antigravity OAuth providers
Google now actively bans accounts for third-party tools that piggyback on
Gemini CLI / Antigravity / Code Assist OAuth, and because abuse prevention
sits at a backend layer the ban can extend to the entire Google account
(Gmail/Drive), with a second violation being permanent.
Ref: https://github.com/google-gemini/gemini-cli/discussions/20632
Removes both OAuth inference providers entirely (modules, provider profiles,
auth/runtime/config/models wiring, the /gquota Code Assist quota command,
the antigravity-cli optional skill, desktop + docs surface in en + zh-Hans).
The API-key 'gemini' provider (GOOGLE_API_KEY/GEMINI_API_KEY against
generativelanguage.googleapis.com) is unaffected and stays fully supported.
* fix(skills): keep the antigravity-cli skill — only the OAuth provider is removed
The antigravity-cli optional skill orchestrates the external `agy` binary as
a coding-agent tool via the terminal tool — it does NOT wrap Hermes inference
through the banned google-antigravity OAuth provider, so it carries none of
the account-ban risk that motivated removing that provider. Restore the skill,
its docs page, the sidebar entry, and the optional-skills catalog row. The
google-antigravity / google-gemini-cli inference providers stay fully removed.