Use Perplexity for Nous-managed search, retaining Firecrawl for extract
and as a per-call search fallback. Explicit search overrides and direct
keys keep their own billing paths; fallback results are never cached.
When the managed route is selected but the Tool Gateway is unavailable
(unentitled account or no Nous token), search reports that selection
error instead of asking for a direct key the user never chose.
The managed search vendor is unannounced, so user-facing copy names the
capability rather than the vendor: status, portal and docs say "managed
web search", and the fallback annotation reads `managed_primary`. Direct-key
configuration docs are unchanged.
Routing, auth, payload, cache and entitlement regressions are covered
through real config loading and local HTTP.
memory-providers.md points the Hindsight section at the catalog entry and the
upstream integration docs, documents the auto-install on `hermes update` /
first agent start (and the allow_lazy_installs=false one-liner), and adds a
"Migrating from bundled Hindsight" subsection. memory.md, overview.md,
integrations/index.md, cli-commands.md, docker.md and nix-setup.md no longer
say hindsight ships in-tree or as the [hindsight] extra.
f50d4b3423 (a docs revert) restored a route-style /user-guide/profiles link in
nous-portal.md (en + zh-Hans); website/scripts/check_doc_links.py rejects those, so
the Docs Site check has been red on main since. Rewritten with the script's --fix.
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.
The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.
Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
before req_reasoning: {}
after req_reasoning: {'reasoning_effort': 'medium'}
agent.reasoning_effort: low -> {'reasoning_effort': 'low'} (unchanged)
A Codex OAuth model (gpt-6-astra, gpt-5.6-sol/terra/luna, gpt-5.5, ...) served
through a proxy — a custom provider with api_mode: codex_responses at
http://127.0.0.1:8317/v1, or the openai-codex provider behind
HERMES_CODEX_BASE_URL / model.base_url — resolved 1,050,000 from
DEFAULT_CONTEXT_LENGTHS instead of the 272,000 Codex enforces. The compressor
then set its trigger at 525,000 tokens, ~2x past the Codex limit and its 272K
billing tier, and a day of Astra sessions burned two Pro weekly windows.
Root cause: step 2 of get_model_context_length treated any non-known hostname
as a generic custom endpoint and never looked at the route's transport, so the
Codex table two steps away was skipped. The decision key is now the route:
_is_codex_route() is true for provider openai-codex on any URL and for a custom
entry whose canonical api_mode is codex_responses. A custom Codex route answers
from the Codex OAuth table first (with the opted-in -900k bump; unknown slugs
still take the endpoint probes), the native provider skips the custom-endpoint
step and keeps its live catalog probe in step 5, and both bypass the persistent
cache so a stale 1.05M entry cannot outlive the fix. Explicit
providers.<name>.models[].context_length / context_length / model.context_length
overrides still win (step 0 runs first); chat_completions proxies are unchanged.
Reported with the exact resolver path by @0xble (#116199); route-keyed lookup
per @JoaoMarcos44 (#116262).
Fixes#116191
Providers page (Codex note), CLI reference, credential-pools command table and
the OAuth-over-SSH port table, so the fixed :1455 listener and its device-code
fallback are discoverable where users look for Codex login help.
model.context_length short-circuits get_model_context_length ahead of every
provider source, but nothing told the user the number they saw was their own
pin rather than provider metadata (#66168). Keep the pin semantics (custom
endpoints depend on it) and make it visible instead:
- agent/context_pin.py: is_context_pinned / context_pin_suffix for renderers,
and warn_once_on_pin_disagreement, called from _resolve_context_length at
agent init. The advertised value comes from LOCAL sources only (endpoint
cache, models.dev disk cache, hardcoded catalog) so the check never adds a
network probe or latency to startup; one warning per (model, pin) per process.
- "(pinned)" label on every CLI surface that renders the window: welcome
banner, /model switch summary, /usage current-context line, wide status bar.
- No new config key or env var; the numeric value used is unchanged.
Fixes#66168
When an OpenAI-compatible local server (llama.cpp, vLLM, ...) serves a
window below MINIMUM_CONTEXT_LENGTH, the startup refusal, the CLI banner
warning and the num_ctx log line were worded for Ollama (/api/show,
model.context_length advice that cannot raise a served window). Name the
served window and the server-agnostic remedies: start the server with a
>=64K context or set model.ollama_num_ctx (honoured on every local
endpoint) to the window it really serves. Hosted routes keep the
model.context_length advice. The floor itself is unchanged.
Part of #87075
The providers page listed every OpenCode caller that carries the affinity header; add the
stateless one-shot case (Desktop commit-message generation with no active session), which
now sends a fresh ephemeral key instead of nothing so the relay no longer rejects it with
400 MissingSessionID (#105841).
`resolve_provider_client("opencode-go", model="gpt-5.6-luna")` built a plain
Chat Completions client, so every auxiliary call (compression, titles,
vision, MoA) on that model went to /chat/completions and OpenCode answered
500 — while the main conversation on the same provider+model worked, because
hermes_cli/runtime_provider re-derives api_mode per model from
opencode_model_api_mode() and the auxiliary path never did.
`agent/opencode_affinity.py::opencode_transport()` is the single per-model
(api_mode, base_url) decision for OpenCode relay targets (built-in families,
`opencode-go-*` custom entries, opencode.ai hosts). `_wrap_transport` — the
transport chokepoint every resolve branch ends in — and the named-custom
branch (whose entry may have persisted the api_mode of whichever model was
selected at save time) consult it, so Responses-only models get
CodexAuxiliaryClient and Anthropic-wire models land on /v1/messages with the
/v1-stripped relay URL. A stale task/provider-level api_mode is ignored for
these targets, exactly like the main runtime.
Live: fake relay — before: OpenAI client → POST /v1/chat/completions → 500;
after: CodexAuxiliaryClient → POST /v1/responses. Control: glm-5 stays a
plain chat client on both.
Fixes#98799
Co-authored-by: Tranquil-Flow <66773372+Tranquil-Flow@users.noreply.github.com>
A fallback_providers entry naming a `providers.<name>` block (or
`custom:<name>`) without its own api_mode was re-detected from the
resolved host: an Anthropic-Messages proxy on a plain host or a
Responses-only relay behind a generic gateway landed on chat_completions
while resolve_provider_client had already built the declared client
(#33062 bottom thread, #81932 provider-level `transport:`). The hint pass
now reads the named block's api_mode/transport as explicit, and entry-level
`transport:` is accepted as an alias of `api_mode` with the same
canonicalisation the providers block uses (`responses` -> codex_responses).
Fixes#33062Fixes#81932
opencode_session_headers derives the key via resolve_affinity_key() (#115904) and keeps the
#115830 oneshot-<uuid> fallback so an OpenCode target never gets {} (OpenCode Go 400 MissingSessionID).
is_output_cap_error already classified "max_completion_tokens is limited to
16384 for glm-5.2" as an output cap, but parse_available_output_tokens_from_error
returned None, so _recover_context_length failed fast instead of clamping.
Add the ("limited to",) signal and the direct-cap regex; the Scaleway string
now yields 16384. Also document model.key_env in the custom provider block.
Trim of the cherry-picked #104566 so it clears the salvage bar and covers #86241:
- The config getter matches entries the way every sibling per-provider knob does
(`_entries_for_route`, route identity only) instead of a second provider-name axis;
returns "" like `get_custom_provider_extra_headers` returns {}.
- `merge_opencode_session_headers` becomes `merge_session_affinity_headers` at both
call sites (main `build_api_kwargs`, auxiliary `_build_call_kwargs`); the alias line
is gone. Both sources merge (an OpenCode target that also declares a header gets both).
- Tests cut from six to two invariants: configured header carries one value per
conversation on chat_completions, anthropic_messages and auxiliary kwargs (different
for another session, caller-pinned wins); unconfigured → no header on any path.
- Docs: `configuring-models.md` per-provider options, `providers.md` entry key list,
`cli-config.yaml.example` — the key is opt-in, default off, so DEFAULT_CONFIG is unchanged.
Why: a session-aware proxy classifies a request with no session id whose last message is a
tool_result as a NEW conversation and re-sends the whole history upstream (cache_write ≈
cache_read). Hermes already derives a rotation-stable conversation key for OpenCode; naming
the header per provider lets any proxy receive it without shipping an identifier by default.
Co-authored-by: 0xAlyDev <agentai891@gmail.com>
The classic CLI (`hermes --resume`, mid-chat `/resume`, oneshot resume) read the
session row's persisted api_mode/base_url verbatim in stored_session_route(), so
a row written while the session ran an anthropic_messages model on opencode-go
(MiniMax) pinned that wire onto a chat_completions model such as
deepseek-v4-flash-vision-exp after the model column moved. This is the CLI twin
of the tui_gateway _rederive_per_model_route() fix already on this branch: for
providers that pick the wire per model (model_derived_api_mode() is not None)
the route follows the stored model and the relay URL is healed; fixed-wire
providers keep honoring their row.
Probe (direct call of _restore_session_model on the PR head, no network): a row
with provider opencode-go, base_url '' and api_mode anthropic_messages for
deepseek-v4-flash-vision-exp resumed on the Go relay (never api.anthropic.com —
the empty base_url is re-resolved to https://opencode.ai/zen/go/v1 on both the
same-provider and provider-changed paths) but with api_mode anthropic_messages;
after: chat_completions on the same relay.
Docs: the OpenCode paragraph in providers.md now states both behaviours users
can observe from #96066 — per-model routing survives resume, and `*-vision*`
OpenCode ids attach images natively without a supports_vision override.
Part of #96066
The providers page listed every OpenCode caller that carries the affinity header; add the
stateless one-shot case (Desktop commit-message generation with no active session), which
now sends a fresh ephemeral key instead of nothing so the relay no longer rejects it with
400 MissingSessionID (#105841).
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.
- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
current users). When false:
- `read_claude_code_credentials()` — the only reader of the borrowed Claude
Code login — returns None, so the resolver fallback, the expired-token
refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
file; the pool prunes a `claude_code` row an earlier adopting process
persisted.
- `_recover_codex_tokens_from_cli` returns None for both automatic recovery
paths (rejected refresh, half-empty singleton); the real AuthError is
surfaced instead. The interactive import offer in `hermes auth add
openai-codex` still asks first and is unaffected.
- One INFO line per process the first time adoption would have happened;
`hermes auth list` / `hermes auth status anthropic|openai-codex` print the
same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.
Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
`python3 website/scripts/check_doc_links.py --fix` over the current tree: 126
route-style links in 9 pages (the six that conflicted with #114784/#114806/
#114851 plus google-gemini, cron and secrets) rewritten to relative file paths.
Check mode is clean afterwards.
Mechanical `check_doc_links.py --fix` pass over website/docs (hand-authored
and generated pages) and the zh-Hans mirror: 1,868 route-style links
(`](/section/page#anchor)`, `](/docs/...)`) become `](../section/page.md#anchor)`.
Every target was asserted to exist on disk; anchors and query strings are
preserved; fenced code blocks and inline-code examples are untouched.
Two dead targets found by the converter were fixed by hand first:
memory-providers.md linked `/user-guide/plugins` (page is
`user-guide/features/plugins`), and the zh-Hans learning-path still linked the
removed `rl-training` page — ported the EN treatment (external Atropos link).
Docusaurus build after: EN locale 0 unresolved Markdown links, 0 broken links,
0 broken anchors.
The salvaged paragraph said a never-signed-in profile "asks you to set one up
(`hermes -p <name> portal`)"; the actual boot error is
`Profile '<name>' is not connected to any AI provider yet` (hermes_cli/auth.py::
resolve_provider) and points at `hermes -p <name> model`. Quote the real
message and state the supported remediation separately.
Also record the two facts the runtime does implement, so the page is not only a
retraction: `hermes -p <name> portal` (alias of `auth add nous --type oauth`) offers a
one-tap import of <hermes-root>/shared/nous_auth.json
(hermes_cli/auth_commands.py::_add_nous_oauth_credential), and `profile create
--clone-all` copies auth.json with the Nous login intact (only
SINGLE_USE_REFRESH_POOL_PROVIDERS grants are stripped, hermes_cli/profiles.py::_clone_all_into).
guides/run-hermes-with-nous-portal.md repeated the same "pick it up automatically"
promise; align it and link to the canonical section with a pinned {#profile-setup}
anchor so the zh-Hans heading slugs the same. EN + zh-Hans in sync.
The Portal "Profile setup" section promised that every profile picks up a
Portal login automatically. The shared token store only merges fresher
tokens into a profile that already has a Nous state; a profile with no
providers.nous entry of its own fails closed at boot and is asked to run
the setup flow (hermes_cli/auth.py:1545, auth_nous.py:1056), because
profiles are independent islands (#111724). Correct the promise to
"sign in once per profile" and describe the fail-closed behavior, in both
the English and zh-Hans copies.
Fixes#114259
_CodexCompletionsAdapter._build_responses_kwargs (title generation, compression,
MoA aggregation) built its own chat->Responses tool list: raw names, no
``strict: False``. On api.perplexity.ai and opencode.ai routes the main loop
already aliased reserved names (hermes_<name>), so every auxiliary call 400ed
while chat looked healthy and MoA silently fell back.
- tools go through codex_responses_adapter._responses_tools (strict: False) and
agent/transports/codex.py::_alias_wire_tools — one reserved-name table for
both paths (OpenCode, Perplexity, xAI)
- replayed history tool_calls are renamed to the alias this request declares
- the alias map rides on the request payload (``_wire_aliases``), popped in
create() and mapped back onto the parsed tool_calls before dispatch; never
instance state (aux adapters are cached and shared)
Ported from #114457's auxiliary half; design per the #114260 report.
Co-authored-by: Tyler Lyon <lyonrt@icloud.com>
The salvaged custom-profile declaration (#114255) gives every ``custom:<name>``
Responses route the OpenAI-compat vocabulary, which closes#114249 but has two
edges the transport must keep:
- ``_profile_declared_efforts`` resolved by provider NAME first, so the new
non-None custom declaration short-circuited the host lookup: a
``custom:my-proxy`` entry pointed at api.router.com stopped inheriting the
Router catalog clamp and would send ``max`` to a gateway that 400s on it.
Resolve by endpoint host first, then by name — the host is the endpoint's
truth; the config-entry name is only a label (the existing Router test now
uses the runtime's real ``custom:<name>`` identity, which is what exposed it).
- A custom entry that merely points at api.openai.com (host-mandated
codex_responses) is still OpenAI: its per-model ladder is known, so skip
profile declarations on the official origin and the Codex backend.
``_is_openai_api_origin`` is the shared exact-host check;
``_is_official_openai_responses_route`` reuses it.
Docs: providers.md states the custom-endpoint effort contract and both
host-following exceptions.
Live: custom:relay deepseek-flash max -> max (was xhigh); custom:oai @
api.openai.com gpt-5.2 max -> xhigh; openai gpt-5.2 -> xhigh, gpt-5.6 -> max;
custom:my-proxy @ api.router.com grok-4.6 max -> xhigh (pick-only: max).
Follow-up to the salvaged #114614 pick:
- `_codex_login_post`: the `for` loop ended in an unreachable `raise … # pragma: no cover`
(needed only to satisfy the return type). A `while True` with the terminal condition
folded into the except branch has no dead path and the same three-attempt bound.
- `_codex_poll_authorization_code`: the pick duplicated the `except KeyboardInterrupt`
handler; the second copy was unreachable.
- Tests: six change-detector tests collapsed into two invariants (poll survives blips but
never retries a non-transport exception; one-shot POST retries once and keeps the typed
AuthError + TLS hint + cause chain at the cap). The existing
test_codex_device_login_ssl_hint.py still pins the poll's terminal hint path.
- Docs: providers.md notes that a single dropped connection during device login is no
longer fatal.
Remove the keyless free tier from the provider table (providers.md,
fallback-providers.md), the .env.example section, and the 'three OpenCode
providers' wording now that only Zen and Go ship built-in.
Closes the remaining atoms of #112600.
A) The CLI startup path already passes `-m` as target_model (c358a6fba0), but
its siblings still resolved credentials against config's `default`: the CLI
auth-fallback rung, `--resume` credential re-resolution, the gateway
provider-override helper (channel overrides, persisted /model switches,
API-server provider refresh), the gateway fallback chain, the TUI /model
switch-from runtime and ACP agent construction. With a `*-free` default the
OpenCode free-tier rung fired first and a Go-only model was built against
the keyless Zen relay ("Model mimo-v2.5 is not supported"). Each now passes
the effective model; `_resolve_runtime_agent_kwargs_for_provider` grows an
optional `target_model` and the two test stubs of it accept the kwarg.
B) normalize_opencode_base_url rewrote the /zen vs /zen/go segment for ANY
provider matched by opencode_provider_family, including custom providers
merely named after a family (`opencode-go-bridge`, #85589) whose relay the
user declared explicitly in `providers:`. The family heal now applies to the
built-in canonical providers only; custom prefix-named providers keep their
per-model api_mode routing and /v1 handling. Documented in the providers
guide.
C) Same function: the official-host check uses parsed.hostname (a port no
longer defeats the heal) and only the path is edited, so query/fragment
round-trip instead of being dropped.
Fixes#112600
A metadata-only model_overrides entry on a custom provider (only context_window set) took the
unknown-model branch of get_model_capabilities and reported max_output_tokens=8192 from the
_UNKNOWN_MODEL_BASE template; /model info showed "Max output: 8,192 tokens" for a model nobody
knows. The output limit is now left unknown (ModelCapabilities.max_output_tokens=None,
ModelInfo.max_output=0), matching the fail-open treatment #113146 gave supports_vision and
supports_reasoning. Consumers already handle a missing value: the dashboard only renders a
positive limit and the /model card only prints a truthy max_output.
There was also no way for a custom provider that resells a catalogued vendor's models (a
gateway/proxy) to reach _BUILTIN_MODEL_METADATA or the models.dev catalog, which is what forced
the metadata-only override in the first place. providers.<name>.catalog_provider (also honoured
on legacy custom_providers[] rows) names that vendor; models_dev resolves the catalog id through
it for capabilities, context, model info and provider info. It affects metadata lookups only,
never routing or credentials, and an explicit model_overrides entry still wins.
Closes the remaining atoms of #112649 (docs: providers page).
Follow-up on the salvaged #112721 commits (@fangliquanflq):
- agent/turn_context.py, tui_gateway/methods_prompt.py: add ``session_id`` to the explicit
snapshot dicts instead of switching them to ``agent._current_main_runtime()``. The titling
prologue is duck-typed (its tests drive a minimal stub) and the snapshot values feed the
``runtime_validator`` equality checks, where ``_current_main_runtime()``'s ``"" `` for a
missing attribute would no longer match the live ``None``. Same outcome — the background
request inherits the conversation's ``x-opencode-session`` — without changing what the
callers read.
- tests/agent/test_opencode_session_affinity.py: the salvaged title test passed on an
unfixed tree because the titler thread republishes the conversation contextvar
(``set_conversation_context``) and the affinity header falls back to it. Replace it with
two invariant tests (sync + async ``call_llm(main_runtime=...)`` with EVERY ambient source
unset via the ``out_of_turn`` fixture): red on origin/main, green here; both also pin
that the explicit binding does not leak past the call.
- website/docs/integrations/providers.md: name the background/out-of-turn auxiliary calls
the header now covers.
Fixes#112717
The config_defaults comment still described the unknown-model template as
"vision/reasoning off", which is exactly the synthesized-False behaviour
#112649 reports and the preceding commit removes. State the new contract
where users read it (config_defaults, providers docs): a context_window-only
override for an uncatalogued model keeps supports_vision / supports_reasoning
unknown so vision_analyze, video_analyze and the reasoning-effort picker stay
available; only an explicit false in the override marks the model text-only.
The dashboard's estimate endpoints make the same headless auxiliary call
as specify/decompose but never bound an affinity scope, so they still sent
no x-opencode-session and the OpenCode Go relay answered 400
MissingSessionID (#112043). Declare kanban:<task_id> for an existing task
and a stable kanban:estimate key for the create dialog (no task yet),
unless a scope is already bound.
Test: _run_estimate captured header None before; now kanban:t_1 /
kanban:estimate and nothing leaks past the call.
Follow-up to the cherry-picked #111876 (@KoNit-K), which shares the design
of the earlier #111875 by the issue author (@ats3v): emit DeepInfra's
top-level ``reasoning_effort`` from the provider profile, ungated on
``supports_reasoning``, ``none`` as the only off switch, ``xhigh`` native,
``ultra`` clamped to ``max`` via the shared vocabulary, unset/unknown omitted.
- drop the constructor/blank-line reformat churn (byte-identical to main)
- replace the 14-case test file with two invariant tests: the profile's
config -> top-level field table, and the transport main-turn path with
``supports_reasoning=False`` (the gate the core allowlist actually passes)
- docs: DeepInfra subsection in integrations/providers.md describing the
two-directional reasoning control
Offline kwargs probe: before every reasoning_config -> ({}, {}) and the
main turn carried no reasoning field; after ``high`` -> ``reasoning_effort:
high``, ``{'enabled': False}`` -> ``none``, ``ultra`` -> ``max``, unset and
unknown levels omitted, aux calls stop emitting the generic
``extra_body.reasoning`` for this provider.
Co-authored-by: Georgi Atsev <georgi@deepinfra.com>
Two invariant tests (red on main): the PKCE key lands as an api_key pool row that
resolve_provider("auto") picks up while the bare --api-key path keeps its default, and a forged
callback path is a 404 while the genuine nonce path yields the code. evals/openrouter_pkce_ab
drives the real auth_add_command against a local fake /api/v1/auth/keys (verifier check,
single-use codes) for legit / wrong-state / replayed-code / malformed-response / api-key-path.
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
__cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
SSLEOFError whose text httpx did not repeat still gets the hint. The
hint now names the TLS 1.2 diagnostic and links the providers docs
note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
classic-groups snippet (EN + existing zh-Hans copy).
Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
Four writers each dropped their own uniquely-named copy of config.yaml next to
the real file and none of them ever deleted anything: hermes setup
(config.yaml.bak.YYYYMMDD_HHMMSS, one per run even with no change), the
corrupt-YAML snapshot (config.yaml.corrupt.<ts>.bak), hermes migrate xai
(config.yaml.bak-pre-migrate-xai-<ts>) and the Docker boot migration
(config.yaml.bak-<ts>, .env.bak-<ts>). A home dir accumulated a dozen variants
with no way to tell which mattered.
hermes_cli/config_backups.py::backup_config is now the single writer:
backups/config/config.yaml.<reason>.<YYYYMMDD-HHMMSS>, skipped when the newest
copy for that reason is byte-identical, rotated to the newest five per reason.
backups/ is already excluded from full backups so nothing nests. Legacy
siblings written by the old schemes are moved into the dir on first use;
hand-named copies (config.yaml.bak-my-note) are left alone.
Live: three `hermes setup --non-interactive` runs against an unchanged config
went from three .bak files in HERMES_HOME to one pre-setup copy under
backups/config/; repeated loads of broken YAML produce one corrupt copy
instead of one per process (deduped by content).
Extend #104477 to the native thinking, vision, metadata, and local header paths identified by #87641. Materialize only at probe boundaries; leave the chat callable and cache ownership untouched. Local-wire A/B: thinking and vision show requests change from 403 to 200, while static credentials and callable chat retain success. Target suites queued.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Port owned-before-borrowed resolution from #104624, crediting the root cause in #104622. Include the synchronous and asynchronous auxiliary fallback recovery sites: forward the failed request key so an unrelated borrowed login never owns that refresh. Local-wire probes preserve the borrowed file and exchange only the owned grant. Broader validation remains queued; do not treat this commit as ready.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: d-bow-dev <24577047+dwb1991@users.noreply.github.com>