Local OpenAI-compatible server aliases (local, vllm, llamacpp, llama.cpp,
llama-cpp) were normalized inconsistently across the three provider-name
tables: hermes_cli.providers mapped them to the orphan "local" id,
hermes_cli.models left them unmapped, while hermes_cli.auth mapped them to
"custom". Align all three on the generic "custom" provider so routing,
the model picker, and credential resolution agree, matching what
resolve_provider already did for vllm/llamacpp. Add a parity contract test.
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
`_credential_token_pair` already maps a non-dict row to (None, None), so the
separate isinstance guard and second early return in
`_merge_pool_row_generation` were the same branch. Adopting a peer generation
now goes through the existing `_replace_entry` swap primitive.
The id->token-pair base map was built three ways: _persist dropped
(None, None) pairs, load_pool and _sync_entry_from_pool_store kept them.
_merge_pool_row_generation treats a missing base as "unknown" (plain
recency merge) but a (None, None) base as a known generation, so a
token-less row got the generation override on the first flush after
load_pool and the plain merge on every later one.
Policy chosen: KEEP blank bases everywhere (new auth_mod._token_pairs_by_id,
built on the existing _entry_ids). A blank base means "no pair when we last
looked", which is exactly the CAS witness the boundary needs: when a peer
lands a pair on that row, every flush of ours keeps the peer's generation
instead of writing our blank tokens back. Dropping blank bases would make
the first flush keep the peer's pair and later flushes overwrite it.
_sync_entry_from_pool_store now records the base before its no-token-
material bail-out so all three sites agree. _update_root_pool_rows reuses
_entry_ids for its incoming map as well.
Also collapse the three copy-from-disk-else-pop loops in
_merge_pool_row_generation into a local _take_from_disk helper.
failure_reason is passed alongside _POOL_STATUS_FIELDS rather than added
to the shared constant, which auth_codex/auth_oauth_grants and
_merge_disk_cooldown_state also consume.
When a stale pool's persist finds the disk pair moved, #120943 copied every
status field from disk, so the stale writer's NEWER account-wide verdict
(402 billing / 429 throttle -> EXHAUSTED) was erased and the rotated pair
re-entered selection immediately. Only a terminal auth death is tied to the
pair it was observed on: discard the stale writer's DEAD onto the peer's
pair, but leave any other status for _merge_disk_cooldown_state's ordinary
recency merge.
Also travel `scope` and `inference_base_url` with the pair (a Nous refresh
rewrites them together with the tokens, so a stale writer must not stamp an
old scope/route onto the newer pair); use the always-initialised
_persisted_token_pairs directly; and re-hydrate an in-memory entry from the
written row only when the store overrode its pair, so unchanged rows keep
runtime-only fields to_dict() omits.
Tests: fold the stale-later-402 witness (rt-1 kept AND last_status ==
exhausted) into the stale-terminal-verdict test and drop the separate 429
rollback test it subsumes.
Fixes#120815
Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
The parse-failure banner, the active-failure record, FailedConfigRead and the
write refusals are one topic; config.py had grown past 4,100 lines with this
PR. Pure move (no behaviour change): config.py drops to 3,999 lines, below
origin/main. Callers outside config.py now import from the new module.
The principal-keyed fingerprint runs resolve_codex_runtime_credentials(read_only=True)
on every cache-only picker read, so the pool-exhausted branch must not probe the usage
endpoint or clear cooldowns from that path: one `not read_only and` guard, kept with
its test. The preserve_corrupt flag threaded through six auth loaders only suppressed
the one-time .json.corrupt sidecar copy on an unparseable store — an edge case the
same read-only resolve already hits on main via _codex_catalog — so it goes.
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.
Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.
- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
Twenty-five call sites told users to run
`python -c "from pm import sync_venv; sync_venv(['x'], explicit=True)"`
because `hermes pm install` only took package names. Add `--extra`
(repeatable; syncs the venv with the named extras and nothing else) and
pm.install_hint(extra), the single builder every site now uses, so the
advice stays correct when the command changes.
A cold PM runtime under allow_lazy_installs:false now reports the extra
the caller wanted and the command that provisions both, instead of a
bare "pm-runtime: not installed".
Reverts 93889b770d ("named profiles no longer inherit the root
profile's auth.json"). After the Desktop update every bot profile that had
relied on the root OpenAI Codex login failed with "No Codex credentials
stored. Run `hermes -p <bot> auth add openai-codex --type oauth`", and users
had to re-run the device-code flow once per bot (5-6 times in the field
report). Sharing one grant across profiles is the intended design: OAuth
refresh tokens are single-use, so ONE grant lives at the root, profiles
resolve it read-only, and a refresh under a profile writes the rotated chain
back to root (Codex / xAI write-through, borrowed-row pool bookkeeping,
forked-grant heal) — never a per-profile copy.
Restored: `_global_auth_file_path` / `_load_global_auth_store` fallback in
`_load_provider_state*` / `read_credential_pool` / `_provider_state_transaction`,
Codex + xAI root write-through, `credential_pool` borrowed-root persistence,
`heal_forked_single_use_oauth_grants`, `share_auth` on profile creation
(Desktop create dialog checkbox), and the docs. `profile_credential_audit.py`
(the `hermes update` "profiles without a provider" notice) is removed with it.
Kept from after #111724: `_save_codex_tokens(set_active=...)` for image gen,
the plugin-auth `status` dispatch and the external-login notice in
`hermes auth list`, and the registry-derived env-var hint in agent_init.
Both rungs return (model_cfg, provider) verbatim, so a single
`provider == "openrouter" or provider in PROVIDER_REGISTRY` condition
reads as one intent: an explicit pin on a known aggregator/registry
provider. The 2-line WHY comment stays because openrouter's absence from
PROVIDER_REGISTRY is deliberate (#109397, #10622) and the merged
condition would otherwise look like a mistake.
Also move `base_url = ...` back down to just above `if base_url:`. The
hoist was only needed while the openrouter rung inspected base_url;
since f57fbfbb46 that rung honours mirrors unconditionally, so the early
read is dead and the assignment belongs next to its sole consumer.
The three picks for #109397 left two ~8-line comment blocks inside
_config_model_provider explaining the openrouter and bare-named-provider
rungs. Cut each to two lines and state both cases in the function docstring
instead, so the intent is visible where the function is documented rather
than buried mid-body. Comment/docstring only; no logic change.
A bare `model.provider` naming an enabled `providers:` / `custom_providers:` entry was
discarded by _config_model_provider(), even though the runtime resolver routes it
(has_named_custom_provider() -> True) and `hermes chat` works against the endpoint.
The two paths disagreed: free_tier_bootstrap._inventory_other_providers() recorded
provider_configured=False, so the dashboard parked sessions on Setup Required for a
working config — the same #109397 symptom one rung over from the openrouter pin.
Reuse the runtime's own lookup after the registry check. has_named_custom_provider()
is narrower than a raw providers-key check: it skips disabled entries and entries
without a usable endpoint, and covers the legacy custom_providers: spelling.
Regression test uses a real temp HERMES_HOME; red on the unfixed rung.
(cherry picked from commit 5567d685131f986bfe7c73485e906a22d7ff1151)
Review follow-up: the explicit model.provider: openrouter pin is provider
intent regardless of the base_url host, because the runtime treats a
non-openrouter base_url under the openrouter pin as a deliberate
mirror/proxy (runtime_provider_backends.py #10622, pinned by
test_explicit_openrouter_honors_config_base_url_mirror). Preserve that
intent instead of rejecting the mirror and parking the dashboard on
Setup Required.
Also correct the stale-base_url regression example: a non-loopback
base_url under a bare (unpinned) provider is the '#14676 stale leftover'
case; an openrouter pin is the one provider where a foreign host is a
legitimate mirror, not staleness.
Add the mirror URL to the bootstrap regression coverage and map the
contributor email for the attribution gate.
Refs #109397
(cherry picked from commit f59e511118861c168da7d5c64f4b5c8640de8b61)
_config_model_provider() treats a configured 'custom' endpoint as explicit
provider intent but not 'openrouter', which (like 'custom') is intentionally
absent from PROVIDER_REGISTRY. An openrouter-pinned install therefore read as
"nothing configured" in the free-tier boot inventory, parking dashboard
sessions on Setup Required while 'hermes chat' worked.
Recognise the openrouter pin the same way, honoring the existing stale-base_url
guard: only a canonical (empty or openrouter.ai) base_url counts; a non-openrouter
base_url under the openrouter pin stays contrived config and is rejected.
Fixes#109397
(cherry picked from commit 10e030812c5d1f6d4b44ec69d3b30c7b13b157e7)
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
get_env_value stops at the first environ hit, so a shell that exports
DEEPSEEK_API_KEY= (empty) hides a real key in .env from the gate while
resolve_api_key_provider_credentials() finds it: the picker omits a provider
the chat path would authenticate with. Read through
get_env_value_prefer_dotenv, the same chain auth.py already uses to resolve
the key, so the two can never disagree (#77007).
Co-authored-by: webtecnica <webtecnica@gmail.com>
os.getenv saw only the launch profile, so a DeepSeek key pasted in
another profile never appeared in Settings → Model until Refresh
models ran against that Bot's own backend.
- hermes_cli.auth.primary_failure_wording(exc) -> (log, user) phrase; reused by
cli_agent_setup_mixin._resolve_fallback_runtime, runtime_provider's fallback
logger and the TUI gateway/Desktop _resolve_runtime_with_fallback (#117482
sibling: 'Primary auth failed' for a 429 on the gateway surface).
- Drop the dead credentials_rate_limited kwarg at the post-turn exit site (the
flag is only True when _ensure_runtime_credentials returned False).
- Fold the new tests: 2 parametrized production-path tests + 1 gateway test.
load_pool("copilot") exchanged the gh CLI token on every load — on installs where copilot is
never selected (main provider deepseek, aux slots auto) that is a network round-trip plus the
"degraded to RAW token" warning on every pool load, ~2 per turn, 1000+ times in five days for
the reporter (#114740). The exchange result only matters once a model is routed to copilot, so
the seeder now keeps the raw token and skips the exchange (and the warning) until
is_provider_explicitly_configured("copilot") is true; the load that follows the user selecting
copilot re-seeds and exchanges as before. auxiliary.<task>.provider now counts as selecting a
provider in _config_selects_provider, like a MoA slot, so an aux slot pinned to copilot still
gets the exchanged token. Contributor tests trimmed to two invariants.
has_usable_secret() only knew the exact string "your_api_key_here", so every
other placeholder this repo ships counted as a configured credential:
- .env.example ships ``your_key_here`` (Xiaomi, Upstage, Ramp Router, Nebius)
and ``your_google_ai_studio_key_here`` / ``your_gemini_key_here`` /
``your_ollama_key_here`` (Google AI Studio, Gemini, Ollama);
- the quickstart and the MCP/skill references use sk-xxx / ghp_xxx / hf_xxx.
A copied-but-unfilled .env therefore resolved as a configured provider and the
placeholder went upstream, producing an opaque 401 instead of the "no usable
credentials — set X" the read point is required to report (AGENTS.md, "Fail
loud at integration boundaries").
Reject both shipped placeholder shapes, and apply the same gate in
_resolve_from_pool(), which handed a pooled placeholder straight to the client.
Tests cover every placeholder .env.example ships, the x-run convention, and a
pooled placeholder that must behave exactly like no credential at all.
A `$HERMES_HOME/plugins/model-providers/<name>/` plugin re-registering a
bundled provider (stepfun with a regional base_url, gmi at a staging host)
wins in `providers._REGISTRY` — register_provider() is last-writer-wins and
the plugin guide promises exactly this — but the runtime reads its endpoint
from `hermes_cli.auth.PROVIDER_REGISTRY`, whose mirror loop skipped every
name already present, so inference kept going to the built-in URL (#48450).
The mirror now applies one explicit precedence rule: when a row core wrote
(built-in or plugin-mirrored) belongs to a name whose profile
currently registered came from a USER plugin,
the row's profile-derived fields are rewritten in place (inference_base_url;
api_key_env_vars / base_url_env_var on api-key rows when the profile declares
env_vars). `providers` records the discovery source per registration
(`provider_source()`), because a bundled profile must never rewrite a
built-in row: several bundled profiles omit the row's `*_BASE_URL` env var
and one differs in auth_type, so an unconditional "profile wins" would have
changed built-in behaviour. With no user plugin PROVIDER_REGISTRY is
byte-identical before/after (78 rows probed). copilot/kimi/zai keep their
bespoke resolution via the existing skip set.
Co-authored-by: xiaoxinova <xiaoxinova@users.noreply.github.com>
Picker admission (#116552) listed out-of-tree external-process and OAuth
plugin providers, but selection and status still dispatched through
provider-name tables:
- `hermes model`: no `_PROVIDER_MODEL_FLOWS` entry and no generic flow, so
picking an admitted plugin row was a silent no-op. One generic flow in
model_setup_flows.py, credential step keyed by the profile's auth_type
(external_process -> launch check; oauth_* -> live pool row, else the
`hermes auth add <name>` hint), catalog via merge_profile_catalog; main.py
falls back to it for any registered profile missing from the table.
- `_STATUS_BY_AUTH_TYPE` had no builder for oauth_device_code/oauth_external,
so `get_auth_status`/`list_available_providers().authenticated` stayed False
with a live pool entry. `get_plugin_oauth_auth_status` (auth_plugin_providers
sibling) reads the pool; gated on PLUGIN_MIRRORED_PROVIDERS so bundled OAuth
providers keep their bespoke status bytes.
- `_external_process_auth_evidence` computed evidence for copilot-acp only, so
`inventory._external_process_signed_in` hid every other ACP row from the
Desktop explicit_only picker. Generic evidence = the binary resolves; the
bundled CLI keeps its token-store chain.
- agent_init Responses-upgrade guard dropped the vendor literal redundant
with the acp:// scheme check.
- `fetch_account_usage` bounds the plugin hook with a shared 10 s deadline
(previously only the CLI wrapped the call; gateway/TUI awaited unbounded),
contextvars-propagated so scoped secrets resolve; overrun -> None.
Part of #116408
(cherry picked from commit 536a4e7fcf2d1ac2b8a70bd62a69707d58ea5d1e)
An out-of-tree ProviderProfile with auth_type oauth_device_code/oauth_external loaded and inferred
but was invisible to hermes auth: _register_plugin_provider skipped every auth_type except
external_process and api_key, so resolve_provider() said "Unknown provider", `hermes auth add`
had nothing to dispatch to, the pool could not refresh its rows (REFRESHABLE_OAUTH_PROVIDERS is a
name set) and PooledCredential.from_dict dropped every extra key outside _EXTRA_KEYS on reload.
- hermes_cli/auth_plugin_providers.py (new sibling; auth.py is at the size cap): the registry
mirror now registers every profile under the auth_type it declares, re-syncs after discovery /
on a miss (salvaged from #101768), and owns the seam lookups: auth_handler dispatch, the
fail-loud error for a non-api-key profile that ships no handler, and refresh eligibility
derived from the profile's refresh_credential hook (never a name set).
- providers/base.py: ProviderProfile.auth_handler(action, args) (salvaged from #111610, sync only)
and refresh_credential(entry) -> rotated fields. A separate hook rather than
auth_handler("refresh", ...) because the pool holds a credential row, not an argparse namespace,
and needs tokens back rather than a bool.
- hermes_cli/auth_commands.py: add/status/logout/refresh (incl. interactive add) consult the
plugin handler before the built-in path; `auth refresh` admits plugin rows via the predicate.
- agent/credential_pool.py: _refresh_entry_impl calls the profile hook; from_dict keeps every
non-field key in extra so plugin metadata survives load -> save -> load (to_dict already wrote
it all; sanitize_borrowed_credential_payload semantics unchanged).
- hermes_cli/auth.py: config import moved below PROVIDER_REGISTRY (salvaged from #94231) so a
plugin imported during discovery never sees a partial auth module.
Built-in providers are untouched: only custom/openrouter were unregistered profiles before and
both stay in the skip list; anthropic/nous/openai-codex auth add/status/refresh output is
byte-identical in the before/after probe.
Part of #116408. Salvages #111610 (@Finn763), #101768 (@zihaofeng2001, absorbing #106361 by
@Finn763) and #94231 (@Kyzcreig).
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: zihaofeng2001 <zihaofeng2001@gmail.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
hermes_cli.auth imports hermes_cli.config at module top; config runs
provider-plugin discovery at import time (module-level provider imports +
_inject_profile_env_vars), so plugins can execute against a partially
initialized auth module during TUI startup — anything they try to
register/inspect on auth silently loses to fallback metadata.
Move the config import below ProviderConfig/PROVIDER_REGISTRY and document
the ordering contract. Regression test spawns a clean interpreter with a
probe plugin and asserts its registration lands (red before this fix,
green after).
hermes_cli.auth mirrors provider-plugin profiles into PROVIDER_REGISTRY
once, at import time, by iterating list_providers(). providers'
_discover_providers() sets its _discovered guard before importing the
plugin directories, so when a plugin's own imports pull hermes_cli.auth in
mid-discovery (e.g. a user plugin importing agent.credential_pool), the
import-time mirror sees a partial profile list and every plugin discovered
afterwards is dropped from the auth registry. resolve_provider() then
rejects those providers with "Unknown provider" even though
get_provider_profile() and resolve_provider_full() know them.
- auth.py: factor the mirror into idempotent sync_plugin_provider_registry()
(returns the number of newly mirrored profiles) and re-sync on a miss via
_registry_lookup() at the resolve_provider() gate, is_known_auth_provider(),
the get_auth_status() dispatch and the credential/status resolvers.
- providers/__init__.py: call back into hermes_cli.auth (through
sys.modules, never importing it, never raising) when discovery finishes
and on any post-discovery register_provider(), so direct
PROVIDER_REGISTRY readers stay correct too. Registrations *during*
discovery are deliberately not mirrored one by one.
- tests: subprocess end-to-end regression (sorted-first plugin imports
hermes_cli.auth, sorted-last plain plugin must still resolve) plus
in-process contract tests: partial snapshot reconciled at discovery end,
post-discovery registration mirrored, sync idempotent / never clobbers,
providers-side hook never imports hermes_cli.auth. All red on main.
Fixes#102123. Absorbs the in-process tests and never-raise hook shape
from #106361 (Finn763). Related: #21685, #69576, #94231.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AwbHcHDBh9vFRHSDj2W2FY
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
A running gateway/chat holds the credential pool in memory. When the CLI
resets a still-binding cooldown from another process, the reset row on
disk has no status at all, so the recency merge in write_credential_pool
(which ranks by last_status_at) could not tell "reset after my cooldown"
from "never had a status" and let the live pool's stale EXHAUSTED entry
win on its next ordinary flush - a rotation, a token refresh, a sibling
429. `hermes auth reset` printed success and the cooldown came back
(#89415, step B in the thread). status_cleared_ids only protects the
resetting write itself.
reset_status/reset_statuses now stamp status_cleared_at on the cleared
row. The disk-cooldown merge adopts the disk row whenever the in-memory
entry is DEAD/EXHAUSTED and that marker postdates its last_status_at, and
_resync_stale_entry lifts the in-memory cooldown from the same marker, so
the live session serves the credential again without a restart. The
marker is sticky but only outranks OLDER statuses: a fresh exhaustion
after the reset stamps a newer last_status_at and still binds.
Live probe (temp HERMES_HOME, fake Codex entry, process A = live pool,
process B = real auth_reset_command): before, disk read `exhausted`
again after A's flush and A re-selected None; after, disk stays clear and
A re-selects the entry (mem ok). Control: reset then a new 402 stays
exhausted through flush and re-select.
_probe_codex_quota_restored was called with the exhausted pool entry's
stored access token. Exhausted entries are skipped by the proactive refresh
chain, so for any cooldown longer than the access-token lifetime (the
weekly-cap case the probe exists for) that token has expired: the usage
endpoint answers 401 token_expired, the probe maps it to None, and the
cooldown is kept. A mid-cooldown credit top-up or plan upgrade was
therefore never discovered until last_error_reset_at elapsed, even though
the refresh token was still valid (#89415).
Add _refresh_expired_codex_probe_token: when the stored token is expired
and a refresh token exists, rotate the pair via refresh_codex_oauth_pure
(cooldown untouched), persist it (refresh tokens are single-use), and probe
with the live token. Both probe callers use it: the pool-only resolve path
(_probe_codex_pool_entry_quota_restored) and
CredentialPool._codex_quota_restored_upstream. A probe that still reports
100% keeps the bench; the rotated tokens are persisted either way.
Fixes#89415
`--provider chatgpt` (and `chatgpt-codex`) now resolves to the existing ChatGPT-backed
Codex OAuth provider in both alias tables (hermes_cli/providers.py::_ALIAS_GROUPS and
hermes_cli/auth.py::resolve_provider), so users do not need to know the internal
`openai-codex` slug. Storage, transport (codex_responses) and base URL are unchanged;
`openai-codex` stays canonical.
Fixes#95794
Co-authored-by: DarRahman <darrahman@users.noreply.github.com>
`hermes status` / `hermes doctor` / the dashboard cards and the `/model` picker
call `get_codex_auth_status()`, whose singleton fallback ran
`resolve_codex_runtime_credentials()` with runtime defaults: a store missing
its refresh_token imported the Codex CLI's single-use token family, and an
expiring token was refreshed and written back. A diagnostic that spends or
borrows a rotating refresh token logs the other program out (#68004).
`resolve_codex_runtime_credentials(read_only=True)` reports the stored state
as-is (no CLI adoption, no refresh, no pool forced refresh, no write) and wins
over `force_refresh`; the Codex status snapshot uses it, and the xAI OAuth
snapshot passes `refresh_if_expiring=False` for the same reason. The pool side
was already an observation (`peek`, 964fbaae2c).
Superseded #68224 (@GauravPatil2515) — same mechanism, re-done on the current
`auth_codex` layout.
Co-authored-by: Gaurav Patil <gauravpatil2516@gmail.com>
A Nous pool entry whose refresh reached the resolver's state-shape raises
("Hermes is not logged into Nous Portal.", "No access token found ...",
"No refresh token is available ...") carried relogin_required=True but no
OAuth code, so `OAuthProviderFlow.is_terminal_refresh_error` classified it
transient: `_recover_failed_refresh` benched the only credential for an
hour with null last_error_* fields and nothing at WARNING (#113718).
Give those raises codes (`nous_auth_missing`, `nous_auth_missing_access_token`,
`nous_auth_missing_refresh_token`) and list them in the Nous flow's
terminal_refresh_codes, mirroring the Codex/xAI `*_auth_missing_refresh_token`
convention. The pool then quarantines, marks the row DEAD with the reason,
and logs the existing WARNING naming `hermes auth add nous`; every consumer
of the classifier (pool, singleton `_refresh_nous_or_quarantine`) agrees.
A failure that says nothing about the login (network error, non-JSON 5xx
body) still benches as before.
Fixes#113718
Salvages #113726 (@whyyagswhy) — tests kept, consumer-side hunk replaced
by the classifier fix.
Co-authored-by: ahrazzle <ahrazzle@users.noreply.github.com>
The three missing-token messages, the refresh_token_reused copy and the
relogin suffix format_auth_error appends all said a bare `hermes auth` /
`hermes model`, which from a named profile re-signs the ROOT store
(93889b770d) — the same loop #114012 measured on the refresh-rejected path.
They now go through oauth_relogin_command('openai-codex') / profile_cli_selector()
so a profile user is told `hermes -p <profile> auth add openai-codex --type oauth`.
Part of #114012
`get_api_key_provider_status` pre-populated configured/logged_in/base_url/
key_source purely so the deleted keyless short-circuit could return them;
every one is now recomputed before the return, so build the snapshot once
instead of overwriting four dead placeholders. Same keys, same order, same
values.
The `_OPENCODE_FREE_EXCLUDED_MODELS` comment also still implied an
anonymous path that no longer exists; it now names the two ids it holds
and why.
The keyless OpenCode free tier was the only provider that ever set
`HermesOverlay.keyless`, so the flag and everything keyed off it is now
unreachable: the `keyless=` field on `HermesOverlay` and
`ProviderDescriptor`, the `_overlay_has_creds` early return, both
`_provider_is_keyless` copies (auth.py and inventory.py), the
`get_api_key_provider_status` keyless short-circuit and its
`key_source: "keyless"` placeholder, and the empty
`_KEYLESS_STABLE_CACHE_PROVIDERS` set whose `_credential_fingerprint`
branch could never match.
Dropping the two catalog-derived test exemptions follows: they computed
the empty set.
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:
- drop the opencode-free provider row, aliases (free/opencode_free), model
catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
clear error naming the removal
Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
Follow-up to the salvaged #111787 commit (@KoNit-K), same mechanism as #75578 (@adikpb).
What:
- Move the per-model cooldown logic out of the 2.7k-line credential_pool.py facade into
agent/credential_pool_model_cooldowns.py (mixin + module helpers), keeping only the
select/has_available/next_available_at hooks and the mark_exhausted_and_rotate branch in the facade.
- The model cooldown uses the same TTL policy as a credential-wide 429 (_exhausted_ttl: provider
reset_at wins, a sole credential keeps its 60s bench) instead of a flat 1h, so a single-credential
user is never benched LONGER for the failed model than before.
- Drop the `quota_scope == "account"` check: nothing in the tree produces that key.
- Also keep billing_unverified 429s credential-wide, matching the credential-wide branch.
- _rebind_primary_credential_pool read `rt` that was not in its scope (NameError on every
post-fallback restore); pass primary_model from the caller instead.
- resolve_anthropic_token(model=...) gates only model-aware callers; model-less diagnostics
(usage display, model discovery) keep the key as before.
- _anthropic_token_or_raise names the cooled model instead of claiming no credentials exist.
- Simplify the salvaged call sites (unconditional select(model=)/resolve_anthropic_token(model=)),
widen test stubs that lacked the new kwargs, trim the tests to two invariants on a real temp store.
- Docs: credential-pools.md documents per-model Anthropic 429 cooldowns.
Why: a generic Anthropic 429 is a per-model rate limit; benching the whole credential took every
other Claude model offline while the env/borrowed token path handed the same benched key straight
back (#111769, #61451).
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: adikpb <67222969+adikpb@users.noreply.github.com>
A named profile with no credentials of its own silently resolved the root
profile's provider state and credential pool, and a token refresh inside
that profile (xAI, Codex, Anthropic PKCE, Nous) wrote the rotated chain back
into the root store. An isolated service profile therefore acted, and
rotated tokens, as the owner with no way to switch it off.
Maintainer ruling: profiles without credentials are asked to set a provider,
never handed another profile's auth. Profiles are independent islands.
What changes
- `hermes_cli/auth.py`: `_load_provider_state*`, `read_credential_pool` and
`_provider_state_transaction` read the active store only; the global-root
resolver, its mtime memo and `_persist_provider_state_to_store` are gone.
- xAI / Codex / Nous-guest / pool refresh paths persist to the active store;
the root write-through, the borrowed-row bookkeeping
(`_borrowed_root_ids`, `persist_pool_entries`, `_update_root_pool_rows`)
and the forked-grant heal are removed. `_write_hermes_oauth_credentials`
loses its root `target`.
- `resolve_provider` / `agent_init` name the profile in the
no-provider error and print `hermes -p <name> model` guidance.
- `hermes update` prints a one-time notice listing every named profile that
has no provider of its own (`hermes_cli/profile_credential_audit.py`) so a
bot never goes quiet unannounced.
- Desktop create dialog: the "Share keys & accounts" checkbox described the
removed inheritance; it now mirrors API keys (`mirror_credentials`) and
says OAuth logins need a sign-in. `share_auth` is accepted from older
clients and ignored; `ProfileMirrored.auth` is a bool again.
- Docs: profiles.md, multi-profile-gateways.md isolation table,
hermes_cli/AGENTS.md.
The fallback was added in 33bf5f62 so kanban/cron workers under a named
profile did not die with "No LLM provider configured" when the credential
lived only at root; that convenience is exactly the isolation hole the
ruling closes, and `--clone` / dashboard mirroring still copy API keys.
Tests: fallback/write-through/heal pins deleted; 4 invariants proven red on
base (profile never reads root; profile refresh never writes root; Nous
connector gate reads only the profile store; Anthropic pool never borrows or
rotates the root grant, root control still refreshes).
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.
The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.
Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.