183 Commits

Author SHA1 Message Date
kshitijk4poor
96cce6843d refactor(auth): drop impossible auth_store dict guard 2026-09-27 18:37:01 +05:30
kshitijk4poor
c8043a3630 fix(auth): drop redundant ownership auth.json reads in nous seeding and load_pool
_seed_nous_singleton re-read auth.json via _profile_owns_pool_provider even
though its only caller (_seed_from_singletons) just loaded the active store
and passed it in; check the passed auth_store through a shared
_store_owns_pool_provider predicate instead (same non-empty-list semantics,
also used by _profile_owns_pool_provider).

In load_pool, borrowing_root_grant repeated the guard that sets
owns_provider (non-None exactly when that guard holds), so test
`owns_provider is False` directly. The tail ownership re-read ran even
with no disk rows, where it could only assign set() -- the constructor
default -- so gate it on disk_ids and re-read only when _persist() ran.
Fix the stale "Computed once" comment.
2026-09-27 18:37:01 +05:30
kshitijk4poor
3dac1b3ae1 fix(auth): read flat refresh_token for forkable pool rows; reuse pool ownership in load_pool
_is_forkable_pool_row only ever receives flat credential_pool rows (the
strip loop over pool entries and heal_pool_rows over _pool_rows), so the
tokens-nesting fallback of _block_tokens was dead weight; read
refresh_token the same way _is_oauth_pool_payload does.

load_pool asked _profile_owns_pool_provider (an uncached auth.json read)
twice. Compute it once after the fork heal and reuse it for the
_borrowed_root_ids check unless _persist() rewrote the store in between
(that write can give the profile its own rows). _seed_nous_singleton keeps
its own call: threading the value through _seed_from_singletons would
change a signature that tests monkeypatch with fixed-arity fakes.
2026-09-27 18:37:01 +05:30
kshitijk4poor
e7bab8eb18 fix(auth): don't reseed root's nous grant into a profile that owns nous rows
A profile left with only an agent_key nous row after the fork strip/heal still
"owns" nous but has no local providers.nous block. The next load_pool('nous')
fell back to the global root block in _seed_nous_singleton and upserted root's
single-use refresh token into the profile pool, recreating the fork one load
later (both device_code and manual:* ak-row shapes). Skip seeding from the
global-root fallback when the profile owns local nous rows; borrowing profiles
(no local rows) are unaffected.

Also drop the redundant try/except around _global_auth_file_path() in
_profile_owns_pool_provider; that function already handles its own failures.
The existing nous strip test now reloads the pool after strip and asserts no
profile row carries root's refresh token.
2026-09-27 18:37:01 +05:30
kshitijk4poor
ddcd845993 fix(auth): skip auth.json re-read for pool ownership in classic mode
nous now takes the single-use path, so every load_pool('nous') (once per
message plus aux calls) re-parsed auth.json in _profile_owns_pool_provider.
In classic mode (_global_auth_file_path() is None) read_credential_pool has
no root fallback and persist_pool_entries cannot route to root, so the
answer is effectively always "owns": return early.

Also point the persist_pool_entries docstring at
SINGLE_USE_REFRESH_POOL_PROVIDERS and document why nous is deliberately
absent from _SINGLE_USE_REFRESH_PROVIDERS (own auth-store locking).
2026-09-27 18:37:01 +05:30
kshitijk4poor
5e2d55ca9c fix(codex): quota probe and /usage pool paths use the pool route base
Pool rows keep the canonical chatgpt.com URL, so the quota-restored probe
(auth_codex + CredentialPool) and the /usage tier-3 and forced-refresh
paths paired a gateway key with chatgpt.com/backend-api/wham/usage. Route
them through _codex_pool_route_base_url, the chat route's rule
(HERMES_CODEX_BASE_URL > model.base_url > row URL).

Refs #121486
2026-09-25 21:27:06 +05:30
ethernet
41934d5015 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-24 06:21:50 -04:00
kshitijk4poor
f24a1d7f92 refactor(credential_pool): drop the loop index _persist no longer reads
The previous commit replaced the positional write with _replace_entry, leaving
the enumerate index unused (ruff B007; flagged by the pre-arm simplify gate).
2026-09-24 15:44:48 +05:30
kshitijk4poor
754ea72538 refactor(credential_pool): collapse the stale-writer early returns
`_credential_token_pair` already maps a non-dict row to (None, None), so the
separate isinstance guard and second early return in
`_merge_pool_row_generation` were the same branch. Adopting a peer generation
now goes through the existing `_replace_entry` swap primitive.
2026-09-24 15:44:48 +05:30
kshitijk4poor
9e66e405cf fix(auth): return the post-persist entry from _adopt; give the adoption guard test teeth
_persist() may swap a peer's newer token generation into self._entries,
but _adopt() returned the pre-persist object. Callers that rebind the live
client from that return value (try_refresh_matching -> _swap_credential,
auxiliary_client, account_usage) would keep using the stale pair for one
guaranteed 401 while the pool already held the rotated one. Re-look the
entry up by id after persisting so the return value matches the pool.

The "adopt only if the written pair differs from the live pair" guard is
kept: without it every ordinary flush would replace the live object and
pull peer cooldown state merged into the written row into memory. The
kept generation test now asserts an unchanged-pair flush preserves object
identity and that a stale writer's _adopt returns the adopted entry;
reverting either the guard or the re-lookup fails both parametrizations.

Also move the reference-only-rows comment onto the blank-pair guard it
describes instead of the rehydration statement.
2026-09-24 15:44:48 +05:30
kshitijk4poor
d525b443f5 refactor(auth): one token-base helper with one blank-row policy; dedupe generation copy loop
The id->token-pair base map was built three ways: _persist dropped
(None, None) pairs, load_pool and _sync_entry_from_pool_store kept them.
_merge_pool_row_generation treats a missing base as "unknown" (plain
recency merge) but a (None, None) base as a known generation, so a
token-less row got the generation override on the first flush after
load_pool and the plain merge on every later one.

Policy chosen: KEEP blank bases everywhere (new auth_mod._token_pairs_by_id,
built on the existing _entry_ids). A blank base means "no pair when we last
looked", which is exactly the CAS witness the boundary needs: when a peer
lands a pair on that row, every flush of ours keeps the peer's generation
instead of writing our blank tokens back. Dropping blank bases would make
the first flush keep the peer's pair and later flushes overwrite it.
_sync_entry_from_pool_store now records the base before its no-token-
material bail-out so all three sites agree. _update_root_pool_rows reuses
_entry_ids for its incoming map as well.

Also collapse the three copy-from-disk-else-pop loops in
_merge_pool_row_generation into a local _take_from_disk helper.
failure_reason is passed alongside _POOL_STATUS_FIELDS rather than added
to the shared constant, which auth_codex/auth_oauth_grants and
_merge_disk_cooldown_state also consume.
2026-09-24 15:44:48 +05:30
kshitijk4poor
badd8f56e4 fix(auth): scope only terminal DEAD verdicts to the stale token generation
When a stale pool's persist finds the disk pair moved, #120943 copied every
status field from disk, so the stale writer's NEWER account-wide verdict
(402 billing / 429 throttle -> EXHAUSTED) was erased and the rotated pair
re-entered selection immediately. Only a terminal auth death is tied to the
pair it was observed on: discard the stale writer's DEAD onto the peer's
pair, but leave any other status for _merge_disk_cooldown_state's ordinary
recency merge.

Also travel `scope` and `inference_base_url` with the pair (a Nous refresh
rewrites them together with the tokens, so a stale writer must not stamp an
old scope/route onto the newer pair); use the always-initialised
_persisted_token_pairs directly; and re-hydrate an in-memory entry from the
written row only when the store overrode its pair, so unchanged rows keep
runtime-only fields to_dict() omits.

Tests: fold the stale-later-402 witness (rt-1 kept AND last_status ==
exhausted) into the stale-terminal-verdict test and drop the separate 429
rollback test it subsumes.

Fixes #120815

Co-authored-by: John Paul Soliva <soliva.johnpaul@icloud.com>
2026-09-24 15:44:48 +05:30
JoaoMarcos44
d2f54de2cc fix(auth): bind stale pool writes to token generation
(cherry picked from commit 1f5c7bcd076cd6184fb048ea1a57c4549f5f4991)
2026-09-24 15:44:48 +05:30
ethernet
c399093de9 merge: reconcile profile-scoped routes and shared desktop backend with PM 2026-09-21 13:18:11 -04:00
teknium1
2632229bcf fix(auth): named profiles read the root auth.json again (revert #111724)
Reverts 93889b770d ("named profiles no longer inherit the root
profile's auth.json"). After the Desktop update every bot profile that had
relied on the root OpenAI Codex login failed with "No Codex credentials
stored. Run `hermes -p <bot> auth add openai-codex --type oauth`", and users
had to re-run the device-code flow once per bot (5-6 times in the field
report). Sharing one grant across profiles is the intended design: OAuth
refresh tokens are single-use, so ONE grant lives at the root, profiles
resolve it read-only, and a refresh under a profile writes the rotated chain
back to root (Codex / xAI write-through, borrowed-row pool bookkeeping,
forked-grant heal) — never a per-profile copy.

Restored: `_global_auth_file_path` / `_load_global_auth_store` fallback in
`_load_provider_state*` / `read_credential_pool` / `_provider_state_transaction`,
Codex + xAI root write-through, `credential_pool` borrowed-root persistence,
`heal_forked_single_use_oauth_grants`, `share_auth` on profile creation
(Desktop create dialog checkbox), and the docs. `profile_credential_audit.py`
(the `hermes update` "profiles without a provider" notice) is removed with it.

Kept from after #111724: `_save_codex_tokens(set_active=...)` for image gen,
the plugin-auth `status` dispatch and the external-login notice in
`hermes auth list`, and the registry-derived env-var hint in agent_init.
2026-09-21 09:55:40 -07:00
teknium1
d03b5f3770 fix(auth): skip the Copilot token exchange while copilot is only an ambient gh-CLI credential
load_pool("copilot") exchanged the gh CLI token on every load — on installs where copilot is
never selected (main provider deepseek, aux slots auto) that is a network round-trip plus the
"degraded to RAW token" warning on every pool load, ~2 per turn, 1000+ times in five days for
the reporter (#114740). The exchange result only matters once a model is routed to copilot, so
the seeder now keeps the raw token and skips the exchange (and the warning) until
is_provider_explicitly_configured("copilot") is true; the load that follows the user selecting
copilot re-seeds and exchanges as before. auxiliary.<task>.provider now counts as selecting a
provider in _config_selects_provider, like a MoA slot, so an aux slot pinned to copilot still
gets the exchanged token. Contributor tests trimmed to two invariants.
2026-09-20 13:22:30 -07:00
Mohamad Kanso
bd2b8124f4 fix(auth): prevent repeated copilot raw token exchange warnings (#114740) 2026-09-20 13:22:30 -07:00
teknium1
bce20d0b1f feat: model-provider plugins classify their own API errors via ProviderProfile.classify_api_error
A kind: model-provider plugin is loaded by providers/ discovery and never enters the
PluginManager hook lifecycle, so transform_api_error_classification was unreachable for it
without shipping a second plugin component. The profile now carries an optional
classify_api_error(error, *, status_code, error_code, message, body, model) callable,
consulted as a classifier stage right after the generic plugin hooks and only for the
provider that produced the error. None or an unknown reason leaves the built-in verdict;
built-in providers are untouched (no name table, no lifecycle change).

Also: a plugin refresh_credential returning None/empty was treated as a successful refresh
(row marked ok, stale bearer replayed up to the refresh cap). It now benches the row like a
failed refresh POST, so the loop rotates or falls to the generic sign-in copy.

Part of #116408

(cherry picked from commit b8129fd6fd6a0cf4eeee6d95a5d5d823668a306d)
2026-09-19 21:24:05 -07:00
teknium1
ec1238fa65 fix(credential-pool): plugin refresh keeps rotated tokens, runs locked, and quarantines dead grants
Independent review of the plugin refresh branch (#116553) found four gaps
between what model-provider-plugin.md promises for `refresh_credential`
and what `_refresh_entry_impl` did:

1. `replace(entry, **hook_result)` raised TypeError on any non-field key
   (`expires_in`, `token_type`, `scope` — the natural token-endpoint shape),
   the except benched the row EXHAUSTED and the pair the server had already
   rotated was dropped: for single-use refresh tokens that is a lost login.
   Field keys now go through `replace()`, everything else merges into
   `entry.extra` (mirrors `from_dict`); `None` = no rotation, mark ok.
2. Plugin providers skipped the locked single-use path, so a gateway and a
   CLI could both POST the same refresh token (`refresh_token_reused`).
   Providers with a hook now take the `_auth_store_lock` path: re-read the
   pool store, adopt a peer's usable rotation and skip the hook, else call
   it and write through. Eligibility derives from `plugin_refresh_hook()`,
   not from extending the built-in name tuple.
3. A raising hook re-benched EXHAUSTED every cooldown forever at DEBUG.
   `AuthError(relogin_required=True)` (or a grant-dead OAuth code) is now
   terminal: the row goes DEAD with a WARNING naming `hermes auth add`.
   Any other exception stays a transient bench (negative test kept).
4. `from_dict`'s extra sweep round-tripped a stray row-level `provider` key
   back onto the row on `to_dict()`; it is bookkeeping, not metadata.

New logic lives in `agent/credential_pool_plugin.py` — credential_pool.py
is at the size cap; the facade only dispatches.

Part of #116408
2026-09-19 20:45:06 -07:00
teknium1
dce233e809 feat(auth): OAuth-shaped provider plugins register, log in and refresh through their profile
An out-of-tree ProviderProfile with auth_type oauth_device_code/oauth_external loaded and inferred
but was invisible to hermes auth: _register_plugin_provider skipped every auth_type except
external_process and api_key, so resolve_provider() said "Unknown provider", `hermes auth add`
had nothing to dispatch to, the pool could not refresh its rows (REFRESHABLE_OAUTH_PROVIDERS is a
name set) and PooledCredential.from_dict dropped every extra key outside _EXTRA_KEYS on reload.

- hermes_cli/auth_plugin_providers.py (new sibling; auth.py is at the size cap): the registry
  mirror now registers every profile under the auth_type it declares, re-syncs after discovery /
  on a miss (salvaged from #101768), and owns the seam lookups: auth_handler dispatch, the
  fail-loud error for a non-api-key profile that ships no handler, and refresh eligibility
  derived from the profile's refresh_credential hook (never a name set).
- providers/base.py: ProviderProfile.auth_handler(action, args) (salvaged from #111610, sync only)
  and refresh_credential(entry) -> rotated fields. A separate hook rather than
  auth_handler("refresh", ...) because the pool holds a credential row, not an argparse namespace,
  and needs tokens back rather than a bool.
- hermes_cli/auth_commands.py: add/status/logout/refresh (incl. interactive add) consult the
  plugin handler before the built-in path; `auth refresh` admits plugin rows via the predicate.
- agent/credential_pool.py: _refresh_entry_impl calls the profile hook; from_dict keeps every
  non-field key in extra so plugin metadata survives load -> save -> load (to_dict already wrote
  it all; sanitize_borrowed_credential_payload semantics unchanged).
- hermes_cli/auth.py: config import moved below PROVIDER_REGISTRY (salvaged from #94231) so a
  plugin imported during discovery never sees a partial auth module.

Built-in providers are untouched: only custom/openrouter were unregistered profiles before and
both stay in the skip list; anthropic/nous/openai-codex auth add/status/refresh output is
byte-identical in the before/after probe.

Part of #116408. Salvages #111610 (@Finn763), #101768 (@zihaofeng2001, absorbing #106361 by
@Finn763) and #94231 (@Kyzcreig).

Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: zihaofeng2001 <zihaofeng2001@gmail.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
2026-09-19 20:45:06 -07:00
rjshrjndrn
7034628c92 fix(codex): proxy override survives credential rotation and model.base_url is honoured
HERMES_CODEX_BASE_URL was applied at pool resolution and on the auxiliary
clients, but two paths still sent the openai-codex provider back to the
default ChatGPT backend:

- credential rotation: client_lifecycle._swap_credential adopts
  PooledCredential.runtime_base_url, which for openai-codex was the pool
  row's stored canonical URL, so the first 401/429 rotation silently left
  the proxy. The override now lives in runtime_base_url, the one place every
  reader of a Codex pool row (resolution and rotation) goes through.
- model.base_url: the openai-codex branch of _pool_entry_mode_and_url
  returned before the generic model.base_url block. It now honours
  model.base_url under model.provider: openai-codex when the pool row still
  carries the canonical URL (env override keeps precedence).

Slim port of #40924 onto the current layout (the original patched
run_agent._swap_credential, the pool seeder and a new auth.py helper; the
seeder half landed in b62bb2a3d5, the helper is replaced by the
profile-scoped get_secret_str read the landed fixes already use).

Fixes #40913
2026-09-19 10:33:08 -07:00
teknium1
a70053d9d7 fix(auth): throttle the pre-probe Codex refresh and pin the pool-side wiring
Review follow-up on the #89415 pre-probe refresh:

- `_refresh_expired_codex_probe_token` ran ahead of `_probe_codex_quota_restored`'s
  5-minute per-token cache with no throttle of its own, so a frozen entry whose
  refresh keeps failing (revoked grant, network down) POSTed to the token endpoint
  under the pool lock on every credential selection. It now consults the same
  per-token probe cache before refreshing and reserves the stale token's slot on
  failure, restoring the pre-PR budget of at most one network call per interval.
- The pool-side refresh adopted the rotated pair into the pool row only; the next
  selection's `_sync_entry_from_auth_store` saw a differing singleton and re-adopted
  the consumed pair (clearing the cooldown with it). Stamp `last_refresh` and write
  the pair back to `providers.openai-codex` exactly as `_refresh_entry` does.
- Tests: one resolver-side refresh test is replaced by one that drives
  `CredentialPool.select()` (probe receives the refreshed bearer, both disk sides
  hold the rotated pair across a second selection) and a control that repeated
  selections with a failing refresh make a single refresh attempt.
2026-09-19 09:40:56 -07:00
teknium1
a564f77746 fix(auth): hermes auth reset survives a live session's next pool flush
A running gateway/chat holds the credential pool in memory. When the CLI
resets a still-binding cooldown from another process, the reset row on
disk has no status at all, so the recency merge in write_credential_pool
(which ranks by last_status_at) could not tell "reset after my cooldown"
from "never had a status" and let the live pool's stale EXHAUSTED entry
win on its next ordinary flush - a rotation, a token refresh, a sibling
429. `hermes auth reset` printed success and the cooldown came back
(#89415, step B in the thread). status_cleared_ids only protects the
resetting write itself.

reset_status/reset_statuses now stamp status_cleared_at on the cleared
row. The disk-cooldown merge adopts the disk row whenever the in-memory
entry is DEAD/EXHAUSTED and that marker postdates its last_status_at, and
_resync_stale_entry lifts the in-memory cooldown from the same marker, so
the live session serves the credential again without a restart. The
marker is sticky but only outranks OLDER statuses: a fresh exhaustion
after the reset stamps a newer last_status_at and still binds.

Live probe (temp HERMES_HOME, fake Codex entry, process A = live pool,
process B = real auth_reset_command): before, disk read `exhausted`
again after A's flush and A re-selected None; after, disk stays clear and
A re-selects the entry (mem ok). Control: reset then a new 402 stays
exhausted through flush and re-select.
2026-09-19 09:40:56 -07:00
Teknium
63bcdcaa43 fix(auth): refresh an expired Codex token before the quota-restored probe
_probe_codex_quota_restored was called with the exhausted pool entry's
stored access token. Exhausted entries are skipped by the proactive refresh
chain, so for any cooldown longer than the access-token lifetime (the
weekly-cap case the probe exists for) that token has expired: the usage
endpoint answers 401 token_expired, the probe maps it to None, and the
cooldown is kept. A mid-cooldown credit top-up or plan upgrade was
therefore never discovered until last_error_reset_at elapsed, even though
the refresh token was still valid (#89415).

Add _refresh_expired_codex_probe_token: when the stored token is expired
and a refresh token exists, rotate the pair via refresh_codex_oauth_pure
(cooldown untouched), persist it (refresh tokens are single-use), and probe
with the live token. Both probe callers use it: the pool-only resolve path
(_probe_codex_pool_entry_quota_restored) and
CredentialPool._codex_quota_restored_upstream. A probe that still reports
100% keeps the bench; the rotated tokens are persisted either way.

Fixes #89415
2026-09-19 09:40:56 -07:00
liuhao1024
8530d81b55 fix(credential-pool): treat rotate-to-self after mark as no-recovery
mark_exhausted_and_rotate benched the failed entry, then _select_unlocked
could hand the very same entry straight back (the auth-store sync adopted
fresher tokens and cleared the bench mid-selection, or a false-positive
quota probe lifted it). The caller took the non-None result as a successful
rotation and retried the same 429 with no attempt cap: a sole-credential
openai-codex pool on a weekly usage_limit_reached spun at ~2 req/s for
hours, ~11k requests, turn never failed (#97315).

If the post-mark selection returns the just-marked entry, return None so
the failure surfaces and provider fallback / error propagation proceed —
mirroring the single-entry guard on the unmatched-identity branch.
Multi-entry pools still rotate to a healthy sibling. Ported from PR #97328.

Fixes #97315
2026-09-19 09:40:56 -07:00
teknium1
4a43cc50ae fix: pin entitlement rotation through recover_with_credential_pool; bench until reset
The rotation test drove pool.mark_exhausted_and_rotate directly, so reverting the
recover_with_credential_pool model_entitlement branch stayed green. It now drives the
production recovery entry point with the classifier verdict and asserts (True, False),
the swap to cred-1 and the model-only bench on cred-0.

A ChatGPT-account entitlement is a plan property, not a quota window: the model bench
now lasts until the explicit reset path (hermes auth reset) clears model_cooldowns,
instead of re-probing the unentitled account hourly on the 429 TTL (#71970).
2026-09-19 09:37:23 -07:00
teknium1
7958cbf2de fix: rotate Codex pool credentials on a ChatGPT-account model entitlement 400
The exact `The '<model>' model is not supported when using Codex with a
ChatGPT account.` 400 classified as format_error, so a two-entry openai-codex
pool never tried its second account: the turn failed as a malformed request
even though the other account was entitled (#71970). #106475 covered the
single-credential case and explicitly left the multi-entry case to rotation,
but no rotation branch existed for it.

- classify the exact normalized text as FailoverReason.model_entitlement
  (rotate + fallback, never retry); arbitrary 400s stay format_error
- recover_with_credential_pool rotates once on that reason; the pool records
  it through the existing model_cooldowns path (Anthropic per-model 429), so
  only (credential, model) is benched: other models keep using the account,
  and `hermes auth` reset clears the marker with everything else
- _mark_entitlement_rejected_model gates on pool.has_available(model=...)
  instead of entry count, so once every account rejects the model it falls
  back to the #106475 session marker (fallback walk skips it, no oscillation)

Salvages the design of #71973 by @kilhyeonjun on the current classifier
tables and the model-cooldown substrate that landed since.

Co-authored-by: kilhyeonjun <kboxstar@gmail.com>
2026-09-19 09:37:23 -07:00
teknium1
c7d0045a8d fix: mid-run credential rotation never rebinds a session to an entry for another endpoint
The delegate fix admits a MIXED same-provider pool (any one entry serves the
child's endpoint) but only guarded the initial lease. A 429/402/401 in the
Azure child later ran recover_with_credential_pool -> mark_exhausted_and_rotate
-> _swap_credential(next_entry) with no endpoint check, and
credential_pool_matches_provider passes a non-custom pool on provider match
alone, so the child was rebound to the api.openai.com entry (#68237 via a
sibling path).

_entry_serves_endpoint moves into agent/credential_pool.py as
credential_pool_entry_serves_endpoint (next to credential_pool_matches_provider;
tools/delegate_tool_config imports it back), and _rotate_and_swap vetoes a
next_entry whose base_url does not serve agent.base_url, returning False exactly
like a rotation that yields nothing: the failed entry stays benched, the
session keeps its credential and the fallback chain takes over.
2026-09-19 09:34:09 -07:00
kshitijk4poor
14a3463454 fix(anthropic): mirror only the Keychain item that held the spent pair; parse both security attribute encodings
Review findings on the mirror, all verified live against throwaway items:

* Identity gate. The service name is fixed but the file path honours
  CLAUDE_CONFIG_DIR, and Claude Code rotates on its own schedule, so the
  item under the service can hold a different login's pair or a newer
  rotation. Overwriting it would be the bug in the other direction. The
  refresh token that was just POSTed is threaded through
  `_write_claude_code_credentials(spent_refresh_token=...)` from both
  callers (the singleton refresher and the pool commit) and the mirror
  only updates an item whose `refreshToken` equals it.
* Attribute parsing. `security` prints an attribute as `"text"` when it
  is plain printable ASCII — UNescaped, an embedded `"` appears raw — and
  as `0x<HEX>  "<echo>"` otherwise. The old regex modelled `\"` escaping
  that never happens: a `"` in the account truncated it and the mirror
  created a second item; any non-ASCII byte made it return "" and the
  mirror silently no-oped. One `find-generic-password -g` call now yields
  account and payload together, parsed line-anchored in both encodings.
* Fail-soft is now total (`except Exception`): the file commit already
  succeeded when the mirror runs, and a raise here made the refresher
  mark a landed rotation as consumed-uncommitted.
* `quoted` lambda -> nested def; `ensure_ascii=True` made explicit since
  `-w` returns non-ASCII payloads as hex.

Two pre-existing test doubles for the writer accept the new keyword.
2026-09-19 11:25:23 +05:30
Tranquil-Flow
20512a47a6 fix(agent): refuse stale singleton adoption over rotated manual:device_code entries (#106705)
A manual:device_code Codex pool entry never writes its rotation back to the
auth.json singleton (independent-credential contract, #39236). The singleton
sync adopted differing singleton tokens with no staleness proof, so after a
pool-side rotation the stale singleton was re-adopted over the pool's fresh
chain and the already-consumed refresh token was POSTed again
(refresh_token_reused).

Gate adoption on the singleton's last_refresh not predating the entry's own
rotation; missing stamps on either side keep the historical
adopt-on-difference behavior (#70111).
2026-09-18 20:57:07 -07:00
Marc Caelier
46503f1672 fix(auth): isolate Codex singleton sync by principal 2026-09-18 20:57:07 -07:00
teknium1
6c7f693473 feat(auth): opt out of borrowing Codex CLI / Claude Code logins (auth.adopt_external_logins)
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.

- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
  current users). When false:
  - `read_claude_code_credentials()` — the only reader of the borrowed Claude
    Code login — returns None, so the resolver fallback, the expired-token
    refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
    file; the pool prunes a `claude_code` row an earlier adopting process
    persisted.
  - `_recover_codex_tokens_from_cli` returns None for both automatic recovery
    paths (rejected refresh, half-empty singleton); the real AuthError is
    surfaced instead. The interactive import offer in `hermes auth add
    openai-codex` still asks first and is unaffected.
  - One INFO line per process the first time adoption would have happened;
    `hermes auth list` / `hermes auth status anthropic|openai-codex` print the
    same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.

Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
2026-09-18 20:55:24 -07:00
teknium1
92bb5b92b8 fix: revert a quota-benched credential through the pool, not a private probe; cover /model and chained rotations
Reshape of the salvaged fix from #114513 (@whyyagswhy):

- ``CredentialPool.reclaim(credential_id, model=)`` is the pool-owned answer to "is the
  benched entry back?": it runs under the pool lock, clears the elapsed cooldown and
  refreshes the token exactly as ``select()`` would, but never bumps ``request_count``
  or round-robin order. The contributor's version called the private
  ``_available_entries()`` outside the lock (its docstring requires the lock: it prunes
  and persists) and left the entry marked ``exhausted`` in the pool after the swap.
- ``_rotate_and_swap`` arms the revert only when nothing is armed yet, so a chained
  429 (preferred → fallback → third) still returns to the PREFERRED entry rather than
  the middle one.
- A deliberate ``/model`` switch (``_finish_switch``) cancels the pending revert with
  the rest of the fallback state; dropped the ``_credential_pool_rotated_to`` "moved
  by hand" heuristic and the ``_provider_fallback_active`` guard (unreachable: the
  hook only runs on the ``not _fallback_activated`` branch).
- Tests trimmed to two invariants on a real ``CredentialPool`` through the real
  ``recover_with_credential_pool`` → ``restore_primary_runtime`` path: (1) stays on the
  fallback while the bench holds, moves back once it lifts, ``_fallback_activated``
  and the model untouched; (2) a 401 bench does not arm a revert.
- Docs: credential-pools "Error Recovery" describes the switch-back.
2026-09-18 10:35:37 -07:00
teknium1
d6add16059 fix(auth): dead OAuth logins are reported once and leave rotation; hints point at hermes auth
A terminally rejected refresh token (invalid_grant / invalid_token /
refresh_token_reused) is the moment a login is lost. Three gaps remained
after #113197 (which added the WARNING for Codex/xAI/Nous):

- Anthropic: the token endpoint's HTTPError carried no classifiable code,
  so a dead Anthropic / Claude Code grant fell through to a transient
  'exhausted' bench at DEBUG and was replayed every hour. The endpoint
  error is now a structured AnthropicOAuthError (status + OAuth error
  code); _recover_failed_refresh logs one WARNING with the repair command
  and marks the row DEAD. A dead grant is not replayed at the fallback
  endpoint. The auxiliary Claude Code refresher (sibling path) warns the
  same way. Claude Code's own credentials file is never touched.
- Codex/xAI/Nous: _quarantine_sources drops only singleton-seeded rows, so
  an independent `hermes auth add` (manual:*) login survived unmarked and
  re-fired the WARNING on every later refresh attempt. The survivor is now
  marked DEAD (leaves rotation until a write-side re-auth clears it).
- Hints on live Hermes paths (anthropic 401 troubleshooting block, the
  no-credentials error) recommended an external CLI's login command,
  which cannot repair Hermes' own login; they now say
  `hermes auth add anthropic` / `hermes auth list anthropic`.

Docs: credential-pools.md documents the dead-login behaviour.
Part of #113023.
2026-09-17 09:02:26 -07:00
teknium1
9bcba9bef7 fix(auth): a terminally rejected OAuth refresh token is logged at WARNING with a re-auth hint
When the pool's refresh of an openai-codex / xai-oauth / nous entry fails
terminally (invalid_grant, revoked family), _recover_failed_refresh clears
the stored tokens and drops the seeded entry -- the moment the user's login
is lost -- but logged it at DEBUG only. At the default log level nothing
explained why the provider suddenly ran on its fallback, so the failure
read as "I signed in once and Hermes keeps failing" (#113023).

Both quarantine sites now log one WARNING carrying the terminal error and
the exact recovery command (`hermes auth add <provider>`). The quarantine
behaviour itself is unchanged.

Part of #113023
2026-09-16 17:24:07 -07:00
teknium1
6de6e6da99 fix(anthropic): model-scoped 429 cooldowns live in a pool sibling; restore path no longer NameErrors
Follow-up to the salvaged #111787 commit (@KoNit-K), same mechanism as #75578 (@adikpb).

What:
- Move the per-model cooldown logic out of the 2.7k-line credential_pool.py facade into
  agent/credential_pool_model_cooldowns.py (mixin + module helpers), keeping only the
  select/has_available/next_available_at hooks and the mark_exhausted_and_rotate branch in the facade.
- The model cooldown uses the same TTL policy as a credential-wide 429 (_exhausted_ttl: provider
  reset_at wins, a sole credential keeps its 60s bench) instead of a flat 1h, so a single-credential
  user is never benched LONGER for the failed model than before.
- Drop the `quota_scope == "account"` check: nothing in the tree produces that key.
- Also keep billing_unverified 429s credential-wide, matching the credential-wide branch.
- _rebind_primary_credential_pool read `rt` that was not in its scope (NameError on every
  post-fallback restore); pass primary_model from the caller instead.
- resolve_anthropic_token(model=...) gates only model-aware callers; model-less diagnostics
  (usage display, model discovery) keep the key as before.
- _anthropic_token_or_raise names the cooled model instead of claiming no credentials exist.
- Simplify the salvaged call sites (unconditional select(model=)/resolve_anthropic_token(model=)),
  widen test stubs that lacked the new kwargs, trim the tests to two invariants on a real temp store.
- Docs: credential-pools.md documents per-model Anthropic 429 cooldowns.

Why: a generic Anthropic 429 is a per-model rate limit; benching the whole credential took every
other Claude model offline while the env/borrowed token path handed the same benched key straight
back (#111769, #61451).

Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Co-authored-by: adikpb <67222969+adikpb@users.noreply.github.com>
2026-09-16 17:16:06 -07:00
KoNit-K
fb358d4379 fix(agent): scope Anthropic 429 cooldowns to model 2026-09-16 17:16:06 -07:00
teknium1
93889b770d fix(auth): named profiles no longer inherit the root profile's auth.json (#111724)
A named profile with no credentials of its own silently resolved the root
profile's provider state and credential pool, and a token refresh inside
that profile (xAI, Codex, Anthropic PKCE, Nous) wrote the rotated chain back
into the root store. An isolated service profile therefore acted, and
rotated tokens, as the owner with no way to switch it off.

Maintainer ruling: profiles without credentials are asked to set a provider,
never handed another profile's auth. Profiles are independent islands.

What changes
- `hermes_cli/auth.py`: `_load_provider_state*`, `read_credential_pool` and
  `_provider_state_transaction` read the active store only; the global-root
  resolver, its mtime memo and `_persist_provider_state_to_store` are gone.
- xAI / Codex / Nous-guest / pool refresh paths persist to the active store;
  the root write-through, the borrowed-row bookkeeping
  (`_borrowed_root_ids`, `persist_pool_entries`, `_update_root_pool_rows`)
  and the forked-grant heal are removed. `_write_hermes_oauth_credentials`
  loses its root `target`.
- `resolve_provider` / `agent_init` name the profile in the
  no-provider error and print `hermes -p <name> model` guidance.
- `hermes update` prints a one-time notice listing every named profile that
  has no provider of its own (`hermes_cli/profile_credential_audit.py`) so a
  bot never goes quiet unannounced.
- Desktop create dialog: the "Share keys & accounts" checkbox described the
  removed inheritance; it now mirrors API keys (`mirror_credentials`) and
  says OAuth logins need a sign-in. `share_auth` is accepted from older
  clients and ignored; `ProfileMirrored.auth` is a bool again.
- Docs: profiles.md, multi-profile-gateways.md isolation table,
  hermes_cli/AGENTS.md.

The fallback was added in 33bf5f62 so kanban/cron workers under a named
profile did not die with "No LLM provider configured" when the credential
lived only at root; that convenience is exactly the isolation hole the
ruling closes, and `--clone` / dashboard mirroring still copy API keys.

Tests: fallback/write-through/heal pins deleted; 4 invariants proven red on
base (profile never reads root; profile refresh never writes root; Nous
connector gate reads only the profile store; Anthropic pool never borrows or
rotates the root grant, root control still refreshes).
2026-09-16 14:34:59 -07:00
teknium1
996f7bc563 feat(credential-pool): numbered env siblings (KEY_2, KEY_3, …) seed rotation
Setting NVIDIA_API_KEY_2 next to NVIDIA_API_KEY is now the whole opt-in
for a second pooled key: _seed_from_env tries VAR_2, VAR_3, … for every
declared var until the first gap, on the generic registry path and the
openrouter branch alike. Secrets stay in the env / secret manager; only
the reference row is persisted. Resolves #76593; supersedes the config-key
approach of #87835.
2026-09-15 18:39:31 -07:00
teknium1
423bc7e4e4 fix(credential-pool): hydrate on-disk env rows on the openrouter branch too
The openrouter branch of _seed_from_env returns before the generic loop,
so a persisted `env:OPENROUTER_API_KEY_2` row stayed empty exactly like
the registry-provider case #103067 fixed. Fold the "declared vars + on-disk
env rows" union into one helper both branches use.
2026-09-15 11:50:48 -07:00
McClean-Newton
23afade67b fix(credential-pool): seed env-source entries not in registry tuple
Restores the env-source seeding loop in _seed_from_env that was dropped
by upstream refactors. Without this, any env-source entry whose env var
name isn't in the provider registry's hardcoded api_key_env_vars tuple
stays empty forever, gets filtered out of rotation by _available_entries,
and round-robin silently degrades to single-key behavior.

Forward-port of ba71e00db07c6263ee8d44b27dfbce2a92e6b39c onto current main
f1ccf436a2. Narrow insertion after final
env_vars list in _seed_from_env: scan only existing entries whose source
begins with env:, extract/dedupe the named variable, and let the existing
seeding path hydrate it via get_env_prefer_dotenv (preserving _Seeder,
secret-scope/profile isolation, borrowed-secret sanitization, and
suppression behavior).

Tests: 10 new tests in test_credential_pool_seed_existing_env_sources.py
covering three-key hydration, round_robin cycling, dotenv/secret-scope
resolution without os.environ bypass, duplicate dedupe, manual rows
untouched, unset remains unavailable, and sanitized persistence.
2026-09-15 11:50:48 -07:00
teknium1
398234748f refactor(agent): one Retry-After parser and one reset-grammar table feed every retry wait
Seven sites hand-rolled `float(headers.get("Retry-After"))` (anon_auth,
shared_metrics_sender, gemini_native_adapter, extract_api_error_context,
nous_rate_guard, skills_hub_github, skills_hub_clawhub x2) and silently
dropped RFC 7231 HTTP-date values that the conversation loop already honours
via agent/retry_utils.py::parse_retry_after_seconds. They now call it; per-site
caps/floors stay at the call site.

The free-text "resets in / quotaResetDelay / retry after N s" regexes lived in
two tables (agent_runtime_helpers vs credential_pool) whose "resets in"
grammars diverged: the pool accepted only integer `Nhr Nmin` while the error
context accepted h/hr/hours + m/min/minutes + s/seconds with decimals. One table
(agent/retry_utils.py::RETRY_DELAY_PATTERNS / reset_delay_from_message) using
the wider grammar, so a pooled credential's cooldown and the UI's reset time
now agree.
2026-09-13 05:09:43 -07:00
Siddharth Balyan
a2db110ccc feat(auth): Nous free tier: free inference and connectors out of the box, one command to sign in (#105258)
* feat(auth): Nous free tier core: anonymous identity minted on first use, welcome inference, shared-store scoping

A fresh install with no provider sets up a free Nous identity (anonymous auth method of the nous
provider) instead of forcing the setup wizard. The identity is persisted through the same path a
real login uses, so the resolver ladder is unchanged. Two seams differ: token acquisition
re-exchanges the anon credential (no refresh token), and routing pins the welcome host's single
model nous/welcome. One identity per shared store; nous.guest: false turns the free tier off.

* test(auth): free tier core contracts: lifecycle, resolver precedence, exchange seam, model pin

* docs(user-guide): free tier and signing in

New page explaining what a fresh install gets before any key or sign-in
(free inference on nous/welcome plus connectors), how the free tier
coexists with a user's own API key, how to sign in with hermes auth
upgrade and keep connectors, how to turn the free tier off with
nous.guest, what hermes logout does in each state, a troubleshooting
table, and a plain privacy note. Wired into the Using Hermes sidebar.

* fix(auth): logout leaves the free tier alone and clears the shared store for a real Nous account

Logging out of the free tier is a no-op: it is not a login, so nothing is cleared and the user
is told they were never signed in. Logging out of a real Nous account now also clears the
cross-profile store, so a profile logout is not silently re-adopted on the next boot.

* fix(model): switching off the free tier points at signing in, never hops providers

* Name the free tier in the gateway startup notice and tell explicit-provider installs about it once

* Render the Nous free tier as free tier on auth status, auth list, hermes status and portal info, short-circuit billing copy for it, and skip the keepalive when there is no refresh token

* fix(auth): free tier is set up where nothing is configured: resolver last rung and first-run check

Both the provider resolver's terminal rung and the CLI first-run check now try to set up the free
tier before declaring nothing configured. On a fresh install the first command lands in chat on
nous/welcome; a failed setup still falls through to the existing guidance.

* Add hermes auth upgrade: sign the free tier into a Nous account while keeping its connectors

The device-code flow runs as usual, with a promotion intent registered on the portal between the
code request and the token poll so the account that approves the code inherits the free tier's
connectors. The promotion status decides the outcome: only a completed one is followed by the
token grant, which is persisted over the free-tier singleton and the shared store. Declined,
superseded, retired and busy outcomes each print their own plain copy, and a retired identity is
cleared so the next use sets up a fresh one. User-facing text never names the free tier's internals.

* Show the Nous free tier as one picker row with nous/welcome and hide it when nous.guest is off

* fix(auth): upgrade opens the consent page for this sign-in; one mint attempt per process; forced free tier wins the first-run check

The browser leg of hermes auth upgrade now prints and opens the promotion claim URL with the
claim code, not the generic device page. A failed mint is attempted once per process so several
bootstrap sites cannot hit a closed gate or a 429 twice; a retired credential resets that so
re-minting still happens. HERMES_FORCE_GUEST is honoured ahead of the first-run provider check.

* fix(auth): pin the welcome model on the selected route, not on profile state; background setup retries after a failure

A credential-pool entry can select a paid Nous key while the profile singleton is still the free
tier. The model pin now keys on the resolved endpoint (welcome host) in agent init and /model, and
the pin in model normalization is removed since it had no route to look at. A failed background
identity setup releases its latch so a later attempt in the same process can try again.

* fix(auth): decide the Nous model together with the route on every credential-pool swap

The credential pool can move a Nous agent between the welcome host and the portal host after
init. One helper, pin_model_for_route, now runs at init and inside every pool swap, so the
welcome host always carries nous/welcome and a paid endpoint always keeps the caller's model.

* fix(auth): apply the route model policy on every wire mode during a pool swap; release the setup latch if the thread cannot start

* fix(auth): free-tier lifecycle takes profile then shared lock, reconciles with the shared store, persists the mint before exchanging, and clears only the identity that died

The shared store is the identity of record for a Hermes root: a profile holding a stale free-tier
identity adopts a sibling's newer sign-in instead of keeping the guest, and never overwrites the
shared account. Locks are taken in the documented order (profile, then shared). A minted credential
is stored as soon as create succeeds, so a rate-limited or timed-out exchange does not lose it and
trigger a second mint. Retiring a dead credential removes only that credential from both stores.
Guest exchange uses the resolver's canonical portal URL.

* fix(auth): a credential rotation never rewrites the conversation model; connectors honour the off switch and replace a retired free-tier credential

The welcome host serves one model, so a rotation onto it is refused for any conversation on another
model instead of silently switching that conversation to nous/welcome (the model pin applies only
when a route is first chosen). The connector token path now treats the free tier as absent when
nous.guest is false, including cached tokens, and shares the one dead-credential rule with
inference: a retired identity is replaced once rather than returning its stale token.

* fix(auth): plain login never imports the free tier as OAuth credentials; the gateway startup line reads persisted state only

A free-tier identity in the shared store is not an OAuth credential to offer for import; a real
sign-in replaces it. The gateway's startup notice now answers provider precedence from persisted
state (no token refresh at boot), so an expired free-tier token cannot stall the online message.
2026-09-11 03:45:31 +05:30
Brian Le
32a59f3bf7 feat: refresh one pooled OAuth grant from the CLI 2026-09-07 08:06:48 -07:00
Brian Le
b86fb277b8 fix: count credential selections across every pool strategy 2026-09-07 08:06:48 -07:00
Teknium
12ad29a48a refactor: isolate credential pool administration methods 2026-09-07 08:06:48 -07:00
kshitijk4poor
52cf39c908 refactor(auth): reuse _CLEAR_STATUS and the merge's None short-circuit in the reset path
Behaviour-preserving cleanup of the salvaged fix:

- reset_statuses goes back to replace(e, **_CLEAR_STATUS, ...) instead of
  re-spelling the six status fields; every other clear site uses the
  constant, so a seventh field would otherwise drift here. The
  count/new_entries/cleared_ids trio collapses to the original stale-list
  shape, and entry.failure_reason is read through the dataclass's
  __getattr__ like the rest of the module.
- Both bypass sites (write_credential_pool, _update_root_pool_rows) pass
  disk_entry=None to _merge_disk_cooldown_state for a cleared id instead of
  adding a parallel branch; the helper already returns the entry unchanged
  for a missing disk row.
- The cleared-id set in _update_root_pool_rows is built once, not per disk
  row (an Iterable argument would also have been drained on the first row).
- persist_pool_entries passes status_cleared_ids plainly; the kwargs splat
  existed only for one test fake, which now accepts the kwarg.
- Docstrings trimmed to the WHY.

Same 116 targeted tests green; reverting the two source files to main still
fails the two regression tests and passes the guard test.
2026-09-05 17:43:45 +05:30
rodrigogs
8378551311 fix(auth): make hermes auth reset actually clear a binding cooldown
Port onto the split layout: persist_pool_entries now forwards
status_cleared_ids to BOTH writers — write_credential_pool (hermes_cli/
auth.py) and the root-store row merge for single-use-refresh providers
(_update_root_pool_rows), which the original fix predates. Both skip the
disk recency merge for deliberately-cleared entries; the kwarg is only
sent when non-empty so upstream fakes with the old signature keep
working. reset_statuses clears failure_reason (lives in extra, replace()
cannot reach it) alongside the status fields.

tests/agent/test_credential_pool.py: 62 passed; refresh-race suite: 3
passed (mutation: the reset/binding tests fail with the fix stashed).
The dashboard-auth-gate / user-providers failures are pre-existing on
pristine upstream/main (reproduced with changes stashed).
2026-09-05 17:43:45 +05:30
Teknium
7a33369e81 simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
2026-09-03 14:00:59 -07:00
Teknium
c93ace77c2 simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files 2026-09-03 14:00:17 -07:00