Commit Graph

53 Commits

Author SHA1 Message Date
kshitijk4poor
56490ca109 refactor(aux): drop the now-uncalled _read_codex_access_token and retarget its test seams
The fold routed every aux Codex read through _resolve_codex_credential_and_base, leaving
_read_codex_access_token with no production callers; three test patches on it had gone inert
(including the 'should use pool token' guard). Point them at the live seams instead.
2026-09-25 21:27:06 +05:30
kshitijk4poor
5e2d55ca9c fix(codex): quota probe and /usage pool paths use the pool route base
Pool rows keep the canonical chatgpt.com URL, so the quota-restored probe
(auth_codex + CredentialPool) and the /usage tier-3 and forced-refresh
paths paired a gateway key with chatgpt.com/backend-api/wham/usage. Route
them through _codex_pool_route_base_url, the chat route's rule
(HERMES_CODEX_BASE_URL > model.base_url > row URL).

Refs #121486
2026-09-25 21:27:06 +05:30
kshitijk4poor
e326520d50 refactor(codex): one credential/route authority for aux + image paths
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
  fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
  directly when the pool yields no token (no second uncached pool load,
  no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
  is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
  helper's error fallback reads the profile-scoped override, not the raw
  process env.
- model setup flow: the confirm guards get the resolved Codex base, not
  the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
2026-09-25 21:27:06 +05:30
kshitijk4poor
e87f673faa fix(codex): send catalog/image credentials only to their own route
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.

- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
  credential actually routes to (runtime_provider._pool_entry_mode_and_url:
  env > model.base_url while the row is canonical > row URL) instead of the
  ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
  resolved with the token from hermes_cli/models.py, the CLI default-model
  swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
  one pool selection; the image plugin, _build_codex_client and the raw
  Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
  chatgpt.com; a gateway key may probe its own gateway's /models.

Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.

Addresses @andrexibiza's review on #121508.
2026-09-25 21:27:06 +05:30
ethernet
920a57e193 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-22 12:11:43 -04:00
kshitijk4poor
a25542c695 refactor(codex): keep the read-only quota guard, drop the preserve_corrupt plumbing
The principal-keyed fingerprint runs resolve_codex_runtime_credentials(read_only=True)
on every cache-only picker read, so the pool-exhausted branch must not probe the usage
endpoint or clear cooldowns from that path: one `not read_only and` guard, kept with
its test. The preserve_corrupt flag threaded through six auth loaders only suppressed
the one-time .json.corrupt sidecar copy on an unparseable store — an edge case the
same read-only resolve already hits on main via _codex_catalog — so it goes.
2026-09-22 20:44:30 +05:30
KCAYAAI
4ae78e3979 fix(models): keep Codex catalogs across token rotation
(cherry picked from commit 92b3a8af51010f0764f264a28bb45e4f19152edb)
2026-09-22 20:44:30 +05:30
ethernet
c399093de9 merge: reconcile profile-scoped routes and shared desktop backend with PM 2026-09-21 13:18:11 -04:00
teknium1
2632229bcf fix(auth): named profiles read the root auth.json again (revert #111724)
Reverts 93889b770d ("named profiles no longer inherit the root
profile's auth.json"). After the Desktop update every bot profile that had
relied on the root OpenAI Codex login failed with "No Codex credentials
stored. Run `hermes -p <bot> auth add openai-codex --type oauth`", and users
had to re-run the device-code flow once per bot (5-6 times in the field
report). Sharing one grant across profiles is the intended design: OAuth
refresh tokens are single-use, so ONE grant lives at the root, profiles
resolve it read-only, and a refresh under a profile writes the rotated chain
back to root (Codex / xAI write-through, borrowed-row pool bookkeeping,
forked-grant heal) — never a per-profile copy.

Restored: `_global_auth_file_path` / `_load_global_auth_store` fallback in
`_load_provider_state*` / `read_credential_pool` / `_provider_state_transaction`,
Codex + xAI root write-through, `credential_pool` borrowed-root persistence,
`heal_forked_single_use_oauth_grants`, `share_auth` on profile creation
(Desktop create dialog checkbox), and the docs. `profile_credential_audit.py`
(the `hermes update` "profiles without a provider" notice) is removed with it.

Kept from after #111724: `_save_codex_tokens(set_active=...)` for image gen,
the plugin-auth `status` dispatch and the external-login notice in
`hermes auth list`, and the registry-derived env-var hint in agent_init.
2026-09-21 09:55:40 -07:00
ethernet
9214174e07 fix(ci): lint legs after the main merge
- check_no_tmp_literals: resolve scratch via tempfile/os.tmpdir; the termux
  container mount point is one marked variable per script
- ruff TID251: desktop E2E fixtures may reach PM internals like tests do;
  the pm.runtime_stage ban message no longer names a module that never existed
- auth_codex: build the capped httpx stream subclass on first use so importing
  hermes_cli.auth_codex no longer forces httpx (the lazy proxy in auth_constants
  was defeated by a module-scope base class; broke lanes without httpx)
- desktop-smoke: launchApp is a parameter; the bundle-env test substitutes a
  refusing launcher instead of letting Playwright spawn a dying binary
  (3 unhandled rejections failed the tests-js lane)
2026-09-19 23:25:18 -04:00
astraltrekkin
de66512ab7 feat(auth): opt-in browser authorization-code + PKCE login for openai-codex
`hermes auth add openai-codex --browser` (or `auth.codex_login_flow: browser`)
signs in through OpenAI's authorize endpoint with PKCE and receives the code on
the loopback listener `http://localhost:1455/auth/callback` — the redirect URI
fixed by the public Codex client registration. Organizations that disable the
device-code grant could not log in at all before (#95743).

Device code stays the default and is never auto-replaced: the browser flow runs
only when the user asks for it, and when :1455 is already taken (a Codex CLI
sign-in in progress) Hermes prints why and falls back to device code instead of
failing. State is a 32-byte nonce compared in constant time; the code, verifier
and tokens are never logged or printed. Credentials land in the existing pool
add path with source `manual:loopback_pkce`, so refresh/rotation treat them like
any other independently added Codex account.

Derived from #97058 by @astraltrekkin (re-homed after the auth_codex.py split;
the fixed registered port replaces the free-port scan, and the flow is opt-in
instead of auto-selected per the maintainer's ruling).

Fixes #95743
2026-09-19 12:11:58 -07:00
teknium1
5857a6fb38 chore: merge origin/main (resolve hermes_cli/auth_codex.py) 2026-09-19 10:51:00 -07:00
teknium1
00e1a55519 fix(setup): Image Generation 'OpenAI (Codex auth)' row starts Codex sign-in and names the real auth command
Selecting Image Generation -> OpenAI (Codex auth) in `hermes setup` / `hermes tools` on a
fresh install saved `image_gen.provider=openai-codex` and printed "no configuration
needed!" without ever signing in, so the backend was unusable until the user guessed the
auth command — and the schema hint pointed at `hermes auth codex`, which does not exist.

Root cause: the row declares `env_vars: []` and only a `post_setup_hint`, a key nothing
consumes; `_configure_provider` runs a hook only for `post_setup`.

- plugins/image_gen/openai-codex: schema declares `post_setup: "openai_codex"`; the hint
  and the `auth_required` error name `hermes auth add openai-codex`.
- hermes_cli/tools_config_post_setup: `_post_setup_openai_codex` in the existing
  `_POST_SETUP_HOOKS` table (sibling of the `xai_grok` credential bootstrap). With
  credentials present it continues; otherwise it offers the device-code sign-in (or skip)
  and saves tokens with `set_active=False`, so picking an image backend never rewrites
  `model.provider` the way the model-provider login does.
- `_POST_SETUP_AUTH_READY` table replaces the `post_setup == "xai_grok"` special case in
  `provider_readiness_status`, so any credential-bootstrap row reports ready/needs_auth
  from the auth store.
- `_save_codex_tokens(set_active=...)` mirrors `_save_xai_oauth_tokens`.

Live probe (real `_configure_provider`, real plugin row, temp HERMES_HOME, OAuth start
stubbed with a recorder): before — post_setup=None, OAuth fired [], hint `hermes auth codex`;
after — no creds: OAuth fired once, logged in, model.provider untouched; with creds: OAuth
not fired.

Fixes #102144
Salvages #102165 (@liuhao1024) — superseded: same direction (post_setup hook), redone on
the split tools_config siblings without calling `_login_openai_codex`, which would have
switched the main model provider.
2026-09-19 10:05:26 -07:00
teknium1
4c3d31eb4d fix(auth): count additional_rate_limits in the Codex quota-restored probe
The probe read only the account-wide `rate_limit` windows, so a model-scoped
allowance under `additional_rate_limits` sitting at 100% (the #97315 reproduction:
top-level pool open, model-specific allowance exhausted) reported "restored" and
lifted the cooldown, feeding one 429 per turn. Fold every additional window into
`worst_used`.
2026-09-19 09:40:56 -07:00
teknium1
a70053d9d7 fix(auth): throttle the pre-probe Codex refresh and pin the pool-side wiring
Review follow-up on the #89415 pre-probe refresh:

- `_refresh_expired_codex_probe_token` ran ahead of `_probe_codex_quota_restored`'s
  5-minute per-token cache with no throttle of its own, so a frozen entry whose
  refresh keeps failing (revoked grant, network down) POSTed to the token endpoint
  under the pool lock on every credential selection. It now consults the same
  per-token probe cache before refreshing and reserves the stale token's slot on
  failure, restoring the pre-PR budget of at most one network call per interval.
- The pool-side refresh adopted the rotated pair into the pool row only; the next
  selection's `_sync_entry_from_auth_store` saw a differing singleton and re-adopted
  the consumed pair (clearing the cooldown with it). Stamp `last_refresh` and write
  the pair back to `providers.openai-codex` exactly as `_refresh_entry` does.
- Tests: one resolver-side refresh test is replaced by one that drives
  `CredentialPool.select()` (probe receives the refreshed bearer, both disk sides
  hold the rotated pair across a second selection) and a control that repeated
  selections with a failing refresh make a single refresh attempt.
2026-09-19 09:40:56 -07:00
Teknium
63bcdcaa43 fix(auth): refresh an expired Codex token before the quota-restored probe
_probe_codex_quota_restored was called with the exhausted pool entry's
stored access token. Exhausted entries are skipped by the proactive refresh
chain, so for any cooldown longer than the access-token lifetime (the
weekly-cap case the probe exists for) that token has expired: the usage
endpoint answers 401 token_expired, the probe maps it to None, and the
cooldown is kept. A mid-cooldown credit top-up or plan upgrade was
therefore never discovered until last_error_reset_at elapsed, even though
the refresh token was still valid (#89415).

Add _refresh_expired_codex_probe_token: when the stored token is expired
and a refresh token exists, rotate the pair via refresh_codex_oauth_pure
(cooldown untouched), persist it (refresh tokens are single-use), and probe
with the live token. Both probe callers use it: the pool-only resolve path
(_probe_codex_pool_entry_quota_restored) and
CredentialPool._codex_quota_restored_upstream. A probe that still reports
100% keeps the bench; the rotated tokens are persisted either way.

Fixes #89415
2026-09-19 09:40:56 -07:00
fangliquanflq
07135710db fix(auth): normalise millisecond last_error_reset_at in Codex pool selection
_pool_codex_access_token compared the persisted last_error_reset_at raw, so
a millisecond epoch read as far-future and the entry was skipped, while
_codex_pool_rate_limit_status normalised the same value to seconds and read
it as elapsed. A usable pool credential became invisible to both paths and
resolve_codex_runtime_credentials raised "quota exhausted (429); retry
after Ns" off the sibling that really was exhausted (#103349).

Use the shared _parse_absolute_timestamp normaliser in selection so the two
readers can never disagree. Ported from PR #103356.

Fixes #103349
2026-09-19 09:40:56 -07:00
teknium1
e5258804df fix: build the dashboard Codex poll client through the shared capped builder
Review follow-up on the 1 MiB auth-body cap:

- `web_routers.oauth`: replace `_codex_body_cap()` with `_codex_client(httpx)`, the one
  builder for the dashboard device-login flow. `_codex_post` and `_codex_poll_authorization`
  both go through it, so the poll loop's long-lived client cannot silently lose the cap.
  The existing oversized-body test now also drives `_codex_poll_authorization`; dropping the
  builder from the poll line fails it (DID NOT RAISE AuthError), restoring it passes.
- `auth_codex._codex_http_client`: drop the `event_hooks` merge — none of its four callers
  pass `event_hooks`, so the client is built with the response hook directly.
2026-09-19 09:33:05 -07:00
teknium1
c6d12dafa0 fix: cap Codex OAuth/device-auth response bodies at 1 MiB in CLI and dashboard
Every Codex auth request (device-code usercode/token, token exchange, token refresh) read a
200 response with an unbounded `client.post(...).json()`, so a hostile or broken endpoint or
proxy answering with megabytes of "JSON" was fully buffered and parsed before the flow could
fail. Real payloads are a few hundred bytes.

`_codex_http_client` now installs an httpx response hook that swaps the body stream for a
capped one: reading past 1 MiB closes the connection and raises a shaped AuthError
(`codex_auth_response_too_large`) instead of buffering on. A hook is used rather than a
streamed reader per call site so `client.post` stays the seam every existing test stubs, and
the cap covers all statuses uniformly; 429 / non-200 classification is untouched. The
dashboard worker builds its own client from an injected `httpx` module, so it shares the hook
(`_codex_body_cap`) rather than the factory.

Fixes #55253
Salvages #55254

Co-authored-by: luyifan <al3060388206@gmail.com>
2026-09-19 09:33:05 -07:00
teknium1
7471586ad0 fix: send the Codex residency header on every JWT-derived Codex request
The Codex models catalog probe (agent/model_metadata.py), the picker catalog
fetch (hermes_cli/codex_models.py), the quota-restored probe
(hermes_cli/auth_codex.py) and the /usage dashboard call
(agent/account_usage.py) each re-decoded the OAuth JWT for
ChatGPT-Account-ID and would 401 on residency-enforced workspaces exactly
like the chat client did (#23896). They now share
agent.codex_headers.codex_account_headers, which emits the account and
x-openai-internal-codex-residency headers from one decode; the three
duplicate `_extract_chatgpt_account_id` decoders are gone.

account_usage keeps auth.json's account_id as the winner over the JWT claim
(pool-only credentials still omit it); header casing unifies on the
codex-rs canonical `ChatGPT-Account-ID` (HTTP header names are
case-insensitive on the wire; the two tests asserting the old casing follow).
2026-09-19 09:28:13 -07:00
teknium1
29bc6343d3 fix(auth): read-only Codex reads take no store lock; /model picker never refreshes
The expiring-singleton status test never reached the singleton resolver:
`load_pool("openai-codex")` mirrors the singleton as a `device_code` pool
entry, so with a still-valid token `pool.peek` answered and the
`get_codex_auth_status()` wiring under test was never exercised (the test
stayed green with `resolve=` reverted). The token is now already expired,
the status result is asserted to come from `hermes-auth-store`, and the
`read_only` beats `force_refresh` call is a secondary assertion on the
resolver itself.

`resolve_codex_runtime_credentials(read_only=True)` reads the store without
`_auth_store_lock`: `_save_auth_store` replaces auth.json atomically, so a
lock-free read never sees a torn file, and materialising `auth.lock` is
itself a write a diagnostic must not make (#68004). Once the pool has
mirrored the singleton a status read leaves the HERMES_HOME manifest
byte-identical.

`_codex_catalog` (the `/model` picker) now reports the stored login
read-only, matching the intent stated on `read_only`; an expired stored
token yields the hardcoded catalog until the runtime lease refreshes it.
2026-09-19 09:24:25 -07:00
teknium1
d43c9622f3 fix(auth): status and doctor Codex reads never adopt, refresh or persist credentials
`hermes status` / `hermes doctor` / the dashboard cards and the `/model` picker
call `get_codex_auth_status()`, whose singleton fallback ran
`resolve_codex_runtime_credentials()` with runtime defaults: a store missing
its refresh_token imported the Codex CLI's single-use token family, and an
expiring token was refreshed and written back. A diagnostic that spends or
borrows a rotating refresh token logs the other program out (#68004).

`resolve_codex_runtime_credentials(read_only=True)` reports the stored state
as-is (no CLI adoption, no refresh, no pool forced refresh, no write) and wins
over `force_refresh`; the Codex status snapshot uses it, and the xAI OAuth
snapshot passes `refresh_if_expiring=False` for the same reason. The pool side
was already an observation (`peek`, 964fbaae2c).

Superseded #68224 (@GauravPatil2515) — same mechanism, re-done on the current
`auth_codex` layout.

Co-authored-by: Gaurav Patil <gauravpatil2516@gmail.com>
2026-09-19 09:24:25 -07:00
teknium1
20751ae92c chore: stack on #115730 to resolve hermes_cli/auth_codex.py conflict
resolve_codex_runtime_credentials: read_only does the lockless read (no recovery
follows); otherwise take _auth_store_lock(), observe the access token and read
with _lock=False so the compare-and-swap recovery still applies.
2026-09-19 03:36:27 -07:00
teknium1
84c77f47e2 fix(auth): read-only Codex reads take no store lock; /model picker never refreshes
The expiring-singleton status test never reached the singleton resolver:
`load_pool("openai-codex")` mirrors the singleton as a `device_code` pool
entry, so with a still-valid token `pool.peek` answered and the
`get_codex_auth_status()` wiring under test was never exercised (the test
stayed green with `resolve=` reverted). The token is now already expired,
the status result is asserted to come from `hermes-auth-store`, and the
`read_only` beats `force_refresh` call is a secondary assertion on the
resolver itself.

`resolve_codex_runtime_credentials(read_only=True)` reads the store without
`_auth_store_lock`: `_save_auth_store` replaces auth.json atomically, so a
lock-free read never sees a torn file, and materialising `auth.lock` is
itself a write a diagnostic must not make (#68004). Once the pool has
mirrored the singleton a status read leaves the HERMES_HOME manifest
byte-identical.

`_codex_catalog` (the `/model` picker) now reports the stored login
read-only, matching the intent stated on `read_only`; an expired stored
token yields the hardcoded catalog until the runtime lease refreshes it.
2026-09-19 01:16:42 -07:00
Tranquil-Flow
a745c35282 fix: Codex CLI recovery no longer replaces a credential from another workspace
Automatic recovery from ~/.codex/auth.json saved the imported pair blindly.
A Codex Desktop/CLI login into a different ChatGPT workspace (same OpenAI
user) therefore silently replaced the Hermes credential being repaired, and
the pool aliases synced with it, so both pool slots pointed at the other
workspace while labels and order looked unchanged (#73667). The save also
raced an explicit re-auth landing between the failed read and the write.

Thread the access_token the caller observed (None on the missing-token path)
into _recover_codex_tokens_from_cli. Under _provider_state_transaction the
save is a compare-and-swap: skip when the stored token no longer equals the
observed one (a concurrent re-auth wins), refuse with a redacted warning and
a relogin hint when the observed and imported tokens carry a different known
principal (agent/credential_pool._codex_principal_identity), else persist.
Unknown identity on either side keeps today's repair behaviour.

Salvage of #73677 (@Tranquil-Flow), reworked per review: the identity guard
and the caller-observed CAS are kept; the second unlocked observation helper
and the 11-test suite are replaced by two invariant tests.
2026-09-19 00:59:44 -07:00
teknium1
784f94fb08 fix(auth): status and doctor Codex reads never adopt, refresh or persist credentials
`hermes status` / `hermes doctor` / the dashboard cards and the `/model` picker
call `get_codex_auth_status()`, whose singleton fallback ran
`resolve_codex_runtime_credentials()` with runtime defaults: a store missing
its refresh_token imported the Codex CLI's single-use token family, and an
expiring token was refreshed and written back. A diagnostic that spends or
borrows a rotating refresh token logs the other program out (#68004).

`resolve_codex_runtime_credentials(read_only=True)` reports the stored state
as-is (no CLI adoption, no refresh, no pool forced refresh, no write) and wins
over `force_refresh`; the Codex status snapshot uses it, and the xAI OAuth
snapshot passes `refresh_if_expiring=False` for the same reason. The pool side
was already an observation (`peek`, 964fbaae2c).

Superseded #68224 (@GauravPatil2515) — same mechanism, re-done on the current
`auth_codex` layout.

Co-authored-by: Gaurav Patil <gauravpatil2516@gmail.com>
2026-09-18 23:34:57 -07:00
teknium1
6c7f693473 feat(auth): opt out of borrowing Codex CLI / Claude Code logins (auth.adopt_external_logins)
Hermes adopts and refreshes the Codex CLI (~/.codex/auth.json) and Claude Code
(~/.claude/.credentials.json) logins automatically whenever its own login is
missing or its refresh is rejected. Both providers hand out single-use, rotating
refresh tokens, so after an adoption two programs hold one token family and
whichever refreshes first logs the other out (#113023). #113816 made the dead
logins visible; this adds the switch the reporter asked for.

- `auth.adopt_external_logins` (config.yaml, default true — nothing changes for
  current users). When false:
  - `read_claude_code_credentials()` — the only reader of the borrowed Claude
    Code login — returns None, so the resolver fallback, the expired-token
    refresh, the pool seed/sync and the auxiliary 401 refresher never touch the
    file; the pool prunes a `claude_code` row an earlier adopting process
    persisted.
  - `_recover_codex_tokens_from_cli` returns None for both automatic recovery
    paths (rejected refresh, half-empty singleton); the real AuthError is
    surfaced instead. The interactive import offer in `hermes auth add
    openai-codex` still asks first and is unaffected.
  - One INFO line per process the first time adoption would have happened;
    `hermes auth list` / `hermes auth status anthropic|openai-codex` print the
    same line so the missing borrowed row is explained.
- Docs: security.md "Borrowed CLI logins" section; providers.md cross-links.

Live (temp HERMES_HOME + CLAUDE_CONFIG_DIR + CODEX_HOME, loopback logging token
endpoint): main with the key set to false still POSTed the Claude Code refresh,
rewrote the file's refresh token, seeded a claude_code pool row and adopted the
Codex CLI pair; on this branch the false arm makes zero Anthropic refresh POSTs,
leaves both external files byte-identical, seeds no row and prints the notice,
while the default arm is byte-for-byte today's behaviour.
2026-09-18 20:55:24 -07:00
teknium1
69deb43372 fix(auth): Codex credential and relogin copy name the failing profile's own sign-in
The three missing-token messages, the refresh_token_reused copy and the
relogin suffix format_auth_error appends all said a bare `hermes auth` /
`hermes model`, which from a named profile re-signs the ROOT store
(93889b770d) — the same loop #114012 measured on the refresh-rejected path.
They now go through oauth_relogin_command('openai-codex') / profile_cli_selector()
so a profile user is told `hermes -p <profile> auth add openai-codex --type oauth`.

Part of #114012
2026-09-18 10:00:29 -07:00
teknium1
227384e332 fix(auth): drop dead code from the Codex login retry; trim to two invariant tests
Follow-up to the salvaged #114614 pick:

- `_codex_login_post`: the `for` loop ended in an unreachable `raise … # pragma: no cover`
  (needed only to satisfy the return type). A `while True` with the terminal condition
  folded into the except branch has no dead path and the same three-attempt bound.
- `_codex_poll_authorization_code`: the pick duplicated the `except KeyboardInterrupt`
  handler; the second copy was unreachable.
- Tests: six change-detector tests collapsed into two invariants (poll survives blips but
  never retries a non-transport exception; one-shot POST retries once and keeps the typed
  AuthError + TLS hint + cause chain at the cap). The existing
  test_codex_device_login_ssl_hint.py still pins the poll's terminal hint path.
- Docs: providers.md notes that a single dropped connection during device login is no
  longer fatal.
2026-09-18 09:18:00 -07:00
liuhao1024
82e5db63bc fix(cli): retry transient transport blips in the Codex device-code login
One dropped connection (e.g. an SSL EOF mid-poll) aborted the whole
device-code login and wasted the browser approval the user had already
completed (#114610). The poll loop now treats transport-level errors as
non-fatal while the device code is still valid (failing only after 6
consecutive blips), and the one-shot device-login POSTs retry twice with
a small linear backoff. Persistent failures keep today's typed AuthError
with the original cause chained and the SSL interop hint; non-transport
exceptions still surface immediately.

Salvaged from #114614. Earlier carriers of the same fix: #78441 (first, on the pre-split
hermes_cli/auth.py loop) and #114611.

Co-authored-by: yunke98888-gif <yunke98888@gmail.com>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:18:00 -07:00
teknium1
93889b770d fix(auth): named profiles no longer inherit the root profile's auth.json (#111724)
A named profile with no credentials of its own silently resolved the root
profile's provider state and credential pool, and a token refresh inside
that profile (xAI, Codex, Anthropic PKCE, Nous) wrote the rotated chain back
into the root store. An isolated service profile therefore acted, and
rotated tokens, as the owner with no way to switch it off.

Maintainer ruling: profiles without credentials are asked to set a provider,
never handed another profile's auth. Profiles are independent islands.

What changes
- `hermes_cli/auth.py`: `_load_provider_state*`, `read_credential_pool` and
  `_provider_state_transaction` read the active store only; the global-root
  resolver, its mtime memo and `_persist_provider_state_to_store` are gone.
- xAI / Codex / Nous-guest / pool refresh paths persist to the active store;
  the root write-through, the borrowed-row bookkeeping
  (`_borrowed_root_ids`, `persist_pool_entries`, `_update_root_pool_rows`)
  and the forked-grant heal are removed. `_write_hermes_oauth_credentials`
  loses its root `target`.
- `resolve_provider` / `agent_init` name the profile in the
  no-provider error and print `hermes -p <name> model` guidance.
- `hermes update` prints a one-time notice listing every named profile that
  has no provider of its own (`hermes_cli/profile_credential_audit.py`) so a
  bot never goes quiet unannounced.
- Desktop create dialog: the "Share keys & accounts" checkbox described the
  removed inheritance; it now mirrors API keys (`mirror_credentials`) and
  says OAuth logins need a sign-in. `share_auth` is accepted from older
  clients and ignored; `ProfileMirrored.auth` is a bool again.
- Docs: profiles.md, multi-profile-gateways.md isolation table,
  hermes_cli/AGENTS.md.

The fallback was added in 33bf5f62 so kanban/cron workers under a named
profile did not die with "No LLM provider configured" when the credential
lived only at root; that convenience is exactly the isolation hole the
ruling closes, and `--clone` / dashboard mirroring still copy API keys.

Tests: fallback/write-through/heal pins deleted; 4 invariants proven red on
base (profile never reads root; profile refresh never writes root; Nous
connector gate reads only the profile store; Anthropic pool never borrows or
rotates the root grant, root control still refreshes).
2026-09-16 14:34:59 -07:00
teknium1
417b707f59 fix: refresh the credential that 401'd on the Codex /usage retry
Two holes in the forced retry. With an explicit live-agent api_key the
retry dropped it and re-resolved from the singleton/pool, so a 401 on pool
entry B rendered account A's usage — the cross-account leak the surrounding
comments say the tiering prevents. The retry now force-refreshes the pool
entry that issued the key (or the singleton when it IS the singleton token)
and fails open when neither matches.

In a pool-only setup (empty singleton) resolve_codex_runtime_credentials
returned the pool token before consulting force_refresh, so the retry resent
the identical revoked bearer. The pool-only branch now rotates the pool
entry through try_refresh_matching when force_refresh is set.

Also: when the forced refresh itself raises inside redeem_codex_reset_credit,
surface the original 401 (re-login hint) instead of the refresh error.

Review finding: explicit-key force refresh re-resolved another account; pool-only tier ignored force_refresh.
2026-09-15 04:53:38 -07:00
kshitijk4poor
ca6b189a7c fix(codex): both transaction locks wait out the endpoint; adopt only a complete pair
Review follow-ups on the refresh transaction:
- `_provider_state_transaction` takes a `timeout_seconds` applied to BOTH the
  active and the root lock. The refresh passes max(default, refresh timeout
  + 5 s); before, only the profile lock used that budget and root's lock kept
  the 15 s default, so the waiting profile raised TimeoutError instead of
  adopting whenever the peer's POST ran long. The regression test now holds
  the endpoint past the lock floor and fails without the passthrough.
- Peer adoption requires a stored access token as well as a rotated refresh
  token; an incomplete stored pair falls through to the refresh.
- `_save_codex_tokens` keeps its body: the per-path lock is reentrant, so the
  refresh calls it inside the open transaction (as the CLI-recovery path
  already did) instead of a split-out helper.
2026-09-14 20:38:07 +05:30
kshitijk4poor
e117e792b6 fix(codex): run the refresh inside the source store's transaction so shared-root peers adopt, not replay
Follow-up to #110024 (ehz0ah's review thread). `_refresh_codex_auth_tokens` POSTed the
single-use refresh token to OpenAI first and only then entered
`_provider_state_transaction("openai-codex")` for the write-back, so root's lock covered the
save alone. `resolve_codex_runtime_credentials` holds only the caller's own profile lock, so
two profiles borrowing the same ROOT grant could both submit `old-rt`; last root save won and
OpenAI answered `refresh_token_reused` / revoked the family — the failure #87503 exists to
prevent.

The transaction now spans re-read -> endpoint refresh -> write-back:
- enter `_provider_state_transaction` first; the yielded state is root's, re-read under root's
  lock. If its refresh token already differs from the one we were about to submit, a peer
  rotated it: adopt the stored pair and return without touching the endpoint.
- otherwise POST and write back through `_store_codex_tokens_in`, the body of
  `_save_codex_tokens` split out so it can run inside an already-open transaction.
  `_save_codex_tokens` keeps its signature for the login/import/CLI-recovery callers.

Holding the advisory flock across the network call is safe here and already the established
shape: `resolve_codex_runtime_credentials` holds the active-store lock across the same POST,
and every waiter's timeout is `max(AUTH_LOCK_TIMEOUT_SECONDS, refresh_timeout + 5)`, i.e. it
outlives one full endpoint timeout. `_load_auth_store` readers never take the lock, so
readers are not blocked; `_file_lock` is reentrant per thread per path, so the nested
transaction inside the caller's lock and the CLI-recovery save inside the transaction both
re-enter cleanly. A release-POST-retake variant would reopen the window it is meant to close.

Test: two refreshers with the same stale pre-read pair against a rotate-once endpoint that
rejects any replay — the endpoint sees `old-rt` exactly once, both callers end with the
rotated pair, root holds it, the profile store stays unshadowed. Red on origin/main
(`refresh_token_reused` surfaces for the second caller).
2026-09-14 20:38:07 +05:30
kshitijk4poor
66ddd5f83c fix(auth): only a Codex token refresh writes through to root
Following the grant's source on every save made a fresh device-code
login (or `hermes auth import`) under a profile that had been borrowing
root's Codex grant overwrite root's account instead of creating the
profile's own. Redirecting a save into another file is the exception, so
it is opt-in: the refresh path passes write_through=True; login, import
and recovery keep saving locally. The two save branches collapse into one
(store, path, set_active) triple.

Test: root discovery on Windows comes from LOCALAPPDATA — set it so the
fixture's root is the resolved root on every host.
2026-09-14 19:49:36 +05:30
liuhao1024
6bd29f26f6 fix(auth): write profile-refreshed Codex tokens through to the global store
Codex refresh tokens are single-use with rotation-family reuse
detection. _save_codex_tokens resolved the state via the profile's
root fallback but always persisted into the ACTIVE (profile) store, so
a profile-scoped refresh left the global store holding the consumed
refresh token — the next process to read it replayed it and OpenAI
revoked the whole rotation family, forcing a manual device-code
re-auth (#87503; observed four times on one multi-profile deployment).

Mirror the xAI source-aware save (#43589/#74339): resolve the state
with _load_provider_state_with_source; when the grant came from the
global root, write the rotated chain back to root only — singleton AND
credential_pool entries, under the root store's own lock, without
creating a shadowing profile key. Best-effort, with the same pytest
seat belt as the xAI path.
Fixes #87503
2026-09-14 19:49:36 +05:30
kshitij
ea42884e99 refactor(auth): codex cooldown clear reuses the pool ownership rule
clear_codex_pool_quota_cooldowns decided "borrow the root?" inside a nested
closure via a tri-state Optional[int] return (None = no rows), which forced
`cleared or 0` and duplicated the rule persist_pool_entries already owns.
Decide once with _profile_owns_pool_provider + _borrowed_single_use_pool_root,
then lock/load/clear/save exactly one store. Behaviour is unchanged for every
(mode x profile rows x root rows) cell; the pre-lock decision races only a
concurrent `hermes auth add` in the profile, whose fresh rows carry no cooldown.

Adds the missing negative invariant: a profile that OWNS Codex rows never has
the root store touched (0 cleared, root byte-identical).
2026-09-12 12:02:55 +05:30
kshitijk4poor
fe6351000f refactor(auth): codex pool reads call read_credential_pool directly; cooldown clear decides ownership under the lock
`_read_codex_pool_entries` had become a pass-through whose only residual
behaviour was taking the active-store lock around a read every other
`read_credential_pool` caller does unlocked (and which never covered the
root file it fell back to). Both consumers now call the helper directly.

`clear_codex_pool_quota_cooldowns` decided "profile owns rows?" with an
unlocked pre-read, then re-read the same file under the lock — a TOCTOU
against a concurrent `hermes auth add` and a wasted parse. It now tries the
active store under its lock and falls back to the borrowed root only when
that store has no codex rows.
2026-09-12 12:02:55 +05:30
kshitijk4poor
65c9d33b14 fix(auth): a borrowed Codex pool clears its quota cooldown in the root store
With reads now inheriting the global-root pool, a profile hitting a stale
root cooldown probes quota, sees it restored, and calls
clear_codex_pool_quota_cooldowns() — which only ever edited the (empty)
profile store, so the next resolve raised quota_exhausted again forever.
Pick the store the same way agent/credential_pool.py persists borrowed
rows (_profile_owns_pool_provider / _borrowed_single_use_pool_root) and
lock/save against that path. The fallback test now also binds
profile-wins precedence; one new test pins the root write.
2026-09-12 12:02:55 +05:30
nekwo
3767c2eacd fix(auth): inherit global Codex pool in profiles 2026-09-12 12:02:55 +05:30
teknium1
d3dcc064df fix(cli): detect ssl.SSLError by type in the Codex login hint; trim tests; add the openssl.cnf snippet to docs
- _ssl_interop_hint: also match ssl.SSLError instances (and one level of
  __cause__/__context__) plus the bare UNEXPECTED_EOF marker, so an
  SSLEOFError whose text httpx did not repeat still gets the hint. The
  hint now names the TLS 1.2 diagnostic and links the providers docs
  note instead of an issue number.
- tests: 3 -> 2 invariants (parametrized login_post/poll SSL case keeps
  the raw text + hint + cause; a plain httpx timeout gets no hint).
- docs: providers.md Codex note carries the reporter's exact openssl.cnf
  classic-groups snippet (EN + existing zh-Hans copy).

Refs #106384. The TLS max-version cap itself stays PR #44392's scope.
2026-09-09 10:14:58 -07:00
liuhao1024
d8eb177c93 fix(cli): keep SSL detail and add middlebox hint on Codex device-login transport errors
Device-login requests on networks whose middlebox rejects the larger
TLS 1.3 ClientHello sent by OpenSSL 3.5+ (post-quantum hybrid groups)
fail with SSLEOFError / handshake timeouts while curl still works, so
they masquerade as a Codex outage (#106384). The polling loop let the
raw httpx error escape unshaped, and _codex_login_post dropped the
exception chain and gave no actionable hint.

- add _ssl_interop_hint() applied to both device-login transport paths
- re-raise _codex_login_post failures with 'from exc' to preserve cause
- wrap the poll POST so transport failures become a shaped AuthError
  (device_code_poll_error) carrying the SSL detail and OPENSSL_CONF
  workaround hint; KeyboardInterrupt handling is unchanged

(cherry picked from commit 8cd94c36ce8437db5b00290b9edbedcd2116c02c)
2026-09-09 10:14:58 -07:00
Teknium
00a3cfe5c9 simplify(compat): hermes_cli split-module docstrings — drop 're-exported there' compat pointers (9 files, comments only) 2026-09-03 13:08:40 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
45bce4faef refactor(hermes_cli): auth_codex — collapse guard ladders, reuse _codex_pool_dicts in pool sync, drop dead resp-None branches 2026-09-02 23:15:03 -07:00
Teknium
55a4b75347 refactor(hermes_cli): auth group B — squeeze body blanks, fold repeated AuthError ctor 2026-09-02 22:27:46 -07:00
Teknium
31e6c6fab4 refactor(hermes_cli): auth_nous/codex — unify refresh_nous_oauth_pure into from_state, codex login POST helper 2026-09-02 22:18:03 -07:00
Teknium
856fd5613c refactor(hermes_cli): auth group B — pack signatures, move heal notice into _HealPass 2026-09-02 21:35:49 -07:00
Teknium
34a6f963f8 refactor(hermes_cli): auth_nous/codex/device_flow/oauth_grants — AST-neutral layout compaction 2026-09-02 21:02:27 -07:00
Teknium
c86e74550a refactor(hermes_cli): auth_codex — dedupe pool reads, reuse _parse_absolute_timestamp, collapse guards 2026-09-02 20:30:06 -07:00