1Password/Bitwarden items can carry several websites, but both backends
collapsed the item's urls[]/uris[] to the first origin that normalizes,
so browser_vault_fill refused every other explicitly saved origin with
origin_mismatch. Reordering the URLs in the manager just moved which
single origin worked.
VaultItemMeta now carries allowed_origins (every normalized, deduped
web origin; origin stays the first/primary one). Fill matching stays
exact-origin against that list — no wildcard, parent-domain or subdomain
inference — and the in-page synchronous check pins the origin actually
matched via build_fill_js(expected_origin=page_origin). App URIs such as
androidapp:// never widen the fill set.
_repair_object_shape recursed over every dict value as a schema node,
including the properties map itself. When one of the map's keys was
literally named properties/required, the missing-type heuristic fired on
the map and injected a bogus "type": "object" string as a parameter,
400ing the whole tool array on strict providers.
Recurse into map values only for the mapping-valued keywords
(properties, patternProperties, $defs, definitions, dependentSchemas),
matching the in-tree precedent in _rewrite_local_refs.
Fixes#110530
Two holes in the forced retry. With an explicit live-agent api_key the
retry dropped it and re-resolved from the singleton/pool, so a 401 on pool
entry B rendered account A's usage — the cross-account leak the surrounding
comments say the tiering prevents. The retry now force-refreshes the pool
entry that issued the key (or the singleton when it IS the singleton token)
and fails open when neither matches.
In a pool-only setup (empty singleton) resolve_codex_runtime_credentials
returned the pool token before consulting force_refresh, so the retry resent
the identical revoked bearer. The pool-only branch now rotates the pool
entry through try_refresh_matching when force_refresh is set.
Also: when the forced refresh itself raises inside redeem_codex_reset_credit,
surface the original 401 (re-login hint) instead of the refresh error.
Review finding: explicit-key force refresh re-resolved another account; pool-only tier ignored force_refresh.
The raw_codex branch of _resolve_openai_codex_branch (main agent built without
explicit creds via agent_init._routed_client_kwargs, and the mid-turn fallback
chain) still hardcoded the official Codex endpoint while the pooled/aux/singleton
paths honoured the override, so a proxy user's main agent silently bypassed it.
Hoist the profile-scoped read into _codex_base_url_override and use it in both
builders; the Cloudflare identity headers follow the resolved base_url.
Review finding: raw_codex client ignored HERMES_CODEX_BASE_URL while pooled/aux/singleton honoured it.
The aux Codex client already resolves its API-key env vars via _scoped_key_env
so a multiplexed profile never borrows a sibling's value; the endpoint override
follows the same rule instead of a bare os.getenv.
_to_async_client rebuilds default_headers from scratch and dropped the blank
Authorization set by _create_openai_client, so every async aux call (the path
aux tasks actually take) for a free-tier OpenCode model still shipped the
keyless placeholder as a bearer and 401'd. Re-apply the same keyless policy on
the async twin.
Review finding: fourth client builder (_to_async_client) missed the keyless header policy.
A free slug (deepseek-v4-flash-free) re-resolved under the paid `opencode` profile — via
`/model --provider opencode`, a fallback switch, or a config default — resolves to the
keyless placeholder, but the primary client builder blanked the Authorization header only
when `agent.provider == "opencode-free"`. The placeholder then went on the wire as a
bearer, the Zen relay 401'd every request, and the credential pool (size 0) had nothing to
rotate: the loop reported in #110831.
Key the header policy on the placeholder as well as the provider, matching agent_init and
auxiliary_client which already do so.
build_model_state dropped every inventory row whose slug name matched a
named `providers:` entry and relabelled the current provider `custom:<key>`
unconditionally. When a `providers:` key shadows a canonical provider name
(providers.openrouter: -> proxy with a models: list), a session genuinely
running on openrouter.ai lost all canonical rows and its current id became
custom:openrouter:..., which resolves to the proxy base_url — picking the
"current" row silently re-routed the session to a different endpoint.
Only rows flagged is_user_defined are replaced by the named catalogs now,
and the current provider is promoted to custom:<key> only when the session
base_url matches that entry's api_url or no canonical row for the raw key
exists (the reported relay/MixedCase class keeps its custom:<key> id).
Review finding: named entry shadowing a canonical provider dropped canonical rows and re-routed the current session to the proxy.
The 700ms re-poll in local-runtime-jobs.ts reached for `window.setTimeout`;
when a test left a running job in the store, the timer fired after vitest had
torn down jsdom and threw `ReferenceError: window is not defined` as an
unhandled error, failing the whole check:test:ui lane on main even though all
7808 tests passed. Use the global timer functions (identical in the renderer)
and drain the running-job state after each local-models-settings test so no
poll is armed across test boundaries.
Stored tool call ids are minted per turn (terminal:0, terminal:1…), so the
same id recurs on later turns of one session. The Responses adapter replayed
each pair verbatim, putting the same call_id on the wire N times; strict
validators (opencode Console, OpenAI) reject the request with 400
"Duplicate function_call_output for call_id" and every retry fails the same
way (#102629, #111231 bug 1).
Every occurrence past the first now gets a `_dup<n>` wire id and the matching
tool output pops the id its function_call was given, in call order, so pairs
stay intact and stored history is untouched.
Re-applied onto the facade split (helpers moved into _replay_tool_call_items /
_tool_output_items) from PR #102634.
With no fallback_chain (or an exhausted one) the ladder tail returned the
_RERAISE_ORIGINAL sentinel and call_llm's bare `raise` re-raised the ORIGINAL
error. After a 401 that the refresh rung already healed, the retry's actionable
"404 requires available credits" was hidden behind "401 Unauthorized", steering
the user to re-login instead of top-up. Raise the narrowed `first_err` from the
ladder tail instead and drop the sentinel; the pool-rotation rung shares the
same tail so it is covered by the same change.
Review finding: exhausted no-chain path re-raised the healed 401 instead of the retry's 404 credits error.
Review feedback on the fall-through fix: the accept boundary and the
retry-fails-with-a-connection-error path through the credential rung were
reachable but untested, and two comments described behavior the code does
not have. Tests and comments only, no behavior change.
Both auth-refresh rungs in the auxiliary recovery ladder performed their
post-refresh retry with a bare `yield` inside a `return` statement, so an
exception from that retry escaped `_aux_recovery_ladder` instead of being
converted into `(None, err)` by `_rung()`. The ladder therefore never reached
`_ladder_provider_fallback` and the configured
`auxiliary.<task>.fallback_chain` was silently skipped.
Observable case: a stale Nous runtime token makes the auxiliary request return
401; the ladder refreshes the credential and retries; when the retried model is
out of credits the resulting 404 escapes the ladder and the task fails (context
compression degrades to truncation) even though a healthy fallback_chain is
configured. With a fresh token the same 404 takes the payment path, which
already falls through correctly.
Guard both rungs with `_rung()` - the helper the sibling payment rung and the
credential-pool rung below already use - so a failed retry resumes the ladder.
The hand-rolled `_AGGREGATOR_MODEL_PREFIX_PROVIDERS = {"openrouter"}` duplicated
aggregator knowledge already owned by hermes_cli.providers.is_routing_aggregator,
so other vendor-prefixed aggregators (ai-gateway, kilocode, custom:* proxies)
kept bypassing the target-profile veto. Use the existing predicate instead of a
second frozenset.
Review finding: aggregator allowlist duplicated is_routing_aggregator and missed other vendor-prefixed aggregators.
Copilot enterprise custom models (BYOK) expose catalog ids shaped
owner/sub/model. normalize_copilot_model_id() only tried the full id and the
id minus its FIRST segment, and returned the stripped form when neither
matched the (unreachable) catalog - corrupting those ids into sub/model,
which the CLI's second normalization pass then stripped again to model. The
Copilot API answered HTTP 400 model_not_supported on every call.
Only accept the strip guess when the remainder is itself a flat id: a result
that still contains "/" cannot be a Copilot id.
Fixes#110597
_is_structured_output_rejection matched several phrasings for a provider refusing the
structured-output field, but not gateways that validate the body with a strict pydantic
model and reject the OBJECT-form json_schema by shape:
HTTP 422: {"detail":[{"loc":["body","response_format","json_schema"],
"msg":"str type expected","type":"type_error.str"}]}
422 already passes the status check; only the phrasing was missing, so the
one-retry-without-response_format rung never engaged and the auxiliary task failed
hard (title generation left 'HTTP 422' in the session title). Treat the shape error as
a rejection: the field is what the provider refuses and the remedy is identical.
Fixes#110631
_is_unsupported_parameter_error matched "does not support" but not the contraction Bedrock
Converse returns for xAI Grok ("This model doesn't support the temperature field"), nor the
inference-profile Claude wording ("`temperature` is deprecated for this model"), so the aux
ladder's retry-without-temperature rung never fired and the ValidationException surfaced on
title/vision/compression calls (#111043, second half).
Bedrock-hosted xAI Grok (us./global. inference-profile prefixes) rejects
temperature/topP in Converse with a hard 400 — the same restriction as
Claude Opus 4.6+, but _forbids_sampling_params is Claude-only, so Grok
needs its own denylist gate in build_converse_kwargs.
Fixes#111043
`hermes_cli/profiles.py` imported SYNC_MANIFEST_NAME from agent_import_sync at
module top, which pulled yaml/utils into every startup that resolves a profile;
the import now happens only inside _bootstrap_profile_dir when --sync-imports
is used (verified: importing hermes_cli.profiles no longer loads
hermes_cli.agent_import_sync).
`profile create --clone-all --sync-imports` accepted the flag (a full copy
carries import-sync.json regardless) but printed no notice; the hint is now
printed for both --clone and --clone-all.
`hermes profile create <name> --clone` copies whatever `hermes import-agent`
had pulled into the source profile, but leaves import-sync.json behind, so
the clone can never run `import-agent --sync` itself: its imported skills and
memories freeze at clone time.
`--sync-imports` (with --clone / --clone-from) also copies the manifest. It
is deliberately narrow: the manifest points at EXTERNAL Claude Code / Codex
trees, never at the source profile, so both profiles remain independent
islands (root AGENTS.md ruling) — config.yaml, SOUL.md and skills are still
one-off copies. Opt-in, one-directional, explicit; --clone-all already
carries the file as part of the full copy. Refused without a clone source.
faulthandler.register(SIGUSR2, ..., chain=True) writes the stack dump and
then re-raises the signal to its previous handler. SIGUSR2's default
disposition is "terminate", so the diagnostic hook added for #70344 kills the
very process an operator is trying to inspect (rc = -12), which is what the
#110437 reporter hit while introspecting a long-lived gateway. Nothing else
in the gateway installs a SIGUSR2 handler, so there is nothing to chain to.
Test: a child interpreter runs the real _start_install_faulthandler, receives
SIGUSR2, and must still be alive with a dump in gateway_faulthandler.log.
Red on origin/main (rc=-12), green with chain=False. Docs: the stack-dump
signal is now documented next to the event-loop watchdog.
#89322 fixed the bridge-local normalizeWhatsAppId, but bridge.js has since
moved its id handling to bridge_helpers.js::normalizeWhatsAppId, which still
turned `<user>:<device>@lid` into the malformed `<user>@<device>@lid` for
mentionedJid / quoted participant / reaction keys, and the Python side
(gateway/platforms/whatsapp_common.py::_normalize_whatsapp_id) did the same
':'->'@' swap on botIds. Drop the local duplicate in bridge.js, import the
helper, and strip the `:<device>` suffix on both layers so the bot's own ids
compare equal to the bare ids WhatsApp sends for mentions and quotes.
One invariant test: device-qualified botIds match a bare mentionedId and a
bare quotedParticipant; a plain group message still does not trigger.
normalizeWhatsAppId did String(value).replace(':','@'), which turns a device-qualified id
like '116342762025117:14@lid' into the malformed '116342762025117@14@lid'. The bot's own id
(sock.user.id / sock.user.lid) carries the :<device> suffix while inbound mentionedJid and
contextInfo.participant (quoted message author) do not, so the bot's id never matches its
botIds set -> @mention and reply-to-bot are never detected in groups. Strip the :<device>
suffix instead so all id forms compare consistently.
Drop tests/gateway/test_slack_conversational_senders.py (12 parametrized
cases driving _handle_slack_message end-to-end for allow_bots policies and
canvas mentions — the allow_bots matrix is already covered by the existing
bot-filter tests in test_slack.py) and fold the remaining coverage into two
_prefilter_inbound tests: every housekeeping subtype is dropped, and every
conversational subtype (absent, file_share, thread_broadcast, me_message,
document_mention, bot_message under allow_bots=all) still passes. Distinct
ts per event: the prefilter dedups by (team, ts) before the subtype gate.
Adds the contributor email mapping for the second author.
Drop the unreachable handle_message assertion in the drop tests
(_prefilter_inbound never calls it) and state the deliberate
file_comment drop decision in the allowlist comment.
Housekeeping subtypes (channel_join/leave/topic/name/purpose,
convert_to_private/public, pins, deletions) are not a person speaking,
yet _prefilter_inbound only rejected message_changed/message_deleted,
so each of them started a full agent turn in free-response channels.
Replace the denylist with an allowlist: a message passes when subtype
is absent, file_share, thread_broadcast or me_message; everything else
is dropped. Fixes#110778.
`_clarify_callback_sync` decided "no answer arrived" by testing whether the
response text starts with '[' (the shape of the timeout / undeliverable
sentinels). A real answer can start with '[' too — a "[A] staging" choice
label picked by number, or "[urgent] ..." free text after Other — so the
clarify resolved and the agent got the answer, yet the Slack card was
rewritten to "This prompt expired" and typing was never re-armed.
`_clarify_send_then_wait` now returns `(response, answered)` and the runner
branches on that flag only.
The Slack click handler popped the retire entry as soon as Other was
clicked, but Other is not terminal: the clarify stays pending for typed
text, so a later timeout or /new reset found nothing to retire and the card
stayed stuck on "Awaiting typed answer". The entry is now popped only on a
terminal outcome (a choice click, or Other on an already-dead entry).
A typed answer to a native card (numeric pick, or text after Other) never
reaches the click handler, so the card kept its buttons forever; the
TEXT_RESOLVED intercept now retires it with the answer.
Review finding: '[' prefix mistaken for the timeout sentinel; Other click dropped the retire entry; typed answers never rewrote the card.
One adapter-facing seam replaces the Slack-only callback: an adapter whose
clarify prompt is a persistent card (Slack Block Kit) defines
`retire_clarify_card(clarify_id, notice)`, and the gateway calls it from
every path that ends a clarify without a button click:
- TurnRunner._clarify_callback_sync: when the bounded wait returns a
sentinel (timeout, /new or run-end clear_session), schedule the retire
with the expired notice on the gateway loop (#110821).
- run_inbound TEXT_REJECTED_PROSE: retire with the cancelled notice before
the prose is routed as a follow-up (#111019). Lookup is on the adapter
class so MagicMock doubles cannot fabricate the method; no platform ==
SLACK special-case.
The Slack map is keyed by clarify_id and popped before the first await, so
a late timer cannot touch a newer prompt and the button handler's ts-keyed
guard makes a racing click a no-op. Gateway-restart-orphaned cards stay
out of scope: nothing is waiting on the new process, and the click path
already renders them expired.
Tests trimmed to invariants: the runner-level timeout probe (card adapter
vs no-card adapter), the inbound prose retire, and one Slack test covering
buttons-dropped + late-click-noop. Docs updated for the new in-place edit.
Aligns the star ranking with the rule the skills index already follows:
GitHub is consulted only by the twice-daily skills-index.yml schedule, whose
artifact every docs deploy reuses. deploy-site.yml now runs
fetch-plugin-stars.py without --probe (reuse-only: artifact → live site copy →
disk → empty) and cannot call the API at all, so a same-day merge train adds
zero requests regardless of cache age.
The probe itself collapses from one REST call per repo (~26 today, growing
with the catalog) to a single GraphQL query with aliased repository fields,
so the scheduled run costs one request no matter how big the catalog gets. A
failed probe (rate limit, renamed repo, bad token) keeps the previous counts.
- Drop the 900s "inbound liveness" INFO line from the salvage: a periodic
log heartbeat is a feature with a separate scoping call (#111211 item 3);
the existing stall watchdog already escalates when getUpdates stops.
- The polling error_callback interpolated the raw exception and left the
redaction call inside the format string, so its lines read
"Telegram network _redact_telegram_error_text(error), scheduling
reconnect: ..." and leaked unredacted text. Redact for real.
- Tests: keep two invariants (recovered wording after network errors;
clean bootstrap stays "confirmed healthy") and the transport test.
httpx timeout exceptions (ConnectTimeout, ReadTimeout, ...) stringify to
"" so every adapter log line built from _redact_telegram_error_text()
ended with a blank reason. Fall back to the exception class name.
Partial salvage of #111222: only the _redact_telegram_error_text hunk;
the transport-layer and polling-recovery hunks are covered by #111221.
Moving the dedup flush and thread lookup from asyncio.to_thread onto the
adapter-owned pool fixed the torn-down default executor, but to_thread also
copies the caller's contextvars and run_in_executor does not. A multiplexed
profile's HERMES_HOME override and secret scope are contextvars, so those
workers silently ran under the launch profile. _run_blocking now runs the call
through contextvars.copy_context().run, matching to_thread semantics.
Review finding: _run_blocking lost the profile HERMES_HOME override / secret scope on the worker.
_connect_websocket now submits the SDK thread to _get_sdk_executor() instead
of the loop default executor; the SimpleNamespace adapter stub in the
profile-scope test predates that contract and lacked the method.
`_fetch_last_message_in_thread` was the last hot-path `asyncio.to_thread`
in the adapter: after a default-executor teardown (#111020) thread-reply
routing would fail the same way the dedup flush did. Route it through
`_run_blocking` like every other blocking SDK call. The remaining
`to_thread` users (`_load_lark_oapi` at connect/onboarding, the voice
transcode with its file-attachment fallback) are cold or degrade cleanly.