Commit Graph

36965 Commits

Author SHA1 Message Date
teknium1
9f5b7ea02e fix(compression): keep the aux ceiling and re-probe feasibility on every main-runtime change
The aux-window clamp installed by _lower_threshold_to_aux_context() was a one-time
assignment to threshold_tokens; ContextCompressor.update_model() recomputed the trigger
from the main model and discarded it, and the _compression_feasibility_checked latch was
never reset, so after a mid-session switch to a larger main model the trigger sat at the
main-model value (450K) while the pinned summariser accepted 272K (#114707).

- ContextCompressor holds the aux window as a durable _aux_context_ceiling that
  _apply_threshold_tokens_cap() honours on every recomputation; update_model() voids it
  only when the main runtime changes (an "auto" aux route follows the main model).
- revalidate_compression_feasibility(agent) resets the latch and re-probes eagerly at
  every runtime change: switch_model (outside the rollback guard), fallback activation
  and primary restore. Symmetric: a runtime whose aux fits restores the main trigger.
- Feasibility notices emit once per distinct verdict so /model --once restores and
  fallback cycles do not re-announce an unchanged verdict.
- Rewrites the switch-time hunk salvaged from #114710: unconditional, outside the
  rollback try, so a catalog hiccup never undoes a good switch. Test kept and extended.

Co-authored-by: KoNit-K <konit.block@protonmail.com>
2026-09-18 09:49:14 -07:00
KoNit-K
8d51d488a4 fix(agent): retain auxiliary compression limit on model switch 2026-09-18 09:49:14 -07:00
teknium1
957bf63dc5 fix(kanban): dashboard /orchestration docstring no longer claims decomposer parity
The decomposer now prefers the root card's assignee for unrouted children and
falls back to the active profile only for cards without an assignee, so the
endpoint's "fallbacks filled the same way the decomposer does" was stale. State
what the endpoint actually resolves; no routing change.
2026-09-18 09:48:41 -07:00
liuhao1024
1ea94b3745 fix(kanban): decomposed cards fall back to the root task's assignee, never the dispatcher's own profile
The decomposer resolves both `kanban.default_assignee` (unrouted children)
and `kanban.orchestrator_profile` (the root after fan-out) as
"explicit config, else the active profile". The active profile is whatever
HERMES_HOME hosts the dispatcher — in the report an incognito `private`
profile with no credentials — so every unrouted child AND the root card
itself were silently re-owned by a profile that can never do the work.

The resolver now tries the root task's own assignee before the active
profile: explicit config → root assignee (if it names an existing profile)
→ active profile. Slim redo of #114303's decompose hunk at the resolver so
it covers the orchestrator fallback too; tests and docs from #114303.

Fixes #114294
Salvages #114303

Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
2026-09-18 09:48:41 -07:00
teknium1
27a30d8515 fix(setup): xAI TTS wizard checks XAI_API_KEY before OAuth to match runtime
_tts_xai_step still announced OAuth-first ordering (docstring and printed
message) while the synthesis and availability paths now prefer an explicit
XAI_API_KEY over the subscription OAuth bearer. Check the key first and fix
the copy; regression test under tests/hermes_cli.
2026-09-18 09:47:06 -07:00
teknium1
ca3c425627 fix(tts): pin _xai_requirements to key-first credential resolution in tests
Reverting tools/tts_tool.py to main left tests/tools/test_tts_streaming.py
green, so the availability probe (_BUILTIN_REQUIREMENTS['xai']) could regress
to OAuth-first silently. Fold one assertion into the existing
prefer_api_key test using the same fake tools.xai_http module.
2026-09-18 09:47:06 -07:00
teknium1
d7865939af fix(tts): xAI TTS availability probe prefers XAI_API_KEY like the synthesis paths
`tools/tts_tool.py::_xai_requirements` (the provider check text_to_speech_tool
dispatch consults) still resolved xAI credentials OAuth-first, while both TTS
synthesis paths now pass `prefer_api_key=True`. Truthiness is unchanged, but a
configured key no longer routes the availability probe through the OAuth pool
(pool select / refresh) for a bearer the metered `/v1/tts` endpoint rejects.

Docs: the streaming-TTS provider table now states the ordering.

Part of the #113727 salvage of #113728 (@beardthelion); sibling site the PR
missed.

Co-authored-by: beardthelion <beardthelion@users.noreply.github.com>
2026-09-18 09:47:06 -07:00
beardthelion
3b0aae77ee fix(tts): prefer explicit XAI_API_KEY over subscription OAuth in streaming TTS
The sync xAI TTS path resolves credentials with prefer_api_key=True
(#87045/#88040): the subscription OAuth bearer authorizes chat but
returns 403 on the metered TTS API. The WebSocket streaming path
(XAIStreamer, salvaged from #47588 before that ordering existed) calls
resolve_xai_http_credentials() with the default OAuth-first order at
both sites, so a user holding both credentials gets the 403ing bearer,
and XAIStreamer.available() reports True on a credential that cannot
stream.

Pass prefer_api_key=True at both sites — identical availability
semantics (OAuth remains the fallback), correct preference order, parity
with _generate_xai_tts.
2026-09-18 09:47:06 -07:00
teknium1
66d79311e1 fix(api_server): resolve the model_routes alias for the in-process wake turn
The HTTP wake self-post resolves the virtual model's model_routes entry via
_select_request_route before running the turn; the in-process wake passed
route=None, so an aliased served profile woke on the global default route.
Resolve it the same way and pass route plus the request overrides.
2026-09-18 09:46:35 -07:00
teknium1
76257d1fd2 fix(api_server): move the in-process wake turn body to api_server_runs
run_internal_session_turn belongs with the run machinery that already owns
_resolve_live_session_id, not in the api_server.py facade. Pure move: the
adapter keeps a thin delegating method, behaviour unchanged.
2026-09-18 09:46:35 -07:00
teknium1
d8edd60399 fix(gateway): wake a served profile's api_server session in-process for background-process completions
A served (multiplexed) profile's api_server turn binds the raw session id
as its session key, so its watch/completion event names no profile and the
wake path self-posted it unprefixed with the primary key — resuming the
session in the DEFAULT profile's store — while an event whose source did
name a route-only served profile resolved no adapter and was deferred
forever.

_self_post_api_server now proves ownership through the served profile's
own session store (the rung the Kanban notifier already applies), scans
served stores when the raw event carries no hint, runs the wake under that
profile's scope via deliver_wake(profile=...), and fails closed for a hinted
profile that does not own the session. The ownership helper moves to
gateway/wake.py so both callers share one function.
2026-09-18 09:46:35 -07:00
teknium1
5dd426b774 test(kanban): trim the served api_server wake suite to two invariants
Replace the fifteen scenario tests from #114680 with two invariant tests:

1. notifier tick: the served profile's owned session wakes in-process under that
   profile's scope with no HTTP self-post and one shared adapter; unknown / foreign-store /
   other-stamped sessions, unserved profiles and a connected-secondary boundary all fail
   closed; the default profile's own api_server subscription still HTTP self-posts.
2. adapter: `run_internal_session_turn` binds the proven profile, adopts the compression
   tip, delivers history as wake-capable, restores the request binding, and refuses both
   a session outside the owner's store and a missing profile.
2026-09-18 09:46:35 -07:00
teknium1
8be66c558d fix(kanban): in-process served wake requires the proven profile; drain fails at once
Follow-up to the salvaged #114680 (@phoebsie):

- `APIServerAdapter.run_internal_session_turn` takes the owning profile as a required
  argument and raises without it. The dropped fallback (ContextVar, then the process-active
  profile) had no caller and would have let a missing ownership proof silently bind a
  wake to whatever profile the process happened to run as — the exact class #114679
  reports.
- A draining gateway refuses the in-process wake immediately, mirroring the HTTP
  self-post's 503 (>= 400 → raise → the notifier rewinds the claim and the next tick
  retries). The concurrent-run cap keeps the 429-style backoff.
- Contributor attribution mapping for builder@phoebie.local -> @phoebsie.
2026-09-18 09:46:35 -07:00
Phoebie Builder
eeada33fa1 review: retry the cap, adopt the compression tip, scope the in-process route to multiplex
Independent review findings addressed:

- adopt the live continuation tip (api_server_runs._resolve_live_session_id)
  before the in-process wake, exactly as the HTTP self-post does: a rotated
  (compressed) origin must be woken on the transcript that is actually live,
  not the retired parent slice (CompressionSessionClosedError otherwise);
- retry a saturated concurrent-run cap with the HTTP path's own backoff
  (gateway.wake._RETRY_DELAYS_SECONDS) instead of failing on the first check,
  so a busy listener no longer burns one notifier failure per tick toward the
  12-strike drop of a durable subscription;
- keep the historical HTTP self-post on a standalone (non-multiplex) gateway:
  the in-process route is now chosen by a served-profile helper that requires
  multiplex_profiles, and the adapter is told the authorized profile explicitly
  instead of re-deriving it from HERMES_HOME;
- skip the ownership store read for the default profile (the fallthrough already
  authorizes it) and run the delivery-time recheck off the event loop;
- tests: standalone named-profile gateway, missing store, failed-connect
  boundary, adapter without in-process delivery, platform-wide api_server route
  vs the default profile's destinations, cap retry/exhaustion, compression tip,
  and the profile handed to the adapter (8 -> 15 cases).
2026-09-18 09:46:35 -07:00
Phoebie Builder
5579fd5cdf fix(kanban): wake a served profile's api_server session in-process under its own scope
A Kanban notify+wake subscription whose destination is a raw session id on the
shared api_server (platform=api_server, chat_id=<session id>, notifier_profile=
served secondary) was never delivered under gateway.multiplex_profiles: no
profile_routes entry can anchor a session id, so _adapter_for_subscription()
returned None and _claim_for_sub skipped the row before claiming (cursor frozen,
no delivery attempt). A platform-wide api_server route is not the fix — it
matches every api_server destination, so the default profile's own api_server
subscriptions would fail closed instead.

The wake leg had the same shape: _self_post_chat_completion() always POSTs to
the unprefixed /v1/chat/completions with the PRIMARY adapter's key, so a served
profile's wake turn would resume the session in the default profile's store.
/p/<profile>/v1/... requires that profile's own API_SERVER_KEY, which a
route-only profile legitimately does not have.

Authorize the shared adapter only when the subscription's session id is a row in
the served profile's own state.db stamped for that profile (NULL legacy stamp =
the store's own profile, the same rule the dashboard session routes apply) and
run the wake in-process under that profile's runtime scope through a new
APIServerAdapter.run_internal_session_turn(). Unknown session, foreign stamp,
unserved/deleted profile and unreadable store all fail closed; the concurrent-run
cap defers the wake (cursor rewind, retry) instead of bypassing it.
2026-09-18 09:46:35 -07:00
liuhao1024
e0c078c958 test(agent): assert pool purity instead of echoing the ambient credential dict
Drop the raw reader-is-None assertion: on a non-hermetic host its failure
message would print the live credential dict into the test output. The
manual-only sources assertion below already pins the invariant the test
cares about, and still fails if the stub ever stops holding (#114431).

Co-authored-by: Nick Coleman <154543891+Digitally-Challenged@users.noreply.github.com>
2026-09-18 09:46:03 -07:00
liuhao1024
3bc6250664 test(agent): keep the auxiliary cache suite hermetic on hosts with an ambient Claude Code login
The isolated_home fixture only redirects HERMES_HOME, but the borrowed
Claude Code reader (agent/anthropic_credentials.read_claude_code_credentials)
still consults ~/.claude/.credentials.json and the macOS Keychain. On a
host with an ambient login, _seed_anthropic_singletons upserts a second
un-cooled-down pool entry, peek() selects it during the cooldown test,
and _pool_cache_hint returns 'anthropic:<id>:<digest>' instead of
'anthropic::<digest>' (fails exactly as reported in #114424).

Stub the reader to None in the fixture and add a regression test that
asserts the stub holds, the pool stays manual-only, and the hint keeps
the 'anthropic::' form.

Fixes #114424
2026-09-18 09:46:03 -07:00
teknium1
9dd36c56cf fix(mcp): remote-session OAuth hint names the configured redirect_host
The loopback SSH hint hard-coded http://127.0.0.1:<port>/callback while a
pre-registered client (Asana) redirects to http://localhost:<port>/callback,
the URL the app must register verbatim. Thread redirect_host from the oauth
config into the redirect handler so the hint prints the same host the
provider will use. Live pass copy nit #6 on #113907.
2026-09-18 09:45:32 -07:00
teknium1
b0cd35e259 fix(mcp): dashboard/Desktop Authorize honours a pre-registered client's pinned loopback redirect
A no-DCR entry (client_id + redirect_port, e.g. the Asana manifest) has
http://localhost:<port>/callback registered with the vendor, which matches
redirect URLs exactly; the dashboard flow overrode it with its own callback
URL, so the in-app Authorize button could never complete for such entries.
The pinned loopback listener now wins (over the dashboard URL and any cached
registration URI); the dashboard flow only publishes the authorization URL,
and the stdin paste reader stays off under a dashboard flow. Docs/post_install
say the approving browser must run on the Hermes machine.
2026-09-18 09:45:32 -07:00
teknium1
e8aa69c1a1 fix(mcp): a running server adopts a changed config.yaml definition on rebuild
run() kept the dict it was started with, so a url/oauth change made while
the process ran (catalog migration, dashboard re-auth to the new endpoint)
left the loop probing the OLD URL — and MCPOAuthManager, seeing a different
URL, evicted the fresh provider the dashboard had just authorised for the
new one, dropping its tokens/DCR client (#113907, atom b). Before each
transport rebuild a remote server re-reads its definition and rebinds when
url/auth/oauth/headers/transport changed.
2026-09-18 09:45:32 -07:00
teknium1
c553df915c fix(mcp): Asana catalog installs a working V2 pre-registered OAuth client
The V1 beta server https://mcp.asana.com/sse is retired and Asana's V2 server
(https://mcp.asana.com/v2/mcp, Streamable HTTP) has no Dynamic Client
Registration: every client must be an Asana "MCP app" the user registers in
the developer console. The salvaged manifest fixed the URL and documented the
manual `oauth:` block; this commit makes the catalog install itself produce
that block so no hand-edit of config.yaml is needed.

- hermes_cli/mcp_catalog.py: manifests may pin a closed `auth.oauth` mapping
  (client_id, client_secret, redirect_host, redirect_port, scope) that
  `_build_server_config` copies to `mcp_servers.<name>.oauth`; every `${VAR}`
  it references must be declared in `auth.env`, mirroring the api_key header
  contract, so a placeholder can never reach the token endpoint as a literal.
  `install_entry` now prompts `auth.env` for OAuth entries too (dashboard and
  Desktop already render `required_env` regardless of auth type).
- optional-mcps/asana/manifest.yaml: declare ASANA_CLIENT_ID/SECRET, pin the
  client + `http://localhost:27890/callback` (Asana matches the redirect URL
  exactly), and rewrite post_install around the MCP-app registration steps.
- tests: shipped-catalog invariant (no manifest installs the retired `/sse`
  URL; the Asana client credentials are declared `${VAR}` refs; callback
  pinned) + install-path invariant for the new `auth.oauth` block, including
  the undeclared-reference rejection.
- docs: catalog section on entries that need a user-owned OAuth app.

Live: `hermes mcp install asana` + `hermes mcp login asana` under a temp
HERMES_HOME. Before: config `url: …/sse`, no oauth block, authorize URL on the
V1 server (mcp.asana.com/authorize) with a DCR client and 127.0.0.1 redirect.
After: config `url: …/v2/mcp` + oauth `${ASANA_CLIENT_ID}` refs, .env holds
the values, authorize URL on app.asana.com/-/oauth_authorize with the
configured client_id, redirect_uri=http://localhost:27890/callback and
resource=https://mcp.asana.com/v2/mcp — the flow Asana's guide documents.
2026-09-18 09:45:32 -07:00
KoNit-K
7954418d90 fix(mcp): migrate Asana catalog to V2 2026-09-18 09:45:32 -07:00
teknium1
b65ec6a3f0 docs: list the DISCORD_MISSED_MESSAGE_BACKFILL_* env fallbacks
The adapter honours six env fallbacks for discord.missed_message_backfill
(including the new MAX_ATTEMPTS one) but none was in the environment
variables reference; document them next to the other DISCORD_* rows.
2026-09-18 09:44:59 -07:00
teknium1
8c6881231a fix(discord): backfill stops re-dispatching a message once its turn delivered, and bounds the rest
Three defects let the missed-message backfill re-run the same inbound message on
every reconnect (24 turns in 12 h for one message on the reporter's install):

- Completion was only ever recorded by the send-path ledger writer, keyed on the
  Discord reply anchor. With reply_to_mode "off", a streamed final delivered via
  edit_message(finalize=True), a fresh-final send or a media-only reply there is no
  anchor, so the row stayed processed/replied=0 and _should_backfill_discord_message
  re-admitted it forever. The adapter already learns the outcome per inbound id in
  _record_discord_processing_complete: on SUCCESS (base confirmed delivery, or the
  stream already delivered) it now writes status='responded', replied=1.
- The scan wrote status='discovered' over every candidate at the top of each pass,
  erasing the queued/processing claim so the 10-minute active-claim guard could never
  fire; two back-to-back scans dispatched the same message twice. The 'discovered'
  write no longer overwrites an existing row's status or updated_at.
- A stored per-channel cursor REPLACED the window floor (12-hour-old messages stayed
  candidates under window_seconds: 3600) and the parent's cursor object was passed
  down to child threads. after = max(cursor, now - window) and threads only see the
  window floor plus their own cursor.
- max_dispatches caps one scan, not one row. New missed_message_backfill.max_attempts
  (default 3; env DISCORD_MISSED_MESSAGE_BACKFILL_MAX_ATTEMPTS) is a lifetime ceiling
  on re-dispatch of a single message; `attempts` now counts dispatches only.

Live repro (real DiscordAdapter, temp HERMES_HOME ledger, fake client yielding the
same message per scan, fresh adapter per scan = reconnect): streamed final with
reply_to_mode=off 3 scans -> before 3 dispatches, after 1; turn that never completes
3 scans -> 3/1; turn that always fails 6 scans -> 6/3; 12h-old cursor with a 1h window
-> before 2 out-of-window dispatches (parent + thread), after 0; control: an
unanswered message still dispatches exactly once, and a cursor newer than the window
still narrows the scan.

Part of #113631 (with the salvaged anchor fallback from #113633 by @KoNit-K).
2026-09-18 09:44:59 -07:00
KoNit-K
dd7af8e765 fix(discord): persist unthreaded recovery replies 2026-09-18 09:44:59 -07:00
KoNit-K
61f99d6d3e fix(dingtalk): guard edits when AI Cards are unavailable 2026-09-18 09:44:26 -07:00
Reinhold
9223a424d8 fix(buzz): pass edit content literally to the CLI
Signed-off-by: Reinhold <310554180+reinhold-ph@users.noreply.github.com>
2026-09-18 09:43:52 -07:00
KoNit-K
e2062866eb fix(buzz): pass edit content as argument 2026-09-18 09:43:52 -07:00
teknium1
f971bbf512 fix(plugins): removed_annotation takes the resolved kill list as a required argument
The Optional default with an on-demand resolved_removed_entries() fallback had no
production caller (plugins_cmd.py and web_server_dashboard.py both pass the list)
and re-opened the per-row live-catalog fetch this PR removes. Make the parameter
required so a future per-row caller fails loudly instead of silently fetching.
2026-09-18 09:43:19 -07:00
teknium1
a89ef71b10 refactor(plugins): kill-list annotation takes the pre-resolved list directly; trim to two invariants
Follow-up to the salvaged #113687 / #113682 commits:

- `removed_annotation(name, dir_path, removed_entries=None)` resolves the kill list once itself
  when no list is passed and matches against it; the `removed_annotation_batch()` wrapper and its
  name-keyed dict are gone (a hub row keyed by `name` could shadow a same-named nested plugin).
  Both surfaces (`plugins list`, `_merged_plugins_hub`) call `resolved_removed_entries()` once
  and pass the list per row.
- The negative cache is one module float deadline instead of a URL-keyed dict plus two helpers;
  the duplicated stale-cache read is one `_stale_live_cache()`.
- Tests trimmed to two invariants: failure memory honours the TTL and still serves a stale copy;
  hub rebuild + `plugins list` each cost one network attempt and still annotate an in-tree removal.
  The #113682 route test's fake gains the new third argument.
- Docs: the live-refresh section says a failed fetch is remembered for a minute.
2026-09-18 09:43:19 -07:00
KoNit-K
6165b6ebd5 fix(dashboard): offload plugins hub rebuild 2026-09-18 09:43:19 -07:00
Konstantin Khlopkov
88b3163742 fix(plugins): one kill-list resolution per plugins hub rebuild and CLI listing 2026-09-18 09:43:19 -07:00
teknium1
f0fe07af0e fix: escape table-cell pipe in Slack autolink example so the docs site compiles
The rich_blocks row's `<url|label>` code span was split at the pipe by the
GFM table parser, leaving a bare `<url` that MDX rejected (Docs Site CI:
"Unexpected end of file in name" at slack.md:474). Escaping the pipe keeps
the rendered text identical.
2026-09-18 09:41:44 -07:00
teknium1
7beee832c9 docs(slack): note that rich_blocks renders <url|label> autolinks in lists
The rich_blocks row only listed structural blocks; state that both GFM links and
Slack mrkdwn autolinks become link elements in rich_text surfaces so operators
know the list/quote/table paths are not mrkdwn-blind.
2026-09-18 09:41:44 -07:00
z23
da25b16662 fix(slack): parse mrkdwn <url|label> inside rich_text lists
rich_blocks turns markdown bullets into rich_text_list. rich_text does
not interpret mrkdwn, so Slack <url|text> was emitted as literal text
while [text](url) became a real link. Section/mrkdwn paragraphs were
unaffected.

Parse Slack autolinks in _inline_elements (lists, quotes, table cells).
Mentions (<@U>, <#C>, <!here>) have no scheme: and stay as text.
2026-09-18 09:41:44 -07:00
teknium1
aca2631d95 fix(test): pin config set's refusal of an invalid list literal
test_invalid_list_literal_warns_and_stores_string asserted the pre-#114471
warn-and-store behaviour that this branch deliberately removed; it was the one
red test in the Python tests job. It now asserts the refusal and that nothing
was written.
2026-09-18 09:41:02 -07:00
teknium1
9c088a7a78 fix(config): fix container type for unseeded roots, top-level lists and bare-name list slots
_expected_container_type only saw a mapping/list for model.aliases when one was
already on disk (no DEFAULT_CONFIG seed), and skipped every single-segment key, so
'config set model.aliases notamap' / 'config set toolsets browser' still stored a
string on a fresh config (#114471 writer atom). Known unseeded container roots
(providers, model.aliases, model_aliases) join custom_providers in a fixed table,
and top-level list keys seeded in DEFAULT_CONFIG are checked too; only a mapping
*section* keeps deferring to _guard_section_overwrite.

agent.disabled_toolsets / skills.disabled readers accept a bare name via
parse_config_string_list, so such a scalar is stored as a one-item list instead
of being refused.
2026-09-18 09:41:02 -07:00
teknium1
7b4e052bda chore(contributors): map Sasni, fotedev, whyyagswhy author emails 2026-09-18 09:41:02 -07:00
teknium1
707922e1a8 test(config): non-list custom_providers keeps the providers view and warns
Two invariants for the malformed-legacy-key fix (salvage of #45730): the v12+ providers view survives a string custom_providers with a warning, and a well-formed list stays silent.
2026-09-18 09:41:02 -07:00
teknium1
0972b4819a fix(config): config set refuses a wrong-shaped value for a list/mapping key
`hermes config set` stored a string where the schema wants a list or a
mapping, with at most a stderr warning ("storing as string"). Every
isinstance-gated reader then ignored the value while `config get` echoed it
back — the reporter's `custom_providers` became a string and Desktop's
Custom Endpoints read "0" with no error anywhere.

Hard guardrail instead: the write path resolves the key's container type
(DEFAULT_CONFIG for nested paths, the legacy `custom_providers` root, or the
list/mapping already on disk) and refuses a plain string or a wrong-shaped
literal, naming the expected type. A value that looks like a list/mapping
but is not valid YAML/JSON is refused too. `--force` keeps its documented
meaning (replace a whole mapping section); a non-list in a list slot has no
override. UPPER_SNAKE names still route to `.env` untouched.

Tests: the two refusals plus a control that valid literals, scalar keys and
`--force` still write. Docs: cli-commands `config set` row.

Config-set atom of #114471.
2026-09-18 09:41:02 -07:00
fotedev
9407b9f1a8 fix(cli): keep api_key/key_env on nested model.aliases dict entries
Top-level `model_aliases:` entries carried credentials into DirectAlias but
the nested `model.aliases:` dict form built it without `api_key`/`key_env`,
so an alias with its own base_url got no credential (401) even though the
URL resolved correctly. Pass both fields through, as the top-level path does.

Source hunk from PR #109834 (earliest filer); invariant tests from the
identical later PR #114472. Nested-alias credential atom of #114471.

Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 09:41:02 -07:00
Sasni
d80639ed66 fix(config): keep the providers view when custom_providers is malformed
`get_compatible_custom_providers()` returned `[]` with no log line whenever
the legacy `custom_providers` key held a non-list (a string written by a bad
`config set`). That emptied the merged view — Desktop Providers → Custom
Endpoints read "0" — even though the v12+ `providers:` entries were valid.

Skip only the legacy list and warn with the key and the received type, so
the `providers:` map still merges and the state is searchable in the log.

Hand-ported from PR #45730 (the function has since moved from
hermes_cli/config.py to hermes_cli/config_providers.py).
2026-09-18 09:41:02 -07:00
teknium1
d7f2644141 fix(codex): profile declarations follow the host and never override OpenAI's own ladder
The salvaged custom-profile declaration (#114255) gives every ``custom:<name>``
Responses route the OpenAI-compat vocabulary, which closes #114249 but has two
edges the transport must keep:

- ``_profile_declared_efforts`` resolved by provider NAME first, so the new
  non-None custom declaration short-circuited the host lookup: a
  ``custom:my-proxy`` entry pointed at api.router.com stopped inheriting the
  Router catalog clamp and would send ``max`` to a gateway that 400s on it.
  Resolve by endpoint host first, then by name — the host is the endpoint's
  truth; the config-entry name is only a label (the existing Router test now
  uses the runtime's real ``custom:<name>`` identity, which is what exposed it).
- A custom entry that merely points at api.openai.com (host-mandated
  codex_responses) is still OpenAI: its per-model ladder is known, so skip
  profile declarations on the official origin and the Codex backend.
  ``_is_openai_api_origin`` is the shared exact-host check;
  ``_is_official_openai_responses_route`` reuses it.

Docs: providers.md states the custom-endpoint effort contract and both
host-following exceptions.

Live: custom:relay deepseek-flash max -> max (was xhigh); custom:oai @
api.openai.com gpt-5.2 max -> xhigh; openai gpt-5.2 -> xhigh, gpt-5.6 -> max;
custom:my-proxy @ api.router.com grok-4.6 max -> xhigh (pick-only: max).
2026-09-18 09:40:31 -07:00
liuhao1024
f111297af1 fix(custom): declare the OpenAI-compat effort vocabulary so Responses keeps a configured max
CustomProfile did not override supported_reasoning_efforts, so on the
Responses transport a custom:<name> relay fell through to the OpenAI
per-model ladder (codex_supported_efforts) and a configured effort=max
was silently clamped to xhigh — while the same provider over
chat-completions forwarded max unchanged, because that path already
clamps onto OPENAI_COMPAT_WIRE_EFFORTS.

Declare the same wire set from the profile so the two transports agree:
max survives, ultra still clamps down to max, and the official OpenAI
backend ladder (gpt-5.5 rejecting max) is untouched.

Fixes #114249
2026-09-18 09:40:31 -07:00
teknium1
454f9c1421 fix(cron): remind once per cooldown after the alert-once gate; migrate alerted_at
Builds on the previous commit (alerted treated like closed): a permanently
silent `alerted` incident would hide a job that stays broken for days, so the
gate now withholds only inside `cron.failure_repeat_alert_hours` (default 6,
0 = re-alert on every failing run) and lets exactly one reminder through, which
re-stamps the window.

- cron/incidents.py: `alerted_at` column (added in place to existing ledgers via
  add_column_if_missing); set_incident_state(..., "alerted") stamps it every time
  so a reminder restarts the cooldown; resolved->detected re-open clears it so
  the same error after a green run alerts immediately.
- cron/scheduler.py: `_repeat_alert_withheld` reads the stamp; a missing or
  unparseable stamp (pre-migration row) delivers rather than swallowing the
  alert; `closed` still wins; the unreadable-ledger fail-open is unchanged. The
  crash path (`_deliver_crash_failure`) shares the gate through
  `_upsert_incident_for_failure`.
- hermes_cli/config_defaults.py + website/docs cron page: document the key.
- tests: fold the contributor's two tests and the old "unacked failures keep
  alerting per run" change-detector into two invariants (unit gate incl. 0 /
  legacy row / closed; end-to-end alert once -> reminder once -> recovery re-arms).

Fixes #113665
Co-authored-by: Yagna Vudathu <yagnavudathu@gmail.com>
2026-09-18 09:40:00 -07:00
Yagna Vudathu
7aa466665b fix(cron): alert once per failure incident instead of every run
A failing job re-pinged the operator on every run: _upsert_incident_for_failure
only suppressed acked (closed) signatures, while the post-delivery alerted mark
had no reader. Treat alerted like closed in the upsert gate so the same
signature goes silent after the first ping. Recovery still re-arms via resolved
-> detected, and a changed error still mints a new incident.

Fixes #113665.
2026-09-18 09:40:00 -07:00
teknium1
b41c2c5d3c fix(docs): custom endpoints do receive top-level reasoning_effort; correct the local-endpoint note
CustomProfile.build_api_kwargs_extras emits the resolved effort (global or
per-model override) as top-level reasoning_effort clamped to the
OpenAI-compatible wire values; only the nested reasoning object is gated to
known-capable hosts. The note previously claimed a plain local base_url
received no effort at all.
2026-09-18 09:39:29 -07:00
teknium1
4b617afcc5 fix(reasoning): trim salvaged tests to two invariants; document escape + local-endpoint gate
Keep one positive/negative invariant (custom-provider prefixed key applies to the bare
runtime slug, unrelated model still misses) and one precedence invariant (direct key beats
the reverse lookup); the resolve_reasoning_config shape is covered by the live repro.

Docs: the "no `hermes config set` support for reasoning_overrides" note predates the
backslash-dot escape (#84064) — point at it. Explain why a plain local OpenAI-compatible
base_url receives no `reasoning` field (gate from 3f0f4a04a9, arbitrary hosts 400 on unknown
fields) and point at the custom-provider `extra_body` opt-in that reaches the wire.

Part of #114073
2026-09-18 09:39:29 -07:00
liuhao1024
1d556a1ce5 fix(reasoning): match provider-prefixed override keys against bare model strings
A per-model agent.reasoning_overrides key written the documented
provider/model form (e.g. ollama-local/qwen3.6:27b-q4_k_m) never matched
once the runtime model string lost the prefix — fallback entries and
custom-provider resolution feed the bare slug, so the resolver fell back
to the global effort instead of the override (#114073).

Add a reverse lookup after the existing variant pass: each prefixed
override key's stripped forms (bare tail, aggregator-stripped) are
matched against the model's variant set. Direct and variant matches
still win, keeping provider-qualified keys most specific.
2026-09-18 09:39:29 -07:00
teknium1
4f519b746e test(update): trim the stale-root-module bridge tests to two invariants
Keep the restart-phase shape (hermes_cli.* purged, root utils stale, a fresh
hermes_cli.config / hermes_cli.managed_scope import heals it) and the negative
(a complete utils is never dropped). The unit eviction test and the "naive
consumer still dies" control were change-detectors for the same mechanism.
The purge fixture now restores the original module graph instead of
setdefault-merging the fresh one over it.
2026-09-18 09:38:34 -07:00