Commit Graph

484 Commits

Author SHA1 Message Date
Yags
61154b6f68 fix(models): route Union Alpha through messages 2026-09-16 11:41:18 -07:00
KoNit-K
91a38622db fix: guard atomic writers after profile deletion 2026-09-16 00:32:15 -07:00
teknium1
9aaa71367b fix(catalog): apply the opencode free exclusion on the live-first keyed Zen/Go picker 2026-09-15 18:23:29 -07:00
teknium1
bb5d745ab1 fix(catalog): keep the delisted deepseek-v4-flash-free out of the live OpenCode picker too
The Zen relay still LISTS deepseek-v4-flash-free in GET /zen/v1/models but no longer
serves it (opencode.ai/docs/zen dropped it; every anonymous POST 400s "Upstream request
failed: Model is unavailable"). Removing it from the offline floor alone leaves the picker
offering it whenever the live fetch succeeds — which is nearly always — so the first turn
400s and the fallback switch strands the session (#111749).

Add it to the live-list exclusion set (renamed from the keyed-twin-only
_OPENCODE_FREE_KEYED_SUFFIX_MODELS to _OPENCODE_FREE_EXCLUDED_MODELS, same semantics for
ox-alpha-free), drop the same id from the opencode-zen keyed catalog (a Zen pick of a free
slug heals to the keyless relay and 400s the same way), move the test fixture's "current
free tier" to the relay's actual state, and pin the filter with one invariant test.
2026-09-15 18:23:29 -07:00
Louis.Anson
c008cfe920 fix(copilot): pass two-slash custom-model ids through unchanged
Copilot enterprise custom models (BYOK) expose catalog ids shaped
owner/sub/model. normalize_copilot_model_id() only tried the full id and the
id minus its FIRST segment, and returned the stripped form when neither
matched the (unreachable) catalog - corrupting those ids into sub/model,
which the CLI's second normalization pass then stripped again to model. The
Copilot API answered HTTP 400 model_not_supported on every call.

Only accept the strip guess when the remainder is itself a flat id: a result
that still contains "/" cannot be a Copilot id.

Fixes #110597
2026-09-15 04:47:37 -07:00
teknium1
fdd5995ecb fix(tui_gateway): serve fails closed when hosting a second profile home; profile RPCs bind the full runtime scope
`hermes serve` / the Desktop backend hosted many profile homes (session
profile_home, the `profile` RPC param, hosted rooms) but never called
agent.secret_scope.set_multiplex_active, so every unscoped get_secret read for
a secondary silently returned the LAUNCH profile's os.environ value, and
@_profile_scoped bound only HERMES_HOME: `config.get full` for profile B
expanded B's `${VAR}` refs to the default profile's plaintext credentials,
model.options listed the default's env-keyed providers, llm.oneshot billed the
default's auxiliary key.

- tui_gateway/launch_profile_policy.py (was launch_terminal_policy.py): the
  first time _profile_home registers a non-launch home the process freezes
  the launch env and flips get_secret to fail closed
  (activate_multi_profile_hosting); launch_secret_scope composes the launch
  profile's .env + external sources over that frozen env so systemd / op-run
  injection survives the flip while a secondary never sees it.
- model_switch._profile_runtime_scope_tokens is the ONE composer for
  home + secret + terminal scope: a named profile binds its own files; the
  launch profile binds its frozen-env scope once multiplexing is active and
  stays unscoped in a single-profile process (legacy os.environ precedence).
  _profile_scoped, _profile_scoped_rpc, _session_profile_runtime_scope,
  _bind_build_profile_scopes and _prepare_turn_input all go through it.
- Hosted-room / Group Chat turns for a DEFAULT-profile member in a
  `multiplex_profiles: true` gateway no longer die at agent build with
  UnscopedSecretError: `profile_home is None` was treated as "no scope"
  in _start_agent_build._build and _prepare_turn_input.
- llm.oneshot runs under the session's (or params.profile's) scope;
  _lap_builtin_rows / _overlay_has_creds / _provider_has_credentials read
  provider keys through _scoped_key_env instead of raw os.environ;
  methods_groups._profile_execution_policy resolves the hosted-room policy
  (which reads provider credentials) under the profile's full scope.

Live repro (real `hermes serve`, two homes, config.get {key: full, profile: b}):
base  a_ref: <A_VALUE>  b_ref: ${B_ONLY_TOKEN}  env_ref: <ENV_INJECTED>
head  a_ref: ${A_ONLY_TOKEN}  b_ref: <B_VALUE>  env_ref: ${ENV_INJECTED_TOKEN}
Control (one home, --single): launch config still resolves env_ref from os.environ.
2026-09-15 03:46:29 -07:00
Teknium
d26dbb2f77 Port from OpenHands/OpenHands#16758: Anthropic model catalogs no longer stop at the first page
Anthropic's /v1/models is cursor-paginated with a default page size of 20.
Both hermes fetchers read a single unpaginated page, so any model past the
first page silently vanished from the /model picker and provider catalogs.

- hermes_cli/models.py _fetch_anthropic_models(): request limit=1000 and
  follow has_more/last_id (bounded, repeated-cursor guarded, de-duped)
- plugins/model-providers/anthropic fetch_models(): same pagination walk,
  and it now honors the base_url argument instead of hardcoding
  api.anthropic.com
- tests: live-HTTP paginated-server regression tests for both fetchers,
  incl. single-page and stuck-cursor termination; updated the two URL-pinning
  pool-discovery tests for the ?limit=1000 contract
2026-09-13 21:07:35 -07:00
kshitijk4poor
db66b21ae4 fix(copilot-acp): honor the fetch_models None-on-failure contract; bound the probe budget; refresh-proof the memo
Gate findings on the #108614 salvage (2c + simplify-code):
- CopilotACPProfile.fetch_models promised None on failure but raised (AuthError on a
  missing CLI, RuntimeError/TimeoutError from the probe) — only the caller in
  hermes_cli/models.py caught it; any other caller got exceptions. Wrap the body and
  return None so the docstring and the base ProviderProfile.fetch_models contract hold.
- timeout_seconds was a PER-REQUEST budget (initialize + session/new each got the full
  timeout => ~30s worst-case foreground stall on a hung CLI). One shared session deadline.
- /model --refresh (clear_provider_models_cache) wiped the disk cache but not the new
  session memo, so a fresh CLI login stayed stale for 5 min past an explicit refresh.
- A failed probe is now memoized for 30s instead of the full 5 min, so signing in to
  the CLI is picked up on the next switch.
- _fresh_acp_memo fixture restores the memo on teardown — no cross-test state leak.
2026-09-13 19:30:48 +05:30
kshitijk4poor
d5ceff958d fix(models): memoize the Copilot ACP session probe; source it from the provider profile
`/model <x>` onto copilot-acp validates through `models_validate._static_catalog`, which
reads `provider_model_ids` with no disk cache. After the session probe landed, every such
switch spawned `copilot --acp`, ran the handshake, and killed it (1-3 s; up to the 15 s
probe timeout when the CLI is installed but the session stalls). The GitHub-API tier that
path used before sat behind a 5-minute in-memory memo; the ACP tier now has the same memo,
and it remembers failures too so a broken CLI is not re-spawned per switch.

The probe itself moves to `CopilotACPProfile.fetch_models` — the slot that already said
"model listing is handled by the ACP subprocess" and returned None — so hermes_cli/models.py
no longer hand-builds `CopilotACPClient` kwargs that `profile.create_client` owns.
Discovery failures are logged at debug instead of swallowed.

Tests: the two picker wiring tests collapse into one parametrized contract; a new test
proves three consecutive switch validations pay one probe and a failed probe is not retried
(fails when the memo read is removed).
2026-09-13 19:17:06 +05:30
Gille
0cd897286a fix(models): discover Copilot ACP session catalog 2026-09-13 19:17:06 +05:30
teknium1
d364473620 refactor(model-providers): fetch_models stubs drop to the ABC; commandcode and AI Gateway catalog GETs go through open_credentialed_url
vertex, bedrock and copilot-acp each overrode ProviderProfile.fetch_models with an identical `return None`, and the overrides could not simply be deleted because the base implementation derives a URL from base_url and would GET e.g. bedrock-runtime.../models or acp://copilot/models. ProviderProfile now carries `supports_model_listing` (default True); the base fetch_models returns None before touching the network when it is False, and the three SDK/subprocess-backed profiles set it in their constructors instead of overriding. Separately, two catalog GETs still went through bare urllib.request.urlopen: commandcode's fetch_models and hermes_cli.models._fetch_ai_gateway_models (which sends the AI Gateway bearer). Both now use the redirect-safe open_credentialed_url path (_urlopen_model_catalog_request in models.py, also applied to the unauthenticated fetch_ai_gateway_models for consistency). This is a security behavior change: a cross-origin redirect from either endpoint no longer forwards the Authorization/attribution headers to the redirect target. The AI Gateway tests that patched the global urlopen are repointed at hermes_cli.models._urlopen_model_catalog_request (the seam tests/hermes_cli/test_models.py already uses), the commandcode test patches open_credentialed_url on the plugin module, and two invariant tests pin the no-network guard and the bearer routing.
2026-09-13 05:19:48 -07:00
Teknium
5ff34f565e fix(multiplex): per-profile catalog, skin and guest-mint state in hermes_cli
DeepInfra catalog (fetched with the launch env's key via os.getenv), Copilot
context limits (api_key ignored on hit), Nous reasoning caps + once-per-process
guards, the curated OpenRouter list, the model-catalog in-process copy (mtime
without path), banner skills, the guest-mint back-off flag and the active skin
were single slots read under per-profile overrides by the gateway and the TUI
gateway; the SWR refresh thread ran without the caller's ContextVars.

Under an override each lives per home key (hermes_cli/models_profile_cache.py
holds the shared slot helper so models.py does not grow), credentials are read
through the scope-aware dotenv reader and keyed by fingerprint, and background
refreshes run under copy_context(). Unscoped behaviour is byte-identical.
2026-09-12 01:35:05 -07:00
Teknium
30e4673776 refactor(models): import openrouter_variant_base from its defining module
Internal moves get no re-export aliases: models_validate reads
hermes_constants.openrouter_variant_base directly and the private
_OPENROUTER_VARIANT_SUFFIXES/_openrouter_variant_base shims are dropped.
2026-09-11 16:50:20 -07:00
jakobdylanc
284ee03df1 fix(models): resolve context length for OpenRouter :nitro/:floor routing variants
`:nitro`, `:floor`, `:exacto`, and `:online` are request-time routing
modifiers, not catalog models — OpenRouter's /models lists only the base
id, and a variant runs the same model with the same context window.

`get_model_context_length()` keyed every lookup on the full suffixed id,
so each one missed and the resolver fell through to a generic family
default or the 256K fallback:

  openai/gpt-5.5:nitro           -> 256K   (real 1.05M)
  x-ai/grok-4.6:nitro            -> 131K   (generic "grok" catch-all)
  anthropic/claude-opus-4.6:nitro-> 200K   (generic "claude" catch-all)

The window silently shrank, triggering early compression and a wrong
/usage readout. f14059fa fixed the sibling half of this bug class in
/model validation; this fixes the metadata half.

Strip a recognized variant suffix for LOOKUP only, keeping the suffixed
id on the wire so the routing opt-in survives. Applied after the explicit
config overrides (steps 0b/0c) so a user-pinned value still wins, and
before every cache/catalog lookup. Gated on the request actually routing
through OpenRouter, so a local Ollama `model:tag` is untouched.

`:free`/`:batch`/`:thinking` are deliberately excluded — those ARE
distinct catalog SKUs with their own windows, so stripping them would
report the wrong number.

The suffix set and base-id split move to hermes_constants (import-safe,
dependency-free) so the metadata layer shares one definition with
hermes_cli.models instead of duplicating it.
2026-09-11 16:50:20 -07:00
Siddharth Balyan
04a76c4109 fix(auth): a sign-in from the free tier settles the default model and route (#105259)
* fix(auth): a sign-in from the free tier settles the default model and route once, for every caller

Picking the free-tier row leaves model.default at nous/welcome pinned to the welcome host. After a
sign-in an account cannot keep either: the welcome host refuses account tokens, and the portal host
serves nous/welcome as a paid model. One completion step, anon_auth.settle_after_upgrade, now runs
after the account is persisted: a config on the free tier's route moves to the account's inference
host and the recommended default for the account's plan, through the same config write a plain
Nous login uses; a config on the user's own model is left alone. The pick is the one
GET /api/model/recommended-default already makes, factored into models.recommended_nous_default_model
so the CLI and the desktop land on the same model. hermes auth upgrade prints the new default.

* fix(auth): a sign-in completion with no eligible recommendation leaves no default model

The static provider-wide default is not narrowed by the account's plan or org policy, so writing it
as a fallback could persist a model the account may not use. When the recommendation cannot yield a
model, the route still moves to the account's host but model.default is left unset; the CLI says so
and points at `hermes model`.

* docs(free-tier): say what happens when no recommendation is available after sign-in

* fix(auth): sign-in completion moves the host and clears the default in one config write

Two writes could fail between them and leave the account host paired with nous/welcome.
_update_config_for_provider gains clear_default so the caller with no model to offer removes
model.default in the same atomic write that sets the host.
2026-09-11 03:45:31 +05:30
Teknium
9a3a15c87a fix(models): a vendor's own id on its first-party provider never re-routes
215fd0ecb9 made the current provider's live catalog outrank static
guesses, but a live catalog that could not be fetched (Codex outage, cold
cache) looked identical to "not served", so `/model gpt-6-astra` on
openai-codex still walked to OpenRouter (which relists every vendor)
whenever an OpenRouter key existed. Same shape for grok-* on xai-oauth
and claude-* on anthropic.

detect_provider_for_model now treats a vendor id typed on that vendor's
single-vendor first-party provider as a selection: stay, let the vendor
accept or reject it (the existing "not found in listing" note still
fires). Cross-vendor remaps (claude id on Codex -> keyed OpenRouter) and
aggregator / custom / multi-vendor reseller sessions are unchanged.
2026-09-10 09:11:00 -07:00
Teknium
2466684db5 fix(models): never auto-switch to a provider the user has no credentials for
/model <name> on provider A, where the name is only known to provider B
(static catalog or OpenRouter), switched the session to B even when B had
no key: an immediate 401 for most vendors, and for OpenRouter — whose
runtime resolves with an EMPTY key instead of raising — a silent switch
onto a metered aggregator. The dashboard's flat Model field had two more
copies of the same guess ("vendor/model on a native provider" → openrouter).

detect_provider_for_model() now walks its ladder as candidates and skips any
target without credentials (env/.env key, auth-store login, or a usable
credential pool entry). Exceptions: the user NAMED the provider (/model nous)
or there is no current provider yet ("auto") — then the guess is handed back
so the credential step fails loudly instead of silently ignoring input. A
vendor/ prefix naming a provider declared in `providers:` is a selection, not
a guess, and always routes. The dashboard fallbacks apply the same gate.

Tests that pinned "switch to OpenRouter/vendor with no key" now grant the
credential they assumed; two new invariants cover the gate.
2026-09-10 03:58:37 -07:00
Teknium
215fd0ecb9 fix(models): a model the current provider serves never re-routes to OpenRouter
detect_provider_for_model() walked static catalogs and then the OpenRouter
catalog. Providers whose live catalog outruns the static list (Codex
early-access ids, Nous Portal slugs, Ollama Cloud) had nothing to stop the
ladder, so `/model gpt-6-astra` from a Codex session, or `/model glm-5.3-flash`
from a Nous-sub-only session, silently rebuilt the agent on metered OpenRouter
(or on a vendor the user has no key for).

The current provider's disk-cached live catalog (cached_provider_model_ids, 1h
TTL + SWR) is now consulted first: an exact or bare-name hit stays put and
resolves to the provider's own spelling. Fixes the whole class at the shared
choke point, covering /model, ACP, oneshot and the dashboard.

Refs #97487 (sj-unit72 diagnosed the ollama-cloud instance of this).
2026-09-10 02:44:26 -07:00
teknium1
53b9615ea6 test(hermes_cli): keep one opencode-go delisting invariant; record the delisting beside the keyed-suffix set
Drop the floor-pin test from #106290 (a snapshot of the static table); the
end-to-end merge test through provider_model_ids with the real curated floor
already fails if the slug is re-added. Note beside
_OPENCODE_FREE_KEYED_SUFFIX_MODELS why ox-alpha-free stays excluded from the
keyless catalog even though it left the curated floor.
2026-09-09 11:51:00 -07:00
kshitijk4poor
e2bd400233 fix(models): only cache unreachability, not HTTP errors; key the entry via base_url_origin
An HTTPError means the host answered — a 401 from a wrong API key must not
be remembered as "unreachable" for the next 60s, or a user who fixes the key
gets a cached empty catalog on the immediate re-probe. Connection-level
failures (timeouts, refused, DNS) are the only thing the cache records.

_probe_neg_key hand-rolled scheme/port defaulting that utils.base_url_origin
already provides; use it.
2026-09-09 21:16:48 +05:30
kshitijk4poor
4e2a871adb fix(models): clear probe negative-cache entry on a successful probe 2026-09-09 21:16:48 +05:30
finn763
48b8528e7c fix(desktop): stop UI freeze on unreachable provider Closes #81123 2026-09-09 21:16:48 +05:30
kshitijk4poor
1eaeb73839 fix(models): OpenRouter disk snapshot follows the configured catalog TTL, not the default constant
refresh_interval_seconds() honours model_catalog.ttl_minutes / legacy ttl_hours;
reading DEFAULT_TTL_MINUTES would let the snapshot and the manifest it is
filtered from go stale on different clocks for anyone who changed the TTL.
2026-09-09 21:16:28 +05:30
kshitijk4poor
8b66bc88df refactor(models): OpenRouter disk cache reuses catalog TTL and atomic JSON cache helpers 2026-09-09 21:16:28 +05:30
Hyperion
4c003069f3 fix(models): persist curated OpenRouter catalog to disk so picker opens fast
Re-derived from PR #96099 (f127ec4e) on current main; the :nitro/:floor validate hunk is omitted because main already handles routing suffixes in hermes_cli/models_validate.py.
2026-09-09 21:16:28 +05:30
Teknium
734461d213 fix(models): same-URL custom endpoints stop evicting each other's cached catalog; no-probe picker opens revalidate
Two picker-freshness defects in cached_fetch_api_models():

1. The disk cache row was keyed on base_url only, with the credential
   fingerprint stored inside the row. N custom_providers entries sharing
   one proxy URL with different keys (#106184) took turns overwriting the
   single slot; every sibling then failed the fingerprint check, got an
   empty catalog, and disappeared from the Desktop pickers (which hide
   zero-model rows). Key on url#fingerprint so each credential owns a row.

2. cache_only opens (Desktop model.options without refresh) served a
   past-TTL row for up to 7 days without ever revalidating, so a model
   loaded on a non-current local endpoint stayed invisible until the user
   found "Refresh Models". Serve the stale row AND spawn the same
   off-thread SWR refresh the blocking path uses; the caller still never
   waits on the network.

Live repro (two rows, one URL, keys A/B; real loopback /v1/models):
  GUI no-probe open  before {'proxy-a': ['model-A1'], 'proxy-b': ['model-B1']}
                     after  {'proxy-a': ['model-A1','model-A2'], 'proxy-b': ['model-B1']}
2026-09-09 03:33:06 -07:00
Teknium
520e63661c fix: keep command-auth model discovery lazy across config and setup 2026-09-07 21:22:49 -07:00
kshitijk4poor
e583d68c14 perf(models): stop revalidating a fresh provider cache just because it lists Astra
``gated_cache`` bypassed both the fresh-hit and the stale-while-revalidate branches of
``cached_provider_model_ids`` whenever the disk entry contained Astra. For an entitled
account that is every entry: live discovery rewrites the entry with Astra → next call is gated
again → a blocking /models round-trip on every picker open, for exactly the providers the user
is most likely on. The parallel prefetch's staleness check didn't know about the gate either, so
it skipped the slug and the blocking call landed in the serial picker loop the prefetch exists to
avoid.

Gate on provenance, not contents: only live discovery ever writes Astra into a same-credential
entry (the static/offline paths filter it out), so a fresh entry IS the entitlement record. The
one filter that matters stays — a stale entry served because the refresh failed drops Astra, and
the entry itself is left intact so the next successful fetch restores it.

The regression test now pins both halves: fresh entry served with Astra and zero live calls;
failed refresh past the fresh window serves the entry minus Astra.
2026-09-07 21:43:54 +05:30
kshitijk4poor
b6852995ed refactor(openai): one is_astra_model predicate; gate Astra at the untrusted Codex inputs
The slug pair {"gpt-6-astra", "gpt-6-astra-900k"} was spelled out five times
(reasoning_effort, transports/codex, models.py x2, codex_models) with five hand-rolled
``.strip().lower().rsplit("/", 1)[-1]`` normalisations. ``is_astra_model`` in
agent/reasoning_effort.py (the module the other four already import) is now the only home, so a
new Astra alias is one edit.

``_finalize_codex_models(..., allow_astra=)`` was a control-coupling flag set True by the two
live callers and left False by the two static ones. The filter now lives where the untrusted
inputs are — ``_drop_undiscovered_astra`` over config.toml default + models_cache.json in
``get_codex_model_ids`` — and ``_finalize_codex_models`` is back to its one-line original.
DEFAULT_CODEX_MODELS never contains Astra, so the static catalog needs no gate.

``_openai_catalog`` appends whatever Astra ids live discovery returned instead of a name-keyed
``if "gpt-6-astra" in live_lower`` branch with a literal fallback that could never fire.
2026-09-07 21:43:54 +05:30
Eva
fdb76ce643 fix(openai): include canonical API picker identity in Astra discovery gate
(cherry picked from commit de57e601193a955a69201e1a274d5ccab83a7634)
2026-09-07 21:43:54 +05:30
Eva
328542d807 fix(openai): revalidate Astra at cached and saved-model picker boundaries
(cherry picked from commit f46e2f4c0609c1543d5636c780c6da0d5ce09a56)
2026-09-07 21:43:54 +05:30
Eva
2c315ff59b feat(openai): add GPT-6 Astra baseline support
(cherry picked from commit a8c53d20c6b16cc35745e364e16bb7259166a1d3)
2026-09-07 21:43:54 +05:30
Teknium
d63e380324 compat(plugins): warn once per name when a plugin resolves an old import path; lint step restored in CI
Every PLUGIN-COMPAT __getattr__ now calls hermes_cli.plugin_compat.warn_once(facade, name, target) before
resolving, emitting a HermesPluginCompatWarning (FutureWarning) once per process per name: old path, new
path, removal target. Importing a facade for its live API stays silent; only resolving a moved name warns.
COMPAT_MANIFEST.md documents the warning and how to silence it during migration.

Verified the runtime never routes through a pointer: every entry point (run_agent, cli, hermes_cli.main,
gateway.run, tui_gateway.server, web_server, model_tools + tool discovery, hermes_state, cron.scheduler,
browser_tool, mcp_tool, kanban, auth) imports clean and `hermes doctor` runs end to end with the warning
promoted to an error.

Also restores the check_compat_pointers CI step to .github/workflows/lint.yml, which a0be177aac dropped
when the compat layer was regenerated (the lint script itself was present; the workflow step was not).

hermes_cli/plugin_compat.py, tests/test_plugin_compat_warning.py and the two-line insert per facade are
part of the compat layer and go away with it.
2026-09-04 00:15:16 -07:00
Teknium
2776813df3 compat(plugins): temporary import-path shims for external plugins — ONE commit, revert on schedule
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:

    git revert <this sha>

removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.

What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
  so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
  tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
  relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)

Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
2026-09-03 17:13:22 -07:00
Teknium
7a33369e81 simplify(compat): interrupt — drop _ThreadAwareEventProxy/_interrupt_event legacy alias, repoint 2 test files
No runtime consumer read the proxy (terminal_tool/environments call is_interrupted()/set_interrupt()
directly); its only users were tests patching tools.interrupt._interrupt_event, which had no effect on
the code under test. tools/terminal_tool.py's own re-export of the name is owned by another worker.
2026-09-03 14:00:59 -07:00
Teknium
c93ace77c2 simplify(compat): config/runtime_provider/plugins/commands/secrets_cli/kanban — drop 96 re-exports (incl. PEP 562 facades) + 3 aliases (get_pre_tool_call_directive/_block_message, get_telegram_handler_factories), repoint 56 callers + 50 test files 2026-09-03 14:00:17 -07:00
Teknium
7b8c11bcf7 simplify(compat): models — drop 52 re-exports from hermes_cli.models, repoint 16 callers + 41 test files 2026-09-03 13:48:49 -07:00
Teknium
d179f28307 simplify(compat): anthropic_adapter — drop 30 re-exports + 1 alias, repoint 22 caller files (32 sites), 38 test files (~125 sites) 2026-09-03 13:16:47 -07:00
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
34abf954bd review-fix(public-api): restore get_session_activity, latest_user_message_row_id, resolve_multiple_toolsets, has_provider, nous_token_has_billing_scope, curated_models_for_provider, clear_edit_approval_requester + tests
All public on BASE 63279301bc, dropped by the simplify refactor (their tests were deleted or
rewritten to the replacement API). Restore each with BASE signature/body as a thin wrapper over the
surviving implementation, and restore the tests at the original call sites: test_message_reactions
again asserts the role=user contract (a newer assistant message is never the default target);
test_hermes_state / test_watchdog_review_76354 go back to get_session_activity(); toolsets, acp auth,
edit_approval, billing-scope and curated-models tests restored/extended.
2026-09-03 09:40:49 -07:00
Teknium
fb14bc4e11 review-fix(whitespace): strip trailing whitespace and EOF blank lines introduced by this PR
Trailing-whitespace-only edits so 'git diff --check BASE HEAD' is clean
(16 diagnostics across 11 files). No code changes.
2026-09-03 09:31:54 -07:00
Teknium
022785a541 Merge origin/main (63279301bc): reasoning-mandatory 400 recovery folded into turn_recovery/error_classifier/models_reasoning_caps 2026-09-03 05:36:21 -07:00
Teknium
f6bd1633f7 fix(reasoning): GLM-5.3 on Nous/OpenRouter no longer 400s when thinking is disabled
Reasoning-mandatory routes answer reasoning: {enabled: false} with HTTP 400
"Reasoning is mandatory for this endpoint and cannot be disabled". Hermes
sends that disable for /reasoning none, agent.reasoning_effort: none, and the
one-shot thinking-exhaustion continuation override (which GLM-5.3-flash
triggers on its own). The Nous profile's catalog guard swallows the disable
only when its per-process capability cache already says mandatory; a gateway
that warmed the cache before the route flipped kept sending it, and the 400
was classified as a non-retryable format_error that aborted the turn.

- error_classifier: new reasoning_mandatory reason (retryable, no fallback,
  no compression), matched before the request-validation branch.
- conversation_loop: one-shot recovery — set agent._reasoning_disable_rejected,
  queue a catalog refresh for the provider, retry.
- chat_completion_helpers: _reasoning_config_for_wire drops every disable
  (configured or ephemeral) once the route has rejected one.
- hermes_cli/models: refresh_reasoning_caps_async(provider) forces a
  background re-fetch of the Nous/OpenRouter catalog.
- openrouter profile: omit a disable when the catalog marks the route
  mandatory (parity with the Nous profile).

Live: z-ai/glm-5.3-flash on the Portal with a poisoned mandatory:false cache.
Before: turn aborted with the 400. After: one retry, thinking stays on, turn
completes.
2026-09-03 05:20:11 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
mr-r0b0t
7e60d0c042 feat(models): add Meta Muse Spark 1.3 family to picker
Add meta/muse-spark-1.3 and meta/muse-spark-1.3-contributor to the
OpenRouter curated list, the meta-ai provider fallback, the
opencode-zen / opencode-free / opencode-go floors, the setup-wizard
shortlist, and regenerate the hosted model catalog.
2026-09-03 00:57:55 -07:00
Christian Reyes
7f2aa70add feat(models): add Muse Spark contributor to OpenRouter 2026-09-03 00:57:55 -07:00
fangliquanflq
401ac7e1f8 fix(gateway): keep first model picker open responsive on cold pricing cache
Picker opens use only process-resident pricing (cached_only) and start a
single-flight daemon prewarm keyed by (profile, endpoint scope); explicit
refresh stays synchronous. Nous fails closed (free_tier_pending) until the
entitlement is known so a free account cannot briefly select paid models.
Free-tier cache becomes per-profile.

Squash of the 5-commit PR #92253 branch (d5b2070ef8..28313ff963), applied
via diff on current main; two adjacent-insertion conflicts resolved by
keeping both sides.
2026-09-03 13:08:31 +05:30
Teknium
2a9481274c refactor(hermes_cli): models.py — flatten provider_model_ids / profile catalog fall-through 2026-09-02 21:34:05 -07:00
Beto de Paola
83ecb6e695 feat(meta-ai): live-first model catalog + generic contributor warning
- Register meta-ai as a live-first picker provider so the /v1/models catalog
  leads the picker; new models appear without a PR
- Override fetch_models to exclude non-chat models (muse-image-*, muse-voice-*)
  from the picker; new chat model families pass through automatically
- Slim fallback_models to a single safety-net entry (muse-spark-1.2), shown
  only when the live fetch fails
- Make data-policy contributor warning model-generic (not hardcoded to 1.2)
  so it covers any future -contributor model
- Update test assertion to match generic warning text

LOCAL ONLY — pre-launch, not for push.
2026-09-02 21:22:53 -07:00
Teknium
43eccd23eb refactor(hermes_cli): models.py — copilot token source table, shared api-key credential lookup 2026-09-02 21:02:21 -07:00