Commit Graph

448 Commits

Author SHA1 Message Date
kshitijk4poor
804cc3c021 fix(gateway): a standalone owner is the per-profile topology, not a refusal
`gateway/host_attach.py::decide()` treated a host owner that answers `multiplex: False`
to the rescan — another profile's standalone gateway — the same as a multiplexer whose
roster excludes us, and issued the permanent REFUSE. On the supervised path that is
exit 78, which `hermes_cli/stderr_timestamp.py` maps to 0 so launchd's
`KeepAlive.SuccessfulExit=false` parks the unit. On a one-process-per-profile fleet
(the topology `multi-profile-gateways.md` documents and `multiplex_profiles: false`
promises to keep) every launchd gateway except the first to claim the host lock was
parked at boot, silently; which ones survived was a boot race against the owner's
record landing.

`request_serve_profile` now returns the owner flagged `standalone` instead of `None`
for a `multiplex: False` answer, and `decide()` turns that into START — this profile
runs its own gateway beside the owner, as before #118097; the host-lock claim still
logs the two-gateway topology and `gateway migrate --multiplex` stays the converge
path. REFUSE is unchanged for a multiplexing owner that excludes the profile.

A/B against a REAL owner process (real host lock, record and control socket answering
the rescan): base → `refuse`, exit 78 → launchd 0 (parked); fix → `start`. Negative
control (owner answers `multiplex: True`, roster excludes the profile): `refuse` on
both.

Field report: debug share 5024e996 (11 launchd profile gateways, 0.21.3 ea0c2b82 —
only 2 of 11 came back after `hermes update`), discussed on #118097.
2026-09-21 08:38:15 -07:00
teknium1
a10620a669 fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.

- ATTACH now requires a LIVE `identify` answer. The claim-time record is
  published with NO served set (the runner settles multiplex a moment later),
  and `host_gateway()` reports `served_known=False` when nothing answers. An
  owner whose served set is unknown yields a TRANSIENT refusal, never an
  attach: previously `default`'s claim published "default,other" before its
  socket bound, `other`'s systemd unit read that as "I am served", exited 78,
  and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
  instead of forcing `multiplex=True`, so a standalone gateway stops claiming
  the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
  `--force` into `_host_attach_or_none`. Both previously exited in the guard
  before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
  config refusal every supervisor parks on; "someone else serves me right now"
  is a runtime observation that ends when that process does. No unit files
  change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
  finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
  would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
  and re-enters with `replace=True`, so it can no longer attach to the corpse
  it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
  `st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
  it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
  generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
  `gateway status`/doctor across N profiles pays one probe, not N.

Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
2026-09-21 05:02:29 -07:00
teknium1
a20a88398f gateway: lifecycle verbs mean "the one host multiplexer"
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:

- `gateway run` for a profile the host gateway already serves ATTACHES: print
  its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
  to re-scan `profiles/` (control socket) and attach once the answer includes
  it. Refuse only when the host gateway cannot be made to serve it. Under a
  service supervisor the attach exits 78 instead of 0 so a redundant unit is
  parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
  landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
  the owner's control socket, so it no longer depends on the claim ordering in
  start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
  process: they restart the host multiplexer and preserve its served set. A
  secondary still running its own gateway is reported with the
  `gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
  argv: a host singleton runs bare/default argv and can never prove it serves
  profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
  multiplexer is whichever profile launched the one host process.

Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
2026-09-21 05:02:29 -07:00
teknium1
bfcd209906 docs(multiplexing): record the config escape hatch and the .env-over-frozen-env precedence
Activation is not purely stricter for the LAUNCH tenant: inside the launch scope its own
.env wins over the frozen boot env, so a key present in both reads differently before and
after activation. Also documents the explicit multiplex_profiles: false escape hatch and
the fail-closed unresolved-profile case.
2026-09-21 02:58:40 -07:00
teknium1
c718267ef3 fix(multiplex): close the profile-scope holes the scope machinery misses
One host process serves every profile, so every execution point must bind the
profile it is acting FOR. These six ran unscoped (or bound only part of a
scope) and resolved get_hermes_home()/credentials against the LAUNCH profile:

- tui_gateway/session_reaper: the idle-reaper and exit-flush transcript writes
  now enter the SESSION's profile scope, the same chokepoint _finalize_session
  already binds. A served profile's transcript was landing in the launch home.
- gateway/run: MCP shutdown tears down per served profile inside that
  profile's scope (mirrors startup discovery and the reconcile chore), with a
  trailing wildcard pass under the launch profile's own scope.
- gateway/run_profile_reconcile: _unserve_profile's adapter teardown, agent
  eviction and state/memory handle release now run inside the deleted
  profile's scope.
- gateway/run_adapters + run_goals + run_notifications: a body with no routed
  profile no longer means "no scope". launch_profile_scope_if_multiplexed()
  binds the launch profile once the process multiplexes; before activation it
  is still literally a nullcontext, so single-profile hosts are unchanged.
- hermes_cli/kanban_db_dispatch: one _worker_profile_scope helper binds the
  assignee's secret AND terminal scope for toolset resolution and the spawn-env
  build, unconditionally instead of only under multiplex.
- hermes_cli/web_server: `hermes serve` activates multi-profile hosting at boot
  when the host has more than one servable profile home, instead of lazily on
  the first ?profile= request after earlier work already ran unscoped.

Secret scope is never widened: a non-launch home resolves from its own .env and
sources only; the launch home keeps its existing env-over-.env precedence.
2026-09-21 02:58:40 -07:00
teknium1
9ecd22e5da fix(cron): release a profile's in-flight claim under the key it registered with
Under one multiplexing ticker, every non-launch profile's in-flight claim
leaked on every run. The cron home scope is a ContextVar: the claim is taken on
the ticker thread inside `_profile_cron_scope`, but the pool worker's `finally`
sits outside `ctx.run`, where the worker thread resolves the LAUNCH home — so
the discard missed the real key. The job then skipped a fire window until the
force-release backstop swept it, and the shutdown drain plus
`hermes.cron.jobs.running` saw phantom work.

`_submit_with_guard` now captures the registering home and passes it to
`release_running_job(job_id, home=...)` on every release path.

Also in this pass:

- Stand down for a profile that runs its OWN gateway (`_cron_profile_gate`,
  the gate the serve/Desktop ticker already passes). On a host pinned to
  per-profile gateways both processes raced that profile's tick lock, and when
  the launch gateway won, delivery went through SharedRouteAdapters/fail-closed
  instead of the profile's live adapters. The gate compares the liveness PID
  against `os.getpid()`: this process holds the launch `gateway.pid` and
  publishes every served profile in `served_profiles`, so a bare liveness answer
  would have stood cron down host-wide.
- Job-liveness consumers ask `is_job_running(job_id, home=...)` instead of the
  host-wide bare-id union, which let profile A's running `daily-brief` report
  profile B's idle one as running and keep B's stale one-shot alive.
- `register_ticked_homes` reaps the parallel pools of homes that leave the
  ticked set; pools lived until `atexit`, so every home ever ticked kept a
  ThreadPoolExecutor and its worker threads.
- `mark_running_jobs_interrupted` reads the real home Path from
  `_inflight_home_path` instead of rebuilding it from the normcased key half.
- The ownership test imports `cron.scheduler_ownership` per test, so the file
  fails behaviourally on base instead of as a collection error.
2026-09-21 02:52:15 -07:00
teknium1
cb647a018f fix(cron): one host ticker owns every profile's cron, per profile
The cron ticker multiplexes N profiles from one process while its ownership
predicates and its in-flight bookkeeping still assumed one profile per process.

- `_should_yield_tick_to_fresh_gateway` asked a process-global boolean
  (`owns_gateway_runtime_lock`) and a launch-home lock probe, so one answer
  covered every profile ticked. It now asks `scheduler_ownership`:
  `owns_cron_tick_for(home)` (this process is the host gateway AND ticks that
  home) and `live_gateway_ticking(home)` (another live host gateway whose
  published served set covers that home).
- In-flight state (`_running_job_ids`, `_running_since`, `_running_futures`,
  `_running_allowance_s`, `_running_worker_pids`, `_running_fire_owners`,
  `_restart_safe_waiter_job_ids`, `_interrupted_job_ids`) is keyed by
  `_inflight_key(job_id)` = `(home key, job id)`; two profiles carrying a
  `daily-brief` no longer read as one job. The public accessors still report
  the host-wide union of bare job ids for the shutdown drain.
- The parallel worker pool is keyed by home: `cron.max_parallel_jobs` is a
  per-profile key, and the single global pool was sized by whichever profile
  ticked first and torn down by the next one.
- `gateway/run.py` no longer gates the cron tick set on
  `gateway.multiplex_profiles`: that flag gates adapters, and with it off every
  non-launch profile's jobs sat in a store no ticker visited.
2026-09-21 02:52:15 -07:00
teknium1
8863b36fd6 fix(cron): stale-code yield reads as an outage in cron status; hermes update restarts proven-stale gateways
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:

* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
  loop) and the success marker; with no success marker on disk it printed
  "✓ Gateway is running — cron jobs will fire automatically", and with a stale one
  it pointed at "Check the gateway log" instead of naming the cause. The persisted
  `CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
  and reported as "Gateway is running STALE code — fires NOTHING", with both
  revisions and the restart command. A fresh heartbeat with a recorded error and no
  success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
  left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
  to the existing drain-first `request_restart` path (SIGUSR1) via
  `hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
  code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
  contract the restart phase already uses for unmapped manual gateways. The drain
  budget computation is shared (`_gateway_drain_budget`).

The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).

Fixes #117275
2026-09-21 01:57:31 -07:00
teknium1
b7803a1763 fix(compression): cap the protected tail at 20% of the context window
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole
rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the
"protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every
compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B
timed out before compaction ever changed anything, and protect_last_n read as an uncompressed
tail rather than a minimum.

TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk /
pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool
groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of
32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%).

Probe (12 tool-heavy turns, 49 rows, 12.8K tokens):
  ctx    8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows;  window [4,10) -> [4,42)
  ctx   16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34)
  ctx   32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22)
  ctx  128K: identical before/after
2026-09-21 01:04:16 -07:00
teknium1
e089be80a1 docs(desktop-sdk): note DecodeText loop is opt-in
DecodeText is a published plugin-SDK export; flipping its `loop` default
to false silently changes third-party plugin visuals, so the SDK doc says
so where the component is listed.
2026-09-20 20:14:20 -07:00
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
bdde0b0c28 docs(cron): describe when a stale-code ticker yields and when it keeps dispatching
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped
code_sha is the on-disk revision; a lock held by an equally stale process never
counts. Placed under "Gateway Integration" so it does not collide with the
"Stale-code yield" section #117501 adds under "Locking".
2026-09-20 20:13:19 -07:00
teknium1
13fe9c7171 feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:

- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
  `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
  declaring profile; standard records still replay on OpenRouter-style routes, strict routes
  drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
  status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
  and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
  one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
  without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".

The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.

Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
41b6ba09d9 docs: name the registered tool in prose that still says cronjob/todo/process
The same rename that left the two reference pages stale (cronjob->cronjob_manage,
todo->todo_list, process->process_manage) left the guides and developer docs
referring to the old tool names; a reader following them gets "no such tool".
Toolset names (`cronjob`, `todo`) are unchanged and left alone.
2026-09-20 12:56:25 -07:00
teknium1
297ebd506e test: trim the ambient-chain refusal tests to two two-home invariants
The 19 cherry-picked tests covered each refusal branch separately. One
A -> B -> A test per adapter over two real homes now proves the whole
contract at the production entry (build_credential / _cached_client): the
launch profile keeps its own credential, the cred-less served profile is
refused before the SDK chain (or boto3) is touched, the launch profile is
unaffected afterwards, and the standalone run keeps today's ambient chain.
Docs: the Azure guide and the multiplexing design page name the refusal.

Superseded #116370 (@JoaoMarcos44) proposed the same mechanism.
2026-09-20 12:09:47 -07:00
teknium1
948753d0f8 docs(plugins): enable prompts for the override grant only when the manifest declares capabilities 2026-09-20 11:52:45 -07:00
teknium1
967a7c94f7 docs(plugins): a provider plugin can register its own api_mode transport 2026-09-20 10:11:40 -07:00
teknium1
e5e7fbcd27 fix(auth): a user provider plugin's endpoint overrides the bundled row too
A `$HERMES_HOME/plugins/model-providers/<name>/` plugin re-registering a
bundled provider (stepfun with a regional base_url, gmi at a staging host)
wins in `providers._REGISTRY` — register_provider() is last-writer-wins and
the plugin guide promises exactly this — but the runtime reads its endpoint
from `hermes_cli.auth.PROVIDER_REGISTRY`, whose mirror loop skipped every
name already present, so inference kept going to the built-in URL (#48450).

The mirror now applies one explicit precedence rule: when a row core wrote
(built-in or plugin-mirrored) belongs to a name whose profile
currently registered came from a USER plugin,
the row's profile-derived fields are rewritten in place (inference_base_url;
api_key_env_vars / base_url_env_var on api-key rows when the profile declares
env_vars). `providers` records the discovery source per registration
(`provider_source()`), because a bundled profile must never rewrite a
built-in row: several bundled profiles omit the row's `*_BASE_URL` env var
and one differs in auth_type, so an unconditional "profile wins" would have
changed built-in behaviour. With no user plugin PROVIDER_REGISTRY is
byte-identical before/after (78 rows probed). copilot/kimi/zai keep their
bespoke resolution via the existing skip set.

Co-authored-by: xiaoxinova <xiaoxinova@users.noreply.github.com>
2026-09-20 10:06:29 -07:00
teknium1
9f0cd9d773 docs: plugin profiles and their aliases resolve in /model --provider and the model picker 2026-09-20 10:01:58 -07:00
teknium1
2182f51d7c fix: name the process holding the state.db write lock when a writer times out
"database is locked (another Hermes process held the state.db write lock for
over 60s)" identified the victim only. The open-descriptor scan cannot single
out the writer because every Hermes process (gateway, CLI sessions, worktree
agents, cron) has the DB open, so an operator hit repeatedly by
session_persistence_failed:locked had nothing to act on.

SQLite's unix VFS takes fcntl byte-range locks whose offset encodes the lock
kind (state.db-shm byte 120 = WAL write, 121 = checkpoint; the pending-byte page
on state.db = PENDING/RESERVED), and the kernel exports them with the owning pid
in /proc/locks. hermes_state_lockowners reads that table at the moment the
patience budget runs out and logs one WARNING per write-class holder with
describe_holder_pid()'s argv summary, for both the transcript write path and
open+init lock patience. The holder stays out of the exception text on purpose:
classify_persistence_error() buckets by phrase and a holder argv such as a
worktree named fix-corrupt-db would flip the bucket.

Docs: the Write Contention section still described attempt-counted retries
(_WRITE_MAX_RETRIES = 15); updated to the time budgets in force and the new log line.
2026-09-20 10:01:42 -07:00
kshitijk4poor
838214c8f7 fix(error-classifier): a profile hook asking for fallback on a terminal reason is non-retryable
turn_api_error enters the fallback walk only for ``retryable=False`` verdicts outside the
retryable-client reasons; the built-in terminal verdicts pin retryable=False while the
rate-limit family stays retryable and cascades after backoff. A ``classify_api_error`` hook
returning ``{"reason": "billing", "should_fallback": True}`` therefore retried the dead
route instead of cascading (the #116408 test passed only because its hook also set
retryable=False). Default retryable to False for such verdicts, leaving rate-limit reasons on
the built-in retry-then-fallback shape; ``RETRYABLE_CLIENT_REASONS`` moves next to the
verdict table so both modules read one set. Contract documented in the plugin guide.
2026-09-20 12:17:52 +05:30
teknium1
bce20d0b1f feat: model-provider plugins classify their own API errors via ProviderProfile.classify_api_error
A kind: model-provider plugin is loaded by providers/ discovery and never enters the
PluginManager hook lifecycle, so transform_api_error_classification was unreachable for it
without shipping a second plugin component. The profile now carries an optional
classify_api_error(error, *, status_code, error_code, message, body, model) callable,
consulted as a classifier stage right after the generic plugin hooks and only for the
provider that produced the error. None or an unknown reason leaves the built-in verdict;
built-in providers are untouched (no name table, no lifecycle change).

Also: a plugin refresh_credential returning None/empty was treated as a successful refresh
(row marked ok, stale bearer replayed up to the refresh cap). It now benches the row like a
failed refresh POST, so the loop rotates or falls to the generic sign-in copy.

Part of #116408

(cherry picked from commit b8129fd6fd6a0cf4eeee6d95a5d5d823668a306d)
2026-09-19 21:24:05 -07:00
teknium1
3e67877e1b docs: declarative OAuth (PKCE) subsection in the model-provider plugin guide
Minimal OAuthPKCEConfig example plus what Hermes owns and the enforced
security boundary, so plugin authors do not hand-roll a browser flow.

(cherry picked from commit fc4e73d5d2067102ab08823aafbb4f31b3cba785)
2026-09-19 21:15:16 -07:00
teknium1
f7240ca980 fix: admitted plugin providers are selectable and status-accurate on every surface
Picker admission (#116552) listed out-of-tree external-process and OAuth
plugin providers, but selection and status still dispatched through
provider-name tables:

- `hermes model`: no `_PROVIDER_MODEL_FLOWS` entry and no generic flow, so
  picking an admitted plugin row was a silent no-op. One generic flow in
  model_setup_flows.py, credential step keyed by the profile's auth_type
  (external_process -> launch check; oauth_* -> live pool row, else the
  `hermes auth add <name>` hint), catalog via merge_profile_catalog; main.py
  falls back to it for any registered profile missing from the table.
- `_STATUS_BY_AUTH_TYPE` had no builder for oauth_device_code/oauth_external,
  so `get_auth_status`/`list_available_providers().authenticated` stayed False
  with a live pool entry. `get_plugin_oauth_auth_status` (auth_plugin_providers
  sibling) reads the pool; gated on PLUGIN_MIRRORED_PROVIDERS so bundled OAuth
  providers keep their bespoke status bytes.
- `_external_process_auth_evidence` computed evidence for copilot-acp only, so
  `inventory._external_process_signed_in` hid every other ACP row from the
  Desktop explicit_only picker. Generic evidence = the binary resolves; the
  bundled CLI keeps its token-store chain.
- agent_init Responses-upgrade guard dropped the vendor literal redundant
  with the acp:// scheme check.
- `fetch_account_usage` bounds the plugin hook with a shared 10 s deadline
  (previously only the CLI wrapped the call; gateway/TUI awaited unbounded),
  contextvars-propagated so scoped secrets resolve; overrun -> None.

Part of #116408

(cherry picked from commit 536a4e7fcf2d1ac2b8a70bd62a69707d58ea5d1e)
2026-09-19 21:11:30 -07:00
teknium1
8b42b6e020 docs: restore the model_capabilities field row dropped in the rebase 2026-09-19 21:08:09 -07:00
Ayush Nangia
ef0bac6cdd feat(providers): plugin-declared per-model capability metadata
Model provider plugins can now declare per-model metadata (capability
booleans, context_window, model_family) through the existing
ProviderProfile registry using the canonical model_overrides schema.
Declarations patch catalog metadata without erasing unknown fields; on a
catalog miss, undeclared capability booleans stay unknown rather than
false.

Precedence (tested): explicit user model_overrides > plugin declarations
> catalog/builtin > _default fill-gap. lookup_models_dev_context shares
the same seam so auto-detected context reflects declarations, including
the dashboard /api/model/info response.

Provider aliases resolve declarations; user override lookup keeps its
existing provider-key rules. Metadata lookup may trigger the existing
lazy provider discovery (imports plugin code); the registry is
process-global and discovered once per process.

Progress on #102115.

(cherry picked from commit a2bf725597efddf5238a8d0447705cf31ac71cc8)
2026-09-19 21:08:09 -07:00
teknium1
3a045623bc docs: resolve field-table rebase conflict (auth hooks rows + catalog/vision rows) 2026-09-19 20:56:31 -07:00
teknium1
08feaa98b5 fix(image-routing): drop the provider-wide supports_vision probe; mirror registry via public types in setup test
ProviderProfile.supports_vision is documented as "the API accepts image content
inside tool-result messages" -- a provider-wide wire capability, not a per-model
user-image verdict. Treating it as the latter flipped every model on the bundled
`router` relay (and models.dev-unknown meta-ai/xiaomi models) from text to native
image parts. Per-model vision now comes from ProviderProfile.model_capabilities
(#116570) through the existing models.dev probe.

The setup-catalog test mirrored the plugin through auth._register_plugin_provider,
which #116553 renames; build the ProviderConfig from public types instead so the
test passes on both trees.
2026-09-19 20:56:31 -07:00
teknium1
9de45d9595 fix(vision): image routing honours ProviderProfile.supports_vision
`decide_image_input_mode` consulted config overrides, the managed runtime,
models.dev and Ollama, but never the registered profile's `supports_vision`
— the field the tool-result media path already trusts. A plugin that
declared vision was therefore native for tool results and text-only for
user-attached images. The declaration is now the last probe in
`_VISION_PROBES`: only an explicit True is a verdict, and per-model catalog
entries still win.

Part of #116408
2026-09-19 20:56:31 -07:00
teknium1
6f12165a94 fix(models): admit every plugin provider to the picker by slug; catalogs fall back to the profile
The CANONICAL_PROVIDERS auto-extend skipped plugin profiles by auth_type
("Non-api-key flows need bespoke picker UX"). Every in-tree non-api-key
profile already owns a hand-written row, so the set never excluded an
in-tree provider - it only hid out-of-tree external-process and OAuth
plugins from /model, hermes model and list_available_providers().
Admission is now by slug (dedupe against the built-in rows); visibility
stays gated downstream by credentials (binary resolves / auth.json /
pool entry), so an admitted row reads authenticated=False until sign-in.

- _profile_live_catalog: external_process profiles use fetch_models(),
  then fallback_models; every other non-api-key profile returns its
  fallback_models instead of None (in-tree ones declare none, so the
  built-in rows are byte-identical).
- _credential_fingerprint: external-process command/argv env overrides
  key the catalog cache, as an API key does for HTTP providers.
- Drop the "[slug] as the model" shim from the canonical picker lap;
  the profile's fallback_models is the honest catalog.
- Tests trimmed to two invariants proven red on base; docs describe how
  picker rows and the authenticated flag are derived.

Co-authored-by: ericmaddox <eric.maddox@outlook.com>
Co-authored-by: Denis Araujo <d.araujo@tjrs.jus.br>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
Co-authored-by: tobenwarrior <lucasyee999@gmail.com>
2026-09-19 20:54:06 -07:00
teknium1
627a12a90e feat(opencode-go): show Go plan windows in /usage through the profile hook
First in-tree consumer of ProviderProfile.fetch_account_usage: the OpenCode Go plan's rolling /
weekly / monthly windows (GET /zen/go/v1/usage) render in every /usage surface without adding the
provider name to the core _USAGE_FETCHERS table. Ported from #113418, which implemented the same
fetch as a core table entry; the literal endpoint (not the runtime base_url, which loses /v1 in
anthropic_messages mode) and the window mapping are theirs.

Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
2026-09-19 20:52:57 -07:00
KoNit-K
400b6c41cd feat(providers): expose plugin account usage hook 2026-09-19 20:52:57 -07:00
teknium1
6621e8aa98 docs: state the two registry-mirror exclusions, the interactive-add namespace shape, and the aux 401 boundary 2026-09-19 20:45:06 -07:00
teknium1
ec1238fa65 fix(credential-pool): plugin refresh keeps rotated tokens, runs locked, and quarantines dead grants
Independent review of the plugin refresh branch (#116553) found four gaps
between what model-provider-plugin.md promises for `refresh_credential`
and what `_refresh_entry_impl` did:

1. `replace(entry, **hook_result)` raised TypeError on any non-field key
   (`expires_in`, `token_type`, `scope` — the natural token-endpoint shape),
   the except benched the row EXHAUSTED and the pair the server had already
   rotated was dropped: for single-use refresh tokens that is a lost login.
   Field keys now go through `replace()`, everything else merges into
   `entry.extra` (mirrors `from_dict`); `None` = no rotation, mark ok.
2. Plugin providers skipped the locked single-use path, so a gateway and a
   CLI could both POST the same refresh token (`refresh_token_reused`).
   Providers with a hook now take the `_auth_store_lock` path: re-read the
   pool store, adopt a peer's usable rotation and skip the hook, else call
   it and write through. Eligibility derives from `plugin_refresh_hook()`,
   not from extending the built-in name tuple.
3. A raising hook re-benched EXHAUSTED every cooldown forever at DEBUG.
   `AuthError(relogin_required=True)` (or a grant-dead OAuth code) is now
   terminal: the row goes DEAD with a WARNING naming `hermes auth add`.
   Any other exception stays a transient bench (negative test kept).
4. `from_dict`'s extra sweep round-tripped a stray row-level `provider` key
   back onto the row on `to_dict()`; it is bookkeeping, not metadata.

New logic lives in `agent/credential_pool_plugin.py` — credential_pool.py
is at the size cap; the facade only dispatches.

Part of #116408
2026-09-19 20:45:06 -07:00
teknium1
8e5a9b16a6 docs+chore: escape bare angle-bracket placeholders in MDX table; map zihaofeng2001 email (#101768 salvage) 2026-09-19 20:45:06 -07:00
teknium1
dce233e809 feat(auth): OAuth-shaped provider plugins register, log in and refresh through their profile
An out-of-tree ProviderProfile with auth_type oauth_device_code/oauth_external loaded and inferred
but was invisible to hermes auth: _register_plugin_provider skipped every auth_type except
external_process and api_key, so resolve_provider() said "Unknown provider", `hermes auth add`
had nothing to dispatch to, the pool could not refresh its rows (REFRESHABLE_OAUTH_PROVIDERS is a
name set) and PooledCredential.from_dict dropped every extra key outside _EXTRA_KEYS on reload.

- hermes_cli/auth_plugin_providers.py (new sibling; auth.py is at the size cap): the registry
  mirror now registers every profile under the auth_type it declares, re-syncs after discovery /
  on a miss (salvaged from #101768), and owns the seam lookups: auth_handler dispatch, the
  fail-loud error for a non-api-key profile that ships no handler, and refresh eligibility
  derived from the profile's refresh_credential hook (never a name set).
- providers/base.py: ProviderProfile.auth_handler(action, args) (salvaged from #111610, sync only)
  and refresh_credential(entry) -> rotated fields. A separate hook rather than
  auth_handler("refresh", ...) because the pool holds a credential row, not an argparse namespace,
  and needs tokens back rather than a bool.
- hermes_cli/auth_commands.py: add/status/logout/refresh (incl. interactive add) consult the
  plugin handler before the built-in path; `auth refresh` admits plugin rows via the predicate.
- agent/credential_pool.py: _refresh_entry_impl calls the profile hook; from_dict keeps every
  non-field key in extra so plugin metadata survives load -> save -> load (to_dict already wrote
  it all; sanitize_borrowed_credential_payload semantics unchanged).
- hermes_cli/auth.py: config import moved below PROVIDER_REGISTRY (salvaged from #94231) so a
  plugin imported during discovery never sees a partial auth module.

Built-in providers are untouched: only custom/openrouter were unregistered profiles before and
both stay in the skip list; anthropic/nous/openai-codex auth add/status/refresh output is
byte-identical in the before/after probe.

Part of #116408. Salvages #111610 (@Finn763), #101768 (@zihaofeng2001, absorbing #106361 by
@Finn763) and #94231 (@Kyzcreig).

Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: zihaofeng2001 <zihaofeng2001@gmail.com>
Co-authored-by: Kyzcreig <9063726+Kyzcreig@users.noreply.github.com>
2026-09-19 20:45:06 -07:00
finn763
749443e292 feat(auth): let model-provider plugins own hermes auth add/status/logout/refresh
A `kind: model-provider` plugin can register a ProviderProfile and a custom
client but could not contribute an interactive login flow: its module is
imported by provider discovery, and the generic command-plugin loader skips
model-provider manifests on purpose, so `register(ctx)` is not a supported way
to add commands. The core login implementations are also provider-name keyed, so
a standards-based device-code provider had to ship a second standalone command
plugin just for login/status/logout.

Add an optional ProviderProfile.auth_handler(action, args) seam and consult it
first from the existing `hermes auth` actions (add/status/logout/refresh) —
before the known-provider gate and before the credential pool:

- truthy return = the provider owned the action; falsy = per-action fallback.
- sync or async handlers (a returned awaitable is awaited).
- a handler exception becomes a readable SystemExit naming provider + action.
- no handler at all = byte-for-byte unchanged built-in behavior, including the
  existing "Unknown provider" exit.
- a registered profile that does not handle the requested action now says so
  instead of reporting its provider as unknown.

No new top-level command, no manifest surface; credentials stay provider-owned.

Tests: tests/hermes_cli/test_provider_auth_seam.py drives the real `hermes auth`
parser with a fixture model-provider plugin — dispatch + argument pass-through
for all four actions, per-action fallback, async handler, handler failure,
registry lookup failure, duplicate registration (last-writer-wins), and an
untouched built-in provider. Docs gain the new hook under the model-provider
plugin guide.
2026-09-19 20:45:06 -07:00
Teknium
b6f8f8eb1f Merge pull request #116345 from NousResearch/boa-w3-small-b
Video generation tools no longer let the agent pick the model; video_gen.model is the only selector (Refs #83080)
2026-09-19 14:32:37 -07:00
teknium1
19b29df13b fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.

The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.

Refs #83080
2026-09-19 12:22:44 -07:00
teknium1
03cfc36f51 docs: describe the session.create model×provider refusal
Hosts speaking the gateway protocol need to know the new -32602 shape
(error.data model/provider/suggestions) and which pairs stay permissive.
2026-09-19 12:12:05 -07:00
Teknium
4fda3bbdca Merge pull request #115846 from NousResearch/fix/boa-tts-stt-voice-tts-config-minlen
feat(tts): streaming TTS speaks a short first sentence sooner via tts.streaming.min_len on every surface (#96927, salvage #96933)
2026-09-19 11:33:15 -07:00
Teknium
0d2cb2cacc Merge pull request #115890 from NousResearch/fix/boa-res-R6-streaming-compaction-astra-oauth-gate
fix(compression): gpt-6-astra on Codex OAuth gets native server-side compaction when opted in (#103720, salvage #103718)
2026-09-19 11:28:10 -07:00
teknium1
a169438178 Merge remote-tracking branch 'origin/main' into HEAD
# Conflicts:
#	hermes_cli/config_defaults.py
2026-09-19 11:22:01 -07:00
Teknium
7feb1af028 Merge pull request #115872 from NousResearch/fix/boa-res-R4-routing-catalog-azure-codex-900k-cap
fix(codex): opted-in -900k aliases cap at the live catalog max_context_window (#105443, salvage #105445)
2026-09-19 11:12:31 -07:00
teknium1
ec2c4eeb80 chore: merge origin/main (resolve apps/desktop/src/lib/voice-client-direct.ts, hermes_cli/config_defaults.py, tools/voice_client_config.py, website/docs/user-guide/features/tts.md) 2026-09-19 10:54:38 -07:00
teknium1
e85cb94da1 chore: merge origin/main (resolve agent/error_classifier.py,tests/agent/test_error_classifier.py) 2026-09-19 10:51:40 -07:00
teknium1
8e3a833365 chore: merge origin/main (resolve website/docs/developer-guide/context-compression-and-caching.md) 2026-09-19 10:47:48 -07:00
teknium1
a516db44e7 Merge remote-tracking branch 'origin/main' into HEAD 2026-09-19 10:46:59 -07:00
teknium1
6c3ff1d732 docs(site): docs and generated skill pages stop suggesting /tmp
Hand-written docs and the generated per-skill mirror pages now show the same
scratch locations the skills and prompts do (~/.hermes/cache/scratch,
$TMPDIR, $HOME/.hermes/cache/scratch/<throwaway-home>) instead of /tmp, and
examples that only needed a placeholder use /path/to/... The mirror pages were
updated in place rather than regenerated: regenerating from the current sources
produces a 200-file unrelated diff (Windows backslash paths, removed skills).

Literals that describe /tmp itself stay and carry a no-tmp marker: the
sandbox tmpfs configuration, the disk-cleanup plugin's scope, the WSL feature
list, the terminal.temp_dir rationale, the Nix container's writable layer and
the Docker Compose in-container pulse-cookie path. One tree-listing line in
nix-setup.md stays unmarked (a marker would render inside the code block).
2026-09-19 10:44:26 -07:00
teknium1
4c4b0b748e docs: describe the summary-provider-overload abort in the compression failure ladder 2026-09-19 10:35:44 -07:00