Commit Graph

5059 Commits

Author SHA1 Message Date
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
12fb5eb476 fix(sessions): make set-journal-mode portable and fail closed where it cannot prove quiescence
Review follow-ups on the new `hermes sessions set-journal-mode` verb:

- The header probe used os.pread, which does not exist on Windows, while the subparser is
  registered unconditionally — the command died there with an uncaught AttributeError. It now
  reads the 20 header bytes through a plain binary open(), and the tests no longer skip on win32.
- foreign_state_db_holders() returns [] unconditionally on Windows (no scan), which made the
  admission gate vacuous: an operator got a silent all-clear and could flip the mode under a
  running gateway. Windows now refuses outright, naming the reason, overridable only by --force.
- Enabling WAL ignored the cross-VM filesystem refusal the runtime enforces
  (apply_wal_with_fallback). target=wal now refuses on virtiofs/9p, where WAL shared memory
  silently corrupts.
- A --db pointing at a garbage file surfaced a raw sqlite3.DatabaseError traceback even though the
  header probe had already read not-a-database, and a directory raised IsADirectoryError. Both now
  bail in the command's own error style; every open/read is guarded.

The admission checks that need no I/O live in a pure _refusal() that takes the platform as data,
so the Windows and cross-VM invariants are tested without faking sys.platform.
2026-09-20 20:13:25 -07:00
teknium1
96da5d97fc feat(sessions): hermes sessions set-journal-mode delete|wal converts an existing WAL store offline (#100896)
`database.journal_mode: delete` can never self-apply to a store that is already WAL:
apply_wal_with_fallback deliberately never live-downgrades (#68545 — other gateway/cron/worker
connections may hold uncheckpointed WAL commits), so operators applying the containment for the
multi-writer corruption class saw one ERROR per process forever and the only escape hatch was an
undocumented hand-run PRAGMA on the file.

The new pre-DB `sessions set-journal-mode` verb is the sanctioned offline path: it refuses while ANY
foreign process holds the file or a sidecar (the same foreign_state_db_holders scan doctor/repair
admission uses, naming each PID), flips through _set_journal_mode_no_wait (busy_timeout=0, so an
opener appearing mid-way makes SQLite refuse instead of racing it), verifies header bytes 18/19,
and reminds the operator when config.yaml disagrees. `--db PATH` covers kanban.db / cron stores
that log the same ERROR. The never-live-downgrade invariant is untouched; the ERROR, doctor hints
and docs now name the command instead of the raw PRAGMA.
2026-09-20 20:13:25 -07:00
Austin Pickett
1a1f4a59e2 fix(auth): explicit-provider gate uses the credential resolver's reader
get_env_value stops at the first environ hit, so a shell that exports
DEEPSEEK_API_KEY= (empty) hides a real key in .env from the gate while
resolve_api_key_provider_credentials() finds it: the picker omits a provider
the chat path would authenticate with. Read through
get_env_value_prefer_dotenv, the same chain auth.py already uses to resolve
the key, so the two can never disagree (#77007).

Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-20 22:46:49 -04:00
xxxigm
4ea57fa61b fix(auth): count profile .env keys in the explicit provider gate
os.getenv saw only the launch profile, so a DeepSeek key pasted in
another profile never appeared in Settings → Model until Refresh
models ran against that Bot's own backend.
2026-09-20 22:46:49 -04:00
Bartok9
6504b665ba fix(kanban): refuse empty complete_task evidence (#117483)
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
2026-09-20 19:03:47 -07:00
teknium1
c545272568 fix(lsp): one stalled request no longer silences a workspace for good — retry window, cold-root warm-up budget, per-root exclusion
A language server that missed its budget once marked its (server, root) pair broken for the
process lifetime, the same 5 s steady-state budget was applied to a cold server that also had
to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched
the server off for every workspace.  Three new keys under the existing `lsp` block, all defaulting
to today's behaviour:

- lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per
  pair; an expired pair gets one more try, and the INFO skip line names the retry time.
- lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client
  waits up to this budget (outer join budget follows); warm requests keep wait_timeout.
- lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path
  also covers everything beneath it); a matching root never spawns, logged once at INFO.  A
  non-list value fails closed — WARNING naming the expected shape, every root skipped — because
  silently excluding nothing would re-pay the stall the key was meant to avoid.

Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo).
2026-09-20 18:54:24 -07:00
teknium1
29c4ec8fa7 fix(update): settle a stuck failed receipt once the live fleet serves the checkout code
A failed update receipt whose plan rows cannot be matched to a live gateway
(unknown profile identity, pre-pull SHAs, empty fleet matrix) could never be
discharged: every up-to-date `hermes update` ran the pending restart, printed
"Pending fleet restart completed." followed by "Fleet restart incomplete",
exited 1 and wrote another failed receipt — and every CLI start kept warning
about mixed sys.modules, even after a manual `hermes gateway restart` had put
the fleet on the checkout code (#117051, atoms 3/4).

In _apply_pending_fleet_restart_catchup, when the restart succeeded but the
receipt still owes gateways, probe the live fleet; if every row is current at
the checkout SHA, persist that matrix as latest.json's post-restart `fleet`
(update_receipt.settle_latest_receipt_fleet) so _receipt_reports_stale_runtime
and the startup warning see the recovery. The "incomplete" line now only
prints when the restart itself failed; the residual case names what is still
off the checkout code instead of contradicting the completion line.

Fixes #117051
2026-09-20 18:52:45 -07:00
teknium1
d6b0d37ece fix(update): pending fleet restart leaves gateways already on the checkout code alone
The catch-up restart stopped every gateway before re-checking whether any was actually on
pre-update code, so the gateway the same `hermes update` had just cold-started was killed
and (on Windows) the stop/start pair printed "No gateway was running" followed by a second
spawn. _run_pending_fleet_restart now skips when the fleet probe shows every live gateway
current on the checkout SHA (identity known; any stale/down/unknown row still restarts).
The Windows direct-spawn report no longer prints the PID twice ("(PID n) (PID: n)").

Part of #117051
2026-09-20 18:52:45 -07:00
teknium1
eaef5ec717 feat(update): hermes update --list-venv-holders prints the venv guard's holders as JSON, exit 3
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].

Fixes #117246
2026-09-20 18:51:36 -07:00
teknium1
2a7658989f test: trim the roster-poll memo tests to the invariants
Keep staleness (edit seen on the next poll), missing-file pickup, parsed-once and the uncached-writer contracts; the per-file mutation-poison and cross-profile cases restate the copier.
2026-09-20 18:45:07 -07:00
John Paul Soliva
55a8f42329 perf(profiles): the listing re-reads a profile's YAML only when that file has changed
`list_profiles()` parsed three YAML files per profile on every call — config.yaml through
`read_user_config_raw` ("no cache", by contract) plus profile.yaml and distribution.yaml through
`_load_yaml_dict`. It is the shared body of `GET /api/profiles` and `profiles.list`, which the Bots
roster polls every 5s per connection, so an idle machine re-parsed every profile's YAML twelve
times a minute to produce the same handful of strings.

config.yaml is the expensive one: the installer seeds it by copying the annotated template verbatim
(`cp cli-config.yaml.example config.yaml`), 119,268 bytes / 2,267 lines of mostly comments, for
`model.default` and `model.provider`.

Measured, 9 profiles, no gateway locks so the liveness path stays out of it: 16.1ms -> 1.6ms with
the installer template, 6.4ms -> 1.6ms hand-trimmed; 27 YAML parses per pass -> 2, and those two
are the default profile's ABSENT files, correctly not cached. The pass is now independent of config
size.

The memo keys on each file's (mtime_ns, size, inode) — the atomic writers rename a temp file into
place, so a rewrite always lands a new inode even inside one mtime tick — and caches only the small
DERIVED values. `read_user_config_raw` and `_load_yaml_dict` keep their uncached contract: callers
write those documents back, and a stale read there could overwrite a newer file. `read_profile_meta`
hands out a copy so a future mutating caller cannot poison the next reader.

Regressions drive the real readers: an edited config/profile/distribution file is seen on the very
next read, `write_profile_meta` round-trips, a missing file is not cached and is picked up when
created, mutating a returned meta does not poison the next reader, an unchanged file is parsed once
across three reads, and `read_user_config_raw` still re-reads. Ignoring the signature fails four.

Fixes #117378
2026-09-20 18:45:07 -07:00
fangliquanflq
c906121a46 test(kanban): cover uncapped runtime output 2026-09-20 18:43:59 -07:00
fangliquanflq
1a4c68ce97 fix(kanban): expose task runtime limit in JSON 2026-09-20 18:43:59 -07:00
chadhouser
7ff57f0cb2 test(kanban): pin the review handoff of a card assigned to its own reviewer
The suite already covers the shape where assignee != reviewer
(test_review_handoff_without_live_run_attributes_run_to_implementer). The
mirror case -- a card created already assigned to its reviewer -- was the gap.

Asserts the honest provenance (no implementer, NULL run profile) and, more to
the point, that request_changes refuses instead of routing the rejection back
to the reviewer. Fails on the parent commit with implementer='reviewer-a'.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-20 18:43:35 -07:00
rodricksz4h5
f50c6f43a2 fix(cli): let the early interface probe follow the profile's home
`_suppress_mouse_residue_early()` runs at import time, before
`_apply_profile_override()` sets HERMES_HOME, so the first read of
`display.interface` always saw the default home. The result was memoised
in a cache keyed on nothing, so every later caller — the Termux fast
paths and the TUI launch decision, which all run after the profile is
applied — was pinned to the default home's interface for the whole run.
`hermes -p coder` with `display.interface: tui` in that profile booted
the classic REPL, and the mirror case booted the TUI for a profile that
asked for cli.

Key the cache on the config path it read. The hot path still parses the
YAML once per home, and the first call after the process is re-homed
re-reads the file it should have read.

Fixes #116902
2026-09-20 18:26:14 -07:00
rodricksz4h5
b3ae818261 fix(cli): attribute a scrubbed-env dashboard to its owner's home
`_hermes_home_for_pid` fell back to the INSPECTING process's `Path.home()`
when the target's environment carried neither HERMES_HOME nor HOME, which
is the normal shape for a systemd/launchd unit with a scrubbed
environment. The dashboard was then attributed to whichever user ran the
command, so `hermes update` and `--stop` read the wrong profile root and
could act on the wrong backend.

A process without HOME resolves its own default home from the password
database, so resolve the target's owner the same way (psutil, then the
/proc owner) before falling back. An unreadable owner or entry keeps the
previous behaviour rather than resolving to nothing.

Fixes #116906
2026-09-20 18:25:48 -07:00
fangliquan
7df20a670d fix(cli): fail closed without bang shell sanitizer 2026-09-20 18:20:44 -07:00
fangliquanflq
3f399c0bd4 fix(cli): align kanban edit and paused cron listing 2026-09-20 16:33:49 -07:00
teknium1
891c1d79c6 fix(constants): one home for managed, container and HERMES_UID policy; drop the config twins
Review follow-up: get_managed_system, the container/chmod-skip check and
the HERMES_UID/GID chown (_resolve_hermes_uid_gid/_chown_to_hermes_uid) now
live only in hermes_constants — the import-safe module apply_secure_dir_policy
already needed them in — and hermes_cli.config re-exports them, so there are
no 'keep in sync' copies and _secure_file skips on the same canonical
_detect_container signal as _secure_dir/get_scratch_dir. The dead config
copies are deleted and test_ensure_hermes_home_uid.py drives the new
symbols; one invariant test pins the single implementation.
2026-09-20 16:32:50 -07:00
teknium1
265e68d784 fix(auth): one primary_failure_wording() helper labels quota vs auth failure at all three fallback surfaces
- hermes_cli.auth.primary_failure_wording(exc) -> (log, user) phrase; reused by
  cli_agent_setup_mixin._resolve_fallback_runtime, runtime_provider's fallback
  logger and the TUI gateway/Desktop _resolve_runtime_with_fallback (#117482
  sibling: 'Primary auth failed' for a 429 on the gateway surface).
- Drop the dead credentials_rate_limited kwarg at the post-turn exit site (the
  flag is only True when _ensure_runtime_credentials returned False).
- Fold the new tests: 2 parametrized production-path tests + 1 gateway test.
2026-09-20 16:23:22 -07:00
686f6c61
89096ba56c fix(cli): treat credential-resolution 429 as quota, not auth
A Codex quota wall at startup was labeled "Primary auth failed" and
exited 1, so Kanban counted a failure and operators went looking for
a bad token. Use is_rate_limited_auth_error() for the fallback notice
and EX_TEMPFAIL (75) when HERMES_KANBAN_TASK is set.
2026-09-20 16:23:22 -07:00
teknium1
214623389a fix(models_validate): one profile-catalog helper decides accept/reject before the generic listing
Fold _profile_owns_catalog and _profile_owned_catalog into _profile_catalog ->
(catalog, authoritative) so _validate_live_listing has a single owned-catalog
branch; drop the redundant catalog-hit test (covered by the owns-catalog accept
tests in test_setup_provider_catalog.py).
2026-09-20 16:22:25 -07:00
Konstantin Khlopkov
57a3bbbcc1 fix(models_validate): a profile's own catalog endpoint decides, not the generic /models listing 2026-09-20 16:22:25 -07:00
teknium1
9aab6ba7d7 test(doctor): strip trailing blank line 2026-09-20 15:57:25 -07:00
joaomarcos
d76c77beed fix(doctor): stop prescribing audit fixes that updates revert (#116774)
The doctor "Browser tools (agent-browser)" row audits the root workspace
tree (npm audit --workspaces=false), whose versions come from the committed
package-lock.json. hermes update reinstalls that exact state via npm ci
(_run_npm_install_deterministic never mutates the lockfile), so the
previously prescribed local `npm audit fix --workspaces=false` was wiped on
the next update and the finding returned - a fix/reintroduce loop (#116774).

- Bump the vulnerable root lockfile pins so the shipped tree audits clean:
  browserslist 4.28.6 -> 4.29.0 (high: GHSA-c83g-rgw3-j3cx,
  GHSA-73wf-gq98-2v4g), baseline-browser-mapping 2.10.43 -> 2.11.25
  (moderate: GHSA-w5vr-8v7q-w6rv), plus updated transitive deps
  (caniuse-lite, electron-to-chromium, node-releases,
  update-browserslist-db); also syncs the stale bootstrap-installer
  lockfile entry to its manifest version. npm audit --workspaces=false now
  reports 0 vulnerabilities.

- Rework the npm-audit remedy for every doctor row: the durable fix is an
  upstream lockfile bump; never prescribe a local mutating fix command
  (workspace-scoped rows already omitted it; the root row prescribed the
  doomed one).

- Invariant tests in tests/hermes_cli/test_doctor_audit_remedy.py: the
  remedy must not prescribe a mutating npm audit fix (proven red on base)
  and must name the lockfile-bump remedy.
2026-09-20 15:57:25 -07:00
tobenwarrior
dff92f3d78 fix(plugins): peel an annotated-tag pin before verifying the checkout
A pin is validated as 40 hex, but that does not make it a commit: a tag
object has a sha of its own, and an entry recorded from 'git rev-parse <tag>'
names the tag object rather than the commit it points at. Git resolves the
ref and detaches at the commit, so the revision guard compared a commit
against a tag-object sha and refused a CORRECT checkout — the entry was
uninstallable, with no way through (see the kiro-acp catalog entry).

Peel the pin to its commit before comparing. The guard keeps its teeth: HEAD
must still be exactly the commit the pin names, and a checkout that lands
anywhere else is still rejected. An unresolvable pin falls back to the raw
string, so a genuinely wrong sha fails exactly as before.

Tests: an annotated-tag pin installs at its commit (and records the commit,
not the tag object, so 'plugins update' drift checks stay exact); a
checkout that lands on another commit is still rejected. Red on base:
'Checked-out revision "..." does not match requested commit "..."'.

(cherry picked from commit 66baa043b70c4f66ae59c0a1147e74ba694140ef)
2026-09-20 15:56:55 -07:00
teknium1
afd2fdf15c fix(tests): exit clear-screen fallback tests use the host OS, no os fake
POSIX case runs natively (skipped on nt); the cmd /c cls + CREATE_NO_WINDOW
assertion moves to a windows_only test that checks the real windows_hide_flags().
2026-09-20 15:56:33 -07:00
Aaron08140
8e4c943477 fix(cli): clear-screen fallback spawns no shell and no console window (#116904)
_clear_terminal_on_exit() fell back to os.system('cls'/'clear') when the
escape sequence was rejected. os.system() spawns a shell: on Windows a
conhost window flashes before the process exits, and in minimal POSIX
containers without `clear` on PATH the shell fails silently while the
except-pass hides it. The repo already standardises on windows_hide_flags()
(CREATE_NO_WINDOW) for short-lived argv spawns; use it here too.

- Windows: subprocess.run(['cmd', '/c', 'cls'], creationflags=windows_hide_flags())
  (cls is a cmd builtin, so probing a `cls` exe would be wrong).
- POSIX: shutil.which('clear') first; skip the spawn entirely when absent
  instead of letting a shell swallow a 127.
- stdin=DEVNULL so a child that reads can never block CLI shutdown.
- Regression tests: parametrised nt/posix argv + creationflags assertions
  (RED on the os.system path), plus the no-`clear` skip case.
2026-09-20 15:56:33 -07:00
teknium1
31017b7344 fix(kanban): a caller asserting run ownership cannot classify a parked card
Review follow-up for #117363: the classify-in-place branch returned before
the expected_run_id guard, so a worker whose run had ended could type a
breaker-parked card. Refuse when expected_run_id is given; the run is over.
2026-09-20 15:55:14 -07:00
teknium1
9850010ff5 fix(kanban): classify only an untyped, run-free breaker block; restore sort order
Tighten the salvaged classify-in-place path to the reporter's contract:
only block_kind IS NULL and current_run_id IS NULL qualify, recurrence
accounting starts at 1, the audit event is the ordinary `blocked` kind so
existing consumers see the reason, and a typed block is never re-typed.
Drops the unrelated VALID_SORT_ORDERS deletion and trims tests to two
invariants (one drives the real enforce_max_runtime breaker path).
2026-09-20 15:55:14 -07:00
allo
6c7e8f5069 test(kanban): add tests for block_task classification of blocked cards
Regression tests for #117363: verify that block_task() can set
block_kind on an already-blocked card, that re-classification works,
and that idempotent calls are no-ops.
2026-09-20 15:55:14 -07:00
teknium1
87f892b8db fix(kanban): cover authorising/authorises in the respawn auth pattern and trim the table
Review follow-up for #117009: the British -ise family now mirrors the -ize
family (e|es|ed|ing|ation); the curated-vocabulary table keeps one
prose-negative per false-positive stem and one auth-positive per family.
2026-09-20 15:53:55 -07:00
Jony
3899401aa6 fix(kanban): ignore auth words in crashed worker output 2026-09-20 15:53:55 -07:00
chelsealong
857709d63b fix(kanban): cover -es/-ing/British spellings in the auth blocker pattern
The curated auth alternation from #117009 dropped authenticates/
authenticating, authorizes/authorizing, and the British authorise/
authorised/authorisation spellings, so those genuine auth-failure
strings would miss the respawn guard and retry immediately instead of
auto-blocking. Extend the alternation to cover them, per review.
2026-09-20 15:53:55 -07:00
chelsealong
6a52c9e41d fix(kanban): curate the respawn guard's auth pattern instead of an open auth\w* stem
_RESPAWN_BLOCKER_RE used auth\w* to detect quota/auth failures in
last_failure_error, which also matches ordinary English words like
"author", "authored", "authoring", "authoritative" and "authors" in
worker progress prose. Because the guard is re-derived from an
unchanged string on every dispatch tick with no terminal state, a
single such word permanently parks an otherwise-healthy ready card
(observed: 437 respawn_guarded events over 7.3h on one card, #117009).

Replace the open stem with a curated set of real auth-failure tokens
(auth, authenticate/authenticated/authentication,
authorize/authorized/authorization, authz) so the false-positive
words stop matching while genuine auth failures still do.
2026-09-20 15:53:55 -07:00
teknium1
56b7a0fb37 fix(local-runtime): an unwritable boot lock degrades to an unlocked boot, never an OSError
Review follow-up: the boot lock's mkdir/os.open ran in the context manager's
__enter__, outside the body's try/except, so a read-only runtimes dir or a
foreign-owned boot.lock raised OSError out of ensure_local_runtime — the
function main documented as never raising into session start. Catch OSError
around acquisition, warn, and yield unlocked (the contention-timeout shape).
2026-09-20 15:53:25 -07:00
chelsealong
4352a7d053 fix(local-runtime): serialize managed runtime boot across processes
Two Hermes backends starting within the same second (e.g. default +
non-default profile) each see no server.json yet and each spawn a
llama-server router on the stable port — the existing per-process
_SUPERVISOR singleton can't rule this out since every process starts
with its own copy. Wrap the state-check/spawn sequence in an OS-held
file lock (fcntl/msvcrt, mirroring managed_uv.py's install lock) so the
second caller waits for the first to publish its state and adopts it
instead of racing to spawn its own.

Fixes #116682
2026-09-20 15:53:25 -07:00
fangliquan
ae908a93ba fix(local-runtime): support MXFP4 GGUF tensors 2026-09-20 15:29:34 -07:00
funky-xamarin
23d3edac4e fix(runtime): resolve missing credential pool endpoints 2026-09-20 15:27:46 -07:00
teknium1
40ac07e42e test(update): drive collect_fleet_versions for the separate-checkout classification
The salvaged suite proved the helpers; this test drives the production entry
point so deleting the code_root wiring in collect_fleet_versions goes red.
2026-09-20 15:23:04 -07:00
fangliquanflq
561658d26d fix(update): exempt gateways from separate checkouts 2026-09-20 15:23:04 -07:00
teknium1
656c1c9bda fix(plugins): isolate SystemExit in event, middleware and prompt-section callbacks too
The hook loop now treats SystemExit like any other callback failure; the
three sibling dispatch loops in this module (event subscribers,
invoke_middleware, system prompt section content) had the same
`except Exception` gap and let a plugin dependency's sys.exit() abandon
the loop and kill the turn. KeyboardInterrupt still propagates on all
of them.
2026-09-20 15:21:14 -07:00
Jony
d1221d43b7 fix(plugins): isolate SystemExit from hook callbacks
(cherry picked from commit 99facb98dc9e9c88af6c4fcc1049c6018871e520)
2026-09-20 15:21:14 -07:00
webdevtodayjason
118984d7a0 fix(plugins): share the loader's SUPPORTED_MANIFEST_VERSION with the installer (#85879)
The installer carried a private manifest-version cap of 1 while the runtime loader (and the
docs) accept manifest_version 2, so 'hermes plugins install' refused every v2 plugin the loader
happily loads. Read the shared constant instead. The alignment landed once in e834ee8971 and
was reverted by the follow-up 5bad3d54ee.

Salvaged from PR #85893 by @webdevtodayjason; rebased onto the decomposed plugins_cmd.py.
2026-09-20 14:29:39 -07:00
teknium1
78004e8463 fix(plugins): hermes plugins doctor understands kind: model-provider
Doctor pushed every plugin through PluginManager._load_plugin, which demands register(ctx);
model-provider plugins register a ProviderProfile at import through providers/ discovery (the
manager skips the kind on purpose), so every valid provider plugin failed with
'no register() function' while 'hermes plugins validate' passed it. Route the kind through
providers._import_plugin_dir, report the registered profiles, and restore the live registry
and module state afterwards.

Same diagnosis and shape as PR #86160 by @oliver-mee (Aug 14), which reached it first; this
version imports through providers._import_plugin_dir (the real discovery loader) instead of
the PluginManager namespace. Credit: #86160.

Co-authored-by: Oliver Mee <102673257+oliver-mee@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
13fe9c7171 feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:

- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
  `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
  declaring profile; standard records still replay on OpenRouter-style routes, strict routes
  drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
  status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
  and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
  one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
  without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".

The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.

Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
9fd49de9d8 test: local-runtime recovery fixture publishes its router record by rename
The parent polls ready.exists() and then json.loads the file, but the child wrote the record in
place, so the file was visible (and empty) the instant it was opened; on a loaded windows-latest
runner that read raced the dump and every parametrized case could die with
JSONDecodeError: Expecting value (seen on #117451's Windows-only lane, legacy-models case).
Write to <ready>.tmp and os.replace it into place, so exists() implies a complete record.
2026-09-20 14:23:15 -07:00
teknium1
5c755d8ee6 test: deflake the Windows-faked orphan reaper and the timed-out-child relay test
Both failed once and passed on retry in a CI run of #117451; each is a real race in the test
harness, reproduced deterministically here and green after the fix under the same simulation.

test_gateway.py: tests that fake is_windows() send the reaper through
_windows_scheduled_task_state, which spawns pwsh whenever one is on PATH (GitHub's ubuntu
runners ship it). On a loaded runner that spawn outlives its 10 s timeout and subprocess.run
kills it through the test's globally patched os.kill, so a foreign PID lands in killed_pids
(assert [(29408, SIGKILL)] == []). A module autouse fixture stands the probe in with None;
the one test of the probe itself undoes it. Repro: a fake `pwsh` that sleeps 15 s on PATH.

test_zombie_process_cleanup.py: the 0.1 s child cap could elapse before the executor's worker
thread had opened the child's turn, turning "late result after timeout" into "child never
started" (assert child_started.is_set()). The settle event the cap rides on now waits for the
child to start before honouring the timeout. Repro: a 0.3 s sleep in the worker before fn().
2026-09-20 14:23:15 -07:00
teknium1
caa2e2dd99 feat(cli): background processes show in the live-work dock and monitor
terminal(background=true) spawns were only a "⚙ N" count in the status bar
while subagents got a dock, a roster, tails and controls. The classic CLI
dock now paints a Processes block under the subagent rows (command,
elapsed, latest output line; exit verdict once finished, retained 60 s),
and the Ctrl+T/F6 monitor lists processes under the agents with Enter =
log tail and x = stop that process. Processes cannot be steered.
2026-09-20 13:55:03 -07:00