Commit Graph

513 Commits

Author SHA1 Message Date
ethernet
8da6c5f386 Merge branch 'ethie/release-machinery' into ethie/pm-clean
# Conflicts:
#	tests/ci/test_desktop_store_eligibility.py
2026-09-22 11:05:14 -04:00
ethernet
207f8fedfd Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/local_runtime/binaries.py
#	hermes_cli/plugins_cmd.py
#	hermes_constants.py
#	tests/test_hermes_constants.py
#	tests/tools/test_clipboard.py
#	tests/tools/test_voice_wsl_pipewire.py
#	tools/computer_use/cua_backend.py
#	tools/voice_mode.py
2026-09-22 09:55:47 -04:00
Siddharth Balyan
4094ab610d hermes_platform.resolver: locate/inspect/probe tiers, AppResolver, gh lookup migrated (NS-921) (#118065)
* feat(platform): resolver core with locate/inspect/probe tiers and ordered candidates

Every resource lookup needs one result shape and one cost contract. `locate` reads
metadata only, `inspect` may open files and call OS APIs in-process, `probe` is fresh
and the only tier that may spawn or connect. `Resolution.candidates` keeps probe order
so fan-out consumers can try every present binary.
Linear NS-921.

* feat(platform): AppResolver over AppDef with plist, PE, registry, and server.json sources

Desktop apps need presence, version, and liveness as separate observations. The runtime
file's bearer token is parsed, used for one request, and discarded inside the probe;
no public type carries it. Endpoints are accepted only when loopback with a numeric port.

* refactor(copilot): gh candidates through locate_command and the Homebrew table

First consumer of the resolver. The gh token probe still tries every present binary in
order; the allowlist loses its two copilot_auth rows.

* feat(platform): availability() over an application declaration

locate() + inspect() only, never probes; the fail-closed _version in
app.py treats a vendor's plist/PE/registry entry as untrusted input.
Salvaged from PR #118122; reads any object with requires_app,
min_version, app_for(os) — nothing here imports the MCP catalog.

* feat(platform): application declarations parsed into AppDef per OS

The parser slice of PR #118122's catalog manifest, re-homed as a
catalog-free module: whoever owns an MCP server declares the app it
fronts per OS and what it needs, and registers it here. Stdlib +
hermes_platform.resolver only. register/lookup/clear are the one seam
the MCP check_fn and the skill gate both read.

* feat(mcp): check_fn honours a registered application declaration

_make_check_fn ANDs the declared app's availability into the
connection-alive check; with nothing registered for the server the
behaviour is the pre-PR3 connection check. Provenance is explicit
registration, not endpoint matching. Returns a plain bool: the registry
caches bool(fn()).

* feat(skills): requires_apps gate through registered declarations

Offer-time filter beside environments:; names resolve through
hermes_platform.declaration, an unknown name hides the skill (fail
closed). The disk snapshot carries requires_apps and the fast path
re-evaluates it (snapshot version bumped to 3): app presence is a host
fact that changes without SKILL.md changing.

* docs: application declarations page

The plugin-facing schema reference: app: and requires: blocks,
availability() states, and the two gates that read the registry.
Registered under Extending > Plugins in the docs sidebar.

* test(platform): declaration parser, availability, gates

The PR3 app-block tests re-homed off the catalog: fixtures are dicts
passed to parse_declaration, the check_fn gate keys on explicit
registration (not endpoint matching), and the import-hygiene probe now
covers hermes_platform.declaration and resolver.availability.
2026-09-22 17:08:30 +05:30
ethernet
c13287c915 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/electron/main.ts
#	hermes_cli/backup.py
#	hermes_cli/config.py
#	hermes_cli/plugin_catalog.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/plugins_cmd_catalog.py
#	hermes_cli/plugins_discovery.py
#	hermes_cli/profiles.py
#	hermes_cli/update_cmd_deps.py
#	pyproject.toml
#	tests/gateway/test_dm_topics.py
#	tests/hermes_cli/test_config.py
#	tests/hermes_cli/test_plugins_cmd.py
#	tests/hermes_cli/test_update_autostash.py
#	tests/tools/test_lazy_deps.py
#	tools/lazy_deps.py
#	tools/skill_ledger.py
#	utils.py
#	website/docs/user-guide/security.md
2026-09-22 05:16:50 -04:00
teknium1
5c0e73eff1 feat(desktop): render plugin-declared settings in the Plugins tab (#46600, #87934)
A plugin manifest's `config_schema` now reaches the Desktop: `plugins.manage list`
returns each plugin's schema with the current `plugins.entries.<id>.settings`
values (`settings_schema`), and a new `settings` action writes edits through
`hermes_cli.plugins_state.save_plugin_setting` — the writer extracted from
`PluginContext.set_config`, so the plugin, the CLI and the Desktop share one
config path, one lock and the same managed-install / managed-key refusals.

The Plugins tab grows a gear per plugin with a schema; the inline form is
table-driven (`FIELD_CONTROLS` / `INITIAL_TEXT` / `COERCE` keyed on the wire
field type) for string / number / boolean / enum / json / secret. Secrets are
declared with `type: secret`: the row carries only the `.env` name and a
presence flag, the client writes the value through the existing `PUT /api/env`
credential route, and the RPC refuses secret keys so nothing lands in
config.yaml.

Contracts regenerated; docs gain a "Settings form in the Desktop" section.
2026-09-22 01:48:18 -07:00
teknium1
0e5809566f feat(plugins): fire pre/post_auxiliary_call events on every auxiliary LLM call (#79733)
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.

- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
  post_api_request payload shape plus `aux_task`, `api_request_id`
  (`aux-...`, shared by every attempt of one logical call), `retry_count`,
  `streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
  turn is in flight; fail-open (a raising/hung subscriber is logged and
  the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
  attempt shares (_relay_sync_completion / _relay_async_completion /
  _relay_sync_stream) run under the hook pair — retries and fallbacks
  included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
  sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
  tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
  aux_task and no api_request events; raising subscriber never breaks
  the call). First is red on origin/main.

Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.

Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-22 01:19:12 -07:00
teknium1
e1a6679399 fix(desktop): runtime plugin loader survives hung imports, leaked timers, duplicate ids and broken reloads
Four error-isolation holes in the disk plugin door
(apps/desktop/src/contrib/runtime-loader.ts), found by a static+live audit of
the loader; none had an issue filed.

- A plugin whose module evaluation never settles (top-level `await` on a dead
  host) hung `import()` forever and, through the scan's sequential loop and
  its re-entrancy guard, froze every later plugin and all future scans until
  restart. `import()` now races a 10 s deadline; the plugin errors on its own
  row ("import timed out") and the scan continues.
- Timers and DOM listeners a plugin took out with bare globals survived
  disable and every hot-reload. `ctx.setTimeout` / `ctx.setInterval` /
  `ctx.addEventListener` are tracked with the plugin and torn down on
  unload; the SDK doc says bare globals are not.
- Two folders exporting one plugin id silently last-wins: the second
  disposed the first's registrations and each hot-reload flipped ownership.
  The first (folder-name sorted, so deterministic) owns the id; the later
  file errors on its own row ("duplicate id, already loaded from <path>").
- A save that no longer loads (syntax error, timeout, duplicate) left the old
  incarnation's contributions and activate handle live beside the error
  row, so the Plugins tab showed a broken file as "loaded" and could
  re-enable stale code. The previous incarnation is unloaded and dropped.

Tests: one invariant per fix in runtime-loader.test.ts, all red on base
(the hang case red by timing out).
2026-09-22 00:46:24 -07:00
teknium1
903c80ff1b docs: document the hooks directory as a trusted-by-placement extension point
`~/.hermes/hooks/` auto-loads every valid `HOOK.yaml` + `handler.py` at
gateway startup with no `plugins.enabled` gate. That is the documented
contract since 3988c3c245 ("Implicit (dir trust)" in the comparison
table), but the plugins page's "disabled by default" promise read as if
it covered gateway hooks too (#37963). Maintainer ruling: keep implicit
dir trust, fix the docs.

- hooks.md: new "Trust model" section stating exactly what loads, when,
  how, and that placing the files is the opt-in; comparison-table cell
  links to it and the plugin-hooks consent cell now says
  `plugins.enabled`.
- plugins.md: note scoping `plugins.enabled` away from gateway hooks.
- developer-guide/plugins: one sentence at the gateway-hook recipe.
- security.md: "Trusted-by-placement extension points" section
  cross-linked from hooks.md.
2026-09-22 00:37:39 -07:00
ethernet
d0d4e91434 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/plugins_cmd.py
#	hermes_cli/plugins_cmd_catalog.py
#	tests/hermes_cli/test_web_plugins_catalog.py
#	website/docs/reference/cli-commands.md
#	website/docs/user-guide/features/plugin-catalog.md
2026-09-22 02:59:10 -04:00
teknium1
913c045c53 fix: resolve user-installed context engines from $HERMES_HOME/plugins/<name>
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.

Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.

Fixes #61839
credit: @giggling-ginger #61995
2026-09-21 23:38:16 -07:00
teknium1
9cbc221dcb fix: clone plugin context engines via clone_for_agent(), not blind deepcopy
A general-plugin context engine is one shared instance; agent init copied it per agent
with copy.deepcopy() only. Engines that hold a SQLite connection or lock (hermes-lcm)
already expose clone_for_agent() for exactly this, but it was never called, so every
init logged "could not be safely copied … falling back to built-in compressor" and the
engine was unusable through the plugin system.

ContextEngine grows clone_for_agent() (default: deepcopy, the previous behaviour) and
_select_context_engine calls it; the failure message now names the hook to override.
Docs: context-engine-plugin.md documents the per-agent clone contract.

Test change (existing on main): tests/agent/test_context_engine.py::
test_agent_init_source_deepcopies_singleton_not_aliases was a source-reading pin on the
literal `copy.deepcopy(_candidate)` line, which this fix intentionally replaces. It is
superseded by tests/agent/test_plugin_context_engine_clone.py, which drives the real
_select_context_engine seam and asserts the invariant it guarded (child update_model()
never mutates the shared singleton) plus the new clone_for_agent() path.

Fixes #99640
credit: @stephenschoettler #62374
credit: @686f6c61 #99677
2026-09-21 23:38:16 -07:00
teknium1
01eae47b3a fix: dispatch only the context kwargs a tool handler declares
ToolRegistry.dispatch spread every injected keyword (task_id, session_id,
user_task, parent_agent, ...) into the handler, so a third-party plugin tool
written as `def handle(args)` raised TypeError on every call. The plugin
contract (plugins/AGENTS.md) says optional kwargs are signature-inspected, not
forwarded unconditionally; hooks already do this in plugins_dispatch. Handlers
taking **kwargs still receive the full payload. Docs updated: the narrow
signature is supported, **kwargs opts into the whole context.

Fixes #68318
credit: @vveerrgg #22146 (slim redo; #68636 is the later duplicate)
2026-09-21 23:38:16 -07:00
ethernet
dcd678d6fd fix(release): close publication custody gaps 2026-09-22 02:19:38 -04:00
teknium1
eba8199ce3 docs(desktop): say plainly that desktop plugins are not sandboxed
The catalog page claimed admission's lint means a marketplace install 'cannot quietly rewire the app'; the lint is a handful of regexes and plugin.js runs in the app realm with the full window.hermesDesktop bridge. The user guide, catalog trust model and SDK security section now describe the real model: human review of a pinned SHA plus two tripwires (lint + loader import allowlist), no isolation.
2026-09-21 22:39:35 -07:00
ethernet
c9bd7459c2 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	scripts/run_tests_parallel.py
#	tui_gateway/model_switch.py
2026-09-22 00:28:30 -04:00
ethernet
8e57bb6c37 docs: document claim-driven stable releases 2026-09-21 23:40:53 -04:00
teknium1
eeb220d40c fix: model-provider plugins resolve per bound profile home (#88143)
Symptom: a model-provider plugin installed with `hermes plugins install`
(e.g. claude-subscription-directsdk) worked from a terminal but the Desktop
app failed every session build with "Unknown provider '<name>'".

Cause: `providers/__init__.py` scanned `$HERMES_HOME/plugins` exactly once
per process, under whichever profile home happened to be bound at the first
lookup, and the registry was global. The Desktop backend and the multiplex
gateway serve several profiles from one process, so any profile other than
the first-discovered one never saw its own plugins, and a plugin installed
while the process ran was invisible until a restart.

Change: bundled, pip and legacy providers stay process-wide; `$HERMES_HOME`
plugins load into a per-home layer keyed by `hermes_home_key()` at lookup
time (`get_provider_profile`, `list_providers`, `provider_source`). The
layer rescans when the plugin directories' mtimes change, so a fresh install
is found on the next lookup. The registration target is a ContextVar so two
turn threads scanning two homes cannot cross-register, and no lock is held
across plugin imports (a lock there could deadlock against a thread
mid-`import hermes_cli.auth`). Per-home module names let two profiles carry
the same plugin. `plugin_dev` reuses the module-name helper.

Live repro (tui_gateway `session.create` on a secondary profile whose config
selects a plugin installed only there): base -> agent_error "Unknown
provider 'fakeprov-b'"; fixed -> provider resolves and the build proceeds to
the plugin's own runtime check.
2026-09-21 18:38:36 -07:00
ethernet
1f38e868ea docs: register the pm/release developer pages in the sidebar
The five pages moved from docs/ into website/docs/developer-guide were
never added to sidebars.ts, so the site did not link them, and they kept
repo-relative links (../tests/..., ../scripts/..., ../website/docs/...)
that only resolved from the old docs/ location. Register them under a
Packaging & Releases category and point the repo-file links at GitHub;
package-management.md now links the moved completion note in place.
2026-09-21 18:37:43 -04:00
ethernet
fea1f34000 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	tests/tools/test_code_execution.py
2026-09-21 18:21:37 -04:00
teknium1
10659d3536 docs(caching): list every machine-paced source for cache_ttl: auto
The config comment and both docs pages named six of the eight sources in
MACHINE_PACED_SOURCES; "tool" and "batch" were missing, so an operator
reading the docs could not predict the tier those sessions get.
2026-09-21 15:07:39 -07:00
teknium1
23301d705c feat(caching): prompt_caching.cache_ttl: auto picks the 1h tier for human-paced sessions, 5m for machine-paced ones
The 1h Anthropic cache tier writes at 2x base (5m: 1.25x) and only pays off when
turns are more than five minutes apart. That is exactly the shape of an interactive
session a person parks and resumes, and exactly not the shape of a subagent, cron
run, one-shot or webhook that calls every few seconds and is gone. A single global
`cache_ttl` cannot be right for both, so operators leave it on 5m and pay a full
context re-write every time they come back to a CLI session after a coffee.

Measured on one install (2 days of per-call API logs, Claude via the Nous route):
63% of interactive cache-write tokens were cold re-writes after a 5-60 minute idle
gap; 1h would cut interactive write cost ~42% while costing ~49% more on subagents
and ~23% more on cron. `auto` resolves once per session from the session source
(`_session_source_for_agent`): 1h for cli/tui/desktop/messaging platforms, 5m for
subagent, cron, oneshot, webhook, kanban, api. Auxiliary/stub calls keep 5m; the
delegate_tool child clamp (#104168) still applies. Default stays "5m".
2026-09-21 15:07:39 -07:00
ethernet
4efb36f81e merge: integrate upstream desktop features through PM preparation
Merge origin/main at 8e806ae1b2. Keep native helper compilation in
prepareDesktopNativeDependencies and keep bundling/beforePack consume-only.
Bind helper sources and headers into preparation identities and cache keys;
copy admitted executable resources beside node_modules and preserve signing
semantics in product freshness checks.

Verified desktop typecheck, focused native/packaging/UI and gateway/cache
tests, and the real Linux preparation/copy/Xvfb execution path. Incoming
upstream anti-slop findings remain unchanged; no baseline was raised.
2026-09-21 15:17:44 -04:00
teknium1
91d1339a8d docs: multi-profile gateway pages match the multiplex-only topology
Every line still promising that `gateway.multiplex_profiles: false` keeps
per-profile gateways, or documenting the deleted `migrate --standalone`
rollback, now describes the shipped behaviour: one gateway per host serves
every profile; `false` no-ops with a warning; `hermes update` folds a fleet
unless a real boundary (different UNIX user, HERMES_HOME outside profiles/)
holds; a blocked fleet keeps `--force` per profile; re-running
`migrate --multiplex` converges a half-migrated host; a manifest on disk is the
resume record, never a rollback. Adds the new "No new per-profile gateways"
section with the exact refusal `hermes -p <name> gateway install` prints.

Files: website/docs/user-guide/multi-profile-gateways.md,
website/docs/reference/cli-commands.md (migrate row),
website/docs/developer-guide/multiplexing-gateway.md (eager activation no
longer honours `false`), hermes_cli/AGENTS.md (migrate + refusal seams).

Completes #118273 (docs item). Part of #109417
2026-09-21 11:40:48 -07:00
ethernet
f622ea76b3 merge: integrate origin/main while preserving PM ownership 2026-09-21 13:09:28 -04:00
kshitijk4poor
804cc3c021 fix(gateway): a standalone owner is the per-profile topology, not a refusal
`gateway/host_attach.py::decide()` treated a host owner that answers `multiplex: False`
to the rescan — another profile's standalone gateway — the same as a multiplexer whose
roster excludes us, and issued the permanent REFUSE. On the supervised path that is
exit 78, which `hermes_cli/stderr_timestamp.py` maps to 0 so launchd's
`KeepAlive.SuccessfulExit=false` parks the unit. On a one-process-per-profile fleet
(the topology `multi-profile-gateways.md` documents and `multiplex_profiles: false`
promises to keep) every launchd gateway except the first to claim the host lock was
parked at boot, silently; which ones survived was a boot race against the owner's
record landing.

`request_serve_profile` now returns the owner flagged `standalone` instead of `None`
for a `multiplex: False` answer, and `decide()` turns that into START — this profile
runs its own gateway beside the owner, as before #118097; the host-lock claim still
logs the two-gateway topology and `gateway migrate --multiplex` stays the converge
path. REFUSE is unchanged for a multiplexing owner that excludes the profile.

A/B against a REAL owner process (real host lock, record and control socket answering
the rescan): base → `refuse`, exit 78 → launchd 0 (parked); fix → `start`. Negative
control (owner answers `multiplex: True`, roster excludes the profile): `refuse` on
both.

Field report: debug share 5024e996 (11 launchd profile gateways, 0.21.3 ea0c2b82 —
only 2 of 11 came back after `hermes update`), discussed on #118097.
2026-09-21 08:38:15 -07:00
ethernet
9d036f8ac9 merge origin/main (15 commits) into ethie/pm-clean
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
2026-09-21 08:16:57 -04:00
teknium1
a10620a669 fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.

- ATTACH now requires a LIVE `identify` answer. The claim-time record is
  published with NO served set (the runner settles multiplex a moment later),
  and `host_gateway()` reports `served_known=False` when nothing answers. An
  owner whose served set is unknown yields a TRANSIENT refusal, never an
  attach: previously `default`'s claim published "default,other" before its
  socket bound, `other`'s systemd unit read that as "I am served", exited 78,
  and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
  instead of forcing `multiplex=True`, so a standalone gateway stops claiming
  the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
  `--force` into `_host_attach_or_none`. Both previously exited in the guard
  before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
  config refusal every supervisor parks on; "someone else serves me right now"
  is a runtime observation that ends when that process does. No unit files
  change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
  finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
  would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
  and re-enters with `replace=True`, so it can no longer attach to the corpse
  it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
  `st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
  it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
  generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
  `gateway status`/doctor across N profiles pays one probe, not N.

Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
2026-09-21 05:02:29 -07:00
teknium1
a20a88398f gateway: lifecycle verbs mean "the one host multiplexer"
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:

- `gateway run` for a profile the host gateway already serves ATTACHES: print
  its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
  to re-scan `profiles/` (control socket) and attach once the answer includes
  it. Refuse only when the host gateway cannot be made to serve it. Under a
  service supervisor the attach exits 78 instead of 0 so a redundant unit is
  parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
  landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
  the owner's control socket, so it no longer depends on the claim ordering in
  start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
  process: they restart the host multiplexer and preserve its served set. A
  secondary still running its own gateway is reported with the
  `gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
  argv: a host singleton runs bare/default argv and can never prove it serves
  profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
  multiplexer is whichever profile launched the one host process.

Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
2026-09-21 05:02:29 -07:00
ethernet
9de542c15b merge origin/main (16 commits) into ethie/pm-clean
main folded the auto-archive housekeeping into per-profile state.db maintenance; the plugin
update-check chore stays. The consolidated-away TestWorkdirParallelPool stays removed. Two new
utf-8 reads in gateway/host_rendezvous.py read utf-8-sig.
2026-09-21 06:09:56 -04:00
teknium1
bfcd209906 docs(multiplexing): record the config escape hatch and the .env-over-frozen-env precedence
Activation is not purely stricter for the LAUNCH tenant: inside the launch scope its own
.env wins over the frozen boot env, so a key present in both reads differently before and
after activation. Also documents the explicit multiplex_profiles: false escape hatch and
the fail-closed unresolved-profile case.
2026-09-21 02:58:40 -07:00
teknium1
c718267ef3 fix(multiplex): close the profile-scope holes the scope machinery misses
One host process serves every profile, so every execution point must bind the
profile it is acting FOR. These six ran unscoped (or bound only part of a
scope) and resolved get_hermes_home()/credentials against the LAUNCH profile:

- tui_gateway/session_reaper: the idle-reaper and exit-flush transcript writes
  now enter the SESSION's profile scope, the same chokepoint _finalize_session
  already binds. A served profile's transcript was landing in the launch home.
- gateway/run: MCP shutdown tears down per served profile inside that
  profile's scope (mirrors startup discovery and the reconcile chore), with a
  trailing wildcard pass under the launch profile's own scope.
- gateway/run_profile_reconcile: _unserve_profile's adapter teardown, agent
  eviction and state/memory handle release now run inside the deleted
  profile's scope.
- gateway/run_adapters + run_goals + run_notifications: a body with no routed
  profile no longer means "no scope". launch_profile_scope_if_multiplexed()
  binds the launch profile once the process multiplexes; before activation it
  is still literally a nullcontext, so single-profile hosts are unchanged.
- hermes_cli/kanban_db_dispatch: one _worker_profile_scope helper binds the
  assignee's secret AND terminal scope for toolset resolution and the spawn-env
  build, unconditionally instead of only under multiplex.
- hermes_cli/web_server: `hermes serve` activates multi-profile hosting at boot
  when the host has more than one servable profile home, instead of lazily on
  the first ?profile= request after earlier work already ran unscoped.

Secret scope is never widened: a non-launch home resolves from its own .env and
sources only; the launch home keeps its existing env-over-.env precedence.
2026-09-21 02:58:40 -07:00
teknium1
9ecd22e5da fix(cron): release a profile's in-flight claim under the key it registered with
Under one multiplexing ticker, every non-launch profile's in-flight claim
leaked on every run. The cron home scope is a ContextVar: the claim is taken on
the ticker thread inside `_profile_cron_scope`, but the pool worker's `finally`
sits outside `ctx.run`, where the worker thread resolves the LAUNCH home — so
the discard missed the real key. The job then skipped a fire window until the
force-release backstop swept it, and the shutdown drain plus
`hermes.cron.jobs.running` saw phantom work.

`_submit_with_guard` now captures the registering home and passes it to
`release_running_job(job_id, home=...)` on every release path.

Also in this pass:

- Stand down for a profile that runs its OWN gateway (`_cron_profile_gate`,
  the gate the serve/Desktop ticker already passes). On a host pinned to
  per-profile gateways both processes raced that profile's tick lock, and when
  the launch gateway won, delivery went through SharedRouteAdapters/fail-closed
  instead of the profile's live adapters. The gate compares the liveness PID
  against `os.getpid()`: this process holds the launch `gateway.pid` and
  publishes every served profile in `served_profiles`, so a bare liveness answer
  would have stood cron down host-wide.
- Job-liveness consumers ask `is_job_running(job_id, home=...)` instead of the
  host-wide bare-id union, which let profile A's running `daily-brief` report
  profile B's idle one as running and keep B's stale one-shot alive.
- `register_ticked_homes` reaps the parallel pools of homes that leave the
  ticked set; pools lived until `atexit`, so every home ever ticked kept a
  ThreadPoolExecutor and its worker threads.
- `mark_running_jobs_interrupted` reads the real home Path from
  `_inflight_home_path` instead of rebuilding it from the normcased key half.
- The ownership test imports `cron.scheduler_ownership` per test, so the file
  fails behaviourally on base instead of as a collection error.
2026-09-21 02:52:15 -07:00
teknium1
cb647a018f fix(cron): one host ticker owns every profile's cron, per profile
The cron ticker multiplexes N profiles from one process while its ownership
predicates and its in-flight bookkeeping still assumed one profile per process.

- `_should_yield_tick_to_fresh_gateway` asked a process-global boolean
  (`owns_gateway_runtime_lock`) and a launch-home lock probe, so one answer
  covered every profile ticked. It now asks `scheduler_ownership`:
  `owns_cron_tick_for(home)` (this process is the host gateway AND ticks that
  home) and `live_gateway_ticking(home)` (another live host gateway whose
  published served set covers that home).
- In-flight state (`_running_job_ids`, `_running_since`, `_running_futures`,
  `_running_allowance_s`, `_running_worker_pids`, `_running_fire_owners`,
  `_restart_safe_waiter_job_ids`, `_interrupted_job_ids`) is keyed by
  `_inflight_key(job_id)` = `(home key, job id)`; two profiles carrying a
  `daily-brief` no longer read as one job. The public accessors still report
  the host-wide union of bare job ids for the shutdown drain.
- The parallel worker pool is keyed by home: `cron.max_parallel_jobs` is a
  per-profile key, and the single global pool was sized by whichever profile
  ticked first and torn down by the next one.
- `gateway/run.py` no longer gates the cron tick set on
  `gateway.multiplex_profiles`: that flag gates adapters, and with it off every
  non-launch profile's jobs sat in a store no ticker visited.
2026-09-21 02:52:15 -07:00
ethernet
a537981a90 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-21 05:07:33 -04:00
ethernet
9ac6af3809 merge origin/main (28 commits) into ethie/pm-clean
main extracted the launchd backend and the setup wizard out of hermes_cli/gateway.py. The
eight launchd functions pm-clean had changed are ported into gateway_launchd.py in its
_gw() style: runtime_command/installation_command for the gateway argv (no VIRTUAL_ENV in
the plist, XML-escaped args), _prepare_service_launcher before every plist write,
utf-8-sig plist reads, and launchd_restart's refresh-first + bounded bootstrap revival.
The systemd service-unit cluster stays in the facade (no gateway_service_unit sibling).
backup.py takes main's browser_profiles backup-only exclusion on our profile_root_entry
shape. main.ts takes main's attach-first backend block with our explicit types.
2026-09-21 05:07:21 -04:00
teknium1
8863b36fd6 fix(cron): stale-code yield reads as an outage in cron status; hermes update restarts proven-stale gateways
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:

* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
  loop) and the success marker; with no success marker on disk it printed
  "✓ Gateway is running — cron jobs will fire automatically", and with a stale one
  it pointed at "Check the gateway log" instead of naming the cause. The persisted
  `CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
  and reported as "Gateway is running STALE code — fires NOTHING", with both
  revisions and the restart command. A fresh heartbeat with a recorded error and no
  success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
  left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
  to the existing drain-first `request_restart` path (SIGUSR1) via
  `hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
  code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
  contract the restart phase already uses for unmapped manual gateways. The drain
  budget computation is shared (`_gateway_drain_budget`).

The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).

Fixes #117275
2026-09-21 01:57:31 -07:00
teknium1
b7803a1763 fix(compression): cap the protected tail at 20% of the context window
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole
rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the
"protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every
compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B
timed out before compaction ever changed anything, and protect_last_n read as an uncompressed
tail rather than a minimum.

TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk /
pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool
groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of
32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%).

Probe (12 tool-heavy turns, 49 rows, 12.8K tokens):
  ctx    8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows;  window [4,10) -> [4,42)
  ctx   16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34)
  ctx   32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22)
  ctx  128K: identical before/after
2026-09-21 01:04:16 -07:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
teknium1
e089be80a1 docs(desktop-sdk): note DecodeText loop is opt-in
DecodeText is a published plugin-SDK export; flipping its `loop` default
to false silently changes third-party plugin visuals, so the SDK doc says
so where the component is listed.
2026-09-20 20:14:20 -07:00
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
bdde0b0c28 docs(cron): describe when a stale-code ticker yields and when it keeps dispatching
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped
code_sha is the on-disk revision; a lock held by an equally stale process never
counts. Placed under "Gateway Integration" so it does not collide with the
"Stale-code yield" section #117501 adds under "Locking".
2026-09-20 20:13:19 -07:00
teknium1
13fe9c7171 feat(providers): external-process provider support for standalone model-provider plugins (from #105863)
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:

- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
  `get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
  declaring profile; standard records still replay on OpenRouter-style routes, strict routes
  drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
  status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
  and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
  one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
  without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".

The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.

Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
2026-09-20 14:29:39 -07:00
teknium1
41b6ba09d9 docs: name the registered tool in prose that still says cronjob/todo/process
The same rename that left the two reference pages stale (cronjob->cronjob_manage,
todo->todo_list, process->process_manage) left the guides and developer docs
referring to the old tool names; a reader following them gets "no such tool".
Toolset names (`cronjob`, `todo`) are unchanged and left alone.
2026-09-20 12:56:25 -07:00
teknium1
297ebd506e test: trim the ambient-chain refusal tests to two two-home invariants
The 19 cherry-picked tests covered each refusal branch separately. One
A -> B -> A test per adapter over two real homes now proves the whole
contract at the production entry (build_credential / _cached_client): the
launch profile keeps its own credential, the cred-less served profile is
refused before the SDK chain (or boto3) is touched, the launch profile is
unaffected afterwards, and the standalone run keeps today's ambient chain.
Docs: the Azure guide and the multiplexing design page name the refusal.

Superseded #116370 (@JoaoMarcos44) proposed the same mechanism.
2026-09-20 12:09:47 -07:00
teknium1
948753d0f8 docs(plugins): enable prompts for the override grant only when the manifest declares capabilities 2026-09-20 11:52:45 -07:00
teknium1
967a7c94f7 docs(plugins): a provider plugin can register its own api_mode transport 2026-09-20 10:11:40 -07:00
teknium1
e5e7fbcd27 fix(auth): a user provider plugin's endpoint overrides the bundled row too
A `$HERMES_HOME/plugins/model-providers/<name>/` plugin re-registering a
bundled provider (stepfun with a regional base_url, gmi at a staging host)
wins in `providers._REGISTRY` — register_provider() is last-writer-wins and
the plugin guide promises exactly this — but the runtime reads its endpoint
from `hermes_cli.auth.PROVIDER_REGISTRY`, whose mirror loop skipped every
name already present, so inference kept going to the built-in URL (#48450).

The mirror now applies one explicit precedence rule: when a row core wrote
(built-in or plugin-mirrored) belongs to a name whose profile
currently registered came from a USER plugin,
the row's profile-derived fields are rewritten in place (inference_base_url;
api_key_env_vars / base_url_env_var on api-key rows when the profile declares
env_vars). `providers` records the discovery source per registration
(`provider_source()`), because a bundled profile must never rewrite a
built-in row: several bundled profiles omit the row's `*_BASE_URL` env var
and one differs in auth_type, so an unconditional "profile wins" would have
changed built-in behaviour. With no user plugin PROVIDER_REGISTRY is
byte-identical before/after (78 rows probed). copilot/kimi/zai keep their
bespoke resolution via the existing skip set.

Co-authored-by: xiaoxinova <xiaoxinova@users.noreply.github.com>
2026-09-20 10:06:29 -07:00
teknium1
9f0cd9d773 docs: plugin profiles and their aliases resolve in /model --provider and the model picker 2026-09-20 10:01:58 -07:00
teknium1
2182f51d7c fix: name the process holding the state.db write lock when a writer times out
"database is locked (another Hermes process held the state.db write lock for
over 60s)" identified the victim only. The open-descriptor scan cannot single
out the writer because every Hermes process (gateway, CLI sessions, worktree
agents, cron) has the DB open, so an operator hit repeatedly by
session_persistence_failed:locked had nothing to act on.

SQLite's unix VFS takes fcntl byte-range locks whose offset encodes the lock
kind (state.db-shm byte 120 = WAL write, 121 = checkpoint; the pending-byte page
on state.db = PENDING/RESERVED), and the kernel exports them with the owning pid
in /proc/locks. hermes_state_lockowners reads that table at the moment the
patience budget runs out and logs one WARNING per write-class holder with
describe_holder_pid()'s argv summary, for both the transcript write path and
open+init lock patience. The holder stays out of the exception text on purpose:
classify_persistence_error() buckets by phrase and a holder argv such as a
worktree named fix-corrupt-db would flip the bucket.

Docs: the Write Contention section still described attempt-counted retries
(_WRITE_MAX_RETRIES = 15); updated to the time budgets in force and the new log line.
2026-09-20 10:01:42 -07:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00