* feat(platform): resolver core with locate/inspect/probe tiers and ordered candidates
Every resource lookup needs one result shape and one cost contract. `locate` reads
metadata only, `inspect` may open files and call OS APIs in-process, `probe` is fresh
and the only tier that may spawn or connect. `Resolution.candidates` keeps probe order
so fan-out consumers can try every present binary.
Linear NS-921.
* feat(platform): AppResolver over AppDef with plist, PE, registry, and server.json sources
Desktop apps need presence, version, and liveness as separate observations. The runtime
file's bearer token is parsed, used for one request, and discarded inside the probe;
no public type carries it. Endpoints are accepted only when loopback with a numeric port.
* refactor(copilot): gh candidates through locate_command and the Homebrew table
First consumer of the resolver. The gh token probe still tries every present binary in
order; the allowlist loses its two copilot_auth rows.
* feat(platform): availability() over an application declaration
locate() + inspect() only, never probes; the fail-closed _version in
app.py treats a vendor's plist/PE/registry entry as untrusted input.
Salvaged from PR #118122; reads any object with requires_app,
min_version, app_for(os) — nothing here imports the MCP catalog.
* feat(platform): application declarations parsed into AppDef per OS
The parser slice of PR #118122's catalog manifest, re-homed as a
catalog-free module: whoever owns an MCP server declares the app it
fronts per OS and what it needs, and registers it here. Stdlib +
hermes_platform.resolver only. register/lookup/clear are the one seam
the MCP check_fn and the skill gate both read.
* feat(mcp): check_fn honours a registered application declaration
_make_check_fn ANDs the declared app's availability into the
connection-alive check; with nothing registered for the server the
behaviour is the pre-PR3 connection check. Provenance is explicit
registration, not endpoint matching. Returns a plain bool: the registry
caches bool(fn()).
* feat(skills): requires_apps gate through registered declarations
Offer-time filter beside environments:; names resolve through
hermes_platform.declaration, an unknown name hides the skill (fail
closed). The disk snapshot carries requires_apps and the fast path
re-evaluates it (snapshot version bumped to 3): app presence is a host
fact that changes without SKILL.md changing.
* docs: application declarations page
The plugin-facing schema reference: app: and requires: blocks,
availability() states, and the two gates that read the registry.
Registered under Extending > Plugins in the docs sidebar.
* test(platform): declaration parser, availability, gates
The PR3 app-block tests re-homed off the catalog: fixtures are dicts
passed to parse_declaration, the check_fn gate keys on explicit
registration (not endpoint matching), and the import-hygiene probe now
covers hermes_platform.declaration and resolver.availability.
A plugin manifest's `config_schema` now reaches the Desktop: `plugins.manage list`
returns each plugin's schema with the current `plugins.entries.<id>.settings`
values (`settings_schema`), and a new `settings` action writes edits through
`hermes_cli.plugins_state.save_plugin_setting` — the writer extracted from
`PluginContext.set_config`, so the plugin, the CLI and the Desktop share one
config path, one lock and the same managed-install / managed-key refusals.
The Plugins tab grows a gear per plugin with a schema; the inline form is
table-driven (`FIELD_CONTROLS` / `INITIAL_TEXT` / `COERCE` keyed on the wire
field type) for string / number / boolean / enum / json / secret. Secrets are
declared with `type: secret`: the row carries only the `.env` name and a
presence flag, the client writes the value through the existing `PUT /api/env`
credential route, and the RPC refuses secret keys so nothing lands in
config.yaml.
Contracts regenerated; docs gain a "Settings form in the Desktop" section.
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.
- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
post_api_request payload shape plus `aux_task`, `api_request_id`
(`aux-...`, shared by every attempt of one logical call), `retry_count`,
`streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
turn is in flight; fail-open (a raising/hung subscriber is logged and
the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
attempt shares (_relay_sync_completion / _relay_async_completion /
_relay_sync_stream) run under the hook pair — retries and fallbacks
included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
aux_task and no api_request events; raising subscriber never breaks
the call). First is red on origin/main.
Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.
Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Four error-isolation holes in the disk plugin door
(apps/desktop/src/contrib/runtime-loader.ts), found by a static+live audit of
the loader; none had an issue filed.
- A plugin whose module evaluation never settles (top-level `await` on a dead
host) hung `import()` forever and, through the scan's sequential loop and
its re-entrancy guard, froze every later plugin and all future scans until
restart. `import()` now races a 10 s deadline; the plugin errors on its own
row ("import timed out") and the scan continues.
- Timers and DOM listeners a plugin took out with bare globals survived
disable and every hot-reload. `ctx.setTimeout` / `ctx.setInterval` /
`ctx.addEventListener` are tracked with the plugin and torn down on
unload; the SDK doc says bare globals are not.
- Two folders exporting one plugin id silently last-wins: the second
disposed the first's registrations and each hot-reload flipped ownership.
The first (folder-name sorted, so deterministic) owns the id; the later
file errors on its own row ("duplicate id, already loaded from <path>").
- A save that no longer loads (syntax error, timeout, duplicate) left the old
incarnation's contributions and activate handle live beside the error
row, so the Plugins tab showed a broken file as "loaded" and could
re-enable stale code. The previous incarnation is unloaded and dropped.
Tests: one invariant per fix in runtime-loader.test.ts, all red on base
(the hang case red by timing out).
`~/.hermes/hooks/` auto-loads every valid `HOOK.yaml` + `handler.py` at
gateway startup with no `plugins.enabled` gate. That is the documented
contract since 3988c3c245 ("Implicit (dir trust)" in the comparison
table), but the plugins page's "disabled by default" promise read as if
it covered gateway hooks too (#37963). Maintainer ruling: keep implicit
dir trust, fix the docs.
- hooks.md: new "Trust model" section stating exactly what loads, when,
how, and that placing the files is the opt-in; comparison-table cell
links to it and the plugin-hooks consent cell now says
`plugins.enabled`.
- plugins.md: note scoping `plugins.enabled` away from gateway hooks.
- developer-guide/plugins: one sentence at the gateway-hook recipe.
- security.md: "Trusted-by-placement extension points" section
cross-linked from hooks.md.
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.
Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.
Fixes#61839
credit: @giggling-ginger #61995
A general-plugin context engine is one shared instance; agent init copied it per agent
with copy.deepcopy() only. Engines that hold a SQLite connection or lock (hermes-lcm)
already expose clone_for_agent() for exactly this, but it was never called, so every
init logged "could not be safely copied … falling back to built-in compressor" and the
engine was unusable through the plugin system.
ContextEngine grows clone_for_agent() (default: deepcopy, the previous behaviour) and
_select_context_engine calls it; the failure message now names the hook to override.
Docs: context-engine-plugin.md documents the per-agent clone contract.
Test change (existing on main): tests/agent/test_context_engine.py::
test_agent_init_source_deepcopies_singleton_not_aliases was a source-reading pin on the
literal `copy.deepcopy(_candidate)` line, which this fix intentionally replaces. It is
superseded by tests/agent/test_plugin_context_engine_clone.py, which drives the real
_select_context_engine seam and asserts the invariant it guarded (child update_model()
never mutates the shared singleton) plus the new clone_for_agent() path.
Fixes#99640
credit: @stephenschoettler #62374
credit: @686f6c61 #99677
ToolRegistry.dispatch spread every injected keyword (task_id, session_id,
user_task, parent_agent, ...) into the handler, so a third-party plugin tool
written as `def handle(args)` raised TypeError on every call. The plugin
contract (plugins/AGENTS.md) says optional kwargs are signature-inspected, not
forwarded unconditionally; hooks already do this in plugins_dispatch. Handlers
taking **kwargs still receive the full payload. Docs updated: the narrow
signature is supported, **kwargs opts into the whole context.
Fixes#68318
credit: @vveerrgg #22146 (slim redo; #68636 is the later duplicate)
The catalog page claimed admission's lint means a marketplace install 'cannot quietly rewire the app'; the lint is a handful of regexes and plugin.js runs in the app realm with the full window.hermesDesktop bridge. The user guide, catalog trust model and SDK security section now describe the real model: human review of a pinned SHA plus two tripwires (lint + loader import allowlist), no isolation.
Symptom: a model-provider plugin installed with `hermes plugins install`
(e.g. claude-subscription-directsdk) worked from a terminal but the Desktop
app failed every session build with "Unknown provider '<name>'".
Cause: `providers/__init__.py` scanned `$HERMES_HOME/plugins` exactly once
per process, under whichever profile home happened to be bound at the first
lookup, and the registry was global. The Desktop backend and the multiplex
gateway serve several profiles from one process, so any profile other than
the first-discovered one never saw its own plugins, and a plugin installed
while the process ran was invisible until a restart.
Change: bundled, pip and legacy providers stay process-wide; `$HERMES_HOME`
plugins load into a per-home layer keyed by `hermes_home_key()` at lookup
time (`get_provider_profile`, `list_providers`, `provider_source`). The
layer rescans when the plugin directories' mtimes change, so a fresh install
is found on the next lookup. The registration target is a ContextVar so two
turn threads scanning two homes cannot cross-register, and no lock is held
across plugin imports (a lock there could deadlock against a thread
mid-`import hermes_cli.auth`). Per-home module names let two profiles carry
the same plugin. `plugin_dev` reuses the module-name helper.
Live repro (tui_gateway `session.create` on a secondary profile whose config
selects a plugin installed only there): base -> agent_error "Unknown
provider 'fakeprov-b'"; fixed -> provider resolves and the build proceeds to
the plugin's own runtime check.
The five pages moved from docs/ into website/docs/developer-guide were
never added to sidebars.ts, so the site did not link them, and they kept
repo-relative links (../tests/..., ../scripts/..., ../website/docs/...)
that only resolved from the old docs/ location. Register them under a
Packaging & Releases category and point the repo-file links at GitHub;
package-management.md now links the moved completion note in place.
The config comment and both docs pages named six of the eight sources in
MACHINE_PACED_SOURCES; "tool" and "batch" were missing, so an operator
reading the docs could not predict the tier those sessions get.
The 1h Anthropic cache tier writes at 2x base (5m: 1.25x) and only pays off when
turns are more than five minutes apart. That is exactly the shape of an interactive
session a person parks and resumes, and exactly not the shape of a subagent, cron
run, one-shot or webhook that calls every few seconds and is gone. A single global
`cache_ttl` cannot be right for both, so operators leave it on 5m and pay a full
context re-write every time they come back to a CLI session after a coffee.
Measured on one install (2 days of per-call API logs, Claude via the Nous route):
63% of interactive cache-write tokens were cold re-writes after a 5-60 minute idle
gap; 1h would cut interactive write cost ~42% while costing ~49% more on subagents
and ~23% more on cron. `auto` resolves once per session from the session source
(`_session_source_for_agent`): 1h for cli/tui/desktop/messaging platforms, 5m for
subagent, cron, oneshot, webhook, kanban, api. Auxiliary/stub calls keep 5m; the
delegate_tool child clamp (#104168) still applies. Default stays "5m".
Merge origin/main at 8e806ae1b2. Keep native helper compilation in
prepareDesktopNativeDependencies and keep bundling/beforePack consume-only.
Bind helper sources and headers into preparation identities and cache keys;
copy admitted executable resources beside node_modules and preserve signing
semantics in product freshness checks.
Verified desktop typecheck, focused native/packaging/UI and gateway/cache
tests, and the real Linux preparation/copy/Xvfb execution path. Incoming
upstream anti-slop findings remain unchanged; no baseline was raised.
Every line still promising that `gateway.multiplex_profiles: false` keeps
per-profile gateways, or documenting the deleted `migrate --standalone`
rollback, now describes the shipped behaviour: one gateway per host serves
every profile; `false` no-ops with a warning; `hermes update` folds a fleet
unless a real boundary (different UNIX user, HERMES_HOME outside profiles/)
holds; a blocked fleet keeps `--force` per profile; re-running
`migrate --multiplex` converges a half-migrated host; a manifest on disk is the
resume record, never a rollback. Adds the new "No new per-profile gateways"
section with the exact refusal `hermes -p <name> gateway install` prints.
Files: website/docs/user-guide/multi-profile-gateways.md,
website/docs/reference/cli-commands.md (migrate row),
website/docs/developer-guide/multiplexing-gateway.md (eager activation no
longer honours `false`), hermes_cli/AGENTS.md (migrate + refusal seams).
Completes #118273 (docs item). Part of #109417
`gateway/host_attach.py::decide()` treated a host owner that answers `multiplex: False`
to the rescan — another profile's standalone gateway — the same as a multiplexer whose
roster excludes us, and issued the permanent REFUSE. On the supervised path that is
exit 78, which `hermes_cli/stderr_timestamp.py` maps to 0 so launchd's
`KeepAlive.SuccessfulExit=false` parks the unit. On a one-process-per-profile fleet
(the topology `multi-profile-gateways.md` documents and `multiplex_profiles: false`
promises to keep) every launchd gateway except the first to claim the host lock was
parked at boot, silently; which ones survived was a boot race against the owner's
record landing.
`request_serve_profile` now returns the owner flagged `standalone` instead of `None`
for a `multiplex: False` answer, and `decide()` turns that into START — this profile
runs its own gateway beside the owner, as before #118097; the host-lock claim still
logs the two-gateway topology and `gateway migrate --multiplex` stays the converge
path. REFUSE is unchanged for a multiplexing owner that excludes the profile.
A/B against a REAL owner process (real host lock, record and control socket answering
the rescan): base → `refuse`, exit 78 → launchd 0 (parked); fix → `start`. Negative
control (owner answers `multiplex: True`, roster excludes the profile): `refuse` on
both.
Field report: debug share 5024e996 (11 launchd profile gateways, 0.21.3 ea0c2b82 —
only 2 of 11 came back after `hermes update`), discussed on #118097.
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.
- ATTACH now requires a LIVE `identify` answer. The claim-time record is
published with NO served set (the runner settles multiplex a moment later),
and `host_gateway()` reports `served_known=False` when nothing answers. An
owner whose served set is unknown yields a TRANSIENT refusal, never an
attach: previously `default`'s claim published "default,other" before its
socket bound, `other`'s systemd unit read that as "I am served", exited 78,
and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
instead of forcing `multiplex=True`, so a standalone gateway stops claiming
the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
`--force` into `_host_attach_or_none`. Both previously exited in the guard
before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
config refusal every supervisor parks on; "someone else serves me right now"
is a runtime observation that ends when that process does. No unit files
change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
and re-enters with `replace=True`, so it can no longer attach to the corpse
it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
`st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
`gateway status`/doctor across N profiles pays one probe, not N.
Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:
- `gateway run` for a profile the host gateway already serves ATTACHES: print
its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
to re-scan `profiles/` (control socket) and attach once the answer includes
it. Refuse only when the host gateway cannot be made to serve it. Under a
service supervisor the attach exits 78 instead of 0 so a redundant unit is
parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
the owner's control socket, so it no longer depends on the claim ordering in
start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
process: they restart the host multiplexer and preserve its served set. A
secondary still running its own gateway is reported with the
`gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
argv: a host singleton runs bare/default argv and can never prove it serves
profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
multiplexer is whichever profile launched the one host process.
Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
main folded the auto-archive housekeeping into per-profile state.db maintenance; the plugin
update-check chore stays. The consolidated-away TestWorkdirParallelPool stays removed. Two new
utf-8 reads in gateway/host_rendezvous.py read utf-8-sig.
Activation is not purely stricter for the LAUNCH tenant: inside the launch scope its own
.env wins over the frozen boot env, so a key present in both reads differently before and
after activation. Also documents the explicit multiplex_profiles: false escape hatch and
the fail-closed unresolved-profile case.
One host process serves every profile, so every execution point must bind the
profile it is acting FOR. These six ran unscoped (or bound only part of a
scope) and resolved get_hermes_home()/credentials against the LAUNCH profile:
- tui_gateway/session_reaper: the idle-reaper and exit-flush transcript writes
now enter the SESSION's profile scope, the same chokepoint _finalize_session
already binds. A served profile's transcript was landing in the launch home.
- gateway/run: MCP shutdown tears down per served profile inside that
profile's scope (mirrors startup discovery and the reconcile chore), with a
trailing wildcard pass under the launch profile's own scope.
- gateway/run_profile_reconcile: _unserve_profile's adapter teardown, agent
eviction and state/memory handle release now run inside the deleted
profile's scope.
- gateway/run_adapters + run_goals + run_notifications: a body with no routed
profile no longer means "no scope". launch_profile_scope_if_multiplexed()
binds the launch profile once the process multiplexes; before activation it
is still literally a nullcontext, so single-profile hosts are unchanged.
- hermes_cli/kanban_db_dispatch: one _worker_profile_scope helper binds the
assignee's secret AND terminal scope for toolset resolution and the spawn-env
build, unconditionally instead of only under multiplex.
- hermes_cli/web_server: `hermes serve` activates multi-profile hosting at boot
when the host has more than one servable profile home, instead of lazily on
the first ?profile= request after earlier work already ran unscoped.
Secret scope is never widened: a non-launch home resolves from its own .env and
sources only; the launch home keeps its existing env-over-.env precedence.
Under one multiplexing ticker, every non-launch profile's in-flight claim
leaked on every run. The cron home scope is a ContextVar: the claim is taken on
the ticker thread inside `_profile_cron_scope`, but the pool worker's `finally`
sits outside `ctx.run`, where the worker thread resolves the LAUNCH home — so
the discard missed the real key. The job then skipped a fire window until the
force-release backstop swept it, and the shutdown drain plus
`hermes.cron.jobs.running` saw phantom work.
`_submit_with_guard` now captures the registering home and passes it to
`release_running_job(job_id, home=...)` on every release path.
Also in this pass:
- Stand down for a profile that runs its OWN gateway (`_cron_profile_gate`,
the gate the serve/Desktop ticker already passes). On a host pinned to
per-profile gateways both processes raced that profile's tick lock, and when
the launch gateway won, delivery went through SharedRouteAdapters/fail-closed
instead of the profile's live adapters. The gate compares the liveness PID
against `os.getpid()`: this process holds the launch `gateway.pid` and
publishes every served profile in `served_profiles`, so a bare liveness answer
would have stood cron down host-wide.
- Job-liveness consumers ask `is_job_running(job_id, home=...)` instead of the
host-wide bare-id union, which let profile A's running `daily-brief` report
profile B's idle one as running and keep B's stale one-shot alive.
- `register_ticked_homes` reaps the parallel pools of homes that leave the
ticked set; pools lived until `atexit`, so every home ever ticked kept a
ThreadPoolExecutor and its worker threads.
- `mark_running_jobs_interrupted` reads the real home Path from
`_inflight_home_path` instead of rebuilding it from the normcased key half.
- The ownership test imports `cron.scheduler_ownership` per test, so the file
fails behaviourally on base instead of as a collection error.
The cron ticker multiplexes N profiles from one process while its ownership
predicates and its in-flight bookkeeping still assumed one profile per process.
- `_should_yield_tick_to_fresh_gateway` asked a process-global boolean
(`owns_gateway_runtime_lock`) and a launch-home lock probe, so one answer
covered every profile ticked. It now asks `scheduler_ownership`:
`owns_cron_tick_for(home)` (this process is the host gateway AND ticks that
home) and `live_gateway_ticking(home)` (another live host gateway whose
published served set covers that home).
- In-flight state (`_running_job_ids`, `_running_since`, `_running_futures`,
`_running_allowance_s`, `_running_worker_pids`, `_running_fire_owners`,
`_restart_safe_waiter_job_ids`, `_interrupted_job_ids`) is keyed by
`_inflight_key(job_id)` = `(home key, job id)`; two profiles carrying a
`daily-brief` no longer read as one job. The public accessors still report
the host-wide union of bare job ids for the shutdown drain.
- The parallel worker pool is keyed by home: `cron.max_parallel_jobs` is a
per-profile key, and the single global pool was sized by whichever profile
ticked first and torn down by the next one.
- `gateway/run.py` no longer gates the cron tick set on
`gateway.multiplex_profiles`: that flag gates adapters, and with it off every
non-launch profile's jobs sat in a store no ticker visited.
main extracted the launchd backend and the setup wizard out of hermes_cli/gateway.py. The
eight launchd functions pm-clean had changed are ported into gateway_launchd.py in its
_gw() style: runtime_command/installation_command for the gateway argv (no VIRTUAL_ENV in
the plist, XML-escaped args), _prepare_service_launcher before every plist write,
utf-8-sig plist reads, and launchd_restart's refresh-first + bounded bootstrap revival.
The systemd service-unit cluster stays in the facade (no gateway_service_unit sibling).
backup.py takes main's browser_profiles backup-only exclusion on our profile_root_entry
shape. main.ts takes main's attach-first backend block with our explicit types.
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:
* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
loop) and the success marker; with no success marker on disk it printed
"✓ Gateway is running — cron jobs will fire automatically", and with a stale one
it pointed at "Check the gateway log" instead of naming the cause. The persisted
`CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
and reported as "Gateway is running STALE code — fires NOTHING", with both
revisions and the restart command. A fresh heartbeat with a recorded error and no
success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
to the existing drain-first `request_restart` path (SIGUSR1) via
`hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
contract the restart phase already uses for unmapped manual gateways. The drain
budget computation is shared (`_gateway_drain_budget`).
The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).
Fixes#117275
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole
rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the
"protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every
compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B
timed out before compaction ever changed anything, and protect_last_n read as an uncompressed
tail rather than a minimum.
TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk /
pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool
groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of
32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%).
Probe (12 tool-heavy turns, 49 rows, 12.8K tokens):
ctx 8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows; window [4,10) -> [4,42)
ctx 16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34)
ctx 32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22)
ctx 128K: identical before/after
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
DecodeText is a published plugin-SDK export; flipping its `loop` default
to false silently changes third-party plugin visuals, so the SDK doc says
so where the component is listed.
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).
The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.
New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped
code_sha is the on-disk revision; a lock held by an equally stale process never
counts. Placed under "Gateway Integration" so it does not collide with the
"Stale-code yield" section #117501 adds under "Locking".
The provider-agnostic half of PR #105863, so a CLI-driven subscription provider can ship as a
standalone `kind: model-provider` plugin instead of a bundled one:
- ProviderProfile: `native_reasoning_details_type`, `model_aliases`, `get_model_context_length`,
`get_usage_cost`, `setup_status`, `discover_models` hooks (all default None / no-op).
- Chat Completions transport: provider-native `reasoning_details` carriers follow only their
declaring profile; standard records still replay on OpenRouter-style routes, strict routes
drop the field wholesale (#70233). Relay/stream accumulate `delta.reasoning_details` verbatim.
- `hermes model`: the generic plugin flow gates an external-process row on the CLI's own login
status (inline `login_command` on a TTY), offers `discover_models()` rows with per-row notes,
and never writes config when the executable is missing.
- `/model` and the pickers: process providers list their live catalog merged with the pinned
one, declared aliases/ids resolve inside the provider, and validation accepts a listed id
without probing `process://`.
- Delegation keeps the selected external-process provider and protocol for the child.
- Model metadata / usage pricing consult the profile's bound and cost hooks first.
- Desktop: `[1m]` renders as a "1M" tag and hyphenated Anthropic versions read "Haiku 4.5".
The bespoke `_model_flow_external_process` and hard-coded `hermes_cli/main.py` paths from the
PR were dropped in favour of main's `_model_flow_plugin_provider`.
Co-authored-by: unsupportedpastels <unsupportedpastels@users.noreply.github.com>
The same rename that left the two reference pages stale (cronjob->cronjob_manage,
todo->todo_list, process->process_manage) left the guides and developer docs
referring to the old tool names; a reader following them gets "no such tool".
Toolset names (`cronjob`, `todo`) are unchanged and left alone.
The 19 cherry-picked tests covered each refusal branch separately. One
A -> B -> A test per adapter over two real homes now proves the whole
contract at the production entry (build_credential / _cached_client): the
launch profile keeps its own credential, the cred-less served profile is
refused before the SDK chain (or boto3) is touched, the launch profile is
unaffected afterwards, and the standalone run keeps today's ambient chain.
Docs: the Azure guide and the multiplexing design page name the refusal.
Superseded #116370 (@JoaoMarcos44) proposed the same mechanism.
A `$HERMES_HOME/plugins/model-providers/<name>/` plugin re-registering a
bundled provider (stepfun with a regional base_url, gmi at a staging host)
wins in `providers._REGISTRY` — register_provider() is last-writer-wins and
the plugin guide promises exactly this — but the runtime reads its endpoint
from `hermes_cli.auth.PROVIDER_REGISTRY`, whose mirror loop skipped every
name already present, so inference kept going to the built-in URL (#48450).
The mirror now applies one explicit precedence rule: when a row core wrote
(built-in or plugin-mirrored) belongs to a name whose profile
currently registered came from a USER plugin,
the row's profile-derived fields are rewritten in place (inference_base_url;
api_key_env_vars / base_url_env_var on api-key rows when the profile declares
env_vars). `providers` records the discovery source per registration
(`provider_source()`), because a bundled profile must never rewrite a
built-in row: several bundled profiles omit the row's `*_BASE_URL` env var
and one differs in auth_type, so an unconditional "profile wins" would have
changed built-in behaviour. With no user plugin PROVIDER_REGISTRY is
byte-identical before/after (78 rows probed). copilot/kimi/zai keep their
bespoke resolution via the existing skip set.
Co-authored-by: xiaoxinova <xiaoxinova@users.noreply.github.com>
"database is locked (another Hermes process held the state.db write lock for
over 60s)" identified the victim only. The open-descriptor scan cannot single
out the writer because every Hermes process (gateway, CLI sessions, worktree
agents, cron) has the DB open, so an operator hit repeatedly by
session_persistence_failed:locked had nothing to act on.
SQLite's unix VFS takes fcntl byte-range locks whose offset encodes the lock
kind (state.db-shm byte 120 = WAL write, 121 = checkpoint; the pending-byte page
on state.db = PENDING/RESERVED), and the kernel exports them with the owning pid
in /proc/locks. hermes_state_lockowners reads that table at the moment the
patience budget runs out and logs one WARNING per write-class holder with
describe_holder_pid()'s argv summary, for both the transcript write path and
open+init lock patience. The holder stays out of the exception text on purpose:
classify_persistence_error() buckets by phrase and a holder argv such as a
worktree named fix-corrupt-db would flip the bucket.
Docs: the Write Contention section still described attempt-counted retries
(_WRITE_MAX_RETRIES = 15); updated to the time budgets in force and the new log line.
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.