782886f87dce91c8af3e85f7bf4e54ff24602168
3483 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
061195fac1 |
FLEET: one host-scoped update-restart obligation, one restart per host
One host runs one multiplexing gateway, but the update pipeline still treated the pull->restart obligation, enumerated units, recovery payloads and the planned-restart notice as per-profile. Two profiles updating meant two outages of the same process, and a served profile's channels were never told. - hermes_cli/update_host_obligation.py: new host-scoped obligation record in gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the unit->live-MainPID collapse rule. The legacy per-home marker stays readable and clearable so an in-flight obligation is still discharged. - update_cmd_fleet: arm/clear/read the host record; the catch-up restart is idempotent per host (a completed restart onto the checkout SHA is never repeated); leftover per-profile units resolving to one MainPID restart once. - update_restart_recovery: payload profiles served by one host process are one restart target, reported under "covered". - gateway notices: owed targets and the online notice span every served profile's home channels; the marker survives until each was reached. |
||
|
|
274bc7b8f6 | docs(desktop): where the Linux Chromium log and minidumps live | ||
|
|
a1398a6f67 |
fix: doctor, cron status, claw and gateway status report the real host topology
Multiplex-only (Teknium ruling) runs exactly ONE `hermes gateway run` per host,
multiplexing every profile. Four reporting surfaces still asked a per-PROFILE
process question ("does MY profile own a gateway process?"), so a SERVED profile
answered "no" and the output lied:
* `hermes -p served cron status` printed "Gateway is not running" and told the
user to run `hermes gateway install` / `gateway run` for that profile — i.e.
to start a SECOND host process, which the architecture forbids.
* `hermes doctor` reported "Per-profile gateways: up/total" from s6 slots and
checked systemd linger against the CURRENT profile's unit, so a doctor run
under a served profile skipped the check entirely.
* `hermes claw`'s destructive-action warning was gated on `get_running_pid()`,
which returns None for a served profile — the token-conflict warning was
SILENTLY SKIPPED (a real safety hole).
* `gateway.status.multiplexer_liveness_for_profile()` returned None for the
default home by construction, so `default` could never be reported as SERVED.
* `hermes doctor`'s state.db holder + WAL wording implied one gateway per
profile.
New `gateway/host_topology.py` resolves "which single process owns the gateway
role on this host, and which profiles does it serve" once, from the host
rendezvous record (`gateway/host_rendezvous.py`), falling back to the default
home's recorded `served_profiles` for a gateway that predates the record.
`default` is just another served profile there.
Root cause: every surface derived gateway identity from per-profile artifacts
(argv `-p <name>`, `gateway.pid`, the active profile's systemd unit, s6 slot
counts) instead of the host record that actually names the owner.
|
||
|
|
a10620a669 |
fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.
- ATTACH now requires a LIVE `identify` answer. The claim-time record is
published with NO served set (the runner settles multiplex a moment later),
and `host_gateway()` reports `served_known=False` when nothing answers. An
owner whose served set is unknown yields a TRANSIENT refusal, never an
attach: previously `default`'s claim published "default,other" before its
socket bound, `other`'s systemd unit read that as "I am served", exited 78,
and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
instead of forcing `multiplex=True`, so a standalone gateway stops claiming
the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
`--force` into `_host_attach_or_none`. Both previously exited in the guard
before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
config refusal every supervisor parks on; "someone else serves me right now"
is a runtime observation that ends when that process does. No unit files
change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
and re-enters with `replace=True`, so it can no longer attach to the corpse
it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
`st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
`gateway status`/doctor across N profiles pays one probe, not N.
Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
|
||
|
|
a20a88398f |
gateway: lifecycle verbs mean "the one host multiplexer"
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this profile's gateway". Under the multiplex-only ruling there is exactly ONE gateway process per host, so they now target that process: - `gateway run` for a profile the host gateway already serves ATTACHES: print its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner to re-scan `profiles/` (control socket) and attach once the answer includes it. Refuse only when the host gateway cannot be made to serve it. Under a service supervisor the attach exits 78 instead of 0 so a redundant unit is parked, not restart-looped. - The attach channel is reachable BEFORE the PID claim: the decision reads the landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to the owner's control socket, so it no longer depends on the claim ordering in start_gateway. - `start --all` / `restart --all` no longer SIGTERM every gateway-looking process: they restart the host multiplexer and preserve its served set. A secondary still running its own gateway is reported with the `gateway migrate --multiplex` one-liner, never killed. - Ownership is decided by the live served set (record + control socket), not by argv: a host singleton runs bare/default argv and can never prove it serves profile X, which rejected every secondary. - The implicit-multiplex verdict no longer requires the DEFAULT profile: the multiplexer is whichever profile launched the one host process. Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the host record is shared per OS user by design, so one test that boots a gateway made every other file's lifecycle code attach to it. |
||
|
|
5d63d3e5c5 |
fix(gateway): launchd gateway reaches the local network on macOS (#71206, #57812)
macOS Local Network Privacy attributes a socket to the executable launchd spawned for the job. A bare venv Python has no application identity and is not platform-entitled, so every LAN connect from the supervised gateway failed with errno 65 (No route to host) while the same URL worked from Terminal — and nehelper never showed a prompt that could grant it, so a headless Mac had no way out. Run the job through /usr/bin/osascript: `do shell script "exec …"` spawns its child as osascript-responsible — an Apple platform binary — and the child is exempt. Verified live on macOS 26.3.1 with a fresh ad-hoc binary under the gui launchd domain: bare → 65, `/bin/sh -c exec` → 65, `/usr/bin/time` → 65, osascript → reachable; an ad-hoc-signed helper .app with NSLocalNetworkUsageDescription (the #115196 approach) stays denied with no prompt even after lsregister, matching the #57812 dead-end table. `do shell script` buffers the child's output until exit, so the shell command redirects stdout/stderr to the log files the plist already routes; `exec` keeps the gateway in the job's process group, so `launchctl bootout` still delivers SIGTERM to it (live: stop → no orphan, restart → exit 75 → KeepAlive respawn, stop → parked). The existing plist-staleness refresh picks the new definition up on `hermes gateway install`/`start`. Mechanism proposed by @leewaiho in #57812. Co-authored-by: leewaiho <18321182+leewaiho@users.noreply.github.com> |
||
|
|
db1f3f4564 |
fix(computer-use): doctor names the stale TCC row on the health_report path too
The first cut appended the recovery only in `_tcc_row`, which is built by the 0.10 fallback probes. Drivers that serve health_report (0.22+, the ones a stale row actually bites) render their own tcc_* rows untouched, so `doctor` never showed it. Apply the hint at the report seam like the display-count guard so both paths get it. Also: one CUA_DRIVER_BUNDLE_ID (permissions.py is the leaf; the daemon imports it), the CLI derives the field set from the same table, and the unverified "--upgrade repairs stale rows" claim is dropped from hint + docs (the user hitting this is already on a current driver). |
||
|
|
42ae8b4410 |
fix(computer-use): name the stale TCC row when CuaDriver shows ON but is denied
`hermes computer-use permissions status` and the doctor's fallback `tcc_*` rows told a user whose System Settings toggle already showed CuaDriver ON to "grant it in System Settings" — the one step that cannot help. macOS keys the TCC row to the app's code-signing requirement; a row written for an earlier CuaDriver build stops matching after a driver update and flipping the toggle does not rewrite it (trycua/cua#3170, repaired by the cua-driver >= 0.22 installer on update). The daemon then reports Accessibility / Screen Recording false while the pane shows ON, and every click silently no-ops (#99732). `permissions.stale_tcc_grant_hint(*missing)` renders the reset for exactly the missing services (`tccutil reset Accessibility|ScreenCapture com.trycua.driver`, then `hermes computer-use permissions grant`); both surfaces append it only when a grant is actually reported False. Docs: the computer-use page still said bounded/unrestricted daemons run "under the Hermes host identity" — they launch through CuaDriver.app since #95381 — and gains the stale-row troubleshooting entry. |
||
|
|
bfcd209906 |
docs(multiplexing): record the config escape hatch and the .env-over-frozen-env precedence
Activation is not purely stricter for the LAUNCH tenant: inside the launch scope its own .env wins over the frozen boot env, so a key present in both reads differently before and after activation. Also documents the explicit multiplex_profiles: false escape hatch and the fail-closed unresolved-profile case. |
||
|
|
c718267ef3 |
fix(multiplex): close the profile-scope holes the scope machinery misses
One host process serves every profile, so every execution point must bind the profile it is acting FOR. These six ran unscoped (or bound only part of a scope) and resolved get_hermes_home()/credentials against the LAUNCH profile: - tui_gateway/session_reaper: the idle-reaper and exit-flush transcript writes now enter the SESSION's profile scope, the same chokepoint _finalize_session already binds. A served profile's transcript was landing in the launch home. - gateway/run: MCP shutdown tears down per served profile inside that profile's scope (mirrors startup discovery and the reconcile chore), with a trailing wildcard pass under the launch profile's own scope. - gateway/run_profile_reconcile: _unserve_profile's adapter teardown, agent eviction and state/memory handle release now run inside the deleted profile's scope. - gateway/run_adapters + run_goals + run_notifications: a body with no routed profile no longer means "no scope". launch_profile_scope_if_multiplexed() binds the launch profile once the process multiplexes; before activation it is still literally a nullcontext, so single-profile hosts are unchanged. - hermes_cli/kanban_db_dispatch: one _worker_profile_scope helper binds the assignee's secret AND terminal scope for toolset resolution and the spawn-env build, unconditionally instead of only under multiplex. - hermes_cli/web_server: `hermes serve` activates multi-profile hosting at boot when the host has more than one servable profile home, instead of lazily on the first ?profile= request after earlier work already ran unscoped. Secret scope is never widened: a non-launch home resolves from its own .env and sources only; the launch home keeps its existing env-over-.env precedence. |
||
|
|
9ecd22e5da |
fix(cron): release a profile's in-flight claim under the key it registered with
Under one multiplexing ticker, every non-launch profile's in-flight claim leaked on every run. The cron home scope is a ContextVar: the claim is taken on the ticker thread inside `_profile_cron_scope`, but the pool worker's `finally` sits outside `ctx.run`, where the worker thread resolves the LAUNCH home — so the discard missed the real key. The job then skipped a fire window until the force-release backstop swept it, and the shutdown drain plus `hermes.cron.jobs.running` saw phantom work. `_submit_with_guard` now captures the registering home and passes it to `release_running_job(job_id, home=...)` on every release path. Also in this pass: - Stand down for a profile that runs its OWN gateway (`_cron_profile_gate`, the gate the serve/Desktop ticker already passes). On a host pinned to per-profile gateways both processes raced that profile's tick lock, and when the launch gateway won, delivery went through SharedRouteAdapters/fail-closed instead of the profile's live adapters. The gate compares the liveness PID against `os.getpid()`: this process holds the launch `gateway.pid` and publishes every served profile in `served_profiles`, so a bare liveness answer would have stood cron down host-wide. - Job-liveness consumers ask `is_job_running(job_id, home=...)` instead of the host-wide bare-id union, which let profile A's running `daily-brief` report profile B's idle one as running and keep B's stale one-shot alive. - `register_ticked_homes` reaps the parallel pools of homes that leave the ticked set; pools lived until `atexit`, so every home ever ticked kept a ThreadPoolExecutor and its worker threads. - `mark_running_jobs_interrupted` reads the real home Path from `_inflight_home_path` instead of rebuilding it from the normcased key half. - The ownership test imports `cron.scheduler_ownership` per test, so the file fails behaviourally on base instead of as a collection error. |
||
|
|
cb647a018f |
fix(cron): one host ticker owns every profile's cron, per profile
The cron ticker multiplexes N profiles from one process while its ownership predicates and its in-flight bookkeeping still assumed one profile per process. - `_should_yield_tick_to_fresh_gateway` asked a process-global boolean (`owns_gateway_runtime_lock`) and a launch-home lock probe, so one answer covered every profile ticked. It now asks `scheduler_ownership`: `owns_cron_tick_for(home)` (this process is the host gateway AND ticks that home) and `live_gateway_ticking(home)` (another live host gateway whose published served set covers that home). - In-flight state (`_running_job_ids`, `_running_since`, `_running_futures`, `_running_allowance_s`, `_running_worker_pids`, `_running_fire_owners`, `_restart_safe_waiter_job_ids`, `_interrupted_job_ids`) is keyed by `_inflight_key(job_id)` = `(home key, job id)`; two profiles carrying a `daily-brief` no longer read as one job. The public accessors still report the host-wide union of bare job ids for the shutdown drain. - The parallel worker pool is keyed by home: `cron.max_parallel_jobs` is a per-profile key, and the single global pool was sized by whichever profile ticked first and torn down by the next one. - `gateway/run.py` no longer gates the cron tick set on `gateway.multiplex_profiles`: that flag gates adapters, and with it off every non-launch profile's jobs sat in a store no ticker visited. |
||
|
|
3917c7dd6a |
fix: RoomLink catalog names the served profile and reads its policy from that profile
`catalog_mapping` took an optional `target_profile` and fell back to the launch process's `HERMES_PROFILE` (then "default"), and `execution_policy_mapping` loaded whatever gateway home was active while labelling the result with the requested `target_profile`. On a multiplexed / Desktop-spawned backend a remote Bot invited for profile B could receive a catalog signed for the launch profile, or B's name over A's approvals/turn-limit/toolsets. - `target_profile` is a required keyword on `catalog_mapping`; an empty or mismatched profile fails loudly naming the field, no env fallback. - `execution_policy_mapping(config=None)` resolves the SERVED profile's own config: a no-op when it is the active home (callers that already scope are unchanged), else `_profile_runtime_scope(get_profile_dir(profile))`; a profile that does not exist is refused instead of silently reading the launch config. - `issue_room_grant`'s default digest inherits the served-profile resolution. - Tests: 27 call sites now name the profile; two invariants red on base. - Docs: multi-profile resolution table row for the RoomLink catalog/policy. Fixes #116900 |
||
|
|
8863b36fd6 |
fix(cron): stale-code yield reads as an outage in cron status; hermes update restarts proven-stale gateways
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:
* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
loop) and the success marker; with no success marker on disk it printed
"✓ Gateway is running — cron jobs will fire automatically", and with a stale one
it pointed at "Check the gateway log" instead of naming the cause. The persisted
`CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
and reported as "Gateway is running STALE code — fires NOTHING", with both
revisions and the restart command. A fresh heartbeat with a recorded error and no
success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
to the existing drain-first `request_restart` path (SIGUSR1) via
`hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
contract the restart phase already uses for unmapped manual gateways. The drain
budget computation is shared (`_gateway_drain_budget`).
The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of
|
||
|
|
661d73f06b |
test+docs: trim doctor exit-status tests to two invariants; document the exit code
Keep the in-process return-code matrix (issues / manual / --fix full and partial repair) and the real-process exit-status check; drop the exception passthrough, --ack and --live cases, which pin behaviour this change does not touch. Document 0/1 exit status under `hermes doctor` in the CLI reference. |
||
|
|
67208caac3 |
fix(memory): replace states and surfaces the whole-entry contract instead of truncating silently (#117952, salvage #109177)
A partial-entry replace (old_text = a span inside an entry, content = its
replacement) silently overwrote the WHOLE entry: every clause outside the
match was destroyed with success=true, most damagingly through the
write-approval replay path (/memory approve -> apply_memory_pending ->
apply_batch) that exists precisely so background agents can stage writes for
human review (#117952, dup of #59184).
The cluster's repeated fix attempt (#66321, #61357, four closed predecessors)
flipped replace to naive span-splicing (entry.replace(old_text, new, 1)).
That corrupts the canonical identifier-style call (replace(old_text="3.11",
content="Python 3.12 project") -> "Python Python 3.12 project project"),
breaks the background-review refine consumer, and split-brains external
memory mirrors, which receive the raw op content, not a merged entry.
The defect is the CONTRACT, not the mechanism: the schema invited patch-style
calls ("mirrors old_text (patch-tool shape)") while the store commits
whole entries. This pins whole-entry semantics end to end and makes the
overwrite visible instead of silent:
- schema + recovery error text + memory_tool docstring now say content is the
COMPLETE new entry and old_text only locates it (website docs updated in the
same PR)
- a successful replace surfaces replaced_entry / replaced_entries (full text
that was overwritten) on single, batch, and approval-replay paths, so the
caller can re-add lost clauses instead of discovering them gone later
- _mutate gains an optional extra payload field; _apply_batch_op returns the
replaced entry text alongside the error
Live repro on origin/main
|
||
|
|
2520eb7c0d | docs(backup): list the browser profile dirs hermes backup excludes | ||
|
|
9a1f06293d |
feat(mcp): make the discovery connect cap configurable (mcp.discovery_concurrency, default 4, 0 = unlimited)
Builds on Halldrix's #117674 (flat cap of 3): the cap now reads config.yaml `mcp.discovery_concurrency` (registered in DEFAULT_CONFIG under the existing `mcp:` section, default 4 per the #117373 ruling). 0 disables the semaphore and restores the unbounded gather; a non-integer / negative value warns and uses the default instead of silently running unbounded. The wave-scaled pass timeout uses the effective cap. Tests trimmed to two invariants driven through `_run_discovery_pass` (the production entry): configured cap bounds in-flight connects while every server still connects and 0 means unlimited; pass timeout stays under the lock waiter budget. Docs: mcp.md runtime section + multi-profile-gateways.md. Fixes #117373 |
||
|
|
2463550c97 |
feat(website): plugin pages render the README from the pinned commit by default
The README section was gated on `readme: true` in the catalog YAML and no entry set it, so all 222 plugin pages shipped without one. READMEs now render for every GitHub/GitLab entry (fetched at the reviewed sha, subdir first then repo root, common casings and docs/README.md as fallbacks); `readme: false` opts an entry out. Build proof: 222/222 READMEs rendered. |
||
|
|
0fd811e687 |
refactor(gateway): drop the command_denied_message knob and silent-denial mode
Maintainer ruling on #117217: a configurable refusal text (and the ""-means-silent mode) is feature creep. The refusal text returned by _check_slash_access is now byte-identical to origin/main again; only the /help and /commands catalog gating for non-admin callers remains in this PR. Removes the run_busy.py override hunk, the two knob tests, and the "Customizing command refusals" docs section. |
||
|
|
2978c23acc |
fix(gateway): /help and /commands show a gated non-admin only the commands they can run
Both catalog handlers called the shared executor without the caller's slash-access policy, so a non-admin under allow_admin_from saw every admin-only command and then hit the refusal on each one. The handlers now pass the policy's runnable set (the /help+/whoami floor plus user_allowed_commands) as `allowed_commands`; gateway_help_lines filters on it and skill commands are hidden for gated users. Admins and ungated scopes are unchanged (option absent -> full catalog). Part of #117217 |
||
|
|
72ba7ae3d6 |
fix(gateway): honor configured slash-command refusal
(cherry picked from commit 8f052e0db1b33a5472473484d0a85aa72d476f73) |
||
|
|
d90c3ad40a |
fix(desktop): group-chat member failure row names the error's first line (#117366)
A member turn that failed without a typed gateway reason rendered as a bare "X hit an error" — a stopped backend, a dead IPC bridge and a provider refusal all looked identical and gave the user nothing to act on. groupFailureReason now falls back to the error message's first non-empty line (trimmed, secret-shaped spans redacted, capped at 200 chars), so the Activity row reads "X hit an error — <cause>". Typed data.reason and the slot-wait classification keep precedence; the roster badge still classifies from the same text. The plugin fence keeps the Electron-side redactSecrets out of reach, so the same shapes (bearer header, ?token=/?api_key= params, vendor key prefixes, user@host:password) live next to the label. Fixes #117366 |
||
|
|
b7803a1763 |
fix(compression): cap the protected tail at 20% of the context window
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the "protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B timed out before compaction ever changed anything, and protect_last_n read as an uncompressed tail rather than a minimum. TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk / pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of 32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%). Probe (12 tool-heavy turns, 49 rows, 12.8K tokens): ctx 8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows; window [4,10) -> [4,42) ctx 16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34) ctx 32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22) ctx 128K: identical before/after |
||
|
|
dec236b214 |
feat(desktop): Uninstall for standalone desktop plugins via Electron IPC
Rows whose only half is a folder in <HERMES_HOME>/desktop-plugins (no agent package behind it) get the same trash button + destructive confirm as the agent rows. The renderer names the FOLDER, never a path; Electron (`hermes:plugin:removeDesktop` -> desktop-plugin-remove.ts) resolves it under the app-level root, refuses anything that is not a single contained segment, refuses unified-package halves (the reconcile would re-copy them; the agent uninstall prunes them), removes a symlinked folder as the link, and deletes the tree. The loader then retires the entry (unload, drop rows, stop watch) so the pane/commands vanish without waiting for the next scan. Why an IPC: these plugins have no agent half, so `plugins.manage remove` on the gateway cannot reach them — they live on this computer, in every connection mode. |
||
|
|
43fe9c1773 | docs(desktop): describe the Plugins hub Uninstall button | ||
|
|
afc3b7c6f3 |
feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash merges ( |
||
|
|
5dd70d7cb6 |
fix(desktop): say when the launch-config reader ignored a desktop block, and pair space-separated argv switches (#77311)
The pre-window config reader is a deliberate YAML subset (two-space keys, four-space `- ` items), so any other valid indentation silently disabled both launch keys — the app just started with no heap ceiling and no way to tell. And the argv scan read `--name=value` only, so `--js-flags --max-old-space-size=4096` lost the value and registered a bogus `max-old-space-size` switch, letting config.yaml's ceiling replace the launcher's instead of merging with it. - one warning naming the supported shape when a `desktop:` block mentions either key and nothing parses; a block with no mention stays silent. - argv pairs a bare `--name` with a following non-flag token, and `--js-flags` with whatever follows it (its value is dash-prefixed). - website/docs/user-guide/desktop.md documents the supported indentation next to the snippet, plus the warning users will see. |
||
|
|
7d2d634708 |
fix(desktop): renderer heap ceiling and desktop.electron_flags reach packaged launches (#77311)
`desktop.electron_flags` was appended to argv only by the `hermes desktop` launcher, so a packaged app started from its Start-menu / .desktop entry never saw it, and nothing set `--js-flags` at all — three Windows reporters on #77311 confirmed `jsflags=NO` on the renderer command line and OOM crash-loops with 15 GB of free RAM. Chromium copies `js-flags` to renderer processes only from the browser's pre-`ready` command line, so main.ts now reads the two launch keys out of config.yaml itself (no YAML dependency in the main bundle: a deliberately narrow subset reader) and applies them via app.commandLine before ready. New typed knob `desktop.renderer_max_old_space_mb` (default 0 = Chromium default) becomes `--js-flags=--max-old-space-size=N`, MERGED with any existing js-flags (appendSwitch replaces, so a naive apply would drop one). The planner is pure (`electron/renderer-heap-flags.ts`) and vitest-covered. |
||
|
|
845fee6cf4 |
feat(website): a page for every catalog plugin and every author
/docs/plugins/<name> and /docs/plugins/by/<maintainer> are generated at build time by a small Docusaurus plugin (website/plugins/plugin-catalog-pages) from the same plugins.json the grid fetches, so a merged catalog PR is the only way a page appears or changes. Plugin page: full description with the Disclosure sentence pulled into a callout, pinned commit / version / platforms / requires facts, tools, hooks and env chips, Desktop install button + CLI command, repository and reviewed source links, optional screenshots gallery, optional README rendered from the PINNED commit (fetched at build, allowlisted HTML, raw HTML dropped, relative links and images resolved against the pinned tree), and a "More by this author" shelf. Author page: everything a maintainer has in the catalog with total stars and a profile link when all repos share one owner. Cards on the grid link through: the title is a link and a card click opens the page (the Desktop picker embed keeps expanding in place). The shared vocabulary (entry type, tier/category taxonomy, link builders) moved to src/components/PluginCatalog/catalog.ts so the three surfaces cannot drift. Docs: field table and catalog README describe screenshots:/readme: and the pages. |
||
|
|
fb9c77e064 |
docs(bot-mode): held messages are replayed after release; Detect stop directives toggle
User-visible behaviour from #117488/#117472 and the quoted-stop-word rule from #117605. |
||
|
|
e089be80a1 |
docs(desktop-sdk): note DecodeText loop is opt-in
DecodeText is a published plugin-SDK export; flipping its `loop` default to false silently changes third-party plugin visuals, so the SDK doc says so where the component is listed. |
||
|
|
968422e7c5 |
fix(gateway): keep the operator restart tail on the home-channel storage notice
The cause-table action is user-phrased ("Send your message again once compression finishes"),
so the OPERATOR notice lost "then `hermes gateway restart`" for store-level failures that stay
broken until the gateway is restarted. The tail is appended for every cause except the
session-scoped ones that clear on their own (compression, compression_closed, turn_lease).
Also: the held-store refusal test is parametrized over optimize / optimize-storage / prune
(optimize-storage, the command the issue names as the field producer, was uncovered) and
asserts the refusal names the same store SessionDB opened — no `hermes sessions` subcommand
can point the command at another database. Docs: doctor refuses the checkpoint only while it
can see a process holding the RETIRED log.
|
||
|
|
6ba45b0e06 |
fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed holder scan doctor and repair use before rewriting the store. While a gateway, Desktop, dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as `PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning, `--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and every agent answered every turn with the retired-WAL refusal until all writers were stopped by hand (#110054, maintainer follow-up 09-20). The DeletedWalGenerationError text is now two layers: a first sentence for the person reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete files while they run, docs link), then the operator detail. The classifier fingerprint "deleted state.db-wal or state.db-shm" is unchanged. The cause table (`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway home-channel notice) and the chat explainer carry the same first steps; the gateway notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause, which for a held retired generation is the second-writer trap. New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the guard text, the developer state-db-recovery page and the sessions guide): the three steps, the do-nots, why maintenance refuses, and what the files beside state.db are (retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups, snapshots). |
||
|
|
12fb5eb476 |
fix(sessions): make set-journal-mode portable and fail closed where it cannot prove quiescence
Review follow-ups on the new `hermes sessions set-journal-mode` verb: - The header probe used os.pread, which does not exist on Windows, while the subparser is registered unconditionally — the command died there with an uncaught AttributeError. It now reads the 20 header bytes through a plain binary open(), and the tests no longer skip on win32. - foreign_state_db_holders() returns [] unconditionally on Windows (no scan), which made the admission gate vacuous: an operator got a silent all-clear and could flip the mode under a running gateway. Windows now refuses outright, naming the reason, overridable only by --force. - Enabling WAL ignored the cross-VM filesystem refusal the runtime enforces (apply_wal_with_fallback). target=wal now refuses on virtiofs/9p, where WAL shared memory silently corrupts. - A --db pointing at a garbage file surfaced a raw sqlite3.DatabaseError traceback even though the header probe had already read not-a-database, and a directory raised IsADirectoryError. Both now bail in the command's own error style; every open/read is guarded. The admission checks that need no I/O live in a pure _refusal() that takes the platform as data, so the Windows and cross-VM invariants are tested without faking sys.platform. |
||
|
|
96da5d97fc |
feat(sessions): hermes sessions set-journal-mode delete|wal converts an existing WAL store offline (#100896)
`database.journal_mode: delete` can never self-apply to a store that is already WAL: apply_wal_with_fallback deliberately never live-downgrades (#68545 — other gateway/cron/worker connections may hold uncheckpointed WAL commits), so operators applying the containment for the multi-writer corruption class saw one ERROR per process forever and the only escape hatch was an undocumented hand-run PRAGMA on the file. The new pre-DB `sessions set-journal-mode` verb is the sanctioned offline path: it refuses while ANY foreign process holds the file or a sidecar (the same foreign_state_db_holders scan doctor/repair admission uses, naming each PID), flips through _set_journal_mode_no_wait (busy_timeout=0, so an opener appearing mid-way makes SQLite refuse instead of racing it), verifies header bytes 18/19, and reminds the operator when config.yaml disagrees. `--db PATH` covers kanban.db / cron stores that log the same ERROR. The never-live-downgrade invariant is untouched; the ERROR, doctor hints and docs now name the command instead of the raw PRAGMA. |
||
|
|
bdde0b0c28 |
docs(cron): describe when a stale-code ticker yields and when it keeps dispatching
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped code_sha is the on-disk revision; a lock held by an equally stale process never counts. Placed under "Gateway Integration" so it does not collide with the "Stale-code yield" section #117501 adds under "Locking". |
||
|
|
050ea53aba |
fix(desktop): honor Applies-to on Custom Endpoints
The page only showed a read-only active-profile note, so endpoint saves followed the left-rail Bot instead of the Settings chips Accounts and API keys already share. |
||
|
|
c545272568 |
fix(lsp): one stalled request no longer silences a workspace for good — retry window, cold-root warm-up budget, per-root exclusion
A language server that missed its budget once marked its (server, root) pair broken for the process lifetime, the same 5 s steady-state budget was applied to a cold server that also had to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched the server off for every workspace. Three new keys under the existing `lsp` block, all defaulting to today's behaviour: - lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per pair; an expired pair gets one more try, and the INFO skip line names the retry time. - lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client waits up to this budget (outer join budget follows); warm requests keep wait_timeout. - lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path also covers everything beneath it); a matching root never spawns, logged once at INFO. A non-list value fails closed — WARNING naming the expected shape, every root skipped — because silently excluding nothing would re-pay the stall the key was meant to avoid. Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo). |
||
|
|
eaef5ec717 |
feat(update): hermes update --list-venv-holders prints the venv guard's holders as JSON, exit 3
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].
Fixes #117246
|
||
|
|
5be70b5e25 | docs(mcp): note Google-hosted OAuth servers get access_type=offline for a refresh token | ||
|
|
86a599cbae |
fix(gateway): human-delay pacing comes from each profile's human_delay config, not process env
BasePlatformAdapter._get_human_delay read HERMES_HUMAN_DELAY_MODE/_MIN_MS/_MAX_MS from the process environment at every send, so under multiplexing the launch profile's pacing applied to every served profile, and the documented `human_delay:` config section (mode/min_ms/max_ms, already in DEFAULT_CONFIG) was never consulted. The runner now resolves `human_delay` per profile through the same seam as the busy-text timings (`_human_delay_from_config`, snapshotted in `_snapshot_profile_busy_modes`, installed by `_wire_adapter_handlers`) and the adapter only consumes the installed range. Invalid `custom` bounds (non-integer, negative, inverted) warn naming the key and fall back to the natural range. Fixes #116895 |
||
|
|
5a074a02cb |
fix(gateway): busy-text debounce/hard-cap come from per-profile config, not process env
BasePlatformAdapter read HERMES_GATEWAY_BUSY_TEXT_{MODE,DEBOUNCE_SECONDS,HARD_CAP_SECONDS}
at construction, freezing the launch profile values into every profile adapter under
multiplexing. The mode was already re-synced per profile by the runner; the two timing
knobs were not. They are now display.busy_text_debounce_seconds /
display.busy_text_hard_cap_seconds, snapshotted per profile next to the busy modes and
installed by _wire_adapter_handlers. Invalid values warn naming the key and fall back.
Fixes #116893
|
||
|
|
f568b860d7 |
fix(agent): pin the reset_at reach through try_activate_fallback and add opt-in fallback.min_switch_reset_seconds
- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback (production entry) on a real AIAgent with a one-entry chain, so dropping the reset_at forwarding goes red (3 failures before, 8 green after). - #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the rate-limited primary's declared reset is sooner than N seconds, try_activate_fallback returns False and no cooldown is armed; docs row added. |
||
|
|
53815e24dc |
fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.
The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.
Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
before req_reasoning: {}
after req_reasoning: {'reasoning_effort': 'medium'}
agent.reasoning_effort: low -> {'reasoning_effort': 'low'} (unchanged)
|
||
|
|
4f0ced3e5f |
fix(desktop): mask typographic double quotes too and document quoted stop words
macOS smart-quote substitution turns the straight quotes a user types into “ ” in the composer, so the quoted-span mask covers both; the user-visible rule (stop words inside code, quotes or blockquotes never hold) now has its sentence in bot-mode.md alongside the behaviour. |
||
|
|
2539a08fd1 | docs(gateway): clarify systemd reload semantics | ||
|
|
97beaeeeb3 |
feat(delegate): warn the child at 80% of its inactivity window before abandoning it
A child that stalls under a configured delegation.child_timeout_seconds used to learn about the budget only by dying, losing its whole context. The liveness wait now queues a one-line "[delegation budget warning]" through the child's steer channel once the idle window is 80% spent (delivered at the child's next iteration boundary), so a slow-but-recoverable child can wrap up and return its summary. The warning fires once per idle window and re-arms when progress resets the window; a progressing child never sees it. Part of #116001 (atom 2A). Semantics of child_timeout_seconds are unchanged. |
||
|
|
03973bd02b |
fix(delegate): child_timeout_seconds bounds inactivity, not total runtime
The configured cap was a dispatch-to-death stopwatch: `await_child` waited on a plain `settled.wait(timeout=child_timeout)`, so any child that outlived the budget was abandoned even while the provider was actively serving it. The report's corpus for #116001 (219 tasks / 75 deaths, 0 of them mid-tool) could not be reproduced here — it needs the reporter's slow OpenAI-compatible endpoint — but the mechanism it names is exactly this gate: a child waiting on an in-flight LLM completion, killed with a nearly-finished context. A slow child is already bounded elsewhere (the per-call stale watchdog, the heartbeat's staleness verdict), so this cap could only ever kill children the runtime had judged healthy. `child_timeout_seconds` now measures time with NO progress: the wait runs in slices and restarts the window on the same signals the heartbeat's stale verdict reads (completed call, tool change, activity-clock tick). A frozen child is still abandoned when the window elapses; a progressing one is never killed for taking long. Timeout entries also carry `last_event_age` (how long the child had been silent), so operators can tell a slow provider from a runaway without transcript forensics. Fixes the mechanism reported in #116001. The budget warning and continuation-respawn items in that issue are separate features and are not part of this change. |
||
|
|
7d8e8f234f | test(gateway): pin the send_voice is_voice contract through the base media dispatch; document Matrix audio attachments |