Commit Graph

3483 Commits

Author SHA1 Message Date
teknium1
061195fac1 FLEET: one host-scoped update-restart obligation, one restart per host
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.

- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
  gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
  unit->live-MainPID collapse rule. The legacy per-home marker stays readable
  and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
  idempotent per host (a completed restart onto the checkout SHA is never
  repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
  restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
  profile's home channels; the marker survives until each was reached.
2026-09-21 07:14:37 -07:00
Austin Pickett
274bc7b8f6 docs(desktop): where the Linux Chromium log and minidumps live 2026-09-21 09:00:05 -04:00
teknium1
a1398a6f67 fix: doctor, cron status, claw and gateway status report the real host topology
Multiplex-only (Teknium ruling) runs exactly ONE `hermes gateway run` per host,
multiplexing every profile. Four reporting surfaces still asked a per-PROFILE
process question ("does MY profile own a gateway process?"), so a SERVED profile
answered "no" and the output lied:

* `hermes -p served cron status` printed "Gateway is not running" and told the
  user to run `hermes gateway install` / `gateway run` for that profile — i.e.
  to start a SECOND host process, which the architecture forbids.
* `hermes doctor` reported "Per-profile gateways: up/total" from s6 slots and
  checked systemd linger against the CURRENT profile's unit, so a doctor run
  under a served profile skipped the check entirely.
* `hermes claw`'s destructive-action warning was gated on `get_running_pid()`,
  which returns None for a served profile — the token-conflict warning was
  SILENTLY SKIPPED (a real safety hole).
* `gateway.status.multiplexer_liveness_for_profile()` returned None for the
  default home by construction, so `default` could never be reported as SERVED.
* `hermes doctor`'s state.db holder + WAL wording implied one gateway per
  profile.

New `gateway/host_topology.py` resolves "which single process owns the gateway
role on this host, and which profiles does it serve" once, from the host
rendezvous record (`gateway/host_rendezvous.py`), falling back to the default
home's recorded `served_profiles` for a gateway that predates the record.
`default` is just another served profile there.

Root cause: every surface derived gateway identity from per-profile artifacts
(argv `-p <name>`, `gateway.pid`, the active profile's systemd unit, s6 slot
counts) instead of the host record that actually names the owner.
2026-09-21 05:06:45 -07:00
teknium1
a10620a669 fix(gateway): a record alone never means "attach", and --replace/--force work
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.

- ATTACH now requires a LIVE `identify` answer. The claim-time record is
  published with NO served set (the runner settles multiplex a moment later),
  and `host_gateway()` reports `served_known=False` when nothing answers. An
  owner whose served set is unknown yields a TRANSIENT refusal, never an
  attach: previously `default`'s claim published "default,other" before its
  socket bound, `other`'s systemd unit read that as "I am served", exited 78,
  and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
  instead of forcing `multiplex=True`, so a standalone gateway stops claiming
  the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
  `--force` into `_host_attach_or_none`. Both previously exited in the guard
  before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
  config refusal every supervisor parks on; "someone else serves me right now"
  is a runtime observation that ends when that process does. No unit files
  change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
  finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
  would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
  and re-enters with `replace=True`, so it can no longer attach to the corpse
  it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
  `st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
  it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
  generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
  `gateway status`/doctor across N profiles pays one probe, not N.

Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
2026-09-21 05:02:29 -07:00
teknium1
a20a88398f gateway: lifecycle verbs mean "the one host multiplexer"
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:

- `gateway run` for a profile the host gateway already serves ATTACHES: print
  its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
  to re-scan `profiles/` (control socket) and attach once the answer includes
  it. Refuse only when the host gateway cannot be made to serve it. Under a
  service supervisor the attach exits 78 instead of 0 so a redundant unit is
  parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
  landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
  the owner's control socket, so it no longer depends on the claim ordering in
  start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
  process: they restart the host multiplexer and preserve its served set. A
  secondary still running its own gateway is reported with the
  `gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
  argv: a host singleton runs bare/default argv and can never prove it serves
  profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
  multiplexer is whichever profile launched the one host process.

Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
2026-09-21 05:02:29 -07:00
kshitijk4poor
5d63d3e5c5 fix(gateway): launchd gateway reaches the local network on macOS (#71206, #57812)
macOS Local Network Privacy attributes a socket to the executable launchd
spawned for the job. A bare venv Python has no application identity and is
not platform-entitled, so every LAN connect from the supervised gateway
failed with errno 65 (No route to host) while the same URL worked from
Terminal — and nehelper never showed a prompt that could grant it, so a
headless Mac had no way out.

Run the job through /usr/bin/osascript: `do shell script "exec …"` spawns
its child as osascript-responsible — an Apple platform binary — and the
child is exempt. Verified live on macOS 26.3.1 with a fresh ad-hoc binary
under the gui launchd domain: bare → 65, `/bin/sh -c exec` → 65,
`/usr/bin/time` → 65, osascript → reachable; an ad-hoc-signed helper .app
with NSLocalNetworkUsageDescription (the #115196 approach) stays denied with
no prompt even after lsregister, matching the #57812 dead-end table.

`do shell script` buffers the child's output until exit, so the shell
command redirects stdout/stderr to the log files the plist already routes;
`exec` keeps the gateway in the job's process group, so `launchctl bootout`
still delivers SIGTERM to it (live: stop → no orphan, restart → exit 75 →
KeepAlive respawn, stop → parked). The existing plist-staleness refresh
picks the new definition up on `hermes gateway install`/`start`.

Mechanism proposed by @leewaiho in #57812.

Co-authored-by: leewaiho <18321182+leewaiho@users.noreply.github.com>
2026-09-21 17:16:23 +05:30
kshitijk4poor
db1f3f4564 fix(computer-use): doctor names the stale TCC row on the health_report path too
The first cut appended the recovery only in `_tcc_row`, which is built by the
0.10 fallback probes. Drivers that serve health_report (0.22+, the ones a
stale row actually bites) render their own tcc_* rows untouched, so `doctor`
never showed it. Apply the hint at the report seam like the display-count
guard so both paths get it.

Also: one CUA_DRIVER_BUNDLE_ID (permissions.py is the leaf; the daemon
imports it), the CLI derives the field set from the same table, and the
unverified "--upgrade repairs stale rows" claim is dropped from hint + docs
(the user hitting this is already on a current driver).
2026-09-21 16:21:00 +05:30
kshitijk4poor
42ae8b4410 fix(computer-use): name the stale TCC row when CuaDriver shows ON but is denied
`hermes computer-use permissions status` and the doctor's fallback `tcc_*`
rows told a user whose System Settings toggle already showed CuaDriver ON to
"grant it in System Settings" — the one step that cannot help. macOS keys the
TCC row to the app's code-signing requirement; a row written for an earlier
CuaDriver build stops matching after a driver update and flipping the toggle
does not rewrite it (trycua/cua#3170, repaired by the cua-driver >= 0.22
installer on update). The daemon then reports Accessibility / Screen Recording
false while the pane shows ON, and every click silently no-ops (#99732).

`permissions.stale_tcc_grant_hint(*missing)` renders the reset for exactly
the missing services (`tccutil reset Accessibility|ScreenCapture
com.trycua.driver`, then `hermes computer-use permissions grant`); both
surfaces append it only when a grant is actually reported False. Docs: the
computer-use page still said bounded/unrestricted daemons run "under the
Hermes host identity" — they launch through CuaDriver.app since #95381 — and
gains the stale-row troubleshooting entry.
2026-09-21 16:21:00 +05:30
teknium1
bfcd209906 docs(multiplexing): record the config escape hatch and the .env-over-frozen-env precedence
Activation is not purely stricter for the LAUNCH tenant: inside the launch scope its own
.env wins over the frozen boot env, so a key present in both reads differently before and
after activation. Also documents the explicit multiplex_profiles: false escape hatch and
the fail-closed unresolved-profile case.
2026-09-21 02:58:40 -07:00
teknium1
c718267ef3 fix(multiplex): close the profile-scope holes the scope machinery misses
One host process serves every profile, so every execution point must bind the
profile it is acting FOR. These six ran unscoped (or bound only part of a
scope) and resolved get_hermes_home()/credentials against the LAUNCH profile:

- tui_gateway/session_reaper: the idle-reaper and exit-flush transcript writes
  now enter the SESSION's profile scope, the same chokepoint _finalize_session
  already binds. A served profile's transcript was landing in the launch home.
- gateway/run: MCP shutdown tears down per served profile inside that
  profile's scope (mirrors startup discovery and the reconcile chore), with a
  trailing wildcard pass under the launch profile's own scope.
- gateway/run_profile_reconcile: _unserve_profile's adapter teardown, agent
  eviction and state/memory handle release now run inside the deleted
  profile's scope.
- gateway/run_adapters + run_goals + run_notifications: a body with no routed
  profile no longer means "no scope". launch_profile_scope_if_multiplexed()
  binds the launch profile once the process multiplexes; before activation it
  is still literally a nullcontext, so single-profile hosts are unchanged.
- hermes_cli/kanban_db_dispatch: one _worker_profile_scope helper binds the
  assignee's secret AND terminal scope for toolset resolution and the spawn-env
  build, unconditionally instead of only under multiplex.
- hermes_cli/web_server: `hermes serve` activates multi-profile hosting at boot
  when the host has more than one servable profile home, instead of lazily on
  the first ?profile= request after earlier work already ran unscoped.

Secret scope is never widened: a non-launch home resolves from its own .env and
sources only; the launch home keeps its existing env-over-.env precedence.
2026-09-21 02:58:40 -07:00
teknium1
9ecd22e5da fix(cron): release a profile's in-flight claim under the key it registered with
Under one multiplexing ticker, every non-launch profile's in-flight claim
leaked on every run. The cron home scope is a ContextVar: the claim is taken on
the ticker thread inside `_profile_cron_scope`, but the pool worker's `finally`
sits outside `ctx.run`, where the worker thread resolves the LAUNCH home — so
the discard missed the real key. The job then skipped a fire window until the
force-release backstop swept it, and the shutdown drain plus
`hermes.cron.jobs.running` saw phantom work.

`_submit_with_guard` now captures the registering home and passes it to
`release_running_job(job_id, home=...)` on every release path.

Also in this pass:

- Stand down for a profile that runs its OWN gateway (`_cron_profile_gate`,
  the gate the serve/Desktop ticker already passes). On a host pinned to
  per-profile gateways both processes raced that profile's tick lock, and when
  the launch gateway won, delivery went through SharedRouteAdapters/fail-closed
  instead of the profile's live adapters. The gate compares the liveness PID
  against `os.getpid()`: this process holds the launch `gateway.pid` and
  publishes every served profile in `served_profiles`, so a bare liveness answer
  would have stood cron down host-wide.
- Job-liveness consumers ask `is_job_running(job_id, home=...)` instead of the
  host-wide bare-id union, which let profile A's running `daily-brief` report
  profile B's idle one as running and keep B's stale one-shot alive.
- `register_ticked_homes` reaps the parallel pools of homes that leave the
  ticked set; pools lived until `atexit`, so every home ever ticked kept a
  ThreadPoolExecutor and its worker threads.
- `mark_running_jobs_interrupted` reads the real home Path from
  `_inflight_home_path` instead of rebuilding it from the normcased key half.
- The ownership test imports `cron.scheduler_ownership` per test, so the file
  fails behaviourally on base instead of as a collection error.
2026-09-21 02:52:15 -07:00
teknium1
cb647a018f fix(cron): one host ticker owns every profile's cron, per profile
The cron ticker multiplexes N profiles from one process while its ownership
predicates and its in-flight bookkeeping still assumed one profile per process.

- `_should_yield_tick_to_fresh_gateway` asked a process-global boolean
  (`owns_gateway_runtime_lock`) and a launch-home lock probe, so one answer
  covered every profile ticked. It now asks `scheduler_ownership`:
  `owns_cron_tick_for(home)` (this process is the host gateway AND ticks that
  home) and `live_gateway_ticking(home)` (another live host gateway whose
  published served set covers that home).
- In-flight state (`_running_job_ids`, `_running_since`, `_running_futures`,
  `_running_allowance_s`, `_running_worker_pids`, `_running_fire_owners`,
  `_restart_safe_waiter_job_ids`, `_interrupted_job_ids`) is keyed by
  `_inflight_key(job_id)` = `(home key, job id)`; two profiles carrying a
  `daily-brief` no longer read as one job. The public accessors still report
  the host-wide union of bare job ids for the shutdown drain.
- The parallel worker pool is keyed by home: `cron.max_parallel_jobs` is a
  per-profile key, and the single global pool was sized by whichever profile
  ticked first and torn down by the next one.
- `gateway/run.py` no longer gates the cron tick set on
  `gateway.multiplex_profiles`: that flag gates adapters, and with it off every
  non-launch profile's jobs sat in a store no ticker visited.
2026-09-21 02:52:15 -07:00
teknium1
3917c7dd6a fix: RoomLink catalog names the served profile and reads its policy from that profile
`catalog_mapping` took an optional `target_profile` and fell back to the launch
process's `HERMES_PROFILE` (then "default"), and `execution_policy_mapping`
loaded whatever gateway home was active while labelling the result with the
requested `target_profile`. On a multiplexed / Desktop-spawned backend a remote
Bot invited for profile B could receive a catalog signed for the launch
profile, or B's name over A's approvals/turn-limit/toolsets.

- `target_profile` is a required keyword on `catalog_mapping`; an empty or
  mismatched profile fails loudly naming the field, no env fallback.
- `execution_policy_mapping(config=None)` resolves the SERVED profile's own
  config: a no-op when it is the active home (callers that already scope are
  unchanged), else `_profile_runtime_scope(get_profile_dir(profile))`; a profile
  that does not exist is refused instead of silently reading the launch config.
- `issue_room_grant`'s default digest inherits the served-profile resolution.
- Tests: 27 call sites now name the profile; two invariants red on base.
- Docs: multi-profile resolution table row for the RoomLink catalog/policy.

Fixes #116900
2026-09-21 02:08:18 -07:00
teknium1
8863b36fd6 fix(cron): stale-code yield reads as an outage in cron status; hermes update restarts proven-stale gateways
After `hermes update` fast-forwards the checkout under a running multiplex gateway,
its cron ticker yields every tick ("stale code: booted on A, disk is at B") for as
long as the process lives. Two things made that a silent total dispatch outage:

* `hermes cron status` weighed only the liveness heartbeat (kept fresh by the yielding
  loop) and the success marker; with no success marker on disk it printed
  "✓ Gateway is running — cron jobs will fire automatically", and with a stale one
  it pointed at "Check the gateway log" instead of naming the cause. The persisted
  `CronTickYielded` error is now recognised (`cron.scheduler.stale_code_yield_labels`)
  and reported as "Gateway is running STALE code — fires NOTHING", with both
  revisions and the restart command. A fresh heartbeat with a recorded error and no
  success marker is no longer green either.
* The post-update fleet version matrix flagged a `stale` gateway and exited 1, but
  left it running. `_verify_fleet_after_update` now hands every proven-stale survivor
  to the existing drain-first `request_restart` path (SIGUSR1) via
  `hermes_cli/update_cmd_stale_survivors.py`: a supervised gateway respawns on the new
  code, a bare `gateway run` is stopped and listed under "Restart manually" — the same
  contract the restart phase already uses for unmapped manual gateways. The drain
  budget computation is shared (`_gateway_drain_budget`).

The yield itself is unchanged: a stale-code process still never dispatches while a
fresher lock holder exists (design of 9a7732b45f).

Fixes #117275
2026-09-21 01:57:31 -07:00
teknium1
661d73f06b test+docs: trim doctor exit-status tests to two invariants; document the exit code
Keep the in-process return-code matrix (issues / manual / --fix full and
partial repair) and the real-process exit-status check; drop the exception
passthrough, --ack and --live cases, which pin behaviour this change does not
touch. Document 0/1 exit status under `hermes doctor` in the CLI reference.
2026-09-21 01:57:05 -07:00
kshitijk4poor
67208caac3 fix(memory): replace states and surfaces the whole-entry contract instead of truncating silently (#117952, salvage #109177)
A partial-entry replace (old_text = a span inside an entry, content = its
replacement) silently overwrote the WHOLE entry: every clause outside the
match was destroyed with success=true, most damagingly through the
write-approval replay path (/memory approve -> apply_memory_pending ->
apply_batch) that exists precisely so background agents can stage writes for
human review (#117952, dup of #59184).

The cluster's repeated fix attempt (#66321, #61357, four closed predecessors)
flipped replace to naive span-splicing (entry.replace(old_text, new, 1)).
That corrupts the canonical identifier-style call (replace(old_text="3.11",
content="Python 3.12 project") -> "Python Python 3.12 project project"),
breaks the background-review refine consumer, and split-brains external
memory mirrors, which receive the raw op content, not a merged entry.

The defect is the CONTRACT, not the mechanism: the schema invited patch-style
calls ("mirrors old_text (patch-tool shape)") while the store commits
whole entries. This pins whole-entry semantics end to end and makes the
overwrite visible instead of silent:

- schema + recovery error text + memory_tool docstring now say content is the
  COMPLETE new entry and old_text only locates it (website docs updated in the
  same PR)
- a successful replace surfaces replaced_entry / replaced_entries (full text
  that was overwritten) on single, batch, and approval-replay paths, so the
  caller can re-add lost clauses instead of discovering them gone later
- _mutate gains an optional extra payload field; _apply_batch_op returns the
  replaced entry text alongside the error

Live repro on origin/main dec236b21 before: approve-apply of a staged
partial-entry replace truncated a 3-clause entry to the replacement span
alone. After: same final entry (whole-entry contract, unchanged semantics)
plus replaced_entries in the response. Grafted naive span-splice on main:
2 existing tests red (test_replace_entry,
test_attended_review_keeps_full_operation_set) — the regression this PR
avoids. Mutation checks: reverting the visibility fields or the exact-match
priority turns the new invariant tests red.
2026-09-21 14:26:37 +05:30
teknium1
2520eb7c0d docs(backup): list the browser profile dirs hermes backup excludes 2026-09-21 01:44:42 -07:00
teknium1
9a1f06293d feat(mcp): make the discovery connect cap configurable (mcp.discovery_concurrency, default 4, 0 = unlimited)
Builds on Halldrix's #117674 (flat cap of 3): the cap now reads config.yaml
`mcp.discovery_concurrency` (registered in DEFAULT_CONFIG under the existing
`mcp:` section, default 4 per the #117373 ruling). 0 disables the semaphore and
restores the unbounded gather; a non-integer / negative value warns and uses
the default instead of silently running unbounded. The wave-scaled pass timeout
uses the effective cap.

Tests trimmed to two invariants driven through `_run_discovery_pass` (the
production entry): configured cap bounds in-flight connects while every server
still connects and 0 means unlimited; pass timeout stays under the lock waiter
budget. Docs: mcp.md runtime section + multi-profile-gateways.md.

Fixes #117373
2026-09-21 01:44:16 -07:00
teknium1
2463550c97 feat(website): plugin pages render the README from the pinned commit by default
The README section was gated on `readme: true` in the catalog YAML and no entry set it, so
all 222 plugin pages shipped without one. READMEs now render for every GitHub/GitLab entry
(fetched at the reviewed sha, subdir first then repo root, common casings and docs/README.md
as fallbacks); `readme: false` opts an entry out. Build proof: 222/222 READMEs rendered.
2026-09-21 01:15:33 -07:00
teknium1
0fd811e687 refactor(gateway): drop the command_denied_message knob and silent-denial mode
Maintainer ruling on #117217: a configurable refusal text (and the ""-means-silent
mode) is feature creep. The refusal text returned by _check_slash_access is now
byte-identical to origin/main again; only the /help and /commands catalog gating
for non-admin callers remains in this PR.

Removes the run_busy.py override hunk, the two knob tests, and the
"Customizing command refusals" docs section.
2026-09-21 01:06:02 -07:00
teknium1
2978c23acc fix(gateway): /help and /commands show a gated non-admin only the commands they can run
Both catalog handlers called the shared executor without the caller's slash-access
policy, so a non-admin under allow_admin_from saw every admin-only command and then hit
the refusal on each one. The handlers now pass the policy's runnable set (the
/help+/whoami floor plus user_allowed_commands) as `allowed_commands`; gateway_help_lines
filters on it and skill commands are hidden for gated users. Admins and ungated scopes
are unchanged (option absent -> full catalog).

Part of #117217
2026-09-21 01:06:02 -07:00
funky-xamarin
72ba7ae3d6 fix(gateway): honor configured slash-command refusal
(cherry picked from commit 8f052e0db1b33a5472473484d0a85aa72d476f73)
2026-09-21 01:06:02 -07:00
teknium1
d90c3ad40a fix(desktop): group-chat member failure row names the error's first line (#117366)
A member turn that failed without a typed gateway reason rendered as a bare
"X hit an error" — a stopped backend, a dead IPC bridge and a provider refusal
all looked identical and gave the user nothing to act on.

groupFailureReason now falls back to the error message's first non-empty
line (trimmed, secret-shaped spans redacted, capped at 200 chars), so the
Activity row reads "X hit an error — <cause>". Typed data.reason and the
slot-wait classification keep precedence; the roster badge still classifies
from the same text.

The plugin fence keeps the Electron-side redactSecrets out of reach, so the
same shapes (bearer header, ?token=/?api_key= params, vendor key prefixes,
user@host:password) live next to the label.

Fixes #117366
2026-09-21 01:05:32 -07:00
teknium1
b7803a1763 fix(compression): cap the protected tail at 20% of the context window
The lean tail budget is max(10K, min(25K, 2.5% of window)) and the boundary walk lets whole
rows overrun it by 1.5x. Neither term knew the window size, so on a small local model the
"protected" tail WAS the request: 10,636 tokens of a 8,192 window (129%), 64% of 16K. Every
compaction pass summarised six rows, kept 39 verbatim, and reclaimed nothing — a Titan RTX 27B
timed out before compaction ever changed anything, and protect_last_n read as an uncompressed
tail rather than a minimum.

TAIL_MAX_CONTEXT_FRACTION (0.20) now bounds both the budget (either tail_mode) and the walk /
pressure-demotion soft ceiling. Required last-user / last-assistant anchors and atomic tool
groups may still exceed it, so the retained tail lands at 22-25% on 8K-32K windows instead of
32-129%. Windows of 128K and above are unchanged (10K lean floor < 20%).

Probe (12 tool-heavy turns, 49 rows, 12.8K tokens):
  ctx    8K: tail 10,636 tok / 39 rows -> 2,116 tok / 7 rows;  window [4,10) -> [4,42)
  ctx   16K: tail 10,636 tok / 39 rows -> 4,246 tok / 15 rows; window [4,10) -> [4,34)
  ctx   32K: tail 10,636 tok / 39 rows -> 7,441 tok / 27 rows; window [4,10) -> [4,22)
  ctx  128K: identical before/after
2026-09-21 01:04:16 -07:00
teknium1
dec236b214 feat(desktop): Uninstall for standalone desktop plugins via Electron IPC
Rows whose only half is a folder in <HERMES_HOME>/desktop-plugins (no agent
package behind it) get the same trash button + destructive confirm as the
agent rows. The renderer names the FOLDER, never a path; Electron
(`hermes:plugin:removeDesktop` -> desktop-plugin-remove.ts) resolves it under
the app-level root, refuses anything that is not a single contained segment,
refuses unified-package halves (the reconcile would re-copy them; the agent
uninstall prunes them), removes a symlinked folder as the link, and deletes
the tree. The loader then retires the entry (unload, drop rows, stop watch)
so the pane/commands vanish without waiting for the next scan.

Why an IPC: these plugins have no agent half, so `plugins.manage remove` on
the gateway cannot reach them — they live on this computer, in every
connection mode.
2026-09-21 00:09:50 -07:00
teknium1
43fe9c1773 docs(desktop): describe the Plugins hub Uninstall button 2026-09-21 00:09:50 -07:00
Siddharth Balyan
afc3b7c6f3 feat(connectors): one backend-owned connection operation, with a setup card on Desktop, TUI and CLI (#111008)
* feat(connectors): the desktop connects apps through one backend-owned operation

Re-based onto main after #109517, #110368, #110574 and #110843 landed as squash
merges (e0ef0eb9c3, d105376b21, ee2f5629b8, 1ab32b212b): the branch's history
no longer shared a base with main, so this is the PR's exact delta against
+3762/-1470, identical to the branch tip 05d7f2d3d5.

The fifteen commits it carried, in order:

--- feat(connectors): the wire follows the connector contract (six-state status, required toolkit metadata, connectionId on mint)

Hermes types exactly what the contract page writes: /tmp/magic/CONTRACT-TOOLKIT-METADATA.resolved.md
(carriers: portal PR2 `sid/connection-api` for the enum and account rows, the portal contract branch
on top of it for the list metadata). No optional-for-compat fields, no fallback branch, no seven-state
word left in the tree. A gateway that does not speak this contract fails validation loudly.

Wire (tools/connectors/gateway/wire.py): `ConnectionStatus` is the six values, `pending` covering the
vendor's INITIALIZING and INITIATED; `ConnectorListItem` requires title, description, https iconUrl and
authKind, and carries activeConnectionId when the session binds an account; `ConnectorListResponse`
types a page whole with its total; `ConnectorAccount` and `ConnectorAccountsResponse` type the account
routes; a mint result names its connectionId on `initiated` and never on `failed`; CONNECTION_REQUIRED
carries connectUrl and connectionId together or not at all; `account` exists on the execute call and is
never sent while multi-account is off.

Client: `list_connectors(search, connected)` refuses a search under three characters before any
request and parses every page whole; `list_accounts` and `account_status` read the account routes,
404 `connection_not_found` is None and 429 raises `RateLimited(retry_after)`.

Watcher: the target keeps the account id the mint named and the toolkit's title and icon from one
list read before the card is emitted; `_status_for` is the one seam between the watcher and the
gateway's status source (the list walk today; the per-account route replaces that body only).
`pending` and a missing status move nothing; revoked and inactive are failed.

Desktop: `ConnectorRow` is the strict list item; `ConnectorRowSeed` is what a tool call's args can
say; `connection.request`/`update` targets carry `connection_id`, `title`, `icon_url`; the mark
ladder is glyph, vendor icon as a plain image, favicon, monogram.

Tests run red first: wire rejects the retired words and the missing metadata; client search
minimum, execute omits account, account_status 404/429; managed targets carry the metadata and the
account id, a pending row moves nothing; the logo ladder order; the store passes the fields through.

--- feat(connectors): onboarding connect-first rides the connection operation

The guided first build had its own connector surface: a 2 s connectors.list poll loop, an auto-open
of every minted link, a hidden "[setup] links opened" note that painted as a user bubble, CONNECT
FIRST rows in the transcript, and a "Start with N apps connected" composer pill that sent the go
signal early. Delete all of it. The build session's first manage_connections connect now shows the
same card as any chat, and the settled tool result is the go signal.

Deleted: store/first-build-connectors.ts, assistant-ui/first-build-connectors.tsx (nothing rendered
it any more), lib/first-build-start.ts, onboarding-chat/start.tsx, the first-build session markers in
handoff-receipt.ts, the connectionRows / latestConnectorPart helpers they were the last readers of,
and the five strings only they used.

The runbook (setup-profile.ts::connectFirstRunbook) now describes the operation's real contract: one
connect call with every picked slug, one card with a row per app, blocks until settled, a per-app
result of connected / skipped / not_connected. No wait, no "Start with", no "already active" branch
(the result never carried it), and no offer to re-mint: Try again and Continue live on the card, and a
Continue means the user moved on.

Tests (red first): the runbook names one connect call and its per-app result and never the retired
model actions; the no-account rule survives an empty pick.

--- feat(connectors): the registry is keyed by profile, every watch read is bounded by the deadline, connectionId is optional, connect and execute carry returnTo and op

Registry (P1-14): tools/connectors/live.py keys an operation by (profile home, session key). Two multiplexed
profiles can carry the same timestamp-based session key; the tool thread opens under its turn's profile
override and the RPC side names the session record's profile home, the default profile resolving to the
process home on both sides.

Watch reads (P1-8 residual): list_connectors takes a per-page timeout and the watcher bounds each read by
the operation's remaining deadline (floor one second), so a stalled gateway page cannot hold the operation
past its deadline.

Contract follow-up from the portal handoff (/tmp/magic/HANDOFF-HERMES-PR3-PORTAL-CONTRACT.md): connectionId
on a mint result and on CONNECTION_REQUIRED is plain Optional (a no-auth toolkit answers active with none; a
failed mint answers with neither link nor id); a target without one has nothing to watch. Connect and
execute requests carry returnTo (hermes-desktop | hermes-desktop-dev | portal) and op, so the vendor's done
page can send the browser back to the app; the call sites and the deep-link handler follow in the next
commit.

Tests, red first: two profiles share a session key without seeing each other; each watch read shrinks with
the deadline; the client honours a per-page timeout; the optional id; the two request fields. The fakes'
list_connectors accept the timeout keyword; two wrapper spies that did not were the cause of a 300 s hang
under the per-file harness.

--- feat(connectors): an MCP target runs on the shared connection operation

manage_connections ran MCP targets as a renderer errand: the desktop card ran
the catalog lookup, the install and the OAuth flow, then told the backend what
had happened, and every surface without a card got `unavailable` plus two
terminal commands. That left the outcome in the renderer's word, and left the
model unable to connect an MCP server anywhere but the desktop.

The backend now owns the work, the way it already owns a managed connector.
`prepare` starts the OAuth flow in process and points the browser at the
backend's own callback route; an install that still needs credentials waits
pending and publishes their names as `required_env`; `observe` reads the flow
or the install worker on every tick. The worker itself moved out of
tui_gateway/mcp_oauth_sessions.py into tools/connectors/mcp_oauth.py, and that
module now calls it, so the Capabilities tab keeps its RPC session table.

The card may only say approved, skipped or continue: a claim of any other
state moves nothing. With that, `renderer_flow` and the `unavailable` state
have no producer left and are gone from the contract. Off the desktop there is
no card, so the action runs at once and the result carries the authorization
URL for the user to open.

Try again reaches the same work: `_reissue` dispatches on `target.kind`, so an
MCP row re-runs its own install, enable or OAuth instead of minting a managed
gateway link.

(cherry picked from commit fa0438c4121807604e7c983ba42b901314a1305b)

--- feat(connectors): the MCP card projects the operation instead of running the flow

The MCP setup card used to own the work: it called the catalog, ran
installMcpCatalogEntry, polled the action, drove the OAuth window, and then told
the backend what had happened through connection.respond. That made the renderer
a second authority on target state, so an install that finished after the window
closed, or a card that never mounted, left the operation with a state nobody
could correct. PR3 moves that work to the backend, so the card has one job left:
show the operation and send the user's consent.

McpSetupPending now renders request.targets the way ConnectorOffer renders them:
one row per target, one verb by (action, state), Continue below the shell.
Install and Enable send {status: 'approved'} and nothing else; an authorize row
opens the link the backend minted, with no OAuth RPC of its own; a running
install holds its verb with the busy mark; a failed row sends
connectors.connect {reconnect: true} on the open operation, the same call the
managed card makes. The owner lookup and that call are now shared helpers on
connector-tool.tsx rather than a second copy here.

ConnectionTargetOutcome loses connected / initiated / failed: those were the
renderer reporting state, which it no longer may do. ConnectionActor loses
renderer_flow for the same reason. ConnectionOperationTarget gains required_env,
the credentials an MCP install is still waiting for; the card renders a field per
entry under the pending row and holds Install until every required one has text,
so the values travel with the approval instead of through a separate catalog call.

(cherry picked from commit 2e42ad19f63e41f274a87199d07eb99a7c995cbe)

--- fix(connectors): the toolkit metadata leaves the wire; the vendor logo is derived from the slug

Sid cancelled portal PR3 (the toolkit metadata on the list item) on 2026-09-14, so the fields the contract
commit typed as required come off the wire: no title, description, iconUrl, authKind or activeConnectionId
on the list item, no total on the page, no search or connected query, no title or icon_url on a target
snapshot. The list item is exactly what portal #1220 emits.

The desktop derives the vendor logo from the toolkit slug instead (`connectorIconUrl`,
https://logos.composio.dev/api/{slug}; the gateway slug is the vendor slug, checked for every lead-order
pick), rendered as a plain image because the host sends no CORS header. The title stays
`connectorTitle(slug)`. The mark ladder is unchanged: glyph, derived vendor icon, favicon, monogram.

Everything else in the contract stands: six-state status, optional connectionId on mint results and on
CONNECTION_REQUIRED, returnTo and op, the account routes. The contract commit's message still names the
metadata; this commit is the correction.

--- fix(connectors): an MCP target survives the card's Continue, and the model is not sent to an action that refuses it

The review of the MCP backend found eight defects; each one is a test first.

- The off-desktop note told the model to confirm with action 'status', which refuses MCP
  names outright. It now says the authorization finishes in the background and the tools
  arrive on the next turn.
- The card's existence is a property of the session. A missing callback no longer routes a
  desktop session down the off-desktop path, where the model would be handed a live link.
- A second failure of a backend attempt was published with actor 'user'. ``refresh`` takes
  the actor, so the frame says who produced the text.
- Continue on the RPC thread can settle the operation between any read of a target's state
  and the transition that follows it. A settled operation has a frozen result, so the lost
  move is dropped; an IllegalTransition no longer escapes into the tool result. The same
  rule covers a worker whose outcome arrives late and a Try again that arrives after Continue.
- Try again on an install carried no credentials, so the second install ran with an empty
  env. The approved map is kept on the runner (never on the target: target fields are
  serialised to the model) and reused.
- An approval that does not cover a required credential no longer reaches the worker, where
  ``install_entry``'s prompt would block on stdin forever; the row waits for the card's
  fields instead.
- ``enable`` writes ``mcp_servers`` under the scope and lock the dashboard's toggle route
  uses, so the two read-modify-write paths in one process cannot drop each other's write.
- Several authorize targets start their flows together and share one wait; one wait per
  target kept the card empty for minutes.
- The off-desktop operation is in no session's registry, so it now emits no connection.update.

Try again is also refused for a target the transition table cannot move back to 'initiated'
(an MCP target has no move out of 'expired') and for an operation that settled during the call.

The contract test that froze the Actor enum is replaced by two behaviour tests: no actor but
the backend watcher can connect an MCP target, and the card's word never moves one.

(cherry picked from commit 081c16412b48667758a77a7e7c29eaf2be6a93dd)

--- fix(connectors): the watcher reads one account route per target, and every frame carries its seq

The watch loop walked the whole toolkit list once per tick to learn whether one
target had connected. That read costs a vendor call per page, cannot tell one
account from another, and forced `awaiting_new_attempt`: after a forced
reconnect the list still reported the OLD account `active`, so the row had to be
disbelieved until it read as something else once. The gateway now serves
`GET /v1/connectors/accounts/{connectionId}`, so each pending target reads its
own account: the id the mint named, one read per target per tick at 1 Hz (the
route's bucket is 180/min per principal), each bounded by the operation's
remaining deadline. A 429 parks that one target until its Retry-After passes and
leaves the others reading. A 404 is "not yet" until the deadline. A target the
mint gave no account for has nothing to read, so it is not read. A forced
reconnect watches the new account, which is why `awaiting_new_attempt` and its
two tests are gone.

`mint` and `run_remote` now name where the browser should come back to
(`returnTo`, plus the operation id on a mint), so the vendor's done page can
hand the user back to the desktop app that asked instead of stranding them on a
web page. Only the desktop registers that URL scheme, so no other surface sends
either field.

`connection.update` frames were built by re-reading the operation after the lock
was released, so a second writer could give an older frame a newer state and the
renderer could not tell which frame was last. Every write now advances a
monotonic `seq` and takes its snapshot under the same lock, and the emitter sends
that snapshot; a renderer that keeps the highest seq per operation can drop a
frame that arrives out of order. `connectors.operation.wake` lets the desktop's
deep-link return ask for a read now instead of at the next tick; it only shortens
the wait and trusts nothing else in the link.

(cherry picked from commit f5423cf72302b92b26c043495184ad6426fa2b36)

--- fix(connectors): the desktop card follows the operation, and says so out loud

The MCP card was a second implementation of the connector card with the
review's defects: it painted for any request on the session, kept its
controls after the operation settled, offered an approve verb while the
backend was still minting an authorize link, and re-enabled Install when
the RPC returned rather than when the state frame moved the row. A second
click in that window sent the consent twice.

The renderer now reads the operation's `seq`: an update or a status frame
whose sequence is not greater than the one already applied changes
nothing, and a resume snapshot neither revives a settled card nor puts
back a row a newer frame has moved. Without it the transport's ordering
decided what the user saw.

`hermes://connections/done?op=…` brings the user back from the browser to
the session that opened the operation and wakes its watcher, so the row
moves at once instead of at the watcher's next tick. Only the op id is
used; the link's status moves no row.

Each row's mark and cue are one polite live region, so a row that flips is
heard and not only seen, focus follows the row the backend moved while the
card holds it, and the waiting mark stops spinning under reduced motion.

(cherry picked from commit 7c30d252556d7d496ddb354b50bcecc41bceef27)

--- fix(connectors): the renderer types seq as the wire carries it

The watcher commit made `seq` a required field on every operation frame. The
desktop store still declared its own optional `seq` so it would compile against a
backend that predates the field; that backend no longer exists on this branch, so
the hedge is dead code and the fixtures were short one field. The store now reads
`seq` from the shared types, and the fixtures count the way the backend does.

The repeat-frame test asserted the whole request keeps its reference. With a real
rising `seq` the request must change; the invariant the test guards is that the
target row keeps its identity so open credential inputs do not remount.

--- fix(connectors): a URL that arrives after the shared wait still lands on its row

The shared prepare wait failed every row still pending when it ran out, while
that row's own thread was still waiting on the provider. When the URL arrived a
moment later the thread's move raised inside the daemon thread, the link was
lost, and Try again started a second flow. The wait now bounds only how long
prepare blocks; a row still pending afterwards is left to its own thread, which
is the only writer of that row and ends with the URL or the flow's own failure.

The comment on the approved credentials said every target field is serialised to
the model; it is not (the snapshot names its keys). The reason they live on the
runner is that they are secrets and the runner's life is exactly the operation's.

--- fix(connectors): the review findings the operation must survive before the fold

The MCP prepare threads and the install worker started with an empty context,
so a named-profile turn's home override never reached them: the flow resolved
the process home's `mcp_servers` and stored the token there. Each thread now
runs in a copy of the calling thread's context. `connection.respond` had the
same gap on the RPC thread: an approval ran the enable, which writes
config.yaml, with no profile bound, so the flag landed in the launch home. The
handler binds the session record's profile the way `_connector_rpc` does;
`config_write_scope(None)` keeps that override, so the enable needs no change.

The watcher raised out of the tool when a per-target Skip landed while that
target's read was in flight: the operation stayed open with `live` closed, and
every later answer got 4004. A read for a row that is no longer live is dropped
at debug; only a refusal on a live row is still a contract violation. The same
skip from the card raced a row the backend had just connected and aborted the
rest of the answer; a skip for a resolved row is ignored and every entry, then
the settle check, still runs.

A read was bounded by the whole remaining deadline, so a hung gateway held the
first read for 300 s and Continue could not return the tool; a read now waits
ten seconds at most. A 429 parked only the target that read, but the budget is
the principal's, so the next target's read in the same tick spent it again:
every live target waits out the one Retry-After.

The MCP surface rule read the platform alone, so a desktop call without the
callback (registry dispatch from execute_code) opened an operation nobody
rendered and blocked for the deadline; it now uses the managed rule, surface
and callback. A write after settlement advanced `seq` while emitting the
frozen frame, so the resume snapshot named a seq no frame carried; the counter
stops at the settle frame. A mint that reports `initiated` with no account id
logs that the watcher cannot read the row.

(cherry picked from commit 587b228016bf8ee2b022e7fa58c2d9f6057809e9)

--- fix(connectors): the desktop card holds a verb until the backend answers, and never takes the keyboard from a credential field

The review of the desktop card found five defects; each one is a test first.

- A resume snapshot was refused whenever the cache held a settled operation, whichever
  operation it was, so a session that opened a second operation after settling the first
  never got its card back from a resume. And the refused snapshot handed the caller the
  settled cache as "the request", so the session was flagged as needing input behind a
  summary with no controls. Only the same operation can refuse the snapshot now, and a
  refused one is no pending card.
- The focus handoff picked the row's first button, which after pending -> initiated is the
  disabled working verb; the focus call was a no-op and the keyboard landed on the document
  body. It picks the first control that can take focus, else the row. It also moved focus out
  of a credential field the user was typing in whenever another row moved; it leaves an
  editable alone. The "focus Continue once every row resolved" branch was dead (Continue
  unmounts the moment nothing is unresolved), so it and its ref plumbing are gone.
- The done link navigated to a settled operation's session and rejected when the wake RPC
  did (4004 once the operation left the live registry). A settled request is ignored, and a
  refused wake is nothing: the wake only shortens the wait, the watcher still ticks.
- Install spun forever when the store refused to send the consent (the operation gone or
  settled under the card): `respondToConnectionRequest` resolves false in that case and the
  verb was only released in `catch`.
- After a partial approval (a required credential missing) the backend answers with a
  same-state frame whose detail names what is missing; nothing released the verb because it
  was held until the row's state moved. The row now remembers the seq the click saw and holds
  the verb only until a frame past it arrives, which is the backend's word on the click
  whether or not the row moved.

(cherry picked from commit 4c534d6e2af979778d9a2423cf97d717edeaa1a9)

--- fix(connectors): a skip that loses the race to the watcher is ignored, not raised

The skip guard read the row's state and then moved it; the watcher can connect the
row between the two, and the refused move aborted the rest of the card's answer.
The refusal itself is now the witness: a move refused for a row that is resolved,
or on a settled operation, is the same nothing-to-do as a row resolved earlier.

Two recording fakes in the managed tests kept their lists on the class; they now
start per instance so a lifted fake cannot share reads between tests.

--- fix(connectors): a resume that lands behind a newer live frame still reports the pending card

The refused-snapshot branch answered "no pending card" for both reasons it can
refuse: the operation settled, or a newer frame already moved a row. Only the
first is no card. For the second the live card is still open and blocking the
turn, so the caller must keep the session flagged as waiting on it.

* test(connectors): defer new connection coverage until implementation settles

Remove PR-added test cases and their unused helpers while retaining
existing tests adapted to the changed connection contract. The three
PR-only renderer test files are removed for now.

Focused behavioral coverage will be added as the final implementation
step before verification. Existing main coverage is not being removed
wholesale, and this does not declare the feature merge-ready.

* fix(connectors): commit MCP authorization at initialize, save setup values after success

The OAuth probe treated one exception as one outcome: any failure after the browser
step restored the token snapshot and manager entry, so a server that accepted the
token but failed tools/list discarded a completed consent. Now the probe reports
whether initialize succeeded (details["initialized"], read from the claimed
MCPServerTask). Failure before that point rolls back as before. Failure after it
saves the server config, keeps the tokens, and reports tools unavailable through
flow.discovery_error; the card can retry discovery without repeating consent.

Catalog install wrote the submitted values to .env before install_entry and the
probe ran. The values now live in the secret scope for the duration of the install
(get_env_value reads through get_secret, so install_entry finds them without a
prompt), and .env is written only after the probe returns tools. A failed probe
removes the server block and writes nothing. _probe_tool_names returns None on a
failed probe instead of an empty list, so failure and a valid empty listing are
distinct.

Failure text is redacted before it reaches target detail: every value the card
submitted for that target is replaced by exact match, then the pattern redactor
runs.

required_env now carries the manifest's secret and default flags; Target carries the
manifest's post_install text as instructions. The wire contract gains secret,
default, instructions and discovery_error; generated TS and OpenRPC regenerated.

* fix(connectors): one OAuth callback receiver picker for the connection card

The card's authorize target built its redirect from the dashboard web server and
raised when none was bound in the process, so a standalone hermes --tui session
could never authorize an MCP server. The receiver is now chosen in one place
(choose_callback_receiver): a pinned pre-registered client keeps the SDK's own
listener on the registered port; a client-advertised loopback URI is used as-is and
its callback arrives through the mcp.servers.oauth.callback relay; otherwise the
backend binds a one-shot loopback listener and feeds it into the flow. The dashboard
route stays with the dashboard web page, which cannot bind a port.

tui_gateway/mcp_oauth_sessions.py had a second copy of the loopback listener and a
_worker that referenced _probe_with_rollback, set_hermes_home_override,
reset_hermes_home_override and Path without importing them, so every RPC-started
flow raised NameError. Both are deleted; start_flow spawns run_worker directly and
uses the same receiver picker. Flow registration is shared (register_flow /
finish_flow) so a card-started flow with a client URI is reachable by the relay.

Under an SSH session with no client listener the attempt's detail carries the
existing paste-the-redirect instructions. Nothing on the card path opens a browser.

* feat(desktop): setup-form modal for MCP connection cards

The connection card rendered an MCP server's setup fields inline: every field as a
password input, no default value, no instructions. A URL such as the n8n MCP
server URL was typed blind, and the manifest's setup text never reached the user.

Two components carry the form now. SetupFieldList renders the ordered fields the
backend declares (a plain field as text prefilled with the manifest default, a
secret field masked and empty). SetupFormDialog composes it with a one-line title,
the manifest instructions, an inline error for a failed attempt, and Cancel /
Connect. The row's Install action opens the dialog when the target has fields;
Connect sends {status: approved, env}; Cancel sends {status: skipped}. A failed
attempt keeps the dialog open with the draft intact. The draft lives in the dialog
component only; nothing reaches the store or the resume snapshot.

Once the backend publishes the authorization URL the dialog shows it as text with
an Open in browser button. Nothing opens a browser on a state change: the Try
again path on both cards used to open the re-minted link at once; it now waits for
the row's update frame and the user's click.

The store types gain the wire's secret, default, instructions and discoveryError
fields and normalise them; a connected target with discoveryError renders as
authorized with tools unavailable.

* feat(tui): connection card in the Ink TUI

The Ink TUI had no client for the connection operation: connection.request,
connection.update, connection.respond and the pending_connection resume snapshot
were unhandled, so a manage_connections call in a TUI session could only print a
link through the model.

connectionOperationStore.ts holds the backend snapshot: a request opens only for a
new operation id, an update applies only to the live operation with a higher seq,
a settled update freezes the id so a late request frame cannot reopen the card.
The gateway event handler feeds it; session resume hydrates it from
pending_connection.

connectionSetupOverlay.tsx renders one callout in the prompt zone: a one-line
title, the manifest instructions, every field (plain rows prefilled with the
default, secret rows masked), then a Connect / Cancel selector. Connect sends the
draft through connection.respond; Cancel skips the active target. When the backend
publishes the authorization URL the callout shows it as text under "Press Enter to
open in browser"; Enter is the only thing that opens it. A failed attempt unlocks
the fields with the draft kept. An accepted secret renders as "Set" and is never
echoed. A target authorized without tools shows that state and Continue.

The overlay joins the existing input-owner set in overlayStore so typing and other
prompts are blocked while it is up.

* feat(cli): connection panel in the classic CLI

The classic hermes CLI passed no connection_callback, so a manage_connections
call could only print an authorization link through the model and could not take
a setup value at all.

The CLI now renders the connection operation as a prompt_toolkit panel, on the
same queue mechanism as the clarify panel: the agent thread's callback opens the
panel, blocks until the first decision, then returns so the operation's watch loop
runs; every later state reaches the panel through ConnectionOperation.on_change
(installed only when no gateway hook is set, restored on close).

Panel: one-line title, the manifest instructions, one row per field (plain rows
prefilled with the default, secret rows masked in the input buffer and rendered as
"Set" once accepted), then Connect / Cancel. Connect sends the draft through
apply_answer; Esc skips the active target; Ctrl-C sets the tool-thread interrupt so
the operation settles as interrupted. Once the backend publishes the authorization
URL the panel shows it with the target detail and "Press Enter to open in browser";
Enter is the only thing that opens it. A failed attempt unlocks the fields with the
draft kept; an authorized target without tools offers Retry discovery / Continue.

Up/Down move between rows, Left/Right toggle the action, and the panel joins the
blocking-overlay guards so chat input, history and voice stay out while it is up.
The single-query (headless) mode passes no callback, as it does for clarify.

* feat(connectors): register a connected MCP server and report its tools in the result

After a successful install or authorization the target carried the probe's tool
names and the model was told the tools "become available on your next turn". MCP
tool schemas are deferrable by construction, so nothing about them lives in the
sent tool array; a server registered in the scoped registry is callable through
tool_describe/tool_call in the same turn. The operation now registers the server
(register_mcp_servers under the owner's home scope) once authorization is
committed, records the registered names on the target, and the settled result
carries a tools_listing block in the deferred-catalog format plus a note that the
tools are callable now. A registration failure keeps the target connected with
tools: [] and a sanitized discovery_error. agent.tools and the system prompt are
untouched; the between-turns refresh updates the catalog block as before.

The card gate no longer asks for the desktop platform. Every surface that renders
the card attaches a connection callback (Desktop, the Ink TUI, the classic CLI);
registry dispatch and messaging sessions attach none and keep the link result.

Pre-commit rollback in probe_with_rollback used restore(only_if_absent=True),
which skips the rollback when a token file exists. On a first-ever authorization
the only file is the one this attempt wrote, so a token the resource rejected was
kept. Live E2E (controlled provider answering 403 to the issued token) showed the
row fail and the token survive; the rollback now restores the snapshot outright,
and the same run shows the token file removed.

* fix(connectors): install an OAuth catalog entry through the card's own flow

The card's install ran install_entry and then a plain probe. For an OAuth entry
that probe has no card flow around it: in the desktop backend it failed at once
("non-interactive environment and no cached tokens"), and in the classic CLI it
saw a TTY, opened a browser by itself and drew install_entry's curses tool
checklist over the panel. Since a probe failure is now an error, the failed
install was rolled back with _remove_mcp_server, and authorize refuses a server
that is not configured. A clean home had no path to a connected OAuth entry, which
is 55 of the 65 catalog entries. Found by the live three-surface run.

An install now builds the entry's configuration in memory (card_install_config:
no prompts, no probe, no checklist; a prior tool selection or the manifest's
curated filter applies). An OAuth entry starts the same flow authorize uses with
that configuration and the setup values in the attempt's secret scope, so the row
reaches the URL step, and the configuration and the setup values are saved
together when initialize accepts the token. Every other entry is probed in memory
and saved after the server answers. A failure writes nothing, so a failed
reinstall keeps the previous configuration, and the failed row asks for its fields
again so the card can reopen the form over the draft it kept.

* fix(agent): make a server connected in this turn callable in this turn

tool_describe and tool_call resolve names inside the agent's toolset selection,
which is fixed when the agent is built. A server that manage_connections had just
registered was therefore "not found" for the rest of the turn whose result calls
its tools available, and stayed out of the next turn's catalog too. The live
Desktop run showed it: the result listed mcp__fx_oauth_fields__echo and the
tool_describe that followed answered not_found.

The executor now adds the MCP servers the call connected to the selection. Only
the selection changes; agent.tools does not, so the sent tool schema bytes stay
the same. A selection of None (every toolset) and the no_mcp sentinel are left
alone.

* fix(cli): reopen the connection form when a required field is still empty

When the backend refuses an answer because a required field is missing it keeps
the row pending and names the fields. The panel mapped that frame to its waiting
phase, which draws the detail line and nothing else, so the user saw "waiting for
FX_API_KEY" with no fields and no buttons until the 300 s deadline. The panel now
returns to the form on the first missing field, over the draft it kept.

* test(connectors): the no-card path is the one with no callback attached

The card gate is "a connection callback is attached" since the classic CLI got
its own panel. This test still attached one under platform "cli" and expected no
card, so it opened a real operation, waited out the 300 s deadline and failed.
It now drives the path that has no card: no callback.

* fix(connectors): order a cancel against the commit and stop deleting a working grant

Found by the live runs on Desktop, the Ink TUI and the classic CLI.

A user's skip did not stop the OAuth attempt. With the worker parked in the token
request, the row settled "skipped" and the token was written 33 s later. A skip
now cancels the attempt, and one lock orders that cancel against the commit: the
attempt is either canceled with the earlier tokens restored, or committed and
kept. When the commit won, the skipped row says the authorization was kept. An
interrupted turn cancels its attempts the same way.

Every attempt began by deleting the saved tokens, so retrying discovery for an
authorized server demanded consent again, and a cancel in between left no grant
at all. The card's flow now connects with the saved tokens first, with no browser
step; only when they do not work does it replace them. The RPC session surface
keeps the old behavior because its caller waits for an authorization URL.

A second Connect from the form a failed row reopened was dropped without a frame,
which left the Desktop dialog with Connect loading and Cancel disabled until the
deadline. An approval on a failed or expired row is now a retry with the new
values.

The reported tool names came from the registration call, which returns nothing
for a server the process already holds. A retry after a failed listing therefore
said "no tools" while the server had them. The names are read from the registry,
and a parked server is woken first.

An authorization that commits after its card closed by deadline was never
registered, so the next turn still could not reach it. The runner keeps such
attempts, and the between-turns refresh adopts the ones that were approved.

An error with no message reached the user as a class name ("CancelledError").
The tool description still said a server's tools arrive on the next turn.

* fix(desktop): let an OAuth install open its link, and keep the form's draft

An install of an OAuth catalog entry with no setup fields reached the URL step
and gave the user nothing to click: an initiated install was always drawn as a
disabled spinner, and the only other place the link is shown is the setup dialog,
which opens for entries with fields. That is most of the catalog. An initiated row
that carries a link now offers Open, whatever the action.

The setup dialog reset its draft whenever the field list changed identity. The
backend sends a fresh list with every frame and an empty one while an attempt
runs, so a failed Connect erased what the user had typed. Fields now only fill in
what the draft lacks, and closing the dialog drops the draft.

The settled summary dropped the "tools unavailable" fact, and the live row offered
a Try again for that state which the gateway refuses for a connected target. The
summary keeps the fact and the row offers no dead control.

* fix(tui): keep the typed draft when a Connect fails

The overlay reset its draft whenever required_env changed identity, and every
backend snapshot delivers a freshly parsed array. A failed Connect therefore came
back as a form with the default region and an empty secret. The draft now resets
per target only, and a snapshot's fields fill in what the draft lacks.

* fix(cli): show the URL step and start an install that has no fields

The panel applied the user's answer and then set its phase to "waiting". The
backend applies the answer on the same thread and its change hook had already set
the next phase, so the URL step of an OAuth install and the reopened form for a
missing field were both overwritten, and the card sat on "Waiting…" until the
deadline. The waiting phase is now set before the answer is applied.

A pending install or enable with no fields opened in the waiting phase, which
sends no approval, so the flow never started. It now opens on Connect/Cancel. A
target that already carries its link opens on the URL step.

Connect on a failed row re-ran the attempt without the values now in the draft; it
sends them. The authorized-without-tools phase offered a "Retry discovery" that a
connected target cannot run inside the same operation; it offers Continue.

* fix(connectors): a newer OAuth attempt replaces the older one for the same server

A card that closed by deadline leaves its worker waiting on the browser for up to
300 s, and a tampered callback leaves one waiting too. A retry or a new operation
for the same server then ran beside it: both wrote the same token files, and the
older one's rollback could write over the newer attempt's grant.

The newest attempt per home and server is recorded. Starting one cancels the older
attempt, takes over its pre-attempt snapshot so a later failure still restores the
state from before either, and the older attempt's rollback leaves the files alone.

* fix(desktop): label an install's link control "Open in browser"

The control that hands an install's authorization link to the browser reused the
action's verb, so the row showed "Install" before the click and "Install" again at
the link step, told apart only by the cue. It now reads "Open in browser", the
label the setup dialog already uses for the same act. Authorize keeps its verb.

* docs(mcp): describe the setup card on the desktop, the terminal UI and the CLI

The MCP guide said the chat install exists only in the desktop app and that the
CLI relays commands. All three surfaces now show the same card: fields, Connect or
Cancel, an authorization link the user opens, one save when the server has
accepted the token, and tools the agent can call in the same turn. The tools
reference gains the result fields (tools, tools_listing, discovery_error) and the
no-card behavior of an OAuth install.

* fix(cli): keep the layout hook callable without a connection widget

The CLI panel commit added connection_widget to _build_tui_layout_children as a
required keyword. That method is the documented override point for wrappers, and
five existing tests (extension hooks, prompt stash, subagent dock) call it without
the new argument, so CI failed with a TypeError. The argument is now optional, like
the other widgets added after the hook was published; a missing widget is left out
of the layout.

The settled tool result also carried the target's setup instructions. Those are
the card's text for the user, and a catalog entry's notes can predate this flow
("restart your session so the tools are loaded"), which contradicts a result that
says the tools are callable now. The model-facing result drops them; the card
payloads keep them. This restores test_connector_local_batches.

* fix(connectors): wait for an in-progress registration before reporting a server's tools

Found with the real Vercel MCP server. Saving the configuration wakes the config
watcher, which starts its own connect for the new server. The operation's
registration call then skips the server as "already connecting" and returns at
once, so the card settled "connected" with no tools, and the 214 tools were
registered three seconds later. With no names in the result the model searched,
found the hosted connector of the same vendor and asked the user to connect that
instead.

The registered names are now read once the registration has finished: while
another task is connecting the server the read waits (30 s at most), a parked
server is woken once, and a server that finished registering with no tools is
still a valid empty list.
2026-09-21 10:04:32 +05:30
teknium1
5dd70d7cb6 fix(desktop): say when the launch-config reader ignored a desktop block, and pair space-separated argv switches (#77311)
The pre-window config reader is a deliberate YAML subset (two-space keys,
four-space `- ` items), so any other valid indentation silently disabled
both launch keys — the app just started with no heap ceiling and no way to
tell. And the argv scan read `--name=value` only, so `--js-flags
--max-old-space-size=4096` lost the value and registered a bogus
`max-old-space-size` switch, letting config.yaml's ceiling replace the
launcher's instead of merging with it.

- one warning naming the supported shape when a `desktop:` block mentions
  either key and nothing parses; a block with no mention stays silent.
- argv pairs a bare `--name` with a following non-flag token, and
  `--js-flags` with whatever follows it (its value is dash-prefixed).
- website/docs/user-guide/desktop.md documents the supported indentation
  next to the snippet, plus the warning users will see.
2026-09-20 20:50:54 -07:00
teknium1
7d2d634708 fix(desktop): renderer heap ceiling and desktop.electron_flags reach packaged launches (#77311)
`desktop.electron_flags` was appended to argv only by the `hermes desktop`
launcher, so a packaged app started from its Start-menu / .desktop entry
never saw it, and nothing set `--js-flags` at all — three Windows reporters
on #77311 confirmed `jsflags=NO` on the renderer command line and OOM
crash-loops with 15 GB of free RAM. Chromium copies `js-flags` to renderer
processes only from the browser's pre-`ready` command line, so main.ts now
reads the two launch keys out of config.yaml itself (no YAML dependency in
the main bundle: a deliberately narrow subset reader) and applies them via
app.commandLine before ready.

New typed knob `desktop.renderer_max_old_space_mb` (default 0 = Chromium
default) becomes `--js-flags=--max-old-space-size=N`, MERGED with any
existing js-flags (appendSwitch replaces, so a naive apply would drop one).
The planner is pure (`electron/renderer-heap-flags.ts`) and vitest-covered.
2026-09-20 20:50:54 -07:00
teknium1
845fee6cf4 feat(website): a page for every catalog plugin and every author
/docs/plugins/<name> and /docs/plugins/by/<maintainer> are generated at
build time by a small Docusaurus plugin (website/plugins/plugin-catalog-pages)
from the same plugins.json the grid fetches, so a merged catalog PR is the
only way a page appears or changes.

Plugin page: full description with the Disclosure sentence pulled into a
callout, pinned commit / version / platforms / requires facts, tools, hooks
and env chips, Desktop install button + CLI command, repository and reviewed
source links, optional screenshots gallery, optional README rendered from the
PINNED commit (fetched at build, allowlisted HTML, raw HTML dropped, relative
links and images resolved against the pinned tree), and a "More by this
author" shelf. Author page: everything a maintainer has in the catalog with
total stars and a profile link when all repos share one owner.

Cards on the grid link through: the title is a link and a card click opens
the page (the Desktop picker embed keeps expanding in place). The shared
vocabulary (entry type, tier/category taxonomy, link builders) moved to
src/components/PluginCatalog/catalog.ts so the three surfaces cannot drift.
Docs: field table and catalog README describe screenshots:/readme: and the
pages.
2026-09-20 20:43:13 -07:00
teknium1
fb9c77e064 docs(bot-mode): held messages are replayed after release; Detect stop directives toggle
User-visible behaviour from #117488/#117472 and the quoted-stop-word rule from #117605.
2026-09-20 20:41:43 -07:00
teknium1
e089be80a1 docs(desktop-sdk): note DecodeText loop is opt-in
DecodeText is a published plugin-SDK export; flipping its `loop` default
to false silently changes third-party plugin visuals, so the SDK doc says
so where the component is listed.
2026-09-20 20:14:20 -07:00
teknium1
968422e7c5 fix(gateway): keep the operator restart tail on the home-channel storage notice
The cause-table action is user-phrased ("Send your message again once compression finishes"),
so the OPERATOR notice lost "then `hermes gateway restart`" for store-level failures that stay
broken until the gateway is restarted. The tail is appended for every cause except the
session-scoped ones that clear on their own (compression, compression_closed, turn_lease).

Also: the held-store refusal test is parametrized over optimize / optimize-storage / prune
(optimize-storage, the command the issue names as the field producer, was uncovered) and
asserts the refusal names the same store SessionDB opened — no `hermes sessions` subcommand
can point the command at another database. Docs: doctor refuses the checkpoint only while it
can see a process holding the RETIRED log.
2026-09-20 20:13:33 -07:00
teknium1
6ba45b0e06 fix(sessions): storage maintenance refuses while a writer holds state.db; human-first retired-WAL guard text + recovery guide
`hermes sessions optimize`, `optimize-storage` and `prune` now run the same fail-closed
holder scan doctor and repair use before rewriting the store. While a gateway, Desktop,
dashboard or cron process holds state.db (or a WAL sidecar) they print each holder as
`PID N (command)` with the stop remedy and exit 1; `--force` overrides with a warning,
`--dry-run` previews are never gated. The Desktop console's `sessions optimize` gets the
same refusal. Why: a user ran `optimize-storage` under a fleet of eight live gateways and
every agent answered every turn with the retired-WAL refusal until all writers were
stopped by hand (#110054, maintainer follow-up 09-20).

The DeletedWalGenerationError text is now two layers: a first sentence for the person
reading a chat bubble or banner (what happened, nothing is lost, quit every Hermes
process on the profile, `hermes doctor` names the holders, never `doctor --fix` or delete
files while they run, docs link), then the operator detail. The classifier fingerprint
"deleted state.db-wal or state.db-shm" is unchanged. The cause table
(`hermes_state_user_copy`, feeding the CLI banner, TUI/Desktop RPC error and the gateway
home-channel notice) and the chat explainer carry the same first steps; the gateway
notice no longer hardcodes `doctor --fix` + `gateway restart` for every non-corrupt cause,
which for a held retired generation is the second-writer trap.

New user-guide page `session-storage-recovery.md` (registered in sidebars, linked from the
guard text, the developer state-db-recovery page and the sessions guide): the three steps,
the do-nots, why maintenance refuses, and what the files beside state.db are
(retired-wal captures + manifest.json, pre-update-emergency backups, corrupt backups,
snapshots).
2026-09-20 20:13:33 -07:00
teknium1
12fb5eb476 fix(sessions): make set-journal-mode portable and fail closed where it cannot prove quiescence
Review follow-ups on the new `hermes sessions set-journal-mode` verb:

- The header probe used os.pread, which does not exist on Windows, while the subparser is
  registered unconditionally — the command died there with an uncaught AttributeError. It now
  reads the 20 header bytes through a plain binary open(), and the tests no longer skip on win32.
- foreign_state_db_holders() returns [] unconditionally on Windows (no scan), which made the
  admission gate vacuous: an operator got a silent all-clear and could flip the mode under a
  running gateway. Windows now refuses outright, naming the reason, overridable only by --force.
- Enabling WAL ignored the cross-VM filesystem refusal the runtime enforces
  (apply_wal_with_fallback). target=wal now refuses on virtiofs/9p, where WAL shared memory
  silently corrupts.
- A --db pointing at a garbage file surfaced a raw sqlite3.DatabaseError traceback even though the
  header probe had already read not-a-database, and a directory raised IsADirectoryError. Both now
  bail in the command's own error style; every open/read is guarded.

The admission checks that need no I/O live in a pure _refusal() that takes the platform as data,
so the Windows and cross-VM invariants are tested without faking sys.platform.
2026-09-20 20:13:25 -07:00
teknium1
96da5d97fc feat(sessions): hermes sessions set-journal-mode delete|wal converts an existing WAL store offline (#100896)
`database.journal_mode: delete` can never self-apply to a store that is already WAL:
apply_wal_with_fallback deliberately never live-downgrades (#68545 — other gateway/cron/worker
connections may hold uncheckpointed WAL commits), so operators applying the containment for the
multi-writer corruption class saw one ERROR per process forever and the only escape hatch was an
undocumented hand-run PRAGMA on the file.

The new pre-DB `sessions set-journal-mode` verb is the sanctioned offline path: it refuses while ANY
foreign process holds the file or a sidecar (the same foreign_state_db_holders scan doctor/repair
admission uses, naming each PID), flips through _set_journal_mode_no_wait (busy_timeout=0, so an
opener appearing mid-way makes SQLite refuse instead of racing it), verifies header bytes 18/19,
and reminds the operator when config.yaml disagrees. `--db PATH` covers kanban.db / cron stores
that log the same ERROR. The never-live-downgrade invariant is untouched; the ERROR, doctor hints
and docs now name the command instead of the raw PRAGMA.
2026-09-20 20:13:25 -07:00
teknium1
bdde0b0c28 docs(cron): describe when a stale-code ticker yields and when it keeps dispatching
The yield predicate now requires a live, fresh-heartbeat gateway whose stamped
code_sha is the on-disk revision; a lock held by an equally stale process never
counts. Placed under "Gateway Integration" so it does not collide with the
"Stale-code yield" section #117501 adds under "Locking".
2026-09-20 20:13:19 -07:00
xxxigm
050ea53aba fix(desktop): honor Applies-to on Custom Endpoints
The page only showed a read-only active-profile note, so endpoint
saves followed the left-rail Bot instead of the Settings chips
Accounts and API keys already share.
2026-09-20 22:46:49 -04:00
teknium1
c545272568 fix(lsp): one stalled request no longer silences a workspace for good — retry window, cold-root warm-up budget, per-root exclusion
A language server that missed its budget once marked its (server, root) pair broken for the
process lifetime, the same 5 s steady-state budget was applied to a cold server that also had
to spawn, initialize and build its program, and the only escape (servers.<id>.disabled) switched
the server off for every workspace.  Three new keys under the existing `lsp` block, all defaulting
to today's behaviour:

- lsp.broken_retry_seconds (0 = lifetime): the broken set stores a monotonic retry deadline per
  pair; an expired pair gets one more try, and the INFO skip line names the retry time.
- lsp.warmup_timeout (0 = wait_timeout): the first request against a root with no running client
  waits up to this budget (outer join budget follows); warm requests keep wait_timeout.
- lsp.exclude_roots ([]): glob patterns matched against the resolved project root (a bare path
  also covers everything beneath it); a matching root never spawns, logged once at INFO.  A
  non-list value fails closed — WARNING naming the expected shape, every root skipped — because
  silently excluding nothing would re-pay the stall the key was meant to avoid.

Part of #116446 (the diagnosability slice landed in #116839, salvage of #116459 by @kokhlo).
2026-09-20 18:54:24 -07:00
teknium1
eaef5ec717 feat(update): hermes update --list-venv-holders prints the venv guard's holders as JSON, exit 3
Scheduled `hermes update --yes` runs on Windows loop against the venv-holder guard when the
Desktop app relaunches its backend, and the refusal text is the only clue. The new read-only
flag runs the same scan (_detect_venv_python_processes, late-bound through hermes_cli.main)
and the same classifiers (pausable-gateway matcher, _hermes_holder_subcommand) and prints
[{pid, exe, argv, kind}], exiting 0 when the venv is free and 3 when holders remain, so
automation can stop exactly those PIDs and retry. Nothing is terminated; the flag is
handled in the update preflight before the lock, backup, or any mutation. Off Windows the
guard never fires and the list is [].

Fixes #117246
2026-09-20 18:51:36 -07:00
teknium1
5be70b5e25 docs(mcp): note Google-hosted OAuth servers get access_type=offline for a refresh token 2026-09-20 18:22:18 -07:00
teknium1
86a599cbae fix(gateway): human-delay pacing comes from each profile's human_delay config, not process env
BasePlatformAdapter._get_human_delay read HERMES_HUMAN_DELAY_MODE/_MIN_MS/_MAX_MS from the
process environment at every send, so under multiplexing the launch profile's pacing applied
to every served profile, and the documented `human_delay:` config section (mode/min_ms/max_ms,
already in DEFAULT_CONFIG) was never consulted. The runner now resolves `human_delay` per
profile through the same seam as the busy-text timings (`_human_delay_from_config`,
snapshotted in `_snapshot_profile_busy_modes`, installed by `_wire_adapter_handlers`) and the
adapter only consumes the installed range. Invalid `custom` bounds (non-integer, negative,
inverted) warn naming the key and fall back to the natural range.

Fixes #116895
2026-09-20 17:01:05 -07:00
teknium1
5a074a02cb fix(gateway): busy-text debounce/hard-cap come from per-profile config, not process env
BasePlatformAdapter read HERMES_GATEWAY_BUSY_TEXT_{MODE,DEBOUNCE_SECONDS,HARD_CAP_SECONDS}
at construction, freezing the launch profile values into every profile adapter under
multiplexing. The mode was already re-synced per profile by the runner; the two timing
knobs were not. They are now display.busy_text_debounce_seconds /
display.busy_text_hard_cap_seconds, snapshotted per profile next to the busy modes and
installed by _wire_adapter_handlers. Invalid values warn naming the key and fall back.

Fixes #116893
2026-09-20 17:01:05 -07:00
teknium1
f568b860d7 fix(agent): pin the reset_at reach through try_activate_fallback and add opt-in fallback.min_switch_reset_seconds
- The direct _arm_rate_limit_cooldown test now drives agent._try_activate_fallback
  (production entry) on a real AIAgent with a one-entry chain, so dropping the
  reset_at forwarding goes red (3 failures before, 8 green after).
- #117484 knob: fallback.min_switch_reset_seconds (DEFAULT_CONFIG 0 = off). When the
  rate-limited primary's declared reset is sooner than N seconds,
  try_activate_fallback returns False and no cooldown is armed; docs row added.
2026-09-20 17:00:43 -07:00
teknium1
53815e24dc fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.

The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.

Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
  before  req_reasoning: {}
  after   req_reasoning: {'reasoning_effort': 'medium'}
  agent.reasoning_effort: low  ->  {'reasoning_effort': 'low'}  (unchanged)
2026-09-20 16:03:38 -07:00
teknium1
4f0ced3e5f fix(desktop): mask typographic double quotes too and document quoted stop words
macOS smart-quote substitution turns the straight quotes a user types into “ ” in the composer,
so the quoted-span mask covers both; the user-visible rule (stop words inside code, quotes or
blockquotes never hold) now has its sentence in bot-mode.md alongside the behaviour.
2026-09-20 15:58:58 -07:00
Jony
2539a08fd1 docs(gateway): clarify systemd reload semantics 2026-09-20 15:58:02 -07:00
teknium1
97beaeeeb3 feat(delegate): warn the child at 80% of its inactivity window before abandoning it
A child that stalls under a configured delegation.child_timeout_seconds used to
learn about the budget only by dying, losing its whole context. The liveness
wait now queues a one-line "[delegation budget warning]" through the child's
steer channel once the idle window is 80% spent (delivered at the child's next
iteration boundary), so a slow-but-recoverable child can wrap up and return its
summary. The warning fires once per idle window and re-arms when progress
resets the window; a progressing child never sees it.

Part of #116001 (atom 2A). Semantics of child_timeout_seconds are unchanged.
2026-09-20 15:50:20 -07:00
finn763
03973bd02b fix(delegate): child_timeout_seconds bounds inactivity, not total runtime
The configured cap was a dispatch-to-death stopwatch: `await_child` waited on a
plain `settled.wait(timeout=child_timeout)`, so any child that outlived the
budget was abandoned even while the provider was actively serving it.

The report's corpus for #116001 (219 tasks / 75 deaths, 0 of them mid-tool) could
not be reproduced here — it needs the reporter's slow OpenAI-compatible endpoint
— but the mechanism it names is exactly this gate: a child waiting on an
in-flight LLM completion, killed with a nearly-finished context. A slow child is
already bounded elsewhere (the per-call stale watchdog, the heartbeat's
staleness verdict), so this cap could only ever kill children the runtime had
judged healthy.

`child_timeout_seconds` now measures time with NO progress: the wait runs in
slices and restarts the window on the same signals the heartbeat's stale verdict
reads (completed call, tool change, activity-clock tick). A frozen child is
still abandoned when the window elapses; a progressing one is never killed for
taking long.

Timeout entries also carry `last_event_age` (how long the child had been silent),
so operators can tell a slow provider from a runaway without transcript
forensics.

Fixes the mechanism reported in #116001. The budget warning and
continuation-respawn items in that issue are separate features and are not part
of this change.
2026-09-20 15:50:20 -07:00
teknium1
7d8e8f234f test(gateway): pin the send_voice is_voice contract through the base media dispatch; document Matrix audio attachments 2026-09-20 15:48:51 -07:00