Desktop love/hate telemetry needs a backend half: six counters
(feature_use, action_use, mode_use, friction, dislike, onboarding) with
closed dimension sets, the v3 JSON schema entries, the public doc rows,
and profile-scoped gateway RPCs the app reports through.
Every value the app sends is collapsed server-side onto the closed set
(unknown -> other); feature use latches per area per UTC day, friction
and dislike are capped per day, the daily report (action counts + mode
time + bot count) latches per day and answers recorded=false on a store
failure so the app keeps the day for a retry. For setting changes the
app sends only the config key: the backend reads the saved value and
compares it to DEFAULT_CONFIG itself, so values never cross the wire;
keys outside DEFAULT_CONFIG read other. All helpers check enabled()
before doing any work.
- m9: shared_metrics.startup_latency had no server-side latch, so any repeat
call added a row. The TUI and Desktop now send an opaque per-launch id
(TUI: one per process; Desktop: minted when the once-per-launch main-process
claim succeeds) and the backend claims (surface, launch_id) once per
process: a reconnect re-sending the same launch counts once, a new Desktop
launch against a long-lived backend still counts. Legacy clients without an
id count once per backend process and surface. The id is only a local latch
key (bounded set, length-capped), never recorded. Only a usable measurement
spends the claim. Contract + apps/shared regenerated.
- m10: relaunch() execs in place, keeping the PID, so `sessions browse` ->
resume counted the picker time as CLI startup. relaunch() stamps its PID in
HERMES_RELAUNCHED_PID right before exec; process-start surfaces skip when it
matches, while children (other PIDs) inheriting the env still count.
- m8 (psutil half): record_process_ready checks the opt-in gate before reading
the process start time (psutil) or starting a thread. The claim is taken
first so a later opt-in never records a stale "startup" mid-session.
Test changes to existing assertions: the TUI/Desktop param assertions now
expect launch_id (a new wire param); the RPC surface test sends distinct
launch ids so its three calls stay three launches under the new latch.
Desktop's packaged updaters (electron-updater, App Installer, Store) never run
`hermes update`, so the receipt cannot count them. Main persists a pending
record before apply (the installer may quit the app before apply returns),
decides a hand-off on the next launch (version changed => success), and the
renderer sends it once the gateway is open; the file is removed only after the
RPC resolved. Checkout hand-offs are excluded: their receipt already counts
them as kind=desktop. The backend buckets raw words into bounded dims.
Two product questions shared metrics could not answer: how long each surface
takes from launch to usable (so startup regressions show per release), and how
current, on which channel and on what class of machine installs run.
hermes.startup.latency {surface, latency_bucket} records once per process start:
cli (process creation -> first rendered prompt, or -q dispatch; Kanban workers
excluded), gateway_boot (-> GatewayRunner.start done), serve_boot (-> hermes
serve listening), and tui / desktop_attach reported by the clients through the
new shared_metrics.startup_latency RPC. The clients declare their surface
because a Desktop on a URL/cloud backend has no HERMES_DESKTOP there; env
detection is the fallback for older clients. In-process surfaces measure from
psutil's process create time, the earliest timestamp available, and hand the
runtime start to a daemon thread under the caller's context so no event loop
waits on it. Everything goes through _emit, so disabled profiles record nothing.
The install snapshot gains release_channel, version_age_bucket, behind_bucket,
ram_bucket, gpu_class and local_model_provider_used. All are read offline:
the installed commit's own date, the channel record / packaged channel / checkout
branch (never the remote URL or branch name), and the update check's existing
cache for this exact revision (never a network call). Rows counted before these
fields existed stay valid as a legacy field set.
Three shared-metrics counters that answer "which models misbehave, frustrate
users, or run out of room", all attributed to catalog provider/model names
(custom endpoints and loopback servers collapse to custom) and all behind the
existing enabled() gate.
hermes.model_tool_quality.count {provider, model, call_role, issue}
Counts every tool call a model emits where the agent validates it, clean calls
as issue=none so the rates have a denominator: invalid_json, unknown_tool,
schema_mismatch (missing required keys / non-object), empty_arguments (only for
tools with required params), repaired (Hermes fixed the name or the streamed
argument JSON and ran the call). Stream assembly marks args it repaired and the
chat transport carries the marker onto the normalized ToolCall, because
normalization otherwise erases it.
hermes.model_friction.count {provider, model, signal}
retry / undo / interrupt / quick_abandon / switch_away, blamed on the model that
produced the turn: the relay session remembers its last primary route, so a
/retry after a /model switch still counts against the retried model. Counted
where the action executes, once: CLI handlers (skipped on the TUI slash worker's
shadow CLI), tui_gateway command.dispatch retry/undo and session.undo (Ink
/retry now sends intent=retry, so it counts as a retry, not an undo), gateway
/retry and /undo (multiplexed runners bind the owning profile home), and every
/model surface via record_model_switch(from_model=...). Interrupts and quick
abandonment (session closed within 60s of a failed turn) come from the runtime's
turn close, for attended entrypoints only; a turn still running when the
session closes is neither.
hermes.context_peak.count {provider, model, peak_fill_bucket, window_bucket, limit_hit}
One row per closed top-level session: the fullest primary context it reached
(post_api_request now carries the compressor's context_length) and whether a
call was rejected as too large (context_overflow / payload_too_large, the
rejections Hermes answers with a forced compression). A session whose every
call overflowed still reports, with unknown buckets.
The gateway counted slash commands at its client boundary (successful
slash.exec / command.dispatch), so every command the Desktop or Ink TUI
handles locally (/new, /branch, /skin, /resume, overlays, pickers) never
reached shared metrics. Counting moves to the clients' dispatcher entry:
the new fire-and-forget RPC records {command, execution_surface} for one
user-typed command (surface desktop|tui from the same client detection),
scoped to the session's profile, always answering {ok: true}.
_count_slash_command is removed from rpc_dispatch so each typed command
lands exactly once, from the client. Contracts regenerated.
Desktop had no way to read or change the shared-metrics opt-ins. The new
profile-scoped RPCs read/write telemetry.shared_metrics.{enabled,send} in the
focused profile's config.yaml through the one config writer, keep the setup
wizard's invariant (send is forced off unless collection is on) and reconcile
the consent windows on every change via the wizard's own helper. first_run
answers also record the desktop setup-completed metric (lazy import until the
events module lands).
Trim the salvaged handler/test prose to the salvage bar and make the
contract match the handler: SessionSetHiddenParams.hidden is required
(no default=True), so the generated TS/OpenRPC types stop advertising an
optional flag whose absence used to hide the whole compression lineage.
Regenerated apps/shared/src/gateway-contract.*.
* feat(docker): publish nousresearch/hermes-sandbox:desktop for terminal backends
The terminal backends (docker, modal, daytona, singularity) all default to
nikolaik/python-nodejs:python3.11-nodejs20, a bare Python+Node base. For Bot
Screen, computer_use and the browser to run INSIDE that sandbox instead of on
the gateway host, the sandbox image needs the display stack.
docker/sandbox-desktop.Dockerfile is that base plus:
- the everyday tools it lacked (jq, ripgrep, fd, tmux, less, nano, vim,
zip, rsync, tree, procps, htop, sudo for the base's uid-1000 `pn`)
- the exact package set the Hermes -desktop image installs (TigerVNC,
Xfce components, dbus, xauth, fonts)
- Playwright's headed Chromium (same build as the -desktop image)
- cua-driver 0.28.2 from its pinned release tarball
No Hermes inside; the default user stays root like the base so nothing
changes for people who just switch docker_image. Desktop processes run as
`pn`. 4.27 GB on amd64.
docker.yml gains a `sandbox` variant with its own cache scope and repository
(nousresearch/hermes-sandbox:desktop, :main-desktop, :<release>-desktop);
the docker-integration suite is skipped for it (no Hermes to test) and
docker/sandbox-desktop-smoke.sh runs instead: as `pn`, every launcher and
cua-driver binary resolves, the real launcher.sh publishes :20, the RFB
socket completes the 3.8 handshake relayed over `docker exec -i` stdio, and
a headed Chromium maps a window on that display. hadolint lints the new
Dockerfile in docker-lint.yml.
* feat(docker): sandbox desktop base on python3.13-nodejs26
Matches the Hermes image (Python 3.13 / Node 26) and the top of requires-python;
the default docker_image tag it inherited was Python 3.11 / Node 20. Same pn
uid 1000, Debian 13; smoke (launcher, RFB relay, headed Chromium) passes.
* feat(docker): bake agent-browser into hermes-sandbox:desktop
The browser tools drive the agent-browser CLI; when the browser follows the
terminal backend that CLI has to exist inside the sandbox. Pinned to the same
^0.26.0 range the gateway resolves, --ignore-scripts like the gateway's npx path.
* feat(bot_desktop): the screen, computer_use and the browser follow the terminal backend
A user who sandboxes `terminal` (docker/ssh/singularity) had the agent's
screen, cua-driver and Chromium running on the gateway HOST beside that
sandbox: Bot Screen gave a headless host a display, the Xfce panel carries
xfce4-terminal, and `computer_use` could open a shell outside the boundary
the sandbox exists for.
Now the desktop lives where the terminal lives:
- tools/environments/streams.py: one primitive per spawn-per-call backend, the
local argv prefix that runs its remainder inside the sandbox with stdio open
(`docker exec -i`, `ssh`, `apptainer exec`). SDK backends (modal, daytona,
vercel) have none and report so.
- tools/bot_desktop/sandbox_host.py: launcher.sh runs inside the sandbox as
the image's `pn`; the pane's RFB bytes ride a 12-line python relay over that
prefix; `cua-driver mcp` is the prefix + the sandbox image's own driver.
- tools/bot_desktop/placement.py + `bot_desktop.placement` (auto|terminal|
gateway). `auto` follows the backend; a sandbox that cannot host a screen
REFUSES with the opt-in named instead of silently using the host.
- runtime.start/stop/status/published_env branch on placement; the pane,
lease, epoch fencing and CLI are unchanged.
- cua_backend: the MCP invocation is the sandbox one when the screen is
there; the host driver's runtime contract is irrelevant then; check_fn is
true under a terminal placement without a host binary.
- browser_tool_session: agent-browser invocations are wrapped in the prefix
with the daemon, socket dir and profile inside the sandbox; screenshots are
fetched back so MEDIA: paths keep working; recycle closes the sandbox
daemon.
- web_routers/display.py: the bridge pumps a relay's stdio when the screen is
in a sandbox, a unix socket otherwise.
Live on docker with nousresearch/hermes-sandbox:desktop: start/observe/RFB
handshake through the dashboard bridge, human takeover fences the agent
(HumanHasControl) and keystrokes reach the sandbox Xvnc, handback restores,
three start/stop rounds leave zero desktop processes; computer_use capture
and list_windows see only the sandbox's Xfce; browser_navigate/snapshot/
vision run with Chromium and agent-browser inside the container and zero
host processes on the bot profile; modal + auto refuses naming the opt-in.
* feat(desktop): Screen pane shows where a sandbox-placed screen runs; Install is host-only
DesktopStatus gains placement ('gateway' | 'terminal:<backend>'). A sandbox
image lacking the stack is a blocker naming hermes-sandbox:desktop, shown in
place of Start; install_command stays None there because the pane's Install
button runs the package manager on the gateway host, the wrong machine, and
display.install refuses for the same reason. The pane header carries
'Screen runs inside the docker sandbox, with the terminal' (4 locales).
* fix(bot_desktop): "is the screen in the sandbox" is a disk check on hot paths, never a config read
Every browser command and CUA spawn asked in_sandbox(), which resolves placement by
loading config, which initializes HERMES_HOME. Under a test's fake home that raised
HomeInitializationError from _run_browser_command; on a real host it read config per
click. Hot paths now ask sandbox_screen_running(): the start marker on disk, written
only by a sandbox start. Policy (in_sandbox) stays for start/install, where config is
the question. display.observe gates on "an RFB endpoint exists" for either placement.
* test(moa): late-accounting sink test asserts the wedged slot's row, not sink order
Under CI load the poll loop can see the interrupt before collecting the fast slot, so the
fast slot also arrives late and first; the test then failed on sink_calls[0]. The
contract is that the wedged slot's real usage reaches the sink.
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* fix(bot_desktop): a sandbox that died under a live screen fails loudly, never falls to the host
sandbox_screen_running() drops a start marker whose terminal environment is no longer
registered (stale after a process restart). When the environment object outlives its
container, the browser's sandbox wrap now checks the published DISPLAY and raises
"the screen inside the terminal backend's sandbox is gone; start it again" instead of
KeyError('AGENT_BROWSER_PROFILE'). Live: fresh sandbox navigate ok; docker rm -f the
container; next navigate returns that error; zero host Chromium either way.
* feat(terminal): nousresearch/hermes-sandbox:desktop is the default container sandbox
Every container backend (docker, modal, daytona, singularity) now defaults to the
sandbox image with the desktop stack, so Bot Screen, computer_use and the browser
run inside the sandbox for everyone who never chose an image; Python 3.13 / Node 26
match the Hermes image. One constant (DEFAULT_SANDBOX_IMAGE) replaces six copies of
the old literal. Migration 47 moves saved configs still holding the OLD default and
never touches an image the user pinned. Docker reuse recreates a container built
from another image, or the flip would silently never take effect for anyone with a
persisted container (live: old container removed, new one on 3.13 / Node 26 with
Xvnc, cua-driver, agent-browser present).
* chore: retrigger CI (zero-job dispatch failure)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
* chore(config): template stamps v47 and shows the new default sandbox image
The template is what install.sh / docker / doctor --fix seed; a stamp behind
DEFAULT_CONFIG makes every fresh install migrate on first run.
* feat(sandbox): the default image change is a decision, not a surprise
A persisted Docker sandbox on another image is kept when docker_image is unset;
only a written docker_image (an explicit pin) recreates it. The pin verdict
travels as TERMINAL_DOCKER_IMAGE_PINNED through both terminal bridges (process
env and per-profile scope) and the container-config allowlist.
Approval surfaces, all through hermes_cli.sandbox_image_switch: the interactive
CLI asks once at startup (y = pin the new image, n = pin the current one, Enter =
ask later); the Screen pane shows the same choice with Switch / Keep buttons via
display.switchSandboxImage; `hermes config set terminal.docker_image …` is the
same answer from any shell. Gateways and cron never decide: they keep the sandbox
and log the notice.
Migration 47 now unsets a saved image equal to the OLD default instead of
rewriting it to the new one — that value was the template copied, not a pin, and
rewriting it would have made the runtime recreate existing sandboxes unasked.
Modal restores its snapshot and Daytona reuses its labeled sandbox regardless of
the configured image, so existing sandboxes there were already untouched.
* fix(config): both plain defaults that preceded the desktop sandbox image are template copies
main pinned nikolaik/python-nodejs:python3.14-nodejs22 (cd0f97f833) without a migration;
a saved config holding either literal is unset by migration 47, so it follows the default
and existing sandboxes get the keep-or-switch decision instead of a silent recreate.
* ci(docker): build the sandbox image on release/dispatch, not every main push
Leaves docker.yml exactly as on main. The sandbox image carries no Hermes code,
so two 4 GB multi-arch builds per merge bought nothing. sandbox-image.yml builds
and smokes on a PR that edits its own Dockerfile/smoke, and publishes only on a
release or a manual dispatch with publish=true. Stable tag stays :desktop.
* fix(config): keep main's config.py/config_defaults.py edits under the sandbox-image delta
The rebase resolved both files wholesale with the branch side, dropping main's move to
hermes_yaml (the 3.14 runtime venv has no PyYAML) and the 3.14 base pin. This is main's
version plus exactly the branch's own changes: DEFAULT_SANDBOX_IMAGE, the pin verdict in the
env bridge, placement defaults and the v47 stamp.
* test: sandbox-image tests read config.yaml through hermes_yaml (no PyYAML on the 3.14 runtime)
* fix(bot_desktop): read the sandbox marker BOM-tolerantly (windows footgun lint)
* fix(bot_desktop): placement is the authority; sandbox screen survives restarts
Review findings on the sandbox-hosted Bot Screen, each reproduced live first.
Authority. The browser preflight and the CUA invocation keyed off screen
LIVENESS, so `placement: terminal` with the screen not yet up handed the tool an
unchanged host command. `runtime.tool_placement()` is now the one resolver:
terminal placement starts the sandbox screen on demand (no auto_start opt-in
inside the user's own sandbox), refused placement raises its reason, and neither
ever yields the host. placement.resolve() answers a local backend from env alone
so the common case costs no config load on the spawn path.
Restart. sandbox_screen_running() deleted the marker whenever the process-local
terminal registry was empty, i.e. after every gateway restart, while Xvnc kept
running in the container; stop() then returned False and left it. The marker
now records the owning container; liveness comes from `docker inspect` on it,
stop/status re-attach to the recorded owner (even after the placement setting
moved), and only a container that is gone drops the marker.
SSH. remote_argv emitted `bash -c <script>` as three words; OpenSSH joins them
and the remote login shell ran `bash -c export` and the rest itself. The script
travels as one quoted word for ssh (remote_command knows the backend); docker
and apptainer keep argv.
CDP reach. agent-browser inside the sandbox reports the sandbox's loopback;
the Browser Use harness, browser_exec and the vault supervisor connect from the
host and got connection refused. streams.forward_port() proxies a local port
over the exec stream (same relay as the RFB bridge) and the CDP URL is rewritten
to the local end.
pids limit. --pids-limit 256 counts threads; measured on the desktop image the
desktop stack is 44, one Chromium tab 212, the agent's browser with two tabs
488. Past the cap every further docker exec died with "procReady not received".
Default is 2048 with the measurements in the comment.
Replacement. An approved image switch force-removed the old container before
`docker run` tried the new image; a private tag or registry outage left nothing.
The image is inspected/pulled first and the old container kept on failure.
Desktop integration. The sandbox start never passed the dock's browser launcher
(no Browser icon) and the thumbnail needed a host launcher pid + host ImageGrab
(always None). The dock runs the sandbox's Playwright Chromium on the shared
profile; the thumbnail is grabbed inside the sandbox (Pillow baked into the
image). The browser profile moves from /tmp — a 512 MB tmpfs emptied on every
container stop — to the desktop user's home, so logins follow the container.
Pin provenance. A TERMINAL_DOCKER_IMAGE written in a routed profile's .env is a
pin even when it spells the default; the scope compared values before.
* docs(bot-screen): no literal tmp path in the profile-location note
* fix(bot_desktop): docker inspect liveness probe closes stdin (TUI subprocess guard)
* fix(bot_desktop): adopting a screen the sandbox kept records the marker
Live ssh probe: after the host's state was lost while the sandbox kept its
Xvnc, start() took the idempotent early return (display already published)
and never wrote the host marker, so status/thumbnail/stop lost the screen.
Record the adopted display like a fresh launch.
Docs: what an ssh host of your own must carry, and why a Dockerfile ENV is
not enough for a login session (PLAYWRIGHT_BROWSERS_PATH via /etc/environment).
* docker(sandbox-desktop): login sessions find the browser (PLAYWRIGHT_BROWSERS_PATH via /etc/environment)
* docs(bot-screen): what the Apptainer path inherits from the image and what it does not
* chore(config): sandbox-image migration is 47→48 (main took 47 for compression.threshold_tokens)
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
Root cause: a draft has no session_id, so use-slash-completions sent commands.catalog /
complete.slash with empty params and tui_gateway's _session_home_scope (session-only
profile_home) scanned the LAUNCH profile's skills; CompleteSlashParams also rejected a
`profile` key outright (extra=forbid -> 4000).
Fix: the hook sends the tile's routed profile when there is no session (cache keyed by
session-or-profile); _session_home_scope resolves a session-less `profile` through
_profile_home; `profile` added to CompleteSlashParams and SkillsReloadParams (same
session_id-or-empty shape in skills.reload) and the TS contract regenerated.
Red-on-base: tests/tui_gateway/test_protocol.py::test_sessionless_slash_palette_follows_profile_param
fails on origin/main (complete.slash -> 4000 "profile: Extra inputs are not permitted");
use-slash-completions.test.tsx "scopes a sessionless catalog and typed lookup to the routed
profile" fails on origin/main (commands.catalog called with {} not { profile }). Both green on head.
_history_dict_text has always had both image renderings, but
_coerce_message_text hard-coded image_urls=True at both call sites, so the
reference-form branch was unreachable from the API: a remote client re-read
the sum of every attachment the conversation ever carried (26.30 MiB on the
measured session.resume, paid again on every reconnect) with no way to ask
for less. session.resume now accepts inline_images (default true) and
threads it through every path that carries messages — cold/lazy/eager
resume, live reattach via _live_session_payload — routing to the existing
[image] branch when false. The REST twin lands in the follow-up commit.
Fixes https://github.com/NousResearch/hermes-agent/issues/116511
Desktop archives sessions via PATCH /api/sessions/{id} -> set_session_archived,
but the TUI gateway had no equivalent JSON-RPC method, so TUI clients had to
keep every session visible or delete it permanently. Open PR #47184 added this
against the pre-split server.py layout and resolves only live runtime sessions;
this lands the method on current main next to session.set_hidden with the same
two-tier resolution: live runtime id first (unpersisted drafts defer via
pending_archived, applied by _ensure_session_db_row like pending_hidden), then
a stored id/key resolved through the profile db, covering the historical rows
the session.list picker shows. Contract declared in tui_gateway/contracts and
rendered into the generated TS/OpenRPC catalog.
Fixes https://github.com/NousResearch/hermes-agent/issues/47168
ensureContrast's accumulating ladder loop trusted its `step` argument: step=0
never advanced (1,000,001 iterations, still unbounded), negative steps walked
away from the pole forever, NaN skipped the loop body entirely so the original
below-threshold color was returned silently, and Infinity/above-1 steps jumped
straight to a full blend. Normalize instead: a finite step in (0.001, 1] is used
as-is (clamped to 1 above that), everything else falls back to the default 0.2
ladder. The desktop 0.2 and TUI 0.05 float sequences are byte-identical.
Fixes#109955
Salvaged from #109975 by kokhlo
projects.db is written by the CLI (hermes projects create), other
Desktop windows, and workspace-settings flows — processes that never
touch this gateway's transports, exactly like state.db. The change
watcher had no projects.db signature, so those writes left the Desktop's
project list and folder tree stale until an unrelated refresh (#53046,
Add a projects.changed watch (2s interval, newest mtime across
projects.db + WAL for the watcher home and every served sibling profile,
mirroring the sessions.changed contract), register the event in the
gateway contract (generated TS/OpenRPC refreshed), and route it through
the desktop lifecycle handler into a $projectsChangeTick that
refreshes both the projects list and the sidebar tree.
Fixes#53046Fixes#56757
Co-authored-by: LeonSGP43 <LeonSGP43@users.noreply.github.com>
The branch-render guard landed earlier (78b133127d); what remained was
provenance: a projected continuation tip rendered as a brand-new
session. list_sessions_rich now stamps continuation_kind='compression'
on projected rows, StoredSessionRow carries it (contracts regenerated),
and the sidebar session row labels continuations with their lineage
root instead of a fresh-conversation affordance. The live tip stays
listable while naming its parent; sealed chain segments stay hidden.
Co-authored-by: wave-2 worker D <wave2d@hermes-triage>
The client rejected its own approval.respond RPC after the generic request
timeout (120s shared default, 30s desktop) while the backend still waits the
full approvals.timeout (300s) for the user: answering later surfaced a false
"request timed out" over an approval the backend then applied anyway.
approval.respond now carries APPROVAL_RESPOND_TIMEOUT_MS (300s, the backend
default); ambientRequestFor forwards the optional deadline so session-routed
approval RPCs can raise their timeout; the hermes-bots group-approval path
rides the same deadline.
Fixes#60654
The web dashboard supports customCSS in theme YAML (PR #14776) but
the desktop app's skin pipeline dropped the field. Users who wanted
custom styling had to hack app.asar, which gets overwritten on every
update.
This adds customCSS passthrough through the full skin pipeline:
- HermesSkin (apps/shared) and DesktopTheme types gain customCSS?: string
- skinToDesktopTheme() passes skin.customCSS through
- applyTheme() injects a scoped <style id="hermes-desktop-custom-css">
tag on theme apply, and removes it when switching to a CSS-less theme
- SkinConfig gets custom_css field, read from YAML as "customCSS"
- _build_skin_config() caps at 32 KiB (same as the web dashboard)
- resolve_skin() emits "customCSS" in the gateway JSON-RPC payload
Users now put CSS in ~/.hermes/skins/<name>.yaml under the customCSS
key — it persists across updates because the skin dir is outside app.asar.
Closes#53013
Refs #53012
The desktop ships cwd_explicit provenance with session.create (#52589),
but the wire contract (extra=forbid) rejected the unknown key, so every
remote-mode create failed with 'out of sync (different versions)' and the
composer never accepted the draft (Desktop core E2E remote-topology).
Local-backend lanes never sent it (a bare new chat is detached, no cwd).
Regenerate the shared contract TS + OpenRPC from the updated registry.
After the locate click, type now refuses unless the located editable is
document.activeElement, so characters are not delivered as page input when
focus stayed on body. The keystroke loop checks an abort signal between
characters; preview.act cancel (timeout or interrupt) and a local Stop both
set it, so queued keystrokes stop. A printable press on body/html is refused
unless allow_shortcut is set.
A non-git project folder still gets an isMain trunk lane from the backend's
path-only heuristic (pinned by test_existing_non_git_workspace_still_becomes_a_project),
but the renderer treated every isMain lane as a branch target, so clicking "+"
ran `git switch <folder-name>` and died with "fatal: not a git repository".
The placement now carries isGit (false only for the heuristic fallback lane);
the wire contract and the desktop lane type pass it through, and the sidebar
skips the pre-create branch switch when the lane is not a git checkout. The
lane itself, its sessions, and the new session landing in the folder are
unchanged.
project_create (desktop_project tool / CLI) persists to the per-profile
projects.db but writes nothing to state.db, so sessions.changed never fires
and the desktop Projects sidebar goes stale until a manual refresh.
The gateway change watcher now watches projects.db and broadcasts
projects.changed when it moves; the desktop live-sync routes the event to a
new $projectsChangeTick that refetches the project list + tree.
Fixes#56757
MessageCompletePayload is extra="forbid", so the forwarded flag raised
ContractViolation under test isolation and logged a wire-contract violation
on every transformed turn. Regenerate the shared TS/OpenRPC contract and
type the desktop field from the generated MessageCompletePayload.
Rebased onto main, where params are validated against tui_gateway/contracts
and unknown keys are rejected: the copy_parent_history / omit_messages flags
on session.create and session.branch now answered 4000. Keep both methods'
wire shape unchanged and give the two new methods their own contracts
(session.branch_stored, session.branch_whole) with messages_omitted results;
the handlers take the behaviour as keyword arguments, not wire flags.
Sabotage showed the barrier's bounds were unguarded: making the
timeout/unsupported-method fallback never settle, or lifting the 10s
replay deadline, left every suite green, as did deleting the pre-read
barrier await in reconcileActiveTranscript.
Restore Bear's original timeout/unsupported it.each from 8ee10f65931
(both must resolve true and release parked frames within the bound), and
add one reconcileActiveTranscript case proving a background refresh does
not read REST until the reconnect replay has dispatched.
A backend restart announces a new replay_epoch over the still-open
socket (gateway.ready or the replay response). adoptReplayEpoch revoked
the pending replay holds by resolving their barriers false, which every
caller treats as "socket lost, abandon the read". Nothing re-arms that
read: the socket never dropped, so no new open follows. The reconnect
backstop read (#94779) was silently discarded and turns the old backend
committed while Desktop was disconnected never reached the transcript.
On an epoch change no replay can cover the old numbering, so REST is the
only recovery left: resolve the waiting barriers true. Their
continuations still run after the parked live frames dispatch. Socket
invalidation keeps resolving false, since the next open re-reads.
Keep the two files that pin the ordering contract: the shared client's
replay barrier (installed before open listeners run, resolves only after
replayed and parked frames dispatch, false when its socket is lost) and
the warm-resume ordering. Drop the renderer hydration suite and the
formatting-only churn in the replay test; the barrier contract they
exercised is covered at the client boundary.
After a reconnect, REST history could paint a completed turn before the
socket dispatched that turn's ordered replay. The replay then appended the
same progress and final rows again, and the unchanged-history signature
shortcut left the duplicates in place. Sequence watermarks order replayed
frames against live frames, but not REST publication against replay.
JsonRpcGatewayClient now installs a per-session replay barrier before
synchronous socket-open listeners run. Each session settles on its own
once its replay and parked live frames have dispatched. Timeout and
unsupported-method fallbacks resolve true. Socket invalidation and epoch
revocation resolve false, so pending reads are abandoned. Connect still
does not await replay, and the existing replay deadline still applies.
pendingSessionReplay() inspects existing sockets without dialing. The
active, tile, fallback and warm-resume hydration paths wait on it before
activating or reading history, check again when the read returns, and
recheck selection and ownership after waiting.
Carved out of #119186; its fold/dedup half is superseded by #120809.
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
The backend stamps `seq` once per event before its transport fan-out, so the
same frame arriving on two of the renderer's sockets to one process carries
the same (session_id, seq); `JsonRpcGateway` is per socket and cannot see the
other copy. `JsonRpcGateway.dispatchEvent` now tags each event with the
`replay_epoch` its socket adopted from `gateway.ready` (per process, not per
socket), and both fan-ins in `use-gateway-boot.ts` (primary + registry
secondaries) pass through one `createGatewayEventDedupe()` gate keyed by
(epoch, session_id, seq) before any store runs.
A duplicate is the same key within 30 s. Not "seq not above the highest
seen": the backend restarts a session's counter at 1 when the session leaves
its 64-session replay ring, and a high-water mark would swallow the whole
restart of any short session. Seq-less / session-less events always pass.
Bounded: 2048 seqs per session, 256 sessions LRU.
Live, with the router fix reverted so two sockets still join the chat: the
DOM never doubled a word (0/38 samples vs 4/40 on main) and the doubled
"reply was cut short" card is back to one.
Closes#120007. Part of #120005.
NS-960. One card, one group ("connectors"): the curated catalog plugins
this OS runs lead the hosted connectors (D1, D4). A plugin whose app is
absent is greyed with the reason and stays pickable (D5). Picks land in
answers.plugins and ride the existing [setup] note to the guide.
The guide's runbook gains the install beat as the last thing before the
handoff card (D2, D3): narrow the picks to what the chosen task needs, then
ONE manage_catalog install call. A new suggested first task sits beside the
existing ones and installs its plugins through that same beat: "Set up my
games and streaming" (NVIDIA App + Broadcast) on a Windows PC with an NVIDIA
GPU, "Help me make something in Blender" everywhere else.
When the guide's install card settles, each plugin row's outcome
(installed / failed / skipped, whatever the user did) is written into the
answers (D6). The build session's runbook names what is ready, what was
offered and not installed, and what was picked but not offered, and tells
the build agent never to install. profiles.remember_onboarding records the
picked plugin names in the default profile's memory, next to the connectors.
The onboarding card needs to list catalog plugins beside the hosted
connectors (NS-960 D1, D4) and grey a plugin whose app is absent (D5).
The catalog had no curated flag, and the manage_catalog row left
app_state empty.
- `onboarding: true` and `title` on catalog entries (loader, validator,
docs); set on blender, nvidia-app and nvidia-broadcast.
- hermes_cli/plugin_catalog_presence.py reads the plugin.json at the
catalog's pinned commit once per pin and judges its app declaration with
the hermes_platform resolver the installer and the Plugins-tab pill use.
No declaration or an unreadable one is `unknown`, never `present`.
- `plugins.manage action=onboarding` lists the curated entries this OS
runs (platform mismatch is the only exclusion) with app_state and the
sentence the card greys the row with.
- manage_catalog plugin rows now carry app_state and the catalog title.
- A live entry that differs from the in-tree entry at the same pin (new
metadata) now follows the same newer-catalog rule as a new pin, so a
checkout that adds `onboarding` is not masked by a published doc that
predates it.
Installing a plugin from any surface now makes its MCP servers and skills
usable in every open chat of that profile on the chat's next turn. There is
no Connect-now button, no /reload-mcp, no relaunch.
- hermes_cli/plugins_activation.py::load_and_go_live: after the forced plugin
rescan, connect the plugin's portable MCP servers one by one
(register_mcp_servers, the connector flow's call), then refresh that
profile's open chats and queue a turn note listing the servers, their
tools and the plugin's skills. activation gains live_now; deferred keeps
only Python tools and prompt sections.
- MCP tools are deferred behind tool_search / tool_call, so appending them
(preserve_prefix) leaves the model-facing tool array and the cached prompt
prefix unchanged; the note rides the existing one-shot turn-note channel,
so the system prompt stays byte-stable. /reload-mcp keeps its consent gate:
it is a full rebuild.
- tui_gateway/methods_tools.py: one session walk (_refresh_live_sessions)
shared by reload.mcp and plugin activation, filtered to the plugin's
profile home.
- hermes plugins install/enable (another process) asks the running Desktop /
dashboard backend to do the in-process half through the new
POST /api/dashboard/agent-plugins/activate, found via the host rendezvous
record and its session token.
- Desktop: the Connect-now toast and its strings are removed in all six
locales; the install toast says what went live ("3 tools connected ·
skill X ready") and warns per server that did not connect.
- Contract: PluginActivation.live_now; regenerated TS/OpenRPC.
- register_skill docstring: plugin skills are listed by skills_list.
* feat(desktop): catalog install card rows for plugin and skill targets
manage_catalog opens the same connection operation manage_connections does,
with rows of kind plugin/skill. The renderer collapsed every non-connector
kind to 'mcp', so those rows would have drawn as MCP servers with no name,
purpose or provenance.
- Contract: ConnectionTargetKind gains plugin/skill; ConnectionOperationTarget
types the catalog fields agreed in CATALOG-ROW-CONTRACT.md (display,
description, tier, platforms, repo, sha, subdir, scan, requirements,
has_desktop_half, target_profile, app_state, skill).
- Store: kinds map through a lookup; catalog rows carry a typed CatalogEntry.
- CatalogRow: kind glyph, name, kind/tier/platform pills, one-line purpose,
then Install / Advanced / Skip. Installing shows the backend's per-row
detail; installed shows "N tools · show names"; failed shows the reason
and Try again (approved again for that row only); skipped stays quiet.
- Advanced: a centred dialog with the Plugins-tab controls (halves, target
profile, enable, force, pin) plus read-only source and scan. Its Install
sends the values in the answer env; the row's Install sends env null.
- A missed tool.start restores the card on a manage_catalog row.
* fix(desktop): catalog row uses the app's sliding progress and one pill style
The installing row drew a full-width pulsing line, which reads as a finished
rule, not work in progress; it now uses the Progress primitive's sliding
indeterminate block (reduced motion honoured there). The kind pill was
all-caps monospace next to two sentence-case pills; all three now share one
shape, matching the V1 offer artboard, and the platform pill stays the one
coloured note. The Advanced modal's Source and Security tables share a fixed
label column, and requirements render as prose, not code.
* fix(desktop): a skill's Advanced shows only the controls a skill has
A skill has no agent/desktop halves, no enable step and no commit pin, so
its Advanced dialog offered three controls that do nothing. It now shows the
target profile, force reinstall, source, security and credentials, and sends
only target_profile, force and the credential values. The preparing line
takes the card's width instead of the whole chat column.
* feat: setup profile is minted by the backend and found by role, not by name
The guided onboarding runs in a profile the desktop used to create itself
(profiles.create with a soul, "already exists" treated as success) and
recognise by the literal "hermes-setup". The upcoming setup toolset grants
catalog installs to that profile, so the marker that grants it must be
written only by the backend.
- profile.yaml carries `role: setup`; read_profile_meta / write_profile_meta /
ProfileInfo know it; profiles.list and GET /api/profiles report it.
- hermes_cli/setup_profile.py: ensure (find by role, adopt a pre-role
hermes-setup dir, else clone default + soul + role) and reset (soul,
memories, skills back to the created state, in place). The soul text moves
here from the renderer.
- tui_gateway/methods_onboarding.py: onboarding.ensure_setup_profile and
onboarding.reset_setup_profile. Neither takes a name; profiles.create and
profiles.configure already reject `role` (unknown key, 4000).
- Copies never inherit the role: --clone-all, profile import, and a
distribution that ships profile.yaml drop it.
- setup.status for a named profile reports `ready` once the boot bootstrap
settled. Since one host backend serves every profile (#118246) the
desktop's setup-profile probe lands on this branch, which never set
`ready`, and the kickoff waited forever.
- Desktop: SETUP_PROFILE, ensureSetupProfile(profiles.create) and
composeSetupSoul are gone. store/setup-profile.ts holds the name the
backend returned (or the roster's role row after a relaunch); kickoff,
handoff and the build card use it. The dev reset calls the reset RPC.
* fix: write the setup soul as bytes so Windows keeps \n line endings
* refactor(desktop): drop the renderer's setup-profile store; the backend is the only owner
Kickoff reads the name straight from onboarding.ensure_setup_profile and records it on
$setupSession, which every later step already carries. The handoff recovery check reads
the roster row's role. No renderer module holds a setup-profile name or a fallback lookup.
A plugin that finished loading after an adapter connected never got its platform
handlers (slash commands, button callbacks, inbound transforms) registered until a
gateway restart, silently. Three pieces, one seam shared by every surface:
1. Discovery listener: PluginManager.on_plugin_loaded(cb) fires from INSIDE
discover_and_load for the plugins a sweep newly loaded (diff of the loaded set),
with a per-plugin activation summary (hermes_cli/plugins_activation.py):
activated_now {gateway_commands, gateway_transforms, hooks, callbacks} vs
deferred {tools, prompt, mcp_servers}. Every mid-run load path now performs a real
discover_plugins(force=True): CLI install/enable (via the gateway), Desktop/TUI
plugins.manage install/toggle/update, dashboard REST install, tool-triggered
force re-discovery, the new `reload-plugins` control-socket verb. A non-forced
discover_plugins() short-circuits on _discovered, which is why reload.mcp after
a mid-run install used to reload the OLD server set.
2. Idempotent re-wire: BasePlatformAdapter.rewire_plugin_handlers() runs only
factories not yet wired on the live native client (keyed (plugin, qualname);
a force reload hands back new function objects). Telegram hoists late handlers
ahead of core's catch-all filters.COMMAND / CallbackQueryHandler (PTB dispatches
the first match per group) and re-wires on the transient-init rebuild; Slack
dedupes register_slack_action_handler per AsyncApp. The gateway runner
subscribes per served profile and re-wires on the loop.
3. Scope limit + honest messaging: handlers only. Tools/prompt stay deferred to
the next session (prompt-cache invariant), MCP servers to mcp.reload; the CLI
hint and plugins.manage results (activation, gateway_reloaded,
restart_required only when no gateway answered) say exactly that.
`dashboard_install_plugin` has returned `python_dependencies` in its ok payload since 96e8a23222,
but `PluginsManageResult` (extra=forbid) was never given the field. Every Desktop install therefore
logged `result of 'plugins.manage' violates its wire contract` and the client never saw a clean
install result. Add the field; regenerate the TS/OpenRPC contracts.
The existing install test now sends the installer's real ok payload key for key, so a payload/contract
drift fails under HERMES_TEST_ISOLATION instead of only in a user's errors.log. Red on main, green here.
Live repro: x64 desktop, main 571ceec006, `hermes://plugin/install?repo=NousResearch/hermes-nvidia`
→ errors.log 21:08:42 "python_dependencies Extra inputs are not permitted".
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours
The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.
- `tools/connectors/portal/`: a client for the portal's tool-list route and a
JSON cache under the Hermes home, one file per portal origin and connector.
An entry is fresh for 24 hours. After that the read revalidates with the
stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
`tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
live in `tui_gateway/methods_connectors_account.py`.
The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.
* feat(connectors): catalog, accounts and member tool rules by RPC
The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.
- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
first. The body is a union on `mode`, so a reader can name who turned a
tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
connector, or one connector on or off), with the revision the user saw. A
stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
page can show one card per app.
* feat(connectors): connect an app without a chat session
Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.
- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
`connectors.operation.wake` and `connection.respond` take `owner`, a union
on `type`: `session` (today's behaviour and authorization) or `account`
(routed by `profile`, authorized by the live transport like `mcp.*`).
`session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
thread, under the profile's scope, so the watcher reads the account and
settles the operation. A second connect for an app that is already
connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
address, so its updates go out on the session-less broadcast path.
* feat(mcp-catalog): eighteen more bundled entries name their hosted connector
A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.
* refactor(connectors): the account handlers share one gate, one params model and one write table
The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.
The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.
The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.
An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.
Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.
* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"
The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.
Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.
* fix(connectors): a connect from the page returns to the app after sign-in
The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.
Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.
* test(connectors): defer the new connector RPC coverage
The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.
Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.
Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.
* fix(cli): the connection panel hands the tool thread back at once
The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.
For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.
The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.
Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.
* fix(connectors): "run it again" lives in the library, so the classic CLI can use it
Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.
`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.
Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.
* feat(connectors): the account list and disconnect go through the portal
`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.
The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.
`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.
* fix(connectors): the account RPCs answer what the portal really sends
Checked against the portal source and against the staging and production
services.
- Errors are read from the upstream error code, not the HTTP status. A rule
write answered 409 for a stale revision and for a user with no organisation;
both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
portal's own result for this user, with its stamp and without provider or
subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.
Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.
* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link
Found by two adversarial reviews of the RPC layer and its types.
- `connectors.connect` from a chat session with no open operation is refused
(`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
registry with no card: it made a link nobody watched, returned a reply
without the required `settled` field, and named an operation that was never
registered. There is one way into an operation: the agent's call, or the
account owner's `connectors.connect`. "Run it again" inside an open
operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
profiles with the same key no longer cross-deliver a sign-in link. The event
payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
`connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
The last one is display data: the gateway enforces the rules, the backend
only passes the list on. The phantom `name` and `description` are gone, and
the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
produces it. The contract generator now fails when a contract enum and its
domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
`optional-mcps/` are untouched by this PR again.
anti-slop: no net-new findings (15 touched files).
* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect
The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.
- The agent turn now declares how a link can reach the user
(`tools/connectors/turn.py`): CARD when the agent was built with a
connection callback, SIDE for a subagent or a background turn, LINK for a
headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
tool batch in the agent loop and read by the connector dispatch path, which
never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
tool on any path that derives the tool list, and a connector call on an
unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
does, so no `connection.update` is emitted for an operation no client asked
for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
`respondToConnectionRequest` share one guard and send nothing for a settled
or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
repeated it to the user.
Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.
* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients
A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.
- `tools/tool_labels.py` is the one place that turns a bridged call into a
label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
→ "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
keeps its own emoji, verb and primary-argument preview. A batch gets exactly
one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
failure text on the row of the call that failed. With friendly labels off
it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
carry a typed `labels` field. It does not depend on the classic CLI's
display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
labels, one row per call. The labels reach the row under a key no tool
argument can use. The connect card it drew under a failed tool result is
gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
`manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
"Reading tool details · N tools".
Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.
* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly
- A failed hosted search or describe used to return nothing, by design, so the
model saw only local tools and told the user that a connected app was
missing. The local results are unchanged; when the hosted leg failed, the
`tool_search` and `tool_describe` results carry
`connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
and one hint line. A rejected token is `sign_in_expired`; an entitlement
refusal or a shut gate adds nothing. `tool_describe` no longer lists those
names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
is a hosted connector account; `mcp: true` only when the user asks for an MCP
server, a local server or an install, or when the name exists only in the
catalog; connect and reconnect are hosted verbs, install, enable and
authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
gateway does not know the connector (confirmed on that failure path) and the
name is a catalog entry does the target fail with "X is a local MCP server.
Call manage_connections with action install ...". It is a per-target
outcome: other targets of the same call keep their links and their card. A
vendor failure on a name both sides know stays an ordinary failed row. The
MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
USER asks for that app again; the description and the settled-result notes
say so. A builder saw the model refuse a direct user request.
Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.
* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled
Reproduced on the real Ink TUI with the rig, then fixed:
- The keyboard was dead during the sign-in wait: the card kept a `submitting`
flag that the normal OAuth path never cleared, and Esc went through the same
guard. The in-flight state now belongs to the answered row and clears when
that row moves, when any later frame of the operation arrives, or after
five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
(the input handler had no branch for this overlay); Shift+arrows scroll the
transcript and the card ignores them; arrow keys no longer move the text
cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
operation stayed in the store, and a resume dropped the pending card. The
flag survives idle, a resume shows the pending card again, a session switch
clears it.
- States with no branch: `not_connected` and a row with no link fell into the
credential form; `expired` vanished with no note. The title and the row text
now name the action (connect, reconnect, install, enable, authorize); a
failed or expired row with no fields offers Try again / Skip; a failed row
WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
per app states the outcome. A settled or dismissed operation id is
remembered, so no replay or resume can reopen its card. Esc in the last
"Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
the card in one sentence.
Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.
* chore(connectors): remove the comments and docstrings this branch added
Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.
Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.
* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures
Found by the end-to-end runs on the pushed head.
- Desktop: after a window reload, Continue on the restored card sent nothing.
The answer looked up the backend that holds the session with the runtime
session id, the lookup wants the stored id, and a failed lookup returned
silently. When the lookup gives no owner the answer now goes out on the
window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
a refused scope, a non-member and a missing organisation alike: the handler
runs with the gateway's globals and did not import the reason enum, so its
own error mapping raised. `connectors.accounts.remove` caught auth failures
in its generic branch. `org_required` was mapped on `policy.set` only. All
six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
`ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
* wip(desktop): port the Connectors tab files and wiring onto the #115191 head
* wip(desktop): Connectors tab on the #115191 contract, catalog arm removed, audit defects fixed
* wip(desktop): Connectors tab passes the anti-slop ratchet; dormant two-ways code and the Available collapse removed
* wip(mcp): every server row says whether config or a plugin provides it; writes refuse plugin rows
* wip(desktop): Connectors tab, the owner's first live round (custom MCP form, kind words, compact dialog)
* wip(desktop): the connector dialog fits its content
* wip(desktop): catalog MCPs show on the Connectors tab until the catalog dies; connector_slug pairs a manifest with its managed app; the closed-gate state
* wip(desktop): connectors cache v3, the seed shape gained connector_slug
* wip(desktop): the owner's answers on the connectors page
A plugin-provided server now shows its tool list: the dialog probes it
through the existing read-only test endpoint, shows the tools without
switches (the plugin owns them), and shows the probe's error with a
Retry when the server cannot start. Its card is named after the server
key in the plugin's mcp.json, not the namespaced runtime key.
The paste box no longer parses `--header` on a `hermes mcp add` line;
the CLI has no such flag.
The rule write sends the member layer's revision only. The portal
always returns a member layer (baseline revision when no row exists)
and compares the write against that row, so the effective revision was
never the right guess. Verified live on staging: two writes in a row,
both accepted, policy restored.
The page cache keeps every read for signed-in accounts too and only
clears itself when the account is signed out. The storage version moves
to v4 so old blobs are ignored.
* chore(desktop): strip the prose comments the connectors page branch added
Comments and docstrings this branch added relative to main are gone;
tool directives, SAFETY lines and the contract docstrings that feed the
generated OpenRPC stay. Guards: Python AST and TypeScript printer output
are identical before and after; ruff, tsc, eslint, the ratchet and the
generated contracts are unchanged.
A plugin manifest's `config_schema` now reaches the Desktop: `plugins.manage list`
returns each plugin's schema with the current `plugins.entries.<id>.settings`
values (`settings_schema`), and a new `settings` action writes edits through
`hermes_cli.plugins_state.save_plugin_setting` — the writer extracted from
`PluginContext.set_config`, so the plugin, the CLI and the Desktop share one
config path, one lock and the same managed-install / managed-key refusals.
The Plugins tab grows a gear per plugin with a schema; the inline form is
table-driven (`FIELD_CONTROLS` / `INITIAL_TEXT` / `COERCE` keyed on the wire
field type) for string / number / boolean / enum / json / secret. Secrets are
declared with `type: secret`: the row carries only the `.env` name and a
presence flag, the client writes the value through the existing `PUT /api/env`
credential route, and the RPC refuses secret keys so nothing lands in
config.yaml.
Contracts regenerated; docs gain a "Settings form in the Desktop" section.