Prerequisites is the first stage, so on a fresh Windows host
<HermesHome>\tools does not exist when Get-PinnedGit runs, and Move-Item
throws DirectoryNotFoundException (reported against the source path).
The uv stager already creates its slot with New-Item -Force.
install.ps1 already kept uv's bootstrap Python out of the Windows
registry, but setup-hermes.ps1 did not, so every Windows dev checkout
registered that interpreter under HKCU. The flag is Windows-only; the
POSIX installers take it too so every bootstrap issues the same command.
PM prepares the ARM64 build environment again inside callers that already
carry a matching developer environment. It is reused as-is, including PATH
order, and a Git Bash caller puts coreutils' link.exe first, so the guard
aborted setup ("MSVC link.exe is shadowed by ...\Git\usr\bin\link.exe").
VsDevCmd's VCToolsInstallDir names the MSVC linker exactly, whichever way
the environment was obtained. PATH order stays the caller's.
installer-tests.yml predates nothing it still owned. Its pytest step
(test_source_launcher_stages.py) is platforms("windows") and already runs in
both tests-os Windows lanes, so every installer PR ran it twice. The
`installer` lane never gated anything on its own either: every path that set
it also sets `python`, which gates tests-os.
The two standalone scripts/tests/*.ps1 suites become one platforms("windows")
pytest file parametrized over Windows PowerShell 5.1 and pwsh 7, so
list_os_marked_tests picks them up with everything else. The `installer` lane
goes away from the classifier, detect-changes, ci.yaml and the
all-checks-pass gate; the classifier contract now pins that install.ps1 and
its suites turn `python` on.
install.ps1 now passes --no-registry to `uv python install` (the review fix that
keeps a registry Python from satisfying the request); the exact-argv contract
moves with it. Both scripts pass under Windows PowerShell 5.1 on a Win11 arm64 host.
Review findings against the pm-clean installers, each reproduced first:
- install.sh: `curl | bash` aborted before main under `set -u` (empty
BASH_SOURCE). The entry guard falls back to $0.
- install.sh: setup/gateway read stdin, which under `curl | bash` is the
script itself. They open /dev/tty when a terminal can be opened, and
otherwise skip with guidance.
- Both: any uv on PATH was trusted. uv 0.6.17 has no `python install
--no-bin`. A PATH uv now has to run and be at least the pinned version,
otherwise the pin is staged.
- install.sh: the staged uv went under ~/.hermes/tools even with a custom
--hermes-home. It now goes to pm's store_root() default,
$HERMES_HOME/tools.
- Both: when a stash failed, the script logged "overwritten below" and ran
`reset --hard` anyway. Local work is now parked before checkout, and a
stash failure stops the install.
- Both: reruns ignored an explicit HERMES_REPO_URL. It now repoints origin.
- Both: --commit had no ancestor guard. The pin must be on the installed
branch.
- install.sh: the blobless fallback was `--depth 1 --single-branch`, so a
non-tip --commit could not check out. It now keeps full history with
blobs fetched on demand.
- Both: ported main's recovery for a commit-less .git (moved aside, #40998)
and for an unmerged index (reset -q before the stash, #4735). Stashing
before checkout makes both reachable.
- install.ps1: on Windows PowerShell 5.1, `native 2>$null` / `2>&1` under
Stop turns stderr into a terminating NativeCommandError, verified on a
Win11 host. Every native call now goes through Invoke-Native, which
relaxes the preference only for that call.
- install.ps1: clone publishes from a staging dir, with retries and a
blobless fallback, and refuses a non-empty destination (mirrors
install.sh). UV_NO_CONFIG and `--no-registry` are restored. pwsh 7
HttpRequestException falls back to the mirror, except TLS trust failures.
tar resolves from System32. A literal CR/LF in the desktop failure
message is removed.
Deletes six tests that regex-extracted main's legacy install.sh functions
(node, browsers, PATH block, lockfile churn; pm runs `npm ci` whenever a
lockfile exists). The two behaviours still relevant are covered by new
behavioural tests.
cryptography ships no win_arm64 wheel, so every Windows ARM64 venv sync
compiles it from the sdist and needs MSVC, Clang, Rust and static OpenSSL.
Only setup-hermes.ps1 (and so activate.ps1) prepared that environment,
between a `pm install --tools-only` and the real sync. install.ps1,
`hermes update` and repair ran the same sync without it and failed in
openssl-sys.
PM owns the sync, so PM prepares it. pm/native_build.py holds the adapter
(moved from scripts/build/windows_deps.py) plus source_build_environment(),
which prepares only on win32-arm64 when the synced project carries the
provider script. A payload has prebuilt dependencies and needs no compiler.
VenvPackage.apply and build_environment pass the result to uv children
only. It carries the bridged pip index settings, which managed_environment
applies only to the ambient environment. The state root stays the store
parent, so existing vcpkg/OpenSSL builds are reused.
setup-hermes.ps1 collapses to one `pm install`: the tools-only split existed
only for this preparation, and pm install already puts its tools on PATH
before the venv sync (pm/cli.py activate check).
Not yet verified live on Windows ARM64.
The products stage (source_completion) writes install-stamp.json, and
this lane deliberately runs only prerequisites/repository/config/complete,
so the checkout never has one. The lane now says so with
--no-source-stamp; full-install callers (windows-e2e.ps1) stay strict.
Replayed locally: the four stages leave no stamp, the old invocation
fails exactly like CI, the lane invocation verifies.
The source completion tail stamps baseVersion from the checkout's
reachable release tag, which is null on a PR checkout, and the bootstrap
verifier demanded a non-null baseVersion, so install.sh protocol failed
on every upstream PR. It now requires the stamp and commit == HEAD, the
identity the runtime actually reports.
Install/update completion rewrites the root install-stamp.json (fresh
builtAt) after products are built, and the stamp was still a shared
web/desktop source input, so every install left its own web_dist
"missing, stale, or damaged" and the post-install probe failed on every
installer E2E leg. Nothing in either build reads the root stamp;
desktop's baked stamp is already tracked as a prepared input.
`& ([scriptblock]::Create((irm .../install.ps1)))` is the documented
Windows install form, and on this line it failed twice before doing any
work:
- The MSIX rewrite re-added a UTF-8 BOM that had been stripped twice
before. Windows PowerShell 5.1 irm keeps it as a literal U+FEFF, so
param( was no longer the first statement: "The assignment expression is
not valid".
- Initialize-ResolvedPaths read and wrote $script:HermesHome and
$script:InstallDir. Under -File, script scope is where param() binds.
Under the scriptblock form param() binds in the scriptblock scope and
$script: names the caller session, so -HermesHome read as empty and
Join-Path threw. Without -HermesHome, the normalized paths never reached
the bare $HermesHome/$InstallDir every stage reads.
The paths are now computed locally and written to scope 1, where param()
bound under -File, the scriptblock form, and dot-sourcing.
Verified on Windows 11 arm64 (PS 5.1): the old file reproduces both
errors, -ShowResolvedPaths is correct under all three entry forms, and a
fresh install clones into the requested home and reaches pm install.
Upstream carries only CalVer tags, which version_from_tag rejects on
purpose, so every PR image build died in "Write install stamp" with
"no reachable release tag". The runtime already reads a tagless source
checkout as base "unknown" plus its commit; the image now records the
same: the workflow admits GITHUB_SHA via --commit, adds version args only
when a release is reachable, and write_install_stamp accepts a missing
base version when the caller supplied the commit (a local tree still
stays unstamped).
pwsh has no Get-WmiObject; it proxies the cmdlet through a Windows
PowerShell compat session, and the ManagementClass comes back
deserialized with its methods stripped, so Win32_ShadowCopy.Create
failed ("does not contain a method named 'Create'") and Delete would
have too. Invoke-CimMethod / Get-CimInstance / Remove-CimInstance are
native in both 5.1 and 7. Under CIM, Win32_ShadowStorage.Volume is a
CimInstance reference, so match on its DeviceID instead of the WMI
path string.
The smoke preferred powershell.exe over pwsh, so it never ran the kit
under pwsh; it now runs the kit under the host running the smoke.
./scripts/<name>.py now sources activate when the inherited environment is
missing or stale (scripts/_hermes-python) and runs under the pinned
interpreter, so these work from a bare shell and any cwd. The fast path from
an activated shell costs ~5ms over bare python.
These all import repo or locked third-party code. CI invokes them as
`python3 scripts/<name>.py` under setup-pm's environment, where the shebang
is inert, so no require_activation() guard: that would fail those lanes,
which run without __HERMES_ACTIVATED.
Left out on purpose: keystroke_diagnostic.py (run by users inside their own
install, where activation would provision their home) and
docker_config_migrate.py (runs in the container through its venv python).
The _hermes-python prologue compared uv.lock / pyproject.toml / pm/lock.json
against facts.json with `-nt`. facts.json is only rewritten on a real sync,
so a checkout that rewrites an input without changing it left the input
newer forever: every run from an activated shell re-sourced activate
(~370ms instead of ~13ms).
pm now records, after every successful full-closure install (no-op syncs
included), a stamp per input under installs/<key>/inputs carrying the exact
mtime the install was verified against, snapshotted before installing so
an input edited mid-install still reads as stale. The prologue re-activates
when any input's mtime differs from its stamp in either direction (branch
switches can move it backwards), when an input is missing, or when there
are no stamps. It globs the stamp dir, so pm owns the only input list.
Also read pyvenv.cfg as utf-8-sig (footgun lane, same file).
The egress developer guide says `HERMES_RUN_E2E=1 scripts/run_tests.sh
tests/agent/test_iron_proxy_e2e.py`, but the runner's `env -i` dropped the
variable, so all three cases skipped silently. Forward it next to the other
opt-in knobs (HERMES_RUN_SLOW_PET_TESTS, HERMES_E2E_BROWSER); before: 3
skipped, after: 3 passed. The test docstring no longer names a --run-e2e
option that never existed.
cp -c clones file by file: 21.5s for a real 148k-entry HERMES_HOME. APFS
clones a whole directory tree in a single clonefileat(2), which a stock Mac
can call through /usr/bin/perl: 5.9s for the same home, identical in name,
type, size, mode, mtime and link target. The kernel stamps cloned
directories with the current time, so a second pass restores each
directory's atime/mtime. On any failure the entry is removed and cloned
with cp -c instead, with a warning the smoke test asserts is absent.
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
pre tarred the entire HERMES_HOME and post extracted it over an emptied
home: two full copies of a tree that is mostly node_modules and venvs.
pre now takes a `hermes backup` of the user's data, then clones every
top-level entry of HERMES_HOME except cache/ (and the Electron userData)
into the backup dir: copy-on-write where the filesystem can (clonefile on
APFS, reflink on btrfs/xfs), a plain copy elsewhere. post moves the live
entries aside and the clones in, by rename only, so the home directory
itself never moves (it may be a mountpoint or symlink target) and cache/
stays where it is. It then deletes the moved-aside post-update trees;
hermes-backup.zip is never deleted.
Windows deletes a VSS snapshot once the old copies of rewritten blocks
exceed the volume's shadow-storage cap, which defaults to ~1% of the disk
(17.8 GB on a 1.8 TB drive). A long rehearsal with other disk activity can
blow that and cost the tester the rollback.
pre now raises the cap to 128 GB (a ceiling, not a reservation) right after
taking the snapshot, recording the exact original in shadowstorage.txt; post
puts it back, also on the snapshot-gone path. An already-larger cap is left
alone.
pre tarred the entire HERMES_HOME, which is slow on a real home (hundreds of
thousands of files) and silently left out files other processes held open:
on a live 33GB home the tar came out at 4.7GB with no error.
pre now takes a `hermes backup` of the user's data, then a Volume Shadow
Copy snapshot of the volume(s) holding HERMES_HOME and the Electron userData.
The snapshot is copy-on-write: 2s, nothing copied, locked files included.
post checks every snapshot still exists before touching anything, then
robocopy /MIR's both trees back from it (only changed files are copied,
files the update added are deleted; HERMES_HOME\cache is left alone), and
deletes the snapshot. If Windows dropped the snapshot, post changes nothing
and points at the hermes backup zip instead.
pre and post now need an elevated PowerShell.
The salvaged restart (#119815) exited 9 when `hermes gateway start --all`
failed, which makes the hand-off's finally block write ok:false, show the
error finale and log "detached update FAILED" in the Desktop even though
the code, venv and Desktop build all verified. Move the restart after the
outcome is settled and surface a restart failure through Write-Result's
existing manual flag instead: the Desktop's boot dialog shows the
`hermes gateway start --all` follow-up, and the update stays a success.
Publish a UI stage so the marquee names the step.
Extend the salvaged test with the exit-code invariant (red on the PR's
own head) and correct the update_cmd_windows docstring that said the
Desktop never restarts the messaging gateway.
Closes#119809
Restore the cheap repo-wide guards the per-file triage classed as source reads
but that protect recurring bug classes (<2s total):
- subprocess env scrubbing near spawn sites (credential leakage)
- gateway UTF-8 encoding= on file I/O (Windows mojibake)
- no raw yaml.safe_load of config.yaml (lost ${ENV} expansion)
- CLI subprocess.run timeouts (hung CLI)
- no locked readers on the shared state.db connection (#99349 segfault)
- CI classifier outputs / live-comment watch list match real workflows
- relay imports no platform crypto (relay trust boundary)
- Desktop relay deliver budget mirrors the Python deadlines (#93911)
- no native title= on Desktop buttons (DESIGN.md rule)
Drop _BASELINE entries in check_os_marker_fakes.py for files that no longer
fake macOS (the checker fails on stale entries), and remove doc/comment
pointers to deleted tests.
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
The payload snapshot omits apps/ (INERT_SNAPSHOT_DIRS), so validate_bootstrap_version
raised FileNotFoundError on every native bundle leg since dd72571178.
(cherry picked from commit 8411fdb333)
The script's .EXAMPLE and HANDOFF.md told Windows users to pass --source,
--ref and --yes, which PowerShell rejects as extra positional arguments.
Use -Source, -Ref and -Yes.
The onboarding card needs to list catalog plugins beside the hosted
connectors (NS-960 D1, D4) and grey a plugin whose app is absent (D5).
The catalog had no curated flag, and the manage_catalog row left
app_state empty.
- `onboarding: true` and `title` on catalog entries (loader, validator,
docs); set on blender, nvidia-app and nvidia-broadcast.
- hermes_cli/plugin_catalog_presence.py reads the plugin.json at the
catalog's pinned commit once per pin and judges its app declaration with
the hermes_platform resolver the installer and the Plugins-tab pill use.
No declaration or an unreadable one is `unknown`, never `present`.
- `plugins.manage action=onboarding` lists the curated entries this OS
runs (platform mismatch is the only exclusion) with app_state and the
sentence the card greys the row with.
- manage_catalog plugin rows now carry app_state and the catalog title.
- A live entry that differs from the in-tree entry at the same pin (new
metadata) now follows the same newer-catalog rule as a new pin, so a
checkout that adds `onboarding` is not masked by a published doc that
predates it.
Windows PowerShell 5.1 returns every match from Get-Command. .Source on
that array joins the paths with a space, and the call operator then treats
the joined string as one program name. Git for Windows ships git.exe in
cmd\ and bin\, so setup died with CommandNotFoundException before pm
install ran.
setup-hermes.ps1 now installs the tool closure first (`pm install
--tools-only`), then prepares the ARM64 compiler environment, then syncs
the venv. The sync inherits that compiler environment. A bare `pm install`
and the update takeover path publish tools and put them on PATH before
uv sync. A missing tool stops the sync. A missing venv does not.
Verified: scripts/run_tests.sh on test_install_default_closure.py,
test_install_extra.py, and test_windows_build_deps.py — 13 passed.
release printed the claim URL and stopped. The other commands printed a
record or a one-line status. Each command now says what happened, what to
wait for, and the next action.
release names the claim, the workflow run, the notes page, and the publish
command when autopublish is off. publish and canary name their workflow.
abandon says the version is spent. A commit build names its page and says
it moves no channel.
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours
The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.
- `tools/connectors/portal/`: a client for the portal's tool-list route and a
JSON cache under the Hermes home, one file per portal origin and connector.
An entry is fresh for 24 hours. After that the read revalidates with the
stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
`tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
live in `tui_gateway/methods_connectors_account.py`.
The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.
* feat(connectors): catalog, accounts and member tool rules by RPC
The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.
- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
first. The body is a union on `mode`, so a reader can name who turned a
tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
connector, or one connector on or off), with the revision the user saw. A
stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
page can show one card per app.
* feat(connectors): connect an app without a chat session
Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.
- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
`connectors.operation.wake` and `connection.respond` take `owner`, a union
on `type`: `session` (today's behaviour and authorization) or `account`
(routed by `profile`, authorized by the live transport like `mcp.*`).
`session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
thread, under the profile's scope, so the watcher reads the account and
settles the operation. A second connect for an app that is already
connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
address, so its updates go out on the session-less broadcast path.
* feat(mcp-catalog): eighteen more bundled entries name their hosted connector
A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.
* refactor(connectors): the account handlers share one gate, one params model and one write table
The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.
The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.
The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.
An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.
Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.
* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"
The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.
Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.
* fix(connectors): a connect from the page returns to the app after sign-in
The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.
Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.
* test(connectors): defer the new connector RPC coverage
The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.
Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.
Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.
* fix(cli): the connection panel hands the tool thread back at once
The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.
For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.
The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.
Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.
* fix(connectors): "run it again" lives in the library, so the classic CLI can use it
Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.
`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.
Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.
* feat(connectors): the account list and disconnect go through the portal
`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.
The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.
`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.
* fix(connectors): the account RPCs answer what the portal really sends
Checked against the portal source and against the staging and production
services.
- Errors are read from the upstream error code, not the HTTP status. A rule
write answered 409 for a stale revision and for a user with no organisation;
both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
portal's own result for this user, with its stamp and without provider or
subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.
Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.
* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link
Found by two adversarial reviews of the RPC layer and its types.
- `connectors.connect` from a chat session with no open operation is refused
(`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
registry with no card: it made a link nobody watched, returned a reply
without the required `settled` field, and named an operation that was never
registered. There is one way into an operation: the agent's call, or the
account owner's `connectors.connect`. "Run it again" inside an open
operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
profiles with the same key no longer cross-deliver a sign-in link. The event
payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
`connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
The last one is display data: the gateway enforces the rules, the backend
only passes the list on. The phantom `name` and `description` are gone, and
the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
produces it. The contract generator now fails when a contract enum and its
domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
`optional-mcps/` are untouched by this PR again.
anti-slop: no net-new findings (15 touched files).
* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect
The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.
- The agent turn now declares how a link can reach the user
(`tools/connectors/turn.py`): CARD when the agent was built with a
connection callback, SIDE for a subagent or a background turn, LINK for a
headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
tool batch in the agent loop and read by the connector dispatch path, which
never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
tool on any path that derives the tool list, and a connector call on an
unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
does, so no `connection.update` is emitted for an operation no client asked
for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
`respondToConnectionRequest` share one guard and send nothing for a settled
or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
repeated it to the user.
Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.
* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients
A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.
- `tools/tool_labels.py` is the one place that turns a bridged call into a
label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
→ "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
keeps its own emoji, verb and primary-argument preview. A batch gets exactly
one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
failure text on the row of the call that failed. With friendly labels off
it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
carry a typed `labels` field. It does not depend on the classic CLI's
display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
labels, one row per call. The labels reach the row under a key no tool
argument can use. The connect card it drew under a failed tool result is
gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
`manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
"Reading tool details · N tools".
Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.
* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly
- A failed hosted search or describe used to return nothing, by design, so the
model saw only local tools and told the user that a connected app was
missing. The local results are unchanged; when the hosted leg failed, the
`tool_search` and `tool_describe` results carry
`connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
and one hint line. A rejected token is `sign_in_expired`; an entitlement
refusal or a shut gate adds nothing. `tool_describe` no longer lists those
names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
is a hosted connector account; `mcp: true` only when the user asks for an MCP
server, a local server or an install, or when the name exists only in the
catalog; connect and reconnect are hosted verbs, install, enable and
authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
gateway does not know the connector (confirmed on that failure path) and the
name is a catalog entry does the target fail with "X is a local MCP server.
Call manage_connections with action install ...". It is a per-target
outcome: other targets of the same call keep their links and their card. A
vendor failure on a name both sides know stays an ordinary failed row. The
MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
USER asks for that app again; the description and the settled-result notes
say so. A builder saw the model refuse a direct user request.
Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.
* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled
Reproduced on the real Ink TUI with the rig, then fixed:
- The keyboard was dead during the sign-in wait: the card kept a `submitting`
flag that the normal OAuth path never cleared, and Esc went through the same
guard. The in-flight state now belongs to the answered row and clears when
that row moves, when any later frame of the operation arrives, or after
five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
(the input handler had no branch for this overlay); Shift+arrows scroll the
transcript and the card ignores them; arrow keys no longer move the text
cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
operation stayed in the store, and a resume dropped the pending card. The
flag survives idle, a resume shows the pending card again, a session switch
clears it.
- States with no branch: `not_connected` and a row with no link fell into the
credential form; `expired` vanished with no note. The title and the row text
now name the action (connect, reconnect, install, enable, authorize); a
failed or expired row with no fields offers Try again / Skip; a failed row
WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
per app states the outcome. A settled or dismissed operation id is
remembered, so no replay or resume can reopen its card. Esc in the last
"Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
the card in one sentence.
Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.
* chore(connectors): remove the comments and docstrings this branch added
Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.
Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.
* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures
Found by the end-to-end runs on the pushed head.
- Desktop: after a window reload, Continue on the restored card sent nothing.
The answer looked up the backend that holds the session with the runtime
session id, the lookup wants the stored id, and a failed lookup returned
silently. When the lookup gives no owner the answer now goes out on the
window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
a refused scope, a non-member and a missing organisation alike: the handler
runs with the gateway's globals and did not import the reason enum, so its
own error mapping raised. `connectors.accounts.remove` caught auth failures
in its generic branch. `org_required` was mapped on `policy.set` only. All
six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
`ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
- tests/hermes_cli/test_config_yaml_comment_preservation.py: a hand-commented config goes
through config set, config unset, save_config (plugin enable + memory.provider), a
_config_version migration bump and a direct atomic_config_write; every comment, the key
order and the quoted "off" must survive, and the boilerplate is appended only on create.
6/8 red on base.
- scripts/check_config_yaml_writers.py (wired into the lint workflow): AST scan that fails
on any atomic_yaml_write / yaml.dump / yaml.safe_dump of a config path, or any PyYAML dump
inside the config system, outside the writer module. Flags all nine base-tree writers.
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.
Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.
- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.