A fresh dashboard /chat always spawned the TUI in the dashboard process's
launch directory, so from a phone or any browser there was no way to aim
a new session at a specific repository. Amp's runners now serve many
directories and the web composer offers a picker of the runner's projects
and discovered git checkouts; this ports that mechanism onto the surface
Hermes already has: the dashboard is the phone/web front, the host's
projects.db + repo-discovery cache are the served directories.
- GET /api/chat/workspaces: the profile's projects (with folders) and
discovered repos (session-derived + scanned), default_cwd, home;
?scan=1 rescans desktop.repo_scan_roots on the host, so headless
installs (no Desktop to populate the cache) discover repos too.
- /api/pty?cwd=<dir>: validated (existing directory, fail-closed 400
via the PTY error path) and forwarded to the TUI child as HERMES_CWD
(self-spawned gateway cwd) + HERMES_TUI_CWD (explicit cwd on
session.create for the dashboard's in-memory gateway, whose own cwd
is the launch dir). Resumed sessions ignore it.
- ui-tui: session.create carries cwd when HERMES_TUI_CWD is set, so
/new inside a dashboard chat stays in the picked workspace too.
- web: workspace selector in the Chat rail above "New chat" (projects,
repos by recency, Other path…, rescan), remembered per profile in
localStorage; 17 locales.
- docs: web-dashboard.md rail + REST sections.
Live E2E: real `hermes dashboard` under a scratch HOME with two git
repos under desktop.repo_scan_roots -> /api/chat/workspaces?scan=1
lists both; /api/pty?cwd=repo-a -> TUI status bar shows ~/code/repo-a;
/api/pty?cwd=<missing> -> "Working directory does not exist" + close.
When a provider's live catalog fetch failed, provider_model_ids() degraded to the
curated static list and cached_provider_model_ids() wrote that list to
provider_models_cache.json with a fresh 1h TTL, exactly as if it were the
account's real catalog. The "only non-empty results are cached" guard never
fired because the fallback is non-empty. A transient Copilot outage therefore
replaced an account's 10 enabled models with the 17-model static list on every
picker surface until the TTL lapsed, and a same-credentials restart, re-auth or
Refresh Models could not shake it (#107391).
Mark the curated list as CuratedFallbackModels at the sites that serve it for a
missing live catalog (the Copilot fetcher, the generic profile merge, and the
static tail when a live fetcher declined). The cache layer then treats it as a
placeholder: it never replaces a same-credentials live row (the account's real
catalog is served instead), it is stored flagged with a 60s TTL when there is
nothing better, and it is never served through the stale-while-revalidate
window. A provider with no live source at all is unaffected: its static list is
its catalog and caches as before.
Hindsight now ships from its maintainer's repo (vectorize-io/hindsight,
hindsight-integrations/hermes) through plugin-catalog/hindsight.yaml, so the
in-tree copy under plugins/memory/hindsight goes away. Homes that still name
`memory.provider: hindsight` are migrated by hermes_cli/memory_provider_migration.py
(the `hermes update` hook and the first agent start install the catalog plugin);
that path is untouched here.
Core-side special cases that only made sense with the bundled copy go with it:
the `memory.hindsight` LAZY_DEPS feature, `_provider_pip_dependencies`'s
`hindsight-all` expansion in `hermes memory setup` (the plugin's own
`_ensure_local_runtime()` installs the embedded runtime now), the compat-manifest
pointers of the deleted module, and prose that listed hindsight among the in-tree
providers. Generic provider-name lists, the HINDSIGHT_* env hints and the API-key
redaction pattern stay — a catalog-installed hindsight still uses them.
Revert this commit alone to restore the bundled provider.
A memory provider installed from the catalog under $HERMES_HOME/plugins/
(the honcho handoff package) loads under the loader's synthetic user
namespace, but the host-side surfaces still imported the bundled path
`plugins.memory.honcho.*` by name: the Desktop GET/PUT
/api/memory/providers/honcho/config 500'd, /oauth/start|status 404'd
("does not support OAuth connect"), doctor reported "honcho-ai not
installed", and profile clone / post-update sync silently skipped.
Add `plugins.memory.import_provider_module(name, submodule=None)`: the
provider package (or one of its submodules) of whichever copy
`find_provider_dir` resolves — bundled today, user dir after the removal
PR — under the module name the loader already owns, so both copies
behave identically. Route every host-side absolute import through it:
the dashboard host-block resolvers and honcho.json writer, the OAuth
route resolver, doctor's honcho/mem0 checks, `profile create --clone`,
the update hook's profile sync and the holographic store release.
In-process, no lock, no child process.
Two invariant tests (user-dir honcho with the bundled copy gone:
/config?surface=declared is 200 with the schema; the oauth_flow module
resolves from the user directory), red on base. The honcho write test's
`_honcho_resolvers` stub gains the provider-name argument the seam now
carries.
Shape follows the host-module resolution slice of #116566 by @erosika;
the contract module, installer lock and child-process OAuth runner from
that PR are not needed once the modules resolve in-process.
Co-authored-by: Erosika <eri@plasticlabs.ai>
#119410 registered gpt-6-terra alongside sol and luna from the request text, and the Codex
forward-compat synthesis put it (and -900k) in the live /model picker on every surface.
Nothing serves it: not the Codex account catalog (astra, sol, luna), not OpenRouter (same
three, plus -pro), and OpenAI's model page 404s. Forward-compat is for published tiers the
account catalog has not listed yet, not for guessed names. Removed from the Codex fallback
list and template chain, the context/900k tables, the effort ladder prefixes, the aux-client
family list, and the Nous/OpenRouter static catalogs; catalog JSON regenerated. A contract
test pins the synthesized GPT-6 set to the published tiers.
The onboarding install card installs its rows concurrently, and each install
runs load_and_go_live: a forced plugin rediscovery, then a read of the plugin's
MCP server configs and skills, then an MCP connect. A forced rediscovery
unloads every plugin first (PluginManager.unload clears _portable_mcp_servers,
_plugin_skills and each server's liveness declaration) and only then loads them
again. When NVIDIA App and NVIDIA Broadcast finished installing together, the
Broadcast pass unloaded the manager while the NVIDIA App go-live was reading
it: NVIDIA App went live with no MCP servers and no skill (the card row said
"Installed" with no tool count, and the build chat had no NVIDIA App tools
until the app restarted).
load_and_go_live now runs one at a time in the process (_GO_LIVE_LOCK), so a
second install's rediscovery cannot run while the first reads or connects.
The connect is inside that lock because it reads the server's liveness
declaration (tools/mcp_tool_transport.py::_live_endpoint), which a forced pass
also clears. The reads share the manager's discovery lock with the pass that
produced them, for forced passes that do not go through load_and_go_live, and
connect_plugin_mcp takes the server configs read there instead of reading the
manager again. The MCP connect stays outside the discovery lock, so chats that
call discover_plugins() are not held for the length of a connect.
Each install read .install-metadata.json before its clone and wrote that
snapshot back after it, so the Desktop install card's second row erased the
first row's record (nvidia-app then showed source 'user'). Every writer now
re-reads the sidecar and changes only its own plugin's record under a
cross-process file lock; install rollback no longer rewrites a stale snapshot.
Live run: the onboarding card listed no plugins. The published
plugin-catalog.json is built from main and stamps generated_at at build
time, so it outranked the checkout that added `onboarding: true`, and every
entry arrived without the flag. The same happens on a bundle built before
the catalog change reaches the docs site.
At the same pin, `onboarding` and `title` now come from the checkout when
the doc lacks the key. A doc that carries the key decides (so a curator can
turn it off upstream), and a different pin still follows the newer-catalog
rule unchanged. The previous commit's widened tie-break is reverted.
The onboarding card needs to list catalog plugins beside the hosted
connectors (NS-960 D1, D4) and grey a plugin whose app is absent (D5).
The catalog had no curated flag, and the manage_catalog row left
app_state empty.
- `onboarding: true` and `title` on catalog entries (loader, validator,
docs); set on blender, nvidia-app and nvidia-broadcast.
- hermes_cli/plugin_catalog_presence.py reads the plugin.json at the
catalog's pinned commit once per pin and judges its app declaration with
the hermes_platform resolver the installer and the Plugins-tab pill use.
No declaration or an unreadable one is `unknown`, never `present`.
- `plugins.manage action=onboarding` lists the curated entries this OS
runs (platform mismatch is the only exclusion) with app_state and the
sentence the card greys the row with.
- manage_catalog plugin rows now carry app_state and the catalog title.
- A live entry that differs from the in-tree entry at the same pin (new
metadata) now follows the same newer-catalog rule as a new pin, so a
checkout that adds `onboarding` is not masked by a published doc that
predates it.
Multiplex-only is the direction. This key exists so fleets that lost per-profile gateways in
the switch keep working while the remaining gaps (per-profile stop/restart, WhatsApp
bridge/relay on secondaries, dashboard scoping) are closed, and it goes away once they are.
Every surface that names the key now says so through one shared notice
(STANDALONE_DEPRECATION_NOTICE): the install/start refusal hint, per-profile and default
`gateway status`, the migrate plan, and the host boot log (WARNING, not INFO). The user
guide carries a deprecation admonition where the key is introduced and no longer describes
it as opting out "for good".
Also: the Windows cold-start test fixture stubbed profiles_to_serve(multiplex) without the
new include_standalone kwarg (the PR's one red on CI).
A named profile that authors `gateway.standalone: true` in its own config.yaml
runs its own gateway again, the pre-multiplex topology, while the default
gateway keeps serving every other profile. Topology becomes something the
operator authors per profile instead of something the box infers from boot
state, which is what a fleet running per-profile gateways lost when
`gateway.multiplex_profiles: false` was retired.
Changed
- hermes_cli/profiles.py: `profile_is_standalone(home)` reads the profile's
own config.yaml (memo by file signature, tolerant of malformed yaml, always
False for the default profile with one warning). `profiles_to_serve()`
excludes standalone profiles; roster callers that mean "every installed
profile" (plugin deps, Windows update, launch policy, dashboard listing and
topology) pass `include_standalone=True`.
- gateway/host_attach.py: `standalone_attach_decision` starts a standalone
profile's gateway beside the host multiplexer once every live gateway
confirms it does not serve that profile; refuses with a rescan message while
one still does. Used by the initial attach check and the lock-losing race.
- hermes_cli/gateway_multiplex_mode.py: a standalone launcher never becomes
the host multiplexer (`STANDALONE_PROFILE_REASON`), including callers that
supply an explicit GatewayConfig.
- hermes_cli/gateway.py, web_server_gateway.py: `hermes -p X gateway
install/start/run` proceeds without --force for a standalone profile; the
refusal text for other profiles points at the opt-out; status shows
"standalone (gateway.standalone: true)" and the default lists skipped
profiles.
- gateway/run_profile_reconcile.py: the host does not re-adopt a profile whose
own gateway is live (removing the key while it runs no longer double-binds).
- hermes_cli/gateway_migrate.py: standalone profiles are neither blocker nor
fold target; the plan lists them as "standalone by config".
- gateway/run.py: one INFO line per standalone profile at host boot.
Tests: two-home E2E through real loaders and resolve_multiplex_mode, decide()
with fake host records for both arms, lock-losing branch, reconcile guard,
migrate plan, refusal predicate both ways, topology, memo and malformed-yaml
contracts. All red on base.
Installing a plugin from any surface now makes its MCP servers and skills
usable in every open chat of that profile on the chat's next turn. There is
no Connect-now button, no /reload-mcp, no relaunch.
- hermes_cli/plugins_activation.py::load_and_go_live: after the forced plugin
rescan, connect the plugin's portable MCP servers one by one
(register_mcp_servers, the connector flow's call), then refresh that
profile's open chats and queue a turn note listing the servers, their
tools and the plugin's skills. activation gains live_now; deferred keeps
only Python tools and prompt sections.
- MCP tools are deferred behind tool_search / tool_call, so appending them
(preserve_prefix) leaves the model-facing tool array and the cached prompt
prefix unchanged; the note rides the existing one-shot turn-note channel,
so the system prompt stays byte-stable. /reload-mcp keeps its consent gate:
it is a full rebuild.
- tui_gateway/methods_tools.py: one session walk (_refresh_live_sessions)
shared by reload.mcp and plugin activation, filtered to the plugin's
profile home.
- hermes plugins install/enable (another process) asks the running Desktop /
dashboard backend to do the in-process half through the new
POST /api/dashboard/agent-plugins/activate, found via the host rendezvous
record and its session token.
- Desktop: the Connect-now toast and its strings are removed in all six
locales; the install toast says what went live ("3 tools connected ·
skill X ready") and warns per server that did not connect.
- Contract: PluginActivation.live_now; regenerated TS/OpenRPC.
- register_skill docstring: plugin skills are listed by skills_list.
* fix(mcp): ACP, `hermes tools` and the desktop connector card use the one enabled reader
#119567 left four readers of `mcp_servers.<name>.enabled` on their own rules.
@webtecnica's #119560 found the ACP one:
- `acp_adapter/session.py::_make_agent` used `is not False`, so an ACP
session kept a server with `enabled: "false"` or `enabled: 0` in its
toolset while the MCP client skipped it.
- `hermes tools` MCP picker (`tools_config_mcp._configure_mcp_tools_interactive`)
read `enabled: 0` as on and crashed on a non-dict entry.
- Desktop connector card (`connectors/data/join.ts`) and the MCP health sweep
(`store/mcp-health.ts`) used `enabled !== false`.
All four now call `mcp_server_enabled` (Python) or `serverEnabled` (renderer),
which share one case table.
Co-authored-by: webtecnica <webtecnica@gmail.com>
* chore: retrigger CI (zero-job dispatch failure, auto-heal)
---------
Co-authored-by: webtecnica <webtecnica@gmail.com>
The `enabled` key had four parsers. The MCP client (`_parse_boolish`) read
`enabled: 0` as on; the toolset resolver and editor (`_parse_enabled_flag`)
read it as off. The server list (`summarize_server`, `/api/mcp/servers`) read
any non-`False` value as on, so `enabled: "false"` showed on while the agent
skipped it. The catalog and `hermes mcp list` accepted only true/1/yes, so
`enabled: on` showed off while the server ran.
`tools/mcp_tool_common.py::mcp_server_enabled` is now the only reader, and
every surface calls it. `_parse_boolish` treats YAML numbers by truthiness
(0 off, other numbers on). Everything else keeps the client's semantics:
the off words are off, absent / null / junk stay on, with the existing
warning for junk.
The desktop MCP page mirrors the rule in `serverEnabled`
(`apps/desktop/src/lib/mcp-servers.ts`). One case table
(`mcp-enabled-cases.json`) drives the Python invariant test and the vitest
test, so the page and the runtime cannot drift apart again.
* fix(profiles): toggle MCP servers via the enabled key the runtime reads
The profile editor (profiles.describe / profiles.configure) tracked an MCP
server's on/off state through a `disabled` key on `mcp_servers.<name>`, but
every runtime resolver keys off `enabled` (enabled_mcp_server_names,
coding_context, oneshot, the gateway's tool-name resolver). Disabling a
server in the editor wrote `disabled: true` and the server stayed live at
runtime; a server with `enabled: false` in config showed as enabled in the
editor.
describe now reports `enabled` (default true) and still honours a legacy
`disabled: true`. _save_mcp_toggles writes `enabled: true/false` for every
server and drops the legacy `disabled` key, so old configs migrate on the
next save.
Tests pin describe's reading of `enabled`, `enabled: false`, legacy
`disabled: true` and an unset flag; configure's written `enabled` flags and
the removal of `disabled`; and that a server disabled through the editor is
excluded by enabled_mcp_server_names.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
* test(profiles): trim MCP toggle tests to two invariants
The describe-only and configure-only tests are subsumed by the round trip:
configure writes the key the runtime resolver reads, and describe agrees
with that resolver.
* fix(config): migrate legacy MCP `disabled: true` to `enabled: false`
Older profile editors switched an MCP server off by writing `disabled: true`,
a key no runtime reader consults, so those servers kept running while the
editor showed them off. The editor now writes `enabled`, but configs saved
before the fix still carry the old key until the user saves again.
Config v46 converts each truthy `disabled` to `enabled: false` and drops the
key; a falsy `disabled` is inert and left alone. `hermes update` runs it for
every profile, so the runtime turns off what the user already turned off.
The runtime readers stay on one key; they do not learn `disabled`.
`disabled: true` wins over an explicit `enabled: true`: `hermes mcp add`
writes `enabled: true`, and the old editor only added `disabled`, so the pair
on disk means "the user turned this off in the editor".
---------
Co-authored-by: BowmanStephen <34071312+BowmanStephen@users.noreply.github.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
The ChatGPT Codex models endpoint gates entries on client_version against each
model's minimal_client_version. Hermes sent the "0.0.0" sentinel, which used to
return the whole account catalog; since the GPT-6 Sol/Luna rollout it returns a
frozen legacy list (astra + the 5.6 trio) while any version at or above the newest
minimal_client_version (0.155.0, 1.0.0, 99.0.0 alike, live 2026-09-22) returns
everything the account is entitled to. So the picker and the context-window probe
hid gpt-6-sol / gpt-6-luna that the vendor's own client showed for the same account.
Both request sites now go through fetch_codex_catalog_entries(): try the
newest-client URL first and fall back to the "0.0.0" sentinel only when that
answer is non-200 or empty (the backend used to reject out-of-sequence versions
that way). No local ~/.codex/models_cache.json is consulted: a missing or stale cache (this
host records 0.147.0, which hides gpt-6-sol at 0.155.0) must not decide what the
account can see.
Report by @Stone441 (#119412); supersedes #119420 by @KoNit-K, whose fix read the
compatibility version from the local `~/.codex` cache.
* feat: setup profile is minted by the backend and found by role, not by name
The guided onboarding runs in a profile the desktop used to create itself
(profiles.create with a soul, "already exists" treated as success) and
recognise by the literal "hermes-setup". The upcoming setup toolset grants
catalog installs to that profile, so the marker that grants it must be
written only by the backend.
- profile.yaml carries `role: setup`; read_profile_meta / write_profile_meta /
ProfileInfo know it; profiles.list and GET /api/profiles report it.
- hermes_cli/setup_profile.py: ensure (find by role, adopt a pre-role
hermes-setup dir, else clone default + soul + role) and reset (soul,
memories, skills back to the created state, in place). The soul text moves
here from the renderer.
- tui_gateway/methods_onboarding.py: onboarding.ensure_setup_profile and
onboarding.reset_setup_profile. Neither takes a name; profiles.create and
profiles.configure already reject `role` (unknown key, 4000).
- Copies never inherit the role: --clone-all, profile import, and a
distribution that ships profile.yaml drop it.
- setup.status for a named profile reports `ready` once the boot bootstrap
settled. Since one host backend serves every profile (#118246) the
desktop's setup-profile probe lands on this branch, which never set
`ready`, and the kickoff waited forever.
- Desktop: SETUP_PROFILE, ensureSetupProfile(profiles.create) and
composeSetupSoul are gone. store/setup-profile.ts holds the name the
backend returned (or the roster's role row after a relaunch); kickoff,
handoff and the build card use it. The dev reset calls the reset RPC.
* fix: write the setup soul as bytes so Windows keeps \n line endings
* refactor(desktop): drop the renderer's setup-profile store; the backend is the only owner
Kickoff reads the name straight from onboarding.ensure_setup_profile and records it on
$setupSession, which every later step already carries. The handoff recovery check reads
the roster row's role. No renderer module holds a setup-profile name or a fallback lookup.
OpenAI shipped gpt-6-sol / gpt-6-terra / gpt-6-luna as the successors of the
gpt-5.6 tier line (Sol and Luna live on OpenRouter + the Nous Portal today).
The curated aggregator catalogs (OPENROUTER_MODELS and the derived nous list,
plus the published website model-catalog.json) now carry the gpt-6 tiers and
their -pro variants instead of the 5.6 ones; the openai-api curated fallback
lists them ahead of 5.6.
Codex OAuth support mirrors the 5.6 + Astra contract for every gpt-6 tier:
curated fallback + forward-compat synthesis (from the 5.6 twin or 5.5),
272K advertised fallback, the opt-in -900k picker variants with the
live-verified 900K bump (still capped by the catalog's max_context_window),
dated-snapshot eligibility, wire-suffix stripping, the compaction auto-raise
on the base slug, and the gpt-5.6 effort ladder (max allowed, minimal
rejected). Pricing rows for gpt-6-sol / gpt-6-luna come from OpenAI's model
pages (272K whole-request tier like Astra); Terra has no published page yet
so it deliberately has none.
/model gpt keeps resolving to the flagship: "astra" joins the rank-0 suffix
set so gpt-6-astra sorts above gpt-6-sol.
Anthropic released Claude Opus 5.5. Both the OpenRouter catalog and the
Nous Portal /v1/models endpoint serve it (verified with a live max_tokens=16
completion on each route — echoed model matches, usage.cost billed).
- hermes_cli/models_catalog_static.py: opus-5.5 in OPENROUTER_MODELS, above
opus-5 and below the fable-5 flagship pair. The nous list is derived from
the OpenRouter tuple, so the Portal picker picks it up from the same edit.
- website/static/api/model-catalog.json: regenerated via
scripts/build_model_catalog.py.
Provider-agnostic metadata already resolves for the new slug — no edits
needed: DEFAULT_CONTEXT_LENGTHS key claude-opus-5 substring-matches to
1,000,000 (matches live OpenRouter metadata), and the claude-opus-5
reasoning-timeout prefix matches through the "." separator to the 240s
floor. Both routes bill via official_models_api live pricing, so no
_OFFICIAL_DOCS_PRICING snapshot entry.
A plugin that finished loading after an adapter connected never got its platform
handlers (slash commands, button callbacks, inbound transforms) registered until a
gateway restart, silently. Three pieces, one seam shared by every surface:
1. Discovery listener: PluginManager.on_plugin_loaded(cb) fires from INSIDE
discover_and_load for the plugins a sweep newly loaded (diff of the loaded set),
with a per-plugin activation summary (hermes_cli/plugins_activation.py):
activated_now {gateway_commands, gateway_transforms, hooks, callbacks} vs
deferred {tools, prompt, mcp_servers}. Every mid-run load path now performs a real
discover_plugins(force=True): CLI install/enable (via the gateway), Desktop/TUI
plugins.manage install/toggle/update, dashboard REST install, tool-triggered
force re-discovery, the new `reload-plugins` control-socket verb. A non-forced
discover_plugins() short-circuits on _discovered, which is why reload.mcp after
a mid-run install used to reload the OLD server set.
2. Idempotent re-wire: BasePlatformAdapter.rewire_plugin_handlers() runs only
factories not yet wired on the live native client (keyed (plugin, qualname);
a force reload hands back new function objects). Telegram hoists late handlers
ahead of core's catch-all filters.COMMAND / CallbackQueryHandler (PTB dispatches
the first match per group) and re-wires on the transient-init rebuild; Slack
dedupes register_slack_action_handler per AsyncApp. The gateway runner
subscribes per served profile and re-wires on the loop.
3. Scope limit + honest messaging: handlers only. Tools/prompt stay deferred to
the next session (prompt-cache invariant), MCP servers to mcp.reload; the CLI
hint and plugins.manage results (activation, gateway_reloaded,
restart_required only when no gateway answered) say exactly that.
* fix(plugins): portable MCP servers get a readable name so tool names fit the 64-char cap
A portable plugin's MCP server was named `<skill_namespace>__<server>`, i.e.
`agent-plugin-<slug>-<sha8>__<server>`. That prefix is right for plugin-data and skill names
(collision-free without coordination, and persisted on disk) but it costs ~40 chars of every
`mcp__<server>__<tool>` name. Providers cap function names at 64, so the registry clamped every
tool of the NVIDIA plugin to a hash-suffixed stub with the verb cut off:
`mcp__agent_plugin_hermes_nvidia_72eb26b1__nvidia_app__n_3fa9c2d1`.
Server names only need to be unique among loaded portable servers, and the loader already refuses
a clash. `portable_mcp_server_name(key, server)` = `<plugin-slug>__<server>`, collapsed to
`<plugin-slug>` when the two match (the one-server package). The loader and the desktop card's
server rows both call it (the card recomputed the f-string on its own before). Skill and plugin-data
namespaces are unchanged; nothing persisted refers to the server name, so no migration.
Tests: the existing portable-load test now pins the relationship (server name = plugin slug +
server; a long vendor tool name reaches the wire unclamped), and a new test covers the clash the
digest used to hide: two enabled packages folding to one slug, second server skipped, first served.
Both red on main. Docs: developer-guide/plugins/index.md.
* fix(plugins): a portable MCP server is named what its mcp.json calls it, nothing prepended
Drop the plugin segment too. A user's own config.yaml server named `nvidia-app` yields
`mcp__nvidia_app__<tool>`; a portable plugin's server of the same name now yields the same. Duplicates
are refused at load (config.yaml first, then first-loaded plugin) with a warning naming both owners.
Spike, isolated HERMES_HOME, real session: a portable plugin with server `acme-tools` and tool
`acme_client_get_driver_status_report` registers as
`mcp__acme_tools__acme_client_get_driver_status_report`; the model found it by tool_search, called it,
and reported that exact name. Same plugin on main: `mcp__agent_plugin_acme_tools_88e7456f__acme_tools__acme_153862de`.
An expired stored token makes _codex_catalog serve the static fallback (no Astra). Under
the principal-only key that fallback outlived the token refresh for the whole cache TTL —
before this stack the auth.json rewrite busted it. The identity now has an "expired" state
so the refresh to a live token for the same principal is a cache miss, as it was.
The helper also hashed the principal itself; _credential_fingerprint blake2b-hashes the
joined parts one frame up (they already carry raw API-key env values), so the second digest
bought nothing. It returns the principal (or the opaque token) directly, and the empty-token
case is one early return.
The principal-keyed fingerprint runs resolve_codex_runtime_credentials(read_only=True)
on every cache-only picker read, so the pool-exhausted branch must not probe the usage
endpoint or clear cooldowns from that path: one `not read_only and` guard, kept with
its test. The preserve_corrupt flag threaded through six auth loaders only suppressed
the one-time .json.corrupt sidecar copy on an unparseable store — an edge case the
same read-only resolve already hits on main via _codex_catalog — so it goes.
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours
The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.
- `tools/connectors/portal/`: a client for the portal's tool-list route and a
JSON cache under the Hermes home, one file per portal origin and connector.
An entry is fresh for 24 hours. After that the read revalidates with the
stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
`tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
live in `tui_gateway/methods_connectors_account.py`.
The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.
* feat(connectors): catalog, accounts and member tool rules by RPC
The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.
- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
first. The body is a union on `mode`, so a reader can name who turned a
tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
connector, or one connector on or off), with the revision the user saw. A
stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
page can show one card per app.
* feat(connectors): connect an app without a chat session
Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.
- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
`connectors.operation.wake` and `connection.respond` take `owner`, a union
on `type`: `session` (today's behaviour and authorization) or `account`
(routed by `profile`, authorized by the live transport like `mcp.*`).
`session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
thread, under the profile's scope, so the watcher reads the account and
settles the operation. A second connect for an app that is already
connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
address, so its updates go out on the session-less broadcast path.
* feat(mcp-catalog): eighteen more bundled entries name their hosted connector
A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.
* refactor(connectors): the account handlers share one gate, one params model and one write table
The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.
The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.
The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.
An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.
Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.
* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"
The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.
Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.
* fix(connectors): a connect from the page returns to the app after sign-in
The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.
Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.
* test(connectors): defer the new connector RPC coverage
The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.
Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.
Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.
* fix(cli): the connection panel hands the tool thread back at once
The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.
For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.
The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.
Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.
* fix(connectors): "run it again" lives in the library, so the classic CLI can use it
Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.
`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.
Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.
* feat(connectors): the account list and disconnect go through the portal
`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.
The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.
`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.
* fix(connectors): the account RPCs answer what the portal really sends
Checked against the portal source and against the staging and production
services.
- Errors are read from the upstream error code, not the HTTP status. A rule
write answered 409 for a stale revision and for a user with no organisation;
both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
portal's own result for this user, with its stamp and without provider or
subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.
Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.
* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link
Found by two adversarial reviews of the RPC layer and its types.
- `connectors.connect` from a chat session with no open operation is refused
(`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
registry with no card: it made a link nobody watched, returned a reply
without the required `settled` field, and named an operation that was never
registered. There is one way into an operation: the agent's call, or the
account owner's `connectors.connect`. "Run it again" inside an open
operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
profiles with the same key no longer cross-deliver a sign-in link. The event
payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
`connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
The last one is display data: the gateway enforces the rules, the backend
only passes the list on. The phantom `name` and `description` are gone, and
the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
produces it. The contract generator now fails when a contract enum and its
domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
`optional-mcps/` are untouched by this PR again.
anti-slop: no net-new findings (15 touched files).
* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect
The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.
- The agent turn now declares how a link can reach the user
(`tools/connectors/turn.py`): CARD when the agent was built with a
connection callback, SIDE for a subagent or a background turn, LINK for a
headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
tool batch in the agent loop and read by the connector dispatch path, which
never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
tool on any path that derives the tool list, and a connector call on an
unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
does, so no `connection.update` is emitted for an operation no client asked
for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
`respondToConnectionRequest` share one guard and send nothing for a settled
or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
repeated it to the user.
Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.
* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients
A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.
- `tools/tool_labels.py` is the one place that turns a bridged call into a
label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
→ "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
keeps its own emoji, verb and primary-argument preview. A batch gets exactly
one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
failure text on the row of the call that failed. With friendly labels off
it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
carry a typed `labels` field. It does not depend on the classic CLI's
display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
labels, one row per call. The labels reach the row under a key no tool
argument can use. The connect card it drew under a failed tool result is
gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
`manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
"Reading tool details · N tools".
Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.
* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly
- A failed hosted search or describe used to return nothing, by design, so the
model saw only local tools and told the user that a connected app was
missing. The local results are unchanged; when the hosted leg failed, the
`tool_search` and `tool_describe` results carry
`connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
and one hint line. A rejected token is `sign_in_expired`; an entitlement
refusal or a shut gate adds nothing. `tool_describe` no longer lists those
names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
is a hosted connector account; `mcp: true` only when the user asks for an MCP
server, a local server or an install, or when the name exists only in the
catalog; connect and reconnect are hosted verbs, install, enable and
authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
gateway does not know the connector (confirmed on that failure path) and the
name is a catalog entry does the target fail with "X is a local MCP server.
Call manage_connections with action install ...". It is a per-target
outcome: other targets of the same call keep their links and their card. A
vendor failure on a name both sides know stays an ordinary failed row. The
MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
USER asks for that app again; the description and the settled-result notes
say so. A builder saw the model refuse a direct user request.
Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.
* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled
Reproduced on the real Ink TUI with the rig, then fixed:
- The keyboard was dead during the sign-in wait: the card kept a `submitting`
flag that the normal OAuth path never cleared, and Esc went through the same
guard. The in-flight state now belongs to the answered row and clears when
that row moves, when any later frame of the operation arrives, or after
five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
(the input handler had no branch for this overlay); Shift+arrows scroll the
transcript and the card ignores them; arrow keys no longer move the text
cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
operation stayed in the store, and a resume dropped the pending card. The
flag survives idle, a resume shows the pending card again, a session switch
clears it.
- States with no branch: `not_connected` and a row with no link fell into the
credential form; `expired` vanished with no note. The title and the row text
now name the action (connect, reconnect, install, enable, authorize); a
failed or expired row with no fields offers Try again / Skip; a failed row
WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
per app states the outcome. A settled or dismissed operation id is
remembered, so no replay or resume can reopen its card. Esc in the last
"Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
the card in one sentence.
Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.
* chore(connectors): remove the comments and docstrings this branch added
Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.
Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.
* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures
Found by the end-to-end runs on the pushed head.
- Desktop: after a window reload, Continue on the restored card sent nothing.
The answer looked up the backend that holds the session with the runtime
session id, the lookup wants the stored id, and a failed lookup returned
silently. When the lookup gives no owner the answer now goes out on the
window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
a refused scope, a non-member and a missing organisation alike: the handler
runs with the gateway's globals and did not import the reason enum, so its
own error mapping raised. `connectors.accounts.remove` caught auth failures
in its generic branch. `org_required` was mapped on `policy.set` only. All
six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
`ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
* wip(desktop): port the Connectors tab files and wiring onto the #115191 head
* wip(desktop): Connectors tab on the #115191 contract, catalog arm removed, audit defects fixed
* wip(desktop): Connectors tab passes the anti-slop ratchet; dormant two-ways code and the Available collapse removed
* wip(mcp): every server row says whether config or a plugin provides it; writes refuse plugin rows
* wip(desktop): Connectors tab, the owner's first live round (custom MCP form, kind words, compact dialog)
* wip(desktop): the connector dialog fits its content
* wip(desktop): catalog MCPs show on the Connectors tab until the catalog dies; connector_slug pairs a manifest with its managed app; the closed-gate state
* wip(desktop): connectors cache v3, the seed shape gained connector_slug
* wip(desktop): the owner's answers on the connectors page
A plugin-provided server now shows its tool list: the dialog probes it
through the existing read-only test endpoint, shows the tools without
switches (the plugin owns them), and shows the probe's error with a
Retry when the server cannot start. Its card is named after the server
key in the plugin's mcp.json, not the namespaced runtime key.
The paste box no longer parses `--header` on a `hermes mcp add` line;
the CLI has no such flag.
The rule write sends the member layer's revision only. The portal
always returns a member layer (baseline revision when no row exists)
and compares the write against that row, so the effective revision was
never the right guess. Verified live on staging: two writes in a row,
both accepted, policy restored.
The page cache keeps every read for signed-in accounts too and only
clears itself when the account is signed out. The storage version moves
to v4 so old blobs are ignored.
* chore(desktop): strip the prose comments the connectors page branch added
Comments and docstrings this branch added relative to main are gone;
tool directives, SAFETY lines and the contract docstrings that feed the
generated OpenRPC stay. Guards: Python AST and TypeScript printer output
are identical before and after; ruff, tsc, eslint, the ratchet and the
generated contracts are unchanged.
* feat(platform): resolver core with locate/inspect/probe tiers and ordered candidates
Every resource lookup needs one result shape and one cost contract. `locate` reads
metadata only, `inspect` may open files and call OS APIs in-process, `probe` is fresh
and the only tier that may spawn or connect. `Resolution.candidates` keeps probe order
so fan-out consumers can try every present binary.
Linear NS-921.
* feat(platform): AppResolver over AppDef with plist, PE, registry, and server.json sources
Desktop apps need presence, version, and liveness as separate observations. The runtime
file's bearer token is parsed, used for one request, and discarded inside the probe;
no public type carries it. Endpoints are accepted only when loopback with a numeric port.
* refactor(copilot): gh candidates through locate_command and the Homebrew table
First consumer of the resolver. The gh token probe still tries every present binary in
order; the allowlist loses its two copilot_auth rows.
* feat(platform): availability() over an application declaration
locate() + inspect() only, never probes; the fail-closed _version in
app.py treats a vendor's plist/PE/registry entry as untrusted input.
Salvaged from PR #118122; reads any object with requires_app,
min_version, app_for(os) — nothing here imports the MCP catalog.
* feat(platform): application declarations parsed into AppDef per OS
The parser slice of PR #118122's catalog manifest, re-homed as a
catalog-free module: whoever owns an MCP server declares the app it
fronts per OS and what it needs, and registers it here. Stdlib +
hermes_platform.resolver only. register/lookup/clear are the one seam
the MCP check_fn and the skill gate both read.
* feat(mcp): check_fn honours a registered application declaration
_make_check_fn ANDs the declared app's availability into the
connection-alive check; with nothing registered for the server the
behaviour is the pre-PR3 connection check. Provenance is explicit
registration, not endpoint matching. Returns a plain bool: the registry
caches bool(fn()).
* feat(skills): requires_apps gate through registered declarations
Offer-time filter beside environments:; names resolve through
hermes_platform.declaration, an unknown name hides the skill (fail
closed). The disk snapshot carries requires_apps and the fast path
re-evaluates it (snapshot version bumped to 3): app presence is a host
fact that changes without SKILL.md changing.
* docs: application declarations page
The plugin-facing schema reference: app: and requires: blocks,
availability() states, and the two gates that read the registry.
Registered under Extending > Plugins in the docs sidebar.
* test(platform): declaration parser, availability, gates
The PR3 app-block tests re-homed off the catalog: fixtures are dicts
passed to parse_declaration, the check_fn gate keys on explicit
registration (not endpoint matching), and the import-hygiene probe now
covers hermes_platform.declaration and resolver.availability.
`capture()` sent `get_window_state` with only `pid`, `window_id` and `session`, so the driver
walked the target's entire accessibility tree before Hermes trimmed the surfaced element list
to `_DEFAULT_MAX_ELEMENTS` (100) and spilled the rest to a cache file. Every node past the
first ~100 was paid for and discarded.
Measured on macOS (cua-driver 0.28.2, M-series), same window, bound fixed at 200:
| Target | Unbounded walk | Bounded |
|---|---|---|
| 1,444-node Chrome window | 540 ms | 83 ms |
| 456-node Finder window (pathologically slow AX surface) | 6.9 s | 0.6 s |
The bound is lossless for the response: the bounded element list is a *prefix* of the
unbounded walk (checked by role/label/depth at 200/400/600/1000), so nothing the model sees
changes.
The bound is internal and config-driven — `computer_use.ax_max_elements`, default 200,
`0` disables (driver default) — and deliberately NOT a model-facing schema parameter:
`max_elements` was removed from the tool schema on purpose, so a capture's surfaced window
stays at its fixed default.
(cherry picked from commit 9f07ff428bd1fcfbcc0e5089d834e1f33d8d1b76)
With the lock-free `_load_config_impl` fast path (049576c679, same
contributor), `load_config_readonly()` on a cache hit is already one
`_load_config_cache_sig` stat plus a dict get with no `_CONFIG_LOCK`, so
the `_HOOK_TIMEOUT_CACHE` memo in front of it saved nothing measurable:
probes/S3-quality-memo-cost.py (20k iters, real temp HERMES_HOME) memo
hit 8.40 us vs uncached fast path 8.95 us. It did add a second staleness
rule: the memo keyed on the file sig alone and skipped load_config's
env-snapshot check, so `hook_callback_timeout: ${VAR}` refreshed in
`load_config()` but not in the memo. It is the lock-free fast path that
makes the memo redundant, so remove the cache dict, the wrapper, the
per-path keying (ccb9e36746) and the tuple publish (b1b932a549): the
resolver is the former `_uncached` body calling `load_config_readonly()`
directly, renamed back to `_resolve_hook_callback_timeout`.
PROOF: probes/S3-quality-resolver-cost.py (memo gone, same harness):
`_resolve_hook_callback_timeout` 9.10 us/call. Dropped the memo-only test
`test_hook_timeout_memo_never_pairs_a_new_sig_with_an_old_value` and its
`_bump_mtime` helper; `scripts/run_tests.sh
tests/hermes_cli/test_config_lock_free_cache_hit.py
tests/hermes_cli/test_plugins.py` green (test_plugins patches
`_resolve_hook_callback_timeout` wholesale). Real import asserts
`_HOOK_TIMEOUT_CACHE` and `_resolve_hook_callback_timeout_uncached` are
gone; ruff + check-windows-footguns clean.
The cache-hit predicate was written twice per impl in hermes_cli/config.py:
once lock-free and once verbatim inside `_CONFIG_LOCK`. Extract the pure
lookup+predicate into `_load_config_cache_hit(path_key, cache_sig)` (sig
compare + ${VAR} env-snapshot check) and `_raw_config_cache_hit(path_key,
cache_key)` and call each from both sides. The helpers only look up and
compare; deepcopy-vs-identity stays with the caller, and the fast path keeps
its own try/except around the helper while the locked path stays
undefensive, so exception behaviour on each side is unchanged.
Not folded: passing the fast-path sig into the locked path to save the
second `_load_config_cache_sig` stat. The post-lock re-stat is the
double-checked-locking freshness guarantee; a write landing between the
fast path and lock acquisition would otherwise publish new contents under a
stale sig. Also the fast path only computes the sig when an entry exists, so
the double stat only happens on a sig-mismatch miss (a real re-parse).
PROOF: tests/hermes_cli/test_config_lock_free_cache_hit.py 3 passed on the
refactor. Mutation A (`_load_config_cache_hit` ignores the sig compare) ->
test_hook_timeout_memo_never_pairs_a_new_sig_with_an_old_value red (1 failed,
2 passed). Mutation B (`_raw_config_cache_hit` always returns None) ->
test_cached_raw_read_completes_while_another_thread_holds_the_config_lock red
("blocked 30.0s behind a held _CONFIG_LOCK"). Both reverted. ruff clean,
check-windows-footguns clean, real import under PYTHONSAFEPATH=1 ok.
Shape gate: #117440 carried 4 added tests in
tests/hermes_cli/test_config_lock_free_cache_hit.py. The call-count proxy
`test_hook_timeout_does_not_read_config_on_every_invocation` rebound
`cfgmod.load_config_readonly` by attribute assignment and only counted calls;
the memo torn-read test plus the two lock-held reads are the invariants and
already cover the behaviour. Its sole helper `_reset_hook_callback_timeout_cache`
(plugins.py) had no other caller in tests/ or hermes_cli/, so it goes too.
Module docstring unit fixed: "0.024us" -> "0.024ms" (config.py already said ms).
PROOF: grep -n _reset_hook_callback_timeout_cache hermes_cli/*.py tests/**/*.py
-> only the deleted test + definition. `python -c "import hermes_cli.plugins"`
ok with PYTHONSAFEPATH=1. scripts/run_tests.sh on the test file: 3 passed.
- plugins.py: the `isinstance(cached_value, float)` guard defended against
nothing — `_resolve_hook_callback_timeout_uncached` only ever returns a
float and the hit test already requires `sig is not None`. Type the cache
value as `float`. `_reset_hook_callback_timeout_cache` is kept because
tests/hermes_cli/test_config_lock_free_cache_hit.py calls it; docstring
now says test-only (no production caller).
- config.py: the 16-line narrative comment in `_load_config_impl` restated
the raw-config comment and said 0.024us where the test says 0.024ms; cut
to 4 lines with the right unit and a pointer to `_read_raw_config_impl`.
- NOT folded: extracting the duplicated cache-hit check into a helper. The
fast path swallows every exception and the locked path does not, and the
load path also re-derives the sig, so a shared helper is not a pure
extraction at the same semantics; left in place.
PROOF: no behaviour change; tests/hermes_cli/test_config_lock_free_cache_hit.py
4 passed, probes/S3-fold-thrash.py still 2/100 uncached resolves.
_HOOK_TIMEOUT_CACHE was one process-global (sig, value) slot. Under
multiplex_profiles every profile switch saw a different config signature
and re-resolved, so the memo only ever served the last-used profile.
Key it by the scope-resolved config path like _LOAD_CONFIG_CACHE:
a dict path_key -> (sig, value). Each value is still one tuple stored by a
single dict item assignment (atomic under the GIL), so a lock-free reader
cannot pair a new sig with an old value. The sentinel initial slot goes
away with the slot itself.
PROOF: probes/S3-fold-thrash.py alternates HERMES_HOME A/B 100 times and
counts uncached resolves: before 100/100 (thrash), after 2/100 (one per
home, 2 cache entries). test_hook_timeout_memo_never_pairs_a_new_sig_with_an_old_value
now iterates the published tuples and stays green (4 passed).
`_read_raw_config_impl` had the same lock-on-hit shape #117440 removed
from `_load_config_impl`: a microsecond cache hit queued behind
`_CONFIG_LOCK`, which `save_config()` holds across an atomic YAML write,
so one background config write stalled every raw read (per-turn policy
checks on the gateway). `_RAW_CONFIG_CACHE` already publishes each entry
as one `(*sig, data)` tuple replaced wholesale, so the lock never
protected the read; check the signature first without it and fall
through to the locked re-parse on a miss.
#117440's `_HOOK_TIMEOUT_CACHE` was a two-field dict written under a lock
but read lock-free, so a reader could observe the new `sig` paired with
the previous `value` (torn read) and serve a stale timeout for a fresh
config.yaml. Hold `(sig, value)` in a single module global rebound in one
STORE instead; readers unpack one published tuple, and the lock goes away
because nothing else needed it.
A cache hit in `_load_config_impl` costs microseconds, but it was served
from inside `_CONFIG_LOCK` — which `save_config()` holds across an atomic
YAML write. Measured on a clean checkout, driving the real functions
against a temp HERMES_HOME:
cache-hit read, uncontended median 0.0237ms
the SAME cached read while another
thread holds _CONFIG_LOCK 10010.2ms
On a gateway this lands on the event loop. `invoke_hook` calls
`_resolve_hook_callback_timeout`, which reads config, and a gateway fires
hooks on every inbound message — so one background config write stalls
every message for the full duration of that write. The same probe showed
the hook path doing 100 config reads across 100 hook invocations.
Two changes:
1. `_load_config_impl` gets a lock-free fast path for cache hits. The lock
never protected the cache dict: CPython dict get/setitem are atomic under
the GIL, and the cached tuple is replaced wholesale rather than mutated
in place, so a reader observes either the complete old tuple or the
complete new one. The lock's real job is serializing the rebuild
(parse + merge + expand) and the writers. Worst case on a race is a
redundant rebuild, which the locked path re-checks and collapses. The
existing `_load_config_cache_sig()` helper is reused, so the fast and
locked paths cannot drift on freshness.
2. `_resolve_hook_callback_timeout` is memoized on that same signature. The
value only changes when config.yaml does; every other call is a dict
lookup. Validation and clamping move unchanged into
`_resolve_hook_callback_timeout_uncached`.
After: the blocked read returns in 0.0ms and the hook path does 1 config
read per 100 invocations.
Tests (tests/hermes_cli/test_config_lock_free_cache_hit.py, 11 tests) drive
the real functions against a temp HERMES_HOME — no mocks of the code under
test. They cover the blocked-read case for both the readonly and deepcopy
entry points, that the fast path still sees a changed file, that
`load_config()` still returns an isolated object, a 4-reader + 1-writer
concurrency arm asserting no torn observation, and the hook memo's
freshness plus its clamp/fallback/zero-disables contract.
RED/GREEN on this base, impl reverted via git stash:
without the change : 8 failed, 3 passed (blocked read: 30.0s)
with the change : 11 passed
Neighbours green: tests/hermes_cli/test_config.py, test_plugins.py,
test_config_loader_e2e.py, test_managed_scope_loaders.py,
test_read_raw_config_readonly.py — 266 passed, 4 skipped.
(cherry picked from commit b787427bdaef81d8b114194779fe1de3b4efe7bd)
restore_heartbeat_watches entered _profile_scope_for_source for every routed
session on every poll. Each entry hydrated the profile secret scope and rebuilt
the terminal policy, and both re-parsed the profile config.yaml from disk, so N
routed sessions cost 2N YAML parses per poll even though nothing changed.
- Group entries by resolved profile home and enter the scope once per group.
- Add utils.load_yaml_file_readonly (file_signature-keyed cache) and use it in
env_loader._load_secrets_config and terminal_scope.build_profile_terminal_scope,
which were both open()+fast_safe_load per scope entry. Present-but-unparseable
still fails closed: parse errors propagate and are never cached.
Measured on a 3-profile host: one _profile_runtime_scope enter/exit 1.80 ms -> 0.11 ms.
(cherry picked from commit c6b16629bd38799bbf166206c73a6140a9559a61)
do_install ran the same _full_identifier + _resolve_source_meta_and_bundle
sequence as _resolve_identifier but outside skills_hub_http_session(), so the
hottest hub path — GitHubSource.fetch's tree GET + SKILL.md + one GET per
support file, all via _github_get -> hub()._skills_hub_http_get — opened a
fresh guarded client/TLS handshake per request while inspect got the pool.
Wrap the resolve + fetch block in the session; the early returns still close
the client via the context manager's finally.
PROOF: extended test_inspect_reuses_one_ssrf_safe_client_for_metadata_and_bundle
to drive do_install (bundle-name step stubbed to return False) and assert one
client; with the `with skills_hub_http_session():` wrap removed it fails
("index fetch bypassed the pool" — the unpooled fallback hits httpx.get).
Gates: ruff clean, check-windows-footguns --all clean, real import OK,
scripts/run_tests.sh tests/hermes_cli/test_skills_hub.py green.
Each failed candidate used to claim "falling back to default certificates"
even when the next candidate (certifi on macOS) loaded fine. Per-failure
warnings now say "trying the next bundle"; the default-certificates warning
is emitted once at the final (None, None) return. Level unchanged (WARNING).
PROOF: test_default_certificates_fallback_is_logged_once_after_all_bundles_fail
fails on the previous per-candidate wording and passes with this change;
tests/hermes_cli/test_urllib_security.py 22 passed.
Only a context built from candidates[0] is ever memoised, so keying on every
candidate's signature cost one extra stat per request on macOS and let a
certifi change invalidate a context that never read certifi.
PROOF: test_fallback_bundle_change_does_not_invalidate_the_memo fails on the
previous tuple key (memo miss -> second parse) and passes with this change;
tests/hermes_cli/test_urllib_security.py 21 passed.
`_build_https_context` returns `(None, None)` when nothing loads, so
`used_path == candidates[0]` already implies `context is not None`; the extra
check was dead. The cache annotation still allowed a None context, which is
never stored since the memo only takes preferred-bundle contexts.
PROOF: import + ruff + check-windows-footguns clean;
scripts/run_tests.sh tests/hermes_cli/test_urllib_security.py → 20 passed
(test_a_failed_bundle_load_is_not_memoised still pins the None-not-cached rule).
`_bundle_signature` hand-rolled `(path, mtime_ns, size)`; `utils.file_signature`
is the sibling convention (hermes_cli/config.py, auth.py, auth_oauth_grants.py)
and adds inode + ctime, so a timestamp-preserving replacement of the bundle
(`cp -p`, `rsync -t`, `os.utime`) also invalidates the memo. Superset key:
invalidates more, never less, so the memo semantics are unchanged. utils.py is
stdlib-only and already imported from hermes_cli/*.py, so no import cycle.
PROOF: `PYTHONSAFEPATH=1 python -c "import hermes_cli.urllib_security as m,
utils; assert m.file_signature is utils.file_signature"` → True;
`_bundle_signature('/nope') == ('/nope', None)`;
scripts/run_tests.sh tests/hermes_cli/test_urllib_security.py → 20 passed;
ruff + check-windows-footguns clean.
On macOS `_ca_bundle_candidates` appends certifi after the configured
bundle, so when the configured file fails to load (chmod 000, NFS blip,
AV lock) `_build_https_context` returns a certifi-only context and the
memo stored it under the (bundle sig, certifi sig) key. Because the key
is mtime/size-based, a transient failure that leaves the file unchanged
pinned the fallback forever: the corporate CA was never used again until
the bundle was rewritten.
`_build_https_context` now reports which candidate actually loaded and
`_resolved_https_context` only memoises when that is candidates[0]. A
fallback (or None) result is still returned, just not cached, so the
preferred bundle is retried on the next request. Single-candidate
behaviour is unchanged.
PROOF: tests/hermes_cli/test_urllib_security.py::test_a_fallback_context_is_not_memoised
(two candidates; first fails once). On the pre-fold prod file the second
resolve performed 0 load_verify_locations calls (loads stayed
[corporate, cacert.pem]) -> AssertionError; with the fix the primary is
re-attempted ([corporate, cacert.pem, corporate]) and then memoised.
Also reproduced by probes/S2-2b-unreadable.txt (gate S2-2ab).
The memo assigned `_HTTPS_CONTEXT_CACHE = (key, context)` even when every
candidate failed to load, so one transient failure (EIO, a bundle mid-rewrite
that still stats the same) pinned the process to default certificates until the
file's mtime or size changed. Only cache a successfully built context; a None
result is retried on the next request, which is what the pre-memo code did.
Also trims the rationale comment to the WHY.
_resolved_https_context() built a fresh SSLContext for every credentialed urllib
request. No opener is ever installed globally, so _secure_opener_from_installed_policy
always took the branch that builds one, and each build parsed the whole CA bundle
(certifi: 231 KB, 119 certificates). Peer DMs, model-catalog probes, Azure detection
and the model-provider plugins all pay it per request.
Memoise the context on the (path, mtime_ns, size) of the bundles that actually get
read, so an edited, rotated or reconfigured bundle is still picked up on the next
request — resolving the key is a stat, not a parse.
Bundle resolution moves into _ca_bundle_candidates() so the loader and the cache key
share one precedence rule rather than keeping two copies that can drift. It returns
the ordered candidates, which also keeps the existing behaviour where a configured
bundle that exists but fails to parse still falls back to certifi on macOS.
agent.ssl_verify._context_for_ca_bundle already shares one context per bundle for the
httpx clients; this is the urllib half of the same fix.
Measured on macOS, per opener build: 4.24 ms -> 0.57 ms. Ten requests now parse the
bundle zero times instead of ten.
(cherry picked from commit 067552a825f184ba3220ef75d2ad4ab200e6c935)
`_platform_plugin_manifests()` scanned only the repo's `plugins/platforms/*`, so a
third-party platform plugin under `<HERMES_HOME>/plugins/` never reached
`OPTIONAL_ENV_VARS`: the Desktop Gateway form and `hermes config` showed bare
variable names with no prompt, description or password masking. It now also
walks `<HERMES_HOME>/plugins/platforms/*` and flat `<HERMES_HOME>/plugins/*`
manifests that declare `kind: platform`.
Slim redo of #46964 by @LeonSGP43 onto the refactored helper (the original
predates `_platform_plugin_manifests` and replaced `fast_safe_load`).
Co-authored-by: LeonSGP43 <cine.dreamer.one@gmail.com>