The stop, start and restart verbs on a profile served by the host multiplexer,
the gateway.parked marker (provisioning may pre-create it), reconcile timing,
and what parked means in gateway status.
Symptom: a host that serves several profiles from one process and mirrors
the active turn's profile into `os.environ["HERMES_HOME"]` for legacy
readers (Hermes WebUI does this on every chat turn, next to the
context-local override) makes every launch-home decision see the served
profile as the launch profile. Two profiles that both configure `atlassian`
with different credentials share whichever MCP connection came first: a
READ_ONLY_MODE=false profile ends up calling a read-only server
(nesquena/hermes-webui#7721). The same misjudgement leaves the launch
residue in the served profile's child env, seeds the launch profile's
bridged allow-all grant into the served profile's secret scope, and lets
the served profile's `terminal.*` config bridge into the shared process env.
Cause: four launch-home checks compare the task's override with
`get_process_hermes_home()`, which reads `HERMES_HOME` live:
`agent.secret_scope.serves_routed_profile` (keys the MCP ledger via
`_mcp_registry_scope`, #108352 / #111481, and the check_fn cache, #111151),
`agent.secret_scope._is_process_home`, `tools.environments.local._is_routed_home`
and `hermes_cli.env_loader._process_hermes_home`. Under the mirror the two
sides are equal for every turn.
Change: `hermes_constants.pin_process_hermes_home(path | None)` lets the
host record the home it serves as its own; `get_routing_process_hermes_home()`
returns the pin when set, else `get_process_hermes_home()`; the four checks
compare against it. The pin is deliberately NOT folded into
`get_process_hermes_home()`: `get_hermes_home()` falls back to it for tasks
carrying no override (MCP loop, spawners), and the host's mirror exists
precisely so those readers see the served profile. Only "is this task
routed / is this the launch home" changes. Unpinned, behaviour is
byte-for-byte the old one; hosts that never mutate `HERMES_HOME` need not
call it. `activate_multi_profile_hosting()` is not the seam for this: it
flips `get_secret` fail-closed process-wide and freezes the launch env,
which an embedding host cannot adopt as a bug fix.
Tests (2 invariants, parametrized over the four checks plus the MCP ledger
key; red on main, green here): pinned + mirrored env -> the served home is
routed and the launch home is not, the MCP key is `(home_key, name)`,
`get_process_hermes_home()` still follows the env var; never pinned or
pinned-then-cleared -> old semantics, including "a mirrored env var IS the
launch home". `tests/conftest.py` resets the pin per test so the
module-global cannot leak between files.
Live repro (WebUI + a stdio FastMCP server named `atlassian` in two
profiles, one gated by READ_ONLY_MODE): base -> one ledger key
`'atlassian'`, the write profile lists only the read-only tools; fixed ->
`(<read_home_key>, 'atlassian')` and `(<write_home_key>, 'atlassian')`,
each profile lists its own tools.
Docs: `gateway/AGENTS.md` § Profile scope (one launch-home identity) and the
isolation table in `website/docs/user-guide/multi-profile-gateways.md`.
Also maps the author e-mail under contributors/emails/ (attribution check).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A multiplexed gateway connects each served profile's adapter inside
_profile_runtime_scope(<profile home>). The receive loop an adapter starts
while connecting inherits that home override, so every final reply the bot
sends is recorded from it. The ledger resolved its path through
get_hermes_home(), which follows the override, and the rows landed in
profiles/<name>/state.db. The boot sweep (sweep_recoverable) and the boot
flood-timer arming (pending_retries) run in the launch context and open the
launch state.db, so they never saw those rows. A served bot's reply cut off by
a crash or SIGKILL between finalize and platform ACK was never redelivered, a
flood-refused reply that spanned a restart was never retried, and
resume_pending was not cleared for a session whose answer sat in the ledger.
The ledger is meant to be one shared store: the boot sweep already scopes rows
by (platform, adapter_profile), and the profile purge terminalizes rows in the
shared store. _db_path now resolves from get_process_hermes_home(), as the
gateway's other process-level files do (gateway.status). It deliberately skips
the get_hermes_home() fallback that lifecycle_ledger uses when HERMES_HOME is
unset: a default gateway started in the foreground has no HERMES_HOME, and that
fallback would follow the override again.
Rows an earlier build already wrote to a profile's state.db stay where they are.
EmailAdapter._sender_accepted runs before any MessageEvent exists and
read only EMAIL_ALLOWED_USERS. Unset, it dropped every sender unless
allow-all was on; set, it dropped everyone not listed. The gateway's
own handling therefore never ran for email:
platforms.email.unauthorized_dm_behavior "pair" (the setup wizard's
"Use DM pairing") and "decline" sent nothing, and a sender admitted by
GATEWAY_ALLOWED_USERS or an approved pairing was dropped. bb304b4914
turned the empty-allowlist branch into drop-all after #50568 had made
"pair" email's explicit opt-in.
The gate now keeps a sender listed by address in EMAIL_ALLOWED_USERS
or GATEWAY_ALLOWED_USERS, a sender the registered gateway
authorization check admits (that is the only reader of the pairing
store), and, under an explicit pair or decline, an unknown sender the
gateway will answer. The default "ignore" still drops unknown senders
before a MessageEvent exists, so the mail-loop guard from fd9c32c0f2
holds.
Three guards keep the wider gate from widening access, and close two
forged-From: paths main already had:
- A sender admitted only so the gateway can answer it (pair or
decline) must authenticate its From:, open access or not: the
pairing code or refusal is mailed back to that address. A granted
sender still needs it short of open access, since a pairing grant
keys on From: just as the allowlist does. Open access follows the
gateway's own order: EMAIL_ALLOW_ALL_USERS wins over a list, while
GATEWAY_ALLOW_ALL_USERS beside a list admits nobody extra, so it no
longer exempts a listed address from From: authentication either
(on main a forged From: of a listed address got through there).
- Open access comes from the gateway's own verdict when a check is
registered. GATEWAY_ALLOW_ALL_USERS beside a GATEWAY_ALLOWED_USERS
list grants a stranger nothing there, so the env flag alone no
longer exempts one from From: authentication (that path mailed a
pairing code to a forged From: on main too).
- A sender whose local part alone matches an allowlist entry is
dropped. The gateway's check also matches an address by its bare
local part (#119446), so without this, GATEWAY_ALLOWED_USERS=alice
(a chat username) would admit or pair alice@<any domain>. The lists
are parsed as the gateway parses them, JSON list literals included,
or '["alice"]' would slip past this guard.
_allowlist_in_effect only served the old condition and is removed.
The scope tests now assert the same scoped reads through
_sender_accepted, with GATEWAY_ALLOWED_USERS covered as well.
Measured end to end with the real GatewayRunner callback wired
(adapter -> gateway ingress):
- pair, decline, GATEWAY_ALLOWED_USERS and an approved pairing each
went from 0 events reaching the gateway to 1. pair mails a pairing
code, decline mails one refusal.
- An unauthenticated From: in pair mode, for a paired address or for a
GATEWAY_ALLOWED_USERS address still reaches nothing.
- A bare GATEWAY_ALLOWED_USERS=stranger entry lets nothing from
stranger@<domain> through, under ignore or pair. Without the
local-part guard that mail reached the gateway in both.
- The same holds for a JSON-literal list, and a pair-mode stranger
with a forged From: under allow-all beside an EMAIL_ or
GATEWAY_ALLOWED_USERS list reaches nothing.
- The default still drops.
On macOS, `hermes gateway install --no-start-now` started the gateway anyway.
`_cmd_install` forwarded the start flags to the systemd and Windows backends but
called `launchd_install(force)` alone, and `launchd_install` always ran
`launchctl bootstrap`. The plist sets RunAtLoad, so bootstrapping it starts the
gateway immediately, and the command still printed "Service installed and
loaded!". The setup wizard had the same gap: answering No to "Start the gateway
now?" and Yes to login auto-start still called `launchd_install(force=False)` and
started the gateway on the spot.
launchd_install now takes start_now. When it is False and launchd is not already
running the gateway, the install writes the plist and does not load it. It also
boots out any idle registration left from before, such as a job parked after a
clean exit, because `hermes gateway start` would kickstart that registration's
old definition instead of loading the new plist. The outdated-plist repair takes
the same path, since its bootout/bootstrap reload would start a stopped gateway.
A gateway that launchd already runs is reloaded as before, not stopped. With
the plist in ~/Library/LaunchAgents, the gateway starts at the next login or on
`hermes gateway start`. `_cmd_install` and the wizard now pass the answer
through.
Measured with launchctl recorded rather than run: `install --no-start-now` went
from 1 bootstrap to 0, the wizard's No went from 1 bootstrap to 0, and the
repair of an outdated plist with a stopped gateway went from a reload to a
rewrite only. The --no-start-on-login half on launchd is #91549 and is not
touched here.
* refactor(fallback): share the pinned-owner chain rule
delegate_task's _resolve_child_fallback_chain decides which fallback chain
a child may walk: a pinned child never borrows the parent chain, an explicit
[] disables fallback, a declared list is the child's own. Cron needs the
same rule for pinned jobs (#100437), so the body moves to
hermes_cli.fallback_config.scoped_fallback_chain and the delegation helper
becomes a thin caller. Behaviour is unchanged; the delegation matrix test
still pins every cell.
* fix(cron): a pinned job never falls back to the global chain
A job with its own provider, model or base_url is an explicit operator pin
(since 0469740ab3 unpinned jobs store none of these). It still walked the
global fallback_providers chain in two places, so a pinned job could run
on a different provider and model than the one chosen:
- _resolve_job_runtime walked the chain on an AuthError or transient
network failure while resolving the pinned primary;
- _resolve_cron_agent_setup handed the global chain to every cron agent as
fallback_model, so the conversation loop's provider ladder could swap a
pinned job mid-run.
Both now read _job_fallback_chain(job, cfg), which returns no chain for a
pinned job through the same scoped_fallback_chain rule delegate_task uses
for pinned children. The pre-dispatch key check reads it too: the global
chain used to skip that check for every job, so a pinned job with a
missing key now blocks before the agent is built instead of failing in the
resolver. The transient-failure notice for a pinned job says it does not
fall back and names --unpin, instead of "No backup provider succeeded".
Unpinned jobs (including legacy *_snapshot records) and same-provider
credential-pool rotation are unchanged. The two scheduler tests that
asserted atomic provider+model fallback swaps used pinned jobs; they now
use unpinned jobs and keep the same assertions.
No per-job fallback_providers list: jobs have no generic override field
(create_job/update_job, the cronjob tool schema and the CLI enumerate each
field), so an opt-in chain would be a new surface on all of them. The
escape hatch is to leave the job unpinned and pick its model with
cron.model / cron.model_provider.
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
* docs(cron): pinned jobs do not use fallback_providers
cron.md "Provider recovery" and the pre-dispatch key check, the cron rows
and section in fallback-providers.md, and the developer notes in
cron-internals.md / provider-runtime.md said every cron job inherits the
global chain. State the new rule, the compatibility note for users who
relied on a pinned job landing on the chain, and the unpinned + cron.model
alternative.
---------
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
Widen the two salvaged fixes to the whole class and make the refusal
supervisor-safe:
- `process_identity.is_desktop_owned_backend()` is the single discriminator
(HERMES_DESKTOP=1 AND the per-spawn HERMES_DASHBOARD_SESSION_TOKEN). The
attach bypass, the named-profile reroute, the env sanitizer, the MCP
discovery timing and the host-rendezvous publish all key on it now; a shell
that merely inherited the flag from Desktop is treated as a normal launch
(#119210). The publish skip from #119832 keyed on the bare flag, which
would have hidden a supervised service launched from a Desktop shell.
- `_attach_to_host_backend` refuses with exit 78 (EX_CONFIG) instead of 1.
Exit 1 under `Restart=always` was an infinite restart loop with nothing
listening on the ingress port (2035 restarts, #119824); 78 is the
deliberate-refusal code `RestartPreventExitStatus=78` parks on, the same
contract the gateway unit already uses.
- Docs: the systemd example carries RestartPreventExitStatus=78 and the
user-side remedy (stop the owner, or `--isolated`).
- Tests trimmed to ≤2 invariants per fix; the control case in the desktop
tests sets the token so it models a real pool child.
The host lock is the only arbiter of two gateways starting together (the
attach check reads a record published a moment after the owner's claim),
and its loser exits 75 so the supervisor retries into attach. The PID
claim passed `force or replace`, so `--replace` skipped that refusal too.
Every unit Hermes generates runs `gateway run --replace`, which turned
the arbiter off for exactly the race it exists for: a unit that saw no
owner yet started a second gateway beside the multiplexer, and the two
fought over the same bot tokens (a --replace token handoff SIGTERMs the
holder).
Only --force skips it now. A --replace that took the owner over already
freed the lock with that process, so real takeovers are unchanged.
The api_server platform builds a fresh AIAgent per request (per-request
callbacks, model route, ephemeral prompt), so the memory provider was
re-initialised on every request. External providers deliver recall as the
PREVIOUS turn's background prefetch held on the provider instance, so a
continued session (X-Hermes-Session-Id, previous_response_id, declared
session key) never received automatic recall, and for hindsight
local_embedded each init also restarted the embedded daemon, killing the
retain still in flight. Pre-existing: the same probe fails on main before
the hindsight catalog migration (526d135a96, bundled provider).
ApiServerMemorySessions parks the session's initialised MemoryManager
between requests (exclusive check-out/check-in, keyed by profile home +
session id, LRU/idle eviction under the owning profile's scope) and
AIAgent(memory_manager=...) adopts it instead of loading and initialising
the provider again. /v1/chat/completions, /v1/responses, session chat and
/v1/runs all go through the same two seams (_create_agent, turn finally).
--clone carried memory.provider (e.g. hindsight) in config.yaml but not the
provider's own config under the profile home (hindsight/config.json,
mem0.json, ...), so the clone booted with the provider selected and silently
unavailable. Copy the ACTIVE provider's <provider>/ dir and/or <provider>.json
by the same convention the dashboard memory-provider routers read, guard the
name against path traversal, tighten copies to 0600 (they can hold an API
key), and say so in the CLI notice. Convention-based on purpose: hindsight is
a catalog plugin now, so a hook the plugin must implement could not fix the
reported case, and no provider module is imported during profile create.
Supersedes #43107 (credit @bionicbutterfly13 for the direction).
Watching the bot drive its screen is the point of Bot Screen, but until now
you learned it had happened only afterwards. With "Open Screen when the bot
uses it" checked on a bot's row menu, the first live tool.start for a screen
tool (computer_use, browser_*) on a session the bot owns brings its Screen
tab forward.
Why opt-in and fenced: Desktop's rule is offer, don't hijack. The raise is
per bot (BotMeta.screenAutoOpen, rides profile ui_meta like pin/hide), never
moves keyboard focus (openBotScreen reveals the pane; the viewer grabs keys
only on Take over), never fires for replayed history (the reconnect replay
re-dispatches parked frames, so the wake is rate-limited to one per bot per
30 s rather than trusting seq), does nothing while the tab is open, and a
manual Close mid-run holds until the run has been quiet for a cooldown.
Live (isolated headless Electron + real serve backend, CDP): opt-in off →
tool.start opens nothing; toggled via the real menu → toast + aria-checked
true → tool.start opens "Hermes · Screen" (Screen is off state), focus stays
on the row; second call is a no-op; real Close → held at +2 s and +29 s,
raised again at +31 s.
Rule borrowed from thomasbek3/hermes-bot-kit computer-viewer's auto-connect.
A fresh dashboard /chat always spawned the TUI in the dashboard process's
launch directory, so from a phone or any browser there was no way to aim
a new session at a specific repository. Amp's runners now serve many
directories and the web composer offers a picker of the runner's projects
and discovered git checkouts; this ports that mechanism onto the surface
Hermes already has: the dashboard is the phone/web front, the host's
projects.db + repo-discovery cache are the served directories.
- GET /api/chat/workspaces: the profile's projects (with folders) and
discovered repos (session-derived + scanned), default_cwd, home;
?scan=1 rescans desktop.repo_scan_roots on the host, so headless
installs (no Desktop to populate the cache) discover repos too.
- /api/pty?cwd=<dir>: validated (existing directory, fail-closed 400
via the PTY error path) and forwarded to the TUI child as HERMES_CWD
(self-spawned gateway cwd) + HERMES_TUI_CWD (explicit cwd on
session.create for the dashboard's in-memory gateway, whose own cwd
is the launch dir). Resumed sessions ignore it.
- ui-tui: session.create carries cwd when HERMES_TUI_CWD is set, so
/new inside a dashboard chat stays in the picked workspace too.
- web: workspace selector in the Chat rail above "New chat" (projects,
repos by recency, Other path…, rescan), remembered per profile in
localStorage; 17 locales.
- docs: web-dashboard.md rail + REST sections.
Live E2E: real `hermes dashboard` under a scratch HOME with two git
repos under desktop.repo_scan_roots -> /api/chat/workspaces?scan=1
lists both; /api/pty?cwd=repo-a -> TUI status bar shows ~/code/repo-a;
/api/pty?cwd=<missing> -> "Working directory does not exist" + close.
memory-providers.md points the Hindsight section at the catalog entry and the
upstream integration docs, documents the auto-install on `hermes update` /
first agent start (and the allow_lazy_installs=false one-liner), and adds a
"Migrating from bundled Hindsight" subsection. memory.md, overview.md,
integrations/index.md, cli-commands.md, docker.md and nix-setup.md no longer
say hindsight ships in-tree or as the [hindsight] extra.
The onboarding card needs to list catalog plugins beside the hosted
connectors (NS-960 D1, D4) and grey a plugin whose app is absent (D5).
The catalog had no curated flag, and the manage_catalog row left
app_state empty.
- `onboarding: true` and `title` on catalog entries (loader, validator,
docs); set on blender, nvidia-app and nvidia-broadcast.
- hermes_cli/plugin_catalog_presence.py reads the plugin.json at the
catalog's pinned commit once per pin and judges its app declaration with
the hermes_platform resolver the installer and the Plugins-tab pill use.
No declaration or an unreadable one is `unknown`, never `present`.
- `plugins.manage action=onboarding` lists the curated entries this OS
runs (platform mismatch is the only exclusion) with app_state and the
sentence the card greys the row with.
- manage_catalog plugin rows now carry app_state and the catalog title.
- A live entry that differs from the in-tree entry at the same pin (new
metadata) now follows the same newer-catalog rule as a new pin, so a
checkout that adds `onboarding` is not masked by a published doc that
predates it.
Multiplex-only is the direction. This key exists so fleets that lost per-profile gateways in
the switch keep working while the remaining gaps (per-profile stop/restart, WhatsApp
bridge/relay on secondaries, dashboard scoping) are closed, and it goes away once they are.
Every surface that names the key now says so through one shared notice
(STANDALONE_DEPRECATION_NOTICE): the install/start refusal hint, per-profile and default
`gateway status`, the migrate plan, and the host boot log (WARNING, not INFO). The user
guide carries a deprecation admonition where the key is introduced and no longer describes
it as opting out "for good".
Also: the Windows cold-start test fixture stubbed profiles_to_serve(multiplex) without the
new include_standalone kwarg (the PR's one red on CI).
Flag list entry, the opt-out path in "No new per-profile gateways", what the
host does on rescan (<=30s, no restart), what happens when the key is removed
while the profile's own gateway runs, and that a standalone profile's cron,
webhook ingress and kanban notifications run only under its own gateway.
The setup profile's `setup` toolset was empty. It now carries one tool, manage_catalog:
- search: catalog plugins (the Plugins tab's live catalog resolver) and hub skills, with
whether each is already installed in the default profile. Read-only.
- install: opens the same connection operation manage_connections opens, with rows of kind
plugin / skill. Nothing installs until the user approves a row. An approved row installs
into `default` (or the profile the Advanced modal named) through dashboard_install_plugin /
the hub's headless install, so the catalog pin, kill list, security scan and live
activation (#119644) are the host's. The row settles with the live MCP tool names and the
plugin's skill.
The model sends catalog ids and an action only; every other key is refused before anything
runs. An unknown id or a plugin this OS cannot run is drawn failed with the installer's own
text. Anywhere a catalog card cannot be drawn (TUI, CLI, messaging, registry dispatch) the
result is the `hermes plugins install` / `hermes skills install` pointer.
- contract: plugin/skill targets follow the MCP transitions.
- run.apply_answer / reissue route a card answer to the module that owns the operation.
- tool_search: `setup` joins the direct-surface toolsets, so the guide's one tool is never
deferred behind tool_search.
- docs: tools reference, toolsets reference, plugin catalog page.
Linear NS-964.
The layout editor gets an Interface mode section above the templates —
Simple "For talking to Hermes. Sidebar and chat; no terminal, file or diff
panes." / Advanced "For developers. Terminal, files, diffs, statusbar and
layouts, as you set them." In Simple the shelf is Sidebar left / Sidebar
right and there is nothing to arrange or save. Settings → Appearance →
Window & layout carries the same control; rows whose value Simple decides
say so. ⌘K "Simple mode" and the rebindable `view.toggleSimpleMode` are the
guaranteed way back. First launch: Basic applies Simple, Elite Advanced.
Six locales, docs, and a Playwright round trip.
`snapshot_paths` / `_store_blob` write blobs OUTSIDE `_ledger_lock`, seconds
before the row that references them is appended. Since `_maintain_size` now
runs `gc_blobs()` on any append that trims, a sweep in another process during
that window deleted the in-flight capture: its row then landed with sha256
values `read_blob` cannot resolve and `rollback_entry` failed.
`_gc_blobs_locked` now skips any unreferenced blob file (incl. `.tmp-*`) whose
st_mtime is newer than `_BLOB_GC_GRACE_SECS = 3600`; no new config key.
Docs (curator.md) and the `gc_blobs` docstring say so.
PROOF: probes/s5_final_inflight_blob.py before -> `P1 blob survives P2's
sweep: False ... read_blob: None`; after -> `True ... read_blob: b'P1
in-flight skill file'`. test_concurrent_appends_never_lose_a_middle_row gained
a fresh + 2h-aged unreferenced blob pair: red with the grace set to 0 (fresh
blob deleted), green at 3600 (aged deleted, fresh kept). ruff clean,
check-windows-footguns --all clean, run_tests.sh ledger + AX files 29 passed.
_trim_oldest never parses lines: it drops the OLDEST ones whatever they
are, so 'malformed lines always survive' was false. Reword docstring,
curator.md and the kept test's docstring to the true guarantee: lines in
the retained tail are never rewritten or parsed, so a malformed line there
survives verbatim. No behaviour change.
PROOF: probes/s5_trim_probe.py (cap 4096, malformed first line, 4 fat
appends) -> old malformed survived: False, contradicting the old wording;
the kept test asserts a LAST-line malformed row, which is what the new
wording promises.
* feat(connectors): the backend serves a connector's tool list, cached for 24 hours
The Connectors page opens one app and shows every tool it has. The backend
had no way to read that list.
- `tools/connectors/portal/`: a client for the portal's tool-list route and a
JSON cache under the Hermes home, one file per portal origin and connector.
An entry is fresh for 24 hours. After that the read revalidates with the
stored ETag: 304 keeps the list, 404 deletes the entry, an upstream failure
serves the stored list marked stale, and a 401 never serves the cache.
- `connectors.tools {slug, refresh}`: account-level, routed by `profile`, no
chat session. Errors carry a fixed `reason` from one closed set on the rail.
- Every connector model that is not operation state moves into
`tui_gateway/contracts/connectors.py`. Handlers that no chat session owns
live in `tui_gateway/methods_connectors_account.py`.
The wire model is tolerant: an unknown facet reads as unclassified and one odd
tool never blanks a connector.
* feat(connectors): catalog, accounts and member tool rules by RPC
The Connectors page needs the app catalog, the connected account of one app,
a way to disconnect it, and the member's own on/off rules. None had an RPC.
- `connectors.catalog`: name, description, category and logo of each app.
- `connectors.accounts`, `connectors.accounts.remove`: read the accounts at
the tool gateway and remove one by id.
- `connectors.policy.get`: the rule layers that apply to the member, widest
first. The body is a union on `mode`, so a reader can name who turned a
tool off.
- `connectors.policy.set`: one change, a union on `type` (the tools of one
connector, or one connector on or off), with the revision the user saw. A
stale revision answers `POLICY_CONFLICT`. The backend composes the upstream
write in one pure function, so no renderer learns the upstream rules.
- Bundled MCP manifests can name their hosted twin with `connector:`, so the
page can show one card per app.
* feat(connectors): connect an app without a chat session
Every connector RPC took a `session_id`, and a connect that did not come from
the model's tool call minted a link with no watcher. The Connectors page has
no chat session, and its card must flip to connected by itself.
- `connectors.list`, `connectors.connect`, `connectors.operation.status`,
`connectors.operation.wake` and `connection.respond` take `owner`, a union
on `type`: `session` (today's behaviour and authorization) or `account`
(routed by `profile`, authorized by the live transport like `mcp.*`).
`session_id` is gone from these params; every desktop caller sends `owner`.
- An account connect runs the same operation lifecycle on a background
thread, under the profile's scope, so the watcher reads the account and
settles the operation. A second connect for an app that is already
connecting returns the open operation and mints nothing.
- `connection.update` carries `owner`. An account operation has no session to
address, so its updates go out on the session-less broadcast path.
* feat(mcp-catalog): eighteen more bundled entries name their hosted connector
A bundled MCP entry and a hosted connector for the same app are one card
on the Connectors page only when the manifest names its hosted twin.
Linear and Notion had the field. These entries get it too: airtable,
asana, attio, calendly, dropbox, figma, railway, supabase, todoist,
betterstack, canva, cloudflare, datadog, intercom, neon, sentry, stripe
and vercel. Atlassian maps to two hosted connectors and Prisma Postgres
is not clearly the same app, so both stay without one.
* refactor(connectors): the account handlers share one gate, one params model and one write table
The six account-level handlers each repeated the availability gate, the
auth catch and the catch-all reply. One decorator now owns that, and each
handler validates its params with its contract model instead of a ladder
of isinstance checks. The five connection RPCs share one guard for the
unexpected-failure reply.
The four write composers for the member rules were the same function
with a different list key and polarity. They are one table now.
The owner union lives in contracts/common.py, so the params side and the
event side stop declaring it twice and the import cycle is gone.
An account operation start carries one event and a flag, so the wait for
the sign-in link blocks instead of polling every 50 ms. run_operation
loses its two account-only parameters; drive_operation is the second
entry point.
Tests: four deleted (they exercised pydantic or the mock), three merged
into tables, two added (a client that still sends the old top-level
session_id is refused; all six account RPCs run off the server loop).
The shared reply helper and the HTTP and managed-client fakes move to
one place each. Comments are one line or gone.
* fix(connectors): a missing tool-list route reads as "unavailable", not "connector gone"
The tool-list read treated every 404 as the portal's "this connector is
not in the catalog" answer. It deleted the cache entry and answered
CONNECTOR_NOT_FOUND, so a page would offer to remove an app that is
connected and works. A portal that does not serve the route yet answers
a bare 404 for every app.
Only the portal's own {"error": "connector_not_found"} means the
connector is gone. Any other 404 is now a tool-list outage: the cached
list is served as stale, or the RPC answers TOOLS_UNAVAILABLE.
* fix(connectors): a connect from the page returns to the app after sign-in
The sign-in link carries a return target only when the session's surface
is the desktop. A chat session binds that surface. An account-owned call
has no chat session, so nothing bound it: the link was minted without a
return target and the browser ended on the portal's done page instead of
coming back to Hermes.
Every account-owned call now runs with the process's own surface bound,
next to its profile scope. The operation thread copies that context, so
the first link and every reissued link carry the return target and the
operation id.
* test(connectors): defer the new connector RPC coverage
The tests for the new account RPCs, the portal client, the tool-list cache
and the rule composer leave this PR and come back in one later change, after
the API is settled. The same was done for #111008.
Kept: the edits that existing tests need because the five connection RPCs
now take `owner` instead of `session_id`, and the rename of the managed
client seam.
Removed: six new test files, their two fakes and the gateway conftest, and
the new cases in test_mcp_catalog.py, test_connectors_gateway_client.py,
gateway-rpc.test.ts and notifications.test.ts. Reverting this commit restores
all of them.
* fix(cli): the connection panel hands the tool thread back at once
The classic CLI's connection callback waited on a queue for the user's first
decision. The operation's watcher starts only after the callback returns, and
the watcher is what polls a hosted account, runs the 300-second deadline and
sees Ctrl+C.
For a hosted connector the panel opens on the sign-in link, where the only
key that filled the queue was Cancel. The account was never polled: the user
signed in, the panel never changed, and Esc reported the app as skipped.
Ctrl+C set the interrupt flag but left the thread parked on the queue, so the
turn never ended.
The callback now opens the panel and returns, as the gateway's callback does
for the desktop and the Ink TUI. The panel's actions already reach the
operation through apply_answer on the UI thread, so the queue is removed. An
install with a form still waits for Connect, because the backend starts no
work for a pending row. Ctrl+C now settles the operation as `interrupt`, and
open rows become `not_connected`.
Checked on the e2e rig with the fake tool gateway: hosted connect completes on
the third status read; Ctrl+C ends the turn and the polling stops; an MCP
install with a plain and a secret field still saves config and both values.
* fix(connectors): "run it again" lives in the library, so the classic CLI can use it
Making a new sign-in link for a failed or expired hosted connector was
implemented only in the JSON-RPC layer (`_reissue`). The classic CLI does not
go through JSON-RPC: its Connect button on a failed row called apply_answer,
which does nothing for a hosted operation because it has no MCP runner. The
panel showed "Waiting…" until the deadline.
`tools.connectors.run.reissue(operation, names)` now holds the checks and the
per-kind action, and returns a refusal reason or None. The gateway maps each
reason to the same JSON-RPC error as before. The CLI calls it for a hosted
row; a refusal is shown on the row. MCP rows keep their path, because Connect
on a failed MCP row re-sends the form values.
Checked on the e2e rig: a scripted failed sign-in, then Connect: a second mint
with `reinitiate: true`, a new link with a new connection id, then connected.
* feat(connectors): the account list and disconnect go through the portal
`connectors.accounts` and `connectors.accounts.remove` called the tool
gateway. They now call the portal's account-management routes
(`GET /api/v1/connectors/accounts`, `DELETE /api/v1/connectors/accounts/{id}`),
which apply the organisation membership checks and write the disconnect audit
row. There is no fallback to the gateway when the portal is unavailable, and a
removal is never retried.
The read of ONE account stays on the gateway (`GET v1/connectors/accounts/{id}`):
the portal has no such route, and the operation watcher polls it once per second.
`ConnectorClient.list_accounts` and `delete_account` are removed. The removed
account's reply model carries `connector`, which both services send.
* fix(connectors): the account RPCs answer what the portal really sends
Checked against the portal source and against the staging and production
services.
- Errors are read from the upstream error code, not the HTTP status. A rule
write answered 409 for a stale revision and for a user with no organisation;
both read as "the policy changed". `org_required` is now `ORG_REQUIRED` and
403 `no_access` is `ORG_ACCESS_DENIED` on every account RPC; only a rejected
sign-in is `NEEDS_NOUS_AUTH`. `connectors.list` and `connectors.connect` with
the account owner map these too.
- `connectors.policy.get` and `connectors.policy.set` carry `effective`: the
portal's own result for this user, with its stamp and without provider or
subject ids. Nothing is recomputed locally.
- A rule write needs the revision the user saw: `expected_revision` is required
and must be a revision string; a bad one is refused before any HTTP call.
- A tool row carries `no_auth`; a list without the upstream flag is an invalid
answer, not `false`.
- `connectors.accounts.remove` returns the app of the removed account. An
invalid id is `INVALID_PARAMS`.
- The tool-list cache is per signed-in member (a hash of the token's `sub`),
so two Nous accounts on one profile do not share entries.
- A malformed slug is a local error, not a 404 from a server nobody called.
Live, staging: no revision and a malformed revision refused locally; a good
revision wrote one disabled Gmail tool and returned it in `effective`; the
same revision again answered `POLICY_CONFLICT`; the list row showed the tool;
the restore brought the member rules back to the start. Live, staging and
production, read-only: all 60 tool lists (5483 tools) parse.
* fix(connectors): the operation RPCs match their contract; a settled card cannot start a new link
Found by two adversarial reviews of the RPC layer and its types.
- `connectors.connect` from a chat session with no open operation is refused
(`UNKNOWN_OPERATION`). It used to call `manage_connections` through the tool
registry with no card: it made a link nobody watched, returned a reply
without the required `settled` field, and named an operation that was never
registered. There is one way into an operation: the agent's call, or the
account owner's `connectors.connect`. "Run it again" inside an open
operation is unchanged.
- `connection.update` for a session is routed by session key AND profile; two
profiles with the same key no longer cross-deliver a sign-in link. The event
payload gets the same redaction as the RPC replies.
- `connection.respond` runs on the long-handler pool: an approval can start MCP
OAuth discovery, which blocked every RPC of the gateway while it ran.
- `connectors.list` rows are a closed snake_case model: `connector`, `enabled`,
`connected`, `connection_status`, `status_reason`, `gateway_disabled_tools`.
The last one is display data: the gateway enforces the rules, the backend
only passes the list on. The phantom `name` and `description` are gone, and
the desktop uses the generated types instead of hand-written copies.
- `tools_listing` (model-only data) no longer rides on `connectors.operation.status`.
- `unavailable` is removed from the target states and settle reasons: nothing
produces it. The contract generator now fails when a contract enum and its
domain enum differ.
- `ConnectorErrorReason` is part of the generated TypeScript and OpenRPC.
- The desktop sends `connection.respond` on the socket that holds the session,
as wake and reissue already did.
- Contract violations are logged every time, at error level.
- An account connect whose prepare step is slow returns the live operation
instead of an error while the operation keeps running.
- The MCP-manifest `connector` field leaves this PR (it moves to a later one
on top of the catalog-reader change). `hermes_cli/mcp_catalog.py` and
`optional-mcps/` are untouched by this PR again.
anti-slop: no net-new findings (15 touched files).
* fix(connectors): the model gets no sign-in link wherever a card exists; side agents cannot connect
The flag that tells the model "a connection card exists" was the session
platform (`== "desktop"`). The Ink TUI and the classic CLI also draw a card,
so there a connector call on an unconnected app handed the model the raw
`connect_url` and told it to pass the link to the user.
- The agent turn now declares how a link can reach the user
(`tools/connectors/turn.py`): CARD when the agent was built with a
connection callback, SIDE for a subagent or a background turn, LINK for a
headless run (`-q`, cron, ACP, api_server, messaging). It is set once per
tool batch in the agent loop and read by the connector dispatch path, which
never sees the agent. The session platform decides return-to-app only.
- CARD: the result carries `connect_card_available` and our hint, never the
link and never the gateway's own hint.
- SIDE: subagents (`delegate_tool`), gateway background turns and the classic
CLI `/bg` are built with `side_agent=True`. They hold no `manage_connections`
tool on any path that derives the tool list, and a connector call on an
unconnected app gets no link, only "report this to the main agent".
- LINK is unchanged.
- The hosted path with no card builds a detached operation, as the MCP path
does, so no `connection.update` is emitted for an operation no client asked
for. Names and docstrings that said "off desktop" now say "no card".
- A settled card is dead on the desktop: `reissueConnectionTarget` and
`respondToConnectionRequest` share one guard and send nothing for a settled
or unknown operation.
- The model-facing settled result no longer carries `connection_id`; the model
repeated it to the user.
Shown on the real clients with a real model (rig, fake tool gateway): Ink TUI
and classic CLI get `connect_card_available` and no link, the model opens the
card, the account connects, the retried call succeeds; `-q` still gets the
link; a subagent and a background turn have no `manage_connections` and get
the no-link hint; on the desktop a card settled with Continue has no enabled
control and sends no RPC.
* feat(tools): every call made through tool_search + tool_call shows a real label on all three clients
A bridged call showed as a generic `tool_call` row in the Ink TUI and as
`⚡ tool_call` in the classic CLI, because the display looked the name up in
the tool registry and bridged names are made at run time. The desktop labelled
only batches that were all hosted connector calls, by parsing names itself.
- `tools/tool_labels.py` is the one place that turns a bridged call into a
label: kind, app, action, emoji and text. Hosted: `connectors__gmail__GMAIL_SEND_EMAIL`
→ "Gmail · send email". MCP: "Linear · list issues". A local deferred tool
keeps its own emoji, verb and primary-argument preview. A batch gets exactly
one label per entry, always; an entry with no name gets a generic label.
- Classic CLI: one row per inner call; the duration on the last row; the
failure text on the row of the call that failed. With friendly labels off
it prints what it printed before.
- Gateway: tool start, progress and complete events and stored transcript rows
carry a typed `labels` field. It does not depend on the classic CLI's
display setting. Clients no longer parse tool names.
- Ink TUI: rows from the labels; the verbose trail keeps Args and Result.
- Desktop: `ConnectorExecution` renders hosted, MCP and mixed turns from the
labels, one row per call. The labels reach the row under a key no tool
argument can use. The connect card it drew under a failed tool result is
gone: after `CONNECTION_REQUIRED` the one way in is the agent's own
`manage_connections` call.
- `tool_search` and `tool_describe` rows read "Searching tools · <query>" and
"Reading tool details · N tools".
Shown on the real desktop (video and screenshots), the Ink TUI and the classic
CLI with the rig: hosted rows, MCP rows, a two-entry batch, a failed entry, a
`CONNECTION_REQUIRED` row with no card under it, labels after a reload, and the
desktop rows with the classic CLI setting off.
* fix(connectors): the model can tell "hosted tools unavailable" from "no such tool"; manage_connections routes MCP names correctly
- A failed hosted search or describe used to return nothing, by design, so the
model saw only local tools and told the user that a connected app was
missing. The local results are unchanged; when the hosted leg failed, the
`tool_search` and `tool_describe` results carry
`connectors: {status: "unavailable", reason: "unreachable" | "sign_in_expired"}`
and one hint line. A rejected token is `sign_in_expired`; an entitlement
refusal or a shut gate adds nothing. `tool_describe` no longer lists those
names under `not_found` next to "search again".
- NS-932. The description now says which side a name belongs to: a bare name
is a hosted connector account; `mcp: true` only when the user asks for an MCP
server, a local server or an install, or when the name exists only in the
catalog; connect and reconnect are hosted verbs, install, enable and
authorize are MCP verbs. It names the three clients that draw a card.
- A misrouted target is refused with the call that works. Only when the
gateway does not know the connector (confirmed on that failure path) and the
name is a catalog entry does the target fail with "X is a local MCP server.
Call manage_connections with action install ...". It is a per-target
outcome: other targets of the same call keep their links and their card. A
vendor failure on a name both sides know stays an ordinary failed row. The
MCP side mirrors it, and never for an entry that is only not installed.
- "Do not re-ask after a skip or a timeout" no longer stops the model when the
USER asks for that app again; the description and the settled-result notes
say so. A builder saw the model refuse a direct user request.
Shown on the Ink TUI and the classic CLI with a real model: a dead gateway and
a 401; "connect fxmail" goes hosted; "install the fx-noauth MCP server" goes
MCP; "connect fx-noauth" reaches the MCP install card in one corrective round
with no hosted mint; a two-target call where one is misrouted still connects
the other with exactly one mint.
* fix(tui): the connection card answers every key, shows what is happening, and is dead once settled
Reproduced on the real Ink TUI with the rig, then fixed:
- The keyboard was dead during the sign-in wait: the card kept a `submitting`
flag that the normal OAuth path never cleared, and Esc went through the same
guard. The in-flight state now belongs to the answered row and clears when
that row moves, when any later frame of the operation arrives, or after
five seconds. Esc skips the row in every phase; Ctrl+C interrupts the turn
(the input handler had no branch for this overlay); Shift+arrows scroll the
transcript and the card ignores them; arrow keys no longer move the text
cursor and the field focus at once.
- The card was lost at turn idle: the overlay flag was cleared while the
operation stayed in the store, and a resume dropped the pending card. The
flag survives idle, a resume shows the pending card again, a session switch
clears it.
- States with no branch: `not_connected` and a row with no link fell into the
credential form; `expired` vanished with no note. The title and the row text
now name the action (connect, reconnect, install, enable, authorize); a
failed or expired row with no fields offers Try again / Skip; a failed row
WITH fields reopens the form over the typed draft, with the failure above it.
- A settled card is dead: at settle the overlay closes and one transcript line
per app states the outcome. A settled or dismissed operation id is
remembered, so no replay or resume can reopen its card. Esc in the last
"Finishing…" moment hides the card and still writes the outcome lines.
- A failed `connection.respond` and a browser that did not open are shown on
the card in one sentence.
Also: `tui_gateway/connector_payload.py` redacted the BOOLEAN `secret` flag of
a credential field to the string "[REDACTED]". On the desktop every credential
field therefore rendered as a password and lost its prefilled default. A
boolean is no longer redacted.
* chore(connectors): remove the comments and docstrings this branch added
Deletions only. Kept: tool directives (`# noqa`, `// eslint-disable`, ...),
`// SAFETY:` lines, and the docstrings of the contract models under
`tui_gateway/contracts/`, which become the descriptions in the generated
OpenRPC and TypeScript.
Checked that no code changed: every Python file has the same AST as before
once docstrings and `pass` are ignored (62 files), and every TypeScript file
prints the same with comments stripped by the TypeScript printer (32 files).
The generated contract files are unchanged.
* fix(connectors): a card restored after a reload answers again; every account RPC names auth and org failures
Found by the end-to-end runs on the pushed head.
- Desktop: after a window reload, Continue on the restored card sent nothing.
The answer looked up the backend that holds the session with the runtime
session id, the lookup wants the stored id, and a failed lookup returned
silently. When the lookup gives no owner the answer now goes out on the
window's active socket, which is what main does.
- `connectors.policy.get` answered `POLICY_UNAVAILABLE` for a rejected sign-in,
a refused scope, a non-member and a missing organisation alike: the handler
runs with the gateway's globals and did not import the reason enum, so its
own error mapping raised. `connectors.accounts.remove` caught auth failures
in its generic branch. `org_required` was mapped on `policy.set` only. All
six account RPCs now answer `NEEDS_NOUS_AUTH`, `FORBIDDEN_SCOPE`,
`ORG_ACCESS_DENIED` and `ORG_REQUIRED` for those four upstream answers.
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.
- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
post_api_request payload shape plus `aux_task`, `api_request_id`
(`aux-...`, shared by every attempt of one logical call), `retry_count`,
`streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
turn is in flight; fail-open (a raising/hung subscriber is logged and
the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
attempt shares (_relay_sync_completion / _relay_async_completion /
_relay_sync_stream) run under the hook pair — retries and fallbacks
included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
aux_task and no api_request events; raising subscriber never breaks
the call). First is red on origin/main.
Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.
Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
A plugin whose import or register() never returns (an infinite loop, a blocking
network call) held PluginManager.discover_and_load() forever, and with it every
synchronous caller: `hermes chat`, gateway startup, ACP session/new (#108139).
Each plugin's import + register() now runs under `plugins.load_timeout_seconds`
(default 10, 0 disables, max 600) on a daemon worker. On overrun the plugin is
recorded as failed with "load timed out after Ns" (same channel as every other
load failure: startup WARNING, `/plugins`, `list_plugins()`), its pre-hang
registrations are disposed, and discovery continues with the next plugin. The
abandoned worker's later `ctx.register_*`/`subscribe`/`on_unload` calls are
refused with a WARNING (the context is marked abandoned), so a late registration
can never land in a registry the failure path already swept. Abandoned loaders
are capped per process (8); past the cap further loads are refused with a named
reason rather than run inline, which would recreate the hang (#98382 shape).
Because the worker cannot own the caller's RLocks: the deferred-platform eager
fallback now runs outside the replacement transaction, discovery re-entered from
a loader worker returns on the already-set discovered flag instead of blocking on
the sweep's lock, and such a worker never joins the background discovery thread
that is waiting on it.
website/docs/user-guide/features/kanban.md:987-988 listed the two gc
retention flags without saying what the edge values do; add "(negative N
is rejected; 0 disables that sweep)" so users don't have to read the code.
hermes_cli/kanban_db.py:4279,4293 — gc_events/gc_worker_logs accept
older_than_seconds=0 as "everything older than now" while _cmd_gc maps
days=0 to "disabled" before calling them. The docstrings did not state
that split, so a library caller could assume 0 is a no-op. One line each.
Gate findings: D.2ab.md:27, D.2c.md:32.
Since the first-strike escalation, a closed transport (socket_closed /
client_closed) forces the reconnect on the first unhealthy sample and the
failure threshold only applies to soft signals (ack staleness, latency,
event silence). The user guide (website/docs/user-guide/messaging/discord.md:89)
and the env-var reference (website/docs/reference/environment-variables.md:780)
still described the threshold as gating every unhealthy sample, so an
operator reading `1/2` followed by a forced reconnect would think the
knob was ignored. One sentence each, citing #118487.
Annotated-tag pins (F8): a catalog `sha` recorded as `git rev-parse <tag>` names
the TAG object, while HEAD can only ever be the commit it points at. The scan
trust check compared HEAD against the unpeeled sha (so every tag-pinned entry
lost the reviewed-pin bypass and prompted on caution findings) and the sidecar
recorded the peeled commit, so `update_available` was true forever and every
`update` re-installed. The installer now peels the pin (`<sha>^{commit}`) for
trust, records `pin` on the catalog block only when the checkout satisfies it
(empty for an off-pin `--ref` install), and every at-pin check goes through
`at_catalog_pin(sidecar, entry_sha)` (repin, dashboard payload, TUI rows).
Re-pin consent (F10): `hermes plugins update` on a catalog install replaced the
tree without asking, even when the new pin declared new tools, hooks, Python
dependencies, host capabilities or a Desktop half. `repin_catalog_plugin` now
diffs the installed manifest against the staged clone BEFORE anything moves
(`_install_plugin_core(before_swap=...)`) and, on a widening:
- CLI: prints the delta and asks y/N (non-interactive → not applied, fail
closed); after a changed re-pin it runs the same `_run_capability_consent`
grant path as the git-pull `update`.
- `plugins.manage update` RPC and the dashboard REST route answer
`{ok: false, consent_required: true, delta, delta_lines}` with nothing
changed; a retry with `accept_capabilities: true` applies it. Desktop shows
the delta in its confirm dialog; the web dashboard uses `window.confirm`.
- Gateway contract regenerated (`accept_capabilities` param; `consent_required`,
`delta`, `delta_lines`, `error` result fields).
Catalog audit findings F8 and F10 (low severity, no issue filed).
The cached live catalog (`HERMES_HOME/cache/plugin-catalog.json`) won over the
in-tree copy unconditionally: right after `hermes update` bumped an in-tree pin
a fresh (<6 h) cache still installed the previous sha, and offline a cache of
ANY age (90 days in the audit probe) outranked the checkout's catalog.
- For an entry both sources carry at different pins the NEWER catalog wins:
the checkout's last `plugin-catalog/` commit time vs the doc's
`generated_at`; when neither resolves (release install, no timestamp) the
entries' `version` labels break the tie, else live wins as before.
- A cache older than LIVE_CATALOG_MAX_STALE_SECONDS (24 h) stops supplying
pins (in-tree takes over) but its removals still count — a kill-list entry
never expires.
- The cache is written via temp file + rename: a concurrent reader (gateway,
TUI, a second CLI) can no longer see a half-written document, which read as
a fetch failure and started a 60 s failure window in that process.
Catalog audit findings F7 and F11 (low severity, no issue filed).
Builds on #72026 (@PRATHAMESH75): list content carries the turn's memory-prefetch /
pre_llm_call context as a durable text part appended once in the prologue, in every
api mode (MoA and codex_app_server included), so the request, the persisted row,
compaction and a later resume all see the message the model saw.
Persistence gap from the #72026 review: in-place preflight compaction (and a
close/early flush that races the prologue) writes the current user row BEFORE the
part exists and the crash persist identity-skips that dict, so a resumed session
replayed the turn without the context. The list branch now pushes the appended part
into that row via set_user_message_content under the same _row_id-under-lock
protocol as the string sidecar backfill, keeping the writer's shape (compaction: raw
parts; flush: text projection).
Titling moves before the injection step so a list turn's title is derived from the
user's ask, not the injected tail. Tests trimmed to one invariant per layer: hook
edit reaches the wire on a list turn and replays after reload; in-place compaction
+ reload keeps the part (red without the backfill); memory query flattens parts.
`~/.hermes/hooks/` auto-loads every valid `HOOK.yaml` + `handler.py` at
gateway startup with no `plugins.enabled` gate. That is the documented
contract since 3988c3c245 ("Implicit (dir trust)" in the comparison
table), but the plugins page's "disabled by default" promise read as if
it covered gateway hooks too (#37963). Maintainer ruling: keep implicit
dir trust, fix the docs.
- hooks.md: new "Trust model" section stating exactly what loads, when,
how, and that placing the files is the opt-in; comparison-table cell
links to it and the plugin-hooks consent cell now says
`plugins.enabled`.
- plugins.md: note scoping `plugins.enabled` away from gateway hooks.
- developer-guide/plugins: one sentence at the gateway-hook recipe.
- security.md: "Trusted-by-placement extension points" section
cross-linked from hooks.md.
skill_view loads SKILL.md whole and the content then rides in context for
every later call of the session, so body size is paid per turn. The only
size signal was the 100k hard cap in skill_manage, and agent-authored skills
grew by small patches until they sat right under it (31 of 432 local skills
over 40k chars, 7 at 100-115k; skill_view results averaging 32k chars).
- skill_linter: advisory `oversized-body` past _BODY_SOFT_BUDGET_CHARS (24k,
~3x the ~200-line standard; bundled skills average ~20k) naming the size,
the token estimate and the references/ split.
- skill_manage patch: attach lint findings the write INTRODUCED (diff of
rules before/after), so the crossing patch reports it once and a clean
patch on an already-large skill stays quiet. Create keeps reporting all.
- curator prompt: a body over the budget is itself a consolidation target.
- docs: skills.md linter paragraph.
A plugin's transform_llm_output replacement reached final_response only.
agent/turn_finalizer.py::finalize_turn fired the hook after the assistant
row had already been persisted — and the row is first written even earlier,
in agent/turn_final_response.py::finish_text_response's durable flush — so
messages[-1], the SQLite/JSON session, /resume and the next turn's replay all
kept the raw model text while the user had seen the rewritten one.
Writing the transformed text back after that flush cannot work: SQLite treats
a non-blank assistant row as settled (resolve_and_repair_transcript_batch
adopts the stored content instead of overwriting it), so the only correct seam
is BEFORE the row is first persisted. apply_llm_output_transform (new, in
turn_finalizer) fires the hook once per turn_id and records the outcome;
finish_text_response calls it ahead of append+flush and writes the result into
the row (api_content for the promoted-reasoning sidecar), finalize_turn's
_persist_step calls it ahead of the recovery-path tail close, and
_apply_output_hooks reads the recorded outcome (firing only when no earlier
seam saw a response) before post_llm_call. Only the current turn's not-yet-
written text changes — earlier turns and the system prompt are untouched.
post_llm_call is unchanged: it is an observer whose return is ignored by
contract, so there is nothing of it to persist (#14913/#44253's premise).
Fixes#44239
Slim redo of #44244 (AIalliAI, earliest; same sync-then-persist idea, moved to
the pre-flush seam) — also supersedes #65921 (SingleVirgin, sibling fix).
Co-authored-by: AIalliAI <285906080+AIalliAI@users.noreply.github.com>
(cherry picked from commit 5fe02a0aeecef422a4ffb4ff4385015f6a1528a0)
A plugin registered on transform_terminal_output only ever saw foreground
`terminal` output: tools/terminal_tool_result.py::_apply_output_transform_hook
runs from finalize_foreground_result and nowhere else. Background output
reached the model through a different seam — process_manage poll/wait/log/kill
results, `list` previews and the completion/heartbeat/watch notifications all
pass through tools/process_registry.py::_redact_process_result — which redacted
but never transformed, so a fleet redaction or summarising plugin silently did
nothing for backgrounded commands.
Apply the same hook helper at that shared seam (one new
transform_process_output wrapper) and at the two gateway agent-notify sites
that read session.output_buffer directly. The order matches the foreground
path and teknium1's review note on #71401: hook first, redaction after, so a
replacement the plugin returns is still masked. returncode is None while the
process runs; env_type is not recorded per process and is passed empty.
Not changed: the spawn acknowledgement ("Background process started") carries
no command output, and the persistent local shell already goes through
_run_foreground and was transformed — the issue's reading of that branch was
wrong; the real gap was the process_manage/notification seam.
Fixes#70760
Slim redo of #71401 (Christopher-Schulze): same seam and ordering, without
the ANSI-stripping relocation and render helper.
Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com>
(cherry picked from commit 84661de52f78ccb84054e1ded92b07166e40546d)
Follow-up trim of the #110265 salvage. `ainvoke_hook` logged raising callbacks
with a bare warning; route them through `_report_hook_failure` (warn-once per
distinct failure, #111922) and, for `_HOOK_TIMEOUT_FAIL_CLOSED_HOOKS`, append
the same named block directive the sync path emits (#109624), so the async twin
cannot drift into a fail-open policy path. Tests trimmed to the salvage bar: the
in-loop await is proven once through the real `_handle_message` path
(`test_async_hook_callback_is_awaited_on_the_gateway_loop`); the manager-level
duplicate is dropped and the narrowing test also pins failure isolation. Docs:
`pre_gateway_dispatch` callbacks may be `async def` and stay unbounded.
Credit order for the three PRs fixing this gap: #102485 (dmspark, earliest,
pre-decomposition `gateway/run.py`), #110253 (KoNit-K, bounded the hook —
rejected by design: neither fail mode is acceptable for a policy gate), #110265
(twidtwid, reporter; cherry-picked because it matches the ainvoke_hook shape,
keeps the hook unbounded, and adapts the existing sync test seams honestly).
Part of #110241
Supersedes #102485
Supersedes #110253
Co-authored-by: David Marcus <dmspark@users.noreply.github.com>
Co-authored-by: KoNit-K <konit.block@protonmail.com>
`agent/shell_hooks.py::_parse_pre_tool_call` translated only the block and
modify dialects, so a shell hook printing the documented
`{"action": "approve", ...}` parsed to None and the tool ran with no approval
prompt — silently, with exit 0, valid JSON and `hermes hooks doctor` green.
The Python-plugin side already accepts approve and routes it through
`_resolve_block_from_details` → `request_tool_approval`; the shell parser now
yields the same `{"action": "approve", "message"?, "rule_key"?}` shape (optional
fields kept only as non-empty stripped strings), so `hermes hooks test` prints
it under `parsed:` and the dispatcher escalates it. the `decision` dialect's
`{"decision": "approve"}` means auto-ALLOW, not "ask a human", so it is
deliberately not mapped; that dialect has no top-level ask dialect to mirror.
Slim redo with credit: #92562 (earliest) bundled a larger policy-authority
rework; #110325 carried the same parser change plus an unrelated rule_key
default change and 10+ tests.
Fixes#92553
Supersedes #92562
Supersedes #110325
Co-authored-by: fangliquanflq <fangliquan@qq.com>
`_get_pre_tool_call_directive_details` returned the first valid block-or-approve
in registration order, so a plugin registered earlier that returned `approve`
hid a later security plugin's `block`; under `approvals.mode: off` an approve
means no prompt at all, so the veto was dropped silently. Precedence is now
`block` > `approve` > none: a valid block still returns immediately (modify
directives seen before it stay attached, as before), a valid approve is held
back until the whole result list has been scanned for a veto, and among approves
the first valid one (with its rule_key) still wins. Modify accumulation is
unchanged and now also keeps modify directives that follow the winning approve,
since the scan no longer stops there. Docstring and hooks.md no longer describe
"first valid directive wins".
Slim redo of #68644 (earliest) and #87449 against the modify-aware shape of the
function on main; both PRs predate it and could not be cherry-picked.
Fixes#87420
Supersedes #68644
Supersedes #87449
Co-authored-by: synscott <1563043+synscott@users.noreply.github.com>
Co-authored-by: Jack Lau <72348727+jackulau@users.noreply.github.com>
`_run_hook_callback_bounded` treated any live abandoned worker for a callback
(`bool(self._hook_abandoned.get(suppression_key))`) as "still running", so one
never-returning `pre_tool_call` callback made every later tool call fail closed
with the timeout message until the process restarted. The timeout path
self-heals through the 60s suppression window; the abandoned path never did.
Policy now: while the suppression window is open the callback is skipped as
before. After it expires a fresh call id may start a new worker even though the
abandoned one is still alive — capped at `_HOOK_MAX_ABANDONED_WORKERS` (3) live
abandoned workers per callback so a hung plugin cannot leak a thread per call
(the #98382 constraint). At the cap the callback keeps being skipped (fail-closed
for pre_tool_call) with a WARNING naming the callback and its module, until one
of its workers finishes and frees a slot.
Test changes: `test_hung_worker_blocks_new_call_identity_after_suppression`
encoded the removed behaviour (exactly one worker, forever); it becomes
`test_hung_worker_caps_new_call_identities_after_suppression`, which pins the
same invariant it was protecting — bounded leak, never one per call — at the new
bound and checks the warning. `test_hung_worker_does_not_fail_closed_forever`
is the #105223 regression (red on base: call-c returned the block directive).
Redone slim against the per-call-id gate that landed in #111177; #105241
targeted the pre-#111177 shape and needed plugins_ledger/__init__ changes for a
one-retry-then-quarantine policy. Its analysis and shape informed this fix.
Fixes#105223
Supersedes #105241
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Catalog trust bugs from the 2026-09-21 plugin audit (lane 3, F1-F6, F9):
- F1 (high): a URL-installed repo shipping its own .hermes-catalog.json rendered
as catalog:official everywhere and marked the real entry "installed". Provenance
now lives on the installer-owned .install-metadata.json record (catalog block
written by _install_plugin_core, sha = checked-out commit); read_catalog_sidecar
never reads the tree. Pre-fix installs are adopted once when the installer
record agrees (pinned at the sidecar sha, cloned from the entry's repo).
- F2: removed.yaml bypassed by git@/ssh:///http:///www. spellings. _normalize_repo
canonicalises to host/owner/repo (scheme, user, www., .git, slashes dropped).
- F6: removed.yaml consulted for INSTALLED plugins too: git-pull update, enable and
gate_manifest (load) refuse recalled plugins, offline (in-tree + cached live
list). --allow-removed is recorded on the install record and exempts it.
- F5: install NAME --ref X recorded the catalog pin, so list/TUI/update claimed
the reviewed pin while HEAD differed. Recorded sha is the checked-out one.
- F3: re-pin replaced the whole tree, losing the installer-created config.yaml,
data and user patches with no warning. Untracked/ignored files are carried
into the new tree; edits to tracked files are copied to
<HERMES_HOME>/plugins-backup/<name>-<sha8>/ with a warning.
- F4: a manifest rename between pins left the OLD dir installed and enabled.
The stale dir is removed and the enabled flag follows the new name.
- F9: dashboard payload `removed` list now includes live removals.
repin_catalog_plugin returns RepinResult(sha, changed, installed_name, warnings);
CLI, dashboard and TUI callers surface the warnings and the new name.
The catalog page claimed admission's lint means a marketplace install 'cannot quietly rewire the app'; the lint is a handful of regexes and plugin.js runs in the app realm with the full window.hermesDesktop bridge. The user guide, catalog trust model and SDK security section now describe the real model: human review of a pinned SHA plus two tripwires (lint + loader import allowlist), no isolation.
Deleting an idle scratch entry left two things behind. Headless browsers
started by a lane's e2e run kept running for days with a "(deleted)" cwd
(about 20 on one host; the process registry never knew them because they
were grandchildren of a shell that had exited). Repos whose linked
worktree lived in the entry kept a dangling registration until someone
ran `git worktree prune` by hand (10 in one repo).
The prune now lives in hermes_constants_scratch (hermes_constants keeps the
entry point). Before an idle entry is removed, same-user processes whose
cwd is inside it, or inside any scratch path that no longer exists, are
TERMed then KILLed; the deleted-cwd sweep runs on every pass so orphans
from earlier deletions are caught too. `.git` files found in the entry
name their repo, which gets `git worktree prune` after the rmtree.
cache/terminal used its own 72h fixed-age sweep; it now shares the 24h
idle rule and the subtree check. `hermes doctor` warns about cache-root
directories over 1 GiB that neither pruner covers, since finished campaign
trees parked there sat for weeks (95 GB on one host). System-prompt
scratch line updated to match.
The scratch pruner deleted top-level entries whose own mtime was older than
72h. A directory's mtime only moves when a direct child is added or removed,
so a lane writing deep inside its tree looked untouched and could lose a live
worktree at the deadline, while finished trees (7 GB clones with their own
venv per campaign lane) sat for three days: 791 entries / 57 GB after four
days of campaigns on one host.
Retention is now idle-based: an entry stays while anything anywhere in its
subtree was written in the last 24h and goes a day after the last write. The
walk short-circuits at the first recent mtime, so live trees cost one stat
and only a truly idle tree pays for a full walk, once, right before deletion.
Symlinks are not followed so a link into the repo cannot keep an entry alive.
The config comment and both docs pages named six of the eight sources in
MACHINE_PACED_SOURCES; "tool" and "batch" were missing, so an operator
reading the docs could not predict the tier those sessions get.
The 1h Anthropic cache tier writes at 2x base (5m: 1.25x) and only pays off when
turns are more than five minutes apart. That is exactly the shape of an interactive
session a person parks and resumes, and exactly not the shape of a subagent, cron
run, one-shot or webhook that calls every few seconds and is gone. A single global
`cache_ttl` cannot be right for both, so operators leave it on 5m and pay a full
context re-write every time they come back to a CLI session after a coffee.
Measured on one install (2 days of per-call API logs, Claude via the Nous route):
63% of interactive cache-write tokens were cold re-writes after a 5-60 minute idle
gap; 1h would cut interactive write cost ~42% while costing ~49% more on subagents
and ~23% more on cron. `auto` resolves once per session from the session source
(`_session_source_for_agent`): 1h for cli/tui/desktop/messaging platforms, 5m for
subagent, cron, oneshot, webhook, kanban, api. Auxiliary/stub calls keep 5m; the
delegate_tool child clamp (#104168) still applies. Default stays "5m".