Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
ACP resolves toolsets like the gateway for the same platform config,
which includes plugin toolsets. An empty platform_toolsets.acp list still
adds enabled plugin toolsets, so the docs point at agent.disabled_toolsets
for removing them.
A fresh ACP agent appended mcp-<server> for every enabled config MCP server
unconditionally, so a platform_toolsets.acp allowlist of server names and the
no_mcp sentinel were ignored on ACP while the gateway honoured both.
The MCP half now comes from the same _get_platform_tools(config, "acp") call
as the base toolsets: its server names (default every enabled server, a listed
allowlist, or none for no_mcp) are keyed as mcp-<server>. Editor-provided
session/new servers are unchanged.
Docs: `hermes tools` has no ACP platform entry, so drop the claim that it
configures platform_toolsets.acp; document the MCP rules with a config example.
A fresh ACP agent hardcoded enabled_toolsets=["hermes-acp"], so
platform_toolsets.acp never narrowed the editor tool surface, unlike the
gateway, cron and api_server which all resolve via
hermes_cli.tools_config._get_platform_tools. Resolve the ACP base the same
way (ACP keeps appending its own mcp-<server> entries), and treat only None,
not an explicit empty list, as "use the hermes-acp default" in the /tools
and MCP-refresh rebuilds so a deny-all list cannot re-widen mid-session.
With the unconfigured default the resolved tool definitions are
byte-identical to the hermes-acp composite, so existing sessions keep the
same tool list and prompt cache.
The fresh-session assertion in test_make_agent_prefers_passed_toolsets_over_config_servers
now checks membership of the config MCP entry: the resolver returns the
expanded toolset keys rather than the bare composite name.
Refs #74582, #79516. Credit: #64045 (@israellot), #80309 (@thatssoheil),
#106834 (@nicolasramos) proposed the resolver routing.
Conflicts:
- scripts/releases/stamping.py, tests/scripts/test_version_stamping.py:
took ethie/pm-clean. The release branch's side was only its base's copy
of "stamping a payload snapshot skips the bootstrap-installer check"
(8411fdb333, same patch-id as 8d34601f47 here); the install-stamp
refactor c13ea774e6 supersedes the rest.
- tests/ci/test_stable_release_graph.py: kept pm-clean's release-epoch
contract (no HERMES_RELEASE_EPOCH on termux-deb, version on docker and
nix only) and the release branch's per-group receipt wiring.
Semantic conflict: the dispatch log step (6e64e961d8) read
inputs.termux_only, which the jobs input replaced. It reads JOBS now, and a
dispatch that selects only some groups has no release.py replay, as a
termux-only one had none before.
Review follow-ups on the crash-left reply adoption:
- A persisted reply is now judged the way live delivery would have judged
it. A bare silence marker ([SILENT] / SILENT / NO_REPLY ...) on an
internal turn, or the reply to a diagnostic wake whose chat policy mutes
diagnostics (read in the routed profile's scope, as the adapter does), is
owed nothing: the marker is cleared and nothing is sent or resumed. A
human turn's bare silence marker becomes the same "returned only a silence
marker" notice the live path sends. Before, the raw "NO_REPLY" reached the
user as a "Recovered reply".
- The active-turn marker start is written as aware UTC and compared as
epoch seconds (startup adoption and recover_interrupted_turns). A naive
local wall clock read by a process in another zone (DST, container vs
unit TZ) was hours off: a fresh in-flight turn was dropped as stale, or a
previous turn's reply could be adopted as this one's. updated_at stays
naive local for the older binary's recency heuristic; a pre-upgrade naive
marker still reads as local time.
- Reply timestamps go through coerce_epoch instead of float(), so one odd
transcript row cannot abort the whole recovery pass.
On an unclean start, suspend_recently_active(120) marked every session
touched in the last 120 s resume_pending/restart_interrupted, so startup
auto-resume ran a fresh model turn for chats whose turn had already
finished and been delivered: one kill re-answered 52 chats in the C12
delivery suite. The durable active-turn markers already name the exact
in-flight turns, so the recency sweep is removed.
That sweep also hid a real window: _handle_message cleared the turn
marker in its finally BEFORE the adapter recorded the delivery
obligation, so a kill in between left neither marker nor ledger row and
the persisted reply was never sent. The adapter now owns the marker for
turns it delivers and clears it right after record_delivery_obligation
(or once nothing more is owed). At unclean startup a marked turn whose
final reply is already in the transcript has that reply adopted into the
delivery ledger (unowned, 'attempting': sent once, marked as a possible
duplicate) instead of being regenerated; a marked turn with no reply
resumes once, as before.
The publication pass checks the held Store submission and prints the
Publish now step; it cannot release it. The developer guide and the draft
warning said the pass released it.
The stable doc now matches the shipped pipeline: rc.<N>-vX.Y.Z attempt refs
with a plain package version, derivation from the published head only, one
outstanding attempt of any version blocking release, publish-time final tags
with draft-state retarget and warning-block stripping, abandon marker refs,
attempt-ref archive and APT pool keys, per-arch receipts feeding each install
arm, the jobs input replacing termux_only, the stable mac feed written at
publication, and the Store submission held until publish. It states that the
tag ruleset is not applied and the locks are honor-system until an admin
applies it. No translated copy of this page exists under website/i18n/
(zh-Hans has no developer-guide/stable-releases.md), so nothing to translate.
After #120386 raised the read-only busy timeout to 5 s, retrying a lock
inside _open_read_only multiplied the wait to ~20 s on blocking callers
(TUI profile loop, exit epilogue, hermes status). The connection already
waited the read budget; only transient disk-I/O errors are retried now.
Probe (30 s exclusive DELETE-mode lock): 20.17 s -> 5.0 s, still
classified as a transient lock.
In rollback-journal (DELETE) mode a sibling process can take the write lock
between schema load and the messages_fts probe. FTS5's xConnect then fails its
%_config read and SQLite reports SQLITE_BUSY with the text "vtable constructor
failed: messages_fts". Every state.db lock classifier matched on the words
"locked"/"busy", so:
- a writable SessionDB() failed after 1s instead of waiting out the lock with
_WRITE_PATIENCE_S, and callers disabled persistence for the run;
- a read-only open (dashboard, `hermes sessions list`, cross-profile readers)
failed on the first busy timeout with no retry at all;
- the error read as not transient (dashboard 500, not 503) and as persistence
cause "unknown" instead of "locked".
Add hermes_state_errors.is_sqlite_lock_error: SQLITE_BUSY/SQLITE_LOCKED by
result code when SQLite supplies one, text only when it does not (our own
re-raised messages, RPC-wrapped strings). Route the writer open patience loop,
the _execute_write retry, the reconcile re-raise, the WAL->DELETE flip, the
maintenance holder probe, is_transient_sqlite_error and
classify_persistence_error through it. The read-only open retries a lock
inside its existing bounded retry budget, next to the transient IOERR case.
Superseding a stale micro marker joins the now-adjacent user turns into one
model-facing row, while the originals stay in display history as compacted
rows, so resumed display history painted every merged input twice (and one
more time per later pass). Flag the join display_metadata.model_only and skip
it in every display projection: resume dedupe, the indexed and legacy
get_messages pages, and the prompt timeline. The model payload is unchanged.
With the pooled local backend (cap 3, one slot reserved for foreground
dials), filing a bot into a section, saving its Configure editor, changing
its avatar or duplicating it ran profiles.configure/set_asset/create at the
background default. Once two warm bots held both background slots, the
cold bot's write queued for 30 s, timed out, entered the background retry
backoff, and was lost: the roster showed the change but the bot's
profile.yaml never got it.
requestForBot now dials those profile writes at foreground priority unless
the caller says otherwise, and the Configure editor's load (describe +
mcp.catalog) is marked foreground because the user just opened it. Polling
and roster warming keep the background default.
The reconnect watcher replaces a failed adapter with a NEW instance, and
every adapter keeps its inbound MessageDeduplicator on the instance. The
rebuilt adapter started with an empty cache, so a platform re-delivering a
recent inbound ID right after the reconnect (websocket resume replay,
webhook retry, unacked poll batch) got it processed and answered again.
The reconnect queue entry now holds the retired adapter's
MessageDeduplicator attributes by reference, and the rebuilt adapter
absorbs their live IDs before it connects. The multiplex secondary-profile
reconnect path gets the same handover. Any adapter using the shared helper
is covered without per-adapter code.
Hermes's 14-day `[tool.uv] exclude-newer` quarantine applies to Hermes's own
dependencies only (uv lock/sync, `hermes update`, LAZY_DEPS extras via
`ensure()`). A plugin's declared `python_dependencies` install under the
PLUGIN's policy: `install_specs(policy="plugin")` runs uv with `--no-config`
from any cwd, still inside the core constraints file.
Reverses item 3 of #118841, which ran the uv tier with cwd=<checkout> for
every install so the quarantine reached plugin deps from any cwd. That made
catalog re-pins floored on a <14-day release uninstallable (#120076:
"only hindsight-client<=0.9.2 is available"; #114530 held on the same gate).
Maintainer ruling (Teknium): "plugins dont have to abide by our 14 day rule
btw. They can have their own security policy on that. Only hermes'
dependencies themselves have to. We should recommend that they do this for
their plugins and we should give guidance to plugin devs that they should
though."
- tools/lazy_deps.py: INSTALL_POLICIES ("core" | "plugin"); `_uv_policy_args`
replaces `_uv_policy_cwd`; `_venv_pip_install(policy=)` defaults to core
(ensure/LAZY_DEPS), `install_specs(policy=)` defaults to plugin.
- hermes_cli/plugin_python_deps.py: `resolve()` passes policy="plugin".
- Docs: developer guide "Dependency security policy" section, catalog README
admission rule 9, AGENTS.md pinning policy — plugin authors are responsible
for their deps and strongly recommended to pin upper bounds, floor on the
oldest API-compatible version and run their own release quarantine
(`uv --exclude-newer` in their CI); operators can set UV_EXCLUDE_NEWER.
- Tests: the #118841 cwd test is replaced by two invariants — a plugin install
carries `--no-config` and no checkout cwd (red on base), a core lazy install
keeps the checkout cwd and no `--no-config`.
* refactor(fallback): share the pinned-owner chain rule
delegate_task's _resolve_child_fallback_chain decides which fallback chain
a child may walk: a pinned child never borrows the parent chain, an explicit
[] disables fallback, a declared list is the child's own. Cron needs the
same rule for pinned jobs (#100437), so the body moves to
hermes_cli.fallback_config.scoped_fallback_chain and the delegation helper
becomes a thin caller. Behaviour is unchanged; the delegation matrix test
still pins every cell.
* fix(cron): a pinned job never falls back to the global chain
A job with its own provider, model or base_url is an explicit operator pin
(since 0469740ab3 unpinned jobs store none of these). It still walked the
global fallback_providers chain in two places, so a pinned job could run
on a different provider and model than the one chosen:
- _resolve_job_runtime walked the chain on an AuthError or transient
network failure while resolving the pinned primary;
- _resolve_cron_agent_setup handed the global chain to every cron agent as
fallback_model, so the conversation loop's provider ladder could swap a
pinned job mid-run.
Both now read _job_fallback_chain(job, cfg), which returns no chain for a
pinned job through the same scoped_fallback_chain rule delegate_task uses
for pinned children. The pre-dispatch key check reads it too: the global
chain used to skip that check for every job, so a pinned job with a
missing key now blocks before the agent is built instead of failing in the
resolver. The transient-failure notice for a pinned job says it does not
fall back and names --unpin, instead of "No backup provider succeeded".
Unpinned jobs (including legacy *_snapshot records) and same-provider
credential-pool rotation are unchanged. The two scheduler tests that
asserted atomic provider+model fallback swaps used pinned jobs; they now
use unpinned jobs and keep the same assertions.
No per-job fallback_providers list: jobs have no generic override field
(create_job/update_job, the cronjob tool schema and the CLI enumerate each
field), so an opt-in chain would be a new surface on all of them. The
escape hatch is to leave the job unpinned and pick its model with
cron.model / cron.model_provider.
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
* docs(cron): pinned jobs do not use fallback_providers
cron.md "Provider recovery" and the pre-dispatch key check, the cron rows
and section in fallback-providers.md, and the developer notes in
cron-internals.md / provider-runtime.md said every cron job inherits the
global chain. State the new rule, the compatibility note for users who
relied on a pinned job landing on the chain, and the unpinned + cron.model
alternative.
---------
Co-authored-by: 686f6c61 <6115107+686f6c61@users.noreply.github.com>
Restore the cheap repo-wide guards the per-file triage classed as source reads
but that protect recurring bug classes (<2s total):
- subprocess env scrubbing near spawn sites (credential leakage)
- gateway UTF-8 encoding= on file I/O (Windows mojibake)
- no raw yaml.safe_load of config.yaml (lost ${ENV} expansion)
- CLI subprocess.run timeouts (hung CLI)
- no locked readers on the shared state.db connection (#99349 segfault)
- CI classifier outputs / live-comment watch list match real workflows
- relay imports no platform crypto (relay trust boundary)
- Desktop relay deliver budget mirrors the Python deadlines (#93911)
- no native title= on Desktop buttons (DESIGN.md rule)
Drop _BASELINE entries in check_os_marker_fakes.py for files that no longer
fake macOS (the checker fails on stale entries), and remove doc/comment
pointers to deleted tests.
OpenAI shipped gpt-6-sol / gpt-6-terra / gpt-6-luna as the successors of the
gpt-5.6 tier line (Sol and Luna live on OpenRouter + the Nous Portal today).
The curated aggregator catalogs (OPENROUTER_MODELS and the derived nous list,
plus the published website model-catalog.json) now carry the gpt-6 tiers and
their -pro variants instead of the 5.6 ones; the openai-api curated fallback
lists them ahead of 5.6.
Codex OAuth support mirrors the 5.6 + Astra contract for every gpt-6 tier:
curated fallback + forward-compat synthesis (from the 5.6 twin or 5.5),
272K advertised fallback, the opt-in -900k picker variants with the
live-verified 900K bump (still capped by the catalog's max_context_window),
dated-snapshot eligibility, wire-suffix stripping, the compaction auto-raise
on the base slug, and the gpt-5.6 effort ladder (max allowed, minimal
rejected). Pricing rows for gpt-6-sol / gpt-6-luna come from OpenAI's model
pages (272K whole-request tier like Astra); Terra has no published page yet
so it deliberately has none.
/model gpt keeps resolving to the flagship: "astra" joins the rank-0 suffix
set so gpt-6-astra sorts above gpt-6-sol.
A plugin that finished loading after an adapter connected never got its platform
handlers (slash commands, button callbacks, inbound transforms) registered until a
gateway restart, silently. Three pieces, one seam shared by every surface:
1. Discovery listener: PluginManager.on_plugin_loaded(cb) fires from INSIDE
discover_and_load for the plugins a sweep newly loaded (diff of the loaded set),
with a per-plugin activation summary (hermes_cli/plugins_activation.py):
activated_now {gateway_commands, gateway_transforms, hooks, callbacks} vs
deferred {tools, prompt, mcp_servers}. Every mid-run load path now performs a real
discover_plugins(force=True): CLI install/enable (via the gateway), Desktop/TUI
plugins.manage install/toggle/update, dashboard REST install, tool-triggered
force re-discovery, the new `reload-plugins` control-socket verb. A non-forced
discover_plugins() short-circuits on _discovered, which is why reload.mcp after
a mid-run install used to reload the OLD server set.
2. Idempotent re-wire: BasePlatformAdapter.rewire_plugin_handlers() runs only
factories not yet wired on the live native client (keyed (plugin, qualname);
a force reload hands back new function objects). Telegram hoists late handlers
ahead of core's catch-all filters.COMMAND / CallbackQueryHandler (PTB dispatches
the first match per group) and re-wires on the transient-init rebuild; Slack
dedupes register_slack_action_handler per AsyncApp. The gateway runner
subscribes per served profile and re-wires on the loop.
3. Scope limit + honest messaging: handlers only. Tools/prompt stay deferred to
the next session (prompt-cache invariant), MCP servers to mcp.reload; the CLI
hint and plugins.manage results (activation, gateway_reloaded,
restart_required only when no gateway answered) say exactly that.
* fix(plugins): portable MCP servers get a readable name so tool names fit the 64-char cap
A portable plugin's MCP server was named `<skill_namespace>__<server>`, i.e.
`agent-plugin-<slug>-<sha8>__<server>`. That prefix is right for plugin-data and skill names
(collision-free without coordination, and persisted on disk) but it costs ~40 chars of every
`mcp__<server>__<tool>` name. Providers cap function names at 64, so the registry clamped every
tool of the NVIDIA plugin to a hash-suffixed stub with the verb cut off:
`mcp__agent_plugin_hermes_nvidia_72eb26b1__nvidia_app__n_3fa9c2d1`.
Server names only need to be unique among loaded portable servers, and the loader already refuses
a clash. `portable_mcp_server_name(key, server)` = `<plugin-slug>__<server>`, collapsed to
`<plugin-slug>` when the two match (the one-server package). The loader and the desktop card's
server rows both call it (the card recomputed the f-string on its own before). Skill and plugin-data
namespaces are unchanged; nothing persisted refers to the server name, so no migration.
Tests: the existing portable-load test now pins the relationship (server name = plugin slug +
server; a long vendor tool name reaches the wire unclamped), and a new test covers the clash the
digest used to hide: two enabled packages folding to one slug, second server skipped, first served.
Both red on main. Docs: developer-guide/plugins/index.md.
* fix(plugins): a portable MCP server is named what its mcp.json calls it, nothing prepended
Drop the plugin segment too. A user's own config.yaml server named `nvidia-app` yields
`mcp__nvidia_app__<tool>`; a portable plugin's server of the same name now yields the same. Duplicates
are refused at load (config.yaml first, then first-loaded plugin) with a warning naming both owners.
Spike, isolated HERMES_HOME, real session: a portable plugin with server `acme-tools` and tool
`acme_client_get_driver_status_report` registers as
`mcp__acme_tools__acme_client_get_driver_status_report`; the model found it by tool_search, called it,
and reported that exact name. Same plugin on main: `mcp__agent_plugin_acme_tools_88e7456f__acme_tools__acme_153862de`.
Activation defined PATH and PYTHONPATH, but `hermes` still resolved to a
global command or an MSIX alias. The shell now defines `hermes` as a
function for this worktree. It runs this checkout's CLI only while the
shell is inside the tree, and it refuses outside it. The prompt gains a
branch prefix and drops it outside the tree. `deactivate` removes the
function and the prefix.
* feat(platform): resolver core with locate/inspect/probe tiers and ordered candidates
Every resource lookup needs one result shape and one cost contract. `locate` reads
metadata only, `inspect` may open files and call OS APIs in-process, `probe` is fresh
and the only tier that may spawn or connect. `Resolution.candidates` keeps probe order
so fan-out consumers can try every present binary.
Linear NS-921.
* feat(platform): AppResolver over AppDef with plist, PE, registry, and server.json sources
Desktop apps need presence, version, and liveness as separate observations. The runtime
file's bearer token is parsed, used for one request, and discarded inside the probe;
no public type carries it. Endpoints are accepted only when loopback with a numeric port.
* refactor(copilot): gh candidates through locate_command and the Homebrew table
First consumer of the resolver. The gh token probe still tries every present binary in
order; the allowlist loses its two copilot_auth rows.
* feat(platform): availability() over an application declaration
locate() + inspect() only, never probes; the fail-closed _version in
app.py treats a vendor's plist/PE/registry entry as untrusted input.
Salvaged from PR #118122; reads any object with requires_app,
min_version, app_for(os) — nothing here imports the MCP catalog.
* feat(platform): application declarations parsed into AppDef per OS
The parser slice of PR #118122's catalog manifest, re-homed as a
catalog-free module: whoever owns an MCP server declares the app it
fronts per OS and what it needs, and registers it here. Stdlib +
hermes_platform.resolver only. register/lookup/clear are the one seam
the MCP check_fn and the skill gate both read.
* feat(mcp): check_fn honours a registered application declaration
_make_check_fn ANDs the declared app's availability into the
connection-alive check; with nothing registered for the server the
behaviour is the pre-PR3 connection check. Provenance is explicit
registration, not endpoint matching. Returns a plain bool: the registry
caches bool(fn()).
* feat(skills): requires_apps gate through registered declarations
Offer-time filter beside environments:; names resolve through
hermes_platform.declaration, an unknown name hides the skill (fail
closed). The disk snapshot carries requires_apps and the fast path
re-evaluates it (snapshot version bumped to 3): app presence is a host
fact that changes without SKILL.md changing.
* docs: application declarations page
The plugin-facing schema reference: app: and requires: blocks,
availability() states, and the two gates that read the registry.
Registered under Extending > Plugins in the docs sidebar.
* test(platform): declaration parser, availability, gates
The PR3 app-block tests re-homed off the catalog: fixtures are dicts
passed to parse_declaration, the check_fn gate keys on explicit
registration (not endpoint matching), and the import-hygiene probe now
covers hermes_platform.declaration and resolver.availability.
A plugin manifest's `config_schema` now reaches the Desktop: `plugins.manage list`
returns each plugin's schema with the current `plugins.entries.<id>.settings`
values (`settings_schema`), and a new `settings` action writes edits through
`hermes_cli.plugins_state.save_plugin_setting` — the writer extracted from
`PluginContext.set_config`, so the plugin, the CLI and the Desktop share one
config path, one lock and the same managed-install / managed-key refusals.
The Plugins tab grows a gear per plugin with a schema; the inline form is
table-driven (`FIELD_CONTROLS` / `INITIAL_TEXT` / `COERCE` keyed on the wire
field type) for string / number / boolean / enum / json / secret. Secrets are
declared with `type: secret`: the row carries only the `.env` name and a
presence flag, the client writes the value through the existing `PUT /api/env`
credential route, and the RPC refuses secret keys so nothing lands in
config.yaml.
Contracts regenerated; docs gain a "Settings form in the Desktop" section.
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.
- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
post_api_request payload shape plus `aux_task`, `api_request_id`
(`aux-...`, shared by every attempt of one logical call), `retry_count`,
`streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
turn is in flight; fail-open (a raising/hung subscriber is logged and
the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
attempt shares (_relay_sync_completion / _relay_async_completion /
_relay_sync_stream) run under the hook pair — retries and fallbacks
included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
aux_task and no api_request events; raising subscriber never breaks
the call). First is red on origin/main.
Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.
Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
Four error-isolation holes in the disk plugin door
(apps/desktop/src/contrib/runtime-loader.ts), found by a static+live audit of
the loader; none had an issue filed.
- A plugin whose module evaluation never settles (top-level `await` on a dead
host) hung `import()` forever and, through the scan's sequential loop and
its re-entrancy guard, froze every later plugin and all future scans until
restart. `import()` now races a 10 s deadline; the plugin errors on its own
row ("import timed out") and the scan continues.
- Timers and DOM listeners a plugin took out with bare globals survived
disable and every hot-reload. `ctx.setTimeout` / `ctx.setInterval` /
`ctx.addEventListener` are tracked with the plugin and torn down on
unload; the SDK doc says bare globals are not.
- Two folders exporting one plugin id silently last-wins: the second
disposed the first's registrations and each hot-reload flipped ownership.
The first (folder-name sorted, so deterministic) owns the id; the later
file errors on its own row ("duplicate id, already loaded from <path>").
- A save that no longer loads (syntax error, timeout, duplicate) left the old
incarnation's contributions and activate handle live beside the error
row, so the Plugins tab showed a broken file as "loaded" and could
re-enable stale code. The previous incarnation is unloaded and dropped.
Tests: one invariant per fix in runtime-loader.test.ts, all red on base
(the hang case red by timing out).
`~/.hermes/hooks/` auto-loads every valid `HOOK.yaml` + `handler.py` at
gateway startup with no `plugins.enabled` gate. That is the documented
contract since 3988c3c245 ("Implicit (dir trust)" in the comparison
table), but the plugins page's "disabled by default" promise read as if
it covered gateway hooks too (#37963). Maintainer ruling: keep implicit
dir trust, fix the docs.
- hooks.md: new "Trust model" section stating exactly what loads, when,
how, and that placing the files is the opt-in; comparison-table cell
links to it and the plugin-hooks consent cell now says
`plugins.enabled`.
- plugins.md: note scoping `plugins.enabled` away from gateway hooks.
- developer-guide/plugins: one sentence at the gateway-hook recipe.
- security.md: "Trusted-by-placement extension points" section
cross-linked from hooks.md.
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.
Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.
Fixes#61839
credit: @giggling-ginger #61995
A general-plugin context engine is one shared instance; agent init copied it per agent
with copy.deepcopy() only. Engines that hold a SQLite connection or lock (hermes-lcm)
already expose clone_for_agent() for exactly this, but it was never called, so every
init logged "could not be safely copied … falling back to built-in compressor" and the
engine was unusable through the plugin system.
ContextEngine grows clone_for_agent() (default: deepcopy, the previous behaviour) and
_select_context_engine calls it; the failure message now names the hook to override.
Docs: context-engine-plugin.md documents the per-agent clone contract.
Test change (existing on main): tests/agent/test_context_engine.py::
test_agent_init_source_deepcopies_singleton_not_aliases was a source-reading pin on the
literal `copy.deepcopy(_candidate)` line, which this fix intentionally replaces. It is
superseded by tests/agent/test_plugin_context_engine_clone.py, which drives the real
_select_context_engine seam and asserts the invariant it guarded (child update_model()
never mutates the shared singleton) plus the new clone_for_agent() path.
Fixes#99640
credit: @stephenschoettler #62374
credit: @686f6c61 #99677
ToolRegistry.dispatch spread every injected keyword (task_id, session_id,
user_task, parent_agent, ...) into the handler, so a third-party plugin tool
written as `def handle(args)` raised TypeError on every call. The plugin
contract (plugins/AGENTS.md) says optional kwargs are signature-inspected, not
forwarded unconditionally; hooks already do this in plugins_dispatch. Handlers
taking **kwargs still receive the full payload. Docs updated: the narrow
signature is supported, **kwargs opts into the whole context.
Fixes#68318
credit: @vveerrgg #22146 (slim redo; #68636 is the later duplicate)