Commit Graph

2553 Commits

Author SHA1 Message Date
ethernet
439dd69a45 fix(memory): honor disabled providers without PyYAML
Parse memory-provider manifests through the declared hermes_yaml layer so manifest-name deny-list entries cannot fail open when PyYAML is absent. Exercise enabled and disabled profiles A→B→A to lock down config and module cache isolation.
2026-09-22 09:28:53 -04:00
ethernet
a49d196b5c Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/env_loader.py
#	hermes_cli/urllib_security.py
#	tools/terminal_scope.py
2026-09-22 07:05:11 -04:00
kshitijk4poor
b9ec39939b fix(discord): pre-seed the starter dedup before awaiting mark_async
_auto_create_thread's dedup pre-seed (is_duplicate(str(thread.id))) stops
the echo MESSAGE_CREATE Discord fires for the starter (id == thread.id)
from re-running the request. This PR turned the tracker persist into
`await self._threads.mark_async(thread_id)` and placed it BEFORE the
pre-seed; to_thread always suspends, so the echo's handler could reach
_discord_message_admission -> is_duplicate during the os.replace window,
claim the id first, and rerun the starter. On main both statements were
sync, so no window existed.

Move the pre-seed (with its comment) directly after
`thread_id = str(thread.id)`; the await now follows it.

Other mark_async call sites checked (discord :4717 slash create-thread,
discord :6140 pre-dispatch, matrix :1438 create_thread, matrix :2083
inbound): nothing after those awaits relies on state a concurrent event
could claim first, so no further reorder.

PROOF: tests/gateway/test_discord_double_dispatch.py::
test_thread_starter_duplicate_dropped now installs a recording mark_async
that asserts thread.id is already in _dedup._seen when it is awaited.
Red with the pre-swap order (1 failed), green after; ruff clean,
check-windows-footguns clean, real import of the adapter from the
worktree ok, scripts/run_tests.sh on the PR's test files green.
2026-09-22 15:54:29 +05:30
kshitijk4poor
7dd6bee33b fix(gateway): rich_sent_store writes run off the event loop
`record`/`record_media` do a read-modify-write of the JSON index ending in
`atomic_json_write`/`os.replace`, and every coroutine on the WhatsApp Cloud,
WhatsApp bridge and Telegram send/inbound paths called them inline on the
loop thread, stalling the whole gateway for a filesystem write nothing
awaits. Add `record_async`/`record_media_async` as thin
`asyncio.to_thread` wrappers (same precedent as
gateway/channel_directory.py) and await them from the seven coroutine call
sites; the sync functions stay for sync callers. The write completes before
the await returns, so lookup-after-send behaviour is unchanged.

Salvaged from #118883 by @Kyzcreig; re-shaped onto asyncio.to_thread (no
writer thread, no in-memory shadow, no queue), plus the missed
plugins/platforms/whatsapp/adapter.py record_media site.
2026-09-22 15:54:29 +05:30
Kyzcreig
ee8a3ded96 fix(gateway): move the thread-participation persist off the event loop
`ThreadParticipationTracker._save` (gateway/platforms/helpers.py) ends in
`atomic_json_write` -> `os.replace`, whose duration is unbounded under
filesystem pressure. All five call sites are coroutines on the inbound-message
or slash-command path:

    plugins/platforms/matrix/adapter.py   _resolve_message_context
                                          create_handoff_thread
    plugins/platforms/discord/adapter.py  _handle_message (x2)
                                          _handle_thread_create_slash

So that rename was paid inline on the running loop, stalling every other
adapter's polling, every in-flight turn and every heartbeat for as long as it
took.

The fix is at the CHOKE POINT rather than at five call sites:

  - `mark_async` does the in-memory insert synchronously and offloads only the
    persist via `asyncio.to_thread`. The insert must stay synchronous because
    both adapters gate on `thread_id in self._threads` immediately after
    marking; deferring it would make mention-gating depend on executor
    availability.
  - all five coroutine call sites now await it.
  - `mark` keeps its exact synchronous contract for the non-loop callers.
  - an RLock is added in the SAME commit that introduces the concurrency: the
    event loop used to serialize every caller by accident, and
    `atomic_json_write` makes each write atomic without making
    check/insert/trim/write atomic. Without it two concurrent marks lose one.

Enforcement is an AST class sweep, not an inventory: no `async def` under
gateway/ or plugins/ may call `<x>._threads.mark(...)`, and every
`mark_async(...)` must be awaited -- an un-awaited one never runs at all, so
the thread is neither persisted nor recorded in memory and mention-gating
re-prompts forever in a thread the bot already joined. A new adapter fails the
gate without anyone remembering a list.

tests/gateway/test_discord_thread_slash_expired_defer.py stubbed the tracker
with `SimpleNamespace(mark=...)`; it now uses the real tracker against a
tmp_path, so the test cannot rot silently the next time this surface moves, and
it additionally asserts the thread really was recorded.

Verified on this exact head, PYTHONPATH pinned to the worktree:
  tests/gateway/test_thread_tracker_mark_off_loop.py             7 passed
  expired-defer + admission-exemption + off-loop                10 passed
  -k "discord or matrix or thread" over tests/gateway  1136 passed, 13 failed
The 13 failures are INHERITED: a clean worktree at upstream/main 75e9567ca7
with none of these changes fails the identical 13.

Every guard is gate-proven -- reverting the RLock, the to_thread, one call
site, one `await`, the sync insert, or the dedupe short-circuit each fails its
own test and only its own.

(cherry picked from commit 6fc8037d292ec051fe76ad213903b2ee5534080a)
2026-09-22 15:54:29 +05:30
daedalus-opus
cb567242d7 fix(gateway): move the sticker-description cache write off the event loop
`gateway.sticker_cache._save_cache` ends in `atomic_json_write` -> `os.replace`,
whose duration is unbounded under filesystem pressure. Its only production
caller is Telegram's `_handle_sticker` -- an inbound-message coroutine -- so
every sticker that missed the cache paid that rename inline on the event loop,
stalling every other adapter and every in-flight turn in the process for its
duration.

Adds `cache_sticker_description_async`, which dispatches the existing sync form
via `asyncio.to_thread`; the Telegram call site awaits it. The sync form keeps
its exact contract and is what the wrapper dispatches to, so it is unchanged for
non-loop callers.

`cache_sticker_description` is a read-modify-write (`_load_cache` -> merge ->
`_save_cache`). `atomic_json_write` makes each WRITE atomic, not the TRIPLE.
While the call was inline the event loop serialized every caller and the race
could not be observed; moving the write to a worker thread introduces real
concurrency, so `_CACHE_LOCK` is added in this same commit rather than deferred.

Tests use no wall-clock thresholds. The liveness witness is ORDERING: the rename
is held open on a barrier released only by a background timer, and the sibling
task must have ticked BEFORE that release. Verified RED first:

- revert `asyncio.to_thread` to an inline call -> 2 failed
  ("the cache write ran on the event-loop thread")
- revert `_CACHE_LOCK` -> the concurrency test fails on the DURABLE FILE:
  "lost an entry ... keys = ['uid_a']"

`test_gate_proof_the_sync_form_does_block_the_loop` drives the identical barrier
through the sync form and asserts the loop DOES starve, so the liveness
assertion cannot pass vacuously.

GREEN: sticker off-loop + existing sticker cache tests -> 10 passed;
`tests/gateway -k "sticker or telegram"` -> 937 passed, 4 skipped; Ruff clean.

(cherry picked from commit 9752a9b1aba7536181d081cad9c1c32917a89680)
2026-09-22 15:54:29 +05:30
ethernet
c13287c915 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/electron/main.ts
#	hermes_cli/backup.py
#	hermes_cli/config.py
#	hermes_cli/plugin_catalog.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/plugins_cmd_catalog.py
#	hermes_cli/plugins_discovery.py
#	hermes_cli/profiles.py
#	hermes_cli/update_cmd_deps.py
#	pyproject.toml
#	tests/gateway/test_dm_topics.py
#	tests/hermes_cli/test_config.py
#	tests/hermes_cli/test_plugins_cmd.py
#	tests/hermes_cli/test_update_autostash.py
#	tests/tools/test_lazy_deps.py
#	tools/lazy_deps.py
#	tools/skill_ledger.py
#	utils.py
#	website/docs/user-guide/security.md
2026-09-22 05:16:50 -04:00
teknium1
0e5809566f feat(plugins): fire pre/post_auxiliary_call events on every auxiliary LLM call (#79733)
Auxiliary LLM calls (titling, compression, MoA advisors/aggregator, vision,
approval, ...) never reached any plugin hook: hook-based observability and
cost plugins were structurally blind to them. Teknium's ruling on #79733:
NEW events rather than reusing the turn-scoped pre/post_api_request pair,
so existing subscribers keep their per-turn semantics.

- agent/auxiliary_hooks.py (new sibling): builds the pre_api_request /
  post_api_request payload shape plus `aux_task`, `api_request_id`
  (`aux-...`, shared by every attempt of one logical call), `retry_count`,
  `streaming`, parent-turn `session_id`/`task_id`/`turn_id` when a main
  turn is in flight; fail-open (a raising/hung subscriber is logged and
  the aux task proceeds); post carries `error`/`error_type` on failure.
- agent/auxiliary_client.py: the three relay funnels every physical
  attempt shares (_relay_sync_completion / _relay_async_completion /
  _relay_sync_stream) run under the hook pair — retries and fallbacks
  included. Main-loop *_api_request events do not fire for aux calls.
- Catalogue: VALID_HOOKS, bounded-timeout hook set, `hermes hooks test`
  sample payloads, hooks.md / plugins index / observer-hooks / plugins.md
  tables, agent + plugins AGENTS.md.
- tests/agent/test_auxiliary_hooks.py: 2 invariants (pair fires with
  aux_task and no api_request events; raising subscriber never breaks
  the call). First is red on origin/main.

Supersedes #32416 (@zrmnelson), #68060 (@JonZal), #77518 (@hsy5571615),
#79826 (@webtecnica) — their relay-boundary placement, usage
normalisation and fail-open policy shaped this implementation.

Co-authored-by: zrmnelson <zacharynelson1@gmail.com>
Co-authored-by: Jonas Zalys <jonas@tryholo.ai>
Co-authored-by: saitsuki <nukuom976228@gmail.com>
Co-authored-by: webtecnica <webtecnica@gmail.com>
2026-09-22 01:19:12 -07:00
kshitijk4poor
c2d7bb89f1 refactor(discord): name the terminal liveness reasons in one constant
Declare `_TERMINAL_HEALTH_REASONS = frozenset({"socket_closed",
"client_closed"})` beside `_read_websocket_health` (the producer of those
literals at adapter.py:1705/:1707/:1715) and use it at the escalation
site in `_liveness_loop` (was `reason in ("socket_closed",
"client_closed")` at :1791).

WHY: the first-strike classification was a magic-string match against
literals produced 80 lines away with no shared definition. A rename in
the producer, or a new hard-closed reason, would silently demote the
escalation back to the threshold path with nothing failing. One named
set next to the producer makes the coupling visible. The
`tuple[bool, str]` return shape is unchanged (existing tests monkeypatch
the sampler with 2-tuple lambdas).

Proof: mutating the set to {"socket_closed"} fails
test_closed_transport_first_strike_forces_reconnect[client_closed];
restored, the liveness file is green.
2026-09-22 13:39:00 +05:30
kshitijk4poor
ac39011c2e refactor(discord): log the liveness-probe exit at INFO unconditionally
Collapse the INFO/WARNING level switch on the probe-exit log
(plugins/platforms/discord/adapter.py:1755-1766) to one unconditional
logger.info and delete `expected_exit`. Drop
test_unexpected_probe_exit_logs_warning, which was the only caller that
could reach the WARNING half.

WHY: the WARNING branch is unreachable in production. It needs
`_running=True and not _disconnecting and _client is None` at the guard,
and `_client` has exactly two production setters to None:

- adapter.py:1269 inside connect(): synchronously followed by
  `self._client = commands.Bot(...)` at :1271 with no await between, so
  the probe coroutine can never observe the None.
- adapter.py:1941 inside disconnect(): runs after `_disconnecting = True`
  (:1915) and after `await self._cancel_liveness_task()` (:1917), so the
  probe is already cancelled (exits via CancelledError at :1752) and the
  flag would read as an expected exit anyway.

No other `_client = None` in plugins/platforms/discord/, gateway/platforms/
base.py or gateway/run.py. `_running=False` and `_disconnecting=True` come
only from teardown paths, all "expected". The deleted test pinned this by
poking `adapter._client = None` directly — a state the code never produces.

This drops the level switch borrowed from #118504 but keeps its intent:
the probe exit is always logged with its full state (#118487, the probe
must never disappear silently).
2026-09-22 13:39:00 +05:30
kshitijk4poor
533796a833 refactor(discord): trim the first-strike comment to its WHY
The six-line comment above the terminal-reason check
(plugins/platforms/discord/adapter.py:1789-1794) restated the resume-swap
mechanism already documented in the test docstring and the issue. Keep the
one sentence that carries the intent (closed transport = confirmed death,
soft signals keep the threshold) and the #118487 pointer. No code change.
2026-09-22 13:39:00 +05:30
kshitijk4poor
fe225b20b1 fix(discord): treat client_closed as terminal like socket_closed
_read_websocket_health reports client_closed when Bot.is_closed() is
true. That is the same transport-dead state as socket_closed, yet it
stayed on the two-strike confirmation path and could show the same
1/2-then-silent-reset pattern if the bot task's done callback were ever
suppressed. Escalate both terminal reasons on the first strike.
2026-09-22 13:39:00 +05:30
kshitijk4poor
43499c4531 fix(discord): log an unexpected liveness-probe exit at WARNING (from #118504)
The guard exit in _liveness_loop fires for two very different reasons:
ordinary teardown (adapter stopped or disconnecting), and a still-running
adapter whose client vanished. The first is noise at INFO; the second
means the gateway keeps running with no watchdog (#118487) and must stand
out in an incident log. Pick the level from the exit cause instead of
logging both at INFO.

Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-22 13:39:00 +05:30
liuhao1024
b1a69d8178 fix(discord): escalate the first socket_closed liveness strike
A closed Gateway transport is a confirmed death, not a suspicion, but the
liveness probe treated socket_closed like every soft signal and waited for
a confirming strike (#118487). That strike never arrives: discord.py swaps
in a fresh socket while resuming, the next transport-side sample reads
healthy, and the strike counter silently resets while a resumed-but-deaf
session stays event-starved until the multi-hour event-silence default
elapses. The bot sits deaf until a manual gateway restart.

Escalate the first socket_closed strike straight to the forced-reconnect
path; soft signals (ack staleness, latency, event silence) keep the
confirmation threshold. Also leave a trace on every previously-silent
probe transition: counter resets now log, and probe exits log the flag
state, so an incident log can no longer confuse a healthy sample with a
dead watchdog task.

(cherry picked from commit a9b796e9e9b113595c049701caea46679430465b)
2026-09-22 13:39:00 +05:30
teknium1
0706dffca1 fix(config): route every config.yaml writer through one comment-preserving writer
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.

Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.

- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
  example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
  untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
  are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
  profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
2026-09-22 01:07:59 -07:00
teknium1
643dea487b fix(plugins): canonical activation keys + remove/toggle bookkeeping across CLI, dashboard and RPC
Symptoms fixed (all live-reproduced on origin/main in a fake HERMES_HOME):
- `hermes plugins disable photon-platform` wrote `platforms/photon` while the loader keyed the
  bundled adapter `photon-platform`, so the disable never applied (#27548). The loader now keys
  bundled platforms `platforms/<dir>` like every other category (one call site in
  plugins_discovery.py); the manifest name stays an accepted alias through gate_manifest.
- `plugins.manage toggle` / dashboard toggle wrote the raw identifier: enabling by bare leaf or
  manifest name returned ok while a stale canonical key in plugins.disabled kept the plugin off.
  dashboard_set_agent_plugin_enabled resolves the canonical key and purges every alias from the
  opposing list (shared _activate_key, also used by cmd_enable/cmd_disable); the RPC reports the key.
- remove/uninstall (CLI, dashboard, RPC) left the plugin in plugins.enabled/disabled and
  plugins.entries (#54336), left memory.provider dangling so the next agent init re-cloned the
  provider from the catalog (uninstall silently reverted), and, for a symlink inside the plugins
  dir, deleted the TARGET plugin and its metadata while the link stayed dangling. _remove_user_plugin
  is the shared tail: unlink a link only, then forget config under every alias and reset
  memory.provider (reported as cleared_memory_provider).
- `hermes plugins list` / dashboard hub / TUI hub called bundled backends, bundled platforms and the
  live memory.provider "not enabled" (#73131, #82898): _plugin_status mirrors gate_manifest.
- A user-installed memory provider parked in plugins.disabled kept loading: load_memory_provider
  honours the deny-list (name, dir name or manifest name) and says so once.
2026-09-22 00:13:50 -07:00
ethernet
d0d4e91434 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/plugins_cmd.py
#	hermes_cli/plugins_cmd_catalog.py
#	tests/hermes_cli/test_web_plugins_catalog.py
#	website/docs/reference/cli-commands.md
#	website/docs/user-guide/features/plugin-catalog.md
2026-09-22 02:59:10 -04:00
teknium1
913c045c53 fix: resolve user-installed context engines from $HERMES_HOME/plugins/<name>
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.

Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.

Fixes #61839
credit: @giggling-ginger #61995
2026-09-21 23:38:16 -07:00
fangliquanflq
72ee40fa68 fix(plugins): remove memory lifecycle hook declarations 2026-09-21 23:26:19 -07:00
fangliquanflq
8ccbe5eb6a fix(plugins): declare bundled plugin hooks 2026-09-21 23:26:19 -07:00
tachyon-r
82d6ab99c4 fix: declare security-guidance hooks in manifest 2026-09-21 23:26:19 -07:00
KoNit-K
de6374194c fix(plugins): clean failed sibling modules 2026-09-21 23:26:19 -07:00
teknium1
77f80d31ee fix(plugins): catalog provenance is installer-owned; kill list covers update/enable/load; re-pin keeps user files
Catalog trust bugs from the 2026-09-21 plugin audit (lane 3, F1-F6, F9):

- F1 (high): a URL-installed repo shipping its own .hermes-catalog.json rendered
  as catalog:official everywhere and marked the real entry "installed". Provenance
  now lives on the installer-owned .install-metadata.json record (catalog block
  written by _install_plugin_core, sha = checked-out commit); read_catalog_sidecar
  never reads the tree. Pre-fix installs are adopted once when the installer
  record agrees (pinned at the sidecar sha, cloned from the entry's repo).
- F2: removed.yaml bypassed by git@/ssh:///http:///www. spellings. _normalize_repo
  canonicalises to host/owner/repo (scheme, user, www., .git, slashes dropped).
- F6: removed.yaml consulted for INSTALLED plugins too: git-pull update, enable and
  gate_manifest (load) refuse recalled plugins, offline (in-tree + cached live
  list). --allow-removed is recorded on the install record and exempts it.
- F5: install NAME --ref X recorded the catalog pin, so list/TUI/update claimed
  the reviewed pin while HEAD differed. Recorded sha is the checked-out one.
- F3: re-pin replaced the whole tree, losing the installer-created config.yaml,
  data and user patches with no warning. Untracked/ignored files are carried
  into the new tree; edits to tracked files are copied to
  <HERMES_HOME>/plugins-backup/<name>-<sha8>/ with a warning.
- F4: a manifest rename between pins left the OLD dir installed and enabled.
  The stale dir is removed and the enabled flag follows the new name.
- F9: dashboard payload `removed` list now includes live removals.

repin_catalog_plugin returns RepinResult(sha, changed, installed_name, warnings);
CLI, dashboard and TUI callers surface the warnings and the new name.
2026-09-21 23:17:05 -07:00
ethernet
727daebce3 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-22 01:32:50 -04:00
Austin Pickett
992f569fc1 fix(telegram): admit replayed updates before PTB dispatch
Keep one adapter-owned update claim across core, observer and native plugin
handlers. Release failed preparation, but retain queued or externally handed-off
work through cancellation and PTB task completion.

Refs https://github.com/NousResearch/hermes-agent/issues/68502

Co-authored-by: Jasmine Naderi <jasmine@smfworks.com>
2026-09-22 00:55:46 -04:00
ethernet
1afa348e1c Merge branch 'rev/pm-core' into ethie/pm-clean 2026-09-21 19:52:58 -04:00
ethernet
029cabb466 feat(pm): hermes pm install --extra NAME; one shared hint for missing extras
Twenty-five call sites told users to run
`python -c "from pm import sync_venv; sync_venv(['x'], explicit=True)"`
because `hermes pm install` only took package names. Add `--extra`
(repeatable; syncs the venv with the named extras and nothing else) and
pm.install_hint(extra), the single builder every site now uses, so the
advice stays correct when the command changes.

A cold PM runtime under allow_lazy_installs:false now reports the extra
the caller wanted and the command that provisions both, instead of a
bare "pm-runtime: not installed".
2026-09-21 19:08:19 -04:00
ethernet
b1c1aa9a28 fix(platforms): raise, don't assert, when an ensured node/npm has no binary
assert is stripped under python -O, so a pm.ensure that recorded no
selected binary would have fallen through to str(None) as the npm/node
path. The four sites now raise pm.InstallError, which the surrounding
handlers already report.
2026-09-21 19:03:21 -04:00
ethernet
94a47320f1 fix(google_chat): request both extras before surfacing the restart
ensure_import raises InstallError("restart Hermes to activate…") after a
SUCCESSFUL install, so the loop aborted on 'google' and never asked for
'google-chat'; the restart landed back on a still-missing extra. Both
are requested in one pass; the first failure is what the registry logs.
2026-09-21 19:02:23 -04:00
ethernet
11bd9b512b fix(photon): declare PHOTON_NODE_BIN again
The adapter still honours PHOTON_NODE_BIN ahead of PM's pinned node, but
the pm hand-merge dropped it from optional_env, so the override became
undocumented and unreachable from setup.
2026-09-21 19:00:14 -04:00
ethernet
a9845b8a04 fix(hindsight): serve the routed profile's env to the daemon; probe once per interpreter
_daemon_subprocess_env copied dict(os.environ), which under multiplex is
the LAUNCH profile's: profile X's daemon started with the default
profile's HERMES_HOME and credentials. It now builds the child env with
served_profile_child_env(inherit_credentials=False) — the manager reads
the LLM keys from the profile's own 0600 env file, so the child needs no
credentials from us.

check_local_runtime spawned 'python -c import ... sentence_transformers'
(a torch cold start) from is_available(), unavailable_reason() and
initialize() on every session. The verdict is now kept per interpreter
path for the process; a reinstall publishes a new PM generation, so a
new path re-probes.
2026-09-21 18:57:44 -04:00
ethernet
8e80f37aad fix(meet): make inherited interactive stdin explicit 2026-09-21 14:36:47 -04:00
ethernet
c69ccc3bf1 merge: refresh upstream models, notification expiry, and desktop controls 2026-09-21 14:01:47 -04:00
teknium1
884b8d980f fix(image_gen): a gateway 429 is reported as a rate limit and retried once, not as a missing model
`_submit_fal_request` (image) and `_submit_fal_video_request` (video plugin)
translated every managed-gateway 4xx into "This model may not yet be enabled
on the Nous Portal's FAL proxy — set FAL_KEY or pick a different model". For
HTTP 429 that remediation is wrong: the gateway body is RATE_LIMIT_EXCEEDED
with a retryAfter, the model is enabled, and agents reading the message
switched models or gave up. On one real install this fired 260 times in a
week (17% of image_generate calls).

Both surfaces now submit through one shared helper
(tools/fal_common.py::submit_managed_fal_with_rate_limit_retry): a 429 whose
Retry-After (header, else body error.retryAfter) fits a 30s cap is waited out
in interrupt-aware 0.5s slices and resubmitted once under a fresh
x-idempotency-key; a second 429, or an unknown/too-long Retry-After, raises
a ValueError that names the rate limit and tells the agent to retry later
rather than switch models. 429 can no longer reach the "may not yet be
enabled" text.
2026-09-21 10:40:03 -07:00
ethernet
9d036f8ac9 merge origin/main (15 commits) into ethie/pm-clean
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
2026-09-21 08:16:57 -04:00
kshitijk4poor
7b660e66ee fix(homeassistant): detect the supervised launch through is_supervised_gateway_launch
gateway.restart already answers "was this gateway launched by a generated
service" (HERMES_SUPERVISED_CHILD or launchd's XPC_SERVICE_NAME, so a plist
that predates the marker still counts); the hint reuses it instead of a
second env read. The negative branch (plain unreachable host without a
supervisor) is now an unmarked test so the Linux lane keeps covering the
function; the macOS-only test holds the positive branch.
2026-09-21 17:16:23 +05:30
kshitijk4poor
eda2432cce fix(gateway): scope the Local Network hint and keep the osascript wrapper out of gateway process scans
Review of the first cut:

- The Home Assistant errno hint matched a bare 65 and two error strings on
  every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
  errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
  genuinely unreachable HA host was told macOS was blocking launchd. Gate on
  darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
  generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
  `gateway run`, so `hermes gateway stop` would have signalled osascript
  alongside the gateway. The canonical matcher now bails on an exact
  argv[0] basename of osascript (the gateway is its child and is matched on
  its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
  carries (osascript's own output) instead of implying it is redundant.
2026-09-21 17:16:23 +05:30
xxxigm
1d871110e1 fix(homeassistant): name the macOS Local Network block behind errno 65 under launchd (#71206)
A launchd-run gateway that macOS Local Network Privacy denies sees every LAN
connect fail with EHOSTUNREACH while the same URL works from Terminal.
Annotate the Home Assistant connect/reconnect log lines with the cause and
the remedy so the failure is actionable instead of a bare "No route to host".

Salvaged from #115196 (remedy text points at the regenerated launchd job).
2026-09-21 17:16:23 +05:30
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
Bartok9
6504b665ba fix(kanban): refuse empty complete_task evidence (#117483)
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
2026-09-20 19:03:47 -07:00
chelsealong
7efacc8f62 fix(a2a): recover the streamed reply text instead of resolving empty
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).

_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".

(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
2026-09-20 18:45:42 -07:00
Forkbert
b113ab6de6 fix(photon): use GUID for liveness probe message id
(cherry picked from commit cdb6ccadb045e599db795823a100ae7b61d377ec)
2026-09-20 18:23:28 -07:00
teknium1
d325e53116 fix(kanban): drop the dead edit_completed_task_result shim and route dashboard priority edits through edit_task
Review follow-up for #117434: edit_completed_task_result had no callers
after edit_task absorbed it; the dashboard's _set_priority kept its own raw
UPDATE + reprioritized INSERT, so edit_task gains a board= passthrough for
the post-commit observer and becomes the single reprioritize primitive.
2026-09-20 16:33:49 -07:00
teknium1
53815e24dc fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.

The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.

Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
  before  req_reasoning: {}
  after   req_reasoning: {'reasoning_effort': 'medium'}
  agent.reasoning_effort: low  ->  {'reasoning_effort': 'low'}  (unchanged)
2026-09-20 16:03:38 -07:00
teknium1
fd94fe9b93 fix(gateway): every adapter send_voice accepts the dispatch's is_voice kwarg
The base media dispatch passes is_voice= to send_voice; line, mattermost and
weixin had explicit signatures without it, so a non-image MEDIA attachment
routed as audio raised TypeError and was dropped — the same class as the
Matrix report (#102221, #116776). Adds a repo-wide signature invariant test.
2026-09-20 15:48:51 -07:00
xiaodu
b6505448f1 fix(matrix): accept the media dispatch's is_voice kwarg in send_voice
The base media dispatch calls send_voice(..., is_voice=is_voice) for every
audio MEDIA attachment (gateway/platforms/base.py _send_one). MatrixAdapter
.send_voice() accepted neither is_voice nor **kwargs, so every non-image
MEDIA delivery raised TypeError and the file was silently dropped — the
failure is visible in rotated logs since 2026-09-08 (never worked).

Accept the flag explicitly: is_voice=False -> plain m.audio in the original
format (no transcode); True or omitted (play_audio legacy callers) -> the
existing MSC3245 voice-bubble path with best-effort Ogg/Opus transcode.

Fixes #116776

(cherry picked from commit d4f89a725498e29a2ee0fe016f8bb98db08f358c)
2026-09-20 15:48:51 -07:00
fangliquan
45a701d8b0 fix(simplex): flatten alpha image thumbnails
(cherry picked from commit 97886801741f4367aadfe83ccb25b10cedb21149)
2026-09-20 15:48:01 -07:00
liuhao1024
7a23b00101 fix(whatsapp): refuse legacy-pidfile kills without a start-time fingerprint
The stale-bridge cleanup accepted a "node" + session-path cmdline substring
as kill evidence for legacy pidfiles (pid line only). A log tail, editor, or
grep that merely mentions the session path matches that same substring, so
the cleanup could SIGTERM a stranger process.

Require the kernel start-time fingerprint and fail closed when the pidfile
lacks one; the bridge-port scan (which verifies a node-executable listener)
reaps the orphan instead. The refusal reason in the warning now distinguishes
a fingerprint-less legacy pidfile from a recycled PID.

Flip the legacy-pidfile regression test to assert the fail-closed outcome.

Fixes #116883
2026-09-20 15:24:51 -07:00
teknium1
1e4cd9ade6 fix(kanban): diagnostic severity colours follow the dashboard theme
The three `--hermes-diag-*` tokens were literals declared on the consuming
elements, so the theme engine's `<html>`-level custom properties could never
reach them and light presets rendered the warning badge at 1.8:1 contrast.
Chain them through the host tokens themes already set (`--color-warning`,
`--color-destructive`) with the shipped literals as fallbacks; same selector
list, so nodes rendered outside `.hermes-kanban` keep a value. Error and
critical share `--color-destructive` (critical keeps its bold weight).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 12:54:48 -07:00
teknium1
34b475e2ce fix(langfuse): warn once per invalid HERMES_LANGFUSE_MAX_DEPTH value, not per payload
#116713 parsed HERMES_LANGFUSE_MAX_DEPTH inside `_safe_value`, i.e. once per
captured prompt, response, tool input and tool output. With an invalid value
(`abc`) every captured field logged the same "Invalid ... Falling back to 4"
WARNING for the life of the process — one multi-tool turn fills agent.log.

Resolve the depth in `_resolve_max_depth`, an lru_cache keyed on the raw env
string: the warning fires once per distinct bad value, a changed env var is
still picked up by a long-lived process (mirrors `_capture_mode`, which also
reads per call and warns once), and valid values skip the int() parse after
the first call.

Follow-up to #116713 (independent review finding).
2026-09-20 12:46:00 -07:00