Commit Graph

1261 Commits

Author SHA1 Message Date
ethernet
2761a55f1b fix(plugins): keep late activation dependency-passive 2026-09-22 13:09:30 -04:00
ethernet
1f48a3d036 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-22 13:03:25 -04:00
teknium1
21d0b12958 feat(plugins): late-loaded plugins wire their platform handlers live (#87770)
A plugin that finished loading after an adapter connected never got its platform
handlers (slash commands, button callbacks, inbound transforms) registered until a
gateway restart, silently. Three pieces, one seam shared by every surface:

1. Discovery listener: PluginManager.on_plugin_loaded(cb) fires from INSIDE
   discover_and_load for the plugins a sweep newly loaded (diff of the loaded set),
   with a per-plugin activation summary (hermes_cli/plugins_activation.py):
   activated_now {gateway_commands, gateway_transforms, hooks, callbacks} vs
   deferred {tools, prompt, mcp_servers}. Every mid-run load path now performs a real
   discover_plugins(force=True): CLI install/enable (via the gateway), Desktop/TUI
   plugins.manage install/toggle/update, dashboard REST install, tool-triggered
   force re-discovery, the new `reload-plugins` control-socket verb. A non-forced
   discover_plugins() short-circuits on _discovered, which is why reload.mcp after
   a mid-run install used to reload the OLD server set.
2. Idempotent re-wire: BasePlatformAdapter.rewire_plugin_handlers() runs only
   factories not yet wired on the live native client (keyed (plugin, qualname);
   a force reload hands back new function objects). Telegram hoists late handlers
   ahead of core's catch-all filters.COMMAND / CallbackQueryHandler (PTB dispatches
   the first match per group) and re-wires on the transient-init rebuild; Slack
   dedupes register_slack_action_handler per AsyncApp. The gateway runner
   subscribes per served profile and re-wires on the loop.
3. Scope limit + honest messaging: handlers only. Tools/prompt stay deferred to
   the next session (prompt-cache invariant), MCP servers to mcp.reload; the CLI
   hint and plugins.manage results (activation, gateway_reloaded,
   restart_required only when no gateway answered) say exactly that.
2026-09-22 09:50:22 -07:00
ethernet
a49d196b5c Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/env_loader.py
#	hermes_cli/urllib_security.py
#	tools/terminal_scope.py
2026-09-22 07:05:11 -04:00
kshitijk4poor
b9ec39939b fix(discord): pre-seed the starter dedup before awaiting mark_async
_auto_create_thread's dedup pre-seed (is_duplicate(str(thread.id))) stops
the echo MESSAGE_CREATE Discord fires for the starter (id == thread.id)
from re-running the request. This PR turned the tracker persist into
`await self._threads.mark_async(thread_id)` and placed it BEFORE the
pre-seed; to_thread always suspends, so the echo's handler could reach
_discord_message_admission -> is_duplicate during the os.replace window,
claim the id first, and rerun the starter. On main both statements were
sync, so no window existed.

Move the pre-seed (with its comment) directly after
`thread_id = str(thread.id)`; the await now follows it.

Other mark_async call sites checked (discord :4717 slash create-thread,
discord :6140 pre-dispatch, matrix :1438 create_thread, matrix :2083
inbound): nothing after those awaits relies on state a concurrent event
could claim first, so no further reorder.

PROOF: tests/gateway/test_discord_double_dispatch.py::
test_thread_starter_duplicate_dropped now installs a recording mark_async
that asserts thread.id is already in _dedup._seen when it is awaited.
Red with the pre-swap order (1 failed), green after; ruff clean,
check-windows-footguns clean, real import of the adapter from the
worktree ok, scripts/run_tests.sh on the PR's test files green.
2026-09-22 15:54:29 +05:30
kshitijk4poor
7dd6bee33b fix(gateway): rich_sent_store writes run off the event loop
`record`/`record_media` do a read-modify-write of the JSON index ending in
`atomic_json_write`/`os.replace`, and every coroutine on the WhatsApp Cloud,
WhatsApp bridge and Telegram send/inbound paths called them inline on the
loop thread, stalling the whole gateway for a filesystem write nothing
awaits. Add `record_async`/`record_media_async` as thin
`asyncio.to_thread` wrappers (same precedent as
gateway/channel_directory.py) and await them from the seven coroutine call
sites; the sync functions stay for sync callers. The write completes before
the await returns, so lookup-after-send behaviour is unchanged.

Salvaged from #118883 by @Kyzcreig; re-shaped onto asyncio.to_thread (no
writer thread, no in-memory shadow, no queue), plus the missed
plugins/platforms/whatsapp/adapter.py record_media site.
2026-09-22 15:54:29 +05:30
Kyzcreig
ee8a3ded96 fix(gateway): move the thread-participation persist off the event loop
`ThreadParticipationTracker._save` (gateway/platforms/helpers.py) ends in
`atomic_json_write` -> `os.replace`, whose duration is unbounded under
filesystem pressure. All five call sites are coroutines on the inbound-message
or slash-command path:

    plugins/platforms/matrix/adapter.py   _resolve_message_context
                                          create_handoff_thread
    plugins/platforms/discord/adapter.py  _handle_message (x2)
                                          _handle_thread_create_slash

So that rename was paid inline on the running loop, stalling every other
adapter's polling, every in-flight turn and every heartbeat for as long as it
took.

The fix is at the CHOKE POINT rather than at five call sites:

  - `mark_async` does the in-memory insert synchronously and offloads only the
    persist via `asyncio.to_thread`. The insert must stay synchronous because
    both adapters gate on `thread_id in self._threads` immediately after
    marking; deferring it would make mention-gating depend on executor
    availability.
  - all five coroutine call sites now await it.
  - `mark` keeps its exact synchronous contract for the non-loop callers.
  - an RLock is added in the SAME commit that introduces the concurrency: the
    event loop used to serialize every caller by accident, and
    `atomic_json_write` makes each write atomic without making
    check/insert/trim/write atomic. Without it two concurrent marks lose one.

Enforcement is an AST class sweep, not an inventory: no `async def` under
gateway/ or plugins/ may call `<x>._threads.mark(...)`, and every
`mark_async(...)` must be awaited -- an un-awaited one never runs at all, so
the thread is neither persisted nor recorded in memory and mention-gating
re-prompts forever in a thread the bot already joined. A new adapter fails the
gate without anyone remembering a list.

tests/gateway/test_discord_thread_slash_expired_defer.py stubbed the tracker
with `SimpleNamespace(mark=...)`; it now uses the real tracker against a
tmp_path, so the test cannot rot silently the next time this surface moves, and
it additionally asserts the thread really was recorded.

Verified on this exact head, PYTHONPATH pinned to the worktree:
  tests/gateway/test_thread_tracker_mark_off_loop.py             7 passed
  expired-defer + admission-exemption + off-loop                10 passed
  -k "discord or matrix or thread" over tests/gateway  1136 passed, 13 failed
The 13 failures are INHERITED: a clean worktree at upstream/main 75e9567ca7
with none of these changes fails the identical 13.

Every guard is gate-proven -- reverting the RLock, the to_thread, one call
site, one `await`, the sync insert, or the dedupe short-circuit each fails its
own test and only its own.

(cherry picked from commit 6fc8037d292ec051fe76ad213903b2ee5534080a)
2026-09-22 15:54:29 +05:30
daedalus-opus
cb567242d7 fix(gateway): move the sticker-description cache write off the event loop
`gateway.sticker_cache._save_cache` ends in `atomic_json_write` -> `os.replace`,
whose duration is unbounded under filesystem pressure. Its only production
caller is Telegram's `_handle_sticker` -- an inbound-message coroutine -- so
every sticker that missed the cache paid that rename inline on the event loop,
stalling every other adapter and every in-flight turn in the process for its
duration.

Adds `cache_sticker_description_async`, which dispatches the existing sync form
via `asyncio.to_thread`; the Telegram call site awaits it. The sync form keeps
its exact contract and is what the wrapper dispatches to, so it is unchanged for
non-loop callers.

`cache_sticker_description` is a read-modify-write (`_load_cache` -> merge ->
`_save_cache`). `atomic_json_write` makes each WRITE atomic, not the TRIPLE.
While the call was inline the event loop serialized every caller and the race
could not be observed; moving the write to a worker thread introduces real
concurrency, so `_CACHE_LOCK` is added in this same commit rather than deferred.

Tests use no wall-clock thresholds. The liveness witness is ORDERING: the rename
is held open on a barrier released only by a background timer, and the sibling
task must have ticked BEFORE that release. Verified RED first:

- revert `asyncio.to_thread` to an inline call -> 2 failed
  ("the cache write ran on the event-loop thread")
- revert `_CACHE_LOCK` -> the concurrency test fails on the DURABLE FILE:
  "lost an entry ... keys = ['uid_a']"

`test_gate_proof_the_sync_form_does_block_the_loop` drives the identical barrier
through the sync form and asserts the loop DOES starve, so the liveness
assertion cannot pass vacuously.

GREEN: sticker off-loop + existing sticker cache tests -> 10 passed;
`tests/gateway -k "sticker or telegram"` -> 937 passed, 4 skipped; Ruff clean.

(cherry picked from commit 9752a9b1aba7536181d081cad9c1c32917a89680)
2026-09-22 15:54:29 +05:30
ethernet
c13287c915 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/electron/main.ts
#	hermes_cli/backup.py
#	hermes_cli/config.py
#	hermes_cli/plugin_catalog.py
#	hermes_cli/plugins_cmd.py
#	hermes_cli/plugins_cmd_catalog.py
#	hermes_cli/plugins_discovery.py
#	hermes_cli/profiles.py
#	hermes_cli/update_cmd_deps.py
#	pyproject.toml
#	tests/gateway/test_dm_topics.py
#	tests/hermes_cli/test_config.py
#	tests/hermes_cli/test_plugins_cmd.py
#	tests/hermes_cli/test_update_autostash.py
#	tests/tools/test_lazy_deps.py
#	tools/lazy_deps.py
#	tools/skill_ledger.py
#	utils.py
#	website/docs/user-guide/security.md
2026-09-22 05:16:50 -04:00
kshitijk4poor
c2d7bb89f1 refactor(discord): name the terminal liveness reasons in one constant
Declare `_TERMINAL_HEALTH_REASONS = frozenset({"socket_closed",
"client_closed"})` beside `_read_websocket_health` (the producer of those
literals at adapter.py:1705/:1707/:1715) and use it at the escalation
site in `_liveness_loop` (was `reason in ("socket_closed",
"client_closed")` at :1791).

WHY: the first-strike classification was a magic-string match against
literals produced 80 lines away with no shared definition. A rename in
the producer, or a new hard-closed reason, would silently demote the
escalation back to the threshold path with nothing failing. One named
set next to the producer makes the coupling visible. The
`tuple[bool, str]` return shape is unchanged (existing tests monkeypatch
the sampler with 2-tuple lambdas).

Proof: mutating the set to {"socket_closed"} fails
test_closed_transport_first_strike_forces_reconnect[client_closed];
restored, the liveness file is green.
2026-09-22 13:39:00 +05:30
kshitijk4poor
ac39011c2e refactor(discord): log the liveness-probe exit at INFO unconditionally
Collapse the INFO/WARNING level switch on the probe-exit log
(plugins/platforms/discord/adapter.py:1755-1766) to one unconditional
logger.info and delete `expected_exit`. Drop
test_unexpected_probe_exit_logs_warning, which was the only caller that
could reach the WARNING half.

WHY: the WARNING branch is unreachable in production. It needs
`_running=True and not _disconnecting and _client is None` at the guard,
and `_client` has exactly two production setters to None:

- adapter.py:1269 inside connect(): synchronously followed by
  `self._client = commands.Bot(...)` at :1271 with no await between, so
  the probe coroutine can never observe the None.
- adapter.py:1941 inside disconnect(): runs after `_disconnecting = True`
  (:1915) and after `await self._cancel_liveness_task()` (:1917), so the
  probe is already cancelled (exits via CancelledError at :1752) and the
  flag would read as an expected exit anyway.

No other `_client = None` in plugins/platforms/discord/, gateway/platforms/
base.py or gateway/run.py. `_running=False` and `_disconnecting=True` come
only from teardown paths, all "expected". The deleted test pinned this by
poking `adapter._client = None` directly — a state the code never produces.

This drops the level switch borrowed from #118504 but keeps its intent:
the probe exit is always logged with its full state (#118487, the probe
must never disappear silently).
2026-09-22 13:39:00 +05:30
kshitijk4poor
533796a833 refactor(discord): trim the first-strike comment to its WHY
The six-line comment above the terminal-reason check
(plugins/platforms/discord/adapter.py:1789-1794) restated the resume-swap
mechanism already documented in the test docstring and the issue. Keep the
one sentence that carries the intent (closed transport = confirmed death,
soft signals keep the threshold) and the #118487 pointer. No code change.
2026-09-22 13:39:00 +05:30
kshitijk4poor
fe225b20b1 fix(discord): treat client_closed as terminal like socket_closed
_read_websocket_health reports client_closed when Bot.is_closed() is
true. That is the same transport-dead state as socket_closed, yet it
stayed on the two-strike confirmation path and could show the same
1/2-then-silent-reset pattern if the bot task's done callback were ever
suppressed. Escalate both terminal reasons on the first strike.
2026-09-22 13:39:00 +05:30
kshitijk4poor
43499c4531 fix(discord): log an unexpected liveness-probe exit at WARNING (from #118504)
The guard exit in _liveness_loop fires for two very different reasons:
ordinary teardown (adapter stopped or disconnecting), and a still-running
adapter whose client vanished. The first is noise at INFO; the second
means the gateway keeps running with no watchdog (#118487) and must stand
out in an incident log. Pick the level from the exit cause instead of
logging both at INFO.

Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-22 13:39:00 +05:30
liuhao1024
b1a69d8178 fix(discord): escalate the first socket_closed liveness strike
A closed Gateway transport is a confirmed death, not a suspicion, but the
liveness probe treated socket_closed like every soft signal and waited for
a confirming strike (#118487). That strike never arrives: discord.py swaps
in a fresh socket while resuming, the next transport-side sample reads
healthy, and the strike counter silently resets while a resumed-but-deaf
session stays event-starved until the multi-hour event-silence default
elapses. The bot sits deaf until a manual gateway restart.

Escalate the first socket_closed strike straight to the forced-reconnect
path; soft signals (ack staleness, latency, event silence) keep the
confirmation threshold. Also leave a trace on every previously-silent
probe transition: counter resets now log, and probe exits log the flag
state, so an incident log can no longer confuse a healthy sample with a
dead watchdog task.

(cherry picked from commit a9b796e9e9b113595c049701caea46679430465b)
2026-09-22 13:39:00 +05:30
teknium1
0706dffca1 fix(config): route every config.yaml writer through one comment-preserving writer
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.

Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.

- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
  example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
  untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
  are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
  profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
2026-09-22 01:07:59 -07:00
ethernet
727daebce3 Merge remote-tracking branch 'origin/main' into ethie/pm-clean 2026-09-22 01:32:50 -04:00
Austin Pickett
992f569fc1 fix(telegram): admit replayed updates before PTB dispatch
Keep one adapter-owned update claim across core, observer and native plugin
handlers. Release failed preparation, but retain queued or externally handed-off
work through cancellation and PTB task completion.

Refs https://github.com/NousResearch/hermes-agent/issues/68502

Co-authored-by: Jasmine Naderi <jasmine@smfworks.com>
2026-09-22 00:55:46 -04:00
ethernet
1afa348e1c Merge branch 'rev/pm-core' into ethie/pm-clean 2026-09-21 19:52:58 -04:00
ethernet
029cabb466 feat(pm): hermes pm install --extra NAME; one shared hint for missing extras
Twenty-five call sites told users to run
`python -c "from pm import sync_venv; sync_venv(['x'], explicit=True)"`
because `hermes pm install` only took package names. Add `--extra`
(repeatable; syncs the venv with the named extras and nothing else) and
pm.install_hint(extra), the single builder every site now uses, so the
advice stays correct when the command changes.

A cold PM runtime under allow_lazy_installs:false now reports the extra
the caller wanted and the command that provisions both, instead of a
bare "pm-runtime: not installed".
2026-09-21 19:08:19 -04:00
ethernet
b1c1aa9a28 fix(platforms): raise, don't assert, when an ensured node/npm has no binary
assert is stripped under python -O, so a pm.ensure that recorded no
selected binary would have fallen through to str(None) as the npm/node
path. The four sites now raise pm.InstallError, which the surrounding
handlers already report.
2026-09-21 19:03:21 -04:00
ethernet
94a47320f1 fix(google_chat): request both extras before surfacing the restart
ensure_import raises InstallError("restart Hermes to activate…") after a
SUCCESSFUL install, so the loop aborted on 'google' and never asked for
'google-chat'; the restart landed back on a still-missing extra. Both
are requested in one pass; the first failure is what the registry logs.
2026-09-21 19:02:23 -04:00
ethernet
11bd9b512b fix(photon): declare PHOTON_NODE_BIN again
The adapter still honours PHOTON_NODE_BIN ahead of PM's pinned node, but
the pm hand-merge dropped it from optional_env, so the override became
undocumented and unreachable from setup.
2026-09-21 19:00:14 -04:00
ethernet
9d036f8ac9 merge origin/main (15 commits) into ethie/pm-clean
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
2026-09-21 08:16:57 -04:00
kshitijk4poor
7b660e66ee fix(homeassistant): detect the supervised launch through is_supervised_gateway_launch
gateway.restart already answers "was this gateway launched by a generated
service" (HERMES_SUPERVISED_CHILD or launchd's XPC_SERVICE_NAME, so a plist
that predates the marker still counts); the hint reuses it instead of a
second env read. The negative branch (plain unreachable host without a
supervisor) is now an unmarked test so the Linux lane keeps covering the
function; the macOS-only test holds the positive branch.
2026-09-21 17:16:23 +05:30
kshitijk4poor
eda2432cce fix(gateway): scope the Local Network hint and keep the osascript wrapper out of gateway process scans
Review of the first cut:

- The Home Assistant errno hint matched a bare 65 and two error strings on
  every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
  errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
  genuinely unreachable HA host was told macOS was blocking launchd. Gate on
  darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
  generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
  `gateway run`, so `hermes gateway stop` would have signalled osascript
  alongside the gateway. The canonical matcher now bails on an exact
  argv[0] basename of osascript (the gateway is its child and is matched on
  its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
  carries (osascript's own output) instead of implying it is redundant.
2026-09-21 17:16:23 +05:30
xxxigm
1d871110e1 fix(homeassistant): name the macOS Local Network block behind errno 65 under launchd (#71206)
A launchd-run gateway that macOS Local Network Privacy denies sees every LAN
connect fail with EHOSTUNREACH while the same URL works from Terminal.
Annotate the Home Assistant connect/reconnect log lines with the cause and
the remedy so the failure is actionable instead of a bare "No route to host".

Salvaged from #115196 (remedy text points at the regenerated launchd job).
2026-09-21 17:16:23 +05:30
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
chelsealong
7efacc8f62 fix(a2a): recover the streamed reply text instead of resolving empty
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).

_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".

(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
2026-09-20 18:45:42 -07:00
Forkbert
b113ab6de6 fix(photon): use GUID for liveness probe message id
(cherry picked from commit cdb6ccadb045e599db795823a100ae7b61d377ec)
2026-09-20 18:23:28 -07:00
teknium1
fd94fe9b93 fix(gateway): every adapter send_voice accepts the dispatch's is_voice kwarg
The base media dispatch passes is_voice= to send_voice; line, mattermost and
weixin had explicit signatures without it, so a non-image MEDIA attachment
routed as audio raised TypeError and was dropped — the same class as the
Matrix report (#102221, #116776). Adds a repo-wide signature invariant test.
2026-09-20 15:48:51 -07:00
xiaodu
b6505448f1 fix(matrix): accept the media dispatch's is_voice kwarg in send_voice
The base media dispatch calls send_voice(..., is_voice=is_voice) for every
audio MEDIA attachment (gateway/platforms/base.py _send_one). MatrixAdapter
.send_voice() accepted neither is_voice nor **kwargs, so every non-image
MEDIA delivery raised TypeError and the file was silently dropped — the
failure is visible in rotated logs since 2026-09-08 (never worked).

Accept the flag explicitly: is_voice=False -> plain m.audio in the original
format (no transcode); True or omitted (play_audio legacy callers) -> the
existing MSC3245 voice-bubble path with best-effort Ogg/Opus transcode.

Fixes #116776

(cherry picked from commit d4f89a725498e29a2ee0fe016f8bb98db08f358c)
2026-09-20 15:48:51 -07:00
fangliquan
45a701d8b0 fix(simplex): flatten alpha image thumbnails
(cherry picked from commit 97886801741f4367aadfe83ccb25b10cedb21149)
2026-09-20 15:48:01 -07:00
liuhao1024
7a23b00101 fix(whatsapp): refuse legacy-pidfile kills without a start-time fingerprint
The stale-bridge cleanup accepted a "node" + session-path cmdline substring
as kill evidence for legacy pidfiles (pid line only). A log tail, editor, or
grep that merely mentions the session path matches that same substring, so
the cleanup could SIGTERM a stranger process.

Require the kernel start-time fingerprint and fail closed when the pidfile
lacks one; the bridge-port scan (which verifies a node-executable listener)
reaps the orphan instead. The refusal reason in the warning now distinguishes
a fingerprint-less legacy pidfile from a recycled PID.

Flip the legacy-pidfile regression test to assert the fail-closed outcome.

Fixes #116883
2026-09-20 15:24:51 -07:00
teknium1
a57a5f90d2 fix: slot-busy skipped Telegram edits no longer count as shown text (review follow-up)
An interim edit skipped because the chat's shared send+edit slot was busy
returned a plain SendResult(success=True), so the stream consumer recorded
the never-shown text as _last_sent_text and reset _flood_strikes. A later
turn-final flood then saw _visible_prefix() == final text and either marked
the turn delivered or entered fallback with an empty continuation — the user
never saw the tail. The adapter now flags the skip in
raw_response={"skipped": True} and _edit_existing leaves the visible prefix
and flood state untouched, so the next tick retries and a flood fallback
re-sends exactly the unseen tail.
2026-09-20 12:17:23 -07:00
Haisam Abbas
fb2f3e213a fix(telegram): pace sends and edits against a shared per-chat slot
Telegram counts an editMessageText against the same per-chat allowance as a
sendMessage, but streaming previews paced only edits (DEFAULT_STREAMING_EDIT
_INTERVAL = 0.8s = 1.25 msg/s into one chat before any reply was sent) —
83% of measured flood penalties. One shared slot per chat: a send WAITS for
its slot (skipping would drop a message), an interim edit is SKIPPED (the
next tick shows the same text anyway), and the final edit is never gated
(the answer is never withheld). A per-adapter tuning knob keeps the slot
available to tests that model instantaneous bursts.
2026-09-20 12:17:23 -07:00
teknium1
1d6fbd0d1f fix: explicit empty backfill channel list disables the Discord scan (review follow-up)
`missed_message_backfill.channels: []` (YAML list or JSON-list string) returned
an empty set on main, and the backfill logged "no channels configured" and
skipped. The gate-CSV refactor treated an empty result as "unset" and fell
through to the allowed-union-free-response default, scanning channels the
operator had explicitly disabled. An explicit list is now authoritative even
when empty; only the default "" string still falls through to env/default.
2026-09-20 12:10:27 -07:00
teknium1
587bb10575 fix(platforms): Discord/WhatsApp/DingTalk gates honour allowlists stored as a JSON-list string
`hermes config set KEY '["-100","-200"]'` used to write the literal as one quoted
YAML string; the writer now emits a real list (acaac9a18, #88163) and the Telegram
gate decodes the legacy string shape (122ad719, #110213). The Discord, WhatsApp and
DingTalk gate parsers still comma-split that string into `{'["-100"', '"-200"]'}`,
so a config written before the writer fix silently locks every allowlisted chat,
channel or user out — with no warning.

Route every remaining comma-split gate through the shared
`gateway/platforms/_shared.py::decode_json_list_literal`:

- Discord `_gate_csv_set` (allowed/ignored/no-thread channels, allowed users/roles),
  `_discord_free_response_channels` and `_missed_message_backfill_channels` now share
  the one parser instead of three hand-rolled splits.
- WhatsApp `_coerce_allow_list` (allow_from, group_allow_from, free_response_chats).
- DingTalk `_csv_set` (allowed_users, allowed_chats, free_response_chats).

Plain CSV strings, YAML lists and malformed JSON keep their previous meaning.
2026-09-20 12:10:27 -07:00
Steve Hsu
97051e4194 fix(telegram): attach real video geometry and a thumbnail so large uploads aren't square
`sendVideo` gives back `width=320 height=320 duration=0` with no thumbnail once an
upload is large enough that Telegram skips its own video processing, and clients
then draw the message as a square tile — for portrait reels and 16:9 clips alike,
even though the delivered file itself is correct.

Measured on one 6 s 2560x1440 clip, inspecting the Bot API response: 4.9 MB and
9.8 MB keep `2560x1440` / duration 7 / a 320x180 thumbnail; 14.8 MB, 19.4 MB and
23.0 MB degrade to the square placeholder; the same 23.0 MB file sent with
`width`/`height`/`duration` plus a JPEG `thumbnail` comes back `2560x1440` with a
320x180 thumbnail.

Probe the local file with ffprobe and attach a 320px-wide JPEG frame from it, on
both Telegram send paths: the gateway adapter's `send_video` and the standalone
`hermes send` media sender. Both helpers return nothing when ffmpeg/ffprobe is
unavailable, which keeps the previous behaviour for hosts without them.
2026-09-20 11:20:26 -07:00
teknium1
c060a821d5 fix(discord): accept a pasted channel link at the _resolve_channel chokepoint too
HomeChannel normalization covers env and YAML homes, but cron
`deliver: discord:<link>` targets and thread metadata reach the adapter as
written and still died in int(). Share one helper
(gateway.config.discord_channel_id_from_link) between HomeChannel and the
resolver every outbound Discord target passes through; message links and
non-link strings keep their existing error path. Tests trimmed to two
invariants: config load (env + YAML) and the resolver entry point.
2026-09-20 10:49:49 -07:00
chelsealong
f45ed8e4ec docs(matrix): document the <=2-member DM auto-classification and its config bypass
_resolve_room_identity() classifies any room with <=2 joined members as a DM
regardless of m.direct or an explicit room name, so those rooms silently
bypass MATRIX_ALLOWED_ROOMS, MATRIX_FREE_RESPONSE_ROOMS, and
MATRIX_REQUIRE_MENTION, and use DM threading instead of
MATRIX_AUTO_THREAD/MATRIX_SESSION_SCOPE. This was previously only visible in
an inline code comment, not in the env-var docs an operator would read.

Fixes #114733
2026-09-20 10:42:25 -07:00
chelsealong
0c872b611c fix(telegram): point dm_topics prerequisite text at BotFather Threaded Mode
The dm_topics "not a forum" warning and the matching docs section told
users to tap the bot's name in the DM and toggle "Topics" in chat
settings. That toggle only exists for group forums; a bot DM has no
such control. The actual prerequisite is Threaded Mode, enabled by the
bot owner via the BotFather Mini App (Bot Settings -> Threads
Settings) -- already documented correctly a few sections later in the
same file, under the /topic prerequisites.

Fixes #115019
2026-09-20 10:39:37 -07:00
KeyArgo
e7509fa1a4 fix(telegram): observe sibling-bot wake-word messages dropped by the bot-to-bot gate
With `require_mention: true` + `bots_require_mention: true` + wake words in `mention_patterns`, a message
authored by another Hermes bot that addressed this bot by name matched no dispatch path (the
bot-to-bot loop breaker skips it) and was also refused by the observe gate (`mention_patterns`
matches are assumed dispatched), so it vanished from both paths with no log line.

Factor the loop-breaker predicate into `_bot_sender_suppressed` and consult it in
`_should_observe_unmentioned_group_message`, so every message the dispatcher drops because of
`bots_require_mention` is kept as observed context instead of being lost (#115119).
2026-09-20 10:30:11 -07:00
fangliquan
3cbd3cce0a fix(whatsapp): honor configured reply prefix in bridge replies
whatsapp.reply_prefix from config.yaml was written into the bridge env and then
popped again by the WHATSAPP_* passthrough loop (the key was in
_BRIDGE_PASSTHROUGH_ENV and the scoped env lookup came back empty), so bridge.js
always fell back to its built-in header and the documented reply_prefix: ""
could not disable it. Resolve the prefix once (scoped env first, then the
adapter value) and keep it out of the passthrough loop.

Fixes #116059
2026-09-20 10:20:40 -07:00
hardwork9047
a0e7fc4a9e fix(gateway): consult _in_bot_thread in the Discord admission gate
`_discord_message_admission()` drops a message that mentions someone other than
the bot when `DISCORD_IGNORE_NO_MENTION` is on (the default) and the channel is
not free-response. It did so without consulting `_in_bot_thread()`, unlike the
other two ingress paths — `_dispatch_recovered_message()` (adapter.py:2292) and
`_handle_message()` (adapter.py:5946). Admission runs on both and returns
`False` unconditionally, so it overrode the thread exemption they grant.

The result was an asymmetry with no obvious cause from the outside: in a thread
the bot had joined, a message with no mention at all was admitted (an empty
`message.mentions` skips the enclosing block), while the same message with one
mention of a third party was dropped. A thread the bot is a participant in is
the one place "addressed to someone else" is least likely to hold.

`thread_require_mention` still gates multi-bot threads, since that check lives
inside `_in_bot_thread()`.

The drop also emitted nothing at any log level, leaving `gateway.log` identical
whether the gate fired or the event never arrived; add a debug line so the two
can be told apart.

Fixes #116568
2026-09-20 10:20:03 -07:00
ethernet
925c08ceca fix: CI python-tests backlog — no import-time dependency syncs, CI-shaped test fixtures
Production:
- agent/bedrock_adapter.py, agent/vertex_adapter.py: pm.ensure_import ran at
  module import. In any process that imports these modules without a committed
  PM selection (CI's build_environment test venv, a fresh checkout) that sync
  rebuilt the dependency environment mid-process and replaced sys.path with a
  generation missing the caller's own packages (anthropic, aiohttp vanished).
  The extra is now ensured at first client build / credential request.
- plugins/platforms/matrix/adapter.py: a complete install needs no
  ensure_and_bind round trip; only a partial one syncs.
- tools/browser_tool.py: drop the facade's duplicate warm_agent_browser_npx_cache
  shim; the compat pointer already resolves to browser_tool_install.

Test harness:
- tests/home_io_guard.py: PATH-entry probes (shutil.which) and the running
  interpreter's own installation (stdlib reads, realpath ancestry, fixture
  symlinks into it) are not Hermes state; a patched Path.expanduser must not
  crash the guard. run_tests.sh no longer filters PATH — the guard owns it.
- tests/tui_gateway/conftest.py: import hermes_bootstrap before any file opens
  a MagicMock hermes_constants window (6 files exited the process at boot).
- tests/hermes_cli/conftest.py probe_root: scratch checkouts the import guard
  probes need hermes_bootstrap.py (the launcher imports it).
- tests/pm/_fixtures.py stage_host_python: a copied relocatable python needs
  its stdlib beside it (No module named 'encodings' on CI).
- tests/install/e2e-assets/smoke-env.mjs: dependency-free env shaping so the
  source-build-env probe runs under bare node (main deleted the Playwright
  entry it was imported through).
- adapt main's new tests to branch seams (model_metadata_http, launch
  completion tail, CI toolchain exports uv after python, source_launch
  hermes_cli stub, systemd_notify single marker).
2026-09-20 11:47:06 -04:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00
kshitijk4poor
de083436bd refactor(gateway): drop Telegram's shadowing _coerce_float_extra; module-level math import
Gate review: TelegramAdapter kept a same-named override with a weaker contract (no finite
guard, negatives clamped instead of reset), so the hierarchy had two parsers under one name.
Both Telegram callers pass explicit bounds; the base method is a strict superset for them.
The cadence test now pins the behaviour (≤ 0.5 s / ≤ 1.0 s) instead of echoing the constant.
2026-09-20 18:14:58 +05:30
kshitijk4poor
530fdb5867 refactor(gateway): one text-batch cadence and one extra-float parser on BasePlatformAdapter
Gate review: the fix left three adapter-private copies of `_coerce_float_extra` and the
0.3/2.0/1.0/4.0 cadence literals in three files. The parser and the cadence constants now
live on BasePlatformAdapter beside the delay attrs they configure; WhatsApp and Weixin call
`_configure_text_batch_delays()`, Telegram reads the same constants through its env helper.
The clamp test is parametrized over both adapters and the Weixin docs name the ceilings.
2026-09-20 18:14:58 +05:30
kshitijk4poor
130b596f6f fix(whatsapp,weixin): text-batch delays default to Telegram cadence
WhatsApp debounced text for 5s (10s near a split) and Weixin for 3s/5s
before dispatching, so every reply paid multiple seconds of idle latency
that Telegram never pays (0.3s/1.0s). Default both adapters to Telegram's
cadence and mirror its ceilings (2.0s / 4.0s, split >= base delay) via the
existing _coerce_float_extra seam. The config keys are unchanged; 0 still
dispatches immediately. Docs updated.

Spotted via #44896 (@liuhao1024). Fixes #44883, refs #25056.
2026-09-20 18:14:58 +05:30