Commit Graph

1215 Commits

Author SHA1 Message Date
Michael Versluis (Berry)
aca906df2c fix(gateway): keep status caches bounded across awaits (#87479)
Telegram's per-(chat_id, status_key) status-message cache grew without
bound; give it the same _STATUS_MESSAGE_IDS_MAX=2000 FIFO half-trim the
Slack adapter already has. In both adapters, guard the post-await write-back
after a successful edit with a compare-before-write (only re-store the id if
the cached entry is still the one we edited) so an eviction or replacement
that happened during the await is not undone.

Partial salvage of #87480: kept the Telegram bound and both compare-before-write
guards (re-applied by hand, 17586 behind), defined the max as a class attr
like Slack instead of an instance attr, dropped the 4 new tests.

(cherry picked from commit 43ae95e98e)
2026-09-23 21:47:15 +05:30
devorun
b331e5d282 fix(buzz): persist channel cursors off the event loop
_handle_events called _save_cursors inline whenever a batch moved the
cursor. It ends in atomic_json_write (mkstemp + fsync + os.replace), and
on the WebSocket transport _handle_events runs once per inbound EVENT
frame, so every message paid an fsync on the gateway's event loop,
stalling every other adapter and in-flight turn for its duration.

Split the snapshot from the write: the payload is still built on the
loop (_channel_state is loop-owned), and the write goes through
asyncio.to_thread. The snapshot is taken under an asyncio.Lock so a
slower, older write can never land after a newer one and regress the
durable cursor. connect() keeps the synchronous _save_cursors.

(cherry picked from commit 299929741429262f497b86a67940ffb37faa1694)
2026-09-23 21:47:15 +05:30
John Paul Soliva
b5a300fe34 fix(email): pairing, decline and gateway grants reach the gateway instead of dying in the adapter pre-gate
EmailAdapter._sender_accepted runs before any MessageEvent exists and
read only EMAIL_ALLOWED_USERS. Unset, it dropped every sender unless
allow-all was on; set, it dropped everyone not listed. The gateway's
own handling therefore never ran for email:
platforms.email.unauthorized_dm_behavior "pair" (the setup wizard's
"Use DM pairing") and "decline" sent nothing, and a sender admitted by
GATEWAY_ALLOWED_USERS or an approved pairing was dropped. bb304b4914
turned the empty-allowlist branch into drop-all after #50568 had made
"pair" email's explicit opt-in.

The gate now keeps a sender listed by address in EMAIL_ALLOWED_USERS
or GATEWAY_ALLOWED_USERS, a sender the registered gateway
authorization check admits (that is the only reader of the pairing
store), and, under an explicit pair or decline, an unknown sender the
gateway will answer. The default "ignore" still drops unknown senders
before a MessageEvent exists, so the mail-loop guard from fd9c32c0f2
holds.

Three guards keep the wider gate from widening access, and close two
forged-From: paths main already had:
- A sender admitted only so the gateway can answer it (pair or
  decline) must authenticate its From:, open access or not: the
  pairing code or refusal is mailed back to that address. A granted
  sender still needs it short of open access, since a pairing grant
  keys on From: just as the allowlist does. Open access follows the
  gateway's own order: EMAIL_ALLOW_ALL_USERS wins over a list, while
  GATEWAY_ALLOW_ALL_USERS beside a list admits nobody extra, so it no
  longer exempts a listed address from From: authentication either
  (on main a forged From: of a listed address got through there).
- Open access comes from the gateway's own verdict when a check is
  registered. GATEWAY_ALLOW_ALL_USERS beside a GATEWAY_ALLOWED_USERS
  list grants a stranger nothing there, so the env flag alone no
  longer exempts one from From: authentication (that path mailed a
  pairing code to a forged From: on main too).
- A sender whose local part alone matches an allowlist entry is
  dropped. The gateway's check also matches an address by its bare
  local part (#119446), so without this, GATEWAY_ALLOWED_USERS=alice
  (a chat username) would admit or pair alice@<any domain>. The lists
  are parsed as the gateway parses them, JSON list literals included,
  or '["alice"]' would slip past this guard.

_allowlist_in_effect only served the old condition and is removed.
The scope tests now assert the same scoped reads through
_sender_accepted, with GATEWAY_ALLOWED_USERS covered as well.

Measured end to end with the real GatewayRunner callback wired
(adapter -> gateway ingress):
- pair, decline, GATEWAY_ALLOWED_USERS and an approved pairing each
  went from 0 events reaching the gateway to 1. pair mails a pairing
  code, decline mails one refusal.
- An unauthenticated From: in pair mode, for a paired address or for a
  GATEWAY_ALLOWED_USERS address still reaches nothing.
- A bare GATEWAY_ALLOWED_USERS=stranger entry lets nothing from
  stranger@<domain> through, under ignore or pair. Without the
  local-part guard that mail reached the gateway in both.
- The same holds for a JSON-literal list, and a pair-mode stranger
  with a forged From: under allow-all beside an EMAIL_ or
  GATEWAY_ALLOWED_USERS list reaches nothing.
- The default still drops.
2026-09-23 07:48:32 -07:00
Austin Pickett
a1838ea87a fix(telegram): keep update receipts across adapter rebuilds and restarts (#120257)
Completed update IDs lived only in the adapter's memory. The gateway
reconnect watcher builds a new TelegramAdapter and connects it with
is_reconnect=True, which keeps Telegram's pending queue, and a new PTB
Updater polls from offset 0. Telegram then resends every update whose
acknowledgement (the next getUpdates offset, or the cleanup call in
Updater.stop) never landed, and the fresh adapter admitted them again.

Write completed IDs to telegram_update_receipts_<bot_id>.json in the
adapter's Hermes home and seed admission from it once per bot. Receipts
older than 24h are dropped: the Bot API keeps unconfirmed updates no
longer than that, and it keeps the lookup clear of the random ID restart
Telegram may do after a week without updates. Writes are coalesced and
run off the loop; disconnect waits for the last one.

Refs #68502

Co-authored-by: Joe Githler <5716896+NoTimeforInfinity@users.noreply.github.com>
2026-09-23 10:05:17 -04:00
teknium1
b3624cc3af docs(imessage): hints match the stripped-on-link send path
Photon messages with a link now arrive stripped rather than raw, and BlueBubbles
keeps link URLs, so the hints stop describing the old leaks.
2026-09-23 03:10:15 -07:00
teknium1
73ba639092 fix(imessage): strip markdown on every Photon plain send and keep link URLs on BlueBubbles
Follow-up on the two salvaged commits:
- One _sidecar_payload_text() helper decides the /send text for both the adapter
  path and _standalone_send (cron / send_message), which had the same raw-markdown
  leak on URL-bearing messages and never stripped with PHOTON_MARKDOWN=false.
- BlueBubbles (same iMessage surface) keeps [label](url) targets as bare URLs.
- Tests trimmed to two invariants.
2026-09-23 03:10:15 -07:00
Semir Kabir
ccfb68688e fix(photon): stop advertising fenced code-block support on iMessage
The gateway's tool-progress path emits terminal commands as fenced code
blocks on any adapter whose supports_code_blocks is True. The Photon
adapter set that flag from PHOTON_MARKDOWN, but the sidecar's /send
router (send-format.mjs) silently routes every URL-bearing message
through the plain-text builder — where fences survive as literal backtick
characters — and even the markdown path renders a fence as inline
monospace text, not a block. Net effect: raw fenced terminal commands
(and raw markdown around them) surfaced in iMessage bubbles no matter
what the user's prompt-level style rules said.

Fix the whole class: the adapter never claims code-block support, so the
gateway emits its compact one-line tool preview instead. Prose markdown
passthrough (bold/italic/headings) is unchanged, and no config change is
required — tool_progress can stay at its user's preferred mode.

Tests: new capability + E2E tests pin that the gateway cannot emit a
fence for Photon with real adapter + real display resolution; the old
supports_code_blocks-mirrors-env expectation (which pinned the buggy
behavior) now asserts the flag is always False.
2026-09-23 03:10:15 -07:00
Eddie Wang
449ac2fb8f fix(photon): strip markdown when the sidecar downgrades to the text builder
chooseSendFormat() routes markdown containing a raw http(s) URL through
spectrum-ts' text() builder, because the markdown builder's iMessage data
detection 500s on those messages (#73615). text() ships the payload
verbatim, and nothing strips the markers on the way down, so any reply
that mentions a link arrives in iMessage as literal markdown source:

    **Release 1.2.0** is out
    - **EUR 5** off this month
    https://example.com/releases/1.2.0

Every ** is visible in the bubble. Remove the URL from that same reply
and it renders correctly, which is what makes the URL the trigger rather
than the content.

spectrum-ts documents markdown() as degrading to readable plain text on
platforms without native support "instead of surfacing raw ** markers".
Selecting text() ourselves opts out of that guarantee, so the adapter has
to honour it instead.

Strip in _sidecar_send() when the payload is markdown and the sidecar will
downgrade it. The format key is deliberately preserved: the sidecar owns
the builder choice (test_rich_links.py pins that contract), and an older
sidecar without chooseSendFormat must keep rendering natively.

_send_plain_fallback() selects the text builder explicitly via
markdown=False and had the same leak, so it strips too.

Reuses the shared strip_markdown() helper rather than adding a second
implementation, so the PHOTON_MARKDOWN=false path and this path produce
identical output and both inherit any future fix to that helper.

Diagnosis previously reported in #85733, which was closed unmerged.

Stripping must not take the URL with it, though. The shared helper
collapsed [label](url) to label alone, which on this path is worse than
raw markdown: iMessage auto-links bare URLs and nothing else, so the
reply arrives with a description and no way to reach the link.

    Open the itinerary on Google Flights     <- URL gone entirely

So strip_markdown() takes keep_link_targets, which rewrites
[label](https://url) as "label\nurl" (own line, because these URLs are
often long) and leaves non-http targets such as mailto: or relative
paths label-only, since iMessage won't linkify those either. The default
is unchanged, so the SMS, IRC, Feishu and QQ callers keep dropping the
target as before.

plugins/platforms/line/adapter.py already carries a private
strip_markdown_preserving_urls() for exactly this reason ("LINE
auto-links bare URLs only"). This moves the behaviour behind the shared
helper instead, so Photon's three plain-text paths -- the downgrade,
_send_plain_fallback(), and PHOTON_MARKDOWN=false -- all agree.

gateway/platforms/bluebubbles.py is the same iMessage surface with the
same loss; left alone here to keep this change to one platform.
2026-09-23 03:10:15 -07:00
teknium1
6fa6df75f2 fix(imessage): platform hints steer replies to plain texting style
The Photon hint told the model "Markdown is rendered (bold, italics, lists,
code)", so replies came back with headers, code fences and backticks. That
is wrong on the real delivery path: the sidecar sends any message containing
a URL as raw text (every *, #, ``` and | shows literally), and even on the
markdown path spectrum-ts flattens headings to bold, tables to "a | b"
rows, and turns code into Unicode math-monospace glyphs that break when
copied. The hint also claimed attachments are metadata-only, which stopped
being true when native media send landed.

Both iMessage hints (Photon plugin, BlueBubbles built-in) now ask for a
texting register: short, answer first, no headers/tables/fences/backticks,
commands on their own plain line so they copy, bare URLs. BlueBubbles also
notes that strip_markdown drops the URL of [text](url) links.
2026-09-23 02:36:34 -07:00
teknium1
21d0b12958 feat(plugins): late-loaded plugins wire their platform handlers live (#87770)
A plugin that finished loading after an adapter connected never got its platform
handlers (slash commands, button callbacks, inbound transforms) registered until a
gateway restart, silently. Three pieces, one seam shared by every surface:

1. Discovery listener: PluginManager.on_plugin_loaded(cb) fires from INSIDE
   discover_and_load for the plugins a sweep newly loaded (diff of the loaded set),
   with a per-plugin activation summary (hermes_cli/plugins_activation.py):
   activated_now {gateway_commands, gateway_transforms, hooks, callbacks} vs
   deferred {tools, prompt, mcp_servers}. Every mid-run load path now performs a real
   discover_plugins(force=True): CLI install/enable (via the gateway), Desktop/TUI
   plugins.manage install/toggle/update, dashboard REST install, tool-triggered
   force re-discovery, the new `reload-plugins` control-socket verb. A non-forced
   discover_plugins() short-circuits on _discovered, which is why reload.mcp after
   a mid-run install used to reload the OLD server set.
2. Idempotent re-wire: BasePlatformAdapter.rewire_plugin_handlers() runs only
   factories not yet wired on the live native client (keyed (plugin, qualname);
   a force reload hands back new function objects). Telegram hoists late handlers
   ahead of core's catch-all filters.COMMAND / CallbackQueryHandler (PTB dispatches
   the first match per group) and re-wires on the transient-init rebuild; Slack
   dedupes register_slack_action_handler per AsyncApp. The gateway runner
   subscribes per served profile and re-wires on the loop.
3. Scope limit + honest messaging: handlers only. Tools/prompt stay deferred to
   the next session (prompt-cache invariant), MCP servers to mcp.reload; the CLI
   hint and plugins.manage results (activation, gateway_reloaded,
   restart_required only when no gateway answered) say exactly that.
2026-09-22 09:50:22 -07:00
kshitijk4poor
b9ec39939b fix(discord): pre-seed the starter dedup before awaiting mark_async
_auto_create_thread's dedup pre-seed (is_duplicate(str(thread.id))) stops
the echo MESSAGE_CREATE Discord fires for the starter (id == thread.id)
from re-running the request. This PR turned the tracker persist into
`await self._threads.mark_async(thread_id)` and placed it BEFORE the
pre-seed; to_thread always suspends, so the echo's handler could reach
_discord_message_admission -> is_duplicate during the os.replace window,
claim the id first, and rerun the starter. On main both statements were
sync, so no window existed.

Move the pre-seed (with its comment) directly after
`thread_id = str(thread.id)`; the await now follows it.

Other mark_async call sites checked (discord :4717 slash create-thread,
discord :6140 pre-dispatch, matrix :1438 create_thread, matrix :2083
inbound): nothing after those awaits relies on state a concurrent event
could claim first, so no further reorder.

PROOF: tests/gateway/test_discord_double_dispatch.py::
test_thread_starter_duplicate_dropped now installs a recording mark_async
that asserts thread.id is already in _dedup._seen when it is awaited.
Red with the pre-swap order (1 failed), green after; ruff clean,
check-windows-footguns clean, real import of the adapter from the
worktree ok, scripts/run_tests.sh on the PR's test files green.
2026-09-22 15:54:29 +05:30
kshitijk4poor
7dd6bee33b fix(gateway): rich_sent_store writes run off the event loop
`record`/`record_media` do a read-modify-write of the JSON index ending in
`atomic_json_write`/`os.replace`, and every coroutine on the WhatsApp Cloud,
WhatsApp bridge and Telegram send/inbound paths called them inline on the
loop thread, stalling the whole gateway for a filesystem write nothing
awaits. Add `record_async`/`record_media_async` as thin
`asyncio.to_thread` wrappers (same precedent as
gateway/channel_directory.py) and await them from the seven coroutine call
sites; the sync functions stay for sync callers. The write completes before
the await returns, so lookup-after-send behaviour is unchanged.

Salvaged from #118883 by @Kyzcreig; re-shaped onto asyncio.to_thread (no
writer thread, no in-memory shadow, no queue), plus the missed
plugins/platforms/whatsapp/adapter.py record_media site.
2026-09-22 15:54:29 +05:30
Kyzcreig
ee8a3ded96 fix(gateway): move the thread-participation persist off the event loop
`ThreadParticipationTracker._save` (gateway/platforms/helpers.py) ends in
`atomic_json_write` -> `os.replace`, whose duration is unbounded under
filesystem pressure. All five call sites are coroutines on the inbound-message
or slash-command path:

    plugins/platforms/matrix/adapter.py   _resolve_message_context
                                          create_handoff_thread
    plugins/platforms/discord/adapter.py  _handle_message (x2)
                                          _handle_thread_create_slash

So that rename was paid inline on the running loop, stalling every other
adapter's polling, every in-flight turn and every heartbeat for as long as it
took.

The fix is at the CHOKE POINT rather than at five call sites:

  - `mark_async` does the in-memory insert synchronously and offloads only the
    persist via `asyncio.to_thread`. The insert must stay synchronous because
    both adapters gate on `thread_id in self._threads` immediately after
    marking; deferring it would make mention-gating depend on executor
    availability.
  - all five coroutine call sites now await it.
  - `mark` keeps its exact synchronous contract for the non-loop callers.
  - an RLock is added in the SAME commit that introduces the concurrency: the
    event loop used to serialize every caller by accident, and
    `atomic_json_write` makes each write atomic without making
    check/insert/trim/write atomic. Without it two concurrent marks lose one.

Enforcement is an AST class sweep, not an inventory: no `async def` under
gateway/ or plugins/ may call `<x>._threads.mark(...)`, and every
`mark_async(...)` must be awaited -- an un-awaited one never runs at all, so
the thread is neither persisted nor recorded in memory and mention-gating
re-prompts forever in a thread the bot already joined. A new adapter fails the
gate without anyone remembering a list.

tests/gateway/test_discord_thread_slash_expired_defer.py stubbed the tracker
with `SimpleNamespace(mark=...)`; it now uses the real tracker against a
tmp_path, so the test cannot rot silently the next time this surface moves, and
it additionally asserts the thread really was recorded.

Verified on this exact head, PYTHONPATH pinned to the worktree:
  tests/gateway/test_thread_tracker_mark_off_loop.py             7 passed
  expired-defer + admission-exemption + off-loop                10 passed
  -k "discord or matrix or thread" over tests/gateway  1136 passed, 13 failed
The 13 failures are INHERITED: a clean worktree at upstream/main 75e9567ca7
with none of these changes fails the identical 13.

Every guard is gate-proven -- reverting the RLock, the to_thread, one call
site, one `await`, the sync insert, or the dedupe short-circuit each fails its
own test and only its own.

(cherry picked from commit 6fc8037d292ec051fe76ad213903b2ee5534080a)
2026-09-22 15:54:29 +05:30
daedalus-opus
cb567242d7 fix(gateway): move the sticker-description cache write off the event loop
`gateway.sticker_cache._save_cache` ends in `atomic_json_write` -> `os.replace`,
whose duration is unbounded under filesystem pressure. Its only production
caller is Telegram's `_handle_sticker` -- an inbound-message coroutine -- so
every sticker that missed the cache paid that rename inline on the event loop,
stalling every other adapter and every in-flight turn in the process for its
duration.

Adds `cache_sticker_description_async`, which dispatches the existing sync form
via `asyncio.to_thread`; the Telegram call site awaits it. The sync form keeps
its exact contract and is what the wrapper dispatches to, so it is unchanged for
non-loop callers.

`cache_sticker_description` is a read-modify-write (`_load_cache` -> merge ->
`_save_cache`). `atomic_json_write` makes each WRITE atomic, not the TRIPLE.
While the call was inline the event loop serialized every caller and the race
could not be observed; moving the write to a worker thread introduces real
concurrency, so `_CACHE_LOCK` is added in this same commit rather than deferred.

Tests use no wall-clock thresholds. The liveness witness is ORDERING: the rename
is held open on a barrier released only by a background timer, and the sibling
task must have ticked BEFORE that release. Verified RED first:

- revert `asyncio.to_thread` to an inline call -> 2 failed
  ("the cache write ran on the event-loop thread")
- revert `_CACHE_LOCK` -> the concurrency test fails on the DURABLE FILE:
  "lost an entry ... keys = ['uid_a']"

`test_gate_proof_the_sync_form_does_block_the_loop` drives the identical barrier
through the sync form and asserts the loop DOES starve, so the liveness
assertion cannot pass vacuously.

GREEN: sticker off-loop + existing sticker cache tests -> 10 passed;
`tests/gateway -k "sticker or telegram"` -> 937 passed, 4 skipped; Ruff clean.

(cherry picked from commit 9752a9b1aba7536181d081cad9c1c32917a89680)
2026-09-22 15:54:29 +05:30
kshitijk4poor
c2d7bb89f1 refactor(discord): name the terminal liveness reasons in one constant
Declare `_TERMINAL_HEALTH_REASONS = frozenset({"socket_closed",
"client_closed"})` beside `_read_websocket_health` (the producer of those
literals at adapter.py:1705/:1707/:1715) and use it at the escalation
site in `_liveness_loop` (was `reason in ("socket_closed",
"client_closed")` at :1791).

WHY: the first-strike classification was a magic-string match against
literals produced 80 lines away with no shared definition. A rename in
the producer, or a new hard-closed reason, would silently demote the
escalation back to the threshold path with nothing failing. One named
set next to the producer makes the coupling visible. The
`tuple[bool, str]` return shape is unchanged (existing tests monkeypatch
the sampler with 2-tuple lambdas).

Proof: mutating the set to {"socket_closed"} fails
test_closed_transport_first_strike_forces_reconnect[client_closed];
restored, the liveness file is green.
2026-09-22 13:39:00 +05:30
kshitijk4poor
ac39011c2e refactor(discord): log the liveness-probe exit at INFO unconditionally
Collapse the INFO/WARNING level switch on the probe-exit log
(plugins/platforms/discord/adapter.py:1755-1766) to one unconditional
logger.info and delete `expected_exit`. Drop
test_unexpected_probe_exit_logs_warning, which was the only caller that
could reach the WARNING half.

WHY: the WARNING branch is unreachable in production. It needs
`_running=True and not _disconnecting and _client is None` at the guard,
and `_client` has exactly two production setters to None:

- adapter.py:1269 inside connect(): synchronously followed by
  `self._client = commands.Bot(...)` at :1271 with no await between, so
  the probe coroutine can never observe the None.
- adapter.py:1941 inside disconnect(): runs after `_disconnecting = True`
  (:1915) and after `await self._cancel_liveness_task()` (:1917), so the
  probe is already cancelled (exits via CancelledError at :1752) and the
  flag would read as an expected exit anyway.

No other `_client = None` in plugins/platforms/discord/, gateway/platforms/
base.py or gateway/run.py. `_running=False` and `_disconnecting=True` come
only from teardown paths, all "expected". The deleted test pinned this by
poking `adapter._client = None` directly — a state the code never produces.

This drops the level switch borrowed from #118504 but keeps its intent:
the probe exit is always logged with its full state (#118487, the probe
must never disappear silently).
2026-09-22 13:39:00 +05:30
kshitijk4poor
533796a833 refactor(discord): trim the first-strike comment to its WHY
The six-line comment above the terminal-reason check
(plugins/platforms/discord/adapter.py:1789-1794) restated the resume-swap
mechanism already documented in the test docstring and the issue. Keep the
one sentence that carries the intent (closed transport = confirmed death,
soft signals keep the threshold) and the #118487 pointer. No code change.
2026-09-22 13:39:00 +05:30
kshitijk4poor
fe225b20b1 fix(discord): treat client_closed as terminal like socket_closed
_read_websocket_health reports client_closed when Bot.is_closed() is
true. That is the same transport-dead state as socket_closed, yet it
stayed on the two-strike confirmation path and could show the same
1/2-then-silent-reset pattern if the bot task's done callback were ever
suppressed. Escalate both terminal reasons on the first strike.
2026-09-22 13:39:00 +05:30
kshitijk4poor
43499c4531 fix(discord): log an unexpected liveness-probe exit at WARNING (from #118504)
The guard exit in _liveness_loop fires for two very different reasons:
ordinary teardown (adapter stopped or disconnecting), and a still-running
adapter whose client vanished. The first is noise at INFO; the second
means the gateway keeps running with no watchdog (#118487) and must stand
out in an incident log. Pick the level from the exit cause instead of
logging both at INFO.

Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-22 13:39:00 +05:30
liuhao1024
b1a69d8178 fix(discord): escalate the first socket_closed liveness strike
A closed Gateway transport is a confirmed death, not a suspicion, but the
liveness probe treated socket_closed like every soft signal and waited for
a confirming strike (#118487). That strike never arrives: discord.py swaps
in a fresh socket while resuming, the next transport-side sample reads
healthy, and the strike counter silently resets while a resumed-but-deaf
session stays event-starved until the multi-hour event-silence default
elapses. The bot sits deaf until a manual gateway restart.

Escalate the first socket_closed strike straight to the forced-reconnect
path; soft signals (ack staleness, latency, event silence) keep the
confirmation threshold. Also leave a trace on every previously-silent
probe transition: counter resets now log, and probe exits log the flag
state, so an incident log can no longer confuse a healthy sample with a
dead watchdog task.

(cherry picked from commit a9b796e9e9b113595c049701caea46679430465b)
2026-09-22 13:39:00 +05:30
teknium1
0706dffca1 fix(config): route every config.yaml writer through one comment-preserving writer
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.

Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.

- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
  example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
  untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
  are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
  profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
2026-09-22 01:07:59 -07:00
Austin Pickett
992f569fc1 fix(telegram): admit replayed updates before PTB dispatch
Keep one adapter-owned update claim across core, observer and native plugin
handlers. Release failed preparation, but retain queued or externally handed-off
work through cancellation and PTB task completion.

Refs https://github.com/NousResearch/hermes-agent/issues/68502

Co-authored-by: Jasmine Naderi <jasmine@smfworks.com>
2026-09-22 00:55:46 -04:00
kshitijk4poor
7b660e66ee fix(homeassistant): detect the supervised launch through is_supervised_gateway_launch
gateway.restart already answers "was this gateway launched by a generated
service" (HERMES_SUPERVISED_CHILD or launchd's XPC_SERVICE_NAME, so a plist
that predates the marker still counts); the hint reuses it instead of a
second env read. The negative branch (plain unreachable host without a
supervisor) is now an unmarked test so the Linux lane keeps covering the
function; the macOS-only test holds the positive branch.
2026-09-21 17:16:23 +05:30
kshitijk4poor
eda2432cce fix(gateway): scope the Local Network hint and keep the osascript wrapper out of gateway process scans
Review of the first cut:

- The Home Assistant errno hint matched a bare 65 and two error strings on
  every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
  errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
  genuinely unreachable HA host was told macOS was blocking launchd. Gate on
  darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
  generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
  `gateway run`, so `hermes gateway stop` would have signalled osascript
  alongside the gateway. The canonical matcher now bails on an exact
  argv[0] basename of osascript (the gateway is its child and is matched on
  its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
  carries (osascript's own output) instead of implying it is redundant.
2026-09-21 17:16:23 +05:30
xxxigm
1d871110e1 fix(homeassistant): name the macOS Local Network block behind errno 65 under launchd (#71206)
A launchd-run gateway that macOS Local Network Privacy denies sees every LAN
connect fail with EHOSTUNREACH while the same URL works from Terminal.
Annotate the Home Assistant connect/reconnect log lines with the cause and
the remedy so the failure is actionable instead of a bare "No route to host".

Salvaged from #115196 (remedy text points at the regenerated launchd job).
2026-09-21 17:16:23 +05:30
chelsealong
7efacc8f62 fix(a2a): recover the streamed reply text instead of resolving empty
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).

_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".

(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
2026-09-20 18:45:42 -07:00
Forkbert
b113ab6de6 fix(photon): use GUID for liveness probe message id
(cherry picked from commit cdb6ccadb045e599db795823a100ae7b61d377ec)
2026-09-20 18:23:28 -07:00
teknium1
fd94fe9b93 fix(gateway): every adapter send_voice accepts the dispatch's is_voice kwarg
The base media dispatch passes is_voice= to send_voice; line, mattermost and
weixin had explicit signatures without it, so a non-image MEDIA attachment
routed as audio raised TypeError and was dropped — the same class as the
Matrix report (#102221, #116776). Adds a repo-wide signature invariant test.
2026-09-20 15:48:51 -07:00
xiaodu
b6505448f1 fix(matrix): accept the media dispatch's is_voice kwarg in send_voice
The base media dispatch calls send_voice(..., is_voice=is_voice) for every
audio MEDIA attachment (gateway/platforms/base.py _send_one). MatrixAdapter
.send_voice() accepted neither is_voice nor **kwargs, so every non-image
MEDIA delivery raised TypeError and the file was silently dropped — the
failure is visible in rotated logs since 2026-09-08 (never worked).

Accept the flag explicitly: is_voice=False -> plain m.audio in the original
format (no transcode); True or omitted (play_audio legacy callers) -> the
existing MSC3245 voice-bubble path with best-effort Ogg/Opus transcode.

Fixes #116776

(cherry picked from commit d4f89a725498e29a2ee0fe016f8bb98db08f358c)
2026-09-20 15:48:51 -07:00
fangliquan
45a701d8b0 fix(simplex): flatten alpha image thumbnails
(cherry picked from commit 97886801741f4367aadfe83ccb25b10cedb21149)
2026-09-20 15:48:01 -07:00
liuhao1024
7a23b00101 fix(whatsapp): refuse legacy-pidfile kills without a start-time fingerprint
The stale-bridge cleanup accepted a "node" + session-path cmdline substring
as kill evidence for legacy pidfiles (pid line only). A log tail, editor, or
grep that merely mentions the session path matches that same substring, so
the cleanup could SIGTERM a stranger process.

Require the kernel start-time fingerprint and fail closed when the pidfile
lacks one; the bridge-port scan (which verifies a node-executable listener)
reaps the orphan instead. The refusal reason in the warning now distinguishes
a fingerprint-less legacy pidfile from a recycled PID.

Flip the legacy-pidfile regression test to assert the fail-closed outcome.

Fixes #116883
2026-09-20 15:24:51 -07:00
teknium1
a57a5f90d2 fix: slot-busy skipped Telegram edits no longer count as shown text (review follow-up)
An interim edit skipped because the chat's shared send+edit slot was busy
returned a plain SendResult(success=True), so the stream consumer recorded
the never-shown text as _last_sent_text and reset _flood_strikes. A later
turn-final flood then saw _visible_prefix() == final text and either marked
the turn delivered or entered fallback with an empty continuation — the user
never saw the tail. The adapter now flags the skip in
raw_response={"skipped": True} and _edit_existing leaves the visible prefix
and flood state untouched, so the next tick retries and a flood fallback
re-sends exactly the unseen tail.
2026-09-20 12:17:23 -07:00
Haisam Abbas
fb2f3e213a fix(telegram): pace sends and edits against a shared per-chat slot
Telegram counts an editMessageText against the same per-chat allowance as a
sendMessage, but streaming previews paced only edits (DEFAULT_STREAMING_EDIT
_INTERVAL = 0.8s = 1.25 msg/s into one chat before any reply was sent) —
83% of measured flood penalties. One shared slot per chat: a send WAITS for
its slot (skipping would drop a message), an interim edit is SKIPPED (the
next tick shows the same text anyway), and the final edit is never gated
(the answer is never withheld). A per-adapter tuning knob keeps the slot
available to tests that model instantaneous bursts.
2026-09-20 12:17:23 -07:00
teknium1
1d6fbd0d1f fix: explicit empty backfill channel list disables the Discord scan (review follow-up)
`missed_message_backfill.channels: []` (YAML list or JSON-list string) returned
an empty set on main, and the backfill logged "no channels configured" and
skipped. The gate-CSV refactor treated an empty result as "unset" and fell
through to the allowed-union-free-response default, scanning channels the
operator had explicitly disabled. An explicit list is now authoritative even
when empty; only the default "" string still falls through to env/default.
2026-09-20 12:10:27 -07:00
teknium1
587bb10575 fix(platforms): Discord/WhatsApp/DingTalk gates honour allowlists stored as a JSON-list string
`hermes config set KEY '["-100","-200"]'` used to write the literal as one quoted
YAML string; the writer now emits a real list (acaac9a18, #88163) and the Telegram
gate decodes the legacy string shape (122ad719, #110213). The Discord, WhatsApp and
DingTalk gate parsers still comma-split that string into `{'["-100"', '"-200"]'}`,
so a config written before the writer fix silently locks every allowlisted chat,
channel or user out — with no warning.

Route every remaining comma-split gate through the shared
`gateway/platforms/_shared.py::decode_json_list_literal`:

- Discord `_gate_csv_set` (allowed/ignored/no-thread channels, allowed users/roles),
  `_discord_free_response_channels` and `_missed_message_backfill_channels` now share
  the one parser instead of three hand-rolled splits.
- WhatsApp `_coerce_allow_list` (allow_from, group_allow_from, free_response_chats).
- DingTalk `_csv_set` (allowed_users, allowed_chats, free_response_chats).

Plain CSV strings, YAML lists and malformed JSON keep their previous meaning.
2026-09-20 12:10:27 -07:00
Steve Hsu
97051e4194 fix(telegram): attach real video geometry and a thumbnail so large uploads aren't square
`sendVideo` gives back `width=320 height=320 duration=0` with no thumbnail once an
upload is large enough that Telegram skips its own video processing, and clients
then draw the message as a square tile — for portrait reels and 16:9 clips alike,
even though the delivered file itself is correct.

Measured on one 6 s 2560x1440 clip, inspecting the Bot API response: 4.9 MB and
9.8 MB keep `2560x1440` / duration 7 / a 320x180 thumbnail; 14.8 MB, 19.4 MB and
23.0 MB degrade to the square placeholder; the same 23.0 MB file sent with
`width`/`height`/`duration` plus a JPEG `thumbnail` comes back `2560x1440` with a
320x180 thumbnail.

Probe the local file with ffprobe and attach a 320px-wide JPEG frame from it, on
both Telegram send paths: the gateway adapter's `send_video` and the standalone
`hermes send` media sender. Both helpers return nothing when ffmpeg/ffprobe is
unavailable, which keeps the previous behaviour for hosts without them.
2026-09-20 11:20:26 -07:00
teknium1
c060a821d5 fix(discord): accept a pasted channel link at the _resolve_channel chokepoint too
HomeChannel normalization covers env and YAML homes, but cron
`deliver: discord:<link>` targets and thread metadata reach the adapter as
written and still died in int(). Share one helper
(gateway.config.discord_channel_id_from_link) between HomeChannel and the
resolver every outbound Discord target passes through; message links and
non-link strings keep their existing error path. Tests trimmed to two
invariants: config load (env + YAML) and the resolver entry point.
2026-09-20 10:49:49 -07:00
chelsealong
f45ed8e4ec docs(matrix): document the <=2-member DM auto-classification and its config bypass
_resolve_room_identity() classifies any room with <=2 joined members as a DM
regardless of m.direct or an explicit room name, so those rooms silently
bypass MATRIX_ALLOWED_ROOMS, MATRIX_FREE_RESPONSE_ROOMS, and
MATRIX_REQUIRE_MENTION, and use DM threading instead of
MATRIX_AUTO_THREAD/MATRIX_SESSION_SCOPE. This was previously only visible in
an inline code comment, not in the env-var docs an operator would read.

Fixes #114733
2026-09-20 10:42:25 -07:00
chelsealong
0c872b611c fix(telegram): point dm_topics prerequisite text at BotFather Threaded Mode
The dm_topics "not a forum" warning and the matching docs section told
users to tap the bot's name in the DM and toggle "Topics" in chat
settings. That toggle only exists for group forums; a bot DM has no
such control. The actual prerequisite is Threaded Mode, enabled by the
bot owner via the BotFather Mini App (Bot Settings -> Threads
Settings) -- already documented correctly a few sections later in the
same file, under the /topic prerequisites.

Fixes #115019
2026-09-20 10:39:37 -07:00
KeyArgo
e7509fa1a4 fix(telegram): observe sibling-bot wake-word messages dropped by the bot-to-bot gate
With `require_mention: true` + `bots_require_mention: true` + wake words in `mention_patterns`, a message
authored by another Hermes bot that addressed this bot by name matched no dispatch path (the
bot-to-bot loop breaker skips it) and was also refused by the observe gate (`mention_patterns`
matches are assumed dispatched), so it vanished from both paths with no log line.

Factor the loop-breaker predicate into `_bot_sender_suppressed` and consult it in
`_should_observe_unmentioned_group_message`, so every message the dispatcher drops because of
`bots_require_mention` is kept as observed context instead of being lost (#115119).
2026-09-20 10:30:11 -07:00
fangliquan
3cbd3cce0a fix(whatsapp): honor configured reply prefix in bridge replies
whatsapp.reply_prefix from config.yaml was written into the bridge env and then
popped again by the WHATSAPP_* passthrough loop (the key was in
_BRIDGE_PASSTHROUGH_ENV and the scoped env lookup came back empty), so bridge.js
always fell back to its built-in header and the documented reply_prefix: ""
could not disable it. Resolve the prefix once (scoped env first, then the
adapter value) and keep it out of the passthrough loop.

Fixes #116059
2026-09-20 10:20:40 -07:00
hardwork9047
a0e7fc4a9e fix(gateway): consult _in_bot_thread in the Discord admission gate
`_discord_message_admission()` drops a message that mentions someone other than
the bot when `DISCORD_IGNORE_NO_MENTION` is on (the default) and the channel is
not free-response. It did so without consulting `_in_bot_thread()`, unlike the
other two ingress paths — `_dispatch_recovered_message()` (adapter.py:2292) and
`_handle_message()` (adapter.py:5946). Admission runs on both and returns
`False` unconditionally, so it overrode the thread exemption they grant.

The result was an asymmetry with no obvious cause from the outside: in a thread
the bot had joined, a message with no mention at all was admitted (an empty
`message.mentions` skips the enclosing block), while the same message with one
mention of a third party was dropped. A thread the bot is a participant in is
the one place "addressed to someone else" is least likely to hold.

`thread_require_mention` still gates multi-bot threads, since that check lives
inside `_in_bot_thread()`.

The drop also emitted nothing at any log level, leaving `gateway.log` identical
whether the gate fired or the event never arrived; add a debug line so the two
can be told apart.

Fixes #116568
2026-09-20 10:20:03 -07:00
kshitijk4poor
de083436bd refactor(gateway): drop Telegram's shadowing _coerce_float_extra; module-level math import
Gate review: TelegramAdapter kept a same-named override with a weaker contract (no finite
guard, negatives clamped instead of reset), so the hierarchy had two parsers under one name.
Both Telegram callers pass explicit bounds; the base method is a strict superset for them.
The cadence test now pins the behaviour (≤ 0.5 s / ≤ 1.0 s) instead of echoing the constant.
2026-09-20 18:14:58 +05:30
kshitijk4poor
530fdb5867 refactor(gateway): one text-batch cadence and one extra-float parser on BasePlatformAdapter
Gate review: the fix left three adapter-private copies of `_coerce_float_extra` and the
0.3/2.0/1.0/4.0 cadence literals in three files. The parser and the cadence constants now
live on BasePlatformAdapter beside the delay attrs they configure; WhatsApp and Weixin call
`_configure_text_batch_delays()`, Telegram reads the same constants through its env helper.
The clamp test is parametrized over both adapters and the Weixin docs name the ceilings.
2026-09-20 18:14:58 +05:30
kshitijk4poor
130b596f6f fix(whatsapp,weixin): text-batch delays default to Telegram cadence
WhatsApp debounced text for 5s (10s near a split) and Weixin for 3s/5s
before dispatching, so every reply paid multiple seconds of idle latency
that Telegram never pays (0.3s/1.0s). Default both adapters to Telegram's
cadence and mirror its ceilings (2.0s / 4.0s, split >= base delay) via the
existing _coerce_float_extra seam. The config keys are unchanged; 0 still
dispatches immediately. Docs updated.

Spotted via #44896 (@liuhao1024). Fixes #44883, refs #25056.
2026-09-20 18:14:58 +05:30
kshitijk4poor
18f1706137 refactor(matrix): drop dead is_direct guard and unreachable sender warning
_schedule_invite_join already requires both is_direct and inviter before recording m.direct, so `is_direct and bool(inviter)` at the reconcile call site was redundant; pass is_direct through. The `if is_direct and not inviter` WARNING could only be reached with GATEWAY_ALLOW_ALL_USERS set and a spec-violating stripped m.room.member event lacking `sender`; the info log already prints is_direct, so drop the branch.
2026-09-20 18:00:06 +05:30
kshitijk4poor
b98dd99142 docs(matrix): reconcile gate comment states the real premise
The adapter comment and the test module docstring claimed a reconciled pending invite "never fires _on_invite". It does: _absorb_sync runs _dispatch_sync (which emits INVITE to _on_invite) and then the reconcile pass over rooms.invite, which joined every entry unconditionally — so a live invite _on_invite rejected was joined ms later, and invites that arrived while the gateway was down were joined on restart with no gate. Reword both to state that premise.
2026-09-20 18:00:06 +05:30
Iain Lane
439f2b1cf6 fix(matrix): apply the inviter allowlist to reconciled pending invites
_on_invite only auto-joins a room when the inviter is allow-listed (or
GATEWAY_ALLOW_ALL_USERS is set), so a live invite from an arbitrary
federated user is rejected. A pending invite that arrives while the
gateway is down takes a different path: _schedule_pending_invite_joins
reconciles it from rooms.invite in the sync response and scheduled the
join unconditionally. An unauthorized invite sent during downtime was
therefore auto-joined on restart, bypassing the allowlist.

Extract the gate from _on_invite into _is_authorized_inviter and apply
it during reconciliation too, reading the inviter from the stripped
invite state (the sender of the m.room.member event for our own user,
as _extract_invite_dm_signal already does for the DM signal). An
inviter that cannot be read from the invite state fails closed, exactly
like an empty sender in _on_invite: the invite is skipped with a
warning and left pending.
2026-09-20 18:00:06 +05:30
Iain Lane
0ce3e7b12b fix(matrix): record reconciled direct invites in m.direct
A direct invite that arrives while the gateway is running fires
_on_invite, which passes is_direct and the inviter through
_schedule_invite_join so the room is recorded in m.direct after the
join. An invite that is still pending across a gateway restart takes a
different path: _schedule_pending_invite_joins reconciles it from
rooms.invite in the sync response, but called _schedule_invite_join
without is_direct or inviter. The DM signal was dropped, the room was
never recorded in m.direct, and it was classified as a group until the
user's own client happened to update m.direct.

Read the signal from the stripped invite state instead: the
m.room.member event for our own user carries the original invite's
is_direct flag, and its sender is the inviter. Thread both through to
_schedule_invite_join so a reconciled direct invite is recorded in
m.direct, and thus lands in _dm_rooms, exactly like a live one.

This gap was surfaced by the triage of #62493.
2026-09-20 18:00:06 +05:30
kshitijk4poor
536ce859c4 fix(telegram): bounded sends do not arm the blocked-loop watchdog 2026-09-20 17:57:43 +05:30