Commit Graph

69 Commits

Author SHA1 Message Date
teknium1
7c5296ce1c fix(buzz): a clean relay close backs off and publishes retrying like any other disconnect
A relay that accepted, authenticated and subscribed and then closed cleanly
made the read loop return without raising, so _websocket_loop reconnected
immediately with no backoff and never flipped health to "retrying". The read
loop now raises ConnectionError on StopAsyncIteration so the clean close takes
the same backoff + degraded path as an idle or send-side disconnect.
2026-09-15 18:59:25 -07:00
teknium1
f529986abf fix(buzz): a dead socket ends the WebSocket connection from either side and publishes retrying
Follow-up to the cherry-picked watchdog from #112052 (@KoNit-K), finishing the
class the reporter of #112049 laid out:

- `_websocket_loop` runs the read loop and the discovery sweep as sibling
  tasks and ends the connection when EITHER finishes. The discovery sweep
  re-raises `ConnectionClosed` instead of logging it and retrying next tick:
  a send that sees the socket closed is proof the read the loop is parked on
  will never return. That is exactly the traceback the reporter watched for
  22-86 h while inbound stayed silent.
- Health is invalidated while reconnecting: the first disconnect publishes
  `retrying` (`_mark_degraded`) and a successful re-subscribe publishes
  `connected` again. Before, `connect()` wrote "connected" once and nothing
  ever changed it, so `/health/detailed` claimed delivery during the silence.
- The teardown awaits both tasks with `gather(return_exceptions=True)` instead
  of a bare `except (CancelledError, Exception): pass`, which could swallow a
  `disconnect()` cancellation landing mid-teardown.
- Slims the salvaged read-loop hunk: the extra "receive task remained parked"
  warning and the in-loop `_mark_degraded()` are dropped; the reconnect log
  line and the loop-level health flip cover both.

Docs: the Buzz page still described inbound as poll-only and the WebSocket
transport as a future optimization; it now describes the watchdog and the
`retrying` health state.
2026-09-15 18:59:25 -07:00
KoNit-K
8f1e0ec990 fix(buzz): recover from uninterruptible websocket reads 2026-09-15 18:59:25 -07:00
teknium1
de114b3af1 refactor(platforms): one scoped-secret reader and spec-driven enablement/YAML-bridge boilerplate across all adapters
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:

- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
  docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
  `external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
  Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
  keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
  (SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
  bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
  slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
  `_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.

Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:

- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
  irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
  matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
  nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
  `yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
  mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.

Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
  neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
  seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
  env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
  does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
  all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
  "Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
  reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
  as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
  value (was list/tuple/set only) — same env text for every real YAML shape.

Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.

Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
2026-09-13 05:32:38 -07:00
teknium1
e73f94fa83 refactor(platforms): every plugin setup wizard uses the shared declines_reconfigure gate
Fourteen platform plugins hand-rolled the "already configured? Reconfigure? [y/N]"
gate at the top of interactive_setup (env check + info line + prompt_yes_no(..., False)),
with drifting wording ("X: already configured" vs "X is already configured." vs
"already enabled") and, for LINE and SimpleX, raw input() loops with their own
EOF/KeyboardInterrupt handling and no gate at all. Fixes to the gate (non-interactive
handling, wording, default) therefore reached only the core Telegram/BlueBubbles/webhook
wizards.

- hermes_cli/setup_platforms.py: `_declines_reconfigure` becomes the public
  `declines_reconfigure(label, question, *env_vars)` (any-of env check, so Matrix's
  token-or-password gate fits); `_save_prompted` becomes `save_prompted` alongside it.
  No alias kept; the three core callers are updated.
- buzz, dingtalk, discord, feishu, google_chat, irc, matrix, mattermost, raft, slack,
  teams, wecom: the hand-rolled gate is replaced by one `declines_reconfigure(...)` call;
  post-decline extras (Discord allowlist nudge, Slack manifest refresh, Raft "Keeping"
  line) stay local and unchanged.
- line, simplex: the raw input() loops move onto hermes_cli.cli_output.prompt (masked
  for secrets, "" on Ctrl-C/EOF) and gain the shared gate on their primary env var.

Behavior change: the gate's info line is now uniformly "<Label>: already configured"
(DingTalk/Feishu/WeCom lose the trailing period + inline ID; Buzz/IRC/Google Chat/Raft/
Teams no longer echo the current value in that line). Feishu and WeCom now gate on the
app/bot ID alone instead of ID AND secret. LINE and SimpleX gain a "Reconfigure?" [y/N]
prompt when already configured; their prompts now honour HERMES_NONINTERACTIVE and print
via the CLI helpers instead of bare print(). Prompt defaults (No) are unchanged everywhere.

Not touched: WhatsApp's gate keys on WHATSAPP_ENABLED being truthy (a "false" value must
not count as configured), which the shared any-set gate cannot express — left hand-rolled.

Test: tests/plugins/platforms/test_interactive_setup_reconfigure_gate.py parametrized over
the 14 wizards — with the primary env var set and the user declining, each wizard must have
called declines_reconfigure with that var and returned without prompting or saving.
Sabotage: reverting mattermost's gate fails that row.
2026-09-13 05:32:38 -07:00
teknium1
9b1990583d refactor(gateway): adapters share helpers.cancel_task / MessageDeduplicator / bounded_put
Six adapters defined their own `_cancel_task` and nine more inlined the same
cancel + suppress(CancelledError) + await block; five kept a hand-rolled TTL-dict
`_is_duplicate` next to the existing `helpers.MessageDeduplicator`; three carried a
`_bounded_put`. Each copy fixed the same bugs on its own schedule (self-cancel deadlock,
done-task re-await, eviction under load).

- `helpers.cancel_task`: None/done no-op, never awaits the current task, swallows the
  task's own exception at teardown. Replaces qqbot/signal/yuanbao/buzz/photon/simplex
  definitions and the inline copies in weixin, discord, email, irc, line, mattermost,
  whatsapp and telegram.
- `helpers.MessageDeduplicator` replaces `_is_duplicate` in qqbot, ntfy, photon,
  wecom_callback and LINE's `_MessageDeduplicator`; every site keeps its own
  max_size/TTL (qqbot and ntfy 1000/300s, photon 4000/48h, wecom_callback 2000/300s,
  LINE 1000/no TTL).
- `helpers.bounded_put` replaces photon/wecom/whatsapp_cloud copies; a re-put now
  refreshes the key to the newest slot at every site.
- telegram gmail-triage scripts resolve under `get_hermes_home()` instead of a hard
  `~/.hermes`, so profiles with HERMES_HOME set find them.

Not changed: `get_chat_info` stays `@abstractmethod` because
tests/gateway/test_relay_capability_surface.py locks the abstract set to exactly
{connect, disconnect, send, get_chat_info} as a cross-repo contract, so the ~17 no-op
overrides remain.

Behavior change: whatsapp_cloud `_bounded_put` was a pure FIFO (no refresh on re-put);
it now refreshes like the other two sites. Task cancellation at the migrated sites
swallows a task's terminal exception where a few copies previously only suppressed
CancelledError (all are shutdown/disconnect paths).
2026-09-13 05:32:38 -07:00
teknium1
73eadd54f5 fix(platforms): standalone senders return redacted error envelopes
20 `plugins/platforms/*/adapter.py::_standalone_send` paths (the out-of-process cron /
send_message delivery) built `{"error": f"... {e}"}` by hand — 83 literals. The exception text
of an httpx/aiohttp failure can carry the Authorization header, a signed URL or a response body
with the token in it, and that string became the tool result the model reads. Only sms went
through the redacting `tools.send_message_senders._error`; discord kept a private regex that
only knew `Authorization: Bot`.

`gateway.platforms._shared.send_error(message)` wraps that helper (agent.redact +
URL-secret scrub) and every standalone literal now goes through it, including the three
envelopes that carry extra keys (discord warnings, photon error_class/retryable, whatsapp's
`(None, err)` tuple). The sms and discord local wrappers are deleted. Telegram already
delegated to the core sender and is untouched.

Behavior change (security): vendor exception text in standalone-send failures is redacted
before reaching the model.
2026-09-13 05:21:39 -07:00
teknium1
74a315bd32 fix(platforms): irc/line/buzz identity-lock conflicts actually fire
gateway.status.acquire_scoped_lock returns (acquired, existing_record). The irc, line and
buzz adapters tested `if not acquire_scoped_lock(...)`, and a non-empty tuple is always
truthy, so two profiles could drive one IRC nick / LINE channel / Buzz identity in
parallel. Route the three through BasePlatformAdapter._acquire_platform_lock (the seam
the other 8 adapters use), which unpacks the tuple, names the owning profile + PID in the
fatal error and honours the `--replace` takeover. Release goes through
_release_platform_lock; the private _lock_key bookkeeping is gone.

Error code changes from `lock_conflict` to `{scope}_lock` — both families are already
matched by gateway.restart.is_global_startup_conflict.

The buzz test mocked acquire_scoped_lock as a bare False, which masked the bug; it now
returns the real (False, record) contract, and irc/line gain the same conflict test.
2026-09-13 05:21:39 -07:00
teknium1
719cb67bdb fix(platforms): carry the inbound message id into the session source everywhere
Same bug class as the Slack/Feishu picks: buzz, dingtalk, email, google_chat,
line, ntfy, photon, sms, teams, wecom and whatsapp already had the platform
message id on the MessageEvent but built the SessionSource without it, so
source.message_id consumers (reply anchor in run.py, /sethome synthetic-thread
check, relay _event_ids fallback, shutdown notice anchor) saw None. Only sites
where the id variable was already in scope are widened.
2026-09-12 08:26:39 -07:00
Teknium
199c66f70f fix(gateway): adapter settings resolve per profile under multiplex, not from the default's env
Under gateway.multiplex_profiles a served secondary profile's adapter is built and
connected inside _profile_runtime_scope while os.environ still holds the DEFAULT
profile's .env. Credentials and allowlists were already read through the profile
scope (get_scoped_secret / _platform_gate_env), but the non-credential SETTINGS the
adapters read with bare os.getenv were not, so a served profile silently ran with the
default profile's values: webhook listener host/port/URL (SMS, Teams, LINE, Feishu,
BlueBubbles), Signal's connect URL/account gate, mention gating and reactions (Slack,
Matrix, Signal, Feishu, BlueBubbles, Discord), Matrix thread/session/E2EE policy and
message-length limits, Discord backfill/command-sync/attachment caps, Buzz reply mode
and env enablement seed, A2A agent name/port/description/toolsets, and the
api_server model alias.

Every such read now goes through the existing scoped reader (get_scoped_secret, or the
adapter's own scope-aware helper): under a secondary's scope the profile's own .env is
authoritative and a miss yields the default -- never another profile's value; the
default profile and single-profile gateways keep reading os.environ exactly as before.
Buzz and A2A previously short-circuited to "extra only / built-in default" under a
scope, which also dropped the profile's OWN .env; they now read the scope so a served
profile matches its standalone gateway.

The parity harness (temp HERMES_HOME, default + 2 secondaries with distinct values for
every env var each adapter reads, real load_gateway_config + adapter factory in both
topologies) went from 70 raw process-env bypass sites across 14 adapters to only the
HERMES_<PLATFORM>_* perf knobs and the api_server listener vars, which are process-
global by design (agent.secret_scope._GLOBAL_ENV_*).
2026-09-11 19:37:59 -07:00
kshitijk4poor
ab2f4602de refactor: MessageEvent to gateway/platforms/event.py; ElicitationHandler takes a call_context thunk
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.

gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).

tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.

ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)

Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
2026-09-07 22:47:33 +05:30
Teknium
e83816a4d1 review-fix(comments): restore lost #NNNN rationale comments across non-test source (mechanical sweep, condensed, code unchanged)
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
2026-09-03 09:44:26 -07:00
Teknium
0071ba9965 Merge origin/main (561b053f79) into simp/forwardport: forward-port 220 main commits into the simplified tree 2026-09-03 03:31:03 -07:00
Teknium
192058fda4 refactor(platforms): a2a/buzz/dingtalk/email/google_chat/feishu-aux/discord-aux 11440->8812; dead code, dispatch tables, unified helpers 2026-09-02 23:34:38 -07:00
Teknium
8f1ef66b1f refactor(adapters/p2p_group): 12752->8794; buzz/photon/a2a/raft dedupe (sidecar paths unified in photon package, JSON-RPC helpers, Nostr event builders), dead a2a security getters + stream_message removed 2026-09-02 14:06:36 -07:00
kshitijk4poor
e0567770f3 fix(gateway): off-loop review follow-ups for #101603 / #101605
Findings from the efficiency review pass on the two salvages, applied as one
small follow-up:

- copilot_auth: check the negative cache BEFORE taking the per-fingerprint
  exchange lock. During the 60 s post-failure window, dashboard polls now
  raise immediately instead of parking an executor thread behind the
  in-flight holder (up to ~50 s) to learn the same answer. Test hangs
  without the check (timeout 124), passes with it.
- buzz _localize_inbound_media: download_path.read_bytes() was still
  evaluated on the loop as the argument to the offloaded cache call — up to
  the 128 MiB inbound cap. Read off the loop too.
- test_list_credential_pool_keeps_loop_responsive: 0.5 s block / 0.25 s
  threshold (2x margin) so runner descheduling cannot false-fail it while a
  real regression still trips it.
2026-09-03 02:07:05 +05:30
kshitijk4poor
568b16122d fix(gateway): finish the media-cache offload sweep (teams NameError, buzz, photon)
The previous commit removed cache_media_bytes from teams/adapter.py's import
block but left one bare call at the Bot-Framework image path — a NameError
that the surrounding 'except Exception' swallowed, so every BF image
attachment was silently dropped (ruff F821). Route it through
cache_media_bytes_async like the file's two other sites.

Same bug class at the sites the sweep did not reach:
- buzz _download_attachment / _localize_inbound_media: cache_media_bytes
  called from async methods → cache_media_bytes_async.
- photon _dispatch_inbound: _normalize_binary_payload (base64 decode +
  cache write of possibly multi-MB payloads) ran on the loop → closure made
  async, helper offloaded via asyncio.to_thread.

Tests: cover cache_media_bytes_async (thread-id + kwarg forwarding);
teams/buzz tests re-pointed at the async seam.
2026-09-03 01:58:56 +05:30
Teknium
c816957a43 fix(buzz): reconcile media pipeline with landed dispatch + threading contracts
Post-rebase composition over #99431/#99429/#99427: file-attachment sends
route through _run_message_send so the mention-recovery ladder covers
media captions; _send_file_attachment/_send_local_file honor the
resolved thread-root anchor and reply_to_mode opt-out; send() records
event_meta on the verified receipt id (#75826); test fakes gain the
auth_tag kwarg and accepted-receipt shape.
2026-08-31 10:06:34 -07:00
Teknium
c731500cf6 fix(buzz): route shared attachment sender through redacted receipt errors
Follow-up reconciliation: _send_file_attachment (the merged #95688/#74999
helper) now uses #78046's strict _parse_send_receipt contract and
redact_path error bounding, so CLI failures never leak host filesystem
paths and zero-exit unverified receipts are rejected on every outbound
media path.
2026-08-31 10:06:34 -07:00
EmpireOperating
a55d66b79a fix(buzz): redact media paths before bounding errors 2026-08-31 10:06:34 -07:00
EmpireOperating
9c25704257 fix(buzz): verify live media delivery receipts 2026-08-31 10:06:34 -07:00
EmpireOperating
37c943997b fix(buzz): support media in standalone sends 2026-08-31 10:06:34 -07:00
Teknium
4c496d3cf6 fix(buzz): reconcile probe-race contract with shared file-attachment sender
#95688's _send_file_attachment refactor re-probed file existence, which
#74999's tests prove can race into a false 'not found' when the file
disappears between the caller's check and the helper's. Callers that
already verified the file pass probe=False; unverified document/video/
voice callers keep the guard.
2026-08-31 10:06:34 -07:00
phil baker
fafc3ddce5 fix: deliver Buzz media as native attachments 2026-08-31 10:06:34 -07:00
Riyaaz
ebda1fbf8e fix(buzz): deliver local images through native upload 2026-08-31 10:06:34 -07:00
Teknium
73071d852c fix(buzz): merge URL-localization and imeta attachment paths in dispatch
Reconciles #84113 (authenticated same-relay URL localization) with #78051
(native imeta ingestion): _dispatch_message now merges caller-provided
verified imeta attachments with text-localized relay media instead of
clobbering them, dedupes paths, and downgrades mixed-source media to
DOCUMENT semantics so audio members are not routed through STT.
2026-08-31 10:06:34 -07:00
EmpireOperating
d15cbbcbf3 fix(buzz): gate inbound attachment side effects 2026-08-31 10:06:34 -07:00
EmpireOperating
00394acfae fix(buzz): ingest verified native attachments 2026-08-31 10:06:34 -07:00
Mathias Gorf
aaad054330 fix(buzz): gate authenticated inbound media on explicit authorization
Localizing inbound relay media spends the agent's own Buzz credentials on
a URL chosen by the sender, so it must not run on the strength of the
adapter's local allow-list alone. Require the gateway's authorization
callback to return an explicit True before any `buzz media get` runs; a
denial, a missing callback, or a raising callback fails closed and leaves
the message text exactly as it arrived.

`_is_sender_authorized` previously wrapped the callback result in
`bool()`, so a truthy non-boolean (a status string, a sentinel) would
satisfy an `is True` gate's intent while bypassing its guarantee. Only
the literal booleans now propagate; anything else is "unknown", which the
existing Slack and Discord callers already treat as trust-unknown.

Reviewers asked for this boundary on the sibling inbound-media PRs
(#77734, #78051); it applies equally to the retrieval path in #75614,
which this change builds on.
2026-08-31 10:06:34 -07:00
joelbrilliant
55136adcc4 fix(buzz): preserve inbound media captions 2026-08-31 10:06:34 -07:00
joelbrilliant
bce94cc1b9 fix(buzz): localize inbound relay media 2026-08-31 10:06:34 -07:00
Teknium
c84c6e2382 fix(buzz): open fresh WS subscriptions from the beginning and discover conversations on a timer (#78429, #93557, #75107)
Three sibling gaps in the WebSocket transport's conversation lifecycle:

- #78429: _send_channel_subscription defaulted a zero last_ts to
  'since ~ now', so the message that CREATED a new conversation (created_at
  fractionally before the subscription) was never delivered. A channel with
  no high-water mark now subscribes from the beginning with
  limit=_FETCH_LIMIT instead; seeded channels still resume from last_ts-1.

- #93557: relays do not guarantee a kind-44100 membership event per new
  conversation, so WS-transport deployments never discovered DMs opened
  mid-session until a reconnect. The WS loop now runs the same
  _discover_dms sweep the poll transport uses, on the same cadence
  (poll_interval * _DM_DISCOVERY_EVERY), via a companion task that is
  cancelled with the connection.

- #75107: _discover_dms only ever adopted DM-shaped conversations, so a
  real community channel the agent joined mid-run was never subscribed
  until restart. In watch-all mode (no explicit channels list) newly
  listed real channels are now adopted and seeded from their newest events
  (history predating the join is not replayed). Explicit watch lists stay
  authoritative.
2026-08-31 09:49:30 -07:00
Teknium
b907b7eb85 fix(buzz): compose #97502's membership-rejection matching into the per-subscription CLOSED handler
- Widen the permanent-rejection match to the exact relay phrasings seen in
  production (#97502): 'not a channel member' and 'auth-required', alongside
  'restricted'.
- Close the re-adoption hole called out in review: _discover_dms() (both the
  dms-list path and the channels-list fallback) now skips channels in
  _restricted_channels, so a restricted channel dropped at runtime cannot be
  silently re-added by the next discovery sweep and re-trigger the rejection.
- Credit: runtime CLOSED matching terms from PR #97502 by @repfigit; the
  per-subscription drop + restricted set is PR #76850 by @xozai.
2026-08-31 09:49:30 -07:00
José Leos
d3730a3fa9 fix(buzz): handle restricted CLOSED per-subscription, stop reconnect flood
When a Buzz relay sends a CLOSED frame for a single subscription with a
'restricted: not a channel member' error, the adapter was raising
ConnectionError, tearing down the entire WebSocket connection, and
immediately reconnecting — causing a ~1.6 s flood in gateway.log.

Root cause: the CLOSED handler unconditionally raised ConnectionError
regardless of whether the error was permanent (restricted) or transient
(e.g. server shutdown).

Fix:
- On a 'restricted' CLOSED, drop only the offending subscription and
  record the channel in a new _restricted_channels set instead of
  tearing down the whole connection.
- Skip restricted channels during connect() seeding and
  _subscribe_websocket() so reconnects don't re-trigger the same error.
- Non-restricted CLOSED frames still raise ConnectionError and reconnect
  as before.

Adds three regression tests:
- test_websocket_loop_drops_restricted_channel_without_reconnect
- test_websocket_loop_reconnects_on_non_restricted_closed
- test_restricted_channels_skipped_during_subscribe

Tested on macOS against buzz.xozai.com: gateway.log shows zero
'restricted' errors and stable 'watching N channel(s) via websocket'
after the fix.
2026-08-31 09:49:30 -07:00
Alessandro Boni
d36827b086 fix(buzz): resume watched channels from a durable cursor across restarts (#90464)
`connect()` calls `_seed_channel()` unconditionally, and seeding marks every
event currently in the channel as seen so a start never replays history at the
agent. A message that arrives after the process starts but before the seed
completes — or at any point while the gateway is down — sits in exactly that
history, so the seed swallows it permanently even though the Buzz relay still
has it. The `seen` set and `last_ts` lived only in memory, so there was nothing
to distinguish "already handled" from "never seen".

Each watched channel's cursor (`chat_type`, `last_ts`, and the bounded `seen`
id list) is now persisted under `HERMES_HOME/buzz/channel-cursors.json` and
restored at connect. Where a cursor exists the channel resumes from it and the
history fetch is skipped entirely; where none exists the old seed-from-history
behaviour is unchanged, so a first-ever run still never replays a backlog.

Details worth noting:

- The file records the identity and relay it was written for. A cursor from a
  different bot or relay is ignored rather than trusted — the channel ids
  would collide while the event stream behind them is a different one.
- Any read or parse failure leaves the cursors empty, which degrades to
  seeding instead of failing the connect. Writes go through
  `utils.atomic_json_write` (temp + fsync + replace), so a crash mid-write
  cannot leave a truncated cursor behind.
- The restored `seen` list is trimmed to `_SEEN_CAP` on load, keeping the
  newest ids, so a hand-edited or legacy file cannot grow the de-dupe set
  without bound.
- Saves are gated on the cursor actually moving, so an idle channel does not
  rewrite the file every poll interval. Both inbound transports are covered:
  the poll sweep and the WebSocket event path share the same check.

Tests: six new cases in `TestChannelCursorPersistence` — the cursor is written
on seed, a restart resumes without spending a CLI call on history and then
delivers the mention that landed while the gateway was down, a foreign
identity or relay is ignored, a corrupt file falls back to seeding, the
restored `seen` set stays bounded, and an idle poll leaves the file untouched.
All six fail on main.

Tested on: Windows 11, Python 3.12. `python -m pytest
tests/gateway/test_buzz_adapter.py tests/gateway/test_buzz_websocket.py -q` —
33 passed (23 pre-existing + 6 new here, plus 4 WebSocket). Requires
`pytest-asyncio` (pinned at 1.3.0 in pyproject) — without it the async cases
in this file error out as unknown marks.
2026-08-31 09:49:30 -07:00
liuhao1024
94d86fa4de fix(buzz): bound WebSocket read idle time to force reconnect on silent relays
A relay-side close the transport never surfaces (observed as a CLOSE_WAIT
socket behind Cloudflare, #98097) parks the read loop forever while the
gateway keeps reporting connected: inbound stops, gateway_state.json stays
healthy, and only a restart recovers. The library keepalive should catch
this first, but as a last resort the read side now waits at most
_WS_READ_IDLE_TIMEOUT (300s) for a frame before raising into the existing
reconnect path, which re-authenticates and re-subscribes with per-channel
since filters intact.

Fixes #98097
2026-08-31 09:49:30 -07:00
Matt Lutz
a74080437a fix(buzz): resolve @mentions to member pubkeys so agent-to-agent pings work
Salvaged from PR #83414 (4 commits squashed to final state) and composed
with the presentation-mention escape retry from PR #82646 already on this
branch: send() now resolves @Name tokens to channel-member pubkeys
(membership-accurate via `channels members`, TTL-cached, Unicode token
boundaries, ambiguous names stay presentation-only) and passes explicit
--mention args; recovery ladder handles membership drift, unresolvable
prose @tokens (escape retry, #78797), and a final self-mention downgrade.
2026-08-31 09:05:41 -07:00
blunkjamie-dev
1885a40ad3 fix(buzz): preserve literal mentions and exact UUID targets 2026-08-31 09:05:41 -07:00
Cameron Aragon
b302ae3f3e docs(buzz): clarify reaction-only user precedence 2026-08-31 09:05:41 -07:00
Cameron Aragon
743f86ed36 test(buzz): clarify reaction-only precedence 2026-08-31 09:05:41 -07:00
Cameron Aragon
c10c77d577 fix(buzz): acknowledge trusted agent tags without dispatch 2026-08-31 09:05:41 -07:00
Elmar Conradie
f40edee459 fix(buzz): trust explicit DM metadata fallback 2026-08-31 09:05:41 -07:00
Elmar Conradie
722209bb51 fix(buzz): require explicit group addressing 2026-08-31 09:05:41 -07:00
arimu1
ef2be55025 fix(buzz): treat NIP-10 replies to own messages as mentions
require_mention gated only on visible text, so Desktop thread replies
(e.g. /approve session) to the agent's own prompts were dropped with no
log. Cache event_id→(author, snippet) from seed/poll/WS/send, resolve
the direct e-tag parent, and dispatch when that parent is ours; also
populate reply_to_* on MessageEvent for gateway context injection.

Fixes #75826
2026-08-31 09:05:41 -07:00
liuhao1024
306dc874c1 fix(buzz): dispatch forum-channel kinds instead of chat kind 9 only
The inbound path hardcoded Nostr kind 9 at both the WebSocket
subscription filter and the dispatch gate (which runs before mention
gating), so Buzz forum channels — kind 45001 thread roots and 45003
comment replies — were silently never dispatched to the agent; chat and
stream channels worked, making the gap invisible (#90309). Block's own
ACP harness documents the forum kinds explicitly.

Introduce _DISPATCH_KINDS = {9, 45001, 45003} for the subscription
filter and dispatch gate. The stream kinds (46010/40007/45002) stay out
of scope until their dispatch semantics are confirmed.
_is_direct_message_event deliberately keeps its kind-9-only check:
widening it would let a p-tagged forum post be reclassified as a DM and
bypass mention gating. The send path already works unchanged (send()
omits --kind and threads via --reply-to).

Fixes #90309
2026-08-31 09:05:41 -07:00
Teknium
972f0314de feat(buzz): compose thread-topology cluster — reply_in_thread opt-out, NIP-10 root anchoring on all send paths, _PLATFORM_DEFAULTS tier
Compose/fix-up on top of the cherry-picked cluster commits:

- Unify the config surface: platforms.buzz.reply_to_mode: off (PlatformConfig
  field, as Discord/Telegram) and extra.reply_in_thread: false (the key Slack
  users know; env BUZZ_REPLY_IN_THREAD) are equivalent opt-outs, bridged
  through _apply_yaml_config and honored by send(), send_image(), and
  _standalone_send (cron delivery).
- Progress/status bubbles honor the opt-out too: gateway/run.py resolves
  _progress_reply_in_thread from the Buzz adapter (mirroring the Slack path)
  so the synthetic-thread fallback and the progress reply anchor are both
  suppressed when the user asked for flat replies (#75082, #95842).
- Deduplicate NIP-10 parsing: inbound session thread_id now reuses
  _extract_thread_root (marked root > reply > legacy positional e-tag)
  instead of a second inline root-marker-only scan.
- display_config: add buzz to _PLATFORM_DEFAULTS at TIER_MEDIUM — with
  edit_message now implemented, accumulate-style progress works, but without
  the entry Buzz inherited the verbose _GLOBAL_DEFAULTS and every interim
  update became a permanent channel post (#95841).
- plugin.yaml optional_env + platform docs for the new keys.
- contributors/emails mappings for the cherry-picked authors.
2026-08-31 07:30:44 -07:00
NanakoIce
66fa6e41c4 fix(buzz): reply in-thread instead of flat channel posts
Buzz has no native thread_id; channel threading is entirely --reply-to on
the triggering event. Interim commentary and progress bubbles only passed
the anchor via metadata.reply_to_message_id (or not at all), so most Kathy
posts landed as new top-level messages and cluttered channels.

- Honor metadata.reply_to_message_id in BuzzAdapter.send
- Pass reply_to on stream commentary sends
- Treat buzz like slack/mattermost for progress thread resolution
- Set _progress_reply_to to the trigger event for buzz
- Add unit tests for adapter metadata and progress routing
2026-08-31 07:30:44 -07:00
Tom Watts
cbeb925f0a fix(buzz): reply into the existing thread instead of nesting a new one
Every Buzz reply opened a fresh thread, including when the user was
already replying inside one. A threaded client fills up with an endless
ladder of one-message threads and the conversation becomes unreadable.

The adapter itself never had threading logic; the behaviour comes from
the generic gateway default. `_reply_anchor_for_event()` in
gateway/platforms/base.py returns `event.message_id`, which is right for
reply-style platforms (Telegram/Discord "reply to this message") but
wrong for a thread-style one: anchoring to the message you are answering
nests a new sub-thread under every single turn.

Buzz threads are NIP-10, so the information needed is already on the
inbound event. The adapter now records each inbound message's thread
root from its `e` tags and resolves the outbound anchor to that root, so
a reply joins the thread the user is typing in. When the trigger was
itself top-level there is no root and the anchor passes through
unchanged, preserving the existing behaviour of opening exactly one
thread from a top-level message.

Fixed in the adapter rather than in `_reply_anchor_for_event()`: Buzz is
a plugin-supplied platform, and its NIP-10 tag semantics do not belong in
core. Root extraction prefers an explicit `root` marker, falls back to a
lone `reply` marker (a message bearing only `reply` started the thread,
so that parent is the root for everything after it), and treats a legacy
unmarked `e` tag as the parent. A mention-only `p` tag is not a reply.

The root cache is an OrderedDict bounded at 512 entries with FIFO
eviction so a long-lived gateway cannot leak, and the resolver is applied
to the image send path as well as `send()`. Both helpers tolerate a
missing `_thread_roots` attribute, since the standalone/cron send path
constructs an adapter without running `__init__`.

Tests cover root extraction (top-level, thread opener, nested, legacy
unmarked tag), the top-level passthrough that guards the existing
behaviour, unknown/None anchors, cache bounding and eviction, and an
end-to-end assertion through `send()` that `--reply-to` carries the root.
Verified against the real event shapes returned by a live hosted relay.
2026-08-31 07:30:44 -07:00
al9000-max
5d73a11a97 Buzz adapter: honor reply_to_mode instead of always threading replies
The Buzz adapter appended --reply-to unconditionally, so every agent reply
threaded onto its parent event id with no way to turn it off.

reply_to_mode is already a generic PlatformConfig field, parsed for any
platform from gateway.platforms.<name>.reply_to_mode, and the Discord and
Telegram adapters both honor it. Buzz never read it, so setting it was a
silent no-op.

Read it in __init__ (BUZZ_REPLY_TO_MODE overrides config.yaml, matching how
require_mention and transport already work in this adapter) and skip the
--reply-to append when it is "off", at all three send paths: send(),
send_image(), and the out-of-process _standalone_send() used for
deliver=buzz cron delivery.

Default is unchanged ("first"), so existing installs keep threading.
2026-08-31 07:30:44 -07:00
yuvalfis
09cbce43e0 fix(buzz): preserve stable thread roots 2026-08-31 07:30:44 -07:00