Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.
Findings from the efficiency review pass on the two salvages, applied as one
small follow-up:
- copilot_auth: check the negative cache BEFORE taking the per-fingerprint
exchange lock. During the 60 s post-failure window, dashboard polls now
raise immediately instead of parking an executor thread behind the
in-flight holder (up to ~50 s) to learn the same answer. Test hangs
without the check (timeout 124), passes with it.
- buzz _localize_inbound_media: download_path.read_bytes() was still
evaluated on the loop as the argument to the offloaded cache call — up to
the 128 MiB inbound cap. Read off the loop too.
- test_list_credential_pool_keeps_loop_responsive: 0.5 s block / 0.25 s
threshold (2x margin) so runner descheduling cannot false-fail it while a
real regression still trips it.
The previous commit removed cache_media_bytes from teams/adapter.py's import
block but left one bare call at the Bot-Framework image path — a NameError
that the surrounding 'except Exception' swallowed, so every BF image
attachment was silently dropped (ruff F821). Route it through
cache_media_bytes_async like the file's two other sites.
Same bug class at the sites the sweep did not reach:
- buzz _download_attachment / _localize_inbound_media: cache_media_bytes
called from async methods → cache_media_bytes_async.
- photon _dispatch_inbound: _normalize_binary_payload (base64 decode +
cache write of possibly multi-MB payloads) ran on the loop → closure made
async, helper offloaded via asyncio.to_thread.
Tests: cover cache_media_bytes_async (thread-id + kwarg forwarding);
teams/buzz tests re-pointed at the async seam.
Post-rebase composition over #99431/#99429/#99427: file-attachment sends
route through _run_message_send so the mention-recovery ladder covers
media captions; _send_file_attachment/_send_local_file honor the
resolved thread-root anchor and reply_to_mode opt-out; send() records
event_meta on the verified receipt id (#75826); test fakes gain the
auth_tag kwarg and accepted-receipt shape.
Follow-up reconciliation: _send_file_attachment (the merged #95688/#74999
helper) now uses #78046's strict _parse_send_receipt contract and
redact_path error bounding, so CLI failures never leak host filesystem
paths and zero-exit unverified receipts are rejected on every outbound
media path.
#95688's _send_file_attachment refactor re-probed file existence, which
#74999's tests prove can race into a false 'not found' when the file
disappears between the caller's check and the helper's. Callers that
already verified the file pass probe=False; unverified document/video/
voice callers keep the guard.
Reconciles #84113 (authenticated same-relay URL localization) with #78051
(native imeta ingestion): _dispatch_message now merges caller-provided
verified imeta attachments with text-localized relay media instead of
clobbering them, dedupes paths, and downgrades mixed-source media to
DOCUMENT semantics so audio members are not routed through STT.
Localizing inbound relay media spends the agent's own Buzz credentials on
a URL chosen by the sender, so it must not run on the strength of the
adapter's local allow-list alone. Require the gateway's authorization
callback to return an explicit True before any `buzz media get` runs; a
denial, a missing callback, or a raising callback fails closed and leaves
the message text exactly as it arrived.
`_is_sender_authorized` previously wrapped the callback result in
`bool()`, so a truthy non-boolean (a status string, a sentinel) would
satisfy an `is True` gate's intent while bypassing its guarantee. Only
the literal booleans now propagate; anything else is "unknown", which the
existing Slack and Discord callers already treat as trust-unknown.
Reviewers asked for this boundary on the sibling inbound-media PRs
(#77734, #78051); it applies equally to the retrieval path in #75614,
which this change builds on.
Three sibling gaps in the WebSocket transport's conversation lifecycle:
- #78429: _send_channel_subscription defaulted a zero last_ts to
'since ~ now', so the message that CREATED a new conversation (created_at
fractionally before the subscription) was never delivered. A channel with
no high-water mark now subscribes from the beginning with
limit=_FETCH_LIMIT instead; seeded channels still resume from last_ts-1.
- #93557: relays do not guarantee a kind-44100 membership event per new
conversation, so WS-transport deployments never discovered DMs opened
mid-session until a reconnect. The WS loop now runs the same
_discover_dms sweep the poll transport uses, on the same cadence
(poll_interval * _DM_DISCOVERY_EVERY), via a companion task that is
cancelled with the connection.
- #75107: _discover_dms only ever adopted DM-shaped conversations, so a
real community channel the agent joined mid-run was never subscribed
until restart. In watch-all mode (no explicit channels list) newly
listed real channels are now adopted and seeded from their newest events
(history predating the join is not replayed). Explicit watch lists stay
authoritative.
- Widen the permanent-rejection match to the exact relay phrasings seen in
production (#97502): 'not a channel member' and 'auth-required', alongside
'restricted'.
- Close the re-adoption hole called out in review: _discover_dms() (both the
dms-list path and the channels-list fallback) now skips channels in
_restricted_channels, so a restricted channel dropped at runtime cannot be
silently re-added by the next discovery sweep and re-trigger the rejection.
- Credit: runtime CLOSED matching terms from PR #97502 by @repfigit; the
per-subscription drop + restricted set is PR #76850 by @xozai.
When a Buzz relay sends a CLOSED frame for a single subscription with a
'restricted: not a channel member' error, the adapter was raising
ConnectionError, tearing down the entire WebSocket connection, and
immediately reconnecting — causing a ~1.6 s flood in gateway.log.
Root cause: the CLOSED handler unconditionally raised ConnectionError
regardless of whether the error was permanent (restricted) or transient
(e.g. server shutdown).
Fix:
- On a 'restricted' CLOSED, drop only the offending subscription and
record the channel in a new _restricted_channels set instead of
tearing down the whole connection.
- Skip restricted channels during connect() seeding and
_subscribe_websocket() so reconnects don't re-trigger the same error.
- Non-restricted CLOSED frames still raise ConnectionError and reconnect
as before.
Adds three regression tests:
- test_websocket_loop_drops_restricted_channel_without_reconnect
- test_websocket_loop_reconnects_on_non_restricted_closed
- test_restricted_channels_skipped_during_subscribe
Tested on macOS against buzz.xozai.com: gateway.log shows zero
'restricted' errors and stable 'watching N channel(s) via websocket'
after the fix.
`connect()` calls `_seed_channel()` unconditionally, and seeding marks every
event currently in the channel as seen so a start never replays history at the
agent. A message that arrives after the process starts but before the seed
completes — or at any point while the gateway is down — sits in exactly that
history, so the seed swallows it permanently even though the Buzz relay still
has it. The `seen` set and `last_ts` lived only in memory, so there was nothing
to distinguish "already handled" from "never seen".
Each watched channel's cursor (`chat_type`, `last_ts`, and the bounded `seen`
id list) is now persisted under `HERMES_HOME/buzz/channel-cursors.json` and
restored at connect. Where a cursor exists the channel resumes from it and the
history fetch is skipped entirely; where none exists the old seed-from-history
behaviour is unchanged, so a first-ever run still never replays a backlog.
Details worth noting:
- The file records the identity and relay it was written for. A cursor from a
different bot or relay is ignored rather than trusted — the channel ids
would collide while the event stream behind them is a different one.
- Any read or parse failure leaves the cursors empty, which degrades to
seeding instead of failing the connect. Writes go through
`utils.atomic_json_write` (temp + fsync + replace), so a crash mid-write
cannot leave a truncated cursor behind.
- The restored `seen` list is trimmed to `_SEEN_CAP` on load, keeping the
newest ids, so a hand-edited or legacy file cannot grow the de-dupe set
without bound.
- Saves are gated on the cursor actually moving, so an idle channel does not
rewrite the file every poll interval. Both inbound transports are covered:
the poll sweep and the WebSocket event path share the same check.
Tests: six new cases in `TestChannelCursorPersistence` — the cursor is written
on seed, a restart resumes without spending a CLI call on history and then
delivers the mention that landed while the gateway was down, a foreign
identity or relay is ignored, a corrupt file falls back to seeding, the
restored `seen` set stays bounded, and an idle poll leaves the file untouched.
All six fail on main.
Tested on: Windows 11, Python 3.12. `python -m pytest
tests/gateway/test_buzz_adapter.py tests/gateway/test_buzz_websocket.py -q` —
33 passed (23 pre-existing + 6 new here, plus 4 WebSocket). Requires
`pytest-asyncio` (pinned at 1.3.0 in pyproject) — without it the async cases
in this file error out as unknown marks.
A relay-side close the transport never surfaces (observed as a CLOSE_WAIT
socket behind Cloudflare, #98097) parks the read loop forever while the
gateway keeps reporting connected: inbound stops, gateway_state.json stays
healthy, and only a restart recovers. The library keepalive should catch
this first, but as a last resort the read side now waits at most
_WS_READ_IDLE_TIMEOUT (300s) for a frame before raising into the existing
reconnect path, which re-authenticates and re-subscribes with per-channel
since filters intact.
Fixes#98097
Salvaged from PR #83414 (4 commits squashed to final state) and composed
with the presentation-mention escape retry from PR #82646 already on this
branch: send() now resolves @Name tokens to channel-member pubkeys
(membership-accurate via `channels members`, TTL-cached, Unicode token
boundaries, ambiguous names stay presentation-only) and passes explicit
--mention args; recovery ladder handles membership drift, unresolvable
prose @tokens (escape retry, #78797), and a final self-mention downgrade.
require_mention gated only on visible text, so Desktop thread replies
(e.g. /approve session) to the agent's own prompts were dropped with no
log. Cache event_id→(author, snippet) from seed/poll/WS/send, resolve
the direct e-tag parent, and dispatch when that parent is ours; also
populate reply_to_* on MessageEvent for gateway context injection.
Fixes#75826
The inbound path hardcoded Nostr kind 9 at both the WebSocket
subscription filter and the dispatch gate (which runs before mention
gating), so Buzz forum channels — kind 45001 thread roots and 45003
comment replies — were silently never dispatched to the agent; chat and
stream channels worked, making the gap invisible (#90309). Block's own
ACP harness documents the forum kinds explicitly.
Introduce _DISPATCH_KINDS = {9, 45001, 45003} for the subscription
filter and dispatch gate. The stream kinds (46010/40007/45002) stay out
of scope until their dispatch semantics are confirmed.
_is_direct_message_event deliberately keeps its kind-9-only check:
widening it would let a p-tagged forum post be reclassified as a DM and
bypass mention gating. The send path already works unchanged (send()
omits --kind and threads via --reply-to).
Fixes#90309
Compose/fix-up on top of the cherry-picked cluster commits:
- Unify the config surface: platforms.buzz.reply_to_mode: off (PlatformConfig
field, as Discord/Telegram) and extra.reply_in_thread: false (the key Slack
users know; env BUZZ_REPLY_IN_THREAD) are equivalent opt-outs, bridged
through _apply_yaml_config and honored by send(), send_image(), and
_standalone_send (cron delivery).
- Progress/status bubbles honor the opt-out too: gateway/run.py resolves
_progress_reply_in_thread from the Buzz adapter (mirroring the Slack path)
so the synthetic-thread fallback and the progress reply anchor are both
suppressed when the user asked for flat replies (#75082, #95842).
- Deduplicate NIP-10 parsing: inbound session thread_id now reuses
_extract_thread_root (marked root > reply > legacy positional e-tag)
instead of a second inline root-marker-only scan.
- display_config: add buzz to _PLATFORM_DEFAULTS at TIER_MEDIUM — with
edit_message now implemented, accumulate-style progress works, but without
the entry Buzz inherited the verbose _GLOBAL_DEFAULTS and every interim
update became a permanent channel post (#95841).
- plugin.yaml optional_env + platform docs for the new keys.
- contributors/emails mappings for the cherry-picked authors.
Buzz has no native thread_id; channel threading is entirely --reply-to on
the triggering event. Interim commentary and progress bubbles only passed
the anchor via metadata.reply_to_message_id (or not at all), so most Kathy
posts landed as new top-level messages and cluttered channels.
- Honor metadata.reply_to_message_id in BuzzAdapter.send
- Pass reply_to on stream commentary sends
- Treat buzz like slack/mattermost for progress thread resolution
- Set _progress_reply_to to the trigger event for buzz
- Add unit tests for adapter metadata and progress routing
Every Buzz reply opened a fresh thread, including when the user was
already replying inside one. A threaded client fills up with an endless
ladder of one-message threads and the conversation becomes unreadable.
The adapter itself never had threading logic; the behaviour comes from
the generic gateway default. `_reply_anchor_for_event()` in
gateway/platforms/base.py returns `event.message_id`, which is right for
reply-style platforms (Telegram/Discord "reply to this message") but
wrong for a thread-style one: anchoring to the message you are answering
nests a new sub-thread under every single turn.
Buzz threads are NIP-10, so the information needed is already on the
inbound event. The adapter now records each inbound message's thread
root from its `e` tags and resolves the outbound anchor to that root, so
a reply joins the thread the user is typing in. When the trigger was
itself top-level there is no root and the anchor passes through
unchanged, preserving the existing behaviour of opening exactly one
thread from a top-level message.
Fixed in the adapter rather than in `_reply_anchor_for_event()`: Buzz is
a plugin-supplied platform, and its NIP-10 tag semantics do not belong in
core. Root extraction prefers an explicit `root` marker, falls back to a
lone `reply` marker (a message bearing only `reply` started the thread,
so that parent is the root for everything after it), and treats a legacy
unmarked `e` tag as the parent. A mention-only `p` tag is not a reply.
The root cache is an OrderedDict bounded at 512 entries with FIFO
eviction so a long-lived gateway cannot leak, and the resolver is applied
to the image send path as well as `send()`. Both helpers tolerate a
missing `_thread_roots` attribute, since the standalone/cron send path
constructs an adapter without running `__init__`.
Tests cover root extraction (top-level, thread opener, nested, legacy
unmarked tag), the top-level passthrough that guards the existing
behaviour, unknown/None anchors, cache bounding and eviction, and an
end-to-end assertion through `send()` that `--reply-to` carries the root.
Verified against the real event shapes returned by a live hosted relay.
The Buzz adapter appended --reply-to unconditionally, so every agent reply
threaded onto its parent event id with no way to turn it off.
reply_to_mode is already a generic PlatformConfig field, parsed for any
platform from gateway.platforms.<name>.reply_to_mode, and the Discord and
Telegram adapters both honor it. Buzz never read it, so setting it was a
silent no-op.
Read it in __init__ (BUZZ_REPLY_TO_MODE overrides config.yaml, matching how
require_mention and transport already work in this adapter) and skip the
--reply-to append when it is "off", at all three send paths: send(),
send_image(), and the out-of-process _standalone_send() used for
deliver=buzz cron delivery.
Default is unchanged ("first"), so existing installs keep threading.
The gateway already streams by sending a first partial message and re-editing
it as tokens arrive, falling back to that path when an adapter does not support
native drafts. The Buzz adapter never implemented edit_message, so it inherited
the base stub that returns success=False and every reply was delivered in one
block when the turn finished, however long the turn took.
buzz-cli already exposes `messages edit` and `messages delete`, so no new
mechanism is needed.
One detail worth calling out for review: buzz-cli reports a NEW event id for
each edit, but the edit TARGET stays the original id, and the stream consumer
holds a single message_id for the whole stream. edit_message therefore returns
the id it was given rather than the one the CLI reports. Returning the CLI's id
would make every edit after the first address a message that was never sent.
delete_message is included because the consumer's fresh-final cleanup path
calls it when it replaces a preview rather than editing in place.
Tested: 10 new cases in tests/gateway/test_buzz_adapter.py covering the edit
target, stdin content, the returned id, echo suppression, finalize being inert,
both no-op guards, retryable vs non-retryable CLI failures, and delete. The
file goes from 33 passing to 6 failing if the adapter change is reverted while
the tests stay.
The WS NIP-42 auth path now prefers the connect()-resolved _auth_tag
(credentials-file aware, #79514) and falls back to a lazy scope-aware
_resolve_auth_tag() so a bare adapter re-auth stays profile-correct
(#98738): scoped multiplex profiles fail closed instead of borrowing
the default profile's tag from os.environ. _exec_buzz fakes updated
for the auth_tag kwarg introduced by the #83155 salvage.
Remove the stray 'return val if val is not None else default' tail left
in _unscoped_profile_secrets() when the new return was added, and note
in the docstring that the process-global cache is startup-gate-only
(review feedback on #95224).
check_requirements() runs at gateway startup before any per-profile
secret scope is installed, and the scope-less get_secret path reads
only os.environ -- so a Bitwarden-managed BUZZ_PRIVATE_KEY (only
BWS_ACCESS_TOKEN in .env) was invisible to the platform gate and Buzz
was silently skipped with a misleading install hint (#95216). When no
scope is active and the process env has no value, consult a cached
one-shot build of the profile secret mapping (build_profile_secret_scope
resolves external secret sources); an active scope still shadows this
rung entirely, so multiplexed cross-profile isolation is unchanged.
BUZZ_RELAY_URL reads in the gate now go through the same helper so an
externally managed relay passes too.
One BUZZ_* read survived the #98738 sweep unscoped: the NIP-42 WebSocket
auth path read BUZZ_AUTH_TAG with a bare os.getenv. Under
gateway.multiplex_profiles the process env holds the default profile's
bridge/.env output, so a scoped secondary profile without its own tag
signed its relay auth event with the default profile's NIP-OA
owner-attestation tag. Reproduced on f3845a72af before the fix; the same
repro now attaches no tag (fail-closed).
The read goes through _get_scoped_secret: scoped multiplex profiles fail
closed to "", while single-profile and unscoped default-profile reads
keep the legacy env behavior. Adversarial coverage added for the leak
itself, the scoped positive control, unscoped precedence, partial-extra
adapter config, scoped validate_config, scoped standalone-send target
resolution, central-authz wildcard/blank-entry/normalization semantics,
and adapter-intake vs central-authz agreement on the same allowlist.
Fixes#98738
Signed-off-by: Kosta Gorod <35299380+KostaGorod@users.noreply.github.com>
Under gateway.multiplex_profiles the default profile's YAML-to-env bridge
writes BUZZ_* values into os.environ, and every Buzz read gave that env
precedence over the secondary profile's PlatformConfig — so each secondary
adapter connected as the default identity, watched its channels, and
resolved its credentials file (#98738).
- Add _profile_scoped()/_scoped_platform_setting(): inside a secondary
profile scope extra is authoritative and env is not consulted (a missing
key fails closed to its default instead of borrowing the default
profile's value); single-profile and unscoped/default-profile reads keep
the legacy env-over-config precedence.
- Apply the scoped read to BuzzAdapter.__init__ (relay, CLI path, channels,
home channel, poll interval, require_mention, transport, allowed users),
_resolve_private_key (BUZZ_CREDENTIALS_FILE), validate_config,
_standalone_send, and check_requirements (which now consults the
profile's own config.yaml via the scoped home override).
- _env_enablement() returns None inside a profile scope and
_apply_yaml_config() skips the env bridge there, so the default profile's
env cannot fabricate Buzz for a profile that never configured it and a
secondary profile's YAML cannot be pinned into the process env
(first-writer-wins, #72348 Telegram/Discord mirror).
- Central authorization now consults a plugin platform's live-adapter
config.extra.allowed_users (gated on the registry entry declaring
allowed_users_env, with an optional normalize_user_id hook so Buzz npub
entries match hex-pubkey user ids) — under multiplex only the default
profile's list ever reached the env var, so listed secondary-profile
users were default-denied (#82871). Empty/absent lists change nothing;
default-deny is preserved.
Platforms added to main after the original branch was cut; keeps the
source invariant (every connectable adapter calls _wire_plugin_handlers)
true, and adds qqbot to the invariant test's gateway list.