`SlackAdapter._is_interactive_user_authorized` (approval / slash-confirm /
clarify Block Kit clicks) and the early pre-fetch gate in the message
handler recovered the runner via `_message_handler.__self__`, which is
None on a multiplexed adapter (closure handler) — so both fell to env-only
auth. The fallback read `SLACK_ALLOW_ALL_USERS` raw from `os.environ` and
its `_env` helper fell through to `os.environ` on a scoped miss: the
DEFAULT profile's allow-all flag / allowlist authorized callers on every
other profile's bot.
- Prefer the wired `set_authorization_check` callback (profile-bound
`_make_adapter_auth_check`) at both sites; keep `__self__` introspection
only for adapters wired without one.
- Env-only fallback reads go through `authz_mixin._platform_gate_env`
(scoped miss under multiplex → "", never os.environ); drop the raw
`os.getenv("SLACK_ALLOW_ALL_USERS")` pre-read.
Reapplies #72657 onto current main (original commit carried a bot
co-author trailer). Same class as Telegram #86296 / #65589.
Co-authored-by: MilaArtyNew <261982280+MilaArtyNew@users.noreply.github.com>
Under `multiplex_profiles` the primary adapter's message handler is a
profile closure, so the Telegram inline-button gate (and the early
message prefilter) cannot recover the runner via `_message_handler.__self__`
and fell to env-only auth. #65589 made the gate prefer the injected
`_authorization_check`, but `_make_adapter_auth_check` built a bare
`(user_id, chat_type, chat_id)` source: never route-stamped, never
`is_bot`.
- `_make_adapter_auth_check`: for the shared primary adapter under
multiplex, mirror the inbound message path exactly — stamp the
`profile_routes` match so the routed profile's pairing store is
consulted, and authorize under the TRANSPORT home via
`_is_user_authorized_for_source` (same split as
`_make_default_profile_message_handler`, 2afed50863). A rejected route
fails closed like the ingress gate. Retain the receiving adapter as
`_transport_adapter_ref` so config.yaml policy reads stay on it.
Accept `is_bot` / `thread_id` keywords. (#86296)
- `BasePlatformAdapter._is_sender_authorized`: forward `is_bot` /
`thread_id` as keywords only when set, so legacy 3-positional callbacks
keep working.
- Telegram `_source_from_message_for_auth` carries `from_user.is_bot`;
the prefilter forwards it so `TELEGRAM_ALLOW_BOTS=mentions|all` is
honored at the early gate under multiplex. (#92840)
- Telegram `_should_pass_unauthorized_dm_for_pairing`: same `__self__`
introspection class — fall back to the injected `gateway_runner` and
the adapter's owner profile.
Fixes#86296Fixes#92840
Co-authored-by: PRATHAMESH75 <118293218+PRATHAMESH75@users.noreply.github.com>
Co-authored-by: Ahmett101 <297889955+Ahmett101@users.noreply.github.com>
_is_callback_user_authorized resolved the gateway's auth chain through
_message_handler.__self__. For a secondary multiplexed adapter the
message handler is a per-profile closure with no __self__, so the
introspection silently fell through to the env-only fallback -- which
knows nothing about config allowlists or the pairing store, denying
every button caller on that profile (fail-closed, but wrong).
Prefer the auth callback GatewayRunner already injects at connection
time via set_authorization_check (registered for primary and multiplexed
adapters alike, delegating to the full _is_user_authorized chain), and
keep the introspection plus env fallback for adapters wired without it.
Same resolution pattern the admin-tier gate uses.
Address hermes-sweeper review on #76487:
- Prefer hermes_profile from send metadata when pruning stale topic
bindings so profile_routes cannot delete the transport adapter's
namespace instead of the routed runtime's
- Namespace lobby/capability cooldowns and /topic off cleanup by
(profile, chat_id)
- Document profile_name PKs and scoped cleanup SQL in telegram.md
- Regression: primary-adapter stamp + routed metadata prune isolation
_build_slash_event and _dispatch_thread_session built their SessionSource
without guild_id/parent_chat_id, while on_message passes both. Route
matching in build_source keys off exactly those fields, so under
gateway.multiplex_profiles a guild- or channel-routed profile never matched
a native slash command: /new, /reset, /model, /profile, /status ... all ran
against the default profile and reset the wrong session (#69178, #91633).
Pass guild_id (interaction.guild_id, falling back to channel.guild like the
message path) and the thread's parent channel id into build_source at both
sites. One test pins channel + thread routing parity with messages.
Fixes#69178Fixes#91633
Co-authored-by: Sora-bluesky <179361977+Sora-bluesky@users.noreply.github.com>
Co-authored-by: jondgilbert <42873618+jondgilbert@users.noreply.github.com>
Co-authored-by: tensorbit89-netizen <257030052+tensorbit89-netizen@users.noreply.github.com>
The two sibling is_bot computations (SessionSource construction and the
thread-history fetch) re-derived bot-ness from bot_id/subtype only, so an
api_human_users post would be human at the drop gate but still flagged
is_bot downstream. Use the single predicate everywhere.
Posts made with a user token (xoxp-) arrive with app_id and no
client_msg_id, so _event_declares_bot_sender dropped them as app traffic;
the only workaround was allow_bots: all. Adds
platforms.slack.extra.api_human_users (SLACK_API_HUMAN_USERS fallback), a
users-only allowlist consulted inside the predicate.
Salvaged from #100964 (users only: an app-id allowlist would also admit
the app's own xoxb bot posts, which share the user+app_id shape).
Follow-up to the #99641 salvage:
- One module-level _normalize_security() (ssl/tls/implicit -> tls, starttls,
plain/none -> plain; unknown -> WARNING + secure default) replaces the three
copies of the alias set; _connect_imap/_connect_smtp/_standalone_send all
compare against the canonical value. Unknown modes no longer raise.
- _tls_context(verify, host) is module-level and shared by all sites; when
verification is disabled for a non-loopback host it logs a WARNING.
- _esecret_bool: an unset/empty env var now yields the caller's default
(previously is_truthy_value('') returned False, silently disabling TLS
verification whenever EMAIL_*_TLS_VERIFY was unset).
- Documented surface is platforms.email.extra.{imap,smtp}_security and
{imap,smtp}_tls_verify in config.yaml; env vars remain an internal bridge
and are NOT added to plugin.yaml (optional_env feeds hermes setup prompts).
- Docs: Proton Mail Bridge / local relays recipe in user-guide/messaging/email.md.
- Tests: starttls builds IMAP4 then .starttls(); unknown mode falls back to
tls/starttls with verification still on.
Adds EMAIL_IMAP_SECURITY / EMAIL_SMTP_SECURITY and EMAIL_IMAP_TLS_VERIFY /
EMAIL_SMTP_TLS_VERIFY (env or platforms.email.extra.*) so the adapter can
talk to local relays such as Proton Mail Bridge (IMAP 1143 / SMTP 1025 with
STARTTLS and a self-signed certificate) instead of hardcoding IMAP4_SSL and
SMTP+STARTTLS with a verified default context.
Salvaged from #99641 (adapter.py only).
- alibaba-coding-plan-cn / alibaba-token-plan-cn keep the shared intl key vars
as ordered fallbacks after their dedicated *_CN_API_KEY, so users who set
ALIBABA_CODING_PLAN_API_KEY / ALIBABA_TOKEN_PLAN_API_KEY for the CN endpoint
keep working (the PR as filed dropped them).
- list_authenticated_providers hides a '-cn' row whose only lit key vars are
ones it shares with its non-CN sibling, unless that CN provider is the
configured model.provider. With only the shared key: one row, not two;
DASHSCOPE_API_KEY alone: 3 alibaba rows, not 4.
- Docs: environment-variables.md, providers.md.
ALIBABA_CODING_PLAN_CN_API_KEY is checked first for the China Coding Plan
endpoint (mirroring kimi-coding-cn), so the intl and CN rows no longer
light off the same key. Fixes#101122.
Every gateway/plugin platform adapter hard-coded aiohttp.ClientSession(trust_env=True)
(~20 sites), so a gateway launched by a Windows Scheduled Task that inherits a stale
HTTP_PROXY (Clash/V2Ray on 127.0.0.1:7890) looped on 'Cannot connect to host' with no
way to opt out short of NO_PROXY hacks per vendor host.
- gateway/platforms/base.py: gateway_trust_env() reads gateway.trust_env (default true);
resolve_proxy_url() skips generic HTTP(S)_PROXY/ALL_PROXY + macOS system-proxy
auto-detect when false (explicit per-platform vars still win).
- All aiohttp ClientSession sites in weixin, qqbot, matrix, line, wecom, slack, sms,
teams, google_chat now pass trust_env=gateway_trust_env(); mattermost + homeassistant
bare sessions gain the same kwarg (intent of #70119 / #56229).
- DEFAULT_CONFIG + cli-config.yaml.example + messaging docs.
- tests/gateway/test_gateway_trust_env.py: config flip + no-bare-literal sweep.
Reported-by: @ranlingfeng (#48820), @frontnopipe-cloud (#76309)
Co-authored-by: rcarrata <rcarratalasanchez@gmail.com>
Co-authored-by: Backroads4Me <TEDLANHAM@GMAIL.COM>
NousDashboardAuthProvider._verify_jwt (and the identical hunk in the
self-hosted OIDC provider) folded EVERY PyJWKClient failure into
ProviderError, which the gate translates to HTTP 503
{"detail":"Auth provider 'nous' unreachable"}. That branch fires for
jwt.DecodeError('Not enough segments') — i.e. the bearer is not a JWT at all
(an opaque peer key, a legacy token, garbage) — and for PyJWKSetError (JWKS
fetched fine, foreign kid). Neither involves reaching Portal, which is why
the hosted sjc agents in #94558 returned a fast, well-formed 503 that
survived token re-mint and instance restart while Portal was healthy.
Add one shared classifier, hermes_cli.dashboard_auth.classify_jwks_lookup_error:
only PyJWKClientConnectionError (transport) and an unexpected bare
PyJWKClientError stay ProviderError; DecodeError / PyJWKSetError /
InvalidTokenError become InvalidCodeError so verify_session() returns None
and the middleware proceeds to the next provider / refresh / 401 exactly as
the protocol documents. Both providers now use it.
Live repro (real NousDashboardAuthProvider against a local reachable JWKS
server; and the real gated web_server app): before — opaque bearer ->
ProviderError "JWKS lookup failed: DecodeError('Not enough segments')" ->
503 unreachable; after — verify_session() -> None, gated GET /api/auth/me
with the opaque bearer -> 401; a real JWT against an unreachable JWKS still
-> ProviderError (503).
This does not add /api/v1/message to the public-path allowlist (#94579):
that route has no verifier in this repo, so bypassing the gate would leave a
state-changing ingress fail-open. The correct fix is classification, which
also covers every other opaque-bearer surface.
Refs #94558
- generate() now passes kwargs.get("model") into _resolve_model(), so the
user's hermes tools pick (forwarded by the dispatcher as top-level
image_gen.model) is honored instead of silently dropped (#55893 class;
matches xai/krea/openrouter).
- Setup schema badge "internal" -> "paid" to match every other paid
image backend in the hermes tools picker.
- Tests: caller-model precedence, unknown caller model falls through,
model kwarg reaches the API payload, badge contract.
Adds a bundled image-generation backend for the Meta Model API
(https://api.meta.ai/v1), which is OpenAI-compatible. Exposes the
muse-image-1.0 model via the standard image_generate tool. This is the
image-gen companion to the already-bundled meta-ai chat provider
(plugins/model-providers/meta-ai, PR #88565).
- plugins/image_gen/meta-ai/ — provider registered as `meta-ai`, matching
the chat provider's id. Reuses the openai SDK pointed at Meta's base URL.
- Auth mirrors the chat provider: MODEL_API_KEY (Meta's documented var),
with META_API_KEY / META_MODEL_API_KEY aliases and a META_BASE_URL
override.
- Text-to-image only for now (capabilities gated); base64 (WebP) and URL
responses both handled and saved under $HERMES_HOME/cache/images/.
- Auto-loads as `kind: backend` and appears in `hermes tools` with no
central list edits, matching the other bundled providers.
- tests/plugins/image_gen/test_meta_ai_provider.py — 27 tests (metadata,
auth-alias resolution, base-url override, model resolution, generate
paths incl. b64 save, aspect mapping, URL caching, error handling).
- docs: image-generation feature page + provider-plugin built-in list.
A record-less delivery flag (final_response_sent /
final_content_delivered set with no recorded turn-final payload) was
trusted blindly by delivered_final_matches (None -> legacy trust), so a
first-edit prefix or a truncated finalize suppressed the gateway's
corrective send — silent partial delivery.
- delivered_final_matches: record-less flags are now reconciled against
the FINAL content via has_delivered_text; only the explicitly-marked
ambiguous-timeout path (_delivery_ambiguous) keeps legacy trust.
- _try_fresh_final and the native-streaming optimistic finalize now
record their delivered payload (the last record-less flag setters);
the optimistic record rolls back on definitive dispatch failure.
- Discord adapter: dead-transport send failures (client gone, WS
closed/reset) are classified as send_path_degraded (retryable) so the
delivery-obligation ledger's reconnect sweep replays the stranded
final response instead of losing it until a process restart.
Fixes#95382; closes the #98552 false-positive class.
The "Connecting to Telegram (attempt N/8)…" line logs at WARNING and
reaches the gateway's default stderr handler, but the matching
"Connected to Telegram (… mode)" line was INFO and went to the log file
only. A healthy startup therefore looked permanently hung at
"attempt 1/8" on the terminal — the logging-illusion half of #90835.
Promote the success line to WARNING so both sides of the connect
transition share the same console sink; a genuine hang is now the
absence of the success line. Adds an AST-level regression test pinning
the level pairing. Sibling adapters (homeassistant, wecom) log both
sides at INFO, so they don't have this asymmetry.
Fixes#90835
Every conversation-affinity hint Hermes sends is derived from the PHYSICAL
session id: prompt_cache_key on both OpenAI-wire transports, OpenRouter's and
Nous Portal's sticky session_id, and xAI's x-grok-conv-id. A host that mints
one physical session per RESPONSE re-keys all four on every reply, so the
conversation never lands back on the routing bucket it just warmed (#96811).
Two hosts do exactly that. Hermes Studio's group chat mints
gc_run_<room>_<profile>_<name>_<uuid4hex> per reply and destroys it after,
and POST /v1/responses with client-managed history mints str(uuid4()) per
request — while parsing X-Hermes-Session-Key one screen earlier and handing
it to the agent.
Hermes must not infer the logical conversation from the id's syntax: that
rule merges independent client-supplied ids and Studio members truncated past
its 96-character boundary (the #79017 failure class). It does not have to.
gateway_session_key is already the "stable per-chat key" built by
gateway.session.build_session_key from that header, and branching
deliberately does not key off it. The affinity path simply never consulted it.
- agent/prompt_cache_scope.py: declared_conversation_scope() resolves the key
into gwk_<sha256[:24]> and outranks the lineage walk (it is stable across
rotation AND across per-response ids). Hashed because, unlike a session id,
the key embeds platform/chat/user identifiers and leaves the process
verbatim as a sticky id and as x-grok-conv-id.
- agent/portal_tags.py: a separate ambient scope for ROUTING, published only
when a host declared one. The providers read the attribution id when it is
unset, so delegate trees keep sharing their parent's sticky key and every
host that keeps one id per conversation is byte-identical to before.
- hermes_state.py: is_explicit_fork_child() — the public view of the marker
rules that keep /branch children, delegate subagents and tool children off
their parent's chat key. Background-review forks clone the live runtime, so
_persist_disabled excludes them for the same reason (#79161).
Refs #96570Fixes#96811
Prevent the dedicated getUpdates pool from reusing server-closed connections and replace a polling HTTP client left open after a timed-out CLOSE-WAIT drain. Keep the general Bot API pool reusable so concurrent sends and edits are unaffected. Add regression coverage for both transport limits and stale-client replacement. Fixes#87057
After updater.stop() times out, HTTPXRequest.initialize() is a no-op unless
the client is already closed, so start_polling reused the wedged socket and
the gateway stayed alive but deaf. Rebuild the polling client after a hung
drain, watch getUpdates I/O independently of get_me(), and enable TCP
keepalive on the fallback transport.
The synchronous card-action handlers and the update-prompt resolver
authorized clicks with _allow_group_message(), which answers "may this
sender chat in this group?" — with group_policy=open it returns True
for everyone. The approval resolver already used the correct operator
gate (_is_interactive_operator_authorized), so the three code paths
disagreed: with an open group policy an out-of-allowlist click on an
update-prompt card was fully executed, and approval clicks returned a
resolved-looking card before being rejected asynchronously.
Authorize all three paths with _is_interactive_operator_authorized(),
which checks membership of admins ∪ allowed_group_users (wildcard and
the empty pairing-mode allowlist keep their existing allow semantics,
matching _admit's DM pairing default). A missing operator identity now
fails closed on the update-prompt resolver instead of skipping the
check.
Fixes#96045
Salvage hardening on top of the three cherry-picked contributor commits
(#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia),
closing the remaining unverified-PID kill sites as one class (#98814, #89614):
- pid_is_hermes: token-boundary 'hermes' match (no more loose substring
false-positives), and an explicit start-time expectation is now honored
on POSIX too (a mismatched fingerprint is a recycled PID on any platform).
- kill_process_tree: drop the guard on our OWN retained Popen child — a
retained handle pins the PID, so the check could only false-refuse.
- gateway.status.terminate_pid: POSIX force-kills also refuse when a
caller-provided expected_start_time no longer matches.
- kill_gateway_processes: re-verify the LIVE cmdline at kill time (the
scan-time match is a TOCTOU window).
- _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and
require a still-matching identity before the delayed SIGKILL escalation.
- whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID
unless the live process is actually a node bridge (was a stranger-kill).
- browser daemon reap/close paths: pass the start-time fingerprint into
ProcessRegistry._terminate_host_pid (previously unverified), and the
session-close path now runs the same daemon identity verification as
the orphan reaper.
- tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows
probes (real spawned processes, real psutil ancestry) wired into the
on-demand windows-latest wine2e lane.
Fixes#98814Fixes#89614
Post-rebase composition over #99431/#99429/#99427: file-attachment sends
route through _run_message_send so the mention-recovery ladder covers
media captions; _send_file_attachment/_send_local_file honor the
resolved thread-root anchor and reply_to_mode opt-out; send() records
event_meta on the verified receipt id (#75826); test fakes gain the
auth_tag kwarg and accepted-receipt shape.
Follow-up reconciliation: _send_file_attachment (the merged #95688/#74999
helper) now uses #78046's strict _parse_send_receipt contract and
redact_path error bounding, so CLI failures never leak host filesystem
paths and zero-exit unverified receipts are rejected on every outbound
media path.
#95688's _send_file_attachment refactor re-probed file existence, which
#74999's tests prove can race into a false 'not found' when the file
disappears between the caller's check and the helper's. Callers that
already verified the file pass probe=False; unverified document/video/
voice callers keep the guard.
Reconciles #84113 (authenticated same-relay URL localization) with #78051
(native imeta ingestion): _dispatch_message now merges caller-provided
verified imeta attachments with text-localized relay media instead of
clobbering them, dedupes paths, and downgrades mixed-source media to
DOCUMENT semantics so audio members are not routed through STT.
Localizing inbound relay media spends the agent's own Buzz credentials on
a URL chosen by the sender, so it must not run on the strength of the
adapter's local allow-list alone. Require the gateway's authorization
callback to return an explicit True before any `buzz media get` runs; a
denial, a missing callback, or a raising callback fails closed and leaves
the message text exactly as it arrived.
`_is_sender_authorized` previously wrapped the callback result in
`bool()`, so a truthy non-boolean (a status string, a sentinel) would
satisfy an `is True` gate's intent while bypassing its guarantee. Only
the literal booleans now propagate; anything else is "unknown", which the
existing Slack and Discord callers already treat as trust-unknown.
Reviewers asked for this boundary on the sibling inbound-media PRs
(#77734, #78051); it applies equally to the retrieval path in #75614,
which this change builds on.
Three sibling gaps in the WebSocket transport's conversation lifecycle:
- #78429: _send_channel_subscription defaulted a zero last_ts to
'since ~ now', so the message that CREATED a new conversation (created_at
fractionally before the subscription) was never delivered. A channel with
no high-water mark now subscribes from the beginning with
limit=_FETCH_LIMIT instead; seeded channels still resume from last_ts-1.
- #93557: relays do not guarantee a kind-44100 membership event per new
conversation, so WS-transport deployments never discovered DMs opened
mid-session until a reconnect. The WS loop now runs the same
_discover_dms sweep the poll transport uses, on the same cadence
(poll_interval * _DM_DISCOVERY_EVERY), via a companion task that is
cancelled with the connection.
- #75107: _discover_dms only ever adopted DM-shaped conversations, so a
real community channel the agent joined mid-run was never subscribed
until restart. In watch-all mode (no explicit channels list) newly
listed real channels are now adopted and seeded from their newest events
(history predating the join is not replayed). Explicit watch lists stay
authoritative.
- Widen the permanent-rejection match to the exact relay phrasings seen in
production (#97502): 'not a channel member' and 'auth-required', alongside
'restricted'.
- Close the re-adoption hole called out in review: _discover_dms() (both the
dms-list path and the channels-list fallback) now skips channels in
_restricted_channels, so a restricted channel dropped at runtime cannot be
silently re-added by the next discovery sweep and re-trigger the rejection.
- Credit: runtime CLOSED matching terms from PR #97502 by @repfigit; the
per-subscription drop + restricted set is PR #76850 by @xozai.
When a Buzz relay sends a CLOSED frame for a single subscription with a
'restricted: not a channel member' error, the adapter was raising
ConnectionError, tearing down the entire WebSocket connection, and
immediately reconnecting — causing a ~1.6 s flood in gateway.log.
Root cause: the CLOSED handler unconditionally raised ConnectionError
regardless of whether the error was permanent (restricted) or transient
(e.g. server shutdown).
Fix:
- On a 'restricted' CLOSED, drop only the offending subscription and
record the channel in a new _restricted_channels set instead of
tearing down the whole connection.
- Skip restricted channels during connect() seeding and
_subscribe_websocket() so reconnects don't re-trigger the same error.
- Non-restricted CLOSED frames still raise ConnectionError and reconnect
as before.
Adds three regression tests:
- test_websocket_loop_drops_restricted_channel_without_reconnect
- test_websocket_loop_reconnects_on_non_restricted_closed
- test_restricted_channels_skipped_during_subscribe
Tested on macOS against buzz.xozai.com: gateway.log shows zero
'restricted' errors and stable 'watching N channel(s) via websocket'
after the fix.
`connect()` calls `_seed_channel()` unconditionally, and seeding marks every
event currently in the channel as seen so a start never replays history at the
agent. A message that arrives after the process starts but before the seed
completes — or at any point while the gateway is down — sits in exactly that
history, so the seed swallows it permanently even though the Buzz relay still
has it. The `seen` set and `last_ts` lived only in memory, so there was nothing
to distinguish "already handled" from "never seen".
Each watched channel's cursor (`chat_type`, `last_ts`, and the bounded `seen`
id list) is now persisted under `HERMES_HOME/buzz/channel-cursors.json` and
restored at connect. Where a cursor exists the channel resumes from it and the
history fetch is skipped entirely; where none exists the old seed-from-history
behaviour is unchanged, so a first-ever run still never replays a backlog.
Details worth noting:
- The file records the identity and relay it was written for. A cursor from a
different bot or relay is ignored rather than trusted — the channel ids
would collide while the event stream behind them is a different one.
- Any read or parse failure leaves the cursors empty, which degrades to
seeding instead of failing the connect. Writes go through
`utils.atomic_json_write` (temp + fsync + replace), so a crash mid-write
cannot leave a truncated cursor behind.
- The restored `seen` list is trimmed to `_SEEN_CAP` on load, keeping the
newest ids, so a hand-edited or legacy file cannot grow the de-dupe set
without bound.
- Saves are gated on the cursor actually moving, so an idle channel does not
rewrite the file every poll interval. Both inbound transports are covered:
the poll sweep and the WebSocket event path share the same check.
Tests: six new cases in `TestChannelCursorPersistence` — the cursor is written
on seed, a restart resumes without spending a CLI call on history and then
delivers the mention that landed while the gateway was down, a foreign
identity or relay is ignored, a corrupt file falls back to seeding, the
restored `seen` set stays bounded, and an idle poll leaves the file untouched.
All six fail on main.
Tested on: Windows 11, Python 3.12. `python -m pytest
tests/gateway/test_buzz_adapter.py tests/gateway/test_buzz_websocket.py -q` —
33 passed (23 pre-existing + 6 new here, plus 4 WebSocket). Requires
`pytest-asyncio` (pinned at 1.3.0 in pyproject) — without it the async cases
in this file error out as unknown marks.
A relay-side close the transport never surfaces (observed as a CLOSE_WAIT
socket behind Cloudflare, #98097) parks the read loop forever while the
gateway keeps reporting connected: inbound stops, gateway_state.json stays
healthy, and only a restart recovers. The library keepalive should catch
this first, but as a last resort the read side now waits at most
_WS_READ_IDLE_TIMEOUT (300s) for a frame before raising into the existing
reconnect path, which re-authenticates and re-subscribes with per-channel
since filters intact.
Fixes#98097
Salvaged from PR #83414 (4 commits squashed to final state) and composed
with the presentation-mention escape retry from PR #82646 already on this
branch: send() now resolves @Name tokens to channel-member pubkeys
(membership-accurate via `channels members`, TTL-cached, Unicode token
boundaries, ambiguous names stay presentation-only) and passes explicit
--mention args; recovery ladder handles membership drift, unresolvable
prose @tokens (escape retry, #78797), and a final self-mention downgrade.