Inside a multiplexed secondary profile scope, `_profile_buzz_extra` read the buzz block only from
`gateway.platforms.buzz`, while the runtime loader (`gateway/config_loader.py::platform_section`,
the same seam that feeds this plugin's `apply_yaml_config_fn`) also resolves a top-level `buzz:`
block and the documented `platforms.buzz` shape. A fully configured secondary profile therefore
failed `check_requirements` closed on every boot ("Platform 'Buzz' requirements not met").
Resolve the section through `platform_section` so the gate sees exactly what the loader sees.
Row L1 (profile-scoped gate reads config), topology T2 (multiplexed secondary profile).
Salvaged from #126060 by @liuhao1024 (authorship kept); the test is trimmed to one A->B->A
invariant across two profile homes. Supersedes #126067 (same mechanism, later).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Follow-up trim to @JE4NVRG's salvaged commit (is_connected → env_is_connected("A2A_PORT")):
- plugins/platforms/a2a/tools.py::_a2a_tools_available read os.getenv("A2A_PORT") too — the
same unscoped gate in the same plugin, so every secondary profile also paid for the a2a
toolset whenever the launch profile set A2A_PORT. Now get_scoped_secret, same policy as
the adapter's port read (scoped miss fails closed; default/T1 reads its own environ).
- Tests trimmed to two invariants: A→B→A over two temp profile homes under
set_multiplex_active(True) with the launch value in os.environ (profile B with no
A2A_PORT of its own is NOT connected and gets no tools; red on origin/main at index 1),
and the standalone T1 case where os.environ IS the profile.
Row L1 (boot-time platform enablement probe), topologies T2/T3 (multiplexed host serving
secondary profiles). No other plugins/platforms/*/__init__.py is_connected/check_requirements
does an unscoped env read (buzz #125985 is handled separately).
Co-authored-by: JE4NVRG <jean.v1803@gmail.com>
is_connected() gated on a raw os.getenv("A2A_PORT"), which under multiplexing
holds the launch (default) profile's value. Every secondary profile therefore
instantiated the inbound A2A server, all of them fell back to the module
default port 9900, and all but one died with "could not bind ... Address
already in use" (fatal / bind_failed in gateway_state.json) on every boot.
Resolve the gate through gateway.platforms._shared.env_is_connected, the same
scope-aware idiom the sibling adapters use -- the adapter's own port read
already goes through _get_scoped_secret for exactly this reason.
Tests: tests/plugins/test_a2a_plugin.py pins the behavior. A profile with no
A2A_PORT in its own scope is not connected and never reads os.environ (a raw
os.getenv fails the test), a scoped port connects, and extra.enabled still
short-circuits.
Fixes#122126
The r3 fold scoped only the dmarc verdict to its clause. SPF and DKIM
still came from a whole-string, last-match-wins regex scan, so the same
quoted/comment smuggle closed for dmarc still authenticated a spoofed
From (GHSA-rxqh-5572-8m77), e.g.
spf=fail smtp.mailfrom="x spf=pass smtp.mailfrom=example.com "@evil.test
spf=fail (spf=pass) smtp.mailfrom=a@example.com
spf=fail smtp.mailfrom=a.spf=pass@example.com
dkim=pass header.d=evil.test header.i="x header.d=example.com y"@evil.test
Every verdict now comes from the leading method=result token of its
own clause (from _ar_clauses, comments dropped), and its domains only
from that clause. Properties are read by a token scanner that consumes
quoted-strings and other key=value tokens whole, so quoted contents are
never read as properties while a quoted value still is
(header.from="example.com"). SPF fails closed on more than one spf
clause; DKIM accepts any single dkim=pass clause whose own header.d
aligns (multi-signature mail is normal), never mixing clauses. The
whole-string methods/props (methods["dmarc"] was dead) are gone.
Also: a stray ')' at depth 0 is now unbalanced (it split header.from
out of the dmarc clause); a From with more than 64 '(' takes the silent
empty-sender drop instead of a parseaddr RecursionError logged as an
error; the cap test asserts the cap directly instead of wall-clock
timing; the empty-From drop assertion moves next to the other
_extract_email_address rejects; and the untested >1-dmarc, unbalanced
and backslash-escape rules get reject strings.
Round-3 review of the #124322 salvage:
- stdlib parseaddr is pure Python and superlinear on hostile input: a
100KB `From: <a@a@...` held the GIL ~1s per message (main ~0), and any
remote sender reaches it before auth. Refuse values over _MAX_FROM_LEN
(2048; RFC 5322 lines cap at 998) right after unfolding so they take the
existing empty-sender drop. The timing assertion now times
_extract_email_address itself instead of a private regex.
- The dmarc clause split ignored quoted-strings and stripped comments one
level only, so an attacker-controlled SPF-passing envelope sender like
smtp.mailfrom="x;dmarc=pass header.from=example.com x"@evil.test planted
a fake dmarc=pass ahead of the real dmarc=fail and authenticated
admin@example.com. Split clauses with a quote- and nested-comment-aware
scanner (comments dropped, quoted values kept so header.from="x" is still
read and unquoted), and fail closed on an unbalanced value or when more
than one clause starts with dmarc=. Tests pin the smuggled clause,
reason="a;b", the nested comment, a quoted aligned header.from, and that
every header.from in the dmarc clause must align (all->any goes red).
- `Doe, John (CEO) <j@x>` was dropped because the fallback refused any
paren. Strip (comments) from the display part first; `attacker@evil.test
(c) <victim>` stays rejected by the '@' rule.
- Results without '@' (`John` -> `john`, `a\"b <v@x>` -> `a\`) were
dispatched as sender ids; return "" so they take the drop path.
The r1 malformed-From fallback regex ([^<>\s]+@[^<>\s]+) backtracked
catastrophically on `From: <a@a@a@...`: 32KB held the GIL ~23s, and any
remote sender reaches it before auth. Capture <([^<>\s]+)> instead and
check the '@' in Python (same accepted set, 100KB in ~3ms).
The fallback also mapped multi-mailbox / group / comment-prefixed From
values (`attacker@evil.test, <victim@x>`, `Grp: a@evil; <victim@x>`,
`attacker@evil.test (c) <victim@x>`) to the bracketed victim, which a
dmarc=pass without header.from then authenticated. Only fall back when the
display part has no ; : ( ) and no '@' unless it is exactly the bracketed
address; `Doe, John <j@x>` and `j@x <j@x>` still resolve. The rule now lives
only in the docstring (the old inline comment was wrong).
dmarc: strip (comments) before splitting clauses, and take the verdict and
header.from from the same first clause that starts with dmarc=, so
`dmarc=pass (p=none; sp=none) header.from=evil.test`, a later dmarc=pass
clause after dmarc=fail, and duplicate misaligned header.from are rejected,
while `arc=pass (dmarc=fail ...); dmarc=pass header.from=<ours>` passes.
Property clean-up is shared via _auth_props.
The empty-sender drop now runs right after address parsing and gets an
assertion (it was untested). Tests extend existing ones (no new tests).
Strict parseaddr returns '' for common RFC-invalid From values that the
old regex resolved: an unquoted address as display name
(john@example.com <john@example.com>), an unquoted comma (Doe, John
<j@example.com>), or a@x.test <b@y.test>. Allowlisted senders using such
clients were silently dropped, and under open access every one of them
shared an empty chat_id. Fall back to the bracketed address only for the
unambiguous shape (no quotes, exactly one <...> pair), so the quoted
display-name spoof still resolves to the attacker. _parse_fetched_message
now drops messages whose From yields no address instead of dispatching
an empty identity.
The dmarc header.from alignment check read header.from from props merged
across all clauses, so a later dkim clause's header.from could override
the dmarc clause's value. Read it from the dmarc clause only, and pin the
misaligned dmarc=pass rejection in the existing dmarc test (previously
no test covered it). Also drop a redundant _domain_of ternary, trim the
_extract_email_address docstring and a duplicate test assertion.
_verify_sender_authentication accepted any dmarc=pass verdict, even one
issued for a different domain than the From we parsed. That is what let
the quoted-display-name spoof ride on the attacker's own truthful DMARC
pass. When the trusted Authentication-Results names header.from, require
it to align with the From domain; otherwise fall through to the aligned
SPF/DKIM checks. Defense in depth for the #124322 parser fix.
_extract_email_address took the FIRST <...> pair, so
From: "Victim <victim@example.com>" <attacker@evil.test> resolved to the
victim. The attacker's own domain passes DMARC truthfully, so the
allowlist (EMAIL_ALLOWED_USERS), pairing and session identity were all
evaluated against an address the sender does not control.
Use email.utils.parseaddr, which keeps the quoted text as the display
name and returns the real addr-spec. RFC 5322 folding is unfolded first
so a folded quoted display name is not mistaken for the mailbox.
Salvages #124322.
Both reconnect paths (gateway/run_adapters.py watcher and multiplex
secondary) build a FRESH SimplexAdapter before connect(is_reconnect=True),
so the per-instance _allowlist_warned flag never suppressed anything: a
daemon-down cold boot re-logged the warning on every backoff retry. Gating
on `not is_reconnect` would instead lose the warning when the first connect
fails. Dedup at module level keyed on (hermes_home_key(), frozenset(names)),
still checked before the connectivity probe. No shared warn-once helper
exists (plugin_compat.warn_once is compat-specific).
Read the value with platform_gate_env (the reader authz uses; differs from
get_scoped_secret when a scope is installed with multiplex off) and decode
JSON list literals with decode_json_list_literal like _coerce_allow_set, so
'["4","9"]' written by `hermes config set` no longer warns that valid IDs
are ignored.
The caplog test now builds two fresh adapters (first connect fails, second
succeeds) and asserts exactly one warning naming only 'alice'; it fails with
2 warnings against the pre-fold adapter.
The name-entry warning in connect() read SIMPLEX_ALLOWED_USERS via raw
os.getenv, while authz reads it profile-scoped. Under multiplexing a
secondary profile would warn about (or stay silent on) the default
profile's list rather than the one actually enforced. Use the module's
_get_scoped_secret + _parse_comma_list like __init__ does.
It also only fired on a successful non-reconnect connect: if the daemon
was down at cold boot the first connect() failed and every retry came in
with is_reconnect=True, so the warning never appeared. Evaluate it before
the connectivity probe, once per adapter via an instance flag.
Test: two connects (first fails) -> exactly one warning naming only
'alice' for scoped '4, alice' while os.environ holds 'bob'. Red on the
pre-fold adapter (0 warnings) and on a raw-os.getenv variant (names bob).
After #44729 SIMPLEX_ALLOWED_USERS matches only the numeric contactId, but
the docs still told operators display names work, and existing name
entries would silently stop matching. Update the docs and log a one-time
warning at first connect listing non-numeric entries that are now ignored.
The SimpleX sender allowlist (SIMPLEX_ALLOWED_USERS) previously matched
against both the stable numeric contactId (user_id) and the mutable
display name (user_name). Since any SimpleX contact can change their
localDisplayName / profile.displayName to match another user's, this
allowed an unauthorized contact to bypass the allowlist by setting a
colliding display name.
Remove the user_name check so that SIMPLEX_ALLOWED_USERS only matches
on the immutable contactId. Operators must use numeric contact IDs in
the allowlist.
Fixes#44729
(cherry picked from commit b4aa29da1567d45920f79aabdb36b44c5f87bde5)
The r3 fold made park refuse an older voice once a newer one from the same
sender had parked. With one /sync batch holding [v1 (slow gate), m1 bare
mention, v2 (fast)], v2 parked first, v1 was then refused, and m1 claimed
v2 -- a voice sent after the mention. v1 was lost and v2's own mention was
answered as empty text.
The bare mention now takes an arrival limit (ParkedVoices.mark) before it
settles, and claim pops the newest parked voice that began before that
limit, dropping older ones. park no longer refuses by arrival; parked
voices are kept per sender ordered by seq (bounded to 4). A claim made
while gates are still in flight records a floor, so a late older voice
(seq <= claimed) cannot park and outlive the claim; the floor clears when
the sender's in-flight list empties. One answer per bare mention, newest
before the mention wins, and the r3 orphan case still leaves nothing
parked.
Also: pending() ignores entries past CLAIM_WINDOW_SECONDS so an expired
voice no longer sends every later text through the mention regexes; the
m.thread root lookup is one _thread_root helper used by both the park
decision and its pre-check; voice_gate is annotated.
The kept test gains a [voice slow, mention, voice2 fast] + mention2 row:
red on a932dc031c (['$voice2', '$text2']) and on f1ea71de7c
(['$voice', '$text2']).
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
The same-sync-batch in-flight mark was one slot per (room, sender), so a
second voice's begin overwrote the first. When the older voice finished
gating after the newer one, the bare mention settled on (and claimed) the
newer voice only, and the older one parked afterwards as an orphan: the
sender's next unrelated bare mention within 120s re-dispatched it, so a
stale voice got downloaded, transcribed and answered.
The in-flight mark is now a list of gates per key. release removes only
its own gate (idempotent) and drops the key once it is empty; settle
waits on all current gates under the same 5s cap. Each gate carries an
arrival sequence and park refuses an older voice once a newer voice from
the same sender has parked, so a late older voice can neither replace
nor outlive the claim of the newer one (still one parked voice per
sender, newest wins).
The mark was also held across the whole _resolve_message_context,
including the display-name fetch and thread mark, and taken for voices
that can never park (the voice itself @mentions the bot, free rooms, bot
threads, non-allowlisted rooms). A bare mention racing such a voice
waited for the voice's display-name fetch (0.61s vs 0.31s on main with a
0.3s fetch). begin now runs only when a synchronous pre-check says the
voice can park, and the mark is released as soon as the park decision
is made (before the display-name fetch; DM voices release there too),
with the finally still covering exceptions and cancellation.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
mautrix dispatches every event of one /sync batch as its own task
(wait_sync=True). The voice only parks after awaiting the room identity,
which is a homeserver round-trip whenever the 60s identity cache is stale,
i.e. in any idle room. The bare-mention text checked the park without
awaiting, found nothing, dispatched an empty text, and the voice then
parked and expired unclaimed.
A parkable voice is now marked in-flight per (room, sender) right before
its first await and released in a finally once it is gated. A bare
mention that sees an in-flight voice waits for it (bounded, 5s) before
claiming; ordinary text only pays two dict lookups and never waits.
After a claim the bare-mention event also gets its read receipt, so the
read marker is not left one event short of the pre-claim behaviour.
Cleanups: ParkedVoices drops its unused window param, stores
(parked_at, voice) instead of a flat tuple, the voice_mention import
moves to the top import block, and the redundant require_mention check
on the claim path is dropped (parking/in-flight only happen under it).
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
Review cleanups on the parked-voice fix:
- The claim bypass is now an explicit mention_claimed parameter threaded
_handle_media_message -> _resolve_message_context. The old _claimed set was
only drained inside the require_mention branch, so a voice whose thread became
a bot thread while parked leaked its id for the adapter's lifetime.
- One MSC3245 predicate (has_voice_marker, `.get(...) is not None`) shared by
is_voice_event and _classify_inbound_media; the two used to disagree on a
null marker (parked as voice, classified as AUDIO).
- One mention helper (_content_mentions_bot) for both the gate and the
bare-mention claim, instead of two copies of the m.mentions extraction.
- ParkedVoices.has() dict lookup gates the claim, so ordinary text messages
no longer pay _strip_mention/regex work when nothing is parked.
- Test: a same-room mention WITH text stays a text (gives the bare-only guard
teeth) and the parked voice is asserted never downloaded.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
Element X ships an MSC3245 voice event with an empty m.mentions block and
sends the mention the user typed while recording as a separate m.text
event right after. With require_mention on, the gate dropped the voice and
the follow-up bare @bot was dispatched as an empty text, so the bot never
answered the voice.
Park an unmentioned voice keyed by (room_id, sender) instead of forgetting
it; a bare mention from the same sender in the SAME room within 120s claims
it and the voice is processed in place of the empty text. The claimed event
id passes the mention gate exactly once, so it is not re-parked; expired
entries are pruned on every access. Parked voices are never downloaded or
transcribed (no wake-word/STT of unmentioned audio), and a mention in
another room never pulls a voice across rooms.
The state lives in a small sibling module (voice_mention.py) because
adapter.py is already far past the file-size budget.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
The model-picker fix added a second copy of the choice picker's auth
check, and each copy rebuilt the callback ctx that _handle_callback_query
had already computed as `cb`. The chat-id picker loop is the only route
into both handlers. Gate there once with the existing cb, so a picker
added to that loop later is covered without anyone having to copy the
check again. Fail-closed behaviour is unchanged: _callback_authorized
returns True only for an authorized tapper and answers the denial text
otherwise. The back-button test calls the handler directly and no longer
needs its auth stub.
Co-authored-by: Dusk1e <yusufalweshdemir@gmail.com>
Every other ModelPickerView callback (provider/model select, expensive
confirm, back) runs through the shared _HermesView._gate auth check, but
_on_cancel did not. A non-allowlisted member could tap Cancel on the
owner's picker, marking it resolved and clearing the view. No model
switch was possible, so impact is low, but the gate should be uniform.
Reuse the same _gate call the sibling handlers use.
Model picker taps (mp:/mpg:/mpv:/mm:/mc:/mb/mx/mg:) never went through the
callback allowlist, so a group member who is not allowlisted could tap the
owner's /model picker and switch the owner's session or config model. Gate
the handler on _callback_authorized like the choice-picker, approval and
clarify buttons, before any picker state is read.
Ported from #33859 (gateway/platforms/telegram.py) to the plugin adapter.
(cherry picked from commit b002be8feacd09fde0ee9752de7b396d0aac69ba)
- salvage summary cap and _bound_oversized_record still composed bare
truncation idioms in model-visible text; route them through elide /
elide_middle so a copied marker is guard-visible.
- the active-task line repr()'d the elided text, escaping the marker's
apostrophe when the user text held both quote kinds and hiding it from
the guard; elide after repr instead (text within the cap stays whole).
- a leftover budget smaller than the marker produced a content-free,
over-budget marker line in _build_verbatim_user_section and the Slack
nested-attachment path; skip the item instead.
- _build_verbatim_user_section elided twice, reporting the wrong total;
one elide at min(cap, remaining).
- drop redundant len() pre-checks before elide() and name verification
stop's repeated 1200.
Co-authored-by: ahisblessed <ahisblessed@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
The Slack Block Kit payload dump and the nested-attachment text budget
are both fed to the agent, and trajectory_compressor's summarizer input
becomes training data; all three still used the imitable bare
"... [truncated]" idiom. Route them through elide()/elide_middle() with
module-level imports and extend the no-idiom invariant to scan
plugins/platforms/slack and trajectory_compressor.py.
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
With reply_in_thread: false the whole channel is one session and the bot
answers top-level, so an unmentioned top-level message there is a
follow-up in a conversation the bot is part of, like a thread reply.
reply_expected is now False for a free-channel message only when it starts
its own session (a new top-level thread), else None. The bot-id set is
built inside _slack_reply_expected, as _channel_gate_allows does.
The Slack rule marked every admitted message that was not a DM, a mention
or a command as not addressed, so a plain "done?" in a thread the bot is part
of, or a reaction trigger, could end on a bare silence marker and vanish,
the case #111624 fixed (#110952).
reply_expected is now False only for a message that opens by @mentioning
someone else, or a top-level message a free-response channel admitted
without a mention. Reaction triggers and pipe-form self mentions count as
addressed; other thread replies are None (visible fallback). The
free-channel predicate moves into _slack_is_free_channel so the gate and
the rule read the same one. The test drives the real _handle_slack_message.
Docs describe the rule in its own note instead of the
ignore_other_user_mentions tip, and the messaging index documents the
human-turn fallback.
Since 5ea8fb2b78 (#111624, for #110952) the gateway rejects a bare silence
marker on any human turn and delivers "The model returned only a silence
marker for a message that needed a reply" instead. That protects a human
who asked this bot something and got nothing back. It also fires on every
human message the adapter admitted without the bot being addressed at all:
a free-response channel, a thread follow-up under
`thread_require_mention: false`, or a message @-mentioning another person
or bot with `ignore_other_user_mentions: false`. A bot whose SOUL declines
peer-addressed turns with a deliberate marker now posts that notice on
every such message. A fleet running several bots in shared Slack threads
reported it as spam on v2026.9.21. #37940 established that intentional
silence must not be re-inflated. Both contracts hold once the turn knows
whether a reply was expected.
`MessageEvent.reply_expected` (True, False, None) is set by the adapter
where the message is admitted. Slack (`slack_reply_expected`): a 1:1 DM,
an @mention of this bot or a command is True, anything else it admits is
False. Other adapters leave None, which keeps today's behaviour, so nothing
changes for them until they are ported. `response_filters.silence_allowed`
holds the one rule (machinery turn, or reply not expected) and both call
sites use it: the live turn in `run_turn._hmwa_shape_agent_response` and
the crash-recovery redelivery from #120377 (1136f135dd), which reads the
flag back from the persisted turn metadata. The suppressed case logs one
DEBUG line naming platform and chat.
Operator workaround until this lands: `platforms.slack.extra.
ignore_other_user_mentions: true` drops peer-addressed messages before a
turn exists.
(cherry picked from commit 094439776ab898cccde303a1c2c911c8ab5bfb75)
Four surfaces build a throwaway AIAgent and never call close() — the
owner boundary that releases memory-provider sessions, tool
subprocesses and httpx clients. In long-lived processes each run leaked
all of them until exit:
- batch_runner._process_single_prompt: one agent per prompt, N prompts
per batch process.
- feishu_comment._run_comment_agent: one agent per comment run in the
gateway process.
- tui_gateway prompt.background: one side agent per background turn.
- cli /bg: one agent per background task in the CLI process.
Wrap each run in try/finally with a suppressed close(), mirroring
gateway/run.py's owner pattern. preview.restart stays deliberately
unclosed (its task exists to leave a detached server running), and the
prompt.background side agent is safe to close: its session_id is the bg
task id, so close() reaps only its own task resources.
Fixes#50197
- A mid-turn notify reply (/status, /approve, clarify answer) shares the
stream's thread key; it no longer seals and overwrites the half-streamed
answer. In-place replacement now requires the final to match the stream
after normalizing mrkdwn markers and whitespace; anything else posts fresh
and leaves the stream open.
- One _commit_stream helper for both seal-then-commit paths, so the rewrite
path also falls back to chat.update when stopStream fails.
- A stream reopened after a server-side seal is seeded with only the text
past the sealed message (tracked as 'base'), not the whole segment.
- Streams older than 15 min are sealed and dropped on the next start.
- _stream_key reuses _workspace_thread_key/scope_id_for_chat; the stream
dict no longer duplicates chat/team ids.
5648f81431 fixed this exact server-side seal (Slack closes a native stream
after a few minutes of a long turn, live-observed at ~5m20s; the lifetime
is not documented) for the native task-card stream: on
message_not_in_streaming_state from appendStream, drop the dead ts and
start a fresh stream, seeded with the full current content so nothing is
lost.
send_draft — the plain-text native streaming path used when task cards
are not enabled — hits the identical seal but never got the fix: its
generic except block only recognizes the feature-gate markers
(not_allowed, missing_scope, ...) and otherwise just logs debug and
returns failure. gateway/stream_consumer_transport.py's
_send_draft_frame() docstring is explicit that "any failure permanently
disables drafts for this run" — so a long turn streaming as plain text
degrades to the edit-based fallback for its remainder exactly the way
the task-card bug did before 5648f81431.
Mirror the task-card fix: on message_not_in_streaming_state from
chat.appendStream, drop the dead ts and _start_stream() a fresh one
seeded with the full accumulated text (not just the delta), so the next
frame's delta still resumes correctly. One reopen per frame; a second
rejection propagates as a real failure, matching the twin's behavior.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 2b4ff4e23bdf684d2fde1a9512f11dcd7ef99c42)
A mrkdwn-rewritten turn-final (e.g. *Done:* -> _Done:_) no longer continues the
streamed text, so it was classified unrelated: the stale stream was sealed and
send() posted a second message (#95430 cause B). Seal, then chat.update the
sealed ts with the final; post fresh only if the in-place update fails.
Co-authored-by: liguoyu <guoyu.li@lcfuturecenter.com>
Problem: with native streaming (chat.startStream/appendStream/stopStream)
the same answer could land twice in a thread — once as the streamed
message, once as a fresh chat.postMessage — while the streamed message
kept its live-typing indicator.
Mechanism: `_try_finalize_stream` matched the turn-final against the
streamed text with a raw `startswith`. The agent strips `final_response`
and joins footers with `rstrip()`, so any surrounding-whitespace
difference made the finalize fall through to a plain post although the
open stream already showed the whole answer (and was never sealed). A
`chat.stopStream` failure took the same fresh-post path even when the
streamed text equalled the final. Streams were also keyed per `chat_id`
only, so two concurrent turns in two threads of one channel sealed or
overwrote each other's stream.
Fix:
- Key native streams per `(team_id, chat_id, thread_ts)`; the stream
consumer stamps the same `thread_id` on every draft frame and on the
turn-final `send()`, so both resolve to the same key.
- Honor the streaming contract (gateway/AGENTS.md): sends carrying
`_interim_send` or `expect_edits` never seal a stream.
- Classify the final against the streamed text as equal / extends /
unrelated with edge-whitespace tolerance (`_stream_relation`). The
stopStream delta is sliced from the RAW final, so nothing inside the
answer (blank lines, fences, tables) is dropped or repeated.
- Commit rule: one `chat.stopStream`, one retry only when no tail is
appended (`markdown_text` APPENDS, so an ambiguous failure must not
repeat it), then `edit_message(finalize=True)` on the stream ts as the
idempotent in-place commit — it already owns format/truncate/Block Kit
and the block-rejection retry. Only when both fail does `send()` post
a fresh message (a duplicate beats a lost answer).
- Oversized tails and rewritten finals (`notify=True`) seal the stale
stream on what is visible before falling back, so no stream is left
with a live-typing indicator.
- `_seal_stream` takes the exact unsent delta instead of recomputing it
from `final_text`; `disconnect()` and the stream API calls route
through the stream's own team client.
Tests: tests/gateway/test_slack_native_streaming.py covers the
whitespace-only difference, the stopStream-failure commit path, the
bounded retry, the uncommittable fallback, interim/preview sends,
per-thread keying, oversized tails, rewritten finals and the
GatewayStreamConsumer end-to-end path.
(cherry picked from commit a64d10071ce7816b124467e29407d7d47bfdde8d)
dingtalk-stream 0.24.3 start() retries forever inside the SDK, logging a
malformed logger.exception() every 3s, so the adapter breaker never saw the
error. Install a dedup filter on the SDK logger that can't raise on bad
format args, detect the websockets incompatibility (bare or chained
TypeError), log one ERROR with a pin-consistent hint, and hand off via
_set_fatal_error(retryable=False) + _notify_fatal_error(): only a
reinstall of the pinned versions and a restart fixes it. Breaker stays
tripped until the error type changes; constants moved to module level.
The reconnect storm in #24851 is driven by a dingtalk-stream/websockets
incompatibility that raises TypeError on every start(); backoff never
recovers it. Log the first occurrence per error run at ERROR with an
upgrade hint instead of a generic WARNING.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
DingTalk stream-mode reconnection storms the gateway: when start() raises
the same error every cycle, _run_stream logs a WARNING and reconnects at the
60s cap forever, generating hundreds of MB of identical log lines and hanging
the gateway.
Add a per-error-type circuit breaker: after 5 consecutive identical errors,
suppress the repeated WARNING (one ERROR summarises) and pause 300s instead of
spinning at 60s. Reset backoff + counters after a clean start() so a recovered
connection is treated fresh.
Salvage of #24881 re-implemented on current main: the original fix targeted
gateway/platforms/dingtalk.py, which has since been refactored to
plugins/platforms/dingtalk/adapter.py. Same logic, new path; tests import the
new module.
Closes#24851.
(cherry picked from commit 84201c58255ae3d3c9b269ac777ea0ff444d9229)
Surface code/msg at debug when the recall API returns non-success
(matches the exception path) and trim the docstring. Drop the trivial
disconnected-client test.
Feishu had no delete_message, so a failed finalize-edit plus fallback
send left the truncated edit bubble next to the full final. Implement
the SDK delete and thread the fallback send to the originating message.
(cherry picked from commit c61add84ad40b1bc39288405a0a05b4f621e00fc)
a74e0155b6 made attachments[].blocks[] reach the agent through
_append_link_unfurls, but rendered each attachment's blocks with no ceiling.
Slack allows 20 attachments per message, so one alert could project 20x what
a single attachment does (measured: 3,247 chars for 1 -> 64,855 for 20 with
8x400-char rich_text sections each), while the top-level blocks path caps
once at 6000.
Share one budget (_SLACK_UNFURL_BLOCKS_MAX_CHARS, the same 6000 the top-level
path uses) across the array: the first attachment keeps its body, later ones
are truncated against the remainder, and a spent budget still leaves every
header visible. After: 3,247 -> 6,659 chars at 20 attachments.
(cherry picked from commit 3efc1e4c532c38a220d4c6ea53745d59d334aa3e)
Telegram's per-(chat_id, status_key) status-message cache grew without
bound; give it the same _STATUS_MESSAGE_IDS_MAX=2000 FIFO half-trim the
Slack adapter already has. In both adapters, guard the post-await write-back
after a successful edit with a compare-before-write (only re-store the id if
the cached entry is still the one we edited) so an eviction or replacement
that happened during the await is not undone.
Partial salvage of #87480: kept the Telegram bound and both compare-before-write
guards (re-applied by hand, 17586 behind), defined the max as a class attr
like Slack instead of an instance attr, dropped the 4 new tests.
(cherry picked from commit 43ae95e98e)
_handle_events called _save_cursors inline whenever a batch moved the
cursor. It ends in atomic_json_write (mkstemp + fsync + os.replace), and
on the WebSocket transport _handle_events runs once per inbound EVENT
frame, so every message paid an fsync on the gateway's event loop,
stalling every other adapter and in-flight turn for its duration.
Split the snapshot from the write: the payload is still built on the
loop (_channel_state is loop-owned), and the write goes through
asyncio.to_thread. The snapshot is taken under an asyncio.Lock so a
slower, older write can never land after a newer one and regress the
durable cursor. connect() keeps the synchronous _save_cursors.
(cherry picked from commit 299929741429262f497b86a67940ffb37faa1694)
Runtime identity resolved through hermes_cli.__version__ (a static 0.0.0
on source installs, rewritten by release stamping) leaked v0.0.0 into
About, /api/health, User-Agents, and plugin compat, and source updates
showed "couldn't reach update server" because identity and channel
authority disagreed with the checkout.
Now: get_version_info() resolves install stamp -> live git -> unknown,
never pyproject metadata, never a package constant. Source checkouts
derive identity from their reachable release tag; the completion tail of
every successful install/update/historical takeover atomically rewrites
install-stamp.json with that identity; a stale source stamp whose commit
no longer matches HEAD defers to live git. ACP/TUI use derived_version
for display and base_version for protocol fields; all ~44 runtime
__version__ consumers migrated; hermes_cli.__version__ and generated
_version.py are gone; release stamping only touches the native manifests
external builders consume (nix/tauri/cargo) and passes release identity
straight into write_install_stamp.py; pyproject.toml stays inert 0.0.0.
Desktop no longer synthesizes a competing install-stamp.json: the
checkout owns its stamp, and desktop-bootstrap classification keys on
the bootstrap-complete marker. verify-bootstrap-version-stamp.py now
cross-checks the checkout's stamp (baseVersion + commit == HEAD).
Validation: 31-file focused suite green (version identity, stamping,
adoption, providers, gateway, acp/tui runtime identity, api server via
extras env, release graph); desktop tsc + 25 vitest green; real-repo
probe: base=unknown derived=git.0635606.dirty source=git on this
checkout; clean-env imports resolve entirely from this tree; windows
footgun + compat-pointer scans clean.
EmailAdapter._sender_accepted runs before any MessageEvent exists and
read only EMAIL_ALLOWED_USERS. Unset, it dropped every sender unless
allow-all was on; set, it dropped everyone not listed. The gateway's
own handling therefore never ran for email:
platforms.email.unauthorized_dm_behavior "pair" (the setup wizard's
"Use DM pairing") and "decline" sent nothing, and a sender admitted by
GATEWAY_ALLOWED_USERS or an approved pairing was dropped. bb304b4914
turned the empty-allowlist branch into drop-all after #50568 had made
"pair" email's explicit opt-in.
The gate now keeps a sender listed by address in EMAIL_ALLOWED_USERS
or GATEWAY_ALLOWED_USERS, a sender the registered gateway
authorization check admits (that is the only reader of the pairing
store), and, under an explicit pair or decline, an unknown sender the
gateway will answer. The default "ignore" still drops unknown senders
before a MessageEvent exists, so the mail-loop guard from fd9c32c0f2
holds.
Three guards keep the wider gate from widening access, and close two
forged-From: paths main already had:
- A sender admitted only so the gateway can answer it (pair or
decline) must authenticate its From:, open access or not: the
pairing code or refusal is mailed back to that address. A granted
sender still needs it short of open access, since a pairing grant
keys on From: just as the allowlist does. Open access follows the
gateway's own order: EMAIL_ALLOW_ALL_USERS wins over a list, while
GATEWAY_ALLOW_ALL_USERS beside a list admits nobody extra, so it no
longer exempts a listed address from From: authentication either
(on main a forged From: of a listed address got through there).
- Open access comes from the gateway's own verdict when a check is
registered. GATEWAY_ALLOW_ALL_USERS beside a GATEWAY_ALLOWED_USERS
list grants a stranger nothing there, so the env flag alone no
longer exempts one from From: authentication (that path mailed a
pairing code to a forged From: on main too).
- A sender whose local part alone matches an allowlist entry is
dropped. The gateway's check also matches an address by its bare
local part (#119446), so without this, GATEWAY_ALLOWED_USERS=alice
(a chat username) would admit or pair alice@<any domain>. The lists
are parsed as the gateway parses them, JSON list literals included,
or '["alice"]' would slip past this guard.
_allowlist_in_effect only served the old condition and is removed.
The scope tests now assert the same scoped reads through
_sender_accepted, with GATEWAY_ALLOWED_USERS covered as well.
Measured end to end with the real GatewayRunner callback wired
(adapter -> gateway ingress):
- pair, decline, GATEWAY_ALLOWED_USERS and an approved pairing each
went from 0 events reaching the gateway to 1. pair mails a pairing
code, decline mails one refusal.
- An unauthenticated From: in pair mode, for a paired address or for a
GATEWAY_ALLOWED_USERS address still reaches nothing.
- A bare GATEWAY_ALLOWED_USERS=stranger entry lets nothing from
stranger@<domain> through, under ignore or pair. Without the
local-part guard that mail reached the gateway in both.
- The same holds for a JSON-literal list, and a pair-mode stranger
with a forged From: under allow-all beside an EMAIL_ or
GATEWAY_ALLOWED_USERS list reaches nothing.
- The default still drops.
Completed update IDs lived only in the adapter's memory. The gateway
reconnect watcher builds a new TelegramAdapter and connects it with
is_reconnect=True, which keeps Telegram's pending queue, and a new PTB
Updater polls from offset 0. Telegram then resends every update whose
acknowledgement (the next getUpdates offset, or the cleanup call in
Updater.stop) never landed, and the fresh adapter admitted them again.
Write completed IDs to telegram_update_receipts_<bot_id>.json in the
adapter's Hermes home and seed admission from it once per bot. Receipts
older than 24h are dropped: the Bot API keeps unconfirmed updates no
longer than that, and it keeps the lookup clear of the random ID restart
Telegram may do after a week without updates. Writes are coalesced and
run off the loop; disconnect waits for the last one.
Refs #68502
Co-authored-by: Joe Githler <5716896+NoTimeforInfinity@users.noreply.github.com>