The base media dispatch calls send_voice(..., is_voice=is_voice) for every
audio MEDIA attachment (gateway/platforms/base.py _send_one). MatrixAdapter
.send_voice() accepted neither is_voice nor **kwargs, so every non-image
MEDIA delivery raised TypeError and the file was silently dropped — the
failure is visible in rotated logs since 2026-09-08 (never worked).
Accept the flag explicitly: is_voice=False -> plain m.audio in the original
format (no transcode); True or omitted (play_audio legacy callers) -> the
existing MSC3245 voice-bubble path with best-effort Ogg/Opus transcode.
Fixes#116776
(cherry picked from commit d4f89a725498e29a2ee0fe016f8bb98db08f358c)
_resolve_room_identity() classifies any room with <=2 joined members as a DM
regardless of m.direct or an explicit room name, so those rooms silently
bypass MATRIX_ALLOWED_ROOMS, MATRIX_FREE_RESPONSE_ROOMS, and
MATRIX_REQUIRE_MENTION, and use DM threading instead of
MATRIX_AUTO_THREAD/MATRIX_SESSION_SCOPE. This was previously only visible in
an inline code comment, not in the env-var docs an operator would read.
Fixes#114733
_schedule_invite_join already requires both is_direct and inviter before recording m.direct, so `is_direct and bool(inviter)` at the reconcile call site was redundant; pass is_direct through. The `if is_direct and not inviter` WARNING could only be reached with GATEWAY_ALLOW_ALL_USERS set and a spec-violating stripped m.room.member event lacking `sender`; the info log already prints is_direct, so drop the branch.
The adapter comment and the test module docstring claimed a reconciled pending invite "never fires _on_invite". It does: _absorb_sync runs _dispatch_sync (which emits INVITE to _on_invite) and then the reconcile pass over rooms.invite, which joined every entry unconditionally — so a live invite _on_invite rejected was joined ms later, and invites that arrived while the gateway was down were joined on restart with no gate. Reword both to state that premise.
_on_invite only auto-joins a room when the inviter is allow-listed (or
GATEWAY_ALLOW_ALL_USERS is set), so a live invite from an arbitrary
federated user is rejected. A pending invite that arrives while the
gateway is down takes a different path: _schedule_pending_invite_joins
reconciles it from rooms.invite in the sync response and scheduled the
join unconditionally. An unauthorized invite sent during downtime was
therefore auto-joined on restart, bypassing the allowlist.
Extract the gate from _on_invite into _is_authorized_inviter and apply
it during reconciliation too, reading the inviter from the stripped
invite state (the sender of the m.room.member event for our own user,
as _extract_invite_dm_signal already does for the DM signal). An
inviter that cannot be read from the invite state fails closed, exactly
like an empty sender in _on_invite: the invite is skipped with a
warning and left pending.
A direct invite that arrives while the gateway is running fires
_on_invite, which passes is_direct and the inviter through
_schedule_invite_join so the room is recorded in m.direct after the
join. An invite that is still pending across a gateway restart takes a
different path: _schedule_pending_invite_joins reconciles it from
rooms.invite in the sync response, but called _schedule_invite_join
without is_direct or inviter. The DM signal was dropped, the room was
never recorded in m.direct, and it was classified as a group until the
user's own client happened to update m.direct.
Read the signal from the stripped invite state instead: the
m.room.member event for our own user carries the original invite's
is_direct flag, and its sender is the inviter. Thread both through to
_schedule_invite_join so a reconciled direct invite is recorded in
m.direct, and thus lands in _dm_rooms, exactly like a live one.
This gap was surfaced by the triage of #62493.
`_handle_media_message` fetched the mxc:// payload (`_download_and_cache_media`)
before `_build_inbound_event` ran `_resolve_message_context`, so an unmentioned
or non-allowlisted room's m.image/m.file was pulled onto the host and then
dropped. Resolve the context first and hand it to `_build_inbound_event` via a
new `ctx` kwarg so the read receipt / thread mark are still applied once.
Same class as the Teams fix in this PR (review follow-up). One invariant test:
unmentioned group media → 0 downloads, mentioned → 1.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
A Matrix `attach_to_session` cron delivery and a CLI→Matrix `/handoff` now land in their own
thread AND a human reply in that thread continues the seeded session (#112918).
The salvaged `MatrixAdapter.create_handoff_thread` (ea-s21, #108367) gives both seeders a
thread root. The second half of the bug is a key-shape mismatch: the handoff watcher and
`_seed_cron_thread_session` seeded `matrix🧵<room>:<root>` while the adapter keys every
in-thread reply on the ROOM's type (`matrix:group|dm:<room>:<root>`), so seed and reply never
met. Following the direction landed for Slack in 1f3f45e87b (#111896), the seeders now mirror
the adapter instead of the adapter moving onto a `thread` slot: rekeying inbound Matrix threads
would orphan every existing Matrix thread session and drop `is_group` in authz for in-thread
messages (`_GROUP_CHAT_TYPES` has no `thread`).
- gateway/run_startup.py: Matrix destinations key `dm`/`group` from the adapter's
`get_chat_info` (`_handoff_home_is_dm`).
- cron/scheduler_delivery.py: `_THREAD_REPLY_CHAT_TYPE` — Slack and Matrix non-DM thread seeds
use `group`. Slack channel cron threads had the same mismatch (adapter `build_source` keys
`group`; the seed said `thread`); the two cron tests that pinned `thread` for Slack channels
asserted the wrong shape and now assert the adapter's.
- Dropped from #108367: the `chat_type="thread"` inbound rekey (see above) and its tests;
contributor tests trimmed to two invariants.
- Docs: Matrix listed among thread-capable handoff/cron platforms.
Live probe (in-process, real MatrixAdapter + real handoff destination + real cron seeder):
before create_handoff_thread -> None; handoff key matrix🧵… ≠ inbound matrix:group:…;
cron seed matrix🧵… ≠ inbound
after create_handoff_thread -> '$seed'; handoff == inbound == cron seed (room and DM cases)
Co-authored-by: ea-s21 <190767603+ea-s21@users.noreply.github.com>
The Matrix adapter inherited the base create_handoff_thread (returns None), so
cron attach_to_session and the gateway handoff watcher silently no-op on Matrix
while Telegram/Discord/Slack support them: the alert/handoff is delivered flat
instead of as a replyable thread, and the scheduler logs thread_id=None ->
"Mirror: no session found". A human reply then can't resume the session with
its seeded context.
Implement it Slack-style. Matrix has no channel-level create-thread API -- a
thread is just events whose m.relates_to / rel_type: m.thread reference a root
event's event_id -- so seed a root message and return its event_id as the thread
handle. The implementation reuses the adapter's own send() (chunking + E2EE
retry + formatting) and registers the root with self._threads.mark(), mirroring
inbound thread handling; _apply_relation_metadata already threads later sends off
a supplied thread_id, so the returned id is immediately usable. Returns None on
missing client / failed send (callers fall back to parent_chat_id).
Adds tests/gateway/test_matrix_handoff_thread.py (seed id returned; default seed
on blank name; None on no client, failed send, and missing event id).
The reply-pill fix split the body on a leading "> " so the strip would not
rewrite the "> <@bot:srv>" pill. That keyed the exemption on the body shape,
not on the relation: a hand-typed blockquote in a plain (non-reply) message
that mentions the bot inside the quote reached the agent with the raw
"@hermes:example.org" text, which main used to strip. Split around the quote
only when m.relates_to carries m.in_reply_to; otherwise strip the whole body.
Review finding: quote-block exemption keyed on body.startswith('> ') instead of m.in_reply_to.
The Matrix adapter strips the bot's mention from the inbound body before
_extract_reply_context parses the inline reply fallback, and that strip is a
blind whole-body replace. A reply to the bot names the bot in the fallback
pill ("> <@bot:server> quoted"), which is exactly what makes the message
count as a mention under the default MATRIX_REQUIRE_MENTION=true -- so the
strip runs on every reply-to-the-bot and rewrites the pill to "> <>", after
which the pill regex no longer matches. reply_to_author_id is lost,
reply_to_text becomes the mangled "<> quoted" remnant, and the prompt
renders "[Replying to: "<> ..."]" instead of "[Replying to your previous
message: ...]".
Split the body into (quote block, reply text) and strip the mention from the
reply text only. The mention gate still sees the raw body, so a reply to the
bot keeps waking the bot; the visible reply text and the quote-block strip
are unchanged.
Fixes#111233
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
The stack inserted the errcode table directly after _strip_reply_fallback's
return with no blank lines, which reads as if the constant belongs to the
function body and trips E305. Two blank lines restore the module-level
boundary; no behaviour change.
The inherited fixture coordinate "40.4302" does not contain the substring
"403" (the dot splits it), so the old substring classifier also passed on
it and the case proved nothing. Use a coordinate that genuinely embeds the
digits so the test is red on the pre-fix classifier.
The pinned mautrix 0.21.1 raises MatrixRequestError (carrying errcode and
http_status) from HTTPAPI._send on every non-2xx and sync() returns only
the parsed JSON dict, so the result-object auth branch in _sync_loop was
unreachable; it dated from the nio client whose SyncError objects were
real. Drop it together with the nio-mock test that pinned it.
With structured attributes guaranteed, the leading-status regex and the
bounded keyword scan over the message text were the only remaining ways
for body digits or HTML words to leak into the verdict, so drop them too:
no errcode/http_status auth signal means retry. Trim the contributor's
17 tests to the two loop-level invariants: both production repros (502
HTML body embedding "403" via an SVG coordinate; timeout echoing a since
token embedding "401") keep looping, and a 401/M_UNKNOWN_TOKEN stops.
The comment above the result-object branch in _sync_loop claimed mautrix's
Client.sync() returns an object carrying a message string for auth failures.
That is wrong. In the pinned mautrix 0.21.0, HTTPAPI._send raises
make_request_error() for any non-2xx and otherwise returns parsed JSON, so a
real M_FORBIDDEN arrives as an exception and is handled by the except branch.
The claim was introduced by this PR, which rewrote an accurate comment about
the earlier matrix-nio client (whose SyncError result objects were genuine).
The branch itself is kept as defense in depth against a future client swap,
but it now classifies with the same errcode/http_status logic as the
exception path instead of a lone "unknown_token" substring test, which
silently missed M_MISSING_TOKEN and M_FORBIDDEN and resynced forever
against a credential that can never succeed.
A structured errcode/http_status is authoritative; the message text is only
consulted when the object exposes neither, since str(object) is an opaque
repr. The text scan deliberately cannot override a structured verdict, so a
transient 502 whose HTML body contains "Forbidden" is still retried.
Adds four tests. Three are discriminating RED/GREEN cases that fail against
the old substring branch (M_MISSING_TOKEN errcode, http_status=401 with no
keyword in the message, and an unstructured object whose only signal is
.message). The fourth pins the precedence rule and passes either way.
Verified: 136 passed / 1 failed in tests/gateway/test_matrix.py; the single
failure (test_password_login_uses_device_id) fails identically at the
pristine PR head and is unrelated.
(cherry picked from commit bc9e6a8dafcf349a4e6b20a261fb2449603c0239)
The independent-verifier caught that my first loop-level test did not
actually prove anything. The 502/SVG coordinate fixture I reused from
gmoranxyz's unit-level test does not contain the substring 403 once
case-folded, so the old naive substring classifier already treated it
as transient. A test that passes under both the buggy code and the
fix proves nothing about the fix.
I replaced the fixture with a plain connection timeout whose message
wraps the real Matrix sync pagination token, an arbitrary digit
string that happens to contain 401. I verified this directly: with
the pre-fix classifier restored, the retry test now fails (the old
code stops the loop on this fixture), and with the fix in place it
passes (the loop retries as it should). That is the RED/GREEN proof
the maintainer originally asked for.
I also documented in the stop test's docstring that it does not
discriminate old from new, since the word forbidden in its message
trips the old naive check too. It is still worth keeping as a
regression test proving genuine auth errors stop the loop, just not
as proof of this specific fix.
While I was in there I also fixed a stale comment above the
M_UNKNOWN_TOKEN sync-object pre-check. It said nio returns SyncError
objects, but the dependency here is mautrix, not matrix-nio, and
importing nio raises ModuleNotFoundError in this codebase. The
pre-check logic itself was already correct and untouched.
Co-authored-by: gmoranxyz <gmoranxyz@users.noreply.github.com>
(cherry picked from commit ad3aad579a675a5aae544a50f88a82717c0ac3b6)
I added two more classifier unit tests for the attribute narrowing:
a bare .code attribute that happens to be 401, and a bare .status
attribute that happens to be 403, both must stay classified as
transient since only .http_status is trustworthy. I also added a
parametrized test for the five transient exception types the sync
loop now short-circuits on.
On top of that I added two tests that exercise _sync_loop directly
instead of just the classifier function in isolation. One replays the
real 502 Umbrel repro string through a mocked client.sync and confirms
the loop retries with the 5s backoff. The other raises a genuine
M_FORBIDDEN error and confirms the loop stops on the first call with
no retry sleep. These catch a regression in how the loop wires the
classifier in, not just a regression in the classifier itself.
(cherry picked from commit f747bb4b5a6e6da1bb9136168f08d6e7af5ea64b)
I hit a bug where the Matrix sync loop treated a passing 502 from
Umbrel's app proxy as a permanent auth failure and stopped syncing for
good. The old check did a naive "403" in str(exc) substring match, and
the 502 HTML error body embedded an SVG path with the coordinate
40.4302, which contains the digit sequence 403.
I replaced the substring check with a layered classifier. Transport
exceptions like TimeoutError, ConnectionError, and OSError are always
treated as transient regardless of their message text. Structured
signals take priority next: the errcode attribute against a known set
of permanent Matrix error codes, then the http_status attribute
against 401/403 specifically (not status, status_code, or code, which
belong to unrelated exception shapes and risk coincidental integer
matches). Only when none of those are present does it fall back to a
bounded, word-boundary-safe text scan on the first 200 characters.
Added tests covering the attribute narrowing, the transient exception
types, and two loop-level tests exercising _sync_loop directly to
confirm it retries on a transient error and stops on a genuine 401/403.
(cherry picked from commit 96d3363e45a63e08d9f07518ed949df334a33b3c)
One reader (gateway.platforms._shared.extra_or_secret) now implements the
precedence every per-profile setting follows for the OWNING profile:
explicit scoped env/.env → that profile's config.yaml (PlatformConfig.extra)
→ the adapter's default. A scoped miss returns the default, never the launch
process's os.environ; single-profile / default-profile installs keep the
documented env-over-YAML contract.
Why: 545e74d0ea (#108705) stopped bridging a secondary's YAML into the
process env and moved readers to config.extra, but the shared reader and the
hand-rolled helpers in Discord/Slack/Matrix/Telegram consulted YAML FIRST and
then fell back to a scoped env read. Two bug classes followed (#108440
post-merge review by andrexibiza, #109032):
- an explicit env value could no longer beat YAML for the owning profile
(DISCORD_ALLOW_MENTION_EVERYONE=false lost to allow_mentions.everyone: true;
TELEGRAM_REACTIONS=true lost to the stock reactions: false);
- a secondary that OMITTED a key inherited the launch profile's bridged env
through the fallback (Matrix process_notices/session_scope, Discord
auto_thread/reactions/mentions, Slack reactions/ignored_channels).
Consumers migrated to the shared reader: Discord _build_allowed_mentions and
_extra_or_env_flag; Slack _slack_allow_bots, _reactions_enabled (the
_extra_or_env_* getters already used it); Matrix _extra_truthy, _extra_csv_set,
session_scope, reactions, require_mention parsers, and — new — the
allowed_users / ignore_user_patterns consumers that never read the seeded YAML
lists; Telegram _extra_bool, _extra_str_set, _reactions_enabled; Feishu
allow_bots; WhatsApp dm_policy/group_policy.
Refs #108440, #109032
#69090 scoped MATRIX_RECOVERY_KEY itself (via _scoped_recovery_key())
so a secondary profile resolves its own recovery key under multiplex,
but left its sibling, MATRIX_RECOVERY_KEY_OUTPUT_FILE, on a bare
os.getenv(). _recovery_key_output_path() is called from inside
_verify_or_bootstrap_cross_signing(), which runs fully inside
_profile_runtime_scope for a secondary profile: when that profile
bootstraps a new recovery key, it either doesn't get written to a
file at all, or gets written to the default profile's configured
path, depending on which one has the env var set.
Route it through the same _get_scoped_secret() helper _scoped_recovery_key()
already uses.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit fb765ee49b2f1a1853e52fb901d25f767180a2fc)
545e74d0 (post-0.21.2) made the Matrix YAML bridge seed its values into
PlatformConfig.extra so secondary multiplex profiles read their own config. The
"csv" bridge kind seeds any non-None value, so `free_response_rooms: ''` now
reaches extra as '' — and the readers' `if raw is None` fallback no longer fires,
so MATRIX_FREE_RESPONSE_ROOMS is ignored and require_mention drops every
un-mentioned message. Before that commit the bridge only wrote env and the key
was absent from extra, so the env value applied.
Route the three identity-check readers (_extra_csv_set, _extra_truthy,
_resolve_max_message_length — the last a three-tier chain where '' also
short-circuited the plugin-registry default) through the shared
gateway.platforms._shared.extra_or_secret, whose default already treats a blank
string as unset (the idiom mattermost/dingtalk/slack readers use). Explicit
scalars, bools and lists (including []) stay authoritative.
Two invariant tests replace the salvaged suite (moved to
tests/plugins/platforms/matrix/ to mirror the source path): blank falls through
for all three readers; explicit values still beat env.
Fixes#109358
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
Scoped secrets — `gateway.platforms._shared.get_scoped_secret` is the single implementation of
the "scope authoritative, unscoped default-profile falls back to os.environ" read:
- plugins/platforms/buzz/adapter.py::_get_scoped_secret (113 LOC, ~100 of which were one
docstring paragraph pasted 16x) -> 3-line forwarder over the canonical with
`external_fallback=True`. Its one genuine extra rung (one-shot profile-scope build so a
Bitwarden-managed key is visible to the startup gate, #95216) moves into `_shared` as that
keyword plus `_unscoped_profile_secrets`.
- weixin::_wx_secret, matrix::_startup_env_secret, the inline try/except copies in slack
(SLACK_APP_TOKEN) and telegram (TELEGRAM_WEBHOOK_SECRET/_URL) -> canonical.
- The "extra-first, then scoped env" reader written 11x under 6 names (weixin._extra_or_env,
bluebubbles/ntfy/photon/wecom `_setting`, dingtalk `_extra_get`, mattermost `_extra_or_env`,
slack `_extra_or_env_flag/_channel_set`, feishu closures) -> `_shared.extra_or_secret`.
- `authz_mixin._platform_gate_env` -> `_shared.platform_gate_env`; discord/telegram drop their
`_scoped_gate_env` twins; run.py / run_config_loaders.py / slack import it directly.
Boilerplate — three table-driven helpers in `_shared` replace the pasted docs template:
- `seed_extra_from_env(spec, home_env=)` replaces 8 `_env_enablement` bodies (buzz, google_chat,
irc, line, ntfy, photon, simplex, teams; raft is a one-liner and untouched).
- `apply_yaml_bridge(cfg, spec)` replaces 7 `_apply_yaml_config` bodies (buzz, dingtalk, feishu,
matrix, mattermost, slack, whatsapp); discord/telegram keep bespoke bridges (alias keys,
nested `platforms.*.extra`, generic-key exclusions). buzz and mattermost previously bypassed
`yaml_env_setter` with hand-rolled `os.environ` writes.
- `env_is_connected(*vars)` replaces 5 identical `_is_connected` (discord, homeassistant,
mattermost, slack, sms).
- 8 identity `_build_adapter` wrappers deleted; `adapter_factory=<Class>`.
Behavior change:
- buzz `_apply_yaml_config` returned None, so under multiplex a secondary Buzz profile got
neither env (correctly skipped) nor `extra` for relay_url/channels/allow_all_users/...; it now
seeds `extra` like every other hook. It also wrote reply_in_thread/reply_to_mode to the process
env even inside a secondary profile's scope (first-writer-wins leak, #80099 class); it no longer
does. BUZZ_POLL_INTERVAL is bridged through the same table.
- `home_channel.name` default when `<X>_HOME_CHANNEL_NAME` is unset is now the literal "Home" for
all plugins (irc/ntfy/buzz used the chat id; simplex/teams/photon/google_chat already used
"Home", as do the built-in platforms in gateway/config_env.py).
- weixin's non-secret tunables (send_chunk_*, rate_limit_circuit_*) now read through the scoped
reader instead of raw os.getenv — a secondary profile no longer inherits the default's values.
- `extra_or_secret` treats a blank string in extra as unset (falls to env) and an explicit False
as a real value, the strictest of the merged copies.
- slack `reaction_trigger_target` bridges via str(); `reaction_triggers` comma-joins any list-ish
value (was list/tuple/set only) — same env text for every real YAML shape.
Docs: website/docs/developer-guide/adding-platform-adapters.md (the template the copies were
pasted from) and gateway/platforms/ADDING_A_PLATFORM.md now show the helpers and the scoped
reader; gateway/AGENTS.md points at the one implementation.
Tests: tests/gateway/test_shared_platform_boilerplate.py — every plugin `_env_enablement`
reads only through the scoped getter (parametrized over the 8 plugins, spy on the seam, raw
`os.getenv`/`get_env_value` asserted untouched); buzz bridge seeds `extra` for a secondary
profile and still bridges env for the default; one home-name rule; extra_or_secret contract;
external_fallback rung. Existing tests repointed: tests/agent/test_secret_scope_tier1_migration.py,
tests/plugins/platforms/buzz/test_buzz_unscoped_requirement_gate.py.
Fourteen platform plugins hand-rolled the "already configured? Reconfigure? [y/N]"
gate at the top of interactive_setup (env check + info line + prompt_yes_no(..., False)),
with drifting wording ("X: already configured" vs "X is already configured." vs
"already enabled") and, for LINE and SimpleX, raw input() loops with their own
EOF/KeyboardInterrupt handling and no gate at all. Fixes to the gate (non-interactive
handling, wording, default) therefore reached only the core Telegram/BlueBubbles/webhook
wizards.
- hermes_cli/setup_platforms.py: `_declines_reconfigure` becomes the public
`declines_reconfigure(label, question, *env_vars)` (any-of env check, so Matrix's
token-or-password gate fits); `_save_prompted` becomes `save_prompted` alongside it.
No alias kept; the three core callers are updated.
- buzz, dingtalk, discord, feishu, google_chat, irc, matrix, mattermost, raft, slack,
teams, wecom: the hand-rolled gate is replaced by one `declines_reconfigure(...)` call;
post-decline extras (Discord allowlist nudge, Slack manifest refresh, Raft "Keeping"
line) stay local and unchanged.
- line, simplex: the raw input() loops move onto hermes_cli.cli_output.prompt (masked
for secrets, "" on Ctrl-C/EOF) and gain the shared gate on their primary env var.
Behavior change: the gate's info line is now uniformly "<Label>: already configured"
(DingTalk/Feishu/WeCom lose the trailing period + inline ID; Buzz/IRC/Google Chat/Raft/
Teams no longer echo the current value in that line). Feishu and WeCom now gate on the
app/bot ID alone instead of ID AND secret. LINE and SimpleX gain a "Reconfigure?" [y/N]
prompt when already configured; their prompts now honour HERMES_NONINTERACTIVE and print
via the CLI helpers instead of bare print(). Prompt defaults (No) are unchanged everywhere.
Not touched: WhatsApp's gate keys on WHATSAPP_ENABLED being truthy (a "false" value must
not count as configured), which the shared any-set gate cannot express — left hand-rolled.
Test: tests/plugins/platforms/test_interactive_setup_reconfigure_gate.py parametrized over
the 14 wizards — with the primary env var set and the user declining, each wizard must have
called declines_reconfigure with that var and returned without prompting or saving.
Sabotage: reverting mattermost's gate fails that row.
Matrix, WhatsApp Cloud and the TTS tool each ran their own ffmpeg argv for the same
speech-tuned libopus encode; they predate the shared helper and never migrated, so the
codec flags, timeout handling and error reporting drifted (matrix 48k/30s, whatsapp_cloud
async subprocess with no timeout and no `-ac 1`, tts an in-place sidecar repair).
`transcode_to_ogg_opus` gains `timeout=` and `output_path=` (sibling-file and in-place
writes go through a `.tmp.ogg` sidecar so a failed encode never truncates the source).
Deleted: matrix `_matrix_transcode_voice_to_ogg`, tts `_ffmpeg_transcode_to_opus`;
whatsapp_cloud `_convert_to_opus` keeps only its warn-once ffmpeg install hint and calls
the helper via `asyncio.to_thread`. Matrix and tts keep their 48k bitrate.
Behavior change: whatsapp_cloud transcodes now use mono (`-ac 1`), `-compression_level 10`
and a 60s timeout like every other voice bubble; a failed encode logs at WARNING for all
three sites (matrix previously DEBUG).
20 `plugins/platforms/*/adapter.py::_standalone_send` paths (the out-of-process cron /
send_message delivery) built `{"error": f"... {e}"}` by hand — 83 literals. The exception text
of an httpx/aiohttp failure can carry the Authorization header, a signed URL or a response body
with the token in it, and that string became the tool result the model reads. Only sms went
through the redacting `tools.send_message_senders._error`; discord kept a private regex that
only knew `Authorization: Bot`.
`gateway.platforms._shared.send_error(message)` wraps that helper (agent.redact +
URL-secret scrub) and every standalone literal now goes through it, including the three
envelopes that carry extra keys (discord warnings, photon error_class/retryable, whatsapp's
`(None, err)` tuple). The sms and discord local wrappers are deleted. Telegram already
delegated to the core sender and is untouched.
Behavior change (security): vendor exception text in standalone-send failures is redacted
before reaching the model.
Nine surfaces (feishu, teams, slack, telegram, whatsapp_cloud, qqbot, matrix, discord, relay)
each re-derived the approval choice set — [Allow Once]; session + always unless smart-denied;
[Deny] — and four of them (discord, slack, teams, whatsapp_cloud) never adopted
base._format_exec_approval, so header/reason/smart-deny wording and truncation budgets drifted
per adapter. Three separate commits had to touch 5–9 adapters for one semantic fix.
BasePlatformAdapter.send_exec_approval now builds an ExecApprovalPrompt (shared text via
_format_exec_approval, shared `(label, choice, style)` rows via _exec_approval_actions) and
hands it to the `_send_exec_approval_prompt` hook. Each adapter keeps only its widget mapping
(~10–20 LOC); platform wording stays via the existing `_EA_*` class attrs, and a new
`_exec_approval_cmd_budget` hook lets Slack/Discord budget the command against their hard
message caps (3000-char section / 2000-char message) instead of computing it inline.
`_EA_REASON_BUDGET` covers Slack's 500 / Discord's 300 reason caps.
The runner used to detect button support by `hasattr(type(adapter), "send_exec_approval")`;
that is now true for every adapter, so `_renders_exec_approval_buttons` asks
`supports_exec_approval_buttons()` (hook overridden?) and keeps the duck-typed check for
non-BasePlatformAdapter classes.
Visible text changes (button semantics unchanged everywhere):
- Discord: the smart-deny line now follows the reason (was inside the header before the
fence); the truncation marker is "..." not "\n... [truncated]".
- Slack: smart-deny line follows the reason instead of the header.
- Teams: unchanged (same 2000-char preview, same smart-deny block).
- WhatsApp Cloud: identical text; body still capped at 1024.
- QQBot/relay: unchanged.
Eight adapters (discord, telegram, wecom, matrix, whatsapp, simplex, feishu, weixin) each kept a
copy of the delayed text-batch flush that base._enqueue_text_event schedules. Two correctness
fixes had landed in single copies only: Discord's asyncio.shield around the dispatch (#12444 —
a late chunk cancelling the flush task aborted the in-flight agent turn) and WeCom/Weixin's
synchronous task-identity check before the pop (a superseded task waking late popped the event
and the successor found nothing). The other adapters carried both bugs latent.
The base now owns `_flush_text_batch` with both fixes, plus the `_pending_text_batches` /
`_pending_text_batch_tasks` dicts and `_SPLIT_THRESHOLD` / delay attrs (defaults; adapters set
their own). Platform policy goes through three small hooks instead of a copied body:
`_text_batch_delay_for(pending)` (Telegram's fast/short tiers, WeCom's attachment-only wait),
`_pop_text_batch(key)` (Feishu's side count table) and `_dispatch_text_batch(event)` (Feishu's
per-chat lock). Telegram keeps its `_flush_buffered` body because its contract differs on
purpose — a cancel after the pop must hold-and-re-raise so teardown can stop a flush — and gains
the identity check there. Matrix's `_split_threshold` is renamed to the shared `_SPLIT_THRESHOLD`;
SimpleX exposes its single delay through the shared attr names.
`helpers.TextBatchAggregator` (zero users) sits inside the revert-scheduled PLUGIN-COMPAT block
and is left for that revert.
The Matrix docs described six agent-exposed matrix_* tools and three MATRIX_TOOLS_ALLOW_* env gates that were never implemented — the tool names and gates appear nowhere in code. Docs now describe actual behavior: no Matrix-specific agent tools; reactions/redactions are internal to approval prompts and pickers; MATRIX_ALLOWED_ROOMS scopes responses. Fixes#100535.
Under gateway.multiplex_profiles a served secondary profile's adapter is built and
connected inside _profile_runtime_scope while os.environ still holds the DEFAULT
profile's .env. Credentials and allowlists were already read through the profile
scope (get_scoped_secret / _platform_gate_env), but the non-credential SETTINGS the
adapters read with bare os.getenv were not, so a served profile silently ran with the
default profile's values: webhook listener host/port/URL (SMS, Teams, LINE, Feishu,
BlueBubbles), Signal's connect URL/account gate, mention gating and reactions (Slack,
Matrix, Signal, Feishu, BlueBubbles, Discord), Matrix thread/session/E2EE policy and
message-length limits, Discord backfill/command-sync/attachment caps, Buzz reply mode
and env enablement seed, A2A agent name/port/description/toolsets, and the
api_server model alias.
Every such read now goes through the existing scoped reader (get_scoped_secret, or the
adapter's own scope-aware helper): under a secondary's scope the profile's own .env is
authoritative and a miss yields the default -- never another profile's value; the
default profile and single-profile gateways keep reading os.environ exactly as before.
Buzz and A2A previously short-circuited to "extra only / built-in default" under a
scope, which also dropped the profile's OWN .env; they now read the scope so a served
profile matches its standalone gateway.
The parity harness (temp HERMES_HOME, default + 2 secondaries with distinct values for
every env var each adapter reads, real load_gateway_config + adapter factory in both
topologies) went from 70 raw process-env bypass sites across 14 adapters to only the
HERMES_<PLATFORM>_* perf knobs and the api_server listener vars, which are process-
global by design (agent.secret_scope._GLOBAL_ENV_*).
Under gateway.multiplex_profiles every secondary profile's config loads inside
_profile_runtime_scope, yet every apply_yaml_config_fn hook (feishu, matrix,
whatsapp, slack, dingtalk, discord non-gate keys, telegram non-gate keys) and
gateway/config_loader.py::bridge_core_env_settings still wrote os.environ there.
First-writer-wins: the first secondary with a require_mention / allowlist /
allow_bots / reactions block made that policy the DEFAULT profile's (live:
TELEGRAM_REQUIRE_MENTION written from a secondary load), and secondaries read
the default's env for the same keys.
- gateway/platforms/_shared.py::yaml_env_setter: the one env-write shape for
YAML->env bridges — env wins, skipped under an active secondary scope.
- Every hook now seeds its values into the profile's PlatformConfig.extra and
uses yaml_env_setter; bridge_core_env_settings seeds telegram/signal
require_mention into extra and skips the env write under scope.
- Readers that bypassed extra/scope now consult extra first (matrix flags +
session_scope, slack reactions, telegram reactions/mention_patterns/_extra_bool,
discord reactions/auto_thread/history_backfill/approval_mentions/allow_mentions,
feishu allow_bots, dingtalk mention_patterns, signal require_mention).
Salvages the shape of PR #100604 (whatsapp, earliest report #80099), #100435
(discord) and #100448 (telegram) by @nftpoetrist on top of current main.
Fixes#80099
Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env. Several
adapter-owned authorization gates still read GATEWAY_ALLOW_ALL_USERS, GATEWAY_ALLOWED_USERS
or their platform allowlist/allow-all raw from os.environ, so the default profile opting
into open access opened every secondary email/QQ/WhatsApp/Matrix/Teams/Slack/LINE/DingTalk
bot to any sender (email additionally skipped From: authentication), the default's Matrix
allowlist decided who may approve tool calls on a secondary bot, and a secondary that
opted in only in its own .env was silently deny-all.
Every such read now goes through the adapter's existing module-local scoped reader
(gateway.platforms._shared.get_scoped_secret / matrix _startup_env_secret): profile
scope first, scoped miss = default, never os.environ; the unscoped default-profile and
single-profile paths keep the environ read, where it IS the profile's own value.
Sites: email _allow_all_senders/_allowlist_in_effect; qqbot _open_dm_opted_in;
whatsapp_common _open_dm_opted_in/_live_dm_allow_from; teams _card_action_denied;
matrix _is_authorized_user, MATRIX_ALLOWED_USERS, MATRIX_IGNORE_USER_PATTERNS,
_extra_csv_set (allowed/free-response rooms); slack _slack_allow_bots/_slack_api_human_users;
line _truthy_env/allowlist (allow-all, user/group/room allowlists); dingtalk _extra_get
(allowed_users/chats, free-response chats, require_mention).
Live repro (temp HERMES_HOME, multiplex on, default env GATEWAY_ALLOW_ALL_USERS=true,
secondary scope without opt-in): EmailAdapter._allow_all_senders() True -> False,
QQAdapter._open_dm_opted_in() True -> False, Matrix _is_authorized_user('@stranger')
True -> False, Teams card action allowed -> denied.
Co-authored-by: Drexuxux <drexux0@gmail.com>
Co-authored-by: MoonsvnLyn <FirmamentalSpring@users.noreply.github.com>
Co-authored-by: svector-anu <anuoluwakolapo94@gmail.com>
Co-authored-by: babatorik <durgun.ismail@gmail.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's
values and a secondary profile's .env exists only in its secret scope.
Matrix, Home Assistant, Teams, DingTalk, SMS and Telegram already read
their *secret* scope-aware, but read the paired endpoint/identity raw from
os.environ — so a secondary's token/secret/password was paired with the
default profile's host or app identity:
- Matrix: MATRIX_HOMESERVER / MATRIX_USER_ID / MATRIX_DEVICE_ID in
__init__ and MATRIX_HOMESERVER in _standalone_send -> the secondary's
access token was sent to the default's homeserver (and its E2EE device
id reused).
- Home Assistant: HASS_URL in __init__ and _standalone_send -> the
secondary's HASS_TOKEN was posted to the default's HA instance.
- Teams: TEAMS_CLIENT_ID / TEAMS_TENANT_ID (env-first, even over extra),
TEAMS_SERVICE_URL, TEAMS_HOME_CHANNEL(_NAME) in _credentials,
_env_enablement, _standalone_send and __init__ -> Bot Framework token
requested for the default's app with the secondary's secret; cron
home channel was the default's conversation.
- DingTalk: DINGTALK_CLIENT_ID in _credentials, DINGTALK_WEBHOOK_URL in
_standalone_send -> wrong app id; cron output posted to the default's
robot webhook (URL carries its access_token).
- Telegram: TELEGRAM_WEBHOOK_URL in connect() -> the default's public
webhook URL registered on the secondary bot, mixing update traffic.
Each read now goes through the same scoped reader its sibling secret
already uses (_get_scoped_secret / _startup_env_secret / get_secret with
the UnscopedSecretError fallback), so a scoped miss is empty rather than
the default profile's value; the unscoped default-profile path and
single-profile deployments keep reading os.environ, which there IS the
profile's own value. Teams _credentials now lets extra (per-profile
config.yaml) win over env, matching every other adapter and its own
__init__.
Live repro (/tmp/mux_audit/fix-endpoint-identity/repro.py, secondary
scope "bot2" with the default profile in os.environ): 19 of 24 reads
returned the default's endpoint/identity before; 0 after.
Salvage: SMS from-number fix cherry-picked from #100625 (tests trimmed to
two invariants). Teams identity (#100622) and DingTalk (#100615)
reimplemented on the minimal shape — the Teams PR added a _profile_scoped()
gate to _env_enablement and the DingTalk PR only covered the YAML bridge,
not DINGTALK_CLIENT_ID / DINGTALK_WEBHOOK_URL. Matrix, Home Assistant and
Telegram parts have no prior PR.
Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
_standalone_send builds formatted_body through its own markdown call and
never went through _markdown_to_html, so cron-delivered equations still
arrived as raw dollars. Tokenize/expand around that conversion as well.
Keep the two tests that fail without the fix: inline + display math reach
formatted_body as data-mx-maths markup with the TeX HTML-escaped while the
plain body keeps the raw TeX; unpaired dollars and text colliding with the
sentinel format pass through unchanged (no IndexError). The 17 unit tests
from #106452 are dropped per the salvage bar (≤ 2 invariant tests).
Also drop the dead `_tex_store = []` pre-assignment and the `_fb`/`html`
temporaries in `_markdown_to_html` — pure tidy, no behaviour change.
Element (feature_latex_maths) typesets <div|span data-mx-maths="TEX">
elements at display time, but the outbound HTML sanitizer allowlists tags
and attributes, so data-mx-maths markup sent by the gateway never reaches
Element intact - messages containing $...$ render as raw dollars.
Convert $...$ (inline) and $$...$$ (display) to opaque sentinel tokens
before Markdown conversion and expand them to data-mx-maths markup after
sanitization. Tokens are printable text with no HTML/Markdown meaning, so
neither the converter nor the sanitizer touches the TeX. Unpaired dollars
(prices, literals) are untouched, and adversarial text colliding with the
sentinel format passes through verbatim (index-checked expansion).
(cherry picked from commit eb73aafe0e08502c54e63dbd03e99796c3b7fee3)
#106167 widened the base contract to SendResult but left six native-batch
overrides (Discord, Email, Matrix, Mattermost, Slack, Telegram) returning
None, which (a) meant a media-only reply on those platforms still reported
FAILURE because _record_delivery(None) records nothing, and (b) produced new
`ty` invalid-method-override diagnostics against the widened base (#106192).
Make the contract honest instead of annotating it Optional: each override
now rolls its batches (and any per-image fallback) into one SendResult, so
the turn-outcome accounting works on every platform, not just Signal and
the base loop. The "legacy overrides return None" comment in
_send_image_batch goes away with the legacy.
ty on the 8 touched files: origin/main 241 diagnostics / 20 override,
this branch 241 / 20 — byte-identical diagnostic set; the intermediate
`-> SendResult` head without this commit had 246 / 25.
Refs #106192
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
Review findings on #102117 (independent reviewer + itsflownium):
* hermes_cli.kanban_db.connect / connect_closing pointed at hermes_cli.projects_db (different DB, no
board= parameter). The compat generator ranked candidate homes by path proximity when a name is
defined in several modules. Now it requires shape compatibility with the BASE definition (same
literal for constants, superset of parameter names for defs) and prefers the facade's own
<stem>_* sibling. Same class fixed for tools.tts_tool.DEFAULT_XAI_BASE_URL (-> tts_tool_providers),
and 17 constants/defs that had been pointed at same-named strangers (Matrix MAX_MESSAGE_LENGTH ->
Signal's 8000, tts MAX_TEXT_LENGTH -> BlueBubbles', honcho/retaindb/supermemory *_SCHEMA -> another
plugin's schema, ...) are now restored from BASE verbatim instead.
* send_yuanbao_direct (restored-def): body called adapter._outbound.send_direct, which HEAD moved to
the sender; rewritten to adapter._outbound.sender.send_direct.
* COMPAT_MANIFEST.md states the scope explicitly: public top-level names only; private names and
test monkeypatch seams are not preserved.
* scripts/check_subprocess_stdin.py: _splat_carries_stdin looked 30 lines ahead in the file text
and was satisfied by an unrelated later stdin=; it now finds the splatted name's definition via AST
and requires stdin inside that expression/body.
Tests: tests/test_compat_manifest_targets.py (pointer identity vs the facade's sibling; kanban
connect(board=) opens a Kanban DB, not projects.db; both FAIL on the previous layer),
test_subprocess_stdin_guard gains the false-negative probe, and the MoA -Q quiet-output contract
tests are back (tests/agent/test_moa_quiet_reference_output.py) against build_moa_facade.
The Sep 2026 decomposition (PR #102117) makes internal import paths a non-API: names now live in
the focused modules that define them. This commit is the ONLY thing keeping the old paths alive,
so external plugins have time to update. It is deliberately a single, unsquashed commit:
git revert <this sha>
removes every shim, stub and manifest at once on the announced date. Nothing in-tree may depend on
these pointers: scripts/check_compat_pointers.py (wired into lint.yml) fails CI if it does.
What it adds (see COMPAT_MANIFEST.md, compat_manifest.json):
- 332 facade modules get one delimited `PLUGIN-COMPAT` block appended at the end of the file
- 1,172 moved names resolved lazily via a module `__getattr__` (PEP 562) — never a top-level import,
so no import cycles; facades that already had `__getattr__` get a chained one
- 592 third-party/stdlib names the old modules used to expose, with their original import statements
- 266 public definitions that had been deleted as unused, restored byte-for-byte from the pre-decomposition
tree (+40 private helpers and 16 imports pulled in only because a restored definition needs them)
- 3 deleted modules recreated as re-export stubs (gateway/startup_watchdog, hermes_cli/observability/
relay_runtime, tools/environments/modal_utils)
- private names (`_x`) get no pointer: they were never API (3,792 skipped)
Verified: all 335 touched modules import under a fresh HERMES_HOME and every manifest name resolves;
the lint reports zero in-tree uses; ruff clean; targeted suites unchanged.
dingtalk: drop DINGTALK_TYPE_MAPPING/EXT_MAP re-exports. google_chat: card_spec_to_cards_v2 test -> .cards.
matrix: drop module-level MAX_MESSAGE_LENGTH alias (no importers). teams: drop TeamsSummaryWriter re-export
(teams_pipeline/runtime + tests -> summary_writer). wecom: drop WeComStreamExpiredError/STREAM_EXPIRED_ERRCODE/
MAX_INTERMEDIATE_FRAMES re-exports (tests -> .streaming). parallel: drop _get_parallel_client/_get_async_parallel_client
aliases (tests -> _get_sync_client). email: drop stale 'alias' comment (_esecret_int is the only name).
photon: re-remove credential_summary() (shim-only, cb9b7c36f3); its no-leak test now drives print_credential_summary.
For each issue anchor present in BASE 63279301bc non-test .py and absent on HEAD, the BASE comment/docstring block was re-attached at the HEAD location of the code it explained (matched by the distinctive code line / enclosing def). Sentences already covered by an existing HEAD comment were deduped; the issue number always survives. Insert-only: no code lines changed.