A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.
Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.
Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.
Phase 2 of #88715.
The bridge now gates groups by WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
(#73465), but only the DM policy and DM allowlist were exported from the adapter's
resolved config — a YAML group_allow_from never reached the bridge and, under a
multiplexed gateway, the launch process's WHATSAPP_GROUP_* values did. _bridge_env
exports both like it already did for the DM pair, so unlisted-group traffic is cut
before media download instead of after the POST to Python.
whatsapp_common._select_allowlist generalises the DM precedence reader (config key
presence wins, an explicit empty list stays authoritative, then the first truthy env
carrier); the adapter's group list and the Cloud sibling's group list use it instead
of hand-rolled `or` chains, so `group_allow_from: []` means the same everywhere.
Tests: bridge env carries group policy + group allowlist from the secondary profile's
YAML (scoped precedence file); env fallback test trimmed to the two invariants (red on
the merge base); bare test adapters seed _group_policy/_group_allow_from, which
_bridge_env now reads.
The Node bridge receives WHATSAPP_GROUP_ALLOWED_USERS via
_BRIDGE_PASSTHROUGH_ENV and the YAML bridge maps config.yaml
whatsapp.group_allow_from onto the same env carrier, but
WhatsAppAdapter.__init__ seeded _group_allow_from from config.extra
alone — an install that set the group allowlist only through the
documented env var silently ran group gating with an empty allowlist:
with group_policy=allowlist every group chat was rejected at intake
(the second failure mode triaged on #72529, the half with 'no PR yet').
Mirror the DM path's precedence: config keys win by key presence
(explicit empty list stays authoritative), then the profile-scoped
WHATSAPP_GROUP_ALLOW_FROM / WHATSAPP_GROUP_ALLOWED_USERS env CSV —
multiplex-safe through _wenv, like every other WHATSAPP_* read.
Tests pin the full precedence matrix: env-only seeding, the
allow_from-style alias, config-wins-over-env, the legacy groupAllowFrom
key, empty-by-default, the scoped .env read under multiplex, and the
group-policy gating outcome for an env-seeded allowlist.
Follow-up to the salvaged #92440 commit, shape-gate cleanup only; behaviour is unchanged
(one mention-bearing payload per logical send, captioned media keeps its caption).
- Fold `_send_whatsapp_with_mentions` into `_send_plugin_standalone` (it was a line-for-line
copy of the caption split + `_send_chunks` loop); `mentions` is attached to the first
payload only via a one-shot kwarg dict.
- Drop the `inspect.signature(sender)` probe: the only registered WhatsApp standalone sender
is the in-tree `_standalone_send`, which gains `mentions` in the same change; a foreign
sender already surfaces as a TypeError through `_handle_send`'s error path.
- Collapse `_normalize_outbound_mentions` to dedupe-only; its input is argparse `list[str]`
already validated by the CLI.
- Revert the `\d` -> `[0-9]` edit to `_BARE_PHONE_RE` / `to_whatsapp_jid`: unrelated to the
feature (`normalize_whatsapp_mention_jid` already rejects non-ASCII via `isascii()`) and it
changed output for nine other `to_whatsapp_jid` callers.
- Split the single 137-line test into a `whatsapp_bridge` fixture + two invariants
(rejections never reach the bridge; mentions ride the first payload only, stale bridge
fails closed); drop the fabricated legacy-sender branch that only existed to cover the
deleted probe. Mutation check: removing the first-payload gate turns the new test red.
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
A reply past 4,096 chars goes out as several sendMessage calls. When chunk 2 was
refused by flood control (RetryAfter past the 5s inline cap) send() returned the
bare flood_control result, so _send_with_retry re-sent the WHOLE payload after
the wait: the user saw chunk 1 twice (reporter: 4 messages, 686 duplicated words),
and paths without a ledger row lost the tail outright.
- send() now reports a mid-split refusal through the existing partial_overflow
contract (the key _edit_overflow_split already sets and the stream consumer
reads): delivered_chunks / total_chunks / last_message_id, plus
undelivered_chunks + delivered_message_ids ONLY when non-delivery is certain
(flood cap, Bot API rejection, connect/pool timeout) — an ambiguous TimedOut may
have reached Telegram and is never resumed from.
- BasePlatformAdapter._send_with_retry resumes from the remainder via a new
_resume_partial_send hook (default None = keep the partial failure, never
re-send the head; the plain-text fallback is skipped for partials too). The
Telegram override sends the leftover formatted chunks, continuing the id sequence.
- Per-chat FIFO send gate, reentrant per asyncio task (media paths nest), held
only around the API calls — never across the reconnect wait — on send() and the
media funnel, so concurrent replies to one chat no longer interleave chunks.
- A flood refusal arms a per-chat cooldown (mirrors the sendChatAction cooldown,
capped 300s); sends inside the window fail closed locally with the same
flood_control:<s> result and no API call, so ledger recognition and redelivery
timing are unchanged.
- Edit path: log "refusing (retry_after Ns > cap)" after the cap check instead of
"waiting Ns" followed by no wait.
Live against a local fake Telegram Bot API with a fake token (RetryAfter=7 on
chunk 2 of a 3-chunk reply): before 4 messages / 324 duplicated words; after 3
messages, 853/853 words, 0 duplicated, 0 lost. Two concurrent 3-chunk sends:
before 9 source switches, after 1. Five sends inside a refused window: before 5
API calls, after 0.
Fixes#114396
Co-authored-by: AStrnbrg <45151087+AStrnbrg@users.noreply.github.com>
Co-authored-by: whyyagswhy <166958865+whyyagswhy@users.noreply.github.com>
The Feishu fix on this branch populates MessageEvent.media_text_inlined so
run_inbound's document note stops claiming "Its content has been included
below" when a text attachment was NOT inlined (>100 KB gate or decode
failure); run_inbound treats a missing flag as inlined. Telegram, Discord,
Slack, the WhatsApp bridge adapter and whatsapp_cloud inline "[Content of
…]" the same way but never set the flag, so their notes lied on the
skip path. Mirror the Feishu/buzz per-attachment contract in each: False
for every cached attachment, flipped to True only when the text was
actually injected.
One parametrized (small→True / large→False) test per adapter in the
existing per-platform test files.
The exit-notify wrap reported every receive-loop exception at ERROR with a
traceback, including the ConnectionClosedOK that follows our own CLOSE frame in
disconnect(). Gate on the adapter's _running flag (published through the WS
thread-local next to on_link_up): a live link's death stays ERROR, an
intentional shutdown logs at DEBUG. Live pass side-effect #4 on #113662.
The ``retrying`` publication landed only on the supervisor-rebuild path (WS
thread dies). On the live link ``_auto_reconnect`` is on, so a receive-loop
error runs lark-oapi's ``_reconnect()`` ladder *inside* the receive loop — the
thread never dies and ``gateway_state.json`` kept saying ``connected`` while
the link was down.
Set the SDK's ``Client.on_reconnecting`` observer (lark-oapi 1.6.8
ws/client.py, fired first thing in ``_reconnect()``) from
``_apply_runtime_ws_overrides``; it hops to the adapter loop and publishes
``retrying`` through the same runtime-status call the supervisor uses, gated
on the client still being the live one. ``connected`` is re-stamped by the
existing ``_ws_link_up`` hook once the ladder's ``_connect()`` schedules a new
receive loop, so no ``on_reconnected`` hook is needed.
Docs: the WebSocket-mode section now says a dead link is rebuilt by Hermes'
supervisor and shows ``retrying`` in gateway status until re-established.
Test: test_sdk_reconnect_ladder_publishes_retrying (red before: FeishuAdapter
has no _ws_link_retrying / observer never set; green after).
Follow-up to the salvaged #113668 commits:
- The wrap stopped the worker loop in a ``finally``, i.e. also on a *normal*
return. On the live link ``_auto_reconnect`` is the SDK default (True — the
adapter only flips it off at teardown), and a successful SDK reconnect makes
the old ``_receive_message_loop`` return after ``_connect()`` scheduled a fresh
one. Stopping the loop there would tear down the healthy rebuilt link (no
CLOSE frame, see #10202) on every transient blip. Stop only when the coroutine
exits with an exception: ladder disabled, or the ladder's
``ClientException`` / ``ServerUnreachableException`` re-raise.
- Do not re-raise after logging: the wrap is now the owner of that exception,
so asyncio no longer prints "Task exception was never retrieved" beside it.
- Wrap the real symbol unconditionally: ``lark_oapi.ws.client.Client
._receive_message_loop`` exists on every shipped SDK (1.6.x and 1.7.3, ws
module unchanged between them); a getattr fallback would silently reopen the
gap on a rename.
- ``gateway_state.json`` kept saying ``connected`` while the supervisor
rebuilt: publish ``retrying`` when the WS thread dies (the adapter is still
running, so ``_mark_disconnected`` is the wrong tool) and re-stamp
``connected`` from ``_ws_link_up`` — fired through a thread-local hook when
the SDK schedules a receive loop, the only in-thread proof the handshake
succeeded — hopped onto the adapter loop and gated on the client still being
the live one.
- Tests: the fake ``lark_oapi.ws.client`` modules carry ``Client`` like the real
module; second invariant now covers the normal-return (ladder success) case;
the existing supervisor-restart test asserts the ``retrying`` publication.
Live: real lark-oapi 1.6.8 ``ws.Client`` + the real adapter worker/supervisor
against a local fake Feishu endpoint + websocket server with fake credentials.
The wrapper re-raised the SDK receive-loop exception into the same
unretrieved bare task. Log it with logger.exception first so the root
cause lands next to the supervisor's rebuild line (suggested in review).
lark_oapi parks Client.start() in run_until_complete(_select()) and runs
the receive loop as a bare create_task whose exception nobody retrieves.
With Hermes disabling the SDK's reconnect ladder, a mid-life receive-loop
death leaves a deaf-but-ESTABLISHED socket: start() never returns, the
supervisor's executor future never completes, and gateway_state.json keeps
reporting connected (#113662).
Wrap Client._receive_message_loop in the existing isolation installer so
its exit stops the worker loop; start() then raises, the future completes,
and _supervise_websocket_thread rebuilds the link with its capped backoff.
Deliberate disconnects stay unaffected: they nil _ws_client first and the
supervisor exits without restarting.
Fixes#113662
The bootstrap racer covers sync connects only. The gateway's WebSocket dials
(relay connector, Yuanbao, Buzz) go through ``websockets.connect`` →
``loop.create_connection``, whose ``happy_eyeballs_delay`` defaults to ``None``:
a serial walk that burns the full connect timeout on every blackholed AAAA
record before IPv4 answers — the same stall class #114265 reports, one layer up.
``websockets`` forwards unknown kwargs to ``loop.create_connection``, so each
call site passes ``happy_eyeballs_delay=0.25`` (the RFC 8305 delay anyio and
the sync racer already use). One invariant test per call site captures the
kwargs at a mocked ``websockets.connect``.
`_on_sidecar_message` ran `_normalize_content` (which base64-decodes and writes
inline attachment/voice bytes into the media cache via `_cache_inbound_attachment`)
before the `chat_type == "group" and self.require_mention` check, so an
unmentioned group attachment was persisted and then dropped. Gate on the
user-typed text extracted without touching attachment bytes (`_mention_gate_text`:
text / richlink / group text items), then normalise and strip the wake word only
for messages that pass.
Same class as the Teams fix in this PR (review follow-up). One invariant test:
unmentioned group attachment -> 0 cache writes, mentioned -> 1 and dispatched.
`_handle_media_message` fetched the mxc:// payload (`_download_and_cache_media`)
before `_build_inbound_event` ran `_resolve_message_context`, so an unmentioned
or non-allowlisted room's m.image/m.file was pulled onto the host and then
dropped. Resolve the context first and hand it to `_build_inbound_event` via a
new `ctx` kwarg so the read receipt / thread mark are still applied once.
Same class as the Teams fix in this PR (review follow-up). One invariant test:
unmentioned group media → 0 downloads, mentioned → 1.
Follow-up to the salvaged #113591 gate:
- Teams writes the bot's conversation identity as `28:<app id>` (activity.recipient.id,
mention entities) while App.id is the bare app id. The mention match and the own-message
filter now accept both spellings, so a real tenant mention no longer fails silently.
- A payload that mentions only other people is no longer treated as a bot mention; the
`<at>` text fallback applies only when the activity carries no mention entities at all.
- `self._extra` holds `platforms.teams.extra` on the instance so keys read after
construction (require_mention today) are not dropped with a constructor local (#114366).
- Reply exemption uses one bounded deque instead of deque + mirror set.
- Tests trimmed to two invariants: gate table (drop before attachment download / keep
mention, reply-to-bot, personal) and env-over-YAML precedence. TEAMS_REQUIRE_MENTION
documented in plugin.yaml and the Teams docs page.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: Fernando Muñoz Suazo <fernandrewm@gmail.com>
Three defects let the missed-message backfill re-run the same inbound message on
every reconnect (24 turns in 12 h for one message on the reporter's install):
- Completion was only ever recorded by the send-path ledger writer, keyed on the
Discord reply anchor. With reply_to_mode "off", a streamed final delivered via
edit_message(finalize=True), a fresh-final send or a media-only reply there is no
anchor, so the row stayed processed/replied=0 and _should_backfill_discord_message
re-admitted it forever. The adapter already learns the outcome per inbound id in
_record_discord_processing_complete: on SUCCESS (base confirmed delivery, or the
stream already delivered) it now writes status='responded', replied=1.
- The scan wrote status='discovered' over every candidate at the top of each pass,
erasing the queued/processing claim so the 10-minute active-claim guard could never
fire; two back-to-back scans dispatched the same message twice. The 'discovered'
write no longer overwrites an existing row's status or updated_at.
- A stored per-channel cursor REPLACED the window floor (12-hour-old messages stayed
candidates under window_seconds: 3600) and the parent's cursor object was passed
down to child threads. after = max(cursor, now - window) and threads only see the
window floor plus their own cursor.
- max_dispatches caps one scan, not one row. New missed_message_backfill.max_attempts
(default 3; env DISCORD_MISSED_MESSAGE_BACKFILL_MAX_ATTEMPTS) is a lifetime ceiling
on re-dispatch of a single message; `attempts` now counts dispatches only.
Live repro (real DiscordAdapter, temp HERMES_HOME ledger, fake client yielding the
same message per scan, fresh adapter per scan = reconnect): streamed final with
reply_to_mode=off 3 scans -> before 3 dispatches, after 1; turn that never completes
3 scans -> 3/1; turn that always fails 6 scans -> 6/3; 12h-old cursor with a 1h window
-> before 2 out-of-window dispatches (parent + thread), after 0; control: an
unanswered message still dispatches exactly once, and a cursor newer than the window
still narrows the scan.
Part of #113631 (with the salvaged anchor fallback from #113633 by @KoNit-K).
rich_blocks turns markdown bullets into rich_text_list. rich_text does
not interpret mrkdwn, so Slack <url|text> was emitted as literal text
while [text](url) became a real link. Section/mrkdwn paragraphs were
unaffected.
Parse Slack autolinks in _inline_elements (lists, quotes, table cells).
Mentions (<@U>, <#C>, <!here>) have no scheme: and stay as text.
send_slash_confirm cut the RAW message to 3800 chars and only then ran format_message,
whose MarkdownV2 escaping expands text (3800 dots -> 7600 chars), so the button card could
still exceed the 4096 cap and fail exactly like the approval card this branch fixed. It now
fits the message with the shared _ea_fit bisection measured after format_message (the
suffix's own rendering is reserved from the budget).
_ea_fit takes an optional `escape` renderer (default _ea_escape) and measures in the
adapter's message_len_fn units, and the Telegram command budget sums the framing with
utf16_len, so astral chars (emoji) are budgeted the way the adapter's 4096 chunker counts
them (len() undercounted them by half).
Tests: slash-confirm preview for a 3800-char message and the approval card for an
emoji-dense command both stay <= 4096 UTF-16 units (red before, green after).
Telegram truncated the command to 3800 RAW chars and only then html.escape()d it, so a
command heavy in `&`/`<`/`>` rendered far past 4096 chars; the framing (header, <pre>,
reason label, deadline, smart-deny line) sat outside that budget and the reason itself
was unbounded. Telegram answered "Message is too long" and the gateway fell back to the
text /approve prompt.
- `BasePlatformAdapter._ea_fit`: the shared exec-approval core now measures both the
command preview and the reason after `_ea_escape` (bisection on the raw prefix; the
escaped length is monotonic) — identity for non-escaping platforms.
- Telegram computes the command budget from `MAX_MESSAGE_LENGTH` minus the rendered
framing and escaped reason (the same idiom Discord and Slack already use) and bounds
the reason at 500 escaped chars.
Live against a local fake Telegram Bot API (fake token) rejecting text > 4096:
before, 3700 x "&" sent 18670 chars → 400 → /approve fallback; after, 4093 chars,
inline keyboard delivered.
Co-authored-by: kelvmg <16878572+kelvmg@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The exec approval prompt already keeps the command/reason in plain `content`
and sends a header-only embed, so the body renders exactly once (#114693).
`send_slash_confirm` and `send_clarify` still restated the message/question
(and the clarify reply hint) in `embed.description`/fields while also sending
the self-contained `content`, so embed-rendering clients saw the body twice —
the same defect class the PR closes. Apply the same rule to both siblings,
drop the now-unused `_embed_body` helper, and pin each with one invariant test.
#114699 removed `_EA_REASON_BUDGET = 300` together with the embed field that used it,
but the shared `_format_exec_approval` core also applies that budget to the plain
content. Unbounded, a long reason starves the command preview to zero and pushes the
content past Discord's 2000-char cap (probe: 5216 chars for a 5000-char reason), which
the API rejects. Restore the budget with its WHY; pin the cap with one invariant test
and condense the embed assertions.
Same class as the Bot Chat drain wedge already on this branch: every JSON-file
scan guarded "did it parse?" and then assumed the value was a dict. A file
holding `42`, `"oops"` or `[1,2,3]` (corruption, truncated write, foreign tool)
passed the guard and raised AttributeError/TypeError at the first `.get()`,
usually before a single healthy sibling was processed. Each site now treats a
non-object payload like a corrupt file under that subsystem's existing policy:
- tools/bot_relay.py::_expire_if_stale / claim_pending_envelopes — the
envelope is skipped by the sweep and not claimed (same as unparseable).
- tools/browser_lightpanda.py::reap_orphaned_lightpanda — record unlinked,
scan continues.
- tools/write_approval.py::list_pending / get_pending — record skipped with
the existing "unreadable pending record" warning / None.
- tui_gateway/methods_session.py::_legacy_spawn_tree_entry / spawn_tree.load —
scalar snapshot reads as empty / returns the existing 5000 error instead of
violating the SpawnTreeLoadResult contract.
- hermes_cli/local_runtime/binaries.py::manifest_verified — False.
- plugins/platforms/a2a/protocol.py::load_conversation — non-dict lines are
dropped, keeping the declared list[dict] return.
- batch_runner.py::_load_dataset / _scan_completed_prompts_by_content /
_combine_batch_files — line skipped and counted as filtered.
- trajectory_compressor.py::process_entry_async — scalar entry passed through
unchanged.
Ported from the source hunks of PR #114241; its gateway/shutdown_flush.py
drain_transcript_spool hunk is left to open PR #84785, and its
recover_pending_to_db / cron / bot_live_delivery / bot_mode_dm hunks are
already on this branch or on main.
(cherry picked from commit d4b54568887e69b3ee3d363ebe4dcd657ccf64f9)
Slack seals a native stream server-side after a few minutes (live-observed
at ~5m20s on three independent long turns, 2026-09-15/16; the lifetime is
not documented). The next chat.appendStream on the card fails with
message_not_in_streaming_state. The adapter returned a bare failure, the
TurnRunner latched native_failed, and the rest of the turn rendered as an
edited text bullet list. Long autonomous turns lost the card UX exactly
when it mattered.
On message_not_in_streaming_state from appendStream, drop the dead
stream_ts and chat.startStream a fresh plan-mode card in the same thread,
then append the current frame there. Every frame already carries the full
visible task projection, so no task state is lost. One reopen per update;
a second rejection surfaces as a real failure. The sealed card is a plain
message now, so no stopStream is sent to it; the turn-final stop targets
the reopened card. Error matching reads SlackApiError.response["error"],
never the message text.
Tests assert the wire sequence (start, append, rejected append, start,
append on the new ts), the reopened frame's task states, the cache pointing
at the new card, and the stop targeting it; plus the one-reopen bound.
Mutation: forcing the expiry branch off turns both tests red.
Four places changed behaviour for users who never touched the setting:
- `_interim_send` was stamped on every `warn` status and media-failure notice, and the
Slack/relay egress doors learned to skip stream sealing for it. Main's status sends carry
no interim mark at all, so the gap is class-wide (every status kind), and fixing it for
warnings alone is an undeclared streaming-contract change. Reverted here; the whole-class
fix belongs in its own PR against gateway/AGENTS.md rule 3.
- The entire post-handler delivery (unwrap, TTS, final text, attachments, delivery-ledger
writes) ran inside `_media_delivery_scope`. Under multiplex that binds the routed home, so
delivery obligations landed in the routed profile's state.db while boot-time
`_claim_pending_obligations` still reads the launch home. Only the policy reads
(`diagnostic_wake_muted`, `warning_text`) bind the routed scope now; delivery stays where
main ran it.
- The turn-crash notice is rebuilt the same way: scope around the policy read, send outside.
- The "delivery failed after multiple attempts" notice is unconditional again: the requested
result itself was lost and this line is its only signal, so it is not a diagnostic.
Tests that asserted the reverted behaviours are removed; the reviewer-round test file is
renamed for what it covers.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
Discord paces a bot's sends at roughly one per second, so chunk 3 of a long
handoff lands after the 2s window anchored on the tag and was still dropped —
the symptom the continuation window exists to fix. Each admitted continuation
now re-arms the window; the gateway bot loop guard bounds a bot that never
stops. The flush-delay override moves onto the BasePlatformAdapter seam
(_text_batch_delay_for) that main relocated the batcher to.
Tests trimmed to the invariants: reply-ping-only bot message rejected by
default (and admitted with the explicit opt-out), and a 3-chunk tagged
handoff paced past the original window arrives as one batched event; the
defensive getattr fallbacks that only served test construction are gone.
_schedule_polling_recovery promised the gateway 'stays alive and will retry' for every error, but a _PollingStallError goes straight to _go_fatal_network (supervisor rebuild). Branch the wording on the error type, drop the watchdog's own pre-log so a stall yields exactly one error-level line (from _go_fatal_network), and carry stalled_for/generation in the stall error text instead. Test module docstring and test name updated to match the hand-off semantics.
Hoist the `_PollingStallError` check in `_handle_polling_network_error` to
right after the teardown/fatal guard, before the retry counter increment,
the exponential sleep, `_stop_updater_or_go_fatal` and both connection
drains. Use a plain `isinstance` (both raise sites construct the error
directly; nothing wraps it). The stall test now also asserts no sleep, no
retry-counter bump, no in-place `updater.stop()` and no drain.
WHY: the check sat after `await asyncio.sleep(delay)` (5-60 s), the
counter bump and a bounded `updater.stop()` (up to 15 s), so a gateway
already confirmed deaf stayed deaf 5-75 s longer and consumed a retry
slot for something that is not a retry.
The in-place `updater.stop()` before the handoff is dropped deliberately:
`_go_fatal_network` -> `_handoff_polling_fatal_error` -> supervisor
rebuild runs `disconnect()`, which performs the same bounded
`updater.stop()` (`_UPDATER_STOP_TIMEOUT`, falls through on timeout) plus
`app.stop()/shutdown()`, so stopping here only duplicated that work on the
slow path. Docstrings for `_handle_polling_network_error` and
`_check_polling_stall` no longer describe the stall as a reconnect-ladder
escalation.
`_check_polling_stall` hand-rolled the body of `_schedule_polling_recovery`
(set `_send_path_degraded`, `_mark_degraded()` when running, spawn
`_handle_polling_network_error`). Call the helper instead, matching the
post-reconnect verifier stall site.
WHY: one recovery entry point keeps degraded-state marking and the
"polling degraded (reason)" log line consistent across every scheduler.
The helper's early-return guards (`_teardown_started or has_fatal_error`,
`_recovery_in_flight()`) are already asserted at the top of
`_check_polling_stall` with no await in between, so they are no-ops here
and behaviour is unchanged.
Replace `_looks_like_polling_stall` (a three-substring classifier over
`str(error)`) with `_PollingStallError(RuntimeError)`. The watchdog and the
post-reconnect verifier raise it at their two stall sites; the reconnect
ladder checks `isinstance` through `_iter_exception_graph` before handing
the adapter to the supervisor.
WHY: a substring match couples recovery routing to log wording (rename a
message, silently lose the fatal handoff) and can misfire on unrelated
errors that quote the same words. A typed exception keeps #113657's
behaviour with no behavioural change beyond the classifier's source.
A Matrix `attach_to_session` cron delivery and a CLI→Matrix `/handoff` now land in their own
thread AND a human reply in that thread continues the seeded session (#112918).
The salvaged `MatrixAdapter.create_handoff_thread` (ea-s21, #108367) gives both seeders a
thread root. The second half of the bug is a key-shape mismatch: the handoff watcher and
`_seed_cron_thread_session` seeded `matrix🧵<room>:<root>` while the adapter keys every
in-thread reply on the ROOM's type (`matrix:group|dm:<room>:<root>`), so seed and reply never
met. Following the direction landed for Slack in 1f3f45e87b (#111896), the seeders now mirror
the adapter instead of the adapter moving onto a `thread` slot: rekeying inbound Matrix threads
would orphan every existing Matrix thread session and drop `is_group` in authz for in-thread
messages (`_GROUP_CHAT_TYPES` has no `thread`).
- gateway/run_startup.py: Matrix destinations key `dm`/`group` from the adapter's
`get_chat_info` (`_handoff_home_is_dm`).
- cron/scheduler_delivery.py: `_THREAD_REPLY_CHAT_TYPE` — Slack and Matrix non-DM thread seeds
use `group`. Slack channel cron threads had the same mismatch (adapter `build_source` keys
`group`; the seed said `thread`); the two cron tests that pinned `thread` for Slack channels
asserted the wrong shape and now assert the adapter's.
- Dropped from #108367: the `chat_type="thread"` inbound rekey (see above) and its tests;
contributor tests trimmed to two invariants.
- Docs: Matrix listed among thread-capable handoff/cron platforms.
Live probe (in-process, real MatrixAdapter + real handoff destination + real cron seeder):
before create_handoff_thread -> None; handoff key matrix🧵… ≠ inbound matrix:group:…;
cron seed matrix🧵… ≠ inbound
after create_handoff_thread -> '$seed'; handoff == inbound == cron seed (room and DM cases)
Co-authored-by: ea-s21 <190767603+ea-s21@users.noreply.github.com>
The Matrix adapter inherited the base create_handoff_thread (returns None), so
cron attach_to_session and the gateway handoff watcher silently no-op on Matrix
while Telegram/Discord/Slack support them: the alert/handoff is delivered flat
instead of as a replyable thread, and the scheduler logs thread_id=None ->
"Mirror: no session found". A human reply then can't resume the session with
its seeded context.
Implement it Slack-style. Matrix has no channel-level create-thread API -- a
thread is just events whose m.relates_to / rel_type: m.thread reference a root
event's event_id -- so seed a root message and return its event_id as the thread
handle. The implementation reuses the adapter's own send() (chunking + E2EE
retry + formatting) and registers the root with self._threads.mark(), mirroring
inbound thread handling; _apply_relation_metadata already threads later sends off
a supplied thread_id, so the returned id is immediately usable. Returns None on
missing client / failed send (callers fall back to parent_chat_id).
Adds tests/gateway/test_matrix_handoff_thread.py (seed id returned; default seed
on blank name; None on no client, failed send, and missing event id).
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.
The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.
Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
A relay that accepted, authenticated and subscribed and then closed cleanly
made the read loop return without raising, so _websocket_loop reconnected
immediately with no backoff and never flipped health to "retrying". The read
loop now raises ConnectionError on StopAsyncIteration so the clean close takes
the same backoff + degraded path as an idle or send-side disconnect.