Same class as the Bot Chat drain wedge already on this branch: every JSON-file
scan guarded "did it parse?" and then assumed the value was a dict. A file
holding `42`, `"oops"` or `[1,2,3]` (corruption, truncated write, foreign tool)
passed the guard and raised AttributeError/TypeError at the first `.get()`,
usually before a single healthy sibling was processed. Each site now treats a
non-object payload like a corrupt file under that subsystem's existing policy:
- tools/bot_relay.py::_expire_if_stale / claim_pending_envelopes — the
envelope is skipped by the sweep and not claimed (same as unparseable).
- tools/browser_lightpanda.py::reap_orphaned_lightpanda — record unlinked,
scan continues.
- tools/write_approval.py::list_pending / get_pending — record skipped with
the existing "unreadable pending record" warning / None.
- tui_gateway/methods_session.py::_legacy_spawn_tree_entry / spawn_tree.load —
scalar snapshot reads as empty / returns the existing 5000 error instead of
violating the SpawnTreeLoadResult contract.
- hermes_cli/local_runtime/binaries.py::manifest_verified — False.
- plugins/platforms/a2a/protocol.py::load_conversation — non-dict lines are
dropped, keeping the declared list[dict] return.
- batch_runner.py::_load_dataset / _scan_completed_prompts_by_content /
_combine_batch_files — line skipped and counted as filtered.
- trajectory_compressor.py::process_entry_async — scalar entry passed through
unchanged.
Ported from the source hunks of PR #114241; its gateway/shutdown_flush.py
drain_transcript_spool hunk is left to open PR #84785, and its
recover_pending_to_db / cron / bot_live_delivery / bot_mode_dm hunks are
already on this branch or on main.
(cherry picked from commit d4b54568887e69b3ee3d363ebe4dcd657ccf64f9)
Slack seals a native stream server-side after a few minutes (live-observed
at ~5m20s on three independent long turns, 2026-09-15/16; the lifetime is
not documented). The next chat.appendStream on the card fails with
message_not_in_streaming_state. The adapter returned a bare failure, the
TurnRunner latched native_failed, and the rest of the turn rendered as an
edited text bullet list. Long autonomous turns lost the card UX exactly
when it mattered.
On message_not_in_streaming_state from appendStream, drop the dead
stream_ts and chat.startStream a fresh plan-mode card in the same thread,
then append the current frame there. Every frame already carries the full
visible task projection, so no task state is lost. One reopen per update;
a second rejection surfaces as a real failure. The sealed card is a plain
message now, so no stopStream is sent to it; the turn-final stop targets
the reopened card. Error matching reads SlackApiError.response["error"],
never the message text.
Tests assert the wire sequence (start, append, rejected append, start,
append on the new ts), the reopened frame's task states, the cache pointing
at the new card, and the stop targeting it; plus the one-reopen bound.
Mutation: forcing the expiry branch off turns both tests red.
Four places changed behaviour for users who never touched the setting:
- `_interim_send` was stamped on every `warn` status and media-failure notice, and the
Slack/relay egress doors learned to skip stream sealing for it. Main's status sends carry
no interim mark at all, so the gap is class-wide (every status kind), and fixing it for
warnings alone is an undeclared streaming-contract change. Reverted here; the whole-class
fix belongs in its own PR against gateway/AGENTS.md rule 3.
- The entire post-handler delivery (unwrap, TTS, final text, attachments, delivery-ledger
writes) ran inside `_media_delivery_scope`. Under multiplex that binds the routed home, so
delivery obligations landed in the routed profile's state.db while boot-time
`_claim_pending_obligations` still reads the launch home. Only the policy reads
(`diagnostic_wake_muted`, `warning_text`) bind the routed scope now; delivery stays where
main ran it.
- The turn-crash notice is rebuilt the same way: scope around the policy read, send outside.
- The "delivery failed after multiple attempts" notice is unconditional again: the requested
result itself was lost and this line is its only signal, so it is not a diagnostic.
Tests that asserted the reverted behaviours are removed; the reviewer-round test file is
renamed for what it covers.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
Discord paces a bot's sends at roughly one per second, so chunk 3 of a long
handoff lands after the 2s window anchored on the tag and was still dropped —
the symptom the continuation window exists to fix. Each admitted continuation
now re-arms the window; the gateway bot loop guard bounds a bot that never
stops. The flush-delay override moves onto the BasePlatformAdapter seam
(_text_batch_delay_for) that main relocated the batcher to.
Tests trimmed to the invariants: reply-ping-only bot message rejected by
default (and admitted with the explicit opt-out), and a 3-chunk tagged
handoff paced past the original window arrives as one batched event; the
defensive getattr fallbacks that only served test construction are gone.
_schedule_polling_recovery promised the gateway 'stays alive and will retry' for every error, but a _PollingStallError goes straight to _go_fatal_network (supervisor rebuild). Branch the wording on the error type, drop the watchdog's own pre-log so a stall yields exactly one error-level line (from _go_fatal_network), and carry stalled_for/generation in the stall error text instead. Test module docstring and test name updated to match the hand-off semantics.
Hoist the `_PollingStallError` check in `_handle_polling_network_error` to
right after the teardown/fatal guard, before the retry counter increment,
the exponential sleep, `_stop_updater_or_go_fatal` and both connection
drains. Use a plain `isinstance` (both raise sites construct the error
directly; nothing wraps it). The stall test now also asserts no sleep, no
retry-counter bump, no in-place `updater.stop()` and no drain.
WHY: the check sat after `await asyncio.sleep(delay)` (5-60 s), the
counter bump and a bounded `updater.stop()` (up to 15 s), so a gateway
already confirmed deaf stayed deaf 5-75 s longer and consumed a retry
slot for something that is not a retry.
The in-place `updater.stop()` before the handoff is dropped deliberately:
`_go_fatal_network` -> `_handoff_polling_fatal_error` -> supervisor
rebuild runs `disconnect()`, which performs the same bounded
`updater.stop()` (`_UPDATER_STOP_TIMEOUT`, falls through on timeout) plus
`app.stop()/shutdown()`, so stopping here only duplicated that work on the
slow path. Docstrings for `_handle_polling_network_error` and
`_check_polling_stall` no longer describe the stall as a reconnect-ladder
escalation.
`_check_polling_stall` hand-rolled the body of `_schedule_polling_recovery`
(set `_send_path_degraded`, `_mark_degraded()` when running, spawn
`_handle_polling_network_error`). Call the helper instead, matching the
post-reconnect verifier stall site.
WHY: one recovery entry point keeps degraded-state marking and the
"polling degraded (reason)" log line consistent across every scheduler.
The helper's early-return guards (`_teardown_started or has_fatal_error`,
`_recovery_in_flight()`) are already asserted at the top of
`_check_polling_stall` with no await in between, so they are no-ops here
and behaviour is unchanged.
Replace `_looks_like_polling_stall` (a three-substring classifier over
`str(error)`) with `_PollingStallError(RuntimeError)`. The watchdog and the
post-reconnect verifier raise it at their two stall sites; the reconnect
ladder checks `isinstance` through `_iter_exception_graph` before handing
the adapter to the supervisor.
WHY: a substring match couples recovery routing to log wording (rename a
message, silently lose the fatal handoff) and can misfire on unrelated
errors that quote the same words. A typed exception keeps #113657's
behaviour with no behavioural change beyond the classifier's source.
A Matrix `attach_to_session` cron delivery and a CLI→Matrix `/handoff` now land in their own
thread AND a human reply in that thread continues the seeded session (#112918).
The salvaged `MatrixAdapter.create_handoff_thread` (ea-s21, #108367) gives both seeders a
thread root. The second half of the bug is a key-shape mismatch: the handoff watcher and
`_seed_cron_thread_session` seeded `matrix🧵<room>:<root>` while the adapter keys every
in-thread reply on the ROOM's type (`matrix:group|dm:<room>:<root>`), so seed and reply never
met. Following the direction landed for Slack in 1f3f45e87b (#111896), the seeders now mirror
the adapter instead of the adapter moving onto a `thread` slot: rekeying inbound Matrix threads
would orphan every existing Matrix thread session and drop `is_group` in authz for in-thread
messages (`_GROUP_CHAT_TYPES` has no `thread`).
- gateway/run_startup.py: Matrix destinations key `dm`/`group` from the adapter's
`get_chat_info` (`_handoff_home_is_dm`).
- cron/scheduler_delivery.py: `_THREAD_REPLY_CHAT_TYPE` — Slack and Matrix non-DM thread seeds
use `group`. Slack channel cron threads had the same mismatch (adapter `build_source` keys
`group`; the seed said `thread`); the two cron tests that pinned `thread` for Slack channels
asserted the wrong shape and now assert the adapter's.
- Dropped from #108367: the `chat_type="thread"` inbound rekey (see above) and its tests;
contributor tests trimmed to two invariants.
- Docs: Matrix listed among thread-capable handoff/cron platforms.
Live probe (in-process, real MatrixAdapter + real handoff destination + real cron seeder):
before create_handoff_thread -> None; handoff key matrix🧵… ≠ inbound matrix:group:…;
cron seed matrix🧵… ≠ inbound
after create_handoff_thread -> '$seed'; handoff == inbound == cron seed (room and DM cases)
Co-authored-by: ea-s21 <190767603+ea-s21@users.noreply.github.com>
The Matrix adapter inherited the base create_handoff_thread (returns None), so
cron attach_to_session and the gateway handoff watcher silently no-op on Matrix
while Telegram/Discord/Slack support them: the alert/handoff is delivered flat
instead of as a replyable thread, and the scheduler logs thread_id=None ->
"Mirror: no session found". A human reply then can't resume the session with
its seeded context.
Implement it Slack-style. Matrix has no channel-level create-thread API -- a
thread is just events whose m.relates_to / rel_type: m.thread reference a root
event's event_id -- so seed a root message and return its event_id as the thread
handle. The implementation reuses the adapter's own send() (chunking + E2EE
retry + formatting) and registers the root with self._threads.mark(), mirroring
inbound thread handling; _apply_relation_metadata already threads later sends off
a supplied thread_id, so the returned id is immediately usable. Returns None on
missing client / failed send (callers fall back to parent_chat_id).
Adds tests/gateway/test_matrix_handoff_thread.py (seed id returned; default seed
on blank name; None on no client, failed send, and missing event id).
A long-lived serve process keeps a deleted profile as the context home of threads
that outlive the delete. A bare `mkdir(parents=True)` right before an atomic write
brings `profiles/<name>/` back after `hermes profile delete` has written the
tombstone and removed the tree.
The writers in `utils` and the seven callers named in #112592 are guarded by the
preceding commits; this one applies the same `mkdir_under_hermes_home` idiom to the
other pre-write directory creations found by the same mechanical rule (auth,
personality, plugin catalog, skills sync, tool discovery cache, platform adapters,
memory plugins, local runtime supervisor, process identity, breadcrumbs). The two
sites that pass `mode=` keep their mkdir behind `assert_named_profile_home_live`.
The guard is a no-op unless the target has a provable `profiles/<name>` ancestor.
Salvaged from #112596 (30-file sweep) on top of #112594 / #112601; the overlapping
files were resolved to the already-landed versions.
A relay that accepted, authenticated and subscribed and then closed cleanly
made the read loop return without raising, so _websocket_loop reconnected
immediately with no backoff and never flipped health to "retrying". The read
loop now raises ConnectionError on StopAsyncIteration so the clean close takes
the same backoff + degraded path as an idle or send-side disconnect.
Follow-up to the cherry-picked watchdog from #112052 (@KoNit-K), finishing the
class the reporter of #112049 laid out:
- `_websocket_loop` runs the read loop and the discovery sweep as sibling
tasks and ends the connection when EITHER finishes. The discovery sweep
re-raises `ConnectionClosed` instead of logging it and retrying next tick:
a send that sees the socket closed is proof the read the loop is parked on
will never return. That is exactly the traceback the reporter watched for
22-86 h while inbound stayed silent.
- Health is invalidated while reconnecting: the first disconnect publishes
`retrying` (`_mark_degraded`) and a successful re-subscribe publishes
`connected` again. Before, `connect()` wrote "connected" once and nothing
ever changed it, so `/health/detailed` claimed delivery during the silence.
- The teardown awaits both tasks with `gather(return_exceptions=True)` instead
of a bare `except (CancelledError, Exception): pass`, which could swallow a
`disconnect()` cancellation landing mid-teardown.
- Slims the salvaged read-loop hunk: the extra "receive task remained parked"
warning and the in-loop `_mark_degraded()` are dropped; the reconnect log
line and the loop-level health flip cover both.
Docs: the Buzz page still described inbound as poll-only and the WebSocket
transport as a future optimization; it now describes the watchdog and the
`retrying` health state.
Under multiplex a secondary profile reads FEISHU_GROUP_POLICY from its own
secret scope only (deliberate 0.21.3 isolation), so a profile whose .env
carries no FEISHU_* policy keys falls back to `allowlist` with an empty
FEISHU_ALLOWED_USERS and every human group message is rejected while DMs
keep working. That deny was logged only at DEBUG, making it look like the
events never arrived (#111420).
Keep the scoped read as is — no environ fallthrough. Instead, the first
group drop caused by the untouched allowlist default logs once at WARNING
naming the chat and the keys to set (FEISHU_GROUP_POLICY /
FEISHU_ALLOWED_USERS in the profile's own .env, or group_rules in its
config.yaml). Operator-configured denies (populated allowlist, per-chat
rule, non-allowlist policy) and later drops stay at DEBUG. The predicate
lives in the topical sibling feishu_admission_diagnostics.py.
Docs: the Group Message Policy section now states the per-profile read and
where to put the keys under a multiplexed gateway.
Co-authored-by: bear0328 <bear0328@users.noreply.github.com>
Co-authored-by: NanPan <111261006+poijygfdyy@users.noreply.github.com>
scope_id_for_chat only consulted the channel→team map, which is empty right after boot (and
after a reconnect) until an inbound event from that channel arrives. A /handoff into a Slack home
without a stored scope_id (SLACK_HOME_CHANNEL env homes, or config homes never re-set via
/sethome) therefore built a key without the team while every thread reply carries it — the
handed-off thread was still orphaned across a restart (#111896).
When the map has no entry and the channel is not known to be shared across workspaces, fall back
to the single authenticated workspace (filled by auth.test at connect); multi-workspace installs
keep returning None.
Follow-up to the salvaged #111938 commit: `_slack_response_payload` already normalizes a
SlackResponse/dict body, so the new `_slack_api_error_code` helper and the two-branch
logger.error were redundant. One log line now always carries `api_error=<code|none>` so an
HTTP 200 + ok=false failure (e.g. message_not_found) is readable without exc_info.
Tests trimmed to one invariant per fix (session key on the response-ready line; API error
code on the edit failure); the `session=unknown` fallback test was a change-detector.
The installer drops uv in $HERMES_HOME/bin without exporting it, so the
bare `uv pip install --python ...` tip failed with `uv: command not found`
for installer-only users. The four copies of the tip (QQ Bot, Feishu, WeCom,
managed Telegram bot) now render through one helper, managed_uv.pip_install_hint,
which names the managed binary when present and falls back to `uv` otherwise.
The standard Hermes install is a `uv venv`, which ships no `pip` module:
`<venv>/bin/python -m pip install qrcode` fails with "No module named pip"
(the exact console output in #111695). Switch all four QR-fallback tips
(Feishu, WeCom, QQ onboarding, Telegram managed bot) to
`uv pip install --python <sys.executable> qrcode`, the form the in-tree
plugin install hints already use (hindsight, mem0), so the printed command
works as-is and still targets the active profile's interpreter.
Adds the Feishu-surface invariant test from #111696 and tightens the
Telegram test to the working command form.
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The Feishu, WeCom, QQ onboarding and Telegram managed-bot flows printed a
hard-coded 'pip install qrcode' tip when the qrcode package was missing. In
Hermes' isolated venv the bare pip either doesn't exist or targets an
unrelated system Python. Print '{sys.executable} -m pip install qrcode'
instead, matching the existing codebase convention for install hints.
Fixes#111695
Fold the three re-arm helpers (_typing_retrigger_state, _clear_typing_retrigger,
_send_typing_quietly) into _retrigger_typing itself; the semantic change from #111886
is unchanged: schedule sendChatAction as a tracked background task instead of awaiting
it on the send path, one in-flight re-arm per chat, at most one per
typing_retrigger_min_interval_seconds (2s default, matching _keep_typing), and honour
typing_indicator: false, which previously only gated the refresh loop.
Tests: keep the two invariants that are red on origin/main — an intermediate send()
returns while sendChatAction is stalled, and a 20-chunk stream costs one sendChatAction
per chat — and drop the eight change-detector variants.
Co-authored-by: aurel282 <aurelien.gek@gmail.com>
`_retrigger_typing` awaited `sendChatAction` inline on the send path, and
streaming re-arms after *every* intermediate send. `sendChatAction` is a
fire-and-forget UI hint whose result nobody reads, but awaiting it ran its
TLS round-trip on the same event loop as the `getUpdates` long-polls.
With several agents streaming concurrently the loop stayed pinned, the
long-polls were never serviced, and they decayed into CLOSE-WAIT while the
adapter still reported `connected` — a gateway that is deaf but healthy, which
`Restart=always` cannot recover because the process never exits.
py-spy put 20 of 20 MainThread samples in `send_typing` -> `send_chat_action`
-> `start_tls`. The handshakes are what cost: with `max_keepalive_connections=4`,
a re-arm per chunk churns the pool so most calls pay a fresh TLS handshake on
the loop thread.
Three changes, all in the re-arm path:
- Schedule the re-arm as a tracked task rather than awaiting it, so a
round-trip never delays a send or a poll. It joins `_background_tasks`, so
shutdown cancels it and it cannot outlive the adapter.
- One in-flight re-arm per chat, and at most one per
`typing_retrigger_min_interval_seconds` (default 2s, `extra` knob; 0 restores
a call per send). Telegram's bubble lasts ~5s and `_keep_typing` already
refreshes every 2s, so the re-arm only has to cover the gap left by a landed
message.
- Honour `typing_indicator: false`. Only `_keep_typing` consulted it, so the
documented workaround still paid for a `sendChatAction` on every
intermediate send.
Simulating 200 streamed chunks with a 10ms loop-blocking handshake:
200 `sendChatAction` calls and 2037ms of send-path time before, 1 call and
13ms after.
Fixes#111727
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
connect() called _register_slash_commands inline on every (re)connect, and
/reload-skills called refresh_skill_group inline; both run
discord_skill_commands_by_category, the same per-skill path-resolution walk
the Telegram menu paid in #110707. On a 1.5k-skill install that holds the
loop past the liveness watchdog. Registration now hops through
asyncio.to_thread from connect(); refresh_skill_group is a coroutine that
hops the rescan the same way (the reload handler already awaits an awaitable
result). Contextvar-scoped profile overrides propagate through to_thread.
One invariant test: the loop keeps ticking while the scan blocks, from both
sites. Red on the PR head, green here.
Review finding: Discord _register_slash_commands/refresh_skill_group ran the skill catalog disk scan synchronously on the event loop.
The gateway's idle-command path resolved skill slash commands inline on the
loop: a cold skill scan, skill file loads and the unavailable-skill rglob over
every skills dir. On a 1.5k-skill install that held the loop ~2 minutes, the
loop-liveness watchdog fired and the gateway exited mid-session (#111091).
_hm_skill_slash_rewrite now runs through _run_in_executor_with_context so the
profile contextvars the scan is scoped to survive the hop. Known commands still
short-circuit before any I/O (previous commit).
Same class in the Telegram inline picker: build_inline_results rebuilds the
command/skill catalog per keystroke via _collect_gateway_skill_entries, the
same path-resolution pass #110707 traced in the command menu.
One invariant test: a known command triggers no scan; an unknown command's scan
runs while the loop keeps ticking. Red on origin/main and on the reorder-only
tree, green here.
Follow-up to the salvaged #110716 commit. Drops the module-level
_build_telegram_command_menu wrapper (telegram_menu_max_commands is a config
read and stays on the loop; asyncio.to_thread takes kwargs directly), and
applies the same hop to _ensure_forum_commands, which rebuilds the menu on the
inbound-message path for every forum chat until registration succeeds — the
second live fire site named in #110707.
Replaces the contributor's test with one that proves the invariant for both
sites: the loop keeps ticking while telegram_menu_commands blocks. Red on
origin/main, green here.
The reply-pill fix split the body on a leading "> " so the strip would not
rewrite the "> <@bot:srv>" pill. That keyed the exemption on the body shape,
not on the relation: a hand-typed blockquote in a plain (non-reply) message
that mentions the bot inside the quote reached the agent with the raw
"@hermes:example.org" text, which main used to strip. Split around the quote
only when m.relates_to carries m.in_reply_to; otherwise strip the whole body.
Review finding: quote-block exemption keyed on body.startswith('> ') instead of m.in_reply_to.
The Matrix adapter strips the bot's mention from the inbound body before
_extract_reply_context parses the inline reply fallback, and that strip is a
blind whole-body replace. A reply to the bot names the bot in the fallback
pill ("> <@bot:server> quoted"), which is exactly what makes the message
count as a mention under the default MATRIX_REQUIRE_MENTION=true -- so the
strip runs on every reply-to-the-bot and rewrites the pill to "> <>", after
which the pill regex no longer matches. reply_to_author_id is lost,
reply_to_text becomes the mangled "<> quoted" remnant, and the prompt
renders "[Replying to: "<> ..."]" instead of "[Replying to your previous
message: ...]".
Split the body into (quote block, reply text) and strip the mention from the
reply text only. The mention gate still sees the raw body, so a reply to the
bot keeps waking the bot; the visible reply text and the quote-block strip
are unchanged.
Fixes#111233
Drop the unreachable handle_message assertion in the drop tests
(_prefilter_inbound never calls it) and state the deliberate
file_comment drop decision in the allowlist comment.
Housekeeping subtypes (channel_join/leave/topic/name/purpose,
convert_to_private/public, pins, deletions) are not a person speaking,
yet _prefilter_inbound only rejected message_changed/message_deleted,
so each of them started a full agent turn in free-response channels.
Replace the denylist with an allowlist: a message passes when subtype
is absent, file_share, thread_broadcast or me_message; everything else
is dropped. Fixes#110778.
`_clarify_callback_sync` decided "no answer arrived" by testing whether the
response text starts with '[' (the shape of the timeout / undeliverable
sentinels). A real answer can start with '[' too — a "[A] staging" choice
label picked by number, or "[urgent] ..." free text after Other — so the
clarify resolved and the agent got the answer, yet the Slack card was
rewritten to "This prompt expired" and typing was never re-armed.
`_clarify_send_then_wait` now returns `(response, answered)` and the runner
branches on that flag only.
The Slack click handler popped the retire entry as soon as Other was
clicked, but Other is not terminal: the clarify stays pending for typed
text, so a later timeout or /new reset found nothing to retire and the card
stayed stuck on "Awaiting typed answer". The entry is now popped only on a
terminal outcome (a choice click, or Other on an already-dead entry).
A typed answer to a native card (numeric pick, or text after Other) never
reaches the click handler, so the card kept its buttons forever; the
TEXT_RESOLVED intercept now retires it with the answer.
Review finding: '[' prefix mistaken for the timeout sentinel; Other click dropped the retire entry; typed answers never rewrote the card.
One adapter-facing seam replaces the Slack-only callback: an adapter whose
clarify prompt is a persistent card (Slack Block Kit) defines
`retire_clarify_card(clarify_id, notice)`, and the gateway calls it from
every path that ends a clarify without a button click:
- TurnRunner._clarify_callback_sync: when the bounded wait returns a
sentinel (timeout, /new or run-end clear_session), schedule the retire
with the expired notice on the gateway loop (#110821).
- run_inbound TEXT_REJECTED_PROSE: retire with the cancelled notice before
the prose is routed as a follow-up (#111019). Lookup is on the adapter
class so MagicMock doubles cannot fabricate the method; no platform ==
SLACK special-case.
The Slack map is keyed by clarify_id and popped before the first await, so
a late timer cannot touch a newer prompt and the button handler's ts-keyed
guard makes a racing click a no-op. Gateway-restart-orphaned cards stay
out of scope: nothing is waiting on the new process, and the click path
already renders them expired.
Tests trimmed to invariants: the runner-level timeout probe (card adapter
vs no-card adapter), the inbound prose retire, and one Slack test covering
buttons-dropped + late-click-noop. Docs updated for the new in-place edit.
- Drop the 900s "inbound liveness" INFO line from the salvage: a periodic
log heartbeat is a feature with a separate scoping call (#111211 item 3);
the existing stall watchdog already escalates when getUpdates stops.
- The polling error_callback interpolated the raw exception and left the
redaction call inside the format string, so its lines read
"Telegram network _redact_telegram_error_text(error), scheduling
reconnect: ..." and leaked unredacted text. Redact for real.
- Tests: keep two invariants (recovered wording after network errors;
clean bootstrap stays "confirmed healthy") and the transport test.
httpx timeout exceptions (ConnectTimeout, ReadTimeout, ...) stringify to
"" so every adapter log line built from _redact_telegram_error_text()
ended with a blank reason. Fall back to the exception class name.
Partial salvage of #111222: only the _redact_telegram_error_text hunk;
the transport-layer and polling-recovery hunks are covered by #111221.
Moving the dedup flush and thread lookup from asyncio.to_thread onto the
adapter-owned pool fixed the torn-down default executor, but to_thread also
copies the caller's contextvars and run_in_executor does not. A multiplexed
profile's HERMES_HOME override and secret scope are contextvars, so those
workers silently ran under the launch profile. _run_blocking now runs the call
through contextvars.copy_context().run, matching to_thread semantics.
Review finding: _run_blocking lost the profile HERMES_HOME override / secret scope on the worker.
`_fetch_last_message_in_thread` was the last hot-path `asyncio.to_thread`
in the adapter: after a default-executor teardown (#111020) thread-reply
routing would fail the same way the dedup flush did. Route it through
`_run_blocking` like every other blocking SDK call. The remaining
`to_thread` users (`_load_lark_oapi` at connect/onboarding, the voice
transcode with its file-attachment fallback) are cold or degrade cleanly.
A dead background event loop tears the loop's default executor down, and
after that every inbound message was dropped inside the dedup gate with a
RuntimeError out of asyncio.to_thread — the adapter went permanently deaf
while the gateway process, websocket and service all stayed healthy.
#10849 already moved the outbound SDK calls onto an adapter-owned,
self-healing pool; this gives the inbound dedup-state flush the same
treatment, so a default-executor teardown can no longer wedge message
intake.