A short Telegram drop used to fail the final reply immediately. The
answer then sat in the delivery ledger until the next gateway boot.
Wait up to 15s for the bot (or a replacement adapter) so a brief
blip delivers now, matching QQBot.
find_spec("microsoft_teams") can be true from sibling namespace packages
while App is still None, so a failed lazy-install crashed connect with
'NoneType' object is not callable instead of a missing-SDK error.
Probe microsoft_teams.apps via the parent first — a dotted find_spec
raises ModuleNotFoundError on 3.11 when the namespace is absent.
The adapter's private thread-deadline helper was the ancestor of the
unified deadline layer's run_bounded_async (#85147 was extracted from
it, plus the caller-cancellation leak fix the original still lacked).
Consolidate: the helper body becomes a thin wrapper mapping
BoundedResult.timed_out back to the asyncio.TimeoutError its 9 call
sites (the PTB retry ladder) expect. ~90 duplicated lines die, along
with the adapter-local copies of the abandon-cleanup runner and the
blocked-loop faulthandler diagnostics (both live in agent/deadline.py).
Everything the call sites rely on is preserved by the unified layer:
- thread-timer deadline that survives a blocked event loop (#63309)
- abandonment of cancellation-shielded tasks (PTB/httpcore anyio init)
- detached best-effort on_abandon cleanup (no httpx pool leak per retry)
- off-loop stack dump when the loop never processes the expiry
Plus one behavior IMPROVEMENT inherited from the shared copy: a caller
cancelling the wrapper no longer leaks the inner task unobserved (the
telegram original had that leak; the extraction fixed it).
test_telegram_init_deadline.py: the #63309 diagnostics probe now pins
the shared layer's dump hook (label "telegram-init") — same contract,
new seam. Wedge + cleanup-crash tests pass unchanged.
Under gateway.multiplex_profiles, secondary profiles are constructed
inside _profile_runtime_scope and their .env lives in the profile's
secret scope - gateway/run.py explicitly does NOT mutate os.environ with
it. Four adapters still read their AUTHORIZATION config via raw
os.getenv, so every secondary profile either (a) silently missed its own
env-only allowlists/policies (fail-closed: all DMs dropped at intake) or
(b) inherited the default profile's GATEWAY_ALLOW_ALL_USERS=true /
allowlists from the shared process env (fail-open admissions):
- weixin.py: WEIXIN_DM_POLICY / WEIXIN_ALLOWED_USERS /
WEIXIN_GROUP_ALLOWED_USERS / WEIXIN_ALLOW_ALL_USERS +
GATEWAY_ALLOW_ALL_USERS in _open_dm_opted_in
- yuanbao.py: YUANBAO_DM_POLICY / DM_ALLOW_FROM / GROUP_POLICY /
GROUP_ALLOW_FROM / ALLOW_ALL_USERS (new _yb_secret helper; AccessPolicy
hard-gates intake)
- signal.py: SIGNAL_GROUP_ALLOWED_USERS / SIGNAL_ALLOWED_USERS (new
_sig_secret helper; empty scoped group list previously meant "drop all
groups" silently)
- wecom/adapter.py: WECOM_DM_POLICY / WECOM_ALLOWED_USERS /
WECOM_GROUP_POLICY / WECOM_ALLOW_ALL_USERS + GATEWAY_ALLOW_ALL_USERS -
while credentials one line above already used _get_scoped_secret
- gateway/run.py::_own_policy_open_startup_violation: the open-policy
startup guard validated GATEWAY_ALLOW_ALL_USERS via raw os.getenv even
though its sibling dm/group reads already used the scoped _getenv
All reads now go through the canonical fail-closed scoped shape QQ's
_resolve_qq_secret already used (scope hit wins; unscoped single-profile
callers keep legacy os.environ behavior). Regression suite drives the
real scope contextvar across all four helpers plus the admission gates
and the startup guard, asserting both directions: profile values are
visible under multiplex, default-profile values never leak.
Fixes#93522
WEIXIN_DM_POLICY/ALLOWED_USERS/GROUP_ALLOWED_USERS, YUANBAO's equivalents,
WECOM_DM_POLICY/ALLOWED_USERS/GROUP_POLICY, and the startup guard's
GATEWAY_ALLOW_ALL_USERS check still read raw os.getenv at adapter
construction time. Under gateway.multiplex_profiles that reads the process
env instead of the per-profile secret scope, so a secondary profile either
silently drops every DM (its own env-only allowlist is invisible) or
inherits the default profile's allow-all/allowlist config.
Route these reads through the existing scoped helpers (_wx_secret,
_get_scoped_secret, gateway.authz_mixin._platform_gate_env, and
gateway.config._getenv) already used for the adjacent credential reads in
the same adapters.
Fixes#93522.
Extract _FLOOD_INLINE_WAIT_CAP_SECS + _flood_cap_result so the 5s cap
and the flood_control:{wait} error contract cannot drift between the
edit path and the send path #92173 added.
Final-review follow-up: swap the bare [:25] slice for the new constant.
Behavior-identical (same 25); removes the last bare option-cap literal
in the file. ChoicePickerView feeds finite /reasoning and /fast choice
lists, so no functional change.
Final-review follow-up: replace the bare 75 in the shown-count with
_DISCORD_MODEL_SELECT_CAPACITY so it can never desync from what the
partitioned menus actually render.
Telegram RetryAfter on send() slept the server retry_after with no
ceiling, so a 97-minute penalty pinned the coroutine. Mirror the edit
path: waits over 5s return immediately; short waits still retry inline.
Review follow-ups for the salvaged picker partitioning:
- Drop the _model_chunks stash — it went stale when navigating to a
provider with an empty model list (early return skipped the
reassignment), producing a wrong 'N more available' count. Derive
shown = min(len(models), 75) directly instead.
- Add _DISCORD_SELECT_MAX_OPTIONS / _DISCORD_SELECT_MAX_ROWS constants
per the file's named-limits convention; replaces 4 bare literals.
- Remove dead total_rows variable.
The Discord /model picker built a single discord.ui.Select filled with
models[:25], silently dropping any models beyond the first 25. Discord
caps a single select at 25 options but allows up to 5 component rows, so
partition the list across up to 3 select menus (25 each; Back/Cancel use
the other 2) instead of truncating.
This fixes providers like Nous (curated list + Portal recommendations
exceed 25) whose tail — including free-tier :free Portal picks — was
previously clipped on Discord while showing fine in the Portal UI / CLI.
- _build_model_select: slice into <=25-option chunks, one select per chunk
(custom_id model_model_select_<i>), all via _on_model_selected. Multi-row
menus get a (n/total) placeholder suffix.
- _on_provider_selected: 'N more available' count reflects models actually
rendered across the partitioned menus.
- Add regression test covering the 37-model Nous case (no truncation/dupes,
per-menu 25 cap holds).
prefers_fresh_final_streaming read only the raw direct_messages_topic_id
key; the adapter's canonical accessor _metadata_direct_messages_topic_id
also accepts the documented telegram_direct_messages_topic_id alias
(treated as equivalent in gateway/delivery.py), so an alias-only lane
would still flatten tables. Route the gate through the accessor and pin
the alias with a regression (mutation-checked: raw-key gate fails it).
Also reshape the happy-path endpoint assertion into the actual invariant
(sendRichMessage present, no rich draft frames) instead of a frozen call
list. Surfaced during review of PR #91436.
#91241 stopped root-DM tables collapsing to bullets by keeping native
draft transport when rich_drafts is off. Private Telegram topics still
reject sendMessageDraft (string thread ids, forum-style thread fields),
so the stream consumer falls back to edit-in-place. Telegram then
rejects a rich edit of that plain MarkdownV2 preview and format_message
permanently rewrites pipe tables into bullet lists — the remaining
report after that merge.
Route drafts through the same integer topic kwargs as send(), and on
that degraded topic path prefer a fresh sendRichMessage (then delete
the preview) instead of the table-to-bullets formatter.
The network-error reconnect path (PR #91524) was the only site converted
from asyncio.wait_for to _await_with_thread_deadline. The same
cancellation-shielding vulnerability exists at two more updater.stop()
sites:
- Conflict-retry path: asyncio.wait_for could hang forever if PTB/AnyIO
cleanup swallowed CancelledError, stalling the conflict-retry ladder.
Now uses _await_with_thread_deadline and escalates to fatal on timeout
(same reasoning: cannot safely reuse an Updater whose lifecycle lock
may still be held).
- Conflict-exhausted fatal path: asyncio.wait_for could hang before the
fatal notification fired. Now uses _await_with_thread_deadline; the
timeout handler already proceeds to fatal notify, so no behavior change
beyond the deadline mechanism.
All three asyncio.wait_for(updater.stop()) sites now use the
thread-deadline helper consistently.
Use the existing wall-clock deadline helper for updater.stop() during network recovery. If PTB cleanup remains cancellation-shielded past the deadline, escalate to retryable fatal recovery so the runner builds a fresh adapter instead of calling start_polling() while the old Updater may still hold its lifecycle lock.
Add regression coverage with stop() swallowing cancellation while holding the same lock start_polling() needs, and verify the old Updater is never reused.
Both polling reconnect paths end on the same 'health pending getUpdates
progress' line, and _record_polling_progress completed silently — so the
log stream for 'reconnected and healthy' was byte-identical to
'reconnected and hung', and a wedged long-poll (#87057 / #69314 /
#71239 class) stayed invisible until a user noticed silence. The only
detection method was sending the bot a test message (#90504).
Emit one INFO on the first confirmed getUpdates round-trip of each
generation, inside the existing event-set branch so steady-state polling
adds no log volume. This turns the pending line into a resolvable pair
('health pending' -> 'confirmed healthy') whose absence after a
reconnect is a reliable hung-poll signature.
Fixes#90504
Adapter ingress derives a session key BEFORE the runner stamps
source.profile in _make_profile_message_handler, so the namespace fell
back to the active profile and every bot in a multiplexed gateway
produced agent:main:<platform>:<chat>. A Telegram private chat reports
the user's own id as chat.id, identical for every bot, so two profiles
sharing one human collapsed onto a single lane: _pending_text_batches,
_active_sessions, the busy-session guard and _post_delivery_callbacks are
all keyed on that string. A day of production logs across two bots shows
60 flushes, none carrying the secondary profile's namespace.
set_owner_profile records credential ownership on the adapter and
_session_key_profile resolves the namespace as source.profile ->
_owner_profile -> the session store's resolver, so a secondary adapter
keys into its own namespace even before the source is stamped. Stamped
sources keep priority, so relay/connector ingress, which routes per event
rather than per credential, is unchanged. _configure_profile_adapter
installs the owner alongside the other handlers, covering startup and
reconnect.
Every candidate is type-checked as a non-blank str, and every attribute
read goes through getattr: adapters are routinely built without
BasePlatformAdapter.__init__, and a duck-typed session store returns a
truthy non-string that would otherwise be interpolated into the key as
agent:<MagicMock ...>:.
Also routes the four call sites that passed no profile at all (feishu
media batches, raft, slack _session_key_for_source, telegram photo
batches) through the same resolver.
test_multiplex_busy_input_mode's secondary-adapter busy case seeded
_active_sessions with the unstamped agent:main: key, asserting the
pre-fix collapse. It now seeds the lane the profile-owned adapter
actually derives.
A primary adapter has no owner and an unstamped source, so it resolves
exactly as before; with multiplex_profiles off the resolver returns None
and every key is byte-identical to today's.
Healthy IPv4-first connect is the new default path, so two transports
were warning on every successful initialize. Keep warning only when a
literal actually failed first. Also restates the transport docstring
and docs to match IPv4-first, hostname last.
A blackholed IPv6 path to api.telegram.org never errors, so
_await_with_thread_deadline never fires and connect hangs at
"attempt 1/8". Known A-record IPs connect over IPv4 immediately.
DoH timeout now fail-opens to the seed IPv4 list instead of the
hostname. Hostname stays last for IPv6-only hosts.
Closes#87015
Matrix m.audio/m.file/m.video events populate content.body with the
uploaded filename when the sender adds no caption. The adapter already
blanks that for m.image (PR #16821, issue #13482) but not for audio,
file, or video msgtypes, so the filename survives into event.text and
is appended after the transcript, where the model reads it as the
user message rather than as transport noise.
Extend the existing adapter-level blanking: add
_looks_like_matrix_media_filename() with the same conservative
heuristic (single token, no whitespace, no path separators, known
media suffix or mimetypes audio/video match) and apply it at the
media message handler for m.audio, m.file, and m.video.
Salvage of #87968 by @AiwendilInTheWoods — reworked from shared
gateway path to adapter-level fix for consistency with the existing
m.image blanking.
`check_telegram_requirements()` re-imports python-telegram-bot after a
lazy install and rebinds the module-level aliases that the top-level
`except ImportError` block set to `typing.Any`. TypeHandler was left out
of all three places: the `global` declaration, the
`from telegram.ext import (...)` list, and the assignments.
So whenever the top-level import fails and the deferred path runs, every
other alias is restored and TELEGRAM_AVAILABLE flips to True, while
TypeHandler stays `Any`. Handler registration then raises
`TypeError: Any cannot be instantiated` and the gateway reports:
[Telegram] Failed to connect to Telegram: Any cannot be instantiated
Gateway started with no connected platforms
The 22.6 -> 22.8 pin bump named in #85272 is the trigger rather than the
defect: it makes the top-level import fail, which is what routes the
module through the deferred path where the omission has always been.
With gateway.multiplex_profiles enabled, the primary Telegram message
handler is the closure returned by _make_default_profile_message_handler(),
so its __self__ is absent. The early intake filter
(_is_user_authorized_from_message) recovered the GatewayRunner via
self._message_handler.__self__ and, finding none, fell back to env-only
authorization — never evaluating the configured chat allowlist through
GatewayRunner._is_user_authorized(). Every non-global sender was then
default-denied in an explicitly allowlisted group.
Prefer the platform-bound authorization callback registered via
set_authorization_check(): it routes through the runner's full auth chain
(platform + group allowlists, pairing store, allow-all) and survives the
closure wrapping, whereas the bound-handler lookup does not. The bound
handler remains the fallback for setups without a registered callback, and
the pairing-passthrough guard for unknown DMs is preserved.
Fixes#87132
Rebased onto current main. `hermes_cli/plugins.py` grew 103KB -> 265KB
across 49 commits since the original branch point, and the attribution
mechanism this change hooks into was replaced along the way: the
`_tools_before` / `_plugin_tool_names` snapshot diff is now a
registration ledger sliced from `registration_start`, and `_plugin_id`
is `plugin_key`.
Re-anchored accordingly:
- Discovery-time pre-registration, module reuse, and the `provides_tools`
opt-in are unchanged.
- Attribution credits `_predeclared_tools` ahead of the ledger slice,
since those tools registered before `registration_start` and the slice
cannot see them.
- A failed materialization no longer carries attribution across. The
failure path now sweeps the whole ownership ledger for the plugin key,
not just the `registration_start:` slice, so the pre-registered tools
are disposed along with the adapter. Attribution and the registry now
agree at zero instead of reporting tools the process is not serving.
tests/hermes_cli/test_deferred_platform_client_tools.py 13/13.
test_plugins.py, test_plugins_cmd_list.py, test_plugin_cli_registration.py
65/65.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On macOS (256 soft fd limit), routing the weixin/email pollers through a
local HTTP proxy leaked one TCP socket per failed poll/connect cycle
until the gateway hit `[Errno 24] Too many open files` and crashed
(launchd respawn loop). Live capture showed 216 of 256 fds pinned on
connections to the proxy, ~214 of them abandoned.
Code-side gaps fixed:
- email adapter, `connect()`: no try/finally around the IMAP test
connection — a failure in login/ID/select/search abandoned the
connected socket with no owner. Every reconnect-watcher retry builds
a fresh adapter, so each retry against an unreachable/proxied host
leaked another fd. Teardown now runs in `finally`.
- email adapter, IMAP teardown: `imaplib.IMAP4.logout()` only swallows
`OSError` internally; on a broken connection `LOGOUT` raises
`IMAP4.abort` before the internal `shutdown()`, leaving the socket
open. New `_close_imap()` helper chases a failed `logout()` with an
unconditional `shutdown()`; used in `connect()` and
`_fetch_new_messages()`.
- weixin adapter: repeated poll failures through a proxy strand
sockets in the aiohttp connector where the tight keepalive reaper
never sees them. The poll loop now recycles its ClientSession
(swap-then-close, safe for concurrent `_process_message` tasks)
after each MAX_CONSECUTIVE_FAILURES streak, tearing down the
connector and every socket it holds.
Targeted tests: tests/gateway/test_poller_fd_lifecycle.py (9 tests).
Reported by @EthanHunter1229 with measured fd captures.
Pre-existing inconsistency: _flush_media_group_event used return
in CancelledError while _flush_text_batch and _flush_photo_batch
used raise. Changed to raise for consistency and to properly
propagate task cancellation. Made more visible by the hold-queue
changes in #83878.
Address review on #83878:
- Permanent fatal fences all hold producers and discards pending maps on
teardown instead of re-populating a queue that can never drain.
- Any hold created while connected schedules a tracked redispatch (cancel-
after-pop no longer orphans until a future reconnect).
- Redispatch failures re-hold current + remainder without tight-looping.
Regression coverage for the three residual paths, plus the interaction
with OOF-156's connect-failure classification: the retryable network
path (telegram_connect_error) must NOT clear the hold queue — reconnect
is precisely what drains it; only non-retryable fatals discard.
The disconnect drop-guard (#55971) correctly prevents dispatch into a
torn-down session. Destroying the event was wrong: by enqueue/flush time
python-telegram-bot has already acked the update and advanced the polling
offset, so Telegram never redelivers. Result: silent permanent loss, no
log, no error.
Hold inbound events (text/photo/media-group) when the drop-guard fires,
salvage pending batch maps on teardown, cancel+await the redispatch task
in the delivery cancel map (lifecycle-tracked), and redispatch from
_mark_connected after reconnect. Cap the hold queue (default 64), dedupe
by object identity, discard on non-retryable fatal. Cancel-after-pop in
flush paths also holds.
Distinct from #72037 (cancel-after-pop during follow-up supersession) and
#81528 (boundary discard). Tests use delay=0 and entered/release Events —
no wall-clock races; includes production terminal-step coverage.
Adds hermes_cli/model_selection_guards.py: a single evaluation point that
runs every selection guard (cost + the new data-policy guard) and returns
the warnings that fired. All seven model-selection surfaces (CLI picker,
cli.py TUI modal, gateway typed /model, dashboard web_server, TUI gateway,
Telegram and Discord pickers) now call the registry instead of importing
model_cost_guard directly — so the data-training-tier warning from
PR #81416 fires everywhere at once, and future guards need zero surface
wiring.
Guard modules keep their public APIs; existing mock patch points
(hermes_cli.model_cost_guard.expensive_model_warning) remain valid.
Slack session keys include the workspace id since #70190, but the kanban
notifier rebuilds the wake source from a subscription row that has no scope
column, so every terminal-event wake keyed without the workspace.
The legacy-key adoption shipped in the same change (`_legacy_slack_session_key`,
`_recovered_row_matches_source_scope`) resolves that unscoped key onto the same
session_id, so the wake passes the busy guards that are keyed by routing key
(`_active_sessions`, `_running_agents`) and only collides afterwards, on session
id, under the per-session turn lease (#64934) — which serializes it behind the
live turn's flush. On a live Slack gateway that shows up as a duplicate run on
one task plus 400+s of waiting before the woken turn starts.
Same failure mode as #56580 / #72191 (chat_type), one field over, and it needs
no schema change: `_thread_metadata_for_source()` already stamps
`slack_team_id`, the notify subscription persists that dict as
`delivery_metadata`, and the notifier already unpacks it. Rows written by
`kanban_tools._maybe_auto_subscribe` carry no workspace, so fall back to the
adapter's channel → workspace map via `scope_id_for_chat()`, read with getattr
so adapters opt in and unscoped platforms' keys stay byte-identical. Slack
answers it from `_remember_channel_team`, which drops channels claimed by two
workspaces, so an unknown or ambiguous channel degrades to today's behavior
instead of guessing wrong.
Also adds the contributor email mapping the attribution check requires.
Co-authored-by: Junie <junie@jetbrains.com>
Adds platforms.slack.extra.native_task_cards: when enabled, live tool
calls render as Slack-native plan/task cards via chat.startStream /
chat.appendStream (task_display_mode: plan, task_update chunks) instead
of text/edit progress bubbles. ID-bearing tool_start/tool_complete
callbacks correlate concurrent same-name tool calls correctly; any
native API failure falls back to one continuously edited text update.
The stream is stopped exactly once when the turn finalizes.
Salvaged from PR #29496 onto current main (TurnRunner/TurnContext seam);
closes#29483.
Slack's Agents & AI Apps feature ships a native streaming surface that
renders a live-typing message instead of the edit-based progressive
updates the adapter used until now.
The adapter now implements the existing draft-streaming interface:
- supports_draft_streaming() opts in whenever the app is connected and
native streaming hasn't been detected as unavailable.
- send_draft() starts a stream on the first frame (chat.startStream,
anchored to the resolved thread_ts, with recipient_team_id/user_id
for channel streams) and appends only the delta on subsequent frames
(chat.appendStream is append-only). The consumer's trailing cursor
glyph is stripped before delta computation.
- Unlike Telegram drafts (ephemeral, replaced by a real sendMessage),
a Slack stream IS the final message. send() therefore intercepts the
turn-final delivery for a chat with an active stream whose streamed
text is a prefix of the final content, and seals it via
chat.stopStream with the remaining delta instead of posting a
duplicate. Rich Block Kit (when enabled) is applied to the sealed
message via chat_update, mirroring the finalize path in edit_message.
- Feature-gate errors from chat.startStream (not_allowed,
missing_scope, unknown_method, ...) are cached on the adapter so
subsequent runs skip straight to edit-based streaming with a single
warning naming the fix (enable Agents & AI Apps for the app);
transient errors only disable drafts for the current run via the
consumer's existing send_draft failure handling.
- Segment breaks (new draft_id) and disconnect() seal any open stream
so chats are never left with a dangling live-typing indicator.
No consumer or config changes: streaming.transport auto/draft now
lights up native streaming on Slack through the same interface
Telegram drafts use, and the edit-based path remains the fallback.
_resolve_thread_id() falls back to _last_inbound_thread[chat_id] when no
explicit thread is present. That fallback exists for interactive DMs, where
Google Chat spawns a fresh thread per top-level user message and the adapter
drops thread_id to keep the session key stable. It also fired for cron
deliveries, which carry job_id in their metadata but no thread: the output
landed as a reply inside the last inbound thread instead of starting a new
top-level message.
Bypass the _last_inbound_thread fallback when metadata has a job_id (i.e. the
message is an automated cron delivery), so cron output posts at top level
unless an explicit thread is requested.
Slack Web API calls return `SlackResponse`/`AsyncSlackResponse`, which are
mapping-like but not `dict` subclasses, so every `isinstance(resp, dict)`
gate took its "unexpected shape" branch at runtime: user and channel names
collapsed to raw IDs, every user resolved as a non-bot (defeating the
allow_bots loop guard), ephemeral replies were reported as failures, and
uploads/caption fallbacks lost their message_id.
Normalize responses through a single `_slack_response_payload()` helper
(dict passes through, SDK response yields `.data`, anything else yields
`{}` so callers keep their fallbacks) and use it at every call site.
Existing Slack tests injected plain dicts, which is why the defect was
invisible; the new tests run each behavioral case against a real
`AsyncSlackResponse` as well.
Follow-up to the salvaged #80489 substring-fallback removal: the
structured 401/403 branch still exited with a bare return, leaving
_running True — dead listener, healthy-looking is_connected(), gateway
never told (the zombie half of the bug, OOF-156 class). It now sets a
non-retryable mattermost_auth_error with token guidance and notifies
the gateway fatal handler.
Also: pytest.importorskip for aiohttp in the verifier probe file
(module-level import crashed collection in envs without the optional
dep), and probe fixtures updated for the escalation attributes.
The WS reconnect loop had a fallback check that looked for "401", "403",
or "unauthorized" as substrings anywhere in an exception's string form.
A transient error whose message happens to contain those digits (a proxy
body, a stack trace, anything) got treated as a permanent auth failure
and stopped reconnection for good.
I removed the substring fallback and kept only the structured check:
aiohttp.WSServerHandshakeError with status in {401, 403}. That's the only
signal that reliably means the server rejected our credentials.
Added two regression tests: one proving a transient error containing
"401" in its text still retries, and one confirming the existing
_closing early-return path is untouched by the removal.
Follow-ups to the salvaged #80032 fatal-error escalation, closing the
gaps its review thread identified plus a sibling of the same class:
1. Partial-batch loss: _check_inbox now dispatches whatever the fetch
returned BEFORE escalating a failure — the early-return dropped
already-fetched messages whose UIDs were marked seen.
2. Seen-after-fetch: UIDs enter _seen_uids only after their fetch
returns a response, so a mid-batch connection failure leaves the
remaining UIDs eligible for the next poll. Per-message processing
moved to _parse_fetched_message behind a poison guard: a message
that fails parsing/auth-verification is marked seen, logged with
its UID, and skipped once — never an eternal crash loop.
3. Reconnect mail loss: connect(is_reconnect=True) restores the
account's seen-UID baseline from a class-level snapshot instead of
re-marking the entire mailbox seen — mail that arrived during an
outage is now processed after the reconnect the escalation triggers.
7 new regression tests.
_fetch_new_messages() wrapped the whole IMAP connect/login/select/search/
fetch sequence in a bare except that logged and returned an empty list —
indistinguishable from a genuinely empty inbox. The adapter never invoked
its fatal-error handler, so the gateway's reconnect/backoff/status
machinery never learned the mailbox was unreachable; outages lasted until
a manual restart.
Track fetch failure on the adapter and, when the poll loop observes it,
set a retryable fatal error (email_imap_fetch_failed) and notify the
gateway handler so the platform enters the reconnect queue just like a
startup connection failure.
The cherry-picked #79448 predated #85049's _classify_connect_exception,
so it added a parallel PrivilegedIntentsRequired branch ahead of the
classifier (plus its own _is_privileged_intents_required detector).
Fold the tailored guidance into the classifier's existing intents arm
instead: one classification path, one error code (discord_intents_required),
and the message now names exactly the intents Hermes requested (Message
Content always; Server Members only when username/role allowlists need it).
Wizard callout, docs corrections, and tests from #79448 kept as-is.
PrivilegedIntentsRequired is a Developer Portal config error; surface which
intents Hermes requested as a non-retryable fatal and teach setup/docs.
Co-authored-by: Cursor <cursoragent@cursor.com>
- Collapse the duplicated discord LoginFailure/PrivilegedIntentsRequired
classification (name-match + isinstance blocks repeated the same code/
message tuples) into a single _is() helper — one message per failure.
- Replace the user-facing HERMES_RECONNECT_ATTENTION_AFTER_SECONDS env var
with agent.reconnect_attention_after in config.yaml (default 7200, 0
disables), bridged internally like gateway_timeout. .env is for secrets.
- Use _float_env for robust parsing instead of bare int(os.getenv(...)).
- Document terminal classification + needs_attention escalation in
website/docs/user-guide/configuration.md.