Follow-up to the salvaged #92440 commit, shape-gate cleanup only; behaviour is unchanged
(one mention-bearing payload per logical send, captioned media keeps its caption).
- Fold `_send_whatsapp_with_mentions` into `_send_plugin_standalone` (it was a line-for-line
copy of the caption split + `_send_chunks` loop); `mentions` is attached to the first
payload only via a one-shot kwarg dict.
- Drop the `inspect.signature(sender)` probe: the only registered WhatsApp standalone sender
is the in-tree `_standalone_send`, which gains `mentions` in the same change; a foreign
sender already surfaces as a TypeError through `_handle_send`'s error path.
- Collapse `_normalize_outbound_mentions` to dedupe-only; its input is argparse `list[str]`
already validated by the CLI.
- Revert the `\d` -> `[0-9]` edit to `_BARE_PHONE_RE` / `to_whatsapp_jid`: unrelated to the
feature (`normalize_whatsapp_mention_jid` already rejects non-ASCII via `isascii()`) and it
changed output for nine other `to_whatsapp_jid` callers.
- Split the single 137-line test into a `whatsapp_bridge` fixture + two invariants
(rejections never reach the bridge; mentions ride the first payload only, stale bridge
fails closed); drop the fabricated legacy-sender branch that only existed to cover the
deleted probe. Mutation check: removing the first-payload gate turns the new test red.
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
The OpenRouter video content URL is built from OPENROUTER_BASE_URL, not from
a provider response, yet save_url ran the full SSRF check on it. An operator
pointing base_url at a LAN/loopback relay could submit and poll (raw
requests) but the final download was refused as SSRF unless they set
security.allow_private_urls — contradicting the issue scope ("operator-
configured endpoints out of scope") and the security.md sentence that the
operator's own base_url is unaffected.
save_url grows a `trusted_origin` flag (plumbed through save_url_video and
set only by OpenRouterVideoGenProvider._save_completed_video): the first hop
skips the private-address class check and uses a plain client, but the
cloud-metadata floor (is_always_blocked_url) still applies, auth headers
stay on hop 1 only, and every redirect target is re-validated in full so a
relay cannot bounce us to another internal address. Provider-returned result
URLs (fal, xai, image providers) keep the full guard — default is False.
security.md now states the precise scope: only the direct base_url hop is
exempt; result URLs from a LAN-hosted provider still need allow_private_urls.
Provider response URLs, model-supplied image refs, manifest-derived pet
URLs, and remote sitemap <loc> entries were fetched with raw
requests/httpx/urllib — bypassing tools/url_safety while every platform
media path already uses it. A hostile or compromised provider/manifest
endpoint could steer a server-side fetch at internal or metadata
addresses; several sites cache the body where it is deliverable back.
Apply the canonical is_safe_url + create_ssrf_safe_client pattern at
every site: per-hop revalidation at TCP connect (closing the
DNS-rebinding window), bounded redirect chains that fail closed on
missing Location, and caller headers scoped to the first hop only —
matching the openrouter provider's own documented contract that its
bearer key must never leave the operator-selected host.
Operator-configured endpoints and pinned release assets are out of
scope — those URLs are operator-selected, not remote-party-controlled.
Fixes#114468Closes#44728
Generated with [Devin](https://devin.ai)
Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
Co-authored-by: Ray <rayjun0412@gmail.com>
Co-authored-by: zapabob <1920071390@campus.ouj.ac.jp>
The create/settings dialogs (salvaged from #114664) let a user set or clear
a board's project_id, but the binding was invisible outside Settings.
GET /boards already annotates every board with project_id + project_name,
so the switcher now renders a "Project: <name>" badge whose × sends
PATCH {project_id: ""} — the same clear the settings dialog uses — without
touching default_workdir.
Settings also stops sending default_workdir: "" next to a chosen project
when the directory field is blank: the explicit "" suppressed the server's
project → default_workdir mirror, so binding via Settings left the board
without the workspace default the create dialog would have seeded.
Trim the salvaged test to the invariants (selector wiring, payload shapes,
badge unbind) and document the control on the Kanban docs page.
The REST API already accepts project_id on board create/update and
GET /boards annotates project_id/project_name, but the dashboard UI
never wired any of it: a board's project binding could only be set
through raw API calls. Add a project selector to the New board and
Board settings dialogs, populated from GET /projects, mirroring the
existing default_workdir wiring. Settings PATCHes send project_id
unconditionally when the selector rendered ("" clears the binding,
server-validated); when the projects store is unreachable the field
is omitted so saving unrelated settings cannot wipe a binding.
Fixes#114652
A reply past 4,096 chars goes out as several sendMessage calls. When chunk 2 was
refused by flood control (RetryAfter past the 5s inline cap) send() returned the
bare flood_control result, so _send_with_retry re-sent the WHOLE payload after
the wait: the user saw chunk 1 twice (reporter: 4 messages, 686 duplicated words),
and paths without a ledger row lost the tail outright.
- send() now reports a mid-split refusal through the existing partial_overflow
contract (the key _edit_overflow_split already sets and the stream consumer
reads): delivered_chunks / total_chunks / last_message_id, plus
undelivered_chunks + delivered_message_ids ONLY when non-delivery is certain
(flood cap, Bot API rejection, connect/pool timeout) — an ambiguous TimedOut may
have reached Telegram and is never resumed from.
- BasePlatformAdapter._send_with_retry resumes from the remainder via a new
_resume_partial_send hook (default None = keep the partial failure, never
re-send the head; the plain-text fallback is skipped for partials too). The
Telegram override sends the leftover formatted chunks, continuing the id sequence.
- Per-chat FIFO send gate, reentrant per asyncio task (media paths nest), held
only around the API calls — never across the reconnect wait — on send() and the
media funnel, so concurrent replies to one chat no longer interleave chunks.
- A flood refusal arms a per-chat cooldown (mirrors the sendChatAction cooldown,
capped 300s); sends inside the window fail closed locally with the same
flood_control:<s> result and no API call, so ledger recognition and redelivery
timing are unchanged.
- Edit path: log "refusing (retry_after Ns > cap)" after the cap check instead of
"waiting Ns" followed by no wait.
Live against a local fake Telegram Bot API with a fake token (RetryAfter=7 on
chunk 2 of a 3-chunk reply): before 4 messages / 324 duplicated words; after 3
messages, 853/853 words, 0 duplicated, 0 lost. Two concurrent 3-chunk sends:
before 9 source switches, after 1. Five sends inside a refused window: before 5
API calls, after 0.
Fixes#114396
Co-authored-by: AStrnbrg <45151087+AStrnbrg@users.noreply.github.com>
Co-authored-by: whyyagswhy <166958865+whyyagswhy@users.noreply.github.com>
The Feishu fix on this branch populates MessageEvent.media_text_inlined so
run_inbound's document note stops claiming "Its content has been included
below" when a text attachment was NOT inlined (>100 KB gate or decode
failure); run_inbound treats a missing flag as inlined. Telegram, Discord,
Slack, the WhatsApp bridge adapter and whatsapp_cloud inline "[Content of
…]" the same way but never set the flag, so their notes lied on the
skip path. Mirror the Feishu/buzz per-attachment contract in each: False
for every cached attachment, flipped to True only when the text was
actually injected.
One parametrized (small→True / large→False) test per adapter in the
existing per-platform test files.
The exit-notify wrap reported every receive-loop exception at ERROR with a
traceback, including the ConnectionClosedOK that follows our own CLOSE frame in
disconnect(). Gate on the adapter's _running flag (published through the WS
thread-local next to on_link_up): a live link's death stays ERROR, an
intentional shutdown logs at DEBUG. Live pass side-effect #4 on #113662.
The ``retrying`` publication landed only on the supervisor-rebuild path (WS
thread dies). On the live link ``_auto_reconnect`` is on, so a receive-loop
error runs lark-oapi's ``_reconnect()`` ladder *inside* the receive loop — the
thread never dies and ``gateway_state.json`` kept saying ``connected`` while
the link was down.
Set the SDK's ``Client.on_reconnecting`` observer (lark-oapi 1.6.8
ws/client.py, fired first thing in ``_reconnect()``) from
``_apply_runtime_ws_overrides``; it hops to the adapter loop and publishes
``retrying`` through the same runtime-status call the supervisor uses, gated
on the client still being the live one. ``connected`` is re-stamped by the
existing ``_ws_link_up`` hook once the ladder's ``_connect()`` schedules a new
receive loop, so no ``on_reconnected`` hook is needed.
Docs: the WebSocket-mode section now says a dead link is rebuilt by Hermes'
supervisor and shows ``retrying`` in gateway status until re-established.
Test: test_sdk_reconnect_ladder_publishes_retrying (red before: FeishuAdapter
has no _ws_link_retrying / observer never set; green after).
Follow-up to the salvaged #113668 commits:
- The wrap stopped the worker loop in a ``finally``, i.e. also on a *normal*
return. On the live link ``_auto_reconnect`` is the SDK default (True — the
adapter only flips it off at teardown), and a successful SDK reconnect makes
the old ``_receive_message_loop`` return after ``_connect()`` scheduled a fresh
one. Stopping the loop there would tear down the healthy rebuilt link (no
CLOSE frame, see #10202) on every transient blip. Stop only when the coroutine
exits with an exception: ladder disabled, or the ladder's
``ClientException`` / ``ServerUnreachableException`` re-raise.
- Do not re-raise after logging: the wrap is now the owner of that exception,
so asyncio no longer prints "Task exception was never retrieved" beside it.
- Wrap the real symbol unconditionally: ``lark_oapi.ws.client.Client
._receive_message_loop`` exists on every shipped SDK (1.6.x and 1.7.3, ws
module unchanged between them); a getattr fallback would silently reopen the
gap on a rename.
- ``gateway_state.json`` kept saying ``connected`` while the supervisor
rebuilt: publish ``retrying`` when the WS thread dies (the adapter is still
running, so ``_mark_disconnected`` is the wrong tool) and re-stamp
``connected`` from ``_ws_link_up`` — fired through a thread-local hook when
the SDK schedules a receive loop, the only in-thread proof the handshake
succeeded — hopped onto the adapter loop and gated on the client still being
the live one.
- Tests: the fake ``lark_oapi.ws.client`` modules carry ``Client`` like the real
module; second invariant now covers the normal-return (ladder success) case;
the existing supervisor-restart test asserts the ``retrying`` publication.
Live: real lark-oapi 1.6.8 ``ws.Client`` + the real adapter worker/supervisor
against a local fake Feishu endpoint + websocket server with fake credentials.
The wrapper re-raised the SDK receive-loop exception into the same
unretrieved bare task. Log it with logger.exception first so the root
cause lands next to the supervisor's rebuild line (suggested in review).
lark_oapi parks Client.start() in run_until_complete(_select()) and runs
the receive loop as a bare create_task whose exception nobody retrieves.
With Hermes disabling the SDK's reconnect ladder, a mid-life receive-loop
death leaves a deaf-but-ESTABLISHED socket: start() never returns, the
supervisor's executor future never completes, and gateway_state.json keeps
reporting connected (#113662).
Wrap Client._receive_message_loop in the existing isolation installer so
its exit stops the worker loop; start() then raises, the future completes,
and _supervise_websocket_thread rebuilds the link with its capped backoff.
Deliberate disconnects stay unaffected: they nil _ws_client first and the
supervisor exits without restarting.
Fixes#113662
The bootstrap racer covers sync connects only. The gateway's WebSocket dials
(relay connector, Yuanbao, Buzz) go through ``websockets.connect`` →
``loop.create_connection``, whose ``happy_eyeballs_delay`` defaults to ``None``:
a serial walk that burns the full connect timeout on every blackholed AAAA
record before IPv4 answers — the same stall class #114265 reports, one layer up.
``websockets`` forwards unknown kwargs to ``loop.create_connection``, so each
call site passes ``happy_eyeballs_delay=0.25`` (the RFC 8305 delay anyio and
the sync racer already use). One invariant test per call site captures the
kwargs at a mocked ``websockets.connect``.
`_on_sidecar_message` ran `_normalize_content` (which base64-decodes and writes
inline attachment/voice bytes into the media cache via `_cache_inbound_attachment`)
before the `chat_type == "group" and self.require_mention` check, so an
unmentioned group attachment was persisted and then dropped. Gate on the
user-typed text extracted without touching attachment bytes (`_mention_gate_text`:
text / richlink / group text items), then normalise and strip the wake word only
for messages that pass.
Same class as the Teams fix in this PR (review follow-up). One invariant test:
unmentioned group attachment -> 0 cache writes, mentioned -> 1 and dispatched.
`_handle_media_message` fetched the mxc:// payload (`_download_and_cache_media`)
before `_build_inbound_event` ran `_resolve_message_context`, so an unmentioned
or non-allowlisted room's m.image/m.file was pulled onto the host and then
dropped. Resolve the context first and hand it to `_build_inbound_event` via a
new `ctx` kwarg so the read receipt / thread mark are still applied once.
Same class as the Teams fix in this PR (review follow-up). One invariant test:
unmentioned group media → 0 downloads, mentioned → 1.
Follow-up to the salvaged #113591 gate:
- Teams writes the bot's conversation identity as `28:<app id>` (activity.recipient.id,
mention entities) while App.id is the bare app id. The mention match and the own-message
filter now accept both spellings, so a real tenant mention no longer fails silently.
- A payload that mentions only other people is no longer treated as a bot mention; the
`<at>` text fallback applies only when the activity carries no mention entities at all.
- `self._extra` holds `platforms.teams.extra` on the instance so keys read after
construction (require_mention today) are not dropped with a constructor local (#114366).
- Reply exemption uses one bounded deque instead of deque + mirror set.
- Tests trimmed to two invariants: gate table (drop before attachment download / keep
mention, reply-to-bot, personal) and env-over-YAML precedence. TEAMS_REQUIRE_MENTION
documented in plugin.yaml and the Teams docs page.
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Co-authored-by: finn763 <165816600+finn763@users.noreply.github.com>
Co-authored-by: Fernando Muñoz Suazo <fernandrewm@gmail.com>
The decomposer now prefers the root card's assignee for unrouted children and
falls back to the active profile only for cards without an assignee, so the
endpoint's "fallbacks filled the same way the decomposer does" was stale. State
what the endpoint actually resolves; no routing change.
Three defects let the missed-message backfill re-run the same inbound message on
every reconnect (24 turns in 12 h for one message on the reporter's install):
- Completion was only ever recorded by the send-path ledger writer, keyed on the
Discord reply anchor. With reply_to_mode "off", a streamed final delivered via
edit_message(finalize=True), a fresh-final send or a media-only reply there is no
anchor, so the row stayed processed/replied=0 and _should_backfill_discord_message
re-admitted it forever. The adapter already learns the outcome per inbound id in
_record_discord_processing_complete: on SUCCESS (base confirmed delivery, or the
stream already delivered) it now writes status='responded', replied=1.
- The scan wrote status='discovered' over every candidate at the top of each pass,
erasing the queued/processing claim so the 10-minute active-claim guard could never
fire; two back-to-back scans dispatched the same message twice. The 'discovered'
write no longer overwrites an existing row's status or updated_at.
- A stored per-channel cursor REPLACED the window floor (12-hour-old messages stayed
candidates under window_seconds: 3600) and the parent's cursor object was passed
down to child threads. after = max(cursor, now - window) and threads only see the
window floor plus their own cursor.
- max_dispatches caps one scan, not one row. New missed_message_backfill.max_attempts
(default 3; env DISCORD_MISSED_MESSAGE_BACKFILL_MAX_ATTEMPTS) is a lifetime ceiling
on re-dispatch of a single message; `attempts` now counts dispatches only.
Live repro (real DiscordAdapter, temp HERMES_HOME ledger, fake client yielding the
same message per scan, fresh adapter per scan = reconnect): streamed final with
reply_to_mode=off 3 scans -> before 3 dispatches, after 1; turn that never completes
3 scans -> 3/1; turn that always fails 6 scans -> 6/3; 12h-old cursor with a 1h window
-> before 2 out-of-window dispatches (parent + thread), after 0; control: an
unanswered message still dispatches exactly once, and a cursor newer than the window
still narrows the scan.
Part of #113631 (with the salvaged anchor fallback from #113633 by @KoNit-K).
rich_blocks turns markdown bullets into rich_text_list. rich_text does
not interpret mrkdwn, so Slack <url|text> was emitted as literal text
while [text](url) became a real link. Section/mrkdwn paragraphs were
unaffected.
Parse Slack autolinks in _inline_elements (lists, quotes, table cells).
Mentions (<@U>, <#C>, <!here>) have no scheme: and stay as text.
CustomProfile did not override supported_reasoning_efforts, so on the
Responses transport a custom:<name> relay fell through to the OpenAI
per-model ladder (codex_supported_efforts) and a configured effort=max
was silently clamped to xhigh — while the same provider over
chat-completions forwarded max unchanged, because that path already
clamps onto OPENAI_COMPAT_WIRE_EFFORTS.
Declare the same wire set from the profile so the two transports agree:
max survives, ultra still clamps down to max, and the official OpenAI
backend ladder (gpt-5.5 rejecting max) is untouched.
Fixes#114249
send_slash_confirm cut the RAW message to 3800 chars and only then ran format_message,
whose MarkdownV2 escaping expands text (3800 dots -> 7600 chars), so the button card could
still exceed the 4096 cap and fail exactly like the approval card this branch fixed. It now
fits the message with the shared _ea_fit bisection measured after format_message (the
suffix's own rendering is reserved from the budget).
_ea_fit takes an optional `escape` renderer (default _ea_escape) and measures in the
adapter's message_len_fn units, and the Telegram command budget sums the framing with
utf16_len, so astral chars (emoji) are budgeted the way the adapter's 4096 chunker counts
them (len() undercounted them by half).
Tests: slash-confirm preview for a 3800-char message and the approval card for an
emoji-dense command both stay <= 4096 UTF-16 units (red before, green after).
Telegram truncated the command to 3800 RAW chars and only then html.escape()d it, so a
command heavy in `&`/`<`/`>` rendered far past 4096 chars; the framing (header, <pre>,
reason label, deadline, smart-deny line) sat outside that budget and the reason itself
was unbounded. Telegram answered "Message is too long" and the gateway fell back to the
text /approve prompt.
- `BasePlatformAdapter._ea_fit`: the shared exec-approval core now measures both the
command preview and the reason after `_ea_escape` (bisection on the raw prefix; the
escaped length is monotonic) — identity for non-escaping platforms.
- Telegram computes the command budget from `MAX_MESSAGE_LENGTH` minus the rendered
framing and escaped reason (the same idiom Discord and Slack already use) and bounds
the reason at 500 escaped chars.
Live against a local fake Telegram Bot API (fake token) rejecting text > 4096:
before, 3700 x "&" sent 18670 chars → 400 → /approve fallback; after, 4093 chars,
inline keyboard delivered.
Co-authored-by: kelvmg <16878572+kelvmg@users.noreply.github.com>
Co-authored-by: KoNit-K <124019182+KoNit-K@users.noreply.github.com>
The exec approval prompt already keeps the command/reason in plain `content`
and sends a header-only embed, so the body renders exactly once (#114693).
`send_slash_confirm` and `send_clarify` still restated the message/question
(and the clarify reply hint) in `embed.description`/fields while also sending
the self-contained `content`, so embed-rendering clients saw the body twice —
the same defect class the PR closes. Apply the same rule to both siblings,
drop the now-unused `_embed_body` helper, and pin each with one invariant test.
#114699 removed `_EA_REASON_BUDGET = 300` together with the embed field that used it,
but the shared `_format_exec_approval` core also applies that budget to the plain
content. Unbounded, a long reason starves the command preview to zero and pushes the
content past Discord's 2000-char cap (probe: 5216 chars for a 5000-char reason), which
the API rejects. Restore the budget with its WHY; pin the cap with one invariant test
and condense the embed assertions.
complete_task and request_review return bare False when the dependency
gate refuses, and _patch_status only enriched the 409 for status=ready,
so PATCH status=done/review on a gated card answered "not valid from
current state" and the bulk entry said "transition refused". Consult
unsatisfied_parents on a refused done/review and name the parents in
the 409 detail and the bulk entry error, matching the CLI/tool wording.
The tracked-item delete path in quick() called shutil.rmtree() on any tracked
directory without consulting the protection list, which only the empty-dir
sweep used. guess_category() files every path under cache/ as "temp", so a
terminal command that merely mentioned $HERMES_HOME/cache tracked the
directory itself, and 7 days later quick() removed cache/ wholesale — taking
cache/terminal (terminal snapshots) with it and breaking every later command
with a mktemp "No such file or directory".
Fix: _is_protected_dir() — a tracked DIRECTORY that is HERMES_HOME itself or
sits under an _EMPTY_DIR_PROTECTED_TOP_LEVEL tree is never tracked by
guess_category(), never listed by dry_run(), and skipped (logged SKIPPED,
entry dropped) by quick(). Files under cache/ still age out as before.
kanban/ (task attachments and workspaces have their own lifecycle) is added to
both _NEVER_TRACK_TOP_LEVEL and _EMPTY_DIR_PROTECTED_TOP_LEVEL; stale pre-fix
"test" entries under it are dropped by the existing re-validation instead of
deleted. The kanban row was first proposed in #80842 (@nicha16).
Fixes#114552
Same class as the Bot Chat drain wedge already on this branch: every JSON-file
scan guarded "did it parse?" and then assumed the value was a dict. A file
holding `42`, `"oops"` or `[1,2,3]` (corruption, truncated write, foreign tool)
passed the guard and raised AttributeError/TypeError at the first `.get()`,
usually before a single healthy sibling was processed. Each site now treats a
non-object payload like a corrupt file under that subsystem's existing policy:
- tools/bot_relay.py::_expire_if_stale / claim_pending_envelopes — the
envelope is skipped by the sweep and not claimed (same as unparseable).
- tools/browser_lightpanda.py::reap_orphaned_lightpanda — record unlinked,
scan continues.
- tools/write_approval.py::list_pending / get_pending — record skipped with
the existing "unreadable pending record" warning / None.
- tui_gateway/methods_session.py::_legacy_spawn_tree_entry / spawn_tree.load —
scalar snapshot reads as empty / returns the existing 5000 error instead of
violating the SpawnTreeLoadResult contract.
- hermes_cli/local_runtime/binaries.py::manifest_verified — False.
- plugins/platforms/a2a/protocol.py::load_conversation — non-dict lines are
dropped, keeping the declared list[dict] return.
- batch_runner.py::_load_dataset / _scan_completed_prompts_by_content /
_combine_batch_files — line skipped and counted as filtered.
- trajectory_compressor.py::process_entry_async — scalar entry passed through
unchanged.
Ported from the source hunks of PR #114241; its gateway/shutdown_flush.py
drain_transcript_spool hunk is left to open PR #84785, and its
recover_pending_to_db / cron / bot_live_delivery / bot_mode_dm hunks are
already on this branch or on main.
(cherry picked from commit d4b54568887e69b3ee3d363ebe4dcd657ccf64f9)
OpenCode's free tier now returns HTTP 403 for anonymous traffic outside
the OpenCode client, so the built-in keyless provider is dead weight:
- drop the opencode-free provider row, aliases (free/opencode_free), model
catalog, keyless runtime ladder rung, header wiring, and cached slugs
- delete the model-providers/opencode-free plugin
- update tests and the compat manifest for the removed symbols
- keep a migration hint in auth.py so users who had it configured see a
clear error naming the removal
Existing opencode-free configs can move to opencode-zen (pay-as-you-go)
or opencode-go (flat subscription).
Slack seals a native stream server-side after a few minutes (live-observed
at ~5m20s on three independent long turns, 2026-09-15/16; the lifetime is
not documented). The next chat.appendStream on the card fails with
message_not_in_streaming_state. The adapter returned a bare failure, the
TurnRunner latched native_failed, and the rest of the turn rendered as an
edited text bullet list. Long autonomous turns lost the card UX exactly
when it mattered.
On message_not_in_streaming_state from appendStream, drop the dead
stream_ts and chat.startStream a fresh plan-mode card in the same thread,
then append the current frame there. Every frame already carries the full
visible task projection, so no task state is lost. One reopen per update;
a second rejection surfaces as a real failure. The sealed card is a plain
message now, so no stopStream is sent to it; the turn-final stop targets
the reopened card. Error matching reads SlackApiError.response["error"],
never the message text.
Tests assert the wire sequence (start, append, rejected append, start,
append on the new ts), the reopened frame's task states, the cache pointing
at the new card, and the stop targeting it; plus the one-reopen bound.
Mutation: forcing the expiry branch off turns both tests red.
Four places changed behaviour for users who never touched the setting:
- `_interim_send` was stamped on every `warn` status and media-failure notice, and the
Slack/relay egress doors learned to skip stream sealing for it. Main's status sends carry
no interim mark at all, so the gap is class-wide (every status kind), and fixing it for
warnings alone is an undeclared streaming-contract change. Reverted here; the whole-class
fix belongs in its own PR against gateway/AGENTS.md rule 3.
- The entire post-handler delivery (unwrap, TTS, final text, attachments, delivery-ledger
writes) ran inside `_media_delivery_scope`. Under multiplex that binds the routed home, so
delivery obligations landed in the routed profile's state.db while boot-time
`_claim_pending_obligations` still reads the launch home. Only the policy reads
(`diagnostic_wake_muted`, `warning_text`) bind the routed scope now; delivery stays where
main ran it.
- The turn-crash notice is rebuilt the same way: scope around the policy read, send outside.
- The "delivery failed after multiple attempts" notice is unconditional again: the requested
result itself was lost and this line is its only signal, so it is not a diagnostic.
Tests that asserted the reverted behaviours are removed; the reviewer-round test file is
renamed for what it covers.
Squash of the 54 commits on victor-kyriazakos:feat/user-channel-warning-suppression
(PR #112302, head f45c640e55) so the contributor's authorship survives a rebase-merge;
the commits interleave with a cron delivery-ledger rework that the salvage removes in
follow-up commits, so per-commit cherry-picks were not practical.
Adds display.suppress_warning_notifications (global + per-platform, default false):
one resolver (gateway/warning_notifications.py), BasePlatformAdapter.emit_warning /
emit_media_warning / warning_text, a notification_category classification carried
through wakes, queues and persistence, and render/present boundaries for CLI/TUI.
Discord paces a bot's sends at roughly one per second, so chunk 3 of a long
handoff lands after the 2s window anchored on the tag and was still dropped —
the symptom the continuation window exists to fix. Each admitted continuation
now re-arms the window; the gateway bot loop guard bounds a bot that never
stops. The flush-delay override moves onto the BasePlatformAdapter seam
(_text_batch_delay_for) that main relocated the batcher to.
Tests trimmed to the invariants: reply-ping-only bot message rejected by
default (and admitted with the explicit opt-out), and a 3-chunk tagged
handoff paced past the original window arrives as one batched event; the
defensive getattr fallbacks that only served test construction are gone.
Two open atoms of #83390 (DeepSeek "This response_format type is unavailable now"):
* `_call_fallback_candidate_sync/_async` only special-cased auth errors, so when the primary
aux provider failed (timeout, rate limit, payment) and the fallback landed on a provider that
rejects `json_schema`, the 400 re-raised and the whole task died — the primary-path rung from
#89589 never applied there. Both fallback paths now retry once without `response_format`.
* Every structured aux call (titles, kanban decomposer, goal judge, plugin structured calls)
paid a guaranteed-fail request on providers that lack `json_schema` before the retry. A
provider profile can now declare `unsupported_response_formats` (DeepSeek: json_schema, per
https://api-docs.deepseek.com/guides/json_mode) and the recovery ladder remembers any route
that rejected a type once (host:port scoped), so `_build_call_kwargs` — shared by the primary
and fallback paths — omits the field before the first request. Dropping rather than
downgrading to json_object matches the end state the retry already produced; json_object
needs a JSON-mentioning prompt and some relays return empty content under it.
New logic lives in agent/auxiliary_structured_output.py; the facade only gains the fallback rung
next to the predicate it uses. tests/agent/conftest.py resets the process-level memo per test.
Fixes#83390, #105191. Closes duplicates #84976, #88830, #102849, #113064.
Co-authored-by: Legion-is-life <Legion-is-life@users.noreply.github.com>
_schedule_polling_recovery promised the gateway 'stays alive and will retry' for every error, but a _PollingStallError goes straight to _go_fatal_network (supervisor rebuild). Branch the wording on the error type, drop the watchdog's own pre-log so a stall yields exactly one error-level line (from _go_fatal_network), and carry stalled_for/generation in the stall error text instead. Test module docstring and test name updated to match the hand-off semantics.
Hoist the `_PollingStallError` check in `_handle_polling_network_error` to
right after the teardown/fatal guard, before the retry counter increment,
the exponential sleep, `_stop_updater_or_go_fatal` and both connection
drains. Use a plain `isinstance` (both raise sites construct the error
directly; nothing wraps it). The stall test now also asserts no sleep, no
retry-counter bump, no in-place `updater.stop()` and no drain.
WHY: the check sat after `await asyncio.sleep(delay)` (5-60 s), the
counter bump and a bounded `updater.stop()` (up to 15 s), so a gateway
already confirmed deaf stayed deaf 5-75 s longer and consumed a retry
slot for something that is not a retry.
The in-place `updater.stop()` before the handoff is dropped deliberately:
`_go_fatal_network` -> `_handoff_polling_fatal_error` -> supervisor
rebuild runs `disconnect()`, which performs the same bounded
`updater.stop()` (`_UPDATER_STOP_TIMEOUT`, falls through on timeout) plus
`app.stop()/shutdown()`, so stopping here only duplicated that work on the
slow path. Docstrings for `_handle_polling_network_error` and
`_check_polling_stall` no longer describe the stall as a reconnect-ladder
escalation.
`_check_polling_stall` hand-rolled the body of `_schedule_polling_recovery`
(set `_send_path_degraded`, `_mark_degraded()` when running, spawn
`_handle_polling_network_error`). Call the helper instead, matching the
post-reconnect verifier stall site.
WHY: one recovery entry point keeps degraded-state marking and the
"polling degraded (reason)" log line consistent across every scheduler.
The helper's early-return guards (`_teardown_started or has_fatal_error`,
`_recovery_in_flight()`) are already asserted at the top of
`_check_polling_stall` with no await in between, so they are no-ops here
and behaviour is unchanged.