`_submit_fal_request` (image) and `_submit_fal_video_request` (video plugin)
translated every managed-gateway 4xx into "This model may not yet be enabled
on the Nous Portal's FAL proxy — set FAL_KEY or pick a different model". For
HTTP 429 that remediation is wrong: the gateway body is RATE_LIMIT_EXCEEDED
with a retryAfter, the model is enabled, and agents reading the message
switched models or gave up. On one real install this fired 260 times in a
week (17% of image_generate calls).
Both surfaces now submit through one shared helper
(tools/fal_common.py::submit_managed_fal_with_rate_limit_retry): a 429 whose
Retry-After (header, else body error.retryAfter) fits a 30s cap is waited out
in interrupt-aware 0.5s slices and resubmitted once under a fresh
x-idempotency-key; a second 429, or an unknown/too-long Retry-After, raises
a ValueError that names the rate limit and tells the agent to retry later
rather than switch models. 429 can no longer reach the "may not yet be
enabled" text.
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
gateway.restart already answers "was this gateway launched by a generated
service" (HERMES_SUPERVISED_CHILD or launchd's XPC_SERVICE_NAME, so a plist
that predates the marker still counts); the hint reuses it instead of a
second env read. The negative branch (plain unreachable host without a
supervisor) is now an unmarked test so the Linux lane keeps covering the
function; the macOS-only test holds the positive branch.
Review of the first cut:
- The Home Assistant errno hint matched a bare 65 and two error strings on
every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
genuinely unreachable HA host was told macOS was blocking launchd. Gate on
darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
`gateway run`, so `hermes gateway stop` would have signalled osascript
alongside the gateway. The canonical matcher now bails on an exact
argv[0] basename of osascript (the gateway is its child and is matched on
its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
carries (osascript's own output) instead of implying it is redundant.
A launchd-run gateway that macOS Local Network Privacy denies sees every LAN
connect fail with EHOSTUNREACH while the same URL works from Terminal.
Annotate the Home Assistant connect/reconnect log lines with the cause and
the remedy so the failure is actionable instead of a bare "No route to host".
Salvaged from #115196 (remedy text points at the regenerated launchd job).
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).
_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".
(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
Review follow-up for #117434: edit_completed_task_result had no callers
after edit_task absorbed it; the dashboard's _set_priority kept its own raw
UPDATE + reprioritized INSERT, so edit_task gains a board= passthrough for
the post-commit observer and becomes the single reprioritize primitive.
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.
The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.
Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
before req_reasoning: {}
after req_reasoning: {'reasoning_effort': 'medium'}
agent.reasoning_effort: low -> {'reasoning_effort': 'low'} (unchanged)
The base media dispatch passes is_voice= to send_voice; line, mattermost and
weixin had explicit signatures without it, so a non-image MEDIA attachment
routed as audio raised TypeError and was dropped — the same class as the
Matrix report (#102221, #116776). Adds a repo-wide signature invariant test.
The base media dispatch calls send_voice(..., is_voice=is_voice) for every
audio MEDIA attachment (gateway/platforms/base.py _send_one). MatrixAdapter
.send_voice() accepted neither is_voice nor **kwargs, so every non-image
MEDIA delivery raised TypeError and the file was silently dropped — the
failure is visible in rotated logs since 2026-09-08 (never worked).
Accept the flag explicitly: is_voice=False -> plain m.audio in the original
format (no transcode); True or omitted (play_audio legacy callers) -> the
existing MSC3245 voice-bubble path with best-effort Ogg/Opus transcode.
Fixes#116776
(cherry picked from commit d4f89a725498e29a2ee0fe016f8bb98db08f358c)
The stale-bridge cleanup accepted a "node" + session-path cmdline substring
as kill evidence for legacy pidfiles (pid line only). A log tail, editor, or
grep that merely mentions the session path matches that same substring, so
the cleanup could SIGTERM a stranger process.
Require the kernel start-time fingerprint and fail closed when the pidfile
lacks one; the bridge-port scan (which verifies a node-executable listener)
reaps the orphan instead. The refusal reason in the warning now distinguishes
a fingerprint-less legacy pidfile from a recycled PID.
Flip the legacy-pidfile regression test to assert the fail-closed outcome.
Fixes#116883
The three `--hermes-diag-*` tokens were literals declared on the consuming
elements, so the theme engine's `<html>`-level custom properties could never
reach them and light presets rendered the warning badge at 1.8:1 contrast.
Chain them through the host tokens themes already set (`--color-warning`,
`--color-destructive`) with the shipped literals as fallbacks; same selector
list, so nodes rendered outside `.hermes-kanban` keep a value. Error and
critical share `--color-destructive` (critical keeps its bold weight).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
#116713 parsed HERMES_LANGFUSE_MAX_DEPTH inside `_safe_value`, i.e. once per
captured prompt, response, tool input and tool output. With an invalid value
(`abc`) every captured field logged the same "Invalid ... Falling back to 4"
WARNING for the life of the process — one multi-tool turn fills agent.log.
Resolve the depth in `_resolve_max_depth`, an lru_cache keyed on the raw env
string: the warning fires once per distinct bad value, a changed env var is
still picked up by a long-lived process (mirrors `_capture_mode`, which also
reads per call and warns once), and valid values skip the int() parse after
the first call.
Follow-up to #116713 (independent review finding).
An interim edit skipped because the chat's shared send+edit slot was busy
returned a plain SendResult(success=True), so the stream consumer recorded
the never-shown text as _last_sent_text and reset _flood_strikes. A later
turn-final flood then saw _visible_prefix() == final text and either marked
the turn delivered or entered fallback with an empty continuation — the user
never saw the tail. The adapter now flags the skip in
raw_response={"skipped": True} and _edit_existing leaves the visible prefix
and flood state untouched, so the next tick retries and a flood fallback
re-sends exactly the unseen tail.
Telegram counts an editMessageText against the same per-chat allowance as a
sendMessage, but streaming previews paced only edits (DEFAULT_STREAMING_EDIT
_INTERVAL = 0.8s = 1.25 msg/s into one chat before any reply was sent) —
83% of measured flood penalties. One shared slot per chat: a send WAITS for
its slot (skipping would drop a message), an interim edit is SKIPPED (the
next tick shows the same text anyway), and the final edit is never gated
(the answer is never withheld). A per-adapter tuning knob keeps the slot
available to tests that model instantaneous bursts.
`missed_message_backfill.channels: []` (YAML list or JSON-list string) returned
an empty set on main, and the backfill logged "no channels configured" and
skipped. The gate-CSV refactor treated an empty result as "unset" and fell
through to the allowed-union-free-response default, scanning channels the
operator had explicitly disabled. An explicit list is now authoritative even
when empty; only the default "" string still falls through to env/default.
`hermes config set KEY '["-100","-200"]'` used to write the literal as one quoted
YAML string; the writer now emits a real list (acaac9a18, #88163) and the Telegram
gate decodes the legacy string shape (122ad719, #110213). The Discord, WhatsApp and
DingTalk gate parsers still comma-split that string into `{'["-100"', '"-200"]'}`,
so a config written before the writer fix silently locks every allowlisted chat,
channel or user out — with no warning.
Route every remaining comma-split gate through the shared
`gateway/platforms/_shared.py::decode_json_list_literal`:
- Discord `_gate_csv_set` (allowed/ignored/no-thread channels, allowed users/roles),
`_discord_free_response_channels` and `_missed_message_backfill_channels` now share
the one parser instead of three hand-rolled splits.
- WhatsApp `_coerce_allow_list` (allow_from, group_allow_from, free_response_chats).
- DingTalk `_csv_set` (allowed_users, allowed_chats, free_response_chats).
Plain CSV strings, YAML lists and malformed JSON keep their previous meaning.
The cherry-picked ownership check walked every ancestor up to `/`, so a
HERMES_HOME kept inside a dotfiles checkout (`~/.git`) turned every
root-level test_* scratch file into a protected "git-owned" file and
silently disabled the plugin's core contract. Cut the walk at
HERMES_HOME for in-home paths; out-of-home (/tmp/hermes-*) trees keep
the full walk since is_safe_path already bounds them.
Tests trimmed to the two invariants: quick() drops a stale tracked entry
for a committed test inside a linked worktree (.git pointer FILE) instead
of deleting it, and root-level scratch is still deleted even with a .git
above HERMES_HOME. Dropped the contributor's literal /tmp test (the repo
never writes /tmp) and the guess_category-only case the quick() test
already drives.
guess_category() matched test_*/tmp_* by basename alone, so a committed
regression test inside a git worktree under $HERMES_HOME/worktrees/ or a
/tmp/hermes-* checkout was tracked and auto-deleted by quick() at session
end (#115295; the protected-top-level-dir half landed in #114770).
Classify such files as non-disposable whenever a .git entry (directory or
linked-worktree pointer file) exists on the directory chain. quick() and
dry_run() already re-validate stored "test" entries through
guess_category(), so stale pre-fix tracked.json entries are dropped from
tracking instead of deleted — no separate migration needed. Scratch
test_* files outside git-owned trees keep aging out as before.
Fixes#115295
`sendVideo` gives back `width=320 height=320 duration=0` with no thumbnail once an
upload is large enough that Telegram skips its own video processing, and clients
then draw the message as a square tile — for portrait reels and 16:9 clips alike,
even though the delivered file itself is correct.
Measured on one 6 s 2560x1440 clip, inspecting the Bot API response: 4.9 MB and
9.8 MB keep `2560x1440` / duration 7 / a 320x180 thumbnail; 14.8 MB, 19.4 MB and
23.0 MB degrade to the square placeholder; the same 23.0 MB file sent with
`width`/`height`/`duration` plus a JPEG `thumbnail` comes back `2560x1440` with a
320x180 thumbnail.
Probe the local file with ffprobe and attach a 320px-wide JPEG frame from it, on
both Telegram send paths: the gateway adapter's `send_video` and the standalone
`hermes send` media sender. Both helpers return nothing when ffmpeg/ffprobe is
unavailable, which keeps the previous behaviour for hosts without them.
HomeChannel normalization covers env and YAML homes, but cron
`deliver: discord:<link>` targets and thread metadata reach the adapter as
written and still died in int(). Share one helper
(gateway.config.discord_channel_id_from_link) between HomeChannel and the
resolver every outbound Discord target passes through; message links and
non-link strings keep their existing error path. Tests trimmed to two
invariants: config load (env + YAML) and the resolver entry point.
_resolve_room_identity() classifies any room with <=2 joined members as a DM
regardless of m.direct or an explicit room name, so those rooms silently
bypass MATRIX_ALLOWED_ROOMS, MATRIX_FREE_RESPONSE_ROOMS, and
MATRIX_REQUIRE_MENTION, and use DM threading instead of
MATRIX_AUTO_THREAD/MATRIX_SESSION_SCOPE. This was previously only visible in
an inline code comment, not in the env-var docs an operator would read.
Fixes#114733
The dm_topics "not a forum" warning and the matching docs section told
users to tap the bot's name in the DM and toggle "Topics" in chat
settings. That toggle only exists for group forums; a bot DM has no
such control. The actual prerequisite is Threaded Mode, enabled by the
bot owner via the BotFather Mini App (Bot Settings -> Threads
Settings) -- already documented correctly a few sections later in the
same file, under the /topic prerequisites.
Fixes#115019
attachTouchDrag() armed a drag on ANY touch pointerdown and immediately
called preventDefault(), which suppresses the synthesized click
TaskCard.handleClick relies on to call props.onOpen(). There was no
movement threshold, so a finger drifting even ~2-3px on a normal tap --
which is universal on real touch hardware -- was enough to arm the
drag and swallow the open.
Fix: defer starting the drag proxy and calling preventDefault() until
the pointer has actually moved past an 8px threshold (matches the
common native drag-affordance convention). A stationary tap never
crosses the threshold, dragging is never armed, and the click fires
normally. A real drag still claims the gesture identically to before,
just after the same few pixels of travel every touch drag implementation
already tolerates.
The bundle (plugins/kanban/dashboard/dist/index.js) has no build step --
it is hand-maintained directly, as established by prior kanban dashboard
PRs (#114882, #108694) -- so the fix is applied there.
Closes#115568.
Testing: no jsdom/vitest harness exists for this bundle (confirmed by
PR #114882's review follow-up, which explicitly rejected turning a
"live-repro jsdom harness" into a pytest because jsdom/react aren't
declared in the root package.json and the Python CI job has no
node_modules -- such a test would be vacuous in CI). Per that
precedent and the "never read source code in tests" rule (no
regex/substring pin on the bundle text), this PR instead extracts
attachTouchDrag() verbatim at test time via Node (already present:
tests-js/ + vitest are in the repo) and drives it through real
pointerdown/pointermove/pointerup sequences against a minimal DOM
stub -- a behavioral test, not a source-shape test. Proven red on the
unfixed bundle (asserts preventDefault is called on a stationary tap)
and green on the fix; skips cleanly via shutil.which("node") if Node
is unavailable in a given lane.
Verification:
- node tests/plugins/fixtures/kanban_touch_drag_probe.js against the
ORIGINAL (unfixed) bundle: fails with "FAIL: a stationary tap called
preventDefault (suppresses the click)", exit 1 -- confirms the probe
reproduces the reported bug
- Same probe against the fixed bundle: "PASS", exit 0
- scripts/run_tests.sh tests/plugins/test_kanban_dashboard_plugin.py --
42/42 passed (1 new, 41 unchanged)
- node --check plugins/kanban/dashboard/dist/index.js -- syntax OK
With `require_mention: true` + `bots_require_mention: true` + wake words in `mention_patterns`, a message
authored by another Hermes bot that addressed this bot by name matched no dispatch path (the
bot-to-bot loop breaker skips it) and was also refused by the observe gate (`mention_patterns`
matches are assumed dispatched), so it vanished from both paths with no log line.
Factor the loop-breaker predicate into `_bot_sender_suppressed` and consult it in
`_should_observe_unmentioned_group_message`, so every message the dispatcher drops because of
`bots_require_mention` is kept as observed context instead of being lost (#115119).
whatsapp.reply_prefix from config.yaml was written into the bridge env and then
popped again by the WHATSAPP_* passthrough loop (the key was in
_BRIDGE_PASSTHROUGH_ENV and the scoped env lookup came back empty), so bridge.js
always fell back to its built-in header and the documented reply_prefix: ""
could not disable it. Resolve the prefix once (scoped env first, then the
adapter value) and keep it out of the passthrough loop.
Fixes#116059
`_discord_message_admission()` drops a message that mentions someone other than
the bot when `DISCORD_IGNORE_NO_MENTION` is on (the default) and the channel is
not free-response. It did so without consulting `_in_bot_thread()`, unlike the
other two ingress paths — `_dispatch_recovered_message()` (adapter.py:2292) and
`_handle_message()` (adapter.py:5946). Admission runs on both and returns
`False` unconditionally, so it overrode the thread exemption they grant.
The result was an asymmetry with no obvious cause from the outside: in a thread
the bot had joined, a message with no mention at all was admitted (an empty
`message.mentions` skips the enclosing block), while the same message with one
mention of a third party was dropped. A thread the bot is a participant in is
the one place "addressed to someone else" is least likely to hold.
`thread_require_mention` still gates multi-bot threads, since that check lives
inside `_in_bot_thread()`.
The drop also emitted nothing at any log level, leaving `gateway.log` identical
whether the gate fired or the event never arrived; add a debug line so the two
can be told apart.
Fixes#116568
Follow-up to the salvaged #116041: the board's `done` column is now newest-
completed-first and `hermes kanban list --sort completed-desc` exists, so the
feature doc and the CLI usage block say so; the stale "per-column ordering
comes from list_tasks" comment in `get_board` now describes the queue
columns only (wording from #116051).
Co-authored-by: MohamadKanso <91088196+MohamadKanso@users.noreply.github.com>
get_board() buckets one list_tasks() fetch, so the done column
inherited the shared priority DESC, created_at ASC order — creation
order, which says nothing about when work finished. Sort the done
bucket newest-completed-first (completed_at DESC NULLS LAST, id DESC)
and expose that as a completed-desc list_tasks sort key; queue lanes
keep the FIFO dispatch default.
Production:
- agent/bedrock_adapter.py, agent/vertex_adapter.py: pm.ensure_import ran at
module import. In any process that imports these modules without a committed
PM selection (CI's build_environment test venv, a fresh checkout) that sync
rebuilt the dependency environment mid-process and replaced sys.path with a
generation missing the caller's own packages (anthropic, aiohttp vanished).
The extra is now ensured at first client build / credential request.
- plugins/platforms/matrix/adapter.py: a complete install needs no
ensure_and_bind round trip; only a partial one syncs.
- tools/browser_tool.py: drop the facade's duplicate warm_agent_browser_npx_cache
shim; the compat pointer already resolves to browser_tool_install.
Test harness:
- tests/home_io_guard.py: PATH-entry probes (shutil.which) and the running
interpreter's own installation (stdlib reads, realpath ancestry, fixture
symlinks into it) are not Hermes state; a patched Path.expanduser must not
crash the guard. run_tests.sh no longer filters PATH — the guard owns it.
- tests/tui_gateway/conftest.py: import hermes_bootstrap before any file opens
a MagicMock hermes_constants window (6 files exited the process at boot).
- tests/hermes_cli/conftest.py probe_root: scratch checkouts the import guard
probes need hermes_bootstrap.py (the launcher imports it).
- tests/pm/_fixtures.py stage_host_python: a copied relocatable python needs
its stdlib beside it (No module named 'encodings' on CI).
- tests/install/e2e-assets/smoke-env.mjs: dependency-free env shaping so the
source-build-env probe runs under bare node (main deleted the Playwright
entry it was imported through).
- adapt main's new tests to branch seams (model_metadata_http, launch
completion tail, CI toolchain exports uv after python, source_launch
hermes_cli stub, systemd_notify single marker).
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).
Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
Gate review: TelegramAdapter kept a same-named override with a weaker contract (no finite
guard, negatives clamped instead of reset), so the hierarchy had two parsers under one name.
Both Telegram callers pass explicit bounds; the base method is a strict superset for them.
The cadence test now pins the behaviour (≤ 0.5 s / ≤ 1.0 s) instead of echoing the constant.
Gate review: the fix left three adapter-private copies of `_coerce_float_extra` and the
0.3/2.0/1.0/4.0 cadence literals in three files. The parser and the cadence constants now
live on BasePlatformAdapter beside the delay attrs they configure; WhatsApp and Weixin call
`_configure_text_batch_delays()`, Telegram reads the same constants through its env helper.
The clamp test is parametrized over both adapters and the Weixin docs name the ceilings.
WhatsApp debounced text for 5s (10s near a split) and Weixin for 3s/5s
before dispatching, so every reply paid multiple seconds of idle latency
that Telegram never pays (0.3s/1.0s). Default both adapters to Telegram's
cadence and mirror its ceilings (2.0s / 4.0s, split >= base delay) via the
existing _coerce_float_extra seam. The config keys are unchanged; 0 still
dispatches immediately. Docs updated.
Spotted via #44896 (@liuhao1024). Fixes#44883, refs #25056.
_schedule_invite_join already requires both is_direct and inviter before recording m.direct, so `is_direct and bool(inviter)` at the reconcile call site was redundant; pass is_direct through. The `if is_direct and not inviter` WARNING could only be reached with GATEWAY_ALLOW_ALL_USERS set and a spec-violating stripped m.room.member event lacking `sender`; the info log already prints is_direct, so drop the branch.
The adapter comment and the test module docstring claimed a reconciled pending invite "never fires _on_invite". It does: _absorb_sync runs _dispatch_sync (which emits INVITE to _on_invite) and then the reconcile pass over rooms.invite, which joined every entry unconditionally — so a live invite _on_invite rejected was joined ms later, and invites that arrived while the gateway was down were joined on restart with no gate. Reword both to state that premise.
_on_invite only auto-joins a room when the inviter is allow-listed (or
GATEWAY_ALLOW_ALL_USERS is set), so a live invite from an arbitrary
federated user is rejected. A pending invite that arrives while the
gateway is down takes a different path: _schedule_pending_invite_joins
reconciles it from rooms.invite in the sync response and scheduled the
join unconditionally. An unauthorized invite sent during downtime was
therefore auto-joined on restart, bypassing the allowlist.
Extract the gate from _on_invite into _is_authorized_inviter and apply
it during reconciliation too, reading the inviter from the stripped
invite state (the sender of the m.room.member event for our own user,
as _extract_invite_dm_signal already does for the DM signal). An
inviter that cannot be read from the invite state fails closed, exactly
like an empty sender in _on_invite: the invite is skipped with a
warning and left pending.
A direct invite that arrives while the gateway is running fires
_on_invite, which passes is_direct and the inviter through
_schedule_invite_join so the room is recorded in m.direct after the
join. An invite that is still pending across a gateway restart takes a
different path: _schedule_pending_invite_joins reconciles it from
rooms.invite in the sync response, but called _schedule_invite_join
without is_direct or inviter. The DM signal was dropped, the room was
never recorded in m.direct, and it was classified as a group until the
user's own client happened to update m.direct.
Read the signal from the stripped invite state instead: the
m.room.member event for our own user carries the original invite's
is_direct flag, and its sender is the inviter. Thread both through to
_schedule_invite_join so a reconciled direct invite is recorded in
m.direct, and thus lands in _dm_rooms, exactly like a live one.
This gap was surfaced by the triage of #62493.
Reusing `_MEDIA_SEND_READ_TIMEOUT` (60 s) as the whole-call deadline turned
httpx's per-phase stall budget into a bandwidth cap: a 20 MB video on a
~2 Mbit/s uplink (~80 s) that succeeds today would fail. `_MEDIA_SEND_DEADLINE`
= 300 s covers the 50 MB Bot API cap at ~2 Mbit/s plus connect and sendVideo
transcoding, and is >2x the summed httpx budgets (pool 8 + connect 10 +
media_write 60 + read 60 = 138 s), so it only fires on a socket that has
stopped raising. Abandon-inside-lock semantics documented at the constant.
No Bot-API write in the adapter was under a wall-clock cap — text sends,
edits, drafts and media uploads relied solely on httpx socket timeouts,
which do not fire when a shielded httpcore socket wedges (same class as the
getUpdates hang in #92991). A stuck send then pinned `_chat_send_lock` and
the loop. Wrap every send/edit/draft in `_await_with_thread_deadline`
(`_TEXT_SEND_DEADLINE`) and both media paths in
`_send_with_dm_topic_reply_anchor_retry`; the helper gains a `label` and a
descriptive TimeoutError message.
Half A of #115280; Half B superseded by #116134 / contradicts #75017.
_fetch_discovery followed redirects but only pinned the document's
self-asserted issuer field, so one cleartext or attacker-hosted hop
could serve a forged document claiming the configured issuer with
attacker jwks_uri and token_endpoint. Verify then accepted
attacker-signed ID tokens and the code exchange POSTed the client
secret to the attacker's token endpoint.
The resolved response.url must now share the configured issuer's
origin (scheme, host, port with default-port normalisation) before
the body is parsed. Same-origin canonicalisation redirects still
pass, and the issuer-field pin remains as the misconfig check it is.