The r3 fold scoped only the dmarc verdict to its clause. SPF and DKIM
still came from a whole-string, last-match-wins regex scan, so the same
quoted/comment smuggle closed for dmarc still authenticated a spoofed
From (GHSA-rxqh-5572-8m77), e.g.
spf=fail smtp.mailfrom="x spf=pass smtp.mailfrom=example.com "@evil.test
spf=fail (spf=pass) smtp.mailfrom=a@example.com
spf=fail smtp.mailfrom=a.spf=pass@example.com
dkim=pass header.d=evil.test header.i="x header.d=example.com y"@evil.test
Every verdict now comes from the leading method=result token of its
own clause (from _ar_clauses, comments dropped), and its domains only
from that clause. Properties are read by a token scanner that consumes
quoted-strings and other key=value tokens whole, so quoted contents are
never read as properties while a quoted value still is
(header.from="example.com"). SPF fails closed on more than one spf
clause; DKIM accepts any single dkim=pass clause whose own header.d
aligns (multi-signature mail is normal), never mixing clauses. The
whole-string methods/props (methods["dmarc"] was dead) are gone.
Also: a stray ')' at depth 0 is now unbalanced (it split header.from
out of the dmarc clause); a From with more than 64 '(' takes the silent
empty-sender drop instead of a parseaddr RecursionError logged as an
error; the cap test asserts the cap directly instead of wall-clock
timing; the empty-From drop assertion moves next to the other
_extract_email_address rejects; and the untested >1-dmarc, unbalanced
and backslash-escape rules get reject strings.
Round-3 review of the #124322 salvage:
- stdlib parseaddr is pure Python and superlinear on hostile input: a
100KB `From: <a@a@...` held the GIL ~1s per message (main ~0), and any
remote sender reaches it before auth. Refuse values over _MAX_FROM_LEN
(2048; RFC 5322 lines cap at 998) right after unfolding so they take the
existing empty-sender drop. The timing assertion now times
_extract_email_address itself instead of a private regex.
- The dmarc clause split ignored quoted-strings and stripped comments one
level only, so an attacker-controlled SPF-passing envelope sender like
smtp.mailfrom="x;dmarc=pass header.from=example.com x"@evil.test planted
a fake dmarc=pass ahead of the real dmarc=fail and authenticated
admin@example.com. Split clauses with a quote- and nested-comment-aware
scanner (comments dropped, quoted values kept so header.from="x" is still
read and unquoted), and fail closed on an unbalanced value or when more
than one clause starts with dmarc=. Tests pin the smuggled clause,
reason="a;b", the nested comment, a quoted aligned header.from, and that
every header.from in the dmarc clause must align (all->any goes red).
- `Doe, John (CEO) <j@x>` was dropped because the fallback refused any
paren. Strip (comments) from the display part first; `attacker@evil.test
(c) <victim>` stays rejected by the '@' rule.
- Results without '@' (`John` -> `john`, `a\"b <v@x>` -> `a\`) were
dispatched as sender ids; return "" so they take the drop path.
The r1 malformed-From fallback regex ([^<>\s]+@[^<>\s]+) backtracked
catastrophically on `From: <a@a@a@...`: 32KB held the GIL ~23s, and any
remote sender reaches it before auth. Capture <([^<>\s]+)> instead and
check the '@' in Python (same accepted set, 100KB in ~3ms).
The fallback also mapped multi-mailbox / group / comment-prefixed From
values (`attacker@evil.test, <victim@x>`, `Grp: a@evil; <victim@x>`,
`attacker@evil.test (c) <victim@x>`) to the bracketed victim, which a
dmarc=pass without header.from then authenticated. Only fall back when the
display part has no ; : ( ) and no '@' unless it is exactly the bracketed
address; `Doe, John <j@x>` and `j@x <j@x>` still resolve. The rule now lives
only in the docstring (the old inline comment was wrong).
dmarc: strip (comments) before splitting clauses, and take the verdict and
header.from from the same first clause that starts with dmarc=, so
`dmarc=pass (p=none; sp=none) header.from=evil.test`, a later dmarc=pass
clause after dmarc=fail, and duplicate misaligned header.from are rejected,
while `arc=pass (dmarc=fail ...); dmarc=pass header.from=<ours>` passes.
Property clean-up is shared via _auth_props.
The empty-sender drop now runs right after address parsing and gets an
assertion (it was untested). Tests extend existing ones (no new tests).
Strict parseaddr returns '' for common RFC-invalid From values that the
old regex resolved: an unquoted address as display name
(john@example.com <john@example.com>), an unquoted comma (Doe, John
<j@example.com>), or a@x.test <b@y.test>. Allowlisted senders using such
clients were silently dropped, and under open access every one of them
shared an empty chat_id. Fall back to the bracketed address only for the
unambiguous shape (no quotes, exactly one <...> pair), so the quoted
display-name spoof still resolves to the attacker. _parse_fetched_message
now drops messages whose From yields no address instead of dispatching
an empty identity.
The dmarc header.from alignment check read header.from from props merged
across all clauses, so a later dkim clause's header.from could override
the dmarc clause's value. Read it from the dmarc clause only, and pin the
misaligned dmarc=pass rejection in the existing dmarc test (previously
no test covered it). Also drop a redundant _domain_of ternary, trim the
_extract_email_address docstring and a duplicate test assertion.
_verify_sender_authentication accepted any dmarc=pass verdict, even one
issued for a different domain than the From we parsed. That is what let
the quoted-display-name spoof ride on the attacker's own truthful DMARC
pass. When the trusted Authentication-Results names header.from, require
it to align with the From domain; otherwise fall through to the aligned
SPF/DKIM checks. Defense in depth for the #124322 parser fix.
_extract_email_address took the FIRST <...> pair, so
From: "Victim <victim@example.com>" <attacker@evil.test> resolved to the
victim. The attacker's own domain passes DMARC truthfully, so the
allowlist (EMAIL_ALLOWED_USERS), pairing and session identity were all
evaluated against an address the sender does not control.
Use email.utils.parseaddr, which keeps the quoted text as the display
name and returns the real addr-spec. RFC 5322 folding is unfolded first
so a folded quoted display name is not mistaken for the mailbox.
Salvages #124322.
Both reconnect paths (gateway/run_adapters.py watcher and multiplex
secondary) build a FRESH SimplexAdapter before connect(is_reconnect=True),
so the per-instance _allowlist_warned flag never suppressed anything: a
daemon-down cold boot re-logged the warning on every backoff retry. Gating
on `not is_reconnect` would instead lose the warning when the first connect
fails. Dedup at module level keyed on (hermes_home_key(), frozenset(names)),
still checked before the connectivity probe. No shared warn-once helper
exists (plugin_compat.warn_once is compat-specific).
Read the value with platform_gate_env (the reader authz uses; differs from
get_scoped_secret when a scope is installed with multiplex off) and decode
JSON list literals with decode_json_list_literal like _coerce_allow_set, so
'["4","9"]' written by `hermes config set` no longer warns that valid IDs
are ignored.
The caplog test now builds two fresh adapters (first connect fails, second
succeeds) and asserts exactly one warning naming only 'alice'; it fails with
2 warnings against the pre-fold adapter.
The name-entry warning in connect() read SIMPLEX_ALLOWED_USERS via raw
os.getenv, while authz reads it profile-scoped. Under multiplexing a
secondary profile would warn about (or stay silent on) the default
profile's list rather than the one actually enforced. Use the module's
_get_scoped_secret + _parse_comma_list like __init__ does.
It also only fired on a successful non-reconnect connect: if the daemon
was down at cold boot the first connect() failed and every retry came in
with is_reconnect=True, so the warning never appeared. Evaluate it before
the connectivity probe, once per adapter via an instance flag.
Test: two connects (first fails) -> exactly one warning naming only
'alice' for scoped '4, alice' while os.environ holds 'bob'. Red on the
pre-fold adapter (0 warnings) and on a raw-os.getenv variant (names bob).
After #44729 SIMPLEX_ALLOWED_USERS matches only the numeric contactId, but
the docs still told operators display names work, and existing name
entries would silently stop matching. Update the docs and log a one-time
warning at first connect listing non-numeric entries that are now ignored.
The SimpleX sender allowlist (SIMPLEX_ALLOWED_USERS) previously matched
against both the stable numeric contactId (user_id) and the mutable
display name (user_name). Since any SimpleX contact can change their
localDisplayName / profile.displayName to match another user's, this
allowed an unauthorized contact to bypass the allowlist by setting a
colliding display name.
Remove the user_name check so that SIMPLEX_ALLOWED_USERS only matches
on the immutable contactId. Operators must use numeric contact IDs in
the allowlist.
Fixes#44729
(cherry picked from commit b4aa29da1567d45920f79aabdb36b44c5f87bde5)
The r3 fold made park refuse an older voice once a newer one from the same
sender had parked. With one /sync batch holding [v1 (slow gate), m1 bare
mention, v2 (fast)], v2 parked first, v1 was then refused, and m1 claimed
v2 -- a voice sent after the mention. v1 was lost and v2's own mention was
answered as empty text.
The bare mention now takes an arrival limit (ParkedVoices.mark) before it
settles, and claim pops the newest parked voice that began before that
limit, dropping older ones. park no longer refuses by arrival; parked
voices are kept per sender ordered by seq (bounded to 4). A claim made
while gates are still in flight records a floor, so a late older voice
(seq <= claimed) cannot park and outlive the claim; the floor clears when
the sender's in-flight list empties. One answer per bare mention, newest
before the mention wins, and the r3 orphan case still leaves nothing
parked.
Also: pending() ignores entries past CLAIM_WINDOW_SECONDS so an expired
voice no longer sends every later text through the mention regexes; the
m.thread root lookup is one _thread_root helper used by both the park
decision and its pre-check; voice_gate is annotated.
The kept test gains a [voice slow, mention, voice2 fast] + mention2 row:
red on a932dc031c (['$voice2', '$text2']) and on f1ea71de7c
(['$voice', '$text2']).
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
The same-sync-batch in-flight mark was one slot per (room, sender), so a
second voice's begin overwrote the first. When the older voice finished
gating after the newer one, the bare mention settled on (and claimed) the
newer voice only, and the older one parked afterwards as an orphan: the
sender's next unrelated bare mention within 120s re-dispatched it, so a
stale voice got downloaded, transcribed and answered.
The in-flight mark is now a list of gates per key. release removes only
its own gate (idempotent) and drops the key once it is empty; settle
waits on all current gates under the same 5s cap. Each gate carries an
arrival sequence and park refuses an older voice once a newer voice from
the same sender has parked, so a late older voice can neither replace
nor outlive the claim of the newer one (still one parked voice per
sender, newest wins).
The mark was also held across the whole _resolve_message_context,
including the display-name fetch and thread mark, and taken for voices
that can never park (the voice itself @mentions the bot, free rooms, bot
threads, non-allowlisted rooms). A bare mention racing such a voice
waited for the voice's display-name fetch (0.61s vs 0.31s on main with a
0.3s fetch). begin now runs only when a synchronous pre-check says the
voice can park, and the mark is released as soon as the park decision
is made (before the display-name fetch; DM voices release there too),
with the finally still covering exceptions and cancellation.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
mautrix dispatches every event of one /sync batch as its own task
(wait_sync=True). The voice only parks after awaiting the room identity,
which is a homeserver round-trip whenever the 60s identity cache is stale,
i.e. in any idle room. The bare-mention text checked the park without
awaiting, found nothing, dispatched an empty text, and the voice then
parked and expired unclaimed.
A parkable voice is now marked in-flight per (room, sender) right before
its first await and released in a finally once it is gated. A bare
mention that sees an in-flight voice waits for it (bounded, 5s) before
claiming; ordinary text only pays two dict lookups and never waits.
After a claim the bare-mention event also gets its read receipt, so the
read marker is not left one event short of the pre-claim behaviour.
Cleanups: ParkedVoices drops its unused window param, stores
(parked_at, voice) instead of a flat tuple, the voice_mention import
moves to the top import block, and the redundant require_mention check
on the claim path is dropped (parking/in-flight only happen under it).
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
Review cleanups on the parked-voice fix:
- The claim bypass is now an explicit mention_claimed parameter threaded
_handle_media_message -> _resolve_message_context. The old _claimed set was
only drained inside the require_mention branch, so a voice whose thread became
a bot thread while parked leaked its id for the adapter's lifetime.
- One MSC3245 predicate (has_voice_marker, `.get(...) is not None`) shared by
is_voice_event and _classify_inbound_media; the two used to disagree on a
null marker (parked as voice, classified as AUDIO).
- One mention helper (_content_mentions_bot) for both the gate and the
bare-mention claim, instead of two copies of the m.mentions extraction.
- ParkedVoices.has() dict lookup gates the claim, so ordinary text messages
no longer pay _strip_mention/regex work when nothing is parked.
- Test: a same-room mention WITH text stays a text (gives the bare-only guard
teeth) and the parked voice is asserted never downloaded.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
Element X ships an MSC3245 voice event with an empty m.mentions block and
sends the mention the user typed while recording as a separate m.text
event right after. With require_mention on, the gate dropped the voice and
the follow-up bare @bot was dispatched as an empty text, so the bot never
answered the voice.
Park an unmentioned voice keyed by (room_id, sender) instead of forgetting
it; a bare mention from the same sender in the SAME room within 120s claims
it and the voice is processed in place of the empty text. The claimed event
id passes the mention gate exactly once, so it is not re-parked; expired
entries are pruned on every access. Parked voices are never downloaded or
transcribed (no wake-word/STT of unmentioned audio), and a mention in
another room never pulls a voice across rooms.
The state lives in a small sibling module (voice_mention.py) because
adapter.py is already far past the file-size budget.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
RunClock ticked from tasks.started_at (the task's first-ever start), so after
a review timeout + retry a healthy current run showed cumulative card age
(e.g. working · 2h for a minutes-old run).
_task_dict() now emits current_run_started_at from the run row
tasks.current_run_id points at (batched single query, same as latest
summaries), and RunClock prefers it, falling back to started_at for older
backends. Fixes#99819.
_git_tracks hand-rolled `git ls-files` without windows_hide_flags, an
isolated git env or stdin=DEVNULL. It runs from the synchronous
post_tool_call hook, so on a windowless Windows host each test_*/tmp_*
candidate could flash a console (#54220/#56747 class), and an inherited
GIT_DIR/GIT_WORK_TREE would point it at the wrong repo. Reuse
hermes_cli.source_check._git_ok, which already does all three and returns
False on any failure. The timeout drops to 5s (ls-files needs no more).
git reads the argument as a pathspec, so an untracked `test_[1].py` or
`tmp_*` glob-matched a tracked sibling and was never cleaned. Prefix it
with `:(literal)`.
Only spawn git when a .git exists at or above HERMES_HOME (HERMES_HOME is
a checkout, or sits inside a dotfiles repo). Without one, no repo can
track the file, so a stock install now does a few stats and skips the
process spawn on every qualifying tool call.
Drop the docstring paragraph that repeated _git_tracks' rationale.
The tracked-file guard only ran when HERMES_HOME/.git existed, so a
HERMES_HOME nested in an enclosing repo (a ~/.git dotfiles repo tracking
~/.hermes/scripts/test_x.py) still had that committed file deleted by
quick() - the same bug class the guard was added for.
guess_category only reaches this check for test_*/tmp_* names, so a
per-candidate 'git -C <parent> ls-files --error-unmatch -- <name>' is
cheap. It finds whichever repo encloses the file, and it replaces the two
lru_caches plus the index-stat signature: with no cache there is nothing
that can outlive the index, and ls-files output no longer needs decoding.
An exact tracked check stays safe above HERMES_HOME, unlike the bare .git
probe that was dropped for being too broad.
Co-authored-by: David Crandall <david@convergentdesign.dev>
_git_tracks wrapped the lookup in a bare except Exception. The only real
failure it hid was UnicodeDecodeError from text=True on non-UTF-8 paths in
the ls-files output, which escapes the (OSError, SubprocessError) catch.
Decode with surrogateescape (matching how Python decodes filesystem paths)
and drop the catch-all so real bugs surface.
Co-authored-by: David Crandall <david@convergentdesign.dev>
`_git_tracked_index()` cached one `git ls-files` set per process, so a
long-lived process could keep classifying a file as untracked after git had
taken ownership of it. `guess_category()` runs on every post-tool-call, so the
gateway can cache the index while `test_scratch.py` is still untracked; an
auto-snapshot then stages and commits it, `quick()` re-validates the stored
"test" entry through `guess_category()`, the stale cache still reports the file
as untracked, and `_delete_item()` unlinks a file git now owns (`git status`
shows `D test_scratch.py`).
Key the cached set on the git index's stat signature as well as the home, so a
staged, committed or unstaged change is a new cache key and no invalidation
hook is needed anywhere. The index path comes from `git rev-parse --git-path
index`, cached per home — so the per-call cost stays a single `os.stat`, which
also covers a linked worktree (where `.git` is a pointer file) and an explicit
`GIT_INDEX_FILE`. Fail-open behaviour is unchanged: no index and no git mean the
empty set, and the guard never raises.
Regression test: classify an untracked `test_scratch.py`, then `git add` and
`git commit` it in the same process with no `cache_clear()`, and assert
`quick()` leaves it on disk while an untracked control beside it is still
cleaned.
(cherry picked from commit a3271750608ca4e02e5205311a09b906e16d33d5)
The `_inside_git_worktree()` guard sliced its parent chain at HERMES_HOME, so
`.git` entries strictly below the home were the only ones consulted. That misses
the case where HERMES_HOME IS a git checkout: on this install `~/.hermes` is the
userfiles repo, so `~/.hermes/scripts/` and any top-level `tmp_*` file sit inside
a worktree yet resolve to no `.git` below the home.
Observed live: the bundled disk-cleanup plugin classified two COMMITTED
regression tests in `~/.hermes/scripts/` as disposable session scratch and
unlinked them — `scripts/test_analyze_upstream_opportunities.py` and
`scripts/test_customization_protocol_v2.py`; the deletion was then committed by
the next auto-snapshot (ce8dc99), so the suite that would have caught a scanner
regression was silently gone. 13 tracked `test_*` files under `scripts/` plus 17
tracked top-level `tmp_*` files were in the same blast radius.
Fix: when HERMES_HOME itself carries `.git`, ask git whether it TRACKS the exact
path (`git ls-files`, one cached subprocess per process keyed on the home) rather
than treating the whole home as protected — a scratch file merely living beside
tracked ones must still be cleanable, or cleanup would be disabled entirely.
Tests: two regression tests — a git-tracked `test_*` file is never classified
disposable and a stale pre-fix tracked.json entry is dropped by quick()'s
re-validation instead of deleted; plus the control that an UNTRACKED `test_*`
file in the same repo is still cleaned. Existing suite 30 -> 32 passed.
(cherry picked from commit c5555940108286c4c1070e84d6d694ad9a35bb76)
(cherry picked from commit c7c39e60d4dc467bf9681c3dcb5781ac96557724)
The model-picker fix added a second copy of the choice picker's auth
check, and each copy rebuilt the callback ctx that _handle_callback_query
had already computed as `cb`. The chat-id picker loop is the only route
into both handlers. Gate there once with the existing cb, so a picker
added to that loop later is covered without anyone having to copy the
check again. Fail-closed behaviour is unchanged: _callback_authorized
returns True only for an authorized tapper and answers the denial text
otherwise. The back-button test calls the handler directly and no longer
needs its auth stub.
Co-authored-by: Dusk1e <yusufalweshdemir@gmail.com>
Every other ModelPickerView callback (provider/model select, expensive
confirm, back) runs through the shared _HermesView._gate auth check, but
_on_cancel did not. A non-allowlisted member could tap Cancel on the
owner's picker, marking it resolved and clearing the view. No model
switch was possible, so impact is low, but the gate should be uniform.
Reuse the same _gate call the sibling handlers use.
Model picker taps (mp:/mpg:/mpv:/mm:/mc:/mb/mx/mg:) never went through the
callback allowlist, so a group member who is not allowlisted could tap the
owner's /model picker and switch the owner's session or config model. Gate
the handler on _callback_authorized like the choice-picker, approval and
clarify buttons, before any picker state is read.
Ported from #33859 (gateway/platforms/telegram.py) to the plugin adapter.
(cherry picked from commit b002be8feacd09fde0ee9752de7b396d0aac69ba)
A substring check on the raw base URL also matches proxies such as
https://proxy.test/api.x.ai/v1, so is_slot_keyed_cache_route now compares
utils.base_url_hostname() exactly, like chat_completion_helpers and
codex_responses_adapter already do.
The aggregator Grok prefixes were spelled out in both prompt_cache_scope and
the OpenRouter profile; one GROK_AGGREGATOR_MODEL_PREFIXES constant keeps the
fork-scope check and the x-grok-conv-id header on the same model set.
#109964 made same-model cache-parity forks (background review, /btw) inherit
the parent's resolved cache scope. That is the right call for
content-addressed caches (Anthropic, DeepSeek, Gemini) and for OpenAI's
routing-only prompt_cache_key. xAI is different: x-grok-conv-id /
prompt_cache_key select ONE server-side conversation slot, so a fork of
similar size that diverges early evicts the parent's slot and the parent's
next call reads cold.
Measured on grok-4.7 (xai-oauth Responses), ~160k context, fork of similar
size diverging right after the system prompt, parent's next call after the
fork:
shared scope (current): 1,152 / 162,239 read (2/2 runs)
derived scope (this PR): 162,176 / 162,239 read (2/2 runs)
no fork (control): 177,536 / 177,639 read (2/2 runs)
build_cache_parity_fork tags the fork (_prompt_cache_fork_tag). The
resolver derives "<scope>::<tag>" for tagged agents on slot-keyed routes
only (xai, xai-oauth, api.x.ai, OpenRouter x-ai/grok-*). The inherited
scope is kept everywhere else. The tag is applied outside the memo, so a
mid-run provider fallback re-evaluates. OpenRouter's Grok x-grok-conv-id
now honours a fork scope over the ambient affinity/conversation scope
(chat_completions threads cache_scope_id to the profile). session_id and
transcript identity are unchanged.
(cherry picked from commit 40fbb39d978648c76940cfb7835f77a42d3719b0)
- salvage summary cap and _bound_oversized_record still composed bare
truncation idioms in model-visible text; route them through elide /
elide_middle so a copied marker is guard-visible.
- the active-task line repr()'d the elided text, escaping the marker's
apostrophe when the user text held both quote kinds and hiding it from
the guard; elide after repr instead (text within the cap stays whole).
- a leftover budget smaller than the marker produced a content-free,
over-budget marker line in _build_verbatim_user_section and the Slack
nested-attachment path; skip the item instead.
- _build_verbatim_user_section elided twice, reporting the wrong total;
one elide at min(cap, remaining).
- drop redundant len() pre-checks before elide() and name verification
stop's repeated 1200.
Co-authored-by: ahisblessed <ahisblessed@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
The Slack Block Kit payload dump and the nested-attachment text budget
are both fed to the agent, and trajectory_compressor's summarizer input
becomes training data; all three still used the imitable bare
"... [truncated]" idiom. Route them through elide()/elide_middle() with
module-level imports and extend the no-idiom invariant to scan
plugins/platforms/slack and trajectory_compressor.py.
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
Follow-up to the #103857 salvage. The Copilot provider profile and the
main-agent GitHub reasoning path each carried the same clamp-then-medium
fallback. Both now call hermes_cli.models.clamp_github_reasoning_effort.
The offline Astra tests move next to the other Copilot effort tests, along
with a check that a structured catalog entry still wins. The main-agent
clamp now has a test pinning max/ultra -> high on the GPT-5 ladder. The
transport test no longer writes config.yaml. The two effort comments now
say the same thing.
With reply_in_thread: false the whole channel is one session and the bot
answers top-level, so an unmentioned top-level message there is a
follow-up in a conversation the bot is part of, like a thread reply.
reply_expected is now False for a free-channel message only when it starts
its own session (a new top-level thread), else None. The bot-id set is
built inside _slack_reply_expected, as _channel_gate_allows does.
The Slack rule marked every admitted message that was not a DM, a mention
or a command as not addressed, so a plain "done?" in a thread the bot is part
of, or a reaction trigger, could end on a bare silence marker and vanish,
the case #111624 fixed (#110952).
reply_expected is now False only for a message that opens by @mentioning
someone else, or a top-level message a free-response channel admitted
without a mention. Reaction triggers and pipe-form self mentions count as
addressed; other thread replies are None (visible fallback). The
free-channel predicate moves into _slack_is_free_channel so the gate and
the rule read the same one. The test drives the real _handle_slack_message.
Docs describe the rule in its own note instead of the
ignore_other_user_mentions tip, and the messaging index documents the
human-turn fallback.
Since 5ea8fb2b78 (#111624, for #110952) the gateway rejects a bare silence
marker on any human turn and delivers "The model returned only a silence
marker for a message that needed a reply" instead. That protects a human
who asked this bot something and got nothing back. It also fires on every
human message the adapter admitted without the bot being addressed at all:
a free-response channel, a thread follow-up under
`thread_require_mention: false`, or a message @-mentioning another person
or bot with `ignore_other_user_mentions: false`. A bot whose SOUL declines
peer-addressed turns with a deliberate marker now posts that notice on
every such message. A fleet running several bots in shared Slack threads
reported it as spam on v2026.9.21. #37940 established that intentional
silence must not be re-inflated. Both contracts hold once the turn knows
whether a reply was expected.
`MessageEvent.reply_expected` (True, False, None) is set by the adapter
where the message is admitted. Slack (`slack_reply_expected`): a 1:1 DM,
an @mention of this bot or a command is True, anything else it admits is
False. Other adapters leave None, which keeps today's behaviour, so nothing
changes for them until they are ported. `response_filters.silence_allowed`
holds the one rule (machinery turn, or reply not expected) and both call
sites use it: the live turn in `run_turn._hmwa_shape_agent_response` and
the crash-recovery redelivery from #120377 (1136f135dd), which reads the
flag back from the persisted turn metadata. The suppressed case logs one
DEBUG line naming platform and chat.
Operator workaround until this lands: `platforms.slack.extra.
ignore_other_user_mentions: true` drops peer-addressed messages before a
turn exists.
(cherry picked from commit 094439776ab898cccde303a1c2c911c8ab5bfb75)
Four surfaces build a throwaway AIAgent and never call close() — the
owner boundary that releases memory-provider sessions, tool
subprocesses and httpx clients. In long-lived processes each run leaked
all of them until exit:
- batch_runner._process_single_prompt: one agent per prompt, N prompts
per batch process.
- feishu_comment._run_comment_agent: one agent per comment run in the
gateway process.
- tui_gateway prompt.background: one side agent per background turn.
- cli /bg: one agent per background task in the CLI process.
Wrap each run in try/finally with a suppressed close(), mirroring
gateway/run.py's owner pattern. preview.restart stays deliberately
unclosed (its task exists to leave a detached server running), and the
prompt.background side agent is safe to close: its session_id is the bg
task id, so close() reaps only its own task resources.
Fixes#50197
Desktop opened /events with no since, and a missing cursor was read as
0, so every open replayed task_events history. Seed the socket from the
snapshot or this connection's last frame, and start a cursorless stream
at MAX(id). An explicit since still replays from there.
Fixes#81537
Gate round-1 follow-ups on the #121486 fix:
- auxiliary_client: inline the pool route lookup (no dead try/except or
fallbacks; HERMES_CODEX_BASE_URL short-circuits once) and read auth.json
directly when the pool yields no token (no second uncached pool load,
no re-select race pairing a new pool key with chatgpt.com).
- image plugin: _read_codex_credential() is the single source for both
is_available() and generate(); _post_image_request requires base_url.
- auth_codex: drop the unused _pool_codex_access_token wrapper; the route
helper's error fallback reads the profile-scoped override, not the raw
process env.
- model setup flow: the confirm guards get the resolved Codex base, not
the chatgpt.com constant.
- cli_model_switch_mixin: self.base_url is always set.
Follow-up to the two contributor commits for #121486. The picker, the
image plugin and the auxiliary Codex client still composed a pooled
gateway key with a base re-read from ambient state (HERMES_CODEX_BASE_URL
or the chatgpt.com default), so a model.base_url-only gateway (env unset)
still sent its key to chatgpt.com.
- auth_codex: resolve_codex_runtime_credentials reports the host a pooled
credential actually routes to (runtime_provider._pool_entry_mode_and_url:
env > model.base_url while the row is canonical > row URL) instead of the
ambient default; get_codex_auth_status carries the same bound base_url.
- picker: get_codex_model_ids(access_token, base_url=) now receives the base
resolved with the token from hermes_cli/models.py, the CLI default-model
swap (self.base_url) and the `hermes model` Codex flow.
- aux/image: _resolve_codex_credential_and_base() returns (token, base) from
one pool selection; the image plugin, _build_codex_client and the raw
Codex client use it (profile-scoped override from #121497 still wins).
- model_metadata: the non-JWT refusal now applies only when the target is
chatgpt.com; a gateway key may probe its own gateway's /models.
Adversarial regressions: model.base_url with env unset, env/route mismatch,
opaque + JWT gateway keys, pool-selected credential, pool row with its own
gateway URL, direct-ChatGPT positive control.
Addresses @andrexibiza's review on #121508.
- A mid-turn notify reply (/status, /approve, clarify answer) shares the
stream's thread key; it no longer seals and overwrites the half-streamed
answer. In-place replacement now requires the final to match the stream
after normalizing mrkdwn markers and whitespace; anything else posts fresh
and leaves the stream open.
- One _commit_stream helper for both seal-then-commit paths, so the rewrite
path also falls back to chat.update when stopStream fails.
- A stream reopened after a server-side seal is seeded with only the text
past the sealed message (tracked as 'base'), not the whole segment.
- Streams older than 15 min are sealed and dropped on the next start.
- _stream_key reuses _workspace_thread_key/scope_id_for_chat; the stream
dict no longer duplicates chat/team ids.
5648f81431 fixed this exact server-side seal (Slack closes a native stream
after a few minutes of a long turn, live-observed at ~5m20s; the lifetime
is not documented) for the native task-card stream: on
message_not_in_streaming_state from appendStream, drop the dead ts and
start a fresh stream, seeded with the full current content so nothing is
lost.
send_draft — the plain-text native streaming path used when task cards
are not enabled — hits the identical seal but never got the fix: its
generic except block only recognizes the feature-gate markers
(not_allowed, missing_scope, ...) and otherwise just logs debug and
returns failure. gateway/stream_consumer_transport.py's
_send_draft_frame() docstring is explicit that "any failure permanently
disables drafts for this run" — so a long turn streaming as plain text
degrades to the edit-based fallback for its remainder exactly the way
the task-card bug did before 5648f81431.
Mirror the task-card fix: on message_not_in_streaming_state from
chat.appendStream, drop the dead ts and _start_stream() a fresh one
seeded with the full accumulated text (not just the delta), so the next
frame's delta still resumes correctly. One reopen per frame; a second
rejection propagates as a real failure, matching the twin's behavior.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 2b4ff4e23bdf684d2fde1a9512f11dcd7ef99c42)
A mrkdwn-rewritten turn-final (e.g. *Done:* -> _Done:_) no longer continues the
streamed text, so it was classified unrelated: the stale stream was sealed and
send() posted a second message (#95430 cause B). Seal, then chat.update the
sealed ts with the final; post fresh only if the in-place update fails.
Co-authored-by: liguoyu <guoyu.li@lcfuturecenter.com>
Problem: with native streaming (chat.startStream/appendStream/stopStream)
the same answer could land twice in a thread — once as the streamed
message, once as a fresh chat.postMessage — while the streamed message
kept its live-typing indicator.
Mechanism: `_try_finalize_stream` matched the turn-final against the
streamed text with a raw `startswith`. The agent strips `final_response`
and joins footers with `rstrip()`, so any surrounding-whitespace
difference made the finalize fall through to a plain post although the
open stream already showed the whole answer (and was never sealed). A
`chat.stopStream` failure took the same fresh-post path even when the
streamed text equalled the final. Streams were also keyed per `chat_id`
only, so two concurrent turns in two threads of one channel sealed or
overwrote each other's stream.
Fix:
- Key native streams per `(team_id, chat_id, thread_ts)`; the stream
consumer stamps the same `thread_id` on every draft frame and on the
turn-final `send()`, so both resolve to the same key.
- Honor the streaming contract (gateway/AGENTS.md): sends carrying
`_interim_send` or `expect_edits` never seal a stream.
- Classify the final against the streamed text as equal / extends /
unrelated with edge-whitespace tolerance (`_stream_relation`). The
stopStream delta is sliced from the RAW final, so nothing inside the
answer (blank lines, fences, tables) is dropped or repeated.
- Commit rule: one `chat.stopStream`, one retry only when no tail is
appended (`markdown_text` APPENDS, so an ambiguous failure must not
repeat it), then `edit_message(finalize=True)` on the stream ts as the
idempotent in-place commit — it already owns format/truncate/Block Kit
and the block-rejection retry. Only when both fail does `send()` post
a fresh message (a duplicate beats a lost answer).
- Oversized tails and rewritten finals (`notify=True`) seal the stale
stream on what is visible before falling back, so no stream is left
with a live-typing indicator.
- `_seal_stream` takes the exact unsent delta instead of recomputing it
from `final_text`; `disconnect()` and the stream API calls route
through the stream's own team client.
Tests: tests/gateway/test_slack_native_streaming.py covers the
whitespace-only difference, the stopStream-failure commit path, the
bounded retry, the uncommittable fallback, interim/preview sends,
per-thread keying, oversized tails, rewritten finals and the
GatewayStreamConsumer end-to-end path.
(cherry picked from commit a64d10071ce7816b124467e29407d7d47bfdde8d)
Use Perplexity for Nous-managed search, retaining Firecrawl for extract
and as a per-call search fallback. Explicit search overrides and direct
keys keep their own billing paths; fallback results are never cached.
When the managed route is selected but the Tool Gateway is unavailable
(unentitled account or no Nous token), search reports that selection
error instead of asking for a direct key the user never chose.
The managed search vendor is unannounced, so user-facing copy names the
capability rather than the vendor: status, portal and docs say "managed
web search", and the fallback annotation reads `managed_primary`. Direct-key
configuration docs are unchanged.
Routing, auth, payload, cache and entitlement regressions are covered
through real config loading and local HTTP.
dingtalk-stream 0.24.3 start() retries forever inside the SDK, logging a
malformed logger.exception() every 3s, so the adapter breaker never saw the
error. Install a dedup filter on the SDK logger that can't raise on bad
format args, detect the websockets incompatibility (bare or chained
TypeError), log one ERROR with a pin-consistent hint, and hand off via
_set_fatal_error(retryable=False) + _notify_fatal_error(): only a
reinstall of the pinned versions and a restart fixes it. Breaker stays
tripped until the error type changes; constants moved to module level.
The reconnect storm in #24851 is driven by a dingtalk-stream/websockets
incompatibility that raises TypeError on every start(); backoff never
recovers it. Log the first occurrence per error run at ERROR with an
upgrade hint instead of a generic WARNING.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
DingTalk stream-mode reconnection storms the gateway: when start() raises
the same error every cycle, _run_stream logs a WARNING and reconnects at the
60s cap forever, generating hundreds of MB of identical log lines and hanging
the gateway.
Add a per-error-type circuit breaker: after 5 consecutive identical errors,
suppress the repeated WARNING (one ERROR summarises) and pause 300s instead of
spinning at 60s. Reset backoff + counters after a clean start() so a recovered
connection is treated fresh.
Salvage of #24881 re-implemented on current main: the original fix targeted
gateway/platforms/dingtalk.py, which has since been refactored to
plugins/platforms/dingtalk/adapter.py. Same logic, new path; tests import the
new module.
Closes#24851.
(cherry picked from commit 84201c58255ae3d3c9b269ac777ea0ff444d9229)