The #127047 comment said a supervisor's stop signal reaches the bridge
through the shared process group. The bridge is spawned with
start_new_session=True; what reaches it is a container-wide broadcast
such as s6-overlay's stage-3 SIGTERM.
Same race as the WhatsApp bridge (#127047): a container/supervisor stop
signals the whole process tree, so the Photon sidecar (its own session)
can die before the gateway's stop flow reaches disconnect(), while
_inbound_running is still True. _supervise_sidecar() then reported
SIDECAR_CRASHED and queued a reconnect during shutdown.
Consult the runner's _stop_requested_by_signal, which the signal handler
sets before any stop work runs. A sidecar exit while the gateway keeps
running is unchanged: still a retryable fatal error.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
A supervisor's stop signal (docker stop, s6-supervise, systemd) reaches the
bridge child in the same process group before the gateway's stop flow reaches
WhatsAppAdapter.disconnect(), so _shutting_down is still False when the poll
loop observes the -15 exit. _check_managed_bridge_exit() then reports a
retryable fatal adapter error, the stranded-platform check exits the gateway
with failure, and the supervisor restarts it into a crash loop (#127047).
Also consult the runner's _stop_requested_by_signal, which the signal handler
flips before any stop work runs. A bridge SIGTERM/SIGINT while the gateway
keeps running is unchanged: still a retryable fatal error so the reconnect
watcher revives it.
(cherry picked from commit 252ae311f06e402f9a8ef1482c87aaefe8d9ac39)
* feat(i18n): layered catalogs — plugin packs and user overlay over bundled locales
* feat(tui): i18n layer — en catalog, nanostore runtime, RPC pack loader, _keys.tui.json emitter
ui-tui/src/i18n/: en.ts (facade over topical siblings under en/), types.ts
(Translations + dotted TranslationKey derived from en), runtime.ts ($locale/
$catalog atoms, translateFrom active→en→key, pack merge with string→fn
wrapping for {0}/{1} placeholders), loader.ts (display.language →
i18n.catalog {lang, surface:'tui'}, English when the method is missing),
useT()/useLocale() hooks, t() for non-React code. useConfigSync feeds the
loader from the existing config.get full hydration. `npm run i18n:keys`
writes locales/_keys.tui.json (sorted flat key list) and runs before build.
* feat(plugins): provides_locales manifest field, ctx.register_locale/register_locale_dir, manifest-only language packs
* chore(tui): split en catalog siblings by lane (slash sibling)
* feat(plugins): validate language packs — parse, text-only, key-subset WARN against en / _keys exports
* feat(tui_gateway): i18n.languages / i18n.catalog RPC + regenerated contracts
* feat(config): display.language accepts any supported_languages() id, refuses unknown ids with the list
* docs(i18n): language packs user guide, pluggable display.language, plugin developer section, AGENTS notes
* feat(plugins): report language-pack layers in the mid-run activation summary
* feat(tui): wire status bar, composer placeholders, hotkey help and approval/clarify/confirm prompts through i18n
StatusRule maps compared state values (ready/running…/summoning) to catalog
text at render via displayStatus(); hotkeys()/placeholder() resolve lazily so
a pack that arrives after boot applies. Catalog grows to 81 keys.
* feat(desktop): pluggable app locales — registry, host.i18n.registerAppLocale, backend packs, keys emitter
- Locale widens to string (BundledLocale keeps the union); TRANSLATIONS stays
the bundled record and every consumer resolves through the registry.
- src/i18n/registry.ts: registerAppLocale(id, {endonym, rtl, translations})
layers partial packs (nested or flat dotted) over bundled/en via
mergeTranslations; a string over a function-valued en entry becomes a
positional {0}/{1} formatter; $appLocaleVersion bumps so translators
re-render; per-source disposers + replaceAppLocaleSource for atomic swaps.
- Backend packs: i18n.languages + i18n.catalog {surface:'desktop'} feed the
registry as source 'backend' (method-not-found is silent); re-synced on
socket open, display.language change and profile switch. A saved pack-only
language is promoted once its pack registers.
- SDK: host.i18n.registerAppLocale / languageOptions; ctx.i18n.registerAppLocale
tracked for unload. Docs in the desktop plugin SDK guide + skill reference.
- Language switcher lists bundled ∪ registered ∪ backend, endonym-only; RTL
from the registry (applyDocumentLocale takes rtl).
- npm run i18n:keys emits locales/_keys.desktop.json (wired into build).
* i18n(cli): route /topup + /subscription copy through t() (cli.billing.*, cli.subscription.*)
Module-level copy tables and modal choice tuples in cli_billing_mixin.py froze
English at import, before display.language was known. They are now key tables /
builder functions evaluated at call time; every user-facing line in the /usage
balance block, /subscription and the five /topup screens reads the catalog.
Choice VALUES stay English identifiers. Fragment-assembled status lines
(Plan: … → cancels · $x left · renews …) become full templates.
* i18n(gateway): exec-approval card contract + base/run/run_busy/run_inbound replies through t()
- base_exec_approval: EA_* English constants stay; add ea_header_text()/ea_reason_label_text()/
ea_smart_deny_line_text()/ea_default_reason_text()/ea_action_labels()/approval_timed_out_notice()
accessors; deadline + timed-out notice resolve via gateway.exec_approval.*
- BasePlatformAdapter._EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/_EA_ACTION_LABELS become
properties (adapters still shadow them with markup class attrs)
- run.py: provider error replies table holds catalog keys; _CONTEXT_OVERFLOW_REPLY -> _context_overflow_reply()
- run_busy/run_inbound: typed approval + slash-confirm matchers accept English ∪ approval.inputs.* (t())
- locales/en.yaml: gateway.exec_approval/busy/errors/... namespaces
* i18n(cli): wire modal, loops, agent-setup mixins through t() (cli.* keys)
* i18n(platforms): route Slack, Matrix and Feishu user-facing text through t()
Exec-approval markup overrides (_EA_HEADER/_EA_REASON_LABEL/_EA_SMART_DENY_LINE/
_EA_ACTION_LABELS) become per-call properties over the shared
gateway.exec_approval.* contract keys, so Slack's 3000-char section budget
measures the resolved template. Slack _APPROVAL_DECISIONS/_CONFIRM_DECISIONS,
Feishu _APPROVAL_LABEL_MAP and Matrix _EA_LEGEND/_EA_TYPED_HINT turn into
key tables resolved at click time; the Matrix typed hints become whole
sentences per offered tier instead of spliced fragments. Slack button labels
are cut to 75 chars and select placeholders to 150 after translation; the
model-facing clarify fallback answer ('choice N') stays English while the
card copy localizes.
locales/en.yaml gains the gateway.exec_approval.* contract keys plus the
platform.shared.* / platform.slack.* / platform.matrix.* / platform.feishu.*
namespaces (and the keys for the other adapters wired in follow-up commits).
* i18n(gateway): run_turn / run_turn_runner / approval-settle copy through t()
- status hints, proxy errors, background task notices, progress heartbeats, session info lines
- tool progress chrome (tool_head/tool_pending/tool_preview/tool_verbose) shared by base.format_tool_event
- run_turn_runner:1406 Chinese clarify placeholder -> gateway.clarify.native_stream_placeholder (zh text kept in zh.yaml)
- _UNEXPECTED_SILENCE_REPLY/_CLARIFY_EXPIRED_NOTICE -> accessor functions
* i18n(platforms): route Google Chat and Teams user-facing text through t()
Google Chat clarify card, typing placeholder, orphan-card labels and the whole
/setup-files reply set (module constants become platform.google_chat.setup_files.*
keys resolved at reply time). The attachment-fallback notice that shipped
hardcoded in Spanish is keyed with an English en value; es.yaml carries the
original Spanish text for those four keys.
Teams approval card header/reason use the gateway.exec_approval.* contract,
_APPROVAL_LABELS becomes a key table resolved at click time, and the meeting
summary writer resolves its section headings/fallbacks per render.
* i18n(platforms): route LINE, WeCom, email, DingTalk, IRC and Home Assistant text through t()
LINE default copy constants become catalog keys resolved in __init__ (the
LINE_*_TEXT / extra.* operator overrides still win); the busy-ack bypass
matcher keys on the leading emoji marker only, so it keeps firing once the
gateway busy heads are localized. WeCom media size/format notices that shipped
hardcoded in Chinese are keyed with English en values and zh.yaml carries the
original Chinese text. DingTalk emotion bubbles resolve per send.
* i18n(cli): route /model switch output and -q status lines through t() (cli.model.*, cli.single_query.*)
Switch-summary labels shared with the gateway reuse gateway.model.* keys
(provider/context/max-output/capabilities/prompt-caching); CLI-only variants
(glyph or no-backtick forms) live under cli.model.*. The hand-padded /model usage
block becomes a (form, description-key) table padded at render time so the
command syntax stays fixed while descriptions translate. -q 'Error:' reuses
gateway.model.error_prefix.
* i18n(cli): route TUI panel/hint/placeholder copy through t() (cli.tui.*)
_APPROVAL_CHOICE_LABELS and _TUI_MODAL_HINTS become key tables resolved at
render time; vault/sudo panel bodies are one catalog value per panel split on
newline; inline plurals use <key>_one/<key>_other. Adds the cli.* namespace
(shared/tui/voice/render/subagents/dock) to locales/en.yaml.
* i18n(cli): voice/wake-word CLI copy through t() (cli.voice.*)
RuntimeError texts raised in _voice_start_recording are human copy (callers
print {e}) and are keyed; the Termux requirement-check match stays English.
Wake state ids stay internal, only their labels localize.
* i18n(cli): live-work dock, subagent monitor and render copy through t()
cli.subagents.* / cli.dock.* / cli.render.*; count fragments pluralize via
_one/_other keys, verdict table holds keys resolved at paint time so width
clipping measures the translated text.
* i18n(gateway): unauthorized/pairing, voice, topics, shutdown, startup, notifications, kanban pings through t()
* i18n(cli): move tips + composer placeholders into the catalog (tips.tNNN / tips.placeholder.pNN)
get_random_tip()/get_random_composer_placeholder() pick a key from the English catalog
(the parity baseline, probed once per process) and resolve it through t() for the active
language, so language packs translate tips like any other string. Also lands the cli.*
en.yaml namespace consumed by the CLI info/help/error-copy wiring in the next commit.
* i18n(cli): wire chat-turn + session mixins through t(); kanban log trimmer matches t() output
* test(cli): assert TUI/dock/voice copy via t(key); prove labels resolve at render time
Pinned-English assertions in the approval-UI, live-work dock and voice tests now
go through the catalog. New test swaps the catalog after import and checks the
approval panel + hint row follow it (the reason _APPROVAL_CHOICE_LABELS and
_TUI_MODAL_HINTS became key tables).
* i18n(cli): route CLI info/help/error copy through t() (cli.* namespace)
cli_info_mixin: /help consumes CommandDef.describe() (added to commands.py: slash.<name>.description
with fallback to .description), section titles/skill/quick-command headers, /tools, /toolsets,
/usage labels, /context, /whoami, /insights, /gateway status, tool-progress labels, bang-shell
denials, MCP config-watch + /reload-mcp confirm/reload lines, /reload-skills, and the session-store
warning all read the catalog at call time (module-level label tables became functions so the
active language is honoured after startup). cli.py: worktree cleanup, tirith warning, show_config
(labels re-padded at print time), quick/plugin/skill slash-command errors, ambiguous-command hint,
stdin error, gateway start, profile warning. cli_chat_error_copy / cli_unknown_command /
cli_output / cli_init_mixin: chat error panel copy, did-you-mean lines, n-more / yes-no prompt
(localized affirmative initial alongside 'y'), unknown-toolsets warning.
* i18n(w1a): wire agent display/explainers/approval + slash registry/help through t()
- hermes_cli/commands.py: CommandDef.describe() resolves slash.<name>.description
at call time; category labels via slash.category.*; help/alias/usage suffixes
via slash.shared.*; gateway_help_lines and commands_platforms/slash_exec use them.
- agent/display.py: display.verb.* resolved at call time via get_tool_verb();
bridge/spinner/thinking-verb/cute-row/failure/preview/diff text via display.*.
- agent/turn_explainers.py: exit-reason / persistence-cause tables become call-time
lookups (explainer.exit.*, explainer.persistence.*, explainer.file_mutation.*).
- agent/background_review.py, session_activity.py, context_breakdown.py,
status_output.py: review summaries, iteration progress, context notices.
- tools/approval.py, approval_context.py: approval.summary.*, approval.noun.*,
approval.window.* pluralized keys.
- gateway/slash_commands*.py: remaining raw strings (busy, whoami, platform,
bundles, memory, skills, approvals, set_home, diff, update, debug, profile,
heartbeat, refine, review, subgoal, loop, retry, compress codex path, save,
sessions, model guard/errors, agents rows, topup, login). HISTORY_UNREADABLE
keeps its English constant; callers use history_unreadable() ->
gateway.shared.history_unreadable.
- locales/en.yaml: new approval/display/explainer/slash blocks + gateway leftovers.
* i18n(telegram): route adapter chat copy through t()
Approval card (header/reason/smart-deny as HTML-escaped properties), inline
button labels, callback toasts (cut at Telegram's 200-char cap), model/choice
pickers, clarify/update/slash-confirm prompts, gmail-triage labels and the
inbound-media failure notice now come from the catalog. _UNAUTHORIZED is a
lazy _unauthorized() so the import no longer binds a language. The command
menu carries a language+payload fingerprint (forum scopes re-register on
change) and BotCommand descriptions are cut at 256.
Adds gateway.exec_approval.* (WAVE2 contract), platform.telegram.*,
platform.discord.* and the slash.*.description keys the Discord table shares
with the CLI registry to locales/en.yaml.
* i18n(gateway/platforms): whatsapp_cloud, yuanbao, weixin, signal, api_server copy through t()
- whatsapp_cloud: clarify list/buttons, approve/deny + slash-confirm labels via platform.whatsapp.* (t()-then-truncate at 20/24/72 caps); _EA_HEADER becomes a property wrapping ea_header_text()
- yuanbao: SLOW_RESPONSE_MESSAGE -> slow_response_message() (platform.yuanbao.slow_response_notice; zh keeps the original text); cron-wrapper markers centralized as module constants for strip_cron_wrapper
- api_server: PROVIDER_AUTH_FAILED_LABEL/PROVIDER_RATE_LIMITED_LABEL stay English for run.py matchers; user_text() renders via t()
- signal/_format_wait, weixin voice caption, openai_routes transformed notice
- run_turn: second _UNEXPECTED_SILENCE_REPLY consumer -> accessor
* i18n(discord): route adapter chat copy through t()
Native slash-command table becomes _NATIVE_SLASH_COMMAND_SPECS holding catalog
keys; _native_slash_commands() resolves descriptions, parameter descriptions
and Choice names for the active language, each cut at Discord's 100-char cap,
and the app-command sync fingerprint now includes get_language() so a
display.language change re-syncs. Exec-approval card (gateway.exec_approval.*
contract), slash-confirm / clarify / update views, model+choice pickers,
thread creation, forum titles, voice acks, the response-truncation notice,
the unauthorized-slash security alert and the media upload-size notices all
read from platform.discord.*. Decorator-declared button labels are relabelled
in __init__ (80-char cap); embed titles cut at 256, select placeholders at
150, option label/description at 100. _UNAUTHORIZED is a lazy _unauthorized().
* i18n(cli): wire status-bar, stream, terminal mixins + terminal_input through t(); rename kwargs that shadow t(key)
* i18n: wire hermes_cli/cli_commands_mixin.py slash-command copy through t()
- 431 new leaves under cli.commands.<cmd>.* in locales/en.yaml; 12 rows reuse
existing gateway.* keys (rollback, diff, resume, branch, btw, model, reasoning)
via a _gt() helper so CLI and gateway replies stay identical.
- Module-level English tables (_BUSY_MODE_*, _REASONING_TOGGLES, _HATCH_PROGRESS,
_DIFF_LABELS, _LOCAL_ENGINE_LINES) become call-time catalog lookups keyed by id.
- Verb tables (Enabling/Disabling, Paused/Resumed/Triggered, planned/done,
Updating/Generating) are one full template per variant; plurals use
<key>_one/<key>_other via _tn(); hand-padded column labels (/snapshot list)
translate the value and re-pad at the call site.
- Multi-line usage blocks are single catalog values split with _lines().
- Model-facing system notes and DB-stored reasons stay English (EXCLUDED).
* tests: assert /handoff, /worktree, /login CLI copy via t(key) instead of pinned English
* test(i18n): pin Telegram/Discord adapter catalog wiring
Lazy unauthorized notice, exec-approval contract keys, HTML escaping before
Telegram <b> wrapping, 200-char toast / 256-char BotCommand caps, Discord
100-char app-command text and 80-char button caps, and language-bearing
command-menu fingerprints on both platforms.
* i18n: reconcile cli.shared on/off vs enabled/disabled after lane merge
* i18n: describe() in TUI-gateway slash listings; localize TUI exit resume hint
* i18n(tr): translate bundled catalog + tui pack
* i18n(ja): translate bundled catalog + tui pack
* i18n(ko): translate bundled catalog + tui pack
* i18n(zh): translate bundled catalog + tui pack
* i18n(fr): translate bundled catalog + tui pack
* i18n(af): translate bundled catalog + tui pack
* i18n(uk): translate bundled catalog + tui pack
* i18n(ar): translate bundled catalog + tui pack
* i18n(pt): translate bundled catalog + tui pack
* i18n(it): translate bundled catalog + tui pack
* i18n(es): translate bundled catalog + tui pack
* i18n(zh-hant): translate bundled catalog + tui pack
* i18n(ru): translate bundled catalog + tui pack
* i18n(hu): translate bundled catalog + tui pack
* i18n(hu): translate pre-existing English-valued leftovers (kanban wake, /context, /status, fast labels)
* i18n(de): translate bundled catalog + tui pack
* i18n(ga): translate bundled catalog + tui pack
* test(i18n): fixture matches _normalize_lang(lang, home) signature
* i18n(tui): scaffold userMessages/slashCmd en siblings
* i18n(tui): wire secure prompts + content tables
* feat(tui): i18n — wire billing, subscription, connection-setup and journey overlays
Adds en siblings billing.ts / subscription.ts / connection.ts (namespaces
billing, subscription, connection, journey) and routes every user-facing
literal in billingOverlay, subscriptionOverlay, connectionSetupOverlay and
journey through useT()/messages(). Module-level label tables became lazy
(scopeStillDeniedResult(), verbOf(T, action)); auto-reload rows dispatch on
stable ids instead of label text. Regenerates locales/_keys.tui.json.
* i18n(tui): wire slash ops/wake replies
* i18n(tui): wire pickers (modelPicker, activeSessionSwitcher, petPicker)
* i18n(tui): wire slash core/debug/setup replies
* i18n(tui): wire hubs (agents overlay/panel/controls, skills, plugins)
* i18n(tui): wire slash session/topup/subscription replies
* i18n(tui): wire chat bits (branding, thinking, messageLine, loaders, todo, queued, banner, entry)
* i18n(tui): register t3 siblings (pickers, hubs, secure, content, chatBits) and regenerate keys
* i18n(tui): wire userMessages copy through the userMessages namespace
* i18n(tui): lazy-copy test for userMessages, regenerate _keys.tui.json
* i18n(tui): wire session/gateway/lib text through the TUI catalog (lane t2)
Adds en siblings session.ts, gatewayMsg.ts, libText.ts (namespaces session,
gatewayMsg, libText) and routes user-facing literals in app/{useMainApp,
useSessionLifecycle,useInputHandlers,turnController,createServerRequestHandler,
setupHandoff,createGatewayEventHandler}.ts, gatewayClient displayed reasons,
lib/*, domain/*, hooks/* through t()/messages(). Status-bar state values that
code compares against, backend-matched strings, log lines, model-bound text,
and machine 'error:' prefixes stay literal. Regenerates locales/_keys.tui.json
(232 keys).
* i18n: serve bundled locales/<lang>.tui.yaml under overlay/packs; TUI pack parity test; regen _keys.tui.json (1250)
* i18n: translate pre-existing English stubs in bundled locales (424 leaves, 14 locales)
* tui: i18n-export-en script (English templates for pack translators)
* docs(i18n): bundled TUI packs are the bottom layer of the tui surface
* i18n(ru): translate TUI pack
* i18n(ar): translate TUI pack
* i18n(es): translate TUI pack
* i18n(pt): translate TUI pack
* i18n(ko): translate TUI pack
* i18n(de): translate TUI pack
* i18n(ja): translate TUI pack
* i18n(fr): translate TUI pack
* i18n(tr): translate TUI pack
* i18n(it): translate TUI pack
* i18n(zh): translate TUI pack
* i18n(zh-hant): translate TUI pack
* i18n(hu): translate TUI pack
* i18n(uk): translate TUI pack
1,169 missing keys translated; 81 pre-existing kept byte-identical. Parity OK missing=0 extra=0 placeholder_mismatch=0 empty=0.
Deliberately identical to en: chatBits.branding.mcpSummary ({0} MCP), chatBits.thinking.agentsHint ((/agents)), session.main.voiceStt (◉ STT), session.main.voiceTtsSuffix ( [tts]), slashCmd.core.help.tuiSection (TUI), slashCmd.core.history.hermesTag (Hermes #{0}), slashCmd.debug.heapdump.heapPath (heapdump: {0}), slashCmd.debug.mem.rss (rss), subscription.stepUp.title (Remote Spending — product feature name, as in core catalog), content.faces.* (glyph-only kaomoji).
* i18n(ga): translate TUI pack
* i18n(af): translate TUI pack
* plugin_guard: locale catalogs in language packs step down the agent-config family
A translated status line such as "Updating AGENTS.md" in locales/<lang>.yaml is UI text the loader
reads as a string leaf; it cannot edit a file. The bundled en.yaml itself tripped agent_config_mod
at critical, making any faithful language pack uninstallable. Injection shapes keep full severity.
* plugin_validate_locales: read key exports with utf-8-sig (Windows footgun lint)
* i18n(relay): route relay adapter prompt copy through t(); drop dead import-bound approval header
Adds platform.relay.* (5 keys) to en and all 16 bundled locales, reusing the sibling platform
translations for the confirm buttons and the Other option.
* ci: fix TUI import order, MDX table pipe, main's overflow-warning wording in all locales; fresh-install fixture carries the i18n kernel
- ui-tui/src/i18n/en.ts: perfectionist/sort-imports (slash before slashCmd)
- docs plugins/index.md: escape the | inside the provides_locales table cell (MDX parsed <id> as JSX)
- display.notice.uncompressed_context_overflow: adopt main's wording (names compression.enabled: false
and /compact) in en + 16 locales; the guardrail test pins that phrase
- tests/scripts/test_fresh_source_install.py: the installer tail now resolves CLI text through
agent.i18n, so the fixture tree carries the i18n kernel + en.yaml (not the agent runtime)
* docs(desktop-plugin-sdk): double-backtick the template-literal example (MDX evaluated ${n})
* test(e2e): display.language is validated against the live language set; exclude it from the arbitrary-string set property
* commands: keep the localized COMMANDS/COMMANDS_BY_CATEGORY module __getattr__ after the compat block removal
* build: never write locales/_keys.*.json from the desktop/TUI builds; regenerate the committed desktop key export
The desktop build regenerated locales/_keys.desktop.json in the checkout, so a
hermes update that rebuilt the app left the tree dirty (Desktop update E2E:
'M locales/_keys.desktop.json'). The key exports are committed artifacts pinned
to en.ts by apps/desktop/scripts/i18n-keys.test.mjs and ui-tui i18n:keys:check;
builds read them, never write them. Regenerated after main's new desktop strings.
* test: unbreak two main-red timing tests the PR merge-ref inherits
- test_local_runtime racing fake publishes the modern state record (legacy pid-only
records are rejected since 65ff3ad353; main has been red on this test since)
- test_run_progress_topics ManyProgressLinesAgent waits for the first bubble instead of
a fixed 0.35s, which a loaded CI runner does not always meet
* chore(i18n): regenerate desktop key catalog for main's new strings (model pricing, copy changelog)
* test(e2e): torture-chamber fd monitor confirms a deleted sidecar is still held before calling it a leak
SQLite's WAL last-close unlinks -shm before closing its descriptor (unixShmUnmap, then
unixShmPurge), so a healthy close shows a (deleted) -shm for microseconds; the 20ms poll
occasionally caught that window on the short-lived opener and failed the episode.
* chore(i18n): regenerate desktop key catalog for main's telemetry/consent strings
* chore(i18n): regenerate desktop key catalog after main sync
---------
Co-authored-by: Teknium <teknium@nousresearch.com>
Two more spawns inherited the gateway's full environment. The bot
desktop launcher copied os.environ minus the display variables; the
agent drives that desktop and its dock opens xfce4-terminal, so every
provider key and bot token was one click away. The Telegram gmail-triage
buttons ran HERMES_HOME/scripts/gmail-triage/*.sh with no env at all, a
path the agent can write. The desktop now builds from
served_profile_child_env() like the agent browser (keeping the user's
HOME); the triage scripts use build_subprocess_env like cron and
webhook-filter scripts.
The scrubbed builder rewrites HOME to the profile home in containers and
profile home mode (the official image does this by default). These are
third-party programs whose own config and logins live under the user's
HOME: ~/.openviking, the raft CLI profile, ~/.config/buzz. The openviking
change copied the parent's HOME back, which is itself a profile home
when Hermes runs under one; use HERMES_REAL_HOME, which the builder
already computes, for all three.
The shape rule (a platform prefix plus _TOKEN/_SECRET/_PASSWORD/_KEY) also
matched variables Hermes never reads: the platform list holds plain words
(LOCAL, GATEWAY, WEBHOOK, SLACK), so a user's SLACK_USER_TOKEN,
LOCAL_LLM_API_KEY or GATEWAY_API_KEY vanished from the terminal and
terminal.env_passthrough could not bring them back. The prefix census it
leaned on also failed open (an unreadable plugins dir cached an empty set).
Adapter secrets now come from what adapters declare: password entries of
the messaging OPTIONAL_ENV_VARS (built-ins plus every platform plugin
manifest) and secret-named keys of the gateway env-override table, both
Tier 1 and refused by passthrough; plus the secret-named required_env of
adapters registered in the current profile scope, read per spawn without
loading deferred adapters, Tier 2 only because required_env is an
unchecked setup list. Secrets nothing declared get a manifest entry
(TELEGRAM_WEBHOOK_SECRET, PHOTON_SIDECAR_TOKEN, A2A_PUSH_SECRET,
TEAMS_GRAPH_ACCESS_TOKEN, TEAMS_INCOMING_WEBHOOK_URL) or join the
policy's read-in-code list (QQ_STT_API_KEY, the two MSGRAPH names).
_exec_buzz built the CLI env from os.environ.copy(), so every buzz call
carried Hermes' gateway tokens and provider keys. It now starts from
hermes_subprocess_env() and adds only the relay URL, key and auth tag.
The raft bridge was spawned with {**os.environ, RAFT_CHANNEL_TOKEN},
handing the raft CLI every Hermes credential. It now gets
hermes_subprocess_env() plus RAFT_PROFILE (the profile-resolved value,
not the launch env's) and the bridge token.
Ported from the raft half of #56245 onto current main.
* refactor(plugins): remove the Sep 2026 decomposition compat layer on schedule
The PLUGIN-COMPAT layer (2776813df3 + d63e380324 + 0a5164cebe) kept pre-#102117 import paths
alive for external plugins until 2026-09-14. That window closed two weeks ago; since then the loader
has already been skipping plugins that use the old paths. This removes the layer itself:
- 328 appended `PLUGIN-COMPAT` blocks (lazy `__getattr__` pointer tables, re-exported third-party
names, restored dead definitions) and the three re-export stub modules
(gateway/startup_watchdog, hermes_cli/observability/relay_runtime, tools/environments/modal_utils)
- COMPAT_MANIFEST.md, compat_manifest.json, scripts/check_compat_pointers.py and its lint step
- the reporting surfaces: CLI banner notice, `hermes plugins compat`, the `hermes doctor` section,
the post-update notice, the Desktop one-time dialog, the loader's pre-import skip and the
`plugins.allow_deprecated_imports` escape hatch
An external plugin that still imports an old path now fails to load with its ImportError as the
reason in `hermes plugins list`, the same path as any broken plugin.
hermes_cli/plugin_compat.py stays as three inert stubs (compat_report, removal_in_effect,
summary_lines): an already-running pre-removal `hermes update` lazy-imports them after the checkout
swap (tests/compat/old_updater_surface.json).
In-tree fallout, both already dead: hermes_cli/setup.py::_check_espeak_ng (no callers; its
`shutil` came from a compat block) and gateway/config.py::SessionResetPolicy ("retained solely for
the scheduled plugin-compat window"). Two test_run_agent patches targeted the removed
`run_agent.handle_function_call` pointer; they now patch `model_tools.handle_function_call`, the
seam production reads, like every sibling test in that file.
* chore: retrigger CI (zero-job startup_failure phantom)
* test: drop resolution allowlist rows for the two deleted which() sites
hermes_cli/setup.py::_check_espeak_ng (dead) and tools/skillevaluator_scan.py::scanner_available
(a restored definition inside a PLUGIN-COMPAT block) no longer exist; the stale-row gate requires
their allowlist entries go with them.
Inside a multiplexed secondary profile scope, `_profile_buzz_extra` read the buzz block only from
`gateway.platforms.buzz`, while the runtime loader (`gateway/config_loader.py::platform_section`,
the same seam that feeds this plugin's `apply_yaml_config_fn`) also resolves a top-level `buzz:`
block and the documented `platforms.buzz` shape. A fully configured secondary profile therefore
failed `check_requirements` closed on every boot ("Platform 'Buzz' requirements not met").
Resolve the section through `platform_section` so the gate sees exactly what the loader sees.
Row L1 (profile-scoped gate reads config), topology T2 (multiplexed secondary profile).
Salvaged from #126060 by @liuhao1024 (authorship kept); the test is trimmed to one A->B->A
invariant across two profile homes. Supersedes #126067 (same mechanism, later).
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Follow-up trim to @JE4NVRG's salvaged commit (is_connected → env_is_connected("A2A_PORT")):
- plugins/platforms/a2a/tools.py::_a2a_tools_available read os.getenv("A2A_PORT") too — the
same unscoped gate in the same plugin, so every secondary profile also paid for the a2a
toolset whenever the launch profile set A2A_PORT. Now get_scoped_secret, same policy as
the adapter's port read (scoped miss fails closed; default/T1 reads its own environ).
- Tests trimmed to two invariants: A→B→A over two temp profile homes under
set_multiplex_active(True) with the launch value in os.environ (profile B with no
A2A_PORT of its own is NOT connected and gets no tools; red on origin/main at index 1),
and the standalone T1 case where os.environ IS the profile.
Row L1 (boot-time platform enablement probe), topologies T2/T3 (multiplexed host serving
secondary profiles). No other plugins/platforms/*/__init__.py is_connected/check_requirements
does an unscoped env read (buzz #125985 is handled separately).
Co-authored-by: JE4NVRG <jean.v1803@gmail.com>
is_connected() gated on a raw os.getenv("A2A_PORT"), which under multiplexing
holds the launch (default) profile's value. Every secondary profile therefore
instantiated the inbound A2A server, all of them fell back to the module
default port 9900, and all but one died with "could not bind ... Address
already in use" (fatal / bind_failed in gateway_state.json) on every boot.
Resolve the gate through gateway.platforms._shared.env_is_connected, the same
scope-aware idiom the sibling adapters use -- the adapter's own port read
already goes through _get_scoped_secret for exactly this reason.
Tests: tests/plugins/test_a2a_plugin.py pins the behavior. A profile with no
A2A_PORT in its own scope is not connected and never reads os.environ (a raw
os.getenv fails the test), a scoped port connects, and extra.enabled still
short-circuits.
Fixes#122126
The r3 fold scoped only the dmarc verdict to its clause. SPF and DKIM
still came from a whole-string, last-match-wins regex scan, so the same
quoted/comment smuggle closed for dmarc still authenticated a spoofed
From (GHSA-rxqh-5572-8m77), e.g.
spf=fail smtp.mailfrom="x spf=pass smtp.mailfrom=example.com "@evil.test
spf=fail (spf=pass) smtp.mailfrom=a@example.com
spf=fail smtp.mailfrom=a.spf=pass@example.com
dkim=pass header.d=evil.test header.i="x header.d=example.com y"@evil.test
Every verdict now comes from the leading method=result token of its
own clause (from _ar_clauses, comments dropped), and its domains only
from that clause. Properties are read by a token scanner that consumes
quoted-strings and other key=value tokens whole, so quoted contents are
never read as properties while a quoted value still is
(header.from="example.com"). SPF fails closed on more than one spf
clause; DKIM accepts any single dkim=pass clause whose own header.d
aligns (multi-signature mail is normal), never mixing clauses. The
whole-string methods/props (methods["dmarc"] was dead) are gone.
Also: a stray ')' at depth 0 is now unbalanced (it split header.from
out of the dmarc clause); a From with more than 64 '(' takes the silent
empty-sender drop instead of a parseaddr RecursionError logged as an
error; the cap test asserts the cap directly instead of wall-clock
timing; the empty-From drop assertion moves next to the other
_extract_email_address rejects; and the untested >1-dmarc, unbalanced
and backslash-escape rules get reject strings.
Round-3 review of the #124322 salvage:
- stdlib parseaddr is pure Python and superlinear on hostile input: a
100KB `From: <a@a@...` held the GIL ~1s per message (main ~0), and any
remote sender reaches it before auth. Refuse values over _MAX_FROM_LEN
(2048; RFC 5322 lines cap at 998) right after unfolding so they take the
existing empty-sender drop. The timing assertion now times
_extract_email_address itself instead of a private regex.
- The dmarc clause split ignored quoted-strings and stripped comments one
level only, so an attacker-controlled SPF-passing envelope sender like
smtp.mailfrom="x;dmarc=pass header.from=example.com x"@evil.test planted
a fake dmarc=pass ahead of the real dmarc=fail and authenticated
admin@example.com. Split clauses with a quote- and nested-comment-aware
scanner (comments dropped, quoted values kept so header.from="x" is still
read and unquoted), and fail closed on an unbalanced value or when more
than one clause starts with dmarc=. Tests pin the smuggled clause,
reason="a;b", the nested comment, a quoted aligned header.from, and that
every header.from in the dmarc clause must align (all->any goes red).
- `Doe, John (CEO) <j@x>` was dropped because the fallback refused any
paren. Strip (comments) from the display part first; `attacker@evil.test
(c) <victim>` stays rejected by the '@' rule.
- Results without '@' (`John` -> `john`, `a\"b <v@x>` -> `a\`) were
dispatched as sender ids; return "" so they take the drop path.
The r1 malformed-From fallback regex ([^<>\s]+@[^<>\s]+) backtracked
catastrophically on `From: <a@a@a@...`: 32KB held the GIL ~23s, and any
remote sender reaches it before auth. Capture <([^<>\s]+)> instead and
check the '@' in Python (same accepted set, 100KB in ~3ms).
The fallback also mapped multi-mailbox / group / comment-prefixed From
values (`attacker@evil.test, <victim@x>`, `Grp: a@evil; <victim@x>`,
`attacker@evil.test (c) <victim@x>`) to the bracketed victim, which a
dmarc=pass without header.from then authenticated. Only fall back when the
display part has no ; : ( ) and no '@' unless it is exactly the bracketed
address; `Doe, John <j@x>` and `j@x <j@x>` still resolve. The rule now lives
only in the docstring (the old inline comment was wrong).
dmarc: strip (comments) before splitting clauses, and take the verdict and
header.from from the same first clause that starts with dmarc=, so
`dmarc=pass (p=none; sp=none) header.from=evil.test`, a later dmarc=pass
clause after dmarc=fail, and duplicate misaligned header.from are rejected,
while `arc=pass (dmarc=fail ...); dmarc=pass header.from=<ours>` passes.
Property clean-up is shared via _auth_props.
The empty-sender drop now runs right after address parsing and gets an
assertion (it was untested). Tests extend existing ones (no new tests).
Strict parseaddr returns '' for common RFC-invalid From values that the
old regex resolved: an unquoted address as display name
(john@example.com <john@example.com>), an unquoted comma (Doe, John
<j@example.com>), or a@x.test <b@y.test>. Allowlisted senders using such
clients were silently dropped, and under open access every one of them
shared an empty chat_id. Fall back to the bracketed address only for the
unambiguous shape (no quotes, exactly one <...> pair), so the quoted
display-name spoof still resolves to the attacker. _parse_fetched_message
now drops messages whose From yields no address instead of dispatching
an empty identity.
The dmarc header.from alignment check read header.from from props merged
across all clauses, so a later dkim clause's header.from could override
the dmarc clause's value. Read it from the dmarc clause only, and pin the
misaligned dmarc=pass rejection in the existing dmarc test (previously
no test covered it). Also drop a redundant _domain_of ternary, trim the
_extract_email_address docstring and a duplicate test assertion.
_verify_sender_authentication accepted any dmarc=pass verdict, even one
issued for a different domain than the From we parsed. That is what let
the quoted-display-name spoof ride on the attacker's own truthful DMARC
pass. When the trusted Authentication-Results names header.from, require
it to align with the From domain; otherwise fall through to the aligned
SPF/DKIM checks. Defense in depth for the #124322 parser fix.
_extract_email_address took the FIRST <...> pair, so
From: "Victim <victim@example.com>" <attacker@evil.test> resolved to the
victim. The attacker's own domain passes DMARC truthfully, so the
allowlist (EMAIL_ALLOWED_USERS), pairing and session identity were all
evaluated against an address the sender does not control.
Use email.utils.parseaddr, which keeps the quoted text as the display
name and returns the real addr-spec. RFC 5322 folding is unfolded first
so a folded quoted display name is not mistaken for the mailbox.
Salvages #124322.
Both reconnect paths (gateway/run_adapters.py watcher and multiplex
secondary) build a FRESH SimplexAdapter before connect(is_reconnect=True),
so the per-instance _allowlist_warned flag never suppressed anything: a
daemon-down cold boot re-logged the warning on every backoff retry. Gating
on `not is_reconnect` would instead lose the warning when the first connect
fails. Dedup at module level keyed on (hermes_home_key(), frozenset(names)),
still checked before the connectivity probe. No shared warn-once helper
exists (plugin_compat.warn_once is compat-specific).
Read the value with platform_gate_env (the reader authz uses; differs from
get_scoped_secret when a scope is installed with multiplex off) and decode
JSON list literals with decode_json_list_literal like _coerce_allow_set, so
'["4","9"]' written by `hermes config set` no longer warns that valid IDs
are ignored.
The caplog test now builds two fresh adapters (first connect fails, second
succeeds) and asserts exactly one warning naming only 'alice'; it fails with
2 warnings against the pre-fold adapter.
The name-entry warning in connect() read SIMPLEX_ALLOWED_USERS via raw
os.getenv, while authz reads it profile-scoped. Under multiplexing a
secondary profile would warn about (or stay silent on) the default
profile's list rather than the one actually enforced. Use the module's
_get_scoped_secret + _parse_comma_list like __init__ does.
It also only fired on a successful non-reconnect connect: if the daemon
was down at cold boot the first connect() failed and every retry came in
with is_reconnect=True, so the warning never appeared. Evaluate it before
the connectivity probe, once per adapter via an instance flag.
Test: two connects (first fails) -> exactly one warning naming only
'alice' for scoped '4, alice' while os.environ holds 'bob'. Red on the
pre-fold adapter (0 warnings) and on a raw-os.getenv variant (names bob).
After #44729 SIMPLEX_ALLOWED_USERS matches only the numeric contactId, but
the docs still told operators display names work, and existing name
entries would silently stop matching. Update the docs and log a one-time
warning at first connect listing non-numeric entries that are now ignored.
The SimpleX sender allowlist (SIMPLEX_ALLOWED_USERS) previously matched
against both the stable numeric contactId (user_id) and the mutable
display name (user_name). Since any SimpleX contact can change their
localDisplayName / profile.displayName to match another user's, this
allowed an unauthorized contact to bypass the allowlist by setting a
colliding display name.
Remove the user_name check so that SIMPLEX_ALLOWED_USERS only matches
on the immutable contactId. Operators must use numeric contact IDs in
the allowlist.
Fixes#44729
(cherry picked from commit b4aa29da1567d45920f79aabdb36b44c5f87bde5)
The r3 fold made park refuse an older voice once a newer one from the same
sender had parked. With one /sync batch holding [v1 (slow gate), m1 bare
mention, v2 (fast)], v2 parked first, v1 was then refused, and m1 claimed
v2 -- a voice sent after the mention. v1 was lost and v2's own mention was
answered as empty text.
The bare mention now takes an arrival limit (ParkedVoices.mark) before it
settles, and claim pops the newest parked voice that began before that
limit, dropping older ones. park no longer refuses by arrival; parked
voices are kept per sender ordered by seq (bounded to 4). A claim made
while gates are still in flight records a floor, so a late older voice
(seq <= claimed) cannot park and outlive the claim; the floor clears when
the sender's in-flight list empties. One answer per bare mention, newest
before the mention wins, and the r3 orphan case still leaves nothing
parked.
Also: pending() ignores entries past CLAIM_WINDOW_SECONDS so an expired
voice no longer sends every later text through the mention regexes; the
m.thread root lookup is one _thread_root helper used by both the park
decision and its pre-check; voice_gate is annotated.
The kept test gains a [voice slow, mention, voice2 fast] + mention2 row:
red on a932dc031c (['$voice2', '$text2']) and on f1ea71de7c
(['$voice', '$text2']).
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
The same-sync-batch in-flight mark was one slot per (room, sender), so a
second voice's begin overwrote the first. When the older voice finished
gating after the newer one, the bare mention settled on (and claimed) the
newer voice only, and the older one parked afterwards as an orphan: the
sender's next unrelated bare mention within 120s re-dispatched it, so a
stale voice got downloaded, transcribed and answered.
The in-flight mark is now a list of gates per key. release removes only
its own gate (idempotent) and drops the key once it is empty; settle
waits on all current gates under the same 5s cap. Each gate carries an
arrival sequence and park refuses an older voice once a newer voice from
the same sender has parked, so a late older voice can neither replace
nor outlive the claim of the newer one (still one parked voice per
sender, newest wins).
The mark was also held across the whole _resolve_message_context,
including the display-name fetch and thread mark, and taken for voices
that can never park (the voice itself @mentions the bot, free rooms, bot
threads, non-allowlisted rooms). A bare mention racing such a voice
waited for the voice's display-name fetch (0.61s vs 0.31s on main with a
0.3s fetch). begin now runs only when a synchronous pre-check says the
voice can park, and the mark is released as soon as the park decision
is made (before the display-name fetch; DM voices release there too),
with the finally still covering exceptions and cancellation.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
mautrix dispatches every event of one /sync batch as its own task
(wait_sync=True). The voice only parks after awaiting the room identity,
which is a homeserver round-trip whenever the 60s identity cache is stale,
i.e. in any idle room. The bare-mention text checked the park without
awaiting, found nothing, dispatched an empty text, and the voice then
parked and expired unclaimed.
A parkable voice is now marked in-flight per (room, sender) right before
its first await and released in a finally once it is gated. A bare
mention that sees an in-flight voice waits for it (bounded, 5s) before
claiming; ordinary text only pays two dict lookups and never waits.
After a claim the bare-mention event also gets its read receipt, so the
read marker is not left one event short of the pre-claim behaviour.
Cleanups: ParkedVoices drops its unused window param, stores
(parked_at, voice) instead of a flat tuple, the voice_mention import
moves to the top import block, and the redundant require_mention check
on the claim path is dropped (parking/in-flight only happen under it).
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
Review cleanups on the parked-voice fix:
- The claim bypass is now an explicit mention_claimed parameter threaded
_handle_media_message -> _resolve_message_context. The old _claimed set was
only drained inside the require_mention branch, so a voice whose thread became
a bot thread while parked leaked its id for the adapter's lifetime.
- One MSC3245 predicate (has_voice_marker, `.get(...) is not None`) shared by
is_voice_event and _classify_inbound_media; the two used to disagree on a
null marker (parked as voice, classified as AUDIO).
- One mention helper (_content_mentions_bot) for both the gate and the
bare-mention claim, instead of two copies of the m.mentions extraction.
- ParkedVoices.has() dict lookup gates the claim, so ordinary text messages
no longer pay _strip_mention/regex work when nothing is parked.
- Test: a same-room mention WITH text stays a text (gives the bare-only guard
teeth) and the parked voice is asserted never downloaded.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
Element X ships an MSC3245 voice event with an empty m.mentions block and
sends the mention the user typed while recording as a separate m.text
event right after. With require_mention on, the gate dropped the voice and
the follow-up bare @bot was dispatched as an empty text, so the bot never
answered the voice.
Park an unmentioned voice keyed by (room_id, sender) instead of forgetting
it; a bare mention from the same sender in the SAME room within 120s claims
it and the voice is processed in place of the empty text. The claimed event
id passes the mention gate exactly once, so it is not re-parked; expired
entries are pruned on every access. Parked voices are never downloaded or
transcribed (no wake-word/STT of unmentioned audio), and a mention in
another room never pulls a voice across rooms.
The state lives in a small sibling module (voice_mention.py) because
adapter.py is already far past the file-size budget.
Co-authored-by: miregal89 <142085869+miregal89@users.noreply.github.com>
The model-picker fix added a second copy of the choice picker's auth
check, and each copy rebuilt the callback ctx that _handle_callback_query
had already computed as `cb`. The chat-id picker loop is the only route
into both handlers. Gate there once with the existing cb, so a picker
added to that loop later is covered without anyone having to copy the
check again. Fail-closed behaviour is unchanged: _callback_authorized
returns True only for an authorized tapper and answers the denial text
otherwise. The back-button test calls the handler directly and no longer
needs its auth stub.
Co-authored-by: Dusk1e <yusufalweshdemir@gmail.com>
Every other ModelPickerView callback (provider/model select, expensive
confirm, back) runs through the shared _HermesView._gate auth check, but
_on_cancel did not. A non-allowlisted member could tap Cancel on the
owner's picker, marking it resolved and clearing the view. No model
switch was possible, so impact is low, but the gate should be uniform.
Reuse the same _gate call the sibling handlers use.
Model picker taps (mp:/mpg:/mpv:/mm:/mc:/mb/mx/mg:) never went through the
callback allowlist, so a group member who is not allowlisted could tap the
owner's /model picker and switch the owner's session or config model. Gate
the handler on _callback_authorized like the choice-picker, approval and
clarify buttons, before any picker state is read.
Ported from #33859 (gateway/platforms/telegram.py) to the plugin adapter.
(cherry picked from commit b002be8feacd09fde0ee9752de7b396d0aac69ba)
- salvage summary cap and _bound_oversized_record still composed bare
truncation idioms in model-visible text; route them through elide /
elide_middle so a copied marker is guard-visible.
- the active-task line repr()'d the elided text, escaping the marker's
apostrophe when the user text held both quote kinds and hiding it from
the guard; elide after repr instead (text within the cap stays whole).
- a leftover budget smaller than the marker produced a content-free,
over-budget marker line in _build_verbatim_user_section and the Slack
nested-attachment path; skip the item instead.
- _build_verbatim_user_section elided twice, reporting the wrong total;
one elide at min(cap, remaining).
- drop redundant len() pre-checks before elide() and name verification
stop's repeated 1200.
Co-authored-by: ahisblessed <ahisblessed@users.noreply.github.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
The Slack Block Kit payload dump and the nested-attachment text budget
are both fed to the agent, and trajectory_compressor's summarizer input
becomes training data; all three still used the imitable bare
"... [truncated]" idiom. Route them through elide()/elide_middle() with
module-level imports and extend the no-idiom invariant to scan
plugins/platforms/slack and trajectory_compressor.py.
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
With reply_in_thread: false the whole channel is one session and the bot
answers top-level, so an unmentioned top-level message there is a
follow-up in a conversation the bot is part of, like a thread reply.
reply_expected is now False for a free-channel message only when it starts
its own session (a new top-level thread), else None. The bot-id set is
built inside _slack_reply_expected, as _channel_gate_allows does.
The Slack rule marked every admitted message that was not a DM, a mention
or a command as not addressed, so a plain "done?" in a thread the bot is part
of, or a reaction trigger, could end on a bare silence marker and vanish,
the case #111624 fixed (#110952).
reply_expected is now False only for a message that opens by @mentioning
someone else, or a top-level message a free-response channel admitted
without a mention. Reaction triggers and pipe-form self mentions count as
addressed; other thread replies are None (visible fallback). The
free-channel predicate moves into _slack_is_free_channel so the gate and
the rule read the same one. The test drives the real _handle_slack_message.
Docs describe the rule in its own note instead of the
ignore_other_user_mentions tip, and the messaging index documents the
human-turn fallback.
Since 5ea8fb2b78 (#111624, for #110952) the gateway rejects a bare silence
marker on any human turn and delivers "The model returned only a silence
marker for a message that needed a reply" instead. That protects a human
who asked this bot something and got nothing back. It also fires on every
human message the adapter admitted without the bot being addressed at all:
a free-response channel, a thread follow-up under
`thread_require_mention: false`, or a message @-mentioning another person
or bot with `ignore_other_user_mentions: false`. A bot whose SOUL declines
peer-addressed turns with a deliberate marker now posts that notice on
every such message. A fleet running several bots in shared Slack threads
reported it as spam on v2026.9.21. #37940 established that intentional
silence must not be re-inflated. Both contracts hold once the turn knows
whether a reply was expected.
`MessageEvent.reply_expected` (True, False, None) is set by the adapter
where the message is admitted. Slack (`slack_reply_expected`): a 1:1 DM,
an @mention of this bot or a command is True, anything else it admits is
False. Other adapters leave None, which keeps today's behaviour, so nothing
changes for them until they are ported. `response_filters.silence_allowed`
holds the one rule (machinery turn, or reply not expected) and both call
sites use it: the live turn in `run_turn._hmwa_shape_agent_response` and
the crash-recovery redelivery from #120377 (1136f135dd), which reads the
flag back from the persisted turn metadata. The suppressed case logs one
DEBUG line naming platform and chat.
Operator workaround until this lands: `platforms.slack.extra.
ignore_other_user_mentions: true` drops peer-addressed messages before a
turn exists.
(cherry picked from commit 094439776ab898cccde303a1c2c911c8ab5bfb75)
Four surfaces build a throwaway AIAgent and never call close() — the
owner boundary that releases memory-provider sessions, tool
subprocesses and httpx clients. In long-lived processes each run leaked
all of them until exit:
- batch_runner._process_single_prompt: one agent per prompt, N prompts
per batch process.
- feishu_comment._run_comment_agent: one agent per comment run in the
gateway process.
- tui_gateway prompt.background: one side agent per background turn.
- cli /bg: one agent per background task in the CLI process.
Wrap each run in try/finally with a suppressed close(), mirroring
gateway/run.py's owner pattern. preview.restart stays deliberately
unclosed (its task exists to leave a detached server running), and the
prompt.background side agent is safe to close: its session_id is the bg
task id, so close() reaps only its own task resources.
Fixes#50197
- A mid-turn notify reply (/status, /approve, clarify answer) shares the
stream's thread key; it no longer seals and overwrites the half-streamed
answer. In-place replacement now requires the final to match the stream
after normalizing mrkdwn markers and whitespace; anything else posts fresh
and leaves the stream open.
- One _commit_stream helper for both seal-then-commit paths, so the rewrite
path also falls back to chat.update when stopStream fails.
- A stream reopened after a server-side seal is seeded with only the text
past the sealed message (tracked as 'base'), not the whole segment.
- Streams older than 15 min are sealed and dropped on the next start.
- _stream_key reuses _workspace_thread_key/scope_id_for_chat; the stream
dict no longer duplicates chat/team ids.
5648f81431 fixed this exact server-side seal (Slack closes a native stream
after a few minutes of a long turn, live-observed at ~5m20s; the lifetime
is not documented) for the native task-card stream: on
message_not_in_streaming_state from appendStream, drop the dead ts and
start a fresh stream, seeded with the full current content so nothing is
lost.
send_draft — the plain-text native streaming path used when task cards
are not enabled — hits the identical seal but never got the fix: its
generic except block only recognizes the feature-gate markers
(not_allowed, missing_scope, ...) and otherwise just logs debug and
returns failure. gateway/stream_consumer_transport.py's
_send_draft_frame() docstring is explicit that "any failure permanently
disables drafts for this run" — so a long turn streaming as plain text
degrades to the edit-based fallback for its remainder exactly the way
the task-card bug did before 5648f81431.
Mirror the task-card fix: on message_not_in_streaming_state from
chat.appendStream, drop the dead ts and _start_stream() a fresh one
seeded with the full accumulated text (not just the delta), so the next
frame's delta still resumes correctly. One reopen per frame; a second
rejection propagates as a real failure, matching the twin's behavior.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
(cherry picked from commit 2b4ff4e23bdf684d2fde1a9512f11dcd7ef99c42)
A mrkdwn-rewritten turn-final (e.g. *Done:* -> _Done:_) no longer continues the
streamed text, so it was classified unrelated: the stale stream was sealed and
send() posted a second message (#95430 cause B). Seal, then chat.update the
sealed ts with the final; post fresh only if the in-place update fails.
Co-authored-by: liguoyu <guoyu.li@lcfuturecenter.com>
Problem: with native streaming (chat.startStream/appendStream/stopStream)
the same answer could land twice in a thread — once as the streamed
message, once as a fresh chat.postMessage — while the streamed message
kept its live-typing indicator.
Mechanism: `_try_finalize_stream` matched the turn-final against the
streamed text with a raw `startswith`. The agent strips `final_response`
and joins footers with `rstrip()`, so any surrounding-whitespace
difference made the finalize fall through to a plain post although the
open stream already showed the whole answer (and was never sealed). A
`chat.stopStream` failure took the same fresh-post path even when the
streamed text equalled the final. Streams were also keyed per `chat_id`
only, so two concurrent turns in two threads of one channel sealed or
overwrote each other's stream.
Fix:
- Key native streams per `(team_id, chat_id, thread_ts)`; the stream
consumer stamps the same `thread_id` on every draft frame and on the
turn-final `send()`, so both resolve to the same key.
- Honor the streaming contract (gateway/AGENTS.md): sends carrying
`_interim_send` or `expect_edits` never seal a stream.
- Classify the final against the streamed text as equal / extends /
unrelated with edge-whitespace tolerance (`_stream_relation`). The
stopStream delta is sliced from the RAW final, so nothing inside the
answer (blank lines, fences, tables) is dropped or repeated.
- Commit rule: one `chat.stopStream`, one retry only when no tail is
appended (`markdown_text` APPENDS, so an ambiguous failure must not
repeat it), then `edit_message(finalize=True)` on the stream ts as the
idempotent in-place commit — it already owns format/truncate/Block Kit
and the block-rejection retry. Only when both fail does `send()` post
a fresh message (a duplicate beats a lost answer).
- Oversized tails and rewritten finals (`notify=True`) seal the stale
stream on what is visible before falling back, so no stream is left
with a live-typing indicator.
- `_seal_stream` takes the exact unsent delta instead of recomputing it
from `final_text`; `disconnect()` and the stream API calls route
through the stream's own team client.
Tests: tests/gateway/test_slack_native_streaming.py covers the
whitespace-only difference, the stopStream-failure commit path, the
bounded retry, the uncommittable fallback, interim/preview sends,
per-thread keying, oversized tails, rewritten finals and the
GatewayStreamConsumer end-to-end path.
(cherry picked from commit a64d10071ce7816b124467e29407d7d47bfdde8d)
dingtalk-stream 0.24.3 start() retries forever inside the SDK, logging a
malformed logger.exception() every 3s, so the adapter breaker never saw the
error. Install a dedup filter on the SDK logger that can't raise on bad
format args, detect the websockets incompatibility (bare or chained
TypeError), log one ERROR with a pin-consistent hint, and hand off via
_set_fatal_error(retryable=False) + _notify_fatal_error(): only a
reinstall of the pinned versions and a restart fixes it. Breaker stays
tripped until the error type changes; constants moved to module level.
The reconnect storm in #24851 is driven by a dingtalk-stream/websockets
incompatibility that raises TypeError on every start(); backoff never
recovers it. Log the first occurrence per error run at ERROR with an
upgrade hint instead of a generic WARNING.
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
DingTalk stream-mode reconnection storms the gateway: when start() raises
the same error every cycle, _run_stream logs a WARNING and reconnects at the
60s cap forever, generating hundreds of MB of identical log lines and hanging
the gateway.
Add a per-error-type circuit breaker: after 5 consecutive identical errors,
suppress the repeated WARNING (one ERROR summarises) and pause 300s instead of
spinning at 60s. Reset backoff + counters after a clean start() so a recovered
connection is treated fresh.
Salvage of #24881 re-implemented on current main: the original fix targeted
gateway/platforms/dingtalk.py, which has since been refactored to
plugins/platforms/dingtalk/adapter.py. Same logic, new path; tests import the
new module.
Closes#24851.
(cherry picked from commit 84201c58255ae3d3c9b269ac777ea0ff444d9229)
Surface code/msg at debug when the recall API returns non-success
(matches the exception path) and trim the docstring. Drop the trivial
disconnected-client test.
Feishu had no delete_message, so a failed finalize-edit plus fallback
send left the truncated edit bubble next to the full final. Implement
the SDK delete and thread the fallback send to the originating message.
(cherry picked from commit c61add84ad40b1bc39288405a0a05b4f621e00fc)