Commit Graph

2414 Commits

Author SHA1 Message Date
teknium1
19b29df13b fix: video generation tools no longer let the agent pick the model
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.

The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.

Refs #83080
2026-09-19 12:22:44 -07:00
kshitijk4poor
6d8badcc6e refactor(telegram): share the cold-boot queue decision and log it
The cold-boot queue policy was computed inline at both start paths, and the
default (drop) was silent — a command sent while the gateway was offline
vanished with no log line (the second half of #71811).

One helper now owns the decision and logs it on every cold boot, so an
operator can tell which policy applied. Behaviour is unchanged.
2026-09-20 00:19:28 +05:30
justcarlosm
a4a0447b04 feat(telegram): drop_pending_on_cold_boot knob to preserve offline queue 2026-09-20 00:19:28 +05:30
teknium1
349f67778f chore: merge origin/main (resolve agent/transports/codex.py) 2026-09-19 10:51:50 -07:00
teknium1
4586cde64d chore: mark the deliberate /tmp literals and shrink the lint baseline to one code block
The seventeen remaining literals are container-side paths, AF_UNIX socket-path-limit
candidates on darwin, detection needles, guard regexes and guidance text that tells the
model to avoid /tmp. Each carries an inline `no-tmp: ok — <why>` so the reason lives
next to the line; the baseline keeps only a fenced tree listing where a marker would render.
2026-09-19 10:44:26 -07:00
teknium1
e16fee4db1 refactor: resolve runtime temp paths via tempfile/TMPDIR instead of literal /tmp
Hermes now routes scratch space through HERMES_HOME/cache/scratch (exported as
TMPDIR), so every production path that still spelled out /tmp bypassed that and
kept teaching the agent the habit. Fallbacks in tool_result_storage,
code_execution_tool, process_registry, the ACP child HOME, mini_swe_runner's
local cwd, and the CI/profiling scripts now use tempfile.gettempdir(); shell
installers fall back to $TMPDIR (then HERMES_HOME) when mktemp is missing, and
repro/eval shells use `mktemp -d -t`. User-facing help text and sample payloads
(hermes send, approvals test, hooks test, voice-mode WSL hints, meet_bot debug
line) no longer suggest /tmp.

Container-side paths (mini_swe_runner docker cwd, sandbox base env, remote
sync tarballs) keep the literal because they name the sandbox filesystem,
not the host.
2026-09-19 10:44:26 -07:00
teknium1
0a1bc07f15 docs(honcho): say that initOnSessionStart blocks tools-mode startup until Honcho answers
What: the `initOnSessionStart` setup-schema description and the Honcho docs page
now state that in `recallMode: tools` the eager init runs synchronously during
agent construction, that it should stay false for Desktop or a local Honcho that
may be down, and that `timeout` in honcho.json caps each SDK call (both keys
added to the Full Config Reference table).

Why: #51492 reports Desktop `session.resume` / `prompt.submit` timeouts when a
local Honcho is down with `initOnSessionStart: true`. The synchronous eager path
is intentional design (#51562 was closed for that reason:
"ready-before-first-tool-call semantics"), so the fix for users is knowing the
trade-off and the two knobs that bound it.
2026-09-19 10:39:32 -07:00
teknium1
78cee6bcbd fix(google_meet): stop the PCM tail thread from swallowing every exception
_pcm_tail_loop caught bare Exception, so any bug in the tail thread died
silently while the bot stayed mute. Only the expected closed-pipe errors
(OSError incl. BrokenPipeError, ValueError on a closed stdin) are quiet now;
anything else surfaces through the thread's excepthook.
2026-09-19 10:07:42 -07:00
teknium1
eb02bec644 test(google_meet): pin late Join button, mic unmute and stdin-fed PCM pump; document caption-derived input
Two invariant tests through the real plugin helpers: a fake page whose Join button
renders after a delay must still be clicked and a muted toggle unmuted (control: a live
mic is reported, never toggled); a `cat` stand-in for paplay must receive bytes appended
to speaker.pcm after the pump started and the tail thread must be gone after teardown.
Both fail on the previous meet_bot.py.

Docs: realtime mode is speak-only - incoming speech is Meet caption scraping, meeting
audio is never sent to the Realtime session (the docs page wrongly said the meeting audio
was transcribed through the configured STT provider); micState is documented.

Fixes #80875
2026-09-19 10:07:42 -07:00
DavidMetcalfe
03874cd461 fix(google_meet): keep the realtime PCM pump alive and unmute the mic after admission
The Linux/macOS pump (paplay / ffmpeg) was started against speaker.pcm while the
file was still empty, read it to EOF and exited before OpenAI Realtime appended any
audio - audioBytesOut grew while nobody heard the bot. The pump now reads raw PCM
from stdin, fed by a tail thread that follows the file as it grows; teardown joins it.

Meet can also seat an authenticated bot muted. On admission the bot now checks the
in-call mic toggle, clicks it when the label says "Turn on microphone", and reports
the outcome as micState in status.json (unmuted_clicked / unmuted / unknown).

Slim redo of #81488 onto the refactored _start_pcm_pump / _drain_loop helpers.
Part of #80875 (atoms 2 and 3).
2026-09-19 10:07:42 -07:00
minh-alpon
5381c6c2c3 fix(google_meet): poll for the Join button instead of clicking once after domcontentloaded
Meet renders "Join now" / "Ask to join" asynchronously after the page reaches
domcontentloaded; the one-shot _join() check could run before the button existed
and the bot sat silently in the pre-join screen. Poll for up to 30 s (0.5 s cadence),
returning as soon as a button was clicked.

Slim port of #37945 onto the refactored _join() helper.
Part of #80875 (atom 1).
2026-09-19 10:07:42 -07:00
teknium1
00e1a55519 fix(setup): Image Generation 'OpenAI (Codex auth)' row starts Codex sign-in and names the real auth command
Selecting Image Generation -> OpenAI (Codex auth) in `hermes setup` / `hermes tools` on a
fresh install saved `image_gen.provider=openai-codex` and printed "no configuration
needed!" without ever signing in, so the backend was unusable until the user guessed the
auth command — and the schema hint pointed at `hermes auth codex`, which does not exist.

Root cause: the row declares `env_vars: []` and only a `post_setup_hint`, a key nothing
consumes; `_configure_provider` runs a hook only for `post_setup`.

- plugins/image_gen/openai-codex: schema declares `post_setup: "openai_codex"`; the hint
  and the `auth_required` error name `hermes auth add openai-codex`.
- hermes_cli/tools_config_post_setup: `_post_setup_openai_codex` in the existing
  `_POST_SETUP_HOOKS` table (sibling of the `xai_grok` credential bootstrap). With
  credentials present it continues; otherwise it offers the device-code sign-in (or skip)
  and saves tokens with `set_active=False`, so picking an image backend never rewrites
  `model.provider` the way the model-provider login does.
- `_POST_SETUP_AUTH_READY` table replaces the `post_setup == "xai_grok"` special case in
  `provider_readiness_status`, so any credential-bootstrap row reports ready/needs_auth
  from the auth store.
- `_save_codex_tokens(set_active=...)` mirrors `_save_xai_oauth_tokens`.

Live probe (real `_configure_provider`, real plugin row, temp HERMES_HOME, OAuth start
stubbed with a recorder): before — post_setup=None, OAuth fired [], hint `hermes auth codex`;
after — no creds: OAuth fired once, logged in, model.provider untouched; with creds: OAuth
not fired.

Fixes #102144
Salvages #102165 (@liuhao1024) — superseded: same direction (post_setup hook), redone on
the split tools_config siblings without calling `_login_openai_codex`, which would have
switched the main model provider.
2026-09-19 10:05:26 -07:00
teknium1
fc08fdb171 feat(image_gen): OpenAI image provider sends custom model ids verbatim and can reuse a named custom endpoint
Custom model ids (#97928): a value of image_gen.openai.model / OPENAI_IMAGE_MODEL outside the
gpt-image-2 / gpt-image-2.5 catalog used to fall through silently to the default tier, so a
gateway serving its own image model names always received `model: gpt-image-2` +
`quality: medium`. resolve_static_model(passthrough=True) now returns the unknown id as the
API model with quality=None, and the request omits `quality` for it — OpenAI-compatible
gateways reject enum values they do not know, and a foreign model has no OpenAI quality
tiers. Only the provider-scoped key and env var pass through; the shared top-level
image_gen.model can hold another provider's id (a FAL path) and is never forwarded.

Named custom endpoint (#83080, config-reuse half): image_gen.openai.provider names a
providers:/custom_providers: entry; its base_url and api_key/key_env fill in whatever
image_gen.openai.base_url / key_env leave unset, so a gateway already declared for chat is
not re-declared with a duplicated key. Explicit image_gen.openai values keep precedence; an
unknown name logs a warning and falls back to the OpenAI variables. The lookup goes through
hermes_cli.runtime_provider._get_named_custom_provider, the same resolver the auxiliary
clients use, so aliases and legacy list entries behave identically.

image_gen is an open config root (deliberately absent from DEFAULT_CONFIG), so
`image_gen.openai.provider` already validates as a known key; docs cover both behaviours.

Live probe (fake /v1/images/generations gateway recording the body, real provider on a temp
HERMES_HOME): before — `custom-image-model` sent as model=gpt-image-2 quality=medium; named
provider → is_available False / auth_required. After — model=custom-image-model, no quality;
named provider → request to the entry's URL with `Bearer <its key_env>`; the catalog model
still maps to gpt-image-2 + quality=medium.
2026-09-19 09:51:12 -07:00
teknium1
3fa2c45bd8 feat(image_gen): OpenAI image provider takes base_url/key_env from config, blanks OpenAI-Project, bypasses macOS system proxy
What: plugins/image_gen/openai resolves its endpoint and credential through one
resolver — image_gen.openai.base_url → OPENAI_BASE_URL → SDK default, and the env
var named by image_gen.openai.key_env → OPENAI_API_KEY — shared by is_available()
and generate() so the two cannot disagree. The client is built on
build_keepalive_http_client (env-only proxy policy) and sends a blank
OpenAI-Project header.

Why: the image endpoint could only be routed via the process-wide OPENAI_BASE_URL /
OPENAI_API_KEY, so a local or third-party image gateway could not be configured
independently of the chat provider (#65309, #97928, #13798). openai.OpenAI() with
trust_env routed localhost endpoints through a macOS system proxy whose
ExceptionsList httpx never sees (#64888). An OPENAI_PROJECT_ID set for chat made
/v1/images/generations 403 model_not_found on projects with a model allow-list
even though the key already carries the project (#60748).

Slim redo of the contributor direction in #18796 (@y0shua1ee), #37208/#37209
(@charzhou), #65312/#65323/#64893 (@asdlem), #60749 (@perelin); all predate the
StaticImageGenProvider refactor and no longer apply.
2026-09-19 09:51:12 -07:00
teknium1
981cb83020 fix(custom): clamp top-level reasoning_effort to Groq's none/default vocabulary on api.groq.com
Groq's OpenAI-compatible wire accepts top-level reasoning_effort only as
"none" or "default" (#75089, Defect 1); the custom profile forwarded the
configured graded level ("medium"/"high") and the request still 400'd on
both the main transport and the auxiliary path. The clamp lives in
CustomProfile.build_api_kwargs_extras, keyed on the resolved base_url host,
so both paths share it. The aux bare-key case now drives call_llm with a
fake client instead of the private _build_call_kwargs.
2026-09-19 09:43:52 -07:00
Chinmayrawat15
1447c83b46 fix(acp): resolve reasoning_config for ACP and Feishu comment agents
`SessionManager._make_agent` and the Feishu doc-comment agent built their
`AIAgent` without `reasoning_config`, so `agent.reasoning_effort: none`
never reached those sessions: the transport applied its default effort,
which non-reasoning models such as gpt-4o-mini reject with HTTP 400 and
which silently re-enables thinking everywhere else. Both surfaces now go
through `hermes_constants.resolve_reasoning_config`, the same chokepoint
the CLI, gateway, TUI, cron and `hermes -p` already use, resolved against
the model the session actually runs so per-model overrides apply.

Ported from PR #85164 by @Chinmayrawat15 (oneshot hunk already on main).

Fixes #85153
2026-09-19 09:39:50 -07:00
kshitijk4poor
0c7738dfd6 refactor(discord): read the free-response auto-thread flag through a getter
Mirrors the _discord_require_mention()/_discord_max_attachment_bytes() shape, makes the
lookup lazy (it only runs on the free-channel path now), and keeps the in-code key
manifests in _handle_message and register() listing the new key.
2026-09-19 21:43:59 +05:30
Xunjin ZHENG
774e070731 feat(discord): add free_response_auto_thread opt-in
Free-response channels skip auto-threading by default so the bot replies
inline (lightweight chat mode). This prevented users who wanted BOTH
mention-free replies AND per-conversation threads from getting either.

Add a new opt-in `discord.free_response_auto_thread` (env:
`DISCORD_FREE_RESPONSE_AUTO_THREAD`, default false) that, when true,
re-enables auto-threading in free-response channels. Voice-linked
channels continue to skip auto-thread regardless, and the flag is
gated behind the global `DISCORD_AUTO_THREAD=true`.

Default behavior is unchanged; all 291 existing discord tests pass.
2026-09-19 21:43:59 +05:30
teknium1
80f4c6d6bf fix(acp): warm numpy with the memory provider so hindsight is covered too (#58083)
The warm-up imported only the provider module. holographic / mnemosyne import
numpy at module top, but hindsight defers the ML stack to is_available() ->
_check_local_runtime() (importlib of hindsight / sentence_transformers), which
ran later on a to_thread worker racing acp-mcp-discovery — the reported hindsight
stack was still reachable. The deadlock partner is numpy's lazy _core init in
every reporter's dump, and a plain `import numpy` up front was every reporter's
workaround, so import_memory_provider_module now also imports numpy (best-effort)
once the provider module is in.

Also: import_memory_provider_module() defaults to the configured memory.provider,
so entry.py drops its duplicate config resolver and outer try; the "ONLY thread"
comment is reworded — hermes_cli's plugin-discovery thread is already running
when hermes acp dispatches.
2026-09-19 02:43:28 -07:00
teknium1
6a93afa0a8 fix(acp): pre-import the memory provider on the main thread on Windows (#58083)
Every faulthandler dump in the thread shows session/new stuck in numpy's
create_module on the main thread while another thread (MCP discovery / ACP
stdin reader) sits in the same lazy import chain — a first-time native
extension import racing another thread deadlocks on Windows (holographic,
mnemosyne and hindsight all reproduce it; a sitecustomize `import numpy`
before any thread exists resolves it every time).

hermes acp now imports the configured memory.provider's module on the main
thread before the MCP-discovery thread and asyncio.run() start (Windows only —
the deadlock is Windows-specific and the import is paid once either way).
plugins.memory.import_memory_provider_module imports the module without
constructing a provider or running register(); the agent build later finds it
in sys.modules.

Trimmed from #91775 (@tigercraft4): same placement and gating; reuses the
existing plugin loader instead of a second module-import routine.

Co-authored-by: tigercraft4 <tigercraft4@tigercraft4.com>
2026-09-19 02:43:28 -07:00
teknium1
396013b17e test: one invariant per bounded cache (LSP docs/baselines, tool-call logger, fuzzy roots, hindsight turns)
Each test is red on origin/main and green on the fix; each also asserts the
control case (evicted LSP file re-opens with diagnostics, every log line
written once, unexpired fuzzy root kept, every hindsight turn shipped once).
Trims the hindsight comment to the why.
2026-09-19 01:55:38 -07:00
Tikkanaditya Siddartha Jyothi
6e1de4850e fix(hindsight): bound the append-mode session turn buffer
_session_turns accumulated every turn's text for the whole session, reset
only at session boundaries, so a never-ending session grew without bound
independent of context compaction.

In append mode each retain ships only the delta since the last watermark
(_session_turns[_last_retained_turn_count:]), and a retained turn is never
read again — sync_turn always slices from the watermark and flush-on-switch
flushes what's left. So after an append retain the buffer drops the retained
prefix (clear() + reset the watermark), bounding it to the un-retained tail.

Overwrite mode is deliberately untouched: legacy/overwrite APIs resend the
whole session each retain because each retain replaces the document, so that
path must keep every turn.

Tests: append trims the retained prefix while shipping every turn exactly
once (no loss, no duplication), stays bounded across 100+ turns, and
overwrite mode keeps the full buffer.

(cherry picked from commit cb77e006754dd4ad5c880f22310d3ac7b3472739)
2026-09-19 01:55:38 -07:00
teknium1
44e6ea3ceb fix(web): openai-native backend docs name the real login command, English picker tag, registry test lists the plugin
Trim of the salvaged openai-native marker provider (#107377):
- `hermes auth --provider openai-codex` does not exist; the login command is
  `hermes auth add openai-codex` (hermes_cli/subcommands/auth.py). Fixed in
  plugin.yaml, provider.py docstring and the web-search docs (2 spots).
- `hermes tools` picker tag was Chinese; now English like every other provider.
- tests/plugins/web/test_web_search_provider_plugins.py enumerates the bundled
  registry exactly, so the new plugin made it fail; it now lists openai-native
  with capability flags (search=True, extract=False) — the search-only
  invariant that keeps web_extract on its own backend.
- contributors/emails mapping for the salvaged author.
2026-09-19 01:51:23 -07:00
lyswty
82c77ff9b9 feat(web): add openai-native backend for Codex server-side web_search
Declares OpenAI's provider-executed Responses `web_search` built-in in place of
the client-side `web_search` function, mirroring the existing xAI native-search
path. Selected via `web.search_backend: openai-native`; search-only, so
`web_extract` keeps resolving to its own backend.

The Responses adapter already recognises built-in tool types
(`_RESPONSES_BUILTIN_TOOL_TYPES`) and preflight passes them through, so the only
missing piece was the swap itself plus a provider name for the config to point at.

Gating is deliberate: two-sided (Codex backend AND a selected openai-native
backend) and fail-closed, so a custom OpenAI-compatible endpoint or an
unresolved provider leaves the client tool untouched.
2026-09-19 01:44:49 -07:00
teknium1
1654575c6c fix(gateway): canonicalize identity first at every adapter ingress path
Every adapter-side lane (text/photo/album batch dicts, _active_sessions, the
busy guard, /stop /new /reset and clarify replies) derived its session key
before the receiving bot's identity was known, so two bots seeing the same
Telegram chat.id == user.id collided on one agent:main lane and a route to an
unserved profile still reached the default lane's running task.

BasePlatformAdapter._canonicalize pins the RoutingIdentity (via the new
session_identity.canonical_identity seam) as the FIRST statement of
handle_message, _enqueue_text_event, _handle_message_while_active, Telegram
_route_photo_event and every _source_session_key; _drop_unresolved drops a
rejected route at the first seam with one WARNING. Topic recovery copies the
source through replace_source so the identity travels; Slack thread sources go
through build_source so the thread key carries the same provenance.
2026-09-19 00:30:29 -07:00
teknium1
c07708671d fix(gateway): every adapter session key goes through one seam (+ lint)
A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.

Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.

Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.

Phase 2 of #88715.
2026-09-18 22:04:43 -07:00
teknium1
e56ec8c9f7 fix(chronos): a 403 invalid_client from NAS hands cron fires to the built-in ticker (#97494)
NAS maps the agent-cron bearer to a provisioned instance via an `agent:*` client or the
`hermes-cli-vps` bootstrap session (hermes-portal `server/agent-cron/instance-auth.ts`). A
container whose auth.json holds a plain `hermes-cli` user login is refused with 403
invalid_client on every arm, re-arm and list, for the life of that credential - and the
re-login users try first (`hermes auth logout nous` + device code) replaces the bootstrap
session, making it permanent. Chronos previously logged one bare warning per job and left
the jobs with no trigger at all: they only ran through the misfire sweep, minutes late.

NasCronClientError now carries the HTTP status and the OAuth `error` code. On an identity
rejection the provider logs ONE warning that names the real remedy (restore the hosted
credential from the Nous Portal; re-login cannot fix it), stops calling NAS, and starts the
built-in ticker with the gateway's own adapters/loop so scheduled jobs keep firing on time.
Transient 5xx/transport failures keep retrying on the next reconcile.

Supersedes #97566 (@wesleysimplicio), whose runtime-credential swap resolves to the same
bearer on main (`agent_key` is the access token) and so could not change the 403.
2026-09-18 20:07:21 -07:00
teknium1
255b4fd9da fix(image-gen): record token usage as soon as the billed HTTP 200 lands
OpenRouter (_generate_via_chat, _generate_via_image_api) and OpenAI gpt-image
called record_token_usage only after image extraction and save succeeded. A
token-billed 200 that returned text but no image (the "returned no image"
fallback case), an empty `data` list, or a failed save therefore consumed
billed tokens that never reached session_model_usage; on the fallback chain
only the fallback model's tokens were recorded.

Move the call to immediately after a successful post_json / SDK call, before
extraction/save, in all three paths. One call per HTTP success, so no double
count; failure responses (non-2xx, timeouts) still record nothing because
there is no body to read usage from.

Tests: the parametrized invariant in test_openrouter_compat_provider gains the
no-image-with-usage case for both surfaces (result is `empty_response`, one row
still recorded); the OpenAI test is parametrized on has_image the same way.
2026-09-18 19:56:51 -07:00
teknium1
6babdc96b8 fix(image-gen): every token-billed image backend records session usage (chat path, Image API, OpenAI gpt-image)
Widen the salvaged Image API hunk (#114340) to the whole class:

- plugins/image_gen/_common.py: record_token_usage() — one helper feeding the
  aux accounting chokepoint (agent.aux_accounting.record_aux_usage) with task
  "image_generation", the billing provider and the priced model id. Dict and
  SDK usage objects alike; a body without tokens is a no-op, as is a call
  outside a turn.
- plugins/image_gen/openrouter: the /chat/completions path (the DEFAULT model
  chain — openai/gpt-5.4-image-2, google/gemini-3-pro-image — is chat-only and
  token-billed) now records too; the Image API path uses the shared helper
  (the contributor's local _record_image_api_usage is folded into it) and both
  pass base_url so pricing resolves the route. Task renamed image_gen ->
  image_generation to match the other aux task names.
- plugins/image_gen/openai: gpt-image bills per text/image token; record the
  Images API usage block under the API model (gpt-image-2), not the Hermes
  quality-tier label.
- FAL, xAI, Krea, DeepInfra, Meta, openai-codex return no token usage and stay
  unrecorded (nothing to bill per token).
- tests: the contributor's two tests folded into one invariant parametrized
  over chat / Image API / no-usage control; one OpenAI invariant.
- docs: image-generation.md "How It Works Internally" gains the accounting step.

Fixes #114324
2026-09-18 19:56:51 -07:00
Yagna Vudathu
98ad0d8148 fix(image-gen): record token-billed OpenRouter calls to session usage
OpenRouter token-billed image calls parsed usage into extra only, so
session_model_usage never saw them. Record prompt/completion tokens via the
ambient aux accounting on success; flat-fee responses without token usage
write nothing.

Fixes #114324.
2026-09-18 19:56:51 -07:00
LeonSGP43
f59c451f81 fix(feishu): avoid threading regular replies 2026-09-19 03:31:00 +05:30
kshitijk4poor
35b38f1f4e fix(whatsapp): bridge gets the adapter's resolved group policy and group allowlist; one allowlist reader
The bridge now gates groups by WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
(#73465), but only the DM policy and DM allowlist were exported from the adapter's
resolved config — a YAML group_allow_from never reached the bridge and, under a
multiplexed gateway, the launch process's WHATSAPP_GROUP_* values did. _bridge_env
exports both like it already did for the DM pair, so unlisted-group traffic is cut
before media download instead of after the POST to Python.

whatsapp_common._select_allowlist generalises the DM precedence reader (config key
presence wins, an explicit empty list stays authoritative, then the first truthy env
carrier); the adapter's group list and the Cloud sibling's group list use it instead
of hand-rolled `or` chains, so `group_allow_from: []` means the same everywhere.

Tests: bridge env carries group policy + group allowlist from the secondary profile's
YAML (scoped precedence file); env fallback test trimmed to the two invariants (red on
the merge base); bare test adapters seed _group_policy/_group_allow_from, which
_bridge_env now reads.
2026-09-19 03:15:40 +05:30
sal
477afd2420 fix(whatsapp): adapter reads WHATSAPP_GROUP_ALLOWED_USERS for the group allowlist (#72529)
The Node bridge receives WHATSAPP_GROUP_ALLOWED_USERS via
_BRIDGE_PASSTHROUGH_ENV and the YAML bridge maps config.yaml
whatsapp.group_allow_from onto the same env carrier, but
WhatsAppAdapter.__init__ seeded _group_allow_from from config.extra
alone — an install that set the group allowlist only through the
documented env var silently ran group gating with an empty allowlist:
with group_policy=allowlist every group chat was rejected at intake
(the second failure mode triaged on #72529, the half with 'no PR yet').

Mirror the DM path's precedence: config keys win by key presence
(explicit empty list stays authoritative), then the profile-scoped
WHATSAPP_GROUP_ALLOW_FROM / WHATSAPP_GROUP_ALLOWED_USERS env CSV —
multiplex-safe through _wenv, like every other WHATSAPP_* read.

Tests pin the full precedence matrix: env-only seeding, the
allow_from-style alias, config-wins-over-env, the legacy groupAllowFrom
key, empty-by-default, the scoped .env read under multiplex, and the
group-policy gating outcome for an env-seeded allowlist.
2026-09-19 03:15:40 +05:30
kshitijk4poor
8503ee4459 refactor(send): route WhatsApp mentions through the existing standalone chunker
Follow-up to the salvaged #92440 commit, shape-gate cleanup only; behaviour is unchanged
(one mention-bearing payload per logical send, captioned media keeps its caption).

- Fold `_send_whatsapp_with_mentions` into `_send_plugin_standalone` (it was a line-for-line
  copy of the caption split + `_send_chunks` loop); `mentions` is attached to the first
  payload only via a one-shot kwarg dict.
- Drop the `inspect.signature(sender)` probe: the only registered WhatsApp standalone sender
  is the in-tree `_standalone_send`, which gains `mentions` in the same change; a foreign
  sender already surfaces as a TypeError through `_handle_send`'s error path.
- Collapse `_normalize_outbound_mentions` to dedupe-only; its input is argparse `list[str]`
  already validated by the CLI.
- Revert the `\d` -> `[0-9]` edit to `_BARE_PHONE_RE` / `to_whatsapp_jid`: unrelated to the
  feature (`normalize_whatsapp_mention_jid` already rejects non-ASCII via `isascii()`) and it
  changed output for nine other `to_whatsapp_jid` callers.
- Split the single 137-line test into a `whatsapp_bridge` fixture + two invariants
  (rejections never reach the bridge; mentions ride the first payload only, stale bridge
  fails closed); drop the fabricated legacy-sender branch that only existed to cover the
  deleted probe. Mutation check: removing the first-payload gate turns the new test red.
2026-09-19 00:24:50 +05:30
Andrey
b34ebc084a feat(send): add WhatsApp native mentions
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
2026-09-19 00:24:50 +05:30
teknium1
b44a481334 fix(video-gen): trust the operator-configured origin on the first download hop
The OpenRouter video content URL is built from OPENROUTER_BASE_URL, not from
a provider response, yet save_url ran the full SSRF check on it. An operator
pointing base_url at a LAN/loopback relay could submit and poll (raw
requests) but the final download was refused as SSRF unless they set
security.allow_private_urls — contradicting the issue scope ("operator-
configured endpoints out of scope") and the security.md sentence that the
operator's own base_url is unaffected.

save_url grows a `trusted_origin` flag (plumbed through save_url_video and
set only by OpenRouterVideoGenProvider._save_completed_video): the first hop
skips the private-address class check and uses a plain client, but the
cloud-metadata floor (is_always_blocked_url) still applies, auth headers
stay on hop 1 only, and every redirect target is re-validated in full so a
relay cannot bounce us to another internal address. Provider-returned result
URLs (fal, xai, image providers) keep the full guard — default is False.

security.md now states the precise scope: only the direct base_url hop is
exempt; result URLs from a LAN-hosted provider still need allow_private_urls.
2026-09-18 10:22:14 -07:00
beardthelion
3933fdf63b fix(security): route remote-party-supplied URL fetches through the SSRF guard
Provider response URLs, model-supplied image refs, manifest-derived pet
URLs, and remote sitemap <loc> entries were fetched with raw
requests/httpx/urllib — bypassing tools/url_safety while every platform
media path already uses it. A hostile or compromised provider/manifest
endpoint could steer a server-side fetch at internal or metadata
addresses; several sites cache the body where it is deliverable back.

Apply the canonical is_safe_url + create_ssrf_safe_client pattern at
every site: per-hop revalidation at TCP connect (closing the
DNS-rebinding window), bounded redirect chains that fail closed on
missing Location, and caller headers scoped to the first hop only —
matching the openrouter provider's own documented contract that its
bearer key must never leave the operator-selected host.

Operator-configured endpoints and pinned release assets are out of
scope — those URLs are operator-selected, not remote-party-controlled.

Fixes #114468
Closes #44728

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: AlexFucuson9 <AlexFucuson9@users.noreply.github.com>
Co-authored-by: Ray <rayjun0412@gmail.com>
Co-authored-by: zapabob <1920071390@campus.ouj.ac.jp>
2026-09-18 10:22:14 -07:00
teknium1
0985aeb732 fix(kanban): board switcher shows the bound project with an unbind action
The create/settings dialogs (salvaged from #114664) let a user set or clear
a board's project_id, but the binding was invisible outside Settings.
GET /boards already annotates every board with project_id + project_name,
so the switcher now renders a "Project: <name>" badge whose × sends
PATCH {project_id: ""} — the same clear the settings dialog uses — without
touching default_workdir.

Settings also stops sending default_workdir: "" next to a chosen project
when the directory field is blank: the explicit "" suppressed the server's
project → default_workdir mirror, so binding via Settings left the board
without the workspace default the create dialog would have seeded.

Trim the salvaged test to the invariants (selector wiring, payload shapes,
badge unbind) and document the control on the Kanban docs page.
2026-09-18 10:19:54 -07:00
liuhao1024
bc3da64fe8 fix(kanban): expose the board's project binding in the dashboard dialogs
The REST API already accepts project_id on board create/update and
GET /boards annotates project_id/project_name, but the dashboard UI
never wired any of it: a board's project binding could only be set
through raw API calls. Add a project selector to the New board and
Board settings dialogs, populated from GET /projects, mirroring the
existing default_workdir wiring. Settings PATCHes send project_id
unconditionally when the selector rendered ("" clears the binding,
server-validated); when the projects store is unreachable the field
is omitted so saving unrelated settings cannot wipe a binding.

Fixes #114652
2026-09-18 10:19:54 -07:00
teknium1
595f3a289c fix(telegram): split replies resume from the refused chunk, never re-send the head; per-chat send order + flood cooldown
A reply past 4,096 chars goes out as several sendMessage calls. When chunk 2 was
refused by flood control (RetryAfter past the 5s inline cap) send() returned the
bare flood_control result, so _send_with_retry re-sent the WHOLE payload after
the wait: the user saw chunk 1 twice (reporter: 4 messages, 686 duplicated words),
and paths without a ledger row lost the tail outright.

- send() now reports a mid-split refusal through the existing partial_overflow
  contract (the key _edit_overflow_split already sets and the stream consumer
  reads): delivered_chunks / total_chunks / last_message_id, plus
  undelivered_chunks + delivered_message_ids ONLY when non-delivery is certain
  (flood cap, Bot API rejection, connect/pool timeout) — an ambiguous TimedOut may
  have reached Telegram and is never resumed from.
- BasePlatformAdapter._send_with_retry resumes from the remainder via a new
  _resume_partial_send hook (default None = keep the partial failure, never
  re-send the head; the plain-text fallback is skipped for partials too). The
  Telegram override sends the leftover formatted chunks, continuing the id sequence.
- Per-chat FIFO send gate, reentrant per asyncio task (media paths nest), held
  only around the API calls — never across the reconnect wait — on send() and the
  media funnel, so concurrent replies to one chat no longer interleave chunks.
- A flood refusal arms a per-chat cooldown (mirrors the sendChatAction cooldown,
  capped 300s); sends inside the window fail closed locally with the same
  flood_control:<s> result and no API call, so ledger recognition and redelivery
  timing are unchanged.
- Edit path: log "refusing (retry_after Ns > cap)" after the cap check instead of
  "waiting Ns" followed by no wait.

Live against a local fake Telegram Bot API with a fake token (RetryAfter=7 on
chunk 2 of a 3-chunk reply): before 4 messages / 324 duplicated words; after 3
messages, 853/853 words, 0 duplicated, 0 lost. Two concurrent 3-chunk sends:
before 9 source switches, after 1. Five sends inside a refused window: before 5
API calls, after 0.

Fixes #114396

Co-authored-by: AStrnbrg <45151087+AStrnbrg@users.noreply.github.com>
Co-authored-by: whyyagswhy <166958865+whyyagswhy@users.noreply.github.com>
2026-09-18 10:17:23 -07:00
teknium1
7f2b64f7a2 fix(platforms): set media_text_inlined in the sibling document paths
The Feishu fix on this branch populates MessageEvent.media_text_inlined so
run_inbound's document note stops claiming "Its content has been included
below" when a text attachment was NOT inlined (>100 KB gate or decode
failure); run_inbound treats a missing flag as inlined. Telegram, Discord,
Slack, the WhatsApp bridge adapter and whatsapp_cloud inline "[Content of
…]" the same way but never set the flag, so their notes lied on the
skip path. Mirror the Feishu/buzz per-attachment contract in each: False
for every cached attachment, flipped to True only when the text was
actually injected.

One parametrized (small→True / large→False) test per adapter in the
existing per-platform test files.
2026-09-18 10:14:38 -07:00
fangliquan
51fa2d51a2 fix(feishu): retain attachments in text batches 2026-09-18 10:14:38 -07:00
fangliquan
2e58dbc9c5 fix(feishu): retain attachment inline flags in batches 2026-09-18 10:14:38 -07:00
fangliquan
29490a0e66 fix(feishu): preserve post file attachments 2026-09-18 10:14:38 -07:00
teknium1
0d0ccb8834 fix(feishu): a receive loop ending during disconnect() is not logged as a died link
The exit-notify wrap reported every receive-loop exception at ERROR with a
traceback, including the ConnectionClosedOK that follows our own CLOSE frame in
disconnect(). Gate on the adapter's _running flag (published through the WS
thread-local next to on_link_up): a live link's death stays ERROR, an
intentional shutdown logs at DEBUG. Live pass side-effect #4 on #113662.
2026-09-18 10:07:27 -07:00
teknium1
88a9ca7be4 fix(feishu): publish retrying when the SDK's own reconnect ladder runs
The ``retrying`` publication landed only on the supervisor-rebuild path (WS
thread dies). On the live link ``_auto_reconnect`` is on, so a receive-loop
error runs lark-oapi's ``_reconnect()`` ladder *inside* the receive loop — the
thread never dies and ``gateway_state.json`` kept saying ``connected`` while
the link was down.

Set the SDK's ``Client.on_reconnecting`` observer (lark-oapi 1.6.8
ws/client.py, fired first thing in ``_reconnect()``) from
``_apply_runtime_ws_overrides``; it hops to the adapter loop and publishes
``retrying`` through the same runtime-status call the supervisor uses, gated
on the client still being the live one. ``connected`` is re-stamped by the
existing ``_ws_link_up`` hook once the ladder's ``_connect()`` schedules a new
receive loop, so no ``on_reconnected`` hook is needed.

Docs: the WebSocket-mode section now says a dead link is rebuilt by Hermes'
supervisor and shows ``retrying`` in gateway status until re-established.

Test: test_sdk_reconnect_ladder_publishes_retrying (red before: FeishuAdapter
has no _ws_link_retrying / observer never set; green after).
2026-09-18 10:07:27 -07:00
teknium1
6d7b53a48a fix(feishu): stop the worker loop only on receive-loop exceptions; publish retrying/connected
Follow-up to the salvaged #113668 commits:

- The wrap stopped the worker loop in a ``finally``, i.e. also on a *normal*
  return. On the live link ``_auto_reconnect`` is the SDK default (True — the
  adapter only flips it off at teardown), and a successful SDK reconnect makes
  the old ``_receive_message_loop`` return after ``_connect()`` scheduled a fresh
  one. Stopping the loop there would tear down the healthy rebuilt link (no
  CLOSE frame, see #10202) on every transient blip. Stop only when the coroutine
  exits with an exception: ladder disabled, or the ladder's
  ``ClientException`` / ``ServerUnreachableException`` re-raise.
- Do not re-raise after logging: the wrap is now the owner of that exception,
  so asyncio no longer prints "Task exception was never retrieved" beside it.
- Wrap the real symbol unconditionally: ``lark_oapi.ws.client.Client
  ._receive_message_loop`` exists on every shipped SDK (1.6.x and 1.7.3, ws
  module unchanged between them); a getattr fallback would silently reopen the
  gap on a rename.
- ``gateway_state.json`` kept saying ``connected`` while the supervisor
  rebuilt: publish ``retrying`` when the WS thread dies (the adapter is still
  running, so ``_mark_disconnected`` is the wrong tool) and re-stamp
  ``connected`` from ``_ws_link_up`` — fired through a thread-local hook when
  the SDK schedules a receive loop, the only in-thread proof the handshake
  succeeded — hopped onto the adapter loop and gated on the client still being
  the live one.
- Tests: the fake ``lark_oapi.ws.client`` modules carry ``Client`` like the real
  module; second invariant now covers the normal-return (ladder success) case;
  the existing supervisor-restart test asserts the ``retrying`` publication.

Live: real lark-oapi 1.6.8 ``ws.Client`` + the real adapter worker/supervisor
against a local fake Feishu endpoint + websocket server with fake credentials.
2026-09-18 10:07:27 -07:00
liuhao1024
38f79aff9f log the receive-loop root cause before the supervisor rebuild
The wrapper re-raised the SDK receive-loop exception into the same
unretrieved bare task. Log it with logger.exception first so the root
cause lands next to the supervisor's rebuild line (suggested in review).
2026-09-18 10:07:27 -07:00
liuhao1024
5743dbb703 fix(feishu): stop the WS worker loop when the SDK receive loop dies
lark_oapi parks Client.start() in run_until_complete(_select()) and runs
the receive loop as a bare create_task whose exception nobody retrieves.
With Hermes disabling the SDK's reconnect ladder, a mid-life receive-loop
death leaves a deaf-but-ESTABLISHED socket: start() never returns, the
supervisor's executor future never completes, and gateway_state.json keeps
reporting connected (#113662).

Wrap Client._receive_message_loop in the existing isolation installer so
its exit stops the worker loop; start() then raises, the future completes,
and _supervise_websocket_thread rebuilds the link with its capped backoff.
Deliberate disconnects stay unaffected: they nil _ws_client first and the
supervisor exits without restarting.

Fixes #113662
2026-09-18 10:07:27 -07:00
teknium1
3ffee76761 fix(gateway): race IPv6/IPv4 on every cold-start WebSocket dial
The bootstrap racer covers sync connects only. The gateway's WebSocket dials
(relay connector, Yuanbao, Buzz) go through ``websockets.connect`` →
``loop.create_connection``, whose ``happy_eyeballs_delay`` defaults to ``None``:
a serial walk that burns the full connect timeout on every blackholed AAAA
record before IPv4 answers — the same stall class #114265 reports, one layer up.

``websockets`` forwards unknown kwargs to ``loop.create_connection``, so each
call site passes ``happy_eyeballs_delay=0.25`` (the RFC 8305 delay anyio and
the sync racer already use). One invariant test per call site captures the
kwargs at a mocked ``websockets.connect``.
2026-09-18 09:56:28 -07:00