Commit Graph

2473 Commits

Author SHA1 Message Date
kshitijk4poor
533796a833 refactor(discord): trim the first-strike comment to its WHY
The six-line comment above the terminal-reason check
(plugins/platforms/discord/adapter.py:1789-1794) restated the resume-swap
mechanism already documented in the test docstring and the issue. Keep the
one sentence that carries the intent (closed transport = confirmed death,
soft signals keep the threshold) and the #118487 pointer. No code change.
2026-09-22 13:39:00 +05:30
kshitijk4poor
fe225b20b1 fix(discord): treat client_closed as terminal like socket_closed
_read_websocket_health reports client_closed when Bot.is_closed() is
true. That is the same transport-dead state as socket_closed, yet it
stayed on the two-strike confirmation path and could show the same
1/2-then-silent-reset pattern if the bot task's done callback were ever
suppressed. Escalate both terminal reasons on the first strike.
2026-09-22 13:39:00 +05:30
kshitijk4poor
43499c4531 fix(discord): log an unexpected liveness-probe exit at WARNING (from #118504)
The guard exit in _liveness_loop fires for two very different reasons:
ordinary teardown (adapter stopped or disconnecting), and a still-running
adapter whose client vanished. The first is noise at INFO; the second
means the gateway keeps running with no watchdog (#118487) and must stand
out in an incident log. Pick the level from the exit cause instead of
logging both at INFO.

Co-authored-by: Konstantin Khlopkov <47825603+kokhlo@users.noreply.github.com>
2026-09-22 13:39:00 +05:30
liuhao1024
b1a69d8178 fix(discord): escalate the first socket_closed liveness strike
A closed Gateway transport is a confirmed death, not a suspicion, but the
liveness probe treated socket_closed like every soft signal and waited for
a confirming strike (#118487). That strike never arrives: discord.py swaps
in a fresh socket while resuming, the next transport-side sample reads
healthy, and the strike counter silently resets while a resumed-but-deaf
session stays event-starved until the multi-hour event-silence default
elapses. The bot sits deaf until a manual gateway restart.

Escalate the first socket_closed strike straight to the forced-reconnect
path; soft signals (ack staleness, latency, event silence) keep the
confirmation threshold. Also leave a trace on every previously-silent
probe transition: counter resets now log, and probe exits log the flag
state, so an incident log can no longer confuse a healthy sample with a
dead watchdog task.

(cherry picked from commit a9b796e9e9b113595c049701caea46679430465b)
2026-09-22 13:39:00 +05:30
teknium1
0706dffca1 fix(config): route every config.yaml writer through one comment-preserving writer
`hermes_cli.config.atomic_config_write` is now THE config.yaml writer: it delegates to
`utils.atomic_roundtrip_yaml_save` (ruamel round-trip), which merges the new state onto the
on-disk document so user comments, key order, quoting and blank lines survive every write.

Why: config.yaml is hand-edited and commented, and every writer that re-serialised the parsed
dict through PyYAML (`save_config`, `config set/unset`, migrations, plugin bookkeeping, auth
provider reset, credential scrub, channel strip, backup restore, profile seed, telegram topic
persistence) destroyed those comments — and `save_config` re-appended the stock boilerplate on
top (#92554, #63039, #50698, #109611, #107511, #66752). The round-trip writer existed
(tui_gateway only) but nothing else used it, so each new writer regressed the class.

- save_config / _write_user_config / atomic_config_write -> round-trip merge; the commented
  example blocks are appended only when the file is created.
- round-trip merge only reassigns nodes whose value changed (element-wise for lists), so an
  untouched scalar/list keeps its inline comments; YAML 1.1-ambiguous strings (off/yes/no...)
  are force-quoted at every depth; duplicate keys are tolerated like PyYAML.
- direct PyYAML writers in auth.py, credential_lifecycle.py, profile_channels.py, backup.py,
  profiles.py, telegram adapter and tui_gateway/server.py now call atomic_config_write.
2026-09-22 01:07:59 -07:00
teknium1
643dea487b fix(plugins): canonical activation keys + remove/toggle bookkeeping across CLI, dashboard and RPC
Symptoms fixed (all live-reproduced on origin/main in a fake HERMES_HOME):
- `hermes plugins disable photon-platform` wrote `platforms/photon` while the loader keyed the
  bundled adapter `photon-platform`, so the disable never applied (#27548). The loader now keys
  bundled platforms `platforms/<dir>` like every other category (one call site in
  plugins_discovery.py); the manifest name stays an accepted alias through gate_manifest.
- `plugins.manage toggle` / dashboard toggle wrote the raw identifier: enabling by bare leaf or
  manifest name returned ok while a stale canonical key in plugins.disabled kept the plugin off.
  dashboard_set_agent_plugin_enabled resolves the canonical key and purges every alias from the
  opposing list (shared _activate_key, also used by cmd_enable/cmd_disable); the RPC reports the key.
- remove/uninstall (CLI, dashboard, RPC) left the plugin in plugins.enabled/disabled and
  plugins.entries (#54336), left memory.provider dangling so the next agent init re-cloned the
  provider from the catalog (uninstall silently reverted), and, for a symlink inside the plugins
  dir, deleted the TARGET plugin and its metadata while the link stayed dangling. _remove_user_plugin
  is the shared tail: unlink a link only, then forget config under every alias and reset
  memory.provider (reported as cleared_memory_provider).
- `hermes plugins list` / dashboard hub / TUI hub called bundled backends, bundled platforms and the
  live memory.provider "not enabled" (#73131, #82898): _plugin_status mirrors gate_manifest.
- A user-installed memory provider parked in plugins.disabled kept loading: load_memory_provider
  honours the deny-list (name, dir name or manifest name) and says so once.
2026-09-22 00:13:50 -07:00
teknium1
913c045c53 fix: resolve user-installed context engines from $HERMES_HOME/plugins/<name>
plugins/context_engine.load_context_engine scanned only the bundled directory. An engine
dropped into $HERMES_HOME/plugins/<name> with `context.engine: <name>` was reachable only
through the general plugin system, which skips any user plugin not listed in
plugins.enabled — so every agent init logged "Context engine '<name>' not found — falling
back to built-in compressor" although the engine was installed and named in config.

Live probe on base (fake HOME, plugins/ctx_demo with register(ctx), context.engine:
ctx_demo): the warning fired on EVERY init, not only the first; adding the plugin to
plugins.enabled made it load through the general fallback. `context.engine` is the
activation signal (as memory.provider / cron.provider are for their kinds), so the engine
loader now resolves bundled then user dirs the way plugins/cron_providers does: same
`user_plugins_dir()` seam, cheap source heuristic (register_context_engine / ContextEngine),
user engines imported under a synthetic namespace, bundled wins on collision, and
discover_context_engines() lists them for `hermes plugins` / the dashboard.

Fixes #61839
credit: @giggling-ginger #61995
2026-09-21 23:38:16 -07:00
fangliquanflq
72ee40fa68 fix(plugins): remove memory lifecycle hook declarations 2026-09-21 23:26:19 -07:00
fangliquanflq
8ccbe5eb6a fix(plugins): declare bundled plugin hooks 2026-09-21 23:26:19 -07:00
tachyon-r
82d6ab99c4 fix: declare security-guidance hooks in manifest 2026-09-21 23:26:19 -07:00
KoNit-K
de6374194c fix(plugins): clean failed sibling modules 2026-09-21 23:26:19 -07:00
teknium1
77f80d31ee fix(plugins): catalog provenance is installer-owned; kill list covers update/enable/load; re-pin keeps user files
Catalog trust bugs from the 2026-09-21 plugin audit (lane 3, F1-F6, F9):

- F1 (high): a URL-installed repo shipping its own .hermes-catalog.json rendered
  as catalog:official everywhere and marked the real entry "installed". Provenance
  now lives on the installer-owned .install-metadata.json record (catalog block
  written by _install_plugin_core, sha = checked-out commit); read_catalog_sidecar
  never reads the tree. Pre-fix installs are adopted once when the installer
  record agrees (pinned at the sidecar sha, cloned from the entry's repo).
- F2: removed.yaml bypassed by git@/ssh:///http:///www. spellings. _normalize_repo
  canonicalises to host/owner/repo (scheme, user, www., .git, slashes dropped).
- F6: removed.yaml consulted for INSTALLED plugins too: git-pull update, enable and
  gate_manifest (load) refuse recalled plugins, offline (in-tree + cached live
  list). --allow-removed is recorded on the install record and exempts it.
- F5: install NAME --ref X recorded the catalog pin, so list/TUI/update claimed
  the reviewed pin while HEAD differed. Recorded sha is the checked-out one.
- F3: re-pin replaced the whole tree, losing the installer-created config.yaml,
  data and user patches with no warning. Untracked/ignored files are carried
  into the new tree; edits to tracked files are copied to
  <HERMES_HOME>/plugins-backup/<name>-<sha8>/ with a warning.
- F4: a manifest rename between pins left the OLD dir installed and enabled.
  The stale dir is removed and the enabled flag follows the new name.
- F9: dashboard payload `removed` list now includes live removals.

repin_catalog_plugin returns RepinResult(sha, changed, installed_name, warnings);
CLI, dashboard and TUI callers surface the warnings and the new name.
2026-09-21 23:17:05 -07:00
Austin Pickett
992f569fc1 fix(telegram): admit replayed updates before PTB dispatch
Keep one adapter-owned update claim across core, observer and native plugin
handlers. Release failed preparation, but retain queued or externally handed-off
work through cancellation and PTB task completion.

Refs https://github.com/NousResearch/hermes-agent/issues/68502

Co-authored-by: Jasmine Naderi <jasmine@smfworks.com>
2026-09-22 00:55:46 -04:00
teknium1
884b8d980f fix(image_gen): a gateway 429 is reported as a rate limit and retried once, not as a missing model
`_submit_fal_request` (image) and `_submit_fal_video_request` (video plugin)
translated every managed-gateway 4xx into "This model may not yet be enabled
on the Nous Portal's FAL proxy — set FAL_KEY or pick a different model". For
HTTP 429 that remediation is wrong: the gateway body is RATE_LIMIT_EXCEEDED
with a retryAfter, the model is enabled, and agents reading the message
switched models or gave up. On one real install this fired 260 times in a
week (17% of image_generate calls).

Both surfaces now submit through one shared helper
(tools/fal_common.py::submit_managed_fal_with_rate_limit_retry): a 429 whose
Retry-After (header, else body error.retryAfter) fits a 30s cap is waited out
in interrupt-aware 0.5s slices and resubmitted once under a fresh
x-idempotency-key; a second 429, or an unknown/too-long Retry-After, raises
a ValueError that names the rate limit and tells the agent to retry later
rather than switch models. 429 can no longer reach the "may not yet be
enabled" text.
2026-09-21 10:40:03 -07:00
kshitijk4poor
7b660e66ee fix(homeassistant): detect the supervised launch through is_supervised_gateway_launch
gateway.restart already answers "was this gateway launched by a generated
service" (HERMES_SUPERVISED_CHILD or launchd's XPC_SERVICE_NAME, so a plist
that predates the marker still counts); the hint reuses it instead of a
second env read. The negative branch (plain unreachable host without a
supervisor) is now an unmarked test so the Linux lane keeps covering the
function; the macOS-only test holds the positive branch.
2026-09-21 17:16:23 +05:30
kshitijk4poor
eda2432cce fix(gateway): scope the Local Network hint and keep the osascript wrapper out of gateway process scans
Review of the first cut:

- The Home Assistant errno hint matched a bare 65 and two error strings on
  every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
  errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
  genuinely unreachable HA host was told macOS was blocking launchd. Gate on
  darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
  generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
  `gateway run`, so `hermes gateway stop` would have signalled osascript
  alongside the gateway. The canonical matcher now bails on an exact
  argv[0] basename of osascript (the gateway is its child and is matched on
  its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
  carries (osascript's own output) instead of implying it is redundant.
2026-09-21 17:16:23 +05:30
xxxigm
1d871110e1 fix(homeassistant): name the macOS Local Network block behind errno 65 under launchd (#71206)
A launchd-run gateway that macOS Local Network Privacy denies sees every LAN
connect fail with EHOSTUNREACH while the same URL works from Terminal.
Annotate the Home Assistant connect/reconnect log lines with the cause and
the remedy so the failure is actionable instead of a bare "No route to host".

Salvaged from #115196 (remedy text points at the regenerated launchd job).
2026-09-21 17:16:23 +05:30
Bartok9
6504b665ba fix(kanban): refuse empty complete_task evidence (#117483)
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
2026-09-20 19:03:47 -07:00
chelsealong
7efacc8f62 fix(a2a): recover the streamed reply text instead of resolving empty
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).

_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".

(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
2026-09-20 18:45:42 -07:00
Forkbert
b113ab6de6 fix(photon): use GUID for liveness probe message id
(cherry picked from commit cdb6ccadb045e599db795823a100ae7b61d377ec)
2026-09-20 18:23:28 -07:00
teknium1
d325e53116 fix(kanban): drop the dead edit_completed_task_result shim and route dashboard priority edits through edit_task
Review follow-up for #117434: edit_completed_task_result had no callers
after edit_task absorbed it; the dashboard's _set_priority kept its own raw
UPDATE + reprioritized INSERT, so edit_task gains a board= passthrough for
the post-commit observer and becomes the single reprioritize primitive.
2026-09-20 16:33:49 -07:00
teknium1
53815e24dc fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.

The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.

Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
  before  req_reasoning: {}
  after   req_reasoning: {'reasoning_effort': 'medium'}
  agent.reasoning_effort: low  ->  {'reasoning_effort': 'low'}  (unchanged)
2026-09-20 16:03:38 -07:00
teknium1
fd94fe9b93 fix(gateway): every adapter send_voice accepts the dispatch's is_voice kwarg
The base media dispatch passes is_voice= to send_voice; line, mattermost and
weixin had explicit signatures without it, so a non-image MEDIA attachment
routed as audio raised TypeError and was dropped — the same class as the
Matrix report (#102221, #116776). Adds a repo-wide signature invariant test.
2026-09-20 15:48:51 -07:00
xiaodu
b6505448f1 fix(matrix): accept the media dispatch's is_voice kwarg in send_voice
The base media dispatch calls send_voice(..., is_voice=is_voice) for every
audio MEDIA attachment (gateway/platforms/base.py _send_one). MatrixAdapter
.send_voice() accepted neither is_voice nor **kwargs, so every non-image
MEDIA delivery raised TypeError and the file was silently dropped — the
failure is visible in rotated logs since 2026-09-08 (never worked).

Accept the flag explicitly: is_voice=False -> plain m.audio in the original
format (no transcode); True or omitted (play_audio legacy callers) -> the
existing MSC3245 voice-bubble path with best-effort Ogg/Opus transcode.

Fixes #116776

(cherry picked from commit d4f89a725498e29a2ee0fe016f8bb98db08f358c)
2026-09-20 15:48:51 -07:00
fangliquan
45a701d8b0 fix(simplex): flatten alpha image thumbnails
(cherry picked from commit 97886801741f4367aadfe83ccb25b10cedb21149)
2026-09-20 15:48:01 -07:00
liuhao1024
7a23b00101 fix(whatsapp): refuse legacy-pidfile kills without a start-time fingerprint
The stale-bridge cleanup accepted a "node" + session-path cmdline substring
as kill evidence for legacy pidfiles (pid line only). A log tail, editor, or
grep that merely mentions the session path matches that same substring, so
the cleanup could SIGTERM a stranger process.

Require the kernel start-time fingerprint and fail closed when the pidfile
lacks one; the bridge-port scan (which verifies a node-executable listener)
reaps the orphan instead. The refusal reason in the warning now distinguishes
a fingerprint-less legacy pidfile from a recycled PID.

Flip the legacy-pidfile regression test to assert the fail-closed outcome.

Fixes #116883
2026-09-20 15:24:51 -07:00
teknium1
1e4cd9ade6 fix(kanban): diagnostic severity colours follow the dashboard theme
The three `--hermes-diag-*` tokens were literals declared on the consuming
elements, so the theme engine's `<html>`-level custom properties could never
reach them and light presets rendered the warning badge at 1.8:1 contrast.
Chain them through the host tokens themes already set (`--color-warning`,
`--color-destructive`) with the shipped literals as fallbacks; same selector
list, so nodes rendered outside `.hermes-kanban` keep a value. Error and
critical share `--color-destructive` (critical keeps its bold weight).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 12:54:48 -07:00
teknium1
34b475e2ce fix(langfuse): warn once per invalid HERMES_LANGFUSE_MAX_DEPTH value, not per payload
#116713 parsed HERMES_LANGFUSE_MAX_DEPTH inside `_safe_value`, i.e. once per
captured prompt, response, tool input and tool output. With an invalid value
(`abc`) every captured field logged the same "Invalid ... Falling back to 4"
WARNING for the life of the process — one multi-tool turn fills agent.log.

Resolve the depth in `_resolve_max_depth`, an lru_cache keyed on the raw env
string: the warning fires once per distinct bad value, a changed env var is
still picked up by a long-lived process (mirrors `_capture_mode`, which also
reads per call and warns once), and valid values skip the int() parse after
the first call.

Follow-up to #116713 (independent review finding).
2026-09-20 12:46:00 -07:00
teknium1
a57a5f90d2 fix: slot-busy skipped Telegram edits no longer count as shown text (review follow-up)
An interim edit skipped because the chat's shared send+edit slot was busy
returned a plain SendResult(success=True), so the stream consumer recorded
the never-shown text as _last_sent_text and reset _flood_strikes. A later
turn-final flood then saw _visible_prefix() == final text and either marked
the turn delivered or entered fallback with an empty continuation — the user
never saw the tail. The adapter now flags the skip in
raw_response={"skipped": True} and _edit_existing leaves the visible prefix
and flood state untouched, so the next tick retries and a flood fallback
re-sends exactly the unseen tail.
2026-09-20 12:17:23 -07:00
Haisam Abbas
fb2f3e213a fix(telegram): pace sends and edits against a shared per-chat slot
Telegram counts an editMessageText against the same per-chat allowance as a
sendMessage, but streaming previews paced only edits (DEFAULT_STREAMING_EDIT
_INTERVAL = 0.8s = 1.25 msg/s into one chat before any reply was sent) —
83% of measured flood penalties. One shared slot per chat: a send WAITS for
its slot (skipping would drop a message), an interim edit is SKIPPED (the
next tick shows the same text anyway), and the final edit is never gated
(the answer is never withheld). A per-adapter tuning knob keeps the slot
available to tests that model instantaneous bursts.
2026-09-20 12:17:23 -07:00
teknium1
1d6fbd0d1f fix: explicit empty backfill channel list disables the Discord scan (review follow-up)
`missed_message_backfill.channels: []` (YAML list or JSON-list string) returned
an empty set on main, and the backfill logged "no channels configured" and
skipped. The gate-CSV refactor treated an empty result as "unset" and fell
through to the allowed-union-free-response default, scanning channels the
operator had explicitly disabled. An explicit list is now authoritative even
when empty; only the default "" string still falls through to env/default.
2026-09-20 12:10:27 -07:00
teknium1
587bb10575 fix(platforms): Discord/WhatsApp/DingTalk gates honour allowlists stored as a JSON-list string
`hermes config set KEY '["-100","-200"]'` used to write the literal as one quoted
YAML string; the writer now emits a real list (acaac9a18, #88163) and the Telegram
gate decodes the legacy string shape (122ad719, #110213). The Discord, WhatsApp and
DingTalk gate parsers still comma-split that string into `{'["-100"', '"-200"]'}`,
so a config written before the writer fix silently locks every allowlisted chat,
channel or user out — with no warning.

Route every remaining comma-split gate through the shared
`gateway/platforms/_shared.py::decode_json_list_literal`:

- Discord `_gate_csv_set` (allowed/ignored/no-thread channels, allowed users/roles),
  `_discord_free_response_channels` and `_missed_message_backfill_channels` now share
  the one parser instead of three hand-rolled splits.
- WhatsApp `_coerce_allow_list` (allow_from, group_allow_from, free_response_chats).
- DingTalk `_csv_set` (allowed_users, allowed_chats, free_response_chats).

Plain CSV strings, YAML lists and malformed JSON keep their previous meaning.
2026-09-20 12:10:27 -07:00
teknium1
8cc183e1e1 fix(disk-cleanup): only .git entries below HERMES_HOME mark a file git-owned
The cherry-picked ownership check walked every ancestor up to `/`, so a
HERMES_HOME kept inside a dotfiles checkout (`~/.git`) turned every
root-level test_* scratch file into a protected "git-owned" file and
silently disabled the plugin's core contract. Cut the walk at
HERMES_HOME for in-home paths; out-of-home (/tmp/hermes-*) trees keep
the full walk since is_safe_path already bounds them.

Tests trimmed to the two invariants: quick() drops a stale tracked entry
for a committed test inside a linked worktree (.git pointer FILE) instead
of deleting it, and root-level scratch is still deleted even with a .git
above HERMES_HOME. Dropped the contributor's literal /tmp test (the repo
never writes /tmp) and the guess_category-only case the quick() test
already drives.
2026-09-20 12:01:55 -07:00
liuhao1024
b3946bb032 fix(disk-cleanup): never classify test_* files inside git worktrees as disposable
guess_category() matched test_*/tmp_* by basename alone, so a committed
regression test inside a git worktree under $HERMES_HOME/worktrees/ or a
/tmp/hermes-* checkout was tracked and auto-deleted by quick() at session
end (#115295; the protected-top-level-dir half landed in #114770).

Classify such files as non-disposable whenever a .git entry (directory or
linked-worktree pointer file) exists on the directory chain. quick() and
dry_run() already re-validate stored "test" entries through
guess_category(), so stale pre-fix tracked.json entries are dropped from
tracking instead of deleted — no separate migration needed. Scratch
test_* files outside git-owned trees keep aging out as before.

Fixes #115295
2026-09-20 12:01:55 -07:00
Steve Hsu
97051e4194 fix(telegram): attach real video geometry and a thumbnail so large uploads aren't square
`sendVideo` gives back `width=320 height=320 duration=0` with no thumbnail once an
upload is large enough that Telegram skips its own video processing, and clients
then draw the message as a square tile — for portrait reels and 16:9 clips alike,
even though the delivered file itself is correct.

Measured on one 6 s 2560x1440 clip, inspecting the Bot API response: 4.9 MB and
9.8 MB keep `2560x1440` / duration 7 / a 320x180 thumbnail; 14.8 MB, 19.4 MB and
23.0 MB degrade to the square placeholder; the same 23.0 MB file sent with
`width`/`height`/`duration` plus a JPEG `thumbnail` comes back `2560x1440` with a
320x180 thumbnail.

Probe the local file with ffprobe and attach a 320px-wide JPEG frame from it, on
both Telegram send paths: the gateway adapter's `send_video` and the standalone
`hermes send` media sender. Both helpers return nothing when ffmpeg/ffprobe is
unavailable, which keeps the previous behaviour for hosts without them.
2026-09-20 11:20:26 -07:00
teknium1
c060a821d5 fix(discord): accept a pasted channel link at the _resolve_channel chokepoint too
HomeChannel normalization covers env and YAML homes, but cron
`deliver: discord:<link>` targets and thread metadata reach the adapter as
written and still died in int(). Share one helper
(gateway.config.discord_channel_id_from_link) between HomeChannel and the
resolver every outbound Discord target passes through; message links and
non-link strings keep their existing error path. Tests trimmed to two
invariants: config load (env + YAML) and the resolver entry point.
2026-09-20 10:49:49 -07:00
chelsealong
f45ed8e4ec docs(matrix): document the <=2-member DM auto-classification and its config bypass
_resolve_room_identity() classifies any room with <=2 joined members as a DM
regardless of m.direct or an explicit room name, so those rooms silently
bypass MATRIX_ALLOWED_ROOMS, MATRIX_FREE_RESPONSE_ROOMS, and
MATRIX_REQUIRE_MENTION, and use DM threading instead of
MATRIX_AUTO_THREAD/MATRIX_SESSION_SCOPE. This was previously only visible in
an inline code comment, not in the env-var docs an operator would read.

Fixes #114733
2026-09-20 10:42:25 -07:00
chelsealong
0c872b611c fix(telegram): point dm_topics prerequisite text at BotFather Threaded Mode
The dm_topics "not a forum" warning and the matching docs section told
users to tap the bot's name in the DM and toggle "Topics" in chat
settings. That toggle only exists for group forums; a bot DM has no
such control. The actual prerequisite is Threaded Mode, enabled by the
bot owner via the BotFather Mini App (Bot Settings -> Threads
Settings) -- already documented correctly a few sections later in the
same file, under the /topic prerequisites.

Fixes #115019
2026-09-20 10:39:37 -07:00
Tyler Lyon
88d2d90fe1 fix(kanban): tapping a card on touch opens it instead of moving it
attachTouchDrag() armed a drag on ANY touch pointerdown and immediately
called preventDefault(), which suppresses the synthesized click
TaskCard.handleClick relies on to call props.onOpen(). There was no
movement threshold, so a finger drifting even ~2-3px on a normal tap --
which is universal on real touch hardware -- was enough to arm the
drag and swallow the open.

Fix: defer starting the drag proxy and calling preventDefault() until
the pointer has actually moved past an 8px threshold (matches the
common native drag-affordance convention). A stationary tap never
crosses the threshold, dragging is never armed, and the click fires
normally. A real drag still claims the gesture identically to before,
just after the same few pixels of travel every touch drag implementation
already tolerates.

The bundle (plugins/kanban/dashboard/dist/index.js) has no build step --
it is hand-maintained directly, as established by prior kanban dashboard
PRs (#114882, #108694) -- so the fix is applied there.

Closes #115568.

Testing: no jsdom/vitest harness exists for this bundle (confirmed by
PR #114882's review follow-up, which explicitly rejected turning a
"live-repro jsdom harness" into a pytest because jsdom/react aren't
declared in the root package.json and the Python CI job has no
node_modules -- such a test would be vacuous in CI). Per that
precedent and the "never read source code in tests" rule (no
regex/substring pin on the bundle text), this PR instead extracts
attachTouchDrag() verbatim at test time via Node (already present:
tests-js/ + vitest are in the repo) and drives it through real
pointerdown/pointermove/pointerup sequences against a minimal DOM
stub -- a behavioral test, not a source-shape test. Proven red on the
unfixed bundle (asserts preventDefault is called on a stationary tap)
and green on the fix; skips cleanly via shutil.which("node") if Node
is unavailable in a given lane.

Verification:
- node tests/plugins/fixtures/kanban_touch_drag_probe.js against the
  ORIGINAL (unfixed) bundle: fails with "FAIL: a stationary tap called
  preventDefault (suppresses the click)", exit 1 -- confirms the probe
  reproduces the reported bug
- Same probe against the fixed bundle: "PASS", exit 0
- scripts/run_tests.sh tests/plugins/test_kanban_dashboard_plugin.py --
  42/42 passed (1 new, 41 unchanged)
- node --check plugins/kanban/dashboard/dist/index.js -- syntax OK
2026-09-20 10:30:47 -07:00
KeyArgo
e7509fa1a4 fix(telegram): observe sibling-bot wake-word messages dropped by the bot-to-bot gate
With `require_mention: true` + `bots_require_mention: true` + wake words in `mention_patterns`, a message
authored by another Hermes bot that addressed this bot by name matched no dispatch path (the
bot-to-bot loop breaker skips it) and was also refused by the observe gate (`mention_patterns`
matches are assumed dispatched), so it vanished from both paths with no log line.

Factor the loop-breaker predicate into `_bot_sender_suppressed` and consult it in
`_should_observe_unmentioned_group_message`, so every message the dispatcher drops because of
`bots_require_mention` is kept as observed context instead of being lost (#115119).
2026-09-20 10:30:11 -07:00
fangliquan
3cbd3cce0a fix(whatsapp): honor configured reply prefix in bridge replies
whatsapp.reply_prefix from config.yaml was written into the bridge env and then
popped again by the WHATSAPP_* passthrough loop (the key was in
_BRIDGE_PASSTHROUGH_ENV and the scoped env lookup came back empty), so bridge.js
always fell back to its built-in header and the documented reply_prefix: ""
could not disable it. Resolve the prefix once (scoped env first, then the
adapter value) and keep it out of the passthrough loop.

Fixes #116059
2026-09-20 10:20:40 -07:00
hardwork9047
a0e7fc4a9e fix(gateway): consult _in_bot_thread in the Discord admission gate
`_discord_message_admission()` drops a message that mentions someone other than
the bot when `DISCORD_IGNORE_NO_MENTION` is on (the default) and the channel is
not free-response. It did so without consulting `_in_bot_thread()`, unlike the
other two ingress paths — `_dispatch_recovered_message()` (adapter.py:2292) and
`_handle_message()` (adapter.py:5946). Admission runs on both and returns
`False` unconditionally, so it overrode the thread exemption they grant.

The result was an asymmetry with no obvious cause from the outside: in a thread
the bot had joined, a message with no mention at all was admitted (an empty
`message.mentions` skips the enclosing block), while the same message with one
mention of a third party was dropped. A thread the bot is a participant in is
the one place "addressed to someone else" is least likely to hold.

`thread_require_mention` still gates multi-bot threads, since that check lives
inside `_in_bot_thread()`.

The drop also emitted nothing at any log level, leaving `gateway.log` identical
whether the gate fired or the event never arrived; add a debug line so the two
can be told apart.

Fixes #116568
2026-09-20 10:20:03 -07:00
teknium1
21e96c9593 docs(kanban): document done-column completion order and the completed-desc sort
Follow-up to the salvaged #116041: the board's `done` column is now newest-
completed-first and `hermes kanban list --sort completed-desc` exists, so the
feature doc and the CLI usage block say so; the stale "per-column ordering
comes from list_tasks" comment in `get_board` now describes the queue
columns only (wording from #116051).

Co-authored-by: MohamadKanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-20 10:18:14 -07:00
liuhao1024
71c614b6e3 fix(kanban): order the board's done column by completion time
get_board() buckets one list_tasks() fetch, so the done column
inherited the shared priority DESC, created_at ASC order — creation
order, which says nothing about when work finished. Sort the done
bucket newest-completed-first (completed_at DESC NULLS LAST, id DESC)
and expose that as a completed-desc list_tasks sort key; queue lanes
keep the FIFO dispatch default.
2026-09-20 10:18:14 -07:00
kshitijk4poor
de083436bd refactor(gateway): drop Telegram's shadowing _coerce_float_extra; module-level math import
Gate review: TelegramAdapter kept a same-named override with a weaker contract (no finite
guard, negatives clamped instead of reset), so the hierarchy had two parsers under one name.
Both Telegram callers pass explicit bounds; the base method is a strict superset for them.
The cadence test now pins the behaviour (≤ 0.5 s / ≤ 1.0 s) instead of echoing the constant.
2026-09-20 18:14:58 +05:30
kshitijk4poor
530fdb5867 refactor(gateway): one text-batch cadence and one extra-float parser on BasePlatformAdapter
Gate review: the fix left three adapter-private copies of `_coerce_float_extra` and the
0.3/2.0/1.0/4.0 cadence literals in three files. The parser and the cadence constants now
live on BasePlatformAdapter beside the delay attrs they configure; WhatsApp and Weixin call
`_configure_text_batch_delays()`, Telegram reads the same constants through its env helper.
The clamp test is parametrized over both adapters and the Weixin docs name the ceilings.
2026-09-20 18:14:58 +05:30
kshitijk4poor
130b596f6f fix(whatsapp,weixin): text-batch delays default to Telegram cadence
WhatsApp debounced text for 5s (10s near a split) and Weixin for 3s/5s
before dispatching, so every reply paid multiple seconds of idle latency
that Telegram never pays (0.3s/1.0s). Default both adapters to Telegram's
cadence and mirror its ceilings (2.0s / 4.0s, split >= base delay) via the
existing _coerce_float_extra seam. The config keys are unchanged; 0 still
dispatches immediately. Docs updated.

Spotted via #44896 (@liuhao1024). Fixes #44883, refs #25056.
2026-09-20 18:14:58 +05:30
kshitijk4poor
18f1706137 refactor(matrix): drop dead is_direct guard and unreachable sender warning
_schedule_invite_join already requires both is_direct and inviter before recording m.direct, so `is_direct and bool(inviter)` at the reconcile call site was redundant; pass is_direct through. The `if is_direct and not inviter` WARNING could only be reached with GATEWAY_ALLOW_ALL_USERS set and a spec-violating stripped m.room.member event lacking `sender`; the info log already prints is_direct, so drop the branch.
2026-09-20 18:00:06 +05:30
kshitijk4poor
b98dd99142 docs(matrix): reconcile gate comment states the real premise
The adapter comment and the test module docstring claimed a reconciled pending invite "never fires _on_invite". It does: _absorb_sync runs _dispatch_sync (which emits INVITE to _on_invite) and then the reconcile pass over rooms.invite, which joined every entry unconditionally — so a live invite _on_invite rejected was joined ms later, and invites that arrived while the gateway was down were joined on restart with no gate. Reword both to state that premise.
2026-09-20 18:00:06 +05:30
Iain Lane
439f2b1cf6 fix(matrix): apply the inviter allowlist to reconciled pending invites
_on_invite only auto-joins a room when the inviter is allow-listed (or
GATEWAY_ALLOW_ALL_USERS is set), so a live invite from an arbitrary
federated user is rejected. A pending invite that arrives while the
gateway is down takes a different path: _schedule_pending_invite_joins
reconciles it from rooms.invite in the sync response and scheduled the
join unconditionally. An unauthorized invite sent during downtime was
therefore auto-joined on restart, bypassing the allowlist.

Extract the gate from _on_invite into _is_authorized_inviter and apply
it during reconciliation too, reading the inviter from the stripped
invite state (the sender of the m.room.member event for our own user,
as _extract_invite_dm_signal already does for the DM signal). An
inviter that cannot be read from the invite state fails closed, exactly
like an empty sender in _on_invite: the invite is skipped with a
warning and left pending.
2026-09-20 18:00:06 +05:30