Commit Graph

2522 Commits

Author SHA1 Message Date
ethernet
8e80f37aad fix(meet): make inherited interactive stdin explicit 2026-09-21 14:36:47 -04:00
ethernet
c69ccc3bf1 merge: refresh upstream models, notification expiry, and desktop controls 2026-09-21 14:01:47 -04:00
teknium1
884b8d980f fix(image_gen): a gateway 429 is reported as a rate limit and retried once, not as a missing model
`_submit_fal_request` (image) and `_submit_fal_video_request` (video plugin)
translated every managed-gateway 4xx into "This model may not yet be enabled
on the Nous Portal's FAL proxy — set FAL_KEY or pick a different model". For
HTTP 429 that remediation is wrong: the gateway body is RATE_LIMIT_EXCEEDED
with a retryAfter, the model is enabled, and agents reading the message
switched models or gave up. On one real install this fired 260 times in a
week (17% of image_generate calls).

Both surfaces now submit through one shared helper
(tools/fal_common.py::submit_managed_fal_with_rate_limit_retry): a 429 whose
Retry-After (header, else body error.retryAfter) fits a 30s cap is waited out
in interrupt-aware 0.5s slices and resubmitted once under a fresh
x-idempotency-key; a second 429, or an unknown/too-long Retry-After, raises
a ValueError that names the rate limit and tells the agent to retry later
rather than switch models. 429 can no longer reach the "may not yet be
enabled" text.
2026-09-21 10:40:03 -07:00
ethernet
9d036f8ac9 merge origin/main (15 commits) into ethie/pm-clean
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
2026-09-21 08:16:57 -04:00
kshitijk4poor
7b660e66ee fix(homeassistant): detect the supervised launch through is_supervised_gateway_launch
gateway.restart already answers "was this gateway launched by a generated
service" (HERMES_SUPERVISED_CHILD or launchd's XPC_SERVICE_NAME, so a plist
that predates the marker still counts); the hint reuses it instead of a
second env read. The negative branch (plain unreachable host without a
supervisor) is now an unmarked test so the Linux lane keeps covering the
function; the macOS-only test holds the positive branch.
2026-09-21 17:16:23 +05:30
kshitijk4poor
eda2432cce fix(gateway): scope the Local Network hint and keep the osascript wrapper out of gateway process scans
Review of the first cut:

- The Home Assistant errno hint matched a bare 65 and two error strings on
  every platform. 65 is ENOPKG on Linux and "No route to host" is Linux's
  errno 113, so a systemd gateway (or a Terminal-run gateway on macOS) with a
  genuinely unreachable HA host was told macOS was blocking launchd. Gate on
  darwin + errno.EHOSTUNREACH + the HERMES_SUPERVISED_CHILD marker the
  generated plist already sets; no new env var.
- With a `"` in the home path the wrapper's own ps line tokenized as
  `gateway run`, so `hermes gateway stop` would have signalled osascript
  alongside the gateway. The canonical matcher now bails on an exact
  argv[0] basename of osascript (the gateway is its child and is matched on
  its own command line). Asserted in the existing hostile-path test.
- Log paths spelled once; docstring now says what StandardOutPath still
  carries (osascript's own output) instead of implying it is redundant.
2026-09-21 17:16:23 +05:30
xxxigm
1d871110e1 fix(homeassistant): name the macOS Local Network block behind errno 65 under launchd (#71206)
A launchd-run gateway that macOS Local Network Privacy denies sees every LAN
connect fail with EHOSTUNREACH while the same URL works from Terminal.
Annotate the Home Assistant connect/reconnect log lines with the cause and
the remedy so the failure is actionable instead of a bare "No route to host".

Salvaged from #115196 (remedy text points at the regenerated launchd job).
2026-09-21 17:16:23 +05:30
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
Bartok9
6504b665ba fix(kanban): refuse empty complete_task evidence (#117483)
Completing a card with no result/summary (or whitespace-only) left
done rows with no handover. Gate before the write txn, audit
completion_blocked_empty_result, raise EmptyCompletionError.
Review approvals stay exempt.
2026-09-20 19:03:47 -07:00
chelsealong
7efacc8f62 fix(a2a): recover the streamed reply text instead of resolving empty
When streaming already delivered the final reply, gateway/run_turn.py's
_hmwa_deliver_turn_response suppresses the normal adapter.send() and returns
None, so A2AAdapter.send() — the only path that ever carries reply text —
never runs. on_processing_complete() then resolves the pending A2A task
future through its SUCCESS default, which was hardcoded to "", so every
streamed A2A reply lands as TASK_STATE_COMPLETED with no status.message and
no artifacts (#116944).

_hmwa_deliver_turn_response already stashes the true final text on
event._streamed_final_response for exactly this situation (the same stash
_final_text_for_post_turn_hooks reads for /goal and /loop). Read it as the
SUCCESS-path fallback text instead of "".

(cherry picked from commit 638041af046ab149a356a7e5107d52c2e6a3c9a8)
2026-09-20 18:45:42 -07:00
Forkbert
b113ab6de6 fix(photon): use GUID for liveness probe message id
(cherry picked from commit cdb6ccadb045e599db795823a100ae7b61d377ec)
2026-09-20 18:23:28 -07:00
teknium1
d325e53116 fix(kanban): drop the dead edit_completed_task_result shim and route dashboard priority edits through edit_task
Review follow-up for #117434: edit_completed_task_result had no callers
after edit_task absorbed it; the dashboard's _set_priority kept its own raw
UPDATE + reprioritized INSERT, so edit_task gains a board= passthrough for
the post-commit observer and becomes the single reprioritize primitive.
2026-09-20 16:33:49 -07:00
teknium1
53815e24dc fix: send reasoning_effort=medium on custom endpoints when agent.reasoning_effort is unset
An unset agent.reasoning_effort already resolves to medium on the Nous Portal,
OpenRouter, AI Gateway and Copilot routes (each profile fills it in
build_api_kwargs_extras). The custom / OpenAI-compatible profile — every
`providers.<name>` block and `--provider custom` — omitted the field instead,
so the endpoint's own default applied; for moonshotai/kimi-k3 that is `max`:
3x the reasoning tokens and ~3x the latency of medium, measured live.

The default is resolved at request time in _reasoning_config_for_wire via
ProviderProfile.default_reasoning_config (the custom profile answers medium),
so it is recorded as what actually went out and the reasoning-rejection
ladder keeps working: a 400 on the field turns the rest of the session back
to "omit". It never touches an explicit effort (low stays low, none stays
none), stays off non chat-completions transports (the Anthropic adapter's
unset = no thinking kwargs stands), off models the catalog or model_overrides
mark supports_reasoning: false, and off local Ollama models pulled without
the thinking capability. Auxiliary calls are untouched: they hand the profile
reasoning_config=None directly, which still omits the field.

Live wire capture (token-injecting proxy, providers.probe -> kimi-k3):
  before  req_reasoning: {}
  after   req_reasoning: {'reasoning_effort': 'medium'}
  agent.reasoning_effort: low  ->  {'reasoning_effort': 'low'}  (unchanged)
2026-09-20 16:03:38 -07:00
teknium1
fd94fe9b93 fix(gateway): every adapter send_voice accepts the dispatch's is_voice kwarg
The base media dispatch passes is_voice= to send_voice; line, mattermost and
weixin had explicit signatures without it, so a non-image MEDIA attachment
routed as audio raised TypeError and was dropped — the same class as the
Matrix report (#102221, #116776). Adds a repo-wide signature invariant test.
2026-09-20 15:48:51 -07:00
xiaodu
b6505448f1 fix(matrix): accept the media dispatch's is_voice kwarg in send_voice
The base media dispatch calls send_voice(..., is_voice=is_voice) for every
audio MEDIA attachment (gateway/platforms/base.py _send_one). MatrixAdapter
.send_voice() accepted neither is_voice nor **kwargs, so every non-image
MEDIA delivery raised TypeError and the file was silently dropped — the
failure is visible in rotated logs since 2026-09-08 (never worked).

Accept the flag explicitly: is_voice=False -> plain m.audio in the original
format (no transcode); True or omitted (play_audio legacy callers) -> the
existing MSC3245 voice-bubble path with best-effort Ogg/Opus transcode.

Fixes #116776

(cherry picked from commit d4f89a725498e29a2ee0fe016f8bb98db08f358c)
2026-09-20 15:48:51 -07:00
fangliquan
45a701d8b0 fix(simplex): flatten alpha image thumbnails
(cherry picked from commit 97886801741f4367aadfe83ccb25b10cedb21149)
2026-09-20 15:48:01 -07:00
liuhao1024
7a23b00101 fix(whatsapp): refuse legacy-pidfile kills without a start-time fingerprint
The stale-bridge cleanup accepted a "node" + session-path cmdline substring
as kill evidence for legacy pidfiles (pid line only). A log tail, editor, or
grep that merely mentions the session path matches that same substring, so
the cleanup could SIGTERM a stranger process.

Require the kernel start-time fingerprint and fail closed when the pidfile
lacks one; the bridge-port scan (which verifies a node-executable listener)
reaps the orphan instead. The refusal reason in the warning now distinguishes
a fingerprint-less legacy pidfile from a recycled PID.

Flip the legacy-pidfile regression test to assert the fail-closed outcome.

Fixes #116883
2026-09-20 15:24:51 -07:00
teknium1
1e4cd9ade6 fix(kanban): diagnostic severity colours follow the dashboard theme
The three `--hermes-diag-*` tokens were literals declared on the consuming
elements, so the theme engine's `<html>`-level custom properties could never
reach them and light presets rendered the warning badge at 1.8:1 contrast.
Chain them through the host tokens themes already set (`--color-warning`,
`--color-destructive`) with the shipped literals as fallbacks; same selector
list, so nodes rendered outside `.hermes-kanban` keep a value. Error and
critical share `--color-destructive` (critical keeps its bold weight).

Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-20 12:54:48 -07:00
teknium1
34b475e2ce fix(langfuse): warn once per invalid HERMES_LANGFUSE_MAX_DEPTH value, not per payload
#116713 parsed HERMES_LANGFUSE_MAX_DEPTH inside `_safe_value`, i.e. once per
captured prompt, response, tool input and tool output. With an invalid value
(`abc`) every captured field logged the same "Invalid ... Falling back to 4"
WARNING for the life of the process — one multi-tool turn fills agent.log.

Resolve the depth in `_resolve_max_depth`, an lru_cache keyed on the raw env
string: the warning fires once per distinct bad value, a changed env var is
still picked up by a long-lived process (mirrors `_capture_mode`, which also
reads per call and warns once), and valid values skip the int() parse after
the first call.

Follow-up to #116713 (independent review finding).
2026-09-20 12:46:00 -07:00
teknium1
a57a5f90d2 fix: slot-busy skipped Telegram edits no longer count as shown text (review follow-up)
An interim edit skipped because the chat's shared send+edit slot was busy
returned a plain SendResult(success=True), so the stream consumer recorded
the never-shown text as _last_sent_text and reset _flood_strikes. A later
turn-final flood then saw _visible_prefix() == final text and either marked
the turn delivered or entered fallback with an empty continuation — the user
never saw the tail. The adapter now flags the skip in
raw_response={"skipped": True} and _edit_existing leaves the visible prefix
and flood state untouched, so the next tick retries and a flood fallback
re-sends exactly the unseen tail.
2026-09-20 12:17:23 -07:00
Haisam Abbas
fb2f3e213a fix(telegram): pace sends and edits against a shared per-chat slot
Telegram counts an editMessageText against the same per-chat allowance as a
sendMessage, but streaming previews paced only edits (DEFAULT_STREAMING_EDIT
_INTERVAL = 0.8s = 1.25 msg/s into one chat before any reply was sent) —
83% of measured flood penalties. One shared slot per chat: a send WAITS for
its slot (skipping would drop a message), an interim edit is SKIPPED (the
next tick shows the same text anyway), and the final edit is never gated
(the answer is never withheld). A per-adapter tuning knob keeps the slot
available to tests that model instantaneous bursts.
2026-09-20 12:17:23 -07:00
teknium1
1d6fbd0d1f fix: explicit empty backfill channel list disables the Discord scan (review follow-up)
`missed_message_backfill.channels: []` (YAML list or JSON-list string) returned
an empty set on main, and the backfill logged "no channels configured" and
skipped. The gate-CSV refactor treated an empty result as "unset" and fell
through to the allowed-union-free-response default, scanning channels the
operator had explicitly disabled. An explicit list is now authoritative even
when empty; only the default "" string still falls through to env/default.
2026-09-20 12:10:27 -07:00
teknium1
587bb10575 fix(platforms): Discord/WhatsApp/DingTalk gates honour allowlists stored as a JSON-list string
`hermes config set KEY '["-100","-200"]'` used to write the literal as one quoted
YAML string; the writer now emits a real list (acaac9a18, #88163) and the Telegram
gate decodes the legacy string shape (122ad719, #110213). The Discord, WhatsApp and
DingTalk gate parsers still comma-split that string into `{'["-100"', '"-200"]'}`,
so a config written before the writer fix silently locks every allowlisted chat,
channel or user out — with no warning.

Route every remaining comma-split gate through the shared
`gateway/platforms/_shared.py::decode_json_list_literal`:

- Discord `_gate_csv_set` (allowed/ignored/no-thread channels, allowed users/roles),
  `_discord_free_response_channels` and `_missed_message_backfill_channels` now share
  the one parser instead of three hand-rolled splits.
- WhatsApp `_coerce_allow_list` (allow_from, group_allow_from, free_response_chats).
- DingTalk `_csv_set` (allowed_users, allowed_chats, free_response_chats).

Plain CSV strings, YAML lists and malformed JSON keep their previous meaning.
2026-09-20 12:10:27 -07:00
teknium1
8cc183e1e1 fix(disk-cleanup): only .git entries below HERMES_HOME mark a file git-owned
The cherry-picked ownership check walked every ancestor up to `/`, so a
HERMES_HOME kept inside a dotfiles checkout (`~/.git`) turned every
root-level test_* scratch file into a protected "git-owned" file and
silently disabled the plugin's core contract. Cut the walk at
HERMES_HOME for in-home paths; out-of-home (/tmp/hermes-*) trees keep
the full walk since is_safe_path already bounds them.

Tests trimmed to the two invariants: quick() drops a stale tracked entry
for a committed test inside a linked worktree (.git pointer FILE) instead
of deleting it, and root-level scratch is still deleted even with a .git
above HERMES_HOME. Dropped the contributor's literal /tmp test (the repo
never writes /tmp) and the guess_category-only case the quick() test
already drives.
2026-09-20 12:01:55 -07:00
liuhao1024
b3946bb032 fix(disk-cleanup): never classify test_* files inside git worktrees as disposable
guess_category() matched test_*/tmp_* by basename alone, so a committed
regression test inside a git worktree under $HERMES_HOME/worktrees/ or a
/tmp/hermes-* checkout was tracked and auto-deleted by quick() at session
end (#115295; the protected-top-level-dir half landed in #114770).

Classify such files as non-disposable whenever a .git entry (directory or
linked-worktree pointer file) exists on the directory chain. quick() and
dry_run() already re-validate stored "test" entries through
guess_category(), so stale pre-fix tracked.json entries are dropped from
tracking instead of deleted — no separate migration needed. Scratch
test_* files outside git-owned trees keep aging out as before.

Fixes #115295
2026-09-20 12:01:55 -07:00
Steve Hsu
97051e4194 fix(telegram): attach real video geometry and a thumbnail so large uploads aren't square
`sendVideo` gives back `width=320 height=320 duration=0` with no thumbnail once an
upload is large enough that Telegram skips its own video processing, and clients
then draw the message as a square tile — for portrait reels and 16:9 clips alike,
even though the delivered file itself is correct.

Measured on one 6 s 2560x1440 clip, inspecting the Bot API response: 4.9 MB and
9.8 MB keep `2560x1440` / duration 7 / a 320x180 thumbnail; 14.8 MB, 19.4 MB and
23.0 MB degrade to the square placeholder; the same 23.0 MB file sent with
`width`/`height`/`duration` plus a JPEG `thumbnail` comes back `2560x1440` with a
320x180 thumbnail.

Probe the local file with ffprobe and attach a 320px-wide JPEG frame from it, on
both Telegram send paths: the gateway adapter's `send_video` and the standalone
`hermes send` media sender. Both helpers return nothing when ffmpeg/ffprobe is
unavailable, which keeps the previous behaviour for hosts without them.
2026-09-20 11:20:26 -07:00
teknium1
c060a821d5 fix(discord): accept a pasted channel link at the _resolve_channel chokepoint too
HomeChannel normalization covers env and YAML homes, but cron
`deliver: discord:<link>` targets and thread metadata reach the adapter as
written and still died in int(). Share one helper
(gateway.config.discord_channel_id_from_link) between HomeChannel and the
resolver every outbound Discord target passes through; message links and
non-link strings keep their existing error path. Tests trimmed to two
invariants: config load (env + YAML) and the resolver entry point.
2026-09-20 10:49:49 -07:00
chelsealong
f45ed8e4ec docs(matrix): document the <=2-member DM auto-classification and its config bypass
_resolve_room_identity() classifies any room with <=2 joined members as a DM
regardless of m.direct or an explicit room name, so those rooms silently
bypass MATRIX_ALLOWED_ROOMS, MATRIX_FREE_RESPONSE_ROOMS, and
MATRIX_REQUIRE_MENTION, and use DM threading instead of
MATRIX_AUTO_THREAD/MATRIX_SESSION_SCOPE. This was previously only visible in
an inline code comment, not in the env-var docs an operator would read.

Fixes #114733
2026-09-20 10:42:25 -07:00
chelsealong
0c872b611c fix(telegram): point dm_topics prerequisite text at BotFather Threaded Mode
The dm_topics "not a forum" warning and the matching docs section told
users to tap the bot's name in the DM and toggle "Topics" in chat
settings. That toggle only exists for group forums; a bot DM has no
such control. The actual prerequisite is Threaded Mode, enabled by the
bot owner via the BotFather Mini App (Bot Settings -> Threads
Settings) -- already documented correctly a few sections later in the
same file, under the /topic prerequisites.

Fixes #115019
2026-09-20 10:39:37 -07:00
Tyler Lyon
88d2d90fe1 fix(kanban): tapping a card on touch opens it instead of moving it
attachTouchDrag() armed a drag on ANY touch pointerdown and immediately
called preventDefault(), which suppresses the synthesized click
TaskCard.handleClick relies on to call props.onOpen(). There was no
movement threshold, so a finger drifting even ~2-3px on a normal tap --
which is universal on real touch hardware -- was enough to arm the
drag and swallow the open.

Fix: defer starting the drag proxy and calling preventDefault() until
the pointer has actually moved past an 8px threshold (matches the
common native drag-affordance convention). A stationary tap never
crosses the threshold, dragging is never armed, and the click fires
normally. A real drag still claims the gesture identically to before,
just after the same few pixels of travel every touch drag implementation
already tolerates.

The bundle (plugins/kanban/dashboard/dist/index.js) has no build step --
it is hand-maintained directly, as established by prior kanban dashboard
PRs (#114882, #108694) -- so the fix is applied there.

Closes #115568.

Testing: no jsdom/vitest harness exists for this bundle (confirmed by
PR #114882's review follow-up, which explicitly rejected turning a
"live-repro jsdom harness" into a pytest because jsdom/react aren't
declared in the root package.json and the Python CI job has no
node_modules -- such a test would be vacuous in CI). Per that
precedent and the "never read source code in tests" rule (no
regex/substring pin on the bundle text), this PR instead extracts
attachTouchDrag() verbatim at test time via Node (already present:
tests-js/ + vitest are in the repo) and drives it through real
pointerdown/pointermove/pointerup sequences against a minimal DOM
stub -- a behavioral test, not a source-shape test. Proven red on the
unfixed bundle (asserts preventDefault is called on a stationary tap)
and green on the fix; skips cleanly via shutil.which("node") if Node
is unavailable in a given lane.

Verification:
- node tests/plugins/fixtures/kanban_touch_drag_probe.js against the
  ORIGINAL (unfixed) bundle: fails with "FAIL: a stationary tap called
  preventDefault (suppresses the click)", exit 1 -- confirms the probe
  reproduces the reported bug
- Same probe against the fixed bundle: "PASS", exit 0
- scripts/run_tests.sh tests/plugins/test_kanban_dashboard_plugin.py --
  42/42 passed (1 new, 41 unchanged)
- node --check plugins/kanban/dashboard/dist/index.js -- syntax OK
2026-09-20 10:30:47 -07:00
KeyArgo
e7509fa1a4 fix(telegram): observe sibling-bot wake-word messages dropped by the bot-to-bot gate
With `require_mention: true` + `bots_require_mention: true` + wake words in `mention_patterns`, a message
authored by another Hermes bot that addressed this bot by name matched no dispatch path (the
bot-to-bot loop breaker skips it) and was also refused by the observe gate (`mention_patterns`
matches are assumed dispatched), so it vanished from both paths with no log line.

Factor the loop-breaker predicate into `_bot_sender_suppressed` and consult it in
`_should_observe_unmentioned_group_message`, so every message the dispatcher drops because of
`bots_require_mention` is kept as observed context instead of being lost (#115119).
2026-09-20 10:30:11 -07:00
fangliquan
3cbd3cce0a fix(whatsapp): honor configured reply prefix in bridge replies
whatsapp.reply_prefix from config.yaml was written into the bridge env and then
popped again by the WHATSAPP_* passthrough loop (the key was in
_BRIDGE_PASSTHROUGH_ENV and the scoped env lookup came back empty), so bridge.js
always fell back to its built-in header and the documented reply_prefix: ""
could not disable it. Resolve the prefix once (scoped env first, then the
adapter value) and keep it out of the passthrough loop.

Fixes #116059
2026-09-20 10:20:40 -07:00
hardwork9047
a0e7fc4a9e fix(gateway): consult _in_bot_thread in the Discord admission gate
`_discord_message_admission()` drops a message that mentions someone other than
the bot when `DISCORD_IGNORE_NO_MENTION` is on (the default) and the channel is
not free-response. It did so without consulting `_in_bot_thread()`, unlike the
other two ingress paths — `_dispatch_recovered_message()` (adapter.py:2292) and
`_handle_message()` (adapter.py:5946). Admission runs on both and returns
`False` unconditionally, so it overrode the thread exemption they grant.

The result was an asymmetry with no obvious cause from the outside: in a thread
the bot had joined, a message with no mention at all was admitted (an empty
`message.mentions` skips the enclosing block), while the same message with one
mention of a third party was dropped. A thread the bot is a participant in is
the one place "addressed to someone else" is least likely to hold.

`thread_require_mention` still gates multi-bot threads, since that check lives
inside `_in_bot_thread()`.

The drop also emitted nothing at any log level, leaving `gateway.log` identical
whether the gate fired or the event never arrived; add a debug line so the two
can be told apart.

Fixes #116568
2026-09-20 10:20:03 -07:00
teknium1
21e96c9593 docs(kanban): document done-column completion order and the completed-desc sort
Follow-up to the salvaged #116041: the board's `done` column is now newest-
completed-first and `hermes kanban list --sort completed-desc` exists, so the
feature doc and the CLI usage block say so; the stale "per-column ordering
comes from list_tasks" comment in `get_board` now describes the queue
columns only (wording from #116051).

Co-authored-by: MohamadKanso <91088196+MohamadKanso@users.noreply.github.com>
2026-09-20 10:18:14 -07:00
liuhao1024
71c614b6e3 fix(kanban): order the board's done column by completion time
get_board() buckets one list_tasks() fetch, so the done column
inherited the shared priority DESC, created_at ASC order — creation
order, which says nothing about when work finished. Sort the done
bucket newest-completed-first (completed_at DESC NULLS LAST, id DESC)
and expose that as a completed-desc list_tasks sort key; queue lanes
keep the FIFO dispatch default.
2026-09-20 10:18:14 -07:00
ethernet
925c08ceca fix: CI python-tests backlog — no import-time dependency syncs, CI-shaped test fixtures
Production:
- agent/bedrock_adapter.py, agent/vertex_adapter.py: pm.ensure_import ran at
  module import. In any process that imports these modules without a committed
  PM selection (CI's build_environment test venv, a fresh checkout) that sync
  rebuilt the dependency environment mid-process and replaced sys.path with a
  generation missing the caller's own packages (anthropic, aiohttp vanished).
  The extra is now ensured at first client build / credential request.
- plugins/platforms/matrix/adapter.py: a complete install needs no
  ensure_and_bind round trip; only a partial one syncs.
- tools/browser_tool.py: drop the facade's duplicate warm_agent_browser_npx_cache
  shim; the compat pointer already resolves to browser_tool_install.

Test harness:
- tests/home_io_guard.py: PATH-entry probes (shutil.which) and the running
  interpreter's own installation (stdlib reads, realpath ancestry, fixture
  symlinks into it) are not Hermes state; a patched Path.expanduser must not
  crash the guard. run_tests.sh no longer filters PATH — the guard owns it.
- tests/tui_gateway/conftest.py: import hermes_bootstrap before any file opens
  a MagicMock hermes_constants window (6 files exited the process at boot).
- tests/hermes_cli/conftest.py probe_root: scratch checkouts the import guard
  probes need hermes_bootstrap.py (the launcher imports it).
- tests/pm/_fixtures.py stage_host_python: a copied relocatable python needs
  its stdlib beside it (No module named 'encodings' on CI).
- tests/install/e2e-assets/smoke-env.mjs: dependency-free env shaping so the
  source-build-env probe runs under bare node (main deleted the Playwright
  entry it was imported through).
- adapt main's new tests to branch seams (model_metadata_http, launch
  completion tail, CI toolchain exports uv after python, source_launch
  hermes_cli stub, systemd_notify single marker).
2026-09-20 11:47:06 -04:00
ethernet
9f837d298b Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Conflicts resolved toward the PM model: main's lazy_deps/update_cmd_deps/npm
stamp machinery stays deleted (PM + scripts/build/node-deps.mjs own it), the
systemd ExecStop stop-mark rides the installation launcher, legacy
linux_only/macos_only/windows_only markers are rewritten to platforms(), and
finalize_update_receipt carries pending manual-serve obligations forward
again (lost when the ContextVar receipt rewrite crossed c0aa3ce354).

Test harness: the real-home I/O guard exempts /proc/<pid>/fd metadata reads
(deleted-WAL holder scans) and run_tests.sh drops ~/.hermes PATH entries so
shutil.which() cannot trip the tripwire.
2026-09-20 10:07:50 -04:00
kshitijk4poor
de083436bd refactor(gateway): drop Telegram's shadowing _coerce_float_extra; module-level math import
Gate review: TelegramAdapter kept a same-named override with a weaker contract (no finite
guard, negatives clamped instead of reset), so the hierarchy had two parsers under one name.
Both Telegram callers pass explicit bounds; the base method is a strict superset for them.
The cadence test now pins the behaviour (≤ 0.5 s / ≤ 1.0 s) instead of echoing the constant.
2026-09-20 18:14:58 +05:30
kshitijk4poor
530fdb5867 refactor(gateway): one text-batch cadence and one extra-float parser on BasePlatformAdapter
Gate review: the fix left three adapter-private copies of `_coerce_float_extra` and the
0.3/2.0/1.0/4.0 cadence literals in three files. The parser and the cadence constants now
live on BasePlatformAdapter beside the delay attrs they configure; WhatsApp and Weixin call
`_configure_text_batch_delays()`, Telegram reads the same constants through its env helper.
The clamp test is parametrized over both adapters and the Weixin docs name the ceilings.
2026-09-20 18:14:58 +05:30
kshitijk4poor
130b596f6f fix(whatsapp,weixin): text-batch delays default to Telegram cadence
WhatsApp debounced text for 5s (10s near a split) and Weixin for 3s/5s
before dispatching, so every reply paid multiple seconds of idle latency
that Telegram never pays (0.3s/1.0s). Default both adapters to Telegram's
cadence and mirror its ceilings (2.0s / 4.0s, split >= base delay) via the
existing _coerce_float_extra seam. The config keys are unchanged; 0 still
dispatches immediately. Docs updated.

Spotted via #44896 (@liuhao1024). Fixes #44883, refs #25056.
2026-09-20 18:14:58 +05:30
kshitijk4poor
18f1706137 refactor(matrix): drop dead is_direct guard and unreachable sender warning
_schedule_invite_join already requires both is_direct and inviter before recording m.direct, so `is_direct and bool(inviter)` at the reconcile call site was redundant; pass is_direct through. The `if is_direct and not inviter` WARNING could only be reached with GATEWAY_ALLOW_ALL_USERS set and a spec-violating stripped m.room.member event lacking `sender`; the info log already prints is_direct, so drop the branch.
2026-09-20 18:00:06 +05:30
kshitijk4poor
b98dd99142 docs(matrix): reconcile gate comment states the real premise
The adapter comment and the test module docstring claimed a reconciled pending invite "never fires _on_invite". It does: _absorb_sync runs _dispatch_sync (which emits INVITE to _on_invite) and then the reconcile pass over rooms.invite, which joined every entry unconditionally — so a live invite _on_invite rejected was joined ms later, and invites that arrived while the gateway was down were joined on restart with no gate. Reword both to state that premise.
2026-09-20 18:00:06 +05:30
Iain Lane
439f2b1cf6 fix(matrix): apply the inviter allowlist to reconciled pending invites
_on_invite only auto-joins a room when the inviter is allow-listed (or
GATEWAY_ALLOW_ALL_USERS is set), so a live invite from an arbitrary
federated user is rejected. A pending invite that arrives while the
gateway is down takes a different path: _schedule_pending_invite_joins
reconciles it from rooms.invite in the sync response and scheduled the
join unconditionally. An unauthorized invite sent during downtime was
therefore auto-joined on restart, bypassing the allowlist.

Extract the gate from _on_invite into _is_authorized_inviter and apply
it during reconciliation too, reading the inviter from the stripped
invite state (the sender of the m.room.member event for our own user,
as _extract_invite_dm_signal already does for the DM signal). An
inviter that cannot be read from the invite state fails closed, exactly
like an empty sender in _on_invite: the invite is skipped with a
warning and left pending.
2026-09-20 18:00:06 +05:30
Iain Lane
0ce3e7b12b fix(matrix): record reconciled direct invites in m.direct
A direct invite that arrives while the gateway is running fires
_on_invite, which passes is_direct and the inviter through
_schedule_invite_join so the room is recorded in m.direct after the
join. An invite that is still pending across a gateway restart takes a
different path: _schedule_pending_invite_joins reconciles it from
rooms.invite in the sync response, but called _schedule_invite_join
without is_direct or inviter. The DM signal was dropped, the room was
never recorded in m.direct, and it was classified as a group until the
user's own client happened to update m.direct.

Read the signal from the stripped invite state instead: the
m.room.member event for our own user carries the original invite's
is_direct flag, and its sender is the inviter. Thread both through to
_schedule_invite_join so a reconciled direct invite is recorded in
m.direct, and thus lands in _dm_rooms, exactly like a live one.

This gap was surfaced by the triage of #62493.
2026-09-20 18:00:06 +05:30
kshitijk4poor
536ce859c4 fix(telegram): bounded sends do not arm the blocked-loop watchdog 2026-09-20 17:57:43 +05:30
kshitijk4poor
d37f59af4c fix(telegram): deadline label default no longer claims init at non-init sites 2026-09-20 17:57:43 +05:30
kshitijk4poor
e2a54239fc docs(telegram): deadline comment names what is bounded 2026-09-20 17:57:43 +05:30
kshitijk4poor
85e673b4a5 fix(telegram): give media uploads their own 300 s wall-clock deadline
Reusing `_MEDIA_SEND_READ_TIMEOUT` (60 s) as the whole-call deadline turned
httpx's per-phase stall budget into a bandwidth cap: a 20 MB video on a
~2 Mbit/s uplink (~80 s) that succeeds today would fail. `_MEDIA_SEND_DEADLINE`
= 300 s covers the 50 MB Bot API cap at ~2 Mbit/s plus connect and sendVideo
transcoding, and is >2x the summed httpx budgets (pool 8 + connect 10 +
media_write 60 + read 60 = 138 s), so it only fires on a socket that has
stopped raising. Abandon-inside-lock semantics documented at the constant.
2026-09-20 17:57:43 +05:30
jollyroger1480
04057250a3 fix(telegram): wall-clock deadline on every Bot API send (text + media)
No Bot-API write in the adapter was under a wall-clock cap — text sends,
edits, drafts and media uploads relied solely on httpx socket timeouts,
which do not fire when a shielded httpcore socket wedges (same class as the
getUpdates hang in #92991). A stuck send then pinned `_chat_send_lock` and
the loop. Wrap every send/edit/draft in `_await_with_thread_deadline`
(`_TEXT_SEND_DEADLINE`) and both media paths in
`_send_with_dm_topic_reply_anchor_retry`; the helper gains a `label` and a
descriptive TimeoutError message.

Half A of #115280; Half B superseded by #116134 / contradicts #75017.
2026-09-20 17:57:43 +05:30
beardthelion
07c1953ea1 fix(dashboard-auth): pin OIDC discovery to the configured issuer origin
_fetch_discovery followed redirects but only pinned the document's
self-asserted issuer field, so one cleartext or attacker-hosted hop
could serve a forged document claiming the configured issuer with
attacker jwks_uri and token_endpoint. Verify then accepted
attacker-signed ID tokens and the code exchange POSTed the client
secret to the attacker's token endpoint.

The resolved response.url must now share the configured issuer's
origin (scheme, host, port with default-port normalisation) before
the body is parsed. Same-origin canonicalisation redirects still
pass, and the issuer-field pin remains as the misconfig check it is.
2026-09-20 00:09:52 -07:00