Commit Graph

2276 Commits

Author SHA1 Message Date
liuhao1024
ff307aea5e fix(telegram): name empty-string transport errors in adapter log lines
httpx timeout exceptions (ConnectTimeout, ReadTimeout, ...) stringify to
"" so every adapter log line built from _redact_telegram_error_text()
ended with a blank reason. Fall back to the exception class name.

Partial salvage of #111222: only the _redact_telegram_error_text hunk;
the transport-layer and polling-recovery hunks are covered by #111221.
2026-09-15 04:39:16 -07:00
fangliquan
e1e943a9df fix(telegram): expose polling transport recovery 2026-09-15 04:39:16 -07:00
teknium1
2298dc8122 fix: propagate caller contextvars into the Feishu adapter-owned executor
Moving the dedup flush and thread lookup from asyncio.to_thread onto the
adapter-owned pool fixed the torn-down default executor, but to_thread also
copies the caller's contextvars and run_in_executor does not. A multiplexed
profile's HERMES_HOME override and secret scope are contextvars, so those
workers silently ran under the launch profile. _run_blocking now runs the call
through contextvars.copy_context().run, matching to_thread semantics.

Review finding: _run_blocking lost the profile HERMES_HOME override / secret scope on the worker.
2026-09-15 04:38:26 -07:00
teknium1
b447554ce5 fix(feishu): thread-reply lookup uses the adapter-owned pool too
`_fetch_last_message_in_thread` was the last hot-path `asyncio.to_thread`
in the adapter: after a default-executor teardown (#111020) thread-reply
routing would fail the same way the dedup flush did. Route it through
`_run_blocking` like every other blocking SDK call. The remaining
`to_thread` users (`_load_lark_oapi` at connect/onboarding, the voice
transcode with its file-attachment fallback) are cold or degrade cleanly.
2026-09-15 04:38:26 -07:00
JackJin
5fbef868a9 fix(gateway): keep Feishu websocket off default executor 2026-09-15 04:38:26 -07:00
liuhao1024
fa2c72746c fix(feishu): run the inbound dedup flush on the adapter-owned pool
A dead background event loop tears the loop's default executor down, and
after that every inbound message was dropped inside the dedup gate with a
RuntimeError out of asyncio.to_thread — the adapter went permanently deaf
while the gateway process, websocket and service all stayed healthy.

#10849 already moved the outbound SDK calls onto an adapter-owned,
self-healing pool; this gives the inbound dedup-state flush the same
treatment, so a default-executor teardown can no longer wedge message
intake.
2026-09-15 04:38:26 -07:00
fangliquan
992b517fdd fix(photon): preserve Unicode NDJSON separators 2026-09-15 04:34:13 -07:00
wang2
fa12d7556c fix(telegram): allow opt-in CJK rich messages 2026-09-15 04:24:53 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1
624777617c fix(discord): filter obfuscated channels on the explicit-id backfill path too
Review finding on #90154: only the wildcard ("*") branch of
_iter_missed_message_backfill_candidates skipped obfuscated channels. A
channel the bot previously talked in that later lost VIEW_CHANNEL still
entered the explicit-id branch and produced a failing history read on
every startup. Apply the predicate once, after both branches build the
candidate list, and import it at module level like the other helpers.

(During the rebase the predicate was also tightened, resolved in the original
commit: is_discord_channel_obfuscated catches AttributeError (the precise
expected failure) instead of a bare Exception so a genuine attribute bug
is not hidden behind the name fallback, and its docstring states the
deliberate bias of the sentinel-name fallback (a visible channel literally
named ___hidden___ is also skipped on discord.py builds without the flag).)
2026-09-15 03:50:10 -07:00
Teknium
1f34742552 fix(discord): skip obfuscated channels in directory and backfill enumeration
Discord's Channel Obfuscation change (announced Aug 12 2026, HTTP
enforcement Nov 16 2026) dispatches channels the bot lacks VIEW_CHANNEL
on with name "___hidden___", flag 1 << 17 (CHANNEL_OBFUSCATED), and
nulled fields. Without filtering, the channel directory lists phantom
"___hidden___" entries the agent can never post to, and wildcard
missed-message backfill wastes history reads on channels that always
403.

Adds is_discord_channel_obfuscated() to gateway/platforms/helpers.py
(checks the flag bit plus the sentinel name for discord.py builds that
don't expose the new flag) and applies it at both enumeration sites.
2026-09-15 03:50:10 -07:00
teknium1
7abe9502ee fix(kanban): worker liveness and kills require the spawn-time start fingerprint
tasks.worker_pid outlives a reboot; afterwards the number can belong to any
process. _pid_alive answered from bare existence, so reclaim_stale_claims kept
extending the claim of a "live" stranger (stuck running task), and
enforce_max_runtime / _terminate_reclaimed_worker SIGTERM'd then SIGKILL'd it.

_set_worker_pid now records gateway.status.get_process_start_time(pid) as
tasks.worker_started_at (additive column, NULL on legacy rows). _worker_alive
(pid, started_at) is the liveness check every reader uses (reclaim, defer,
reconcile, crash sweep, max-runtime, archive, reopen invalidation); a live pid
whose fingerprint disagrees is a recycled PID: treated as dead, never
signalled (termination reports pid_recycled). Legacy rows without a
fingerprint keep the existence answer until their next spawn.

Row hermes_cli/kanban_db_dispatch.py:458 (lane4_high) confirmed by tracing:
the SIGKILL at :461 was gated only on _pid_alive.
2026-09-15 03:47:15 -07:00
teknium1
8b181940b4 fix(multiplex): background threads and teardown paths carry the turn's profile scope
The profile scope (HERMES_HOME override, secret scope, terminal policy) is a
contextvar bundle bound per turn. A bare threading.Thread / Timer / gRPC
callback starts with an empty context and resolves the LAUNCH profile:

- agent/title_generator.py: the auto-title thread read
  auxiliary.title_generation (model, language, provider key) from the default
  profile's config and billed the default's key for a secondary's session.
  Spawn via agent.memory_provider.spawn_context_thread (copy_context).
- tui_gateway/session_lifecycle.py: every teardown caller is a bare Timer
  (ws-orphan reap), the idle-reaper thread, atexit _shutdown_sessions, the
  session.close pool RPC, superseded_by_resume or compute_host flush - none
  carries a scope, yet on_session_end / commit_memory_session / agent.close ->
  shutdown_memory_provider read the provider's config + credentials at call
  time. Under multiplex they failed closed (tail never committed, #110622
  class); on the Desktop backend a secondary's transcript went to the launch
  profile's memory tenant. _finalize_session and _teardown_session now bind
  _session_profile_runtime_scope(session) around those blocks, which covers
  every spawn site through the single chokepoint.
- plugins/platforms/google_chat/adapter.py: Pub/Sub callbacks run on the gRPC
  SubscriberClient's threads and run_coroutine_threadsafe copies THAT empty
  context onto the loop task, so _dispatch_message and everything under it
  (attachment cache, per-user OAuth token store via _acquire_user_chat_api ->
  _load_per_user_chat_api, TTS keys, delivery ledger, bot-id cache) resolved
  the launch profile. connect() captures its scope; _on_pubsub_message and
  _submit_on_loop run under a per-callback copy of it.

spawn_context_thread gains a kwargs passthrough for the title thread's
callbacks.
2026-09-15 03:47:15 -07:00
teknium1
d1794d5539 fix(multiplex): children spawned for a served profile start from that profile's env
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:

- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
  base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
  BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.

tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.

Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
2026-09-15 03:47:15 -07:00
Sudhir Patil
ec11359b44 fix(discord): clear fatal status on successful reconnect
DiscordAdapter.connect() set self._running = True directly instead of
calling self._mark_connected(), unlike every other platform adapter
(Telegram, WeCom, Matrix, Feishu, Google Chat, IRC, LINE, Mattermost,
ntfy, photon, raft, simplex, a2a, buzz, dingtalk, whatsapp).

_mark_connected() clears _fatal_error_code/_fatal_error_message/
_fatal_error_retryable and rewrites the runtime status file as
"connected". Bypassing it meant a transient connect failure (e.g. a
one-off DNS blip: "Cannot connect to host discord.com:443 ssl:default
[Temporary failure in name resolution]") left the platform reported as
permanently fatal in gateway_state.json / the dashboard, even after the
adapter successfully reconnected and was actively serving messages for
hours.

Reproduced under gateway.multiplex_profiles: true with a secondary
profile's Discord bot (sarathi:discord) — the bot reconnected
repeatedly ("Connected as ..." logged many times over 18+ hours) while
the dashboard kept showing the original fatal error the entire time.

Fixes #102554.

Added a regression test asserting connect() clears a previously
recorded fatal error.
2026-09-15 03:44:36 -07:00
teknium1
fd303c0137 fix(gateway): 'decline' survives the config.yaml load path; Telegram forwards it; wizard offers it
config_loader._dm_behavior_choice still normalized against {"pair","ignore"},
so `unauthorized_dm_behavior: decline` in config.yaml (top level or a
platform block) was coerced back to "pair" on the real startup path
(load_gateway_config), and `unauthorized_dm_decline_message` was never
bridged into gw_data. Both now go through gateway.config.UNAUTHORIZED_DM_BEHAVIORS
(single source) and the presence bridge. The round-trip test exercises
load_gateway_config with a real config.yaml (top-level decline, telegram
override, custom message) instead of GatewayConfig.from_dict.

Telegram's intake prefilter only forwarded unauthorized DMs when the
behavior was exactly "pair", so with an allowlist configured a decline was
never sent. Anything that needs an outbound reply (!= "ignore") passes.

`hermes gateway setup` gains a "Politely decline unknown senders" choice
that writes platforms.<platform>.unauthorized_dm_behavior: decline; docs
mention it. Upstream-source references dropped from docstrings.
2026-09-15 03:44:13 -07:00
Teknium
5705b68f70 fix: memory-plugin and Qwen-CLI config JSON survives Windows BOM
Port from earendil-works/pi#8337 (UTF-8 BOM normalization in text inputs):
sibling sites the merged #81967 BOM sweep missed. json.loads hard-fails on
a leading U+FEFF and every one of these loaders swallows the exception and
silently falls back to defaults — a user who edited mem0.json, honcho.json,
hindsight/config.json, or supermemory.json in Notepad lost their whole
config with no error, and Qwen CLI OAuth creds saved with a BOM raised
qwen_auth_read_failed.

- plugins/memory/{honcho,mem0,hindsight,supermemory}: 13 read sites -> utf-8-sig
- hermes_cli/auth.py: _read_qwen_cli_tokens -> utf-8-sig
- tests: BOM regression tests per loader (sabotage-proven) + plain-UTF-8 guard
2026-09-15 03:38:29 -07:00
teknium1
d8053f4806 fix(video_gen): cap LTX 2.5 at 10s for 1440p/2160p; omit unset enum durations
fal's LTX 2.5 fast endpoints accept 6-20s only up to 1080p — "At 1440p and
2160p, all frame rates support up to 10 seconds" — so a 4K request with the
family's 20s ceiling was rejected by the vendor. Families can now declare
`duration_cap_by_resolution`, applied after the enum snap / range clamp on the
resolved resolution enum.

An unset duration on a duration_enum family also snapped to enum[0] (6s),
silently overriding the endpoint's own "auto" default; None now omits the key
for enum families exactly as it already did for range families.

test_managed_media_gateways asserts the alibaba/happy-horse/ namespace by
prefix rather than the exact v1.1 literal so the next version bump doesn't
flip an unrelated gateway test.
2026-09-15 03:37:11 -07:00
teknium1
1c1980dc1f fix(video_gen): make the duration enum explicit instead of sniffing tuple shape
`durations` carried two meanings told apart only by len==2 and gap>1: a
(min, max) range to clamp, or an enum to snap. A family with exactly two
legal values would have been misread as a range (review finding on #91311).
`durations` is now always the (min, max) window (what capabilities()/
list_models() read) and families with discrete values add `duration_enum`;
_clamp_duration takes the family and branches on the key, not the shape.

Also: restore the exact v1.1 endpoint assertion in the gateway namespace
test (a startswith/endswith check would not catch a silent version drift),
add the ltx-2.5 i2v snap case, and keep happy-horse on audio_native (the
schema test forbids audio+audio_native together, and v1.1 audio is always on).
2026-09-15 03:37:11 -07:00
Teknium
37286d3064 feat(video_gen): LTX 2.5 + Kling O3 families; Happy Horse upgraded to v1.1
Adds two new FAL video families and upgrades one:

- ltx-2.5 (cheap tier): lightricks/ltx-2.5/{text,image}-to-video/fast.
  Lightricks' open-source audio-video model. Native audio, 6-20s integer
  duration enum, 720p-2160p (i2v), $0.09/s at 720p. duration_int + 2k/4k
  resolution aliases; no seed key in the schema.
- kling-o3 (premium tier): fal-ai/kling-video/o3/standard/{text,image}-to-video.
  Kuaishou's frontier multi-shot model, 3-15s, optional native audio
  ($0.084/s off, $0.112/s on). String durations, i2v drops aspect_ratio,
  no seed/resolution keys.
- happy-horse upgraded from the sparse-docs 1.0 endpoints to
  alibaba/happy-horse/v1.1/{text,image}-to-video with the full published
  schema: nine aspect ratios, 720p/1080p, 3-15s integer durations, seed
  supported, audio native (no generate_audio key), i2v drops aspect_ratio.

All flags derived from each endpoint's llms.txt schema. Payload builder
asserted locally against the schemas; test for the old Happy Horse
"prompt-only" contract updated to pin the v1.1 schema, plus new payload
tests for ltx-2.5 and kling-o3.
2026-09-15 03:37:11 -07:00
kshitijk4poor
d793f7b9fb fix(supermemory): explicit failure sentinel and bounded pending-turn buffer
Two review findings on the salvage stack:

- _write_turns' _quietly consolidation keyed failure on a None result,
  implicitly assuming add_memory never legitimately returns None. A None-
  returning stub (the most common mock idiom) would mark every successful
  write failed and re-append the batch forever. Module-level _FAILED
  sentinel: only a raised exception re-queues.
- _pending_turns had no bound: a persistently failing service accumulated
  one entry per turn for the process lifetime (gateway runs never
  re-initialize), and every retry re-sent the whole accumulated payload —
  O(n^2) upload bytes, and a size-rejected batch could never shrink. Cap
  the buffer (50 turns / 256 KiB, drop oldest, warn once per trim).
2026-09-15 11:55:10 +05:30
kshitijk4poor
bc1d4776db docs(supermemory): purge stale session-ingest / /v4/conversations references
The per-turn capture rewrite removed the raw urllib /v4/conversations
ingest, but stale references survived outside the diff hunks:

- README "Behavior" still carried the "written once via the conversations
  endpoint" paragraph contradicted by the new bullets right above it.
- website memory-providers (EN + zh-Hans) still listed full-session ingest,
  session-end /v4/conversations ingest, and ingest in the base-url and
  api_timeout rows — the PR had updated one line per file but missed the
  rest of the section.

Now every surface describes per-turn documents.add capture with retry.
2026-09-15 11:55:10 +05:30
kshitijk4poor
3cad3c6db8 fix(supermemory): route capture-write logging through _quietly; drop dead default in custom id
Two cleanups on the salvaged per-turn capture path (review of #109359):

- _write_turns hand-rolled the try/except/log that the file's own _quietly
  helper already provides (same exception class, level, exc_info shape);
  the deleted _ingest path used _quietly for the identical call. Consolidate:
  add_memory always returns a dict, so a None result is the failure sentinel.
- _capture_custom_id's `or 'hermes'` never fires: _sanitize_tag already
  returns _DEFAULT_CONTAINER_TAG on empty input, and the literal duplicated
  the constant the helper owns. Verified behavior-preserving (tests green
  with the fallback artificially restored).
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
c6da9e0788 fix(supermemory): serialize capture writes and document at-least-once retry
sync_turn runs on the MemoryManager worker, but on_session_switch and
shutdown run on the caller thread, so two _write_turns() calls could
snapshot the same pending batch and each replace the whole list (duplicate
append or lost pending turn). A capture lock now covers the
snapshot/write/replace sequence; _write_turns() reads pending turns inside
the lock instead of taking a caller-built list.

Retries are at-least-once: the documents API appends on a shared custom_id
and does not dedupe by content, so a write it accepted but whose response
was lost is appended again. Stated in the docstring and README instead of
implied.

Addresses review on #109359.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
1b76cfff83 fix(supermemory): keep pending turns across session switch until written
A failed flush at on_session_switch() restored the pending buffer and
then cleared it on the next line, so an unavailable service at the switch
boundary still lost every pending turn. Pending turns now carry their own
session_id; _write_turns() batches per session, so a later retry (next
turn, session end, shutdown) writes old-session turns under the old
session's custom_id even after the switch.

Addresses review on #109359.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
03627dbf95 fix(supermemory): write turns via documents API instead of session-end conversations ingest
sync_turn now writes each completed turn through the SDK's documents.add,
keyed by custom_id "<session>_<date>_b<0-5>" so all turns of a session in
one 4-hour window append to a single document. This matches the capture
shape of the other Supermemory agent integrations and removes the raw
urllib POST to /v4/conversations, which the self-hosted server does not
implement (#101270).

Failed turn writes stay pending and are retried with the next turn, at
session end, on session switch, and at shutdown. Previously a failed
session-end ingest was logged once and the whole session was lost.

Inline base64 data URIs in captured text are replaced with "[image]" so
pasted screenshots no longer land in the document as megabytes of text.

Metadata stays type/session_id/timestamp plus the existing sm_source.
2026-09-15 11:55:10 +05:30
kshitijk4poor
82d61165d0 style(matrix): separate _MATRIX_PERMANENT_ERRCODES from the preceding function
The stack inserted the errcode table directly after _strip_reply_fallback's
return with no blank lines, which reads as if the constant belongs to the
function body and trips E305. Two blank lines restore the module-level
boundary; no behaviour change.
2026-09-15 10:50:45 +05:30
kshitijk4poor
257288ede1 test(matrix): make the 502 SVG fixture actually embed "403"
The inherited fixture coordinate "40.4302" does not contain the substring
"403" (the dot splits it), so the old substring classifier also passed on
it and the case proved nothing. Use a coordinate that genuinely embeds the
digits so the test is red on the pre-fix classifier.
2026-09-15 10:50:45 +05:30
kshitijk4poor
d5430fb0c1 refactor(matrix): classify sync auth errors on errcode/http_status only
The pinned mautrix 0.21.1 raises MatrixRequestError (carrying errcode and
http_status) from HTTPAPI._send on every non-2xx and sync() returns only
the parsed JSON dict, so the result-object auth branch in _sync_loop was
unreachable; it dated from the nio client whose SyncError objects were
real. Drop it together with the nio-mock test that pinned it.

With structured attributes guaranteed, the leading-status regex and the
bounded keyword scan over the message text were the only remaining ways
for body digits or HTML words to leak into the verdict, so drop them too:
no errcode/http_status auth signal means retry. Trim the contributor's
17 tests to the two loop-level invariants: both production repros (502
HTML body embedding "403" via an SVG coordinate; timeout echoing a since
token embedding "401") keep looping, and a 401/M_UNKNOWN_TOKEN stops.
2026-09-15 10:50:45 +05:30
Stephen Chin
9969a995f6 fix(matrix): correct sync result-object comment and route it through the classifier
The comment above the result-object branch in _sync_loop claimed mautrix's
Client.sync() returns an object carrying a message string for auth failures.
That is wrong. In the pinned mautrix 0.21.0, HTTPAPI._send raises
make_request_error() for any non-2xx and otherwise returns parsed JSON, so a
real M_FORBIDDEN arrives as an exception and is handled by the except branch.
The claim was introduced by this PR, which rewrote an accurate comment about
the earlier matrix-nio client (whose SyncError result objects were genuine).

The branch itself is kept as defense in depth against a future client swap,
but it now classifies with the same errcode/http_status logic as the
exception path instead of a lone "unknown_token" substring test, which
silently missed M_MISSING_TOKEN and M_FORBIDDEN and resynced forever
against a credential that can never succeed.

A structured errcode/http_status is authoritative; the message text is only
consulted when the object exposes neither, since str(object) is an opaque
repr. The text scan deliberately cannot override a structured verdict, so a
transient 502 whose HTML body contains "Forbidden" is still retried.

Adds four tests. Three are discriminating RED/GREEN cases that fail against
the old substring branch (M_MISSING_TOKEN errcode, http_status=401 with no
keyword in the message, and an unstructured object whose only signal is
.message). The fourth pins the precedence rule and passes either way.

Verified: 136 passed / 1 failed in tests/gateway/test_matrix.py; the single
failure (test_password_login_uses_device_id) fails identically at the
pristine PR head and is unrelated.

(cherry picked from commit bc9e6a8dafcf349a4e6b20a261fb2449603c0239)
2026-09-15 10:50:45 +05:30
Stephen Chin
6bd9cf8018 fix(matrix): use a genuinely discriminating fixture for the sync-loop test
The independent-verifier caught that my first loop-level test did not
actually prove anything. The 502/SVG coordinate fixture I reused from
gmoranxyz's unit-level test does not contain the substring 403 once
case-folded, so the old naive substring classifier already treated it
as transient. A test that passes under both the buggy code and the
fix proves nothing about the fix.

I replaced the fixture with a plain connection timeout whose message
wraps the real Matrix sync pagination token, an arbitrary digit
string that happens to contain 401. I verified this directly: with
the pre-fix classifier restored, the retry test now fails (the old
code stops the loop on this fixture), and with the fix in place it
passes (the loop retries as it should). That is the RED/GREEN proof
the maintainer originally asked for.

I also documented in the stop test's docstring that it does not
discriminate old from new, since the word forbidden in its message
trips the old naive check too. It is still worth keeping as a
regression test proving genuine auth errors stop the loop, just not
as proof of this specific fix.

While I was in there I also fixed a stale comment above the
M_UNKNOWN_TOKEN sync-object pre-check. It said nio returns SyncError
objects, but the dependency here is mautrix, not matrix-nio, and
importing nio raises ModuleNotFoundError in this codebase. The
pre-check logic itself was already correct and untouched.

Co-authored-by: gmoranxyz <gmoranxyz@users.noreply.github.com>
(cherry picked from commit ad3aad579a675a5aae544a50f88a82717c0ac3b6)
2026-09-15 10:50:45 +05:30
Stephen Chin
f8085f2e03 test(matrix): add loop-level and attribute-narrowing coverage
I added two more classifier unit tests for the attribute narrowing:
a bare .code attribute that happens to be 401, and a bare .status
attribute that happens to be 403, both must stay classified as
transient since only .http_status is trustworthy. I also added a
parametrized test for the five transient exception types the sync
loop now short-circuits on.

On top of that I added two tests that exercise _sync_loop directly
instead of just the classifier function in isolation. One replays the
real 502 Umbrel repro string through a mocked client.sync and confirms
the loop retries with the 5s backoff. The other raises a genuine
M_FORBIDDEN error and confirms the loop stops on the first call with
no retry sleep. These catch a regression in how the loop wires the
classifier in, not just a regression in the classifier itself.

(cherry picked from commit f747bb4b5a6e6da1bb9136168f08d6e7af5ea64b)
2026-09-15 10:50:45 +05:30
Stephen Chin
104889ace6 fix(matrix): tighten sync error classifier
I hit a bug where the Matrix sync loop treated a passing 502 from
Umbrel's app proxy as a permanent auth failure and stopped syncing for
good. The old check did a naive "403" in str(exc) substring match, and
the 502 HTML error body embedded an SVG path with the coordinate
40.4302, which contains the digit sequence 403.

I replaced the substring check with a layered classifier. Transport
exceptions like TimeoutError, ConnectionError, and OSError are always
treated as transient regardless of their message text. Structured
signals take priority next: the errcode attribute against a known set
of permanent Matrix error codes, then the http_status attribute
against 401/403 specifically (not status, status_code, or code, which
belong to unrelated exception shapes and risk coincidental integer
matches). Only when none of those are present does it fall back to a
bounded, word-boundary-safe text scan on the first 200 characters.

Added tests covering the attribute narrowing, the transient exception
types, and two loop-level tests exercising _sync_loop directly to
confirm it retries on a transient error and stops on a genuine 401/403.

(cherry picked from commit 96d3363e45a63e08d9f07518ed949df334a33b3c)
2026-09-15 10:50:45 +05:30
kshitijk4poor
7959e3b0ff test(discord): bind the "0 disables event-silence" half of the knob test
The zero-knob test only asserted that ack_stale still trips with the knob
at 0. Because _read_websocket_health evaluates ack age before the
event-silence dimension, the test stayed green even with the
`_event_max_silence_seconds > 0` guard deleted: it never observed the
stale stamp being ignored. Assert (True, "healthy") with knob=0, a stale
stamp and a green transport first, then make the ACK stale and keep the
ack_stale assertion. Dropping the guard now fails this test.

Also correct the on_socket_event_type comment: discord.py dispatches
socket_event_type before the op-code switch but only for a non-null `t`;
heartbeat ACK frames carry `t: null`, they do not "return before" it.
2026-09-15 10:48:58 +05:30
kshitijk4poor
909731dc71 refactor(discord): state the dispatch-liveness rationale once
The #109521 explanation (socket_event_type fires for every parsed
DISPATCH frame and is not debug-gated unlike on_socket_raw_receive;
op-11 ACKs carry no event type) was written out four times: knob init,
stamp init, the on_socket_event_type handler, and _read_websocket_health.
Keep the authoritative paragraph on the handler that actually stamps, and
point the other three sites at it so a future edit has one place to go.

Drop the math.isfinite(event_silence) half of the silence guard. Both
operands are our own perf_counter floats, so their difference cannot be
non-finite; the ack_age guard is different because _last_ack comes from
discord.py. math stays imported for the remaining finiteness checks.
2026-09-15 10:48:58 +05:30
kshitijk4poor
b4a5ce62e7 fix(discord): scope the event-silence knob warning to its dimension
`_warn_liveness_config_disabled` told operators an unusable value turns
off "the websocket liveness probe". That is true for the interval,
threshold, ack-age and latency knobs, which sit in the probe's startup
guard, but `websocket_event_max_silence_seconds` is only checked inside
`_read_websocket_health`, so ack-age/latency keep guarding. Say so, or
an operator reading the log would believe the whole watchdog is down.
2026-09-15 10:48:58 +05:30
salch-cred
101861c7f7 fix(discord): dispatch-side liveness dimension detects an ACKing-but-deaf gateway socket (#109521)
Incident 2 of #109521: a Gateway socket can stay ESTABLISHED and keep
ACKing heartbeats while zero DISPATCH events are parsed, so every
transport-side liveness sample (ready/open/ack-age/latency) reads
healthy for hours. The merged #109963 deliberately dropped the
event_silence dimension: a raw-frame stamp is debug-gated
(on_socket_raw_receive needs enable_debug_events) and, since heartbeat
ACKs are frames, ack_stale always fires first by construction.

This adds the dispatch-side signal that was requested instead:

- stamp on on_socket_event_type, which discord.py 2.7.1 dispatches for
  every parsed DISPATCH frame with no debug gate (verified live against
  the real received_message path: 4/4 frames fired with
  enable_debug_events=False, on_socket_raw_receive 0/4)
- new knob websocket_event_max_silence_seconds (default 4h, the
  incident report's field-proven operator bound); 0 opts out of this
  dimension ONLY — the #109782 review failure put the knob in
  _start_liveness_probe's all-or-nothing guard, killing the whole
  watchdog; it is gated strictly inside _read_websocket_health here
- the stamp resets per connection (connect() clears it), and a None
  stamp (no event parsed yet on this connection) is not silence
- docs (en + zh-Hans) cover the new knob and the per-dimension opt-out

Fixes #109521

(cherry picked from commit b4baa97fc45794209711a45e052111d7d44d5f90)
2026-09-15 10:48:58 +05:30
Andrew Madson
1ad89ac018 feat(web): send Hermes identity headers on Perplexity requests
Add HTTP-Referer, X-Title, User-Agent (HermesAgent/<version>) and
X-Pplx-Integration: hermes-agent to the Perplexity web provider's
/search and /sdk/content/snippets requests. Same static headers Hermes
already sends Kimi and OpenCode; nothing per-user, no extra requests.
2026-09-14 13:04:39 -07:00
teknium1
3275ca88ec fix(image_gen): Codex-auth images use the native images endpoints, no chat host model
The openai-codex image provider rode a Responses call with a hosted
image_generation tool on a pinned chat model (gpt-5.5). Two failure
classes came with that shape: when OpenAI withdrew gpt-5.5 from an
account cohort every image call 404'd while chat kept working
(#105398, #107076), and the host model was free to answer in text
instead of calling the tool, so we streamed SSE, kept partial frames
and retried on empty streams.

Post to chatgpt.com/backend-api/codex/images/generations and
images/edits instead - the route the official Codex client uses
(codex-rs/ext/image-generation). No host model, no SSE, no
partial-frame handling; the response is a plain JSON body with
b64_json. Remote source URLs are fetched client-side and inlined as
data URLs because the backend's own downloader 400s on ordinary
public images.

The backend treats model/quality/size as advisory (#107233), so the
result now reports reported_quality/reported_size next to the
requested values plus the x-codex-imagegen-request-id for support.
GPT Image 2.5 is deliberately not added to this catalog: the backend
accepts any model id, including nonexistent ones, and generates with
its server-managed engine (C2PA reports gpt-image 2.0), so a 2.5 tier
here would be a label with no effect (#106708).
2026-09-14 10:18:13 -07:00
teknium1
939a2f64b4 fix: long-lived processes stop minting duplicate state.db writer handles
Gateway, dashboard, ACP server and the CLI already hold one registry-shared
SessionDB per state.db path, yet several in-process call paths still opened
a bare SessionDB() beside it. Each one is a full writer: schema init, write
lock, token-writer thread and a close-time WAL checkpoint. On a dashboard
serving overlapping requests that stacked up to the "5 live SessionDB
handles" precursor within seconds; per #110544 they are now harmless to each
other's WAL generation, but the leak itself remained.

Pure readers attach read_only=True (no writer connection, no write lock):
  - plugins/hermes-achievements/dashboard/plugin_api.py::scan_sessions
    (dashboard, per background scan and per /rescan; highest-frequency site)
  - hermes_cli/console_engine.py::_session_db (dashboard console; list,
    stats and export are reads; rename/optimize opt in to a writer)
  - tools/process_registry_results.py::_owns_result (gateway, per retained
    result load)
  - hermes_cli/main.py::_session_db (last-session / title / cwd lookups)
  - hermes_cli/terminal_breadcrumbs.py, hermes_cli/status.py,
    hermes_cli/main_tui_launch.py (lookup one-shots)

Writers share the process's registry handle (hermes_state_registry.acquire;
close()/release_or_close release one refcount):
  - acp_adapter/session.py::SessionManager._get_db — the AIAgent it builds
    acquires the same path, so the ACP server held two writers per process
  - hermes_cli/kanban_db_dispatch.py::_retag_legacy_worker_sessions
    (gateway dispatcher tick)
  - hermes_cli/main.py::_create_titled_session, hermes_cli/oneshot.py,
    hermes_cli/foreign_sessions.py — the CLI acquires the same handle a
    moment later

The "N live SessionDB handles" warning now counts only writable members:
read-only attaches are the sanctioned per-request shape for dashboard
routers and CLI lookups, and counting them turned a healthy topology into
an operator alarm (#100896 field reports of restart loops keyed on it).

Live repro (one registry writer + 4 overlapping dashboard/gateway paths in
one process): before 5 writable opens, 5 live handles, warning fired;
after 1 writable open (the registry handle), 0 from the request paths,
no warning.

Refs #100896 #103339
2026-09-14 08:10:36 -07:00
Victor Kyriazakos
a4e2a82a6d fix(gateway): preflight task-card destination before transport fallback 2026-09-14 07:46:51 -07:00
Teknium
aecee6f66a feat(video): Gemini Omni Flash 1.1 — text-to-video now works (was image-only)
FAL shipped google/gemini-omni-flash/v1.1/* (Aug 2026): the family gains a
text-to-video endpoint, 360p/720p/1080p/4k resolution enum, and keeps
3-10s integer durations with always-on native audio.

- plugins/video_gen/fal: bump gemini-omni-flash to the v1.1 endpoints,
  declare the resolution enum, refresh display/strengths copy
- tests: family test asserts the versioned dual-modality endpoints; the
  i2v-only clean-error guard survives via a synthetic family; catalog
  invariant now requires both endpoints on every family

Schema verified against FAL OpenAPI (t2v: prompt required, 16:9/9:16,
360p-4k, duration 3-10 int; i2v adds image_url + optional end_image_url).
Pricing: $0.03/s 360p, $0.10/s 720p, $0.15/s 1080p, $0.30/s 4K.
2026-09-13 21:40:26 -07:00
Teknium
80e8de8c34 feat(video-gen): add Wan 3.0 + Wan 3.0 Prime FAL video families
Alibaba's Wan 3.0 generation went live on FAL (was "Coming Soon" since
the Wan 2.7 rollout). Adds two families to FAL_FAMILIES:

- wan-3.0: alibaba/wan-3.0/{text,image}-to-video — 2-30s integer
  durations, 480p/720p/1080p, native audio, adaptive aspect handling
  ($0.05-$0.20/s by resolution)
- wan-3.0-prime: alibaba/wan-3.0-prime/{text,image}-to-video — premium
  tier, same surface ($0.068-$0.28/s)

Schema quirks mapped from the FAL OpenAPI/llms.txt:
- i2v takes `start_image_url` (image_param_key, Kling-4K pattern)
- duration is a JSON integer (duration_int)
- audio toggle key is `audio`, not `generate_audio` — new generic
  `audio_param_key` family flag in _build_payload (veo3.1 regression
  asserted unchanged)
2026-09-13 21:28:11 -07:00
Teknium
d26dbb2f77 Port from OpenHands/OpenHands#16758: Anthropic model catalogs no longer stop at the first page
Anthropic's /v1/models is cursor-paginated with a default page size of 20.
Both hermes fetchers read a single unpaginated page, so any model past the
first page silently vanished from the /model picker and provider catalogs.

- hermes_cli/models.py _fetch_anthropic_models(): request limit=1000 and
  follow has_more/last_id (bounded, repeated-cursor guarded, de-duped)
- plugins/model-providers/anthropic fetch_models(): same pagination walk,
  and it now honors the base_url argument instead of hardcoding
  api.anthropic.com
- tests: live-HTTP paginated-server regression tests for both fetchers,
  incl. single-page and stuck-cursor termination; updated the two URL-pinning
  pool-discovery tests for the ?limit=1000 contract
2026-09-13 21:07:35 -07:00
teknium1
82e4ac4b01 fix(opencode-free): aux default is mimo-v2.5-free — the only surviving free model that answers promptly
Live anonymous probes (2026-09-13, HermesAgent UA, 4 rounds): nemotron-3.5-lightning-free
returned zero bytes for >90s on every attempt and nemotron-3-ultra-free took ~40s, while
mimo-v2.5-free answered 200 in 2-4s each time. An aux model (titles, compression, memory)
that hangs is worse than a delisted one, so the default moves to the model that works.
2026-09-13 21:06:39 -07:00
Teknium
cc81e436ce fix(models): delist hy3-free and laguna-s-2.1-free — OpenCode relay dropped them (anon 401)
The OpenCode Zen relay no longer serves hy3-free (since ~2026-08-31) or
laguna-s-2.1-free (new, verified 2026-09-09): both are gone from the live
GET /zen/v1/models catalog and anonymous chat completions return
401 {"type":"ModelError","message":"Model <id> is not supported"}
(2 probes >=60s apart, x-opencode-session header present).

- hermes_cli/models_catalog_static.py: remove both slugs from the
  opencode-free offline floor and the opencode-zen discovery floor;
  document the delist dates in the catalog comment.
- plugins/model-providers/opencode-free: default_aux_model moves from the
  dead laguna-s-2.1-free to nemotron-3.5-lightning-free (fastest surviving
  anonymous model).
- tests: swap fixtures off the dead slugs; extend the floor-exclusion
  invariant to cover both.

The live revalidation path already hides them when the relay is reachable;
this fixes the OFFLINE floor and the aux default, which would otherwise
offer/route to models that 401.
2026-09-13 21:06:39 -07:00
Teknium
3304d205be feat(video-gen): Kling 3.0 Standard + Pro families on the FAL backend
Adds kling-v3 (fal-ai/kling-video/v3/standard/*) and kling-v3-pro
(fal-ai/kling-video/v3/pro/*) to FAL_FAMILIES: start_image_url i2v key,
aspect_ratio dropped on i2v, string duration 3-15s, generate_audio and
negative_prompt real, no seed/resolution keys per the published llms.txt
schemas. Payload shapes pinned in tests; docs mention updated.
2026-09-13 21:01:32 -07:00
teknium1
aea84cbb1a test(slack): collapse pasted-table tests to five invariants, cover live unfurl path
The rebase moved the inbound attachment loop into SlackAdapter._append_link_unfurls,
so the nested-table hunk now lives there and is asserted directly. Drop the source-
provenance references and duplicate cell-level cases; one ragged/malformed-row test
covers raw_text, rich_text, None and unknown cell types.
2026-09-13 20:59:55 -07:00
Teknium
a74e0155b6 feat(slack): pasted tables now reach the agent instead of silently vanishing
Port from qwibitai/nanoclaw#3666: Slack represents a pasted table as
'table' blocks — usually nested in attachments[].blocks[], sometimes
top-level. They appear in neither the message text nor the file list,
so the agent received the sentence before the table and nothing else.

- _render_slack_table_block(): projects rows as 'cell | cell' lines,
  collecting text leaves from raw_text/rich_text cell subtrees; capped
  at 20k chars with a visible '[table truncated]' marker.
- Wired into all three ingestion paths: _extract_text_from_slack_blocks
  (thread history + attachment-nested blocks), the live inbound
  attachment loop, and _extract_additional_text_from_slack_blocks
  (top-level blocks on live messages).
- _serialize_slack_blocks_for_agent skips 'table' blocks — the
  allowlist drops 'rows', so it only emitted an empty husk.
2026-09-13 20:59:55 -07:00
teknium1
8ed7c180d7 fix(discord): preflight send_voice uploads too; move the size gate into adapter_media
send_voice built its own discord.File from the audio bytes and never ran
the size preflight, so an oversized audio attachment still burned the
doomed 413 round-trip that #50846 is about. Route it through the same
_reject_oversized_upload helper as _send_file_attachment.

The limit constant and _discord_upload_limit_bytes lived on the adapter
facade while every consumer is in adapter_media.py; the facade+sibling
layout puts topic code in the sibling, so they move there.
2026-09-13 20:59:17 -07:00