Commit Graph

2299 Commits

Author SHA1 Message Date
teknium1
e3a90079b4 fix(plugins): one unreadable plugin child no longer aborts iter_plugin_dirs or memory-provider discovery
iter_plugin_dirs stat'd <child>/__init__.py inside every plugin dir, so a
single mode-000 / ACL-denied $HERMES_HOME/plugins/<x> still raised
PermissionError out of the loader, memory-provider discovery (dashboard
memory settings, hermes memory setup, plugins memory picker) and the
user cron-provider scan. Catch OSError per child and log the same
'Skipping unreadable plugin directory' warning the list path already emits.

Part of #111804
2026-09-15 18:48:59 -07:00
teknium1
4b00842536 fix(web): keyless MCP honours SSE CR terminators and declared charsets
Review follow-up on the keyless Exa UTF-8 fix:

- `_parse_mcp_body` split SSE frames on `\n` only, so a body using bare CR
  line terminators (permitted by the SSE spec, and parsed fine before via
  `splitlines()`) failed with "Unrecognized MCP response shape". Split on
  the SSE terminators CRLF / CR / LF instead.
- `mcp_call` decoded the body as UTF-8 unconditionally, ignoring a charset
  the server did declare. Decode with the declared charset when the
  Content-Type carries one and fall back to UTF-8 only when it is absent.
- The HTTP >= 400 branch and `_keenable_request` still surfaced
  `response.text`, so a non-ASCII error body from a charset-less text/*
  response reached the user as ISO-8859-1 mojibake. Both now go through the
  same decode helper as the success path.
2026-09-15 18:43:28 -07:00
teknium1
28cdd72815 fix(web): keyless Exa decodes its SSE body as UTF-8 so CJK searches work
Exa's MCP endpoint answers `text/event-stream` without a charset, so
`requests` decoded `.text` as ISO-8859-1: every non-ASCII character came back
as mojibake, and for CJK results the UTF-8 continuation byte 0x85 became
U+0085, which `str.splitlines()` treats as a line break — the `data:` JSON
line was cut in two and the perfectly valid result surfaced as "Unrecognized
MCP response shape", after which the ring fell through to keyless Firecrawl
(403). Decode the bytes as UTF-8 (JSON-RPC and SSE are UTF-8 by spec) and
split SSE frames on newlines only. A parsed envelope that really carries no
text is now reported as "no text content" instead of an unrecognized shape.
2026-09-15 18:43:28 -07:00
KoNit-K
ffd02f37e8 fix(web): cache extracts by returned URL 2026-09-15 18:42:38 -07:00
teknium1
0959224313 fix(kanban): claim-less complete no longer closes a live worker's run
complete_task authorised a terminal transition by task status alone; the
`current_run_id = ?` fence only applied when the caller volunteered
expected_run_id (derived from HERMES_KANBAN_* env). A human at the CLI, an
orchestrator session or any env-less caller therefore marked a `running`
card done and _end_run closed the dispatcher worker's run row while that
worker kept executing (#111764).

Mirror the fence request_review already carries: a `running` task under a
live claim needs expected_run_id (worker ownership) or force=True (explicit
operator override), otherwise LiveClaimError. `hermes kanban complete
--force` and the dashboard's "mark done" (a human action) carry the override;
the kanban_complete tool reports a structured refusal. Completing `ready`,
`blocked` or `review` cards without a claim is unchanged, so the manual /
orchestrator flows PR #73188 pinned keep working.

Fixes #111764
2026-09-15 18:34:40 -07:00
teknium1
f7ea39481a fix(kanban): dashboard estimate calls declare a relay-affinity key too
The dashboard's estimate endpoints make the same headless auxiliary call
as specify/decompose but never bound an affinity scope, so they still sent
no x-opencode-session and the OpenCode Go relay answered 400
MissingSessionID (#112043). Declare kanban:<task_id> for an existing task
and a stable kanban:estimate key for the create dialog (no task yet),
unless a scope is already bound.

Test: _run_estimate captured header None before; now kanban:t_1 /
kanban:estimate and nothing leaks past the call.
2026-09-15 18:33:43 -07:00
teknium1
b027a4658e fix: trim DeepInfra reasoning salvage to the invariant and document it
Follow-up to the cherry-picked #111876 (@KoNit-K), which shares the design
of the earlier #111875 by the issue author (@ats3v): emit DeepInfra's
top-level ``reasoning_effort`` from the provider profile, ungated on
``supports_reasoning``, ``none`` as the only off switch, ``xhigh`` native,
``ultra`` clamped to ``max`` via the shared vocabulary, unset/unknown omitted.

- drop the constructor/blank-line reformat churn (byte-identical to main)
- replace the 14-case test file with two invariant tests: the profile's
  config -> top-level field table, and the transport main-turn path with
  ``supports_reasoning=False`` (the gate the core allowlist actually passes)
- docs: DeepInfra subsection in integrations/providers.md describing the
  two-directional reasoning control

Offline kwargs probe: before every reasoning_config -> ({}, {}) and the
main turn carried no reasoning field; after ``high`` -> ``reasoning_effort:
high``, ``{'enabled': False}`` -> ``none``, ``ultra`` -> ``max``, unset and
unknown levels omitted, aux calls stop emitting the generic
``extra_body.reasoning`` for this provider.

Co-authored-by: Georgi Atsev <georgi@deepinfra.com>
2026-09-15 18:22:40 -07:00
KoNit-K
4fe3d288eb fix(providers): send DeepInfra reasoning effort 2026-09-15 18:22:40 -07:00
teknium1
3272fb35aa docs: profile-scope invariant in AGENTS.md — one process serves many profiles; out-of-turn code binds its scope
Root AGENTS.md § Code Shape Rules replaces "module-level constants are fine — they cache after
_apply_profile_override() sets HERMES_HOME" (true for `hermes -p x <cmd>`, inverted under the
multiplex gateway and the Desktop/dashboard `serve` backend, where os.environ holds the LAUNCH
profile) with the invariant: a profile = home + secret scope + terminal scope, bound per profile
ACTIVITY, and every execution point with no turn on the stack binds it explicitly. Names the real
seams: gateway/run.py::_profile_runtime_scope, tui_gateway @_profile_scoped +
_session_profile_runtime_scope (+ _profile_runtime_scope_tokens, launch_profile_policy ->
set_multiplex_active), cron/scheduler_provider.py::_profile_cron_scope,
gateway/run_agent_cache.py::_run_release_in_profile_scope, tools/environments/local.py::
served_profile_child_env, agent/memory_provider.py::spawn_context_thread. Adds a routing-table row
for profiles / multiplex / secret scope.

Area AGENTS.md paragraphs, one per seam, for gateway/ (activity-not-turn binding, hooks per
profile, adapter YAML never reaches os.environ, unserved shared-ingress reported via
_note_unserved_secondary_platform + needs_attention at the single writer), tui_gateway/ (RPC
binding is home AND secret AND terminal; HOME-only is half-bound; teardown chokepoint), cron/
(per-home tick lock, ticker scope incl. pre-loop code, kanban notifier routing, worker liveness by
(pid, worker_started_at) fingerprint, descendant fence as a path), hermes_cli/ (DEFAULT_CONFIG
key <-> reader parity, service-install matrix, -p vs multiplex home binding), tools/ (check_fn
reads through get_secret and is cached per hermes_home_key, one env builder per spawn, MCP trust
per profile), plugins/ (lifecycle hooks are bound by the caller; never cache the home from
initialize()), apps/desktop/src/ (pooled serve per (connection, profile); remote topologies),
agent/ (end-of-session flush is caller-bound; set_multiplex_active gates fail-closed).

Corrects the statements the multiplex model made wrong, in the same PR: root module-constant
sentence; hermes_cli "sets HERMES_HOME before any import" (+ cli-internals.md);
ADDING_A_PLATFORM.md §2 raw os.getenv loader (now an _ENV_STEPS row through config.py::_getenv)
and §4 platform_env_map in gateway/run.py (now _PLATFORM_ALLOWLIST_ENV in pairing.py + registry
allowed_users_env); platform_registry.py "may set os.environ (guard with not os.getenv)";
cron/AGENTS.md hardcoded ~/.hermes/cron/.tick.lock; gateway-internals.md agent:main as THE key
format, ~/.hermes/hooks/, single-profile `gateway stop`, plus a new "Multiplexed profiles"
section; tools/AGENTS.md os.getenv check_fn sample; "installed per turn" wording; "one temp
HERMES_HOME" E2E wording; multi-profile-gateways.md intro lists system units, Windows tasks, s6
and the Desktop backend.
2026-09-15 10:59:22 -07:00
teknium1
9c311038fa fix(discord): slash registration and /skill refresh scan the catalog off the event loop
connect() called _register_slash_commands inline on every (re)connect, and
/reload-skills called refresh_skill_group inline; both run
discord_skill_commands_by_category, the same per-skill path-resolution walk
the Telegram menu paid in #110707. On a 1.5k-skill install that holds the
loop past the liveness watchdog. Registration now hops through
asyncio.to_thread from connect(); refresh_skill_group is a coroutine that
hops the rescan the same way (the reload handler already awaits an awaitable
result). Contextvar-scoped profile overrides propagate through to_thread.

One invariant test: the loop keeps ticking while the scan blocks, from both
sites. Red on the PR head, green here.

Review finding: Discord _register_slash_commands/refresh_skill_group ran the skill catalog disk scan synchronously on the event loop.
2026-09-15 06:35:04 -07:00
teknium1
ab782a7685 fix(gateway): skill-slash fallthrough and the Telegram inline picker run off the event loop
The gateway's idle-command path resolved skill slash commands inline on the
loop: a cold skill scan, skill file loads and the unavailable-skill rglob over
every skills dir. On a 1.5k-skill install that held the loop ~2 minutes, the
loop-liveness watchdog fired and the gateway exited mid-session (#111091).
_hm_skill_slash_rewrite now runs through _run_in_executor_with_context so the
profile contextvars the scan is scoped to survive the hop. Known commands still
short-circuit before any I/O (previous commit).

Same class in the Telegram inline picker: build_inline_results rebuilds the
command/skill catalog per keystroke via _collect_gateway_skill_entries, the
same path-resolution pass #110707 traced in the command menu.

One invariant test: a known command triggers no scan; an unknown command's scan
runs while the loop keeps ticking. Red on origin/main and on the reorder-only
tree, green here.
2026-09-15 06:35:04 -07:00
teknium1
660d4b8d87 fix(telegram): hop only the menu build off the loop; one invariant test covers both menu sites
Follow-up to the salvaged #110716 commit. Drops the module-level
_build_telegram_command_menu wrapper (telegram_menu_max_commands is a config
read and stays on the loop; asyncio.to_thread takes kwargs directly), and
applies the same hop to _ensure_forum_commands, which rebuilds the menu on the
inbound-message path for every forum chat until registration succeeds — the
second live fire site named in #110707.

Replaces the contributor's test with one that proves the invariant for both
sites: the loop keeps ticking while telegram_menu_commands blocks. Red on
origin/main, green here.
2026-09-15 06:35:04 -07:00
KoNit-K
0a5846d37f fix(telegram): build command menu off event loop 2026-09-15 06:35:04 -07:00
teknium1
89ef145254 fix(kanban): dashboard link reports the gate; docs; trim to two invariant tests
The dashboard's POST /links is the fourth writer of link_tasks (CLI, tool,
dashboard, plus the graph builder); return the same ``gated`` flag so every
surface that can create the deadlock can see it. Document the
``dependency_wait`` payload the link path emits and the delegation rule the
reporter derived (never link a support card under the card it unblocks).

Drops the CLI output test (a change-detector on prose); the two DB-level
invariants (event emitted on demotion / none for a done parent) stay.
2026-09-15 06:25:42 -07:00
teknium1
060e2a4286 fix: exempt the Matrix quote block from the mention strip only on real replies
The reply-pill fix split the body on a leading "> " so the strip would not
rewrite the "> <@bot:srv>" pill. That keyed the exemption on the body shape,
not on the relation: a hand-typed blockquote in a plain (non-reply) message
that mentions the bot inside the quote reached the agent with the raw
"@hermes:example.org" text, which main used to strip. Split around the quote
only when m.relates_to carries m.in_reply_to; otherwise strip the whole body.

Review finding: quote-block exemption keyed on body.startswith('> ') instead of m.in_reply_to.
2026-09-15 06:19:19 -07:00
KeyArgo
0813a2276c fix(gateway): keep the Matrix reply fallback pill out of the mention strip
The Matrix adapter strips the bot's mention from the inbound body before
_extract_reply_context parses the inline reply fallback, and that strip is a
blind whole-body replace. A reply to the bot names the bot in the fallback
pill ("> <@bot:server> quoted"), which is exactly what makes the message
count as a mention under the default MATRIX_REQUIRE_MENTION=true -- so the
strip runs on every reply-to-the-bot and rewrites the pill to "> <>", after
which the pill regex no longer matches. reply_to_author_id is lost,
reply_to_text becomes the mangled "<> quoted" remnant, and the prompt
renders "[Replying to: "<> ..."]" instead of "[Replying to your previous
message: ...]".

Split the body into (quote block, reply text) and strip the mention from the
reply text only. The mention gate still sees the raw body, so a reply to the
bot keeps waking the bot; the visible reply text and the quote-block strip
are unchanged.

Fixes #111233
2026-09-15 06:19:19 -07:00
liuhao1024
a019976430 test(slack): address review nits on the subtype allowlist
Drop the unreachable handle_message assertion in the drop tests
(_prefilter_inbound never calls it) and state the deliberate
file_comment drop decision in the allowlist comment.
2026-09-15 04:40:49 -07:00
teo-nex
90aa5e3511 fix(slack): preserve allowed bot posts and canvas mentions 2026-09-15 04:40:49 -07:00
liuhao1024
cac288a0c3 fix(slack): gate inbound turns on a conversational-subtype allowlist
Housekeeping subtypes (channel_join/leave/topic/name/purpose,
convert_to_private/public, pins, deletions) are not a person speaking,
yet _prefilter_inbound only rejected message_changed/message_deleted,
so each of them started a full agent turn in free-response channels.
Replace the denylist with an allowlist: a message passes when subtype
is absent, file_share, thread_broadcast or me_message; everything else
is dropped. Fixes #110778.
2026-09-15 04:40:49 -07:00
teknium1
6a22abe5ba fix: retire clarify cards on an explicit no-answer signal, not the '[' prefix
`_clarify_callback_sync` decided "no answer arrived" by testing whether the
response text starts with '[' (the shape of the timeout / undeliverable
sentinels). A real answer can start with '[' too — a "[A] staging" choice
label picked by number, or "[urgent] ..." free text after Other — so the
clarify resolved and the agent got the answer, yet the Slack card was
rewritten to "This prompt expired" and typing was never re-armed.
`_clarify_send_then_wait` now returns `(response, answered)` and the runner
branches on that flag only.

The Slack click handler popped the retire entry as soon as Other was
clicked, but Other is not terminal: the clarify stays pending for typed
text, so a later timeout or /new reset found nothing to retire and the card
stayed stuck on "Awaiting typed answer". The entry is now popped only on a
terminal outcome (a choice click, or Other on an already-dead entry).

A typed answer to a native card (numeric pick, or text after Other) never
reaches the click handler, so the card kept its buttons forever; the
TEXT_RESOLVED intercept now retires it with the answer.

Review finding: '[' prefix mistaken for the timeout sentinel; Other click dropped the retire entry; typed answers never rewrote the card.
2026-09-15 04:40:04 -07:00
teknium1
71cac9426d fix(gateway): retire native clarify cards on timeout, reset and prose cancel
One adapter-facing seam replaces the Slack-only callback: an adapter whose
clarify prompt is a persistent card (Slack Block Kit) defines
`retire_clarify_card(clarify_id, notice)`, and the gateway calls it from
every path that ends a clarify without a button click:

- TurnRunner._clarify_callback_sync: when the bounded wait returns a
  sentinel (timeout, /new or run-end clear_session), schedule the retire
  with the expired notice on the gateway loop (#110821).
- run_inbound TEXT_REJECTED_PROSE: retire with the cancelled notice before
  the prose is routed as a follow-up (#111019). Lookup is on the adapter
  class so MagicMock doubles cannot fabricate the method; no platform ==
  SLACK special-case.

The Slack map is keyed by clarify_id and popped before the first await, so
a late timer cannot touch a newer prompt and the button handler's ts-keyed
guard makes a racing click a no-op. Gateway-restart-orphaned cards stay
out of scope: nothing is waiting on the new process, and the click path
already renders them expired.

Tests trimmed to invariants: the runner-level timeout probe (card adapter
vs no-card adapter), the inbound prose retire, and one Slack test covering
buttons-dropped + late-click-noop. Docs updated for the new in-place edit.
2026-09-15 04:40:04 -07:00
KoNit-K
a41187e187 fix(gateway): retire Slack clarify cards on prose cancellation 2026-09-15 04:40:04 -07:00
teknium1
1847ad2883 fix(telegram): trim the liveness heartbeat, fix the polling error-callback log lines
- Drop the 900s "inbound liveness" INFO line from the salvage: a periodic
  log heartbeat is a feature with a separate scoping call (#111211 item 3);
  the existing stall watchdog already escalates when getUpdates stops.
- The polling error_callback interpolated the raw exception and left the
  redaction call inside the format string, so its lines read
  "Telegram network _redact_telegram_error_text(error), scheduling
  reconnect: ..." and leaked unredacted text. Redact for real.
- Tests: keep two invariants (recovered wording after network errors;
  clean bootstrap stays "confirmed healthy") and the transport test.
2026-09-15 04:39:16 -07:00
liuhao1024
ff307aea5e fix(telegram): name empty-string transport errors in adapter log lines
httpx timeout exceptions (ConnectTimeout, ReadTimeout, ...) stringify to
"" so every adapter log line built from _redact_telegram_error_text()
ended with a blank reason. Fall back to the exception class name.

Partial salvage of #111222: only the _redact_telegram_error_text hunk;
the transport-layer and polling-recovery hunks are covered by #111221.
2026-09-15 04:39:16 -07:00
fangliquan
e1e943a9df fix(telegram): expose polling transport recovery 2026-09-15 04:39:16 -07:00
teknium1
2298dc8122 fix: propagate caller contextvars into the Feishu adapter-owned executor
Moving the dedup flush and thread lookup from asyncio.to_thread onto the
adapter-owned pool fixed the torn-down default executor, but to_thread also
copies the caller's contextvars and run_in_executor does not. A multiplexed
profile's HERMES_HOME override and secret scope are contextvars, so those
workers silently ran under the launch profile. _run_blocking now runs the call
through contextvars.copy_context().run, matching to_thread semantics.

Review finding: _run_blocking lost the profile HERMES_HOME override / secret scope on the worker.
2026-09-15 04:38:26 -07:00
teknium1
b447554ce5 fix(feishu): thread-reply lookup uses the adapter-owned pool too
`_fetch_last_message_in_thread` was the last hot-path `asyncio.to_thread`
in the adapter: after a default-executor teardown (#111020) thread-reply
routing would fail the same way the dedup flush did. Route it through
`_run_blocking` like every other blocking SDK call. The remaining
`to_thread` users (`_load_lark_oapi` at connect/onboarding, the voice
transcode with its file-attachment fallback) are cold or degrade cleanly.
2026-09-15 04:38:26 -07:00
JackJin
5fbef868a9 fix(gateway): keep Feishu websocket off default executor 2026-09-15 04:38:26 -07:00
liuhao1024
fa2c72746c fix(feishu): run the inbound dedup flush on the adapter-owned pool
A dead background event loop tears the loop's default executor down, and
after that every inbound message was dropped inside the dedup gate with a
RuntimeError out of asyncio.to_thread — the adapter went permanently deaf
while the gateway process, websocket and service all stayed healthy.

#10849 already moved the outbound SDK calls onto an adapter-owned,
self-healing pool; this gives the inbound dedup-state flush the same
treatment, so a default-executor teardown can no longer wedge message
intake.
2026-09-15 04:38:26 -07:00
fangliquan
992b517fdd fix(photon): preserve Unicode NDJSON separators 2026-09-15 04:34:13 -07:00
wang2
fa12d7556c fix(telegram): allow opt-in CJK rich messages 2026-09-15 04:24:53 -07:00
teknium1
23036e20a6 fix(ux): plain-language, actionable user-facing messages (core)
Squashed integration of the user-facing message audit for this surface set.
Full per-finding receipts: /tmp/ux-audit/lanes/*-receipt.md (campaign artifacts).
2026-09-15 04:12:13 -07:00
teknium1
624777617c fix(discord): filter obfuscated channels on the explicit-id backfill path too
Review finding on #90154: only the wildcard ("*") branch of
_iter_missed_message_backfill_candidates skipped obfuscated channels. A
channel the bot previously talked in that later lost VIEW_CHANNEL still
entered the explicit-id branch and produced a failing history read on
every startup. Apply the predicate once, after both branches build the
candidate list, and import it at module level like the other helpers.

(During the rebase the predicate was also tightened, resolved in the original
commit: is_discord_channel_obfuscated catches AttributeError (the precise
expected failure) instead of a bare Exception so a genuine attribute bug
is not hidden behind the name fallback, and its docstring states the
deliberate bias of the sentinel-name fallback (a visible channel literally
named ___hidden___ is also skipped on discord.py builds without the flag).)
2026-09-15 03:50:10 -07:00
Teknium
1f34742552 fix(discord): skip obfuscated channels in directory and backfill enumeration
Discord's Channel Obfuscation change (announced Aug 12 2026, HTTP
enforcement Nov 16 2026) dispatches channels the bot lacks VIEW_CHANNEL
on with name "___hidden___", flag 1 << 17 (CHANNEL_OBFUSCATED), and
nulled fields. Without filtering, the channel directory lists phantom
"___hidden___" entries the agent can never post to, and wildcard
missed-message backfill wastes history reads on channels that always
403.

Adds is_discord_channel_obfuscated() to gateway/platforms/helpers.py
(checks the flag bit plus the sentinel name for discord.py builds that
don't expose the new flag) and applies it at both enumeration sites.
2026-09-15 03:50:10 -07:00
teknium1
7abe9502ee fix(kanban): worker liveness and kills require the spawn-time start fingerprint
tasks.worker_pid outlives a reboot; afterwards the number can belong to any
process. _pid_alive answered from bare existence, so reclaim_stale_claims kept
extending the claim of a "live" stranger (stuck running task), and
enforce_max_runtime / _terminate_reclaimed_worker SIGTERM'd then SIGKILL'd it.

_set_worker_pid now records gateway.status.get_process_start_time(pid) as
tasks.worker_started_at (additive column, NULL on legacy rows). _worker_alive
(pid, started_at) is the liveness check every reader uses (reclaim, defer,
reconcile, crash sweep, max-runtime, archive, reopen invalidation); a live pid
whose fingerprint disagrees is a recycled PID: treated as dead, never
signalled (termination reports pid_recycled). Legacy rows without a
fingerprint keep the existence answer until their next spawn.

Row hermes_cli/kanban_db_dispatch.py:458 (lane4_high) confirmed by tracing:
the SIGKILL at :461 was gated only on _pid_alive.
2026-09-15 03:47:15 -07:00
teknium1
8b181940b4 fix(multiplex): background threads and teardown paths carry the turn's profile scope
The profile scope (HERMES_HOME override, secret scope, terminal policy) is a
contextvar bundle bound per turn. A bare threading.Thread / Timer / gRPC
callback starts with an empty context and resolves the LAUNCH profile:

- agent/title_generator.py: the auto-title thread read
  auxiliary.title_generation (model, language, provider key) from the default
  profile's config and billed the default's key for a secondary's session.
  Spawn via agent.memory_provider.spawn_context_thread (copy_context).
- tui_gateway/session_lifecycle.py: every teardown caller is a bare Timer
  (ws-orphan reap), the idle-reaper thread, atexit _shutdown_sessions, the
  session.close pool RPC, superseded_by_resume or compute_host flush - none
  carries a scope, yet on_session_end / commit_memory_session / agent.close ->
  shutdown_memory_provider read the provider's config + credentials at call
  time. Under multiplex they failed closed (tail never committed, #110622
  class); on the Desktop backend a secondary's transcript went to the launch
  profile's memory tenant. _finalize_session and _teardown_session now bind
  _session_profile_runtime_scope(session) around those blocks, which covers
  every spawn site through the single chokepoint.
- plugins/platforms/google_chat/adapter.py: Pub/Sub callbacks run on the gRPC
  SubscriberClient's threads and run_coroutine_threadsafe copies THAT empty
  context onto the loop task, so _dispatch_message and everything under it
  (attachment cache, per-user OAuth token store via _acquire_user_chat_api ->
  _load_per_user_chat_api, TTS keys, delivery ledger, bot-id cache) resolved
  the launch profile. connect() captures its scope; _on_pubsub_message and
  _submit_on_loop run under a per-callback copy of it.

spawn_context_thread gains a kwargs passthrough for the title thread's
callbacks.
2026-09-15 03:47:15 -07:00
teknium1
d1794d5539 fix(multiplex): children spawned for a served profile start from that profile's env
Under gateway.multiplex_profiles (and the Desktop/dashboard backend serving named
profiles) os.environ holds the LAUNCH profile's .env. Five spawn sites built a
child's env from it while acting for another profile, so the child saw the
launch profile's HERMES_HOME (bot_relay, key_cmd), its credentials, HERMES_MODEL
and TERMINAL_* policy, and none of the served profile's own .env:

- tui_gateway/server.py _SlashWorker: pinned HERMES_HOME but kept the launch
  base with tier-2 credentials + settings.
- tools/bot_relay.py delivery_env (relay RPC + --run-delivery): dict(os.environ).
- tools/browser_tool.py _build_browser_env: re-added BROWSERBASE/FIRECRAWL/
  BROWSER_USE keys from os.environ after the scrub.
- plugins/platforms/a2a/adapter.py _forward_to_profile: {**os.environ}.
- agent/command_token_source.py _mint: key_cmd helper inherited os.environ.

tools.environments.local.served_profile_child_env is the one builder: pin the
target home, drop the launch profile's .env residue and bridged TERMINAL_*
(strip_launch_profile_env), and for children that legitimately run with the
profile's credentials (agent worker, token helper) overlay the target profile's
own secrets - what a standalone `hermes -p X` loads itself, never a sibling's.
The browser keeps the provider scrub and re-adds only its passthrough keys via
get_secret. Outside multiplex the env is unchanged.

Live proof from inside the child (launch A, served B, multiplex on): all five
children print HERMES_HOME == B, see B_MARKER=b from B's .env and do not see
A_MARKER; the browser child gets B's FIRECRAWL_API_KEY. On base every one leaked
A_MARKER and lacked B_MARKER; bot_relay and key_cmd also had A's HERMES_HOME.
2026-09-15 03:47:15 -07:00
Sudhir Patil
ec11359b44 fix(discord): clear fatal status on successful reconnect
DiscordAdapter.connect() set self._running = True directly instead of
calling self._mark_connected(), unlike every other platform adapter
(Telegram, WeCom, Matrix, Feishu, Google Chat, IRC, LINE, Mattermost,
ntfy, photon, raft, simplex, a2a, buzz, dingtalk, whatsapp).

_mark_connected() clears _fatal_error_code/_fatal_error_message/
_fatal_error_retryable and rewrites the runtime status file as
"connected". Bypassing it meant a transient connect failure (e.g. a
one-off DNS blip: "Cannot connect to host discord.com:443 ssl:default
[Temporary failure in name resolution]") left the platform reported as
permanently fatal in gateway_state.json / the dashboard, even after the
adapter successfully reconnected and was actively serving messages for
hours.

Reproduced under gateway.multiplex_profiles: true with a secondary
profile's Discord bot (sarathi:discord) — the bot reconnected
repeatedly ("Connected as ..." logged many times over 18+ hours) while
the dashboard kept showing the original fatal error the entire time.

Fixes #102554.

Added a regression test asserting connect() clears a previously
recorded fatal error.
2026-09-15 03:44:36 -07:00
teknium1
fd303c0137 fix(gateway): 'decline' survives the config.yaml load path; Telegram forwards it; wizard offers it
config_loader._dm_behavior_choice still normalized against {"pair","ignore"},
so `unauthorized_dm_behavior: decline` in config.yaml (top level or a
platform block) was coerced back to "pair" on the real startup path
(load_gateway_config), and `unauthorized_dm_decline_message` was never
bridged into gw_data. Both now go through gateway.config.UNAUTHORIZED_DM_BEHAVIORS
(single source) and the presence bridge. The round-trip test exercises
load_gateway_config with a real config.yaml (top-level decline, telegram
override, custom message) instead of GatewayConfig.from_dict.

Telegram's intake prefilter only forwarded unauthorized DMs when the
behavior was exactly "pair", so with an allowlist configured a decline was
never sent. Anything that needs an outbound reply (!= "ignore") passes.

`hermes gateway setup` gains a "Politely decline unknown senders" choice
that writes platforms.<platform>.unauthorized_dm_behavior: decline; docs
mention it. Upstream-source references dropped from docstrings.
2026-09-15 03:44:13 -07:00
Teknium
5705b68f70 fix: memory-plugin and Qwen-CLI config JSON survives Windows BOM
Port from earendil-works/pi#8337 (UTF-8 BOM normalization in text inputs):
sibling sites the merged #81967 BOM sweep missed. json.loads hard-fails on
a leading U+FEFF and every one of these loaders swallows the exception and
silently falls back to defaults — a user who edited mem0.json, honcho.json,
hindsight/config.json, or supermemory.json in Notepad lost their whole
config with no error, and Qwen CLI OAuth creds saved with a BOM raised
qwen_auth_read_failed.

- plugins/memory/{honcho,mem0,hindsight,supermemory}: 13 read sites -> utf-8-sig
- hermes_cli/auth.py: _read_qwen_cli_tokens -> utf-8-sig
- tests: BOM regression tests per loader (sabotage-proven) + plain-UTF-8 guard
2026-09-15 03:38:29 -07:00
teknium1
d8053f4806 fix(video_gen): cap LTX 2.5 at 10s for 1440p/2160p; omit unset enum durations
fal's LTX 2.5 fast endpoints accept 6-20s only up to 1080p — "At 1440p and
2160p, all frame rates support up to 10 seconds" — so a 4K request with the
family's 20s ceiling was rejected by the vendor. Families can now declare
`duration_cap_by_resolution`, applied after the enum snap / range clamp on the
resolved resolution enum.

An unset duration on a duration_enum family also snapped to enum[0] (6s),
silently overriding the endpoint's own "auto" default; None now omits the key
for enum families exactly as it already did for range families.

test_managed_media_gateways asserts the alibaba/happy-horse/ namespace by
prefix rather than the exact v1.1 literal so the next version bump doesn't
flip an unrelated gateway test.
2026-09-15 03:37:11 -07:00
teknium1
1c1980dc1f fix(video_gen): make the duration enum explicit instead of sniffing tuple shape
`durations` carried two meanings told apart only by len==2 and gap>1: a
(min, max) range to clamp, or an enum to snap. A family with exactly two
legal values would have been misread as a range (review finding on #91311).
`durations` is now always the (min, max) window (what capabilities()/
list_models() read) and families with discrete values add `duration_enum`;
_clamp_duration takes the family and branches on the key, not the shape.

Also: restore the exact v1.1 endpoint assertion in the gateway namespace
test (a startswith/endswith check would not catch a silent version drift),
add the ltx-2.5 i2v snap case, and keep happy-horse on audio_native (the
schema test forbids audio+audio_native together, and v1.1 audio is always on).
2026-09-15 03:37:11 -07:00
Teknium
37286d3064 feat(video_gen): LTX 2.5 + Kling O3 families; Happy Horse upgraded to v1.1
Adds two new FAL video families and upgrades one:

- ltx-2.5 (cheap tier): lightricks/ltx-2.5/{text,image}-to-video/fast.
  Lightricks' open-source audio-video model. Native audio, 6-20s integer
  duration enum, 720p-2160p (i2v), $0.09/s at 720p. duration_int + 2k/4k
  resolution aliases; no seed key in the schema.
- kling-o3 (premium tier): fal-ai/kling-video/o3/standard/{text,image}-to-video.
  Kuaishou's frontier multi-shot model, 3-15s, optional native audio
  ($0.084/s off, $0.112/s on). String durations, i2v drops aspect_ratio,
  no seed/resolution keys.
- happy-horse upgraded from the sparse-docs 1.0 endpoints to
  alibaba/happy-horse/v1.1/{text,image}-to-video with the full published
  schema: nine aspect ratios, 720p/1080p, 3-15s integer durations, seed
  supported, audio native (no generate_audio key), i2v drops aspect_ratio.

All flags derived from each endpoint's llms.txt schema. Payload builder
asserted locally against the schemas; test for the old Happy Horse
"prompt-only" contract updated to pin the v1.1 schema, plus new payload
tests for ltx-2.5 and kling-o3.
2026-09-15 03:37:11 -07:00
kshitijk4poor
d793f7b9fb fix(supermemory): explicit failure sentinel and bounded pending-turn buffer
Two review findings on the salvage stack:

- _write_turns' _quietly consolidation keyed failure on a None result,
  implicitly assuming add_memory never legitimately returns None. A None-
  returning stub (the most common mock idiom) would mark every successful
  write failed and re-append the batch forever. Module-level _FAILED
  sentinel: only a raised exception re-queues.
- _pending_turns had no bound: a persistently failing service accumulated
  one entry per turn for the process lifetime (gateway runs never
  re-initialize), and every retry re-sent the whole accumulated payload —
  O(n^2) upload bytes, and a size-rejected batch could never shrink. Cap
  the buffer (50 turns / 256 KiB, drop oldest, warn once per trim).
2026-09-15 11:55:10 +05:30
kshitijk4poor
bc1d4776db docs(supermemory): purge stale session-ingest / /v4/conversations references
The per-turn capture rewrite removed the raw urllib /v4/conversations
ingest, but stale references survived outside the diff hunks:

- README "Behavior" still carried the "written once via the conversations
  endpoint" paragraph contradicted by the new bullets right above it.
- website memory-providers (EN + zh-Hans) still listed full-session ingest,
  session-end /v4/conversations ingest, and ingest in the base-url and
  api_timeout rows — the PR had updated one line per file but missed the
  rest of the section.

Now every surface describes per-turn documents.add capture with retry.
2026-09-15 11:55:10 +05:30
kshitijk4poor
3cad3c6db8 fix(supermemory): route capture-write logging through _quietly; drop dead default in custom id
Two cleanups on the salvaged per-turn capture path (review of #109359):

- _write_turns hand-rolled the try/except/log that the file's own _quietly
  helper already provides (same exception class, level, exc_info shape);
  the deleted _ingest path used _quietly for the identical call. Consolidate:
  add_memory always returns a dict, so a None result is the failure sentinel.
- _capture_custom_id's `or 'hermes'` never fires: _sanitize_tag already
  returns _DEFAULT_CONTAINER_TAG on empty input, and the literal duplicated
  the constant the helper owns. Verified behavior-preserving (tests green
  with the fallback artificially restored).
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
c6da9e0788 fix(supermemory): serialize capture writes and document at-least-once retry
sync_turn runs on the MemoryManager worker, but on_session_switch and
shutdown run on the caller thread, so two _write_turns() calls could
snapshot the same pending batch and each replace the whole list (duplicate
append or lost pending turn). A capture lock now covers the
snapshot/write/replace sequence; _write_turns() reads pending turns inside
the lock instead of taking a caller-built list.

Retries are at-least-once: the documents API appends on a shared custom_id
and does not dedupe by content, so a write it accepted but whose response
was lost is appended again. Stated in the docstring and README instead of
implied.

Addresses review on #109359.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
1b76cfff83 fix(supermemory): keep pending turns across session switch until written
A failed flush at on_session_switch() restored the pending buffer and
then cleared it on the next line, so an unavailable service at the switch
boundary still lost every pending turn. Pending turns now carry their own
session_id; _write_turns() batches per session, so a later retry (next
turn, session end, shutdown) writes old-session turns under the old
session's custom_id even after the switch.

Addresses review on #109359.
2026-09-15 11:55:10 +05:30
Mahesh Sanikommu
03627dbf95 fix(supermemory): write turns via documents API instead of session-end conversations ingest
sync_turn now writes each completed turn through the SDK's documents.add,
keyed by custom_id "<session>_<date>_b<0-5>" so all turns of a session in
one 4-hour window append to a single document. This matches the capture
shape of the other Supermemory agent integrations and removes the raw
urllib POST to /v4/conversations, which the self-hosted server does not
implement (#101270).

Failed turn writes stay pending and are retried with the next turn, at
session end, on session switch, and at shutdown. Previously a failed
session-end ingest was logged once and the whole session was lost.

Inline base64 data URIs in captured text are replaced with "[image]" so
pasted screenshots no longer land in the document as megabytes of text.

Metadata stays type/session_id/timestamp plus the existing sm_source.
2026-09-15 11:55:10 +05:30
kshitijk4poor
82d61165d0 style(matrix): separate _MATRIX_PERMANENT_ERRCODES from the preceding function
The stack inserted the errcode table directly after _strip_reply_fallback's
return with no blank lines, which reads as if the constant belongs to the
function body and trips E305. Two blank lines restore the module-level
boundary; no behaviour change.
2026-09-15 10:50:45 +05:30