Commit Graph

2100 Commits

Author SHA1 Message Date
ethernet
8b56a9306f refactor(pm): own group-only builds and finish dependency hints 2026-09-11 18:25:05 -04:00
ethernet
6ee8610b67 refactor(pm): finish consumer contracts and enforce private engine imports 2026-09-11 18:18:58 -04:00
ethernet
80273b4507 refactor(pm): route Python tool installs through PM
Use PM-selected interpreters and tool entrypoints for Browser Use, Hindsight and Python language servers. Sync the declared Google Chat extras instead of changing the active environment with pip.

Verified the affected 11-file Nix test subset: 359 passed, 7 skipped. The full suite and real third-party package installation were not run.
2026-09-11 18:05:43 -04:00
ethernet
284dbaf537 fix(pm): isolate bootstrap dependencies and unify YAML on ruamel
Activation reaches plugin discovery before the application dependencies
exist. Give PM its own locked Python project and runtime so it can install
or repair the application without importing that dependency tree.

Keep PM outside the application workspace. A shared uv workspace resolves
the application graph and cannot provide this isolation. Route mutations
through an isolated worker and preserve transaction callbacks, cancellation,
custom package registrations, and correlated receipts.

Use the same runtime builder for source installs and packaged payloads.
Keep offline wheelhouse support in that builder. Nix builds the independent
PM lock as a separate derivation. Refuse lazy-disabled bootstrap before
installing tools or dependencies.

Move first-party YAML readers and writers to ruamel. Keep the application
lock's transitive PyYAML requirements for third-party packages.

Verification:
- Focused canonical Python suite: 177 passed, 1 host-gated skip.
- Electron backend probes: 12 passed. Electron typecheck passed.
- Both uv locks, scoped lint, Bash syntax, and whitespace checks passed.
- Cold activation, corrupt-app repair, offline staging, and relocation ran.
- Built and exercised the Nix PM runtime and standalone YAML merge script.

Six broader caller test files retain the same 24 failing test IDs as an
archive of HEAD. The existing real-home guard blocks those tests before
they can exercise the affected paths. No full-suite pass is claimed.
Native Windows signing and full Bionic package execution remain unverified.
2026-09-11 12:23:51 -04:00
ethernet
b3bfc3afe5 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	apps/desktop/electron/backend-connection-state.test.ts
#	apps/desktop/electron/backend-connection-state.ts
#	apps/desktop/electron/backend-exit.test.ts
#	apps/desktop/electron/main.ts
#	apps/desktop/electron/pool-spawn-coordinator.test.ts
#	apps/desktop/electron/pool-stop.ts
#	apps/desktop/electron/preload.ts
#	apps/desktop/src/app/settings/about-settings.tsx
#	apps/desktop/src/app/updates-overlay.tsx
#	apps/desktop/src/global.d.ts
#	apps/desktop/src/store/notifications.ts
#	apps/desktop/src/store/updates.ts
#	gateway/config_loader.py
#	hermes_cli/banner.py
#	plugins/platforms/dingtalk/adapter.py
#	tests/hermes_cli/test_plugins_cmd.py
#	tests/test_live_system_guard.py
#	tui_gateway/server.py
#	website/docs/user-guide/desktop.md
2026-09-11 09:36:04 -04:00
Teknium
cbd03e6e4c fix(gateway): secondary-profile adapters no longer inherit the default's allow-all / allowlists
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's .env. Several
adapter-owned authorization gates still read GATEWAY_ALLOW_ALL_USERS, GATEWAY_ALLOWED_USERS
or their platform allowlist/allow-all raw from os.environ, so the default profile opting
into open access opened every secondary email/QQ/WhatsApp/Matrix/Teams/Slack/LINE/DingTalk
bot to any sender (email additionally skipped From: authentication), the default's Matrix
allowlist decided who may approve tool calls on a secondary bot, and a secondary that
opted in only in its own .env was silently deny-all.

Every such read now goes through the adapter's existing module-local scoped reader
(gateway.platforms._shared.get_scoped_secret / matrix _startup_env_secret): profile
scope first, scoped miss = default, never os.environ; the unscoped default-profile and
single-profile paths keep the environ read, where it IS the profile's own value.

Sites: email _allow_all_senders/_allowlist_in_effect; qqbot _open_dm_opted_in;
whatsapp_common _open_dm_opted_in/_live_dm_allow_from; teams _card_action_denied;
matrix _is_authorized_user, MATRIX_ALLOWED_USERS, MATRIX_IGNORE_USER_PATTERNS,
_extra_csv_set (allowed/free-response rooms); slack _slack_allow_bots/_slack_api_human_users;
line _truthy_env/allowlist (allow-all, user/group/room allowlists); dingtalk _extra_get
(allowed_users/chats, free-response chats, require_mention).

Live repro (temp HERMES_HOME, multiplex on, default env GATEWAY_ALLOW_ALL_USERS=true,
secondary scope without opt-in): EmailAdapter._allow_all_senders() True -> False,
QQAdapter._open_dm_opted_in() True -> False, Matrix _is_authorized_user('@stranger')
True -> False, Teams card action allowed -> denied.

Co-authored-by: Drexuxux <drexux0@gmail.com>
Co-authored-by: MoonsvnLyn <FirmamentalSpring@users.noreply.github.com>
Co-authored-by: svector-anu <anuoluwakolapo94@gmail.com>
Co-authored-by: babatorik <durgun.ismail@gmail.com>
Co-authored-by: salch-cred <salch-cred@users.noreply.github.com>
2026-09-11 02:24:55 -07:00
shannonsands
a6ee31f55a feat(wisdom): add Hermes Collective Wisdom Agent V1 (#94266)
* feat(wisdom): add trusted publish and install foundation

* feat(wisdom): add private contribution loop

* feat(wisdom): add managed consumption workflows

* fix(wisdom): close cross-repository safety gaps

* fix(wisdom): align local package and lifecycle policy

* fix(wisdom): require explicit profile setup

* docs(wisdom): repin reconciled gateway head

* fix(wisdom): fence content downloads and approval receipts

* docs(wisdom): record generation-fenced downloads

* docs(wisdom): record unified delivery PR

* fix(ci): stop passing invalid classifier inputs

* docs(wisdom): remove internal requirements ledger

* feat(wisdom): localize dashboard and desktop copy

* feat(wisdom): complete local contribution and consumption UX

* style(wisdom): satisfy desktop lint

* chore(wisdom): refresh requirements pin

* test(dashboard): allow formatted profile copy

* test(wisdom): stabilize desktop interaction coverage

* fix(wisdom): surface dashboard action failures

* fix(wisdom): add repeatable Portal demo login

* feat(wisdom): add actionable skill notifications

* feat(wisdom): add notification install and update actions

* fix(wisdom): make Telegram skill alerts actionable

* fix(wisdom): always refresh demo Agent login

* feat(wisdom): embed Telegram notification actions

* fix(wisdom): preserve Telegram notifications after actions

* fix(wisdom): keep Telegram notification cards readable

* feat(wisdom): add Telegram candidate approval flow

* feat(wisdom): explain Telegram qualification reasons

* fix(wisdom): reconcile cross-surface candidate actions

* feat(telegram): add Collective Wisdom management command

* chore(wisdom): refresh Gateway contract pin

* chore(wisdom): advance Gateway contract pin

* feat(wisdom): align command UX across clients

* feat(slack): add Collective Wisdom management parity

* feat(wisdom): add security and professionalism reviews

* feat(wisdom): add first-time qualification guidance

* feat(wisdom): simplify qualification sharing choices

* feat(skills): add optional editorial metadata

* feat(wisdom): enrich legacy skill presentation

* fix(wisdom): harden review and update boundaries

* fix(wisdom): emit canonical review timestamps

* fix(wisdom): align with merged gateway and main

* wisdom: add agent-led sharing core (policy, evidence, schemas, templates, delivery, weekly job, share/install flows)

- hermes_wisdom/agent_led/: policy resolution (server > local > defaults),
  7-day evidence builder that excludes bundled/hub/managed skills and
  dismissed/handled/recently-suggested content hashes, strict pydantic
  schemas for agent output with repair-or-reject, fixed copy templates
  (Share / Teammate / Published / Update / Mute), idempotent retried
  delivery ledger with stale-action resolution, weekly review job,
  resumable Share and Install flows.
- prompts/: candidate review, recipient recommendation, share packaging.
- tests/wisdom/test_agent_led.py: 30 tests.

* wisdom: agent-led renderers and button action dispatcher

- render.py: Telegram HTML, Slack blocks, Desktop payload; editorial name
  is the emphasized line, product label stays separate.
- actions.py: resolve opaque wa:<action>:<dedup> targets via the delivery
  ledger; Not now -> dismissal, Mute -> fixed options, Share -> resumable
  packaging flow, Install/Update -> plan command. Never publishes/installs.

* wisdom: CLI verbs, agent_led config default, conversational catalog skill

- hermes wisdom browse/review-week/act/share/dismiss/mute (all --json).
- wisdom.agent_led config block, default enabled.
- SKILL.md rewritten so natural-language catalog questions map to the CLI
  verbs, share/install flows and fixed notification templates.

* wisdom: wire agent-led weekly review into gateway tick and Telegram buttons

- gateway housekeeping tick calls maybe_run_weekly_review with a home
  channel sender when a Telegram adapter is available.
- Telegram: wa: callbacks resolved through the ledger (stale-safe), mute
  duration keyboard, send_wisdom_agent_recommendation rich card + fallback.

* fix(wisdom): integrate local mediation and harden model and setup boundaries

* fix(wisdom): honor authoritative recommendation policy and defer on failure

* fix(wisdom): synchronize opaque suppression and recheck delivery preferences

* feat(wisdom): route weekly selection through the session-owned assessment queue

* fix(wisdom): prepare and submit the reviewed generated share package

* feat(wisdom): separate native Share preparation from publication consent

* feat(wisdom): sync native mute choices through a leased preference outbox

* feat(wisdom): bind native mute controls to durable preference choices

* feat(wisdom): add scoped desktop and dashboard notification settings

* fix(wisdom): revalidate feed recommendations before assessment and delivery

* fix(wisdom): persist validated delivery receipts before completing notices

* feat(wisdom): add private notification claim and receipt client

* Persist Wisdom send reservations and recover delivery acknowledgements

* Route legacy Wisdom controls through current native review

* Add typed private Wisdom operation outcome client

* fix(wisdom): make agent-led advice usable in the local demo

* fix(wisdom): keep requested consent outside proactive limits

* fix(wisdom): distinguish unavailable assessments and preserve digest text

* fix(wisdom): assess ongoing usefulness beyond the current task

* fix(wisdom): restore immediate qualification sharing controls

* fix(wisdom): separate qualification review from installation advice

* fix(wisdom): collapse review checklists and simplify sharing copy

* fix(wisdom): show compact sharing progress and publication receipts

* fix(wisdom): require credential prefixes rather than matching skill names

* fix(wisdom): finish package checks before presenting sharing consent

* fix(wisdom): scan local skills before qualification cards

* fix(wisdom): update moderation results on existing sharing cards

* fix(wisdom): keep sharing review accessible from receipt cards

* fix(wisdom): align mediated review cards and collapsible checks

* fix(wisdom): clarify clean security summary wording

* fix(wisdom): normalize consent plans and add explicit recheck

* fix(wisdom): keep install and update receipts concise

* fix(wisdom): collapse assessments and deduplicate operation cards

* fix(wisdom): restore private Portal review from native cards

* fix(wisdom): sync Portal publication to original consent card

* fix(wisdom): show local skill version on sharing cards

* fix(wisdom): skip agent recommendations for self-published versions

* fix(wisdom): simplify candidate notices and local-edit recovery copy

* feat(wisdom): submit locally reviewed packages with one confirmation

* feat(wisdom): expose safe receipt and outcome sync recovery

* wisdom: onboarding notice says detect and share, names the user's own skill

Copy review from the product owner on the first and returning
qualification notices (fixed delivery mode):
- the feature blurb now says the org enabled detection *and sharing*
- both notices say the detected skill is one the user created
- both close with an exclamation mark

Applied identically to hermes_wisdom.notice, the desktop and web i18n
strings, and the tests that assert the sentences.

* wisdom: one opener, no approval line, ask to share after the skill is shown

Product owner review of the candidate card.

- The Hermes written card now opens with the same sentence as the fixed card
  ("Your organisation has enabled Collective Wisdom, a feature designed to
  automatically detect and share useful skills across all team members.")
  instead of its own blurb, so there is one first time message.
- "Nothing is shared without your approval." removed from Telegram, Slack
  and Desktop. The buttons already make the permission explicit.
- "Would you like to share?" no longer appears before the skill is named.
  It is now the last line, after the skill name, description, why suggested
  and the checks, and reads "Would you like to share it?" (matching the
  agent led template wording).

Tests updated for the new order; proposalNotice removed from all desktop locales.

* wisdom: American spelling, organization

Product owner decision: user facing copy uses American spelling.
Changes "Your organisation" to "Your organization" in the chat notice,
the Hermes written card opener, the desktop and web strings, and the
tests that assert them. Identifiers such as nas_organisation:* and the
German and French locales are untouched.

* wisdom: candidate card copy round 4 (owner review)

Apply the product owner's round 4 copy decisions to the Hermes Collective
Wisdom candidate card on Telegram, Slack, Desktop and the shared views:

1. Hermes-written cards are titled "Hermes Collective Wisdom" instead of
   the bare "Collective Wisdom".
2. The "Reusable skill ready to review" line is gone from the candidate
   card (Telegram rich card and plain fallback, legacy agent-led share
   template).
3. The skill name and description are labelled: "Skill name: <name>" and
   "What it does: <description>" (Telegram, Slack, Desktop).
4. "Why suggested:" is now "Why others might benefit:".
5. A passing professionalism review reads "Safe to share at work ✓ (no
   inappropriate content found)" with no per-check bullets and no "Pass";
   a failed review reads "Needs a look before sharing at work (possible
   inappropriate content)" and lists only the checks that flagged
   something. Pending/unavailable wording is unchanged.
6. Telegram button toasts: "Will ask later...", "Preparing more
   details...", "Sharing...".
7. Qualification reasons: "You used this skill consistently across many
   days." and "You've really refined this skill."
8. prompts/wisdom_candidate_review.md asks for a compelling
   editorial_name, a simple one_line_description and a compelling
   why_coworkers_benefit under 300 characters; "Be concise and
   convincing." becomes "Be concise and compelling: the goal is that the
   user wants to share it."

Tests updated for the new strings; review_text() gains direct coverage.

* wisdom: re-apply owner copy after rebase

- Native share cards (advice_view/interaction_view): drop the approval line, ask "Would you like to share it?" as the last line after the checks
- Hermes-written completion card titled "Hermes Collective Wisdom"
- Qualification reasons use the owner wording (consistently across many days / really refined)
- American spelling (organization) in remaining English copy
- Desktop test asserts the current Share button; web test matches the returning notice

* fix(wisdom): pin reconciled Gateway and verify Unicode hash vectors

Pin Gateway 60cd2d6b613ae3cd4a6e65155d1142006d907e78 and byte-identical producer artifacts. Verify every content-order case and package-manifest binding. Validation: 186 focused Python tests, Ruff and contract verifier.

* fix(wisdom): reconcile optional SDK tests and frontend lint

* fix(wisdom): default to agent-written notification summaries

* fix(wisdom): restore deferred install review and browse controls

* feat(wisdom): inspect installed setup with exact package provenance

* feat(wisdom): run native-approved installed setup steps with durable evidence

* fix(wisdom): recover interrupted setup with explicit native consent

* feat(wisdom): hand native installs into guided setup review

* fix(wisdom): continue requested setup with fixed notification copy

* fix(wisdom): preserve setup while waiting for a session model

* fix(wisdom): expose canonical setup review controls on desktop

* fix(wisdom): resume setup after recorded automatic updates

* fix(wisdom): make missing setup prerequisites recheckable

* chore(wisdom): align Agent with verified Gateway contract

* fix(wisdom): stop guessing team slugs in portal links

* fix(wisdom): retire pending advice on account sign-out

* fix(wisdom): cancel advice after terminal account revocation

* fix(wisdom): fence feed responses across account sign-out

* fix(wisdom): checkpoint signed-out feed before reactivation

* fix(wisdom): link proactive advice to scoped notification settings

* fix(wisdom): coalesce queued publication recommendations by version

* fix(wisdom): keep package review navigation local and deferable

* fix(wisdom): reflect installed state in discovery controls

* fix(wisdom): show exact checks before command confirmation

* chore(wisdom): pin bounded analytics privacy contract

* chore(wisdom): pin retired legacy notification contract

* feat(wisdom): review publisher usage with exact sharing copy

* fix(wisdom): align discovery and review check summaries

* fix(wisdom): show expired consent before confirmation

* fix(wisdom): require fresh review for legacy install controls

* fix(wisdom): preserve review expiry across check toggles

* fix(wisdom): retain update policy in native install reviews

* fix(wisdom): surface failed native card edits

* fix(wisdom): persist local command approval reviews

* fix(wisdom): use saved approvals for messaging commands

* test(wisdom): provide scan result in setup handoff fixture

* test(wisdom): exercise Telegram approvals with saved review state

* fix(wisdom): retain suppression policy for offline deferral

* fix(wisdom): reconsider candidates after deferred suppression expires

* fix(wisdom): bind review checks and report verified readiness separately

* fix(wisdom): persist accepted publication intent and recover exact outcomes

* fix(sync): pin UTF-8 tree ordering across writers

* chore(wisdom): pin organisation-scoped Gateway authorization

* fix(wisdom): restrict consent delivery to user-facing sessions

* chore(wisdom): refresh reviewed Gateway contract pin

* fix(wisdom): preserve kept tools in Blank Slate exclusions

* test(auth): reset anonymous fixture with a profile-scoped cache

* fix(wisdom): gate local surfaces and work on current profile entitlement

* fix(wisdom): invalidate quiet tool cache on entitlement changes

* test(wisdom): authorize local consent gateway fixtures

* fix(wisdom): keep entitlement decoding free of native crypto imports

* test(wisdom): provide local entitlement to demo CLI subprocess

* ci: leave upstream workflow unchanged in Wisdom PR

* fix(wisdom): ship package and contracts in Nix wheels

---------

Co-authored-by: hbizi <36184542+hbizi@users.noreply.github.com>
2026-09-11 19:04:06 +10:00
Teknium
45a6101f36 fix(gateway): secondary-profile send_message, notices and /loop wakeups go out via their own bot
Under gateway.multiplex_profiles a turn running for a secondary profile P resolved
its live adapter by bare platform from runner.adapters — the DEFAULT profile's map —
so P's send_message tool calls (send/react/media on slack, matrix, wecom, buzz, ntfy,
every plugin platform), its "Gateway shutting down/restarted" and /update notices,
its /loop wakeups, and its Discord bot's unauthorized-slash operator alert all left
through the default bot (Telegram DMs landed in the user's chat with the other bot).

Every such door now resolves through the profile-aware, fail-closed resolver already
used by the inbound reply path (authz_mixin: _adapters_for_profile /
_authorization_adapter / _adapter_for_source): P's own adapter, or None → a clear
error, never the default bot.

- tools/send_message_senders.py::_live_adapter — resolve via
  runner._authorization_adapter(platform, get_active_profile_name()); shared by
  _send_via_adapter, _handle_react/unreact, media sends, matrix E2EE fast path and
  the WeCom standalone sender.
- gateway/authz_mixin.py — extract _adapters_for_profile (the whole map, for relay-
  aware resolve_delivery_transport callers); _authorization_adapter reuses it.
- gateway/run_shutdown.py — shutdown/restart notice for a running session uses the
  session's source transport / agent:<profile>: key lane, never self.adapters.
- gateway/slash_commands.py + run_notifications.py — /restart and /update markers
  persist `profile`; the restart notice, update result and update prompt resolve the
  requester's own adapter (legacy markers fall back to the session_key lane).
  /goal, /heartbeat, /approve, /deny confirmations use the source's own transport.
- gateway/slash_commands_goals.py + run_goals.py — /loop persists `profile` in its
  route; the wakeup watcher scans every served profile's store under its own scope
  (same shape as _handoff_watcher) and fires through that profile's adapter map.
- plugins/platforms/discord/adapter.py::_notify_unauthorized_slash — alert stays in
  the owning profile (its adapters and its home channels).
- plugins/platforms/wecom/adapter.py::_standalone_send — via _live_adapter.
- docs: website/docs/user-guide/multi-profile-gateways.md (outbound identity).

Tests (red on base): tests/tools/test_send_message_multiplex_profile_adapter.py,
tests/gateway/test_multiplex_notice_egress_profile_adapter.py.

Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com>
2026-09-10 18:17:22 -07:00
Teknium
e253ea0233 fix(gateway): secondary-profile adapters no longer send their credentials to the default profile's host/identity
Under gateway.multiplex_profiles, os.environ holds the DEFAULT profile's
values and a secondary profile's .env exists only in its secret scope.
Matrix, Home Assistant, Teams, DingTalk, SMS and Telegram already read
their *secret* scope-aware, but read the paired endpoint/identity raw from
os.environ — so a secondary's token/secret/password was paired with the
default profile's host or app identity:

- Matrix: MATRIX_HOMESERVER / MATRIX_USER_ID / MATRIX_DEVICE_ID in
  __init__ and MATRIX_HOMESERVER in _standalone_send -> the secondary's
  access token was sent to the default's homeserver (and its E2EE device
  id reused).
- Home Assistant: HASS_URL in __init__ and _standalone_send -> the
  secondary's HASS_TOKEN was posted to the default's HA instance.
- Teams: TEAMS_CLIENT_ID / TEAMS_TENANT_ID (env-first, even over extra),
  TEAMS_SERVICE_URL, TEAMS_HOME_CHANNEL(_NAME) in _credentials,
  _env_enablement, _standalone_send and __init__ -> Bot Framework token
  requested for the default's app with the secondary's secret; cron
  home channel was the default's conversation.
- DingTalk: DINGTALK_CLIENT_ID in _credentials, DINGTALK_WEBHOOK_URL in
  _standalone_send -> wrong app id; cron output posted to the default's
  robot webhook (URL carries its access_token).
- Telegram: TELEGRAM_WEBHOOK_URL in connect() -> the default's public
  webhook URL registered on the secondary bot, mixing update traffic.

Each read now goes through the same scoped reader its sibling secret
already uses (_get_scoped_secret / _startup_env_secret / get_secret with
the UnscopedSecretError fallback), so a scoped miss is empty rather than
the default profile's value; the unscoped default-profile path and
single-profile deployments keep reading os.environ, which there IS the
profile's own value. Teams _credentials now lets extra (per-profile
config.yaml) win over env, matching every other adapter and its own
__init__.

Live repro (/tmp/mux_audit/fix-endpoint-identity/repro.py, secondary
scope "bot2" with the default profile in os.environ): 19 of 24 reads
returned the default's endpoint/identity before; 0 after.

Salvage: SMS from-number fix cherry-picked from #100625 (tests trimmed to
two invariants). Teams identity (#100622) and DingTalk (#100615)
reimplemented on the minimal shape — the Teams PR added a _profile_scoped()
gate to _env_enablement and the DingTalk PR only covered the YAML bridge,
not DINGTALK_CLIENT_ID / DINGTALK_WEBHOOK_URL. Matrix, Home Assistant and
Telegram parts have no prior PR.

Co-authored-by: nftpoetrist <264138787+nftpoetrist@users.noreply.github.com>
2026-09-10 18:13:53 -07:00
nftpoetrist
8099745a4e fix(sms): scope TWILIO_PHONE_NUMBER per profile
__init__ and _standalone_send read TWILIO_PHONE_NUMBER via raw os.getenv,
right next to the already-scope-aware TWILIO_ACCOUNT_SID/TWILIO_AUTH_TOKEN
(_get_scoped_secret, established by merged PR #76664). Under
gateway.multiplex_profiles, a secondary profile with its own Twilio account
would silently send replies from the default profile's bridged
TWILIO_PHONE_NUMBER instead -- Twilio rejects a from-number not owned by
that profile's account (auth failure), or the message gets attributed to
the wrong sender.

__init__'s path is currently unreachable via normal secondary-profile
adapter construction ("sms" is in gateway/config.py's
PORT_BINDING_PLATFORM_VALUES, so _start_one_profile_adapters refuses to
construct a port-binding platform for a secondary profile) -- fixed anyway
for consistency with its sibling reads and to avoid a latent bug if that
guard's scope ever changes. _standalone_send's path IS reachable: it's the
generic cron/home-channel out-of-process delivery hook, which runs inside
_profile_runtime_scope.

Swaps both reads to the adapter's own existing _get_scoped_secret() helper,
matching its sibling account_sid/auth_token reads exactly -- no new
mechanism needed.

Adds a TestMultiplexProfileScope test class to tests/gateway/test_sms.py
mirroring the established Buzz/Discord/Telegram/WhatsApp/LINE/DingTalk/
Teams coverage for this bug class.

(cherry picked from commit 2aeb3148367ecf87598d4458221a5fbf30c0af01)
2026-09-10 18:13:53 -07:00
Teknium
2952dc62bc fix(feishu): drive-comment turns and WS-thread callbacks run under their own profile, not the launch profile
Under `gateway.multiplex_profiles`, a Feishu adapter is built and connected inside
`_profile_runtime_scope` (HERMES_HOME override + secret scope as contextvars), but two
hops started from an EMPTY context and so executed under the LAUNCH profile:

- `feishu_comment.handle_drive_comment_event` ran the whole comment AIAgent turn on a bare
  `loop.run_in_executor(None, ...)`: model/credential resolution raised
  `UnscopedSecretError` (silent empty reply), or — when the default profile held the same
  key — used the default profile's config/model/state for a secondary profile's doc.
- `FeishuAdapter._connect_websocket` ran the lark WS client on a bare executor thread. The
  SDK fires every event/card callback on that thread and they hop back to the adapter loop via
  `run_coroutine_threadsafe`, which copies the CALLER's context — so all pre-handler work
  (inbound media caching, `.update_response` marker, FEISHU_REACTIONS env, drive comments)
  ran unscoped. The pending-inbound drainer thread spawned from that callback had the same
  shape.

Carry the scope across each hop with `contextvars.copy_context().run`. For the SDK-owned
thread the snapshot is taken once at `_connect_websocket` (inside the profile scope; the
restart supervisor task inherits it too), so no per-callback re-scoping is needed.

Supersedes the `_submit_on_loop` re-scoping approach of #63962 (nateEc, earliest fix):
scoping the WS thread at its source covers every callback without rebuilding the secret
scope per call.

Co-authored-by: Nathan Shan <nathanielcrush51@gmail.com>
2026-09-10 18:12:02 -07:00
ethernet
6756d11b5f fix(pm): ship full chromium without headless shell
Full Chromium serves both headed and headless sessions. The separate
shell duplicates the browser payload and is not needed for either mode.

Remove the shell from PM and Docker. Select the managed Chromium
executable for agent-browser and the full Chromium channel for direct
Playwright callers. Route setup through PM and remove retired packages
from cached bundle stores without changing the user's tool store.

Update signing, architecture checks, launch probes and install guidance.
Leave llama packages and Docker archive cleanup unchanged.

Verification:
- Real agent-browser navigation, clicks, DOM reads and screenshots pass
  in headed and headless modes with the same Chromium executable.
- The direct Playwright doctor probe passes.
- Focused Python and desktop packaging tests pass, as do six Docker
  checks and both real-browser task-scroll tests.
- The built linux/amd64 image is 1.393 GB compressed, 223.6 MB smaller.
- The broader PM suite and two unrelated setup tests still fail.
  Those failures reproduce on unchanged HEAD.
- Five updated eval scripts parse; their full scenarios were not run.
2026-09-10 16:08:03 -04:00
Justin Bennington
d7b0a72c2a fix(providers): route Actual through chat completions (E-1047) 2026-09-10 14:56:55 -04:00
Erosika
d74f13e4a0 fix(honcho): include a2aSessions in identity_signature
sync_turn reads a2a_sessions from the config bound when the provider was built. A cached gateway provider kept the old value after honcho.json flipped it, because the signature that busts that cache did not carry the flag.
2026-09-10 10:45:57 -07:00
Erosika
431cd9084b fix(honcho): bound the joined author peer memory by session count
_joined_author_peers kept an entry for every honcho session the manager ever wrote to. It now holds at most _SESSION_CACHE_MAX_SIZE sessions and drops the oldest past that, so a forgotten session's authors rejoin on their next write. A failed join no longer leaves an empty entry behind.
2026-09-10 10:45:57 -07:00
Erosika
bcac7e9465 fix(honcho): read an author join's observation flags through one manager method
The join read the manager-wide user_observe_me and user_observe_others directly. It now asks _join_observation_flags(honcho_session_id), which returns the same values today. #103889 stores the effective flags per session and replaces the body of that method.
2026-09-10 10:45:57 -07:00
Erosika
2ac7fcddf8 fix(honcho): a bot author never lands on the session's human runtime peer
_generated_runtime_peer_id takes a reserved set, and the bot path passes the session's human peer ids: each runtime id and the peer _resolve_user_peer_id returns for the key. bot:coder with a runtime human coder and no runtimePeerPrefix now gets the digest suffix. _explicit_user_peer_ids keeps its meaning for prefixed runtime users.
2026-09-10 10:45:57 -07:00
Erosika
d96a6d9ab9 fix(honcho): include the workspace in identity_signature
identity_signature now carries cfg.workspace_id. The gateway folds these values into its agent cache key, and a workspace change in honcho.json reused a cached agent that was still bound to the old workspace.
2026-09-10 10:45:57 -07:00
Erosika
8704e9ca4c fix(honcho): put this agent's aiPeer in the a2a session key
_a2a_session_key now names the session <session>:a2a:<aiPeer>:<sender id>-<digest>. Two profiles that share a workspace and a session key wrote one sender's DMs into one Honcho session. The recipient peer comes from the same aiPeer derivation the session builder uses, moved into session_peers.assistant_peer_id_for so the two cannot drift.
2026-09-10 10:45:57 -07:00
Erosika
b4d7a33b51 fix(honcho): derive a bot author's peer from its full id with the runtime digest rule
_peer_id_for_runtime_id now looks up userPeerAliases by the full bot id and otherwise passes everything after bot: through _generated_runtime_peer_id. A digest suffix is added when sanitizing changed the id or the result equals peerName or an alias target, so bot:eri never resolves to the operator's peer and bot:a.b stays apart from bot:a-b. The docstring and README no longer claim a cloned profile's aiPeer defaults to the profile name.
2026-09-10 10:45:57 -07:00
Erosika
bdeb6f1b77 refactor(honcho): name the a2a session from core's a2a_key
The plugin spelled the `a2a:` prefix itself. `agent.turn_author.a2a_key` is the shared name
for a bot author's turns, so the session key now derives from it and every reader that files
bot turns apart agrees on the prefix. The resulting key is unchanged.
2026-09-10 10:45:57 -07:00
Erosika
170589dd10 fix(honcho): every bot author gets its own peer or its turn is skipped
A gateway platform marks a bot sender with its raw user id and a bot flag, never a `bot:` id.
`resolve_author_peer_id` treated that author as a human, so `pinUserPeer` collapsed a bot onto
`peerName` and an unresolved peer opened the a2a session under the human's peer. The resolver now
takes `is_bot` and gives every bot its own peer. `sync_turn` skips the turn when no peer resolves
or when the peer equals this agent's `aiPeer`.

The a2a session key carries an eight-character digest of the author id, so two ids that sanitize
alike stay in separate sessions. During a bot-authored turn `honcho_conclude` and `honcho_profile`
refuse writes and the built-in memory mirror is skipped, because conclusions and cards describe
the human. The README paragraph on bot DMs now matches the code.
2026-09-10 10:45:57 -07:00
Erosika
9f2a9384dc fix(honcho): pinUserPeer collapses the operator's accounts, not bot authors
with pinUserPeer on, resolve_author_peer_id returned None for every author,
so a bot dm's words were written under the human's pinned peer inside the
a2a session. the pin exists to unify one person's platform accounts. a bot
is not one of them.

bot: authors now resolve to their peer before the pin check, so a pinned
operator still gets bot speech attributed to the bot.
2026-09-10 10:45:57 -07:00
Erosika
df1513b728 docs(honcho): document a2aSessions and bot dm attribution 2026-09-10 10:45:57 -07:00
Erosika
f7d5ac3230 feat(honcho): write bot dms into their own a2a session
A DM relayed from another Hermes profile ran as a turn in the recipient's
Bot Chat session. sync_turn wrote the bot's words and the recipient's reply
into that session, and before per-author writes they landed under the
human's peer. The human's representation absorbed conversations the human
never had.

The turn context now marks such turns with scope a2a:<bot id>.
sync_turn routes a bot-authored turn into a separate Honcho session keyed
<session>:a2a:<sanitized bot id>, created with the sender bot as its user
peer, and never writes it into the human's session. The key is deterministic
so every turn from the same bot reaches the same session, and it stays
inside Honcho's 100 character session id limit. Recall still reads the
human's session only.

a2aSessions (host block, then root, default true) turns the routing on.
With it off, bot-authored turns are skipped. A bot turn that names no
author id is skipped as well, because nothing can key its session. Human
turns are unchanged.

get_or_create takes a user_peer_id override so the a2a session's roster is
the bot and the assistant, not the runtime human.
2026-09-10 10:45:57 -07:00
Erosika
88fc9402d8 feat(honcho): declare identity_signature and drop the gateway's honcho keys
The gateway agent cache read honcho.json itself through a honcho-named
block in gateway/run.py and gateway/run_agent_cache.py. Every other memory
provider had no way to bust the cache when its identity mapping changed.

HonchoMemoryProvider.identity_signature() now returns the same values under
provider-neutral keys: user_identity, agent_identity, pin_user_identity,
runtime_identity_prefix, user_identity_aliases, session_prefixing. The
gateway files them under memory.<key> through the MemoryProvider hook. The
hook reads config only, memoizes on the file's mtime and size, and returns
an empty dict when the file cannot be read.

The honcho-specific extractor, its memo and its key tuple are gone from the
gateway. The pinPeerName cache-busting test now asserts on
memory.pin_user_identity.
2026-09-10 10:45:57 -07:00
Erosika
46d625b097 feat(honcho): map bot authors onto their profile peer
The bot-mode dispatcher names another profile as bot:<profile>. The
resolver treated that like a human runtime id, so a configured
runtimePeerPrefix produced peers like telegram_bot:coder and a profile
that already owns an AI peer in the same workspace got a second one.

A bot:<profile> author now resolves in this order: a userPeerAliases entry
for the full bot id, else the sanitized profile name. A cloned profile's
aiPeer defaults to the profile name, so a same-workspace sender lands on
its existing AI peer. Prefixes never apply to bot ids. pinUserPeer still
collapses bot authors onto the pinned peer, the same as every other author.
2026-09-10 10:45:57 -07:00
Erosika
20b117ad3b fix(honcho): read the turn author from sync_turn and treat the alt id as the session peer
sync_turn only knew the author through the on_turn_start stash. The memory
manager now passes turn_author and scope with the turn, and a caller
that skips on_turn_start left the stash empty or stale.

sync_turn takes both keywords and reads the author from turn_author first.
The stash stays as the fallback for callers that never pass it.

resolve_author_peer_id compared the author against the primary runtime id
only. A transport that names the participant by the alt id (Telegram
username instead of UID) got a second peer for the same person. Either
runtime id now counts as the session's own participant.
2026-09-10 10:45:57 -07:00
Erosika
6aff2fc65b feat(honcho): write each turn under its author's peer
The manager resolved one user peer in `get_or_create` and froze it onto the
session, then `_flush_session` chose between it and the assistant peer by
role. Every user turn in a shared session landed on that one peer, so the
first person to message the agent collected everyone else's facts — and a
Honcho conclusion, once derived, is not self-correcting.

`resolve_author_peer_id` maps the turn's author onto its own peer using the
alias-then-prefix order `_resolve_user_peer_id` already applies, so an
aliased account reaches the same peer whichever turn it wrote. `sync_turn`
resolves it before starting the write thread, so a following turn cannot
retag a queued write. `_flush_session` then writes each user message under
that peer.

A shared session's roster is open — people and other agents arrive after
the session exists — so `_author_peer_for_session` joins a peer when it
first writes instead of enumerating participants at init. Joins are
remembered per session, and a failed join still writes under the right
peer, losing only the observe config.

Three cases return None and keep the session's own peer: no author named,
the author IS the session's peer, and `pinPeerName` set — that flag is an
explicit request to unify identities, so it still collapses authors in a
shared chat.

An unnamed author stays unattributed rather than defaulting to the owner.
That preserves today's behavior for the transports that send no author, so
those turns still reach the session peer; #83500 owner-gated the memory-file
migration for the same reason.

Display names never become peer IDs — they are attacker-influenceable on
any platform where participants set their own name.
2026-09-10 10:45:57 -07:00
Teknium
aeecb110f8 fix(deepseek): deepseek-flash is the canonical Flash id; retired names fold onto it
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.

Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.

Builds on YipTszkwan's #107126 (earliest fix in the cluster).
2026-09-10 02:44:26 -07:00
YipTszkwan
8435a3ae00 fix(deepseek): recognise the version-less deepseek-flash model id
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:

* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
  omitted `extra_body.thinking`. The server then defaults to thinking-on, so
  the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
  V-series regex), so the id a user picked never reached the wire and the
  config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
  instead of the real 1M window, capping the model at an eighth of its
  context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
  detector at its 180s default instead of the 600s reasoning-model floor.

Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".

Adds the id to all four gates plus regression coverage for each site.
2026-09-10 02:44:26 -07:00
ethernet
f2a19b0e06 refactor: remove unused extraction copies
Use the existing session-export, transcription and DingTalk owners.
Remove unused setup/watchdog helpers and the no-op package migration
hook. No package overrides it; user-state migrations keep their own
existing owners. Keep version-change installation coverage and the
scheduled external plugin compatibility blocks.

Verified with the real adapter, transcription, session-snapshot and
PM core suites. No live messaging service or user-state operation.
2026-09-09 18:21:39 -04:00
ethernet
e4cc7f09d9 merge: integrate upstream catalog with PM publication
Keep upstream's reviewed catalog as the only plugin name index.
Catalog pins and custom update sources share staged PM validation.
Publish code and dependencies with recovery after process death.
Reject a concurrent enablement change before publishing disabled code.

Use the manifest loader's supported version in the installer. Keep
probe cooldowns for timeouts, not TLS failures that a CA change fixes.
Preserve the backup, uninstall, browser and memory-provider repairs.

Verified with the canonical runner on native Windows ARM64, real Git
repositories, local TLS endpoints and UV dependency generations.
Desktop catalog tests and both TypeScript checks pass. The full suite
and native release builds were not run. No remote push.
2026-09-09 16:49:27 -04:00
Teknium
b7bef04861 fix(kanban): an explicit scratch workspace never inherits the board's project
Move the "explicit scratch means no project" decision into the one resolver
every surface funnels through, `kanban_db.create_task`: board-project
inheritance now runs only when the caller left `workspace_kind` open
(`None`), and `workspace_kind` defaults to scratch after that check. The
tool handler keeps the #106347 fix for `project=""` (no `or` collapse) and
`board=` scoping but drops its handler-local sentinel logic, since the
resolver now owns the rule; the `self_task` project inheritance for
dispatcher-owned workers is unchanged.

Sibling surfaces had the same bug through the same line and are fixed by
the same change:
- CLI `hermes kanban create --workspace scratch` on a project-scoped board
  produced a project worktree; `--workspace` no longer defaults in
  argparse so the resolver can tell "omitted" from "scratch".
- Dashboard `POST /tasks` with `workspace_kind: "scratch"` did the same;
  `CreateTaskBody.workspace_kind` defaults to `None` for the same reason.
- `kanban_swarm.create_swarm` threads `None` through for consistency.

Tests: one resolver invariant in test_kanban_board_project.py (explicit
scratch stays scratch, omitted still inherits) and the salvaged tool test
folded into a single parametrized matrix over scoped/unscoped target boards.
2026-09-09 12:18:58 -07:00
Tranquil-Flow
77c994bab2 fix(line): keep progress heartbeats and stale mappings out of the postback cache (#106446)
(cherry picked from commit 4b0dc3c58cfe9771745384f20d75ca48d9b25408)
2026-09-09 11:47:55 -07:00
teknium1
3dc3809d07 refactor(fal): render the managed billing message once in fal_common
Image and video callers were each formatting the same four-field dict into
the same sentence. Return the rendered tail from _managed_fal_billing_error
so the wording lives in one place; output is byte-identical.
2026-09-09 11:46:36 -07:00
Matt Earls
0c6b94e499 fix(image-gen): preserve managed FAL billing errors
Avoid retrying idempotent managed FAL submissions because the retry can mask the initial billing failure. Surface structured Nous billing diagnostics consistently for image and video paths, with hermetic regression coverage.

(cherry picked from commit 289ce039e9a522dc8016ae4a512214c05d0a8bc0)
2026-09-09 11:46:36 -07:00
teknium1
322905e91b feat(memory): make the mem0 sync char cap configurable via mem0.json
A flat 450-char cap fits 512-token embedders (bge-small-zh-v1.5,
all-minilm) but stores only ~5% of the window on 8192-token models
(text-embedding-3-small, jina-embeddings-v3, bge-m3), degrading memory
quality for users those models served fine before truncation existed.

Read `sync_max_chars` from mem0.json once in initialize() (450 default)
and pass it to _truncate_for_sync(). Config over auto-detection: the
Ollama /api/show probe + known-model table proposed in #37427 adds a
network call and a curated list for a number the operator already knows
from their embedder choice; the setup wizard's mem0.json is the plugin's
behavioral-settings surface (no new HERMES_* env var). Documented in the
plugin README and the memory-providers docs page.

Dynamic-cap requirement and measurements (450 OK / 600 -> HTTP 500 on
bge-small-zh-v1.5:f16) by @szicely in #106235.

Refs #37421 #106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
2026-09-09 10:55:02 -07:00
liuhao1024
040742d06b fix(memory): truncate oversized mem0 sync messages at the source
sync_turn() sent the whole turn to backend.add() untruncated. OSS
embedding models with small context windows (Ollama bge-small-zh-v1.5:
512 tokens) reject the request with HTTP 500, and hosted APIs answer
INPUT_TOKEN_LIMIT_EXCEEDED — in both cases _try() only logs, silently
dropping the turn's memory extraction after long conversations.

Cap each synced message at its last sentence boundary within 450 chars
before ingestion: short turns pass through unchanged, long turns keep a
coherent statement for fact extraction. This replaces the previous
retry-on-error approach, which could not match Ollama's HTTP 500 shape.

Salvaged from #37427 with the test suite trimmed to the one invariant
(oversized turn still reaches a small-context backend, short text
untouched, no breaker failure).

Refs #37421 #106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
2026-09-09 10:55:02 -07:00
Teknium
1d651b3bb2 feat(matrix): render LaTeX in the standalone (cron) sender too
_standalone_send builds formatted_body through its own markdown call and
never went through _markdown_to_html, so cron-delivered equations still
arrived as raw dollars. Tokenize/expand around that conversion as well.
2026-09-09 10:43:58 -07:00
Teknium
e22e04692a test(matrix): trim LaTeX salvage to two invariant tests, tidy _markdown_to_html
Keep the two tests that fail without the fix: inline + display math reach
formatted_body as data-mx-maths markup with the TeX HTML-escaped while the
plain body keeps the raw TeX; unpaired dollars and text colliding with the
sentinel format pass through unchanged (no IndexError). The 17 unit tests
from #106452 are dropped per the salvage bar (≤ 2 invariant tests).

Also drop the dead `_tex_store = []` pre-assignment and the `_fb`/`html`
temporaries in `_markdown_to_html` — pure tidy, no behaviour change.
2026-09-09 10:43:58 -07:00
romanovzky
199f4d5b96 feat(matrix): render LaTeX math via Element data-mx-maths markup
Element (feature_latex_maths) typesets <div|span data-mx-maths="TEX">
elements at display time, but the outbound HTML sanitizer allowlists tags
and attributes, so data-mx-maths markup sent by the gateway never reaches
Element intact - messages containing $...$ render as raw dollars.

Convert $...$ (inline) and $$...$$ (display) to opaque sentinel tokens
before Markdown conversion and expand them to data-mx-maths markup after
sanitization. Tokens are printable text with no HTML/Markdown meaning, so
neither the converter nor the sanitizer touches the TeX. Unpaired dollars
(prices, literals) are untouched, and adversarial text colliding with the
sentinel format passes through verbatim (index-checked expansion).

(cherry picked from commit eb73aafe0e08502c54e63dbd03e99796c3b7fee3)
2026-09-09 10:43:58 -07:00
Konstantin Khlopkov
1a981d8e83 fix(telegram): bots_require_mention gates bot quote-replies behind an explicit @mention (closes #106430)
(cherry picked from commit a67fa8c4b21e31a44019d850aedc5be9015a546d)
2026-09-09 10:37:37 -07:00
teknium1
91adf584a4 fix(gateway): every send_multiple_images override returns the aggregate SendResult
#106167 widened the base contract to SendResult but left six native-batch
overrides (Discord, Email, Matrix, Mattermost, Slack, Telegram) returning
None, which (a) meant a media-only reply on those platforms still reported
FAILURE because _record_delivery(None) records nothing, and (b) produced new
`ty` invalid-method-override diagnostics against the widened base (#106192).

Make the contract honest instead of annotating it Optional: each override
now rolls its batches (and any per-image fallback) into one SendResult, so
the turn-outcome accounting works on every platform, not just Signal and
the base loop. The "legacy overrides return None" comment in
_send_image_batch goes away with the legacy.

ty on the 8 touched files: origin/main 241 diagnostics / 20 override,
this branch 241 / 20 — byte-identical diagnostic set; the intermediate
`-> SendResult` head without this commit had 246 / 25.

Refs #106192
2026-09-09 09:45:54 -07:00
Teknium
e74c4a00ca Merge pull request #69446 from NousResearch/feat/plugin-catalog
feat: plugin catalog — curated SHA-pinned plugin index (CLI, admission CI, docs, dashboard)
2026-09-09 09:22:21 -07:00
Gianpietro Dal Zio
653418d842 fix(dashboard): coalesce expensive reads before worker admission 2026-09-09 21:16:40 +05:30
Teknium
41b4555ed9 feat(plugins): catalog is the sole discovery system — re-port onto main's layout
- hermes_cli/plugins_cmd_catalog.py: new sibling owning resolution, the
  .hermes-catalog.json provenance sidecar, search/info/validate, re-pin on
  update, and the dashboard/TUI payload builders. plugins_cmd.py only
  gains the hooks (cmd_install catalog branch, cmd_update / dashboard
  update re-pin, dashboard_install_plugin catalog_name + kill list,
  dispatch entries); the community index (plugin_index.py) is gone.
- hermes_cli/plugin_catalog.py: catalog_dir parameter replaces the
  test-only HERMES_PLUGIN_CATALOG_DIR env var; live refresh reads ONE
  published document (/docs/api/plugin-catalog.json, 6h cache, in-tree
  fallback) instead of the unauthenticated GitHub contents API (60 req/h,
  1 request per entry); in-tree and live removals are unioned so a stale
  cache can never un-block.
- Catalog route lives in web_routers/dashboard_ui.py (the facade is off
  limits); _plugin_runtime_status shared from web_server_dashboard.py;
  hub rows carry removed_reason. TUI plugins.manage gains catalog_name
  install, catalog row fields and an update action.
- plugin_validate: the probe context honours ctx.get_config defaults
  (real plugins do int(ctx.get_config("timeout", 180)) in register()).
- Installed-state merge matches through the sidecar's catalog_name
  first — catalog names rarely equal manifest names.
- extract-plugins.py emits plugin-catalog.json; deploy-site triggers on
  plugin-catalog/** so entry merges republish it.
2026-09-09 04:38:01 -07:00
ericmaddox
bee840bc8c fix(providers,agent): handle strict-string tool message validation and 422 on opencode-go (fixes #104731)
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
2026-09-09 03:52:47 -07:00
Yuan Chenglu (袁成路)
6a0f519d1e fix(opencode-go): set supports_vision_tool_messages=False for Xiaomi MiMo backend
## Problem

When using the opencode-go provider with Xiaomi MiMo models (e.g.
mimo-v2.5, mimo-v2.5-pro), the Hermes agent intermittently fails with:

    Error code: 400 - {'error': {'code': '400',
      'message': 'Error from provider (Xiaomi): Param Incorrect',
      'param': 'text is not set', 'type': ''}}

This occurs specifically when tool results contain multipart content
with image_url parts (e.g. browser screenshots). The opencode-go relay
forwards these as-is to the Xiaomi MiMo backend, which rejects list-type
tool message content while still accepting multimodal user messages.

## Root Cause

The OpenCodeGoProfile inherits supports_vision_tool_messages=True from
ProviderProfile (the default). When this flag is True, the agent sends
tool results with image parts directly to the model. However, Xiaomi
MiMo's API rejects this format:

> "Set to False for providers that accept multimodal user messages but
> reject list-type tool content (e.g. Xiaomi MiMo, which returns 400
> 'text is not set')."
>   — providers/base.py, line 73

The direct 'xiaomi' provider profile already correctly sets this to
False (plugins/model-providers/xiaomi/__init__.py, line 13), but the
opencode-go relay profile was missing this safeguard.

The relevant code path is in run_agent.py:_tool_result_content_for_active_model()
(line 4543), which checks _provider_supports_vision_tool_messages() when
deciding whether to embed images in tool-result messages.

## Fix

Add supports_vision_tool_messages=False to the OpenCodeGoProfile
instantiation in plugins/model-providers/opencode-zen/__init__.py.

This single-line change prevents tool-result images from being sent
as multipart content to the MiMo backend, while preserving the model's
image recognition capability through user messages and vision tool
invocations (both of which use different code paths unaffected by this
flag).

## Testing

Verified with the mimo-v2.5 model via opencode-go provider:

1. Browser tool + screenshot recognition
   → Navigated to https://www.baidu.com, took screenshot, identified
     top 3 trending topics from the image
   → Result: PASSED, recognized all topics correctly

2. Direct image as user message
   → Sent a screenshot PNG directly via --image flag, asked model to
     describe the content
   → Result: PASSED, model correctly read text from the image

3. Provider profile verification
   → Confirmed get_provider_profile('opencode-go').supports_vision_tool_messages
     returns False at runtime
   → Result: PASSED

4. No regression on non-MiMo models
   → opencode-zen provider retains supports_vision_tool_messages=True
     (unaffected)

---

fix(opencode-go): 为 Xiaomi MiMo 后端设置 supports_vision_tool_messages=False

## 问题描述

使用 opencode-go provider 搭配 Xiaomi MiMo 模型(如 mimo-v2.5、
mimo-v2.5-pro)时,Hermes agent 间歇性地抛出以下错误:

    Error code: 400 - {'error': {'code': '400',
      'message': 'Error from provider (Xiaomi): Param Incorrect',
      'param': 'text is not set', 'type': ''}}

该错误发生在工具返回结果包含 image_url 类型的 multipart 内容的场景下
(如浏览器截图)。opencode-go 中继层将这些内容原样转发给 Xiaomi MiMo
后端,而 MiMo 接受多模态用户消息,但拒绝 list-type tool message 内容。

## 根因分析

OpenCodeGoProfile 继承了 ProviderProfile 的默认值
supports_vision_tool_messages=True。当此标志为 True 时,agent 会将含
图片的工具结果直接发送给模型。但 Xiaomi MiMo API 拒绝此格式:

> providers/base.py 第 73 行注释明确指出:
> "Set to False for providers that accept multimodal user messages but
> reject list-type tool content (e.g. Xiaomi MiMo, which returns 400
> 'text is not set')."

直接的 'xiaomi' provider profile 已正确设置了该值为 False
(plugins/model-providers/xiaomi/__init__.py 第 13 行),但
opencode-go 中继 profile 遗漏了这一安全设置。

相关代码路径:run_agent.py 的 _tool_result_content_for_active_model()
方法(第 4543 行),该方法通过检查
_provider_supports_vision_tool_messages() 来决定是否在 tool-result
消息中嵌入图片。

## 修复方案

在 plugins/model-providers/opencode-zen/__init__.py 的
OpenCodeGoProfile 实例化中添加 supports_vision_tool_messages=False。

这一行改动阻止了 tool-result 图片以 multipart 格式发送给 MiMo 后端,
同时通过用户消息和 vision tool 调用的路径(使用不同代码路径,不受
此标志影响)保留了模型的图像识别能力。

## 测试验证

使用 mimo-v2.5 模型通过 opencode-go provider 验证:

1. 浏览器截图 + 图像识别
   → 导航至 https://www.baidu.com,截取首页截图,从图片中识别出
     热搜榜前三条
   → 结果:通过,正确识别所有热搜话题

2. 用户消息直接传图
   → 通过 --image 参数直接发送截图 PNG,要求模型描述图片内容
   → 结果:通过,模型正确读取图片中的文字

3. Provider profile 运行时验证
   → 确认 get_provider_profile('opencode-go')
     .supports_vision_tool_messages 在运行时返回 False
   → 结果:通过

4. 非 MiMo 模型无回归
   → opencode-zen provider 保持 supports_vision_tool_messages=True
     不受影响

## 修改文件

  plugins/model-providers/opencode-zen/__init__.py (+5 lines)

Signed-off-by: Yuan Chenglu (袁成路) <ycl_pj@163.com>
2026-09-09 03:52:47 -07:00
ethernet
8b7eae99ef fix(pm): own interpreter selection and dependency recovery
Pin uv and uvx to the PM interpreter instead of ambient Python discovery.
A matching dependency stamp cannot prove that installed files still exist.
Repair now rebuilds the recorded workspace and lock in a fresh generation,
checks startup imports, and publishes the selection only after success.

Run startup recovery before dependency activation. Keep manual PM repair
reachable when the selected environment is damaged. Preserve plugin
selection, retry ownership, and the previous generation on failure.
Remove the separate pip, ensurepip, per-extra, and install-time quarantine
ladders. Keep orphan launcher restoration.

Verification: 717 targeted tests passed on native Windows ARM64, with
56 skipped. Ruff, diff checks, and the source-scoped compat check passed.
A disposable real Hermes install recovered deleted YAML and dotenv files,
then printed CLI help with exit 0. Its lock and stamp stayed unchanged.
The full suite and a release build were not run for this change.
2026-09-08 23:39:55 -04:00