A DM relayed from another Hermes profile ran as a turn in the recipient's
Bot Chat session. sync_turn wrote the bot's words and the recipient's reply
into that session, and before per-author writes they landed under the
human's peer. The human's representation absorbed conversations the human
never had.
The turn context now marks such turns with scope a2a:<bot id>.
sync_turn routes a bot-authored turn into a separate Honcho session keyed
<session>:a2a:<sanitized bot id>, created with the sender bot as its user
peer, and never writes it into the human's session. The key is deterministic
so every turn from the same bot reaches the same session, and it stays
inside Honcho's 100 character session id limit. Recall still reads the
human's session only.
a2aSessions (host block, then root, default true) turns the routing on.
With it off, bot-authored turns are skipped. A bot turn that names no
author id is skipped as well, because nothing can key its session. Human
turns are unchanged.
get_or_create takes a user_peer_id override so the a2a session's roster is
the bot and the assistant, not the runtime human.
DeepSeek retired deepseek-v4-flash on 2026-09-10 (V4.1-Flash release); the API's
model name is now `deepseek-flash` and /v1/models lists only it. Hermes still
folded every non-V-series name onto deepseek-v4-flash, so `/model deepseek-flash`
on the DeepSeek provider was rewritten, then the validator "auto-corrected" it
back against the live listing: "Auto-corrected deepseek-v4-flash -> deepseek-flash"
on every switch.
Retired aliases (deepseek-chat / -reasoner and other fuzzy names) now fold onto
deepseek-flash; the curated catalog, profile fallback list, aux default, goal-judge
hint and pricing snapshot (2026-09-10 off-peak USD) follow the docs. Dated
deepseek-v4-* ids still pass through untouched.
Builds on YipTszkwan's #107126 (earliest fix in the cluster).
DeepSeek's 2026-09 Flash refresh introduced a version-less canonical id:
GET /v1/models now returns `deepseek-flash` (alongside `deepseek-v4-pro`), the
API accepts it directly, and the older `deepseek-v4-flash` is server-side
aliased onto it. Every DeepSeek model-id gate in Hermes keys off the
`deepseek-v<N>` prefix, so the new id silently missed all four:
* DeepSeekProfile.build_api_kwargs_extras classified it as non-thinking and
omitted `extra_body.thinking`. The server then defaults to thinking-on, so
the user's thinking toggle and `reasoning_effort` were quietly ignored.
* `_normalize_for_deepseek` folded it onto `deepseek-v4-flash` (it misses the
V-series regex), so the id a user picked never reached the wire and the
config stored a different model than the picker advertised.
* `DEFAULT_CONTEXT_LENGTHS` fell through to the 128K `deepseek` catch-all
instead of the real 1M window, capping the model at an eighth of its
context before compaction kicked in.
* `_REASONING_STALE_TIMEOUT_FLOORS` had no entry, leaving the stale-stream
detector at its 180s default instead of the 600s reasoning-model floor.
Verified live against api.deepseek.com: `deepseek-flash` answers 200 with
`model: deepseek-flash`, accepts image input (the refresh folds vision into
the Flash model), and the in-between id `deepseek-v4.1-flash` is rejected
with "The supported API model names are deepseek-flash, deepseek-v4-pro".
Adds the id to all four gates plus regression coverage for each site.
Keep the two tests that fail on main without the fix:
- test_fal_common: a keyed managed submit makes exactly one POST (plus the
negative arm: an unkeyed submit still goes through the SDK retry ladder)
- test_image_generation: the 409 BILLING_ERROR body surfaces
`unsupported_pricing_meter` instead of the generic "not yet enabled" text
Dropped from #106484: the duplicate video-plugin billing test (same helper,
same assertion), the `_fal_client = fake` / `import_fal_client` stub churn
and the `tools.lazy_deps` stub — fal-client is installed in CI (`--extra fal`)
so those fixtures were not needed; the `_load_fal_client` no-op fixture on
TestManagedGatewayErrorTranslation for the same reason.
Also drop the redundant `retry_request is None` re-check in
`_ManagedFalSyncClient.submit` — `__init__` already raises when the helper
is missing.
Avoid retrying idempotent managed FAL submissions because the retry can mask the initial billing failure. Surface structured Nous billing diagnostics consistently for image and video paths, with hermetic regression coverage.
(cherry picked from commit 289ce039e9a522dc8016ae4a512214c05d0a8bc0)
A flat 450-char cap fits 512-token embedders (bge-small-zh-v1.5,
all-minilm) but stores only ~5% of the window on 8192-token models
(text-embedding-3-small, jina-embeddings-v3, bge-m3), degrading memory
quality for users those models served fine before truncation existed.
Read `sync_max_chars` from mem0.json once in initialize() (450 default)
and pass it to _truncate_for_sync(). Config over auto-detection: the
Ollama /api/show probe + known-model table proposed in #37427 adds a
network call and a curated list for a number the operator already knows
from their embedder choice; the setup wizard's mem0.json is the plugin's
behavioral-settings surface (no new HERMES_* env var). Documented in the
plugin README and the memory-providers docs page.
Dynamic-cap requirement and measurements (450 OK / 600 -> HTTP 500 on
bge-small-zh-v1.5:f16) by @szicely in #106235.
Refs #37421#106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
sync_turn() sent the whole turn to backend.add() untruncated. OSS
embedding models with small context windows (Ollama bge-small-zh-v1.5:
512 tokens) reject the request with HTTP 500, and hosted APIs answer
INPUT_TOKEN_LIMIT_EXCEEDED — in both cases _try() only logs, silently
dropping the turn's memory extraction after long conversations.
Cap each synced message at its last sentence boundary within 450 chars
before ingestion: short turns pass through unchanged, long turns keep a
coherent statement for fact extraction. This replaces the previous
retry-on-error approach, which could not match Ollama's HTTP 500 shape.
Salvaged from #37427 with the test suite trimmed to the one invariant
(oversized turn still reaches a small-context backend, short text
untouched, no breaker failure).
Refs #37421#106235
Co-authored-by: szicely <140148567+szicely@users.noreply.github.com>
Drop the bare "tool.content" pattern (any 400 mentioning tool.content in a
non-list context would be sent through the image-strip path) and the profile
flag snapshot test; the behaviour tests (classifier verdict + proactive
downgrade) already pin the contract.
- Declare `supports_vision_tool_messages=False` and `supports_vision=True` on `opencode_go` provider profile in `plugins/model-providers/opencode-zen/__init__.py`
- Route HTTP 422 errors through `_IMAGE_TOOL_RULES` and add `tool.content.str`, `tool.content`, and `input should be a valid string` patterns to `_MULTIMODAL_TOOL_CONTENT_PATTERNS` in `agent/error_classifier.py`
- Add unit tests for OpenCode Go proactive tool result downgrade, HTTP 422 Console Go classification, and profile capability contract in `tests/run_agent/test_multimodal_tool_content_recovery.py` and `tests/plugins/model_providers/test_opencode_go_profile.py`
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
Move the general-vs-memory hook ownership logic out of the memory collector into
PluginLedgerMixin (_drop_fallback_hooks / _register_fallback_hook) so the collector
and the loader each call one manager method instead of reaching into manager privates.
Hoist hashlib to module scope. Trim the new suite to the three invariant cases
(run-once across load orders, distinct sources not suppressed, re-exported register).
The plugin loader execs sibling modules before the package __init__, so a
module-level `from . import _read_mem0_json` in _setup.py failed against the
empty parent shell and the whole module silently dropped out — every later
`hermes memory setup mem0` died with "cannot import name 'post_setup'".
The test loads mem0 through the real discovery path from a cold sys.modules
and asserts the cached _setup module exposes post_setup (red on base).
Also maps the contributor email for #103078 credit.
Campaign tracker: https://github.com/NousResearch/hermes-agent/issues/104154
Adds plugins/web/perplexity — a keyed-only WebSearchProvider over httpx:
- search: POST https://api.perplexity.ai/search (documented Search API),
search_context_size=low so `snippet` stays description-sized;
results[].snippet -> description, max_results capped at the API's 20.
- extract: POST /sdk/content/snippets — the query-relevant page-excerpt
route behind `pplx content snippets` (the CLI's `content fetch` is
deprecated upstream). web_extract has no query, so the URLs' path words
serve as the relevance query; per-URL `error` entries survive a 200.
- Wired into the same touchpoints as the other keyed vendors: legacy
backend set + credential ladder + availability probe (web_tools),
registry preference walk, OPTIONAL_ENV_VARS, `hermes config`/status/
dump key lists, nous_subscription direct-credential detection, setup
summary, test conftests, docs.
Not a keyless-ring member (Perplexity has no anonymous tier). Related
closed PRs #9192 / #23981 / #45225 predate the plugin ABC.
dingtalk: drop DINGTALK_TYPE_MAPPING/EXT_MAP re-exports. google_chat: card_spec_to_cards_v2 test -> .cards.
matrix: drop module-level MAX_MESSAGE_LENGTH alias (no importers). teams: drop TeamsSummaryWriter re-export
(teams_pipeline/runtime + tests -> summary_writer). wecom: drop WeComStreamExpiredError/STREAM_EXPIRED_ERRCODE/
MAX_INTERMEDIATE_FRAMES re-exports (tests -> .streaming). parallel: drop _get_parallel_client/_get_async_parallel_client
aliases (tests -> _get_sync_client). email: drop stale 'alias' comment (_esecret_int is the only name).
photon: re-remove credential_summary() (shim-only, cb9b7c36f3); its no-leak test now drives print_credential_summary.
ethernet8023: the no-leak test was deleted alongside the function it pinned; the
display-only credential status is live in print_credential_summary. credential_summary
restored byte-identical to BASE so the leak contract is CI-pinned again.
PhotonAdapter.__init__, check_requirements, validate_config,
_env_enablement, _markdown_enabled, _reactions_enabled, and
_standalone_send in adapter.py, plus load_project_credentials and
load_dashboard_project_id in auth.py, all read PHOTON_PROJECT_ID/
PHOTON_NODE_BIN/PHOTON_SIDECAR_PORT/PHOTON_SIDECAR_AUTOSTART/
PHOTON_PROBE_*/PHOTON_REQUIRE_MENTION/PHOTON_MENTION_PATTERNS/
PHOTON_REACTIONS/PHOTON_MARKDOWN/PHOTON_HOME_CHANNEL(_NAME)/
PHOTON_DASHBOARD_PROJECT_ID via raw os.getenv -- only
PHOTON_PROJECT_SECRET and PHOTON_SIDECAR_TOKEN were already scoped via
_get_scoped_secret.
Notably __init__'s project_id read was a stronger variant of the bug
(like the IRC fix in this series, item 11): the original
`os.getenv("PHOTON_PROJECT_ID") or extra.get("project_id") or stored_id`
ordering let a raw env read override even an explicitly configured
config.yaml extra -- a secondary profile that set its own project_id via
extra would still silently authenticate against the default profile's
Spectrum project, because the default profile's project id is always
bridged to os.environ under multiplex and env was checked first.
_reactions_enabled() and the require_mention/mention_patterns reads in
__init__ are exercised on every live inbound message / tapback, not just
at construction, so a secondary profile's reaction/mention-gating
behavior would be driven by the default profile's settings for the
adapter's entire runtime lifetime.
Switch every raw PHOTON_* read (except the two already scoped) to
_get_scoped_secret(), matching the module's existing helper (already
defined identically in both adapter.py and auth.py). Left
_dashboard_host()/_spectrum_host() and the interactive device-login flow
functions in auth.py untouched -- these are CLI-only management-plane
calls (`hermes photon login`/`setup`), not part of the gateway's
per-profile adapter construction/connection lifecycle, so they are not
reachable under a multiplexed secondary profile's scope; noted as a
"Scope note" in the PR body rather than silently expanding scope to
unreachable call sites.
Adds a new tests/plugins/platforms/photon/test_multiplex_profile_scope.py
(9 tests, two classes covering auth.py and adapter.py separately)
mirroring the fixture/assertion style established in
tests/gateway/test_line_plugin.py's TestMultiplexProfileScope, reusing
test_auth.py's tmp_hermes_home isolation pattern so tests don't depend on
the real ~/.hermes/auth.json fallback. Mutation-verified: stashed the
production fix and confirmed 7 of 9 new tests fail against pre-fix code
(the other 2 are non-differentiating regression guards -- unscoped-
default-profile-precedence, one per class -- which correctly pass either
way). Restored the fix; all 140 tests in tests/plugins/platforms/photon/,
the 10 photon-related parametrized tests in
test_adapter_startup_secret_scope.py, and the broader
test_multiplex_adapter_registry.py / test_adapter_connect_classification.py
suites (45 tests) pass.
Under multiplex_profiles the Hindsight provider's writer, daemon-start and
prefetch threads were spawned as bare threading.Thread, so they started with
an empty contextvars Context: no profile secret scope and no HERMES_HOME
override. get_secret() fails closed there, so the local_embedded daemon never
booted and every retain raised UnscopedSecretError, even though the spawning
thread (initialize()/sync_turn() inside the gateway's copy_context'd turn) had
the scope all along.
Spawn each thread with contextvars.copy_context().run so the child inherits
the spawner's scope + home override. No environ fallback, no re-parsed .env.
The shared hindsight-loop thread needs no wrap: coroutines submitted via
run_coroutine_threadsafe already run in the submitter's context per call.
Fixes#92608Fixes#94933
Co-authored-by: KIAgent01 <297567825+KIAgent01@users.noreply.github.com>
Co-authored-by: Parker Fawcett <259203091+Parker-Fawcett@users.noreply.github.com>
ThreadingHTTPServer request threads do not inherit the gateway's profile
ContextVars, so security.authenticate()/is_trusted_peer() read the
process-global A2A_* env on every request — every secondary profile's
listener authenticated against the default profile's tokens. Capture an
immutable A2ASecurityContext at adapter construction (which runs inside
_profile_runtime_scope for secondary profiles) and have the request
handler consult it instead of re-reading env per request.
Salvage note: the original `get_secret() / except UnscopedSecretError:
os.getenv()` fallthrough in _startup_env was replaced by _profile_scoped()
gating (Buzz/Raft pattern) — inside a secondary profile's scope the scope
is authoritative and a miss never falls through to os.environ.
Salvaged from #80956.
A2AAdapter.__init__ / _default_agent_name / _load_served_agents read
A2A_PORT, A2A_AGENT_NAME and A2A_AGENT_DESCRIPTION from raw os.environ,
so a secondary multiplex profile borrowed the default profile's port and
Agent Card identity. Skip the env read when constructed inside a
secondary profile's scope (_profile_scoped(), #98738 pattern) and fall
to config.extra / module defaults instead. The default profile keeps its
unscoped env precedence.
Salvaged from #100382 (tests trimmed to two).
NousDashboardAuthProvider._verify_jwt (and the identical hunk in the
self-hosted OIDC provider) folded EVERY PyJWKClient failure into
ProviderError, which the gate translates to HTTP 503
{"detail":"Auth provider 'nous' unreachable"}. That branch fires for
jwt.DecodeError('Not enough segments') — i.e. the bearer is not a JWT at all
(an opaque peer key, a legacy token, garbage) — and for PyJWKSetError (JWKS
fetched fine, foreign kid). Neither involves reaching Portal, which is why
the hosted sjc agents in #94558 returned a fast, well-formed 503 that
survived token re-mint and instance restart while Portal was healthy.
Add one shared classifier, hermes_cli.dashboard_auth.classify_jwks_lookup_error:
only PyJWKClientConnectionError (transport) and an unexpected bare
PyJWKClientError stay ProviderError; DecodeError / PyJWKSetError /
InvalidTokenError become InvalidCodeError so verify_session() returns None
and the middleware proceeds to the next provider / refresh / 401 exactly as
the protocol documents. Both providers now use it.
Live repro (real NousDashboardAuthProvider against a local reachable JWKS
server; and the real gated web_server app): before — opaque bearer ->
ProviderError "JWKS lookup failed: DecodeError('Not enough segments')" ->
503 unreachable; after — verify_session() -> None, gated GET /api/auth/me
with the opaque bearer -> 401; a real JWT against an unreachable JWKS still
-> ProviderError (503).
This does not add /api/v1/message to the public-path allowlist (#94579):
that route has no verifier in this repo, so bypassing the gate would leave a
state-changing ingress fail-open. The correct fix is classification, which
also covers every other opaque-bearer surface.
Refs #94558
- generate() now passes kwargs.get("model") into _resolve_model(), so the
user's hermes tools pick (forwarded by the dispatcher as top-level
image_gen.model) is honored instead of silently dropped (#55893 class;
matches xai/krea/openrouter).
- Setup schema badge "internal" -> "paid" to match every other paid
image backend in the hermes tools picker.
- Tests: caller-model precedence, unknown caller model falls through,
model kwarg reaches the API payload, badge contract.
Adds a bundled image-generation backend for the Meta Model API
(https://api.meta.ai/v1), which is OpenAI-compatible. Exposes the
muse-image-1.0 model via the standard image_generate tool. This is the
image-gen companion to the already-bundled meta-ai chat provider
(plugins/model-providers/meta-ai, PR #88565).
- plugins/image_gen/meta-ai/ — provider registered as `meta-ai`, matching
the chat provider's id. Reuses the openai SDK pointed at Meta's base URL.
- Auth mirrors the chat provider: MODEL_API_KEY (Meta's documented var),
with META_API_KEY / META_MODEL_API_KEY aliases and a META_BASE_URL
override.
- Text-to-image only for now (capabilities gated); base64 (WebP) and URL
responses both handled and saved under $HERMES_HOME/cache/images/.
- Auto-loads as `kind: backend` and appears in `hermes tools` with no
central list edits, matching the other bundled providers.
- tests/plugins/image_gen/test_meta_ai_provider.py — 27 tests (metadata,
auth-alias resolution, base-url override, model resolution, generate
paths incl. b64 save, aspect mapping, URL caching, error handling).
- docs: image-generation feature page + provider-plugin built-in list.
check_requirements() runs at gateway startup before any per-profile
secret scope is installed, and the scope-less get_secret path reads
only os.environ -- so a Bitwarden-managed BUZZ_PRIVATE_KEY (only
BWS_ACCESS_TOKEN in .env) was invisible to the platform gate and Buzz
was silently skipped with a misleading install hint (#95216). When no
scope is active and the process env has no value, consult a cached
one-shot build of the profile secret mapping (build_profile_secret_scope
resolves external secret sources); an active scope still shadows this
rung entirely, so multiplexed cross-profile isolation is unchanged.
BUZZ_RELAY_URL reads in the gate now go through the same helper so an
externally managed relay passes too.