The gateway already streams by sending a first partial message and re-editing
it as tokens arrive, falling back to that path when an adapter does not support
native drafts. The Buzz adapter never implemented edit_message, so it inherited
the base stub that returns success=False and every reply was delivered in one
block when the turn finished, however long the turn took.
buzz-cli already exposes `messages edit` and `messages delete`, so no new
mechanism is needed.
One detail worth calling out for review: buzz-cli reports a NEW event id for
each edit, but the edit TARGET stays the original id, and the stream consumer
holds a single message_id for the whole stream. edit_message therefore returns
the id it was given rather than the one the CLI reports. Returning the CLI's id
would make every edit after the first address a message that was never sent.
delete_message is included because the consumer's fresh-final cleanup path
calls it when it replaces a preview rather than editing in place.
Tested: 10 new cases in tests/gateway/test_buzz_adapter.py covering the edit
target, stdin content, the returned id, echo suppression, finalize being inert,
both no-op guards, retryable vs non-retryable CLI failures, and delete. The
file goes from 33 passing to 6 failing if the adapter change is reverted while
the tests stay.
The WS NIP-42 auth path now prefers the connect()-resolved _auth_tag
(credentials-file aware, #79514) and falls back to a lazy scope-aware
_resolve_auth_tag() so a bare adapter re-auth stays profile-correct
(#98738): scoped multiplex profiles fail closed instead of borrowing
the default profile's tag from os.environ. _exec_buzz fakes updated
for the auth_tag kwarg introduced by the #83155 salvage.
Remove the stray 'return val if val is not None else default' tail left
in _unscoped_profile_secrets() when the new return was added, and note
in the docstring that the process-global cache is startup-gate-only
(review feedback on #95224).
check_requirements() runs at gateway startup before any per-profile
secret scope is installed, and the scope-less get_secret path reads
only os.environ -- so a Bitwarden-managed BUZZ_PRIVATE_KEY (only
BWS_ACCESS_TOKEN in .env) was invisible to the platform gate and Buzz
was silently skipped with a misleading install hint (#95216). When no
scope is active and the process env has no value, consult a cached
one-shot build of the profile secret mapping (build_profile_secret_scope
resolves external secret sources); an active scope still shadows this
rung entirely, so multiplexed cross-profile isolation is unchanged.
BUZZ_RELAY_URL reads in the gate now go through the same helper so an
externally managed relay passes too.
One BUZZ_* read survived the #98738 sweep unscoped: the NIP-42 WebSocket
auth path read BUZZ_AUTH_TAG with a bare os.getenv. Under
gateway.multiplex_profiles the process env holds the default profile's
bridge/.env output, so a scoped secondary profile without its own tag
signed its relay auth event with the default profile's NIP-OA
owner-attestation tag. Reproduced on f3845a72af before the fix; the same
repro now attaches no tag (fail-closed).
The read goes through _get_scoped_secret: scoped multiplex profiles fail
closed to "", while single-profile and unscoped default-profile reads
keep the legacy env behavior. Adversarial coverage added for the leak
itself, the scoped positive control, unscoped precedence, partial-extra
adapter config, scoped validate_config, scoped standalone-send target
resolution, central-authz wildcard/blank-entry/normalization semantics,
and adapter-intake vs central-authz agreement on the same allowlist.
Fixes#98738
Signed-off-by: Kosta Gorod <35299380+KostaGorod@users.noreply.github.com>
Under gateway.multiplex_profiles the default profile's YAML-to-env bridge
writes BUZZ_* values into os.environ, and every Buzz read gave that env
precedence over the secondary profile's PlatformConfig — so each secondary
adapter connected as the default identity, watched its channels, and
resolved its credentials file (#98738).
- Add _profile_scoped()/_scoped_platform_setting(): inside a secondary
profile scope extra is authoritative and env is not consulted (a missing
key fails closed to its default instead of borrowing the default
profile's value); single-profile and unscoped/default-profile reads keep
the legacy env-over-config precedence.
- Apply the scoped read to BuzzAdapter.__init__ (relay, CLI path, channels,
home channel, poll interval, require_mention, transport, allowed users),
_resolve_private_key (BUZZ_CREDENTIALS_FILE), validate_config,
_standalone_send, and check_requirements (which now consults the
profile's own config.yaml via the scoped home override).
- _env_enablement() returns None inside a profile scope and
_apply_yaml_config() skips the env bridge there, so the default profile's
env cannot fabricate Buzz for a profile that never configured it and a
secondary profile's YAML cannot be pinned into the process env
(first-writer-wins, #72348 Telegram/Discord mirror).
- Central authorization now consults a plugin platform's live-adapter
config.extra.allowed_users (gated on the registry entry declaring
allowed_users_env, with an optional normalize_user_id hook so Buzz npub
entries match hex-pubkey user ids) — under multiplex only the default
profile's list ever reached the env var, so listed secondary-profile
users were default-denied (#82871). Empty/absent lists change nothing;
default-deny is preserved.
The Discord adapter resolves username allowlist entries to numeric IDs at
connect and mirrors them into os.environ — but the gateway's per-turn .env
hot-reload (load_hermes_dotenv(override=True)) restores the raw usernames
from the file. From the second agent turn onward, _is_user_authorized
compared numeric user_ids against username strings and dropped every
message from the operator as 'Unauthorized user' while the adapter layer
still admitted them (bot reacted, never replied).
Fix: gateway authz unions the adapter's resolved numeric IDs
(DiscordAdapter.resolved_allowlist_user_ids()) into the env-derived
allowlist. Union only fires when an env allowlist is configured (never a
widening; fail-closed branch unchanged), is duck-typed + isinstance-guarded
against mock adapters, and filters non-numeric entries so unresolved
usernames and '*' can't leak through adapter memory.
Live repro: symptom fired on origin/main (authorized=False after reload),
passes with fix; stranger + empty-allowlist + raising-resolver negatives
hold. Sabotage run: incident test fails on unfixed authz_mixin.
Remove the implicit hermes peer and the peer question from new connection setup. Preserve explicit peer settings and keep memory paths consistent with the captured client identity.
Add setup, configuration, request, recall, and session regression tests, plus upgrade guidance.
- Tavily plugin deleted (plugins/web/tavily), keyless endpoints and
ring entry removed from keyless_mcp, legacy backend set / credential
ladder / preference walks / rescue key map scrubbed.
- TAVILY_API_KEY deregistered across config, setup, status, dump, and
nous_subscription surfaces. The tvly- redaction pattern stays --
legacy keys in user envs still deserve masking.
- Sibling test pins migrated (keenable/exa stand in where tavily was
the fixture vendor); tavily test suite deleted.
- Docs updated: web-search, configuration, integrations,
environment-variables, tools-reference, web-dashboard, provider
plugin dev guide.
Live-verified from an isolated HERMES_HOME with all web creds blanked:
zero-config resolution lands in the 4-vendor ring, live keyless ring
search succeeds, no tavily anywhere in resolution order.
_looks_like_connect_timeout and _looks_like_pool_timeout carried two
copies of the same 15-line DFS skeleton (seen-set, stack, __cause__/
__context__ descent) differing only in the one-line match predicate —
follow-up to the #98094 review.
Extract _iter_exception_graph() and collapse both classifiers onto it.
Behavior is byte-identical (subprocess parity vs origin/main on real PTB
error fixtures: 6/6 identical), and the two classifiers gain direct unit
tests for the first time, including the cycle/diamond chain shapes the
inline copies had no coverage for.
_drain_polling_connections still bounded its shutdown()/initialize() with
asyncio.wait_for (#66377), while its sibling the general-pool drain moved
to _await_with_thread_deadline (#98094). httpcore's pool close runs under
AsyncShieldCancellation, so a cancellation-resistant close keeps wait_for
pending forever even after its timeout fires — the tracked
_polling_error_task wedges and every escalation gate behind it stalls.
Use the same wall-clock deadline helper (cancel + abandon, no cancel-await)
on both polling-drain awaits, and add a regression test whose close
swallows cancellation — the shape the existing cancellable-hang test
cannot catch.
Follow-ups on top of #98964's cherry-pick:
- PHOTON_READ_RECEIPTS env toggle (default true) so users can keep
messages at Delivered; declared in plugin.yaml optional_env
- adapter drops both 'read' and 'read_receipt' content types (alias
coverage from #91759 by @mooserini) + regression test
- docs: photon.md feature note + environment-variables.md row
Telegram's BotCommand menu is hard-capped (100/scope, ~4KB payload; Hermes
defaults to 60 slots), so most skill commands can never appear in the /
menu. Inline mode has no such cap: typing @botname <query> in any chat now
returns a live, searchable picker over EVERY core command, plugin command,
and installed skill — results computed per keystroke, paginated 50 at a
time. The Telegram analog of Discord's dynamic /skill autocomplete
(#18741).
- plugins/platforms/telegram/inline_picker.py: PTB-free catalog/rank/
pagination logic (unit-testable without python-telegram-bot). First
query token filters; the remainder is carried into the sent command as
its argument (@bot plan migrate auth → sends /plan migrate auth).
- adapter: InlineQueryHandler registration (inert until the bot owner
enables inline mode via BotFather /setinline) + _handle_inline_query
with the same auth path as inline-button callbacks — unauthorized users
get an empty list, so the installed-skill catalog is not leaked to
arbitrary users (inline queries arrive from any chat).
- Tap-to-send dispatches through the existing command path: the sent
message starts with /, which reaches the bot even under default privacy
mode. Zero new dispatch code.
- Docs: telegram.md inline-picker section incl. the one-time BotFather
/setinline setup.
Sweeper review, all three points:
- Placement: no new plugins/model-providers/ directory. The Token Plan
profiles register from the existing alibaba plugin module — one module
per vendor, matching how the kimi module carries both of its endpoint
variants. Token Plan is the same vendor/service (Model Studio), same
OpenAI-compatible protocol, its own key + endpoints; splitting to a
standalone repo remains a 5-minute change if maintainers prefer.
- Runtime coverage: TestRuntimeAlibabaRegionalAndTokenPlan exercises
resolve_runtime_provider() for all four variants — provider, api_key,
api_mode, base_url — alongside the existing zai/minimax/kilocode
runtime regressions.
- Docs: providers.md, environment-variables.md, cli-commands.md updated
with the bundled variants and their env keys/base-url overrides.
The models.dev catalog advertises alibaba-cn, alibaba-coding-plan-cn, and
alibaba-token-plan(-cn), and resolve_provider_full()'s catalog chain lets
the CLI --provider path resolve them — but auth.resolve_provider() (the
credential/runtime path used by 'hermes chat') consults only
PROVIDER_REGISTRY and raised "Unknown provider 'alibaba-coding-plan-cn'"
(hermes_cli/auth.py:1937). PROVIDER_REGISTRY auto-extends from provider
profiles (auth.py:461-490), so the fix registers the missing profiles at
that chokepoint: alibaba-cn joins the alibaba plugin, alibaba-coding-plan-cn
joins alibaba-coding-plan, and a new alibaba-token-plan plugin registers
both regional token-plan tiers. Names match the catalog keys exactly.
No core edits — plugins/model-providers is the designed extension path.
The chat_completions and transport-parity Mistral tests pinned the
same branch with different reasoning_config shapes ({effort:none} vs
{enabled:False, effort:none}) — behaviorally identical since effort
short-circuits first, but the drift reads as a semantic difference.
Align both to the explicit form. Also scope the port-guard comment to
the try/except shape it actually shares with hermes_cli/models.py.
urlparse raises ValueError on non-integer / out-of-range ports, and
http://myhost:99999/v1 passes OpenAI-client construction (only httpx
rejects it later), so the crash was reachable from build_kwargs on
every request for such a URL. Wrap the parsed.port check in the same
try/except ValueError guard hermes_cli.models already uses around its
11434 check, and pin it with parametrized tests.
reasoning_effort: none was injecting extra_body.think=false for every
custom provider. Mistral (and other extra=forbid hosts) reject that
field with HTTP 422. Keep think=false on Ollama URLs only; still send
top-level reasoning_effort=none so /v1 thinking-off keeps working.
Review finding on salvaged #28253: the hand-rolled mapping sent
ultra -> medium while xhigh -> high (stronger request, weaker wire
value). Declare NEBIUS_EFFORTS in agent/reasoning_effort.py and use
clamp_effort like the zai/kimi/tokenhub call sites; disable detection
stays ahead of the clamp since clamp_effort('none', ...) returns the
floor, not off. Adds a monotonicity regression test.
Review findings on salvaged #93548: _warm_efforts_async now returns
early under PYTEST_CURRENT_TEST (matching the canonical OpenRouter caps
warmer) so a test that forgets to monkeypatch it can't fire live HTTP
when RAMP_ROUTER_API_KEY is set; the codex transport's fail-open
except in _profile_declared_efforts logs at debug instead of silently
swallowing profile-hook bugs.
Router shipped a minimal /v1/chat/completions compatibility surface
(translated onto Responses) after this PR was written, so the
'does not exist and 404s' wording is stale. Responses remains the
native wire — per-model reasoning-effort validation, reasoning
summaries, and prompt caching live there — so the api.router.com
host mandate is unchanged; only the comments and docs are updated.
Addresses the automated review on this PR:
- _profile_declared_efforts falls back from provider name to the
endpoint's host (via model_metadata's URL->provider map), so a named
custom provider pointed at api.router.com — which the host mandate
already routes onto this transport — gets the catalog clamp instead
of the default vocabulary and a Router 400.
- _parse_efforts validates catalog levels against EFFORT_LADDER at
ingest, logging and dropping unrecognized tiers; a model whose whole
vocabulary is unrecognized stays out of the map (transport defaults)
instead of passing the requested effort through unclamped.
- fetch_models dedupes ids while preserving Router's deliberate listing
order.
- plugin.yaml credits the human contributor per repo convention.
Router attributes coding-agent clients by User-Agent prefix (the way it
already recognizes OpenCode's versioned UA), and its WAF rejects
default/blank client UAs. Mirror the xai profile: declare
User-Agent: Hermes-Agent/<version> in default_headers, which
agent_init's profile-headers fallback applies at client construction.
Re-verified live: one-shot chat through api.router.com still works.
Ramp Router is an OpenAI Responses-compatible LLM gateway at
https://api.router.com/v1 that routes each request across upstream
providers (OpenAI, Anthropic, xAI, Fireworks, ...) with server-side
fallbacks and spend controls. Nous asked for a PR adding it as a
provider, so:
- plugins/model-providers/router/: RouterProfile plugin —
api_mode=codex_responses, RAMP_ROUTER_API_KEY auth,
RAMP_ROUTER_BASE_URL override, live account-scoped catalog via
GET /v1/models (no hardcoded fallback_models: IDs are key-scoped and
Router's docs mandate runtime catalog reads).
- hermes_cli/providers.host_mandated_api_mode +
runtime_provider._detect_api_mode_for_url: api.router.com ->
codex_responses. The host is Responses-only — POST /v1/chat/completions
does not exist and 404s — so this is a genuine host mandate (exact
hostname match per #32243, mirroring the api.meta.ai precedent).
- providers/base.py: new overrideable supported_reasoning_efforts(model)
hook (tri-state: None=defer, ()=model takes no reasoning params,
tuple=clamp set). Router validates reasoning.effort per model and
returns HTTP 400 invalid-argument on levels outside the model's
published vocabulary, and 400 unsupported_parameter when a
non-reasoning model receives any reasoning field (both verified live).
The profile answers from a cached copy of the catalog's
router.capabilities.reasoning block: cache-only on the hot path,
seeded for free by fetch_models(), disk-mirrored across processes
(/cache/router_catalog.json), background-warmed when cold
— same design as the OpenRouter reasoning-caps clamp on the chat path.
- agent/transports/codex.py: consult the profile-declared vocabulary in
the generic effort-clamp branch (xai/actual/github branches untouched;
profiles that do not override the hook see no behavior change).
- cli-config.yaml.example + adding-providers.md + providers/README.md:
document the provider, the host mandate, and the new hook.
- tests: behavior contracts for the host mandate/URL detection/spoof
rejection, profile registration + auth auto-registry wiring, catalog
parsing, and transport clamp/suppression/fallback paths.
Verified live against api.router.com (Aug 2026): one-shot chat,
streaming SSE, tool calls + parallel_tool_calls, encrypted-reasoning
replay on OpenAI-served models, function_call_output follow-up turns on
OpenAI- and Fireworks-served models; store:false / prompt_cache_key /
include:[reasoning.encrypted_content] / reasoning.summary accepted
across backends; effort clamp confirmed to convert a would-be 400
(xhigh on o3) into a successful request via the disk mirror.
/bg (formerly /background, which is retired) keeps the existing semantics:
spawn a fresh, independent agent session in the background.
/btw is now its own command matching the convention other harnesses use:
ask a quick side question ABOUT the current conversation without
interrupting it. A one-shot auxiliary LLM call (main model by default,
overridable via auxiliary.side_question.* in config.yaml) answers from a
read-only transcript snapshot — the live session's history, role
alternation, and prompt cache are untouched, and the current turn keeps
running.
Surfaces wired: CLI (inline mid-run dispatch), gateway (all messengers,
busy-dispatch table + idle dispatch, i18n across all 17 locales), TUI
(prompt.btw RPC + btw.complete event), Discord native slash, relay
command manifest, desktop exec routing, docs (EN + zh-Hans).
Second follow-up for salvaged PR #94547, folding in review findings from the
duplicate-PR cluster (#47015, #55054, #58476, #72977, #73685 all fix the
same 401) and the sweeper review of #73685:
- Replace the dot-anchored suffix predicate with exact-match against the
existing _ALLOWED_TEAMS_SERVICE_HOSTS allowlist (two of the five
duplicate PRs converged on this independently). Any Azure customer can
register <name>.trafficmanager.net profiles, so suffix matching was not
safe. Also requires https on the default port — :444 on an allowlisted
host no longer receives the bearer (sweeper finding on #73685).
- Stream _fetch_attachment_bytes through _read_httpx_body_with_limit
instead of buffering response.content — the shared inbound media cap now
applies to authenticated downloads too (sweeper finding: a lying
Content-Length must not OOM the gateway).
- Serialize token refresh with a lazily-bound asyncio.Lock so concurrent
attachments share one STS POST (review finding on #94547).
- Token expiry now uses time.monotonic() (from #55054) — wall-clock jumps
can't extend a stale token.
- Tests updated: exact-allowlist predicate (lookalike/subdomain/port/scheme
negatives), streaming fake client, and a concurrent-cold-cache lock test.
Mutation-checked: suffix match, silent drop, no-lock, and unbounded
buffer each fail a test.
Follow-up for salvaged PR #94547 (Sibbern's Bot Framework attachment auth):
- The host check used bare endswith('trafficmanager.net') /
endswith('botframework.com'), which matched attacker lookalike hosts
(evil-trafficmanager.net) — sending the bot's bearer token off-platform.
Same threat model _ALLOWED_TEAMS_SERVICE_HOSTS already documents. Now a
single dot-anchored predicate, _is_botframework_attachment_host(),
used by both _fetch_attachment_bytes and the _on_message image branch
(was copy-pasted in two places).
- BF images whose bytes fail image validation were silently dropped with
no log (cache_media_bytes returns None; old path logged). Add the
missing else-warning, mirroring the document branch.
- Record cached_m.media_type instead of the raw content_type so the
MessageEvent MIME matches what was actually cached.
- Init _bf_token_cache in __init__ (was masked by getattr).
- Tests: host predicate dot-anchoring (attacker lookalikes blocked),
BF routing vs generic helper, bearer attach + attacker-host block end
to end, token acquire + cache reuse, token-failure degradation, and
the silent-drop regression guard. Mutation-checked: reverting the
dot-anchor, the else-warning, or the token cache each fails a test.
Inline/pasted images in Teams arrive with a contentUrl on
smba.trafficmanager.net (/v3/attachments/...). Unlike file uploads
(pre-authenticated SharePoint downloadUrls), these connector URLs
require the bot's own bearer token; fetching them anonymously fails
with 401 Unauthorized and the image is silently dropped.
- add _get_botframework_token(): client-credentials token for
https://api.botframework.com/.default, cached until ~5 min before
expiry
- _fetch_attachment_bytes(): attach the token when the attachment
host is *.trafficmanager.net / *.botframework.com; SharePoint and
other URLs remain auth-free as before
- route image/* attachments with Bot Framework contentUrls through
the authenticated fetch instead of cache_image_from_url (which
sends no Authorization header)
SSRF guards unchanged. Verified on a live Teams personal-scope bot:
pasted images previously logged '[teams] Failed to cache image
attachment: 401 Unauthorized' and now cache and deliver correctly.
POST /boards/{slug}/export and POST /boards/import, so the desktop and
dashboard can drive board transfer. Both exchange filesystem paths
rather than bytes, the same contract profile export/import uses and for
the same reason: the client runs its native save/open dialog on the
machine that hosts the backend, so a path is all either side needs, and
a board carrying a few hundred megabytes of attachments never has to
cross the renderer heap.
Also covers the pre-existing rename (PATCH) and delete endpoints, which
had no tests — including that delete archives to a restorable directory
and refuses to touch `default`.
* feat(a2a): outbound client tools are config-gated — served only when a2a_agents configured, inbound platform enabled, or A2A_PORT set (-561 tok/call on unconfigured installs)
* ci: retrigger after runner startup_failure on rerun attempt
* refactor(video_generate): capability-gated dynamic schema — 6 optional args render only when the active provider/model honors them; fleet capability declarations + declaration<->implementation contract tests
* fix(video_gen): H3/Grok/Happy-Horse/Gemini audio is ALWAYS-ON native, not absent — new audio_native family key + audio_always_on capability surfaces as description line (maintainer catch)
* test(video_gen): duration-span test pins the active-model contract — resolved family's real window, short families not inflated, union fallback still spans 30s
Platforms added to main after the original branch was cut; keeps the
source invariant (every connectable adapter calls _wire_plugin_handlers)
true, and adds qqbot to the invariant test's gateway list.
ctx.register_platform_handler(platform, factory) — the generic surface for
plugins to wire native handlers into any platform adapter at connect()
time. Factories receive (native, adapter): the platform's client/app
object (PTB Application, discord.py Bot, slack_bolt AsyncApp, Teams App,
DingTalkStreamClient, aiohttp web.Application) or None for adapters with
no separate native object.
- BasePlatformAdapter._wire_plugin_handlers(native): shared, isolated
invocation helper — a raising plugin cannot block a platform connect.
- All 27 connectable adapters call it: telegram/slack/teams/line/
api_server/msgraph_webhook wire before their dispatch tables freeze;
the rest hook at connect success.
- register_telegram_handler and get_telegram_handler_factories retained
as thin back-compat aliases over the telegram bucket.
- Source-invariant test guarantees every adapter with connect() keeps
calling the hook.
Mirrors the Slack precedent (register_slack_action_handler): plugins queue
a factory at register() time; the Telegram adapter invokes each factory
with (application, adapter) at connect() time, before the core handlers
register, so pattern-scoped plugin handlers take precedence for their own
updates while everything else falls through unchanged. Factories are
isolated — a raising plugin cannot prevent Telegram from connecting.
Unblocks standalone plugins that need PTB update types the core adapter
doesn't route (Telegram Business API secretary bots, custom callback
prefixes, chat-member events) without touching core files.
Fixes two related native-streaming bubble defects surfaced in production:
1. Duplicate bubble on long turns — when keep-alive already refreshed the
6-min reply window, the Layer-2 clock fallback still declined the finalize
frame and forced a proactive send(), duplicating the message. Skip the
clock fallback while keep-alive is active; intermediate-frame failures are
now fully fire-and-forget (only a failed FINAL frame falls back to send()).
2. Split / mini bubbles ('Cla' + 'ude ...') — two compounding root causes:
a) In native streaming a mid-turn commentary (e.g. a Hindsight recall
notice) called _reset_segment_state(), clearing the cumulative
_accumulated so the next delta + finalize frame carried only the few
chars accumulated after the reset. Native streaming now skips that
reset (commentary still posts as its own message via send()).
b) The adapter-side _BlockChunker.update() 'only grow' guard silently
dropped any cumulative snapshot shorter than its high-water mark, so
after a baseline reset the leading characters were stranded before
_emitted_len. Removed the _BlockChunker sentence-alignment + idle-flush
layer entirely; intermediate frames are pure identity-dedup, matching
the fire-and-forget model.
Also removes ~232 lines of now-dead code (_BlockChunker class, idle-flush
machinery, block-stream constants) and aligns the test suite with the
fire-and-forget frame model, including a regression test that locks the
native-commentary-no-reset behavior.
Tests: 177 passed, 3 skipped (wecom + stream_consumer suites).