- Key _user_space_cache on _conn_snapshot instead of client object identity,
so _new_client() results from the same connection share the cached user
(previously every on_memory_write triggered an uncached /api/v1/system/status
probe with a 30s default timeout)
- Thread a short timeout (0.05s) through the write-path identity probe
- Harden _tool_remember to snapshot the client before URI construction + POST,
matching the pattern already established in on_memory_write
- Remove dead instance method _user_scoped_uri (zero callers; all call sites
use the module-level function directly)
Co-authored-by: ehz0ah <haozhe4547@gmail.com>
Review follow-up to the viking://~ migration: the ~ home alias only
expands for USER/ADMIN roles. The DEFAULT dev auth mode (no
server.auth_mode, no root_api_key) resolves every request as ROOT,
which bypasses current-user expansion — the canonical parser rejects
viking://~ with 400 'Home alias URI is not canonical' (verified on a
live 0.4.16 server). A deployment upgrading to 0.4.16 with an
untouched ov.conf is in dev mode, so the ~ spelling would break
exactly the way the old uid-less one will.
Mirror the upstream first-party plugin pattern instead: resolve the
user space client-side from /api/v1/system/status (result.user,
'default' fallback) and emit explicit-uid
viking://user/<user>/memories/... URIs, which are canonical under
every auth mode (dev/ROOT, trusted/USER, api-key) and every server
version. viking://~/... input typed by the user keeps passing through
untouched. (#91995)
Upstream OpenViking removed the uid-less viking://user/<segment>
shorthand (#4196, merged 2026-08-21): reserved segments like memories
and peers no longer expand to the caller's space and the server
rejects them with HTTP 400 (NamespaceShapeError). First-party clients
were migrated to viking://~ in the same change; the Hermes plugin was
not (#91995).
Migrate every URI the plugin constructs — the profile/preferences/
entities session-start reads, the _build_memory_uri memory-mirroring
write path, and the tool-schema example — to viking://~/... README
uid-less references updated to match; canonical user-scoped forms
(viking://user/default/...) are unchanged. The ~ alias requires
OpenViking server >= 0.4.16 (#4167).
A short Telegram drop used to fail the final reply immediately. The
answer then sat in the delivery ledger until the next gateway boot.
Wait up to 15s for the bot (or a replacement adapter) so a brief
blip delivers now, matching QQBot.
find_spec("microsoft_teams") can be true from sibling namespace packages
while App is still None, so a failed lazy-install crashed connect with
'NoneType' object is not callable instead of a missing-SDK error.
Probe microsoft_teams.apps via the parent first — a dotted find_spec
raises ModuleNotFoundError on 3.11 when the namespace is absent.
The adapter's private thread-deadline helper was the ancestor of the
unified deadline layer's run_bounded_async (#85147 was extracted from
it, plus the caller-cancellation leak fix the original still lacked).
Consolidate: the helper body becomes a thin wrapper mapping
BoundedResult.timed_out back to the asyncio.TimeoutError its 9 call
sites (the PTB retry ladder) expect. ~90 duplicated lines die, along
with the adapter-local copies of the abandon-cleanup runner and the
blocked-loop faulthandler diagnostics (both live in agent/deadline.py).
Everything the call sites rely on is preserved by the unified layer:
- thread-timer deadline that survives a blocked event loop (#63309)
- abandonment of cancellation-shielded tasks (PTB/httpcore anyio init)
- detached best-effort on_abandon cleanup (no httpx pool leak per retry)
- off-loop stack dump when the loop never processes the expiry
Plus one behavior IMPROVEMENT inherited from the shared copy: a caller
cancelling the wrapper no longer leaks the inner task unobserved (the
telegram original had that leak; the extraction fixed it).
test_telegram_init_deadline.py: the #63309 diagnostics probe now pins
the shared layer's dump hook (label "telegram-init") — same contract,
new seam. Wedge + cleanup-crash tests pass unchanged.
Adds an optional occurred_at (ISO-8601 date/datetime) parameter to the
hindsight_retain tool schema, threaded into the retain item's timestamp
field. When absent, the item timestamp defaults to the configured event
clock (base from PR #82928 by @ragingbulld, authorship preserved) so the
Hindsight server can resolve relative time phrases; previously no item
timestamp was ever sent and temporal memories landed with null
occurred_start/occurred_end.
Fixes#93568. Salvages #82928.
Use Hermes timezone-aware timestamps for retained events and turn messages. Pass the public timestamp field supported by hindsight-client 0.6.1 and cover the final serialized request field.
Under gateway.multiplex_profiles, secondary profiles are constructed
inside _profile_runtime_scope and their .env lives in the profile's
secret scope - gateway/run.py explicitly does NOT mutate os.environ with
it. Four adapters still read their AUTHORIZATION config via raw
os.getenv, so every secondary profile either (a) silently missed its own
env-only allowlists/policies (fail-closed: all DMs dropped at intake) or
(b) inherited the default profile's GATEWAY_ALLOW_ALL_USERS=true /
allowlists from the shared process env (fail-open admissions):
- weixin.py: WEIXIN_DM_POLICY / WEIXIN_ALLOWED_USERS /
WEIXIN_GROUP_ALLOWED_USERS / WEIXIN_ALLOW_ALL_USERS +
GATEWAY_ALLOW_ALL_USERS in _open_dm_opted_in
- yuanbao.py: YUANBAO_DM_POLICY / DM_ALLOW_FROM / GROUP_POLICY /
GROUP_ALLOW_FROM / ALLOW_ALL_USERS (new _yb_secret helper; AccessPolicy
hard-gates intake)
- signal.py: SIGNAL_GROUP_ALLOWED_USERS / SIGNAL_ALLOWED_USERS (new
_sig_secret helper; empty scoped group list previously meant "drop all
groups" silently)
- wecom/adapter.py: WECOM_DM_POLICY / WECOM_ALLOWED_USERS /
WECOM_GROUP_POLICY / WECOM_ALLOW_ALL_USERS + GATEWAY_ALLOW_ALL_USERS -
while credentials one line above already used _get_scoped_secret
- gateway/run.py::_own_policy_open_startup_violation: the open-policy
startup guard validated GATEWAY_ALLOW_ALL_USERS via raw os.getenv even
though its sibling dm/group reads already used the scoped _getenv
All reads now go through the canonical fail-closed scoped shape QQ's
_resolve_qq_secret already used (scope hit wins; unscoped single-profile
callers keep legacy os.environ behavior). Regression suite drives the
real scope contextvar across all four helpers plus the admission gates
and the startup guard, asserting both directions: profile values are
visible under multiplex, default-profile values never leak.
Fixes#93522
WEIXIN_DM_POLICY/ALLOWED_USERS/GROUP_ALLOWED_USERS, YUANBAO's equivalents,
WECOM_DM_POLICY/ALLOWED_USERS/GROUP_POLICY, and the startup guard's
GATEWAY_ALLOW_ALL_USERS check still read raw os.getenv at adapter
construction time. Under gateway.multiplex_profiles that reads the process
env instead of the per-profile secret scope, so a secondary profile either
silently drops every DM (its own env-only allowlist is invisible) or
inherits the default profile's allow-all/allowlist config.
Route these reads through the existing scoped helpers (_wx_secret,
_get_scoped_secret, gateway.authz_mixin._platform_gate_env, and
gateway.config._getenv) already used for the adjacent credential reads in
the same adapters.
Fixes#93522.
Extract _FLOOD_INLINE_WAIT_CAP_SECS + _flood_cap_result so the 5s cap
and the flood_control:{wait} error contract cannot drift between the
edit path and the send path #92173 added.
Final-review follow-up: swap the bare [:25] slice for the new constant.
Behavior-identical (same 25); removes the last bare option-cap literal
in the file. ChoicePickerView feeds finite /reasoning and /fast choice
lists, so no functional change.
Final-review follow-up: replace the bare 75 in the shown-count with
_DISCORD_MODEL_SELECT_CAPACITY so it can never desync from what the
partitioned menus actually render.
Telegram RetryAfter on send() slept the server retry_after with no
ceiling, so a 97-minute penalty pinned the coroutine. Mirror the edit
path: waits over 5s return immediately; short waits still retry inline.
Review follow-ups for the salvaged picker partitioning:
- Drop the _model_chunks stash — it went stale when navigating to a
provider with an empty model list (early return skipped the
reassignment), producing a wrong 'N more available' count. Derive
shown = min(len(models), 75) directly instead.
- Add _DISCORD_SELECT_MAX_OPTIONS / _DISCORD_SELECT_MAX_ROWS constants
per the file's named-limits convention; replaces 4 bare literals.
- Remove dead total_rows variable.
The Discord /model picker built a single discord.ui.Select filled with
models[:25], silently dropping any models beyond the first 25. Discord
caps a single select at 25 options but allows up to 5 component rows, so
partition the list across up to 3 select menus (25 each; Back/Cancel use
the other 2) instead of truncating.
This fixes providers like Nous (curated list + Portal recommendations
exceed 25) whose tail — including free-tier :free Portal picks — was
previously clipped on Discord while showing fine in the Portal UI / CLI.
- _build_model_select: slice into <=25-option chunks, one select per chunk
(custom_id model_model_select_<i>), all via _on_model_selected. Multi-row
menus get a (n/total) placeholder suffix.
- _on_provider_selected: 'N more available' count reflects models actually
rendered across the partitioned menus.
- Add regression test covering the 37-model Nous case (no truncation/dupes,
per-menu 25 cap holds).
prefers_fresh_final_streaming read only the raw direct_messages_topic_id
key; the adapter's canonical accessor _metadata_direct_messages_topic_id
also accepts the documented telegram_direct_messages_topic_id alias
(treated as equivalent in gateway/delivery.py), so an alias-only lane
would still flatten tables. Route the gate through the accessor and pin
the alias with a regression (mutation-checked: raw-key gate fails it).
Also reshape the happy-path endpoint assertion into the actual invariant
(sendRichMessage present, no rich draft frames) instead of a frozen call
list. Surfaced during review of PR #91436.
#91241 stopped root-DM tables collapsing to bullets by keeping native
draft transport when rich_drafts is off. Private Telegram topics still
reject sendMessageDraft (string thread ids, forum-style thread fields),
so the stream consumer falls back to edit-in-place. Telegram then
rejects a rich edit of that plain MarkdownV2 preview and format_message
permanently rewrites pipe tables into bullet lists — the remaining
report after that merge.
Route drafts through the same integer topic kwargs as send(), and on
that degraded topic path prefer a fresh sendRichMessage (then delete
the preview) instead of the table-to-bullets formatter.
GLM-5.3 accepts a graded low/medium/high/max reasoning_effort scale
(verified live in #91789: monotonic reasoning-token scaling, no 400s),
but the effort mapper reused GLM-5.2's two-level vocabulary, silently
rewriting low/medium to high. Adds GLM53_EFFORTS/GLM53_OVERRIDES and a
per-model vocabulary pick in the zai plugin; 5.2 keeps its high/max
clamp. Closes#91789. Also covers the gap noted when closing #86947
(credit @santhanakrishnan-d and @terje1965 for the graded-scale finding).
The network-error reconnect path (PR #91524) was the only site converted
from asyncio.wait_for to _await_with_thread_deadline. The same
cancellation-shielding vulnerability exists at two more updater.stop()
sites:
- Conflict-retry path: asyncio.wait_for could hang forever if PTB/AnyIO
cleanup swallowed CancelledError, stalling the conflict-retry ladder.
Now uses _await_with_thread_deadline and escalates to fatal on timeout
(same reasoning: cannot safely reuse an Updater whose lifecycle lock
may still be held).
- Conflict-exhausted fatal path: asyncio.wait_for could hang before the
fatal notification fired. Now uses _await_with_thread_deadline; the
timeout handler already proceeds to fatal notify, so no behavior change
beyond the deadline mechanism.
All three asyncio.wait_for(updater.stop()) sites now use the
thread-deadline helper consistently.
Use the existing wall-clock deadline helper for updater.stop() during network recovery. If PTB cleanup remains cancellation-shielded past the deadline, escalate to retryable fatal recovery so the runner builds a fresh adapter instead of calling start_polling() while the old Updater may still hold its lifecycle lock.
Add regression coverage with stop() swallowing cancellation while holding the same lock start_polling() needs, and verify the old Updater is never reused.
GLM-5.3 is live on api.z.ai (coding plan endpoint) but had no entries in
Hermes, so it silently fell back to the generic 202K GLM context —
triggering premature context compression on a 1M-window model.
- model_metadata: 'glm-5.3': 1_048_576 (same base model as 5.2; 1M
context / 128K max output per docs.z.ai/guides/llm/glm-5.3, verified
2026-08-14)
- auth: add glm-5.3 to coding-plan probe lists (global + CN)
- models: add glm-5.3 to picker/model lists (6 sites)
- zai provider: reasoning_effort mapping covers glm-5.3 (accepted live
by the endpoint, HTTP 200)
Live verification (2026-08-21): big-pickle and mimo-v2.5-free return 429
FreeUsageLimitError for ANY User-Agent except the opencode CLI's own
'opencode/latest' — same IP, no cooldown effect, while the other six free
models serve our honest HermesAgent UA freely. Hermes sends deliberate
attribution headers and does not impersonate other clients, so these two
models are broken for our users by policy on OpenCode's side; delist them
rather than ship dead picker entries.
- opencode-free catalog: 8 -> 6 models (both curated lists)
- plugin default_aux_model: big-pickle -> laguna-s-2.1-free (fastest
non-gated free model)
- keyless predicate keeps big-pickle (it IS free-tier; correct routing if
a user enters it manually — they get the relay's own 429, not our 401)
E2E: picker shows 6, laguna aux default completes a keyless agent turn.
Widens the salvaged #91323 fix (@vinsew): the effort vocabulary moves to
agent.reasoning_effort (OX_ALPHA_EFFORTS/OVERRIDES, the declared-policy
home every other model vocabulary lives in), and the translation is
shared between the opencode-zen profile and the keyless opencode-free
profile — Ox Alpha is reachable through both, and the free profile
previously dropped effort entirely.
Live-verified: medium clamps to low (raw medium 400s: 'This model always
engages in thinking... use low, high, or max'), xhigh rounds to max, and
full agent turns with effort=medium complete on BOTH providers.
OpenCode documents x-preview-f-free as accepting low, high, and max reasoning effort on its Zen Chat Completions endpoint. Hermes previously resolved the user's per-model override to max but the plain Zen provider profile discarded it, so successful calls silently ran at the server default.
Introduce an OpenCodeZenProfile scoped only to x-preview-f-free. It forwards the normalized top-level reasoning_effort, maps xhigh to max, preserves server defaults when unset or disabled, and leaves every other Zen model untouched.
Add profile and full transport tests that prove max reaches the outgoing request and that non-target models are unaffected. Also correct the nanoid security-pin comment to match the already-locked 3.3.18 release.
Reworks the salvaged OpenCode Free provider to match the tier's real
auth contract (verified live 2026-08-21): the Zen relay serves free
models ANONYMOUSLY and 401s any unrecognized bearer, so the provider now
declares no credentials at all and routes every model through the shared
keyless machinery from the Ox Alpha fix (empty Authorization default
header overriding the SDK bearer).
On top of the salvaged base:
- auth.py: no api_key_env_vars; drop the keyed-auth special case
- runtime_provider.py: restore the plain fail-closed path (opencode-free
never reaches it — the keyless runtime resolves first)
- models.py: opencode-free joins the opencode family (prefix stripping,
Zen endpoint routing incl. muse->responses); keyless predicate extended
with unsuffixed free slugs (big-pickle); free runtime pins EVERY
opencode-free model keyless; curated catalog replaces the models.dev
cost==0 filter (it lags reality: deepseek-v4-flash-free stayed 'free'
there after its promo ended and the relay began 401ing it — delisted)
- agent_runtime_helpers.py: replace the httpx transport-sharing auth-strip
wrapper with the shared header policy (no proxy-mount loss)
- model_setup_flows.py: skip the API-key prompt for opencode-free
- plugin profile: keyless headers, no env vars
- .env.example + providers.md: keyless docs (no OPENCODE_FREE_API_KEY)
- tests rewritten to the keyless contract, incl. catalog-membership
invariant (every curated model must satisfy the keyless predicate)
E2E: full AIAgent turns with zero keys complete on x-preview-f-free via
provider opencode-free and alias 'free', incl. a real terminal tool
round-trip; muse routes to /v1/responses; picker lists 8 keyless models.
Adds an OpenCode Free provider plugin. Free model discovery uses models.dev
(cost.input == 0 AND status != "deprecated"), matching opencode CLI's exact
filter logic.
The free tier requires a real account API key and throttles third-party
clients by User-Agent:
- With OPENCODE_FREE_API_KEY configured, the key is sent as a Bearer token
and requests identify as "opencode/latest".
- Without a key, the keyless fallback strips the SDK's always-injected empty
Authorization header and still sends the opencode User-Agent.
- The credential resolver no longer blanks OPENCODE_FREE_API_KEY
unconditionally (the stale keyless-tier assumption), and credential-pool
exhaustion no longer surfaces the misleading "Set OPENCODE_FREE_API_KEY"
message.
Co-authored-by: Jean-François <jfm@laposte.net>
Signed-off-by: Rudraksh Chahal <131520192+rudrakshchahal@users.noreply.github.com>
getAllTranscripts @odata.id often uses users('{organizer}')/onlineMeetings('{id}'),
which the slash-only parser missed, so new jobs still stored the transcript id
and hit the refuse-GET guard. Accept the quoted-user form and re-parse the
stored notification on run so replay picks the meeting id.
- Add x-preview-f-free (Ox Alpha: free, 1M context, ZDR) plus all newly
listed Zen models (gpt-5.6 sol/terra/luna, claude-opus-5, gemini-3.7/3.6
flash + lite, grok-4.6/4.5, muse-spark-1.2, kimi-k3, qwen3.7-max,
hy3-free, laguna-s-2.1-free, nemotron-3.5-lightning-free,
muse-spark-1.2-contributor-free) and Go models (gpt-5.6-luna, grok-4.5,
glm-5.3, qwen3.8-max, hy3, hy3-preview, muse-spark-1.2-contributor).
- Drop delisted north-mini-code-free from Zen.
- Route grok-* on Zen and Go through /v1/responses per the published
endpoint tables (grok-4.6/4.5/build-0.1 on Zen, grok-4.5 on Go).
- 1M context fallback for x-preview-f (Ox Alpha).
- Refresh hermes setup provider samples for both providers.
Catalogs verified against live GET /zen/v1/models and /zen/go/v1/models
plus https://opencode.ai/docs/zen/ and /docs/go/ endpoint tables (2026-08-20).
OpenCode Go and Zen serve muse-spark* only on /v1/responses.
Hermes was sending /chat/completions, which returns HTTP 503
with an empty assistant message. Match the published endpoint
table and the existing gpt-* routing.
- Route muse-spark* to codex_responses on opencode-go and opencode-zen
- Add regression assertions next to the gpt-5.6-luna cases
Fresh installs with zero web credentials now rotate web_search/
web_extract across FIVE vendors' public free tiers — Exa, Parallel,
Tavily, Firecrawl, Keenable — instead of a 2-vendor 50/50 split, with
next-in-line ring failover on rate limits (multi-hop until a vendor
serves or the ring is exhausted; served_by marks the actual vendor).
- plugins/web/keenable/: new bundled provider (search via /v1/search,
fetch via /v1/fetch; keyed Bearer or keyless with the mandatory
X-Keenable-Title app header). Credit: integration proposed by
Ilya Gusev (Keenable) in #49758; Free/Paid picker rows included.
- keyless_mcp: tavily/firecrawl/keenable keyless search+extract
wrappers, _KEYLESS_RING + per-process round-robin cursor (seeded by
the random session id, advances per unpinned request), pinned-vendor
entry (pin = start there; rotation off), paid-pinned vendors excluded
from the ring entirely.
- Tavily/Firecrawl providers route keyless traffic through the ring;
both are now default-on ring members (no longer selection-gated).
- web_tools/registry: keenable in backend sets, auto-detect, availability
probes; _keyless_preference() delegates to the ring cursor.
- KEENABLE_API_KEY in OPTIONAL_ENV_VARS; docs updated (ring semantics).
Live E2E: all 10 vendorXcapability paths (5 search + 5 extract) served
real results keyless; rotation cycled all five vendors over 5 dispatch
calls; double-throttle failover walked exa->parallel->tavily.
Salvaged from #78434 by @Slobaka (also the issue reporter; earlier than
the competing #78436). hermes doctor no longer paints a green web check
when the explicitly selected provider cannot initialize — web splits
into per-capability rows (web search / web extract) resolved through
the same registry resolvers the dispatchers use, with readiness from a
true availability probe (_provider_is_ready).
Keyless-tier integration on top of the salvage:
- _provider_is_ready counts is_keyless_available() as ready — keyless
mode is a working state, not a misconfiguration (zero-config installs
and selected-keyless Tavily/Firecrawl show ok, not warn)
- Tavily/Firecrawl gain is_keyless_available() (True only when
explicitly selected — they stay out of the zero-config fallback)
- doctor triggers plugin discovery before reading the registry (fresh
doctor processes saw an empty registry and warned on everything)
E2E: searxng-selected-without-URL warns (the #78412 repro);
zero-config, tavily-keyless, firecrawl-keyless all read ok;
parallel pinned paid without a key warns.
Both polling reconnect paths end on the same 'health pending getUpdates
progress' line, and _record_polling_progress completed silently — so the
log stream for 'reconnected and healthy' was byte-identical to
'reconnected and hung', and a wedged long-poll (#87057 / #69314 /
#71239 class) stayed invisible until a user noticed silence. The only
detection method was sending the bot a test message (#90504).
Emit one INFO on the first confirmed getUpdates round-trip of each
generation, inside the existing event-set branch so steady-state polling
adds no log volume. This turns the pending line into a resolvable pair
('health pending' -> 'confirmed healthy') whose absence after a
reconnect is a reliable hung-poll signature.
Fixes#90504
- Cross-vendor failover: when Exa's or Parallel's keyless free tier
returns a rate-limit-shaped error, the request retries once on the
other vendor's free endpoint (search + whole-batch extract). Result
notes served_by; a peer pinned to its paid tier is never used;
non-throttle errors never fail over.
- Docs: failover note + Tavily/Firecrawl keyless-when-selected rows.
- Firecrawl keyless test expectations aligned with the keyless tier.
Salvaged from #50659 by @LeonSGP43 onto current main (the client
resolver was rewritten for strict-selection semantics since the PR;
reapplied the keyless mode as a third client_mode inside the new
resolver). An explicit firecrawl selection with no FIRECRAWL_API_KEY /
FIRECRAWL_API_URL now routes through a minimal REST client (v2 search +
scrape, no Authorization header) instead of erroring. Unconfigured
installs never route here — the keyless path requires the explicit
selection. Fixes#49912.
- Updated the Tavily API key description to clarify that it is optional and keyless access is supported.
- Modified the Tavily plugin and provider to handle requests with or without an API key, using Bearer authentication when the key is provided.
- Enhanced documentation to reflect the new keyless functionality and updated environment variable descriptions.
- Added tests to ensure correct behavior for both keyed and keyless requests.
Both the wire path and the picker only consulted the catalog's
`mandatory` flag, so a route the Portal lists as accepting no reasoning
parameter at all still got sent a disable, and still offered a Thinking
toggle in the model picker.
For a route it serves, the aggregator's own catalog outranks the
models.dev inference: `supports_reasoning: false` now suppresses the
disable on the wire and drops reasoning controls from the picker
entirely, so there is no disable left to describe.