attachTouchDrag() armed a drag on ANY touch pointerdown and immediately
called preventDefault(), which suppresses the synthesized click
TaskCard.handleClick relies on to call props.onOpen(). There was no
movement threshold, so a finger drifting even ~2-3px on a normal tap --
which is universal on real touch hardware -- was enough to arm the
drag and swallow the open.
Fix: defer starting the drag proxy and calling preventDefault() until
the pointer has actually moved past an 8px threshold (matches the
common native drag-affordance convention). A stationary tap never
crosses the threshold, dragging is never armed, and the click fires
normally. A real drag still claims the gesture identically to before,
just after the same few pixels of travel every touch drag implementation
already tolerates.
The bundle (plugins/kanban/dashboard/dist/index.js) has no build step --
it is hand-maintained directly, as established by prior kanban dashboard
PRs (#114882, #108694) -- so the fix is applied there.
Closes#115568.
Testing: no jsdom/vitest harness exists for this bundle (confirmed by
PR #114882's review follow-up, which explicitly rejected turning a
"live-repro jsdom harness" into a pytest because jsdom/react aren't
declared in the root package.json and the Python CI job has no
node_modules -- such a test would be vacuous in CI). Per that
precedent and the "never read source code in tests" rule (no
regex/substring pin on the bundle text), this PR instead extracts
attachTouchDrag() verbatim at test time via Node (already present:
tests-js/ + vitest are in the repo) and drives it through real
pointerdown/pointermove/pointerup sequences against a minimal DOM
stub -- a behavioral test, not a source-shape test. Proven red on the
unfixed bundle (asserts preventDefault is called on a stationary tap)
and green on the fix; skips cleanly via shutil.which("node") if Node
is unavailable in a given lane.
Verification:
- node tests/plugins/fixtures/kanban_touch_drag_probe.js against the
ORIGINAL (unfixed) bundle: fails with "FAIL: a stationary tap called
preventDefault (suppresses the click)", exit 1 -- confirms the probe
reproduces the reported bug
- Same probe against the fixed bundle: "PASS", exit 0
- scripts/run_tests.sh tests/plugins/test_kanban_dashboard_plugin.py --
42/42 passed (1 new, 41 unchanged)
- node --check plugins/kanban/dashboard/dist/index.js -- syntax OK
With `require_mention: true` + `bots_require_mention: true` + wake words in `mention_patterns`, a message
authored by another Hermes bot that addressed this bot by name matched no dispatch path (the
bot-to-bot loop breaker skips it) and was also refused by the observe gate (`mention_patterns`
matches are assumed dispatched), so it vanished from both paths with no log line.
Factor the loop-breaker predicate into `_bot_sender_suppressed` and consult it in
`_should_observe_unmentioned_group_message`, so every message the dispatcher drops because of
`bots_require_mention` is kept as observed context instead of being lost (#115119).
whatsapp.reply_prefix from config.yaml was written into the bridge env and then
popped again by the WHATSAPP_* passthrough loop (the key was in
_BRIDGE_PASSTHROUGH_ENV and the scoped env lookup came back empty), so bridge.js
always fell back to its built-in header and the documented reply_prefix: ""
could not disable it. Resolve the prefix once (scoped env first, then the
adapter value) and keep it out of the passthrough loop.
Fixes#116059
`_discord_message_admission()` drops a message that mentions someone other than
the bot when `DISCORD_IGNORE_NO_MENTION` is on (the default) and the channel is
not free-response. It did so without consulting `_in_bot_thread()`, unlike the
other two ingress paths — `_dispatch_recovered_message()` (adapter.py:2292) and
`_handle_message()` (adapter.py:5946). Admission runs on both and returns
`False` unconditionally, so it overrode the thread exemption they grant.
The result was an asymmetry with no obvious cause from the outside: in a thread
the bot had joined, a message with no mention at all was admitted (an empty
`message.mentions` skips the enclosing block), while the same message with one
mention of a third party was dropped. A thread the bot is a participant in is
the one place "addressed to someone else" is least likely to hold.
`thread_require_mention` still gates multi-bot threads, since that check lives
inside `_in_bot_thread()`.
The drop also emitted nothing at any log level, leaving `gateway.log` identical
whether the gate fired or the event never arrived; add a debug line so the two
can be told apart.
Fixes#116568
Follow-up to the salvaged #116041: the board's `done` column is now newest-
completed-first and `hermes kanban list --sort completed-desc` exists, so the
feature doc and the CLI usage block say so; the stale "per-column ordering
comes from list_tasks" comment in `get_board` now describes the queue
columns only (wording from #116051).
Co-authored-by: MohamadKanso <91088196+MohamadKanso@users.noreply.github.com>
get_board() buckets one list_tasks() fetch, so the done column
inherited the shared priority DESC, created_at ASC order — creation
order, which says nothing about when work finished. Sort the done
bucket newest-completed-first (completed_at DESC NULLS LAST, id DESC)
and expose that as a completed-desc list_tasks sort key; queue lanes
keep the FIFO dispatch default.
Gate review: TelegramAdapter kept a same-named override with a weaker contract (no finite
guard, negatives clamped instead of reset), so the hierarchy had two parsers under one name.
Both Telegram callers pass explicit bounds; the base method is a strict superset for them.
The cadence test now pins the behaviour (≤ 0.5 s / ≤ 1.0 s) instead of echoing the constant.
Gate review: the fix left three adapter-private copies of `_coerce_float_extra` and the
0.3/2.0/1.0/4.0 cadence literals in three files. The parser and the cadence constants now
live on BasePlatformAdapter beside the delay attrs they configure; WhatsApp and Weixin call
`_configure_text_batch_delays()`, Telegram reads the same constants through its env helper.
The clamp test is parametrized over both adapters and the Weixin docs name the ceilings.
WhatsApp debounced text for 5s (10s near a split) and Weixin for 3s/5s
before dispatching, so every reply paid multiple seconds of idle latency
that Telegram never pays (0.3s/1.0s). Default both adapters to Telegram's
cadence and mirror its ceilings (2.0s / 4.0s, split >= base delay) via the
existing _coerce_float_extra seam. The config keys are unchanged; 0 still
dispatches immediately. Docs updated.
Spotted via #44896 (@liuhao1024). Fixes#44883, refs #25056.
_schedule_invite_join already requires both is_direct and inviter before recording m.direct, so `is_direct and bool(inviter)` at the reconcile call site was redundant; pass is_direct through. The `if is_direct and not inviter` WARNING could only be reached with GATEWAY_ALLOW_ALL_USERS set and a spec-violating stripped m.room.member event lacking `sender`; the info log already prints is_direct, so drop the branch.
The adapter comment and the test module docstring claimed a reconciled pending invite "never fires _on_invite". It does: _absorb_sync runs _dispatch_sync (which emits INVITE to _on_invite) and then the reconcile pass over rooms.invite, which joined every entry unconditionally — so a live invite _on_invite rejected was joined ms later, and invites that arrived while the gateway was down were joined on restart with no gate. Reword both to state that premise.
_on_invite only auto-joins a room when the inviter is allow-listed (or
GATEWAY_ALLOW_ALL_USERS is set), so a live invite from an arbitrary
federated user is rejected. A pending invite that arrives while the
gateway is down takes a different path: _schedule_pending_invite_joins
reconciles it from rooms.invite in the sync response and scheduled the
join unconditionally. An unauthorized invite sent during downtime was
therefore auto-joined on restart, bypassing the allowlist.
Extract the gate from _on_invite into _is_authorized_inviter and apply
it during reconciliation too, reading the inviter from the stripped
invite state (the sender of the m.room.member event for our own user,
as _extract_invite_dm_signal already does for the DM signal). An
inviter that cannot be read from the invite state fails closed, exactly
like an empty sender in _on_invite: the invite is skipped with a
warning and left pending.
A direct invite that arrives while the gateway is running fires
_on_invite, which passes is_direct and the inviter through
_schedule_invite_join so the room is recorded in m.direct after the
join. An invite that is still pending across a gateway restart takes a
different path: _schedule_pending_invite_joins reconciles it from
rooms.invite in the sync response, but called _schedule_invite_join
without is_direct or inviter. The DM signal was dropped, the room was
never recorded in m.direct, and it was classified as a group until the
user's own client happened to update m.direct.
Read the signal from the stripped invite state instead: the
m.room.member event for our own user carries the original invite's
is_direct flag, and its sender is the inviter. Thread both through to
_schedule_invite_join so a reconciled direct invite is recorded in
m.direct, and thus lands in _dm_rooms, exactly like a live one.
This gap was surfaced by the triage of #62493.
Reusing `_MEDIA_SEND_READ_TIMEOUT` (60 s) as the whole-call deadline turned
httpx's per-phase stall budget into a bandwidth cap: a 20 MB video on a
~2 Mbit/s uplink (~80 s) that succeeds today would fail. `_MEDIA_SEND_DEADLINE`
= 300 s covers the 50 MB Bot API cap at ~2 Mbit/s plus connect and sendVideo
transcoding, and is >2x the summed httpx budgets (pool 8 + connect 10 +
media_write 60 + read 60 = 138 s), so it only fires on a socket that has
stopped raising. Abandon-inside-lock semantics documented at the constant.
No Bot-API write in the adapter was under a wall-clock cap — text sends,
edits, drafts and media uploads relied solely on httpx socket timeouts,
which do not fire when a shielded httpcore socket wedges (same class as the
getUpdates hang in #92991). A stuck send then pinned `_chat_send_lock` and
the loop. Wrap every send/edit/draft in `_await_with_thread_deadline`
(`_TEXT_SEND_DEADLINE`) and both media paths in
`_send_with_dm_topic_reply_anchor_retry`; the helper gains a `label` and a
descriptive TimeoutError message.
Half A of #115280; Half B superseded by #116134 / contradicts #75017.
_fetch_discovery followed redirects but only pinned the document's
self-asserted issuer field, so one cleartext or attacker-hosted hop
could serve a forged document claiming the configured issuer with
attacker jwks_uri and token_endpoint. Verify then accepted
attacker-signed ID tokens and the code exchange POSTed the client
secret to the attacker's token endpoint.
The resolved response.url must now share the configured issuer's
origin (scheme, host, port with default-port normalisation) before
the body is parsed. Same-origin canonicalisation redirects still
pass, and the issuer-field pin remains as the misconfig check it is.
First in-tree consumer of ProviderProfile.fetch_account_usage: the OpenCode Go plan's rolling /
weekly / monthly windows (GET /zen/go/v1/usage) render in every /usage surface without adding the
provider name to the core _USAGE_FETCHERS table. Ported from #113418, which implemented the same
fetch as a core table entry; the literal endpoint (not the runtime base_url, which loses /v1 in
anthropic_messages mode) and the window mapping are theirs.
Co-authored-by: Angello Picasso <angello.picasso@devsu.com>
video_generate advertised an optional `model` argument (and the xAI edit/extend
tools a model override) so the LLM could route a single call to a different
model family — a different endpoint and billing tier — than the one the user
selected in `hermes tools`. image_generate never exposed this, and #83080 asked
to extend it there; the ruling is the opposite: models do not choose models.
The `model` property is gone from the static and dynamic video_generate schema
and from xai_video_edit / xai_video_extend; a `model` smuggled into the call is
ignored and the configured `video_gen.model` (then the provider default) is what
reaches the request. Config-side selection (`video_gen.model`,
`video_gen.<provider>.model`, `<PROVIDER>_VIDEO_MODEL`) is unchanged, and the
xAI plugin's explicit-model branch is no longer reachable from the tool layer.
Refs #83080
The cold-boot queue policy was computed inline at both start paths, and the
default (drop) was silent — a command sent while the gateway was offline
vanished with no log line (the second half of #71811).
One helper now owns the decision and logs it on every cold boot, so an
operator can tell which policy applied. Behaviour is unchanged.
The seventeen remaining literals are container-side paths, AF_UNIX socket-path-limit
candidates on darwin, detection needles, guard regexes and guidance text that tells the
model to avoid /tmp. Each carries an inline `no-tmp: ok — <why>` so the reason lives
next to the line; the baseline keeps only a fenced tree listing where a marker would render.
Hermes now routes scratch space through HERMES_HOME/cache/scratch (exported as
TMPDIR), so every production path that still spelled out /tmp bypassed that and
kept teaching the agent the habit. Fallbacks in tool_result_storage,
code_execution_tool, process_registry, the ACP child HOME, mini_swe_runner's
local cwd, and the CI/profiling scripts now use tempfile.gettempdir(); shell
installers fall back to $TMPDIR (then HERMES_HOME) when mktemp is missing, and
repro/eval shells use `mktemp -d -t`. User-facing help text and sample payloads
(hermes send, approvals test, hooks test, voice-mode WSL hints, meet_bot debug
line) no longer suggest /tmp.
Container-side paths (mini_swe_runner docker cwd, sandbox base env, remote
sync tarballs) keep the literal because they name the sandbox filesystem,
not the host.
What: the `initOnSessionStart` setup-schema description and the Honcho docs page
now state that in `recallMode: tools` the eager init runs synchronously during
agent construction, that it should stay false for Desktop or a local Honcho that
may be down, and that `timeout` in honcho.json caps each SDK call (both keys
added to the Full Config Reference table).
Why: #51492 reports Desktop `session.resume` / `prompt.submit` timeouts when a
local Honcho is down with `initOnSessionStart: true`. The synchronous eager path
is intentional design (#51562 was closed for that reason:
"ready-before-first-tool-call semantics"), so the fix for users is knowing the
trade-off and the two knobs that bound it.
_pcm_tail_loop caught bare Exception, so any bug in the tail thread died
silently while the bot stayed mute. Only the expected closed-pipe errors
(OSError incl. BrokenPipeError, ValueError on a closed stdin) are quiet now;
anything else surfaces through the thread's excepthook.
Two invariant tests through the real plugin helpers: a fake page whose Join button
renders after a delay must still be clicked and a muted toggle unmuted (control: a live
mic is reported, never toggled); a `cat` stand-in for paplay must receive bytes appended
to speaker.pcm after the pump started and the tail thread must be gone after teardown.
Both fail on the previous meet_bot.py.
Docs: realtime mode is speak-only - incoming speech is Meet caption scraping, meeting
audio is never sent to the Realtime session (the docs page wrongly said the meeting audio
was transcribed through the configured STT provider); micState is documented.
Fixes#80875
The Linux/macOS pump (paplay / ffmpeg) was started against speaker.pcm while the
file was still empty, read it to EOF and exited before OpenAI Realtime appended any
audio - audioBytesOut grew while nobody heard the bot. The pump now reads raw PCM
from stdin, fed by a tail thread that follows the file as it grows; teardown joins it.
Meet can also seat an authenticated bot muted. On admission the bot now checks the
in-call mic toggle, clicks it when the label says "Turn on microphone", and reports
the outcome as micState in status.json (unmuted_clicked / unmuted / unknown).
Slim redo of #81488 onto the refactored _start_pcm_pump / _drain_loop helpers.
Part of #80875 (atoms 2 and 3).
Meet renders "Join now" / "Ask to join" asynchronously after the page reaches
domcontentloaded; the one-shot _join() check could run before the button existed
and the bot sat silently in the pre-join screen. Poll for up to 30 s (0.5 s cadence),
returning as soon as a button was clicked.
Slim port of #37945 onto the refactored _join() helper.
Part of #80875 (atom 1).
Selecting Image Generation -> OpenAI (Codex auth) in `hermes setup` / `hermes tools` on a
fresh install saved `image_gen.provider=openai-codex` and printed "no configuration
needed!" without ever signing in, so the backend was unusable until the user guessed the
auth command — and the schema hint pointed at `hermes auth codex`, which does not exist.
Root cause: the row declares `env_vars: []` and only a `post_setup_hint`, a key nothing
consumes; `_configure_provider` runs a hook only for `post_setup`.
- plugins/image_gen/openai-codex: schema declares `post_setup: "openai_codex"`; the hint
and the `auth_required` error name `hermes auth add openai-codex`.
- hermes_cli/tools_config_post_setup: `_post_setup_openai_codex` in the existing
`_POST_SETUP_HOOKS` table (sibling of the `xai_grok` credential bootstrap). With
credentials present it continues; otherwise it offers the device-code sign-in (or skip)
and saves tokens with `set_active=False`, so picking an image backend never rewrites
`model.provider` the way the model-provider login does.
- `_POST_SETUP_AUTH_READY` table replaces the `post_setup == "xai_grok"` special case in
`provider_readiness_status`, so any credential-bootstrap row reports ready/needs_auth
from the auth store.
- `_save_codex_tokens(set_active=...)` mirrors `_save_xai_oauth_tokens`.
Live probe (real `_configure_provider`, real plugin row, temp HERMES_HOME, OAuth start
stubbed with a recorder): before — post_setup=None, OAuth fired [], hint `hermes auth codex`;
after — no creds: OAuth fired once, logged in, model.provider untouched; with creds: OAuth
not fired.
Fixes#102144
Salvages #102165 (@liuhao1024) — superseded: same direction (post_setup hook), redone on
the split tools_config siblings without calling `_login_openai_codex`, which would have
switched the main model provider.
Custom model ids (#97928): a value of image_gen.openai.model / OPENAI_IMAGE_MODEL outside the
gpt-image-2 / gpt-image-2.5 catalog used to fall through silently to the default tier, so a
gateway serving its own image model names always received `model: gpt-image-2` +
`quality: medium`. resolve_static_model(passthrough=True) now returns the unknown id as the
API model with quality=None, and the request omits `quality` for it — OpenAI-compatible
gateways reject enum values they do not know, and a foreign model has no OpenAI quality
tiers. Only the provider-scoped key and env var pass through; the shared top-level
image_gen.model can hold another provider's id (a FAL path) and is never forwarded.
Named custom endpoint (#83080, config-reuse half): image_gen.openai.provider names a
providers:/custom_providers: entry; its base_url and api_key/key_env fill in whatever
image_gen.openai.base_url / key_env leave unset, so a gateway already declared for chat is
not re-declared with a duplicated key. Explicit image_gen.openai values keep precedence; an
unknown name logs a warning and falls back to the OpenAI variables. The lookup goes through
hermes_cli.runtime_provider._get_named_custom_provider, the same resolver the auxiliary
clients use, so aliases and legacy list entries behave identically.
image_gen is an open config root (deliberately absent from DEFAULT_CONFIG), so
`image_gen.openai.provider` already validates as a known key; docs cover both behaviours.
Live probe (fake /v1/images/generations gateway recording the body, real provider on a temp
HERMES_HOME): before — `custom-image-model` sent as model=gpt-image-2 quality=medium; named
provider → is_available False / auth_required. After — model=custom-image-model, no quality;
named provider → request to the entry's URL with `Bearer <its key_env>`; the catalog model
still maps to gpt-image-2 + quality=medium.
What: plugins/image_gen/openai resolves its endpoint and credential through one
resolver — image_gen.openai.base_url → OPENAI_BASE_URL → SDK default, and the env
var named by image_gen.openai.key_env → OPENAI_API_KEY — shared by is_available()
and generate() so the two cannot disagree. The client is built on
build_keepalive_http_client (env-only proxy policy) and sends a blank
OpenAI-Project header.
Why: the image endpoint could only be routed via the process-wide OPENAI_BASE_URL /
OPENAI_API_KEY, so a local or third-party image gateway could not be configured
independently of the chat provider (#65309, #97928, #13798). openai.OpenAI() with
trust_env routed localhost endpoints through a macOS system proxy whose
ExceptionsList httpx never sees (#64888). An OPENAI_PROJECT_ID set for chat made
/v1/images/generations 403 model_not_found on projects with a model allow-list
even though the key already carries the project (#60748).
Slim redo of the contributor direction in #18796 (@y0shua1ee), #37208/#37209
(@charzhou), #65312/#65323/#64893 (@asdlem), #60749 (@perelin); all predate the
StaticImageGenProvider refactor and no longer apply.
Groq's OpenAI-compatible wire accepts top-level reasoning_effort only as
"none" or "default" (#75089, Defect 1); the custom profile forwarded the
configured graded level ("medium"/"high") and the request still 400'd on
both the main transport and the auxiliary path. The clamp lives in
CustomProfile.build_api_kwargs_extras, keyed on the resolved base_url host,
so both paths share it. The aux bare-key case now drives call_llm with a
fake client instead of the private _build_call_kwargs.
`SessionManager._make_agent` and the Feishu doc-comment agent built their
`AIAgent` without `reasoning_config`, so `agent.reasoning_effort: none`
never reached those sessions: the transport applied its default effort,
which non-reasoning models such as gpt-4o-mini reject with HTTP 400 and
which silently re-enables thinking everywhere else. Both surfaces now go
through `hermes_constants.resolve_reasoning_config`, the same chokepoint
the CLI, gateway, TUI, cron and `hermes -p` already use, resolved against
the model the session actually runs so per-model overrides apply.
Ported from PR #85164 by @Chinmayrawat15 (oneshot hunk already on main).
Fixes#85153
Mirrors the _discord_require_mention()/_discord_max_attachment_bytes() shape, makes the
lookup lazy (it only runs on the free-channel path now), and keeps the in-code key
manifests in _handle_message and register() listing the new key.
Free-response channels skip auto-threading by default so the bot replies
inline (lightweight chat mode). This prevented users who wanted BOTH
mention-free replies AND per-conversation threads from getting either.
Add a new opt-in `discord.free_response_auto_thread` (env:
`DISCORD_FREE_RESPONSE_AUTO_THREAD`, default false) that, when true,
re-enables auto-threading in free-response channels. Voice-linked
channels continue to skip auto-thread regardless, and the flag is
gated behind the global `DISCORD_AUTO_THREAD=true`.
Default behavior is unchanged; all 291 existing discord tests pass.
The warm-up imported only the provider module. holographic / mnemosyne import
numpy at module top, but hindsight defers the ML stack to is_available() ->
_check_local_runtime() (importlib of hindsight / sentence_transformers), which
ran later on a to_thread worker racing acp-mcp-discovery — the reported hindsight
stack was still reachable. The deadlock partner is numpy's lazy _core init in
every reporter's dump, and a plain `import numpy` up front was every reporter's
workaround, so import_memory_provider_module now also imports numpy (best-effort)
once the provider module is in.
Also: import_memory_provider_module() defaults to the configured memory.provider,
so entry.py drops its duplicate config resolver and outer try; the "ONLY thread"
comment is reworded — hermes_cli's plugin-discovery thread is already running
when hermes acp dispatches.
Every faulthandler dump in the thread shows session/new stuck in numpy's
create_module on the main thread while another thread (MCP discovery / ACP
stdin reader) sits in the same lazy import chain — a first-time native
extension import racing another thread deadlocks on Windows (holographic,
mnemosyne and hindsight all reproduce it; a sitecustomize `import numpy`
before any thread exists resolves it every time).
hermes acp now imports the configured memory.provider's module on the main
thread before the MCP-discovery thread and asyncio.run() start (Windows only —
the deadlock is Windows-specific and the import is paid once either way).
plugins.memory.import_memory_provider_module imports the module without
constructing a provider or running register(); the agent build later finds it
in sys.modules.
Trimmed from #91775 (@tigercraft4): same placement and gating; reuses the
existing plugin loader instead of a second module-import routine.
Co-authored-by: tigercraft4 <tigercraft4@tigercraft4.com>
Each test is red on origin/main and green on the fix; each also asserts the
control case (evicted LSP file re-opens with diagnostics, every log line
written once, unexpired fuzzy root kept, every hindsight turn shipped once).
Trims the hindsight comment to the why.
_session_turns accumulated every turn's text for the whole session, reset
only at session boundaries, so a never-ending session grew without bound
independent of context compaction.
In append mode each retain ships only the delta since the last watermark
(_session_turns[_last_retained_turn_count:]), and a retained turn is never
read again — sync_turn always slices from the watermark and flush-on-switch
flushes what's left. So after an append retain the buffer drops the retained
prefix (clear() + reset the watermark), bounding it to the un-retained tail.
Overwrite mode is deliberately untouched: legacy/overwrite APIs resend the
whole session each retain because each retain replaces the document, so that
path must keep every turn.
Tests: append trims the retained prefix while shipping every turn exactly
once (no loss, no duplication), stays bounded across 100+ turns, and
overwrite mode keeps the full buffer.
(cherry picked from commit cb77e006754dd4ad5c880f22310d3ac7b3472739)
Trim of the salvaged openai-native marker provider (#107377):
- `hermes auth --provider openai-codex` does not exist; the login command is
`hermes auth add openai-codex` (hermes_cli/subcommands/auth.py). Fixed in
plugin.yaml, provider.py docstring and the web-search docs (2 spots).
- `hermes tools` picker tag was Chinese; now English like every other provider.
- tests/plugins/web/test_web_search_provider_plugins.py enumerates the bundled
registry exactly, so the new plugin made it fail; it now lists openai-native
with capability flags (search=True, extract=False) — the search-only
invariant that keeps web_extract on its own backend.
- contributors/emails mapping for the salvaged author.
Declares OpenAI's provider-executed Responses `web_search` built-in in place of
the client-side `web_search` function, mirroring the existing xAI native-search
path. Selected via `web.search_backend: openai-native`; search-only, so
`web_extract` keeps resolving to its own backend.
The Responses adapter already recognises built-in tool types
(`_RESPONSES_BUILTIN_TOOL_TYPES`) and preflight passes them through, so the only
missing piece was the swap itself plus a provider name for the config to point at.
Gating is deliberate: two-sided (Codex backend AND a selected openai-native
backend) and fail-closed, so a custom OpenAI-compatible endpoint or an
unresolved provider leaves the client tool untouched.
Every adapter-side lane (text/photo/album batch dicts, _active_sessions, the
busy guard, /stop /new /reset and clarify replies) derived its session key
before the receiving bot's identity was known, so two bots seeing the same
Telegram chat.id == user.id collided on one agent:main lane and a route to an
unserved profile still reached the default lane's running task.
BasePlatformAdapter._canonicalize pins the RoutingIdentity (via the new
session_identity.canonical_identity seam) as the FIRST statement of
handle_message, _enqueue_text_event, _handle_message_while_active, Telegram
_route_photo_event and every _source_session_key; _drop_unresolved drops a
rejected route at the first seam with one WARNING. Topic recovery copies the
source through replace_source so the identity travels; Slack thread sources go
through build_source so the thread key carries the same provenance.
A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.
Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.
Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.
Phase 2 of #88715.
NAS maps the agent-cron bearer to a provisioned instance via an `agent:*` client or the
`hermes-cli-vps` bootstrap session (hermes-portal `server/agent-cron/instance-auth.ts`). A
container whose auth.json holds a plain `hermes-cli` user login is refused with 403
invalid_client on every arm, re-arm and list, for the life of that credential - and the
re-login users try first (`hermes auth logout nous` + device code) replaces the bootstrap
session, making it permanent. Chronos previously logged one bare warning per job and left
the jobs with no trigger at all: they only ran through the misfire sweep, minutes late.
NasCronClientError now carries the HTTP status and the OAuth `error` code. On an identity
rejection the provider logs ONE warning that names the real remedy (restore the hosted
credential from the Nous Portal; re-login cannot fix it), stops calling NAS, and starts the
built-in ticker with the gateway's own adapters/loop so scheduled jobs keep firing on time.
Transient 5xx/transport failures keep retrying on the next reconcile.
Supersedes #97566 (@wesleysimplicio), whose runtime-credential swap resolves to the same
bearer on main (`agent_key` is the access token) and so could not change the 403.
OpenRouter (_generate_via_chat, _generate_via_image_api) and OpenAI gpt-image
called record_token_usage only after image extraction and save succeeded. A
token-billed 200 that returned text but no image (the "returned no image"
fallback case), an empty `data` list, or a failed save therefore consumed
billed tokens that never reached session_model_usage; on the fallback chain
only the fallback model's tokens were recorded.
Move the call to immediately after a successful post_json / SDK call, before
extraction/save, in all three paths. One call per HTTP success, so no double
count; failure responses (non-2xx, timeouts) still record nothing because
there is no body to read usage from.
Tests: the parametrized invariant in test_openrouter_compat_provider gains the
no-image-with-usage case for both surfaces (result is `empty_response`, one row
still recorded); the OpenAI test is parametrized on has_image the same way.
Widen the salvaged Image API hunk (#114340) to the whole class:
- plugins/image_gen/_common.py: record_token_usage() — one helper feeding the
aux accounting chokepoint (agent.aux_accounting.record_aux_usage) with task
"image_generation", the billing provider and the priced model id. Dict and
SDK usage objects alike; a body without tokens is a no-op, as is a call
outside a turn.
- plugins/image_gen/openrouter: the /chat/completions path (the DEFAULT model
chain — openai/gpt-5.4-image-2, google/gemini-3-pro-image — is chat-only and
token-billed) now records too; the Image API path uses the shared helper
(the contributor's local _record_image_api_usage is folded into it) and both
pass base_url so pricing resolves the route. Task renamed image_gen ->
image_generation to match the other aux task names.
- plugins/image_gen/openai: gpt-image bills per text/image token; record the
Images API usage block under the API model (gpt-image-2), not the Hermes
quality-tier label.
- FAL, xAI, Krea, DeepInfra, Meta, openai-codex return no token usage and stay
unrecorded (nothing to bill per token).
- tests: the contributor's two tests folded into one invariant parametrized
over chat / Image API / no-usage control; one OpenAI invariant.
- docs: image-generation.md "How It Works Internally" gains the accounting step.
Fixes#114324