Follow-up to the streaming reasoning commit for #99552 covering the atoms it left open:
- Non-streaming `/v1/chat/completions` returns the turn's reasoning as
`choices[0].message.reasoning_content`; non-streaming `/v1/responses` emits a
completed `reasoning` output item ahead of the message (and ahead of that step's
`function_call` items), so `GET /v1/responses/{id}` replays it too. The text is read
from the assistant messages the agent already persisted (`build_assistant_message`
stores it under `reasoning`) instead of re-accumulating `reasoning_callback` deltas:
in non-stream mode the agent fires the callback per provider delta AND once more with
the full text as the post-response fallback, so accumulating would double it.
- The Responses input parser skips `{type: "reasoning"}` items in `input` and
`conversation_history`. Responses SDK clients replay a prior response's `output`
list as the next `input`; before, the item became an empty `user` message (in
`input`) or a 400 (in `conversation_history`).
- The streaming `response.completed` envelope carries the full reasoning item
(`id`, `status: completed`) rather than a `{type, summary}` stub, matching the
`output_item.done` item and the non-streaming shape.
Live (real route handlers, real AIAgent, fake OpenAI SSE model replaying
`delta.reasoning_content`): before — non-stream chat had no `reasoning_content`,
non-stream responses output was `[message]`, a chained turn echoing the output list
stored an empty user turn and the model saw a repaired placeholder; after — chat
`reasoning_content` equals the streamed text exactly, responses/GET output is
`[reasoning, message]`, the chained turn's model input is `[user, assistant, user]`.
OpenAI-compatible clients (Open WebUI, opencode, LibreChat, the Vercel AI SDK)
saw only answer text from the API server: the streaming writers never wired the
agent's structured ``reasoning_callback``, so reasoning deltas that every native
surface already renders were dropped at the transport boundary (#99552).
- `_spawn_stream_agent` passes a `reasoning_callback` (via `_run_agent` /
`_create_agent`) that tags deltas `("__reasoning__", text)` on the stream queue,
keeping them distinct from answer text. The lossy 500-char `reasoning.available`
progress preview is deliberately not used.
- `/v1/chat/completions`: reasoning rides `choices[0].delta.reasoning_content`
(the DeepSeek-style field those clients render as a thinking block).
- `/v1/responses`: each thinking burst is a spec-native `reasoning` output item
(`output_item.added`, `reasoning_summary_part.added`,
`reasoning_summary_text.delta/done`, `reasoning_summary_part.done`,
`output_item.done`), closed before the next message/function_call item opens
and echoed in `response.completed` output; `sequence_number` stays monotonic.
- `GET /v1/capabilities` advertises `features.reasoning_streaming: true`.
- Docs: api-server page documents the wire fields and the capability flag.
Gating is unchanged: nothing is emitted unless the model produces reasoning under
the resolved `reasoning_config` (`model_options.reasoning.enabled: false` opts out).
`_profile_name_for_source` already accepts `adapter_profile=None`, so the
branch in `_stamp_routed_profile` was two spellings of the same call.
Collapse it and give the two test doubles that replace the method the
same keyword so they keep matching the production signature.
Salvage follow-up to #106019.
A matcher exception fell through to the default profile, so a transient
failure while resolving a sender route silently served the message from the
default profile's runtime. Raise the existing rejection instead, which the
ingress gate already drops on.
Only the failure path changes: a source that simply matches no route keeps
falling through to the default/active profile as before.
Voice input reused the bound text channel's cached source and replaced only
`user_id`, so a second speaker inherited the profile resolved for the first.
The ingress gate does not re-resolve a source that already carries a profile.
Re-resolve at the voice call site instead of clearing the profile: clearing
would drop the receiving bot and could re-home voice arriving on a secondary
profile's own bot. `_voice_input_source` reattaches the transport provenance
`from_dict` discards, and `_stamp_routed_profile` takes the receiving bot's
profile as the fallback when no route matches.
Kanban re-subscription could not repair a row created before sender capture:
`user_id` was only written by the INSERT, so a legacy row stayed senderless and
the notifier's conservative fallback left it undeliverable for good. It now
self-heals like `user_id_alt`.
An explicit `user_id: null` or empty string is rejected instead of widening the
route to every sender, and `to_dict` omits the field when unset so a round-trip
cannot reintroduce it. Numeric `0` from an adapter normalizes to "0" rather than
being dropped, without changing the shared coercion used by the other fields.
Closes#33548. `gateway.profile_routes` could only discriminate on where a
message came from (guild/channel/thread), so giving two people in one shared
chat their own isolated profile meant running two bots. Add `user_id` as a
route discriminator, conjunctive with the existing location fields and
matched on exact equality.
Kanban notifications revalidate a subscription's route before delivering, so
they pass the persisted sender too; a legacy subscription with no sender
identity falls back to the route's own `user_id` rather than skipping a route
that could have won, keeping the notify path fail-closed.
Cron is deliberately left out: it has no authenticated inbound sender, so a
`user_id` route never qualifies a cron delivery target. Operators need a
location-only route for that, which the docs now state.
A secondary-owned Yuanbao bot keyed its per-group dispatch queue and RecallGuard
entries with the free `build_session_key(source)` — no profile, so `agent:main:` —
while `handle_message` popped under `agent:<owner>:`. Two derivations of one
identity: the group queue was shared across bots and the RecallGuard entries
leaked. Weixin, Telegram's photo batch, Slack's thread key and Raft's wake key
each carried their own copy of the call as well.
Every adapter-side key now comes from `BasePlatformAdapter._source_session_key`
/ `_event_session_key` (owner namespace, runner-seeded isolation flags, and —
after the RoutingIdentity PR — the pinned identity). Weixin's `_text_batch_key`
override is deleted (the base does the same). Slack's thread key reads the
isolation flags from the adapter config the runner seeds, not the store's.
Lint: pattern P32 in `scripts/ci/profile_scope_patterns.json` flags
`build_session_key(` / `SessionSource(` under `gateway/platforms/**` and
`plugins/platforms/**` except `platforms/base.py`; the checker gains an optional
`path_regex` per pattern. Advisory, like every other pattern.
Phase 2 of #88715.
A multiplexed gateway answered "which bot received this / who may admit it /
where does it run" in three places (`_transport_owner`,
`_authorization_home_for_source`, `_resolve_profile_home_for_source` +
`_session_key_profile`) that agreed only because they read the same fallback
chain. `gateway/session_identity.py` answers them once: `resolve_identity()`
folds `_admit_primary_source` + `_stamp_routed_profile` + the transport-owner
lookup and pins a frozen `RoutingIdentity` (transport_profile, runtime_profile,
authorization_home, runtime_home, weak transport ref) on the source as a
wire-invisible attribute, like `_transport_adapter_ref`. Under multiplexing a
route to an unserved profile raises `IdentityUnresolved` instead of a
`None`-means-default return; `"default"` is spelled out inside the object.
Additive: the existing helpers become thin readers of the identity when it is
present and keep their fallback chain when it is not, `source.profile` stays
the serialized runtime profile (None on the wire ⇔ default) and every
historical `agent:main` key is byte-identical. `replace_source()` copies a
source without losing its provenance (run_topics used to hand-copy the
transport ref).
Phase 1 of #88715; the gateway rows of #90142 / #93943.
A turn's hard interrupt fans out only to `_active_children`; background
delegate_task units are detached from the parent at dispatch
(`_dispatch_background` / honor_parent_interrupt=False), so /stop left them
running to completion and their result arrived minutes later as a parked
wake. Every stop surface now also calls
`tools.async_delegation.interrupt_for_session` for the session's units:
- gateway: `_interrupt_and_clear_session` (busy /stop, /new share it — /new
already did this in `_handle_reset_command`, the second request is
idempotent) and the idle `_handle_stop_command` tail, which replied
"No active task to stop." while a background child was running; it now
stops them and replies "Stopped".
- tui_gateway `session.interrupt` (Desktop Stop / TUI stop) — own UI sid +
spawner id only, so a viewer tab never kills gateway work.
- acp_adapter `cancel`.
- CLI `/stop` already used `interrupt_all` (process-wide); unchanged.
The stop recurses: depth>0 delegations are always synchronous
(`_model_background_value`), so the child's hard interrupt reaches its
workers through its own `_active_children` fan-out, and each level's
interrupted partial result rolls up as that child's completion.
An interrupted child's entry now carries what it actually had: the loop's
`final_response` is the "Operation interrupted." placeholder (also appended
as the closing assistant row), so `_build_result_entry` takes the child's
last real assistant text as `summary` and keeps the placeholder as `error`.
The unit finalizes normally and re-enters at once as its completion notice
(status=interrupted, "Partial output: ...", "Subagent Task Interrupted" on
TUI/Desktop) instead of the chat waiting for the child's budget to run out.
Docs: delegate_task description, tools/AGENTS.md, delegation.md,
gateway-session-lifecycle.md.
Part of #114456
Gate review: tcp_site.py grew its own `_WILDCARD_HOSTS` that already diverged
from shared_ingress.py (dropped "*" and ""). One `is_wildcard_host` in
shared_ingress, used by both the TIME_WAIT rebind guard and
`listener_base_url`.
On macOS, an exclusive bind (SO_REUSEADDR disabled) refuses a port that is
still held only by a server-side TIME_WAIT socket for 2*MSL (~30s) after a
previous connection close — the gateway's own shutdown, or any
``Connection: close`` request. A ``/restart`` issued within that window
failed to bind, even though nobody was actually listening.
For an explicit host, a refused connect probe proves nobody is listening:
the kernel still rejects an exact duplicate bind even with SO_REUSEADDR, and
a foreign wildcard listener would answer the probe. So one retry with
SO_REUSEADDR is safe in that case. A wildcard host keeps the strict
exclusive path, since a foreign listener on a non-loopback interface can't
be probed this way. A live listener still wins either way.
The bind, probe, and retry logic is shared by webhook.py and api_server.py
through the new gateway/platforms/tcp_site.py, since both adapters had the
same exposure — api_server additionally marked a spurious TIME_WAIT bind
failure as a non-retryable port conflict.
(cherry picked from commit 289fd06148c447829f1e38b66ff78b9e4bebdeb4)
_send_file hit the identical stale -2 after its own tokenless re-send and still
raised the generic "sendmessage error"; it now raises the shared
_session_not_ready_error. The genuine rate-limit cooldown message carries
ret/errcode/errmsg so operators can tell a real -2 frequency limit from the
session case. One parametrised test (stored token / no token) asserts the send
fails once, classify_send_error does not call it rate_limited, and the breaker
stays closed; troubleshooting row in the Weixin docs.
iLink answers a bot-initiated send with ret=-2 errmsg="prepare failed" when the
peer's session is not ready (no inbound message yet, or the context token is
gone). That is deterministic: treating it as a frequency limit spent the retry
budget on an error that never clears and opened the 30 s rate-limit breaker for
the whole account, blocking every later send (#80125, #112709).
After the tokenless re-send (46ab37ce35) has been tried — or when there is no
token to drop — a stale-session -2 now fails fast with a descriptive
"session not ready" error and never reaches the rate-limit path. The text
avoids the "rate limit" substring classify_send_error keys on.
Salvage of #80152 (fail fast + keep the breaker closed); #80156 arrived five
minutes later with the same change.
Co-authored-by: RelaxJonh <92573950+RelaxJonh@users.noreply.github.com>
The bridge now gates groups by WHATSAPP_GROUP_POLICY / WHATSAPP_GROUP_ALLOWED_USERS
(#73465), but only the DM policy and DM allowlist were exported from the adapter's
resolved config — a YAML group_allow_from never reached the bridge and, under a
multiplexed gateway, the launch process's WHATSAPP_GROUP_* values did. _bridge_env
exports both like it already did for the DM pair, so unlisted-group traffic is cut
before media download instead of after the POST to Python.
whatsapp_common._select_allowlist generalises the DM precedence reader (config key
presence wins, an explicit empty list stays authoritative, then the first truthy env
carrier); the adapter's group list and the Cloud sibling's group list use it instead
of hand-rolled `or` chains, so `group_allow_from: []` means the same everywhere.
Tests: bridge env carries group policy + group allowlist from the secondary profile's
YAML (scoped precedence file); env fallback test trimmed to the two invariants (red on
the merge base); bare test adapters seed _group_policy/_group_allow_from, which
_bridge_env now reads.
Gate review: `degraded` became a whole-life serving state but four
sibling predicates only accepted `running`: derive_gateway_busy /
derive_gateway_drainable (the NAS drain gate reported a degraded gateway
with in-flight turns as idle and undrainable), the stale-heartbeat
detectors in `hermes gateway status` and the dashboard, and the Windows
doctor probe. All four accept `degraded`; a dead watchdog-stamped
`degraded` is still excluded by `gateway_running=False`.
The parked-platform ERROR pointed at `/platform resume`, which only
resumes platforms in the retry queue - a non-retryable failure never
enters it. The remedy is `hermes gateway restart`.
Test: a degraded gateway with active agents is busy; not live => not
drainable.
Gate review on the stack:
- `hermes gateway restart` under systemd only accepted `gateway_state == "running"`
as proof of the replacement; a boot with a parked platform stamps `degraded` for
its whole life, so every restart/update on such a host waited out the 60 s+
budget and reported a false failure while the gateway was serving. The verifier
now accepts `degraded` as restarted and prints one DEGRADED warning.
- Drain release and scale-to-zero wake re-stamped `running` unconditionally, wiping
the parked-platform signal after the first `.drain_request.json` cycle. Every
"we are serving" stamp now goes through `_serving_state()`; the mixed fatal +
retryable boot path sets the flag too (it fell through as a plain run before).
`_startup_parked_platforms` is a bool — the joined error text was only ever logged.
- `_wait_for_tcp_port_free`: an unresolvable or unreachable configured host raised
a non-refused OSError on every probe and burned the full 10 s wait; only a
connect timeout means "listener alive", anything else means nothing to wait for.
- Windows `restart()` replaces its blind `time.sleep(1.0)` "let Windows release the
port" with the same configured-address wait.
- The api_server bind retry rebuilds the AppRunner per EADDRINUSE attempt instead of
calling aiohttp's private `_unreg_site`; attempt count is a named constant.
Test: a `degraded` replacement is reported as restarted (red on the previous head).
E2E re-run: predecessor releases the port 0.5 s after the first bind → bound on the
third attempt (+0.61 s), connect() True.
The shared startup gate returned early whenever any sibling platform connected,
so a non-retryable fatal platform failure left one WARNING and a 'running'
runtime state: the gateway served messaging while every API client was gone,
with nothing retried and no observable trace.
Park the failures, log at ERROR and stamp gateway_state 'degraded'. Retryable
peers are untouched (the reconnect watcher heals them). Behaviour and exit code
unchanged; only logging and health reporting. Port-race half of the report is
left to the open PRs #92060/#96419.
(cherry picked from commit edfb98f86911b70aa247ab0f718f0218002e154e)
Folds on the salvaged #92060 (@eliasburlison):
- `_wait_for_api_server_port_free` read only API_SERVER_HOST/PORT from the
environment, so an api_server port set in config.yaml (`platforms.api_server.port`)
was never waited on. The adapter's host/port resolution is now the shared
`gateway.platforms.api_server.listen_address()`, used by both the adapter and the
restart path, and the wait is skipped when api_server is disabled (a foreign
listener on the default port is nobody's race).
- Only ECONNREFUSED means the listener is gone. A timed-out connect (full accept
queue on a draining predecessor) was reported as "free", which would have let the
replacement start straight into EADDRINUSE again.
- The bind retry in `APIServerAdapter.connect` creates a fresh TCPSite per attempt
and unregisters the failed one: aiohttp registers the site before binding, so
re-starting the same object raises "already registered".
- The injected clock/sleeper/connect knobs are gone; the tests bind a real listener.
Tests: a real listener closed 300 ms in → wait returns True; a busy port taken from
config.yaml (env unset) → wait reports busy while enabled and is skipped when disabled.
E2E: with the predecessor releasing the port 0.5 s after the first bind, origin/main
logs `Errno 48 ... address already in use` and connect() returns False; this head
binds on the third attempt (+0.61 s) and connect() returns True.
On macOS api_server cannot SO_REUSEADDR, so replacing the process as
soon as the old PID exits still hits EADDRINUSE. The new process then
stays up with no API. Wait for the listen port to refuse connections,
and retry the bind a few times before treating the conflict as fatal.
Fixes#91547
(cherry picked from commit 5cb9545023646277b2fb59ebfb5e89b0fd161aba)
Gate review: releasing an undispatched runtime claim kept its pre-claim
error ("503 ..."), so the redelivery timer re-armed and claimed/released
the row every tier for as long as the adapter was absent - the loop the
previous commit closed for `send_path_degraded` rows only. Only a flood
row keeps its error (the platform wait must be honoured); every other
row is released as reconnect-only and re-claimed by the reconnect sweep.
`is_reconnect_only()` replaces the three spellings of that predicate.
Gate review on the stack:
- A runtime claim released because the adapter was gone comes back as `failed`
with `send_path_degraded`, a fresh `updated_at` and its attempt refunded.
`pending_retries` listed it as due immediately, so the redelivery worker would
claim and release it every ~2 s for as long as the adapter stayed absent, never
bounded by the attempts cap. Such rows are reconnect-driven: `pending_retries`
skips them and only the reconnect sweep re-claims them.
- With the last attempt reserved for the boot sweep only two in-process retries can
happen, so the 600 s backoff tier was unreachable and the docs promised a
"30 s, 2 min, 10 min" schedule the code never ran. Two tiers, asserted against
MAX_ATTEMPTS, and the docs say what happens.
- `_finish_ledger_delivery` no longer wakes the timer for a dead chat (blocked bot,
deleted group): the ledger would never retry it, so the wake was a wasted task and
a full pending-rows scan per dead send.
- The runtime sweep classifies a row's error only after the ownership filter, so
rows this process can never claim skip the classifier under the ledger lock.
Review on #91655 (@ehz0ah): a fixed 30/120/600 s schedule against an outage
that lasts hours burns the whole MAX_ATTEMPTS budget in ~12 minutes and the
row becomes `abandoned` — worse than before, when it at least stayed `failed`
and was recovered at the next restart.
`retry_not_before` now returns None once a row has spent MAX_ATTEMPTS - 1
attempts on an unclassified rejection, so the in-process timer stops arming
for it and the runtime sweep leaves it `failed` for the boot sweep, whose
adapter reconnect is a real recovery signal. Flood refusals and allowlisted
reconnect errors keep their existing budget handling.
A final reply the platform rejected with anything other than flood control
(a 5xx, a transient parse error, an unclassified failure) sat in the ledger
as `failed` until the next gateway restart, because the runtime sweep only
claimed `send_path_degraded` and flood rows and no timer was armed for it
(#91653). The deadline-driven redelivery worker main already runs for flood
rows now covers every failed row: `retry_not_before` keeps the platform's
wait for flood rows, retries allowlisted reconnect errors at once, backs
any other rejection off by attempts (30 s, 2 min, 10 min) so an outage is
not hammered, and returns None for a whole-chat death (blocked bot, deleted
chat) so a dead target is never resent. `pending_flood_retries` becomes
`pending_retries`; the adapter arms the timer on every rejection.
`release_runtime_claim`'s attempts decrement is untouched: it refunds an
attempt a claim never spent, which the backoff relies on unchanged.
design via #91655 by @dasgltd
Co-authored-by: Daniel Silva <daniel@dasg.ltd>
- HOST_INTERPRETER_KILL_REJECTION is one constant in cron/lifecycle_guard;
terminal_tool_guards and the execute_code lifecycle guard both use it,
so an image-name kill inside a cell now names the proc_* / explicit-PID
route instead of the generic "cannot restart or stop the gateway" text.
One invariant test with the generic path as control.
- gateway/restart.is_supervised_gateway_launch: the callee already maps
None to os.environ; pass environ straight through.
Review follow-up on #113667:
- _pattern_reaches_host_interpreter: when the `-f` head token is not an
interpreter image, require the pattern to plausibly reach the gateway
cmdline (hermes_cli / hermes+gateway tokens, mirroring Branch D). On
the previous head `pkill -f 'hermes-polis/run.sh'` and
`pkill -f my_hermes_bot.py` were hard-blocked although both are allowed
on main and cannot match the gateway; both spellings are now negative
controls in test_kill_forms_that_do_not_reach_the_gateway.
- lifecycle_ledger: the unclean-exit warning enumerated only OS causes
(SIGKILL / OOM / VM death); an agent- or descendant-issued kill of the
host interpreter leaves identical evidence and is now named in the
cause family. One invariant test.
- test file: encoding="utf-8" on the bare write_text calls the footgun
scanner flags.
An agent-issued `taskkill /F /IM python.exe` (or `pkill -9 python3`, `killall python`,
`Stop-Process -Name python`, `taskkill /FI "IMAGENAME eq python.exe"`, `pgrep python | xargs
kill`) from inside the supervised gateway killed the gateway: every branch of
_GATEWAY_LIFECYCLE_PATTERN was anchored on a hermes/gateway token, and the supervised gateway is
literally a `python` process. Branch E is token-aware (not a line regex) so option values are
never read as targets, `-f` cmdline patterns are judged as patterns (`pkill -f 'python
my_script.py'` passes, `pkill -f 'python -m hermes_cli.main'` does not), and other image names
(`taskkill /F /IM agent-browser.exe`) stay killable. Numeric-PID kills stay out of scope: the
explicit PID / `proc_*` id is the ownership-scoped route the terminal rejection now names.
The guard never ran on the Windows Scheduled-Task topology either: the launcher exports only the
generalized HERMES_SUPERVISED_CHILD marker, which gateway/restart.py never read.
is_supervised_gateway_launch() reads it and gates the self-kill guards;
is_gateway_supervisor_process() deliberately keeps ignoring it because it also selects the
exit-75 restart route, which the task has no restart policy to honour (#113670).
Supersedes the narrow `/IM python.exe` regex from #113671 (kept for authorship); the Windows
spellings from #94379 (`hermes.exe gateway restart`, `taskkill`/`Stop-Process` on hermes-gateway
tokens) ride along.
Fixes#113667
reset_file_dedup treats a falsy task_id as "all tasks"; mirror the existing
`or "default"` guard from the hermes-mode /compress site at the three forwarded
sites, and shrink the run_turn comment to the one non-obvious fact.
Sibling site of the previous commit: _compress_codex_app_server_session passed
no task_id either, so a manual /compress on a codex_app_server session reset the
"default" bucket and left the live session's read_file dedup armed. Forward the
session row id; one parametrised test covers both codex entry points through
the real compress_context against the real read tracker.
The gateway's hygiene compaction (generic sweep in run_turn.py and the codex
app-server variant in run.py) called _compress_context without a task_id, so
the read_file/skill_view dedup boundary reset ran under the "default" bucket
while the live turn records tool reads under the session row id — the task_id
the main turn hands to run_conversation. After a hygiene compaction pruned a
skill_view or read_file result, the next call for the same file returned an
"unchanged" stub pointing at content no longer in context (#98206).
Forward session_entry.session_id / session_id as task_id at both sites.
Manual /compress on the three surfaces was already fixed by b002dfc04d.
Follow-up to the salvaged #92440 commit, shape-gate cleanup only; behaviour is unchanged
(one mention-bearing payload per logical send, captioned media keeps its caption).
- Fold `_send_whatsapp_with_mentions` into `_send_plugin_standalone` (it was a line-for-line
copy of the caption split + `_send_chunks` loop); `mentions` is attached to the first
payload only via a one-shot kwarg dict.
- Drop the `inspect.signature(sender)` probe: the only registered WhatsApp standalone sender
is the in-tree `_standalone_send`, which gains `mentions` in the same change; a foreign
sender already surfaces as a TypeError through `_handle_send`'s error path.
- Collapse `_normalize_outbound_mentions` to dedupe-only; its input is argparse `list[str]`
already validated by the CLI.
- Revert the `\d` -> `[0-9]` edit to `_BARE_PHONE_RE` / `to_whatsapp_jid`: unrelated to the
feature (`normalize_whatsapp_mention_jid` already rejects non-ASCII via `isascii()`) and it
changed output for nine other `to_whatsapp_jid` callers.
- Split the single 137-line test into a `whatsapp_bridge` fixture + two invariants
(rejections never reach the bridge; mentions ride the first payload only, stale bridge
fails closed); drop the fabricated legacy-sender branch that only existed to cover the
deleted probe. Mutation check: removing the first-payload gate turns the new test red.
Co-authored-by: google-labs-jules[bot] <161369871+google-labs-jules[bot]@users.noreply.github.com>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: David Metcalfe <80915+DavidMetcalfe@users.noreply.github.com>
(cherry picked from commit ba7fd43826f9e36d886cc294e072378d5d082aa5)
toolsets_for_source recovered the route for the per-route toolsets grant by
splitting the session chat_id "webhook:{route}:{delivery_id}" on ":", while
authentication used the exact URL segment. A route named "build:external"
therefore resolved to route "build" and inherited its toolsets: a caller
holding the weak route's HMAC secret got the privileged sibling's terminal/
file tools. Reproduced on main through the real aiohttp handler with a
signed request.
Key the lookup on source.user_id instead, which _dispatch_agent_run already
stamps as exactly "webhook:{route_name}" from the authenticated segment.
No split at all: delivery_id is caller-supplied (X-GitHub-Delivery, svix-id,
X-Request-ID), so any parse of chat_id, including rsplit, stays attacker-
influenced. Routes without ":" resolve identically before and after; a ":"
route with no colliding sibling now gets the toolsets it configured, where
the old code silently fell back to the platform default.
Tests go through the HTTP handler, HMAC, and the gateway resolver: the ":"
route gets its own toolsets (red on base), and a crafted delivery id on the
privileged route cannot name another route (pins against a future rsplit).
Reported-by: pinarsadioglu (GHSA-2fmg-cjqm-hhrj)
The SQL chain step (#114287) stopped a `_reset_from` child of a compression-ended parent
from winning tip projection. The Python twin had the same blind spot:
`_is_compression_child_row` / `_compression_lineage_root` treated the reset fork as a
continuation, so `get_compression_lineage(tip)` collapsed to `[tip]` (ancestors lost for
prompt-cache scope and export) and the fork shared the lineage's turn-lease key. Both now
ask `_is_explicit_fork_child_row(include_reset=True)`; `get_compression_lineage`'s own
early return keeps excluding only branch/delegate/tool so a reset child that later
compresses still walks forward to its children.
Gateway bare `/resume` lists with `order_by_last_active=True`: a lineage compressed for
days is projected onto its live tip and belongs where the user last touched it, not at
its root's `started_at` (the reporter's tip, active yesterday, was buried under a
September-12 start). Desktop already requests `order=recent`.
Docs: `/resume` row in slash-commands reference. Tests: one lineage-walk invariant, one
/resume ranking invariant, both red on origin/main.
Part of #114271
Widen the container-backend refusal salvaged from #113530 to the sibling
surfaces that render the same host checkpoints: the messaging gateway's
/rollback (restore refused, bare listing prefixed with the reason) and
/diff session, and the CLI's /diff session. The gateway arm follows the
CLI's "default" classification, i.e. the configured terminal backend.
Drop the thin _checkpoint_container_backend wrapper in favour of the
container_backend_for_task predicate it wrapped, trim the salvaged suite
to two invariant tests (one per class: no host store touched by a
container task; every surface refuses a host restore/diff from a
container session, with a local control), and update the docs.
Co-authored-by: fangliquan <fangliquan@qq.com>
GatewayTurnMixin kept _FAILED_TURN_NOTICE / _PARTIAL_FAILED_TURN_NOTICE as
compat aliases of agent.turn_failure_copy; re-point the five gateway reads
and the test references at the module constants instead.
Terminal-failure paths (HTTP-200 content-policy refusal, ``_Trunc.end_turn``, retry
exhaustion, interrupt before any assistant text) persist the accepted user row and return
before ``finalize_turn``, so ``user`` stays the durable conversation tail. The next prompt
appends a second user row, ``repair_message_sequence`` merges the pair, and the provider
is asked to act on the failed request again. The gateway compensates with
``_hmwa_close_failed_turn`` (#108033); standalone ACP, the CLI and the TUI/Desktop hand
``result["messages"]`` straight back as history and had no closer.
Close it once, at ``agent/conversation_loop.py::run_conversation`` — the seam every
envelope leaves through — with a Hermes-authored assistant boundary
(``agent/turn_failure_copy.py::FAILED_TURN_NOTICE`` / ``PARTIAL_FAILED_TURN_NOTICE``,
which the gateway now aliases instead of keeping its own copy). Idempotence is keyed on
``SessionDB.latest_conversation_role`` (durable state, not content), so a redelivery or a
tail another writer already closed is a no-op and the gateway's closer no-ops in turn.
The context-pressure classes (``compression_exhausted``, ``compression_deferred``,
``failure_reason == "context_overflow"``) are excluded: appending to an oversized session
is the #1630 growth loop; their repair is rotation.
Adjacent defect from the same report: ``acp_adapter/server.py::_finish_turn`` called
``final_response.startswith`` on ``None`` for an interrupted turn — the same one-line fix
PR #64471 by @israellot filed first (its wider prompt()-restructure is superseded by the
current ``_finish_turn`` shape).
Slimmer redo of #114168 by @kendrickkester (same seam and invariants; the +1023-line
PR carried a new copy module, an accepted-turn re-anchoring scan and an 859-line suite).
Two invariant tests: the real ACP path (loopback provider, refusal then a new prompt) and
the durable-tail idempotence / overflow exclusion.
Co-authored-by: Kendrick Kester <kendrick.kester@gmail.com>
Co-authored-by: Israel Lot <israel.lot@gmail.com>
`_send_with_flood_retry` retried a short-wait flood result by calling
`adapter.send(**kwargs)` again with the FULL content. When the Telegram
adapter's split send had already landed the head chunks and returned the
`partial_overflow` flood result (its inline attempts exhausted with
retry_after <= the inline cap), that retry duplicated the visible head on
screen. Stop retrying when the failed result carries
`raw_response["partial_overflow"]`; callers already treat the returned
partial/flood failure correctly ("preview" / partial-continuation paths).
Test: a fake adapter returning a partial-overflow flood result on the first
send must not be called again with the full content and no retry sleep runs.
A reply past 4,096 chars goes out as several sendMessage calls. When chunk 2 was
refused by flood control (RetryAfter past the 5s inline cap) send() returned the
bare flood_control result, so _send_with_retry re-sent the WHOLE payload after
the wait: the user saw chunk 1 twice (reporter: 4 messages, 686 duplicated words),
and paths without a ledger row lost the tail outright.
- send() now reports a mid-split refusal through the existing partial_overflow
contract (the key _edit_overflow_split already sets and the stream consumer
reads): delivered_chunks / total_chunks / last_message_id, plus
undelivered_chunks + delivered_message_ids ONLY when non-delivery is certain
(flood cap, Bot API rejection, connect/pool timeout) — an ambiguous TimedOut may
have reached Telegram and is never resumed from.
- BasePlatformAdapter._send_with_retry resumes from the remainder via a new
_resume_partial_send hook (default None = keep the partial failure, never
re-send the head; the plain-text fallback is skipped for partials too). The
Telegram override sends the leftover formatted chunks, continuing the id sequence.
- Per-chat FIFO send gate, reentrant per asyncio task (media paths nest), held
only around the API calls — never across the reconnect wait — on send() and the
media funnel, so concurrent replies to one chat no longer interleave chunks.
- A flood refusal arms a per-chat cooldown (mirrors the sendChatAction cooldown,
capped 300s); sends inside the window fail closed locally with the same
flood_control:<s> result and no API call, so ledger recognition and redelivery
timing are unchanged.
- Edit path: log "refusing (retry_after Ns > cap)" after the cap check instead of
"waiting Ns" followed by no wait.
Live against a local fake Telegram Bot API with a fake token (RetryAfter=7 on
chunk 2 of a 3-chunk reply): before 4 messages / 324 duplicated words; after 3
messages, 853/853 words, 0 duplicated, 0 lost. Two concurrent 3-chunk sends:
before 9 source switches, after 1. Five sends inside a refused window: before 5
API calls, after 0.
Fixes#114396
Co-authored-by: AStrnbrg <45151087+AStrnbrg@users.noreply.github.com>
Co-authored-by: whyyagswhy <166958865+whyyagswhy@users.noreply.github.com>
The Feishu fix on this branch populates MessageEvent.media_text_inlined so
run_inbound's document note stops claiming "Its content has been included
below" when a text attachment was NOT inlined (>100 KB gate or decode
failure); run_inbound treats a missing flag as inlined. Telegram, Discord,
Slack, the WhatsApp bridge adapter and whatsapp_cloud inline "[Content of
…]" the same way but never set the flag, so their notes lied on the
skip path. Mirror the Feishu/buzz per-attachment contract in each: False
for every cached attachment, flipped to True only when the text was
actually injected.
One parametrized (small→True / large→False) test per adapter in the
existing per-platform test files.
The error now names the resolved home's `.env` (and whether the platform's
token key is defined there), `config.yaml` (block absent / `enabled: false` /
no token) and the environment variable(s) checked, so a Windows or profile
home user can fix the file this process actually read. When a gateway started
from the same home already has the platform connected, the message says its
token lives only in that process's environment and which key to add to `.env`.
Docstrings in `gateway/channel_directory.py` and `send_cmd._load_hermes_env`
stop naming `~/.hermes`; the pipe-script-output guide documents the message.
Once a stalled session's backlog is spooled (#114266), the order-preserving
replay attempt before each write hit the still-dead DB and logged a
gateway.shutdown_flush 'Replay of spooled transcript message ... failed'
WARNING per append, on top of the per-append ERROR escalation that already
reports the outage. When the session already has recorded append failures the
replay failure is expected and now logs at DEBUG; the first drain (no recorded
failure yet) still warns, and a recovered DB still replays the spool in order.
'Persisted transcript lagged live cached history' repeated at WARNING on every
turn for 11 days in #114266. _load_turn_history now keeps a per-session-key
streak on the runner (cleared on /new via _CONVERSATION_SCOPED_STATE); after
_TRANSCRIPT_LAG_ESCALATION_TURNS consecutive lagging turns the line logs at
ERROR with an operator-attention suffix, and a caught-up turn resets it.
A SessionStore whose state.db is unavailable (_db is None) early-returned from
append_to_transcript: no failure counter, no log, and the turn was gone. The
missing store now flows through the same serialized path (RuntimeError
'no owning session store ... deferring' is counted), so the WARNING->ERROR
escalation covers the reporter's outage shape.
Once a session hits the escalation threshold its in-memory backlog is moved to
the on-disk pending spool (recover_pending_to_db replays it at boot) instead of
sitting in memory until the 200-message cap or a crash. The spool is drained
BEFORE the next live write so recovery preserves transcript order.
Part of the 'best-effort durable write' atom of #114266.
Addresses review finding on PR #114301: _rebuild_fts_once was stamping
_fts_rebuild_last_attempt_at before the 'db is None or no rebuild_fts'
early-return guard, so a call with no usable DB burned the 5-minute
cooldown window without attempting anything. Move the stamp below the
guard so only real rebuild attempts start the cooldown.
Adds a regression test covering: no DB, DB without rebuild_fts, and the
follow-up real attempt succeeding immediately once a usable DB appears
(not blocked by a phantom cooldown). Full gateway/test_session.py suite:
67 passed.
When the session FTS (full-text search) index rebuild fails, the
gateway previously entered a cooldown and then permanently stopped
retrying once the cooldown expired, leaving search silently broken
for the rest of the process lifetime. This adds a proper retry after
the cooldown window elapses, and escalates to ERROR-level logging
(previously missing 'import logging' meant this path could not even
log the failure) so repeated rebuild failures are visible to
operators instead of being swallowed.
Related but out of scope: while testing the fallback paths we
observed that the JSONL transcript fallback writer can write zero
bytes when the primary DB write path fails partway through - flagging
this for maintainers as a separate follow-up, not fixed here.
Fixes#114266
The gateway exception reply (_STATUS_HINTS[401]) and the cron auth failure
notice still said a literal `hermes auth add <provider>` with no profile
selector — the placeholder this PR removed from the chat copy. Both now
format relogin_command_hint(provider): the exact OAuth command for a known
OAuth slug (turn agent's provider / the job's pinned provider), `hermes auth
add <slug>` for an API-key slug, and a profile-pinned placeholder when the
slug is unknown at the call site.
Part of #114012
The bootstrap racer covers sync connects only. The gateway's WebSocket dials
(relay connector, Yuanbao, Buzz) go through ``websockets.connect`` →
``loop.create_connection``, whose ``happy_eyeballs_delay`` defaults to ``None``:
a serial walk that burns the full connect timeout on every blackholed AAAA
record before IPv4 answers — the same stall class #114265 reports, one layer up.
``websockets`` forwards unknown kwargs to ``loop.create_connection``, so each
call site passes ``happy_eyeballs_delay=0.25`` (the RFC 8305 delay anyio and
the sync racer already use). One invariant test per call site captures the
kwargs at a mocked ``websockets.connect``.
`_handle_webhook` ran `_collect_attachments` (one REST download per attachment)
before the `is_group and self.require_mention` check, so every attachment on an
unmentioned group message was fetched onto the host and then dropped. The gate
only needs the text, which is available pre-download, so it now runs first;
the "(attachment)" placeholder and the missing-fields check follow the download
as before.
Same class as the Teams fix in this PR (review follow-up). One invariant test:
unmentioned group record with an attachment -> 0 `_download_attachment` calls,
mentioned -> 1 and dispatched.
The `[Triggering message id: …]` note that `_prepend_inbound_reply_context`
adds for Discord turns is a model instruction (which id to pass to the
discord tools), not something the user wrote. `_hmwa_apply_message_timestamp`
derived `persist_user_message` from the already-wrapped text, so every
Discord-origin user row stored the note as the first line of `content` —
the desktop transcript (and FTS, memory providers, exports) showed the
envelope instead of the message.
- `run_inbound.py`: the note is now the OUTERMOST prefix (after the reply
pointer) and rendered by one function, `discord_triggering_note`;
`strip_discord_triggering_note` peels exactly that prefix for THIS event
off the persisted text, so the `[Replying to: "…"]` pointer survives.
- `run_turn.py::_hmwa_apply_message_timestamp`: persist the stripped text.
The wrapper keeps riding `message_text`; when the durable row differs, the
live bytes land in the replay-only `api_content` sidecar by the existing
persist-override contract (same as timestamps and per-turn sidecar notes).
- `run_turn.py::_run_agent_queued_followup`: the in-band queued follow-up
runs the same inbound prep but passed no persist override; it now carries
the authored text too.
Live: real inbound prep → persist seam → SessionStore.append_to_transcript on
a temp state.db. Before: content='[Triggering message id: `1550…` — use as
`message_id` …]\n\nCreate a project plan for Q4'. After: content='Create a
project plan for Q4', api_content=<wrapped bytes>; reply-pointer and
no-message_id (desktop relay) controls unchanged.
Slimmer redo of #71309 by @JonthanaHanh (regex strip on a `gateway/run.py`
site that no longer exists; tests exercised a copied regex); #71619
(@calvinnwq, +1303/-40 over 10 files) is over the salvage bar.
Fixes#71304Fixes#114719
Co-authored-by: JonthanaHanh <92574114+JonthanaHanh@users.noreply.github.com>