The invariant is "per-platform streaming stays off when global
streaming.enabled is false"; interim assistant messages are a separate
delivery path, so turn them off to keep the test about streaming.
Co-authored-by: WadydX <65117428+WadydX@users.noreply.github.com>
Addresses @harjothkhara's review comment on PR #53709 — adds a
runtime-path test verifying that display.platforms.telegram.streaming=True
does not re-enable gateway streaming when streaming.enabled=False.
The test exercises the production _run_with_agent path with Telegram
platform, per-platform streaming override True, and global streaming
disabled, asserting no streaming edits are sent.
(cherry picked from commit b8b08912b60daeefa840b376e6926ee54bd0b025)
Session chat streams already register the agent under the run id (stop works),
but guarded tools had no approval notifier and failed closed. Register a
run-id-keyed approval session, emit approval.request on the SSE stream, and drop
the mapping when the turn ends.
Co-authored-by: baleian <baleian90@gmail.com>
Hand-grafted from #51878 onto the split api_server_openai_routes.py; approvals keyed
by completion id and resolved via POST /v1/runs/{id}/approval (#51871).
dingtalk-stream 0.24.3 start() retries forever inside the SDK, logging a
malformed logger.exception() every 3s, so the adapter breaker never saw the
error. Install a dedup filter on the SDK logger that can't raise on bad
format args, detect the websockets incompatibility (bare or chained
TypeError), log one ERROR with a pin-consistent hint, and hand off via
_set_fatal_error(retryable=False) + _notify_fatal_error(): only a
reinstall of the pinned versions and a restart fixes it. Breaker stays
tripped until the error type changes; constants moved to module level.
DingTalk stream-mode reconnection storms the gateway: when start() raises
the same error every cycle, _run_stream logs a WARNING and reconnects at the
60s cap forever, generating hundreds of MB of identical log lines and hanging
the gateway.
Add a per-error-type circuit breaker: after 5 consecutive identical errors,
suppress the repeated WARNING (one ERROR summarises) and pause 300s instead of
spinning at 60s. Reset backoff + counters after a clean start() so a recovered
connection is treated fresh.
Salvage of #24881 re-implemented on current main: the original fix targeted
gateway/platforms/dingtalk.py, which has since been refactored to
plugins/platforms/dingtalk/adapter.py. Same logic, new path; tests import the
new module.
Closes#24851.
(cherry picked from commit 84201c58255ae3d3c9b269ac777ea0ff444d9229)
The Telegram final-delivery suite is the wrong home and its delete
assertion duplicates test_stream_consumer.py. Keep only the new
behaviour: a full resend threads its first chunk to the originating
message; later chunks stay unthreaded.
Surface code/msg at debug when the recall API returns non-success
(matches the exception path) and trim the docstring. Drop the trivial
disconnected-client test.
Feishu had no delete_message, so a failed finalize-edit plus fallback
send left the truncated edit bubble next to the full final. Implement
the SDK delete and thread the fallback send to the originating message.
(cherry picked from commit c61add84ad40b1bc39288405a0a05b4f621e00fc)
Reuse StreamFallbackMixin._fallback_len_budget() (message_len_fn default,
debug-logged per-chat override) instead of a second silent try/except ladder,
so the live budget and the fallback chunker can't drift. Drop the
message_len_fn=None test: message_len_fn is a BasePlatformAdapter property,
so that state is unreachable for real adapters.
Drop the full-preamble reset counterpart; the durable-preamble
on_new_message contract (#17280) is already covered in
tests/gateway/test_stream_consumer.py.
A segment-break finalize never carries the streaming cursor (the cursor
is only appended to mid-stream frames), so the sub-floor
standalone-message guard's cursor-membership test let every 1-2 token
preamble land as its own durable message at each tool boundary. Each
landing fired on_new_message, resetting the gateway's tool-progress
anchor and fragmenting accumulated tool progress into one persistent
message per tool on draft-streaming platforms (#99026).
Extend the guard to cover segment-break finalizes (finalize and not
is_turn_final). Turn finals stay exempt so a complete short answer is
never swallowed. Update the existing per-segment callback test to use
segment text above the floor, and add draft-mode regression coverage:
- short preamble rounds: no durable sends, no progress resets, drafts
and the gateway final path unaffected
- full preamble rounds: still one durable send + one reset per segment
(#17280 chronological-order contract preserved)
- turn-final short answer: always delivered
(cherry picked from commit 5ac896e1f22e6b2164544287c342c7fc7ec2c373)
Ladder message_len_fn_for_chat -> message_len_fn -> len in
_resolve_length_budget (teknium1 review on #72629) and add regression tests
for the hot-reload mixed-version adapter (#72628).
The home-channel/thread assertions in
test_restart_home_channel_notification_not_deduped_across_threads
expected raw thread metadata and None; shutdown sends now carry
_interim_send per #98432. Align both assertions.
(cherry picked from commit c4ec949ae77ae5611684061a4afb8d96af335e26)
Drop the local telegram stub (conftest._ensure_telegram_mock installs it)
and the copied _make_event/_DummyAdapter/_make_adapter in favour of the
helpers in test_active_session_text_merge.py. Merge the debounce and
direct-merge cases into one test parametrized over busy_text_mode.
Queue-mode (debounce) and direct-merge follow-ups whose busy handler
returns False after the running turn released the guard must start a
fresh turn; a live guard keeps the queue-behind behavior.
Test file taken verbatim from 33801e8fb7 (PR #121464). Its base.py hunk is
dropped: the same guard re-check is already provided (more broadly) by the
preceding commit from PR #121003, and the two hunks conflict.
(test file cherry picked from commit 33801e8fb7d0c1553511e470f410b636ed9426b8)
The base adapter sends a message to the runner's busy handler while the
previous turn still holds the session guard. The handler awaits before it
decides: the compression-lock read in the stock interrupt mode, or the
profile secret-scope load when profiles are multiplexed. If the previous
turn reaches _finish_session_task meanwhile, it finds the slot empty,
releases the guard and exits. The follow-up is then queued with no task
to drain it, either by the runner FIFO or by the base adapter when the
handler returns False. It waits until the next inbound message, whose new
turn runs BEFORE it (M1, M3, M2). If nothing else arrives, it never runs.
_handle_message_while_active now checks the guard after the handler
returns. If the guard is gone, it starts what the handler queued. If the
handler left the event to the base path and nothing is queued, it starts
the event itself.
A handler that raises after it stored the event (for example when
composing the busy ack fails) now counts as having handled it, based on
the event's _gateway_accepted receipt. Otherwise the new path would start
the stored event and the base path would queue it again, so it would run
twice. This also stops the older no-race case from merging the same text
into the slot twice ("M2\nM2").
tests/gateway/test_busy_followup_guard_release.py covers the stock
interrupt and multiplexed queue-text routes, and the case where the
handler raises after it stored the follow-up.
Known behavior, unchanged or out of scope:
- During a restart drain with busy_input_mode queue/steer, a follow-up
started this way hits the drain gate ("not accepting new work")
instead of waiting in the slot for the shutdown flush. The non-race
path does the same (#82381).
- A message can still run before an earlier follow-up if it arrives
after the guard is released but before that follow-up's handler
returns (for example during its busy-ack send). Nothing is lost.
- The interrupt decision still uses the running agent read before the
handler awaits.
- The "Interrupting current task" ack can arrive after the turn has
already ended.
(cherry picked from commit bb9f15c7fa29cdffc56790ea3118d8baac2195f1)
linux_only / macos_only / windows_only were replaced by platforms(...)
and conftest now rejects them, but test docstrings and comments still
explained gating in terms of the old names, which points readers at a
marker they cannot use. Reword them to platforms("<os>") and the
-m platforms lane. ci.yaml's comment already says platforms("windows").
The stack's regression tests pasted the same platform-parity loop into four
files and added eight cases for one invariant ("a seeded config must not
pin a global display key"), over the AGENTS.md invariant-test budget.
- assert_keeps_platform_display_defaults() in test_display_config.py is the
single parity helper.
- test_shipped_template_keeps_every_platform_default moves out of
TestYAMLNormalisation (it was attached there only by indentation) into
TestInstallerSeededConfigThroughGatewayResolver, reuses _seed_like_installer
and absorbs the _resolve_gateway_display_bool qqbot hop.
- test_config_edit_seed's test becomes the one parametrized every-seeder test:
config edit (template / no template), `hermes setup agent` with Enter (the
W1 leak fixed in this stack), _apply_default_agent_settings and blank slate.
The per-file copies in test_setup_agent_settings / test_setup_blank_slate
are removed.
- Dropped test_template_seeded_home_keeps_qqbot_reasoning_off (duplicate; its
dead HERMES_HOME set + raising=False patch went with it) and
test_show_reasoning_stays_a_known_config_key (change-detector for the
rejected #121232 approach). The explicit opt-in control (#7148) stays.
On origin/main's production files the template test and all five seeder
cases fail; the control passes.
Mirrors the #121230 report: copy cli-config.yaml.example to config.yaml the way
the installers/doctor/docker do, then resolve show_reasoning for QQBot/Telegram/
Discord via gateway.run._resolve_gateway_display_bool with the exact arguments
run_turn.py uses. Fails on the pre-fix template (True), passes with the fix.
Controls: an operator's explicit global display.show_reasoning: true still
reaches gateway platforms, and display.show_reasoning stays a known key for
`hermes config set` (i.e. the fix must not drop it from DEFAULT_CONFIG).
Co-authored-by: KoNit-K <konit.block@protonmail.com>
The curl installer, the Windows installer, the Docker first boot and
`hermes doctor --fix` copy cli-config.yaml.example into config.yaml byte
for byte. The template had five display keys uncommented: tool_progress,
interim_assistant_messages, long_running_notifications, busy_ack_detail
and show_reasoning. The gateway reads config.yaml without a DEFAULT_CONFIG
merge, and resolve_display_setting takes a global display.<key> ahead of
_PLATFORM_DEFAULTS. So every seeded home ran with those values on every
platform. Telegram and Slack posted every tool call. Signal, email, SMS
and the other no-edit platforms got progress lines, heartbeats and
interim messages. Every messaging reply had the reasoning block prepended.
First-time `hermes setup` (quick and full) and Blank Slate setup also
wrote display.tool_progress: "all". That write was added as a Quick
Install recommended default (79aeaa97e6) nine days before the
per-platform tiers landed (#8006), and it has the same effect for
tool_progress on homes the template never touched.
`hermes config edit` on a home with no config.yaml wrote DEFAULT_CONFIG
unstripped, which pins show_reasoning (all 21 platforms),
interim_assistant_messages (12) and tool_preview_length (16). It now
seeds like the installer and `doctor --fix`: the template when the
checkout has one (a full file to edit, written owner-only), otherwise
DEFAULT_CONFIG with defaults stripped.
Measured through the real gateway loader across the 21 platforms in
_PLATFORM_DEFAULTS, a template-seeded home differed from a bare one on
tool_progress for 19 platforms, show_reasoning for 21, busy_ack_detail
for 14, long_running_notifications for 13 and interim_assistant_messages
for 12. With the pins commented out and the setup writes removed, the
diff is empty, and the same holds for both `config edit` seeds.
The CLI does not depend on these values. It defaults tool_progress to
"all" and show_reasoning to true when the keys are absent, and the TUI
defaults interim_assistant_messages to true.
Homes that were already seeded keep their values. A template value
cannot be told apart from one the operator chose, so there is no
migration. The messaging docs now say which lines to delete.
(cherry picked from commit 96450d4500613ab1ba45c7e972f31de570bc2d71)
The terminal tool stores the container key in ProcessSession.task_id
(`session:<key>` on local/SSH, `default` / `profile:<name>` under
persistent Docker) and the spawning turn's id in owner_task_id. The
abandoned-turn cleanup from #76687 (/stop, /new, eviction, the
inactivity watchdog, API-server disconnect and /v1/runs stop) snapshots
and reaps by the turn's raw session id, and snapshot_running_ids /
kill_all compared it against task_id, so they never matched: the
turn's background jobs kept running after the turn was abandoned, and a
notify_on_complete job woke the session later. process_manage(list)
missed the session's own jobs for the same reason.
Match ownership on owner_task_id in snapshot_running_ids, kill_all
(and so kill_started_since) and list_sessions(task_id=), falling back
to task_id for entries without an owner. has_active_processes stays on
the container key: its caller passes container keys.
The two new tests spawn through the real terminal tool inside a session
context and drive the real /stop core and inactivity watchdog; both
fail on the previous code.
(cherry picked from commit 1735bbf84dae90ea2c68499a8971f6bed1570366)
Two of the three `request_serve_profile` monkeypatches in this file were
unreachable (proved with a side-effecting stub: only the standalone-owner
test's copy ever ran). Both tests publish the record under their OWN pid, so
on the lock-lost path `_owner_is_standalone()` returns False at the
`owner.pid == os.getpid()` shortcut before the wire is consulted, and the
stubs pinned nothing.
test_a_multiplexing_owner_is_still_refused documents "an owner we cannot
interrogate is treated as a multiplexer", which IS _owner_is_standalone's
logic, so it now publishes a different-pid, served-unknown owner (as the
standalone-owner test does) and the never-answers stub is actually asked.
Mutating _owner_is_standalone to answer True now fails this test; it passed
before.
test_a_replace_unit_that_replaced_nothing_is_still_refused_when_it_loses_the_lock
mocks _host_attach_or_none so decide() never runs; it patches
`gateway.run._owner_is_standalone` to False where _claim_host_gateway_role
reads it, pinning the branch its docstring describes (a multiplexing lock
holder must not be rescued by the start-beside carve-out) so it survives a
change to that pid shortcut.
decide() runs twice per start (the CLI guard in hermes_cli.gateway and again
from start_gateway -> _host_attach_or_none), and the lock-losing path in
_claim_host_gateway_role re-derives the same fact via _owner_is_standalone(),
so every boot of a generated unit beside another profile's standalone gateway
logged the "starting beside it" WARNING three times. decide() is a verdict
function; reporting belongs at the action site. Demote host_attach's copy to
INFO and keep run.py's lock-claim WARNING, which fires exactly once on both the
--replace and plain paths and carries the `gateway migrate --multiplex` hint.
The lifecycle test still asserts the converge hint is logged by decide(); it
now captures at INFO.
Allow read-only merged-tag queries through the live-system guard, route checkpoint rekeying to its real ref-deletion owner, and update stale fixtures to exercise current update and PM boundaries. Fix the shutdown test wait by patching the bound server global.
Move retention, orphan pruning, status and clear operations into a topical sibling and route callers and tests directly to their defining module. Keep the shared store paths and git execution in checkpoint_manager; shorten the local child-env WHAT docstring.
Conflict resolutions and semantic fixups:
- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
fence missed, deregister from _live_foreground in a finally) wrapped around
pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
the old early-recovery block stays gone; main's interrupted-pull restore
(auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
and is cleared once git is done. The marker's target is the ref git actually
moves to (a release tag, not always origin/<branch>), since the restore
compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
had dropped and main's auto-merged restore needs (NameError on the first
launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
settings.about block; trim it to `updates` as pm-clean's type and the other
overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
key-leak switch leg needs the SDK, and the api_server two-tenant test needs
aiohttp, both PM runtime extras the test env does not carry.
Follow-up to the salvaged mirror commit.
- Opt-in (`mirror_to_session: true`, or `hermes webhook subscribe --mirror-to-session`),
default off like cron's `mirror_delivery`. The mirrored text lands with user authority in
the target chat, and on `deliver_only` routes it is the raw rendered payload, so a route
author has to ask for it. Only a real boolean `true` opts in (a YAML string "false" no
longer does).
- The mirror runs inside the routed profile's scope, so a `/p/<profile>/` delivery writes
into THAT profile's state.db. A Telegram DM chat_id is the user's id on every bot; an
unscoped mirror landed in the default profile's DM with the same person (live-probed).
- Tests trimmed to two invariants against a real state.db: an opted-in delivery lands in the
routed profile's chat session and not the default profile's; a route without the opt-in
(including `mirror_to_session: "false"`) never touches the target transcript.
- Docs: webhooks guide section "Replying to a delivery" (default, profile scope, trust note),
CLI reference row, bundled hermes-agent skill reference.
A webhook route that delivers to telegram/discord/... runs in an ephemeral
webhook:<route>:<delivery_id> session. The delivered text never reaches
the target chat's own session, so when the user replies there ("so he's
out?") the agent has no idea it just sent them anything and asks what
they mean.
After a successful cross-platform send, mirror the delivered text into the
target chat's transcript as a labelled user turn ("[Webhook delivery:
<route>]\n..."), the same path and role convention cron briefs use
(cron.scheduler_delivery._maybe_mirror_cron_delivery, #2221). Best-effort:
never fails the delivery, skipped when the chat has no session yet.
Route opt-out via mirror_to_session: false. delivery_info now carries the
route name and the flag for both agent-run and deliver_only routes.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CZ15jNgPUHdPsMw4wWon6v
- tools/browser_tool_install.py: keep pm-clean's frozen old-updater stub; main's
UTF-8 decode fix touched only the npx prefetch body it replaces.
- tests/hermes_cli/test_update_scoped_reconciliation.py: keep pm-clean's test
subset (catch-up rides the PM completion owner) and take main's gateway-less
host evidence (#120740): the updated seed that holds the host at a running
gateway, and the two gateway-less matrices for the source change that merged
cleanly into update_cmd_fleet.py.
Independent-review follow-up for the group room failed-turn fix.
- Group room poll: a retained error newer than the pre-submit snapshot
(turn start replaces it with a fresh started_at) is this turn's failure
and wins over transcript text. The core closer writes no failed-turn row
behind a tool row, so "said X, called a tool, provider 401" left X as the
newest assistant row and the room posted it, dropped the 401 and re-drove
the member to the round cap (live repro on the PR head).
- Both pickers end the turn at a failed_turn row instead of scanning past
it; with the retained error gone (backend restart) the row's notice is
reported through the failed path instead of a silent pass / dropped
stranded marker. The stranded harvest treats a retained error as the
stranded turn's own the same way.
- REST cold-load/paging (/api/sessions/{id}/messages, /messages/around)
type legacy untyped notice rows like session.resume does, via one
read-side helper in agent/turn_failure_copy.py.
- Gateway closer: the fresh-session closure test now asserts the row's
display_kind through a real SessionDB round-trip (red without the
run_turn.py stamp).
- E2E: mock trigger that says text + calls a tool, then 401s; spec asserts
the room reports the error and never posts that text.
Invariant for the import-time config bridge: a multiplexed backend's first import of
gateway.run can happen inside a routed profile's session (hermes serve imports it lazily
from an agent build), and the bridge must still write the launch home's agent.max_turns and
terminal.cwd into the process env. Red on the previous gateway.run (HERMES_MAX_ITERATIONS
came from the routed profile), green with get_process_hermes_home().
Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.