Authorize route-only profiles using the ordered canonical route matcher and
served-profile set at both claim and delivery. Preserve secondary credential
boundaries, retry denied routes, and keep scope/parent anchors plus transport
provenance on synthetic wakes. Install the destination runtime scope rather
than inheriting the notifier's scope; a removed profile cannot wake as primary.
Slim forward-port of the direction in #101196/#101397 and #93863 (#93851).
Canonical scope_id takes precedence over the guild alias and live chat cache.
Two route invariants and two runtime-scope invariants reproduced red first.
Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Salvage #82561's idle watch consumer through the existing drain, retaining
notify-off suppression and retrying failed injections. Preserve #102647's
secondary-profile ownership for reconstructed sources, preflight, synthetic
completion turns, and raw process progress/final notices. Unavailable
transports defer rather than borrowing the primary bot or losing the event.
Keep retained transport provenance and alias-aware relay resolution. Add
three invariant tests and an offline real-process/Discord lifecycle eval.
Co-authored-by: trwpang <134544126+trwpang@users.noreply.github.com>
Co-authored-by: holny <holny@foxmail.com>
Preflight source-first transport resolution and parent/API database readiness
before spending any durable delivery attempt, including every batch sibling.
Unavailable targets remain retryable; terminal targets retain their drop
semantics and stateless API completions remain durable display rows, not turns.
Minimal combined salvage of #82704 and #96086; adjacent eligibility, profile
adapter selection and idle watch-consumer work remain separate.
Co-authored-by: Screaming Sun <5552699+eventh0riz0n@users.noreply.github.com>
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Salvage the busy-auth ordering fix from #95602. Gateway-owned wakes have no external identity; keep external and plugin authentication unchanged. Preserve pending human text by appending wakes through the existing FIFO instead of base-adapter text merging.
Co-authored-by: Robert Montelongo <ramontelongo22@gmail.com>
Bind settlement to the exact adapter task and event. Rejected or cancelled preparation refunds the existing claim; cancellation after the agent runner starts remains counted. Keep profile-scoped callback context and the manager replacement guard. A fire count is not outbound delivery proof. Drop departed routes rather than executing their stale schedule.
Slim accounting-invariant salvage of #93174; preserve current direct adapter dispatch instead of reviving its FIFO/inflight implementation. Prior art #92858.
Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com>
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Recover active watches from current persisted session origins and exact route
keys, reading heartbeat state off-loop in each source's profile. Failed scans
leave watches intact for the poller's next retry. Start the heartbeat poller
even when startup restores no watches.
Slim synthesis of #92660, #98310 and #98313, with earlier restart recovery
prior art from #92594. The integration poller calls restore_heartbeat_watches
on every poll, including empty registries.
Co-authored-by: chelsealong <chelsealong@126.com>
Co-authored-by: liuhao1024 <sunsky.lau@gmail.com>
Co-authored-by: fangliquanflq <fangliquan@qq.com>
Clear the departed schedule using the route-owned database when ending a conversation, so stale watches cannot fire after reset and resuming the archived session cannot resurrect the schedule. Compression migration already archives its parent and remains unchanged.
Adjacent lifecycle reports: #80273 by pierrenode and #80208 by 0xGr1mm. This follow-up covers durable reset cleanup, not their poller changes.
Namespace delivery markers and assign fresh keyless turn identities instead
of inferring ownership from IDs or process-local row baselines. Query only
marker existence on the canonical live compression continuation and ancestors.
Preserve raw reply IDs and exclude metadata from provider wire messages.
Expand the two existing invariants with resumed cross-chat ID collisions,
a real independent SQLite writer, reaped siblings, and archived-history
allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass.
The browser and Modal fixtures replaced agent/hermes_cli with packages whose empty search paths hid newly imported production modules. Point those fake package search paths at the real package directories while retaining their explicit collaborators. This exercises real secret-scope and subprocess helpers without weakening the lifecycle or snapshot assertions.
Keep the exact hidden-directory and visible-hit assertions, but avoid an incidental process-killer token in the fixture query. The installed live-system pytest guard flattens argv and mistakes real skill plus a hermes temp path for a process-killer command. Neutral fixture text leaves that guard enabled and the search behavior unchanged.
The splice hazard the tail seed guards against is live in remote-lifecycle's
scrapeReadyPort: the remote spawn log merges stdout and stderr, yet it was
matched with the `^`-anchored READY_RE. Export one READY_IN_MERGED_OUTPUT_RE
(lookbehind token boundary — the hand-rolled `[^0-9A-Z_]` class admitted
lowercase-glued tokens) from backend-ready.ts and use it for both merged
buffers. Comments trimmed to the WHY; 6 new tests collapsed to the two
invariants (splice recovers, prose does not match) plus one remote-log case.
Both callers build the wait in the same synchronous block as
`outputTail.attach(child)`, and stream 'data' is asynchronous, so
`bufferedOutput()` is always empty at the seed and this block recovers
nothing today. Say so at the code, so nobody reads the seed as the live
recovery path for a lost READY sentinel — and so the condition that makes it
load-bearing again (any await reintroduced between attach and this call, the
pre-#100442 ordering) is written down next to the code that depends on it.
The spawn-time output tail feeds stdout AND stderr into ONE buffer, so its
contents are not line-accurate. uvicorn logs to stderr in raw chunks, and a
chunk that ends without a newline is concatenated directly onto the stdout
sentinel that follows it:
INFO Started server process [4711]HERMES_BACKEND_READY port=65238
The late-attach seed scan added for #60323 matched that buffer with the
line-anchored `_READY_RE`, so `^` never lined up and the seed silently
recovered nothing. The live stdout listener cannot cover the gap either: the
sentinel was already consumed by the tail, and flowing-mode streams never
replay chunks to late listeners. The wait then runs out its full 90s deadline
and a healthy, listening backend is SIGTERMed — while desktop.log, which reads
both streams, plainly shows the READY line. That contradiction (READY logged,
boot timed out anyway) is the reported macOS regression.
Scan the tail with a boundary-guarded pattern instead of a line-anchored one.
`port=<digits>` keeps the token unambiguous, so prose naming the sentinel and
longer identifiers ending in it still do not match. The live per-line scanner
keeps `^`: individual stream chunks ARE line-accurate there, and loosening it
would let unrelated child output settle the boot on a bogus port.
Scope: the seed only. Watching stderr on the live path is #97086's change and
is deliberately untouched here — the two compose.
Drop the 3-line facade wrapper (hermes_cli/gateway.py is already 3x the facade
threshold) and call the existing _ensure_user_systemd_env() directly under
`is_linux() and INVOCATION_ID` — the same Linux gate the process_registry seam
uses, instead of os.name == "posix". The fail-closed test now targets
_ensure_user_systemd_env() itself. Hedge the scope-unavailable error text: the
probe also returns False when systemd-run is missing or times out, so the
D-Bus diagnosis is the usual cause, not the only one.
Drives the real run_gateway() boot with the service-manager environment (no bus
vars, INVOCATION_ID set) and asserts the bus address reaches the worker env the
dispatch paths build — including surviving the secret scrubber, without which the
fix would be a silent no-op. The companion test pins the other half: with no user
manager present nothing is fabricated, so the scope probe's refusal stays honest.
A system-level unit (/etc/systemd/system, User=<someone>) is exec'd with neither
XDG_RUNTIME_DIR nor DBUS_SESSION_BUS_ADDRESS, and a process environment is fixed
at exec time. 'systemd-run --user --scope' therefore fails for the whole lifetime
of that gateway even after the user manager is up and /run/user/<uid>/bus is
reachable. That is the seam every restart-safe worker crosses
(restart_safe_gateway_child_argv), and it fails closed by design — so on headless
systemd installs every agent-driven cron job and every Kanban dispatch died at
launch, ~26ms in, with nothing but 'error' on the job row.
_ensure_user_systemd_env() already derives both values from our own uid and adopts
them only when the runtime dir is really ours and the socket really exists; it was
just wired exclusively to the systemctl management paths, never to the gateway's
own boot. Call it from run_gateway() — the single in-process boot every entry point
goes through — so the adoption precedes every worker-environment snapshot (cron
builds its env after the scope check, Kanban before it, so fixing this at the
dispatch seam would only fix one of them).
The fail-closed posture is unchanged: with no user manager at all the probe still
reports unavailable and dispatch still refuses. That refusal now names the remedy
in the message the operator actually reads (it is stored as the cron execution's
error), instead of only the symptom.
Fixes#104893
_root_mount_has_marker duplicated _proc_file_has_marker's open/OSError shell;
both now go through _read_proc(), and the root-mount scan checks every "/"
line instead of returning on the first. The regression test exercises the
helper on both arms (host running containers -> False, containerd rootfs ->
True, missing file -> False) instead of patching builtins.open globally.
Closes#58135
On a cgroup-v2 host with Docker's containerd image store, is_container()
false-positives whenever any container is running. The mountinfo fallback
scanned the entire file for containerd/crio/kubepods substrings, but each
running container contributes an overlay mount whose option string carries
lowerdir=/var/lib/containerd/..., so a plain host was misclassified as a
container. The result is cached per-process, making it depend on whether a
container happened to be running at first call — which then flipped
subprocess HOME and broke browser tool launches (Chrome not found).
Fix: only inspect the root ('/') mount line. Inside a container the root
mount is the runtime's overlay/snapshot and carries the marker; on a host
the root is a real block device and container overlays live at non-root
mount points.
Adds a regression test reproducing the host-running-containers case.
(cherry picked from commit 588f55113ebc16813b8216772feb761b6292d922)
(cherry picked from commit f42295814d7e2b5fe681e6d1b96ea72026058195)
The previous test replaced _resolve_context_length itself with a fake that
re-did the threading, so it could not fail if the production call dropped the
kwarg; the sibling test only called get_model_context_length directly (green on
main, and its list-shaped `models` fixture is silently ignored by the config
loader). One test now patches agent.context_compressor.get_model_context_length
where production reads it and asserts the kwarg arrives; file moved next to the
other compressor tests. The defensive list copy is dropped (no caller mutates).
ContextCompressor._resolve_context_length called get_model_context_length
without custom_providers, so a per-model context_length override in
custom_providers (e.g. 256K for baseten-deepseek-flash) was skipped and the
hardcoded catalog default (128K for 'deepseek') won instead. /context then
reported the wrong window. Thread custom_providers through the compressor
constructor and into resolution.
The new standalone test file duplicated TestGitBashExternalProgramProbe's
harness; one assertion on the recorded kwargs covers the #78820 contract.
The sibling ASLR-probe assertion is dropped (unchanged code, already green
on main). Comment trimmed to the WHY.
`_bash_starts()` in tools/environments/local_gitbash_probe.py ran bash.exe
with capture_output=True but no stdin=, so the MSYS2 child inherited the
TUI gateway's stdin pipe. The MSYS2 runtime switches that shared pipe to
PIPE_NOWAIT; the gateway's next sys.stdin.readline() then fails with
ERROR_NO_DATA, which the CRT maps to OSError(EINVAL), and the gateway
exits with code 1 ("gateway exited") on the first terminal call.
The sibling probe _mandatory_aslr_enabled() already passes
stdin=subprocess.DEVNULL; this brings _bash_starts() in line with it.
Add tests/tools/test_gitbash_probe_stdin.py asserting both probes forward
stdin=DEVNULL to subprocess.run (fails on the pre-fix source).
The added test was a copy of test_approval_click_once_maps_to_once with only the
key's chat_type changed; retarget the original to the key build_source actually
emits. Both branches stay: approval keys carry "dm" (build_source), update-prompt
keys carry event.scene ("c2c") — a comment records why.
PR #30737 added authorization checks for approval button clicks but
hardcoded chat_type == 'c2c'. C2C sessions are actually created with
chat_type='dm' (build_source at L1278, L1488), so the check always
rejects DM approval buttons as unauthorized.
Accept both 'c2c' and 'dm' for direct-message session keys.
_resolve_child_fallback_chain re-implemented hermes_cli.fallback_config's entry
validation to log per-index warnings, wrapped a pure function in try/except, and
carried two names (routing_cfg / fallback_cfg) for one argument. It now delegates
to get_fallback_chain() and keeps the single warning for "no usable routes";
the facade re-export is dropped (import from the defining module). Tests collapse
to the parametrized decision table plus the pin-derivation, real-config-loader,
review-ownership and activation invariants (22 -> 20 cases, 417 -> ~230 lines).
A delegation.model-only pin (provider inherited) took the unpinned
column because pinned was keyed on override_provider alone, so the
child inherited the parent chain and a mid-run failure could silently
swap the pinned model. model at the call site is creds["model"] --
delegation.model config -- at both call sites, so keying on it adds no
false pins. base_url-only pins already resolve to a provider override.
Caught by review on the PR; two matrix tests added (model-only pin with
absent and declared chains).
Adopted from PR #80421 with the author's explicit go-ahead on #80450
('Please proceed!'): config_defaults entry, cli-config.yaml.example
block, and user-guide docs for the delegation-scoped fallback chain.
Co-authored-by: Andrex Ibiza, MBA <84248988+andrexibiza@users.noreply.github.com>
Composes #80465's pin semantics with #80438/#80421's
delegation.fallback_providers semantics (#65038) and settles the
composition cell the three PRs leave undefined (#80450 map).
_resolve_child_fallback_chain implements the whole decision table in
one pure function: pinned children get no chain unless the delegation
section declares one (a delegation-scoped chain honors both explicit
intents; the silent drag in #80450 is specifically the PARENT chain
substituting a pin); an explicit empty list disables fallback; absent
config preserves historical parent inheritance exactly. Malformed
declared values log and fall back PIN-AWARE — never the parent chain on
a pinned child, so the config-error path cannot reintroduce the drag.
Entries normalize through the canonical get_fallback_chain.
The two existing inheritance tests are made hermetic against the real
user config (as #80438 also did).
Stress run with a real orchestrator subagent showed the per-task keys buried in results[] were not relayed
upward; the same prose lines now sit at the top of the sync delegate_task result. evals/subagent_process_handoff/
stress_handoff_live.py runs 12 real parent+child scenarios (handoff, orphan, unread, clean read, cap, exited refusal,
mixed fan-out, parent controlling the inherited process, sibling theft incl. adversarial, nested orchestrator,
5-way burst) and scores them against runtime accounting, never model prose.
A process that finishes while the child is alive needs no handoff, but if the child never
polls/waits/logs it, the result vanished: the completion notice is suppressed in the parent
and the child's summary never mentions it. Finalization now attaches exit code + output
tail as unread_completions, rendered in the parent's delegation notice.
The schema-diet test froze the verb list (a change-detector); it now asserts the enum
matches _SESSION_ACTIONS plus the by-name verbs. The cancellation harness gave a worker
thread 1 s to reach its transport, which a loaded CI runner missed twice this week; the
cancellation-latency assertion is unchanged, only the start-up wait is wider.
A child's background processes are killed at its teardown and their
notify_on_complete notices are suppressed in the parent, yet the child's
terminal result still said `notify_on_complete: true` and the parent's
delegation notice said nothing about processes left behind. Orchestrators
believed "CI watcher running" and waited on a completion that could never
arrive (recurring in the Sep 7 campaign sessions).
- process_manage(action="handoff", session_id, data="<purpose>"), children
only: process_registry.transfer_ownership flips owner_task_id/task_id/
session_key to the parent under the registry lock, so the completion is
stamped with the parent's owner at exit, passes the parent's sa- filter,
and is reaped by the parent, not the child. Cap 3 per child; an exited,
foreign, or non-child request is a tool error. The purpose rides the
event as handoff_note and renders in the parent's notice.
- Child terminal(background=True, notify=True) now returns
notify_on_complete=false plus a note: wait, kill, or hand off.
- _ChildRun.account_background_processes records handed_off_processes and
orphaned_processes on the result before cleanup kills the leftovers; the
parent's delegation block renders both.
`_on_tool_progress` bailed on the tool-progress gate before dispatching
`subagent.*`, so a Desktop/TUI user who hid tool-call chrome also lost the
subagent rows in the status stack and spawn tree. Subagent lifecycle is
application state (like `todo.updated`, clarify and MCP consent cards,
which already bypass the gate); the gate now applies only to the optional
progress chrome (reasoning previews, MoA rows, tool.generating).
The core-tool rename shipped `todo_list` on the wire (legacy alias `todo`
kept for old transcripts), but the Desktop renderer still matched the tool
by the literal `todo` in seven places: the live tool.start/tool.complete
mirror into the composer status stack, the todo-stream router, args
carry-over, the transcript hoist, the silent-tool class, the count noun,
and stored-history hydration. Every live task update therefore went into
the transcript as an ordinary tool row while the task panel stayed empty,
and reopening a chat never restored a finished list.
One predicate (`isTodoToolName`) now owns the wire/legacy name pair and
every site reads it.
Five test files and five evals scripts merged since this branch's base still import the
moved names via gateway.platforms.base. They work (base.py imports the names for its own
use) but the PR's invariant is that in-tree code imports from the defining module.
Same cleanup as the previous commit, applied to the third file that carried the
TYPE_CHECKING + in-function import pair. gateway.run imports none of the telegram/discord/
slack modules the conftest stubs, so import order is not a concern. test_feishu.py keeps its
TYPE_CHECKING import on purpose (FeishuAdapter is gated on optional lark_oapi).
Regression guard for the one-line import fix in commit 1; identical to the test in #102952,
whose /btw half is superseded by this PR. Red on origin/main (NameError), green here.
Self-review follow-ups on the F821 sweep:
- plugins/platforms/sms/adapter.py: the optional-import block now sets AIOHTTP_AVAILABLE like
the homeassistant / webhook / whatsapp_cloud adapters, and both call sites test the flag.
Removes the `if not aiohttp is not None:` double negation left by inlining
`_aiohttp_available()`.
- tools/mcp_tool_sampling.py: `call_context` defaults to `lambda: None` so the use site is a
single call instead of an Optional guard; the only None caller was a test. The
`from __future__ import annotations` was noise (`Context` is a runtime import). Comment
names the actual cycle (mcp_tool_server_run imports this module).
- gateway/platforms/helpers.py: drop the `from __future__ import annotations` — the only
MessageEvent annotations are attribute-target locals, which are never evaluated.
- tests/gateway/test_telegram_audio_vs_voice.py, test_video_context_note.py: module-level
`from gateway.run import GatewayRunner` like the ~100 sibling files; the
TYPE_CHECKING block + `# type: ignore[name-defined]` were contradicting each other.
(tests/e2e/conftest.py and test_feishu.py keep TYPE_CHECKING deliberately: they stub
telegram/discord before importing, and FeishuAdapter is gated on optional lark_oapi.)
Mutation check: neutralising the thunk read (`captured = None`) fails
test_captured_context_is_replayed_in_consent_call; restored → 14/14 green. ty on the three
touched production files vs origin/main: 0 new, 6 resolved.
Breaks the two import cycles that forced Protocol stand-ins in the F821 sweep, so the two
sites now name the real types.
gateway/platforms/event.py (new leaf): MessageType, ProcessingOutcome, MessageEvent moved
out of base.py verbatim. Their only dependency is gateway.session.SessionSource; base.py
imported helpers.py at module level, so helpers could not name MessageEvent. Now
TextBatchAggregator is typed by the real MessageEvent. 249 importers repointed
(`from gateway.platforms.base import` -> `.event`, preserving each import's layout);
gateway.platforms.__init__ re-exports from .event. The three revert-scheduled PLUGIN-COMPAT
pointers that named these symbols (gateway.slash_commands → MessageType, dingtalk → MessageType,
photon → ProcessingOutcome) and their COMPAT_MANIFEST rows now target gateway.platforms.event.
Docs updated: ADDING_A_PLATFORM.md, adding-platform-adapters.md (en + zh-Hans).
tools/mcp_tool_sampling.py: ElicitationHandler no longer holds a back-reference to its
MCPServerTask (mcp_tool imports sampling, so the task type cannot be named there). It only
ever read owner._pending_call_context, so it takes `call_context: Callable[[], Context | None]`
and MCPServerTask passes `lambda: self._pending_call_context`. The consent call is one
`functools.partial`, run directly or inside the captured Context.
ty on the 11 touched production files vs origin/main: 0 new diagnostics, 14 resolved.
(The one `source: SessionSource = None` diagnostic moves with the class; typing it Optional
exposes ~60 unguarded call sites — separate follow-up.)
Tests: tests/gateway + tests/plugins + tests/tools + touched files, 18,235 passed; the 31
failures reproduce identically on origin/main (macOS /private/tmp, systemd socket,
long-path fixtures, live-service tests).
Sweep of `ruff check . --select F821 --target-version py311`: 2,234 hits. 2,201 are left
alone on purpose: tui_gateway (2,169; bind_module rebinds bodies onto server.py globals,
all names verified to resolve there), the Feishu adapter (27; globals().update() SDK
binding) and the godmode script (5; dead standalone script). The other 33 were all
genuine defects. No lint config change; no TYPE_CHECKING escape hatches — every
annotation names a real, imported type; ty on the touched files: 0 new diagnostics.
- gateway/slash_commands.py: HISTORY_UNREADABLE never imported after #102117
→ NameError on the /btw error branch (same one-liner as #102952).
- gateway/platforms/whatsapp_common.py: `-> Path` return annotation with no Path
import (the body uses `_Path`). Never raised at runtime thanks to
`from __future__ import annotations`, but `typing.get_type_hints()` and ty
both fail on it.
- gateway/run.py: ActivityProvenance imported at module level
(agent.session_activity has no gateway deps); stringly annotation and the
lazy in-function import are gone.
- tools/patch_parser.py: PatchResult imported at module level; real return
annotation. The "avoid circular import" lazy import guarded a cycle that
does not exist (file_operations_common never imports patch_parser).
- gateway/platforms/helpers.py: base.py imports helpers at module level, so
MessageEvent cannot be named here; TextBatchAggregator only reads .text and
.source, so it is typed by a BatchableEvent Protocol that MessageEvent
satisfies structurally.
- tools/mcp_tool_sampling.py: mcp_tool imports this module, so MCPServerTask
cannot be named here; ElicitationHandler only reads
owner._pending_call_context, typed by an ElicitationOwner Protocol.
- plugins/platforms/sms/adapter.py: aiohttp is an optional dep ([messaging] extra) →
module-level try/except ImportError binding `aiohttp = web = None`, the pattern the
homeassistant / webhook / whatsapp_cloud adapters already use. Retires three lazy
in-function imports and the `_aiohttp_available()` wrapper; `_handle_webhook` typed
`web.Request -> web.Response`.
- plugins/platforms/teams/summary_writer.py: plain module-level `import httpx` — httpx is a
hard core dependency (pyproject `httpx[socks]==0.28.1`), so the lazy import and the
"imported on every CLI start" docstring premise were both wrong (plugin discovery never
imports this module; it is reached only via the Teams adapter / meeting pipeline).
Tests:
- tests/hermes_cli/test_config.py: a test body orphaned by the wave-1 prune
(6b81590c55) sat inside the class as dead code with self/tmp_path unbound
— header restored, so the v11→12 custom_providers migration is covered.
- tests/tools/test_mcp_tool.py: @staticmethod recursing on `self` in the
win32 branch; call portalocker directly.
- tests/test_background_review_list_shapes.py: main() still ran 3 pruned tests.
- tests/agent/test_cursor_optimizations_parity.py: bench() used names only
imported inside a sibling test.
- GatewayRunner / FeishuAdapter / Dict / Optional: missing imports.