f3ccab1f6067d188cdffb35b7a6ad46292a2ca2f
15118 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
f3ccab1f60 | test: define passive one-line subagent dock contract | ||
|
|
fb5715c19f | fix: preserve subagent controls across attached session peers | ||
|
|
a0fa236ed4 | fix: retain worker tool activity in narrow classic monitor | ||
|
|
8a23f14df3 | fix: keep classic subagent dock compact and skin-colored | ||
|
|
7befa11bf2 | fix: retain subagent control after live session reattachment | ||
|
|
924c5ded2e | fix: scope subagent stops and publish authoritative live progress | ||
|
|
d99a63b645 | fix: yield subagent monitor to incoming CLI prompts | ||
|
|
4cd4f395ea | test: isolate missing first-party import guard fixture | ||
|
|
8b01df963d |
feat: expose session-scoped subagent roster and live tail RPCs
Project live children and background units for shared TUI/Desktop consumers; read a bounded tail from the existing runtime transcript. Reuse existing queued steering semantics rather than claiming delivery. Projection adapted from PR #70899 with exact live session and transport ownership. Co-authored-by: joaomarcos <joaomarcosdias444@gmail.com> |
||
|
|
4c8d380434 | test(cli): synchronize monitor inputs on rendered frames | ||
|
|
8842d4804d | fix(cli): keep monitor controls pinned and resize chrome isolated | ||
|
|
84100efae3 | feat(cli): show live dock and full-screen subagent controls above composer | ||
|
|
6f68394eac | feat(cli): add scoped bounded subagent dock state and controls | ||
|
|
866332bfb5 |
fix(relay): authorize send_message targets and surface egress declines (P5) (#99220)
* fix(relay): authorize send_message targets and surface egress declines
P5 of the relay egress-authorization workstream. The relay path
authenticated the SENDER but never authorized the DESTINATION, and the
gateway compounded it from both ends.
(a) send_message could silently name an arbitrary relay target. Its
`target` parameter is free-form ('platform:chat_id'), so a model could
name ANY chat id and the gateway would emit an outbound frame for it.
gateway/relay/egress.py adds an attestation floor: a relay-routed
destination must have a provenance this gateway can show -- the
operator's home channel, the channel directory, or its own gateway
session origins. Anything else is refused HERE, with a visible tool
error naming the target, before a frame is written. Non-relay platforms
and platforms served by a live native adapter in this process are
untouched (same precedence resolve_delivery_transport applies).
(b) Connector declines were swallowed into apparent successes. The
connector's egress floor answers an unauthorized destination with a
DEFINITE failure whose text is deliberately uniform (F-005). Several
relay lanes degrade a *transport drop* by design and were degrading an
*authorization refusal* the same way:
- _send_media returned None, sending the caller into
BasePlatformAdapter's text fallback -- a DIFFERENT op re-addressed at
the very chat the connector had just refused.
- _send_prompt returned None, so exec-approval / slash-confirm /
clarify reported "relay prompt op unavailable" (a wrong reason) and
ran their numbered-text fallbacks into the refused chat.
- task_card_stop discarded the error entirely.
- typing / delete / react / thread ops degraded silently at debug.
is_egress_decline() classifies THAT a decline happened (never why --
the uniform text is not parsed for reasons) and requires a definite,
non-ambiguous failure, so a lost-ack retry is still a transport
outcome. Lanes with an error-carrying contract now report the decline
verbatim; cosmetic bool/None lanes still degrade but log it at WARNING.
Advisory progress drops that legitimately degrade are unchanged: the
task_card send lane, the draft ambiguous/except branches, and every
transport-exception path keep their existing fail-open behaviour.
Tests: 21 mutations of the production source, all KILLED.
* fix(relay): authorize the RESOLVED target; declines must not fall back
Review round 1 (independently confirmed by a second reviewer) found three
blockers. Two are fixed here; the third (B-2, Telegram @username) is a policy
decision left open deliberately.
B-1 — THE FIX CAUSED THE OUTAGE IT PREVENTED (tools/send_message_tool.py)
The P5(a) guard ran ABOVE Slack user->DM resolution, so it authorized the
internal pseudo-id `_parse_target_ref` emits (`user_name:ben`, `user:U...`).
Provenances only ever hold RESOLVED conversation ids, so a fully attested DM
was compared as a handle against a set of `D...` ids and refused:
base slack:@ben SENT head(before) slack:@ben REFUSED
Every Slack DM by handle was broken. Moved the guard below resolution; it now
authorizes the destination that is actually sent to, and the refusal names the
resolved id. Position is load-bearing, so it is commented as such and pinned:
reverting the move turns exactly the four new cases red.
B-3 — A DECLINE IS NOT A LANE FAILURE (gateway/run.py)
`_approval_send_outcome` had only sent/failed/ambiguous, so a connector
decline collapsed into `failed` — which is the cue to run the plain-text
fallback into the chat the connector had just refused. The adapter fix in the
previous commit improved the error STRING while user-visible behaviour stayed
identical to base; the commit message overstated it. Fixed properly:
- new `declined` verdict, recognised via the shared `is_egress_decline`
contract (not string sniffing at the call site)
- exec-approval returns without the text fallback
- slash-confirm suppresses the text reply AND clears the registration, so a
card that never rendered cannot capture the user's next message
`send_clarify` was already correct (returns early inside the adapter).
MUTATIONS (production source; both directions)
classifier never returns 'declined' -> KILLED (4 cases)
ALL failures classified as 'declined' -> KILLED (2 cases)
guard moved back above Slack resolution -> KILLED (4 cases)
decline CODE changed (review M05) -> KILLED
marker match made case-sensitive (M10) -> KILLED
M05 was a tautology: the test asserted the imported constant against itself,
so changing the constant could not fail it. The wire contract is now pinned as
a literal, because the connector stamps that exact string and a one-sided
change is a silent cross-repo break.
REGRESSION CHECK: the 12 failures + 1 collection error in this test selection
are PRE-EXISTING cross-test contamination — the identical set fails at
|
||
|
|
fef0e16fe1 |
fix(relay): re-dial once with a fresh token before treating a 4401 as revocation (#102602)
* fix(relay): re-dial once with a fresh token before treating a 4401 as revocation A 4401 close after a successful handshake was read unconditionally as the connector having revoked this gateway's per-gateway secret (opt-out), so the transport latched auth_revoked, the adapter went relay_disabled, and all messaging stopped until a manual restart. But the connector sends the same plain 4401 'unauthorized' for an EXPIRED upgrade token (make_upgrade_token TTL is 300s). Scale-to-zero makes that routine: the instance is suspended while a re-dial is in flight, the token was minted before the freeze, the dial completes on resume with a token past its TTL, the connector refuses it with 4401, and the gateway misreads an expired token as a revoked credential. Production incident 2026-09-02 - the connector DB secret was never revoked. Fix, in gateway/relay/ws_transport.py: - The first post-handshake 4401 is provisional. The reader schedules ONE immediate re-dial (_redial_with_fresh_token) that bypasses the reconnect backoff; _dial_and_start mints a fresh token on every call. The backoff supervisor design is untouched. - Only a 4401 against that fresh token (either refused at the upgrade or closed after a descriptor on that connection, tracked by dial generation) latches auth_revoked. Terminal behaviour is otherwise unchanged. - If the fresh dial fails for a non-auth reason, hand off to the normal backoff supervisor. - Read the Close frame reason (_close_reason_of, sibling of _close_code_of). A 4401 whose reason is exactly 'expired' never latches revocation and takes the normal reconnect path - forward-compatible hook for the connector change landing separately. - disconnect() cancels the one-shot retry task alongside the supervisor. Tests (tests/gateway/relay/test_ws_transport.py, real websockets server): - 4401 once, next dial accepted -> reconnected, auth_revoked False, exactly 2 dials. - 4401 on every dial -> auth_revoked True after exactly one retry (2 dials), no supervisor, no further dials. - 4401 reason 'expired' (once and repeated) -> never latched, reconnects via the normal supervisor. - Existing 7d-B tests keep passing (4401 before any handshake stays retryable; the revoking-every-dial stub still latches). * fix(relay): tolerate a partially built transport in disconnect() teardown Three teardown tests construct WebSocketRelayTransport via object.__new__ without __init__, so the new _auth_retry attribute was absent and disconnect() raised AttributeError. Read it with getattr like the other optional teardown handles. * fix(relay): one live dialer — the fresh-token retry is a flag on the next dial, not a second dialer Review (round 1) found a race: a reader that dies with a provisional 4401 while the backoff supervisor is already mid-dial started _redial_with_fresh_token as a SECOND concurrent dialer; both installed sockets/readers and the supervisor could overwrite the accepted retry socket (3 dials, wrong socket). Now a provisional 4401 sets _auth_retry_pending; _dial_and_start consumes it and stamps _auth_retry_generation on whichever dialer performs the next dial. The reader arms a dialer only when none is live (_dialer_running), and an upgrade-time 4401 on that generation latches revocation from either dialer (_latch_if_fresh_token_refused). The reader also iterates its captured ws handle, not self._ws. Regression test reproduces the race with a deterministic fake connect; it fails with the _dialer_running guard removed. * fix(relay): a dial whose reader died mid-hello is a failed dial; the retry marker survives network failures Review round 2: - BLOCKER: with one-live-dialer, a reader that dies while its own dialer is still inside _dial_and_start (hello in flight) arms nothing — that is the dialer's job — but the dialer then returned 'connected', leaving no socket, no reader, no dialer. _dial_and_start now raises ConnectionError when the reader it installed has already finished, so both dialers take their normal failure path (retry -> supervisor; supervisor -> backoff). - MAJOR: _auth_retry_pending was consumed before websockets.connect, so a connect-time network failure un-marked the retry and the NEXT dial's real fresh-token 4401 read as another first strike (revocation never latched). The marker is now consumed only when the token reaches an auth outcome: upgrade accepted (stamp the generation) or upgrade 4401'd (judge it). Two regression tests, each mutation-checked red against its own guard. |
||
|
|
d095f8fb15 | fix(tui): bound subscriber delivery without blocking shared turns | ||
|
|
8b2ef359d1 |
feat(tui): discover cooperative local session owners
Fence owner discovery by profile, lease and loopback endpoint, and hand the authenticated URL to the existing Ink transport without acquiring a competing lease. Distinguish lease age from turn activity in unsupported-owner recovery. Client slice only: requires the integration runtime to advertise shared_runtime_url and provide the session-attach handshake. Classic CLI attachment remains an integration gap. |
||
|
|
b3014688f9 | fix: isolate shared session membership and prune departed viewers | ||
|
|
4d32503f86 |
fix(tui_gateway): skip a dead peer's transport when draining its queued prompt
_drain_queued_prompt attaches the transport pinned to the queued envelope, but the client that queued the prompt may have disconnected while the prompt sat in the queue. Attaching it then pins a dead peer into the session's fan-out, where it costs one failed write before the fan-out prunes it. That is not a regression — the line this replaced rebound the whole slot to that same dead transport, which cost the session for the entire drained turn rather than one write — but the guard is one condition, so it goes in. _transport_is_dead is the predicate already used by the reaper and the orphan check: the parked drop sentinel, or a transport carrying _closed. Drain semantics are otherwise unchanged. The prompt still runs, and the clients already attached keep their stream; only the dead peer's attachment is skipped. tests/tui_gateway/test_multi_client_fanout.py pins that the drained prompt still dispatches while the disconnected queuer stays out of the slot, alongside the existing case where a live queuer is attached and the single-client control where the queuer takes an unoccupied slot. Raised by the automated review on #86784. |
||
|
|
68ed3ffd10 |
fix(tui_gateway): check steer authority by transport membership under fan-out
subagent.steer resolves authority by comparing the request's context-bound transport with the session's transport slot. Once a session mirrors to more than one client that slot holds a FanoutTransport, so the comparison fails for every client, the peer that commissioned the subagent included, and every steer is rejected. Fan-out without this check ships that regression, and no existing test catches it because the suite only exercises single-client sessions. Authority now asks whether the request's transport is attached to the session, directly or through the fan-out. The single-client case is unchanged: a bare slot still compares by identity. This widens authority. Any client attached to a mirrored session can steer that session's subagents, not only the peer that commissioned them. Narrowing it back to the commissioning peer requires recording that peer per subagent, which this change does not do. tests/tui_gateway/test_multi_client_fanout.py pins the commissioning peer's authority inside a fan-out, pins the widened case, and keeps a single-client control in which an unattached client is still refused. The four browser.controller.* handlers gate on the session transport slot exactly as subagent.steer did, through one shared gate in the _controller_method decorator, and the conversion missed them. Once a session mirrors to more than one client the slot holds a FanoutTransport, which is identical to no peer's WSTransport, so browser.controller.register, .result, .heartbeat and .detach all answer "session is not owned by this transport" for every client, the peer that registered the controller included. Browser control is therefore unusable on any mirrored session. No existing test catches it because the browser-control suite only exercises single-client sessions. All four now ask the same question steer asks: is the request's transport attached to this session, directly or through the fan-out. The single-client case is unchanged, because a bare slot still compares by identity. Only registration widens. On .result, .heartbeat and .detach the broker's is_owner check sits below the session gate and compares controller.owner is owner against the transport recorded at attach time (gateway/browser_control_broker.py), so a mirrored peer that did not register the controller is still refused there, now with "controller is not owned by this transport" instead of the session message. browser.controller.register has no such check, so any client attached to a mirrored session may register a controller for it; the broker's principal lane keeps that inside one authenticated identity, and a second identity in the same lane hard-replaces the first. tests/tui_gateway/test_multi_client_fanout.py pins the registering peer's access inside a fan-out, pins the widened and broker-refused cases, and keeps a single-client control in which an unattached client is still refused on all three scope-gated handlers. |
||
|
|
de25545dce |
fix(tui_gateway): fan session events out instead of rebinding the transport slot
A session held exactly one transport, and prompt.submit, session.resume, session.activate, and the queued-prompt drain all rebound that slot. A second client therefore took the stream away from the first: the earlier client stopped receiving the turn it was already rendering, and either client disconnecting parked the whole session on the drop sentinel. FanoutTransport goes in the same slot and satisfies the same Transport protocol, so write_json and every other reader of the slot are unchanged. It delivers each frame to a snapshot of its peers, concurrently when more than one peer is attached and the caller is not on an event loop, and prunes any peer that returns False or raises. A dead client is dropped; a slow one costs the emitter at most one write timeout per frame rather than one per peer. Request/response RPCs are unaffected: they still answer on the request's context-bound transport, so a client only ever sees replies to its own calls. The rebind sites become attach sites through _attach_session_transport, whose ladder keeps the single-client shape identical. The same object already in the slot is a no-op; an empty, stdio, or parked slot is taken outright; only the arrival of a second live client wraps both. The queued-prompt drain is included because it pinned the drained turn to the queuer and silenced everyone else. A non-peer newcomer such as stdio or the drop sentinel never displaces a live client, so an activate dispatched without a bound websocket cannot silence the socket that owns the session. Disconnect detaches first. A session that retains another client keeps streaming and is neither parked nor reaped, and only the clientless ones follow the existing close_on_disconnect and park-sentinel path, so a single-client disconnect, the orphan reaper, and its grace window behave as before. _ws_session_is_orphaned is unchanged: it still asks whether the drop sentinel is in the slot, and a fan-out is never the sentinel, so a session that still has a peer is never reported as orphaned. Attach performs no entitlement check: any authenticated peer may mirror any session. The fan-out architecture follows the approach in #40822 by @OmarB97. What the slot's later history forces. _close_sessions_for_transport drops the #83716 rebind-to-the-most-recent-surviving-viewer, which fan-out membership subsumes — a pop-out window is a peer, so a session that still shows in one is never returned as clientless — and keeps the #77129 revalidation before parking, now expressed as a liveness check under _session_transport_lock so it is race-free against attach and detach. _transport_is_live_peer defers its last answer to _transport_is_dead: a socket that already latched _closed is a departed client, and admitting it would keep a session out of both the park and the reap. _transport_is_dead also learns the fan-out: a FanoutTransport with no live peer is dead, so a session whose peers were all pruned by failed writes cannot outlive the TTL and LRU reapers. The upstream test that pinned the #83716 rebind, test_close_transport_rebinds_session_to_remaining_viewer, is re-expressed in fan-out terms: both windows attached, the pop-out closes, the session stays with the main window unparked and still receiving frames. |
||
|
|
496fa5c9d1 | test: reproduce shared session terminal event stealing | ||
|
|
520e63661c | fix: keep command-auth model discovery lazy across config and setup | ||
|
|
c111ede3e5 |
fix(picker): resolve key_cmd credentials for model discovery
`key_cmd` (#86891) authenticates a provider with a SHORT-LIVED bearer minted by a command — SSO/OIDC brokers, cloud IAM, internal auth proxies. The request path has honoured it since it landed, but the picker resolved probe credentials from `api_key`/`key_env` ONLY, so a key_cmd provider probed `/v1/models` with an EMPTY key. Against an authenticated endpoint the probe 401s, discovery returns nothing, and the provider falls back to its single configured default model. The picker shows ONE model, indistinguishable from an endpoint that genuinely serves one — while inference keeps working, because that path mints correctly. Reproduced against a LiteLLM gateway behind Entra OIDC: 0 models discovered with an empty key, 26 with the minted token. Both picker probe sites already funnel through `_entry_credentials()`, so the fix lands in one place: it now reports a `cmd:<key_cmd>` identity, and each site falls back to `resolve_probe_token()` after api_key/key_env. An explicit static key still wins, so existing configs are unaffected. The identity is keyed on the COMMAND, never the minted token: the token rotates on every refresh, so keying on its value would change the group fingerprint constantly and force a re-probe on every open. Two entries on one URL with different helpers still get distinct rows. `resolve_probe_token()` lives in agent.command_token_source, which already owns key_cmd minting, and shares the CommandTokenSource cache with the request path — a cache read, not a fresh sign-in. Fail-closed: a helper needing an interactive sign-in degrades to today's empty-key behaviour rather than taking down every other provider's row. `_model_flow_named_custom` (the `hermes model` setup flow) is the sibling path — it builds its own `Authorization: Bearer` from the same incomplete resolution — and is fixed the same way, with one ordering constraint: the value persisted to config.yaml is computed BEFORE the mint, so a short-lived bearer can never be written back to shadow the key_cmd meant to re-mint it. Tests drive the real code paths and assert on the credential each probe receives rather than on function source, so a semantics-preserving refactor does not fail them. Verified they fail with the fix reverted. |
||
|
|
0d6106eab8 | test: consolidate Gemini array regressions into two invariants | ||
|
|
bbcf1ee180 |
fix: preserve native Gemini union constraints
Complete the type-array normalization salvaged from #55643: stringify mixed union enum metadata, preserve existing anyOf constraints, and keep array items and object properties/required on the corresponding typed branches. Exercise real native request serialization over loopback and Google SDK validation with a scalar control; no live Google credentials were available. |
||
|
|
6a04ea67c0 |
fix(gemini): collapse array-typed tool schemas instead of crashing translation
The enum-compatibility check evaluated `[...] in {...}` on an array `type`,
raising TypeError: unhashable type: 'list' and aborting translation of the whole
tool catalog rather than the one offending tool.
Addresses both review points:
- Reuses tools.schema_sanitizer._normalize_type_array instead of picking the
first non-null member, so a real union becomes an anyOf of single-type
branches and no branch is dropped.
- Derives the type outside the key loop and sets nullable after it, so the
flag implied by "null" in the array beats an input nullable: false whichever
key the producer emitted first. Both orders are pinned by a parametrized test.
|
||
|
|
5280fe9987 |
fix: cron and local DMs reach an open Desktop Bot Chat
Route local producers to durable owner ingress before attempting the unowned CLI lane. Preserve per-run/per-message IDs and receipt-first retry handling; never fall back after ambiguous admission. Report cron admission as queued, not completed or failed, in job status, the execution ledger and CLI/tool UX. Native isolated Electron validation reproduces SESSION_NOT_OWNED on main for both idle and busy owners. Fixed owner consumes idle cron, busy cron, local DM and mounted-chat cron exactly once, keeps its lease, yields to queued human input, and preserves the prior model-request prefix and tool schema. Inference alone used a deterministic loopback wire stub; no paid model call. |
||
|
|
db96c8ced7 |
fix: admit Bot Chat deliveries through the live session owner
Adapt FalconOrtiz's owner-mailbox proposal from #101564 onto the current notification poller and topical modules. The durable mailbox is cross-process ingress only: the existing owner admits its normal prompt turn after the current turn and human FIFO clear. Retain immutable receipts, stable admission identities, capability and lease fencing, compression lineage, and disable blind recovery replay of imported turns. The original stale server hunks and expiring receipt protocol were rebuilt rather than cherry-picked: the current facade decomposition and durable busy admission contract differ. Credit the earlier owner-mailbox work in #100544 and durable producer work in #100319. Co-authored-by: fangliquanflq <fangliquan@qq.com> Co-authored-by: 686f6c61 <github@00b.tech> |
||
|
|
6178e9f4ee | fix(approvals): honor GNU env split escapes and argv0 operands | ||
|
|
50617d1c75 | fix(approvals): preserve env argv and shell comment boundaries | ||
|
|
58faa10134 |
fix(approvals): match denied executable paths behind shell prefixes
Adapt the command-position, bounded-candidate and launcher-option work from embwl0x's #76063 to the current detection owner, then add executable basename projection from Rohith Pariki's #104338. Parse raw quote state before applying existing text normalization so quoted arguments do not become commands. Cover shell payloads and literal env split-string carriers, retain path-specific rules and whole-command globs, and document the supported normalization rather than claiming an OS capability sandbox. Related: #104308, #76037, #76063, #104338, #78521, #86711. No automatic closing directives: the older carriers also contain broader case syntax and git-option work not included here. Co-authored-by: embwl0x <embwl0x@users.noreply.github.com> Co-authored-by: Rohith Pariki <rohithpariki@gmail.com> |
||
|
|
4810074d73 | fix: retain completions until explicit adapter admission | ||
|
|
561897db59 | fix: keep admitted heartbeats in their owning conversation | ||
|
|
bf1bf7515a | fix: retry Kanban wakes until adapter admission | ||
|
|
3b7ff435fd |
fix(kanban): preserve durable origins for worker-created tasks
Carry the owning task's notification subscriptions independently of dependency edges, within the creation transaction. Prefer its durable session over worker and request-local sessions while preserving explicit overrides. Cover worker CLI create and built-in decomposition, and retain conversation route anchors. Auto-subscribe no longer upgrades an inherited passive subscription. Slim adaptation of Christopher-Schulze's session-precedence fix in #85687, expanded to durable subscription provenance and sibling creation paths. Related: #85575, #85687 Validation: strict RED/GREEN (7 failing cases before; 7 passing after), then 58 Kanban test files: 383 passed, 2 skipped. Real dispatcher-spawn subprocess probe covers direct, linked, unlinked, explicit-session, worker CLI, built-in children and a plain CLI negative control, with recording transport only. Co-authored-by: Christopher <210261288+Christopher-Schulze@users.noreply.github.com> |
||
|
|
a247012dc1 |
fix(gateway): retain conversation route anchors for Kanban callbacks
Carry canonical scope_id and parent_chat_id through session context and reply metadata. Stamp slash-created subscriptions with the routed source profile, not the notifier process profile. The gateway facade only wires the topical metadata helper and existing context binding. Two real gateway/DB invariants reproduced missing anchors and wrong owner before implementation. Related: #101196 and #101397. Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com> |
||
|
|
30b3ca16f4 |
fix(kanban): deliver routed profile notifications on the authorized transport
Authorize route-only profiles using the ordered canonical route matcher and served-profile set at both claim and delivery. Preserve secondary credential boundaries, retry denied routes, and keep scope/parent anchors plus transport provenance on synthetic wakes. Install the destination runtime scope rather than inheriting the notifier's scope; a removed profile cannot wake as primary. Slim forward-port of the direction in #101196/#101397 and #93863 (#93851). Canonical scope_id takes precedence over the guild alias and live chat cache. Two route invariants and two runtime-scope invariants reproduced red first. Co-authored-by: Brooklyn Nicholson <brooklyn.bb.nicholson@gmail.com> Co-authored-by: liuhao1024 <sunsky.lau@gmail.com> |
||
|
|
9ee3d129b4 |
fix: wake idle process watches on their owning profile transport
Salvage #82561's idle watch consumer through the existing drain, retaining notify-off suppression and retrying failed injections. Preserve #102647's secondary-profile ownership for reconstructed sources, preflight, synthetic completion turns, and raw process progress/final notices. Unavailable transports defer rather than borrowing the primary bot or losing the event. Keep retained transport provenance and alias-aware relay resolution. Add three invariant tests and an offline real-process/Discord lifecycle eval. Co-authored-by: trwpang <134544126+trwpang@users.noreply.github.com> Co-authored-by: holny <holny@foxmail.com> |
||
|
|
67161511d1 |
fix: retain async completions while delivery owners are unavailable
Preflight source-first transport resolution and parent/API database readiness before spending any durable delivery attempt, including every batch sibling. Unavailable targets remain retryable; terminal targets retain their drop semantics and stateless API completions remain durable display rows, not turns. Minimal combined salvage of #82704 and #96086; adjacent eligibility, profile adapter selection and idle watch-consumer work remain separate. Co-authored-by: Screaming Sun <5552699+eventh0riz0n@users.noreply.github.com> Co-authored-by: fangliquanflq <fangliquan@qq.com> |
||
|
|
d6867ab459 |
fix: admit trusted busy wakes without changing queued human input
Salvage the busy-auth ordering fix from #95602. Gateway-owned wakes have no external identity; keep external and plugin authentication unchanged. Preserve pending human text by appending wakes through the existing FIFO instead of base-adapter text merging. Co-authored-by: Robert Montelongo <ramontelongo22@gmail.com> |
||
|
|
59eb509f4c |
fix: refund heartbeat admissions that never enter agent execution
Bind settlement to the exact adapter task and event. Rejected or cancelled preparation refunds the existing claim; cancellation after the agent runner starts remains counted. Keep profile-scoped callback context and the manager replacement guard. A fire count is not outbound delivery proof. Drop departed routes rather than executing their stale schedule. Slim accounting-invariant salvage of #93174; preserve current direct adapter dispatch instead of reviving its FIFO/inflight implementation. Prior art #92858. Co-authored-by: Finn763 <165816600+Finn763@users.noreply.github.com> Co-authored-by: fangliquanflq <fangliquan@qq.com> |
||
|
|
27fa0b7eef |
fix: keep heartbeat watches in their route and profile lifecycle
Retry startup restoration on every poll, including an empty registry. Resolve current session identity after compression and scope each watch independently, preserving direct idle adapter admission and busy coalescing. Slim redo of #80208, #80225 and #80273. Co-authored-by: Gr1mmJ4w <ahmetsonersancak@anadolu.edu.tr> Co-authored-by: Drexuxux <drexux0@gmail.com> Co-authored-by: pierrenode <298902573+pierrenode@users.noreply.github.com> |
||
|
|
be154517a8 |
fix: restore gateway heartbeat watches after restart
Recover active watches from current persisted session origins and exact route keys, reading heartbeat state off-loop in each source's profile. Failed scans leave watches intact for the poller's next retry. Start the heartbeat poller even when startup restores no watches. Slim synthesis of #92660, #98310 and #98313, with earlier restart recovery prior art from #92594. The integration poller calls restore_heartbeat_watches on every poll, including empty registries. Co-authored-by: chelsealong <chelsealong@126.com> Co-authored-by: liuhao1024 <sunsky.lau@gmail.com> Co-authored-by: fangliquanflq <fangliquan@qq.com> |
||
|
|
1ec22acca0 |
fix(gateway): stop departed session heartbeats at reset boundaries
Clear the departed schedule using the route-owned database when ending a conversation, so stale watches cannot fire after reset and resuming the archived session cannot resurrect the schedule. Compression migration already archives its parent and remains unchanged. Adjacent lifecycle reports: #80273 by pierrenode and #80208 by 0xGr1mm. This follow-up covers durable reset cleanup, not their poller changes. |
||
|
|
73ddf0672c | test: keep attempt-cap probe on provider-confirmed fallback | ||
|
|
e9313f6458 | fix: let provider evidence adjudicate past-window preflight estimates | ||
|
|
3114916ee4 |
fix(gateway): carry accepted-input ownership through persistence
Namespace delivery markers and assign fresh keyless turn identities instead of inferring ownership from IDs or process-local row baselines. Query only marker existence on the canonical live compression continuation and ancestors. Preserve raw reply IDs and exclude metadata from provider wire messages. Expand the two existing invariants with resumed cross-chat ID collisions, a real independent SQLite writer, reaped siblings, and archived-history allocation controls. All 20 full-handler checkpoints and 63 targeted tests pass. |
||
|
|
f120ea7149 | fix(gateway): exclude observations from accepted-turn ownership | ||
|
|
136d80d040 | fix(gateway): retain one durable owner for failed input turns |