Mock inference gains a one-shot `E2E_CALL(<tool>)[<json>]` script so a spec can make
the REAL tool run from a real Bot Chat; the new spec creates a sender bot, seeds a
`writer` profile titled "Scribe", sends a DM to "Scribe" and asserts the rendered
tool result says "Message dispatched to @writer" (base: "No teammate named
'Scribe'"), plus the sender's state.db tool row.
Three group-room turn-loop defects, each live-reproduced in the real
Electron app against a real gateway with mock inference:
- Hold classifier (group-rounds.ts): a stop/halt/pause token holds a member
only within two words of an @mention, so "@impl go, das ist halt ein Test"
is delivered instead of setting a sticky hold (#103893); addressing the
whole room (@all/@everyone) without a stop word releases every hold, so a
Stopped room wakes on "@all <task>" without the literal "resume" (#97740).
- Retained failure (group-turns.ts): the gateway keeps a failed turn under
session.resume.inflight as {status:'error'}; the room read that as live
work and slid the deadline to the 20-minute cap while the member looked
busy. groupSessionBusy/retainedGroupTurnError classify it as finished;
the foreground poll throws it into the failed-turn path (activity +
roster badge) and the harvester consumes the stranded marker (#92760,
diagnosis from #95103 by @hrnbld).
- Approval card (group-chat-parts.tsx): approval choices submit on click;
the clip-prone footer Respond button is gone for approvals (#91706).
Room message code blocks wrap instead of overflowing (#91857 by
@piskooooo).
Tests: invariant vitest cases in group-rounds/group-turns/group-chat-parts
(all red on base); three Electron specs under apps/desktop/e2e using two
new mock-inference triggers (gated rm -rf for a real approval prompt, a
non-retryable 401 for a retained failure). Docs: bot-mode.md § Groups.
Co-authored-by: hrnbld <hrnbld@users.noreply.github.com>
Co-authored-by: piskooooo <piskooooo@users.noreply.github.com>
Gate reasoning_effort and fast behind the same switch as model/provider in
desktopSessionCreateParams (includeComposerSelection): a Bot-workspace tile
targets a different profile without switching the window's composer, so all
four fields of an unrelated session's pick would otherwise ride into its
session.create. Omitting them lets the bot profile's configured defaults apply.
The vitest added for this PR left the hoisted requestGatewayForAgent mock
with a recorded call (restoreAllMocks only restores spies), which failed the
next test's not-called assertion in CI ("keeps an unlisted named local
legacy-profile tile owned by its bare profile"). The test now resets that mock
and the composer atoms it set, and asserts the two new omitted fields.
New Playwright spec bot-tile-ignores-ambient-composer-model.spec.ts drives the
real app: open a bot's canonical chat, pick a second model from the composer
menu (the mock provider now lists extraModels; receivedModels records the
model of every completion request), Ctrl+T a side chat, send a turn, and
assert the inference request carried the profile default. Fails on main's
index.ts (request carried the ambient pick), passes on head.
Keep failed-member exclusion across the room queue, rather than resetting
it per pending thread. A new user action after failure still permits a new
attempt. Share the drain activity epoch so a skipped queued thread cannot
hide the preceding member failure.
Proven red in real Electron: hold transport refusal, enqueue same-thread
and cross-thread sends, then release; old head submits three times, fixed
head once. Strengthen follow-up evidence with distinct provider replies,
exact public log order/count, and per-input inference counts.
Real-Electron Playwright spec for #100406: a two-member room (primary
profile + code-farmer). The user addresses only @code-farmer; its
scripted reply @mentions hermes; the assertion is a `default`-authored
"B" entry in the persisted room log. On origin/main the room settles
after Code Farmer's line and the spec fails at that assertion; with the
mention-alias fix it passes.
The mock inference server gains a per-speaker script for group rooms:
`E2E_SAY(<handle>)[<line>]` tokens in the user's send answer the member
whose turn prompt opens with `You are @<handle>`; unscripted members
reply "(pass)". `{at}` stands for `@` so the script itself never
mentions anyone and round one only drives the member the user tagged.
Same site class as the e2e fixtures: tests-js/scripts/mock-server.ts writes
the config for `npm run dev:mock`. Entry-level providers.<name>.context_length
is honoured since #98387, so a 4096 window now fails agent init below
MINIMUM_CONTEXT_LENGTH exactly as it did for the Playwright fixtures.
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
one-line compat forwarder) and the new implementation already drains
both pipes asynchronously with bounded abandonment, which supersedes
this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
addition inside the consolidated file; rewired the five upstream specs
still importing './mock-server' to the consolidated path (export sets
verified identical) and dropped the superseded apps/desktop/e2e copy.
The dev:mock script duplicated the e2e mock server. The copy had only
the plain chat reply; every scripted path lived only in the e2e version.
A single mock server now lives in tests-js/scripts/mock-server.ts.
The e2e suite imports it as a library. Running the file directly
starts the server, writes a mock config, and launches the desktop app.
The dev:mock script now runs that file.
The e2e tsconfig lists tests-js/scripts in its include, because the
composite project rule requires every imported file to be listed.