Change-detectors, tautologies, source-reading tests, redundant duplicates,
mock-echo tests and dead/unrunnable tests. Per-test rationale in the lane
ledger (category + reason for every removal).
Mock inference gains a one-shot `E2E_CALL(<tool>)[<json>]` script so a spec can make
the REAL tool run from a real Bot Chat; the new spec creates a sender bot, seeds a
`writer` profile titled "Scribe", sends a DM to "Scribe" and asserts the rendered
tool result says "Message dispatched to @writer" (base: "No teammate named
'Scribe'"), plus the sender's state.db tool row.
Three group-room turn-loop defects, each live-reproduced in the real
Electron app against a real gateway with mock inference:
- Hold classifier (group-rounds.ts): a stop/halt/pause token holds a member
only within two words of an @mention, so "@impl go, das ist halt ein Test"
is delivered instead of setting a sticky hold (#103893); addressing the
whole room (@all/@everyone) without a stop word releases every hold, so a
Stopped room wakes on "@all <task>" without the literal "resume" (#97740).
- Retained failure (group-turns.ts): the gateway keeps a failed turn under
session.resume.inflight as {status:'error'}; the room read that as live
work and slid the deadline to the 20-minute cap while the member looked
busy. groupSessionBusy/retainedGroupTurnError classify it as finished;
the foreground poll throws it into the failed-turn path (activity +
roster badge) and the harvester consumes the stranded marker (#92760,
diagnosis from #95103 by @hrnbld).
- Approval card (group-chat-parts.tsx): approval choices submit on click;
the clip-prone footer Respond button is gone for approvals (#91706).
Room message code blocks wrap instead of overflowing (#91857 by
@piskooooo).
Tests: invariant vitest cases in group-rounds/group-turns/group-chat-parts
(all red on base); three Electron specs under apps/desktop/e2e using two
new mock-inference triggers (gated rm -rf for a real approval prompt, a
non-retryable 401 for a retained failure). Docs: bot-mode.md § Groups.
Co-authored-by: hrnbld <hrnbld@users.noreply.github.com>
Co-authored-by: piskooooo <piskooooo@users.noreply.github.com>
Gate reasoning_effort and fast behind the same switch as model/provider in
desktopSessionCreateParams (includeComposerSelection): a Bot-workspace tile
targets a different profile without switching the window's composer, so all
four fields of an unrelated session's pick would otherwise ride into its
session.create. Omitting them lets the bot profile's configured defaults apply.
The vitest added for this PR left the hoisted requestGatewayForAgent mock
with a recorded call (restoreAllMocks only restores spies), which failed the
next test's not-called assertion in CI ("keeps an unlisted named local
legacy-profile tile owned by its bare profile"). The test now resets that mock
and the composer atoms it set, and asserts the two new omitted fields.
New Playwright spec bot-tile-ignores-ambient-composer-model.spec.ts drives the
real app: open a bot's canonical chat, pick a second model from the composer
menu (the mock provider now lists extraModels; receivedModels records the
model of every completion request), Ctrl+T a side chat, send a turn, and
assert the inference request carried the profile default. Fails on main's
index.ts (request carried the ambient pick), passes on head.
Keep failed-member exclusion across the room queue, rather than resetting
it per pending thread. A new user action after failure still permits a new
attempt. Share the drain activity epoch so a skipped queued thread cannot
hide the preceding member failure.
Proven red in real Electron: hold transport refusal, enqueue same-thread
and cross-thread sends, then release; old head submits three times, fixed
head once. Strengthen follow-up evidence with distinct provider replies,
exact public log order/count, and per-input inference counts.
Real-Electron Playwright spec for #100406: a two-member room (primary
profile + code-farmer). The user addresses only @code-farmer; its
scripted reply @mentions hermes; the assertion is a `default`-authored
"B" entry in the persisted room log. On origin/main the room settles
after Code Farmer's line and the spec fails at that assertion; with the
mention-alias fix it passes.
The mock inference server gains a per-speaker script for group rooms:
`E2E_SAY(<handle>)[<line>]` tokens in the user's send answer the member
whose turn prompt opens with `You are @<handle>`; unscripted members
reply "(pass)". `{at}` stands for `@` so the script itself never
mentions anyone and round one only drives the member the user tagged.
Same site class as the e2e fixtures: tests-js/scripts/mock-server.ts writes
the config for `npm run dev:mock`. Entry-level providers.<name>.context_length
is honoured since #98387, so a 4096 window now fails agent init below
MINIMUM_CONTEXT_LENGTH exactly as it did for the Playwright fixtures.
Keep unknown failures red, rotate evidence per attempt, and emit receipts for signature-confirmed historical cases. Add CI-only diagnostics and an exact-tag input for the unresolved July hand-off.
Follow-ups on the cherry-picked handler:
- a throwing observer can no longer change the decision; the handler returns
an explicit deny regardless of logging failures
- the denied-URL log carries origin only, so query tokens / signed URLs from
attacker-controlled content never reach the persisted desktop log
- the hidden link-title window (loads arbitrary user-linked pages on render,
had no window-open handler at all) now denies too
- tests trimmed to two invariants (proven red against the pre-fix shape)
setWindowOpenHandler opened details.url as a side effect before denying.
Per GHSA-9f4c-93c8-jc8g (CVE-2026-70608, High 7.2), a sandboxed iframe
with no allow-popups and no user gesture can reach this handler via the
OpenURL path -- and the desktop renders untrusted artifact HTML in
<iframe sandbox="allow-scripts">. A malicious artifact could therefore
force the OS browser to an attacker URL with zero interaction. Electron
ships no fixed 40.x release (fix is 41.10.3+/42.0.1), so we close it at
the seam, version-independently.
- electron/window-open-policy.ts: pure decideWindowOpen (always deny) +
createWindowOpenHandler(onDenied) that denies and never opens a URL;
the hook is logging-only.
- main.ts: wireCommonWindowHandlers uses it (covers primary + all
secondary/quick windows); the deny is logged, no side-effect open.
- Trusted external links are unaffected: they already route through the
audited hermes:openExternal IPC channel (openExternalUrl, http/https/
mailto allowlist). Converted the one remaining bare window.open on the
Electron path (env-var docs menu) to openExternalLink; other
window.open sites are bridge-absent web fallbacks.
- tests-js/window-open-policy.test.ts: 4 tests pinning always-deny, the
logging-only hook, and that a throwing hook never degrades to allow.
Conflicts, three, resolved:
- scripts/desktop-update.ps1: upstream's side taken whole. Upstream moved
the hand-off to scripts/desktop-update/windows.ps1 (this file is now a
one-line compat forwarder) and the new implementation already drains
both pipes asynchronously with bounded abandonment, which supersedes
this branch's stderr-drain fix for the same deadlock.
- apps/desktop/e2e/fixtures.ts: kept upstream's resolveElectronBinary
import alongside this branch's consolidated mock-server path.
- tests-js/scripts/mock-server.ts: kept upstream's task-panel trigger
addition inside the consolidated file; rewired the five upstream specs
still importing './mock-server' to the consolidated path (export sets
verified identical) and dropped the superseded apps/desktop/e2e copy.
The JS alignment suite pinned 24.0.0 as a supported Node; with the
engines arm raised to ^24.11.0 (babel 8 requires >=24.11), 24.0.0 and
24.10.x are now correctly rejected and 24.11+/24.18+ accepted.
Batch follow-ups on top of the salvaged privacy declarations:
- tests-js/desktop-mac-usage-descriptions.test.ts: EXPECTED_USAGE_DESCRIPTIONS
now pins the FINAL key set (camera + calendar x2 + reminders x2 + screen
capture + local network) so the drift-protection assertion locks the whole
batch as permanent regression coverage.
- entitlements.mac.plist: add com.apple.security.personal-information.reminders
alongside the calendars entitlement #65220 added — the reminders usage
descriptions need the matching entitlement under hardened runtime (sibling
site the original PR missed).
Address maintainer review feedback (PR #66215, comment by @teknium1):
> `tests/test_desktop_mac_entitlements.py:47` reads `apps/desktop/package.json`
> from pytest. `AGENTS.md:1319-1329` requires assertions about `package.json`
> and JS-side artifacts to be in the JS/Vitest suite; otherwise CI
> classification can skip the regression test on a JS-only change.
The CI change classifier (`scripts/ci/classify_changes.py`) marks
`apps/desktop/package.json` as `_FRONTEND` (in `_PY_SKIP`), so a Python test
that reads it would be skipped on a JS-only PR — regression goes green on
the PR, red on main.
Move the regression to `tests-js/desktop-mac-usage-descriptions.test.ts`,
following the same convention as commit dbf86b923 ("test: port macOS
entitlements test from Python to vitest"), which ports an earlier Python
entitlements regression into `tests-js/desktop-mac-entitlements.test.ts`
for the identical reason. The new file is a sibling of that one — both
pin Desktop macOS manifest contracts, but they assert against different
files (`entitlements.mac.plist` vs `build.mac.extendInfo` in package.json).
The Vitest port mirrors the original assertions 1:1: every
`NS*UsageDescription` key pinned (parametrized over key + required
substring + reason), no leading/trailing whitespace or newlines in any
`extendInfo` string, and a drift-protection assertion that fails when a
new privacy key is added to the build config without a matching row.
A runtime type guard on `extendInfo` ensures a non-string plist scalar
raises a clean assertion error here ("`X` in build.mac.extendInfo must
be a string (got boolean)") rather than crashing the test runner with
`value.trim is not a function` deep in the whitespace test — caught by
Flash + GPT-OSS cross-vendor review.
Verified:
- `cd tests-js && npm run check` → typecheck clean, 14/14 tests pass
(4 files including the new one with 5 tests).
- Mutation: removing `NSAppleMusicUsageDescription` from
`apps/desktop/package.json` flips 1 test red with the exact symptom
("Info.plist privacy usage description \`NSAppleMusicUsageDescription\`
is missing"). Restore → 14/14 green.
- Mutation: adding an unpinned `NSSpeechRecognitionUsageDescription` with
whitespace flips 2 tests red (drift-protection + whitespace).
- Mutation: adding a non-string `CFBundleBooleanTest: true` flips the
whole file red with the clean "must be a string (got boolean)"
assertion (no downstream crash).
- `apps/desktop` Electron Vitest project still passes (42 files,
432 tests + 1 skipped).
Closes the maintainer comment thread on PR #66215.
Fixes#54551
The desktop window opened blank white on a fresh install: React threw
"Minified React error #527" before the first paint, from the
`vendor-react-<hash>.js` chunk.
`apps/desktop` pins react and react-dom to the same exact version, but
`vite.config.ts` aliased both to a hardcoded `../../node_modules/<pkg>` —
straight into the monorepo root, where npm is free to hoist a different
react. `@streamdown/math` is a root dependency whose react peer is
`^18.0.0 || ^19.0.0` and which declares no react-dom peer, so npm hoists
the newest react (19.2.8) to the root while react-dom stays at the
version hoisted from the workspaces (19.2.7). react-dom's own peer is
`react: ^19.2.7`, which 19.2.8 satisfies, so the install reports success
and nothing warns. The bundle then shipped react 19.2.8 with react-dom
19.2.7 and React refused to run.
`npm ci` masks this because the lockfile pins the root react to 19.2.7,
which is why CI is green. The recurrence engine is
`_run_npm_install_deterministic()`: when `npm ci` fails it falls back to
`npm install --no-save`, which re-resolves the whole tree and never
records the result — so the split comes back on the next update and
leaves no trace.
Fix the resolution rather than the hoist. The aliases now resolve both
packages from the desktop workspace itself, where npm guarantees the
declared versions are reachable (it nests a copy under the workspace
exactly when the hoisted one differs), so the pair can only ever match.
Pinning react at the npm layer instead was rejected: every manifest-level
pin tried (root dependency, root `overrides`, a scoped override on
`@streamdown/math`) breaks a fresh install with
`ERESOLVE ... peer ink-text-input@"6.0.0" from @hermes/ink@0.0.1`.
Two guards keep it from silently returning:
- `assert-root-install.mjs` (the existing preflight of `build`,
`dev:renderer` and `preview`) now fails the build when the resolved
react and react-dom versions differ, so a split surfaces as an
actionable error instead of a white window.
- A `tests-js` contract test asserts every workspace pins the two to the
same exact version, and that the desktop bundler no longer points at a
hardcoded `node_modules` path.
Verified on a synthetic split tree (root react 19.2.8 / react-dom 19.2.7,
workspace react 19.2.7): the old aliases resolve 19.2.8 + 19.2.7, the new
ones resolve 19.2.7 + 19.2.7, and the preflight exits 1 when the
workspace itself resolves the mismatch.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(install): time-box the Windows node-deps stage so a stalled npm or Playwright install can't hang setup forever
scripts/install.sh has bounded this same work with run_with_timeout
"$NODE_DEPS_TIMEOUT" (600s default) since #39219, but install.ps1 never got
the guard: Install-NodeDeps ran both `npm install` and `npx playwright
install chromium` unbounded. A stalled registry fetch or a wedged Chromium
archive extraction (#76222, #84614) froze the installer indefinitely -- one
user left it running 12+ hours overnight before asking for help.
Route both invocations through _Invoke-NativeWithTimeout: cmd.exe launches
the native command with its output merged to a log, the parent polls with a
wall-clock deadline and tails new log lines to the console each tick (the
live progress that makes a 3-minute download distinguishable from a hang),
and on timeout taskkill /T /F kills the real process tree and returns 124 --
the same convention as coreutils timeout and bash's run_with_timeout.
Wait-Job was rejected for this: jobs swallow live output and Stop-Job leaves
the npm child running. Windows PowerShell 5.1-safe throughout.
Timeouts surface as a warning with the log path, a note that re-running the
installer resumes (stages are idempotent), and the NODE_DEPS_TIMEOUT env
override for slow links -- mirroring bash.
Fixes#76222.
Closes#84614.
Supersedes #76303.
Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>
* fix(installer): roll stage timers over to hours so an overnight stall doesn't read as "744 hours"
formatElapsed rendered a running stage as m:ss with unbounded minutes: a
node-deps stage left hanging overnight showed "744:38", which the user who
reported the hang understandably read as 744 hours. formatDuration
(completed stages) had the same unbounded-minutes shape.
Move both formatters into src/lib/format.ts (pure, no React) and add the
hour rollover: h:mm:ss live, "Xh Ym" completed. tests-js pins the shapes,
including 744m38s -> 12:24:38.
---------
Co-authored-by: JonthanaHanh <JonthanaHanh@users.noreply.github.com>
`hermes update` was pruning root-level Node dependencies (agent-browser)
because npm ci always wipes and reifies node_modules according to its
active filter -- no root-first/workspace-first ordering or flag
combination (--workspaces=false, --include-workspace-root, etc.) can
reliably keep a root-only package.json dependency from being pruned by
a subsequent workspace-scoped npm ci. Confirmed empirically and via
npm/cli source (isArboristCmd hardcodes includeWorkspaceRoot=false for
ci/install), so no amount of install-order juggling fixes this for good.
Instead of chasing install order, remove the root-only dependencies
that made the npm step fragile in the first place:
- agent-browser is no longer a root package.json dependency. It
resolves lazily via `npx agent-browser` (tools/browser_tool.py
already had this as a fallback; it's now the primary path).
warm_agent_browser_npx_cache() is called fire-and-forget from both
`hermes update` and `hermes doctor --fix` to keep npx's cache warm,
preserving the "available before any session starts" property
agent-browser had as an eager dependency without re-entangling it
with the npm workspace graph.
- @streamdown/math moves to apps/desktop/package.json, where it's
actually imported (markdown-text.tsx, katex-memo.ts) -- it was
never used anywhere else and was subject to the same pruning risk.
- _update_node_dependencies() collapses to a single
`npm ci --workspace ui-tui --workspace web` call now that root has
no dependencies of its own to protect, and keeps its original spot
ahead of `_build_web_ui()` at both call sites in update_cmd.py --
with no root-only dependencies left to protect, there's no reason
for the Node refresh and the web build to run in any particular
order relative to each other.
- hermes_cli/tools_config.py's post-setup Chromium-install path and
hermes_cli/doctor.py's agent-browser check both now resolve through
the same PATH -> Homebrew/Hermes-managed-node -> npx cascade
(_find_agent_browser / _resolve_npx_bin) instead of hand-rolling
their own node_modules/.bin lookups, so they can't diverge from what
browser tools actually invoke at runtime.
- tests-js/package-json-lazy-deps.test.ts gets a lockfile-level check
mirroring the existing camofox one, so a future regression that
reintroduces agent-browser into package-lock.json fails this test
directly instead of relying on manual review to catch it.
Fixes#43564.
The dev:mock script duplicated the e2e mock server. The copy had only
the plain chat reply; every scripted path lived only in the e2e version.
A single mock server now lives in tests-js/scripts/mock-server.ts.
The e2e suite imports it as a library. Running the file directly
starts the server, writes a mock config, and launches the desktop app.
The dev:mock script now runs that file.
The e2e tsconfig lists tests-js/scripts in its include, because the
composite project rule requires every imported file to be listed.
Both halves of this bug were the same failure mode: allowScripts is keyed by
exact name@version, so an entry stops matching the moment a dependency moves
and npm demotes the blocked script to a warning nobody reads. The breakage
surfaces much later as a missing native artifact on one platform.
Assert the two relationships that make the allowlist meaningful — every
versioned pin resolves to a version the lockfile installs, and every package
the lockfile marks as having an install script carries a decision. A
bare-name key stays exempt from the version check so a standing denial like
unicode-animations survives bumps.
Lives in tests-js because the CI change classifier does not run the Python
suite for a manifest-only diff.
pin brace-expansion to 5.0.8
update concurrently to 10.0.4
update electron-builder to 26.15.3
update eslint to 10.8.0
update eslint-plugin-perfectionist to 5.10.0
update @assistant-ui/react to 0.15.0
update @assistant-ui/react-streamdown to 0.3.8
update radix-ui to 1.6.7
update react-router-dom to react-router@8.3.0 - react-router-dom is no longer a standalone package, it just reexports react-router
remove @radix-ui/react-slot: we import this from `radix-ui`
remove eslint-plugin-react: we imported it, but never actually used it!
Every npm workspace package now defines check:* scripts (check:unit,
check:lint, check:bundle, check:typecheck, etc.) that fan out to
separate matrix runners in CI. The check umbrella script chains all
shards for local dev.
The matrix discovery in the workspaces job queries npm workspaces,
finds check:* scripts (in package.json insertion order), falls back to
check when none exist, and emits an include matrix. No hardcoded
package names — the workflow is fully auto-derived from workspace
metadata.
Previously every package ran a single check script on one worker, and
the fix step (lint:fix + prettier) ran as a separate CI step with
special-cased run_fix gating to avoid running on every shard. Now that
lint is just another check:lint shard, the run_fix field and the fix
step are gone entirely — lint runs in its own runner like everything
else.
tests/test_desktop_mac_entitlements.py asserts about
apps/desktop/electron/*.plist — the same CI blind spot as the other
ported tests: the change classifier routes apps/ changes to the
frontend lane, so a PR touching only the plists would skip the Python
suite and the regression would go green on the PR and red on main.
Ported to tests-js/desktop-mac-entitlements.test.ts so it runs in the
correct lane. All three tests carry over: the inherit plist grants
audio-input (regression #37718), every device.* entitlement on the
main app is also inherited, and both files remain well-formed.
Parsing uses the `plist` package (pinned to ^3.1.0, the version
already present in the workspace lockfile, so no new transitive
packages) plus `@types/plist` — Node's stand-in for plistlib.
Verified: tests-js `npm run check` passes (typecheck + 9/9 tests), and
a mutation run (removing audio-input from the inherit plist) turns
both regression tests red.