Windows cold start can hold the backend's GIL for 12-28s while gateway
platform modules import (feishu/lark, weixin, telegram C extensions),
leaving the /api/ws upgrade unanswered well past the probe's fixed 10s
connect budget. The probe then fails, the desktop declares the backend
unhealthy, spawns a second backend, and boot lands in the
'Ignoring stale Hermes backend exit (1)' cascade (#96177).
The probe now supports backend-progress-aware waiting: once the quiet
connect budget passes, a keepWaitingWhile callback is polled and the
deadline only fires when the backend stops reporting progress (hard
capped by maxConnectWaitMs so a hung backend still fails the boot).
- gateway-ws-probe.ts: keepWaitingWhile / progressCheckIntervalMs /
maxConnectWaitMs options; behavior is unchanged when the callback is
omitted (remote-gateway probes keep the fixed timeout).
- main.ts: both local-backend boot paths (primary window + pool) pass
keepWaitingWhile = backendMakingProgress(child, outputTail) with a
90s cap aligned with the existing port-announcement cold-start budget.
- backend-claim.ts: BackendOutputTail tracks lastActivityAt; new
backendMakingProgress() signal: alive child + (no output yet, i.e.
inside the block-buffered import window, or recent output within 30s).
- Regression tests: probe waits past the base budget while progress is
reported, fails once progress stops, respects the hard cap, fails
closed on a throwing callback; tail activity + progress-signal units.
Benchmark (simulated 13s import stall, production budgets):
before: FAIL at 10.0s -> boot cascade after: OK at 13.8s
warm start: OK 0.96s after: OK 0.96s (no delta)
dead backend: FAIL 10.0s after: FAIL 10.1s (same)
hung backend (alive, never ready): after: FAIL 90.8s (hard cap)
Closes#96177
(cherry picked from commit e8dbe9e0a2bfb6853401d55a2bbda02393560de0)