`run_tests.sh` defaults to twice the core count, and the value this branch
started with came from a rule of thumb of 1.5x cores plus a measurement on a
16-core machine. A sweep on the real runner disagrees with both.
Run 32549672063 on the 96-core runner (EPYC 7763, 377GB) timed the whole suite
at six worker counts, two repetitions for each. A warmup run came first, and
retries were off:
workers x cores rep 1 rep 2 mean
48 0.5x 138s 139s 138s
96 1.0x 127s 126s 126s <- fastest
144 1.5x 130s 134s 132s
192 2.0x 132s 133s 132s
240 2.5x 140s 139s 140s
288 3.0x 143s 142s 142s
One worker for each core wins. Both repetitions agree on the order.
The shape is the more useful result. The range is 126s to 142s across a 6x
range of worker counts. The suite has sufficient concurrency at this machine
size, so nothing above the core count buys anything. The remaining time
belongs to the slowest individual files and to the setup. A future gain must
come from those, and not from this number.
The sweep ran from a temporary workflow that this branch does not keep.