refactor(tests): xdist is the standard runner; drop the per-file subprocess machinery

scripts/run_tests.sh now runs pytest-xdist -n <N> --dist loadfile as
the single canonical path on every OS (Linux and Windows CI lanes both
use it). The per-file subprocess model (run_tests_parallel.py, and the
interim run_xdist.sh experiment) is deleted along with its two
self-tests: persistent xdist workers pay the interpreter+import wall
once per worker instead of once per file (~0.5-1.5s x ~3400 files was
a ~6-minute floor on Windows), and --dist loadfile pins a file's tests
to ONE worker, bounding state pollution to co-scheduled files — which
is exactly the class of flake we are now committing to fix properly.

The serial process-killer quarantine phase is dropped too. It existed
to keep process-tree-sweep tests from killing sibling xdist workers;
the durable fix belongs in those tests (sweeps must target their own
children, not enumerate every python process), and keeping a
divergent two-phase path would hide that work.

Kept from the old wrapper: hermetic env -i scrubbing, Windows
location-var forwarding, venv probing, bytecode pre-compile,
-m 'not integration', and the HERMES_TEST_IMAGE docker-knob
allowlist. HERMES_TEST_FILE_TIMEOUT/FILE_RETRIES/SLICE go away with
the runner they parameterized; -j/HERMES_TEST_WORKERS now map to xdist
-n (Linux CI pins 96, Windows 32, default auto).

Docs updated to match: AGENTS.md (runner contract, flake policy,
isolation section), CONTRIBUTING.md, tests/conftest.py comments,
classify_changes docstring + its lane expectation (a .sh runner no
longer trips the supply-chain scan lane), comfyui README,
hermes-agent contributor guide, debugpy skill and its website doc.
This commit is contained in:
ethernet
2026-08-31 13:32:49 -04:00
parent c81c93e771
commit ef4a82eed0
20 changed files with 118 additions and 2021 deletions

View File

@@ -172,10 +172,9 @@ jobs:
NOUS_API_KEY: ""
run: |
# Each of these tests drives a container, so the docker daemon sets
# the limit and not the processor. This pins the workers explicitly
# (nproc) so the comment and behavior stay coupled if the runner's
# core count or the default ever changes.
HERMES_TEST_WORKERS=$(nproc) scripts/run_tests.sh tests/docker/ --file-timeout 600
# the limit and not the processor. This pins the xdist worker count
# to the core count.
HERMES_TEST_WORKERS=$(nproc) scripts/run_tests.sh tests/docker/
# ---------------------------------------------------------------------------
# Rebuild and push each architecture only after the unprivileged build/test

View File

@@ -26,8 +26,7 @@ jobs:
# there has a long tail of files that spawn servers/sockets/subprocesses
# and run minutes-per-file on the CI runner (they pass in seconds on an
# idle dev box) — measured 88.6% completion at 60 minutes on run
# 33338310764. The per-file cap (HERMES_TEST_FILE_TIMEOUT) bounds any
# single hung file; this bounds the whole lane.
# 33338310764.
timeout-minutes: ${{ matrix.timeout_min }}
strategy:
fail-fast: false
@@ -37,19 +36,16 @@ jobs:
runner: ubuntu-latest-96-core
marker: ""
workers: "96"
file_timeout: ""
timeout_min: 30
- os: windows
runner: windows-latest-32-core
marker: ""
workers: "32"
file_timeout: "2400"
timeout_min: 150
- os: macos
runner: macos-latest
marker: macos_only
workers: ""
file_timeout: ""
timeout_min: 30
steps:
- name: Disable Windows Defender real-time scanning
@@ -151,20 +147,17 @@ jobs:
run: uv cache prune --ci
- name: Run tests
# Three shapes:
# Two shapes:
#
# * Windows (EXPERIMENT, PR #96458): xdist with --dist loadfile.
# The per-file subprocess model pays a spawn+import wall of
# ~0.5-1.5s per file x ~3400 files there — a ~6-minute floor
# before any test runs. xdist pays that wall once per worker.
# loadfile keeps each file's tests on ONE worker, so the
# remaining hazard is state shared by files co-scheduled on a
# worker. Failures from that are stateful-test bugs to fix, not
# runner bugs.
# * Linux: the full suite via scripts/run_tests.sh — per-file
# isolation, bounded parallelism. The spawn floor is ~15ms per
# file there, so isolation is nearly free; keep the stronger
# guarantee.
# * Full-suite lanes (Linux + Windows): scripts/run_tests.sh —
# pytest-xdist with --dist loadfile behind a hermetic env.
# Persistent workers pay the interpreter+import wall once per
# worker instead of once per file (a ~6-minute floor on
# Windows under the old per-file subprocess model). loadfile
# keeps each file's tests on ONE worker, so the remaining
# hazard is state shared by files co-scheduled on a worker.
# Failures from that are stateful-test bugs to fix, not runner
# bugs.
# * macOS (marker set): plain pytest over the files that carry the
# marker. list_os_marked_tests.py narrows WHICH FILES are
# imported (collection otherwise drags ~900 unrelated modules
@@ -179,10 +172,6 @@ jobs:
run: |
set -uo pipefail
if [ "${{ matrix.os }}" = "windows" ]; then
bash scripts/run_xdist.sh 32 tests/
exit $?
fi
if [ -n "${{ matrix.marker }}" ]; then
LIST="${RUNNER_TEMP:-.}/selected-tests.txt"
@@ -224,19 +213,11 @@ jobs:
# activation.
scripts/run_tests.sh
env:
# The maximum number of test FILES that run together. Linux pins
# 96 (the measured whole-suite sweep — one worker per core is
# fastest on the 96-core runner, and the curve is shallow).
# Windows and macOS leave this unset so run_tests_parallel.py's
# default (cpu_count) applies.
# xdist worker count (-n). Linux pins 96 (the measured whole-suite
# sweep — one worker per core is fastest on the 96-core runner,
# and the curve is shallow); Windows pins 32. macOS leaves this
# unset (it runs the marked-file lane, not the full suite).
HERMES_TEST_WORKERS: ${{ matrix.workers }}
# Windows CI-only: the full-suite lane showed ~160 files at or
# past the 300s default under runner contention (they pass in
# seconds on an idle box); each kill also burns a retry, so the
# churn — not test time — dominated the lane's wall clock. Give
# the per-file cap real headroom there. Empty on other lanes
# (run_tests.sh drops empty vars before the runner sees them).
HERMES_TEST_FILE_TIMEOUT: ${{ matrix.file_timeout }}
# Ensure tests don't accidentally call real APIs
OPENROUTER_API_KEY: ""
OPENAI_API_KEY: ""

View File

@@ -1579,30 +1579,29 @@ def profile_env(tmp_path, monkeypatch):
### Python
**ALWAYS use `scripts/run_tests.sh`** — do not call `pytest` directly. The script enforces
hermetic environment parity with CI (unset credential vars, TZ=UTC, LANG=C.UTF-8,
per-file subprocess isolation via `scripts/run_tests_parallel.py` — no xdist,
worker count auto-scaled from CPU count). Direct `pytest`
pytest-xdist with `--dist loadfile` so every test of a file runs on ONE worker —
worker count defaults to the CPU count). Direct `pytest`
on a 16+ core developer machine with API keys set diverges from CI in ways
that have caused multiple "works locally, fails in CI" incidents (and the reverse).
```bash
scripts/run_tests.sh # full suite, CI-parity
scripts/run_tests.sh tests/gateway/ # one directory
scripts/run_tests.sh tests/agent/test_foo.py -k test_x # one test (file + -k; the runner is file-granular)
scripts/run_tests.sh tests/agent/test_foo.py -k test_x # one test (file + -k)
scripts/run_tests.sh -v --tb=long # pass-through pytest flags
```
**Flake policy:** the runner auto-retries a failing test FILE once in a fresh
subprocess (`--file-retries`, default 1; `HERMES_TEST_FILE_RETRIES=0` to
disable). Pass-on-retry counts as green but is printed in a `⚠ FLAKY` summary
section with both attempts' output. A FLAKY report is a bug to fix, not noise
to ignore — timing-sensitive tests must not assume a quiet runner (loose
**Flake policy:** the runner is xdist, so a file may pass or fail depending
on which siblings share its worker — any order-dependent failure is a bug to
fix, not noise. Timing-sensitive tests must not assume a quiet runner (loose
wall-clock bounds ≥ 2s, event-based sync, no `assert not _wait_until(...)`
negative-timing races).
#### Subprocess-per-test-file isolation
#### File-pinned xdist workers
Every test file runs in a freshly-spawned Python subprocess via `run_tests_parallel.py`. This means module-level dicts/sets and
ContextVars from one test file cannot leak into the next.
Every test file runs pinned to ONE xdist worker (`--dist loadfile` via `scripts/run_tests.sh`), so module-level
state pollution is bounded to files co-scheduled on the same worker; files that need true isolation should reset
state in fixtures.
#### Why the wrapper

View File

@@ -201,8 +201,8 @@ ln -sf "$(pwd)/venv/bin/hermes" ~/.local/bin/hermes
### Run tests
```bash
# Preferred — matches CI (hermetic `env -i`, per-file subprocess isolation
# via run_tests_parallel.py, worker count auto-scaled); see AGENTS.md
# Preferred — matches CI (hermetic `env -i`, pytest-xdist with --dist
# loadfile, worker count defaults to the CPU count); see AGENTS.md
scripts/run_tests.sh
# Alternative (activate the venv first). The wrapper is still recommended

View File

@@ -45,7 +45,7 @@ When you change a script:
The parent hermes-agent repo used to enable `pytest-xdist` by default
(`-n auto`); the canonical runner has since moved to per-file subprocess
isolation via `scripts/run_tests_parallel.py` and no longer uses xdist.
pytest-xdist with `--dist loadfile` via `scripts/run_tests.sh`.
This suite is small enough that parallelism isn't worth the complexity, and
pytest-xdist isn't always installed in the user's environment. The
`-c tests/pytest.ini -o addopts="-p no:xdist"` flags make the suite run

View File

@@ -144,7 +144,7 @@ def _py_test_only(p: str) -> bool:
Product jobs (Desktop E2E's ``hermes serve`` backend, the Docker image)
run installed code — nothing under ``tests/`` is packaged or importable
there. scripts/run_tests.sh and run_tests_parallel.py are deliberately
there. scripts/run_tests.sh is deliberately
NOT test-only: they are runner infrastructure, and a bad edit there can
mask real failures, so they stay conservative (python_prod=true).
"""

View File

@@ -3,10 +3,11 @@
# `pytest` directly to guarantee your local run matches CI behavior.
#
# What this script enforces:
# * Per-file isolation via scripts/run_tests_parallel.py — each test
# file runs in its own freshly-spawned `python -m pytest <file>`
# subprocess. No xdist, no shared workers, no module-level leakage
# between files.
# * pytest-xdist with --dist loadfile — each test FILE's tests all run on
# ONE worker, so file-internal ordering is preserved and cross-file
# pollution is bounded to files co-scheduled on a worker. Persistent
# workers also pay the interpreter+import wall (~0.5-1.5s on Windows)
# once per worker instead of once per file.
# * TZ=UTC, LANG=C.UTF-8, PYTHONHASHSEED=0 (deterministic)
# * Env vars blanked (conftest.py also does this, but this
# is belt-and-suspenders for anyone running pytest outside our
@@ -15,23 +16,19 @@
#
# Usage:
# scripts/run_tests.sh # full suite
# scripts/run_tests.sh -j 4 # cap parallelism
# scripts/run_tests.sh -j 4 # cap worker count
# scripts/run_tests.sh tests/agent/ # discover only here
# scripts/run_tests.sh tests/agent/ tests/acp/ # multiple roots
# scripts/run_tests.sh tests/foo.py # single file
# scripts/run_tests.sh tests/foo.py -q # path + bare pytest flag
# scripts/run_tests.sh tests/foo.py -v --tb=long # bare flags "just work"
# scripts/run_tests.sh -k 'pattern' # value flags pass through too
# scripts/run_tests.sh tests/foo.py -- --tb=long # explicit '--' still works
#
# Bare pytest flags (anything starting with '-' that isn't one of this
# runner's own options: -j/--jobs, --paths, --slice, --file-timeout, etc.)
# are forwarded to each per-file pytest invocation automatically — no '--'
# separator required. The explicit '--' form still works and stacks with
# bare flags. Positional path arguments override the default discovery
# root (tests/).
# Bare pytest flags (anything starting with '-' that isn't -j/--jobs) are
# forwarded to pytest. Positional path arguments override the default
# discovery root (tests/).
set -euo pipefail
set -uo pipefail
# ── Locate repo root ────────────────────────────────────────────────────────
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
@@ -40,7 +37,7 @@ REPO_ROOT="$(cd "$SCRIPT_DIR/.." && pwd)"
# ── Locate python ───────────────────────────────────────────────────────────
# Probe local venvs first; fall back to the Nix devShell's editable venv
# (HERMES_PYTHON is exported by the devShell hook and ships [dev] extras:
# pytest, pytest-asyncio, pytest-timeout, ruff, ty).
# pytest, pytest-asyncio, pytest-xdist).
#
# A candidate must have pytest INSTALLED, not merely exist. The release venv
# at ~/.hermes/hermes-agent/venv has bin/activate but no pytest, so an
@@ -59,12 +56,7 @@ for candidate in "$REPO_ROOT/.venv" "$REPO_ROOT/venv" "$HOME/.hermes/hermes-agen
break
fi
SKIPPED_VENVS="$SKIPPED_VENVS $candidate"
fi
# Native Windows venv layout: python.exe and activate live under
# Scripts/, and there is no bin/. Anyone running this script from
# Git Bash / MSYS with a `python -m venv`- or uv-created venv hits
# this branch — without it the canonical runner refuses to start.
if [ -f "$candidate/Scripts/activate" ]; then
elif [ -f "$candidate/Scripts/activate" ]; then
if "$candidate/Scripts/python.exe" -c 'import pytest' 2>/dev/null; then
VENV="$candidate"
VENV_PYTHON="$candidate/Scripts/python.exe"
@@ -73,40 +65,35 @@ for candidate in "$REPO_ROOT/.venv" "$REPO_ROOT/venv" "$HOME/.hermes/hermes-agen
SKIPPED_VENVS="$SKIPPED_VENVS $candidate"
fi
done
if [ -n "$SKIPPED_VENVS" ]; then
for skipped in $SKIPPED_VENVS; do
echo "▶ skipping venv without pytest: $skipped" >&2
done
fi
if [ -n "$VENV" ]; then
PYTHON="$VENV_PYTHON"
elif [ -n "${HERMES_PYTHON:-}" ] && [ -x "$HERMES_PYTHON" ] \
&& "$HERMES_PYTHON" -c 'import pytest' 2>/dev/null; then
# Guard with an import check: HERMES_PYTHON may point at the RELEASE
# venv (no pytest) when inherited from a wrapped `hermes` binary rather
# than the devShell hook.
PYTHON="$HERMES_PYTHON"
echo "▶ no local venv — using Nix dev venv via HERMES_PYTHON: $PYTHON"
if [ -z "$VENV_PYTHON" ]; then
if [ -n "${HERMES_PYTHON:-}" ] && "${HERMES_PYTHON}" -c 'import pytest' 2>/dev/null; then
VENV_PYTHON="$HERMES_PYTHON"
else
echo "error: no virtualenv with pytest found in $REPO_ROOT/.venv or $REPO_ROOT/venv," >&2
echo " and HERMES_PYTHON is not a python with pytest (enter the Nix devShell or create a venv)" >&2
echo "✗ No venv with pytest found. Install dev extras:" >&2
echo " uv sync --extra dev" >&2
if [ -n "$SKIPPED_VENVS" ]; then
echo " (skipped for missing pytest:$SKIPPED_VENVS — install dev extras there, or create $REPO_ROOT/.venv)" >&2
fi
exit 1
fi
# ── Live-gateway plugin (computed before we drop env) ───────────────────────
EXTRA_PYTHONPATH=""
EXTRA_PYTEST_PLUGINS=""
if [ -f "$HOME/.hermes/pytest_live_guard.py" ]; then
EXTRA_PYTHONPATH="$HOME/.hermes"
EXTRA_PYTEST_PLUGINS="pytest_live_guard"
fi
PYTHON="$VENV_PYTHON"
# ── Split args: our -j/--jobs vs pytest passthrough ─────────────────────────
N="${HERMES_TEST_WORKERS:-auto}"
PYTEST_ARGS=()
while [ $# -gt 0 ]; do
case "$1" in
-j|--jobs)
N="$2"; shift 2 ;;
-j*)
N="${1#-j}"; shift ;;
--jobs=*)
N="${1#--jobs=}"; shift ;;
*)
PYTEST_ARGS+=("$1"); shift ;;
esac
done
# ── Windows location variables (computed before we drop env) ───────────────
# `env -i` forwards HOME, which is enough on POSIX. Native Windows CPython
@@ -123,27 +110,28 @@ for _win_var in USERPROFILE HOMEDRIVE HOMEPATH LOCALAPPDATA APPDATA SYSTEMROOT T
fi
done
# ── Test-runner knobs (computed before we drop env) ────────────────────────
# The runner's own documented environment knobs must survive the hermetic
# `env -i` below, or they are silent no-ops for anyone invoking this script:
#
# * HERMES_TEST_WORKERS / PATHS / FILE_TIMEOUT / FILE_RETRIES / SLICE are
# read by run_tests_parallel.py at argparse-default time — inside the
# stripped environment.
# ── Live-gateway plugin (computed before we drop env) ───────────────────────
EXTRA_PYTHONPATH=""
EXTRA_PYTEST_PLUGINS=""
if [ -f "$HOME/.hermes/pytest_live_guard.py" ]; then
EXTRA_PYTHONPATH="$HOME/.hermes"
EXTRA_PYTEST_PLUGINS="pytest_live_guard"
fi
# ── Test-runner knobs (computed before we drop env) ──────────────────────────
# * HERMES_TEST_IMAGE is read by tests/docker/conftest.py to skip its
# session-scoped `docker build`. CI's docker.yml sets it to the image
# the build step just loaded; stripping it made every per-file pytest
# subprocess rebuild the 5GB image from a cold builder cache instead
# (~4 min per worker per run, and the rebuilt image lacked the
# HERMES_GIT_SHA build-arg the workflow bakes in).
# the build step just loaded; stripping it made every pytest subprocess
# rebuild the 5GB image from a cold builder cache instead (~4 min per
# worker per run, and the rebuilt image lacked the HERMES_GIT_SHA
# build-arg the workflow bakes in).
#
# These are test-infrastructure knobs, not credentials — same class as the
# HERMES_RUN_SLOW_PET_TESTS / HERMES_E2E_BROWSER opt-ins already forwarded.
# Keep this an explicit allowlist (no HERMES_TEST_* glob) so the "no
# credential can leak" property stays auditable at a glance.
TEST_ENV=()
for _test_var in HERMES_TEST_IMAGE HERMES_TEST_WORKERS HERMES_TEST_PATHS \
HERMES_TEST_FILE_TIMEOUT HERMES_TEST_FILE_RETRIES HERMES_TEST_SLICE; do
for _test_var in HERMES_TEST_IMAGE; do
if [ -n "${!_test_var:-}" ]; then
TEST_ENV+=("$_test_var=${!_test_var}")
fi
@@ -152,20 +140,18 @@ done
# ── Run in hermetic env ──────────────────────────────────────────────────────
# env -i: start with empty environment, opt-in only what we need.
# No credential var can leak — you'd have to explicitly add it here.
echo "▶ running per-file parallel test suite via run_tests_parallel.py"
echo "▶ running pytest-xdist (-n $N --dist loadfile)"
echo " (TZ=UTC LANG=C.UTF-8 PYTHONHASHSEED=0; clean env)"
cd "$REPO_ROOT"
# ── Pre-compile .pyc bytecode cache ─────────────────────────────────────────
# Each test file runs in its own subprocess via run_tests_parallel.py.
# Pre-building the bytecode cache once here (instead of each subprocess
# compiling on first import) avoids redundant work across ~2000 processes.
# Uses git to list tracked .py files (skips venv, node_modules, etc).
# xdist workers import the same modules; pre-building the bytecode cache once
# here avoids every worker compiling on first import.
echo "▶ pre-compiling bytecode cache"
"$PYTHON" -m compileall -q -j 0 -- $(git ls-files '*.py') >/dev/null 2>&1 || true
echo "▶ launching test runner"
echo "▶ pytest -n $N --dist loadfile"
exec env -i \
PATH="$PATH" \
HOME="$HOME" \
@@ -180,4 +166,6 @@ exec env -i \
${HERMES_E2E_BROWSER:+HERMES_E2E_BROWSER="$HERMES_E2E_BROWSER"} \
${EXTRA_PYTHONPATH:+PYTHONPATH="$EXTRA_PYTHONPATH"} \
${EXTRA_PYTEST_PLUGINS:+PYTEST_PLUGINS="$EXTRA_PYTEST_PLUGINS"} \
"$PYTHON" "$SCRIPT_DIR/run_tests_parallel.py" "$@"
"$PYTHON" -m pytest -n "$N" --dist loadfile -p no:cacheprovider \
-m "not integration" -q --tb=line \
${PYTEST_ARGS[@]+"${PYTEST_ARGS[@]}"}

File diff suppressed because it is too large Load Diff

View File

@@ -1,75 +0,0 @@
#!/usr/bin/env bash
# xdist runner: persistent workers sharing one collection pass, instead of
# the per-file subprocess model. EXPERIMENT (PR #96458): the per-file model
# pays a spawn+import wall of ~0.5-1.5s per file x 3400 files (~6 min floor
# on Windows, the dominant share of the lane's runtime). xdist pays that
# import wall once per worker.
#
# TWO PHASES:
#
# 1. xdist -n N --dist loadfile over the whole suite EXCEPT the
# process-killer files below. loadfile keeps a file's tests on ONE
# worker, bounding cross-file pollution to co-scheduled files.
#
# 2. The killer files run SERIALLY, one plain pytest per file. These are
# the tests that spawn and kill REAL process trees (live_system_guard
# bypasses, venv-holder sweeps, taskkill /F /T, stale-process reaps).
# Under the per-file runner each ran alone in its own session
# (start_new_session), so a sweep could only hit its own children.
# Under shared xdist workers, those sweeps enumerate every python
# process — 32 sibling workers and the controller match — and the
# first CI xdist run died with "runner lost communication" (run
# 33389773498, 1h15m, no logs flushed). Serially, a sweep's only
# reachable victims are its own children.
#
# Failures in phase 1 are stateful-test bugs to fix, not runner bugs.
set -uo pipefail
cd "$(dirname "$0")/.."
export PATH="/c/Program Files/Git/bin:$PATH"
unset PYTHONPATH PYTHONPYCACHEPREFIX HERMES_RUNTIME_DIR
export TZ=UTC LANG=C.UTF-8 LC_ALL=C.UTF-8 PYTHONHASHSEED=0 PYTHONUTF8=1
N="${1:-auto}"
shift || true
# Tests that spawn/kill real process trees. Do NOT run these inside xdist
# workers — see the phase-2 comment above.
KILLERS=(
tests/cron/test_cron_script.py
tests/gateway/test_control_socket_windows_live.py
tests/gateway/test_replace_child_reap.py
tests/gateway/test_whatsapp_bridge_pidfile.py
tests/gateway/test_whatsapp_connect.py
tests/gateway/test_whatsapp_stale_bridge.py
tests/hermes_cli/test_dashboard_lifecycle_flags.py
tests/hermes_cli/test_gateway_windows.py
tests/hermes_cli/test_update_orphan_backend_reap.py
tests/hermes_cli/test_update_stale_dashboard.py
tests/test_install_autostash_conflict_recovery.py
tests/test_install_lockfile_churn.py
tests/test_install_ps1_venv_process_tree.py
tests/test_install_unmerged_index.py
tests/tools/test_process_registry.py
)
IGNORES=()
for f in "${KILLERS[@]}"; do
[ -f "$f" ] && IGNORES+=(--ignore "$f")
done
echo "▶ phase 1: xdist -n $N --dist loadfile (bulk)"
PYTEST_STATUS=0
.venv/Scripts/python.exe -m pytest -n "$N" --dist loadfile \
-p no:cacheprovider -m "not integration" -q --tb=line \
"${IGNORES[@]}" "$@" || PYTEST_STATUS=$?
echo "▶ phase 2: serial per-file (process-killer tests)"
for f in "${KILLERS[@]}"; do
[ -f "$f" ] || continue
echo " - $f"
.venv/Scripts/python.exe -m pytest "$f" -p no:cacheprovider \
-m "not integration" -q --tb=line || PYTEST_STATUS=$?
done
exit "$PYTEST_STATUS"

View File

@@ -86,7 +86,7 @@ run_conversation():
Use the canonical runner — it enforces CI-parity (hermetic `env -i`, unset
credentials, TZ=UTC, per-file subprocess isolation via
`scripts/run_tests_parallel.py` — no xdist, worker count auto-scaled):
`scripts/run_tests.sh` — pytest-xdist, `--dist loadfile`, worker count = CPU count):
```bash
scripts/run_tests.sh # full suite

View File

@@ -107,7 +107,7 @@ scripts/run_tests.sh tests/path/to/test_file.py::test_name --trace
scripts/run_tests.sh tests/path/to/test_file.py --showlocals --tb=long
```
Note: `scripts/run_tests.sh` runs each test file in a captured subprocess via `run_tests_parallel.py` (no xdist), so interactive pdb does NOT work under the wrapper. Run pytest directly for `--pdb`:
Note: `scripts/run_tests.sh` runs pytest under xdist workers, so interactive pdb does NOT work under the wrapper. Run pytest directly for `--pdb`:
```bash
source .venv/bin/activate

View File

@@ -176,10 +176,12 @@ CASES = {
_lanes(python=True, scan=True),
),
# Runner infrastructure is NOT tests-only — a bad runner edit can mask
# real failures, so it keeps the conservative full lane set.
# real failures, so it keeps the conservative full lane set. (scan is
# off for the .sh form: the supply-chain lanes scan executable .py/.pth
# payloads, and a shell script isn't one.)
"test runner script → python_prod stays on": (
["scripts/run_tests_parallel.py"],
_lanes(python=True, scan=True),
["scripts/run_tests.sh"],
_lanes(python=True),
),
# Supply-chain lanes
".pth file → scan": (["evil.pth"], _lanes(python=True, scan=True)),

View File

@@ -115,20 +115,18 @@ os.environ["HERMES_TEST_ISOLATION"] = os.environ.get("HERMES_HOME", "") or "1"
HERMES_HOME_AT_CONFTEST_IMPORT = os.environ.get("HERMES_HOME", "")
# ── Per-file process isolation ──────────────────────────────────────────────
# Tests run via ``scripts/run_tests_parallel.py``, which spawns a fresh
# ``python -m pytest <file>`` subprocess per test file. Cross-file state
# leakage (module-level dicts, ContextVars, caches) is impossible: each
# file gets a clean Python interpreter. Intra-file ordering is the test
# author's responsibility — if test A in foo.py mutates state that test B
# in foo.py reads, that's a real bug to fix in the file (it would also
# bite anyone running ``pytest tests/foo.py`` directly).
# ── File-level scheduling isolation ──────────────────────────────────────────
# Tests run via ``scripts/run_tests.sh`` — pytest-xdist with ``--dist
# loadfile``, which pins every test of a FILE to ONE worker. Cross-file
# state leakage is bounded to files co-scheduled on the same worker; the
# historic per-file-subprocess model gave full process isolation but paid
# a spawn+import wall per file that dominated Windows runtimes. Intra-file
# ordering is the test author's responsibility — if test A in foo.py
# mutates state that test B in foo.py reads, that's a real bug to fix in
# the file (it would also bite anyone running ``pytest tests/foo.py``
# directly).
#
# This replaces the historic _reset_module_state autouse fixture (manual
# state clearing) and the brief experiment with subprocess-per-test
# isolation (too slow at ~17k tests).
#
# See ``scripts/run_tests_parallel.py`` for the runner.
# See ``scripts/run_tests.sh`` for the runner.
# ── Credential env-var filter ──────────────────────────────────────────────
@@ -783,16 +781,14 @@ def _state_db_write_guard(request, monkeypatch):
yield
# ── Module-level state reset — replaced by per-file process isolation ──────
# ── Module-level state reset — replaced by file-pinned xdist workers ────────
#
# Each test FILE runs in a freshly-spawned ``python -m pytest <file>``
# subprocess via ``scripts/run_tests_parallel.py``, so module-level dicts /
# sets / ContextVars from tests in one file cannot leak into tests in
# another file. No manual per-module clearing needed.
#
# Within a single file, ordering is the author's responsibility. If your
# tests in the same file share mutable state, either reset it explicitly
# in a fixture or split them across files.
# ``--dist loadfile`` pins each test FILE to ONE worker, so heavy
# co-scheduling pollution (module-level dicts / sets / ContextVars shared
# by many files) is bounded to files that land on the same worker. Within
# a single file, ordering is the author's responsibility. If your tests
# in the same file share mutable state, either reset it explicitly in a
# fixture or split them across files.
#
# The skill ``test-suite-cascade-diagnosis`` documents the cascade patterns
# this replaces; the running example was ``test_command_guards`` failing

View File

@@ -1,462 +0,0 @@
"""Verify scripts/run_tests_parallel.py kills test-spawned grandchildren.
Setup
-----
A test in this file spawns a long-lived Python grandchild that writes
its PID + a nonce to a tempfile, then exits without cleaning up.
With the old ``subprocess.run`` runner, that grandchild would orphan
and outlive the test (and the whole runner). With the current Popen +
``start_new_session`` + ``_kill_tree`` runner, the grandchild gets
SIGKILL'd via process-group kill when its file's pytest exits.
The leaker test always passes — its only job is to spawn a grandchild
and walk away. The verifier runs the runner over the leaker file in a
subprocess, then waits for the grandchild PID to disappear from the
kernel's process table.
POSIX-only: Windows has its own grandchild lifecycle (no shared session,
``taskkill /F /T`` semantics). Marked accordingly.
"""
from __future__ import annotations
import json
import os
import subprocess
import sys
import textwrap
import time
from pathlib import Path
import pytest
# Both tests share the same handoff file: the leaker writes here, the
# verifier reads here. We park it in $TMPDIR with a unique-per-run name
# so concurrent invocations of the suite don't clobber each other.
_HANDOFF_DIR = Path(os.environ.get("TMPDIR", "/tmp")) / "hermes-isolation-probe"
_HANDOFF_DIR.mkdir(exist_ok=True)
def _handoff_path_for(nonce: str) -> Path:
return _HANDOFF_DIR / f"grandchild-{nonce}.json"
def _pid_alive(pid: int) -> bool:
"""POSIX: send signal 0 to probe whether ``pid`` is still alive.
``os.kill(pid, 0)`` raises ``ProcessLookupError`` if the process is
gone, ``PermissionError`` if it exists but we can't signal it
(someone else's pid). We treat PermissionError as "alive" because
the process exists and that's all we need to know.
"""
if sys.platform == "win32": # pragma: no cover — POSIX-only test
# On Windows we'd use OpenProcess + GetExitCodeProcess; this
# test is skipped on Windows so the path is unreachable.
raise RuntimeError("_pid_alive POSIX-only")
try:
os.kill(pid, 0)
except ProcessLookupError:
return False
except PermissionError:
return True
return True
def test_progress_output_tolerates_legacy_stdout_encoding(tmp_path: Path) -> None:
"""Progress glyphs must not crash the runner on non-UTF-8 consoles."""
repo_root = Path(__file__).resolve().parent.parent
runner = repo_root / "scripts" / "run_tests_parallel.py"
probe_dir = tmp_path / "probe"
probe_dir.mkdir()
probe = probe_dir / "test_probe_smoke.py"
probe.write_text("def test_smoke():\n assert True\n", encoding="utf-8")
env = os.environ.copy()
env["PYTHONIOENCODING"] = "cp1252:strict"
proc = subprocess.run(
[
sys.executable,
str(runner),
"--paths",
str(probe_dir),
"-j",
"1",
"--file-timeout",
"30",
],
cwd=repo_root,
env=env,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
text=True,
timeout=60,
)
assert proc.returncode == 0, proc.stdout
assert "UnicodeEncodeError" not in proc.stdout
assert "1 tests passed" in proc.stdout
@pytest.mark.skipif(sys.platform == "win32", reason="POSIX-only probe")
@pytest.mark.live_system_guard_bypass
def test_grandchild_leak_is_killed_by_runner(tmp_path: Path) -> None:
"""Run the parallel runner over a probe file and verify cleanup.
1. Materialize a probe file that spawns a long-lived grandchild and
writes its PID to disk before exiting.
2. Invoke ``scripts/run_tests_parallel.py`` against the probe file.
3. Wait for the grandchild PID to vanish (poll for ~5s).
4. Assert the runner exited cleanly AND the grandchild is dead.
"""
repo_root = Path(__file__).resolve().parent.parent
runner = repo_root / "scripts" / "run_tests_parallel.py"
assert runner.exists(), f"runner missing at {runner}"
# Probe lives in a temp dir, NOT under tests/, so the regular suite
# never picks it up — only our explicit invocation does.
probe_dir = tmp_path / "probe"
probe_dir.mkdir()
probe = probe_dir / "test_probe_leaker.py"
nonce = f"{os.getpid()}-{int(time.time() * 1000)}"
handoff = _handoff_path_for(nonce)
if handoff.exists():
handoff.unlink()
probe_src = textwrap.dedent(f"""
import json, os, subprocess, sys, time
from pathlib import Path
HANDOFF = Path({str(handoff)!r})
def test_spawns_grandchild_and_walks_away():
# Long-lived grandchild: detached, ignores SIGTERM (we want
# SIGKILL or process-group kill to be the only thing that
# works, simulating a misbehaving server).
child = subprocess.Popen(
[
sys.executable, "-c",
"import os, signal, sys, time; "
"signal.signal(signal.SIGTERM, signal.SIG_IGN); "
"sys.stdout.write(f'gc-pgid={{os.getpgid(0)}} gc-pid={{os.getpid()}}\\\\n'); "
"sys.stdout.flush(); "
"time.sleep(600)",
],
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
# IMPORTANT: do NOT pass start_new_session here. We want
# the grandchild to inherit the pytest subprocess's
# process group, so when the runner kills the group the
# grandchild dies too.
)
# Read the first line so we can record gc's pgid in the
# handoff, then walk away — don't close the pipe (would
# signal EOF and let the child see SIGPIPE on next write).
first_line = child.stdout.readline().decode().strip()
HANDOFF.write_text(json.dumps({{
"pid": child.pid,
"diag": first_line,
"test_pid": os.getpid(),
"test_pgid": os.getpgid(0),
}}))
assert child.pid > 0
""").strip()
probe.write_text(probe_src + "\n")
# Run the parallel runner against just the probe file. The runner
# discovers under ``tests/`` by default, so we override via --paths.
proc = subprocess.run(
[
sys.executable,
str(runner),
"--paths",
str(probe_dir),
"-j",
"1",
# Tight per-file timeout: the probe finishes in <1s, no
# need for 10min.
"--file-timeout",
"30",
],
cwd=repo_root,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
# The runner declares its stdio UTF-8 (see _make_stdio_glyph_safe);
# decode the same way so ✓-glyph assertions hold on Windows, where
# text=True alone would decode with the locale codec (cp1252).
encoding="utf-8",
errors="replace",
timeout=60,
)
assert handoff.exists(), (
f"probe never wrote handoff file; runner output:\n{proc.stdout}"
)
handoff_data = json.loads(handoff.read_text())
grandchild_pid = handoff_data["pid"]
diag = handoff_data.get("diag", "(no diag)")
test_pid = handoff_data.get("test_pid")
test_pgid = handoff_data.get("test_pgid")
handoff.unlink()
# The runner must have exited cleanly (probe test passes).
assert proc.returncode == 0, (
f"runner exited {proc.returncode}; output:\n{proc.stdout}"
)
# The grandchild must be gone. Poll for a bit because process-group
# SIGKILL + reaping isn't synchronous; on a loaded box it can take
# a beat.
deadline = time.monotonic() + 5.0
while time.monotonic() < deadline:
if not _pid_alive(grandchild_pid):
break
time.sleep(0.05)
else:
# Test cleanup: kill the leaked grandchild ourselves so a
# FAILED assertion doesn't leave a sleep(600) running.
try:
os.kill(grandchild_pid, 9)
except ProcessLookupError:
pass
pytest.fail(
f"grandchild PID {grandchild_pid} survived runner exit; "
f"diag={diag!r} test_pid={test_pid} test_pgid={test_pgid}; "
f"runner output:\n{proc.stdout}"
)
# ── Bare pytest-flag passthrough ─────────────────────────────────────────────
#
# The runner routes any token starting with ``-`` that isn't one of its own
# options (``-j``/``--jobs``, ``--paths``, ``--slice``, ``--file-timeout``,
# ``--generate-slices``, ``--files``, ``--include-integration``) straight
# through to each per-file pytest invocation — no ``--`` separator required.
# Before this, a bare ``-q`` errored out with "unrecognized arguments",
# forcing a retry on every run. These tests are behavior contracts, not
# snapshots: they assert that bare flags reach pytest and that value-taking
# flags (``-k expr``) keep their value instead of having it stolen by the
# positional-path discovery.
def _make_probe_dir(tmp_path: Path) -> Path:
"""Two trivial passing tests, one named test_alpha, one test_beta."""
probe_dir = tmp_path / "probe"
probe_dir.mkdir()
(probe_dir / "test_flagprobe.py").write_text(
"def test_alpha():\n assert True\n\n"
"def test_beta():\n assert True\n"
)
return probe_dir
def _run_runner(probe_dir: Path, *extra: str) -> subprocess.CompletedProcess:
repo_root = Path(__file__).resolve().parent.parent
runner = repo_root / "scripts" / "run_tests_parallel.py"
return subprocess.run(
[sys.executable, str(runner), "--paths", str(probe_dir),
"-j", "1", "--file-timeout", "30", *extra],
cwd=repo_root,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
# The runner declares its stdio UTF-8 (see _make_stdio_glyph_safe);
# decode the same way so ✓-glyph assertions hold on Windows, where
# text=True alone would decode with the locale codec (cp1252).
encoding="utf-8",
errors="replace",
timeout=60,
)
def test_bare_value_flag_keeps_its_value(tmp_path: Path) -> None:
"""``-k test_alpha`` reaches pytest as a selector, not as a path.
The value token (``test_alpha``) must NOT be swallowed by the runner's
positional-path discovery — if it were, discovery would look for a path
named ``test_alpha``, find nothing, and the run would degrade. We assert
the run succeeds AND only one of the two tests was selected (proving the
``-k`` filter actually applied inside pytest).
"""
probe_dir = _make_probe_dir(tmp_path)
proc = _run_runner(probe_dir, "-k", "test_alpha")
assert proc.returncode == 0, proc.stdout
# Exactly one test selected: the per-file summary shows "1✓" (1 passed).
# test_beta is deselected by the -k filter.
assert "1✓" in proc.stdout or "1 passed" in proc.stdout, proc.stdout
assert "2✓" not in proc.stdout, (
f"both tests ran — -k filter did not apply:\n{proc.stdout}"
)
def test_positional_path_not_treated_as_flag(tmp_path: Path) -> None:
"""A positional path arg still overrides discovery (not routed to pytest)."""
probe_dir = _make_probe_dir(tmp_path)
repo_root = Path(__file__).resolve().parent.parent
runner = repo_root / "scripts" / "run_tests_parallel.py"
# Pass the probe dir positionally (no --paths), plus a bare -q.
proc = subprocess.run(
[sys.executable, str(runner), str(probe_dir), "-j", "1",
"--file-timeout", "30", "-q"],
cwd=repo_root, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
encoding="utf-8", errors="replace", timeout=60,
)
assert proc.returncode == 0, proc.stdout
# Discovery found the probe file (2 tests), proving the positional path
# was consumed as a root, not forwarded to pytest as a bad flag.
assert "test_flagprobe.py" in proc.stdout, proc.stdout
def test_file_retry_self_heals_and_prints_both_attempts(tmp_path: Path) -> None:
"""A pass-on-retry is green, loud, and retains the failing traceback."""
repo_root = Path(__file__).resolve().parent.parent
runner = repo_root / "scripts" / "run_tests_parallel.py"
marker = tmp_path / "ran-once"
probe = tmp_path / "test_flaky_probe.py"
probe.write_text(
textwrap.dedent(
f"""
from pathlib import Path
def test_flaky_once():
marker = Path({str(marker)!r})
if not marker.exists():
marker.write_text("failed once")
assert False, "simulated first-attempt flake"
assert True
"""
),
encoding="utf-8",
)
proc = subprocess.run(
[
sys.executable,
str(runner),
"--files",
str(probe),
"--file-retries",
"1",
"-j",
"1",
"-q",
],
cwd=repo_root,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
text=True,
timeout=60,
)
assert proc.returncode == 0, proc.stdout
assert "FLAKY file" in proc.stdout
assert "simulated first-attempt flake" in proc.stdout
assert "first-attempt output" in proc.stdout
assert "retry output" in proc.stdout
# ---------------------------------------------------------------------------
# Zero-collection is not a pass; node ids are translated, not dropped.
#
# Both behaviors were real foot-guns: a run where NOTHING was collected printed
# "0 tests passed, 0 failed (100% complete)" (reads green), and a pytest node id
# (`file.py::Class::test`) was silently discarded by path discovery so the run
# ended with "No test files to run" while looking like an accepted selector.
def test_zero_collected_across_run_fails_and_says_so(tmp_path: Path) -> None:
"""A -k that matches nothing must FAIL, not report a green summary."""
probe_dir = _make_probe_dir(tmp_path)
proc = _run_runner(probe_dir, "-k", "zzz_matches_nothing")
assert proc.returncode == 1, proc.stdout
assert "NO TESTS RAN" in proc.stdout
assert "NOT a pass" in proc.stdout
def test_node_id_selector_runs_the_named_test(tmp_path: Path) -> None:
"""``file.py::test_alpha`` runs that test instead of discovering nothing."""
probe_dir = _make_probe_dir(tmp_path)
target = probe_dir / "test_flagprobe.py"
repo_root = Path(__file__).resolve().parent.parent
proc = subprocess.run(
[sys.executable, str(repo_root / "scripts" / "run_tests_parallel.py"),
f"{target}::test_alpha", "-j", "1", "--file-timeout", "30"],
cwd=repo_root, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
text=True, timeout=60,
)
assert proc.returncode == 0, proc.stdout
assert "No test files to run" not in proc.stdout
assert "node id" in proc.stdout # explains the translation
# Ran exactly the one selected test, not both in the file.
assert "1 tests passed" in proc.stdout
def test_explicit_k_wins_over_node_id_inference(tmp_path: Path) -> None:
"""A caller's own ``-k`` is not overridden by the node-id translation."""
probe_dir = _make_probe_dir(tmp_path)
target = probe_dir / "test_flagprobe.py"
repo_root = Path(__file__).resolve().parent.parent
proc = subprocess.run(
[sys.executable, str(repo_root / "scripts" / "run_tests_parallel.py"),
f"{target}::test_alpha", "-k", "test_beta",
"-j", "1", "--file-timeout", "30"],
cwd=repo_root, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
text=True, timeout=60,
)
# -k test_beta wins: one test ran, and it wasn't filtered to nothing.
assert proc.returncode == 0, proc.stdout
assert "1 tests passed" in proc.stdout
def test_multiple_absolute_paths_split_on_pathsep(tmp_path: Path) -> None:
"""``--paths`` accepts ``os.pathsep``-joined absolute paths.
On Windows the absolute paths contain drive-letter colons, so a naive
``split(":")`` shreds them into phantom roots and only one (or neither)
of the two probe dirs would be discovered.
"""
dir_a = _make_probe_dir(tmp_path)
dir_b = tmp_path / "probe_b"
dir_b.mkdir()
(dir_b / "test_flagprobe_b.py").write_text(
"def test_gamma():\n assert True\n"
)
repo_root = Path(__file__).resolve().parent.parent
runner = repo_root / "scripts" / "run_tests_parallel.py"
proc = subprocess.run(
[sys.executable, str(runner),
"--paths", os.pathsep.join([str(dir_a), str(dir_b)]),
"-j", "1", "--file-timeout", "30", "-q"],
cwd=repo_root, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
encoding="utf-8", errors="replace", timeout=60,
)
assert proc.returncode == 0, proc.stdout
assert "Discovered 2 test files" in proc.stdout, proc.stdout
@pytest.mark.skipif(sys.platform != "win32", reason="drive-letter paths")
def test_drive_letter_colon_is_not_a_path_separator(tmp_path: Path) -> None:
"""An absolute ``--paths`` value stays one root on Windows.
The naive split used to produce a phantom relative root ``'C'`` (the
drive letter) alongside the real path; discovery only worked by the
accident of ``repo_root / '\\rooted\\rest'`` re-anchoring onto the
repo's drive.
"""
probe_dir = _make_probe_dir(tmp_path)
proc = _run_runner(probe_dir, "-q")
assert proc.returncode == 0, proc.stdout
drive = str(probe_dir)[0]
assert f"['{drive}', " not in proc.stdout, (
f"drive letter split off as a phantom root:\n{proc.stdout}"
)
assert "Discovered 1 test files" in proc.stdout, proc.stdout

View File

@@ -1,69 +0,0 @@
"""The runner's status glyphs must not crash narrow console encodings.
On native Windows, piped or legacy-console stdio defaults to cp1252, which
cannot encode the runner's ✓/✗ progress glyphs — before the fix, the first
per-file status line killed the whole run with UnicodeEncodeError. The
failure depends only on the stream's encoding, so these tests pin it on
every OS by building a cp1252 stream explicitly.
"""
from __future__ import annotations
import importlib.util
import io
import sys
from pathlib import Path
REPO_ROOT = Path(__file__).resolve().parents[1]
_RUNNER_PATH = REPO_ROOT / "scripts" / "run_tests_parallel.py"
def _load_runner():
spec = importlib.util.spec_from_file_location("run_tests_parallel", _RUNNER_PATH)
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)
return mod
def _cp1252_stream() -> tuple[io.TextIOWrapper, io.BytesIO]:
raw = io.BytesIO()
return io.TextIOWrapper(raw, encoding="cp1252", errors="strict"), raw
def test_cp1252_stream_reproduces_the_crash_without_the_fix() -> None:
# Baseline for the bug: a strict cp1252 stream cannot take the glyph.
stream, _raw = _cp1252_stream()
try:
stream.write("✓")
except UnicodeEncodeError:
return
raise AssertionError("expected UnicodeEncodeError on strict cp1252")
def test_glyph_safe_stdio_survives_cp1252(monkeypatch) -> None:
mod = _load_runner()
stream, raw = _cp1252_stream()
monkeypatch.setattr(sys, "stdout", stream)
monkeypatch.setattr(sys, "stderr", stream)
mod._make_stdio_glyph_safe()
print("✓ tests/foo.py (3 tests, 1.2s) ✗")
sys.stdout.flush()
out = raw.getvalue()
assert "✓".encode("utf-8") in out, "stream should now carry UTF-8 glyphs"
assert b"tests/foo.py (3 tests, 1.2s)" in out, "line content must survive"
def test_glyph_safe_stdio_noop_without_reconfigure(monkeypatch) -> None:
# Streams without .reconfigure (e.g. pytest's capture buffers, plain
# StringIO) must pass through untouched instead of raising.
mod = _load_runner()
plain = io.StringIO()
monkeypatch.setattr(sys, "stdout", plain)
monkeypatch.setattr(sys, "stderr", plain)
mod._make_stdio_glyph_safe()
print("✓ still fine")
assert "✓ still fine" in plain.getvalue()

View File

@@ -125,7 +125,7 @@ scripts/run_tests.sh tests/path/to/test_file.py::test_name --trace
scripts/run_tests.sh tests/path/to/test_file.py --showlocals --tb=long
```
Note: `scripts/run_tests.sh` runs each test file in a captured subprocess via `run_tests_parallel.py` (no xdist), so interactive pdb does NOT work under the wrapper. Run pytest directly for `--pdb`:
Note: `scripts/run_tests.sh` runs pytest under xdist workers, so interactive pdb does NOT work under the wrapper. Run pytest directly for `--pdb`:
```bash
source .venv/bin/activate