Commit Graph

59 Commits

Author SHA1 Message Date
ethernet
c9bd7459c2 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	scripts/run_tests_parallel.py
#	tui_gateway/model_switch.py
2026-09-22 00:28:30 -04:00
teknium1
439eb0395e fix(tests): relocated pytest basetemps stop piling up in $HOME; runner sweeps roots killed runs left
Two leaks from the test temp plumbing:

tests/conftest.py relocates pytest's basetemp out of the native Hermes
home (#111101) with mkdtemp(dir=native.parent), which is the operator's
$HOME, and nothing removed it: 123 hermes-pytest-basetemp-* dirs (552 MB)
appeared there in a day, one per bare pytest process. The relocated
basetemp now goes into one prunable root (/var/tmp/hermes-pytest on
POSIX, a non-dotted sibling of the native home elsewhere), is removed at
pytest_unconfigure, and idle siblings from killed runs are swept on entry.

scripts/run_tests_parallel.py deletes each per-file temp root in finally,
but a SIGKILLed runner (tool timeout, stray pkill) never gets there and
leaks one root per in-flight worker: 983 r-* roots (3.4 GB) in three
days. The runner now sweeps 24h-idle roots at start and forces read-only
permission fixtures writable before rmtree instead of skipping them.
2026-09-21 20:05:27 -07:00
ethernet
4fda664b0f tests: resolve platforms() specs in the lane selector and runner note
scripts/ci/list_os_marked_tests.py matched the literal lane word inside a
platforms() string, so 80 files gated platforms("posix") never reached the
macOS lane ("posix = linux or macOS" was false for CI), and "any" files
reached none. The selector now resolves specs the way the conftest gate
does (posix ⊇ linux+macos, any ⊇ all, "not X" admits the rest) and only
looks inside mark.platforms(...) calls, dropping a false positive whose
only "macos" was inside a generated-file string.

scripts/run_tests_parallel.py still grepped the retired linux_only/
macos_only/windows_only names, so its "N files SKIPPED on this host" note
had been silent since the migration. It now shares the selector's
resolver and names the spec and the lane(s) it runs on.

Restore the platforms("windows") mark that
test_suppress_platform_ver_console_stubs_syscmd_ver lost in the
windows_only migration (its docstring still declared it); it passed
vacuously on Linux and was deselected on the Windows lane.
2026-09-21 18:53:24 -04:00
ethernet
14b3232cbd fix(test-runner): key the scratch root by user, not a shared literal
The runner's per-run temp roots live under a fixed `/var/tmp/hermes-pytest`, chosen for
real reasons (disk-backed, not hidden, short enough for AF_UNIX sun_path). But a fixed
literal in a world-writable sticky dir belongs to whoever creates it first: a root-owned
root — a container or system-service run — makes every later `makedirs`/`mkdtemp` there
fail with EPERM for every other user on the host, with no way back that does not need
root. That is exactly what happened on luna: /var/tmp/hermes-pytest is root:root 755, so
every local suite run by the login user died at `runner crashed: PermissionError(13)`
before collecting a single test.

Key the name by uid (with the same non-/var/tmp fallback), so no run can be blocked by
another user's leftovers. Invariant test: two uids never share a scratch root, and the
root is created under the expected parent.
2026-09-21 17:24:29 -04:00
ethernet
9f2ba1b74d merge origin/main (779 commits) into ethie/pm-clean
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).

Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.

uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
2026-09-21 00:58:39 -04:00
ethernet
3b181dd848 merge origin/main into ethie/pm-clean 2026-09-21 00:38:40 -04:00
liuhao1024
8828e356f7 fix(tests): let run_tests_parallel take the explicit file list from a file
--files carries the whole list as one argv element, and Linux caps a
single argument at MAX_ARG_STRLEN (128 KiB) - a much smaller limit than
ARG_MAX. The whole-suite list (~210 KB) dies with E2BIG in execve before
the runner's first line runs, so 'run the whole suite except one file'
cannot be expressed through --files at all.

Add --files-from PATH (or '-' for stdin), one path per line, mutually
exclusive with --files. A bare '-' after --files-from is normalized to
the '='-joined form because argparse treats '-' as a positional.
2026-09-20 10:51:40 -07:00
teknium1
6159bf4d87 runner: drop the RLIMIT_DATA worker cap
On the CI runner the pillow-heif HEIF encode in tests/tools/test_image_source.py hangs under
the cap (2/2 runs, 36/40 then 300 s timeout; 40/40 in 26 s on main). The wheel's encoder spins
on a failed allocation instead of erroring, so a heap cap on C code trades an OOM for a hang.
Keep the leak fix and the no-relaunch-on-kill rule.
2026-09-20 09:00:16 -07:00
teknium1
439ebe0ae9 fix(tests): stop the code_kernel reader-thread leak that OOM-killed test workers; cap worker heap
tests/tools/test_local_env_blocklist.py::TestPythonpathSelectiveStrip::
test_execute_code_composition_strips_inherited_hermes_entries hands code_kernel a MagicMock
process whose stdout/stderr only fake read(); the kernel drains with read1(), which on a bare
MagicMock never returns EOF. _stdout_reader died on `buf += chunk`, but _stderr_reader's
`while chunk := stderr.read1(4096)` spun forever appending mocks (each call growing
mock_calls) in a daemon thread that outlived the test: ~1 GB/min until the kernel killed the
worker. Five OOM incidents on this file (08-30, 09-13, 09-14, 09-16, 09-19), always blamed on
whichever test ran next. The fake now returns EOF from read1() as well.

Runner guardrails so the next runaway is a traceback, not a swap storm:
- each pytest worker runs under RLIMIT_DATA (8 GiB, Linux; HERMES_TEST_WORKER_MEM_GB, 0 = off).
  RLIMIT_AS is avoided on purpose: browsers spawned by tests reserve huge address space.
- a worker killed by signal or the file timeout is never --file-retries relaunched; a runaway
  relaunched while the first tree is still being reaped doubled the damage on 09-16.

Live: the file went from 20 min / 20 GB to 5.6 s / 140 MB. An allocate-forever probe dies with
MemoryError in 5 s; a SIGKILL'd worker launches once on this runner, twice on base.
2026-09-20 09:00:16 -07:00
ethernet
e1576d06a6 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
Resolved toward the branch: PM provisions uv/python (main's install.ps1 uv-shim
salvage + its test and workflow steps dropped), the shim re-exec stays retired,
package.json carries no electron-builder block (afterExtract identity stamp wired
into electron-builder.config.cjs instead; after-pack.mjs keeps signing only),
Desktop workspace-deps helpers stay retired. Main's scratch-dir bootstrap
(export_scratch_tmp_env) is taken and re-run after profile resolution.
2026-09-19 22:57:07 -04:00
teknium1
ece388eca8 fix: test-runner scratch root is /var/tmp/hermes-pytest; two tests follow TMPDIR
~/.cache/hermes-pytest failed seven files: a dot-dir ancestor made the hidden-dir search tests
see every fixture as hidden, and the longer root pushed the PulseAudio/voice AF_UNIX test
sockets past sun_path. /var/tmp is the FHS disk-backed temp root (never tmpfs), non-hidden,
and shorter than the old /tmp root. test_tool_result_storage asserts STORAGE_DIR instead of a
literal, and the zh-Hans bot-mode mirror follows the English code block it must copy.
2026-09-19 10:44:26 -07:00
teknium1
0746903679 fix: test-runner scratch lives in ~/.cache/hermes-pytest; env-less home fallback cannot raise
A runner root under HERMES_HOME/cache made conftest relocate every basetemp out of the live
Hermes home into ~/hermes-pytest-basetemp-*: deeper paths pushed AF_UNIX test sockets past
sun_path, and a profile home under $HOME renders as ~/… (which one test compared verbatim).
The root now sits in $XDG_CACHE_HOME/hermes-pytest (disk-backed, outside the Hermes home,
as short as the old /tmp root) and the test asserts the displayed form.

get_real_home()'s tempfile fallback raised RuntimeError on Windows when a child env carried
no HOME/USERPROFILE; it falls back to the old literal instead of crashing env construction.
2026-09-19 10:44:26 -07:00
teknium1
07df62d604 fix: sockets keep a short temp root; lint skips git-ignored artifacts; test runner scratch leaves /tmp
Chrome puts its SingletonSocket under $TMPDIR and AF_UNIX paths cap at 104/108 bytes, so a
deep HERMES_HOME (profile homes, test homes) made the new scratch TMPDIR kill Chrome at
startup ("Socket path too long") — two browser test files went red on the branch and green
on main. hermes_constants.socket_safe_tmpdir() keeps the scratch root when it fits and falls
back to the OS root for sockets only; the browser env and the code kernel RPC socket use it.

check_no_tmp_literals walked git-ignored runner artifacts (test_durations.json) and failed on
whatever the last test run wrote; it now skips `git ls-files --others --ignored` paths.

run_tests_parallel created its per-file temp roots in the system temp dir and exported no
TMPDIR, so a full-suite run wrote gigabytes of fixtures to tmpfs (3,225 leftover roots, 9.9 GB,
were sitting in /tmp on the dev box). Roots now live under HERMES_HOME/cache/scratch/pytest and
the test process inherits TMPDIR=<root>, so the existing cleanup removes every temp file.
2026-09-19 10:44:26 -07:00
ethernet
d70feca03d Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	hermes_cli/update_cmd.py
#	tests/hermes_cli/test_cmd_update.py
2026-09-19 01:14:17 -04:00
teknium1
0ddba07ad7 fix(ci): report an interpreter crash as CRASHED, not "no tests ran"
When a per-file pytest subprocess dies by signal (the sqlite cross-thread
close in #113186 was a SIGSEGV after every test had passed), faulthandler
prints "Fatal Python error: Segmentation fault" and no summary line, so
every count parses to 0. The runner filed that under "1 file where no
tests ran (collection/import error, ...)" beneath a summary that read
"0 failed" and exited 1 — two wrong diagnoses for one real bug, and it
was misread as a runner problem twice on main.

The runner now detects a signal death or a "Fatal Python error:" banner,
prefixes the captured output with the diagnosis (same convention as the
timeout path), marks the progress line CRASHED, counts "N files CRASHED"
on the summary line, lists the file in its own failure bucket, and no
longer trips the "NO TESTS RAN" guard for a crash that ran tests. The
flake retry already covers crashes (any non-zero rc), so nothing changes
there.
2026-09-18 19:41:51 -07:00
ethernet
a6ae6ace51 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.github/workflows/js-tests.yml
#	agent/model_metadata.py
#	apps/desktop/electron/main.ts
#	apps/desktop/scripts/bundle-electron-main.mjs
#	apps/desktop/src/app/settings/about-settings.tsx
#	apps/desktop/src/app/settings/gateway-settings.test.tsx
#	apps/desktop/src/app/settings/gateway-settings.tsx
#	apps/desktop/src/app/updates-overlay.tsx
#	gateway/shutdown_flush.py
#	hermes_bootstrap.py
#	hermes_cli/local_runtime/binaries.py
#	hermes_cli/main.py
#	hermes_cli/managed_uv.py
#	hermes_cli/update_cmd.py
#	hermes_cli/update_cmd_deps.py
#	hermes_cli/update_cmd_fleet.py
#	hermes_cli/update_cmd_maint.py
#	hermes_cli/update_receipt.py
#	hermes_cli/update_serve_obligations.py
#	hermes_constants.py
#	tests/hermes_cli/test_doctor.py
#	tests/hermes_cli/test_managed_uv.py
#	tests/hermes_cli/test_pending_supervisor_recovery.py
#	tests/hermes_cli/test_startup_fast_guards.py
#	tests/hermes_cli/test_update_desktop_stale_warning.py
#	tests/hermes_cli/test_update_fleet_restart_pending.py
#	tests/hermes_state/test_hermes_state.py
#	tests/tools/test_tirith_security.py
#	tools/bot_relay.py
#	tools/checkpoint_manager.py
#	tools/write_approval.py
#	website/docs/getting-started/updating.md
#	website/docs/reference/environment-variables.md
2026-09-18 17:26:10 -04:00
teknium1
49397cf2b4 fix(tests): parallel runner reports a known flag's missing value as usage, not per file
The bare-flag check asked pytest's parser which tokens it does not know,
but wrapped parse_known_args in the same blanket except that guards
parser construction. A known flag with a missing value (`--tb` alone)
raises pytest.UsageError there, which the except turned into "nothing
unknown", so discovery ran and every per-file pytest died with
"argument --tb: expected one argument".

Keep the fallback around building the parser only; let parse_known_args
run outside it and surface UsageError (and the unknown-token list) as
this runner's own usage error before discovery. One invariant test.
2026-09-18 10:23:21 -07:00
teknium1
d7bedcee1e fix(tests): parallel runner rejects unknown bare flags with usage instead of sweeping
A bare token this runner does not own used to be forwarded to every per-file
pytest, so a typo (`--jbs`, or `--help` before #114065's fix) discovered the
whole suite and each file died with "unrecognized arguments" — hours to learn
about a typo. Validate the bare passthrough tokens against pytest's own
argparse parser (installed plugins loaded), and fail once with this runner's
usage (exit 2) before discovery. argparse handles the attached-short-value
(`-rA`), combined-flag (`-xvs`) and `-k expr` forms, so real pytest flags keep
passing through; tokens after a literal `--` are the caller's explicit choice
and are never validated. If pytest's parser cannot be built, the check is
skipped and behaviour is unchanged.

Follow-up to KoNit-K's `-h`/`--help` interception (#114065). Fixes #114059.
2026-09-18 10:23:21 -07:00
KoNit-K
41332e7851 fix(tests): handle parallel runner help flags 2026-09-18 10:23:21 -07:00
ethernet
b4a294fff9 Merge origin/main; keep PM as plugin dependency owner
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
2026-09-17 13:52:05 -04:00
teknium1
d228013832 fix(ci): feed the timeout scaler only healthy durations, and wire the cache in CI
Greptile's two findings on the original PR were both right.

1. The scaler read test_durations.json from the checkout, but CI ran on
   a fresh runner where that file never exists (it is gitignored and the
   slicing-era artifact/merge job that produced it is gone). The feature
   was inert exactly where the false FLAKY kills happen. tests.yml now
   restores the most recent main-saved cache before the run (PRs read
   only) and saves it after a green push to main, mirroring the
   ci-timings-baseline restore/save pattern already in ci.yaml.

2. _save_durations persisted every file's total subprocess wall,
   including the ~cap of a timed-out attempt and the retry-summed wall
   of a FLAKY file. With the scaler that compounds: a hang cached at
   ~300s earns 900s next run, then ~900s cached earns 2700s, until the
   job timeout is the only bound. _clean_pass_durations drops failed and
   FLAKY files from the write so a file's cached duration is always a
   first-attempt-clean measurement; those files keep their previous
   known-good entry.

Tests trimmed to the salvage bar (<=2 invariants for the scaler plus one
for the cache filter) and moved next to the other runner tests under
tests/scripts/.
2026-09-15 03:47:55 -07:00
Teknium
9a69785790 fix(ci): scale per-file test timeout by cached duration to stop false FLAKY kills
The flat 300s --file-timeout SIGKILL'd known-slow large-collection
files when CI load dilated their runtime past the cap; the automatic
one-shot retry then passed, manufacturing a FLAKY report for a healthy
file. Seen 2026-08-18 on main run 32155223248's sibling PR runs:
tests/test_hermes_state.py (239 tests) killed at 300s on attempt 1,
passed in 205s on retry.

_effective_file_timeout() now gives each file
max(flat_cap, 3 x last cached duration) from test_durations.json.
The bound is only ever raised — genuinely hung files are still killed,
uncached files keep the flat cap, and --file-timeout/HERMES_TEST_FILE_TIMEOUT
semantics are unchanged.

Includes a sabotage-verified unit test (fails without the scaler).
2026-09-15 03:47:55 -07:00
teknium1
087c4f4f4e fix(tests): drop the RLIMIT_AS memory cap; the leak sweep is the fix
The per-process address-space cap was defense-in-depth on top of the
SessionDB leak sweep (nobody sets the knob; the sweep removes the leak).
On the 96-worker CI runner it was also the only PR-specific difference
when tests/tools/test_image_source.py hung to the 600s SIGKILL while the
same file passes in ~30s on every sibling branch: RLIMIT_AS counts virtual
reservations, and image/threading libraries reserve far more address
space than they touch. Keep the sweep, remove the cap and its env knob.
2026-09-15 03:44:50 -07:00
Teknium
d070e480a3 fix(tests): close leaked SessionDB handles suite-wide and cap pytest memory
Root cause of the 2026-08-16 OOM incidents (three runs of
`python -m pytest -o addopts= -q tests/hermes_cli/` ballooning to
16-25 GB RSS and getting killed): ~40 files under tests/hermes_cli/
construct SessionDB() directly and never close it. Each instance keeps
the writer connection (state.db + -wal fds), up to _READ_POOL_MAX pooled
readers with their SQLite page caches, and — once token accounting has
run — an atexit registration that pins the instance alive until
interpreter exit. In one process over 637 files those accumulate without
bound; the sanctioned per-file runner masks it, so CI never saw it.

Fix the class, not the sites:

* hermes_state: register every successfully constructed SessionDB in a
  test-only WeakSet (populated only when HERMES_TEST_ISOLATION is set,
  i.e. under this test suite; production never touches it).
* tests/conftest.py: autouse _close_leaked_session_dbs teardown closes
  everything left in the registry after each test. close() is idempotent
  and unregisters the pinning atexit hook, so instances become
  collectable.
* tests/conftest.py: session-scoped _pytest_memory_cap applies a
  defensive RLIMIT_AS of 12 GiB (Linux only) so any future in-process
  leak fails fast with MemoryError instead of eating the box.
  Overridable/disable-able via HERMES_PYTEST_MEM_CAP (documented in
  scripts/run_tests_parallel.py).
* tests/hermes_state/test_session_db_leak_sweep.py: behavior contract
  for registration, idempotent close, and the cross-test sweep.

Measured (capped single-process `pytest -o addopts= -q tests/hermes_cli/`):
peak RSS 4.16 GiB before -> 1.67 GiB after; per-test open .db fd count
previously climbed monotonically (0 -> 12 -> 17 -> 104 within the
SessionDB-heavy files), now stays bounded (<= 5, transient). Sanctioned
runner over the affected 35 files: 495 passed, 0 failed, no FLAKY.

Incident evidence: ~/.hermes/logs/oom-incidents/20260816-202114
(fd dumps show 100+ open state.db/state.db-wal handles across pytest
tmpdirs; 3rd recurrence that day).
2026-09-15 03:44:50 -07:00
ethernet
612d542281 Merge remote-tracking branch 'origin/main' into ethie/pm-clean
# Conflicts:
#	.gitignore
#	Dockerfile
#	agent/onboarding.py
#	apps/desktop/electron/main.ts
#	apps/desktop/electron/pool-stop.ts
#	apps/desktop/src/components/model-picker.test.tsx
#	apps/desktop/src/store/updates.ts
#	apps/desktop/vite.config.ts
#	datagen-config-examples/run_browser_tasks.sh
#	docs/rca-ssl-cacert-post-git-pull.md
#	gateway/run.py
#	hermes_cli/backup.py
#	hermes_cli/credential_lifecycle.py
#	hermes_cli/dashboard_procs.py
#	hermes_cli/doctor_state.py
#	hermes_cli/env_loader.py
#	hermes_cli/gateway_windows.py
#	hermes_cli/local_runtime/endpoint.py
#	hermes_cli/psutil_android.py
#	hermes_cli/update_cmd.py
#	hermes_cli/update_cmd_windows.py
#	hermes_cli/web_routers/local_models.py
#	hermes_cli/web_server_config.py
#	hermes_cli/web_server_cron.py
#	plugins/memory/hindsight/__init__.py
#	plugins/memory/holographic/__init__.py
#	plugins/memory/honcho/cli.py
#	plugins/memory/mem0/__init__.py
#	plugins/platforms/google_chat/oauth.py
#	plugins/platforms/photon/adapter.py
#	scripts/ci/list_os_marked_tests.py
#	scripts/run_tests.sh
#	tests/agent/test_compression_stall_fallback.py
#	tests/agent/test_create_openai_client_ssl_verify.py
#	tests/gateway/test_google_chat_oauth_dependencies.py
#	tests/hermes_cli/conftest.py
#	tests/hermes_cli/test_cli_init.py
#	tests/hermes_cli/test_gateway_migrate_multiplex.py
#	tests/hermes_cli/test_psutil_android_extract.py
#	tests/hermes_cli/test_relaunch.py
#	tests/hermes_cli/test_update_check.py
#	tests/hermes_cli/test_update_handoff_desktop_rebuild.py
#	tests/hermes_cli/test_worktree_gc.py
#	tests/scripts/desktop_update/test_desktop_update_windows_python_handoff.py
#	tests/scripts/desktop_update/test_desktop_update_windows_retry_policy.py
#	tests/scripts/desktop_update/test_desktop_update_windows_timestamp.py
#	tests/scripts/install/test_install_autostash_conflict_recovery.py
#	tests/scripts/install/test_install_clone_throttle_fallback.py
#	tests/scripts/install/test_install_commit_pin_rollback.py
#	tests/scripts/install/test_install_diverged_update.py
#	tests/scripts/install/test_install_lockfile_churn.py
#	tests/scripts/install/test_install_macos_launcher.py
#	tests/scripts/install/test_install_no_initial_commit.py
#	tests/scripts/install/test_install_ps1_ascii_only.py
#	tests/scripts/install/test_install_ps1_browser_install.py
#	tests/scripts/install/test_install_ps1_managed_node_swap.py
#	tests/scripts/install/test_install_ps1_native_stderr_eap.py
#	tests/scripts/install/test_install_ps1_node_path_for_npm.py
#	tests/scripts/install/test_install_ps1_python_fallback_venv.py
#	tests/scripts/install/test_install_ps1_resolver_strictmode.py
#	tests/scripts/install/test_install_ps1_uv_install_fallback.py
#	tests/scripts/install/test_install_ps1_uv_powershell_host.py
#	tests/scripts/install/test_install_ps1_venv_process_tree.py
#	tests/scripts/install/test_install_ps1_venv_recreate_safety.py
#	tests/scripts/install/test_install_ps1_venv_rename_abort.py
#	tests/scripts/install/test_install_ps1_venv_transaction_boundary.py
#	tests/scripts/install/test_install_ps1_web_server_syntax_probe.py
#	tests/scripts/install/test_install_scripts_computer_use.py
#	tests/scripts/install/test_install_sh_acp_launcher.py
#	tests/scripts/install/test_install_sh_bootstrap_marker.py
#	tests/scripts/install/test_install_sh_browser_install.py
#	tests/scripts/install/test_install_sh_install_method_stamp.py
#	tests/scripts/install/test_install_sh_node_deps_failure.py
#	tests/scripts/install/test_install_sh_node_deps_workspaces.py
#	tests/scripts/install/test_install_sh_node_global_prefix.py
#	tests/scripts/install/test_install_sh_node_npm_check.py
#	tests/scripts/install/test_install_sh_node_prerelease.py
#	tests/scripts/install/test_install_sh_node_probe.py
#	tests/scripts/install/test_install_sh_node_tarball_without_xz.py
#	tests/scripts/install/test_install_sh_pythonpath_sanitization.py
#	tests/scripts/install/test_install_sh_reuse_supported_python.py
#	tests/scripts/install/test_install_sh_root_fhs_uv_python_path.py
#	tests/scripts/install/test_install_sh_setup_wizard_tty_probe.py
#	tests/scripts/install/test_install_sh_symlink_stomp.py
#	tests/scripts/install/test_install_sh_termux_network_prereqs.py
#	tests/scripts/install/test_install_sh_termux_python_bounds.py
#	tests/scripts/install/test_install_sh_uv_lock_config.py
#	tests/scripts/install/test_install_unmerged_index.py
#	tests/scripts/test_run_tests_parallel.py
#	tests/test_managed_runtime_resolution.py
#	tests/test_project_metadata.py
#	tests/tools/test_browser_use_cli.py
#	tests/tools/test_tts_pythonpath_fallback.py
#	tests/tui_gateway/test_hosted_room_driver_runtime.py
#	tests/tui_gateway/test_tui_gateway_server.py
#	tools/lazy_deps.py
#	tools/voice_mode.py
#	uv.lock
#	website/docs/developer-guide/macos-bundle-updates.md
#	website/docs/developer-guide/pm-audit-status.md
#	website/docs/developer-guide/shared-bundle-builds.md
#	website/docs/developer-guide/source-update-completion.md
#	website/docs/developer-guide/stable-releases.md
2026-09-14 15:38:34 -04:00
teknium1
d10bb2ab6f test: make tests/ mirror the source tree; drop issue numbers from filenames
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.

Parallel directories for one source package, folded into the mirror:
  tests/acp        -> tests/acp_adapter   (its __init__/conftest move with it)
  tests/cli        -> tests/hermes_cli    (prompt_toolkit fixture merged into
                                           hermes_cli/conftest.py)
  tests/run_agent  -> tests/agent         (backoff fixture becomes
                                           agent/conftest.py)
  tests/relay      -> tests/gateway/relay
  tests/state      -> tests/hermes_state

246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.

Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.

Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).

Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
2026-09-13 09:18:02 -07:00
ethernet
92686159d1 fix(pm): integrate audited runtime and lifecycle repairs
Prepare dependency generations before selecting them. Keep shipped tool
bytes separate from writable additions, and store facts beside their entries.
Validate proposed plugin sets before config publication. Restore the previous
config if the facts write fails.

Consolidate duplicate updater, backup, setup, and voice helpers. Repair
launcher selection, dependency consumers, download ownership, update feeds,
and native Windows process and file handling.

Verification: 206 changed/prior-failing Python files reported 4630 passed,
one failed, and 330 skipped. Fix the remaining Hindsight fixture boundary.
The final targeted rerun reported 234 passed and two skipped. The store
review regression batch reported 83 passed and one skipped. Desktop
TypeScript checks, 56 selected Electron tests, 24 release tests, and the
removed-import/compatibility guards passed.

This is an integration checkpoint, not full audit acceptance. The complete
Python suite has not run on this fixed tree. Crash-atomic plugin publication,
generation cleanup, receipt correlation, and packaged lifecycle acceptance
remain open in docs/pm-audit-status.md.
2026-09-05 22:36:48 -04:00
ethernet
e97de8d8c3 Merge branch 'ethie/windows-tests' into ethie/pm-clean
# Conflicts:
#	.gitignore
#	agent/deadline.py
#	pyproject.toml
#	scripts/run_tests.sh
#	tests/agent/lsp/test_install_and_lint_fixes.py
#	tests/cron/test_file_permissions.py
#	tests/cron/test_media_delivery_parity.py
#	tests/cron/test_script_claim_heartbeat.py
#	tests/hermes_cli/test_dep_ensure.py
#	tests/hermes_cli/test_gateway_wsl.py
#	tests/hermes_cli/test_linux_desktop_entry.py
#	tests/hermes_cli/test_npm_engine.py
#	tests/hermes_cli/test_runtime_repair.py
#	tests/hermes_cli/test_tui_npm_install.py
#	tests/test_hermes_constants.py
#	tests/test_install_autostash_conflict_recovery.py
#	tests/test_install_macos_launcher.py
#	tests/test_install_sh_acp_launcher.py
#	tests/test_install_sh_bootstrap_marker.py
#	tests/test_install_sh_node_deps_failure.py
#	tests/test_install_sh_symlink_stomp.py
#	tests/test_windows_subprocess_no_window_flags.py
#	tests/tools/test_approval_timeout_overflow.py
#	tests/tools/test_browser_homebrew_paths.py
#	tests/tools/test_find_shell.py
#	tests/tools/test_lazy_deps.py
#	tests/tools/test_lazy_deps_durable_target.py
#	tools/approval.py
#	uv.lock
2026-09-04 01:06:06 -04:00
ethernet
afd64bd374 refactor(tests): host-dispatched runner — per-file isolation on POSIX, xdist on Windows
xdist-as-standard was the wrong shape for Linux: the per-file subprocess
model ran the full suite there in 3m45s with zero cross-file failures,
while the first xdist run took 7m51s and failed 111 tests. The two
costs are structural opposites:

  * xdist multiplies full-tree collection by the worker count — every
    worker imports the whole suite and pays pytest's per-item
    fixture-closure machinery (profiled: ~98M function calls to
    collect 42k items against the conftest's autouse fixtures —
    _matchfactories alone 1.37M calls, traverse_fixture_closure 783k).
    On the 96-core runner that front-loads 106s before any test runs.
  * The per-file model pays a spawn+import wall per file: ~15ms on
    Linux (nothing), 0.5-1.5s on Windows (~6-min floor across 3400
    files — the dominant cost of that lane).

So scripts/run_tests.sh now dispatches on host: run_tests_parallel.py
(restored verbatim from b8696aedea^, with its two self-tests) on POSIX,
pytest-xdist --dist loadfile on Windows. Both paths share one hermetic
env contract (env -i scrub, TZ=UTC, PYTHONHASHSEED=0, venv probing,
bytecode pre-compile). -j/HERMES_TEST_WORKERS feeds either backend.

The 111 xdist co-scheduling failures on Linux were the per-file
model's isolation guarantee surfacing as test bugs — with the POSIX
lane back on per-file, that class disappears where it never existed;
the Windows lane keeps loadfile (with its remaining co-scheduling
hazards as fix-at-the-test work).

Workflow comments, AGENTS.md, CONTRIBUTING.md, conftest comments and
the classify test follow the dual-path truth.
2026-08-31 17:53:45 -04:00
ethernet
ef4a82eed0 refactor(tests): xdist is the standard runner; drop the per-file subprocess machinery
scripts/run_tests.sh now runs pytest-xdist -n <N> --dist loadfile as
the single canonical path on every OS (Linux and Windows CI lanes both
use it). The per-file subprocess model (run_tests_parallel.py, and the
interim run_xdist.sh experiment) is deleted along with its two
self-tests: persistent xdist workers pay the interpreter+import wall
once per worker instead of once per file (~0.5-1.5s x ~3400 files was
a ~6-minute floor on Windows), and --dist loadfile pins a file's tests
to ONE worker, bounding state pollution to co-scheduled files — which
is exactly the class of flake we are now committing to fix properly.

The serial process-killer quarantine phase is dropped too. It existed
to keep process-tree-sweep tests from killing sibling xdist workers;
the durable fix belongs in those tests (sweeps must target their own
children, not enumerate every python process), and keeping a
divergent two-phase path would hide that work.

Kept from the old wrapper: hermetic env -i scrubbing, Windows
location-var forwarding, venv probing, bytecode pre-compile,
-m 'not integration', and the HERMES_TEST_IMAGE docker-knob
allowlist. HERMES_TEST_FILE_TIMEOUT/FILE_RETRIES/SLICE go away with
the runner they parameterized; -j/HERMES_TEST_WORKERS now map to xdist
-n (Linux CI pins 96, Windows 32, default auto).

Docs updated to match: AGENTS.md (runner contract, flake policy,
isolation section), CONTRIBUTING.md, tests/conftest.py comments,
classify_changes docstring + its lane expectation (a .sh runner no
longer trips the supply-chain scan lane), comfyui README,
hermes-agent contributor guide, debugpy skill and its website doc.
2026-08-31 17:52:32 -04:00
ethernet
9a81c291b9 fix(registries): reads resolve None to the active home scope; calm windows CI
normalize_scope unified the two scope contracts that live on the same
registries, which was wrong: WRITE paths (register/snapshot/restore)
treat None as the process-global layer, but READ paths (get_provider/
list_providers/registry_generation) treat None as 'resolve to the
active home's scope' — hermes_home_key(None) is the active default home.
Routing reads through normalize_scope hid every scoped registration
from ambient-home lookups: 52 linux failures across plugin discovery,
secret-source profile isolation, web-search/browser backend selection
(the CI run on e66a627aa5).

- read sites go back to hermes_home_key(scope); write/slot sites keep
  normalize_scope (None preserved); the docstring now spells out both
  contracts so the unification bug can't be reintroduced blind
- run_tests_parallel.py default worker count: cpu_count*2 -> cpu_count.
  Oversubscription made the 32-core windows runner run 64 pytest
  subprocesses; the suite's own linux measurement showed one worker per
  core is the flat optimum, and windows CI showed the cost — 160 files
  past the 300s per-file cap (they pass in seconds on an idle box), each
  burn a kill + retry cycle, and the churn dominated the lane's wall
  clock. docker.yml's explicit nproc pin now just documents its own
  intent instead of fighting a *2 default
- windows CI lane: HERMES_TEST_FILE_TIMEOUT=900 (via a matrix field)
  so contended-but-passing files stop getting killed at 300s
2026-08-31 17:52:32 -04:00
ethernet
9d5fca2488 ci: run the full Python suite on Windows and Linux, macOS runs macos_only
Fold the OS-specific lanes into tests.yml's test job as a three-OS matrix.
Linux and Windows run the full suite (conftest skips foreign-OS markers).
macOS keeps macos_only for now. Delete tests-os.yml.
2026-08-31 17:52:31 -04:00
ethernet
6e294f543d fix: read source files with encoding=utf-8-sig
Read text files with the encoding utf-8-sig so a BOM at the start of a
file does not cause a Unicode decode error (Windows editors add BOMs).

Reconstructed from ethie/pm commits 48a32b135b + 013219e814 onto the
current upstream/main base: only the utf-8 -> utf-8-sig transforms were
carried (370 exact line pairs across 205 files); pm-rename hunks that
rode in the original commit were left to the pm-store commit, and
utf8sig hunks entangled with content changes ride their owning commit.

Rebuilt on ethie/pm-clean off ac6c8028e0 (upstream/main).
2026-08-29 21:23:00 -04:00
ethernet
969094e4d2 fix(tests): remove four shared-state and lifetime faults at high concurrency
The suite now runs as one job with high per-file concurrency. Four tests
depend on state that they share with their siblings, or on a timer that
outlives them. That was safe at 8 workers. It is not safe at 96 or more.
Runs 32547184159 and 32551746525 show them.

1. Every pytest subprocess shared one temp root.

pytest puts tmp_path under <temproot>/pytest-of-<user>/. At the end of a
session it walks that directory with cleanup_dead_symlinks(). The walk lists
the directory. Then it asks whether the `pytest-current` symlink resolves.
Then it unlinks the symlink. A second process replaces that symlink between
the question and the unlink. The first process then raises FileNotFoundError
after all of its tests passed. Two files failed this way and passed on retry.

scripts/run_tests_parallel.py now gives each subprocess its own temp root
through PYTEST_DEBUG_TEMPROOT, and deletes it after the attempt. No two
processes share a directory. The race has no shared object to act on.

Proof: a direct driver of _pytest.pathlib.cleanup_dead_symlinks against one
root, with a second thread that replaces the symlink, raises the same
FileNotFoundError on 'pytest-current' as CI. A private root for each
subprocess removes that condition. A separate check confirms that 5
subprocesses receive 5 distinct roots, that tmp_path lands inside the private
root, and that no root survives the attempt.

2. The config read guard walked directories that other tests were writing.

tests/hermes_cli/test_config_read_guard.py scanned the tree with rglob. rglob
descends into every directory and filters after that, so it calls scandir() on
__pycache__ trees that the guard never inspects. Sibling processes create and
delete those entries during the run. A directory that disappears in the middle
of a walk raises FileNotFoundError out of rglob.

The scan now uses os.walk. It prunes excluded directories before it descends,
and it ignores a directory that disappears. __pycache__ joins the excluded
set, because bytecode is not source.

The guard still catches what it exists to catch. With a planted raw
yaml.safe_load of config.yaml in hermes_cli/, the test fails and names the
planted file. With a clean tree it passes.

3. A PTY test waited for a file to exist, and not for its content.

tests/tools/test_process_registry_write_stdin_surrogates.py spawns a child
that runs open(out,'wb').write(sys.stdin.buffer.readline()). open() creates
the file empty. The bytes arrive only after the PTY delivers the line. The
wait stopped at out.exists(), which the empty file already satisfies, so the
read returned b'' when the parent won that gap. This test failed both attempts
in CI, and did not pass on retry.

The test now waits for the expected bytes, with a bounded deadline.

Proof: the old wait loses 6 times in 25 runs on an idle 16-core machine. The
new wait loses 0 times in 25.

4. A dialog close timer outlived the test that started it.

ConfirmDialog holds the "done" beat for 600ms after a successful confirm, then
calls onClose. The timer had no cleanup, so an unmount inside that window left
it armed. It then called onClose on a tree that is gone, which reaches
setState in the parent. vitest can tear the environment down first, and React
then reads `window` during the update:

    ReferenceError: window is not defined
     at resolveUpdatePriority (react-dom-client.development.js:1308)
     at dispatchSetState
     at Timeout.t4 [as _onTimeout] session-actions-menu.tsx:574

The frame at session-actions-menu.tsx:574 is the `onClose` prop of
DeleteSessionDialog. The owner of the timer is ConfirmDialog, which now keeps
the handle in a ref and clears it on unmount.

Zoomable had the same fault, with a 1500ms timer that clears a "copied" flag.
copy-button.tsx and tooltip.tsx already clear their timers.

Proof: a new test confirms, unmounts inside the 600ms window, then advances
the clock. Against the old code it fails with "expected onClose to not be
called at all, but actually been called 1 times". Against the new code it
passes.

Verification:
- The affected Python files and the tests of the runner itself pass under
  scripts/run_tests.sh.
- The desktop ui suite passes: 566 files, 5382 tests, and no
  "window is not defined".
- eslint reports 0 errors on apps/desktop. The 118 warnings are the state
  before this change. The two cleanup effects carry an eslint-disable line for
  the ref-mirror rule. They write a timer handle, and not a mirror of a
  reactive value. The rule permits this, and its own comment names the case.
- The PTY test cannot run on the NixOS development machine. That machine has
  no python3 outside the nix store, and the test uses the literal `python3`.
  The child exits 127 there. The fix rests on the 25-run measurement above and
  on CI.
2026-08-22 02:25:12 -04:00
ethernet
641d254db4 ci: add macos and windows test lanes for the os-marked tests
the markers from the previous commit skip off-host. without a host to
run them on, every marked test is a silent skip. this commit adds the
hosts.

- tests-os.yml runs -m macos_only on macos-latest and -m windows_only
  on windows-latest. ci.yml requires both lanes in all-checks-pass.
- a lane fails on pytest exit code 5 (zero tests selected). a renamed
  marker cannot produce a green job that ran nothing.
- each lane repeats 'not integration' because a command-line -m
  replaces the addopts filter.
- scripts/ci/list_os_marked_tests.py selects which files each lane
  imports. -m filters after collection, and collection imports every
  module. without this helper, one unrelated ImportError on the
  foreign host fails a job whose own tests passed. the helper exits
  non-zero when a marker matches no file, and writes bytes with
  explicit lf so windows crlf translation cannot corrupt the bash
  file list. it has its own tests in tests/ci/.
- the local runner now reports the skipped count and prints a note:
  macos_only/windows_only tests were skipped on this host, and this
  ci lane runs them. a green local run on linux no longer reads as
  coverage of the other hosts.
- the runner default job count is now #cpu, not #cpu*2.
2026-08-09 22:09:49 -04:00
Jeff Watts
298ef06458 fix(tests): Windows-aware path-list split and UTF-8 progress output in parallel runner
Two Windows bugs in scripts/run_tests_parallel.py:

- --files/--paths/HERMES_TEST_PATHS were split on ':', which shreds
  absolute Windows paths at the drive letter ('C:\repo\tests' ->
  ['C', '\repo\tests']): the drive letter became a phantom discovery
  root and the rooted remainder only resolved by WindowsPath
  re-anchoring it onto repo_root's drive. New _split_pathspec() keeps
  drive-letter colons glued to their path and accepts ';' (os.pathsep)
  on Windows, while ':'-joined lists (CI generate job) keep working.

- With piped stdout (CI, subprocess capture) Windows encodes the
  runner's output as the ANSI code page, so printing the per-file
  progress glyphs raised UnicodeEncodeError inside the executor
  done-callback and every progress line was silently lost -- which is
  also why test_bare_value_flag_keeps_its_value failed on win32 (no
  '1[check]' line, and the summary says '1 tests passed', which does not
  contain '1 passed'). The runner now reconfigures its own
  stdout/stderr to UTF-8 on Windows, and the tests decode the captured
  output as UTF-8.

Adds regression tests: os.pathsep-joined absolute roots (all
platforms) and no-phantom-drive-root (win32).

Fixes #57149
2026-08-08 12:33:19 -07:00
SmokeDev
3e7a11ca2e fix(test-runner): native Windows venv probe + glyph-safe stdio
Two Windows fixes for the canonical test runner, salvaged from #66496:

- scripts/run_tests.sh: probe the native Windows venv layout
  (Scripts/activate → Scripts/python.exe) alongside bin/activate,
  adapted to main's pytest-import-guarded loop with SKIPPED_VENVS
  reporting. Without it a python -m venv / uv venv on Git Bash/MSYS is
  never found and the runner refuses to start.
- scripts/run_tests_parallel.py: _make_stdio_glyph_safe() reconfigures
  stdout/stderr to UTF-8 (errors=replace fallback) so the ✓/✗ progress
  glyphs cannot crash a cp1252 console when the runner is invoked
  directly (run_tests.sh's PYTHONUTF8=1 only covers the wrapped path).
  No-op on UTF-8 stdio. Ships 3 OS-independent cp1252 tests plus
  encoding=utf-8 in the runner-subprocess assertions.

Dropped from the original PR: the USERPROFILE/HOMEDRIVE/HOMEPATH/
SYSTEMROOT env-forwarding hunk — main's WIN_ENV loop already forwards
a superset (#67385/#70813).
2026-07-29 21:30:53 -07:00
teknium1
35b1e57862 fix(tests): a run that collects nothing can no longer look green
Three foot-guns in the canonical test runner, each of which cost real
debugging time by making an unverified run look verified.

1. Zero collection across the whole run reported success-shaped output.
   Per-file rc=5 is rewritten to rc=0 so a platform-gated file (every test
   skipped on this OS) doesn't fail the suite — correct, but it also meant a
   run where NOTHING was collected anywhere printed
   "0 tests passed, 0 failed (100% complete)" and, with no failures
   recorded, could exit 0. Now the run-level guard counts every collected
   outcome (passed/failed/skipped/errors/xfailed/xpassed): an all-skipped
   file still passes, but zero-collected-anywhere prints an explicit
   "✗ NO TESTS RAN — this is NOT a pass" block naming the likely causes and
   returns 1.

2. A venv without pytest was selected merely for existing. The probe
   accepted any directory with bin/activate, so in a checkout/worktree
   without a local .venv it picked the RELEASE venv
   (~/.hermes/hermes-agent/venv, no pytest). Every file then died with
   "No module named pytest" and the run reported 0 tests. Candidates are now
   import-checked for pytest — the same guard the HERMES_PYTHON fallback
   already applied — and a skipped candidate is named on stderr.

3. Pytest node ids were silently discarded. This runner is file-granular,
   so `tests/foo.py::TestBar::test_baz` isn't an existing path: discovery
   dropped it and the run ended "No test files to run" while the selector
   looked accepted. Node ids are now translated to the FILE plus an inferred
   `-k` on the leaf name (parametrized ids reduced to the function name),
   with a note explaining the translation. An explicit caller `-k` wins over
   the inferred one.

Tests: 4 behavior contracts in tests/test_run_tests_parallel.py. Verified
by sabotage — reverting the runner fails 3 of the 4 (the fourth pins the
pre-existing all-skipped tolerance so fix 1 can't regress it).
2026-07-25 07:09:22 -07:00
teknium1
75e0d52034 fix(windows): sweep remaining bare read_text/write_text sites + linter rule
AST-driven pass over every Path.read_text()/write_text() without an
explicit encoding= across non-test code: 71 sites in 34 files
(skills_hub, hermes_cli/main+profiles+service_manager+container_boot,
mem0/hindsight/honcho plugins, achievements dashboard, release/CI
scripts, productivity+comfyui skill helpers, agent/*). Verified zero
positional-encoding collisions before insertion; per-file compile()
check after.

Adds a check-windows-footguns rule flagging bare single-line
read_text/write_text (multi-line forms stay covered by the AST guard
test from #38985). Together with the salvaged contributor commits this
retires the ~169-site bare file-I/O class (#37423's long tail).
2026-07-24 17:10:39 -07:00
teknium1
d4b867cf9f fix(windows): sweep remaining unguarded text-mode subprocess sites codebase-wide
AST-driven pass over every subprocess.run/Popen/check_output/check_call/call
with text=True (or universal_newlines=True) and no explicit encoding=:
append encoding='utf-8', errors='replace' at the kwarg site. 136 call
sites across 28 files (cli.py, hermes_cli/main.py, tools_config.py,
environments, computer_use, gateway, scripts, skills helpers, agent/*).

Together with the salvaged #55339/#60741 commits this closes out issue
#53428's bug class; the salvaged #60751 linter rule in
check-windows-footguns.py now enforces it repo-wide (verified: 807 files
scanned, zero findings).
2026-07-24 11:45:57 -07:00
Teknium
597615ade4 fix(ci): make tests, workflows, and attribution reliable under load (#66373)
* feat(attribution): conflict-free contributor mappings via contributors/emails/ directory

The AUTHOR_MAP dict in scripts/release.py was a merge-conflict magnet:
every concurrent salvage PR appended entries to the same lines of the
same file, so parallel PRs re-conflicted on every merge to main.

New system: one file per email under contributors/emails/ — filename is
the commit-author email, first non-comment line is the GitHub login.
File additions never conflict, so any number of PRs can add mappings
concurrently.

- scripts/release.py: AUTHOR_MAP is now LEGACY_AUTHOR_MAP (frozen)
  merged with the directory at import time (directory wins). All
  existing consumers (resolve_author, contributor_audit.py) unchanged.
- scripts/add_contributor.py: idempotent CLI to add a mapping; refuses
  conflicting reassignments (incl. against the legacy map), validates
  email/login shapes.
- contributor-check.yml: attribution gate now accepts a mapping file OR
  a legacy entry; failure message prints the exact add_contributor
  command. Also auto-resolves bare <login>@users.noreply.github.com
  emails is intentionally NOT added (kept id+login form only, matching
  previous behavior).
- contributor_audit.py: guidance now points at add_contributor.py.
- tests/scripts/test_contributor_map.py: 12 tests covering loader,
  merge precedence, CLI idempotency/conflict/validation, subprocess E2E.

* feat(ci): one-shot per-file flake retry in the parallel test runner

A failing test FILE is re-run once in a fresh subprocess. Pass-on-retry
counts as green but is loudly reported in a '⚠ FLAKY' summary section
(with both attempts' output preserved) so the flake gets fixed instead
of eating a full-run rerun. Deterministic failures fail both attempts —
regressions cannot be laundered green.

- --file-retries N / HERMES_TEST_FILE_RETRIES (default 1, 0 disables)
- E2E verified: simulated first-run-fail flake goes green with banner;
  deterministic failure still exits 1; retries=0 restores old behavior.

This converts the dominant CI failure mode (one timing-sensitive test
flaking a 4600-test shard, requiring a manual 10-minute rerun and an
agent triage loop) into a self-healing retry that costs one file's
runtime.

* test(approval): loosen wall-clock perf bounds 0.15s -> 2.0s

These guard against catastrophic regex backtracking (seconds-to-minutes
class), but 0.15s is within scheduler-stall noise on loaded shared CI
runners — test_max_accepted_separator_free_input_is_fast failed a CI
shard this week on runner load alone. 2.0s still catches the regression
class with zero flake surface.

* fix(ci): job timeouts everywhere + retries on all network installs

Reliability pass over every workflow:
- timeout-minutes on all 21 jobs that lacked one (a hung job previously
  burned the 6-hour default runner budget)
- ./.github/actions/retry wrapped around every network-fetching install
  that lacked it: pip installs (deploy-site, skills-index), npm ci
  (deploy-site website, upload_to_pypi web + ui-tui), uv sync (docker
  test deps). Deterministic build steps (npm run build) deliberately
  NOT retried — split into separate steps so a real build failure fails
  fast instead of retrying 3x.

* docs(agents): document the file-retry flake policy

* fix(ci): curl retries on deploy hook + skills-index probe

* fix(ci): kill the remaining transient-failure classes in workflows + Dockerfile

From the workflow reliability audit:
- tests.yml: duration-cache restore had NO restore-keys while saves use
  run_id-suffixed keys — the cache never matched once, so LPT slicing
  always ran blind and unbalanced slices pushed heavy files toward the
  per-file timeout. One-line restore-keys fixes slice balancing.
- Label gates (lint ci-reviewed, supply-chain mcp-catalog-reviewed):
  'gh pr view || true' turned an API blip into 'label absent' → false
  BLOCKING failure. Now 3x retry, and API failure is reported as an API
  failure instead of a missing label.
- detect-changes action: compare API retried before failing open (was
  silently running all lanes on any blip).
- uv-lockfile-check: 'uv lock --check' resolves against PyPI — retried
  so registry blips don't read as 'lockfile stale'.
- docker.yml merge job: imagetools create retried (Docker Hub eventual
  consistency on just-pushed digests).
- Dockerfile: apt-get Acquire::Retries=3; s6-overlay ADDs converted to
  curl --retry 3 (ADD cannot retry; checksums still enforced); npm
  --fetch-retries=5; playwright chromium fetch retried 3x.
- Advisory artifact uploads (per-slice durations, ci-timings report)
  get continue-on-error so an artifact-service blip can't fail a green
  test slice.

* fix(tests): kill the two root-cause flakes — leaking pre-warm timer + env-dependent provider list

- test_tui_gateway_server.py: session.create / non-eager session.resume
  arm a 50ms threading.Timer (_schedule_agent_build) that outlives its
  test and fires into the NEXT test's _make_agent mock, racily
  corrupting captured state (the recurring session_resume shard
  failures). Replaced the per-test whack-a-mole stub with a module-wide
  autouse fixture; the 3 worker-lifecycle tests that genuinely need the
  deferred build opt back in via @pytest.mark.real_agent_prewarm (new
  marker in pyproject).
- test_api_key_providers.py: PROVIDER_ENV_VARS is now derived from the
  live PROVIDER_REGISTRY instead of a hand-list that had drifted
  (missing HF_TOKEN / DEEPINFRA_API_KEY) — resolve_provider('auto')
  tests failed on any machine with HF_TOKEN exported. E2E-verified with
  HF_TOKEN/DEEPINFRA_API_KEY set: 42/42 pass.

* test: de-flake 30 timing-sensitive test files for loaded CI runners

Root-cause fixes from the flake audit (session-DB mining + repo sweep):

Event-based sync instead of sleep-sync:
- title_generator: mock sets threading.Event, wait(10) replaces
  sleep(0.3) hoping the daemon thread got scheduled
- docker zombie_reaping / profile_gateway: poll-for-state helpers
  replace fixed 1-3s sleeps (s6 transitions + SIGCHLD reaping are async)
- process_registry tree test: select()-bounded readline replaces an
  unbounded blocking read (parent wedge now fails THIS test with a clear
  message instead of an opaque rc=124 file kill); SIGTERM grace 1s->2s
  (the 1s partition window mid-interpreter-startup is how a child PID
  escaped the live-system guard in CI)

Timeout raises (loaded 8-way-sliced runners see ~5s scheduling floors;
all of these complete in ms-to-1s when healthy so the raises cost
nothing on green runs):
- subprocess/thread waits <= 2s raised to 10-15s across mcp_tool,
  mcp_circuit_breaker, mcp_reconnect_retry_reset, mcp_parked_self_probe,
  mcp_cancelled_error_propagation, registry, clarify_gateway, interrupt,
  voice_cli_integration, docker_environment, session_store_lock_io,
  planned_stop_watcher, cli_interrupt_subagent, thread_scoped_output
  (joins now also assert not is_alive() so stragglers fail loudly)
- wall-clock discrimination ceilings loosened where the guarded hang is
  10x larger: local_background_child_hang 4s->10s, interrupt_cleanup
  setup 5s->20s + pgid-exit 30s->60s, mcp_stability grandchild spinup
  5s->15s, protocol/gil-starvation fast-handler 0.5s->2s,
  iso_certify_seam 1.5s->5s, wait_for_mcp_discovery 0.1s->1s
- narrow assertion windows widened: honcho first-turn wait 0.4..0.65 ->
  0.25..2.0 (property is bounded-not-hung, not an exact wall-clock);
  compression fork-lock TTL 1s->3s (12 refresh chances per lease);
  compression-lock expiry margins symmetric (ttl 0.05->0.5, sleep 1.0)
- telegram hung-DNS bound 1.0->1.4 (fake hang is 1.5s — must stay under)

* fix(tests): repair indentation from de-flake batch edit

* fix(tests): harden env isolation and replace remaining sleep-sync races

The full 42k-test run and complete npm check surfaced three more classes:

- Environment isolation: local ~/.honcho defaultHost and SSH_* variables
  leaked into Python/TUI tests. Pin the default Honcho host in the
  hermetic fixture, isolate the one fallback test from ~/.honcho, and
  blank SSH_* around terminalSetup tests. This flipped 20 false failures
  back to deterministic behavior on developer machines.
- Background-thread sleep-sync: Honcho async writer tests patched
  time.sleep globally, then busy-polled with that same mocked sleep. Under
  full-suite load the poller could starve the writer. Each test now waits
  on an Event emitted by the exact flush/retry transition; 30/30 passed
  under 15-way contention.
- Desktop streaming: the test slept 80ms and assumed a 500ms timer could
  not fire before its assertion. A loaded runner descheduled the test for
  >500ms and both chunks arrived. Producer controls now gate second-chunk
  and completion transitions explicitly.

Also make file-retry observability complete: a self-healed flaky file now
prints BOTH attempts' full output in the FLAKY summary. Two behavioral
runner tests prove pass-on-retry is green+loud+traceback-preserving, while
a deterministic failure remains red.

* refactor(ci): use gh bot pat, better retries

refactor(ci): use retry action for PR label fetch
the retry action now captures stdout as a step output, so it can serve
double duty: retry + output capture for commands like 'gh pr view' whose
result must be consumed by later steps.

Retry action gains:
- 'stdout' output (heredoc-delimited to preserve newlines)
- tee to temp file so stdout still streams to the job log
- step id 'retry' for output reference

Both lint.yml and supply-chain-audit.yml now use the retry action
directly with 'command: gh pr view ...' and read
steps.<id>.outputs.stdout.

ci: use AUTOFIX_BOT_PAT for all gh CLI / GitHub API auth

Replace secrets.GITHUB_TOKEN and github.token with
secrets.AUTOFIX_BOT_PAT across all workflows and composite actions
that use the gh CLI or GitHub API. The PAT has consistent permissions
across fork PRs (where GITHUB_TOKEN is read-only), avoids API rate
limit sharing with the default token, and is already used by
js-autofix.yml for the same reasons.

19 sites swapped across 9 files:
- lint.yml (3): label fetch, comment post/edit, comment update
- supply-chain-audit.yml (5): scan, critical comment, unbounded dep
  comment, label fetch, mcp-catalog comment
- lockfile-diff.yml (1): PR comment post/update
- skills-index-freshness.yml (1): issue creation on degraded probe
- skills-index.yml (2): index build, trigger deploy workflow
- upload_to_pypi.yml (2): release view poll, release upload
- ci.yml (1): timings report
- deploy-site.yml (2): skills index crawl
- detect-changes/action.yml (1): compare API call

---------

Co-authored-by: ethernet <arilotter@gmail.com>
2026-07-17 20:55:24 +00:00
ethernet
7fe1cb384e feat(ci): python test speedups 2026-07-13 15:29:20 -04:00
Teknium
55e3ee1ab8 fix: remove dead f-string prefixes via ruff F541 (216 sites) (#52336)
ruff check --fix --select F541 . on current main. Pure prefix removals;
adjacent-string concatenations keep the f only on interpolating fragments.
No string content or live placeholder altered.
2026-07-05 13:42:46 -07:00
Teknium
b508d4296e test(ci): raise per-file timeout 140s → 300s to stop false timeouts (#54143)
* test(ci): raise per-file timeout 140s to 300s to stop false timeouts

The per-file parallel runner caps each test-file subprocess at a flat
wall-clock budget. Combined with per-test subprocess isolation (a fresh
Python process per test), a large-collection file pays N x (interpreter
startup + import) of overhead before any test logic runs. That overhead
dilates under load on shared CI runners, so a file that finishes in
~100s on a quiet box can blow the old 140s cap purely from scheduling
jitter, surfacing as a false 'no tests ran' timeout (rc=124) with zero
actual test failures.

Raise the default to 300s (5 min). The Docker build matrix jobs already
take 7-10 min, so this headroom costs nothing on total CI wall time
while still bounding a genuinely hung file.

* docs: add infographic for CI per-file timeout bump
2026-06-28 02:41:07 -07:00
Teknium
2523917680 fix(tests): bare pytest flags pass through run_tests.sh without a '--' separator (#54008)
The parallel runner only forwarded pytest args after a literal '--', so a
bare 'scripts/run_tests.sh tests/foo.py -q' (or -v/-x/-k/--tb=long) errored
out with 'unrecognized arguments'. This contradicted the docstring's
promise that common pytest flags pass through, and forced a retry on every
run that used pytest muscle-memory.

Now any token starting with '-' that isn't one of the runner's own options
(-j/--jobs, --paths, --slice, --file-timeout, --generate-slices, --files,
--include-integration) is routed to each per-file pytest invocation
automatically. Value-taking flags given space-separated (-k expr, -m mark,
-p plugin, -o name=val, etc.) keep their value instead of having it stolen
by positional-path discovery. The explicit '--' separator still works and
stacks with bare flags.

- scripts/run_tests_parallel.py: argv splitter routes bare unknown flags to
  pytest; value-flag lookahead; updated docstring.
- scripts/run_tests.sh: usage comment reflects bare-flag passthrough.
- tests/test_run_tests_parallel.py: 4 behavior-contract tests (bare -q runs,
  -k keeps its value/filters, '--' still works, positional path stays a root).
2026-06-27 22:43:26 -07:00
ethernet
dd0e4ab81a change(ci): slice files in matrix job
avoid duplicating work, avoid file discovery on each job
2026-06-26 19:15:18 -07:00
ethernet
1a75387fa8 change(ci): log json decode error in durations 2026-06-26 19:15:18 -07:00
ethernet
707ae6e623 change(tests): don't count with pytest collect
it's way too slow. just grep files lol
2026-06-26 19:15:18 -07:00
ethernet
9a861cd0ab change(tests): don't pass pytest args when counting tests 2026-06-26 19:15:18 -07:00
ethernet
fb1dd1bf91 change(ci): docker-publish.yml -> docker.yml 2026-06-26 19:15:18 -07:00