tests/conftest.py had grown to 2,343 lines, past the 2,000-line gate. Move
three self-contained topics out, unchanged:
- env_filter.py: the credential / behavioral env-var name tables the
hermetic fixture blanks;
- live_system_guard.py: the autouse live-system guard fixture, its marks
and the protected checkout roots;
- platform_gating.py: the platforms() marker evaluation used by the
collection hook.
conftest imports them instead of listing them in pytest_plugins: it is not
the rootdir conftest, and pytest fails a run that loads a non-root conftest
carrying pytest_plugins after startup. Imported fixtures register on the
conftest module under their old names, so autouse order is unchanged.
Tests that reached into the moved symbols now import the new modules.
Conflict resolutions and semantic fixups:
- tools/environments/base.py: main's hard-exit kill fence (kill a spawn the
fence missed, deregister from _live_foreground in a finally) wrapped around
pm-clean's output collector.
- pyproject.toml: pm-clean's marker list plus main's new `live` marker.
- hermes_cli/main.py: pm-clean runs startup recovery from hermes_bootstrap, so
the old early-recovery block stays gone; main's interrupted-pull restore
(auto-merged above it) runs right after bootstrap, as on main.
- hermes_cli/update_cmd.py: main's interrupted-pull marker now guards
pm-clean's first tree mutation (release-tag detach, ff-only, or reconcile)
and is cleared once git is done. The marker's target is the ref git actually
moves to (a release tag, not always origin/<branch>), since the restore
compares against it.
- hermes_cli/_early_recovery.py: restore `import subprocess`, which pm-clean
had dropped and main's auto-merged restore needs (NameError on the first
launch after a killed update; test_update_interrupted_pull red -> green).
- apps/desktop/src/i18n/{de,es,fr}.ts: main's new locales carry the full
settings.about block; trim it to `updates` as pm-clean's type and the other
overlays do (tsc: 27 errors -> 0).
- main's new e2e tests: `import yaml` -> hermes_yaml; wake-word import table
names pyopen_wakeword (pm-clean's wake-openwakeword extra); the anthropic
key-leak switch leg needs the SDK, and the api_server two-tenant test needs
aiohttp, both PM runtime extras the test env does not carry.
Re-review findings on the foreground-process registry:
- Spawn vs exit races. A child spawned but not yet registered when the hard-exit kill ran, or a
command launched after the kill took its snapshot, survived under init. The hard-exit kill now
raises a one-way exit fence and waits (bounded) for spawns already past it to register; a spawn
refuses once the fence is up, and one that registers after a timed-out wait kills itself.
- The immediate kill fell back to proc.kill() for every handle. On Modal/Daytona/Vercel that is a
blocking SDK cancel (an 8s cancel made _hard_exit take 8s). Popen handles are still killed
inline (killpg, never blocks); every other handle's kill runs on a daemon thread under one
shared 0.5s deadline.
- The registry lock was a plain Lock: a signal landing on a thread that held it deadlocked the
exit. It is now an RLock (via a Condition) and the hard-exit path only takes it with a timeout,
falling back to a lock-free copy.
- The kanban worker's SIGALRM deadman os._exit()ed without the kill; it now goes through it.
Conflict resolutions and semantic fixups:
- utils.py / hermes_yaml.py: main widened ruamel's round-trip emitter so a long
double-quoted scalar is never folded after an escaped backslash. pm-clean builds
every rt emitter through hermes_yaml.roundtrip_yaml(), so the width lives there
(ROUNDTRIP_YAML_WIDTH moves with it); xai_retirement imports it from hermes_yaml.
- hermes_cli/banner.py: keep pm-clean's removal of the banner update check. Main's
GIT_NO_LAZY_FETCH fix for it applies to its replacement, source_check: every
read-only probe (source_git_env) now refuses promisor lazy fetches, and the
partial-clone test targets that probe (red without the flag).
- .github/workflows/tests.yml: keep setup-pm; main's uv pin bump does not apply.
Main's WAL-capable SQLite gates are kept, run against $HERMES_PYTHON (the
PM-pinned interpreter, SQLite 3.53.1). The e2e step takes main's
--include-integration invocation.
- apps/desktop: package.json has no build block here, so main's macOS locale-marker
restore joins the darwin branch of the existing after-pack.mjs, and its test
loads the hook from electron-builder.config.cjs and imports PlatformPackager
from app-builder-lib's root (electron-builder 27 exports no ./out paths). The
win32 row is dropped: this hook sanitizes and signs PE trees on win32 by design.
- reconciliation.ts: main's rowId hydration (#119326) was merged into the first of
pm-clean's split helpers only; the resolver is now one helper both halves use.
- en.ts: both sides' keys kept. tests/tools/test_lazy_deps.py stays deleted.
- Tests main added with `import yaml` use hermes_yaml, like the rest of the tree.
agent/bedrock_adapter.py calls lazy_deps.ensure("provider.bedrock") at
import time. The HERMES_DISABLE_LAZY_INSTALLS kill-switch was only set by
a per-test fixture, so collecting any test module that imports the
adapter ran a real `uv pip install boto3` into the shared CI venv (the
unit job never synced the bedrock extra). test_bedrock_adapter.py raced
it: when the install had not landed yet, its botocore tests skipped and
test_call_converse_replays_thinking_botocore_accepts failed with
"No module named 'botocore'" (FLAKY on this PR's second CI run).
- tests/conftest.py sets the kill-switch at import, before collection.
- tests.yml syncs --extra bedrock with the other lazy-install extras the
suite exercises, so the botocore tests keep running, deterministically.
- The unguarded botocore test importorskips like its siblings.
- Invariant: test_lazy_deps.py asserts the switch is set at collection
(red on origin/main's conftest, green here).
A turn's auto-title thread printing its failure warning while pytest's
fd capture snapped the call phase SIGSEGV'd the worker (CI: whole
test_api_content_sidecar.py CRASHED on an unrelated PR). The existing
teardown join in _close_leaked_session_dbs runs after that snap. An
innermost pytest_runtest_call wrapper now joins them before capture exits.
Builds on tancou's #119129 (cherry-picked above): the pin now lives in
get_routing_process_hermes_home() and only the four routed-profile DECISIONS read it.
get_process_hermes_home()/get_hermes_home() keep following HERMES_HOME, so an env-only
home switch in a multiplexed process resolves as before.
- set_multiplex_active(True) pins the launch home only when no host pin exists, and
set_multiplex_active(False) releases only the pin it created itself. A transient toggle
(gateway_migrate._multiplex_read_mode, cron external-worker restore) no longer drops an
embedding host's explicit pin_process_hermes_home(launch).
- profiles._cleanup_gateway_service binds set_hermes_home_override(profile_dir) beside the
env write. Under the previous head, DELETE /api/profiles/<x> from a multi-profile dashboard
resolved get_service_name() against the pinned launch home -> bare `hermes-gateway`, and
disabled/stopped/unlinked the HOST multiplexer's unit. Same path serves rename_profile.
Tests (red on the previous head): explicit pin survives True->False; env readers follow the
env while pinned; two-home delete removes hermes-gateway-victim and leaves hermes-gateway.
Symptom: a host that serves several profiles from one process and mirrors
the active turn's profile into `os.environ["HERMES_HOME"]` for legacy
readers (Hermes WebUI does this on every chat turn, next to the
context-local override) makes every launch-home decision see the served
profile as the launch profile. Two profiles that both configure `atlassian`
with different credentials share whichever MCP connection came first: a
READ_ONLY_MODE=false profile ends up calling a read-only server
(nesquena/hermes-webui#7721). The same misjudgement leaves the launch
residue in the served profile's child env, seeds the launch profile's
bridged allow-all grant into the served profile's secret scope, and lets
the served profile's `terminal.*` config bridge into the shared process env.
Cause: four launch-home checks compare the task's override with
`get_process_hermes_home()`, which reads `HERMES_HOME` live:
`agent.secret_scope.serves_routed_profile` (keys the MCP ledger via
`_mcp_registry_scope`, #108352 / #111481, and the check_fn cache, #111151),
`agent.secret_scope._is_process_home`, `tools.environments.local._is_routed_home`
and `hermes_cli.env_loader._process_hermes_home`. Under the mirror the two
sides are equal for every turn.
Change: `hermes_constants.pin_process_hermes_home(path | None)` lets the
host record the home it serves as its own; `get_routing_process_hermes_home()`
returns the pin when set, else `get_process_hermes_home()`; the four checks
compare against it. The pin is deliberately NOT folded into
`get_process_hermes_home()`: `get_hermes_home()` falls back to it for tasks
carrying no override (MCP loop, spawners), and the host's mirror exists
precisely so those readers see the served profile. Only "is this task
routed / is this the launch home" changes. Unpinned, behaviour is
byte-for-byte the old one; hosts that never mutate `HERMES_HOME` need not
call it. `activate_multi_profile_hosting()` is not the seam for this: it
flips `get_secret` fail-closed process-wide and freezes the launch env,
which an embedding host cannot adopt as a bug fix.
Tests (2 invariants, parametrized over the four checks plus the MCP ledger
key; red on main, green here): pinned + mirrored env -> the served home is
routed and the launch home is not, the MCP key is `(home_key, name)`,
`get_process_hermes_home()` still follows the env var; never pinned or
pinned-then-cleared -> old semantics, including "a mirrored env var IS the
launch home". `tests/conftest.py` resets the pin per test so the
module-global cannot leak between files.
Live repro (WebUI + a stdio FastMCP server named `atlassian` in two
profiles, one gated by READ_ONLY_MODE): base -> one ledger key
`'atlassian'`, the write profile lists only the read-only tools; fixed ->
`(<read_home_key>, 'atlassian')` and `(<write_home_key>, 'atlassian')`,
each profile lists its own tools.
Docs: `gateway/AGENTS.md` § Profile scope (one launch-home identity) and the
isolation table in `website/docs/user-guide/multi-profile-gateways.md`.
Also maps the author e-mail under contributors/emails/ (attribution check).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Every "does this task serve a ROUTED home" decision (serves_routed_profile,
_is_process_home, _is_routed_home, env_loader._process_hermes_home) compares the
home override with get_process_hermes_home(), which read os.environ["HERMES_HOME"]
live. A host that mirrors the served profile into that env var per turn (hermes-webui)
made every served profile look like the launch one: MCP registry scope None, bare
cross-profile connection names, launch residue kept in served child envs, the launch
GATEWAY_ALLOW_ALL_USERS grant seeded into the served scope.
set_multiplex_active(True) now pins the launch home (hermes_constants.
pin_process_hermes_home; first pin wins, an embedding host may pin explicitly) and
get_process_hermes_home() returns the frozen value while multiplex is active.
Standalone hermes -p x gateway run (multiplex inactive) keeps following the env.
No os.environ fallthrough is added anywhere.
Closes#119242
Two leaks from the test temp plumbing:
tests/conftest.py relocates pytest's basetemp out of the native Hermes
home (#111101) with mkdtemp(dir=native.parent), which is the operator's
$HOME, and nothing removed it: 123 hermes-pytest-basetemp-* dirs (552 MB)
appeared there in a day, one per bare pytest process. The relocated
basetemp now goes into one prunable root (/var/tmp/hermes-pytest on
POSIX, a non-dotted sibling of the native home elsewhere), is removed at
pytest_unconfigure, and idle siblings from killed runs are swept on entry.
scripts/run_tests_parallel.py deletes each per-file temp root in finally,
but a SIGKILLed runner (tool timeout, stray pkill) never gets there and
leaks one root per in-flight worker: 983 r-* roots (3.4 GB) in three
days. The runner now sweeps 24h-idle roots at start and forces read-only
permission fixtures writable before rmtree instead of skipping them.
_host_matches_platforms turned a misspelt spec (platforms("linx")) into
a skip, so the test vanished on every host with both lanes green — the
exact silent-drop class the collection guard exists to catch. Validate
every spec first and raise pytest.UsageError naming the item; a matching
sibling spec no longer excuses a typo next to it.
The default install checks the repo out INSIDE the Hermes home
(install.sh INSTALL_DIR=$HERMES_HOME/hermes-agent), so the guard tripped
on the checkout's own test data and on traceback source reads from any
default-location run. Exempt the project root like the interpreter
prefixes: the checkout is not Hermes state.
A Hermes-launched shell also hands pytest TMPDIR=<home>/cache/scratch
(tagged HERMES_SCRATCH_DIR), and importing hermes_bootstrap re-applies
that redirect for a custom HERMES_HOME. Both put the session sandbox,
basetemp and every tempfile default inside a guarded root (16 errors in
test_update_completion_routing alone under a bare pytest). conftest now
strips Hermes' own export before allocating temp space and pins the
system default so the import-time hook is a no-op; user-set temp vars
are untouched and the parallel runner's own TMPDIR is unaffected.
Host-scoped update-restart obligation (061195fac1 / 953b6f6f08 / 3be255eca6)
lands on the PM model: the obligation record, its readers and the legacy
per-home marker compat come in as-is. The catch-up restart path
(`_apply_pending_fleet_restart_catchup` / `_run_pending_fleet_restart`) is
retired here (the fleet restart rides the completion owner), so main's
per-host restart-once guard on that path is not carried; its unit→live
MainPID collapse IS ported into the live post-update systemd pass
(`_restart_systemd_gateway_units`), with the two collapse tests rewritten
against that function (red on the pre-port tree: `_unit_main_pid` absent).
Tests that exercised only the retired catch-up path are dropped.
Desktop: main's shared log-rotation planner replaces the inline constants
in main.ts; the merge keeps our machine-profile import beside its import.
utf-8 → utf-8-sig on the three new BOM-intolerant reads (footguns lint).
One host runs one multiplexing gateway, but the update pipeline still treated
the pull->restart obligation, enumerated units, recovery payloads and the
planned-restart notice as per-profile. Two profiles updating meant two outages
of the same process, and a served profile's channels were never told.
- hermes_cli/update_host_obligation.py: new host-scoped obligation record in
gateway.host_rendezvous.host_state_dir() (host-update-restart.json), plus the
unit->live-MainPID collapse rule. The legacy per-home marker stays readable
and clearable so an in-flight obligation is still discharged.
- update_cmd_fleet: arm/clear/read the host record; the catch-up restart is
idempotent per host (a completed restart onto the checkout SHA is never
repeated); leftover per-profile units resolving to one MainPID restart once.
- update_restart_recovery: payload profiles served by one host process are one
restart target, reported under "covered".
- gateway notices: owed targets and the online notice span every served
profile's home channels; the marker survives until each was reached.
launchd ProgramArguments now run through osascript (#71206); the plist test unwraps the
exec argv and keeps the PM launcher-shaped assertions. run_tests.sh forwards
HERMES_GATEWAY_LOCK_DIR alongside the SSL vars.
Review fixes on the lifecycle-verbs PR. Three of them were escape hatches that
looked implemented and were dead code, and one turned a boot race into a
permanently parked unit.
- ATTACH now requires a LIVE `identify` answer. The claim-time record is
published with NO served set (the runner settles multiplex a moment later),
and `host_gateway()` reports `served_known=False` when nothing answers. An
owner whose served set is unknown yields a TRANSIENT refusal, never an
attach: previously `default`'s claim published "default,other" before its
socket bound, `other`'s systemd unit read that as "I am served", exited 78,
and systemd parked it for good.
- `served_profiles()` honours the actual `gateway.multiplex_profiles` setting
instead of forcing `multiplex=True`, so a standalone gateway stops claiming
the whole roster.
- `--replace` is threaded through the CLI guard into `start_gateway`, and
`--force` into `_host_attach_or_none`. Both previously exited in the guard
before the code that implements them ever ran ("nothing to start", rc=0).
- A supervised attach exits 75 (EX_TEMPFAIL), not 78. 78 is the PERMANENT
config refusal every supervisor parks on; "someone else serves me right now"
is a runtime observation that ends when that process does. No unit files
change: systemd already has RestartForceExitStatus=75/RestartSec=5, the s6
finish script passes 75 through, launchd relaunches a non-78 failure. Exit 0
would not do — s6 parks a clean exit too.
- `restart --all` retracts the stopped owner's record (`discard_dead_record`)
and re-enters with `replace=True`, so it can no longer attach to the corpse
it just stopped and exit 0.
- Rendezvous hardening: the dir is created/repaired 0o700, a record whose
`st_uid` is not ours is ignored, liveness is proven BEFORE we dial the home
it names, and a live `identify` must agree about `hermes_home`.
- `-p X gateway restart --all` reaches the `--all`-aware branch instead of the
generic guard's `hermes -p default gateway restart` one-liner.
- `host_gateway()` is memoized (2s TTL, invalidated on every record write), so
`gateway status`/doctor across N profiles pays one probe, not N.
Tests: the two new files build the record as raw JSON, so they COLLECT and RUN
against a tree without the `home` field and fail on the outcome. A/B against
the PR head: 9 failed / 6 passed → 15 passed. conftest's per-test
HERMES_GATEWAY_LOCK_DIR now defers to a caller-supplied value (and
run_tests.sh forwards it through `env -i`), and the per-process dir is a
deterministic self-sweeping per-PID path instead of an atexit-only mkdtemp.
`test_runner_startup_failures.py` stubs the new attach gate and releases the
host role it claims.
`gateway run`, `start --all`, `restart --all` and `stop` each assumed "this
profile's gateway". Under the multiplex-only ruling there is exactly ONE
gateway process per host, so they now target that process:
- `gateway run` for a profile the host gateway already serves ATTACHES: print
its PID + served set, exit 0, spawn nothing. Not served yet -> ask the owner
to re-scan `profiles/` (control socket) and attach once the answer includes
it. Refuse only when the host gateway cannot be made to serve it. Under a
service supervisor the attach exits 78 instead of 0 so a redundant unit is
parked, not restart-looped.
- The attach channel is reachable BEFORE the PID claim: the decision reads the
landed rendezvous record (now carrying the owner's HERMES_HOME) and talks to
the owner's control socket, so it no longer depends on the claim ordering in
start_gateway.
- `start --all` / `restart --all` no longer SIGTERM every gateway-looking
process: they restart the host multiplexer and preserve its served set. A
secondary still running its own gateway is reported with the
`gateway migrate --multiplex` one-liner, never killed.
- Ownership is decided by the live served set (record + control socket), not by
argv: a host singleton runs bare/default argv and can never prove it serves
profile X, which rejected every secondary.
- The implicit-multiplex verdict no longer requires the DEFAULT profile: the
multiplexer is whichever profile launched the one host process.
Tests: per-test HERMES_GATEWAY_LOCK_DIR isolation in tests/conftest.py — the
host record is shared per OS user by design, so one test that boots a gateway
made every other file's lifecycle code attach to it.
- tests/conftest.py: `real_bash` fixture — the Windows runners resolve `bash` to System32's
WSL launcher (UTF-16 "no installed distributions", exit 1); prefer Git for Windows'. Used
by the setup-pin, install stage-frame and source-launcher shell tests.
- source-build-env.ps1: Test-Path before Remove-Item — under $ErrorActionPreference='Stop'
a missing identity variable aborted the try block before the child ran (Windows PS 5.1
raised where pwsh on Unix did not).
- desktop-update/windows.ps1: `--force` precedes the target arguments (the hand-off contract
test reads argv in that order).
- test_install_ps1_desktop_stage: assert the current contract — the shared completion tail
(source_completion.py --desktop) builds the products and -IncludeDesktop selects the desktop
product inside `products` rather than adding a stage. The test predated the completion-tail
refactor and had been red on the Windows lane since.
- test_windows_native_support: the restart watcher argv is `runtime_command` shaped
([python, -I, -c, bootstrap, pid, delay, ...]).
- test_mint_launchers: create the fixture repo's pm/ dir before copying pm/environments.py.
- test_browser_use_pm: console-script launchers report sys.argv[0] without `.exe`.
- test_update_stale_gateway_yield (from main): `_verify_fleet_after_update` has no
`node_failures` here (PM owns node).
Merge fallout (my resolution errors, all caught by CI):
- hermes_cli/backup.py + gateway.py: `theirs` on those hunks re-imported clusters HEAD had
already moved to backup_restore.py / kept in the facade. backup.py loses the 349-line
duplicate (main's #110179 fix is ported into backup_restore._import_db_member); the
systemd service-unit cluster returns to gateway.py (PM's _prepare_service_launcher /
_pm_managed_node_dirs / _systemd_command have no home in main's extraction) with main's
utf-8-sig read. gateway_service_unit.py is dropped.
- gateway/run.py: main's plugin-update chore is not profile-scoped (the housekeeping
ordering test pins the scope/drain sequence).
- pyproject + 30 test files: `import yaml` -> `import hermes_yaml as yaml` (pm-clean has no
pyyaml); gateway/config._bundled_platform_manifest_name reads through hermes_yaml.
- tests re-seamed onto pm-clean's shape: residency admission (installed_engine),
supervisor child env (binary is a constructor argument), update import guard
(update_cmd_deps is gone; our probe already scrubs PYTHONPATH — both #115032 invariants
pass), shallow-count git responses (stash path asks `status --porcelain -z`); dropped
tests for retired code (_run_node_bootstrap/_ensure_tui_node, Windows resume demotion).
- tests/tools/test_local_env_blocklist.py: restore the two helpers the suite-reduction
commit dropped and the blocklist import.
Real fixes:
- pm: classify_uv_failure/ResolutionConflict move beside the uv runner (pm.environment,
stdlib-only). pm.workspace imports tomllib at module level and cannot load on the 3.10
bootstrap python that streams uv output in the Docker arm64 image.
- tools/browser_tool.warm_agent_browser_npx_cache: back as a permanent definition — it is on
the frozen old-updater surface, and the revert-scheduled compat pointer does not count.
- hermes_cli/memory_setup: the dashboard's pip row uses pm.environments.
running_from_selected_environment for installed vs restart_required.
- scripts/windows-build-deps.ps1: export DISTUTILS_USE_SDK/MSSdk so setuptools trusts the
primed MSVC environment instead of asking vswhere (`env -i` test runner on win32-arm64
compiling ruamel-yaml-clib); run_tests.sh forwards them.
- tests/pm/test_windows_build_deps.py: start the protocol test from a parent env without the
toolchain variables the runner job already exports.
- tests/conftest.py scrubs HERMES_BUNDLED_PLUGINS (Nix-wrapped hermes on the dev host);
tests/home_io_guard.py treats sys.path site-packages under the real home as the
interpreter's installation (PM-activated developer shell).
- tests-js: four `curly` lint errors from main's new scripts.
Branch semantics kept where main and PM disagree: update_cmd_deps.py,
constraints-termux.txt, the Electron update-api-check module and the
post-swap hand-off test stay deleted; the pending-fleet-restart catch-up
and the local_runtime tag/download ladder stay retired (PM owns engines).
Ported from main onto the branch's shape: profile_scoped_chore for the
auto-archive and plugin-update housekeeping chores, the local-runtime
cross-process boot lock and residency cap, the checkpoint tmp_pack sweep,
the cua daemon-liveness status probe, the remote-served Desktop update
flag (posix.sh / windows.ps1), sign-in for env-pinned remote gateways
(urlDisabled on RemoteSetupFields), the uvloop extra split (uvicorn
without [standard]), and the umask-scoping spawn test.
uv.lock regenerated with pm.build_env --lock-only; new utf-8 reads from
main switched to utf-8-sig (check-windows-footguns).
tests/tui_gateway/test_tui_gateway_server.py swaps agent.title_generator in
sys.modules for a bare ModuleType; the autouse sweep then called
wait_for_title_upgrades on the stub and errored every test in that file at
teardown. A stub spawned no threads, so a missing helper means nothing to
join.
tests/gateway/test_timestamp_sidecar_replay.py crashed the interpreter on CI
(native fault, green on rerun) on unrelated PRs. Root cause: every
run_conversation turn in its fixture spawns the auto-title upgrade daemon
thread (title_generator.maybe_auto_title). That thread outlives the test,
fails its model call (no provider under CI), and then writes the derived
title into the fixture's SessionDB after the fixture closed it, which
reopens sqlite on the daemon thread (_reopen_after_close_locked) and prints
the auxiliary-failure warning after pytest capture teardown. At the end of
the file the threads are still in native sqlite while the interpreter
finalizes: the check_same_thread=False-at-shutdown SIGSEGV shape of
#113186. Locally the thread finishes in ~250 ms so the race never shows;
on a loaded runner it lands on finalization.
Fix the class, not the file:
- tests/conftest.py: the autouse SessionDB leak sweep now joins the
auto-title upgrade threads (bounded, agent.title_generator.
wait_for_title_upgrades) before closing stores, so no title worker
outlives its test in any file (5 other files spawn them today).
- tests/gateway/test_timestamp_sidecar_replay.py: titling is not under
test; the fixture no-ops maybe_auto_title (same as
tests/agent/test_tool_call_incremental_persistence.py), so its own
db.close() no longer races a worker either.
- tests/hermes_state/test_session_db_leak_sweep.py: handoff pair pinning
the invariant (a slow upgrade thread started in one test is dead by the
next); red on base, green with the fix.
Proof (scratch plugin delaying the title model call by 1 s):
base: 2 auto-title threads alive at interpreter exit, every thread
"SessionDB reopened after close() on thread auto-title"; fixed: no thread
spawned / none alive at exit in this file and the other five.
Conflicts resolved toward the branch: PM owns dependency preparation, the
Windows shim re-exec/hand-off path stays retired (main's shim-parent wait,
gateway-resume env token and update_cmd_deps tests dropped), docs describe
the PM update flow. The docker workflow parks install-stamp.json around the
toolchain step instead of deleting it so tests/docker can compare provenance.
Every unflagged `hermes update` resolves its channel through R2 now, so the
mocked update harnesses reached the network (and 404 today). An autouse
fixture in tests/hermes_cli defaults every channel to a source-branch record
through the documented `_resolve_channel` seam; archive/transport tests keep
the real reader via the `real_release_channels` marker. The stable-identity
test pins the channel's commit through that seam instead of the retired
release-candidates pointer, and the completion-process fixture ships
source_completion.py in its NEW tree.
Desktop: js-tests installs the dev extra (the electron contracts drive real
hermes_cli/pm code through HERMES_PYTHON); the source-backend fixture runs PM's
worker on PM's staged runtime; async canImportHermesCli is awaited; the
feature-flag test pins the platform rule; a mocked profile store exports what
settings-scope reads; the peer-device e2e imports the mock-provider helpers
from tests-js like its siblings.
On macOS the Keychain is Claude Code's authoritative credential store, but
Hermes only ever wrote ~/.claude/.credentials.json. Since the refresh token is
single-use and rotating, every Hermes-initiated refresh left the Keychain
holding an already-invalidated token, which Claude Code then spent into
invalid_grant and discarded ("Login: Expired").
_write_claude_code_credentials now mirrors the committed refresh into the
existing "Claude Code-credentials" entry via security add-generic-password,
merging the rotated token triple over the existing payload so
subscriptionType / rateLimitTier / scopes survive. The payload is fed on stdin
(bare -w), never argv. Fail-soft: a mirror failure is logged, never raised —
the file commit already succeeded and the resolver resolves from it. No-op off
Darwin and when no entry exists (never create one the user has not).
Add a raw payload reader (metadata preserved), a pure merge helper, and the
mirror; extend the conftest keychain guard to neutralize the new writer in any
test that hasn't opted in.
Four origin/main merges brought back `linux_only` / `macos_only` /
`windows_only` marks in 41 test files, along with the pre-platforms()
versions of scripts/ci/list_os_marked_tests.py and check_os_marker_fakes.py.
Because the legacy names are no longer registered, pytest treated them as
unknown marks — a warning — so every Windows- or macOS-only test RAN on
Linux (test_local_runtime_recovery.py tripped the live-system kill guard).
Rewrite the marks, restore the platforms()-aware CI scripts (keeping main's
os.walk fix for vanishing __pycache__ dirs), drop the stale _BASELINE entries,
and make the conftest reject the retired marks outright so the next merge
cannot resurrect them silently.
Move the require_mcp_2_sdk fixture from tests/tools/conftest.py to the root
tests/conftest.py so tests/test_mcp_serve.py can use it too, and read the
required version from the `mcp==X` pin in pyproject's [mcp] extra instead of
a hardcoded "2.0.0" — the pin moves, the guard follows.
Sharpen the remaining presence-only guards named in the issue that still went
red under an installed-but-older SDK on current main: tests/test_mcp_serve.py
(the mcp_server_e2e fixture and TestServerCreation asserted `mcp.server.MCPServer`,
absent in 1.x) and tests/tools/test_mcp_oauth.py::TestCallbackPortReservation
(2.0 AuthorizationCodeResult vs the 1.x tuple). test_mcp_oauth_manager.py and
test_mcp_device_flow.py already pass under 1.28.1 on current main, so they are
left alone; PR #114068's importorskip hunks in those files do not overlap these
edits.
Reconcile plugin declarations and validation through PM's atomic generation publication; preserve external runtimes, target markers, and conflict refusal. Keep one source-update completion owner and port upstream lifecycle changes to the PM desktop/runtime paths.
The conftest session sandbox only recognised ~/.hermes as "production", so a Windows dev shell
exporting HERMES_HOME=%LOCALAPPDATA%\hermes was honoured as a custom home; every import-time
path capture (tui_gateway.server._hermes_home) then pointed at the live install and the state.db
guard tripped on store-touching tests (#112692). Compare against the same platform-default root the
guard itself uses. Adds the invariant test for the lazy resolve: a HERMES_HOME redirected after
import is honoured while the context-local override is still ignored (#102526).
Three caches that read config.yaml still keyed change detection on
st_mtime (or st_mtime_ns + st_size), so a same-size replacement that
keeps the old timestamp (cp -p, rsync -t, a timestamp-pinning writer)
was never noticed:
- model_tools._tool_defs_cache_key: get_tool_definitions kept serving
stale dynamic tool schemas / mcp_servers for the process lifetime.
- CLI mcp_servers auto-reload watcher (cli_tui_mixin seed +
cli_info_mixin._check_config_mcp_changes): the replaced mcp_servers
section was never reloaded. The seed is now _config_sig.
- tui_gateway/server._load_cfg_raw / _save_cfg (_cfg_mtime -> _cfg_sig):
the raw-config cache served the stale document and the next _save_cfg
would write it back over the on-disk file.
All three now use utils.file_signature like the rest of the PR. Tests that
reset the renamed module/instance attributes follow the rename; one
pinned-mtime replacement test per cache, red on the previous head.
Also corrects two stale type/comment annotations in hermes_cli/config.py
(_env_cache key shape, _RAW_CONFIG_CACHE record shape).
Review finding: three sibling config.yaml caches (tool-defs memo, CLI mcp watcher, TUI-gateway raw cfg) still compared mtime/size only.
Follow-up to the salvaged commit from #111169 (@KoNit-K):
- Drop HERMES_KANBAN_TASK_TITLE. The worker already has HERMES_KANBAN_TASK
and HERMES_KANBAN_BOARD/HERMES_KANBAN_DB pinned in its env, so
maybe_auto_title reads the card title from the board itself (no new
HERMES_* env var for non-secret config; the dispatcher and the
delegation scrub list stay untouched).
- Unreadable or missing card: the session is named `Kanban task <id>`
with zero auxiliary calls (the fallback the issue asked for; the
#109743 seed left such workers untitled).
- The card title persists at `llm` authority via set_auto_title, so a
manual /title still wins and the upgrade thread never starts.
- Tests trimmed to two invariants against a real board + SessionDB
(card title, unreadable-card fallback), both red on origin/main.
The fallback root when the system temp dir sits inside the operator's
Hermes home was PROJECT_ROOT/.pytest_cache, but the default install checks
the repo out inside that very home (~/.hermes/hermes-agent,
%LOCALAPPDATA%\hermes\hermes-agent), so the relocated basetemp still
resolved under the native home and get_default_hermes_root() pointed the
sandbox back at the live install. mkdtemp under native.parent is outside
the home by construction, and a loud assertion now fails collection if the
chosen basetemp ever resolves inside it.
Review finding: fallback basetemp under PROJECT_ROOT/.pytest_cache stays inside the native home when the repo lives in ~/.hermes.
Every per-test sandbox is <basetemp>/.../hermes_test, and get_default_hermes_root() prefers
the platform-native home whenever HERMES_HOME sits under it. A basetemp inside ~/.hermes
(pytest --basetemp, or TMPDIR/TEMP pointing there — the default on Windows, where the home is
%LOCALAPPDATA%\hermes) therefore turned every sandbox back into the live install, and any
test that resolves the default profile wrote fixtures over the operator's config.yaml, .env
and MEMORY.md.
Hook into pytest_configure after _pytest.tmpdir has built the TempPathFactory and move a
basetemp that resolves under the native home to a fresh tempdir outside it (falling back to
the repo's ignored .pytest_cache when the system temp dir is itself inside the home). The
fix lives at the basetemp seam so it covers every test, not only the two writers in
test_profiles.py, and needs no per-test fixture or opt-out marker.
Fixes#111101
tests/tui_gateway/test_multi_profile_hosting_fail_closed.py — through the
registered RPC handlers: config.get for a secondary resolves only its own
secrets and flips the process fail-closed; llm.oneshot / model.options bodies
see the profile's home + secrets; a launch-profile agent build binds its own
scope once multiplexing is active; CONTROL: a single-profile serve keeps the
os.environ fall-through.
tests/hermes_cli/test_web_multi_profile_scope.py — through the real FastAPI
app: GET /api/config?profile=b expands only B's refs and never mutates
os.environ; console `send` for B lands B's .env in the request scope, not the
process env.
All red on origin/main by source swap (control green), green on head.
test_profile_terminal_scope_entrypoints follows the launch_profile_policy
rename and releases the launch turn's new secret scope.
tests/conftest.py resets the process-global hosting latch
(_MULTIPLEX_ACTIVE, the frozen launch env, _served_profile_homes) per test:
one request routed to a named profile otherwise left every later test in the
file fail-closed.
The per-process address-space cap was defense-in-depth on top of the
SessionDB leak sweep (nobody sets the knob; the sweep removes the leak).
On the 96-worker CI runner it was also the only PR-specific difference
when tests/tools/test_image_source.py hung to the 600s SIGKILL while the
same file passes in ~30s on every sibling branch: RLIMIT_AS counts virtual
reservations, and image/threading libraries reserve far more address
space than they touch. Keep the sweep, remove the cap and its env knob.
Root cause of the 2026-08-16 OOM incidents (three runs of
`python -m pytest -o addopts= -q tests/hermes_cli/` ballooning to
16-25 GB RSS and getting killed): ~40 files under tests/hermes_cli/
construct SessionDB() directly and never close it. Each instance keeps
the writer connection (state.db + -wal fds), up to _READ_POOL_MAX pooled
readers with their SQLite page caches, and — once token accounting has
run — an atexit registration that pins the instance alive until
interpreter exit. In one process over 637 files those accumulate without
bound; the sanctioned per-file runner masks it, so CI never saw it.
Fix the class, not the sites:
* hermes_state: register every successfully constructed SessionDB in a
test-only WeakSet (populated only when HERMES_TEST_ISOLATION is set,
i.e. under this test suite; production never touches it).
* tests/conftest.py: autouse _close_leaked_session_dbs teardown closes
everything left in the registry after each test. close() is idempotent
and unregisters the pinning atexit hook, so instances become
collectable.
* tests/conftest.py: session-scoped _pytest_memory_cap applies a
defensive RLIMIT_AS of 12 GiB (Linux only) so any future in-process
leak fails fast with MemoryError instead of eating the box.
Overridable/disable-able via HERMES_PYTEST_MEM_CAP (documented in
scripts/run_tests_parallel.py).
* tests/hermes_state/test_session_db_leak_sweep.py: behavior contract
for registration, idempotent close, and the cross-test sweep.
Measured (capped single-process `pytest -o addopts= -q tests/hermes_cli/`):
peak RSS 4.16 GiB before -> 1.67 GiB after; per-test open .db fd count
previously climbed monotonically (0 -> 12 -> 17 -> 104 within the
SessionDB-heavy files), now stays bounded (<= 5, transient). Sanctioned
runner over the affected 35 files: 495 passed, 0 failed, no FLAKY.
Incident evidence: ~/.hermes/logs/oom-incidents/20260816-202114
(fd dumps show 100+ open state.db/state.db-wal handles across pytest
tmpdirs; 3rd recurrence that day).
`scripts/run_tests.sh tests/<dir>/` is how a change gets its regression
coverage run, so a test filed under the wrong directory is a test nobody
runs when that code changes. Two kinds of drift had accumulated.
Parallel directories for one source package, folded into the mirror:
tests/acp -> tests/acp_adapter (its __init__/conftest move with it)
tests/cli -> tests/hermes_cli (prompt_toolkit fixture merged into
hermes_cli/conftest.py)
tests/run_agent -> tests/agent (backoff fixture becomes
agent/conftest.py)
tests/relay -> tests/gateway/relay
tests/state -> tests/hermes_state
246 loose files at tests/ root, routed by the package they import/patch:
hermes_cli, hermes_state, agent, gateway, tools, plugins, tui_gateway, cron.
Installer and desktop-update script tests go to tests/scripts/{install,
desktop_update}/. 43 tests of root-level modules (batch_runner, utils,
hermes_constants, packaging) stay at the root.
Filenames drop their issue numbers (95 files: test_89315_x.py -> test_x.py);
the number stays in the module docstring where it has context.
Collisions: test_cli_skin_integration.py existed in both tests/ and tests/cli
with different subsets — merged into one (10 tests, all kept);
run_agent/test_pre_compress_memory_context.py -> agent/..._handoff.py;
tests/test_account_usage.py -> agent/test_account_usage_fetch.py;
tests/test_web_server.py -> hermes_cli/test_web_server_ws_ping.py.
Deleted: test_minisweagent_path.py (empty since PR #2804),
test_model_picker_scroll.py (tested a private copy of the logic, imported
nothing), test_process_loop_event_loop_warning.py (asserted asyncio behaviour,
imported nothing from Hermes).
Repo-root path arithmetic (Path(__file__).parents[N], dirname chains) is
bumped for the 202 files that changed depth and verified by evaluating every
such expression against the new location. classify_changes' desktop-updater
lane prefix, tests-os.yml's ignore glob and every in-tree path comment follow
the moves. tests/test_tests_tree_layout.py keeps the tree from drifting back.
computer_use kept its own approval decision: two module dicts
(_session_auto_approve / _always_allow) mirroring tools.approval's
session store and _persist_choice, a private verdict vocabulary
(approve_once/approve_session/always_approve) that hermes_cli mapped
back to once/session/always, and — the real problem — `if
_approval_callback is None: return None`. Only the interactive CLI ever
installed that callback, so every other host (gateway turns, cron,
api_server, tui_gateway, ACP) ran destructive desktop input with no
approval at all, ignoring cron_mode / unattended_mode / the permanent
allowlist, and "always" grants were invisible to `is_approved`,
`clear_session` and the messaging-platform approval buttons.
_request_approval now calls tools.approval._run_approval_gate with
pattern_key `cua:<action>:<background|foreground>` (the old scope shape,
so a background grant still never covers the visible foreground variant)
and fail_closed_when_no_human=True, the same posture as
request_tool_approval / the SSH-config write gate. The private dicts,
their release/atexit clearing, the verdict mapping in
hermes_cli/cli_modal_mixin.py and the extra callback install in cli.py
are deleted: the CLI's terminal_tool callback answers computer_use
prompts like any other tool. set_approval_callback stays as an optional
explicit-callback hook with the shared callback contract
(cb(command, description, **kw) -> once|session|always|deny|timeout);
no in-tree host uses it.
Behavior change:
- No approval callback and no gateway (cron, api_server/webhook,
headless -q, plain library use): destructive actions are now REFUSED
with a BLOCKED error and never reach the backend. Previously they
silently ran. cron honors approvals.cron_mode, unattended platforms
approvals.unattended_mode, -q approvals.single_query_mode.
- --yolo / gateway /yolo / approvals.mode: off still allow (unchanged).
- Gateway sessions (Telegram/Discord/Slack/...) now get a real pending
approval with once/session/always buttons instead of default-allow.
- session/always grants live in tools.approval's store; "always" is one
command_allowlist entry (`cua:click:background`) and is scoped to that
action+mode — the old blanket "always_approve unlocks everything for
the session" no longer exists.
- Denial wording is the shared gate's ("BLOCKED: User denied ...",
"BLOCKED: Action timed out ..."); the error JSON keeps `action`.
Tests: tests/tools/test_computer_use_approval_isolation.py
::test_no_callback_refuses_unless_yolo (blocked + no backend call, then
yolo executes) and ::test_always_grant_lands_in_the_shared_store
(is_approved sees the cua:<action>:<mode> key; second call served from
the store). Sabotage: restoring the `callback is None -> allow`
short-circuit fails the first; swapping the shared gate for a private
grant set fails the second plus the three delivery-ladder scope tests.
tests/tools/conftest.py gains `grant_computer_use_approvals` for
dispatch tests that only care about routing.
noninteractive_git_env() now spawns `git config --get-all safe.directory` before
building the env. Eight tests fake subprocess.run/Popen with a fixed sequence of
expected git calls (update check, plugin pull, MCP install, bounded probe) and the
extra spawn tripped them in CI. An autouse fixture stubs the read to "no entries";
the two carve-out invariant tests opt back in with @pytest.mark.real_safe_directory
(and were confirmed to still exercise the real read: the ordering test would fail
against the stub).
contributors/emails: pry@privacydied.net -> privacydied (check-attribution).
The shared fixture isolates HERMES_HOME, but profile-root resolution also
resolves the native default. This trips the real-home guard even for tests
that use a temporary custom home. Base 75a646e5b3 has the same failures.
Isolate the native default in the shared fixture. Capture its parent before
test fixtures run so explicit home overrides keep their own layout. Leave
HOME, Path.home(), production resolver behavior, and the I/O guard intact.
On the original base, this fixture fixes all 24 PM authority failures and
92 update failures/setup errors. The same four unrelated /proc DB-holder
probe failures remain on both base and current code. Current targeted
profile, path, PM, and guard checks pass: 214 passed, 4 skipped.
Keep upstream's reviewed catalog as the only plugin name index.
Catalog pins and custom update sources share staged PM validation.
Publish code and dependencies with recovery after process death.
Reject a concurrent enablement change before publishing disabled code.
Use the manifest loader's supported version in the installer. Keep
probe cooldowns for timeouts, not TLS failures that a CA change fixes.
Preserve the backup, uninstall, browser and memory-provider repairs.
Verified with the canonical runner on native Windows ARM64, real Git
repositories, local TLS endpoints and UV dependency generations.
Desktop catalog tests and both TypeScript checks pass. The full suite
and native release builds were not run. No remote push.
#106623 (ca16cafee4) blocks any test subprocess whose argv resolves to
`hermes gateway run|start|restart`, so the harness cannot spawn a runtime that
outlives the worker and restarts the developer's gateway. The matcher strips the
`-p <profile>` selector and reads through `sh -c`, so it also fired on
`docker exec -u hermes <ctr> sh -c 'hermes -p x gateway start'` and broke
tests/docker/test_profile_gateway.py on every Docker build since. A gateway
launched inside a container cannot reach the host unit or webhook port; the
guard now skips commands whose argv[0] is a container runtime (docker, podman,
nerdctl). Host-side spawns stay blocked.