16 Commits

Author SHA1 Message Date
ethernet
330ff28d44 test: platform lanes after the main merge
- test_windows_build_deps: the fixture Get-Command knows link.exe (MSVC linker
  pin) and asserts the pin; real PowerShell children on a cold runner exceed 30s
- test_runtime sealed worker: symlink the interpreter instead of copying it so
  a framework build still finds its stdlib (macOS)
- test_code_execution_modes: compare realpaths (/var/tmp vs /private/var/tmp)
- test_post_update_expose_cli: --help reaches the boot bootstrap; --version is
  answered on the pre-import fast path and never exposes shims
2026-09-20 00:15:40 -04:00
ethernet
975572451f test: replace child environment ladders with process contracts 2026-09-13 15:10:31 -04:00
ethernet
92686159d1 fix(pm): integrate audited runtime and lifecycle repairs
Prepare dependency generations before selecting them. Keep shipped tool
bytes separate from writable additions, and store facts beside their entries.
Validate proposed plugin sets before config publication. Restore the previous
config if the facts write fails.

Consolidate duplicate updater, backup, setup, and voice helpers. Repair
launcher selection, dependency consumers, download ownership, update feeds,
and native Windows process and file handling.

Verification: 206 changed/prior-failing Python files reported 4630 passed,
one failed, and 330 skipped. Fix the remaining Hindsight fixture boundary.
The final targeted rerun reported 234 passed and two skipped. The store
review regression batch reported 83 passed and one skipped. Desktop
TypeScript checks, 56 selected Electron tests, 24 release tests, and the
removed-import/compatibility guards passed.

This is an integration checkpoint, not full audit acceptance. The complete
Python suite has not run on this fixed tree. Crash-atomic plugin publication,
generation cleanup, receipt correlation, and packaged lifecycle acceptance
remain open in docs/pm-audit-status.md.
2026-09-05 22:36:48 -04:00
ethernet
49945b1402 refactor(tests): one runner everywhere — xdist removed, per-file isolation on every host
pytest-xdist is gone: run_tests.sh no longer dispatches on host, the
Windows xdist (--dist loadfile) arm is deleted, and pytest-xdist is
dropped from the dev extra (uv lock removes it + execnet). The
per-file subprocess runner (run_tests_parallel.py) is THE runner on
every host — the shape the CI linux lane already used.

- tests-os.yml: the macOS/Windows lanes now run scripts/run_tests.sh
  like every other lane, passing the marker-narrowed file list via
  --files and -m after --. The runner natively tolerates per-file
  empty collections (exit 5, platform-gated file) and fails the run
  when NOTHING collected — replacing the hand-rolled exit-5 branch.
- tests/gateway/conftest.py: the hasattr(config, 'workerinput')
  controller-guard was xdist-only dead code; the file-locked cache
  already handled the per-file model. Removed.
- tests/tools/test_browser_supervisor.py: the port came from xdist's
  worker_id (fixed 9225 under per-file isolation — concurrent files
  collide). Binds an ephemeral port instead (bind 0, read back).
- ~35 comment sites named xdist as the isolation mechanism; they now
  describe the shared-process hazard they actually guard against
  (bare pytest runs, same-file ordering) without naming a runner that
  no longer exists.
2026-09-04 19:01:41 -04:00
Teknium
98c140bc4b simplify(compat): code_execution_tool/environments.local — drop 36 re-exports, repoint 5 callers + 14 test files 2026-09-03 13:24:04 -07:00
Teknium
606cb2de92 refactor(tools/code_exec): unify code_kernel local/remote helpers, split checkpoint_manager god methods, compact spill helpers 2026-09-02 14:44:15 -07:00
Teknium
4e7eb39947 refactor(code-execution): session kernels always on — kernel_mode knob retired (#96787)
* refactor(code-execution): retire kernel_mode — session kernels always on for local runs (remote per-call is a tracked gap, not a mode)

* test(code-execution): env-filtering probes use reset=true — kernel env is frozen at spawn, so env rules are only observable on a fresh kernel

* test(code-execution): kernel-aware fixes for mode/pythonpath suites — reset=true on frozen-at-spawn probes, per-test kernel disposal, abort-after-capture fake Popen

* test(code-execution): strict-mode cwd is a behavior contract (staging tmpdir, not session cwd) — kernel stages in hermes_kernel_*, per-call in hermes_sandbox_*
2026-08-27 22:22:39 -07:00
kshitij
222465d847 refactor(tools): unify probe caches and dedupe the exclusion log
/simplify-code findings on the full PR diff:

- _is_usable_python had the same sticky-failure bug the previous commit
  fixed in _python_environment_prefix: lru_cache pinned a transient
  probe failure (fork pressure, timeout) as False forever, silently
  locking project mode to sys.executable. Both probes now share a
  success-only bounded dict cache via _cache_probe_result() with FIFO
  eviction at _PROBE_CACHE_MAX (the old < cap guard stopped caching new
  entries instead of evicting, re-probing entry 33+ on every call).
- The hermes-root-omitted logger.info fired on every external-env call
  in project mode; now deduped once per interpreter path per process
  (matching the tirith/mcp warn-once convention).
- Regression test: _is_usable_python probe failures are retried, not
  cached (mutation-verified).
2026-08-12 17:42:34 +05:30
kshitij
89556c63ac fix(tools): harden interpreter-environment probe for the strict-mode default
Follow-up to the salvaged #81201 commits:

- Short-circuit _uses_hermes_python_environment when the child IS the
  running interpreter (path or realpath match). The default strict-mode
  path no longer spawns a probe subprocess at all, and a flaky probe of
  sys.executable can never drop the hermes root from PYTHONPATH
  (protects the test_repo_root_modules_are_importable invariant). The
  realpath leg also covers uv-style venvs whose bin/python resolves to
  the same binary.
- Stop caching failed probes: _python_environment_prefix now uses a
  success-only dict cache instead of lru_cache, so one transient
  timeout under load no longer sticks for the process lifetime.
- Deduplicate the subprocess probe scaffolding shared with
  _is_usable_python into _probe_python().
- Log once when the hermes root is omitted so import-behavior changes
  are diagnosable from user reports.
- Tests: fail the composition tests loudly if execute_code never
  reaches Popen (was vacuously passing on exceptions); assert the
  staging dir is literally first in PYTHONPATH (was truthiness only);
  add guards for probe-failure retry and the no-probe short-circuit.
2026-08-12 17:42:34 +05:30
Elisa Martinez Abad
ec884fc655 feat(tests): add tests to cover external-venv PYTHONPATH isolation 2026-08-12 17:42:34 +05:30
Teknium
39975613b1 test: prune wave 2 + speed fixes — 28,106 → 19,757 test functions, suite wall 315s → 294s
Second, deeper pass over tools/gateway/hermes_cli plus first pass over
the trees wave 1 missed (acp, acp_adapter, skills, computer_use, docker,
dashboard, conformance, monitoring, secret_sources, hermes_state,
providers). Same rubric as wave 1 (AGENTS.md test policy); security,
alternation/caching invariants, issue-number regressions, and E2E kept.

Real test-quality fixes found and rooted out along the way:
- tests/tools/test_command_guards.py made real auxiliary-LLM HTTPS calls
  (DEFAULT_CONFIG smart-approval leaked in) — pinned approval
  mode=manual via autouse fixture: 17.4s → 0.4s.
- test_model_switch_custom_providers.py / test_user_providers_model_switch.py
  silently probed live provider catalogs (~2s/test) — stubbed
  cached_provider_model_ids/provider_model_ids/fetch_api_models.
- test_telegram_noise_filter.py: 15-platform copy-paste matrix over
  shared gateway.run logic → 3 representative platforms (55s → 3.9s).
- test_gateway_shutdown.py: stop()'s 5s interrupt-deadline loop spun on
  MagicMock agents — interrupt.side_effect now clears _running_agents
  (22s → 1.0s).
- test_gateway_inactivity_timeout.py poll-harness timings shrunk 3-5x
  (24s → 1.1s); test_mcp_stability.py backoff/SIGTERM-grace sleeps
  patched (15.4s → 2.5s); test_async_delegation.py negative-drain wait
  5s → 0.5s.
- test_telegram_init_deadline.py: loop-block margin restored to 1.0s
  with rationale comment — the watchdog-dump assertion needs the loop
  blocked well past deadline+grace under parallel load (flaked once in
  the 40-worker verification run at a 0.2s margin).

Verification: full hermetic suite via scripts/run_tests.sh —
2,438 files, 21,718 tests passed, 0 failed, 293.9s wall.
Suite totals vs original baseline: 46,820 → 19,757 test functions
(−57.8%), wall 583.5s → 293.9s (−50%), subprocess CPU 13,564s → 11,623s.
2026-07-29 13:39:40 -07:00
Teknium
a3c0de1a36 test(execute-code): cover session cwd record precedence in project mode
Follow-up for salvaged PRs #56055 + #56803, adapted to the per-session
cwd record store (PR #65213): the record (live cd state) is rung 1,
the registered session.cwd.set override rung 2, TERMINAL_CWD rung 3 —
the same ladder file tools and the terminal resolve against. Covers
record-over-override, record-only, and stale-record fall-through.
2026-07-16 04:24:07 -07:00
Gaurav Saxena
4d6686c18a fix(execute-code): honor session cwd overrides 2026-07-16 04:24:07 -07:00
kshitij
5fba236644 chore: ruff auto-fix PLR6201 resweep — tuple → set in membership tests (#27355)
Six days after #23937 (608 fixes) the codebase had accumulated 241 new
PLR6201 violations. Same mechanical `x in (...)` → `x in {...}` fix,
same zero-risk profile: set lookup is O(1) vs O(n) for tuple and the
two are semantically equivalent for hashable scalar membership tests.

All 241 instances fixed via `ruff check --select PLR6201 --fix
--unsafe-fixes`, zero remaining. Every changed value is a hashable
scalar (str/int/None/enum/signal); no risk of unhashable runtime
errors. No behavior change.

Test plan:
- 119 files changed, +244/-244 (net zero) — exactly one-line edits
- `ruff check` clean afterward
- Compile checks pass on the largest touched files (cli.py, run_agent.py,
  gateway/run.py, gateway/platforms/discord.py, model_tools.py)
- Subset broad test run on tests/gateway/ tests/hermes_cli/ tests/agent/
  tests/tools/: 18187 passed, 59 pre-existing failures (verified against
  origin/main with the same shape — identical failure count, identical
  category — all xdist test-order flakes unrelated to this change)

Follows the same template as PR #23937 ([tracker: #23972](https://github.com/NousResearch/hermes-agent/issues/23972)).
2026-05-17 02:29:41 -07:00
Teknium
e614e87954 tests: skip POSIX-venv-layout tests on Windows
test_code_execution_modes.py had two test-level failures and two
class-level stale skip reasons on this Windows-native branch:

  - TestResolveChildPython::test_project_with_virtualenv_picks_venv_python
  - TestResolveChildPython::test_project_prefers_virtualenv_over_conda

Both fail on Windows with OSError: [WinError 1314] — they call
pathlib.Path.symlink_to() to build a fake venv, which requires
developer mode or admin on Windows.  They also assume POSIX venv
layout (bin/python) where Windows uses Scripts/python.exe.  Skip
them with a specific, accurate reason.

Also updated two class-level skipif reasons that said
'execute_code is POSIX-only' — no longer true on this branch.
New reason explains it's the test infrastructure (symlinks + POSIX
venv layout) that's the blocker, not execute_code itself.

Results on Windows Python 3.11:
  Before: 41 passed, 10 skipped, 2 failed
  After:  43 passed, 12 skipped, 0 failed
2026-05-08 14:27:40 -07:00
Teknium
285bb2b915 feat(execute_code): add project/strict execution modes, default to project (#11971)
Weaker models (Gemma-class) repeatedly rediscover and forget that
execute_code uses a different CWD and Python interpreter than terminal(),
causing them to flip-flop on whether user files exist and to hit import
errors on project dependencies like pandas.

Adds a new 'code_execution.mode' config key (default 'project') that
brings execute_code into line with terminal()'s filesystem/interpreter:

  project (new default):
    - cwd       = session's TERMINAL_CWD (falls back to os.getcwd())
    - python    = active VIRTUAL_ENV/bin/python or CONDA_PREFIX/bin/python
                  with a Python 3.8+ version check; falls back cleanly to
                  sys.executable if no venv or the candidate fails
    - result    : 'import pandas' works, '.env' resolves, matches terminal()

  strict (opt-in):
    - cwd       = staging tmpdir (today's behavior)
    - python    = sys.executable (today's behavior)
    - result    : maximum reproducibility and isolation; project deps
                  won't resolve

Security-critical invariants are identical across both modes and covered by
explicit regression tests:

  - env scrubbing (strips *_API_KEY, *_TOKEN, *_SECRET, *_PASSWORD,
    *_CREDENTIAL, *_PASSWD, *_AUTH substrings)
  - SANDBOX_ALLOWED_TOOLS whitelist (no execute_code recursion, no
    delegate_task, no MCP from inside scripts)
  - resource caps (5-min timeout, 50KB stdout, 50 tool calls)

Deliberately avoids 'sandbox'/'isolated'/'cloud' language in tool
descriptions (regression from commit 39b83f34 where agents on local
backends falsely believed they were sandboxed and refused networking).

Override via env var: HERMES_EXECUTE_CODE_MODE=strict|project
2026-04-18 01:46:25 -07:00