fc2972b29980da345c835acdb04ac87ff1874021
23 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
cfa327e5dc | refactor(hclib): remaining hermes_cli library modules — dead code, unified helpers, flattened branches | ||
|
|
78335adfec | refactor(hclib): shared roots — hermes_constants, hermes_logging, utils, _subprocess_compat compaction and dedupe | ||
|
|
f6234d00c5 |
fix(security): close GitSpawn RCE class — malicious repo .git/config no longer executes on context gathering (GHSA-7x36-8jrh-v4pw)
Hermes gathers workspace context by running git against the session directory automatically — the coding-workspace snapshot, gateway project-tree build, /diff, @diff|@staged context refs, goal-gate fingerprint, and -w startup worktree add — before any prompt, tool call, approval, or trust gate. Those probes ran the system git without stripping the repository's own config, so a repo delivered as files with its .git directory intact (a shared zip, sync folder, or USB stick; git clone never transfers .git/config) could set an execution-sink git setting and get arbitrary host code execution as the user with nothing on screen. - core.fsmonitor / core.hooksPath / pager / editor / credential helper: neutralized by routing every automatic probe through noninteractive_git_env(), which pins those keys to inert values via GIT_CONFIG_* and ignores global/system config. bounded_git_probe (the reported sink, coding_context._git + tui_gateway.git_probe) now defaults to that env; worktree-add, working_diff, web_git, context_references, goals, and subagent_worktree route through it too. - Attribute-scoped [diff "x"] command=/textconv= drivers: the attacker names the driver in .gitattributes, so GIT_CONFIG_KEY overrides can't enumerate them. Added harden_git_argv(), which inserts --no-ext-diff --no-textconv on diff-rendering subcommands (diff/show/ log/blame) only — status et al reject the flags. Both flags required (verified empirically; each alone leaves the other live). Builds on the noninteractive_git_env config-scrubbing from the gemini-cli #28792 port. Real-git E2E regression suite arms a malicious repo and asserts every automatic path neutralizes fsmonitor, hooks, external-diff, and textconv; a baseline test proves the repo is armed. |
||
|
|
413b6ba3dd |
Port from google-gemini/gemini-cli#28792: harden internal git env
(cherry picked from commit
|
||
|
|
90e916efc9 |
fix(windows): compose the taskkill identity guards into one fail-closed class fix
Salvage hardening on top of the three cherry-picked contributor commits (#91297 gebilaowang404 + AlexMnrs, #96741 burak33bb, #98826 ayushnangia), closing the remaining unverified-PID kill sites as one class (#98814, #89614): - pid_is_hermes: token-boundary 'hermes' match (no more loose substring false-positives), and an explicit start-time expectation is now honored on POSIX too (a mismatched fingerprint is a recycled PID on any platform). - kill_process_tree: drop the guard on our OWN retained Popen child — a retained handle pins the PID, so the check could only false-refuse. - gateway.status.terminate_pid: POSIX force-kills also refuse when a caller-provided expected_start_time no longer matches. - kill_gateway_processes: re-verify the LIVE cmdline at kill time (the scan-time match is a TOCTOU window). - _reap_unsupervised_gateway_orphans: fingerprint orphans at scan time and require a still-matching identity before the delayed SIGKILL escalation. - whatsapp _kill_port_process: never kill a bare netstat/lsof-scanned PID unless the live process is actually a node bridge (was a stranger-kill). - browser daemon reap/close paths: pass the start-time fingerprint into ProcessRegistry._terminate_host_pid (previously unverified), and the session-close path now runs the same daemon identity verification as the orphan reaper. - tests/hermes_cli/test_taskkill_identity_windows_live.py: live Windows probes (real spawned processes, real psutil ancestry) wired into the on-demand windows-latest wine2e lane. Fixes #98814 Fixes #89614 |
||
|
|
cdd063528f |
fix(hermes_cli): fail-closed PID-ownership guard before Windows taskkill
Guard every Windows `taskkill /PID` against stale/recycled PIDs (#89614: 8x 0xEF blue screens; a rebooted PID can be svchost.exe). Adopted the community patch by AlexMnrs (commit 0162465): shared psutil-based (pid, create_time) guard reusing the repo's existing get_process_start_time machinery: - fail closed on invalid/unknown/recycled identities (0/-1/None/bool/non-int) - capture identity at discovery, re-validate at kill time - all three sites through pid_is_hermes; taskkill stays hidden Sites: _subprocess_compat.kill_process_tree, dashboard_procs._kill_stale_dashboard_processes (win32), update_cmd._stop_process_trees. Refs #90471, #89614 Co-authored-by: Alex Monrás <AlexMnrs@users.noreply.github.com> |
||
|
|
547f985286 |
refactor(deadline): consolidate site-local tree-kills onto agent.deadline.kill_process_tree (#85125 4d)
Per-site decisions:
1. hermes_cli/_subprocess_compat.py kill_process_tree(proc) -> None:
MIGRATED. Body now delegates to agent.deadline.kill_process_tree(proc.pid)
via a function-local import; keeps the swallow-everything fail-open
contract and the (proc) -> None signature (agent/shell_hooks.py imports
it by name; _kill_git_process_tree alias preserved). The old body is kept
verbatim as _legacy_kill_process_tree and used as fallback when the
delegation import/call fails. A final proc.kill() is retained on the
happy path so Popen bookkeeping sees the exit (matches old behavior).
2. tools/browser_tool.py _kill_process_tree(proc): MIGRATED, same pattern
(delegate + _legacy_kill_process_tree fallback). Behavior delta: the old
body sent SIGTERM then SIGKILL with zero grace between them; the shared
primitive sends SIGKILL only. With no grace period the observable effect
is identical, and the psutil descendant sweep now also reaches
agent-browser's setsid'd daemon grandchild, which killpg alone missed.
tests/tools/test_browser_npx_warmup.py's TestKillProcessTree repointed at
the legacy fallback (its assertions describe the fallback's internals).
3. tools/code_execution_tool.py _kill_process_group(proc, escalate):
MIGRATED. It was a plain parent+descendants terminate (then wait 5s +
kill when escalate=True) — expressed as two delegated calls:
kill_process_tree(pid, sig=SIGTERM), then on escalate-timeout
kill_process_tree(pid, sig=SIGKILL). Delegation failure degrades to
proc.kill(), mirroring the old psutil-failure fallback. Delta: the old
body terminated children before the parent; the shared primitive
signals the group atomically (child is a session leader via
start_new_session=True) plus an identity-aware descendant sweep —
strictly wider coverage, same signals.
4. gateway/status.py: KEPT BOTH SITES.
- terminate_pid (~l305) taskkill wrapper: NOT migrated. Its contract is
incompatible with the shared primitive — it must RAISE OSError with
taskkill's stderr on non-zero exit (callers branch on that), falls back
to os.kill on FileNotFoundError, and its POSIX branch is deliberately a
single-PID SIGTERM/SIGKILL, not a tree kill. Wrapping the bool-returning
fail-soft primitive would invert the error contract.
- reap_gateway_children (~l2029): NOT migrated. It operates on a
pre-snapshotted child list from a parent that is already dead
(psutil.Process(pid) on the parent would fail), and every signal is
wrapped in identity/ownership checks the primitive lacks: is_running()
identity, zombie skip, and the skip-if-ppid-still-equals-parent guard,
plus SIGTERM -> wait_procs -> SIGKILL staging and a reaped-count return.
The coupling is the feature; migrating would delete the safety logic.
5. scripts/run_tests_parallel.py _kill_process_tree (~l253): NOT migrated.
Dev tooling that intentionally kills by CAPTURED pgid because the direct
child is usually already reaped (psutil/pid-based primitive cannot find
it), and it avoids the psutil import on the test-runner hot path. Its
docstring already documents why psutil is the wrong tool there.
New tests: tests/agent/test_treekill_consolidation.py — delegation +
raise-swallowing tests per migrated wrapper, consumer-identity checks, and
a live end-to-end probe (setsid grandchild dies through the compat wrapper,
zero survivors).
|
||
|
|
4e3de140c1 |
fix(cli): bound the Windows process-scan probes so a slow WMI scan cannot wedge hermes update (#87134)
subprocess.run(capture_output=True, timeout=N) is not hang-safe on Windows: after the timeout fires, run()'s cleanup kills the direct child and then joins the pipe reader threads with an UNBOUNDED communicate(). A descendant (conhost.exe under wmic/powershell) holding duplicated pipe handles keeps the pipes from EOF and the join never returns. _scan_gateway_pids() runs its wmic / Get-CimInstance Win32_Process scans exactly that way, and on machines where the full process scan genuinely exceeds its 10/15s budget (cold WMI on first boot, ARM VMs, heavy Update/AV activity) hermes update wedged forever inside _pause_windows_gateways_for_update() before printing a single line — observed live on a fresh Windows 11 ARM64 VM with a faulthandler stack pinning the main thread in subprocess._communicate and only a conhost.exe child surviving. The single-flight update lock then blocks retries until the wedged process is killed by hand. This is the same deadlock class bounded_git_probe already fixed for git probes (#68609 / #66037). Generalize that proven pattern into a shared bounded_probe_run() — explicit communicate(timeout), kill_process_tree on failure, bounded 1s drain, then abandon the daemonic readers — and migrate the whole call-site class onto it: - hermes_cli/gateway.py _scan_gateway_pids (the site that hung; reached from hermes update, cron, gateway restart/status, dashboard) - hermes_cli/dashboard_procs.py wmic scan (same shape, reached on update) - hermes_cli/claw.py tasklist + PowerShell probes (same shape; its try/except cannot catch a hang because a hang raises nothing) - bounded_git_probe now delegates to bounded_probe_run (identical contract, one copy of the cleanup logic) Unlike bounded_git_probe, bounded_probe_run returns the CompletedProcess (or None) rather than collapsing to stdout, because the gateway scan branches on returncode to trip its wmic -> powershell fallback. Tests: tests/hermes_cli/test_bounded_probe_run.py covers success, nonzero-exit passthrough, spawn failure, bounded timeout (fails against the old unbounded semantics — verified by sabotage), errors= decoding, DEVNULL stdin, POSIX process-group placement, and the bounded_git_probe delegation contract. Existing test_git_probe_tree_kill.py passes unchanged against the delegated implementation. Closes #87134 |
||
|
|
0a1cca5648 | refactor(gateway): share breakaway marker constant | ||
|
|
3b9d1b3cde |
fix(hooks): kill the whole process tree when a shell hook times out
Port from openai/codex#37527: Terminate timed-out hook process trees. A shell hook that forked helpers (scanners, watchers, "cmd &") and then hit its timeout left those descendants running forever — subprocess.run() only kills the direct child. Worse, descendants holding the inherited pipe write ends could stall run()'s post-kill communicate() drain. - agent/shell_hooks.py _spawn(): spawn hooks in their own process group on POSIX (process_group=0, Python >=3.11); on timeout/error, reap the whole tree via the shared kill_process_tree() helper, then drain bounded (1s). Hooks that complete in time keep their descendants, so intentionally detached helpers survive successful runs (mirrors codex semantics). - hermes_cli/_subprocess_compat.py: rename _kill_git_process_tree -> kill_process_tree (it was never git-specific; taskkill /T /F on Windows, ownership-gated os.killpg on POSIX). Backward-compat alias retained. - tests/agent/test_shell_hooks_tree_kill.py: real-subprocess regression tests (descendant killed on timeout, preserved on success, own-group spawn, fast-path contract, fail-open). Sabotage-verified: reverting the process_group spawn fails exactly the two new behavior tests. Gap proven live on main first: a forking hook timed out at 2s and its descendant survived; same probe against this branch shows it reaped. |
||
|
|
ee472a7fdb |
fix: Windows agent-loop papercuts — path splitting, hashing, autocomplete, screenshots, OS detection (#84419)
Sweep of open Windows issues affecting day-to-day agent operation (explicitly excluding install/setup and locale classes): - hermes_cli/_subprocess_compat.py: new split_command_line() — Windows- safe command-line tokenizer (posix=False + quote stripping) so backslash paths survive. POSIX behavior unchanged (plain shlex.split). - hermes_cli/console_engine.py (#83934): console commands like 'sessions export C:\Users\me\out.jsonl' no longer silently mangle the path into a relative filename in the cwd. - agent/shell_hooks.py (#78293): hook commands with backslash paths now spawn, resolve their script path, and pass hooks doctor instead of reporting 'not executable'. All three shlex sites routed through the shared splitter. - agent/prompt_builder.py (#51755): system prompt now reports Windows (11) on Windows 11 — platform.release() returns 10 for both; distinguish via sys.getwindowsversion().build >= 22000. - hermes_cli/commands.py (#42016): @ autocomplete no longer crashes the prompt_toolkit event loop when rg emits a path on a different mount (device paths \.\nul, other drive letters) — relpath ValueError is skipped per-entry. - tools/browser_use_cli.py (#83884): screenshot-path detection now matches Windows drive-letter paths (C:\... and C:/...) in addition to POSIX; Browser Use screenshots attach on Windows. - tools/skills_hub.py + tools/skills_guard.py (#62310): the two 'MUST stay symmetric' skill content hashes actually agree on Windows now. Bundle keys are normalized to POSIX separators before hashing, and the disk digest sorts by rel-posix STRING (case-sensitive) instead of Path objects (case-insensitive on Windows). Fixes permanent false-positive update_available for every installed skill. Tests: tests/tools/test_windows_agent_loop_papercuts.py — 16 cases covering each fix, including a disk-vs-bundle hash symmetry check built with native Windows separators and a mixed-case filename. |
||
|
|
aec331899e | chore: suppress windows-footgun false positive on gated killpg | ||
|
|
42e92c9c09 |
fix(git): kill the whole probe process tree on timeout (port of openai/codex#36793)
Timing out a bounded git probe must not leave helper descendants (credential helpers, git-remote-https, hook children) running after the probe fails open. bounded_git_probe now spawns the child in its own process group on POSIX (process_group=0), and _kill_git_process_tree signals the whole group with os.killpg — gated on the child actually leading its own group (pgid == pid), so a shared-group spawn can never take down unrelated processes. Windows keeps the existing taskkill /T /F tree kill. Proven live on main: a fake git that forks a 300s descendant left the descendant running after the probe timeout; with the fix the descendant dies with the launcher. Fast path and fail-open contract unchanged. Port of openai/codex#36793 (Terminate timed-out Git process trees). |
||
|
|
58708c7066 |
fix(git): never block internal git calls on credential prompts
Port from openai/codex#34540 / #34612 ("detach non-interactive subprocesses from stdin"): internal git invocations that run with nobody attached — MCP catalog installs, plugin install/update, profile distribution staging, worktree base fetches, and the desktop review pane's git/gh backend — could hang on a credential prompt when a remote is private, misconfigured, or requires auth. git prompts on the inherited terminal (or via Git Credential Manager on Windows), so the operation silently waits until its timeout, or forever at sites without one (mcp_catalog clones have no timeout at all and inherit the parent terminal). - Add noninteractive_git_env() to hermes_cli/_subprocess_compat.py: GIT_TERMINAL_PROMPT=0 + GCM_INTERACTIVE=Never on a copy of the environment; GIT_ASKPASS/SSH_ASKPASS deliberately preserved so working non-interactive auth still succeeds. - Wire it + stdin=DEVNULL into: mcp_catalog._do_git_install (clone/ checkout), plugins_cmd (clone + pull), profile_distribution._git_clone, web_git._git/_gh (gh also gets GH_PROMPT_DISABLED=1), and cli.py's worktree base fetch helper. - Tests: env contract, a real-git E2E against a local 401 Basic-auth HTTP server proving fail-fast ("terminal prompts disabled") instead of a hang, and per-call-site plumbing assertions. Sabotage-verified: removing the env from web_git._git fails the site test. |
||
|
|
5c5960d9f9 |
fix(windows): suppress console window flashes in env probes, lazy installs, and platform.win32_ver()
From windowless processes (the pythonw gateway and the kanban workers it
spawns), three spawn paths flash visible console windows on Windows:
1. tools/env_probe.py::_run() ran its interpreter/pip probes
(python3 / python / pip / 'python3 -m pip' / PEP-668 check, ~5 per
worker start) without creationflags — one console flash per probe.
2. tools/lazy_deps.py had four spawn sites with the same defect:
'uv pip install', the 'pip --version' probe, ensurepip, and the
pip install fallback.
Both now pass creationflags=windows_hide_flags() (CREATE_NO_WINDOW on
Windows, 0 on POSIX) — stdio capture still works because the child is
hidden, not detached.
3. CPython 3.11's platform.win32_ver() unconditionally calls
_syscmd_ver(), which runs 'cmd /c ver' via
subprocess.check_output(shell=True) with no window suppression. Any
dependency touching platform.uname()/version()/platform() at import
time flashes one 'cmd' window per windowless process. New helper
_subprocess_compat.suppress_platform_ver_console() (Windows-only,
never raises) stubs platform._syscmd_ver so win32_ver() falls back to
sys.getwindowsversion().platform_version — verified byte-identical
platform.platform() output on CPython 3.11
('Windows-10-10.0.26100-SP0' either way). Called at the top of
hermes_cli/main.py, right after the hermes_bootstrap guard, before
heavyweight imports.
Verified on Windows 11 by polling EnumWindows at ~15 ms and attributing
new visible HWNDs to the suspect process tree (conhost child presence is
NOT evidence of a visible window — it appears even with
CREATE_NO_WINDOW). Tests: tests/tools/test_windows_native_support.py,
test_env_probe.py, test_lazy_deps.py, test_lazy_deps_durable_target.py —
153 passed; the 3 failures are pre-existing on upstream/main in a
Windows environment (POSIX-only assertions and NTFS chmod semantics).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
0dbf639bc8 |
fix(windows): hidden-console daemons — extend the parent-console fix to every detached spawn path (#70205)
Extends the desktop backend's root-cause fix (
|
||
|
|
967e078ae4 |
fix(windows): share one bounded, tree-killing git probe across both call sites (#68997)
subprocess.run(["git", ...], timeout=...) deadlocks on Windows: run()'s
post-timeout cleanup calls an unbounded communicate() after killing git.
Killing the PATH-resolved launcher can leave a suspended descendant git.exe
holding duplicates of the captured stdout/stderr handles, so the pipes never
reach EOF and the reader-thread join blocks forever — leaking a process +
two reader threads per fired timeout (the accumulating git.exe load behind
Windows Defender CPU spikes).
Two fail-open probe call sites had this identical flaw:
- tui_gateway/git_probe.py::run_git — on the Desktop agent-build path
(_start_agent_build -> _session_info -> branch() -> run_git), where the
hang turned an optional branch label into "agent initialization timed
out" (#68609).
- agent/coding_context.py::_git — hangs the agent turn inside
build_coding_workspace_block under an ACP host (#66037).
Consolidate both onto one shared bounded_git_probe() in
hermes_cli/_subprocess_compat.py (both files already import from there, so
no new import surface):
- explicit communicate(timeout), then on ANY failure a tree-kill —
proc.kill() AND, on Windows, best-effort taskkill /T /F so the suspended
descendant that holds the pipe writers dies too — plus a bounded 1s
post-kill drain; if the pipes are still held they're abandoned (the
orphaned reader threads are daemonic and cost nothing).
- fail open to "" on every path: spawn error, timeout, kill() raising
(access denied / already reaped — a raise inside the except handler
previously escaped the contract), and non-timeout communicate() failures
now also terminate the child instead of leaving it running.
- the taskkill spawn can't re-enter the deadlock class: it captures no
pipes (DEVNULL), so its own timeout cleanup has no reader threads to join.
Normal-path spawn contract is preserved byte-for-byte: PIPE/PIPE/DEVNULL,
text + utf-8 errors="replace", hidden-window creationflags on Windows only,
nonzero returncode -> "". Each call site keeps its own timeout (1.5s / 2.5s).
Supersedes #68622 (Sora-bluesky — git_probe fix + tree-kill) and #66038
(iamwongeeeee — coding_context fix), folding both into one shared helper so
the two sites can't drift and every timeout tree-kills the descendant. Tests
consolidated onto the helper, incl. the previously-missing assertion that a
Windows timeout escalates to taskkill /T /F.
Co-authored-by: Sora-bluesky <sora.bluesky.dev@gmail.com>
Co-authored-by: iamwongeeeee <wykim777@naver.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d3d621f7c3 |
revert(windows): roll back terminal-popup PRs #53791 #53810 #53829 (#53853)
* Revert "fix(windows): capture is not a no-window boundary; route flashing spawns through chokepoint (#53829)" This reverts commit |
||
|
|
5db1430af9 |
fix(windows): stop terminal-window popups from background spawns (#53810)
* fix(windows): stop terminal-window popups from background spawns Native-Windows desktop/gateway users saw cmd/conhost windows flash on gateway restart, image paste, the dashboard Projects tree, voice notes, and ~5 min after closing the app (detached cron). Two root causes: - Console-subsystem exes (taskkill, schtasks, wmic, netstat, tasklist, agent-browser, git, ffmpeg, powershell, git-bash) spawned via raw subprocess allocate a fresh console when the launching process has none (pythonw desktop backend / detached gateway) - even with output captured. - uv venv pythonw shims re-exec console python.exe, so Python children get a console regardless of how they're launched. Fixes: - Single hidden-spawn primitive (_subprocess_compat.run/.popen) that ORs CREATE_NO_WINDOW on Windows, no-op on POSIX. Route every Hermes-owned console-exe spawn through it. - FreeConsole() catch-all in hermes_bootstrap: any Python child that exclusively owns an auto-allocated console detaches it at startup (GetConsoleProcessList()==1 gate leaves shared interactive consoles untouched). - Replace PowerShell/wmic gateway PID scans with in-process psutil. - Skip schtasks queries on non-interactive desktop restarts. - Prefer native agent-browser .exe over .cmd shims. - Guard test bans raw subprocess spawns of the Windows-only console tools repo-wide so the popup class can't regress. * fix(windows): scope FreeConsole to background entry points; fix merge fallout Console detach review (per #53810 feedback): GetConsoleProcessList()==1 can't tell a uv pythonw->python phantom console apart from a user opening the interactive CLI/TUI in its own fresh console (double-click, shortcut, ConPTY) — both report a single attached process with a tty. Running FreeConsole() in the import-time bootstrap therefore risked detaching a legitimately-interactive terminal. - Extract FreeConsole into explicit hermes_bootstrap.detach_orphan_console(); remove it from apply_windows_utf8_bootstrap() (import side effect). - Call it only from known background mains: gateway run, dashboard backend (start_server, what the desktop spawns), cron standalone, tui_gateway entry, slash worker. Interactive CLI/TUI never calls it. - Behavior-contract tests: frees only when solo owner, leaves shared console, no-op without console / on POSIX, and asserts it's not an import side effect. Merge fallout from origin/main (#53791): - local.py: 3-way merge left a dangling **_popen_kwargs (NameError crashing every terminal init). _subprocess_compat.popen already hides the window, so drop it. - discord adapter: merge stacked an undefined windows_hide_flags() onto the primitive call; drop the redundant arg. - test_gateway: scan now goes psutil-first (zero spawn); rewrite the case-variant test to drive that production path. * test(claw): mock _subprocess_compat.run seam for Windows process scan claw.py's Windows tasklist/powershell scan routes through the hidden-spawn primitive; the tests still patched claw_mod.subprocess, so on win32 the mock was never hit and real spawns returned nothing. Patch the actual seam. |
||
|
|
fe0b3f2338 |
fix(windows): retry watcher Popen without breakaway when parent job denies it, plus regression tests for the breakaway bit (#40956)
#40909 added `CREATE_BREAKAWAY_FROM_JOB` to `windows_detach_flags()`, which fixed the headline bug (gateway dies after Desktop GUI update and never comes back). The flag's own docstring acknowledges that restrictive parent job objects can still refuse breakaway with `ERROR_ACCESS_DENIED`, surfacing as `OSError` on the `subprocess.Popen` call: "Callers in this codebase already wrap detached spawns in try/except OSError and fall back to a cmd.exe wrapper, so the breakaway-denied case degrades gracefully rather than crashing." That's true for `_spawn_detached` in `gateway_windows.py` (the `hermes gateway start` path), which has both the breakaway bit AND a retry-without-breakaway fallback. It's NOT true for the post-update watcher path in `launch_detached_profile_gateway_restart` (`hermes_cli/gateway.py`), which only has `except OSError: return False` and gives up entirely. If a user's shell/terminal/container wraps Hermes in a breakaway-denying job, the gateway-respawn watcher silently fails to launch instead of trying again without breakaway. This PR closes that gap and adds the regression tests that were missing from the original fix. ## Changes ### `hermes_cli/_subprocess_compat.py` Adds a sibling helper `windows_detach_flags_without_breakaway()` so callers can express the fallback symbolically (via the helper) rather than coding the magic `& ~0x01000000` mask at every site. Documented on `windows_detach_flags` and `windows_detach_flags_without_breakaway` with the recommended try/except pattern. ### `hermes_cli/gateway.py::launch_detached_profile_gateway_restart` Two changes, both aligned with the canonical pattern in `gateway_windows._spawn_detached`: 1. The outer watcher Popen now wraps in `try/except OSError`, and on failure retries with `windows_detach_flags_without_breakaway()` (POSIX never reaches this branch — `start_new_session=True` can't raise OSError). 2. The inlined respawn payload (the `python -c` watcher) also wraps its CreateProcess in try/except OSError and retries with `_flags & ~_CREATE_BREAKAWAY_FROM_JOB` on failure. This matters because the watcher's job-object inheritance is independent of the outer process's — even if the outer Popen succeeds with breakaway, the respawned gateway might inherit a job that doesn't. ### Regression tests in `tests/tools/test_windows_native_support.py` #40909 shipped the fix without any test that the breakaway bit is present (the existing `test_windows_detach_flags_has_expected_win32_bits` asserted only the three legacy bits). Four new tests close that: - `test_windows_detach_flags_includes_breakaway_from_job` — explicit assertion that the breakaway bit is in the default bundle, with the rationale spelled out in the docstring so a future maintainer staring at this test understands why removing it would resurrect the gateway-dies-after-GUI-update bug. - `test_windows_detach_flags_without_breakaway_drops_only_that_bit` — fallback payload keeps the other three detach bits intact. - `test_launch_detached_profile_gateway_restart_inlined_watcher_uses_breakaway` — static-text check on the stringified watcher payload. The inlined Python program isn't reachable via normal import-time inspection because it lives in a `textwrap.dedent("""...""")` literal that gets passed to a separate `python -c` interpreter. Asserting that both `_CREATE_BREAKAWAY_FROM_JOB` (symbolic) and `0x01000000` (hex literal) appear inside the dedent block is a sufficient regression guard against accidental refactors. - `test_launch_detached_profile_gateway_restart_outer_popen_has_access_denied_fallback` — static check that this PR's fallback retry is wired up symbolically. Without standing up a real Windows job object that refuses breakaway, we can't trigger the OSError in a unit test; the text guard catches the case where a future refactor removes the helper import or the `& ~_CREATE_BREAKAWAY_FROM_JOB` retry. Also extends `test_windows_detach_flags_has_expected_win32_bits` to include the breakaway bit assertion and updates `test_windows_flags_zero_on_posix` to cover the new helper. ## Tests Locally on Windows: 8/8 in the `-k "detach or breakaway or popen_kwargs or launch_detached or gateway_run_update or hermes_cli_gateway"` slice pass. Broader `tests/hermes_cli/test_gateway*.py + test_windows_native_support.py`: 172 passed, 10 failed. All 10 failures are pre-existing POSIX-only tests running on a Windows host (os.geteuid, SIGKILL fallback, is_linux fixture mismatches). Stashing this PR and re-running on bare post-#40909 main reproduces all 10 identically — none are regressions. POSIX paths unchanged: `windows_detach_flags()` and `windows_detach_flags_without_breakaway()` both return 0 off Windows, `windows_detach_popen_kwargs()` still yields `{"start_new_session": True}`. ## Out of scope - The other detached-spawn site in `hermes_cli/gateway.py` (around line 3068) also uses `windows_detach_popen_kwargs()` + `except OSError`. It deserves the same fallback treatment but the codepath is different enough (not the update-flow watcher) that it warrants a separate PR with its own scrutiny. - `gateway/run.py` has Windows branches with `windows_detach_popen_kwargs` too — same reasoning. ## Context Follow-up to #40909 (merged). I had a parallel PR (#40934, closed) that duplicated the core breakaway fix; the bits unique to that PR that #40909 didn't cover are the contents of this one. Closing #40934 and opening this slimmed-down version as the focused follow-up. |
||
|
|
fc086da8bd |
fix(gateway,windows): reliability — JOB breakaway + status --deep probes + test-leak fix (#40909)
* fix(gateway,windows): reliability — supervisor task, JOB breakaway, status --deep
Three coordinated fixes for the Windows gateway reliability story:
1. CREATE_BREAKAWAY_FROM_JOB on every detached spawn
The 'hermes update' triggered from the Electron Desktop GUI ran inside
Electron's job object. Without breakaway, the post-update gateway
watcher spawned by update — already DETACHED_PROCESS — was still
reaped when Electron's job tore down, so the gateway never came back
after a GUI-initiated update. Adds CREATE_BREAKAWAY_FROM_JOB (0x01000000)
to:
- hermes_cli/_subprocess_compat.py::windows_detach_flags() — used by
every helper that calls windows_detach_popen_kwargs(), including
launch_detached_profile_gateway_restart()
- The watcher subprocess's own respawn snippet in
hermes_cli/gateway.py (inlined flags so the watcher's child
respawn also breaks away)
_spawn_detached() in gateway_windows.py already had the flag; this
change brings the rest of the codebase to parity.
2. Per-minute supervisor Scheduled Task — Windows equivalent of
systemd Restart=always
Introduces hermes_cli/gateway_supervisor.py and registers it as a
second Scheduled Task ('Hermes_Gateway_Supervisor', SC MINUTE /MO 1,
LIMITED rights) alongside the existing ONLOGON task. Every minute,
the supervisor uses the same gateway.status.get_running_pid() probe
as 'hermes gateway status' and, if no gateway is alive, calls
gateway_windows._spawn_detached() (which now includes BREAKAWAY) to
bring one back.
Covers every crash mode, not just 'machine rebooted': taskkill,
OOM, GUI update SIGTERM, parent job teardown. Cheap — one pythonw
startup per minute when down, one PID-existence check per minute
when up.
Wired into both the schtasks-success and Startup-folder-fallback
install paths via _install_supervisor_best_effort(), and removed in
uninstall(). Best-effort: a failing supervisor install logs a
warning but doesn't roll back the primary install.
3. 'hermes gateway status --deep' shows per-probe PASS/FAIL
Replaces the existing terse '--deep' output (which only printed
paths) with an actual diagnostic table:
[1] PID file present
[2] Lock file held by a live process
[3] get_running_pid() result
[4] _pid_exists(pid) — OS-level liveness
[5] gateway_state.json (state + age)
[6] Last lifecycle event from gateway-exit-diag.log
When the high-level summary disagrees with reality, the user can
see exactly which signal is lying.
Test-leak fix
-------------
tests/hermes_cli/test_gateway_wsl.py::TestGatewayCommandWSLMessages
monkey-patched is_linux/is_wsl/supports_systemd_services to simulate
WSL but did NOT stub is_windows(). On a Windows host, the dispatcher
in _gateway_command_inner takes the is_windows() branch BEFORE the
WSL guidance branch, so the test invoked gateway_windows.install()
for real. install() writes to %APPDATA%\...\Startup\Hermes_Gateway.cmd
— the REAL user Startup folder, never sandboxed by tmp_path — pointing
at the test's pytest-of-<user>/pytest-<N>/.../gateway-service/ wrapper.
When pytest tore down the tmp_path, every subsequent Windows login
flashed a cmd.exe window that failed to find the missing target.
Stubs is_windows=False on all four affected tests:
test_install_wsl_no_systemd
test_start_wsl_no_systemd
test_status_wsl_running_manual
test_status_wsl_not_running
Defense-in-depth: _build_startup_launcher() now prefixes the launcher
with 'if not exist <target> exit /b 0', so any future stale Startup
entry silently no-ops instead of flashing a console window.
Status enhancements
-------------------
- status() now reports supervisor task presence alongside the existing
schtasks/Startup info, and nudges the user to reinstall if the
supervisor isn't registered.
- Deep mode dumps both the supervisor task name + script path.
* fix(gateway,windows): drop the per-minute supervisor task — keep breakaway + deep probes
Earlier in this branch we added a per-minute schtasks-based supervisor to
respawn the gateway after crashes / GUI-update SIGTERMs. The implementation
flashed a brief console window on every firing, which stole window focus.
We tried several variants:
- cmd.exe wrapper invoking pythonw -> flashes (cmd.exe is console-subsystem)
- schtasks /TR pointing at pythonw -> flashes (uv venv launcher pythonw is
actually subsystem=Console, not GUI; it respawns the real pythonw)
- schtasks /TR pointing at base uv -> still flashes (Task Scheduler-side
conhost preallocation; documented Windows quirk)
- XML registration with <Hidden>true> -> still flashes (<Hidden> only hides
the task in the Task Scheduler UI, not the spawned window)
Researched what leading projects do:
- Ollama: GUI-subsystem tray exe + Startup-folder shortcut. No supervisor.
- Tailscale: real Windows Service via SCM. Session 0, no console possible.
- Syncthing: --no-console flag inside the binary + Startup folder.
- openclaw: VBS Run(..., 0, False) wrapper. Suppresses the *window* but
Super User Q971162 confirms focus-steal still occurs in some cases.
None of these use a per-minute polling scheduled task. The 'auto-restart on
crash' responsibility belongs INSIDE the daemon (Tailscale's in-process
recovery / Ollama's monitor+worker pair) OR is delegated to the Windows
Service Control Manager — not Task Scheduler.
So this commit drops the supervisor entirely. The CREATE_BREAKAWAY_FROM_JOB
fix in _subprocess_compat.py (from commit c1e5fa433) survives — that is the
*real* fix for problem #2 (GUI-update kills gateway): the post-update
watcher in launch_detached_profile_gateway_restart() now breaks out of
Electron's job object, so the gateway respawn watcher survives the GUI
quit and successfully respawns the gateway.
Surviving from c1e5fa433:
* CREATE_BREAKAWAY_FROM_JOB in hermes_cli/_subprocess_compat.py (fixes #2)
* Inlined breakaway flag in the watcher respawn snippet in gateway.py
* hermes gateway status --deep PASS/FAIL probes (fixes #1 — visibility)
* 'if not exist <target> exit /b 0' guard in _build_startup_launcher
(fixes #3 — silent no-op for stale Startup entries)
* tests/hermes_cli/test_gateway_wsl.py is_windows=False stubs (root cause
of #3 — pytest WSL tests no longer leak Startup entries on Win hosts)
Removed in this commit:
* hermes_cli/gateway_supervisor.py (entire file)
* Supervisor section in hermes_cli/gateway_windows.py (~180 lines):
get_supervisor_task_name, get_supervisor_script_path,
_build_supervisor_cmd_script, _write_supervisor_script,
_install_supervisor_task, is_supervisor_task_registered,
_install_supervisor_best_effort
* _install_supervisor_best_effort() calls in install() (3 spots)
* supervisor cleanup block in uninstall()
* supervisor display lines in status() / status(deep=True)
Future direction (out of scope for this PR): the right place for Windows
'Restart=always' semantics is a real Windows Service installed via
pywin32's win32serviceutil.ServiceFramework — session-0 isolation, SCM
auto-restart, no console window possible. That's a meaningful next-PR
project, not a band-aid.
Tests: 51 pass / 2 pre-existing failures in
tests/hermes_cli/test_gateway_{windows,wsl}.py (the 2 failures are
TestSupportsSystemdServicesWSL cases that fail on origin/main too —
unrelated to this PR).
|
||
|
|
66827f8947 |
chore: prune unused imports and duplicate import redefinitions
Remove unused imports (F401) and duplicate/shadowed import redefinitions (F811) across the codebase using ruff's safe autofixes. No behavioral changes -- imports only. - ~1400 safe autofixes applied across 644 files (net -1072 lines) - __init__.py re-exports preserved (excluded from F401 removal so public re-export surfaces stay intact) - Re-exports that are imported or monkeypatched by tests but look unused in their defining module are kept with explicit # noqa: F401 (gateway/run.py load_dotenv; run_agent re-exports from agent.message_sanitization, agent.context_compressor, agent.retry_utils, agent.prompt_builder, agent.process_bootstrap, agent.codex_responses_adapter) - Unsafe F841 (unused-variable) fixes deliberately skipped -- those can change behavior when the RHS has side effects - ruff lints remain disabled in pyproject.toml (only PLW1514 is selected); this is a one-time cleanup, not a config change Verification: - python -m compileall: clean - pytest --collect-only: all 27161 tests collect (zero import errors) - core entry points import clean (run_agent, model_tools, cli, toolsets, hermes_state, batch_runner, gateway) - static scan: every name any test imports directly from an edited module still resolves |
||
|
|
e93bfc6c93 |
feat(windows): close remaining POSIX-only landmines — TUI crash, kanban waitpid, AF_UNIX sandbox, /bin/bash, npm .cmd shims, cwd tracking, detach flags
Second pass on native Windows support, driven by a systematic audit across
five areas: POSIX-only primitives (signal.SIGKILL/SIGHUP/SIGPIPE, os.WNOHANG,
os.setsid), path translation bugs (/c/Users → C:\Users), subprocess patterns
(npm.cmd batch shims, start_new_session no-op on Windows), subsystem health
(cron, gateway daemon, update flow), and module-level import guards.
Every change is platform-gated — POSIX (Linux/macOS) behaviour is preserved
bit-identical. Explicit "do no harm" test: test_posix_path_preserved_on_linux,
test_posix_noop, test_windows_detach_popen_kwargs_is_posix_equivalent_on_posix.
## New module
- hermes_cli/_subprocess_compat.py — shared helpers (resolve_node_command,
windows_detach_flags, windows_hide_flags, windows_detach_popen_kwargs).
All no-ops on non-Windows.
## CRITICAL fixes (would crash or silently break on Windows)
- tui_gateway/entry.py: SIGPIPE/SIGHUP referenced at module top level would
AttributeError on import on Windows, breaking `hermes --tui` entirely (it
spawns this module as a subprocess). Guard each signal.signal() call with
hasattr() and add SIGBREAK as Windows' SIGHUP equivalent.
- hermes_cli/kanban_db.py: os.waitpid(-1, os.WNOHANG) in dispatcher tick was
unguarded. os.WNOHANG doesn't exist on Windows. Gate the whole reap loop
behind `os.name != "nt"` — Windows has no zombies anyway.
- tools/code_execution_tool.py: AF_UNIX socket for execute_code RPC fails on
most Windows builds. Fall back to loopback TCP (AF_INET on 127.0.0.1:0
ephemeral port) when _IS_WINDOWS. HERMES_RPC_SOCKET env var now accepts
either a filesystem path (POSIX) or `tcp://127.0.0.1:<port>` (Windows).
Generated sandbox client parses both.
- cron/scheduler.py: `argv = ["/bin/bash", str(path)]` hardcoded. Use
shutil.which("bash") so Windows (Git Bash via MinGit) works, with a
readable error when bash is genuinely absent.
- 6 bare npm/npx spawn sites: tools_config.py x2, doctor.py, whatsapp.py
(npm install + node version probe), browser_tool.py x2. On Windows npm
is npm.cmd / npx is npx.cmd (batch shims); subprocess.Popen(["npm", ...])
fails with WinError 193. shutil.which(...) returns the absolute .cmd
path which CreateProcessW accepts because the extension routes through
cmd.exe /c. POSIX behaviour unchanged (shutil.which still returns the
same path subprocess would resolve itself).
## HIGH fixes (silent misbehaviour on Windows)
- tools/environments/local.py get_temp_dir: hardcoded /tmp returned on
Windows meant `_cwd_file = "/tmp/hermes-cwd-*.txt"`, which bash wrote
via MSYS2's virtual /tmp but native Python couldn't open. Result: cwd
tracking silently broken — `cd` in terminal tool did nothing. Windows
branch now returns `%HERMES_HOME%/cache/terminal` with forward slashes
(works in both bash and Python, guaranteed no spaces).
- tools/environments/local.py _make_run_env PATH injection: `/usr/bin not
in split(":")` heuristic mangles Windows PATH (";" separator). Gate
the injection behind `not _IS_WINDOWS`.
- hermes_cli/gateway.py launch_detached_profile_gateway_restart: outer
Popen + watcher-script Popen both used start_new_session=True, which
Windows silently ignores. Watcher stayed attached to CLI's console,
died when user closed terminal after `hermes update`, left gateway
stale. Now branches through windows_detach_popen_kwargs() helper
(CREATE_NEW_PROCESS_GROUP | DETACHED_PROCESS | CREATE_NO_WINDOW on
Windows, start_new_session=True on POSIX — identical to main).
## MEDIUM fixes
- gateway/run.py /restart and /update handlers: hardcoded bash/setsid
chain crashes on Windows when user triggers /update in-gateway. Now
has sys.platform=="win32" branch using sys.executable + a tiny
Python watcher with proper detach flags. POSIX path is unchanged.
- cli.py _git_repo_root: Git on Windows sometimes returns /c/Users/...
style paths that break subprocess.Popen(cwd=...) and Path().resolve().
Added _normalize_git_bash_path() helper that translates /c/Users,
/cygdrive/c, /mnt/c variants to native C:\Users form. POSIX no-op.
_git_repo_root() now routes every result through it.
- cli.py worktree .worktreeinclude: os.symlink on directories failed
hard on Windows (requires admin or Developer Mode). Falls back to
shutil.copytree with a warning log.
## Tests
- 29 new tests in tests/tools/test_windows_native_support.py covering:
subprocess_compat helpers, TUI entry signal guards, kanban waitpid
guard, code_execution TCP fallback source-level invariants, cron bash
resolution, npm/npx bare-spawn lint per-file, local env Windows temp
dir, PATH injection gating, git bash path normalization, symlink
fallback, gateway detached watcher flags.
- One existing test assertion adjusted in test_browser_homebrew_paths:
it compared captured Popen argv to the BARE `"npx"` literal; after the
shutil.which() change argv[0] is the absolute path. New assertion
checks the shape (two items, second is `agent-browser`) rather than
the exact first-item string. Behaviour unchanged; test was too strict.
All 56 tests pass on Linux (30 from previous commits + 26 new).
267 tests from the affected files/dirs (browser, code_exec, local_env,
process_registry, kanban_db, windows_compat) all pass — zero regressions.
tests/hermes_cli/ (3909 pass) and tests/gateway/ (5021 pass) unchanged;
all pre-existing test failures confirmed unrelated via `git stash` re-run.
## What's still deferred (LOW priority)
- Visible cmd-window flashes on short-lived console apps (~14 sites) —
cosmetic, needs a follow-up pass once we have user reports.
- agent/file_safety.py POSIX-only security deny patterns — separate
hardening task.
- tools/process_registry.py returning "/tmp" as fallback — theoretical;
reachable only when all env-var candidates fail.
|