Non-interactive sessions (hermes chat -q, hermes -z) snapshot the tool
registry at AIAgent construction time. If background MCP discovery hasn't
finished, MCP tools are invisible for the entire session — and unlike
interactive mode, there is no between-turns late-binding refresh to recover.
Root cause: wait_for_mcp_discovery() only joins an already-created discovery
thread, so it no-ops if a direct/single-query path reaches agent construction
before MCP startup created that thread. Oneshot._run_agent() didn't call it
at all.
Fix:
- Add ensure_mcp_discovery_before_agent_build() helper to mcp_startup.py:
idempotently starts discovery if needed + bounded wait. Fail-open on errors.
- Add single_query parameter to _resolve_discovery_timeout/wait_for_mcp_discovery:
uses mcp_single_query_discovery_timeout (default 15s) instead of the
interactive mcp_discovery_timeout (1.5s) because one-shot sessions have no
second turn to recover.
- Wire into CLI _init_agent (single_query from _single_query_mode flag set
in cli.py's single-query path) and oneshot._run_agent (single_query=True).
- Interactive sessions unchanged: keep 1.5s bound (between-turns refresh covers).
Closes#38448, #51316, #37013, #68137
Composite salvage of #60017 (chrishart0), #51322 (Bartok9), #38620 (buptwz),
#43544 (halonke), #36882 (vanhoof).
get_custom_provider_extra_headers() was returning the result of
normalize_extra_headers() on the first matching base_url, even when
that entry had no extra_headers configured. A later providers.<name>
entry sharing the same URL but with headers set was therefore ignored.
Fix: store the normalized headers and only return when non-empty,
otherwise continue searching the remaining entries.
Fixes#74465
Verify that set_config_value and unset_config_value refuse to write
when config.yaml contains YAML syntax errors, and the original file
is left intact.
2de1e86c16 appended updated versions of five doctor tests without
removing the originals; the earlier definitions were silently shadowed
(dead) and tests/test_no_shadowed_test_definitions.py now fails on every
PR slice that runs it. Keep the later (runtime-winning) definitions,
delete the stale earlier ones.
Follow-ups to the previous commit (#74414 by @webtecnica, re #74373):
- When distribution_owned is OMITTED, restore the legacy contract: every
staged entry outside USER_OWNED_EXCLUDE is copied. The cherry-picked
filter consulted owned_paths(), which silently narrowed omitted-list
distributions to DEFAULT_DIST_OWNED and dropped undeclared payload
(extra top-level files/dirs existing distributions legitimately ship).
- Make explicit allowlists path-aware so documented nested entries like
skills/research/ and cron/digest.json select exactly that subtree/file
instead of being dropped by the top-level name comparison. Traversal
segments (.., absolute) and USER_OWNED_EXCLUDE roots are still rejected.
- Regression tests: omitted-list legacy behavior + nested-path allowlist.
_copy_dist_payload() in profile_distribution.py iterated all staged
entries without consulting the manifest's distribution_owned allowlist,
so manifests that restricted distribution_owned only had cosmetic effect.
Fix: compute manifest.owned_paths() at the top of _copy_dist_payload()
and skip entries not in that set, after the USER_OWNED_EXCLUDE check.
The owned_paths() method already existed on DistributionManifest and
correctly falls back to DEFAULT_DIST_OWNED when no explicit
distribution_owned is set, so the new filter preserves backward
compatibility for existing manifests.
Closes#74373
Follow-ups to the previous commit (#74155 by @Drexuxux):
- enrich_model_switch_warnings_for_gateway() -> merge_preflight_compression_warning()
still called the sync resolve_display_context_length() provider probe ladder
inline in both async /model call sites; dispatch it via asyncio.to_thread.
- Replace the inspect.getsource() test (source-reading tests are banned by
AGENTS.md) with behavioral tests that drive the real _handle_model_command:
assert the resolver runs off the loop thread and that the warning enrichment
is dispatched through asyncio.to_thread.
resolve_display_context_length() runs two blocking chains: the route
comparison in should_clear_context_pin() and the provider probe ladder in
get_model_context_length() (blocking requests calls to Anthropic /v1/models,
Copilot, Nous, Codex, GMI, Ollama, models.dev and OpenRouter).
The gateway message path already offloads both via
get_model_context_length_async() and should_clear_context_pin_async(), but
the /model slash-command handlers (_handle_model_command, _finish_switch)
called the sync helper directly, freezing the whole event loop for the
duration of the probe ladder - no messages processed on any platform, and
the Discord heartbeat timeouts that get_model_context_length_async() was
introduced to prevent.
Add resolve_display_context_length_async(), a thin asyncio.to_thread wrapper
mirroring the two existing *_async helpers (no logic duplication), and await
it at both handlers.
7b5a18817 migrated the sibling slug sites to custom_provider_slug, which
keeps a keyed providers: entry's config key as its durable identity. It
covered find_custom_provider_identity_by_model; canonical_custom_identity's
third recovery source - the configured-provider fallback - still built
f"custom:{normalized}" out of whatever string the caller happened to hold.
_get_named_custom_provider matches on either spelling, so a display name
that differs from its config key matches the entry and then heals to
custom:<display-name>. That is a second identity for one endpoint: the
endpoint- and model-based sources of the same function return
custom:<config-key>, and so does everything that persists or restores a
session's provider override. canonical_custom_identity exists precisely to
make a bare "custom" routable again, and tui_gateway calls it on the
session-persist, resume and recovery paths - so the divergence lands in
stored session identity.
Re-resolve through the endpoint the matched entry owns, reusing the
function's own URL-based canonicaliser rather than duplicating the match
logic. Legacy unkeyed custom_providers: entries keep their name identity,
and an unconfigured candidate still returns None.
Ports the negative limit/offset fix onto the current router modules
(hermes_cli/web_routers/sessions.py, profiles.py) since the handlers
moved out of web_server.py in 011ec4513e after this PR was opened.
Per review feedback: only add Query(..., ge=0) — no le=500. The
messages route already clamps oversized requests with min(limit, 500)
and must keep that behavior (succeed + cap) rather than reject them;
the two session-list routes never had a public 500 cap and shouldn't
gain a new rejecting one as a side effect of this fix.
_resolve_explicit_runtime's generic-provider branch accepted model_cfg's
persisted api_mode unconditionally, letting a stale mode from a previous
provider (e.g. anthropic_messages) leak into a newly-switched provider
(e.g. gemini) and break the transport. Reuse the existing
_provider_supports_explicit_api_mode guard, already used by the copilot
and named-custom-provider resolution paths for exactly this case, so the
persisted mode is only honored when model_cfg's provider matches the one
being resolved.
Closes#74318
resolve_entry_api_key() and the duplicated _fallback_entry_api_key()
read key_env via a raw os.getenv(), bypassing per-profile secret
scoping in the multiplexed gateway. Under multiplexing this can hand
a fallback request another profile's credential. Both now resolve
through agent.secret_scope.get_secret(), which reads the active
profile scope when multiplexing is on and falls back to os.environ
unchanged when it's off, so single-profile behavior is preserved.
Closes#74311
_load_auth_store() treated every exception from reading auth.json as
corruption and returned an empty store. EMFILE under fd exhaustion,
EACCES, EIO and a stalled network mount all reached that branch. This
module does read-modify-write in roughly fifteen places, so the empty
store was one _save_auth_store() away from erasing every stored
credential.
Separate OSError from parse failure: a file that exists but cannot be
read now raises, naming the real cause and leaving the file on disk
untouched. Only a genuine parse failure takes the preserve-and-start-
empty branch, which is unchanged.
The backup was also unreliable in exactly the conditions that triggered
it: shutil.copy2 opens a file, so under EMFILE it failed too, its bare
except swallowed that, and the log still said "Corrupt file preserved
at ..." when nothing had been written. Track whether the copy landed
and say so accurately.
_venv_launcher_ancestors() ran after
_wait_for_windows_update_gateway_exit(), but the drain stops tracking a
PID exactly when it dies - for the common graceful-drain case the worker
is gone by the time the wait returns, and a dead pid's parent cannot be
recovered, so the launcher stop never fired on that path. Resolve
launcher ancestors before draining and stop the snapshot afterwards
alongside the survivors; a launcher that already exited with its worker
raises ProcessLookupError at the kill and is skipped.
The set-cover invariant test now marks drained workers uninspectable
(construction raises, like psutil.NoSuchProcess), so a post-drain
launcher lookup can never reappear unnoticed.
_is_pausable_gateway() hand-rolled a second gateway parser and regressed
a valid form: in `--profile gateway gateway run` the profile VALUE
shadowed the subcommand token, so the scan reported that gateway as a
fatal preflight holder. Delegate to
gateway.status.looks_like_gateway_command_line() - profile-selector
aware, shlex-tokenizing, run-only - so the preflight exemption, the
pause discovery, and the updater's guard fallback share one parser.
Non-run gateway subcommands, serve backends, and REPLs still block; the
bare-`gateway` form now classifies as a running gateway, mirroring the
canonical matcher's contract.
The pause stops every gateway its discovery maps, but the venv-holder
guard sees the process table as it is now: a gateway respawned by its
supervisor (Scheduled Task, login watchdog) inside the pause-to-guard
window, or one started through a spawn path discovery does not map,
still holds venv .pyds - and the guard dead-ended the update on exactly
the kind of process the pause machinery exists to stop.
When every remaining holder classifies as a pausable gateway - using the
same _is_pausable_gateway matcher the Desktop preflight uses, so the two
views cannot drift - stop them and re-scan once. Any non-gateway holder
(REPL, stray script, Desktop backend) keeps the hard refusal exactly as
before, and a survivor after the stop still aborts.
The Desktop update preflight (`scanVenvBlockers` -> `python -m
hermes_cli._scan_venv_blockers`) reports every venv-side python as a
blocker and aborts the handoff:
main.ts: scanVenvBlockers(...) <- aborts HERE
return { ok:false, error:'venv-blocked' }
spawnUpdaterProcess(hermes-setup ...) <- never reached
But a *gateway* is not a dead-end holder. `hermes-setup` invokes
`hermes update --yes --gateway`, and the CLI updater's
`_pause_windows_gateways_for_update()` gracefully drains and stops
running gateways before touching the venv — machinery added for exactly
these processes (#50090 and follow-ups). The preflight replicated the
CLI's *guard* without its *pause*, so a Windows service-mode gateway
(e.g. a Scheduled Task running `gateway run`) made every Desktop update
abort forever with
[updates] venv-blocked: N process(es) hold the install
PID ... python.exe ... -m hermes_cli.main gateway run --replace
while the component one layer down was never allowed to run and handle
it. The abort points at a process the updater knows how to stop.
Fix: `_is_pausable_gateway()` exempts `hermes_cli.main ... gateway run`
invocations (both halves of the venv-shim launcher/worker chain match,
since the uv-side worker re-runs the same argv). Everything else keeps
blocking — the Desktop `serve` backend, other `gateway` subcommands,
operator REPLs and stray scripts have no pause machinery downstream.
The CLI updater's own post-pause venv guard is untouched, so a pause
that genuinely fails still aborts before any .pyd mutation.
The JSON gains a diagnostic `pausable_gateways` count. The TS consumer
validates only `ok`/`blocked`/`processes` and ignores unknown keys, so
old and new Desktop builds both accept the new document; no Electron
rebuild is required for the fix to take effect (the scan runs from the
repo's Python).
`import pty` at module scope pulls in `termios`, which does not exist on
Windows. That raised ModuleNotFoundError during *collection*, so pytest
aborted the whole module with
Interrupted: 1 error during collection
before any skip marker could take effect. The single PTY-dependent test
was already correctly marked `skipif(sys.platform == "win32")` — the
import crashed ahead of it and took the module's 13 other, entirely
platform-agnostic tests down as collateral. Windows contributors got zero
gateway coverage and, worse, a collection error that masks real failures
in any batch that includes this file.
Two changes:
- Move `import pty` into the `stdin_is_tty` branch that actually uses it
(the sole `pty.openpty()` call). Nothing else in the module needs it.
- Skip `test_systemd_install_checks_linger_status` on Windows. It drives
`_systemd_linger_enabled()` -> `os.getuid()`, which does not exist on
Windows; the production helper is annotated
"windows-footgun: ok — POSIX systemd helper, never invoked on Windows",
so the test is Linux-only by nature. It was previously hidden behind
the collection crash.
POSIX behaviour is unchanged: both guards are `skipif(win32)`, inactive
off Windows, and the local import resolves exactly where the module-level
one did.
before (Windows): 0 collected, 1 collection error
after (Windows): 14 collected, 9 passed, 5 skipped
On Windows a gateway started through the venv shim is a two-process chain:
venv\Scripts\python.exe (launcher — keeps venv .pyd files mapped)
└─ uv\python\...\python.exe (worker — writes the gateway PID file)
`_pause_windows_gateways_for_update()` builds its pause set from
`find_gateway_pids()`, which reads the PID file and therefore only ever
sees the *worker*. The venv-holder guard immediately downstream
(`_detect_venv_python_processes()`) matches on the venv path prefix, so it
only ever sees the *launcher*.
The two sets are disjoint. A gateway the updater had just gracefully
drained still left its launcher alive, the guard reported that launcher as
a venv holder, and the update aborted — every time. On the Desktop path
this surfaces as the dead-end dialog:
[updates] venv-blocked: 2 process(es) hold the install
PID ... python.exe ...\venv\Scripts\python.exe -m hermes_cli.main gateway run --replace
Note the reported holder is a gateway the updater believes it stopped.
The Desktop path is affected because `hermes-setup.exe` runs
`hermes update --yes --gateway --force`, and `--force` deliberately does
NOT bypass the venv guard (that needs `--force-venv`), so the abort is
correct behaviour reacting to an incomplete pause.
Fix: after the graceful drain, walk one hop up from each mapped gateway
PID and force-kill parents that live under the project venv.
Deliberately additive, not a substitution:
- The planned-stop marker and the graceful drain still target the worker
(the PID that wrote the PID file), so clean shutdown is unchanged and
updates don't get pushed onto the hard-kill path.
- `terminate_pid(force=True)` is `taskkill /T` (tree kill), so killing a
launcher that outlived its worker also reaps stragglers.
- `_resume_windows_gateways_after_update()` needs no change: the mapped
respawn argv is rebuilt from the profile name
(`_gateway_run_args_for_profile`), never from the killed PID, and the
restart watcher's `_pid_exists()` wait still terminates because the
tree kill takes the whole chain down.
- Only the venv-side parent is returned. Unrelated ancestors (a Scheduled
Task's `cmd.exe`, an operator's shell) are ignored, and the caller's own
process chain is excluded so a CLI `hermes update` never nominates
itself.
Tests assert the invariant the two PID-resolution paths must satisfy —
the pause's kill set must cover the guard's abort set — rather than
snapshotting PIDs. Verified to fail without the fix:
AssertionError: pause stopped [] but the venv guard aborts on [400]
— disjoint sets abort the update
--force was silently ignored for 'model' keys — the guard always
redirected to model.default even when the user explicitly asked to
replace the entire section. Now --force triggers a warning and
proceeds with the destructive overwrite for model too, matching
the non-model mapping --force behaviour.
Prevent 'hermes config set <section> <scalar>' from silently destroying
an existing mapping. The bare 'model' shorthand is preserved by
redirecting to 'model.default' — all other mapping sections are refused
with a helpful error unless --force is used.
Closes#74995
- Bound ALL reads of the on-disk JWT store through one _read_jwt_store()
helper (load, eviction, save-merge) — the 1 MiB cap previously only
covered the load path; eviction and save could still parse an
oversized/corrupt store and rewrite it back out (sweeper finding).
- Fix the class, not the site: the recovery gates checked the literal
provider == "copilot" while /model and profile configs can leave the
alias spelling in place (the reporter's own log shows provider=copilot
AND provider=github-copilot in one session — the aliased turns would
have silently skipped recovery). Single owner:
AIAgent._is_copilot_provider() (slug aliases + Copilot base-URL
fallback), used by both run_agent recovery methods and both
conversation_loop gates.
- Update the salvaged 401 test to current main's client-retirement
contract (release deferred to GC — no synchronous .close()).
- Add copilot_stale_cred_retry_attempted to the TurnRetryState field
contract test; add bounded-store and alias-gate regression tests.
The stale-staged-updater deadlock is not Windows-specific: hermes-setup
under ~/.hermes is only refreshed by a full installer run
(copy_self_to_hermes_home no-ops during --update), so every desktop whose
staged updater predates the HERMES_UPDATE_HANDOFF_PID export (8c76fe19)
runs an old parent that never sends the env var against a new child that
demands it — exit 2 ('Hermes is still running') forever, on macOS and
Linux just as on Windows.
Replace the wmic ancestry walk (deprecated, absent on current Win11,
GBK decode juggling) with psutil.Process().parents() — psutil is already
a hard dependency and is the project's canonical no-kill process probe.
Drop the os.name == 'nt' gate so all platforms heal. Add tests: a marker
owned by our parent process is recognized as our orchestrator; a live
non-ancestor holder is still refused.
Independent review pass: utils.base_url_host_matches already owns the
exact-or-dot-suffix hostname contract (userinfo/port stripped, lowercased,
trailing dot removed), so the predicate delegates instead of hand-rolling
a second suffix match to keep in sync. Also locks in the normalization
behavior the review verified empirically: uppercase+port, trailing-dot,
userinfo-stripped, and IPv6-literal cases added to the contract tests.
Three catalog-side defects from the same report, all downstream of the
exact-host assumption and the config/env asymmetry:
- Discovery read only $OPENAI_BASE_URL, so a config-set
model.base_url (the supported way to select a data-residency host)
was ignored and /model listed the catalog of api.openai.com, not the
configured endpoint. New _openai_discovery_base_url() resolves
env override -> matching model.base_url -> canonical default, the same
precedence inference uses.
- _credential_fingerprint() hashed env vars only, so hermes config set
model.base_url kept serving the previous endpoint's cached catalog
until TTL expiry. The effective endpoint is now folded into the
fingerprint for openai/openai-api.
- is_default_openai matched two literal URLs, so regional hosts (which
serve the identical 120+ entry dump) bypassed the curated intersection
and flooded the picker with whisper/tts/embedding/dall-e rows. Now uses
the shared official-host predicate; custom OpenAI-compatible proxies
keep the verbatim live list.
- validate_requested_model's curated-catalog soft-accept (#46850) no
longer applies on official OpenAI hosts: their /v1/models listing is
access-scoped and authoritative, so accepting an absent model
manufactures a selection that 400s at first use. Custom proxies and
other providers keep the #46850 fallback. The #37404
empty-intersection -> curated picker fallback is deliberately left
unchanged.
The P1 from the enterprise data-residency report: with
model.base_url=https://us.api.openai.com/v1, every tool-calling turn 400'd
('Function tools with reasoning_effort are not supported ... use
/v1/responses') because the runtime resolvers hardcoded
api_mode=chat_completions and consulted URL detection only. openai-api
declares codex_responses in its overlay; the declaration was never
consulted, so any OpenAI host that wasn't literally api.openai.com landed
on the wrong wire protocol.
New _fallback_api_mode(provider, base_url, model): URL detection first
(host-mandated wire shapes keep priority), then
providers.determine_api_mode() (the provider's declared transport), then
chat_completions only for genuinely unknown providers. All three runtime
fallback sites route through it: the pool-entry path, the explicit-runtime
path, and the API-key-provider path, so the lanes cannot drift apart.
Blast radius beyond openai-api: minimax, minimax-cn, and copilot-acp were
the other overlays whose declared non-chat transport fell through to
chat_completions on the same paths (same latent bug class). openrouter is
unaffected (declares openai_chat). _detect_api_mode_for_url also now uses
the shared official-host predicate, so regional hosts detect as
codex_responses on the direct-URL lane too.
Pointing openai-api at OpenAI's documented regional hosts
(us.api.openai.com / eu.api.openai.com, mandatory for customers with
data-residency obligations) silently degraded Hermes because three
subsystems tested 'is this OpenAI' with exact-hostname equality against
api.openai.com.
Adds providers.is_official_openai_host(): canonical host plus dot-suffix
subdomains of api.openai.com, hostname-parsed only. Lookalike hosts
(api.openai.com.attacker.test) and path spoofs (proxy.test/api.openai.com/v1)
stay rejected, preserving the #32243 hardening: a genuine *.api.openai.com
subdomain requires control of openai.com DNS.
host_mandated_api_mode() now routes through the predicate, so regional
hosts mandate codex_responses exactly like the canonical host.
The retag gate was global, so once one board reclaimed its legacy rows a
second board on the same state.db never got swept. Key the state_meta gate
on the workspaces root and skip reopening state.db on every spawn via an
in-process set. Align the dispatcher-spawn test with the worker's own
`kanban` source tag and cover the per-board gate.
The repo's .npmrc sets engine-strict=true and package.json pins
engines.npm, so an npm outside that range aborts every npm ci /
npm install we run inside the checkout:
npm error code EBADENGINE
npm error notsup Required: {"npm":"<11.10.0 || >=12.0.0"}
npm error notsup Actual: {"npm":"11.10.0"}
Our callers made that worse: _run_npm_install_deterministic sees
`npm ci` fail and falls through to `npm install`, which fails
identically, so the user got a buried EBADENGINE and no remedy.
React to the failure instead of predicting it. npm states the
required range in its own error, so there is no need for a version
probe on the happy path or a semver range matcher — the recovery
reads the constraint out of the output it just produced, upgrades,
and retries once.
Scope is deliberately narrow. Hermes only upgrades an npm inside its
own managed Node tree ($HERMES_HOME/node), installing with --prefix
so bin/npm keeps resolving to the upgraded lib/node_modules/npm; a
managed install writes prefix=~/.local into node/etc/npmrc, so
without the override the "upgrade" would land elsewhere while the
managed npm stayed stale. A system / nvm / brew / Nix npm belongs to
the user, so that case prints the exact command and lets the original
failure stand.
The upgrade runs from a temp cwd with npm_config_min_release_age=0,
otherwise the checkout's own min-release-age gate would refuse the
npm release we need.
_run_npm_install_deterministic's capture_output=False callers (the
desktop install) streamed npm output and returned stderr=None, which
would leave the recovery nothing to read — stderr is now teed, so
live output is unchanged and the text stays inspectable.
Verified end to end against real npm binaries on copies of a managed
tree: managed npm 11.10.0 -> EBADENGINE -> upgraded to 12.0.2 ->
retry exits 0; a foreign npm 11.10.0 hard-fails with the manual
command and is left untouched.
Use providers keys as the canonical custom-provider identity while accepting legacy bare keys, display-name slugs, bare custom fallback, and doubled custom prefixes across resolution, pickers, doctor, and runtime reverse lookup.
Co-authored-by: Bakhtier Sizhaev <bakhtiersizhaev@users.noreply.github.com>
A running worker now polls its comment thread and folds new operator notes
into the live turn via the OUT-OF-BAND steer channel (list_comments_after +
a heartbeat-driven bridge, watermarked so history isn't re-injected and the
worker's own notes are skipped). No block→comment→unblock dance. Desktop's
composer sends notes live ("delivered within a few seconds") with "Requeue
with note" as the restart option and a help tooltip.
Boards gain an optional project_id. When set, the board's default_workdir
mirrors the project's primary repo and every new task inherits the project —
a deterministic worktree + branch per task — unless it names its own. New
GET /projects; board create/patch/list carry project_id + resolved name; the
create dialog defaults its workspace to the board's and allows a per-task
path override. Desktop: "Board settings…" gains a project picker.
Salvaged from #57016 by @lEWFkRAD:
- cli.py: handle file:///C:/... drive-letter URIs on nt (strip the
leading slash urlparse leaves); join Termux example paths with literal
forward slashes so hints stay POSIX on Windows.
- gateway/status.py + hermes_cli/gateway.py: normalize backslashes to
forward slashes before the HERMES_HOME substring match so separator
style cannot defeat profile ownership detection.
- hermes_cli/banner.py: cprint degrades to plain print when
prompt_toolkit has no console (NoConsoleScreenBufferError on
redirected/absent Windows stdout).
- hermes_cli/browser_connect.py: posixpath.join for WSL /mnt/c/... bases
(os.path.join would emit backslashes on nt).
- Test hardening: symlink skip-guards, USERPROFILE alongside HOME for
ntpath.expanduser, SIGKILL absence skipif fixed via monkeypatch,
drive-letter URI / separator-normalization / banner-fallback coverage.
Dropped from the original PR: tests/cli/conftest.py fixture and the
AppSession _output monkeypatch — main's merged tests/cli/conftest.py
already handles that prompt_toolkit pollution.
The cross-process update lock (fe8e4d93d) made the in-progress marker
mutually exclusive across every update entrypoint — but the Tauri
updater holds that marker for its WHOLE run and then spawns
hermes update as a child stage. The child read the marker, found its
own parent's live pid, refused with exit 2, and the GUI mapped that to
"Hermes is still running. Close all Hermes windows and try the update
again." Retry spawns a fresh updater that deadlocks against itself the
same way, so every GUI-driven update dead-ends on the failure screen
with no winnable retry (observed: three consecutive self-refusals in
bootstrap-installer.log within 90 seconds).
Hand the claim off explicitly: update_child_env exports
HERMES_UPDATE_HANDOFF_PID naming the updater's own pid, and
UpdateLock.acquire treats a live holder matching that pid as the lock
we are already running under — run without claiming, and release
leaves the parent's marker untouched. The env var alone grants
nothing: the pid must also be the live marker owner, so a stale or
forged value cannot bypass the lock, and a dashboard-spawned
hermes update (no handoff env) is still refused exactly as before.
Autouse fixture also resets approval_module._YOLO_MODE_FROZEN so a
HERMES_YOLO_MODE=1 host env can't poison every case (the one
startup-frozen test still patches it back explicitly). Adds the darwin
'ps -o stat=' zombie branch to _is_alive_like_dispatcher, mirroring
production hermes_cli/kanban_db.py — a no-op on Linux.
Salvaged from PR #34069 by @sunwz1115.
Co-authored-by: sunwz1115 <192549904+sunwz1115@users.noreply.github.com>
Guard the .lazy-refresh-incomplete marker writer (update_cmd), launch-time
recovery (main.py), and _early_recovery repair paths behind a two-condition
check: running under pytest AND the target is this live checkout. Sandboxed
tmp_path tests still exercise the real code paths.
Salvaged from PR #72002 by @fcavalcantirj. Fixes#72000.
Co-authored-by: fcavalcantirj <felipe.cavalcanti.rj@gmail.com>